Voxel-based upsampling for context awareness of point cloud processing
By using initial upsampling and context information for voxel pruning in the video encoding and decoding process, the problem of point cloud upsampling and high computational cost in the prior art is solved, and efficient point cloud upsampling and refinement is achieved.
Patent Information
- Application Number
- CN202380073924.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-10
- Filing Date
- 2023-10-17
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively upsample in point cloud processing, resulting in unrefined geometry of point cloud and high computational cost.
Point clouds are upsampled by using initial upsampling during video encoding and decoding, voxel pruning is performed in combination with context information, and pruned point clouds are generated, and feature aggregation and context-aware upsampling are performed during decoding.
Efficient upsampling of point clouds is achieved, the geometry of point clouds is refined, the computing cost is reduced, and the accuracy of point cloud processing is improved.
Smart Images

Figure CN120077647A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application is an international application that claims the benefit of U.S. Provisional Patent Application No. 63 / 438,212, entitled "CONTEXT - AWARE VOXEL - BASED UPSAMPLING FOR POINT CLOUD PROCESSING," filed on January 10, 2023 ("the '212 application") and U.S. Provisional Patent Application No. 63 / 417,284, entitled "CONTEXT - AWARE VOXEL - BASED UPSAMPLING FOR POINT CLOUD PROCESSING," filed on October 18, 2022 ("the '284 application") under 35 U.S.C. § 119(e). The entire contents of these two U.S. provisional patent applications are hereby incorporated by reference into this application. Reference Citation
[0002] This application incorporates by reference in its entirety the following applications: U.S. Provisional Patent Application No. 63 / 291,015, titled "Hybrid Framework for Point Cloud Compression", filed on December 17, 2021 ("the '015 application"); U.S. Provisional Patent Application No. 63 / 297,869, titled "A Scalable Framework for Point Cloud Compression", filed on January 10, 2022 ("the '869 application"); U.S. Provisional Patent Application No. 63 / 388,087, titled "A Scalable Framework for Point Cloud Compression", filed on July 11, 2022 ("the '087 application"); U.S. Provisional Patent Application No. 63 / 252,482, titled "Method and Apparatus for Point Cloud Compression Using Hybrid Deep Entropy Coding", filed on October 5, 2021 ("the '482 application"); U.S. Provisional Patent Application No. 63 / 297,894, titled "Coordinate Refinement and Upsampling from Quantized Point Cloud Reconstruction", filed on January 10, 2022 ("the '894 application"); and U.S. Provisional Patent Application No. 63 / 388,600, titled "Deep Distribution-Aware Point Feature Extractor for AI-Based Point Cloud Compression", filed on July 12, 2022 ("the '600 application"). BACKGROUND OF THE DISCLOSURE
[0003] The point cloud (PC) data format is a common data format across several commercial domains such as autonomous driving, robotics, augmented reality / virtual reality (AR / VR), civil engineering, computer graphics, and the animation / movie industry. 3D LiDAR (Light Detection and Ranging) sensors have been deployed in self-driving cars, and affordable LiDAR sensors are available. With the advancement of sensing technology, 3D point cloud data has become more practical than ever before. SUMMARY OF THE DISCLOSURE
[0004] Embodiments described herein include methods used in video encoding and decoding (collectively referred to as "decoding").
[0005] A first example method / apparatus according to some embodiments may include: upsampling a first point cloud using an initial upsampling to obtain a second point cloud; associating features of the second point cloud with context information to obtain a third point cloud; predicting an occupancy state of at least one voxel of the third point cloud; and removing voxels classified as empty in the third point cloud according to the predicted occupancy state to generate a pruned point cloud.
[0006] For some embodiments of the first example method, the initial upsampling includes nearest-neighbor upsampling.
[0007] For some embodiments of the first example method, associating features includes: concatenating features of the second point cloud with the context information to obtain a third point cloud.
[0008] For some embodiments of the first example method, the context information is voxel-wise context information.
[0009] For some embodiments of the first example method, the context information includes a context point cloud.
[0010] For some embodiments of the first example method, the context information includes information about the second point cloud.
[0011] For some embodiments of the first example method, the context information includes information about the voxel occupancy state of the second point cloud.
[0012] For some embodiments of the first example method, the context information includes information about the position of a child voxel relative to the position of a parent voxel of the first point cloud.
[0013] For some embodiments of the first example method, the context information includes coordinate information about the positions of occupied voxels in at least one of the first point cloud and the second point cloud.
[0014] For some embodiments of the first example method, the context information includes coordinate information, and the coordinate information is in the form of one of Euclidean coordinates, spherical coordinates, and cylindrical coordinates.
[0015] For some embodiments of the first example method, the context information provides known information about the first point cloud in addition to information available for the initial upsampling of the first point cloud.
[0016] For some embodiments of the first example method, the context information includes the bitdepth of the second point cloud.
[0017] Some embodiments of the first example method may also include performing feature decoding on the input point cloud and the first bitstream to generate the first point cloud.
[0018] Some embodiments of the first example method may also include: performing feature aggregation on the trimmed point cloud to generate the aggregated features; and performing a context-aware upsampling process on the aggregated features to generate the decoded point cloud.
[0019] Some embodiments of the first example method may also include: performing a feature to residual conversion on the trimmed point cloud to generate a residual output; and adding the trimmed point cloud to the residual output to generate the decoded point cloud.
[0020] Some embodiments of the first example method may also include performing feature aggregation on the trimmed point cloud to generate the aggregated features, wherein the feature to residual conversion is performed on the aggregated features.
[0021] For some embodiments of the first example method, a first neural network is used to perform the prediction of the occupancy state.
[0022] For some embodiments of the first example method, predicting the occupancy state predicts the ground-truth occupancy state of at least one voxel.
[0023] For some embodiments of the first example method, predicting the occupancy state predicts the likelihood that the at least one voxel is occupied.
[0024] For some embodiments of the first example method, the voxels of the third point cloud are removed using a voxel pruning process.
[0025] Some embodiments of the first example method may also include aggregating at least one feature of the second point cloud.
[0026] For some embodiments of the first example method, predicting the occupancy state of at least one voxel includes: aggregating at least one feature of the third point cloud; processing the aggregated features with a multi-layer perceptron (MLP) layer to generate an MLP layer output; performing a softmax process on the MLP layer output to generate a softmax output value; and performing thresholding of the softmax output value to generate the predicted occupancy state of at least one voxel of the third point cloud.
[0027] For some embodiments of the first example method, thresholding of the softmax output values converts softmax output values greater than 0.5 into an output value of 1, and converts softmax output values equal to 0.5 or less into an output value of 0.
[0028] For some embodiments of the first example method, predicting the occupancy state of at least one voxel includes: aggregating at least one feature of the third point cloud; and generating a predicted occupancy state of at least one voxel of the third point cloud based on the aggregated feature.
[0029] For some embodiments of the first example method, aggregating at least one feature includes: repeating the concatenation process one or more times, the concatenation process including: performing a sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; performing a non-linear activation process on the convolutional output point cloud to generate a non-linear output point cloud; and if there is a next cycle of the concatenation process, preparing the non-linear output point cloud as the input point cloud, where the third point cloud is the input point cloud of the first cycle of the concatenation process, and where the last cycle of the concatenation process generates the aggregated feature.
[0030] Some embodiments of the first example method may further include adding the third point cloud to the ReLU output point cloud of the last cycle of the concatenation process.
[0031] For some embodiments of the first example method, aggregating at least one feature includes: performing a sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; and performing a non-linear activation process on the convolutional output point cloud to generate the aggregated feature.
[0032] For some embodiments of the first example method, the non-linear activation process includes a rectifier linear unit (ReLU) activation process, and the non-linear output point cloud includes a ReLU output point cloud.
[0033] For some embodiments of the first example method, aggregating at least one feature includes: repeating the first cascading process one or more times, the first cascading process including: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first non-linear activation process on the first convolutional output point cloud to generate a first non-linear output point cloud; and if there is a next cycle of the first cascading process, preparing the first non-linear output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; repeating the second cascading process one or more times, the second cascading process including: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second non-linear activation process on the second convolutional output point cloud to generate a second non-linear output point cloud; and if there is a next cycle of the second cascading process, preparing the second non-linear output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process generates a second cascading process output; concatenating the first cascading process output and the second cascading process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0034] For some embodiments of the first example method, aggregating at least one feature includes: repeating the first cascading process one or more times, the first cascading process including: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first rectifier linear unit (ReLU) activation process on the first convolutional output point cloud to generate a first ReLU output point cloud; if there is a next cycle of the first cascading process, preparing the first ReLU output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; repeating the second cascading process one or more times, the second cascading process including: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second rectifier linear unit (ReLU) activation process on the second convolutional output point cloud to generate a second ReLU output point cloud; and if there is a next cycle of the second cascading process, preparing the second ReLU output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process generates a second cascading process output; concatenating the first cascading process output and the second cascading process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0035] For some embodiments of the first example method, aggregating at least one feature includes: performing a self-attention process on a third point cloud; adding the third point cloud to the output of the self-attention process to generate an input to an MLP process; performing an MLP process on the input to the MLP process; and adding the input to the MLP process to the output of the MLP process to generate the aggregated feature;
[0036] For some embodiments of the first example method, the self-attention process generates output features based on the k nearest neighbors of the voxels of the third point cloud.
[0037] For some embodiments of the first example method, aggregating at least one feature of the third point cloud includes performing the feature aggregation process two or more times.
[0038] The first example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to: upsample a first point cloud using initial upsampling to obtain a second point cloud; associate the features of the second point cloud with context information to obtain a third point cloud; predict the occupancy state of at least one voxel of the third point cloud; and remove the voxels classified as empty in the third point cloud according to the predicted occupancy state to generate a trimmed point cloud.
[0039] For some embodiments of the first example apparatus, the initial upsampling includes: nearest neighbor upsampling.
[0040] For some embodiments of the first example apparatus, associating features includes: concatenating the features of the second point cloud with context information to obtain a third point cloud.
[0041] An example device according to some embodiments may include: an apparatus according to the apparatus listed above; and at least one of the following: (i) an antenna configured to receive a signal that includes data representing the image, (ii) a band limiter configured to limit the received signal to a band that includes data representing the image, or (iii) a display configured to display the image.
[0042] Some embodiments of the example device may further include at least one of a TV, a cellular phone, a tablet, and a set-top box (STB).
[0043] Example computer-readable media according to some embodiments can include instructions for causing one or more processors to perform the following operations: upsample a first point cloud using an initial upsampling to obtain a second point cloud; associate features of the second point cloud with context information to obtain a third point cloud; predict an occupancy state of at least one voxel of the third point cloud; and remove voxels classified as empty in the third point cloud based on the predicted occupancy state to generate a trimmed point cloud.
[0044] Example computer program products according to some embodiments can include instructions that, when executed by one or more processors, cause the one or more processors to: upsample a first point cloud using an initial upsampling to obtain a second point cloud; associate features of the second point cloud with context information to obtain a third point cloud; predict an occupancy state of at least one voxel of the third point cloud; and remove voxels classified as empty in the third point cloud based on the predicted occupancy state to generate a trimmed point cloud.
[0045] A second example method according to some embodiments can include performing context-aware upsampling of a first point cloud to determine an upsampled second point cloud, where the context-aware upsampling includes: associating features of a third point cloud with context information, the third point cloud being at least partially based on an initial upsampled version of the first point cloud; and removing voxels of a fourth point cloud predicted to be empty from the third point cloud at least partially based on the context information to generate an upscaled second point cloud.
[0046] A third example method according to some embodiments can include: upsampling a first point cloud using an initial upsampling to obtain a second point cloud; associating features of the second point cloud with context information to obtain a third point cloud; predicting an occupancy state of at least one voxel of the third point cloud, where predicting the occupancy state of at least one voxel includes aggregating at least one feature of the third point cloud, where aggregating at least one feature of the third point cloud includes using a first neural network, and where using the first neural network to aggregate at least one feature of the third point cloud includes using a first set of neural network parameters with the first neural network; removing voxels classified as empty in the third point cloud based on the predicted occupancy state to generate a trimmed point cloud; and performing feature aggregation on the trimmed point cloud to generate an aggregated feature, where performing feature aggregation on the trimmed point cloud includes using a second neural network, where using the second neural network to generate the aggregated feature includes using a second set of neural network parameters with the second neural network, and where the first set of neural network parameters is the same as the second set of neural network parameters.
[0047] Some embodiments of the third example method can further include aggregating at least one feature of the second point cloud.
[0048] For some embodiments of the third example method, wherein aggregating at least one feature of the second point cloud includes using a third neural network, and wherein using the third neural network to aggregate at least one feature of the second point cloud includes: using a third set of neural network parameters with the third neural network, and wherein the third set of neural network parameters is the same as the first set of neural network parameters.
[0049] For some embodiments of the third example method, the initial upsampling includes nearest neighbor upsampling.
[0050] For some embodiments of the third example method, associating features includes concatenating features of the second point cloud with context information to obtain a third point cloud.
[0051] For some embodiments of the third example method, associating features includes concatenating features of the second point cloud with context information to obtain a third point cloud.
[0052] For some embodiments of the third example method, the context information is per-voxel context information.
[0053] Some embodiments of the third example method may further include performing feature decoding on the input point cloud and the first bitstream to generate the first point cloud.
[0054] Some embodiments of the third example method may further include performing a context-aware upsampling process on the aggregated features to generate a decoded point cloud.
[0055] Some embodiments of the third example method may further include performing a feature-to-residual transformation on the trimmed point cloud to generate a residual output; and adding the trimmed point cloud to the residual output to generate a decoded point cloud.
[0056] For some embodiments of the third example method, a feature-to-residual transformation is performed on the aggregated features.
[0057] For some embodiments of the third example method, predicting the occupancy state predicts the true occupancy state of at least one voxel.
[0058] For some embodiments of the third example method, predicting the occupancy state predicts the likelihood that the at least one voxel is occupied.
[0059] For some embodiments of the third example method, voxels of the third point cloud are removed using a voxel pruning process.
[0060] For some embodiments of the third example method, predicting the occupancy state of at least one voxel further includes: processing the aggregated features with a multi-layer perceptron (MLP) layer to generate an MLP layer output; performing a softmax process on the MLP layer output to generate softmax output values; and performing thresholding of the softmax output values to generate the predicted occupancy state of at least one voxel of the third point cloud.
[0061] For some embodiments of the third example method, the thresholding of the softmax output values converts softmax output values greater than 0.5 into an output value of 1, and converts softmax output values equal to 0.5 or less into an output value of 0.
[0062] For some embodiments of the third example method, predicting the occupancy state of at least one voxel includes: aggregating at least one feature of the third point cloud; and generating a predicted occupancy state of at least one voxel of the third point cloud based on the aggregated features.
[0063] For some embodiments of the third example method, aggregating at least one feature of the third point cloud includes: repeating a concatenation process one or more times, the concatenation process including: performing a sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; performing a non-linear activation process on the convolutional output point cloud to generate a non-linear output point cloud; and if there is a next cycle of the concatenation process, preparing the non-linear output point cloud as the input point cloud, where the third point cloud is the input point cloud of the first cycle of the concatenation process, and where the last cycle of the concatenation process generates the aggregated features.
[0064] Some embodiments of the third example method may further include adding the third point cloud to the ReLU output point cloud of the last cycle of the concatenation process.
[0065] For some embodiments of the third example method, aggregating at least one feature includes: performing a sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; and performing a non-linear activation process on the convolutional output point cloud to generate the aggregated features.
[0066] For some embodiments of the third example method, the non-linear activation process includes a rectifier linear unit (ReLU) activation process, and the non-linear output point cloud includes a ReLU output point cloud.
[0067] For some embodiments of the third example method, aggregating at least one feature of the third point cloud includes: repeating the first cascaded process one or more times, the first cascaded process including: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first non-linear activation process on the first convolutional output point cloud to generate a first non-linear output point cloud; and if there is a next cycle of the first cascaded process, preparing the first non-linear output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascaded process, wherein the last cycle of the first cascaded process generates a first cascaded process output; repeating the second cascaded process one or more times, the second cascaded process including: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second non-linear activation process on the second convolutional output point cloud to generate a second non-linear output point cloud; and if there is a next cycle of the second cascaded process, preparing the second non-linear output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascaded process, wherein the last cycle of the second cascaded process generates a second cascaded process output; concatenating the first cascaded process output and the second cascaded process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0068] For some embodiments of the third example method, aggregating at least one feature includes: repeating the first cascading process one or more times, the first cascading process including: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first rectified linear unit (ReLU) activation process on the first convolutional output point cloud to generate a first ReLU output point cloud; if there is a next cycle of the first cascading process, preparing the first ReLU output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; repeating the second cascading process one or more times, the second cascading process including: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second rectified linear unit (ReLU) activation process on the second convolutional output point cloud to generate a second ReLU output point cloud; and if there is a next cycle of the second cascading process, preparing the second ReLU output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process generates a second cascading process output; concatenating the first cascading process output and the second cascading process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0069] For some embodiments of the third example method, aggregating at least one feature of the third point cloud includes: performing a self-attention process on the third point cloud; adding the third point cloud to the self-attention process output to generate an MLP process input; performing an MLP process on the MLP process input; and adding the MLP process input to the MLP process output to generate the aggregated feature;
[0070] For some embodiments of the third example method, the self-attention process generates an output feature based on k nearest neighbors of voxels of the third point cloud.
[0071] For some embodiments of the third example method, aggregating at least one feature of the third point cloud includes performing a feature aggregation process two or more times.
[0072] For some embodiments of the third example method, the first set of neural network parameters and the second set of neural network parameters are the same set of neural network parameters, and the same set of neural network parameters is used by at least the first neural network and the second neural network.
[0073] For some embodiments of the third example method, the first set of neural network parameters and the second set of neural network parameters are distinct but identical sets of neural network parameters.
[0074] A fourth example method according to some embodiments may include: obtaining a first point cloud; determining an occupancy state of at least one voxel of the first point cloud; removing voxels classified as empty in the first point cloud according to the determined occupancy state to generate a second point cloud; correlating features of the second point cloud with context information to obtain a third point cloud; downsampling the third point cloud using an initial downsampling to obtain a fourth point cloud; and outputting the fourth point cloud as an encoded point cloud.
[0075] A fourth example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to: obtain a first point cloud; determine an occupancy state of at least one voxel of the first point cloud; remove voxels classified as empty in the first point cloud according to the determined occupancy state to generate a second point cloud; correlate features of the second point cloud with context information to obtain a third point cloud; downsample the third point cloud using an initial downsampling to obtain a fourth point cloud; and output the fourth point cloud as an encoded point cloud.
[0076] A fifth example method / apparatus according to some embodiments may include: accessing data including a first point cloud; and transmitting the data including the first point cloud.
[0077] A fifth example method / apparatus according to some embodiments may include: an access unit configured to access data including a first point cloud; and a transmitter configured to transmit the data including the first point cloud.
[0078] A sixth example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any one of the methods listed above.
[0079] A seventh example method / apparatus according to some embodiments may include at least one processor configured to perform any one of the methods listed above.
[0080] An eighth example method / apparatus according to some embodiments may include a computer-readable medium storing instructions for causing one or more processors to perform any one of the methods listed above.
[0081] A ninth example method / apparatus according to some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any one of the methods listed above.
[0082] An example signal according to some embodiments may include a bitstream generated according to any one of the methods listed above.
[0083] In additional embodiments, an encoder and a decoder apparatus are provided to perform the methods described herein. The encoder or decoder apparatus may include a processor configured to perform the methods described herein. The apparatus may include a computer-readable medium (e.g., non-transitory medium) storing instructions for performing the methods described herein. In some embodiments, the computer-readable medium (e.g., non-transitory medium) stores video encoded using any one of the methods described herein.
[0084] One or more embodiments of the present invention also provide a computer-readable storage medium having stored thereon instructions for performing bidirectional optical flow, encoding or decoding video data according to any one of the methods described above. Embodiments of the present invention also provide a computer-readable storage medium having stored thereon a bitstream generated according to the methods described above. Embodiments of the present invention also provide a method and apparatus for transmitting a bitstream generated according to the methods described above. Embodiments of the present invention also provide a computer program product including instructions for performing any one of the methods described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1A is a system diagram illustrating an example communication system according to some embodiments.
[0086] Figure 1B is an illustration of an example wireless transmit / receive unit (WTRU) that may be used within the Figure 1A illustrated communication system according to some embodiments.
[0087] Figure 1C is a system diagram showing a set of example interfaces of a system according to some embodiments.
[0088] Figure 2A is a schematic diagram showing an exemplary voxel-based representation of a point cloud.
[0089] Figure 2B is a schematic diagram showing an exemplary sparse voxel-based representation of a point cloud.
[0090] Figure 3 is a schematic process diagram showing an example nearest neighbor (NN) upsampling of a point cloud.
[0091] Figure 4 is a schematic process diagram showing an exemplary voxel-based upsampling with trimming.
[0092] Figure 5 is a schematic process diagram showing an exemplary context-aware voxel-based upsampling with trimming according to some embodiments.
[0093] Figure 6A is a table showing example position values according to some embodiments.
[0094] Figure 6B is a schematic perspective view showing an example sub-voxel position as context information according to some embodiments.
[0095] Figure 7 is a flowchart showing an example process for cascading a number of context-aware upsamplings according to some embodiments.
[0096] Figure 8 is a schematic process diagram showing an exemplary context-aware voxel-based upsampling using initial feature aggregation according to some embodiments.
[0097] Figure 9 is a flowchart showing an example process for binary classification according to some embodiments.
[0098] Figure 10 is a block diagram showing an example process for feature aggregation with cascaded sparse convolutional layers according to some embodiments.
[0099] Figure 11 is a block diagram showing an example ResNet block for feature aggregation according to some embodiments.
[0100] Figure 12 is a block diagram showing an example Inception-ResNet block for feature aggregation according to some embodiments.
[0101] Figure 13 is a block diagram showing an example transformer block for feature aggregation according to some embodiments.
[0102] Figure 14 is a block diagram showing an example architecture of a self-attention module according to some embodiments.
[0103] Figure 15 is a flowchart showing an example process for cascading a number of feature aggregations according to some embodiments.
[0104] Figure 16 is a block diagram showing an example original decoder architecture according to some embodiments.
[0105] Figure 17 is a block diagram showing an example decoder architecture with voxel-based upsampling according to some embodiments.
[0106] Figure 18 is a block diagram showing an example decoder architecture with voxel-based upsampling and feature aggregation according to some embodiments.
[0107] Figure 19 is a block diagram showing an example decoder architecture without a feature-to-residual converter according to some embodiments.
[0108] Figure 20 is a block diagram showing an example decoder architecture with a single progression through voxel-based upsampling and feature aggregation according to some embodiments.
[0109] Figure 21 is a block diagram showing an example sparse tensor operation according to some embodiments.
[0110] Figure 22 is a block diagram showing an example decoder architecture according to some embodiments.
[0111] Figure 23 is a block diagram showing an example decoder architecture according to some embodiments.
[0112] Figure 24 is a flowchart showing an example process of trimmed context-aware voxel-based upsampling according to some embodiments.
[0113] Figure 25 is a flowchart showing an example process of context-aware voxel-based upsampling and feature aggregation according to some embodiments.
[0114] Figure 26 is a flowchart showing an example process of encoding a bitstream according to some embodiments.
[0115] Entities, connections, arrangements, etc. shown and described in connection with the various figures are presented by way of example and not limitation. Accordingly, any and all statements or other indications as to what is "depicted" in a particular figure, what a particular element or entity "is" or "has" in a particular figure, and any and all similar statements that can be read in isolation and out of context as absolute and thus restrictive can only be properly read as being constructively qualified with clauses such as "in at least one embodiment, ...". For the sake of brevity and clarity of presentation, this implicit leading clause is not repeated in the detailed description. Detailed Description
[0116] Figure 1A FIG. 6 is a diagram illustrating an exemplary communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multi-access system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. The communication system 100 may enable a plurality of wireless users to access such content through sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, and filter bank multicarrier (FBMC), etc.
[0117] As Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a radio access network (RAN) 104 / 113, a core network (CN) 106, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, but it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (where any one of the WTRUs may be referred to as a "station" and / or "STA") may be configured to transmit and / or receive wireless signals and may include a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smartphone, a laptop computer, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable device, a head-mounted display (HMD), a vehicle, a drone, a medical device and application (e.g., remote surgery), an industrial device and application (e.g., a robot and / or other wireless devices operating in an industrial and / or automated processing chain environment), a consumer electronic device, a device operating on a commercial and / or industrial wireless network, etc. Any one of the WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0118] The communication system 100 may further include base stations 114a and / or base stations 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks such as the CN 106, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be transceiver base stations (BTSs), Node Bs, evolved Node Bs, home Node Bs, home evolved Node Bs, gNBs, NR Node Bs, site controllers, access points (APs), wireless routers, etc. Although the base stations 114a, 114b are each depicted as a single element, it should be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0119] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of wireless services to a specific geographical area, which may be relatively fixed or may change over time. The cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In an embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0120] Base stations 114a, 114b may communicate with one or more of WTRUs 102a, 102b, 102c, 102d via air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) may be used to establish air interface 116.
[0121] More specifically, as noted above, communication system 100 may be a multiple access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA, etc. For example, base station 114a in RAN 104 / 113 and WTRUs 102a, 102b, 102c may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may use wideband CDMA (WCDMA) to establish air interface 116. WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0122] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as evolved UMTS terrestrial radio access (E-UTRA), which may use Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-A Pro to establish an air interface 116.
[0123] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as NR radio access, which may use New Radio (NR) to establish an air interface 116.
[0124] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may implement LTE radio access and NR radio access together using, for example, the dual connectivity (DC) principle. Accordingly, the air interface used by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions to / from multiple types of base stations (e.g., eNBs and gNBs).
[0125] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0126] Figure 1AThe base station 114b therein can be, for example, a wireless router, a home Node B, a home evolved Node B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area such as a commercial venue, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for drones), and a road. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As Figure 1A shown, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106.
[0127] The RAN 104 / 113 can communicate with the CN 106, which can be any type of network configured to provide voice, data, applications, and / or Internet protocol voice (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data can have different quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements, etc. The CN 106 can provide call control, billing services, location-based services, prepaid calls, Internet connectivity, video distribution, etc., and / or perform advanced security functions, such as user authentication. Although not shown in Figure 1A it, it should be understood that the RAN 104 / 113 and / or the CN 106 can communicate directly or indirectly with other RANs using the same RAT or a different RAT as the RAN 104 / 113. For example, in addition to being connected to the RAN 104 / 113 that can utilize the NR radio technology, the CN 106 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0128] CN 106 can also act as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, Internet 110, and / or other networks 112. PSTN 108 can include a circuit-switched telephone network that provides plain old telephone service (POTS). Internet 110 can include a global system of interconnected computer networks and devices that use common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet protocol suite. Network 112 can include wired communication networks and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 can include another CN connected to one or more RANs, and the one or more RANs can employ the same RAT or a different RAT as RAN 104 / 113.
[0129] Some or all of the WTRUs in communication system 100, such as WTRU 102a, 102b, 102c, 102d, can include multi-mode capabilities (e.g., WTRU 102a, 102b, 102c, 102d can include multiple transceivers for communicating with different wireless networks over different wireless links). For example, Figure 1A the illustrated WTRU 102c can be configured to communicate with a base station 114a that can employ a cellular-based radio technology and with a base station 114b that can employ IEEE 802 radio technology.
[0130] Figure 1B is a system diagram illustrating an example WTRU 102. As Figure 1B shown, WTRU 102 can include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138, etc. It should be understood that while remaining consistent with the embodiments, WTRU 102 can include any sub-combination of the foregoing elements.
[0131] The processor 118 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 can perform signal decoding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to a transceiver 120, which can be coupled to a transmit / receive element 122. Although Figure 1B the processor 118 and the transceiver 120 are depicted as separate components, it should be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0132] The transmit / receive element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via an air interface 116. For example, in one embodiment, the transmit / receive element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 can be a transmitter / detector configured to transmit and / or receive signals such as IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 can be configured to transmit and / or receive both RF signals and optical signals. It should be understood that the transmit / receive element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0133] Although the transmit / receive element 122 is depicted as a single element in Figure 1B the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 116.
[0134] The transceiver 120 can be configured to modulate the signals to be transmitted by the transmit / receive element 122 and demodulate the signals received by the transmit / receive element 122. As noted above, the WTRU 102 can have multi-mode capabilities. For example, thus, the transceiver 120 can include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs (such as NR and IEEE 802.11).
[0135] The processor 118 of the WTRU 102 may be coupled to the speaker / microphone 124, keypad 126, and / or the display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit) and may receive user input data therefrom. The processor 118 may also output user data to the speaker / microphone 124, keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may access information from any type of suitable memory (such as non-removable memory 130 and / or removable memory 132) and store data in any type of suitable memory. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, and a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from a memory that is not physically located on the WTRU 102 (such as on a server or a home computer (not shown)) and store data in that memory.
[0136] The processor 118 may receive power from a power source 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry battery packs (e.g., nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), solar cells, and fuel cells, etc.
[0137] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) via an air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that the WTRU 102 may obtain location information by any suitable location determination method while remaining consistent with the embodiments.
[0138] The processor 118 may also be coupled to other peripheral devices 138, which may include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, Modules, FM radio units, digital music players, media players, video game player modules, Internet browsers, virtual reality and / or augmented reality (VR / AR) devices, and activity trackers, etc. The peripheral device 138 may include one or more sensors, and the sensor may be one or more of the following: gyroscope, accelerometer, Hall effect sensor, magnetometer, orientation sensor, proximity sensor, temperature sensor, time sensor; geographical location sensor; altimeter, light sensor, touch sensor, magnetometer, barometer, gesture sensor, biometric sensor, and / or humidity sensor.
[0139] The WTRU 102 may include a full-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with specific subframes for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference via signal processing performed by hardware (e.g., chokes) or via a processor (e.g., a separate processor (not shown) or via the processor 118). In an embodiment, the WTRU 102 may include a half-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)).
[0140] Although the WTRU is described as a wireless terminal in Figures 1A to 1B it is contemplated that in certain representative embodiments, such a terminal may (e.g., temporarily or permanently) use a wired communication interface with a communication network.
[0141] In a representative embodiment, the other network 112 may be a WLAN.
[0142] In view of Figures 1A to 1B and the corresponding description, one or more or all of the functions described herein may be performed by one or more emulation devices (not shown). The emulation device may be one or more devices configured to mimic one or more or all of the functions described herein. For example, the emulation device may be used to test other devices and / or simulate network and / or WTRU functions.
[0143] A simulation device can be designed to implement one or more tests of other devices in a laboratory environment and / or in an operator network environment. For example, one or more simulation devices can perform one or more functions or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more functions or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The simulation device can be directly coupled to another device for testing purposes and / or can perform tests using over-the-air wireless communication.
[0144] One or more simulation devices can perform one or more (including all) functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the simulation device can be used in a test laboratory and / or in a test scenario of a non-deployed (e.g., test) wired and / or wireless communication network to implement tests of one or more components. One or more simulation devices can be test equipment. Direct RF coupling and / or wireless communication via an RF circuit (e.g., which can include one or more antennas) can be used by the simulation device to transmit and / or receive data.
[0145] Figure 1C FIG. is a system diagram showing a set of example interfaces of a system according to some embodiments. An extended reality display device and its control electronics can be implemented. System 150 can be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of System 150 can be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing element and encoder / decoder element of System 150 are distributed across multiple ICs and / or discrete components. In various embodiments, System 150 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, System 1000 is configured to implement one or more of the aspects described in this document.
[0146] System 150 includes at least one processor 152 configured to execute instructions loaded therein for implementing various aspects as described, for example, in this document. Processor 152 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 150 includes at least one memory 154 (e.g., volatile memory devices and / or non-volatile memory devices). System 150 may include a storage device 158, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 158 may include internal storage devices, attached storage devices (including detachable and non-detachable storage devices), and / or network-accessible storage devices.
[0147] System 150 includes an encoder / decoder module 156 configured to, for example, process data to provide encoded video or decoded video, and encoder / decoder module 156 may include its own processor and memory. Encoder / decoder module 156 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding module and a decoding module. Additionally, encoder / decoder module 156 may be implemented as a separate element of system 150, or may be incorporated within processor 152 as a combination of hardware and software known to those skilled in the art.
[0148] The program code to be loaded onto processor 152 or encoder / decoder 156 to execute the various aspects described in this document may be stored in storage device 158 and subsequently loaded onto memory 154 for execution by processor 152. According to various embodiments, one or more of processor 152, memory 154, storage device 158, and encoder / decoder module 156 may store one or more of the various items during the execution of the processes described in this document. Such stored items may include but are not limited to input video, decoded video or partially decoded video, bitstreams, matrices, variables, and intermediate or final results of processing equations, formulas, operations, and operation logic.
[0149] In some embodiments, the memory internal to the processor 152 and / or the encoder / decoder module 156 is used to store instructions and provide working memory for processing during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 152 or the encoder / decoder module 152) is used for one or more of these functions. Such external memory can be the memory 154 and / or the storage device 158, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, fast external dynamic volatile memory (such as RAM) is used as the working memory for video encoding and decoding operations, such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by the Joint Video Exploration Team (JVET)).
[0150] Input to the elements of the system 150 can be provided through various input devices as indicated in block 172. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, an RF signal transmitted over the air by a broadcaster, (ii) component (COMP) input terminals (or a set of COMP input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 1C Other examples not shown include composite video.
[0151] In various embodiments, the input device of block 172 has corresponding input processing elements associated therewith as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal band to one band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) de-multiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down-converter, a demodulator, an error corrector, and a de-multiplexer. The RF section may include a tuner that performs various functions of these, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band. Various embodiments re-order the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0152] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 150 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, if necessary, for example, within a separate input processing IC or within processor 152. Similarly, various aspects of USB or HDMI interface processing may be implemented, if necessary, within a separate interface IC or within processor 152. The demodulated, error-corrected, and de-multiplexed stream is provided to various processing elements, including, for example, processor 152 and encoder / decoder 156, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0153] The various elements of system 150 may be disposed within an integrated housing in which the various elements may be interconnected using a suitable connection arrangement 174 (e.g., an internal bus known in the art, including an inter-chip (I2C) bus, wiring, and a printed circuit board) and data may be transmitted between these elements.
[0154] System 150 includes communication interface 160 that enables communication with other devices via communication channel 162. Communication interface 160 can include, but is not limited to, a transceiver configured to send and receive data over communication channel 162. Communication interface 160 can include, but is not limited to, a modem or a network card, and communication channel 162 can be implemented, for example, within wired and / or wireless media.
[0155] In various embodiments, data is streamed or otherwise provided to system 150 using a wireless network such as, for example, a Wi-Fi network, such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received via communication channel 162 and communication interface 160 suitable for Wi-Fi communication. The communication channel 162 of these embodiments is typically connected to an access point or a router that provides access to an external network, including the Internet, for allowing streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streamed data to system 150, and the set-top box delivers data via an HDMI connection of input box 172. Still other embodiments use an RF connection of input box 172 to provide streamed data to system 150. As described above, various embodiments provide data in a non-streamed manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0156] System 150 can provide output signals to various output devices, including display 176, speaker 178, and other peripheral devices 180. The display 176 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 176 can be used for a television, a tablet, a laptop, a cellular phone (mobile phone), or other devices. The display 176 can also be integrated with other components (e.g., as in a smart phone) or be standalone (e.g., an external monitor for a laptop). In various examples of embodiments, other peripheral devices 180 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, applicable to both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 180 that provide functions based on the output of system 150. For example, a disc player performs the function of playing the output of system 150.
[0157] In various embodiments, the control signal is communicated between the system 150 and the display 176, the speaker 178, or other peripheral devices 180 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to the system 1000 via dedicated connections through respective interfaces 164, 166, and 168. Alternatively, the output devices may be connected to the system 150 via the communication interface 160 using the communication channel 162. The display 176 and the speaker 178 may be integrated in a single unit with other components of the system 150 in an electronic device such as, for example, a television. In various embodiments, the display interface 164 includes a display driver such as, for example, a timing controller (T Con) chip.
[0158] For example, if the RF portion of the input 172 is part of a stand-alone set-top box, the display 176 and the speaker 178 may alternatively be independent of one or more of the other components. In various embodiments where the display 176 and the speaker 178 are external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0159] The system 150 may include one or more sensor devices 168. Examples of sensor devices that may be used include one or more GPS sensors, gyroscopic sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors may be used to determine information such as the location and orientation of a user. In cases where the system 150 serves as a control module (such as control modules 124, 132) for an extended reality display, the location and orientation of the user may be used to determine how to render the image data such that the user perceives the correct portion of a virtual object or virtual scene from the correct viewing point. In the case of a head-mounted display device, the location and orientation of the device itself may be used to determine the location and orientation of the user for the purpose of rendering virtual content. In the case of other display devices such as a phone, a tablet, a computer monitor, or a television, other inputs may be used to determine the location and orientation of the user in order to render the content. For example, the user may use a touch screen, a keypad or keyboard, a trackball, a joystick, or other inputs to select and / or adjust the desired viewing point and / or viewing direction. In cases where the display device has sensors such as accelerometers and / or gyroscopes, the viewing point and orientation for rendering the content may be selected and / or adjusted based on the movement of the display device.
[0160] These embodiments can be executed by computer software implemented by the processor 152, or by hardware, or by a combination of hardware and software. As a non-limiting example, these embodiments can be implemented by one or more integrated circuits. As a non-limiting example, the memory 154 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory). As a non-limiting example, the processor 152 can be of any type suitable for the technical environment and can include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0161] This application discusses point cloud processing and compression, which includes the processing, compression, representation, analysis, and understanding of point cloud signals. In addition, this application discusses an adaptive, voxel-based point cloud upsampling method based on a deep neural network, which includes some embodiments applied to point cloud processing and compression. This application also discusses performing upsampling in the voxel domain.
[0162] For example, in cars connected via a 5G network and in immersive (e.g., VR / AR / MR) communications, point cloud data can consume a large amount of network traffic. An efficient representation format can be used for point clouds and communications. Specifically, the raw point cloud data can be organized and processed for modeling and sensing of, for example, the world, environment, or scene. Compression of the raw point cloud can be used for storage and transmission of the data.
[0163] In addition, a point cloud can represent sequential scans of the same scene that may contain multiple moving objects. A dynamic point cloud captures moving objects, while a static point cloud captures a static scene and / or static objects. A dynamic point cloud is typically organized into frames, where different frames are captured at different times. Processing and compression of the dynamic point cloud can be performed in real time or with a small delay.
[0164] The automotive industry (including, for example, autonomous vehicles) can use point clouds. An autonomous vehicle "senses" its environment to make driving decisions based on its surroundings. Typically, a LiDAR sensor generates a (dynamic) point cloud for use by a perception engine. In addition, typically these point clouds are dynamic, with a high capture frequency, sparse, not necessarily colored, and not visible to the human eye. Such point clouds can include other attributes, such as the reflectivity provided by LiDAR, which can indicate the material of the sensed object and can be used for decision-making.
[0165] The automotive industry and autonomous vehicles are some of the areas where point clouds can be used. Autonomous vehicles "detect" and sense their environment to make good driving decisions based on the reality of their immediate surroundings. Sensors such as LiDAR generate the (dynamic) point clouds used by the perception engine. These point clouds are generally not intended to be viewed by the human eye, and these point clouds can be or can not be colored and are typically sparse and dynamic, with a very high capture frequency. Such point clouds can have other properties, such as reflectivity provided by LiDAR, as this property indicates the material of the sensed object and can help in making decisions.
[0166] Virtual reality (VR) and immersive worlds have become a hot topic and are seen by many as the future of 2D flat video. The viewer can be immersed in an all-round environment, which is contrary to a standard TV where the viewer can only see the virtual world in front of the viewer. Depending on the degree of freedom of the viewer in the environment, there are several levels of immersion. The point cloud format can be used to distribute VR world and environment data. Such point clouds can be static or dynamic and typically have an average size, such as less than millions of points at a time.
[0167] Point clouds can also be used for a variety of other purposes, such as scanning cultural heritage objects and / or buildings, where objects such as statues or buildings are scanned in 3D form. The spatial configuration data of the object can be shared without transporting or accessing the actual object or building. Moreover, this data can be used to preserve the knowledge of the object in case the object or building is damaged (such as a temple destroyed by an earthquake). Such point clouds are typically static, colored, and very large.
[0168] Another use case is in topographic maps and cartography using 3D representations, where the map is not limited to a flat surface and can include terrain undulations. For example, Google Maps uses a mesh instead of a point cloud for its 3D maps. However, point clouds can be a suitable data format for 3D maps, and such point clouds are also typically static, colored, and very large.
[0169] World modeling and sensing via point clouds can allow machines to record and use spatial configuration data about the 3D world around them, which can be used in the applications discussed above.
[0170] 3D point cloud data consists of discrete samples of the surface of an object or scene. To adequately represent the real world with point samples, a large number of points can be used. For example, a typical VR immersive scene includes millions of points, while point clouds can generally include hundreds of millions of points. Therefore, processing such large-scale point clouds is computationally expensive, especially for consumer devices with limited computing power, such as smartphones, tablets, and automotive navigation systems.
[0171] To store and process an input point cloud at an affordable computational cost, the input point cloud can be downsampled, where the downsampled point cloud summarizes the geometry of the input point cloud while having far fewer points. The downsampled point cloud is then input into subsequent machine tasks for further processing. The downsampled point cloud can be processed by gradually upsampling the point cloud. Specifically, for some embodiments, a learning-based autoencoder architecture can use downsampling for feature extraction and upsampling for reconstruction. For example, such upsampling can be used in conjunction with point cloud compression (e.g., on the decoder) and with point cloud super-resolution. Among other things, this application discusses examples of adaptive point cloud upsampling methods according to some embodiments.
[0172] Figure 2A is a schematic diagram showing an exemplary voxel-based representation of a point cloud. In the voxel-based representation of point cloud data, 3D point coordinates are uniformly quantized through a quantization step. Each point in the representation 200 corresponds to an occupied voxel, the size of which is equal to the quantization step, as Figure 2A shown. The "natural" voxel representation may not be efficient in terms of memory usage because most voxels are typically empty. Then a sparse voxel representation is introduced, where, for efficient storage and processing, the occupied voxels are arranged in a sparse tensor format.
[0173] Figure 2B is a schematic diagram showing an exemplary sparse voxel-based representation of a point cloud. In the Figure 2B example of the sparse voxel representation 250 shown, the empty voxels 252 (represented by dashed lines) do not necessarily consume as much memory or storage as the occupied voxels 254 (with solid diagonals).
[0174] Note that Figure 2A and 2B and including Figure 3 , 4 , 5, 8, and 9 in the remaining figures, the point cloud is shown in 2D only for purposes of explanation and simplification, and these concepts can generally be applied to and extended to 3D.
[0175] By representing a point cloud as 3D voxels, a 3D convolutional neural network can be utilized to process (decompose) the point cloud. Applying a 2D convolutional neural network to 2D images has been successful. With conventional 3D convolution, the 3D kernel is covered at each position specified by the stride step, regardless of whether the voxel is occupied or empty. The stride represents the amount of movement or step on the 3D voxel grid when applying the convolution. If the stride is set to 1, the 3D kernel typically slides over each voxel in 3D space to compute the output. In this case, the dimensions (height, width, depth) of the output voxel space are the same as those of the input voxel space. If the stride step is set to 2, the 3D kernel slides every two voxels to compute the output. In this case, each dimension of the output voxel space becomes half of that of the input voxel space. An empty voxel is a voxel where there is no 3D point at the position of the voxel. To avoid the computational and memory consumption caused by empty voxels, if the point cloud voxels are represented by a sparse tensor, a sparse 3D convolutional layer can be applied.
[0176] Figure 3 is a schematic process diagram showing an example of nearest neighbor (NN) upsampling of a point cloud. To perform, for example, two-times (2X) upsampling in the voxel domain, the occupied voxels are divided into two voxels along all x, y, and z directions. Thus, after upsampling, one (occupied) parent voxel becomes 8 (= 2 3 ) occupied child voxels, and the resolution of the point cloud is upsampled quadratically in each dimension (x, y, and z). According to this example method, if there is any feature vector associated with the parent voxel, that feature vector will be directly inherited by its 8 child voxels. The mechanism of this upsampling method is described in Figure 3 and is called nearest neighbor (NN) upsampling. For Figure 3 example, after the input point cloud 302 is NN upsampled 304, the output point cloud 306 and its features are input into the subsequent pipeline for additional processing.
[0177] This NN upsampling method has some drawbacks. First, after upsampling, the geometry of the point cloud is a natural magnification of the original geometry, i.e., there is no refinement in terms of the shape of the point cloud. Second, the number of occupied voxels after upsampling is always eight times that of the original voxels, which may cause a high computational cost in subsequent processing.
[0178] Figure 4is a schematic process diagram showing an exemplary voxel-based upsampling with pruning. In the article "Multiscale Point Cloud Geometry Compression" by Wang, Jianqiang et al., 2021 DATA COMPRESSION CONFERENCE (DCC) 73-82, IEEE (2021) ("Wang"), the authors proposed a method to further improve the upsampled geometry through binary classification and voxel pruning. Figure 4 shows Wang's method.
[0179] According to the exemplary method 400, the input point cloud PC 0 402 is first upsampled using the NN upsampling block 404 (as Figure 4 shown), which produces the initially upsampled point cloud PC 1 406. PC 1 is input into a neural network-based binary classifier 408, which determines the occupancy status 410 of each occupied voxel in PC 1 . By removing all voxels classified as unoccupied ( Figure 4 the "0" in), the initially upsampled point cloud PC 1 is pruned 412. The refined upsampled PC 2 414 is the output. See the pruning example in Figure 21 .
[0180] This method addresses the above two drawbacks of the Figure 3 shown NN upsampling method. However, in many applications, the success of the binary classifier is crucial for the method to be an accurate geometric refinement. According to some embodiments, the present application improves the performance of binary classification by introducing additional voxel context information.
[0181] According to Figure 4 the feature information of the method in is a high-level abstract descriptor of the geometry generated by a deep neural network. In some cases, such feature information may be too abstract and insufficient to perform accurate classification. Instead, the context information introduced in the present application can include local knowledge of each voxel, which can be a useful clue for binary classification.
[0182] Figure 5 is a schematic process diagram showing an exemplary context-aware voxel-based upsampling with pruning according to some embodiments. The present application introduces a context point cloud carrying additional known information about the initial ("naturally") upsampled point cloud PC 1 . By combining the context point cloud 510 with the upsampled point cloud PC 1The 506 - level cascade 512 (association), subsequent binary classification 516, and voxel pruning 520 processes can better refine the initial up - sampled point cloud PC 1 506, which can result in a more accurate up - sampled point cloud PC 2 522. For some embodiments, the voxel pruning process 520 takes the up - sampled point cloud PC 1 506 and the point cloud PC of the binary classification 1 518 as inputs to generate the output point cloud PC 2 522. Figure 5 The block diagram shows the voxel - based up - sampling method 500 of some embodiments. The voxel - based up - sampling method performs classification 516 and pruning 520 to refine the up - sampled geometry on the point cloud PC 1 (e.g., generated by the NN up - sampling 504 of the input point cloud PC 0 502). In Figure 5 context, a context point cloud PC CTX 510 carrying context information is introduced. The context building block 508 takes the initial up - sampled point cloud PC 1 506 as an input and outputs the context point cloud PC CTX 510. The context point cloud PC CTX 510 is cascaded with the "naturally" up - sampled point cloud PC 1 506 to generate an augmented point cloud 514 as the input to the binary classification stage 516. According to this example, the context point cloud PC CTX 510 shares the same geometry with PC 1 506, while PC CTX 510 is designed to include context information (e.g., voxel - wise discrimination information) for predicting the true occupancy state (in this case, for voxels, whether the sub - voxels are "empty" or "occupied"). By cascading 512 the features from PC CTX 510 and PC 1 506 to produce the augmented point cloud PC' 1 514.
[0183] Cascading is a commonly used operator in deep neural networks. The cascading operator cascades (all) the features in PC CTX with the corresponding features in PC 1 to generate the augmented point cloud PC' 1 . If the occupied voxel (x, y, z) in PC CTX has a context information vector c of length a, and PC 1At the same position (x, y, z) therein, there is a related feature vector f of length b 1 , then the concatenation operator will combine c and f 1 by concatenation and generate another feature vector [c f 1 of length (a + b). The generated feature vector [c f 1 will be assigned to the voxel position (x, y, z) of the augmented point cloud PC’ 1 . This step can be performed for all occupied voxels in PC CTX and PC 1 to generate the augmented point cloud PC’ 1 . This augmented point cloud PC’ 1 replaces PC 1 as the input to the binary classifier.
[0184] According to some embodiments, the context information can be any known knowledge or known context about the voxel. For example, the context information can be the position of the voxel, such as [x, y, z] coordinates. The context information can include, for example, the bit depth of the input point cloud. The context information can be other information, such as, for example, the relative position of the voxel with respect to the parent voxel. However, the context is not limited to position information, and in some embodiments, other types of information can be included in addition to or instead of position information. In some embodiments, the context information (e.g., voxel-by-voxel discrimination information) can be used to predict the true occupancy state of the voxel (e.g., whether the voxel is "empty" or "occupied") because the context information can provide some known information about the voxel. By incorporating such known context information, the deep neural network can better infer the occupancy state.
[0185] For some embodiments, the context information is derived from the processing of the input point cloud. The context information can be any information about the voxel and can even be determined before encoding the voxel. For example, assume that the position [x, y, z] of the voxel and the current bit depth information for the [x, y, z] voxel are known. In this way, the context information can be represented in spherical coordinates. See Equations 1, 2, and 3 below.
[0186] The context information PC CTX is from the input point cloud PC 0 . According to some embodiments, there is no other external source of context information. For example, the context information can be the (x, y, z) position coordinates. The context information vector c = [x y z] can be directly assigned to the voxel position (x, y, z) in PC CTX . For some embodiments, the (x, y, z) coordinate position can be preprocessed and transformed to another coordinate system, such as by Eq. 1 to 3 as described below, and assigned to PC CTXThe voxel position (x, y, z) in
[0187] In some embodiments, the context information includes x, y, and z coordinates. For example, for a PC CTX For an occupied voxel (x, y, z) in 1 having a bit depth N, which means that the PC 1 has a dimension of 2 N x 2 N x 2 N . Thus, the context information vector associated with the voxel (x, y, z) in the PC CTX can be
[0188] In addition, instead of working with Euclidean coordinates, spherical coordinates can be used, which are particularly useful for processing LiDAR scans. For this purpose, Eqs. 1, 2, and 3 can be applied, which are: where N is the bit depth, r is the radial distance, is the elevation angle, and θ is the azimuth angle. The vector c becomes or (if the distance is normalized). In some embodiments, the context information can also be the bit depth of the PC 1 which is N. In this case, the feature is a constant scalar c = N. In addition, instead of working with Euclidean coordinates or spherical coordinates, cylindrical coordinates can be used because cylindrical coordinates can be used to process LiDAR scans. For this purpose, Eqs. 1 and 3 are applied to calculate the radial distance (r) and the azimuth angle (θ) respectively. The Euclidean coordinates of the voxel (x, y, z) are converted to cylindrical coordinates (r, θ, z). Thus, the vector c that holds the context information becomes c = [r θ z] or (if the distance is normalized). Providing context information in different ways can make the binary classification process work more easily. For example, in some cases, if the height z is particularly relevant to the occupancy state of the voxel, including the height z can help with classification. Another example is a LiDAR point cloud, in which case the height is within a reasonable range, e.g., the height is greater than zero because LiDAR cannot sense something underground. And in some embodiments, if the azimuth angle θ (Eq. 3) may be highly relevant to occupancy, including the azimuth angle can help with classification.
[0189] For some embodiments, the encoder can be at Figure 5The encoder may be configured to operate in the opposite direction shown. For example, the encoder may obtain a first point cloud, such as the pruned point cloud 522. The voxel occupancy state of the point cloud may be determined. A second point cloud may be generated by removing voxels determined to be empty from the first point cloud. Features of the second point cloud may be determined and associated with context information to generate a third point cloud. For some embodiments, the features of the second point cloud may be concatenated with the context information. The third point cloud may be downsampled to obtain a fourth point cloud. The fourth point cloud may be output as an encoder output.
[0190] Figure 6A is a table showing example position values according to some embodiments. In some embodiments, the context information may be the position of a child voxel relative to its parent voxel. For example, "front" / "back", "left" / "right", and "top" / "down" may be represented as "0" and "1", respectively, as shown in FIG. Figure 6A 600, 602, 604. In other words, the leftmost value of the feature array indicates the front / back state, the middle value indicates the left / right state, and the rightmost value indicates the up / down state. A zero for the front / back state indicates front, while a one for the front / back state indicates back. A zero for the left / right state indicates left, while a one for the left / right state indicates right. A zero for the up / down state indicates up, while a one for the up / down state indicates down.
[0191] Figure 6B 650 is a schematic perspective diagram showing example child voxel positions as context information according to some embodiments. 1 The child voxel of will have 3 bits of context information c = [0 10], while the PCs behind, to the right and above its parent voxel 650 1 The subvoxel of will have a 3-bit context feature 652c = [11 0], such as Figure 6B In some embodiments, the 3-bit context feature vector may be interpreted as a binary number and converted to a decimal number, for example, c=[0 1 0] becomes a scalar c=2 and c=[1 1 0] becomes a scalar c=6. In some embodiments, the context information may be any portion, combination and / or arrangement of the aforementioned example context features.
[0192] Back to Figure 5 , consider about PC 0 The upper left corner feature f 1 The following example will follow f 1 From the input point cloud PC 0 Expanded point cloud PC′ before binary classification block 1 .
[0193] PC 0 The characteristic f 1and other features f 2 , f 3 , …, f 8 are each 1 - dimensional vectors. For this non - restrictive example, the vector length will be 5. Thus, each feature vector f 1 , f 2 , …, f 8 will be a 1×5 vector with 5 numbers. These feature vectors contain information about the true occupancy status of the PC 2 , and are thus called geometric features. For example, f 1 contains information about the true occupancy status of the upper - left corner of the PC 2 ; while f 4 contains information about the true occupancy of the upper - right corner of the PC 2 .
[0194] In this example, f 1 , f 2 , …, f 8 are abstract high - level features / descriptors generated by a deep neural network. For this example, the values of f 1 , f 2 , …, f 8 do not have a specific physical meaning and are thus abstract and “high - level”. However, these values provide meaning for the neural network itself to perform inference. Assume these vector features are random numbers, for example: f 1 = [2.1 0.3 1.23 4.5 0.1], which is chosen to show the movement of f 1 .
[0195] As part of the PC 0 , f 1 is passed through an NN upsampling block which creates 4 copies of f 1 at the upper - left corner of the PC 1 , as shown in Figure 5 . The PC 1 is passed to a context - building block which creates four context - information vectors corresponding to f CTX at the upper - left corner of the PC 1 . These four corresponding context - information vectors are c 11 , c 12 , c 13 and c 14 . Assume the context - building block uses the coordinates (x, y, z) of the voxels (or just (x, y) in the 2D example of Figure 5 ) to build the context - information vectors, then c 11 = [0 0], c 12 = [1 0], c 13 = [0 1], c 14 = [1 1],
[0196] PC CTX and PC 1 are input into a cascading block, thereby generating a point cloud PC'. 1 For some embodiments, the cascading block can perform the following exemplary cascading. · c 11 is cascaded with f 1 to generate a new vector: [c 11 f 1 = [0 0 2.1 0.3 1.23 4.5 0.1], which is the voxel at the upper left corner (first row, first column) of PC'. 1 of the upper left corner (first row, first column). · c 12 is cascaded with f 1 to generate a new vector: [c 12 f 1 = [1 0 2.1 0.3 1.23 4.5 0.1], which is the voxel within the first row and second column of PC'. 1 of the first row, second column. · c 13 is cascaded with f 1 to generate a new vector: [c 13 f 1 = [0 1 2.1 0.3 1.23 4.5 0.1], which is the voxel within the second row and first column of PC'. 1 of the second row, first column. · c 14 is cascaded with f 1 to generate a new vector: [c 14 f 1 = [1 1 2.1 0.3 1.23 4.5 0.1], which is the voxel within the second row and second column of PC'. 1 of the second row, second column. The resulting point cloud PC' 1 is input into a binary classification block, and based on the determined true point cloud PC'', 1 it is determined whether PC' 1whether a voxel claimed to be occupied is actually occupied.
[0197] Figure 7 is a flowchart showing an example process for cascading several context-aware upsamplings. In some embodiments of the exemplary process 700, context-aware voxel-based upsamplings 702, 706, 710 can be cascaded multiple times to achieve a higher upsampling rate, as Figure 7 shown. In this case, between two consecutive upsamplings 702, 706, 710, feature aggregation blocks 704, 708, 712 can be inserted for refinement and feature aggregation. For example, a feature aggregation block can take as input a sparse tensor of features with N channels. The feature aggregation block modifies the features to better serve the compression task. In particular, to obtain a high-quality reconstruction for point cloud decompression, the feature aggregation block generates descriptive or discriminative geometric features that can represent local geometric details. The output features still have N channels, which means that the feature aggregation module does not change the shape of the sparse tensor. For some embodiments, the feature aggregation can be, for example, a cascaded sparse convolutional layer architecture, a Residual Network (ResNet) architecture, an Initial ResNet (IRN) architecture, or a Transformer block. For some embodiments, the context-aware upsampling can be performed by Figure 5 the process shown.
[0198] For some embodiments, Figure 7 the feature aggregation block shown in Figure 7 can include a weight sharing mechanism. In Figure 7 , multiple context-aware blocks are cascaded (in series). For some embodiments,
[0199] Figure 8 is a schematic process diagram showing an example context-aware voxel-based upsampling using initial feature aggregation. In some embodiments, a feature aggregation block 806 can be inserted right after the NN upsampling block 804 for initial feature refinement, as Figure 8 shown. The PC 0 802 features are continued to the PC 1 808 features based on the NN upsampling 804 assignment. Additionally, a cascading block 814 is inserted before the binary classification block 818, similar to Figure 5 shown. For some embodiments, Figure 8 the block of Figure 5 operates the same as the block described in
[0200] As Figure 5 shown, the context construction block 810 takes the initial upsampled point cloud PC1 takes 808 as input and outputs the context point cloud PC CTX 812. By cascading the context point cloud 812 with the upsampled point cloud PC 1 808 to 814, subsequent binary classification 818 and voxel pruning 822 processes can refine the initial upsampled point cloud PC 1 808, which can produce a more accurate upsampled point cloud PC 2 824. By cascading the features from PC CTX 812 and PC 1 808, an augmented point cloud PC’ 1 816 is produced. For some embodiments of the example process 800, the voxel pruning process 822 takes the upsampled point cloud PC 1 808 and the binary classification point cloud PC” 1 820 as input and generates the output point cloud PC 2 824.
[0201] Figure 9 is a flowchart showing an example process for binary classification according to some embodiments. Figure 9 can be considered a way to perform binary classification using a neural network. The binary classifier can be used to predict the true occupancy status of each occupied voxel in the input point cloud (PC 1 ). The binary classifier classifies each occupied voxel in PC 1 as 1 (occupied) or 0 (empty), so that the geometry of PC 1 can be refined. For some embodiments, Figure 9 the binary classification process 900 of Figure 5 and 8 can be used for the binary classification block of
[0202] In Figure 9 , the cascaded point cloud PC’ 1 is input into a feature aggregation block for feature refinement and extraction, with an output channel size of D 1 . The input point cloud undergoes feature aggregation 902. Then, the aggregated features are input into a multi-layer perceptron (MLP) layer 904 with channel dimensions (D 1 , D 2 , …, 1) for classification. The MLP layer is a neural network layer that applies a linear mapping to the input feature vector. For example, to map an input feature of length D 1 to an output feature of length D 2 , the MLP layer multiplies a matrix of size D 1 ×D 2 by the input feature to obtain an output feature of length D 2Output features. When cascading two MLP layers, a non-linear activation function (e.g., ReLU function) is inserted between the two MLP layers. The output is input to the softmax function 906, which converts the MLP output values to the range from 0 to 1. For values above 0.5 up to 1, the thresholding block 908 converts the value to 1 to indicate an occupied state. For values in the range 0 to 0.5, the thresholding block converts the value to 0 to indicate an empty state as reflected in the binary classification output 910.
[0203] Figures 10 - 13 Four different design choices for feature aggregation of some embodiments are shown. For example, in some embodiments, the feature aggregation block can be a cascaded sparse convolutional layer architecture 1000 (e.g., Figure 10 ), a Residual Network (ResNet) architecture 1100 (e.g., Figure 11 ), an Initial-ResNet (IRN) architecture 1200 (e.g., Figure 12 ), or a Transformer block architecture 1300 (e.g., Figure 13 ).
[0204] Figure 10 is a block diagram showing an example process with cascaded sparse convolutional layers for feature aggregation according to some embodiments. For some embodiments, for example, Figure 10 in the example shown, two blocks are repeated multiple times to form a series. These two blocks are sparse 3D convolutional layers 1002, 1006, 1010 ("CONV D"), followed by ReLU activations 1004, 1008, 1012 ("ReLU"). For Figure 10 the example shown, "CONV D" represents a sparse 3D convolutional layer with D output channels. The "ReLU" activation refers to the rectifier linear unit activation function. For example, the ReLU activation block can output 0 for negative input values and can output the input multiplied by a scalar value for positive input values. In another embodiment, the ReLU activation function can be replaced by other activation functions, such as the tanh() activation function and / or the sigmoid() activation function. For some embodiments, the non-linear activation process can include a rectifier linear unit (ReLU) activation process.
[0205] Figure 11 is a block diagram showing an example ResNet block for feature aggregation according to some embodiments. In some embodiments, the feature aggregation process can use a ResNet architecture, as Figure 11As shown in the article by He, Kaiming et al., "Deep Residual Learning for Image Recognition," PROCEEDINGS OF THE IEEE CONF. ON COMPUTER VISION AND PATTERN RECOGNITION 770 - 778, IEEE (2016) ("He"), an exemplary ResNet architecture is described. For example, see the rightmost processing line on page 4 of He Figure 3 The rightmost processing line.
[0206] Figure 11 The example in shows a ResNet block architecture for aggregating features through D channels. For some embodiments, for example Figure 11 In the example shown in, two blocks are repeated multiple times to form a series. These two blocks are sparse 3D convolutional layers 1102, 1106, 1110 ("CONV D"), followed by ReLU activations 1104, 1108, 1112 ("ReLU"). Compared with Figure 10 Compared with Figure 11 A residual connection 1114 is introduced to the input to add the input to the output of the series of convolutional layers.
[0207] Figure 12 is a block diagram showing an example initial - ResNet block for feature aggregation according to some embodiments. For certain embodiments, feature aggregation can be constructed with an initial - ResNet (IRN) architecture, as Figure 12 shown. Figure 1(b) of Wang also shows the IRN architecture. Figure 12 The example of shows the architecture of an IRN block for aggregating features through D channels. The IRN block divides the feature aggregation process into three parallel paths.
[0208] The path with more convolutional layers ( Figure 12 The left path in) aggregates (more) global information with a larger receptive field. For some embodiments, the left path may include convolutional layer 1202, followed by two sets of ReLU activations 1204, 1208 and convolutional layers 1206, 1210.
[0209] The path with fewer convolutional layers ( Figure 12 The middle path in) aggregates local detailed information with a smaller receptive field. For some embodiments, the middle path may include convolutional layer 1212, followed by ReLU activation 1214 and convolutional layer 1216.
[0210] The last path 1220 ( Figure 12 The right path in) is a residual connection that brings the input directly to the output, similar toFigure 11 The residual connection in. For some embodiments, it may be possible to insert ReLU blocks after the CONV D / 2 blocks and before the concatenation 1218 on each of the left and middle paths of Figure 12 .
[0211] Figure 13 is a block diagram showing an example transformer block for feature aggregation according to some embodiments. Section 3.2 of the article "Voxel Transformer for 3D Object Detection" by Mao, Jiageng et al., PROCEEDINGS OF THE IEEE / CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION 3164 - 3173, IEEE (2021) ("Mao") discusses a voxel transformer. For some embodiments, the transformer architecture of the present application may be similar to the voxel transformer of Mao.
[0212] Figure 13 shows a diagram of a transformer block. The output of the self - attention block 1302 is added to the input of the self - attention block through a residual connection. Connected in series with the result of this addition is the MLP block 1304, where the output of the MLP block 1304 is added to the input of the MLP block 1304 via another residual connection. The MLP block includes a series of multi - layer perceptron (MLP) layers.
[0213] Figure 14 is a block diagram showing an example architecture of a self - attention block according to some embodiments. For some embodiments, Figure 13 the self - attention block 1302 of Figure 14 can be implemented using the architecture 1400 described in A Given the current feature vector f associated with the voxel position A i and its k adjacent features associated with the voxel position A where A i are the k nearest neighbors of A in the input sparse tensor, where 0 ≤ i ≤ (k - 1), the self - attention block 1400 attempts to update the feature f based on all the adjacent features A . These points A i are obtained through a k - nearest neighbor (kNN) search 1402 based on the coordinates of A. The query 1404 representing Q for A A is calculated by Eq.4: Q A = MLP Q (f A ) Eq.4 where MLP Q(·)1404 represents the MLP layer that obtains the said query. is the positional encoding 1410 between voxel A and A i and it is through Eq.5: where MLP P (·)1410 represents the MLP layer used to obtain the said positional encoding. P A and are the 3D coordinates of the centers of voxel A and A i , respectively. The key of voxel A and the values of all the nearest neighbors of voxel A are calculated using Eq.6 and 7. where MLP K (·)1406 and MLP V (·)1408 are the MLP layers to obtain the said key and value respectively. The self-attention block outputs the output feature f′ of position A given by Eq.8 A . where σ(·) is the softmax normalization function 1412, d is the length of the feature vector f A and c is a predefined constant.
[0214] Figure 14 Eq.4 to 8 are shown in Figure 14 . Eq.4 is shown in the upper left corner of Q where the MLP A (·) block 1404 takes the current feature vector f A as input and outputs Q Figure 14 . At the top of A , the kNN block 1402 takes the current feature vector f i as input and performs a k-nearest neighbor (kNN) search based on the coordinates of A. The output of the kNN block 1402 is the coordinates of A Figure 14 where 0 ≤ i ≤ (k - 1). Eq.5 is shown in the upper right corner of P where the MLP i (·) block 1410 takes the difference between voxel A and A as input and outputs Figure 14 The feature vector of Eq.6 is shown in the middle left of where the feature vector K is the input to the MLP K (·) block 1406 and the output of the MLP is added to where 0 ≤ i ≤ (k - 1). Eq. 7 is shown in the Figure 14 right middle part of where the feature vector V is the input to the MLP V (·) block 1408, and the output of the MLP V (·) block 1408 is added to to generate where 0 ≤ i ≤ (k - 1). Applying Eq. 8 to Figure 14 , the input to the softmax normalization function 1412 is the dot product of with . This dot product is divided by to be normalized by the length of the feature vector f A . The output feature (f′ A ) at position A shown at the bottom of Figure 14 is the sum of the k nearest neighbors of the dot product of the softmax function output with .
[0215] In some embodiments, instead of using kNN search to find the k points in the point cloud that are closest to point A, all points within a distance r from A in the point cloud can be used. This operation is called ball query. The value (or radius) r of the ball query can be determined by the quantization step s of the quantizer. For example, given a larger s corresponding to a coarser point cloud, the value of r becomes larger to cover more points from the original point cloud.
[0216] In some embodiments, kNN search can be used to find the k points closest to a query point (such as A). However, after that, only the points within a distance r from A are retained. The value of r can be determined in the same way as the ball query. The distance metric used in kNN search can be any distance metric.
[0217] In some embodiments, a transformer block (such as the example shown in Figure 13 ) updates the features of all occupied positions in the sparse tensor in the same way and outputs the updated sparse tensor. For some embodiments, the MLP Q (·), MLP P (·), MLP K (·), and MLP V (·) may only contain a fully connected layer, which corresponds to a linear projection.
[0218] Figure 15 is a flowchart showing an example process for cascading several feature aggregations according to some embodiments. For some embodiments, several feature aggregation blocks 1502, 1504, 1506, 1508 (e.g., Figures 10 to 13The example feature aggregation blocks shown in are cascaded together to further enhance performance, as shown in the example process 1500 of Figure 15 The feature aggregation blocks can be of the same type. For example, they are all transformer blocks.
[0219] In some embodiments where the feature aggregation blocks are of the same type, the parameters of their neural network layers are shared. By sharing the neural network parameters among the feature aggregation blocks, the total number of parameters of the neural network model can be reduced, which has (for example) the following two benefits. First, the model size of the neural network can be reduced, which makes the storage or transmission of the neural network model easier. Second, reducing the total number of parameters to be learned during the training phase can make the training converge faster. However, the result of sharing the neural network parameters among the feature aggregation blocks may be a reduction in the capacity of the neural network model, which may make the neural network less capable of extracting high-level features and may reduce the performance of the neural network. These potential consequences may be disadvantageous for some applications.
[0220] For some embodiments, each aggregation block can use the same neural network with a separate set of neural network parameters. For some embodiments, each aggregation block can use a separate neural network with a separate set of neural network parameters. For some embodiments, the separate sets of neural network parameters can be the same. For some embodiments, the first set of neural network parameters and the second set of neural network parameters are the same (identical) set of neural network parameters, and the same set of neural network parameters can be used by at least a first neural network and a second neural network. For some embodiments, the first set of neural network parameters and the second set of neural network parameters are distinct but identical sets of neural network parameters.
[0221] Of course, in some embodiments, not all sets of neural network parameters are the same across all neural networks, and not all neural networks are modeled the same across all functional blocks (e.g., feature aggregation blocks). In some embodiments having two or more feature aggregation blocks, two or more feature aggregation blocks including two or more corresponding neural networks can utilize two or more corresponding identical sets of neural network parameters.
[0222] In some embodiments, the feature aggregation blocks can be a mixture of different types of feature aggregation blocks, such as a mixture of IRN blocks and transformer blocks. For some embodiments, a single feature aggregation block can be replaced by two or more cascaded feature aggregation blocks to achieve better compression performance.
[0223] Context-aware, voxel-based upsampling can be applied to point cloud decompression. In some embodiments, context-aware, voxel-based upsampling can be applied to the decoder of Application '087 to generate a point cloud closer to the real one.
[0224] Figure 16 is a block diagram showing an example original decoder architecture according to some embodiments. Figure 16 Shows the architecture 1600 of the decoder of application '087, which includes a base layer and an enhancement layer. The base layer receives the bitstream BS 0 , which is used to perform base decoding 1602 and dequantization 1604, and generate a coarse / simplified point cloud PC 0 . PC 0 is a simplified or low-resolution version of the original input point cloud in a voxel-based representation. From the bitstream BS 1 , the feature decoder block 1606 decodes the per-voxel features of PC 0 , which outfits each occupied voxel in PC 0 with a vector feature that is an abstraction of the local geometry. If BS 1 is not available, which is called the "skip mode" in application '087, the feature decoder still synthesizes a vector feature for each voxel in PC 0 based on the geometry of PC 0 . The resulting feature attached to the point cloud is denoted as PC' 0 In application '087, each feature in PC' 0 is input to the feature-to-residual converter 1608 to decode a set of local 3D points. Specifically, for the voxel A located at (x 0 ,y A ,z A ) in PC' A , its feature f A is input to the feature-to-residual converter, which outputs k sets of 3D points {(x′ 0 ,y′ 0 ,z′ 0 ),(x′ 1 ,y′ 1 ,z′ 1 ),…,(x′ k-1 ,y′ k-1 ,z′ k-1 )}. The feature-to-residual converter 1608 can be a series of MLP layers. Geometric summation ( Figure 16 "⊕" in) translates the decoded point sets by translating them with (x,y,z), as shown in Eqs. 9-11: x′ i =x i +x A ,i = 0,1,…,k-1 Eq.9 y′ i =y i +yA , for i = 0, 1, …, k - 1, Eq. 10 z' i = z i + z A , for i = 0, 1, …, k - 1, Eq. 11 associated with the translated point set for each voxel in the feature set PC' 0 {(x' 0 , y' 0 , z' 0 ), (x' 1 , y' 1 , z' 1 ), …, (x' k-1 , y' k-1 , z' k-1 )} in the formed decoded point cloud PC DEC . If there are M voxels in PC', the decoded point cloud PC 0 contains Mk points. DEC
[0225] In Figure 16 's base layer, the base decoder is used to decode the point cloud from the bitstream BS 0 . The dequantizer is applied to the point cloud to obtain a coarser point cloud PC 0 . For some embodiments, the dequantizer can use a step size s.
[0226] In the enhancement layer, the feature decoder is applied to decode BS 1 and the already decoded coarser point cloud PC 0 to output a set of per - point features PC' 0 . The feature set PC' 0 contains the per - point features of each point in PC 0 . For example, point A in PC 0 has its own feature vector f' A . The decoded feature vector f' A can have a different size from f A , where f A is its corresponding feature vector on the encoder side. However, both f A and f' A are generated to describe the local refinement geometric details of PC 0 close to point A. The decoded feature set PC' 0 is input to the feature - to - residual converter, which generates the residual component of PC DEC . The coarser point cloud PC 0 and the residual are input to the geometric summation block. This summation block adds the residual component to the coarser point cloud PC 0 to produce the final decoded point cloud PC DEC。
[0227] For some embodiments, the base decoder can be any PCC codec. In some embodiments, the base decoder is selected to be a lossy PCC codec, such as the codec in Wang. In some embodiments, the base codec can be a lossless PCC codec, such as the MPEG G-PCC standard, or a depth entropy model with octree representation.
[0228] The number m of decoding points can be a fixed constant, such as m = 5, or the number of decoding points can be adaptively selected, such as based on prior knowledge about the density level of the original point cloud. For example, if the original point cloud is very sparse, m can be set to a small number, such as m = 2.
[0229] The feature-to-residual converter converts the decoded feature set PC’ 0 back to the residual component of PCDEC. Specifically, for some embodiments, the feature-to-residual converter applies a deep neural network to convert each feature vector f’ 0 in PC’ A (associated with point A in PC 0 ) back to the corresponding set of residual points S’ A .
[0230] In some embodiments, the feature-to-residual converter can be a series of MLP layers. In this case, the feature vector, say f’, 0 in PC’ A , is input into a series of MLP layers. The MLP layers directly output a set of m 3D points C 0 , C 1 , …, C m-1 which gives the decoded residual set S’ A . Thus, for a PC 0 with n points A 1 , A n-1 , …, A 0 , the feature-to-residual converter generates the corresponding decoded residual sets, which are represented as S’ 0 , S’ 1 , …, S’ n-1 . These residual sets together constitute the decoded residual component.
[0231] In some embodiments, those 3D points in the residual component that are too far from the origin can be removed. Specifically, for a point C i in the residual component, if its distance to the origin is greater than a threshold t, the point C is considered an outlier and removed from the residual component. The threshold t can be a predetermined constant. The threshold can also be selected according to the quantization step s of the quantizer on the encoder. For example, a larger s means that the PC0 is thicker, and the threshold can be set to a larger value to keep more nodes in the residual component.
[0232] The value of k can be selected based on the density level of the input point cloud. For a dense point cloud, the value of k can be larger (e.g., k = 10). For a sparse point cloud, such as a LiDAR scan, the value of k can be very small, such as k = 1, which can mean Figure 16 of the PC 0 each point in is associated with only one point in the original point cloud.
[0233] Figure 17 is a block diagram showing an example decoder architecture with voxel-based upsampling according to some embodiments. For some embodiments, as Figure 16 and Figure 17 shown, the base layer receives the bitstream BS 0 , which is used to perform base decoding 1702 and dequantization 1704. For some embodiments, as Figure 17 shown, a context-aware upsampling block 1708 is inserted between the feature decoder block 1706 and the feature-to-residual converter 1710. Additionally, the geometric summation block now takes the PC 1 instead of Figure 16 the PC shown in 0 as input. In some embodiments, instead of encoding / decoding the original PC Figure 16 in 0 , now Figure 17 the PC encoded / decoded in 0 is made twice smaller, thus having fewer bits. For some embodiments, Figure 17 the context-aware upsampling block 1708 of Figure 5 can be performed by the example context-aware voxel-based upsampling with a pruning process shown in
[0234] Figure 18 is a block diagram showing an example decoder architecture with voxel-based upsampling and feature aggregation according to some embodiments. In some embodiments of the example process 1800, a feature summarization block 1810 is inserted between the context-aware upsampling block 1808 and the feature-to-residual converter 1812, as Figure 18 shown. In this case, the context-aware upsampling block 1808 can also be cascaded multiple times, for example, in the manner presented in Figure 7 , to achieve a higher upsampling ratio of the PC 0 and additional bit savings. For some embodiments, the base layer receives the bitstream BS 0 , which is used to perform base decoding 1802, dequantization 1804, and feature decoding 1806.
[0235] Figure 19 is a block diagram showing an example decoder architecture without a feature-to-residual converter according to some embodiments. For some embodiments, compared to Figure 18 , only the context-aware upsampling blocks 1908, 1912 are presented, and the feature-to-residual converter is removed, as Figure 19 shown. In this case, the PC 0 is gradually upsampled and refined to obtain the decoded point cloud PC DEC . In the example of Figure 18 , two context-aware upsampling blocks 1908, 1912 are shown on either side of the feature aggregation block 1910. However, in some embodiments, a different number of context-aware upsampling blocks may be used to obtain PC DEC . For some embodiments, the base layer receives the bitstream BS 0 , which is used to perform base decoding 1902, dequantization 1904, and feature decoding 1906.
[0236] In Application '015, a hybrid decoding framework for PCC is used to implement octree-based PCC, voxel-based PCC, and point-based PCC. For some embodiments, one or more of these types of PCCs may be used in the Figure 17 , 18 and the method shown in 19. Application '015 proposes combining two of these types of PCCs. In particular, in one case, only (i) octree-based PCC and (ii) voxel-based PCC methods are used. This configuration corresponds to Figure 19 .
[0237] The processes described in this application can be applied to point cloud super-resolution. For some embodiments, the context-aware upsampling process can be applied to the input point cloud PC 0 and the set of features associated with each of its occupied voxels to achieve 2x super-resolution. In some embodiments, multiple context-aware upsampling processes can be applied to the input point cloud PC 0 and the set of features associated with each of its occupied voxels to achieve more than 2x super-resolution, as Figure 7 shown.
[0238] For some embodiments, the features associated with the voxels can be attributes such as color and intensity. In some embodiments, the features associated with the voxels can be local geometric features of the PC 0 extracted by a neural network layer, such as the feature aggregation process shown in Figures 10 - 12 . In some embodiments, the features associated with the voxels can be a concatenation of attributes and geometric features.
[0239] Figure 20is a block diagram showing an example decoder architecture with a single process of voxel - based upsampling and feature aggregation according to some embodiments. In some embodiments of the example feature decoder process 2000, if the context - aware voxel - based upsampling 2002 is applied only once, feature aggregation 2004 can be appended after the context - aware voxel - based upsampling for feature aggregation and refinement, as Figure 20 shown.
[0240] If context - aware voxel - based upsampling is applied and then feature aggregation (as Figure 20 shown) is applied, then the feature aggregation and all other feature aggregations within the context - aware upsampling block have the same neural network architecture and share the same neural network parameters. For some embodiments, by making all feature aggregation blocks share the same set of weights, the total number of neural network parameters can be reduced. By sharing neural network parameters between feature aggregation blocks, the total number of parameters of the neural network model can be reduced, which has (for example) the following two benefits. First, the model size of the neural network can be reduced, which makes the storage or transmission of the neural network model easier. Second, reducing the total number of parameters to be learned during the training phase can make the training converge faster. However, the result of sharing neural network parameters between feature aggregation blocks may be a reduction in the capacity of the neural network model, which may make the neural network less able to extract high - level features and may reduce the performance of the neural network. These potential consequences may be disadvantageous for certain applications.
[0241] Of course, in some embodiments, not all groups of neural network parameters are the same across all neural networks, and not all neural networks are modeled identically across all functional blocks (e.g., feature aggregation blocks). In some embodiments with two or more feature aggregation blocks, two or more feature aggregation blocks including two or more corresponding neural networks may utilize two or more corresponding identical groups of neural network parameters.
[0242] For some embodiments, if a trimmed context - aware voxel - based upsampling as Figure 5 shown is used, the following feature aggregation blocks can share the same group of neural network parameters: (i) the feature aggregation blocks within the binary classification block (which is shown in Figure 9 ), and (ii) the feature aggregation blocks after the context - aware voxel - based upsampling block (which is shown in Figure 20 ).
[0243] Similarly, for some embodiments, if as Figure 8As shown, when context-aware voxel-based upsampling is used together with initial feature aggregation, the following feature aggregation blocks can share the same set of neural network parameters: (i) the feature aggregation block after the nearest neighbor (NN) upsampling block (which is shown in Figure 8 ), (ii) the feature aggregation block within the binary classification block (which is shown in Figure 9 ), and (iii) the feature aggregation block after the context-aware voxel-based upsampling block (which is shown in Figure 20 ).
[0244] Figure 21 is a block diagram showing an example sparse tensor operation according to some embodiments. Figure 21 Example process 2100 for showing the downsampling 2104, upsampling 2108, coordinate reading / splitting 2116, and coordinate pruning 2112 processes. For simplicity, Figure 21 the operations in are shown in 2D space, and the same basic principle can be applied to 3D space. In this example, the input point cloud A0 2102 occupies voxels at positions (0, 2), (0, 3), (0, 4), (0, 5), (1, 1), (1, 6), (2, 6), (3, 5), (4, 4), (5, 4), (6, 4), and (7, 4) 2118, where the origin is zero-based and at the upper left corner. Thus, the coordinate reader / splitter 2116 outputs occupancy coordinates of (0, 2), (0, 3), (0, 4), (0, 5), (1, 1), (1, 6), (2, 6), (3, 5), (4, 4), (5, 4), (6, 4), and (7, 4) 2118. By downsampling A0 2102 at a ratio of 2, the number of voxels in the downsampled point cloud A1 2106 is reduced by half in each dimension, and a voxel is considered occupied if any one of the corresponding 4 points in A0 2102 is occupied. After upsampling A1 2106, the number of voxels is restored in A2 2110, and a voxel in A2 2110 is considered occupied if the corresponding voxel in A1 2106 is occupied. A2 2110 is denser than the input point cloud A0 2102. To remove / prune the points (voxels) not occupied in A0 2102 from A2 2110, the coordinate pruning block uses the occupancy coordinate information. In the resulting point cloud A3 2114, only the voxels occupied in the original point cloud A0 are considered occupied.
[0245] Figure 22 is a block diagram showing an example decoder architecture according to some embodiments. For some embodiments, an example feature decoder 2200 based on sparse 3D convolution, downsampling, and upsampling is shown in Figure 22 . Bitstream BS 12202 is entropy decoded 2204 and feature dequantized 2206 to generate a downsampled feature set F’ down , followed by sequential upsampling using the geometry of the PC 0 to gradually magnify and refine the features. In Figure 22 , the bitstream BS 1 is decoded by an entropy decoder, followed by a dequantizer, producing a downsampled feature set F’ down .
[0246] As shown in the upper right corner of Figure 22 , a 3D sparse tensor 2244 is constructed (uniquely) based on the geometry (coordinates) of the PC 0 . The tensor is sequentially downsampled 2240, 2236, producing a tensor PC’ down . Figure 22 The PC’ in down and the PC for the feature encoder (not shown) down can have the same geometry, but their features can be different for some embodiments. To upsample F’ down , the geometry of F’ down is converted to the geometry of PC’ down . To convert the geometry of F’ down to the geometry of PC’ down , the feature replacement block 2208 replaces the original features of PC’ down with F’ down , producing another sparse tensor PC” down .
[0247] PC” down is upsampled by two upsampling processing blocks, each of which contains an upsampling operator 2210, 2222 and two sparse 3D convolutional layers. In Figure 22 , “upsampling 2” is a sparse tensor upsampling operator 2210, 2222 with a ratio of 2. For some embodiments, the upsampling 2 block can include sparse tensor upsampling operators 2210, 2222, and two sets of convolutional decoders 2212, 2216, 2224, 2228 and rectified linear units (ReLU) 2214, 2218, 2226, 2230. Similar to the upsampling operator for a “conventional” 2D image, the upsampling 2 block magnifies the size of the sparse tensor by a factor of 2 along each dimension. See the illustrative example in Figure 21 . After each upsampling processing block, the resulting tensor is refined using the corresponding coordinate readers 2238, 2242 and coordinate pruning blocks 2220, 2232. Figure 21Illustrative examples of coordinate pruning are shown, which remove some occupied voxels of the input tensor and retain the remaining voxels based on a set of input coordinates that can be obtained from a coordinate reader. The coordinate pruning block 2220 removes some voxels (and associated features) from the upsampled version of the PC” down and retains only those voxels that also appear in the downsampled version of the PC 0 . The output of the second coordinate pruning block 2232 is a tensor that has the same geometry as the PC 0 . This tensor is input into the feature reader 2234 to obtain the decoded feature set PC’ 0 .
[0248] Figure 23 is a block diagram showing an example decoder architecture according to some embodiments. Compared with Figure 22 , for some embodiments, the coordinate pruning blocks 2318, 2332 (and for some embodiments, the feature aggregation blocks 2320, 2334) are incorporated into the Figure 23 preceding upsampling processing block in
[0249] For some embodiments of the example decoder 2300, the tensor is successively downsampled 2342, 2338, resulting in the tensor PC’ down . In some embodiments, the bitstream is entropy decoded 2302 and feature dequantized 2304 to generate the downsampled feature set F’ down . To convert the geometry of F’ down to the geometry of PC’ down , the feature replacement block 2306 replaces the original features of PC’ down with F’ down , resulting in another sparse tensor PC” down . For some embodiments, the upsampling 2 block may include sparse tensor upsampling operators 2308, 2322, and two sets of convolutional decoders 2310, 2314, 2324, 2328 and rectifier linear units (ReLU) 2312, 2316, 2326, 2330. For some embodiments, after each upsampling processing block, the resulting tensor is refined using the corresponding coordinate readers 2340, 2344 and coordinate pruning blocks 2318, 2332. The output of the second feature aggregation block 2334 is a tensor that has the same geometry as the PC 0 . This tensor is input into the feature reader 2336 to obtain the decoded feature set PC’ 0 .
[0250] Figure 24is a flowchart showing an example process of voxel - based upsampling with trimmed context awareness according to some embodiments. For some embodiments, the example process 2400 may include upsampling 2402 a first point cloud using an initial upsampling to obtain a second point cloud. For some embodiments, the exemplary process 2400 may further include associating 2404 the features of the second point cloud with per - voxel context information to obtain a third point cloud. For some embodiments, the exemplary process 2400 may also include predicting 2406 the occupancy status of at least one voxel of the third point cloud. For some embodiments, the exemplary process 2400 may further include removing 2408 the voxels of the third point cloud classified as empty according to the predicted occupancy status to generate a trimmed point cloud. For some embodiments, the initial upsampling may include nearest - neighbor upsampling. For some embodiments, associating the features may include concatenating the features.
[0251] Figure 25 is a flowchart showing an example process of context - aware voxel - based upsampling and feature aggregation according to some embodiments. For some embodiments, the example process 2500 may include upsampling 2502 a first point cloud using an initial upsampling to obtain a second point cloud. For some embodiments, the exemplary process 2500 may also include associating 2504 the features of the second point cloud with context information to obtain a third point cloud. For some embodiments, the example process 2500 may also include predicting 2506 the occupancy status of at least one voxel of the third point cloud, where predicting the occupancy status of at least one voxel includes aggregating at least one feature of the third point cloud, where aggregating at least one feature of the third point cloud includes using a first neural network, and where using the first neural network to aggregate at least one feature of the third point cloud includes: using a first set of neural network parameters with the first neural network. For some embodiments, the example process 2500 may also include removing 2508 the voxels of the third point cloud classified as empty according to the predicted occupancy status to generate a trimmed point cloud. For some embodiments, the example process 2500 may also include performing 2510 feature aggregation on the trimmed point cloud to generate the aggregated features, where performing feature aggregation on the trimmed point cloud includes using a second neural network, where using the second neural network to generate the aggregated features includes: using a second set of neural network parameters with the second neural network, and where the first set of neural network parameters is the same as the second set of neural network parameters.
[0252] Figure 26FIG. 0 is a flowchart showing an example process for encoding a bitstream according to some embodiments. For some embodiments, example process 2600 may include obtaining 2602 a first point cloud. For some embodiments, example process 2500 may further include determining 2604 an occupancy status of at least one voxel of the first point cloud. For some embodiments, exemplary process 2500 may further include removing 2606 voxels of the first point cloud classified as empty according to the determined occupancy status to generate a second point cloud. For some embodiments, exemplary process 2500 may further include associating 2608 features of the second point cloud with context information to obtain a third point cloud. For some embodiments, exemplary process 2500 may further include downsampling 2610 the third point cloud using an initial downsampling to obtain a fourth point cloud. For some embodiments, exemplary process 2500 may further include outputting 2612 the fourth point cloud as an encoded point cloud.
[0253] For some embodiments, an apparatus may include one or more processors configured to: upsample a first point cloud using nearest neighbor upsampling to obtain a second point cloud; concatenate features of the second point cloud with per-voxel context information to obtain a third point cloud; predict an occupancy status of at least one voxel of the third point cloud; and remove voxels of the third point cloud classified as empty according to the predicted occupancy status.
[0254] Although methods and systems according to some embodiments are discussed in the context of virtual reality (VR), some embodiments may also be applied to the mixed reality (MR) / augmented reality (AR) context. Additionally, although the term "head-mounted display (HMD)" is used herein according to some embodiments, for some embodiments, some embodiments may be applied to wearable devices capable of, for example, VR, AR, and / or MR (which may or may not be attached to the head).
[0255] An example method according to some embodiments may include upsampling a first point cloud using an initial upsampling to obtain a second point cloud; associating features of the second point cloud with context information to obtain a third point cloud; predicting an occupancy status of at least one voxel of the third point cloud; and removing voxels classified as empty in the third point cloud according to the predicted occupancy status to generate a trimmed point cloud.
[0256] For some embodiments of the example method, the initial upsampling may include nearest neighbor upsampling.
[0257] For some embodiments of the example method, associating features may include concatenating features of the second point cloud with context information to obtain a third point cloud.
[0258] For some embodiments of the example method, predicting the occupancy state can be performed using a first neural network.
[0259] For some embodiments of the example method, predicting the occupancy state can predict the true occupancy state of at least one voxel.
[0260] For some embodiments of the example method, predicting the occupancy state can predict the likelihood that at least one voxel is occupied.
[0261] For some embodiments of the example method, removing the voxels of the third point cloud can use a voxel pruning process to remove the voxels.
[0262] Some embodiments of the example method can further include aggregating at least one feature of the second point cloud.
[0263] For some embodiments of the example method, the context information can be per-voxel context information.
[0264] For some embodiments of the example method, predicting the occupancy state of at least one voxel can include: aggregating at least one feature of the third point cloud; processing the aggregated features with a multi-layer perceptron (MLP) layer to generate an MLP layer output; performing a softmax process on the MLP layer output to generate a softmax output value; and performing thresholding of the softmax output value to generate the predicted occupancy state of at least one voxel of the third point cloud.
[0265] For some embodiments of the example method, thresholding of the softmax output value converts softmax output values greater than 0.5 into an output value of 1, and converts softmax output values equal to 0.5 or less into an output value of 0.
[0266] For some embodiments of the example method, predicting the occupancy state of at least one voxel can include: aggregating at least one feature of the third point cloud; and generating the predicted occupancy state of at least one voxel of the third point cloud based on the aggregated features.
[0267] For some embodiments of the example method, aggregating at least one feature can include: repeating the concatenation process one or more times, the concatenation process can include: performing a sparse 3D convolution on the input point cloud to generate a convolutional output point cloud; performing a non-linear activation process on the convolutional output point cloud to generate a non-linear output point cloud; and if there is a next cycle of the concatenation process, preparing the non-linear output point cloud as the input point cloud, where the third point cloud is the input point cloud of the first cycle of the concatenation process, and where the last cycle of the concatenation process generates the aggregated features.
[0268] Some embodiments of the exemplary method may further include: adding a third point cloud to the ReLU output point cloud of the last loop of the cascading process.
[0269] For some embodiments of the exemplary method, aggregating at least one feature may include: performing a sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; and performing a non-linear activation process on the convolutional output point cloud to generate the aggregated feature.
[0270] For some embodiments of the exemplary method, the non-linear activation process may be a rectifier linear unit (ReLU) activation process.
[0271] For some embodiments of the exemplary method, aggregating at least one feature may include: repeating a first cascading process one or more times, the first cascading process may include: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first non-linear activation process on the first convolutional output point cloud to generate a first non-linear output point cloud; and if there is a next loop of the first cascading process, preparing the first non-linear output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first loop of the first cascading process, wherein the last loop of the first cascading process generates a first cascading process output; repeating a second cascading process one or more times, the second cascading process may include: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second non-linear activation process on the second convolutional output point cloud to generate a second non-linear output point cloud; and if there is a next loop of the second cascading process, preparing the second non-linear output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first loop of the second cascading process, wherein the last loop of the second cascading process generates a second cascading process output; concatenating the first cascading process output and the second cascading process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0272] For some embodiments of the example method, aggregating at least one feature may include: repeating the first cascading process one or more times, where the first cascading process may include: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first rectified linear unit (ReLU) activation process on the first convolutional output point cloud to generate a first ReLU output point cloud; if there is a next cycle of the first cascading process, preparing the first ReLU output point cloud as the first input point cloud, where the third point cloud is the first input point cloud of the first cycle of the first cascading process, and where the last cycle of the first cascading process generates a first cascading process output; repeating the second cascading process one or more times, where the second cascading process may include: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second rectified linear unit (ReLU) activation process on the second convolutional output point cloud to generate a second ReLU output point cloud; and if there is a next cycle of the second cascading process, preparing the second ReLU output point cloud as the second input point cloud, where the third point cloud is the second input point cloud of the first cycle of the second cascading process, and where the last cycle of the second cascading process generates a second cascading process output; cascading the first cascading process output and the second cascading process output to generate a cascaded output; and adding the third point cloud to the cascaded output to generate the aggregated feature.
[0273] For some embodiments of the example method, aggregating at least one feature of the third point cloud may include performing a self-attention process on the third point cloud; adding the third point cloud to the self-attention process output to generate an MLP process input; performing an MLP process on the MLP process input; and adding the MLP process input to the MLP process output to generate the aggregated feature.
[0274] For some embodiments of the example method, the self-attention process generates an output feature based on k nearest neighbors of voxels of the third point cloud.
[0275] For some embodiments of the example method, aggregating at least one feature of the third point cloud may include performing the feature aggregation process two or more times.
[0276] Some embodiments of the example method may further include: performing feature decoding on an input point cloud and a first bitstream to generate the first point cloud.
[0277] Some embodiments of the example method may further include: performing a feature-to-residual conversion on the trimmed point cloud to generate a residual output; and adding the trimmed point cloud to the residual output to generate a decoded point cloud.
[0278] Some embodiments of the example method may further include: performing feature aggregation on the trimmed point cloud to generate the aggregated features, wherein the feature-to-residual transformation is performed on the aggregated features.
[0279] Some embodiments of the example method may further include: performing feature aggregation on the trimmed point cloud to generate the aggregated features; and performing a context-aware upsampling process on the aggregated features to generate the decoded point cloud.
[0280] An example apparatus according to some embodiments may include a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to: upsample a first point cloud using nearest neighbor upsampling to obtain a second point cloud; concatenate the features of the second point cloud with context information to obtain a third point cloud; predict the occupancy status of at least one voxel of the third point cloud; and remove the voxels classified as empty in the third point cloud according to the predicted occupancy status to generate a trimmed point cloud.
[0281] An example device according to some embodiments may include the apparatus according to the example apparatus; and at least one of the following: (i) an antenna configured to receive a signal that includes data representing the image, (ii) a band limiter configured to limit the received signal to a band that includes the data representing the image, or (iii) a display configured to display the image.
[0282] Some embodiments of the example method may further include at least one of a TV, a cellular phone, a tablet, and a set-top box (STB).
[0283] An example apparatus according to some embodiments may include an access unit configured to access data including a first point cloud; and a transmitter configured to transmit the data including the first point cloud.
[0284] An example method according to some embodiments may include: accessing data including a first point cloud; and transmitting the data including the first point cloud.
[0285] An example computer-readable medium according to some embodiments may include instructions for causing one or more processors to perform the following operations: upsample a first point cloud using nearest neighbor upsampling to obtain a second point cloud; concatenate the features of the second point cloud with context information to obtain a third point cloud; predict the occupancy status of at least one voxel of the third point cloud; and remove the voxels classified as empty in the third point cloud according to the predicted occupancy status to generate a trimmed point cloud.
[0286] An example computer program product according to some embodiments can include instructions that, when executed by one or more processors, cause the one or more processors to: upsample a first point cloud using nearest neighbor upsampling to obtain a second point cloud; concatenate features of the second point cloud with context information to obtain a third point cloud; predict an occupancy state of at least one voxel of the third point cloud; and remove voxels classified as empty in the third point cloud based on the predicted occupancy state to generate a trimmed point cloud.
[0287] An example method according to some embodiments can include: performing context-aware upsampling of a first point cloud to determine an upsampled second point cloud, where the context-aware upsampling can include: associating features of a third point cloud with context information, the third point cloud being at least partially based on an initial upsampled version of the first point cloud; and removing voxels of a fourth point cloud predicted to be empty from the third point cloud at least partially based on the context information to generate an enlarged second point cloud.
[0288] An additional example method according to some embodiments can include: upsampling a first point cloud using initial upsampling to obtain a second point cloud; associating features of the second point cloud with context information to obtain a third point cloud; predicting an occupancy state of at least one voxel of the third point cloud, where predicting the occupancy state of at least one voxel includes: aggregating at least one feature of the third point cloud, where aggregating at least one feature of the third point cloud includes using a first neural network, and where using the first neural network to aggregate at least one feature of the third point cloud includes: using a first set of neural network parameters with the first neural network; removing voxels classified as empty in the third point cloud based on the predicted occupancy state to generate a trimmed point cloud; and performing feature aggregation on the trimmed point cloud to generate aggregated features, where performing feature aggregation on the trimmed point cloud includes using a second neural network, where using the second neural network to generate the aggregated features includes: using a second set of neural network parameters with the second neural network, and where the first set of neural network parameters is the same as the second set of neural network parameters.
[0289] Some embodiments of the additional example method can further include: aggregating at least one feature of the second point cloud.
[0290] For some embodiments of the additional example method, aggregating at least one feature of the second point cloud can include using a third neural network, and using the third neural network to aggregate at least one feature of the second point cloud can include: using a third set of neural network parameters with the third neural network, and the third set of neural network parameters can be the same as the first set of neural network parameters.
[0291] Some embodiments of the additional example method may further include: performing feature decoding on the input point cloud and the first bitstream to generate the first point cloud.
[0292] Some embodiments of the additional example method may further include: performing a feature-to-residual transformation on the trimmed point cloud to produce a residual output; and adding the trimmed point cloud to the residual output to generate a decoded point cloud.
[0293] For some embodiments of the additional example method, the feature-to-residual transformation may be performed on the aggregated features.
[0294] Some embodiments of the additional example method may further include: performing a context-aware upsampling process on the aggregated features to generate a decoded point cloud.
[0295] For some embodiments of the additional example method, the initial upsampling may include nearest neighbor upsampling.
[0296] For some embodiments of the additional example method, the associated features may include concatenating the features of a second point cloud with context information to obtain a third point cloud.
[0297] For some embodiments of the additional example method, predicting the occupancy state may predict the true occupancy state of at least one voxel.
[0298] For some embodiments of the additional example method, predicting the occupancy state may predict the likelihood that at least one voxel is occupied.
[0299] For some embodiments of the additional example method, removing the voxels of the third point cloud may use a voxel pruning process to remove the voxels.
[0300] For some embodiments of the additional example method, the context information may be per-voxel context information.
[0301] For some embodiments of the additional example method, predicting the occupancy state of at least one voxel may include: processing the aggregated features with a multi-layer perceptron (MLP) layer to generate an MLP layer output; performing a softmax process on the MLP layer output to generate a softmax output value; and performing thresholding of the softmax output value to generate the predicted occupancy state of at least one voxel of the third point cloud.
[0302] For some embodiments of the additional example method, thresholding of the softmax output value may convert softmax output values greater than 0.5 into an output value of 1, and convert softmax output values equal to 0.5 or less into an output value of 0.
[0303] For some embodiments of the additional example method, predicting the occupancy state of at least one voxel may include: aggregating at least one feature of the third point cloud; and generating a predicted occupancy state of at least one voxel of the third point cloud based on the aggregated feature.
[0304] For some embodiments of the additional example method, aggregating at least one feature of the third point cloud may include: repeating the cascading process one or more times, the cascading process may include: performing sparse 3D convolution on the input point cloud to generate a convolutional output point cloud; performing a non-linear activation process on the convolutional output point cloud to generate a non-linear output point cloud; and if there is a next cycle of the cascading process, preparing the non-linear output point cloud as the input point cloud, the third point cloud may be the input point cloud of the first cycle of the cascading process, and the last cycle of the cascading process may generate the aggregated feature.
[0305] Some embodiments of the additional example method may further include: adding the third point cloud to the ReLU output point cloud of the last cycle of the cascading process.
[0306] For some embodiments of the additional example method, aggregating at least one feature may include: performing sparse 3D convolution on the input point cloud to generate a convolutional output point cloud; and performing a non-linear activation process on the convolutional output point cloud to generate the aggregated feature.
[0307] For some embodiments of the additional example method, the non-linear activation process may include a rectifier linear unit (ReLU) activation process, and the non-linear output point cloud includes a ReLU output point cloud.
[0308] For some embodiments of the additional example method, aggregating at least one feature of the third point cloud may include: repeating the first cascading process one or more times, the first cascading process may include: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first non-linear activation process on the first convolutional output point cloud to generate a first non-linear output point cloud; and if there is a next cycle of the first cascading process, preparing the first non-linear output point cloud as the first input point cloud, wherein the third point cloud may be the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process may generate a first cascading process output; repeating the second cascading process one or more times, the second cascading process may include: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second non-linear activation process on the second convolutional output point cloud to generate a second non-linear output point cloud; and if there is a next cycle of the second cascading process, preparing the second non-linear output point cloud as the second input point cloud, wherein the third point cloud may be the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process may generate a second cascading process output; concatenating the first cascading process output and the second cascading process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0309] For some embodiments of the additional example method, aggregating at least one feature may include: repeating the first cascading process one or more times, the first cascading process may include: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first rectified linear unit (ReLU) activation process on the first convolutional output point cloud to generate a first ReLU output point cloud; and if there is a next cycle of the first cascading process, preparing the first ReLU output point cloud as the first input point cloud, wherein the third point cloud may be the first input point cloud of the first cycle of the first cascading process, and wherein the last cycle of the first cascading process may generate a first cascading process output; repeating the second cascading process one or more times, the second cascading process may include: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second rectified linear unit (ReLU) activation process on the second convolutional output point cloud to generate a second ReLU output point cloud; and if there is a next cycle of the second cascading process, preparing the second ReLU output point cloud as the second input point cloud, wherein the third point cloud may be the second input point cloud of the first cycle of the second cascading process, and wherein the last cycle of the second cascading process may generate a second cascading process output; concatenating the first cascading process output and the second cascading process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0310] For some embodiments of the additional example method, aggregating at least one feature may include: performing a self-attention process on the third point cloud; adding the third point cloud to the self-attention process output to generate an MLP process input; performing an MLP process on the MLP process input; and adding the MLP process input to the MLP process output to generate the aggregated feature;
[0311] For some embodiments of the additional example method, the self-attention process may generate an output feature based on k nearest neighbors of voxels of the third point cloud.
[0312] For some embodiments of the additional example method, aggregating at least one feature of the third point cloud may include: performing the feature aggregation process two or more times.
[0313] For some embodiments of the additional example method, the first set of neural network parameters and the second set of neural network parameters may be the same set of neural network parameters used by at least the first neural network and the second neural network.
[0314] For some embodiments of the additional example method, the first set of neural network parameters and the second set of neural network parameters can be significant but identical sets of neural network parameters.
[0315] An additional example apparatus according to some embodiments can include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.
[0316] A first example method / apparatus according to some embodiments can include: upsampling a first point cloud using an initial upsampling to obtain a second point cloud; associating features of the second point cloud with context information to obtain a third point cloud; predicting an occupancy state of at least one voxel of the third point cloud; and removing voxels classified as empty in the third point cloud according to the predicted occupancy state to generate a trimmed point cloud.
[0317] For some embodiments of the first example method, the initial upsampling includes nearest neighbor upsampling.
[0318] For some embodiments of the first example method, associating features includes concatenating features of the second point cloud with context information to obtain a third point cloud.
[0319] For some embodiments of the first example method, the context information is voxel-by-voxel context information.
[0320] For some embodiments of the first example method, the context information includes a context point cloud.
[0321] For some embodiments of the first example method, the context information includes information about the second point cloud.
[0322] For some embodiments of the first example method, the context information includes information about the voxel occupancy state of the second point cloud.
[0323] For some embodiments of the first example method, the context information includes information about the position of a child voxel relative to the position of a parent voxel of the first point cloud.
[0324] For some embodiments of the first example method, the context information includes coordinate information about the positions of occupied voxels in at least one of the first point cloud and the second point cloud.
[0325] For some embodiments of the first example method, the context information includes coordinate information, and the coordinate information is in the form of one of Euclidean coordinates, spherical coordinates, and cylindrical coordinates.
[0326] For some embodiments of the first example method, the context information provides known information about the first point cloud in addition to the information available for the initial upsampling of the first point cloud.
[0327] For some embodiments of the first example method, the context information includes the bit depth of the second point cloud.
[0328] Some embodiments of the first example method may further include: performing feature decoding on the input point cloud and the first bitstream to generate a first point cloud.
[0329] Some embodiments of the first example method may further include: performing feature aggregation on the trimmed point cloud to generate the aggregated features; and performing a context-aware upsampling process on the aggregated features to generate a decoded point cloud.
[0330] Some embodiments of the first example method may further include: performing feature-to-residual conversion on the trimmed point cloud to produce a residual output; and adding the trimmed point cloud to the residual output to generate a decoded point cloud.
[0331] Some embodiments of the first example method may further include: performing feature aggregation on the trimmed point cloud to generate the aggregated features, wherein the feature-to-residual conversion is performed on the aggregated features.
[0332] For some embodiments of the first example method, a first neural network is used to perform the prediction of the occupancy state.
[0333] For some embodiments of the first example method, predicting the occupancy state predicts the true occupancy state of at least one voxel.
[0334] For some embodiments of the first example method, predicting the occupancy state predicts the likelihood that at least one voxel is occupied.
[0335] For some embodiments of the first example method, removing the voxels of the third point cloud uses a voxel pruning process to remove the voxels.
[0336] Some embodiments of the first example method may further include: aggregating at least one feature of the second point cloud.
[0337] For some embodiments of the first example method, predicting the occupancy state of at least one voxel includes: aggregating at least one feature of the third point cloud; processing the aggregated features with a multi-layer perceptron (MLP) layer to generate an MLP layer output; performing softmax processing on the MLP layer output to generate a softmax output value; and performing thresholding on the softmax output value to generate the predicted occupancy state of at least one voxel of the third point cloud.
[0338] For some embodiments of the first example method, thresholding of the softmax output values converts softmax output values greater than 0.5 into an output value of 1, and converts softmax output values equal to 0.5 or less into an output value of 0.
[0339] For some embodiments of the first example method, predicting the occupancy status of at least one voxel includes: aggregating at least one feature of the third point cloud; and generating a predicted occupancy status of at least one voxel of the third point cloud based on the aggregated feature.
[0340] For some embodiments of the first example method, aggregating at least one feature includes: repeating the concatenation process one or more times, the concatenation process including: performing a sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; performing a non-linear activation process on the convolutional output point cloud to generate a non-linear output point cloud; and if there is a next cycle of the concatenation process, preparing the non-linear output point cloud as the input point cloud, where the third point cloud is the input point cloud of the first cycle of the concatenation process, and where the last cycle of the concatenation process generates the aggregated feature.
[0341] Some implementations of the first example method may further include: adding the third point cloud to the ReLU output point cloud of the last cycle of the concatenation process.
[0342] For some embodiments of the first example method, aggregating at least one feature includes: performing a sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; and performing a non-linear activation process on the convolutional output point cloud to generate the aggregated feature.
[0343] For some embodiments of the first example method, the non-linear activation process includes a rectifier linear unit (ReLU) activation process, and the non-linear output point cloud includes a ReLU output point cloud.
[0344] For some embodiments of the first example method, aggregating at least one feature includes: repeating the first cascading process one or more times, the first cascading process including: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first non-linear activation process on the first convolutional output point cloud to generate a first non-linear output point cloud; and if there is a next cycle of the first cascading process, preparing the first non-linear output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; repeating the second cascading process one or more times, the second cascading process including: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second non-linear activation process on the second convolutional output point cloud to generate a second non-linear output point cloud; and if there is a next cycle of the second cascading process, preparing the second non-linear output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process generates a second cascading process output; concatenating the first cascading process output and the second cascading process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0345] For some embodiments of the first example method, aggregating at least one feature includes: repeating the first cascading process one or more times, the first cascading process including: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first rectified linear unit (ReLU) activation process on the first convolutional output point cloud to generate a first ReLU output point cloud; if there is a next cycle of the first cascading process, preparing the first ReLU output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; repeating the second cascading process one or more times, the second cascading process including: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second rectified linear unit (ReLU) activation process on the second convolutional output point cloud to generate a second ReLU output point cloud; and if there is a next cycle of the second cascading process, preparing the second ReLU output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process generates a second cascading process output; concatenating the first cascading process output and the second cascading process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0346] For some embodiments of the first example method, aggregating at least one feature includes: performing a self-attention process on a third point cloud; adding the third point cloud to an output of the self-attention process to generate an MLP process input; performing an MLP process on the MLP process input; and adding the MLP process input to an output of the MLP process to generate the aggregated feature;
[0347] For some embodiments of the first example method, the self-attention process generates output features based on k nearest neighbors of voxels of the third point cloud.
[0348] For some embodiments of the first example method, aggregating at least one feature of a third point cloud includes performing the feature aggregation process two or more times.
[0349] A first example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to: upsample a first point cloud using initial upsampling to obtain a second point cloud; associate features of the second point cloud with context information to obtain a third point cloud; predict an occupancy state of at least one voxel of the third point cloud; and remove voxels classified as empty from the third point cloud according to the predicted occupancy state to generate a trimmed point cloud.
[0350] For some embodiments of the first example apparatus, the initial upsampling includes nearest neighbor upsampling.
[0351] For some embodiments of the first example apparatus, associating features includes concatenating features of the second point cloud with the context information to obtain the third point cloud.
[0352] An example device according to some embodiments may include: an apparatus according to the apparatus listed above; and at least one of the following: (i) an antenna configured to receive a signal that includes data representing the image, (ii) a band limiter configured to limit the received signal to a band that includes data representing the image, or (iii) a display configured to display the image.
[0353] Some embodiments of the example device may further include at least one of a TV, a cellular phone, a tablet, and a set-top box (STB).
[0354] An example computer-readable medium according to some embodiments may include instructions for causing one or more processors to perform the following operations: upsample a first point cloud using an initial upsampling to obtain a second point cloud; associate features of the second point cloud with context information to obtain a third point cloud; predict an occupancy state of at least one voxel of the third point cloud; and remove voxels classified as empty in the third point cloud based on the predicted occupancy state to generate a trimmed point cloud.
[0355] An example computer program product according to some embodiments may include instructions that, when executed by one or more processors, cause the one or more processors to: upsample a first point cloud using an initial upsampling to obtain a second point cloud; associate features of the second point cloud with context information to obtain a third point cloud; predict an occupancy state of at least one voxel of the third point cloud; and remove voxels classified as empty in the third point cloud based on the predicted occupancy state to generate a trimmed point cloud.
[0356] A second example method according to some embodiments may include performing context-aware upsampling of a first point cloud to determine an upsampled second point cloud, where the context-aware upsampling includes: associating features of a third point cloud with context information, the third point cloud being at least partially based on an initial upsampled version of the first point cloud; and removing voxels of a fourth point cloud predicted to be empty from the third point cloud at least partially based on the context information to generate an enlarged second point cloud.
[0357] A third example method according to some embodiments may include: upsampling a first point cloud using an initial upsampling to obtain a second point cloud; associating features of the second point cloud with context information to obtain a third point cloud; predicting an occupancy state of at least one voxel of the third point cloud, where predicting the occupancy state of at least one voxel includes aggregating at least one feature of the third point cloud, where aggregating at least one feature of the third point cloud includes using a first neural network, and where using the first neural network to aggregate at least one feature of the third point cloud includes: using a first set of neural network parameters with the first neural network; removing voxels classified as empty in the third point cloud based on the predicted occupancy state to generate a trimmed point cloud; and performing feature aggregation on the trimmed point cloud to generate an aggregated feature, where performing feature aggregation on the trimmed point cloud includes using a second neural network, where using the second neural network to generate the aggregated feature includes: using a second set of neural network parameters with the second neural network, and where the first set of neural network parameters is the same as the second set of neural network parameters.
[0358] Some embodiments of the third example method may further include aggregating at least one feature of the second point cloud.
[0359] For some embodiments of the third example method, wherein aggregating at least one feature of the second point cloud includes using a third neural network, and wherein using the third neural network to aggregate at least one feature of the second point cloud includes: using a third set of neural network parameters with the third neural network, and wherein the third set of neural network parameters is the same as the first set of neural network parameters.
[0360] For some embodiments of the third example method, the initial upsampling includes nearest neighbor upsampling.
[0361] For some embodiments of the third example method, associating features includes concatenating features of the second point cloud with context information to obtain a third point cloud.
[0362] For some embodiments of the third example method, associating features includes concatenating features of the second point cloud with context information to obtain a third point cloud.
[0363] For some embodiments of the third example method, the context information is per-voxel context information.
[0364] Some embodiments of the third example method may further include: performing feature decoding on the input point cloud and the first bitstream to generate a first point cloud.
[0365] Some embodiments of the third example method may further include: performing a context-aware upsampling process on the aggregated features to generate a decoded point cloud.
[0366] Some embodiments of the third example method may further include: performing a feature-to-residual transformation on the trimmed point cloud to generate a residual output; and adding the trimmed point cloud to the residual output to generate a decoded point cloud.
[0367] For some embodiments of the third example method, performing the feature-to-residual transformation on the aggregated features.
[0368] For some embodiments of the third example method, predicting the occupancy state predicts the true occupancy state of at least one voxel.
[0369] For some embodiments of the third example method, predicting the occupancy state predicts the likelihood that at least one voxel is occupied.
[0370] For some embodiments of the third example method, removing voxels of the third point cloud uses a voxel pruning process to remove voxels.
[0371] For some embodiments of the third example method, predicting the occupancy state of at least one voxel further includes: processing the aggregated features with a multi-layer perceptron (MLP) layer to generate an MLP layer output; performing a softmax process on the MLP layer output to generate softmax output values; and performing thresholding of the softmax output values to generate the predicted occupancy state of at least one voxel of the third point cloud.
[0372] For some embodiments of the third example method, the thresholding of the softmax output values converts softmax output values greater than 0.5 into an output value of 1, and converts softmax output values equal to 0.5 or less into an output value of 0.
[0373] For some embodiments of the third example method, predicting the occupancy state of at least one voxel includes: aggregating at least one feature of the third point cloud; and generating a predicted occupancy state of at least one voxel of the third point cloud based on the aggregated features.
[0374] For some embodiments of the third example method, aggregating at least one feature of the third point cloud includes: repeating a concatenation process one or more times, the concatenation process including: performing a sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; performing a non-linear activation process on the convolutional output point cloud to generate a non-linear output point cloud; and if there is a next cycle of the concatenation process, preparing the non-linear output point cloud as the input point cloud, where the third point cloud is the input point cloud of the first cycle of the concatenation process, and where the last cycle of the concatenation process generates the aggregated features.
[0375] Some embodiments of the third example method may further include: adding the third point cloud to the ReLU output point cloud of the last cycle of the concatenation process.
[0376] For some embodiments of the third example method, aggregating at least one feature includes: performing a sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; and performing a non-linear activation process on the convolutional output point cloud to generate the aggregated features.
[0377] For some embodiments of the third example method, the non-linear activation process includes a rectifier linear unit (ReLU) activation process, and the non-linear output point cloud includes a ReLU output point cloud.
[0378] For some embodiments of the third example method, aggregating at least one feature of the third point cloud includes: repeating the first cascaded process one or more times, the first cascaded process including: performing a first sparse 3D convolution on a first input point cloud to generate a first convolution output point cloud; performing a first non-linear activation process on the first convolution output point cloud to generate a first non-linear output point cloud; and if there is a next cycle of the first cascaded process, preparing the first non-linear output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascaded process, wherein the last cycle of the first cascaded process generates a first cascaded process output; repeating the second cascaded process one or more times, the second cascaded process including: performing a second sparse 3D convolution on a second input point cloud to generate a second convolution output point cloud; performing a second non-linear activation process on the second convolution output point cloud to generate a second non-linear output point cloud; and if there is a next cycle of the second cascaded process, preparing the second non-linear output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascaded process, wherein the last cycle of the second cascaded process generates a second cascaded process output; cascading the first cascaded process output and the second cascaded process output to generate a cascaded output; and adding the third point cloud to the cascaded output to generate the aggregated feature.
[0379] For some embodiments of the third example method, aggregating at least one feature includes: repeating the first cascading process one or more times, the first cascading process including: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first rectified linear unit (ReLU) activation process on the first convolutional output point cloud to generate a first ReLU output point cloud; if there is a next cycle of the first cascading process, preparing the first ReLU output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; repeating the second cascading process one or more times, the second cascading process including: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second rectified linear unit (ReLU) activation process on the second convolutional output point cloud to generate a second ReLU output point cloud; and if there is a next cycle of the second cascading process, preparing the second ReLU output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process generates a second cascading process output; concatenating the first cascading process output and the second cascading process output to generate a concatenated output; and adding the third point cloud to the concatenated output to generate the aggregated feature.
[0380] For some embodiments of the third example method, aggregating at least one feature of the third point cloud includes: performing a self-attention process on the third point cloud; adding the third point cloud to the self-attention process output to generate an MLP process input; performing an MLP process on the MLP process input; and adding the MLP process input to the MLP process output to generate the aggregated feature;
[0381] For some embodiments of the third example method, the self-attention process generates an output feature based on k nearest neighbors of voxels of the third point cloud.
[0382] For some embodiments of the third example method, aggregating at least one feature of the third point cloud includes: performing a feature aggregation process two or more times.
[0383] For some embodiments of the third example method, the first set of neural network parameters and the second set of neural network parameters are the same set of neural network parameters, and the same set of neural network parameters is used by at least the first neural network and the second neural network.
[0384] For some embodiments of the third example method, the first set of neural network parameters and the second set of neural network parameters are a significant but identical set of neural network parameters.
[0385] A fourth example method according to some embodiments may include: obtaining a first point cloud; determining an occupancy status of at least one voxel of the first point cloud; removing voxels classified as empty from the first point cloud according to the determined occupancy status to generate a second point cloud; associating features of the second point cloud with context information to obtain a third point cloud; downsampling the third point cloud using an initial downsampling to obtain a fourth point cloud; and outputting the fourth point cloud as an encoded point cloud.
[0386] A fourth example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to: obtain a first point cloud; determine an occupancy status of at least one voxel of the first point cloud; remove voxels classified as empty from the first point cloud according to the determined occupancy status to generate a second point cloud; associate features of the second point cloud with context information to obtain a third point cloud; downsampling the third point cloud using an initial downsampling to obtain a fourth point cloud; and outputting the fourth point cloud as an encoded point cloud.
[0387] A fifth example method / apparatus according to some embodiments may include: accessing data including a first point cloud; and transmitting the data including the first point cloud.
[0388] A fifth example method / apparatus according to some embodiments may include: an access unit configured to access data including a first point cloud; and a transmitter configured to transmit the data including the first point cloud.
[0389] A sixth example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any one of the methods listed above.
[0390] A seventh example method / apparatus according to some embodiments may include at least one processor configured to perform any one of the methods listed above.
[0391] An eighth example method / apparatus according to some embodiments may include a computer-readable medium storing instructions for causing one or more processors to perform any one of the methods listed above.
[0392] A ninth example method / apparatus according to some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing at least one processor to perform any one of the methods listed above.
[0393] Example signals according to some embodiments may include bitstreams generated according to any of the methods listed above.
[0394] The present disclosure describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described and at least individual characteristics are shown, typically described in a way that may sound limiting. However, this is for clarity in the description and does not limit the disclosure or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Additionally, these aspects can also be combined and interchanged with aspects described in previously filed applications.
[0395] Aspects described and contemplated in the present disclosure can be implemented in many different forms. While some embodiments are specifically shown, other embodiments are contemplated and the discussion of specific embodiments does not limit the breadth of the specific implementations. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These aspects and other aspects can be implemented as methods, apparatuses, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having a bitstream generated according to any of the methods stored thereon.
[0396] In the present disclosure, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably. Generally, but not necessarily, the term "reconstruct" is used on the encoder side, while "decode" is used on the decoder side.
[0397] The terms HDR (High Dynamic Range) and SDR (Standard Dynamic Range) typically convey specific dynamic range values to a person of ordinary skill in the art. However, additional embodiments are also contemplated where a reference to HDR is understood to mean "higher dynamic range", and a reference to SDR is understood to mean "lower dynamic range". Such additional embodiments are not constrained by any specific value of the dynamic range, which may generally be associated with the terms "High Dynamic Range" and "Standard Dynamic Range".
[0398] This document describes various methods, and each method among the methods includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires steps or actions in a specific order, the order and / or usage of specific steps and / or actions can be modified or combined. Additionally, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding, and can occur, for example, before, during, or in an overlapping time period with the second decoding.
[0399] For example, various numerical values can be used in this disclosure. The specific values are for illustrative purposes, and the aspects are not limited to these specific values.
[0400] The embodiments described herein can be executed by a processor or other hardware or by computer software implemented by a combination of hardware and software. As a non-limiting example, these embodiments can be implemented by one or more integrated circuits. As a non-limiting example, the processor can be of any type suitable for the technical environment, and can encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0401] Various specific embodiments relate to decoding. As used in this disclosure, "decoding" can cover, for example, all or part of the process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as, for example, entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes or alternatively includes processes performed by the decoders of the various specific embodiments described in this disclosure, such as, for example, extracting a picture from a tiled (packed) picture, determining an upsampling filter to be used, then upsampling the picture, and flipping the picture back to its intended orientation.
[0402] As another example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally refer to a broader decoding process will be discerned based on the specific context described.
[0403] The various embodiments relate to encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in the present disclosure can encompass, for example, all or part of the process performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also include or alternatively include processes performed by the encoders of the various embodiments described in the present disclosure.
[0404] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in yet another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally refer to a broader encoding process will be discerned based on the context of the specific description.
[0405] When the drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding apparatus. Similarly, when the drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding method / process.
[0406] The embodiments and aspects described herein can be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even when discussed in the context of only a single form of embodiment (e.g., only as a method), the embodiments of the features discussed can be implemented in other forms (e.g., an apparatus or a program). The apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a processor generally referring to a processing device, which includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as, for example, a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices facilitating the communication of information among end users.
[0407] References to "one embodiment" or "an embodiment" or "one specific embodiment" or "specific embodiments" and their other variations mean that the specific features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one specific embodiment" or "in specific embodiments" and any other variations occurring throughout the present disclosure do not necessarily all refer to the same embodiment.
[0408] Additionally, the present disclosure may refer to "determining" various pieces of information. Determining information can include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.
[0409] In addition, the present disclosure may refer to "accessing" each piece of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information, among one or more of these.
[0410] Additionally, the present disclosure may refer to "receiving" each piece of information. Like "accessing", receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Further, "receiving" is typically involved in one way or another during operations such as, for example, storing information, processing information, sending information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information.
[0411] It should be understood that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any of the following, namely " / ", "and / or", and "at least one of", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first-listed option and the second-listed option (A and B), or the selection of only the first-listed option and the third-listed option (A and C), or the selection of only the second-listed option and the third-listed option (B and C), or the selection of all three options (A and B and C). This can be extended to as many items as are listed.
[0412] Moreover, as used herein, the word "signaling" particularly refers to indicating something to a corresponding decoder. For example, in some embodiments, an encoder signals a particular one of a plurality of parameters for region-based filter parameter selection for artifact removal filtering. Thus, in one embodiment, the same parameters are used on both the encoder side and the decoder side. Therefore, for example, the encoder may send (explicit signaling) a particular parameter to the decoder such that the decoder may use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as other parameters, signaling may be used without sending (implicit signaling) to simply allow the decoder to know and select the particular parameter. Bit savings are achieved in various embodiments by avoiding sending any actual functionality. It should be understood that signaling may be implemented in various ways. For example, in various embodiments, information is signaled to a corresponding decoder using one or more syntax elements, flags, etc. Although the foregoing relates to the verb form "signaling" of the word, the word may also be used herein as a noun "signal".
[0413] Specific implementations can generate various signals and format these signals to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the specific implementations. For example, the signal can be formatted to carry a bitstream of the embodiments. Such signals can be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting can include, for example, encoding a data stream and using the encoded data stream to modulate a carrier. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. Signals can be stored on a processor-readable medium.
[0414] Note that various hardware elements of one or more of the embodiments are referred to as “modules” that perform (i.e., execute, implement, etc.) the various functions described herein in connection with the corresponding modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) that a person of ordinary skill in the relevant art deems suitable for a given specific implementation. Each such module can also include executable instructions for performing one or more of the functions described as being performed by the corresponding module, and note that these instructions can take the form of or include the following instructions: hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., and can be stored in any suitable one or more non-transitory computer-readable media (such as those commonly referred to as RAM, ROM, etc.).
[0415] Although the features and elements have been described above in specific combinations, those of ordinary skill in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Additionally, the methods described herein can be implemented in a computer program, software, or firmware that is incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks and digital versatile disks (DVDs)). A processor associated with the software can be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A method, comprising: upsampling a first point cloud using an initial upsampling to obtain a second point cloud; associating the features of the second point cloud with context information to obtain a third point cloud; predicting the occupancy status of at least one voxel of the third point cloud; and removing the voxels classified as empty in the third point cloud according to the predicted occupancy status to generate a trimmed point cloud.
2. The method according to claim 1, wherein the initial upsampling includes nearest neighbor upsampling.
3. The method according to any one of claims 1-2, wherein associating features includes: concatenating the features of the second point cloud with the context information to obtain the third point cloud.
4. The method according to any one of claims 1-3, wherein the context information is per-voxel context information.
5. The method according to any one of claims 1-4, wherein the context information includes a context point cloud.
6. The method according to any one of claims 1-5, wherein the context information includes information about the second point cloud.
7. The method according to any one of claims 1-6, wherein the context information includes information about the voxel occupancy status of the second point cloud.
8. The method according to any one of claims 1-7, wherein the context information includes information about the position of sub-voxels relative to the position of parent voxels with respect to the first point cloud.
9. The method according to any one of claims 1 to 8, wherein the context information includes coordinate information about the positions of occupied voxels in at least one of the first point cloud and the second point cloud.
10. The method according to any one of claims 1-9, wherein the context information includes coordinate information, and wherein the coordinate information is in the form of one of Euclidean coordinates, spherical coordinates, and cylindrical coordinates.
11. The method according to any one of claims 1-10, wherein in addition to the information available for the initial upsampling of the first point cloud, the context information also provides known information about the first point cloud.
12. The method according to any one of claims 1 to 11, wherein the context information includes the bit depth of the second point cloud.
13. The method according to any one of claims 1-21, further comprising: performing feature decoding on an input point cloud and a first bitstream to generate the first point cloud.
14. The method according to claim 13, further comprising: performing feature aggregation on the trimmed point cloud to generate aggregated features; and performing a context-aware upsampling process on the aggregated features to generate a decoded point cloud.
15. The method according to claim 13, further comprising: performing a feature-to-residual transformation on the trimmed point cloud to generate a residual output; and adding the trimmed point cloud to the residual output to generate a decoded point cloud.
16. The method according to claim 15, further comprising: performing feature aggregation on the trimmed point cloud to generate aggregated features, wherein the feature-to-residual transformation is performed on the aggregated features.
17. The method according to any one of claims 1-16, wherein predicting the occupancy state is performed using a first neural network.
18. The method according to any one of claims 1-17, wherein predicting the occupancy state predicts the true occupancy state of at least one voxel.
19. The method according to any one of claims 1-18, wherein, predicting the occupancy state predicts the likelihood that the at least one voxel is occupied.
20. The method according to any one of claims 1-19, wherein the voxels of the third point cloud are removed using a voxel pruning process.
21. The method according to any one of claims 1-20, further comprising: aggregating at least one feature of the second point cloud.
22. The method according to any one of claims 1-21, wherein, predicting the occupancy state of at least one voxel comprises: aggregating at least one feature of the third point cloud; processing the aggregated features with a multi-layer perceptron (MLP) layer to generate an MLP layer output; performing a softmax process on the MLP layer output to generate a softmax output value; and performing thresholding on the softmax output value to generate the predicted occupancy state of at least one voxel of the third point cloud.
23. The method according to claim 22, wherein, the thresholding of the softmax output value converts a softmax output value greater than 0.5 into an output value of 1, and converts a softmax output value equal to or less than 0.5 into an output value of 0.
24. The method according to any one of claims 1-21, wherein, predicting the occupancy state of at least one voxel comprises: aggregating at least one feature of the third point cloud; and generating the predicted occupancy state of at least one voxel of the third point cloud based on the aggregated features.
25. The method according to any one of claims 22-24, wherein aggregating at least one feature comprises: repeating a concatenation process one or more times, the concatenation process comprising: performing sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; performing a non-linear activation process on the convolutional output point cloud to generate a non-linear output point cloud; and if there is a next cycle of the concatenation process, preparing the non-linear output point cloud as the input point cloud, wherein the third point cloud is the input point cloud of the first cycle of the concatenation process, and wherein the last cycle of the concatenation process generates the aggregated features.
26. The method according to claim 25, further comprising: adding the third point cloud to the ReLU output point cloud of the last cycle of the concatenation process.
27. The method according to any one of claims 22-24, wherein aggregating at least one feature comprises: performing sparse 3D convolution on an input point cloud to generate a convolutional output point cloud; and performing a non-linear activation process on the convolutional output point cloud to generate the aggregated features.
28. The method according to any one of claims 25 - 27, wherein the non-linear activation process includes a rectified linear unit (ReLU) activation process, and the non-linear output point cloud includes a ReLU output point cloud.
29. The method according to any one of claims 22 - 24, wherein aggregating at least one feature comprises: repeating the first cascading process one or more times, the first cascading process comprising: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first non-linear activation process on the first convolutional output point cloud to generate a first non-linear output point cloud; and if there is a next cycle of the first cascading process, preparing the first non-linear output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; repeating the second cascading process one or more times, the second cascading process comprising: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second non-linear activation process on the second convolutional output point cloud to generate a second non-linear output point cloud; and if there is a next cycle of the second cascading process, preparing the second non-linear output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process generates a second cascading process output; cascading the first cascading process output and the second cascading process output to generate a cascaded output; and adding the third point cloud to the cascaded output to generate the aggregated feature.
30. The method according to any one of claims 22 - 24, wherein aggregating at least one feature comprises: repeating the first cascading process one or more times, the first cascading process comprising: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first rectified linear unit (ReLU) activation process on the first convolutional output point cloud to generate a first ReLU output point cloud; and if there is a next cycle of the first cascading process, preparing the first ReLU output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; repeating the second cascading process one or more times, the second cascading process comprising: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second rectified linear unit (ReLU) activation process on the second convolutional output point cloud to generate a second ReLU output point cloud; and if there is a next cycle of the second cascading process, preparing the second ReLU output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, The last cycle of the second cascading process generates the second cascading process output; Cascade the first cascading process output and the second cascading process output to generate a cascaded output; and Add the third point cloud to the cascaded output to generate the aggregated feature.
31. The method according to any one of claims 22-24, wherein aggregating at least one feature comprises: Performing a self-attention process on the third point cloud; Adding the third point cloud to the self-attention process output to generate an MLP process input; Performing an MLP process on the MLP process input; and Adding the MLP process input to the MLP process output to generate the aggregated feature.
32. The method according to claim 31, wherein, The self-attention process generates output features based on the k nearest neighbors of the voxels of the third point cloud.
33. The method according to any one of claims 22 to 24, wherein aggregating at least one feature of the third point cloud comprises: Performing the feature aggregation process two or more times.
34. An apparatus, comprising: A processor; and A non-transitory computer-readable medium that stores instructions that, when executed by the processor, operate to cause the apparatus to: Upsample a first point cloud using initial upsampling to obtain a second point cloud; Associate the features of the second point cloud with context information to obtain a third point cloud; Predict the occupancy status of at least one voxel of the third point cloud; and Remove the voxels classified as empty in the third point cloud according to the predicted occupancy status to generate a trimmed point cloud.
35. The apparatus according to claim 34, wherein, The initial upsampling includes nearest neighbor upsampling.
36. The apparatus according to any one of claims 34 to 35, wherein associating features comprises: Cascade the features of the second point cloud with the context information to obtain the third point cloud.
37. A device, comprising: The apparatus according to claim 34; and At least one of the following: (i) an antenna configured to receive a signal that includes data representing the image, (ii) a band limiter configured to limit the received signal to a band that includes the data representing the image, or (iii) a display configured to display the image.
38. The apparatus according to claim 37, further comprising at least one of the following: a TV, a cellular phone, a tablet, and a set-top box (STB).
39. A computer-readable medium comprising instructions for causing one or more processors to perform the following operations: Upsample a first point cloud using initial upsampling to obtain a second point cloud; Associate the features of the second point cloud with context information to obtain a third point cloud; Predict the occupancy status of at least one voxel of the third point cloud; and Remove the voxels classified as empty in the third point cloud according to the predicted occupancy status to generate a trimmed point cloud.
40. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to: Use initial upsampling to upsample a first point cloud to obtain a second point cloud; Associate features of the second point cloud with context information to obtain a third point cloud; Predict an occupancy state of at least one voxel of the third point cloud; And Remove voxels classified as empty in the third point cloud according to the predicted occupancy state to generate a trimmed point cloud.
41. A method, Comprising: Performing context-aware upsampling of a first point cloud to determine an upsampled second point cloud, Wherein the context-aware upsampling comprises: Associating features of a third point cloud with context information, the third point cloud being at least partially based on an initial upsampled version of the first point cloud; and Removing voxels of a fourth point cloud predicted to be empty from the third point cloud at least partially based on the context information to generate an enlarged second point cloud.
42. A method, Comprising: Using initial upsampling to upsample a first point cloud to obtain a second point cloud; Associating features of the second point cloud with context information to obtain a third point cloud; Predicting an occupancy state of at least one voxel of the third point cloud, Wherein predicting the occupancy state of at least one voxel comprises: aggregating at least one feature of the third point cloud, Wherein aggregating at least one feature of the third point cloud comprises: using a first neural network, and Wherein using the first neural network to aggregate at least one feature of the third point cloud comprises: using a first set of neural network parameters with the first neural network; Removing voxels classified as empty in the third point cloud according to the predicted occupancy state to generate a trimmed point cloud; and Performing feature aggregation on the trimmed point cloud to generate aggregated features, Wherein performing the feature aggregation on the trimmed point cloud comprises: using a second neural network, Wherein using the second neural network to generate the aggregated features comprises: using a second set of neural network parameters with the second neural network, and Wherein the first set of neural network parameters is the same as the second set of neural network parameters.
43. The method according to claim 42, further Comprising: Aggregating at least one feature of the second point cloud.
44. The method according to claim 43, Wherein, Wherein aggregating at least one feature of the second point cloud comprises: using a third neural network, Wherein using the third neural network to aggregate at least one feature of the second point cloud comprises: using a third set of neural network parameters with the third neural network, and Wherein the third set of neural network parameters is the same as the first set of neural network parameters.
45. The method according to any one of claims 42-44, wherein the initial upsampling comprises nearest neighbor upsampling.
46. The method according to any one of claims 42 to 45, wherein associating features Comprises: Cascading the features of the second point cloud with the context information to obtain the third point cloud.
47. The method according to any one of claims 42 to 46, wherein the associated feature comprises: cascading the feature of the second point cloud with the context information to obtain the third point cloud.
48. The method according to any one of claims 42-47, wherein, the context information is per-voxel context information.
49. The method according to any one of claims 42 to 48, which further comprises: performing feature decoding on the input point cloud and the first bitstream to generate the first point cloud.
50. The method according to claim 49, further comprises: performing a context-aware upsampling process on the aggregated features to generate a decoded point cloud.
51. The method according to claim 49, further comprises: performing a feature-to-residual transformation on the trimmed point cloud to generate a residual output; and adding the trimmed point cloud to the residual output to generate a decoded point cloud.
52. The method according to claim 51, wherein, the feature-to-residual transformation is performed on the aggregated features.
53. The method according to any one of claims 42-52, wherein predicting the occupancy state predicts the true occupancy state of at least one voxel.
54. The method according to any one of claims 42-53, wherein predicting the occupancy state predicts the likelihood that the at least one voxel is occupied.
55. The method according to any one of claims 42-54, wherein the voxels of the third point cloud are removed using a voxel pruning process to remove voxels.
56. The method according to any one of claims 42-55, wherein, predicting the occupancy state of at least one voxel further comprises: processing the aggregated features with a multi-layer perceptron (MLP) layer to generate an MLP layer output; performing a softmax process on the MLP layer output to generate a softmax output value; and performing thresholding of the softmax output value to generate the predicted occupancy state of at least one voxel of the third point cloud.
57. The method according to claim 56, wherein, the thresholding of the softmax output value converts softmax output values greater than 0.5 into output value 1, and converts softmax output values equal to or less than 0.5 into output value 0.
58. The method according to any one of claims 42-55, wherein predicting the occupancy state of at least one voxel comprises: aggregating at least one feature of the third point cloud; and generating the predicted occupancy state of at least one voxel of the third point cloud based on the aggregated features.
59. The method according to any one of claims 56 to 58, wherein aggregating at least one feature of the third point cloud comprises: repeating the cascading process one or more times, the cascading process comprising: performing sparse 3D convolution on the input point cloud to generate a convolutional output point cloud; performing a non-linear activation process on the convolutional output point cloud to generate a non-linear output point cloud; and If there is a next cycle of the cascading process, prepare the non-linear output point cloud as the input point cloud, wherein the third point cloud is the input point cloud of the first cycle of the cascading process, and wherein the last cycle of the cascading process generates the aggregated feature.
60. The method according to claim 59, further comprising: adding the third point cloud to the ReLU output point cloud of the last cycle of the cascading process.
61. The method according to any one of claims 56 - 58, wherein aggregating at least one feature comprises: performing a sparse 3D convolution on the input point cloud to generate a convolutional output point cloud; and performing a non-linear activation process on the convolutional output point cloud to generate the aggregated feature.
62. The method according to any one of claims 59 - 61, wherein the non-linear activation process comprises a rectifier linear unit (ReLU) activation process, and the non-linear output point cloud comprises a ReLU output point cloud.
63. The method according to any one of claims 56 to 58, wherein aggregating at least one feature of the third point cloud comprises: repeating a first cascading process one or more times, the first cascading process comprising: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first non-linear activation process on the first convolutional output point cloud to generate a first non-linear output point cloud; and if there is a next cycle of the first cascading process, preparing the first non-linear output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; repeating a second cascading process one or more times, the second cascading process comprising: performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; performing a second non-linear activation process on the second convolutional output point cloud to generate a second non-linear output point cloud; and if there is a next cycle of the second cascading process, preparing the second non-linear output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process generates a second cascading process output; cascading the first cascading process output and the second cascading process output to generate a cascaded output; and adding the third point cloud to the cascaded output to generate the aggregated feature.
64. The method according to any one of claims 56 - 58, wherein aggregating at least one feature comprises: repeating a first cascading process one or more times, the first cascading process comprising: performing a first sparse 3D convolution on a first input point cloud to generate a first convolutional output point cloud; performing a first rectifier linear unit (ReLU) activation process on the first convolutional output point cloud to generate a first ReLU output point cloud; and If there is a next cycle of the first cascading process, prepare the first ReLU output point cloud as the first input point cloud, wherein the third point cloud is the first input point cloud of the first cycle of the first cascading process, wherein the last cycle of the first cascading process generates a first cascading process output; Repeat the second cascading process one or more times, the second cascading process including: Performing a second sparse 3D convolution on a second input point cloud to generate a second convolutional output point cloud; Performing a second rectifier linear unit (ReLU) activation process on the second convolutional output point cloud to generate a second ReLU output point cloud; and If there is a next cycle of the second cascading process, prepare the second ReLU output point cloud as the second input point cloud, wherein the third point cloud is the second input point cloud of the first cycle of the second cascading process, wherein the last cycle of the second cascading process generates a second cascading process output; Cascade the first cascading process output and the second cascading process output to generate a cascaded output; and Add the third point cloud to the cascaded output to generate the aggregated feature.
65. The method according to any one of claims 56 - 58, wherein aggregating at least one feature comprises: Performing a self - attention process on the third point cloud; Adding the third point cloud to the output of the self - attention process to generate an MLP process input; Performing an MLP process on the MLP process input; and Adding the MLP process input to the MLP process output to generate the aggregated feature; 66. The method according to claim 65, wherein the self - attention process generates output features based on K nearest neighbors of voxels of the third point cloud.
67. The method according to any one of claims 56 - 58, wherein aggregating at least one feature of the third point cloud comprises: Performing a feature aggregation process two or more times.
68. The method according to any one of claims 42 - 67, wherein, the first set of neural network parameters and the second set of neural network parameters are the same set of neural network parameters, and the same set of neural network parameters is used by at least the first neural network and the second neural network.
69. The method according to any one of claims 42 - 68, wherein, the first set of neural network parameters and the second set of neural network parameters are significantly but identical sets of neural network parameters.
70. A method, comprising: Obtaining a first point cloud; Determining the occupancy status of at least one voxel of the first point cloud; Removing the voxels classified as empty in the first point cloud according to the determined occupancy status to generate a second point cloud; Associating the features of the second point cloud with context information to obtain a third point cloud; Performing downsampling on the third point cloud using an initial downsampling to obtain a fourth point cloud; and Outputting the fourth point cloud as an encoded point cloud.
71. An apparatus, comprising: A processor; and A non - transitory computer - readable medium that stores instructions which, when executed by the processor, operate to cause the apparatus: Obtain a first point cloud; Determine the occupancy status of at least one voxel of the first point cloud; Remove the voxels classified as empty from the first point cloud according to the determined occupancy status to generate a second point cloud; Associate the features of the second point cloud with context information to obtain a third point cloud; Downsample the third point cloud using an initial downsampling to obtain a fourth point cloud; And Output the fourth point cloud as the encoded point cloud.
72. A method, comprising: Access data including a first point cloud; And Transmit the data including the first point cloud.
73. An apparatus, comprising: An access unit configured to access data including a first point cloud; And A transmitter configured to send the data including the first point cloud.
74. An apparatus, comprising: A processor; And A non-transitory computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to perform any of the methods according to claims 1-33 and 41-70.
75. An apparatus comprising at least one processor configured to perform the method according to any one of claims 1-33 and 41-70.
76. An apparatus comprising a computer-readable medium storing instructions for causing one or more processors to perform the method according to any one of claims 1-33 and 41-70.
77. An apparatus comprising at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform the method according to any one of claims 1-33 and 41-70.
78. A signal comprising a bitstream generated according to any one of claims 1-33 and 41-70.