Attribute coding for point cloud compression
By combining a deep learning-based encoder with a traditional attribute compression scheme, the problem of insufficient attribute data efficiency in point cloud compression is solved, achieving more efficient point cloud reconstruction and decoding, and improving compression performance.
Patent Information
- Application Number
- CN202480022407.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-02
- Filing Date
- 2024-04-03
- Publication Date
- 2025-11-14
AI Technical Summary
Existing point cloud compression technologies are inefficient in processing geometric and attribute data of point clouds, especially in the compression of attribute data, resulting in poor compression performance.
It employs a deep learning-based encoder and decoder framework, combined with a traditional attribute compression scheme, to process the geometric data of point clouds through geometric reconstruction and reduction, and uses a recoloring scheme to generate attribute data, allowing for flexible configuration between different encoders and decoders.
It improves the overall performance of point cloud compression, especially in attribute data compression, and achieves more efficient point cloud reconstruction and decoding.
Smart Images

Figure CN120958489A_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application No. 18 / 624,683, filed April 2, 2024, and U.S. Provisional Application No. 63 / 493,806, filed April 3, 2023, the entire contents of each of which are incorporated herein by reference. U.S. Patent Application No. 18 / 624,683, filed April 2, 2024, claims the benefit of U.S. Provisional Application No. 63 / 493,806, filed April 3, 2023. Technical Field
[0002] This disclosure relates to point cloud encoding and decoding. Background Technology
[0003] A point cloud is a collection of points in three-dimensional space. These points can correspond to points on objects within that space. Therefore, point clouds can be used to represent the physical content of three-dimensional space. Point clouds have practical applications in a wide variety of situations. For example, point clouds can be used in the context of autonomous vehicles to represent the position of objects on a road. In another example, point clouds can be used in the context of representing the physical content of an environment to locate virtual objects in augmented reality (AR) or mixed reality (MR) applications. Point cloud compression is the process of encoding and decoding point clouds. Encoding point clouds reduces the amount of data required to store and transmit them. Summary of the Invention
[0004] Generally, this disclosure describes techniques for point cloud decoding (e.g., encoding or decoding), including decoding geometric and attribute data of point cloud data. Decoding may include either or both encoding and / or decoding. Specifically, a deep learning-based encoder may be used to efficiently encode geometric information (e.g., point coordinates within the point cloud). Point attribute data (e.g., color, reflectivity, brightness, surface normals, etc.) typically comprises a larger amount of data compared to the geometric data. Therefore, a point cloud encoder may reconstruct the geometric data, then reduce both the geometric and attribute data before encoding the attribute data. A point cloud decoder may then decode and reconstruct the full-scale geometric data, reduce the geometric data, and use the reduced geometric data to decode the attribute data. Subsequently, the point cloud decoder may enlarge the attribute data and apply the enlarged attribute data to the full-scale geometric data to reconstruct the point cloud.
[0005] In one example, an apparatus for decoding (e.g., reconstructing) point cloud data includes: a memory configured to store point cloud data; and one or more processors implemented in a circuit and configured to: decode encoded point cloud geometry data of the point cloud to reconstruct the point cloud geometry data of the point cloud; reduce the point cloud geometry data to form reduced point cloud geometry data; and decode (e.g., reconstruct) attribute data of the point cloud using the reduced point cloud geometry.
[0006] In another example, a method for decoding (e.g., reconstructing) point cloud data includes: decoding encoded point cloud geometry data of the point cloud to reconstruct the point cloud geometry data of the point cloud; reducing the point cloud geometry data to form reduced point cloud geometry data; and using the reduced point cloud geometry to decode (e.g., reconstruct) attribute data of the point cloud.
[0007] In another example, a computer-readable storage medium stores instructions that, when executed, cause a processor to: decode encoded point cloud geometry data of a point cloud to reconstruct the point cloud geometry data of the point cloud; reduce the point cloud geometry data to form reduced point cloud geometry data; and decode (e.g., reconstruct) attribute data of the point cloud using the reduced point cloud geometry.
[0008] In another example, an apparatus for decoding point cloud data includes: components for decoding encoded point cloud geometry data of the point cloud to reconstruct the point cloud geometry data; components for reducing the point cloud geometry data to form reduced point cloud geometry data; and components for decoding (e.g., reconstructing) attribute data of the point cloud using the reduced point cloud geometry.
[0009] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. Attached Figure Description
[0010] Figure 1 This is a block diagram illustrating an example point cloud encoding and decoding system that can perform the techniques of this disclosure.
[0011] Figure 2 This is a block diagram illustrating an example point cloud encoder according to the technology of this disclosure.
[0012] Figure 3 This is a block diagram illustrating an example point cloud decoder according to the technology of this disclosure.
[0013] Figure 4 This is a conceptual diagram illustrating a specific example of a coding framework based on the technology of this disclosure.
[0014] Figure 5 This is a conceptual diagram illustrating a specific example of a decoding framework according to the technology of this disclosure.
[0015] Figure 6 This is a block diagram illustrating an example point cloud coding framework according to the technology of this disclosure.
[0016] Figure 7 This is a block diagram illustrating an example point cloud decoding framework according to the technology of this disclosure.
[0017] Figure 8A and Figure 8B This is a conceptual diagram illustrating an example of a reduced point cloud voxel.
[0018] Figure 9 This is a block diagram illustrating a set of example stages that can be included in a deep learning-based attribute upsampler.
[0019] Figure 10 This is a block diagram illustrating another set of example stages that can be included in a deep learning-based attribute upsampler.
[0020] Figure 11 This is a flowchart illustrating an example method for encoding point cloud data according to the technology of this disclosure.
[0021] Figure 12 This is a flowchart illustrating an example method for decoding point cloud data according to the technology of this disclosure.
[0022] Figure 13 This is a conceptual diagram illustrating a laser package such as a LIDAR sensor or other system, which includes one or more lasers that scan points in three-dimensional space.
[0023] Figure 14 This is a conceptual diagram illustrating an example ranging system 900 that can be used with one or more technologies disclosed herein.
[0024] Figure 15 This is a conceptual diagram illustrating an example of a vehicle-based scenario in which one or more technologies of this disclosure may be used.
[0025] Figure 16 This is a conceptual diagram illustrating an example extended reality system in which one or more techniques of this disclosure may be used.
[0026] Figure 17 This is a conceptual diagram illustrating an example mobile device system in which one or more technologies of this disclosure may be used.
[0027] Figure 18 and Figure 19This is a flowchart illustrating an example of a deep learning-based geometric encoder and decoder network. Detailed Implementation
[0028] Point clouds (PCs) are 3D data representations used for tasks such as virtual reality (VR) and mixed reality (MR), autonomous driving, and cultural heritage preservation. A point cloud is a collection of points in 3D space, represented by 3D coordinates (x, y, z) called geometry. Each point can also be associated with multiple attributes such as color, normal vector, and reflectivity. Depending on the target application and the point cloud acquisition method, point clouds can be classified into point cloud scenes and point cloud objects. Point cloud scenes can be captured using LiDAR sensors and can be acquired dynamically.
[0029] Point cloud objects can be further divided into static point clouds and dynamic point clouds. A static point cloud is a single object. A dynamic point cloud is a time-varying point cloud that includes a sequence of point cloud instances. Each instance of a dynamic point cloud is a static point cloud. Dynamic time-varying point clouds can be used for AR / VR, volumetric video streaming, and telepresence, and can be generated using 3D models (i.e., CGI) or captured from real-world scenes using various methods, such as multiple cameras with depth sensors around the object. These point clouds are dense, photorealistic point clouds, which can have a large number of points, especially in high-precision or large-scale captures (up to 60 frames per second (FPS), millions of points per frame). Therefore, efficient point cloud compression (PCC) is useful for practical applications in VR and MR.
[0030] The Moving Picture Experts Group (MPEG) has approved two PCC (point cloud compression) standards: (1) S. Schwarz, M. Preda, V. Baroncini, M. Budagavi, P. Cesar, PAChou, RA Cohen, M. Krivoku′ca, S. Lasserre, Z. Li et al., “Emerging MPEG standards for point cloud compression”, IEEE Journal on Emerging and Selected Topics in Circuits and Systems, Vol. 9, No. 1, pp. 133-148, 2018; and (2) D. Graziosi, O. Nakagami, S. Kuma, A. Zaghetto, T. Suzuki and A. Tabatabai, “An overview of ongoing point cloud compression standardization activities: Video-based (v-pcc) and geometry-based (g-pcc)”, APSIPA Transactions on Signal and Information Processing, Vol. 9, 2020. MPEG has approved the Geometry Based Point Cloud Compression (G-PCC) standard: "MPEG-PCC-TMC13: Geometry Based Point Cloud Compression G-PCC", in 2021, available at github.com / MPEGGroup / mpeg-pcc-tmc13. MPEG has also approved the Video Based Point Cloud Compression (V-PCC) standard: "MPEG-PCC-TMC2: Video Based Point Cloud Compression VPCC", in 2022, available at github.com / MPEGGroup / mpeg-pcc-tmc2.
[0031] G-PCC includes octree geometry decoding as a general geometry decoding tool and predictive geometry decoding (tree-based) tools for LiDAR-based point clouds. G-PCC is still developing methods based on triangle meshes or trisoups for approximating surfaces of 3D models. On the other hand, V-PCC encodes dynamic point clouds by projecting 3D points onto a 2D plane and then uses a video codec (e.g., High Efficiency Video Decoding (HEVC)) to encode each frame over time. MPEG has also proposed Common Test Conditions (CTCs) for evaluating test models: S. Schwarz, G. Martin-Cocher, D. Flynn, and M. Budagavi, “Common test conditions for point cloud compression,” document ISO / IEC JTC1 / SC29 / WG11w17766, Ljubljana, Slovenia, 2018.
[0032] As noted above, efficient point cloud compression is useful for applications such as virtual and mixed reality, autonomous driving, and cultural heritage. Some techniques, such as m59617: Anique Akhtar, Zhu Li, Geert Van der Auwera, Adarsh Krishnan Ramasubramonian, Luong Pham Van, Marta Karczewicz, Dynamic PointCloud Geometry Compression using Sparse Convolutions, MPEG-137Online, document m59617, April 2022, and m60307: Anique Akhtar, Zhu Li, Geert Van der Auwera, Adarsh Krishnan Ramasubramonian, Marta Karczewicz, [AI-3DGC][EE5.3 Test 2] Results dynamic point cloud compression, MPEG-139Online, document M60307, July 2022, utilize deep learning-based point cloud compression for dense dynamic point clouds by using a deep learning network consisting of encoder and decoder modules.
[0033] Within the context of this disclosure, deep learning can refer to the use of multiple hidden layers in an artificial neural network. A deep learning-based geometric encoder can refer to a geometric encoder comprising a computer-based neural network including multiple hidden layers. A deep learning-based attribute upsampler can refer to an attribute upsampler comprising a computer-based neural network including multiple hidden layers.
[0034] Research is underway using deep learning solutions to perform point cloud compression. Point clouds typically consist of a set of points and their attributes. These points correspond to locations in three-dimensional space, for example, having X, Y, and Z coordinates. These attributes can include, for example, color, reflectivity, brightness, surface normals, etc. Because point clouds possess both geometry and attributes, solutions have been proposed for point cloud geometry compression, point cloud attribute compression, and joint point cloud geometry and attribute compression.
[0035] Deep learning-based solutions generally perform well when applied to geometry compression. However, they still have shortcomings in point cloud attribute compression. The techniques disclosed herein include geometry compression using one codec and attribute compression using another. For example, deep learning-based geometry compression schemes can be combined with non-deep learning-based attribute compression schemes. These techniques can provide flexibility to the compression framework and create a robust baseline for deep learning-based point cloud attribute compression.
[0036] These techniques can also include multi-scale attribute compression and the use of deep learning-based post-processing. Heuristic tests have shown that most of the compressed bits are consumed by attribute decoding. Therefore, creating efficient attribute compression schemes is important for achieving good compression performance. Multi-scale attribute compression schemes can improve the overall compression performance of the decoding framework.
[0037] Some decoding techniques include lossy point cloud geometry compression schemes based on deep learning for dynamic point cloud compression. Lossy geometry schemes employ a prediction network to predict the latent representation of the current frame using previous frames. This framework performs P-frame inter-frame point cloud encoding, where the current frame is encoded with reference to previously decoded frames. The architecture is implemented using a sparse convolutional neural network (CNN) with sparse tensors. The architecture convolves the target coordinates to map the latent representation of the previous frame to the downsampled coordinates of the current frame, thereby predicting the feature embedding of the current frame. The framework compresses the predicted and actual features using a known probabilistic factorization entropy model, sending the residuals of the predicted and actual features. Compared to G-PCC and V-PCC, these techniques exhibit better geometry compression performance on dense point clouds and have efficient encoding / decoding runtime.
[0038] In one or more examples, this disclosure describes flexible configurations of deep learning-based frameworks where example techniques use a recoloring scheme (as an example) to generate attributes for a reconstructed point cloud and employ a conventional attribute compression scheme, rather than a joint geometry and attribute compression scheme. One or more example techniques described in this disclosure can utilize a G-PCC recoloring scheme to obtain attributes of a target point cloud from a source point cloud. The G-PCC recoloring scheme uses weighted distance-based nearest neighbor calculations in the source point cloud to compute the attributes of the target point cloud. Therefore, in one or more examples, these example techniques can allow the use of attribute compression from different codecs and geometry compression from different codecs.
[0039] Figure 1 This is a block diagram illustrating an example encoding and decoding system 100 capable of implementing the techniques of this disclosure. The techniques of this disclosure generally relate to decoding (encoding and / or decoding) point cloud data, i.e., supporting point cloud compression. Generally, point cloud data includes any data used for processing point clouds. This decoding can be effective in compressing and / or decompressing point cloud data.
[0040] like Figure 1 As shown, system 100 includes a source device 102 and a destination device 116. The source device 102 provides encoded point cloud data for decoding by the destination device 116. Specifically, in Figure 1 In this example, source device 102 provides point cloud data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can include any of a wide range of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, mobile phones (such as smartphones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, land or sea vehicles, spacecraft, aircraft, robots, LiDAR devices, satellites, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication.
[0041] exist Figure 1In the example, source device 102 includes a data source 104, a memory 106, a point cloud encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a point cloud decoder 300, a memory 120, and a data consumer 118. According to this disclosure, the point cloud encoder 200 of source device 102 and the point cloud decoder 300 of destination device 116 can be configured to apply the techniques of this disclosure related to attribute decoding for point cloud compression. Therefore, source device 102 represents an example of an encoding device, while destination device 116 represents an example of a decoding device. In other examples, source device 102 and destination device 116 may include other components or arrangements. For example, source device 102 may receive data (e.g., point cloud data) from an internal or external source. Similarly, destination device 116 may interface with an external data consumer without including the data consumer in the same device.
[0042] like Figure 1 The system 100 shown is merely an example. In general, other digital encoding and / or decoding devices can perform the techniques disclosed herein related to attribute decoding for point cloud compression. Source device 102 and destination device 116 are merely examples of such devices, where source device 102 generates decoded data for transmission to destination device 116. This disclosure refers to a “decoding” device as a device that performs the decoding (e.g., encoding and / or decoding) of data. Thus, point cloud encoder 200 and point cloud decoder 300 represent examples of decoding devices, specifically, encoder and decoder, respectively. In some examples, source device 102 and destination device 116 can operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes both encoding and decoding components. Therefore, system 100 can support one-way or two-way transmission between source device 102 and destination device 116, for example, for streaming, playback, broadcasting, telephone, navigation, and other applications.
[0043] Generally, data source 104 represents the source of data (i.e., raw, unencoded point cloud data) and can provide a series of sequential “frames” of data to point cloud encoder 200, which encodes the data in these frames. Data source 104 of source device 102 may include point cloud capture devices such as any of a variety of cameras or sensors (e.g., a 3D scanner or a light detection and ranging (LIDAR) device, one or more cameras), archives containing previously captured data, and / or data feed interfaces receiving data from data content providers. Alternatively or additionally, point cloud data may be computer-generated from a scanner, camera, sensor, or other data source. For example, data source 104 may generate computer graphics-based data as source data, or produce a combination of real-time data, archived data, and computer-generated data. In each case, point cloud encoder 200 encodes the captured data, pre-captured data, or computer-generated data. Point cloud encoder 200 may rearrange frames from the received order (sometimes referred to as “display order”) to a decoding order for decoding. Point cloud encoder 200 may generate one or more bitstreams including encoded data. Then, the source device 102 can output the encoded data to the computer-readable medium 110 via the output interface 108 for reception and / or retrieval by, for example, the input interface 122 of the destination device 116.
[0044] The memory 106 of the source device 102 and the memory 120 of the destination device 116 can represent general-purpose memory. In some examples, memory 106 and memory 120 can store raw data, such as raw data from data source 104 and raw decoded data from point cloud decoder 300. Additionally or alternatively, memory 106 and memory 120 can store software instructions that can be executed by, for example, point cloud encoder 200 and point cloud decoder 300. Although memory 106 and memory 120 are shown separately from point cloud encoder 200 and point cloud decoder 300 in this example, it should be understood that point cloud encoder 200 and point cloud decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memory 106 and memory 120 can store encoded data, such as output from point cloud encoder 200 and input to point cloud decoder 300. In some examples, portions of memory 106 and memory 120 can be allocated as one or more buffers, for example, for storing raw decoded and / or encoded data. For example, memory 106 and memory 120 can store data representing point clouds.
[0045] Computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium enabling source device 102 to transmit encoded data directly to destination device 116 in real time, for example, via a radio frequency network or a computer-based network. Output interface 108 can modulate the transmitted signal including the encoded data, and input interface 122 can demodulate the received transmitted signal according to a communication standard, such as a wireless communication protocol. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network (such as the Internet). The communication medium can include a router, switch, base station, or any other equipment that may be useful for facilitating communication from source device 102 to destination device 116.
[0046] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded data.
[0047] In some examples, source device 102 may output encoded data to file server 114 or another intermediate storage device that may store the encoded data generated by source device 102. Destination device 116 may access the stored data from file server 114 via streaming or downloading. File server 114 may be any type of server device capable of storing encoded data and sending it to destination device 116. File server 114 may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 may access the encoded data from file server 114 via any standard data connection including an internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing encoded data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming protocol, a downloading protocol, or a combination thereof.
[0048] Output interface 108 and input interface 122 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 can be configured to transmit data (such as encoded data) according to cellular communication standards (such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc.). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured according to specifications such as IEEE 802.11, IEEE 802.15 (e.g., ZigBee). TM ),Bluetooth TM Other wireless standards, such as standard 102, can be used to transmit data (such as encoded data). In some examples, source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include an SoC device for performing functions belonging to point cloud encoder 200 and / or output interface 108, and destination device 116 may include an SoC device for performing functions belonging to point cloud decoder 300 and / or input interface 122.
[0049] The technology disclosed herein can be applied to support encoding and decoding of any of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors and processing devices such as local or remote servers, geographic mapping or other applications.
[0050] The input interface 122 of the destination device 116 receives an encoded bitstream from a computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, etc.). The encoded bitstream may include signaling information defined by the point cloud encoder 200 and also used by the point cloud decoder 300, such as syntax elements having values describing the characteristics and / or processing of the decoding unit (e.g., slice, picture, picture group, sequence, etc.). The data consumer 118 uses the decoded data. For example, the data consumer 118 may use the decoded data to determine the location of a physical object. In some examples, the data consumer 118 may include a display for presenting an image based on the point cloud.
[0051] The point cloud encoder 200 and point cloud decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic components, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Each of the point cloud encoder 200 and point cloud decoder 300 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (encoder-decoder) in the respective device. Devices including point cloud encoder 200 and / or point cloud decoder 300 may include one or more integrated circuits, microprocessors, and / or other types of devices.
[0052] The point cloud encoder 200 and point cloud decoder 300 can operate according to decoding standards such as the Video Point Cloud Compression (V-PCC) standard or the Geometric Point Cloud Compression (G-PCC) standard. This disclosure can generally relate to the decoding (e.g., encoding and decoding) of images, thus including processes for encoding or decoding data. Encoded bitstreams typically include a series of values for syntax elements representing decoding decisions (e.g., decoding modes).
[0053] This disclosure may generally relate to "signaling" certain information, such as syntax elements. The term "signaling" can generally refer to communication of values for syntax elements and / or other data for decoding encoded data. That is, the point cloud encoder 200 may signal the values of syntax elements in the bitstream. Typically, signaling refers to generating values in the bitstream. As noted above, the source device 102 may transmit the bitstream to the destination device 116 substantially in real time or not in real time (such as when syntax elements are stored in storage device 112 for later retrieval by the destination device 116).
[0054] ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) is investigating the potential need for standardization of point cloud decoding techniques with compression capabilities significantly exceeding current methods and will work towards creating such a standard. This exploration is being conducted collaboratively within a team called the 3D Graphics Team (3DG) to evaluate compression technology designs proposed by experts in the field.
[0055] Point cloud compression activities are categorized into two distinct approaches. The first is “Video Point Cloud Compression” (V-PCC), which segments a 3D object and projects these segments onto multiple 2D planes (represented as “patches” in 2D frames), which are then decoded by an older 2D video codec, such as the High Efficiency Video Decoding (HEVC) (ITU-T H.265) codec. The second approach is “Geometry-Based Point Cloud Compression” (G-PCC), which directly compresses the 3D geometry (i.e., the location of a set of points in 3D space) and associated attribute values (for each point associated with the 3D geometry). G-PCC addresses point cloud compression in both Category 1 (static point clouds) and Category 3 (dynamically acquired point clouds). The latest draft of the G-PCC standard is available in the G-PCC DIS (ISO / IEC JTC1 / SC29 / WG11 w19088, Brussels, Belgium, January 2020), and the codec specification is available in the G-PCC Codec Specification v6 (ISO / IEC JTC1 / SC29 / WG11 w19091, Brussels, Belgium, January 2020).
[0056] A point cloud contains a set of points in 3D space and can have attributes associated with those points. Attributes can be color information such as R, G, B or Y, Cb, Cr, or reflectivity information, or other attributes. Point clouds can be captured by various cameras or sensors (such as LiDAR sensors and 3D scanners) and can also be computer-generated. Point cloud data is used in a variety of applications, including but not limited to architecture (modeling), graphics (3D models for visualization and animation), and the automotive industry (LiDAR sensors for navigation aids).
[0057] The 3D space occupied by point cloud data can be enclosed by virtual bounding boxes. The position of a point within the bounding box can be represented with a certain precision; therefore, the position of one or more points can be quantized based on this precision. At the smallest level, the bounding box is divided into voxels, which are the smallest spatial units represented by a unit cube. A voxel in the bounding box can be associated with zero, one, or more points. The bounding box can be divided into multiple cubic / cuboid regions, which can be called tiles. Each tile can be decoded into one or more slices. Dividing the bounding box into slices and tiles can be based on the number of points in each partition, or on other considerations (e.g., a specific region can be decoded into a tile). Slice regions can be further subdivided using splitting decisions similar to those in video codecs.
[0058] Figure 2This is a block diagram illustrating an example point cloud encoder 200. The modules shown are logical and do not necessarily correspond one-to-one with the implementation code in the reference implementation of the G-PCC codec (i.e., the TMC13 test model software studied by ISO / IEC MPEG (JTC 1 / SC 29 / WG 11)).
[0059] In both the point cloud encoder 200 and the point cloud decoder 300, the point cloud locations are decoded first. Attribute decoding depends on the decoding geometry. The compressed geometry is typically represented as an octree with leaf-level access from the root to the individual voxels.
[0060] At each node of the octree, occupancy is signaled for one or more of its child nodes (up to eight nodes) when not inferred. Multiple neighborhoods are specified, including (a) nodes sharing a face with the current octree node, (b) nodes sharing a face, edge, or vertex with the current octree node, etc. Within each neighborhood, occupancy of the current node or its child nodes can be predicted using the occupancy of the node and / or its child nodes. For points sparsely filled in some nodes of the octree, the codec also supports a direct decoding mode, in which the 3D position of the point is directly encoded. Signaling flags can be sent to indicate that direct mode is signaled. At the lowest level, the number of points associated with an octree node / leaf node can also be decoded.
[0061] Once the geometry is decoded, the attributes corresponding to that geometric point are also decoded. When there are multiple attribute points corresponding to a reconstructed / decoded geometric point, the attribute values representing the reconstructed point can be derived.
[0062] There are three attribute decoding methods in G-PCC: Region Adaptive Hierarchical Transform (RAHT) decoding, interpolation-based hierarchical nearest neighbor prediction (prediction transform), and interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform). RAHT and lifting are typically used for Class 1 data, while prediction is typically used for Class 3 data. However, any method can be used for any data, and like the geometry codec in G-PCC, the attribute decoding method used to decode point clouds is specified in the bitstream.
[0063] Attribute decoding can be performed at each level of detail (LOD), where a finer representation of the point cloud attribute is obtained at each level of detail. Each level of detail can be specified based on a distance metric from neighboring nodes or based on the sampling distance.
[0064] At the point cloud encoder 200, the residual obtained as the output of the attribute decoding method can be quantized. The residual can be obtained by subtracting the attribute value from the prediction derived from the attribute values of the points in the neighborhood of the current point and based on the attribute values of the previously encoded points. The quantized residual can be decoded using context-adaptive arithmetic decoding.
[0065] exist Figure 2 In the example, the point cloud encoder 200 includes a deep learning-based geometric encoder 202, a geometric reconstruction unit 206, a geometric reduction unit 210, a recoloring unit 212, a color transformation unit 204, a region adaptive hierarchical transformation (RAHT) unit 218, a LOD generation unit 220, a boosting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.
[0066] like Figure 2 As shown in the example, the point cloud encoder 200 can obtain the set of point locations and attribute sets in a point cloud. The point cloud encoder 200 can obtain data from data source 104 ( Figure 1 The point cloud encoder 200 obtains a set of point locations and a set of attributes. These locations may include the coordinates of the points in the point cloud. These attributes may include information about the points in the point cloud, such as color, reflectivity, intensity, etc., associated with the points. The deep learning-based geometric encoder 202 of the point cloud encoder 200 can generate a geometric bitstream 203 containing an encoded representation of the point locations in the point cloud. The point cloud encoder 200 can also generate an attribute bitstream 205 containing an encoded representation of the attribute set.
[0067] After encoding the geometric information, the geometric reconstruction unit 206 can decode and reconstruct the geometric information. The geometric reduction unit 210 can reduce the geometric information. As noted above, an octree can be used to represent the geometric information. The root node of the octree can be divided into eight child nodes. For each node of the octree, the node can be further divided into eight child nodes, including dividing the node in half along the X, Y, and Z dimensions. Such division can continue until, for example, the node of the smallest size of the octree is reached, which is called the leaf node of the octree. That is, the leaf node has no child nodes and is undivided.
[0068] To reduce an octree, the geometric reduction unit 210 can determine that a node has eight leaf nodes and determine the number of occupied leaf nodes among those eight. If this number exceeds a threshold (e.g., zero), the geometric reduction unit 210 can represent the node as an occupied leaf node in the reduced octree. Otherwise, if the number is less than or equal to the threshold, the geometric reduction unit 210 can represent the node as an unoccupied leaf node in the reduced octree. For example, if the threshold is zero, all leaf nodes of the eight leaf nodes will need to be unoccupied to represent the node as an unoccupied leaf node in the reduced octree; otherwise, the node will be represented as an occupied leaf node in the reduced octree.
[0069] After reducing the geometric information, the recoloring unit 212 can apply attribute information to the points of the reduced octree. For example, the recoloring unit 212 can reduce the original attribute information to a similar degree as the geometric information. Such reduction may include culling attribute data, mixing attribute data, or otherwise reducing attribute data so that the reduced attribute data can be applied to the reduced geometry.
[0070] Color transformation unit 204 can transform the color information of an attribute to different domains. For example, color transformation unit 204 can transform color information from the RGB color space to the YCbCr color space.
[0071] Furthermore, RAHT unit 218 can apply RAHT decoding to the attributes of the reconstructed points. In some examples, according to RAHT, the attributes of the block at the 2x2x2 point location are obtained and transformed along one direction to obtain four low-frequency nodes (L) and four high-frequency nodes (H). Subsequently, the four low-frequency nodes (L) are transformed in a second direction to obtain two low-frequency nodes (LL) and two high-frequency nodes (LH). The two low-frequency nodes (LL) are transformed along a third direction to obtain one low-frequency node (LLL) and one high-frequency node (LLH). The low-frequency node LLL corresponds to the DC coefficients, and the high-frequency nodes H, LH, and LLH correspond to the AC coefficients. The transformation in each direction can be a 1-D transformation with two coefficient weights. The low-frequency coefficients can be obtained as coefficients for the next higher-level 2x2x2 block for the RAHT transformation, and the AC coefficients are encoded without modification; such transformations continue until the top root node. The weights to be used for these coefficients are calculated from top to bottom using a tree traversal for encoding; the transformation order is bottom to top. These coefficients can then be quantized and decoded.
[0072] Alternatively or additionally, LOD generation unit 220 and lifting unit 222 may apply LOD processing and lifting to the attributes of the reconstructed points, respectively. LOD generation is used to break down the attributes into different refinement levels. Each refinement level provides a refinement of the attributes of the point cloud. The first refinement level provides a coarse approximation and contains few points; subsequent refinement levels typically contain more points, and so on. Refinement levels can be constructed using distance-based metrics, or one or more other classification criteria (e.g., subsampling from a specific order). Thus, all reconstructed points can be included in the refinement levels. Each level of detail can be generated by taking the union of all points up to a specific refinement level: for example, LOD1 is obtained based on refinement level RL1, LOD2 is obtained based on RL1 and RL2, and so on, with the union of RL1, RL2, ..., RLN yielding LODN. In some cases, LOD generation may be followed by a prediction scheme (e.g., a prediction transform) in which the attributes associated with each point in the LOD are predicted based on a weighted average of the previous points, and the residuals are quantized and entropy-decoded. The enhancement scheme is built on a predictive transformation mechanism, in which update operators are used to update the coefficients and adaptive quantization of the coefficients is performed.
[0073] RAHT unit 218 and lifting unit 222 can generate coefficients based on these attributes. Coefficient quantization unit 224 can quantize the coefficients generated by RAHT unit 218 or lifting unit 222. Arithmetic encoding unit 226 can apply arithmetic decoding to the syntax elements representing the quantized coefficients. Point cloud encoder 200 can output these syntax elements in attribute bit stream 205. Attribute bit stream 205 may also include other syntax elements, including syntax elements that are not arithmetically encoded.
[0074] In some examples, the geometry reduction unit 210 can reduce geometric data by a certain amount represented by a specific value, such as a reduction factor. The point cloud encoder 200 can encode this value as a parameter in a parameter set, such as a Sequence Parameter Set (SPS) or Attribute Parameter Set (APS), a slice header, a frame header, or another High-Level Syntax (HLS). In some examples, the reduction factor can also indicate the amount of attribute data to be reduced before recoloring. In some examples, the point cloud encoder 200 can encode a second value, separate from the first value, that indicates the amount of attribute data to be reduced.
[0075] Figure 3 This is a block diagram illustrating an example point cloud decoder 300. Figure 3In the example, the point cloud decoder 300 includes a deep learning-based geometric decoder 302, an attribute arithmetic decoding unit 304, a geometric reduction unit 306, an inverse quantization unit 308, a RAHT unit 314, a LoD generation unit 316, an inverse boosting unit 318, an inverse color transformation unit 322, and a point cloud magnification unit 324.
[0076] The point cloud decoder 300 can receive a geometric bitstream 203 and an attribute bitstream 205. A deep learning-based geometric decoder 302 typically decodes the geometric data of the geometric bitstream 203. The attribute arithmetic decoding unit 304 of the decoder 300 can apply arithmetic decoding (e.g., context-adaptive binary arithmetic decoding (CABAC) or other types of arithmetic decoding) to the syntax elements in the attribute bitstream 205 to decode the attribute bitstream 205.
[0077] Generally, attribute bitstream 250 represents a reduced version of attribute data relative to the geometric data in geometry bitstream 203. Therefore, geometry reduction unit 306 can, for example, reduce the reconstructed geometric data from deep learning-based geometry decoder 302 based on a reduction value. Point cloud decoder 300 can decode the reduction value from high-level syntax (HLS) data such as sequence parameter sets (SPS), attribute parameter sets (APS), slice headers, frame headers, etc. Geometry reduction unit 306 can reduce the geometric data based on a reduction factor.
[0078] Additionally, the inverse quantization unit 308 can inverse quantize attribute values. Attribute values may be based on syntax elements obtained from the attribute bitstream 205 (e.g., including syntax elements decoded by the attribute arithmetic decoding unit 304).
[0079] Depending on how the attribute values are encoded, RAHT unit 314 can perform RAHT decoding to determine the color values of points in the point cloud based on the inverse-quantized attribute values. RAHT decoding is performed from the top to the bottom of the tree. At each level, low-frequency and high-frequency coefficients derived from the inverse-quantization process are used to derive composition values. At leaf nodes, the derived values correspond to the attribute values of the coefficients. The point weight derivation process is similar to that used at point cloud encoder 200. Alternatively, LOD generation unit 316 and inverse lifting unit 318 can use level-of-detail techniques to determine the color values of points in the point cloud. LOD generation unit 316 decodes each LOD, thus giving a progressively finer representation of the point's attributes. In the case of prediction transform, LOD generation unit 316 derives the prediction of the point from a weighted sum of points previously reconstructed in the same LOD or earlier. LOD generation unit 316 can add the prediction to the residual (obtained after inverse-quantization) to obtain the reconstructed values of the attributes. When using an enhancement scheme, the LOD generation unit 316 may also include update operators to update the coefficients used to derive attribute values. In this case, the LOD generation unit 316 may also apply inverse adaptive quantization.
[0080] Furthermore, the inverse color transformation unit 322 can apply an inverse color transformation to color values. The inverse color transformation can be the reverse of the color transformation applied by the color transformation unit 204 of the encoder 200. For example, the color transformation unit 204 can transform color information from the RGB color space to the YCbCr color space. Correspondingly, the inverse color transformation unit 322 can transform color information from the YCbCr color space to the RGB color space.
[0081] After decoding both geometric and attribute information, the point cloud upscaling unit 324 can reconstruct the point cloud. Specifically, the point cloud upscaling unit 324 can upscale the attribute information to the scale of the geometric information. The point cloud upscaling unit 324 can upscale the attribute information according to a reduction factor. Alternatively, the point cloud decoder 300 can decode individual values representing the amount of upscaling to be applied to the attribute information. In some examples, the point cloud upscaling unit 324 can be a deep learning-based attribute upsampling unit. Finally, the point cloud upscaling unit 324 can apply the upscaled attribute information to a point in the geometric information to reconstruct that point.
[0082] Example Figure 2 and Figure 3Various units assist in understanding the operations performed by the encoder 200 and decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or combinations thereof. A fixed-function circuit is a circuit that provides specific functionality and is pre-defined for the operations that can be performed. A programmable circuit is a circuit that can be programmed to perform various tasks and provides flexible functionality for the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more units in the unit may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units in the unit may be integrated circuits.
[0083] As described above, machine learning techniques such as deep learning can be used for encoding and decoding point cloud data. The techniques disclosed herein include the use of a deep learning-based lossy point cloud geometry compression scheme for dynamic point cloud compression. The lossy geometry scheme employs a prediction network to predict the latent representation of the current frame using previous frames. This example technique performs P-frame inter-frame point cloud encoding, where the current frame is encoded with reference to previously decoded frames. This architecture can be implemented using a sparse convolutional neural network (CNN) with sparse tensors. This example architecture convolves the target coordinates to map the latent representation of the previous frame to the downsampled coordinates of the current frame, thereby predicting the feature embedding of the current frame. The encoder sends the residuals of the predicted and actual features by compressing the predicted features and actual features using a known probabilistic factorization entropy model. Compared to G-PCC and V-PCC, these machine learning techniques demonstrate better compression performance on dense point clouds with efficient encoding / decoding runtime.
[0084] Using machine learning for encoding / decoding (e.g., compression / decompression) point cloud data can present challenges. Point cloud compression codecs often perform well in geometric decoding but poorly in attribute decoding, and vice versa. This also applies to deep learning-based point cloud compression schemes, which often perform well in geometric compression but fail or are inadequate in attribute compression.
[0085] In one or more examples described in this disclosure, to achieve decoding efficiency, such as deep learning-based decoding efficiency, there may be benefits in that geometric compression and attribute compression are separated from the compression framework, and that a geometric compression method from one codec is then combined with an attribute compression method from another codec to achieve better compression performance. In one or more examples, this disclosure describes a flexible configuration in the compression framework in which the point cloud encoder 200 and the point cloud decoder 300 employ deep learning-based geometric compression and conventional attribute compression methods for point cloud compression.
[0086] Further flexibility can be provided by reducing and subsequently amplifying the attribute information. Most of the bit rate in previous encoding schemes has been consumed by the attribute information. Any improvement in the attribute compression scheme can significantly improve the overall decoding efficiency of the framework. Therefore, according to the techniques of this disclosure, the point cloud encoder 200 can reduce the attribute information before encoding, and the point cloud decoder 300 can decode and then amplify the attribute information.
[0087] High-Level Syntax (HLS) data can be signaled by the point cloud encoder 200 and received by the point cloud decoder 300. The HLS data may include attribute decoding types for decoding recoloring attributes so that the point cloud decoder 300 can reconstruct attribute values. The decoding method (e.g., Region Adaptive Hierarchical Transformation (RAHT) of G-PCC) can be signaled as an identifier in a parameter set (e.g., Sequence Parameter Set (SPS) or Attribute Parameter Set (APS)). The portion of the bit stream carrying the decoded attribute bits is, for example, NALU. A list of decoding methods can be specified, each providing a means of decoding the attributes of the point cloud (optionally, these decoding methods can also decode geometry in a lossless manner). An index of this list can be signaled in the bit stream to indicate the decoding method used to decode the attributes. This index can be signaled in the parameter set (e.g., APS, SPS) or other means. When a deep learning mechanism is used for recoloring or decoding, the parameters / coefficients corresponding to the recoloring or decoding can also be signaled in the bit stream.
[0088] Figure 4 This is a block diagram illustrating example code framework 400. Figure 5 This is a block diagram illustrating the example decoding framework 500. Figure 4 In the encoding framework 400, a deep learning-based geometric encoder 404, a G-PCC recoloring unit 406, and a G-PCC lossless geometric and lossy attribute encoder 408 are included. Generally speaking, the encoding framework 400 can correspond to... Figure 1 and Figure 2The point cloud encoder 200 includes a deep learning-based geometry encoder 404 that may correspond to a deep learning-based geometry encoder 202, a G-PCC recoloring unit 406 that may correspond to a recoloring unit 212, and a G-PCC lossless geometry and lossy attribute encoder 408 that may correspond to any or all of the color transformation unit 204, RAHT unit 218, LOD generation unit 220, lifting unit 222, coefficient quantization unit 224, and arithmetic encoding unit 226. In this example, the original point cloud 402 is provided to both the deep learning-based geometry encoder 404 and the G-PCC recoloring unit 406. The deep learning-based geometry encoder 404 encodes the geometric information of the original point cloud 402 and forms a geometric bitstream 410 that includes the encoded geometric information of the original point cloud 402.
[0089] The deep learning-based geometry encoder 404 also decodes and reconstructs geometric information, such as an octree indicating whether points exist in a node (i.e., whether a node is occupied). Occupied nodes can be divided into eight child nodes, each of which can include an indication of occupancy. The G-PCC recoloring unit 406 can receive the reconstructed geometric information from the deep learning-based geometry encoder 404 and recolor the reconstructed geometry using the geometric and attribute information of the original point cloud 402. Then, the G-PCC lossless geometry and lossy attribute encoder 408 can encode the recolorized, reconstructed geometry to form an attribute bitstream 412.
[0090] Figure 5 The decoding framework 500 includes a deep learning-based geometry decoder 502 (which may correspond to a deep learning-based geometry decoder 302) and a G-PCC decoder 504 (which may correspond to an attribute arithmetic decoding unit 304, an inverse quantization unit 308, a RAHT unit 314, an LOD generation unit 316, an inverse boosting unit 318, and an inverse color transformation unit 322). In this example, the deep learning-based geometry decoder 502 receives a geometry bitstream 508. The deep learning-based geometry decoder 502 decodes the geometry bitstream 508 and reconstructs the geometry information. The G-PCC decoder 504 receives an attribute bitstream 510 and uses the reconstructed geometry information to decode the attribute bitstream 510 and reconstruct the point cloud. For example, the G-PCC decoder 504 can apply the decoded attribute information to the reconstructed geometry information to form a recolored reconstructed point cloud 506.
[0091] exist Figure 4 and Figure 5In this model, encoding framework 400 uses a deep learning-based encoder-decoder to compress the geometry, and decoding framework 500 uses a deep learning-based encoder-decoder to decompress the geometry to obtain a reconstructed point cloud. The reconstructed point cloud differs geometrically from the original point cloud and may not be readily usable as attributes of the original point cloud. In some examples, encoding framework 400 performs compression, and decoding framework 500 can use a G-PCC recoloring scheme to modify the attributes of the original point cloud to create newer attributes for the reconstructed point cloud. The recoloring point cloud has both the geometry of the reconstructed point cloud and attributes derived from the original point cloud. The recoloring point cloud is decoded using G-PCC, where the geometry is encoded losslessly and the attributes are encoded lossily. The reconstructed geometry and reconstructed attributes are combined to obtain the reconstructed point cloud.
[0092] The following changes / flexibility can be added Figure 4 and Figure 5 The frame shown is an example. Although... Figure 4 and Figure 5 Examples of deep learning-based encoders and decoders are shown, but geometric encoders or decoders are not necessarily limited to deep learning-based geometric encoders or decoders; any geometric encoder or decoder can be used.
[0093] For recoloring, the encoding framework 400 performs compression, and the decoding framework 500 can employ a weighted distance-based nearest neighbor search recoloring scheme used in the G-PCC standard (G-PCC Standard: WG 7, MPEG 3D Graphics Coding, G-PCC codec description, document N00271, January 2022). The recoloring scheme can modify attributes to suit newer geometry. Any recoloring scheme that modifies the values of attributes and / or their correspondence with geometry can be used. The "recoloring" algorithm is not limited to "recoloring" of color attributes (e.g., RGB or YCbCr), but more generally can be an algorithm that recalculates attribute values (such as normal vectors, reflectivity, etc.) from a point location in one geometry to a point location in another geometry. Deep learning mechanisms can also be applied to perform recoloring.
[0094] like Figure 4The illustration shows a G-PCC lossless geometric and lossy attribute encoder. These example techniques are not limited to G-PCC lossless geometric and lossy attribute encoding; for example, G-PCC Region Adaptive Hierarchical Transform (RAHT). This method can be employed in conjunction with any encoding scheme that includes another deep learning-based encoding or uses V-PCC. Furthermore, this example can use any “lossy attribute encoder” instead of employing “lossless geometric and lossy attribute encoding.” Deep learning mechanisms can also be applied to perform decoding.
[0095] Encoding framework 400 and decoding framework 500 use a geometry encoder / decoder from a single codec and an attribute encoder / decoder from a separate codec to create a complete codec framework that outperforms two individual codecs. In lossy point cloud compression, the geometry changes after compression, and attaching attributes from individual codecs to the corresponding geometry of those attributes is challenging. Therefore, in encoding framework 400 and point cloud decoding framework 500, a recoloring scheme is used to attach attributes and their corresponding geometry.
[0096] Figure 6 This is a block diagram illustrating an example point cloud encoding framework 600 according to the technology of this disclosure. In this example, the point cloud encoding framework 600 includes a deep learning-based geometry encoder 604 (which may correspond to a deep learning-based geometry encoder 202), a geometry reduction unit 606 (which may correspond to a geometry reduction unit 210), a recoloring unit 608 (which may correspond to a recoloring unit 212), and an attribute encoder 610 (which may correspond to a color transformation unit 204, a RAHT unit 218, a LOD generation unit 220, a boosting unit 222, a coefficient quantization unit 224, and an arithmetic encoding unit 226). The point cloud encoding framework 600 differs from the encoding framework 400 in that the point cloud encoding framework 600 includes a geometry reduction unit 606, and the attribute encoder 610 encodes a reduced version of the attribute information, as discussed in more detail below.
[0097] Generally, a deep learning-based geometry encoder 604 receives geometric information from the original point cloud 602, while a recoloring unit 608 receives both geometric and attribute information from the original point cloud 602. The deep learning-based geometry encoder 604 encodes the geometric information to form a geometry bitstream 612. The deep learning-based geometry encoder 604 decodes and reconstructs the geometric information, and provides the reconstructed geometric information to a geometry reduction unit 606. The geometry reduction unit 606 then reduces the geometric information and provides the reduced geometric information to the recoloring unit 608. The recoloring unit 608 can form a recolorized reduced point cloud from the reduced geometric information and the original geometric and attribute information of the original point cloud 602. Then, an attribute encoder 610 encodes the attribute information of the recolorized reduced point cloud to form an attribute bitstream 614.
[0098] Figure 7 This is a block diagram illustrating an example point cloud decoding framework 700 according to the technology of this disclosure. In this example, the point cloud decoding framework 700 includes a deep learning-based geometry decoder 702, a geometry reduction unit 704, an attribute decoder 706, and a deep learning-based attribute upsampler 708. The deep learning-based geometry decoder 702 may correspond to a deep learning-based geometry decoder 302, the geometry reduction unit 704 may correspond to a geometry reduction unit 306, the attribute decoder 706 may correspond to an attribute arithmetic decoding unit 304, an inverse quantization unit 308, a RAHT unit 314, a LOD generation unit 316, an inverse boosting unit 318, and an inverse color transformation unit 322, and the deep learning-based attribute upsampler 708 may correspond to a point cloud magnification unit 324.
[0099] Generally, a deep learning-based geometry decoder 702 receives a geometry bitstream 710. The deep learning-based geometry decoder 702 decodes the geometry information in the geometry bitstream 710 and reconstructs geometry information from the decoded geometry information. A geometry reduction unit 704 reduces the geometry information to form reduced geometry information and provides the reduced geometry information to an attribute decoder 706. The attribute decoder 706 receives an attribute bitstream 712 including encoded attribute information and decodes the encoded attribute information. The attribute decoder 706 can provide the decoded attribute information to a deep learning-based attribute upsampler 708, which can also receive the original reconstructed geometry information and upscale the attribute information to the scale of the reconstructed geometry information. The deep learning-based attribute upsampler 708 can also apply the upscaled attribute information to the reconstructed geometry information to form a reconstructed point cloud 714.
[0100] According to the technology disclosed herein, the attribute encoder 610 encodes the reduced attributes to save attribute bits, which can improve decoding efficiency. A deep learning-based geometry encoder 604 is used to compress geometric information, which is then decoded and reconstructed to obtain a reconstructed point cloud. The geometry reduction unit 606 can use various reduction factors (e.g., 1, 2, 4, 8, etc.) to reduce the reconstructed geometric information, as discussed in more detail below. Then, the recoloring unit 608 recolors the reduced geometry using the original point cloud attributes according to a recoloring scheme (such as the G-PCC recoloring scheme). The attribute encoder 610 can then encode the recoloring point cloud attributes using an attribute decoding scheme (lossy or lossless) (such as G-PCC RAHT or G-PCC prediction / lifting transform).
[0101] At the point cloud decoding framework 700, the geometry bitstream 710 is decoded by a deep learning-based geometry decoder 702. The reconstructed geometry is reduced by a geometry reduction unit 704 and then provided to an attribute decoder 706. The attribute decoder 706 can perform inverse GPCC RAHT or prediction / boosting transforms. The attribute decoder 706 can append the decoded reduced attributes to the corresponding reduced reconstructed geometry of these attributes to obtain a reduced point cloud. As an example, a deep learning-based point cloud attribute upsampler is used to upsample the attributes and map the upsampled attributes to the reconstructed geometry.
[0102] In some examples, explicit reduction of the reconstructed geometry may not be necessary. For instance, the geometry decoding process itself may include one or more reduced versions of the geometry, which are then magnified / processed to obtain the reconstructed geometry. These reduced versions can be passed to the attribute decoder, and the reconstructed geometry is used instead of the reduced version. This avoids the need to perform explicit reduction on the reconstructed geometry, thus saving processing time and resources.
[0103] Figure 8A and Figure 8B This is a conceptual diagram illustrating an example of a reduced point cloud voxel. Figure 8A An example voxel 800 is depicted, comprising daughter voxels 802A, 802B, and 802C, each of which is occupied, and the others are unoccupied. In this example, reduction of voxel 800 results in a reduced voxel 804 being occupied, since daughter voxels 802A, 802B, and 802C are occupied. Figure 8B An example voxel 810 is depicted, comprising all unoccupied sub-voxels. Thus, in this example, reduction of voxel 810 results in a reduced voxel 812 that is not occupied.
[0104] In some examples, geometric reduction unit 306 or geometric reduction unit 704 can employ KxKxK voxel mesh reduction, where K is a value defining the reduction factor. If K is 2, the point cloud is divided into a 2x2x2 voxel mesh, where each voxel is a point in 3D space. Then, 8 voxels within the 2x2x2 voxel mesh are merged into a single voxel. In some examples, a final voxel is considered occupied if any voxel within the 8 voxels is also occupied. In some examples, a final voxel is considered occupied if the number of occupied sub-voxels exceeds a threshold.
[0105] Figure 9 This is a block diagram illustrating an example set of stages that can be included in a deep learning-based attribute upsampler (such as deep learning-based attribute upsampler 708 or point cloud magnification unit 324). In this example, the set of stages includes a deep learning layer 902, an upsampling layer 910, and a deep learning layer 920. Deep learning layer 902 initially processes the reduced point cloud geometry and attribute information 900, providing the result to upsampling layer 910. Upsampling layer 910 then upsamples the result of deep learning layer 902 using reconstructed geometric data 904. Deep learning layer 920 then processes the upsampled geometry and attribute information to produce a reconstructed point cloud 930.
[0106] Figure 10 This is a block diagram illustrating an example set of stages that can be included in a deep learning-based attribute upsampler (such as deep learning-based attribute upsampler 708 or point cloud magnification unit 324). In this example, the set of stages includes a deep learning layer 1002 (which may correspond to deep learning layer 902), an upsampling layer 1010 (which may correspond to upsampling layer 910), and a deep learning layer 1020 (which may correspond to deep learning layer 920). Specifically, deep learning layer 102 includes a sparse convolutional (SConv) 3x3x3 layer 1004, an initial residual block (IRB) layer 1006, and an SConv 3x3x3 layer 1008. In this example, upsampling layer 1010 is a convolution on a target coordinate 5x5x5 layer. In this example, deep learning layer 1020 includes an SConv 3x3x3 layer 1022, an IRB layer 1024, and an SConv layer 1026. These layers process the reduced point cloud geometry and attributes 1000 and the reconstructed geometric data 1012 in sequence to produce the reconstructed point cloud 1030.
[0107] For illustrative purposes, in Figure 10The example uses sparse convolutional layers with a 3x3x3 kernel size, but other deep learning-based layers can be used additionally or alternatively. Similarly, while a convolution on target coordinates with a 5x5x5 kernel size is used as an example, other upsampling techniques can be used instead. Convolutions on target coordinate layers can be used to map features from reduced geometry to magnified geometry. If the downsampling / upsampling factor is large, multiple consecutive upsampling layers can be used, or multiple consecutive attribute upsamplers can be used to upsampling attributes.
[0108] For attribute upsampling, mean squared error (MSE) loss can be used instead of classification loss during training. MSE loss compares the predicted attribute with the original attribute to aid in training the network. Training is not limited to MSE loss; any loss function that compares and attempts to minimize the distance between the predicted and original attributes can be used. The Adam optimizer can be used to train the network. The network is not limited to the Adam optimizer and can alternatively or additionally include root mean square propagation (RMSProp), stochastic gradient descent (SGD), adaptive gradient algorithm (Adagrad), or any other such optimizer.
[0109] Attribute upsampling does not necessarily have to be based on deep learning. Examples of non-deep learning-based property upsampling include those discussed below: Marc Alexa, Johannes Behr, Daniel Cohen-Or, Shachar Fleishman, David Levin, and Claudio T. Silva, Computing and rendering point set surfaces, IEEE Trans. Vis. & Comp. Graphics, 9(1):3-15, 2003; Yaron Lipman, Daniel Cohen-Or, David Levin, and Hillel Tal-Ezer, Parameterization-free projection for geometry reconstruction, ACM Trans. on Graphics (SIGGRAPH), 26(3):22:1-5, 2007; Hui Huang, Dan Li, Hao Zhang, Uri Ascher, and Daniel Cohen-Or, Consolidation of unorganized point clouds for surface reconstruction, ACM Trans. On Graphics (SIGGRAPH Asia), 28(5):176:1-7, 2009; Hui Huang, Shihao Wu, and Minglun Gong, Daniel Cohen-Or, Uri Ascher, and Hao Zhang, Edge-aware point set resampling, ACM Trans. on Graphics, 32(1):9:1-12, 2013; and Shihao Wu, Hui Huang, Minglun Gong, Matthias Zwicker, and Daniel Cohen-Or, Deep points consolidation, ACM Trans. on Graphics (SIGGRAPH Asia), 34(6):176:1-13, 2015. Figure 10 The example illustrates how different deep learning-based layers can achieve the same effect when using a fully convolutional network with sparse convolutions. Similarly, for upsampling layers, other deep learning layers (such as transposed convolutions, deconvolutions, unpooling, etc.) can be used to achieve similar results.
[0110] The attributes in the framework are not limited to color information. For example, these techniques can be applied to decoding / upsampling surface normal information, reflectivity, intensity, etc. Color space conversion can be applied to individual modules or the entire framework. For example, during recoloring, the YCbCr color space can be employed by an attribute encoder, an attribute decoder, and / or a deep learning-based attribute upsampling unit. The point cloud encoder 200 can encode syntax elements representing the color spaces associated with individual modules or the entire framework in the bitstream, and similarly, the point cloud decoder 300 can decode and use the syntax elements to determine which color spaces should be used in which modules.
[0111] Point cloud encoder 200 and point cloud decoder 309 can decode the reduction factor so that point cloud decoder 300 can determine how much adaptive geometric reduction to perform and how much attribute upsampling to perform. If no reduction is performed at point cloud encoder 200, there is no need for upsampling at decoder, and upsampling can be used. Figure 3 and Figure 4 The architecture is shown. The reduction factor can be signaled in the parameter set, slice, or other syntax structure within the bitstream (e.g., in the Sequence Parameter Set (SPS) or Attribute Parameter Set (APS)). Although referred to as the downsampling factor (from the encoder's perspective), the point cloud decoder 300 can use this factor for other operations, including upsampling. In some examples, two reduction factors can be signaled in the bitstream: one for geometry and one for attributes. In some examples, a single downsampling factor can be signaled for both geometry and attributes.
[0112] The attribute decoding technique type used to decode the recolored attributes can be signaled to the decoder side so that the point cloud decoder 300 can reconstruct the attribute values. The decoding method (e.g., Region Adaptive Hierarchical Transformation (RAHT) of G-PCC) can be signaled as an identifier in the parameter set (e.g., Sequence Parameter Set (SPS) or Attribute Parameter Set (APS)). The portion of the bitstream carrying the decoded attribute bits can be, for example, a Network Abstraction Unit (NALU).
[0113] A list of decoding techniques can be specified. Each decoding technique can provide a way to decode the properties of the point cloud (optionally, these decoding techniques can also decode the geometry in a lossless manner). An index in this list can be signaled in the bitstream to indicate the decoding technique used to decode the properties. This index can be signaled in a parameter set (e.g., APS, SPS) or another HLS.
[0114] When deep learning mechanisms are used for recoloring or decoding, signals can also be sent in the bitstream to notify the parameters / coefficients corresponding to the recoloring or decoding.
[0115] Figure 11 This is a flowchart illustrating an example method for encoding point cloud data according to the technology of this disclosure. About Figure 2 To explain using the point cloud encoder 200 Figure 11 The method. Other point cloud encoding devices (such as those conforming to...) Figure 6 The point cloud encoding devices of the 600 point cloud encoding framework can execute this method or a similar method.
[0116] Initially, the point cloud encoder 200 encodes the geometric information of the point cloud (1100). For example, the point cloud encoder 200 can use a deep learning-based point cloud encoding technique to encode the geometric information. Then, the point cloud encoder 200 can decode and reconstruct the geometric information (1102). The point cloud encoder 200 can also reduce the geometric information (1104). The point cloud encoder 200 can use the attribute information of the point cloud to recolor the reduced geometric information (1106). Then, the point cloud encoder 200 can encode the reduced attribute information (1108). In some examples, the point cloud encoder 200 can further encode HLS information, which specifies, for example, a reduction factor for the geometric information and / or an upsampling factor for the attribute information.
[0117] so, Figure 11 The method represents an example of a method for decoding point cloud information, including: decoding encoded point cloud geometry data of a point cloud to reconstruct point cloud geometry data of the point cloud; reducing the point cloud geometry data to form reduced point cloud geometry data; and using the reduced point cloud geometry to decode attribute data of the point cloud.
[0118] Figure 12 This is a flowchart illustrating an example method for decoding point cloud data according to the technology of this disclosure. About Figure 3 The point cloud decoder 300 is used to interpret Figure 12 The method. Other point cloud decoding devices (such as those conforming to...) Figure 7 The point cloud decoding framework 700 and those point cloud decoding devices can execute this method or a similar method.
[0119] Initially, the point cloud decoder 300 decodes the geometric information (1200). The point cloud decoder 300 can also decode reduction factors in, for example, High-Level Syntax (HLS) data. The point cloud decoder 300 can, for example, reduce the geometric information according to the reduction factor (1202). Then, the point cloud decoder 300 can decode the attribute information (1204). Then, the point cloud decoder 300 can, for example, upsample the attribute information using a deep learning-based attribute upsampler and / or an upsampling factor included in the bitstream (1206).
[0120] so, Figure 12 The method represents an example of a method for decoding point cloud information, including: decoding encoded point cloud geometry data of a point cloud to reconstruct point cloud geometry data of the point cloud; reducing the point cloud geometry data to form reduced point cloud geometry data; and using the reduced point cloud geometry to decode attribute data of the point cloud.
[0121] Figure 13 This is a conceptual diagram illustrating a laser package 1300, such as a LiDAR sensor or other system, which includes one or more lasers for scanning points in three-dimensional space. The laser package 1300 may correspond to... Figure 7 LiDAR 380. Data source 104 ( Figure 1 It may include laser package 1300.
[0122] like Figure 13 As shown, a laser package 1300 (i.e., a sensor that scans points in 3D space) can be used to capture point clouds. However, it should be understood that some point clouds are not generated by an actual LiDAR sensor, but can be encoded as if they were generated by an actual LiDAR sensor. Figure 13 In the example, laser package 1300 includes a LiDAR head 1302 comprising a plurality of lasers 1304A-1304E (collectively referred to as "lasers 1304") arranged at different angles relative to the origin in a vertical plane. Laser package 1300 is rotatable about a vertical axis 1308. Laser package 1300 can use the returned laser light to determine the distance and position of points in a point cloud. The laser beams 1306A-1306E (collectively referred to as "laser beams 1306") emitted by lasers 1304 of laser package 1300 can be characterized by a set of parameters. The distances indicated by arrows 1310 and 1312 represent example laser correction values for lasers 1304B and 1304A, respectively.
[0123] Laser 1300 can be used to acquire both geometric data and attribute data of the points within that geometric data. According to the technology disclosed herein, point cloud geometric data can be reduced, and the reduced point cloud geometric data can then be used to decode attribute data.
[0124] Figure 14 This is a conceptual diagram illustrating an example ranging system 1400 that can be used with one or more technologies disclosed herein. Figure 14In one example, the ranging system 1400 includes an illuminator 1402 and a sensor 1404. The illuminator 1402 may emit light 1406. In some examples, the illuminator 1402 may emit light 1406 as one or more laser beams. Light 1406 may be of one or more wavelengths, such as infrared wavelengths or visible light wavelengths. In other examples, light 1406 is not a coherent laser. When light 1406 encounters an object (such as object 1408), light 1406 produces a return light 1410. The return light 1410 may include backscattered light and / or reflected light. The return light 1410 may pass through a lens 1411, which guides the return light 1410 to generate an image 1412 of object 1408 on the sensor 1404. The sensor 1404 generates a signal 1414 based on the image 1412. The image 1412 may include points (e.g., as shown by...). Figure 14 The set of points represented by the small dots in image 1412.
[0125] In some examples, illuminator 1402 and sensor 1404 may be mounted on a rotating structure, allowing illuminator 1402 and sensor 1404 to capture a 360-degree view of the environment. In other examples, ranging system 1400 may include one or more optical components (e.g., mirrors, collimators, diffraction gratings, etc.) that enable illuminator 1402 and sensor 1404 to detect objects within a specific range (e.g., up to 360 degrees). Although Figure 14 The example shows only a single illuminator 1402 and sensor 1404, but the ranging system 1400 may include multiple sets of illuminators and sensors.
[0126] In some examples, illuminator 1402 generates a structured light pattern. In such examples, ranging system 1400 may include multiple sensors 1404 on which corresponding images of the structured light pattern are formed. Ranging system 1400 can use the differences between the images of the structured light pattern to determine the distance to object 1408 from which the structured light pattern backscatters. When object 1408 is relatively close to sensor 1404 (e.g., 0.2 meters to 2 meters), the structured light-based ranging system can have a high level of accuracy (e.g., sub-millimeter accuracy). This high level of accuracy can be useful in facial recognition applications such as unlocking mobile devices (e.g., mobile phones, tablets, etc.) and for security applications.
[0127] In some examples, the ranging system 1400 is a time-of-flight (ToF) based system. In some examples of the ToF-based ranging system 1400, an illuminator 1402 generates pulses of light. In other words, the illuminator 1402 can modulate the amplitude of the emitted light 1406. In such examples, a sensor 1404 detects the return light 1410 from the pulse of light 1406 generated by the illuminator 1402. The ranging system 1400 can then determine the distance to an object 1408 from which the light 1406 backscatters, based on the delay between the emission of the light 1406 and the detection of the light, and the known speed of light in air. In some examples, the illuminator 1402 can modulate the phase of the emitted light 1406 instead of (or in addition to) modulating the amplitude of the emitted light 1406. In such an example, sensor 1404 can detect the phase of the return light 1410 from object 1408, and use the speed of light and the time difference between when illuminator 1402 generates light 1406 at a particular phase and when sensor 1404 detects the return light 1410 at that particular phase to determine the distance to a point on object 1408.
[0128] In other examples, point clouds can be generated without using illuminator 1402. For example, in some examples, sensor 1404 of ranging system 1400 may include two or more optical cameras. In such examples, ranging system 1400 can use the optical cameras to capture a stereo image of the environment including object 1408. Ranging system 1400 (e.g., point cloud generator 1420) can then calculate differences between positions in the stereo image. Ranging system 1400 can then use these differences to determine distances to positions shown in the stereo image. Based on these distances, point cloud generator 1420 can generate a point cloud.
[0129] Sensor 1404 can also detect other properties of object 1408, such as color and reflectivity information. Figure 14 In the example, point cloud generator 1420 can generate a point cloud based on signal 1418 generated by sensor 1404. Ranging system 1400 and / or point cloud generator 1420 can form data source 104 ( Figure 1 Part of ).
[0130] Figure 15 This is a conceptual diagram illustrating an example of a vehicle-based scenario in which one or more technologies of this disclosure can be used. Figure 15 In the example, vehicle 1500 includes laser package 1502, such as a LIDAR system. Laser package 1502 can be used with laser package 600 (… Figure 13 Implemented in the same way. Although Figure 15The example is not shown, but vehicle 1500 may also include a data source (such as data source 104). Figure 1 )) and G-PCC encoders (such as G-PCC encoder 200 ( Figure 1 )).exist Figure 15 In the example, laser package 1502 emits laser beams 1504, which are reflected from pedestrians 1506 or other objects on the road. The data source of vehicle 1500 can generate a point cloud based on the signal generated by laser package 1502. A G-PCC encoder of vehicle 1500 can encode the point cloud to generate a bit stream 1508, such as... Figure 2 Geometric potential flow and Figure 2 The attribute bitstream. Bitstream 1508 may contain far fewer bits than the unencoded point cloud obtained by the G-PCC encoder. The output interface of vehicle 1500 (e.g., output interface 108) Figure 1 The bit stream 1508 can be sent to one or more other devices. Therefore, the vehicle 1500 may be able to send the bit stream 1508 to other devices faster than uncoded point cloud data. Additionally, the bit stream 1508 may require less data storage capacity.
[0131] The techniques disclosed herein can further reduce the number of bits in bitstream 1508. For example, by reducing the point cloud geometry data and then using the reduced point cloud geometry data to decode the attribute data of the point cloud, the amount of attribute data to be encoded can be significantly reduced, thereby reducing the number of bits in bitstream 1508. However, by subsequently recoloring the full-scale decoded set of geometry data using the reduced attribute data, the resulting reconstructed point cloud can represent a high-resolution reproduction of the decoded original point cloud.
[0132] exist Figure 15 In the example, vehicle 1500 can send bit stream 1508 to another vehicle 1510. Vehicle 1510 may include a G-PCC decoder, such as G-PCC decoder 300. Figure 1 The G-PCC decoder of vehicle 1510 can decode bitstream 1508 to reconstruct a point cloud. Vehicle 1510 can use the reconstructed point cloud for various purposes. For example, vehicle 1510 can determine, based on the reconstructed point cloud, that pedestrian 1506 is on the road ahead of vehicle 1500 and therefore begin to decelerate, for example, even before the driver of vehicle 1510 becomes aware that pedestrian 1506 is on the road. Thus, in some examples, vehicle 1510 can perform autonomous navigation operations, generate notifications or warnings, or perform another action based on the reconstructed point cloud.
[0133] Additionally or alternatively, vehicle 1500 may send bit stream 1508 to server system 1512. Server system 1512 may use bit stream 1508 for various purposes. For example, server system 1512 may store bit stream 1508 for subsequent reconstruction of the point cloud. In this example, server system 1512 may use the point cloud together with other data (e.g., vehicle telemetry data generated by vehicle 1500) to train an autonomous driving system. In other examples, server system 1512 may store bit stream 1508 for subsequent reconstruction for forensic accident investigations (e.g., if vehicle 1500 collides with pedestrian 1506).
[0134] Figure 16 This is a conceptual diagram illustrating an example extended reality system in which one or more technologies of this disclosure may be used. Extended reality (XR) is a term used to cover a range of technologies including augmented reality (AR), mixed reality (MR), and virtual reality (VR). Figure 16 In the example, a first user 1600 is located at a first location 1602. User 1600 wears an XR headset 1604. Alternatively, user 1600 may use a mobile device (e.g., a mobile phone, tablet, etc.). The XR headset 1604 includes a depth sensor, such as a LiDAR system, that detects the position of a point on an object 1606 at location 1602. The data source of the XR headset 1604 may use signals generated by the depth sensor to generate a point cloud representation of the object 1606 at location 1602. The XR headset 1604 may include a G-PCC encoder (e.g., Figure 1 The G-PCC encoder 200 is configured to encode point clouds to generate bit stream 1608.
[0135] The techniques disclosed herein can further reduce the number of bits in bitstream 1608. For example, by reducing the point cloud geometry data and then using the reduced point cloud geometry data to decode the attribute data of the point cloud, the amount of attribute data to be encoded can be significantly reduced, thereby reducing the number of bits in bitstream 1608. However, by subsequently recoloring the full-scale decoded set of geometry data using the reduced attribute data, the resulting reconstructed point cloud can represent a high-resolution reproduction of the decoded original point cloud.
[0136] XR headset 1604 can send bitstream 1608 (e.g., via a network, such as the Internet) to XR headset 1610 worn by user 1612 at second location 1614. XR headset 1610 can decode bitstream 1608 to reconstruct a point cloud. XR headset 1610 can use the point cloud to generate an XR visualization (e.g., AR visualization, MR visualization, VR visualization) representing object 1606 at location 1602. Thus, in some examples, such as when XR headset 1610 generates a VR visualization, user 1612 at location 1614 can have a 3D immersive experience at location 1602. In some examples, XR headset 1610 can determine the location of a virtual object based on the reconstructed point cloud. For example, XR headset 1610 can determine, based on the reconstructed point cloud, that the environment (e.g., location 1602) includes a flat surface, and then determine that a virtual object (e.g., a cartoon character) should be positioned on that flat surface. The XR headset 1610 can generate XR visualizations in which virtual objects are located at a defined position. For example, the XR headset 1610 can show a cartoon character sitting on a flat surface.
[0137] Figure 17 This is a conceptual diagram illustrating an example mobile device system in which one or more technologies of this disclosure may be used. Figure 17 In the example, mobile device 1700, such as a mobile phone or tablet, includes a depth sensor, such as a LiDAR system, which detects the location of points on object 1702 in the environment of mobile device 1700. The data source of mobile device 1700 can use the signals generated by the depth sensor to generate a point cloud representation of object 1702. Mobile device 1700 may include a G-PCC encoder (e.g., Figure 1 The G-PCC encoder 200 is configured to encode point clouds to generate bitstream 1704. Figure 17 In the example, mobile device 1700 can send a bit stream to remote device 1706 (such as a server system or other mobile device). Remote device 1706 can decode the bit stream 1704 to reconstruct a point cloud. Remote device 1706 can use the point cloud for various purposes. For example, remote device 1706 can use the point cloud to generate an environment map of mobile device 1700. For example, remote device 1706 can generate a map of the interior of a building based on the reconstructed point cloud. In another example, remote device 1706 can generate an image (e.g., computer graphics) based on the point cloud. For example, remote device 1706 can use points in the point cloud as vertices of a polygon and use the color attributes of the points as the basis for coloring the polygon. In some examples, remote device 1706 can use the point cloud to perform facial recognition.
[0138] Figure 18 and Figure 19 This is a flowchart illustrating an example of a deep learning-based geometric encoder and decoder network. Specifically, Figure 18 An example deep learning-based geometric encoder network is depicted, and Figure 19 An example deep learning-based geometry decoder network is depicted. Figure 18 In the example, the deep learning-based geometric encoder network includes: processing block 1800, which includes 16x3 convolutions. 3 Layers and Corrected Linear Units (ReLUs); Processing block 1802, which includes 32x3 convolutions. 3 / 2↓ layers, ReLU, and raw residual blocks (IRB); processing block 1804, which includes 32x3 convolutions. 3 / 2 layers, first ReLU, convolution 64x3 3 / 2↓ layer, second ReLU and IRB; and processing block 1806, which includes a 64x3 convolution. 3 Layer, first ReLU, convolution 32x3 3 / 2↓ layer, second ReLU, PRB and convolution 8x3 3 layer.
[0139] exist Figure 19 In the example, the deep learning-based geometry decoder network includes: processing block 1900, which comprises transposed convolutions of 64x3 3 / 2↑ layers, first ReLU, convolution 64x3 3 Layer, second ReLU and IRB; processing block 1902, which includes transposed convolutions 32x3 3 / 2↑ layers, first ReLU, convolution 32x3 3 Layer, second ReLU and IRB; processing block 1904, which includes transposed convolutions of 16x3 3 / 2↑ layers, first ReLU, convolution 16x3 3 Layers, second ReLU and IRB; and classifiers 1906, 1908 and 1910.
[0140] exist Figure 18 and Figure 19 In the example, conv Nx3 3 This refers to a 3x3x3 convolutional filter with N channels (where N can be, for example, 8, 16, 32, etc.). IRB represents the initial residual block. T-Conv represents a transposed convolution with a 2x stride. 2↑ represents a stride of 2 for amplification, while 2↓ represents a stride of 2 for reduction.
[0141] Deep learning-based networks can be trained using loss functions (e.g., three loss functions, one for each scale). Classifiers 1906, 1908, and 1910 can use binary cross-entropy loss as the classification loss to classify each voxel as occupied or unoccupied. For example, the following loss function can be used:
[0142]
[0143] Among them O v It is the baseline truth value of whether voxel v is occupied (1) or unoccupied (0).
[0144] Along with the reconstruction loss (as explained above with binary cross-entropy), the bit rate loss can be used to optimize rate distortion. The overall network can be trained end-to-end with joint reconstruction and bit rate losses. The Adam optimizer can be used, where the learning rate decays from 0.0008 to 0.00001. Figure 18 and Figure 19 In the example, if the classification loss is replaced by the mean squared error (MSE) loss, such a network would form a deep learning-based attribute upsampler / downsampler network. That is, such a network would be trained to process (downsample / upsample) attribute data, rather than being trained to determine the location of points.
[0145] The following clauses illustrate various examples of the technology disclosed herein.
[0146] Clause 1: An apparatus for decoding point cloud data, the apparatus comprising: a memory configured to store point cloud data; and one or more processors implemented in a circuit and configured to: decode encoded point cloud geometry data of the point cloud to reconstruct point cloud geometry data of the point cloud; reduce the point cloud geometry data to form reduced point cloud geometry data; and decode attribute data of the point cloud using the reduced point cloud geometry.
[0147] Clause 2: The device according to Clause 1, wherein, in order to decode the attribute data of the point cloud, the one or more processors are configured to encode the attribute data of the point cloud, and wherein the one or more processors are further configured to encode the point cloud geometry data using a deep learning-based geometry encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data.
[0148] Clause 3: The device according to Clause 2, wherein the one or more processors are further configured to encode a value representing a reduction amount to be applied to the point cloud geometry data, wherein, in order to reduce the point cloud geometry data, the one or more processors are configured to reduce the point cloud geometry data according to the value representing the reduction amount.
[0149] Clause 4: The device according to Clause 1, wherein, in order to decode the attribute data of the point cloud, the one or more processors are configured to decode the attribute data of the point cloud to form reduced point cloud attribute data, and wherein the one or more processors are further configured to amplify the reduced point cloud attribute data.
[0150] Clause 5: The device according to Clause 4, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to use a deep learning-based attribute upsampler to amplify the reduced point cloud attribute data.
[0151] Clause 6: The device according to Clause 4, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to apply a 5x5x5 convolutional target coordinate layer to the reduced point cloud attribute data.
[0152] Clause 7: The device according to Clause 4, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to reconstruct amplified point cloud attribute data of the point cloud, and wherein the one or more processors are further configured to apply the amplified point cloud attribute data to the point cloud geometry to reconstruct the point cloud.
[0153] Clause 8: The device according to Clause 4, wherein the one or more processors are further configured to decode a value representing a reduction amount to be applied to the point cloud geometry data, wherein, in order to reduce the point cloud geometry data, the one or more processors are configured to reduce the point cloud geometry data according to the value representing the reduction amount.
[0154] Clause 9: The device according to Clause 8, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to amplify the reduced point cloud attribute data according to the value representing the amount of reduction to be applied to the point cloud geometry data.
[0155] Clause 10: The device according to Clause 8, wherein the value includes a first value, and wherein the one or more processors are further configured to decode a second value representing an amplification amount to be applied to the reduced point cloud attribute data, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to amplify the reduced point cloud attribute data according to the second value representing the amplification amount to be applied to the point cloud attribute data.
[0156] Clause 11: The device according to Clause 1, wherein the attribute data includes color data in red-green-blue (RGB) format or in one of luminance, blue hue chromaticity and red hue chromaticity (YCbCr) formats.
[0157] Clause 12: The apparatus according to Clause 1, wherein, in order to reduce the point cloud geometry data, the one or more processors are configured to: redefine the node as an occupied leaf node in the reduced octree for each node of an octree comprising eight leaf nodes, wherein at least one of the eight leaf nodes is point-occupied; and redefine the node as an unoccupied leaf node in the reduced octree for each node of the octree comprising eight leaf nodes, wherein none of the eight leaf nodes is point-occupied.
[0158] Clause 13: The apparatus according to Clause 1, wherein, in order to reduce the point cloud geometry data, the one or more processors are configured to: redefine the node as an occupied leaf node in the reduced octree for each node of the octree comprising eight leaf nodes, wherein the number of occupied leaf nodes among the eight leaf nodes is greater than a threshold; and redefine the node as an unoccupied leaf node in the reduced octree for each node of the octree comprising eight leaf nodes, wherein the number of occupied leaf nodes among the eight leaf nodes is less than or equal to the threshold.
[0159] Clause 14: A method for decoding point cloud data, the method comprising: decoding coded point cloud geometry data of a point cloud to reconstruct point cloud geometry data of the point cloud; reducing the point cloud geometry data to form reduced point cloud geometry data; and using the reduced point cloud geometry to decode attribute data of the point cloud.
[0160] Clause 15: The method according to Clause 14, wherein decoding the attribute data of the point cloud includes encoding the attribute data of the point cloud, the method further includes encoding the point cloud geometry data using a deep learning-based geometric encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data.
[0161] Clause 16: The method according to Clause 15 further includes encoding a value representing a reduction amount to be applied to the point cloud geometry, wherein reducing the point cloud geometry includes reducing the point cloud geometry according to the value representing the reduction amount.
[0162] Clause 17: The method according to Clause 14, wherein decoding the attribute data of the point cloud includes decoding the attribute data of the point cloud to form reduced point cloud attribute data, the method further includes enlarging the reduced point cloud attribute data.
[0163] Clause 18: The method according to Clause 17, wherein amplifying the reduced point cloud attribute data includes using a deep learning-based attribute upsampler to amplify the reduced point cloud attribute data.
[0164] Clause 19: The method according to Clause 17, wherein scaling up the reduced point cloud attribute data includes applying a 5x5x5 convolutional target coordinate layer to the reduced point cloud attribute data.
[0165] Clause 20: The method according to Clause 17, wherein amplifying the reduced point cloud attribute data includes applying one or more of a transposed convolutional layer, a deconvolutional layer, or an unpooling layer to the reduced point cloud attribute data.
[0166] Clause 21: The method according to Clause 17, wherein amplifying the reduced point cloud attribute data includes reconstructing the amplified point cloud attribute data of the point cloud, the method further comprising applying the amplified point cloud attribute data to the point cloud geometry to reconstruct the point cloud.
[0167] Clause 22: The method according to Clause 17 further includes decoding a value representing a reduction amount to be applied to the point cloud geometry, wherein reducing the point cloud geometry includes reducing the point cloud geometry based on the value representing the reduction amount.
[0168] Clause 23: The method according to Clause 22, wherein amplifying the reduced point cloud attribute data includes amplifying the reduced point cloud attribute data according to the value representing the amount of reduction to be applied to the point cloud geometry.
[0169] Clause 24: The method according to Clause 22, wherein the value includes a first value, and the method further includes decoding a second value representing an amplification amount to be applied to the reduced point cloud attribute data, wherein amplifying the reduced point cloud attribute data includes amplifying the reduced point cloud attribute data according to the second value representing the amplification amount to be applied to the point cloud attribute data.
[0170] Clause 25: The method according to Clause 14, wherein the point cloud geometry data represents points in the three-dimensional space of the point cloud, and wherein the attribute data represents attributes of the points.
[0171] Clause 26: The method described in Clause 14, wherein the attribute data includes color data in red-green-blue (RGB) format or in one of luminance, blue hue chromaticity, and red hue chromaticity (YCbCr) formats.
[0172] Clause 27: The method according to Clause 14, wherein reducing the point cloud geometry data comprises: redefining each node of an octree comprising eight leaf nodes, wherein at least one of the eight leaf nodes is point-occupied, as an occupied leaf node in the reduced octree; and redefining each node of the octree comprising eight leaf nodes, wherein none of the eight leaf nodes is point-occupied, as an unoccupied leaf node in the reduced octree.
[0173] Clause 28: The method according to Clause 14, wherein reducing the point cloud geometry data comprises: redefining each node of an octree comprising eight leaf nodes, wherein the number of occupied leaf nodes among the eight leaf nodes is greater than a threshold, as an occupied leaf node in the reduced octree; and redefining each node of the octree comprising eight leaf nodes, wherein the number of occupied leaf nodes among the eight leaf nodes is less than or equal to the threshold, as an unoccupied leaf node in the reduced octree.
[0174] Clause 29: A computer-readable storage medium storing instructions that, when executed, cause a processor to: decode encoded point cloud geometry data of a point cloud to reconstruct point cloud geometry data of the point cloud; reduce the point cloud geometry data to form reduced point cloud geometry data; and decode attribute data of the point cloud using the reduced point cloud geometry.
[0175] Clause 30: An apparatus for decoding point cloud data, the apparatus comprising: components for decoding encoded point cloud geometry data of a point cloud to reconstruct point cloud geometry data of the point cloud; components for reducing the point cloud geometry data to form reduced point cloud geometry data; and components for decoding attribute data of the point cloud using the reduced point cloud geometry.
[0176] Clause 31: An apparatus for decoding point cloud data, the apparatus comprising: a memory configured to store point cloud data; and one or more processors implemented in a circuit and configured to: decode encoded point cloud geometry data of the point cloud to reconstruct point cloud geometry data of the point cloud; reduce the point cloud geometry data to form reduced point cloud geometry data; and decode attribute data of the point cloud using the reduced point cloud geometry.
[0177] Clause 32: The apparatus according to Clause 31, wherein, in order to decode the attribute data of the point cloud, the one or more processors are configured to encode the attribute data of the point cloud, and wherein the one or more processors are further configured to encode the point cloud geometry data using a deep learning-based geometry encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data.
[0178] Clause 33: The apparatus of Clause 32, wherein the one or more processors are further configured to encode a value representing a reduction amount to be applied to the point cloud geometry, wherein, in order to reduce the point cloud geometry, the one or more processors are configured to reduce the point cloud geometry according to the value representing the reduction amount.
[0179] Clause 34: The apparatus according to Clause 31, wherein, in order to decode the attribute data of the point cloud, the one or more processors are configured to decode the attribute data of the point cloud to form reduced point cloud attribute data, and wherein the one or more processors are further configured to amplify the reduced point cloud attribute data.
[0180] Clause 35: The device according to Clause 34, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to use a deep learning-based attribute upsampler to amplify the reduced point cloud attribute data.
[0181] Clause 36: The device according to any one of Clauses 34 and 35, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to apply a 5x5x5 convolutional target coordinate layer to the reduced point cloud attribute data.
[0182] Clause 37: The device according to any one of Clauses 34 to 36, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to reconstruct amplified point cloud attribute data of the point cloud, and wherein the one or more processors are further configured to apply the amplified point cloud attribute data to the point cloud geometry to reconstruct the point cloud.
[0183] Clause 38: The device according to any one of Clauses 34 to 37, wherein the one or more processors are further configured to decode a value representing a reduction amount to be applied to the point cloud geometry data, wherein, in order to reduce the point cloud geometry data, the one or more processors are configured to reduce the point cloud geometry data according to the value representing the reduction amount.
[0184] Clause 39: The device according to Clause 38, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to amplify the reduced point cloud attribute data according to the value representing the amount of reduction to be applied to the point cloud geometry data.
[0185] Clause 40: The device according to Clause 38, wherein the value includes a first value, and wherein the one or more processors are further configured to decode a second value representing an amplification amount to be applied to the reduced point cloud attribute data, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to amplify the reduced point cloud attribute data according to the second value representing the amplification amount to be applied to the point cloud attribute data.
[0186] Clause 41: The device pursuant to any one of Clauses 31 to 40, wherein the attribute data comprises color data in red-green-blue (RGB) format or in one of luminance, blue hue chromaticity and red hue chromaticity (YCbCr) formats.
[0187] Clause 42: The device according to any one of Clauses 31 to 41, wherein, in order to reduce the point cloud geometry data, the one or more processors are configured to: redefine the node as an occupied leaf node in the reduced octree for each node of an octree comprising eight leaf nodes, wherein at least one of the eight leaf nodes is point-occupied; and redefine the node as an unoccupied leaf node in the reduced octree for each node of the octree comprising eight leaf nodes, wherein none of the eight leaf nodes is point-occupied.
[0188] Clause 43: The device according to any one of Clauses 31 to 41, wherein, in order to reduce the point cloud geometry data, the one or more processors are configured to: redefine the node as an occupied leaf node in the reduced octree for each node of the octree comprising eight leaf nodes, wherein the number of occupied leaf nodes among the eight leaf nodes is greater than a threshold; and redefine the node as an unoccupied leaf node in the reduced octree for each node of the octree comprising eight leaf nodes, wherein the number of occupied leaf nodes among the eight leaf nodes is less than or equal to the threshold.
[0189] Clause 44: A method for decoding point cloud data, the method comprising: decoding coded point cloud geometry data of a point cloud to reconstruct point cloud geometry data of the point cloud; reducing the point cloud geometry data to form reduced point cloud geometry data; and using the reduced point cloud geometry to decode attribute data of the point cloud.
[0190] Clause 45: The method according to Clause 44, wherein decoding the attribute data of the point cloud includes encoding the attribute data of the point cloud, the method further includes encoding the point cloud geometry data using a deep learning-based geometry encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data.
[0191] Clause 46: The method according to Clause 45 further includes encoding a value representing a reduction amount to be applied to the point cloud geometry, wherein reducing the point cloud geometry includes reducing the point cloud geometry according to the value representing the reduction amount.
[0192] Clause 47: The method according to Clause 44, wherein decoding the attribute data of the point cloud includes decoding the attribute data of the point cloud to form reduced point cloud attribute data, the method further includes enlarging the reduced point cloud attribute data.
[0193] Clause 48: The method according to Clause 47, wherein amplifying the reduced point cloud attribute data includes using a deep learning-based attribute upsampler to amplify the reduced point cloud attribute data.
[0194] Clause 49: The method according to any one of Clauses 47 and 48, wherein scaling up the reduced point cloud attribute data comprises applying a convolutional target coordinate 5x5x5 layer to the reduced point cloud attribute data.
[0195] Clause 50: The method according to any one of Clauses 47 to 49, wherein amplifying the reduced point cloud attribute data comprises applying one or more of a transposed convolutional layer, a deconvolutional layer, or an unpooling layer to the reduced point cloud attribute data.
[0196] Clause 51: The method according to any one of Clauses 47 to 50, wherein amplifying the reduced point cloud attribute data includes reconstructing the amplified point cloud attribute data of the point cloud, and the method further includes applying the amplified point cloud attribute data to the point cloud geometry to reconstruct the point cloud.
[0197] Clause 52: The method according to any one of Clauses 47 to 51, the method further comprising decoding a value representing a reduction amount to be applied to the point cloud geometry, wherein reducing the point cloud geometry includes reducing the point cloud geometry based on the value representing the reduction amount.
[0198] Clause 53: The method according to Clause 52, wherein amplifying the reduced point cloud attribute data includes amplifying the reduced point cloud attribute data according to the value representing the amount of reduction to be applied to the point cloud geometry.
[0199] Clause 54: The method according to Clause 52, wherein the value includes a first value, the method further comprising decoding a second value representing an amplification amount to be applied to the reduced point cloud attribute data, wherein amplifying the reduced point cloud attribute data includes amplifying the reduced point cloud attribute data according to the second value representing the amplification amount to be applied to the point cloud attribute data.
[0200] Clause 55: The method according to any one of Clauses 44 to 54, wherein the point cloud geometric data represents points in a three-dimensional space of the point cloud, and wherein the attribute data represents attributes of the points.
[0201] Clause 56: The method according to any one of Clauses 44 to 55, wherein the attribute data includes color data in red-green-blue (RGB) format or in one of luminance, blue hue chromaticity and red hue chromaticity (YCbCr) formats.
[0202] Clause 57: The method according to any one of Clauses 44 to 56, wherein reducing the point cloud geometry data comprises: redefining each node of an octree comprising eight leaf nodes, wherein at least one of the eight leaf nodes is point-occupied, the node as an occupied leaf node in the reduced octree; and redefining each node of the octree comprising eight leaf nodes, wherein none of the eight leaf nodes is point-occupied, the node as an unoccupied leaf node in the reduced octree.
[0203] Clause 58: The method according to any one of Clauses 44 to 56, wherein reducing the point cloud geometry data comprises: redefining the node as an occupied leaf node in the reduced octree for each node of the octree comprising eight leaf nodes, wherein the number of occupied leaf nodes among the eight leaf nodes is greater than a threshold; and redefining the node as an unoccupied leaf node in the reduced octree for each node of the octree comprising eight leaf nodes, wherein the number of occupied leaf nodes among the eight leaf nodes is less than or equal to the threshold.
[0204] Clause 59: A computer-readable storage medium storing instructions that, when executed, cause a processor to: decode encoded point cloud geometry data of a point cloud to reconstruct point cloud geometry data of the point cloud; reduce the point cloud geometry data to form reduced point cloud geometry data; and decode attribute data of the point cloud using the reduced point cloud geometry.
[0205] Clause 60: An apparatus for decoding point cloud data, the apparatus comprising: components for decoding encoded point cloud geometry data of a point cloud to reconstruct point cloud geometry data of the point cloud; components for reducing the point cloud geometry data to form reduced point cloud geometry data; and components for decoding attribute data of the point cloud using the reduced point cloud geometry.
[0206] It should be recognized that, based on the examples, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all actions or events described are necessary for implementing the technique). Furthermore, in some examples, actions or events may be performed concurrently (e.g., through multithreading, interrupt handling, or multiple processors) rather than sequentially.
[0207] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium (which corresponds to a tangible medium such as a data storage medium) or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. Thus, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.
[0208] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Additionally, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead refer to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs utilize lasers to optically reproduce data. The combinations described above should also be included within the scope of computer-readable media.
[0209] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuit" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.
[0210] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but implementation by different hardware units is not necessarily required. Specifically, as described above, the various units may be combined in a codec hardware unit, or may be provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.
[0211] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. An apparatus for decoding point cloud data, the apparatus comprising: A memory configured to store point cloud data; and One or more processors, said one or more processors being implemented in a circuit and configured to: Decode the encoded point cloud geometric data to reconstruct the point cloud geometric data; Reduce the point cloud geometric data to form reduced point cloud geometric data; as well as The attribute data of the point cloud is decoded using the reduced point cloud geometry data.
2. The apparatus of claim 1, wherein, in order to decode the attribute data of the point cloud, the one or more processors are configured to encode the attribute data of the point cloud, and wherein the one or more processors are further configured to encode the point cloud geometric data using a deep learning-based geometric encoder to form the encoded point cloud geometric data before decoding the encoded point cloud geometric data.
3. The apparatus of claim 2, wherein the one or more processors are further configured to encode a value representing a reduction amount to be applied to the point cloud geometry data, wherein, in order to reduce the point cloud geometry data, the one or more processors are configured to reduce the point cloud geometry data according to the value representing the reduction amount.
4. The apparatus of claim 1, wherein, in order to decode the attribute data of the point cloud, the one or more processors are configured to decode the attribute data of the point cloud to form reduced point cloud attribute data, and wherein the one or more processors are further configured to amplify the reduced point cloud attribute data.
5. The device of claim 4, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to use a deep learning-based attribute upsampler to amplify the reduced point cloud attribute data.
6. The device of claim 4, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to apply a 5x5x5 convolutional target coordinate layer to the reduced point cloud attribute data.
7. The device of claim 4, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to reconstruct amplified point cloud attribute data of the point cloud, and wherein the one or more processors are further configured to apply the amplified point cloud attribute data to the point cloud geometry to reconstruct the point cloud.
8. The apparatus of claim 4, wherein the one or more processors are further configured to decode a value representing a reduction amount to be applied to the point cloud geometry data, wherein, in order to reduce the point cloud geometry data, the one or more processors are configured to reduce the point cloud geometry data according to the value representing the reduction amount.
9. The apparatus of claim 8, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to amplify the reduced point cloud attribute data according to a value representing the amount of reduction to be applied to the point cloud geometry data.
10. The device of claim 8, wherein the value includes a first value, and wherein the one or more processors are further configured to decode a second value representing an amplification amount to be applied to the reduced point cloud attribute data, wherein, in order to amplify the reduced point cloud attribute data, the one or more processors are configured to amplify the reduced point cloud attribute data according to the second value representing the amplification amount to be applied to the point cloud attribute data.
11. The device of claim 1, wherein, in order to reduce the point cloud geometric data, the one or more processors are configured to: For each node in an octree comprising eight leaf nodes, wherein at least one of the eight leaf nodes is occupied, the node is redefined as an occupied leaf node in a reduced octree; and For each node of the octree that includes eight leaf nodes, wherein none of the eight leaf nodes is occupied, the node is redefined as an unoccupied leaf node in the reduced octree.
12. A method for decoding point cloud data, the method comprising: Decode the encoded point cloud geometric data to reconstruct the point cloud geometric data; Reduce the point cloud geometric data to form reduced point cloud geometric data; as well as The attribute data of the point cloud is decoded using the reduced point cloud geometry.
13. The method of claim 12, wherein decoding the attribute data of the point cloud includes encoding the attribute data of the point cloud, the method further comprising: Before decoding the encoded point cloud geometric data, the point cloud geometric data is encoded using a deep learning-based geometric encoder to form the encoded point cloud geometric data. as well as Values representing the amount of reduction to be applied to the point cloud geometry are encoded, wherein reducing the point cloud geometry includes reducing the point cloud geometry according to the values representing the amount of reduction.
14. The method of claim 12, wherein decoding the attribute data of the point cloud comprises decoding the attribute data of the point cloud to form reduced point cloud attribute data, the method further comprising: Amplifying the reduced point cloud attribute data, wherein amplifying the reduced point cloud attribute data includes using a deep learning-based attribute upsampler to amplify the reduced point cloud attribute data.
15. The method of claim 14, wherein amplifying the reduced point cloud attribute data comprises applying at least one of a 5x5x5 convolutional target coordinate layer, a transposed convolutional layer, a deconvolutional layer, or an unpooling layer to the reduced point cloud attribute data.
16. The method of claim 14, wherein amplifying the reduced point cloud attribute data comprises reconstructing amplified point cloud attribute data of the point cloud, and the method further comprises applying the amplified point cloud attribute data to the point cloud geometry to reconstruct the point cloud.
17. The method of claim 14, further comprising decoding a value representing a reduction amount to be applied to the point cloud geometry, wherein reducing the point cloud geometry includes reducing the point cloud geometry based on the value representing the reduction amount.
18. The method of claim 17, wherein amplifying the reduced point cloud attribute data comprises amplifying the reduced point cloud attribute data according to a value representing the amount of reduction to be applied to the point cloud geometry.
19. The method of claim 17, wherein the value includes a first value, and the method further includes decoding a second value representing an amplification amount to be applied to the reduced point cloud attribute data, wherein amplifying the reduced point cloud attribute data includes amplifying the reduced point cloud attribute data according to the second value representing the amplification amount to be applied to the point cloud attribute data.
20. The method of claim 12, wherein reducing the point cloud geometric data comprises: For each node of an octree comprising eight leaf nodes, wherein at least one of the eight leaf nodes is occupied, the node is redefined as an occupied leaf node in a reduced octree. as well as For each node of the octree that includes eight leaf nodes, wherein none of the eight leaf nodes is occupied, the node is redefined as an unoccupied leaf node in the reduced octree.