Attributes coding for point cloud compression

BR112025020208A2Pending Publication Date: 2026-08-11
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025020208
Authority / Receiving Office
BR · BR
Patent Type
Applications
Publication Date
2026-08-11

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

1 / 67 ATTRIBUTE CODING FOR POINT CLOUD COMPRESSION

[0001] This application claims priority to US patent application no. US Patent Application No. 18 / 624,683, filed April 2, 2024, and US Provisional Patent Application No. 63 / 493,806, filed April 3, 2023, each of which has its entire content incorporated into the present invention by reference. US Patent Application No. 18 / 624,683, filed April 2, 2024, claims the benefit of US Provisional Patent Application No. 63 / 493,806, filed April 3, 2023. TECHNICAL FIELD

[0002] This disclosure relates to point cloud encoding and decoding. BACKGROUND

[0003] A point cloud is a collection of points in three-dimensional space. The points can correspond to points on objects in three-dimensional space. In this way, a point cloud can be used to represent the physical content of three-dimensional space. Point clouds can be useful in a wide variety of situations. For example, point clouds can be used in the context of autonomous vehicles to represent the positions of objects on a highway. In another example, point clouds can be used in the context of representing the physical content of an environment for the purposes of positioning virtual objects in an augmented reality (AR) or mixed reality (MR) application. Point cloud compression is a process for encoding and decoding point clouds. Encoding point clouds can reduce the amount of data required for storing and transmitting point clouds. SUMMARY

[0004] In general, this disclosure describes techniques for point cloud coding (e.g., encoding or decoding), including coding geometry and attribute data from point cloud data. Encoding can Petition 870250085531, dated 09 / 22 / 2025, pp. 261 / 359 2 / 67 include one or both of encoding and / or decoding. In particular, geometry information (e.g., point coordinates within the point cloud) can be efficiently encoded using a deep learning-based encoder. Attribute data for the points (e.g., color, reflectance, brightness, surface normals, or the like) generally includes a larger amount of data than geometry data. Therefore, the point cloud encoder can reconstruct the geometry data, and then downscale the geometry data and also downscale the attribute data before encoding the attribute data. A point cloud decoder can then decode and reconstruct the geometry data at full scale, downscale the geometry data, and decode the attribute data using the downscaled geometry data.Subsequently, the point cloud decoder can upscale the attribute data and apply the upscaled attribute data to the full-scale geometry data to reconstruct the point cloud.

[0005] In one example, a device for encoding (e.g., reconstructing) point cloud data includes a memory configured to store point cloud data; and one or more processors implemented in the circuit set and configured to: decode encoded point cloud geometry data to a point cloud in order to reconstruct point cloud geometry data to the point cloud; downscale point cloud geometry data to form downscaled point cloud geometry data; and encode (e.g., reconstruct) attribute data to the point cloud using the downscaled point cloud geometry.

[0006] In another example, a method for encoding (e.g., reconstructing) point cloud data, wherein the method comprises: decoding encoded point cloud geometry data into a point cloud in order to reconstruct point cloud geometry data into the point cloud; downscaling the point cloud geometry data Petition 870250085531, dated 09 / 22 / 2025, pp. 262 / 359 3 / 67 to form scaled-down point cloud geometry data; and to encode (e.g., reconstruct) attribute data for the point cloud using the scaled-down point cloud geometry.

[0007] In another example, a computer-readable storage medium has instructions stored within it that, when executed, cause a processor to: decode encoded point cloud geometry data into a point cloud in order to reconstruct point cloud geometry data into the point cloud; downscale point cloud geometry data to form downscaled point cloud geometry data; and encode (e.g., reconstruct) attribute data into the point cloud using the downscaled point cloud geometry.

[0008] In another example, a device for encoding point cloud data, wherein the device comprises: means for decoding encoded point cloud geometry data to a point cloud in order to reconstruct point cloud geometry data for the point cloud; means for downscaling point cloud geometry data to form downscaled point cloud geometry data; and means for encoding (e.g., reconstructing) attribute data to the point cloud using the downscaled point cloud geometry.

[0009] The details of one or more examples are set forth in the attached drawings and the description below. Other attributes, objectives and advantages will become apparent from the description, drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a block diagram illustrating an example point cloud encoding and decoding system, which can perform the techniques of this disclosure.

[0011] Figure 2 is a block diagram illustrating an example point cloud encoder according to techniques in this disclosure.

[0012] Figure 3 is a block diagram illustrating a decoder. Petition 870250085531, dated 09 / 22 / 2025, pp. 263 / 359 4 / 67 Example point cloud according to the techniques of this disclosure.

[0013] Figure 4 is a conceptual diagram illustrating an example encoding structure according to certain examples of the techniques in this disclosure.

[0014] Figure 5 is a conceptual diagram illustrating an example decoding structure according to certain examples of the techniques in this disclosure.

[0015] Figure 6 is a block diagram illustrating an example point cloud coding structure according to the techniques of this disclosure.

[0016] Figure 7 is a block diagram illustrating an example point cloud decoding structure according to the techniques of this disclosure.

[0017] Figures 8A and 8B are conceptual diagrams illustrating examples of voxel downscaling of a point cloud.

[0018] Figure 9 is a block diagram illustrating a set of example stages that can be included in a deep learning-based feature supersampler.

[0019] Figure 10 is a block diagram illustrating another set of example stages that can be included in a deep learning-based feature supersampler.

[0020] Figure 11 is a flowchart illustrating a method for encoding example point cloud data according to the techniques of this disclosure.

[0021] Figure 12 is a flowchart illustrating a method for decoding example point cloud data according to the techniques of this disclosure.

[0022] Figure 13 is a conceptual diagram illustrating a laser package, such as a LIDAR sensor or other system that includes one or more lasers, scanning points in three-dimensional space.

[0023] Figure 14 is a conceptual diagram illustrating an example 900 range measurement system that can be used with one or more techniques of this disclosure. Petition 870250085531, dated 09 / 22 / 2025, pp. 264 / 359 5 / 67

[0024] Figure 15 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques from this disclosure can be used.

[0025] Figure 16 is a conceptual diagram illustrating an example extended reality system in which one or more techniques from this disclosure can be used.

[0026] Figure 17 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure can be used.

[0027] Figures 18 and 19 are flow diagrams illustrating example deep learning-based geometry encoder and decoder networks. DETAILED DESCRIPTION

[0028] A point cloud (PC) is a three-dimensional data representation for tasks such as virtual reality (VR) and mixed reality (MR), autonomous driving, cultural heritage, etc. Point clouds are a collection of points in three-dimensional space, represented by their three-dimensional coordinates (x, y, z) called the geometry. Each point can also be associated with multiple attributes, such as color, normal vectors, and reflectance. Depending on the target application and point cloud acquisition methods, the point cloud can be categorized into point cloud scenes and point cloud objects. Point cloud scenes can be captured using LiDAR sensors and can be dynamically acquired.

[0029] Point cloud objects can be subdivided into static point clouds and dynamic point clouds. A static point cloud is a single object.A dynamic point cloud is a time-varying point cloud, including a sequence of point cloud instances. Each instance of a dynamic point cloud is a static point cloud. Time-varying dynamic point clouds can be used in AR / VR, volumetric video streaming, and telepresence, and can be generated using three-dimensional models, i.e., CGI, or captured from real-world scenarios using various methods, such as multiple cameras with sensors. Petition 870250085531, dated 09 / 22 / 2025, pp. 265 / 359 6 / 67 depth surrounding the object. These point clouds are dense, photorealistic point clouds that can contain a huge number of points, especially in high-precision or large-scale captures (millions of points per frame at up to 60 frames per second (FPS)). Therefore, efficient point cloud compression (PCC) is useful to enable practical use in VR and MR applications.

[0030] The Moving Picture Experts Group (MPEG - Moving The Picture Experts Group (PICTURE) approved two PCC (point cloud compression) standards: (1) S. Schwarz, M. Preda, V. Baroncini, M. Budagavi, P. Cesar, PA Chou, RA Cohen, M. Krivoku'ca, S. Lasserre, Z. Li et al., Emerging MPEG standards for point cloud compression, IEEE Journal on Emerging and Selected Topics in Circuits and Systems, volume 9, no. 1, pages 133 to 148, 2018, and (2) D. Graziosi, O. Nakagami, S. Kuma, A. Zaghetto, T. Suzuki, and A. Tabatabai, An overview of ongoing point cloud compression standardization activities: Video-based (v-pcc) and geometry-based (g-pcc), APSIPA Transactions on Signal and Information Processing, volume 9, 2020. MPEG approved the geometry-based point cloud compression standard (G-PCC - geometry-based point cloud compression): MPEG-PCC-TMC13: Geometry Based Point Cloud Compression GPCC, 2021, available at github.com / MPEGGroup / mpeg-pcc-tmc13.MPEG has approved video-based point cloud compression (V-PCC): MPEG-PCC-TMC2: Video Based Point Cloud Compression VPCC, 2022, available at github.com / MPEGGroup / mpeg-pcc-tmc2.

[0031] G-PCC includes octree geometry coding as a generic geometry coding tool and a predictive (tree-based) geometry coding tool that targets LiDAR-based point clouds. G-PCC is still developing a method based on triangle meshes or triangle soup (trisoup) to approximate the surface of the three-dimensional model. V-PCC, on the other hand, encodes dynamic point clouds by projecting three-dimensional points onto a two-dimensional plane and then... Petition 870250085531, dated 09 / 22 / 2025, pp. 266 / 359 7 / 67 uses video codecs, for example, high-efficiency video coding (HEVC), to encode each frame over time. MPEG also proposed common test conditions (CTCs) to evaluate test models: S. Schwarz, G. Martin-Cocher, D. Flynn and M. Budagavi, Common test conditions for point cloud compression, ISO / IECJTC1 / SC29 / WG11 w17766 document, Ljubljana, Slovenia, 2018.

[0032] As mentioned above, efficient point cloud compression is useful for applications such as virtual and mixed reality, autonomous driving, and cultural heritage. Some techniques, such as m59617: Anique Akhtar, Zhu Li, Geert Van der Auwera, Adarsh ​​Krishnan Ramasubramonian, Luong Pham Van, Marta Karczewicz, Dynamic Point Cloud Geometry Compression using Sparse Convolutions, MPEG-137 Online, document m59617, April 2022, and m60307: Anique Akhtar, Zhu Li, Geert Van der Auwera, Adarsh ​​Krishnan Ramasubramonian, Marta Karczewicz, [AI-3DGC][EE5.3 Test 2] Results dynamic point cloud compression, MPEG-139 Online, document M60307, ​​July 2022, use deep learning-based point cloud compression for dense dynamic point clouds using a deep learning network consisting of an encoder and decoder module.

[0033] Within the context of this disclosure, deep learning may refer to the use of multiple hidden layers in an artificial neural network. A deep learning-based geometry encoder may refer to a geometry encoder that includes a computer-based neural network that includes multiple hidden layers. A deep learning-based feature supersampler may refer to a feature supersampler that includes a computer-based neural network that includes multiple hidden layers.

[0034] Research on performing point cloud compression using deep learning solutions is ongoing. A point cloud typically includes a collection of points as well as attributes for the points. The points correspond to positions in a three-dimensional space, for example, having Petition 870250085531, dated 09 / 22 / 2025, pp. 267 / 359 8 / 67 X, Y, and Z coordinates. Attributes may include, for example, color, reflectance, brightness, surface normals, or the like. Because point clouds have both geometry and attributes, solutions have been proposed for point cloud geometry compression, point cloud attribute compression, as well as joint compression of point cloud geometry and attributes.

[0035] Deep learning-based solutions typically perform well when applied to geometry compression. However, they still lack point cloud attribute compression. The techniques in this disclosure include using geometry compression from one codec and attribute compression from another codec. For example, a deep learning-based geometry compression scheme can be combined with a non-deep learning-based attribute compression scheme. These techniques can provide flexibility to the compression framework and create a strong baseline for deep learning-based point cloud attribute compression.

[0036] These techniques may additionally include the use of multiscale attribute compression with post-processing based on deep learning. Heuristic tests have shown that most of the compression bits are consumed by attribute encoding. Therefore, creating an effective attribute compression scheme is important to achieve good compression performance. The multiscale attribute compression scheme can improve the overall compression performance of the encoding framework.

[0037] Some coding techniques include a lossy point cloud geometry compression scheme based on deep learning for dynamic point cloud compression. The lossy geometry scheme predicts the latent representation of a current frame using a previous frame by employing a prediction network. The framework performs inter-frame point cloud coding of frame P, where the current frame is coded with reference to the previously decoded frame. The architecture is implemented using a sparse convolutional neural network (CNN). Petition 870250085531, dated 09 / 22 / 2025, pp. 268 / 359 9 / 67 with sparse tensors. The architecture employs convolution on target coordinates to map the latent representation of the previous frame to the subsampled coordinates of the current frame to predict the embedding of current frame attributes. The framework transmits the residual of the predicted attributes and the actual attributes by compressing them using a learned probabilistic factored entropy model. Compared to G-PCC and V-PCC, these techniques demonstrate better geometry compression performance in dense point clouds with efficient encoding / decoding runtime.

[0038] In one or more examples, this disclosure describes a flexible configuration of the deep learning-based framework where, instead of having a joint geometry and attribute compression scheme, the example techniques use a recoloring scheme (as an example) to generate attributes for the reconstructed point cloud and employ traditional attribute compression schemes. One or more of the example techniques described in this disclosure may utilize the G-PCC recoloring scheme to obtain attributes of the target point cloud from the source point cloud. The G-PCC recoloring scheme employs nearest neighbors based on weighted distance in the source point cloud to calculate attributes of the target point cloud. Consequently, in one or more examples, the example techniques may allow the use of attribute compression from a different codec, and geometry compression from a different codec.

[0039] Figure 1 is a block diagram illustrating an example encoding and decoding system 100 that can perform the techniques of this disclosure. The techniques of this disclosure relate generally to the process of encoding (e.g., encoding and / or decoding) point cloud data, i.e., to support point cloud compression. In general, point cloud data includes any data for processing a point cloud. The encoding process can be effective in compressing and / or decompressing point cloud data. Petition 870250085531, dated 09 / 22 / 2025, pages 269 / 359 10 / 67

[0040] As shown in Figure 1, the system 100 includes a source device 102 and a destination device 116. The source device 102 provides encoded point cloud data to be decoded by a destination device 116. In particular, in the example in Figure 1, the source device 102 provides the point cloud data to the destination device 116 via a computer-readable medium 110. The source device 102 and the destination device 116 may comprise any of a wide range of devices, including desktop computers, notebook computers (i.e., laptop computers), tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, land or sea vehicles, spacecraft, aircraft, robots, LIDAR devices, satellites, or similar devices.In some cases, source device 102 and destination device 116 can be equipped for wireless communication.

[0041] In the example in Figure 1, source device 102 includes a data source 104, a memory 106, a point cloud encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a point cloud decoder 300, a memory 120, and a data consumer 118. According to this disclosure, the point cloud encoder 200 of source device 102 and the point cloud decoder 300 of destination device 116 can be configured to apply the techniques of this disclosure related to attribute encoding for point cloud compression. In this way, source device 102 represents an example of an encoding device, while destination device 116 represents an example of a decoding device.In other examples, source device 102 and destination device 116 may include other components or arrangements. For example, source device 102 may receive data (e.g., point cloud data) from an internal or external source. Similarly, destination device 116 may interface with a consumer. Petition 870250085531, dated 09 / 22 / 2025, pp. 270 / 359 11 / 67 of external data, instead of including a data consumer on the same device.

[0042] As shown in Figure 1, system 100 is merely an example. In general, other digital encoding and / or decoding devices can perform the techniques of this disclosure related to attribute encoding for point cloud compression. Source device 102 and destination device 116 are merely examples of devices in which source device 102 generates encoded data for transmission to destination device 116. This disclosure refers to an encoding device as a device that performs the process of encoding (encoding and / or decoding) data. Thus, point cloud encoder 200 and point cloud decoder 300 represent examples of encoding devices, in particular, an encoder and a decoder, respectively.In some examples, the source device 102 and the destination device 116 may operate in a substantially symmetrical manner, such that each of the source device 102 and the destination device 116 includes encoding and decoding components. In this way, the system 100 can support unidirectional or bidirectional transmission between the source device 102 and the destination device 116, for example, for streaming, playback, broadcast, telephony, navigation, and other applications.

[0043] In general, data source 104 represents a data source (i.e., raw, unencoded point cloud data) and can provide a sequential series of data frames to the point cloud encoder 200, which encodes the data into frames. Data source 104 of source device 102 may include a point cloud capture device, such as any of a variety of cameras or sensors, for example, a three-dimensional scanning device or a light detection and ranging (LIDAR) device, one or more video cameras, a file containing previously captured data, and / or a data feed interface for receiving data from a data content provider. Alternatively or additionally, the point cloud data may be computer-generated from a Petition 870250085531, dated 09 / 22 / 2025, pp. 271 / 359 12 / 67 scanner, a camera, a sensor, or other type of data. For example, the data source 104 may generate computer-based graphic data, such as the source data, or produce a combination of live data, archived data, and computer-generated data. In each case, the point cloud encoder 200 encodes the captured, pre-captured, or computer-generated data. The point cloud encoder 200 may rearrange the frames from the received order (sometimes called the display order) into an encoding order for the encoding process. The point cloud encoder 200 may generate one or more bitstreams including encoded data. The source device 102 may then output the encoded data via the output interface 108 to the computer-readable medium 110 for reception and / or retrieval, for example, by the input interface 122 of the destination device 116.

[0044] Memory 106 of source device 102 and memory 120 of destination device 116 may represent general-purpose memories. In some examples, memory 106 and memory 120 may store raw data, for example, raw data from data source 104 and raw and decoded data from point cloud decoder 300. Additionally or alternatively, memory 106 and memory 120 may store executable software instructions, for example, by point cloud encoder 200 and point cloud decoder 300, respectively. Although memory 106 and memory 120 are shown separately from point cloud encoder 200 and point cloud decoder 300 in this example, it should be understood that point cloud encoder 200 and point cloud decoder 300 may also include internal memories for functionally similar or equivalent purposes.Furthermore, memory 106 and memory 120 can store encoded data, for example, the output of point cloud encoder 200 and the input to point cloud decoder 300. In some examples, portions of memory 106 and memory 120 can be allocated as one or more buffers, for example, to store raw, decoded and / or encoded data. Petition 870250085531, dated 09 / 22 / 2025, pp. 272 / 359 13 / 67 For example, memory 106 and memory 120 can store data that represents a point cloud.

[0045] The computer-readable medium 110 may represent any type of medium or device capable of carrying encoded data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium to enable the source device 102 to transmit encoded data directly to the destination device 116 in real time, for example, via a radio frequency network or computer-based network. The output interface 108 may modulate a transmission signal that includes the encoded data, and the input interface 122 may demodulate the received transmission signal according to a communication standard, such as a wireless communication protocol. The communication medium may comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines.The communication medium can be part of a packet-based network, such as a local area network, a wide area network, or a global network, such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication between the source device 102 and the destination device 116.

[0046] In some examples, source device 102 may output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of several distributed or locally accessed data storage media, such as a hard disk, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other digital storage media suitable for storing encoded data.

[0047] In some examples, source device 102 can output data Petition 870250085531, dated 09 / 22 / 2025, pp. 273 / 359 14 / 67 encoded output to file server 114 or other intermediate storage device capable of storing the encoded data generated by source device 102. Destination device 116 can access stored data from file server 114 via streaming or download. File server 114 can be any type of server device capable of storing encoded data and transmitting that encoded data to destination device 116. File server 114 can represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 can access encoded data from file server 114 through any standard data connection, including an internet connection.This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., digital subscriber line (DSL), cable modem, etc.), or a combination of both that is suitable for accessing encrypted data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming protocol, a download protocol, or a combination thereof.

[0048] Output interface 108 and input interface 122 may represent wireless transmitters / receivers, modems, wired network communication components (e.g., Ethernet cards), wireless communication components operating in accordance with any of a variety of Institute of Electrical and Electronics Engineers (IEEE) 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 comprise wireless components, output interface 108 and input interface 122 may be configured to transfer data, such as encoded data, in accordance with a cellular communication standard such as fourth generation (4G). Petition 870250085531, dated 09 / 22 / 2025, pp. 274 / 359 15 / 67 LTE (long-term evolution), LTE-advanced, fifth generation (5G), or similar. In some instances where output interface 108 comprises a wireless transmitter, output interface 108 and input interface 122 may be configured to transfer data, such as encoded data, according to other wireless standards, such as an IEEE 802.11 specification, an IEEE 802.15 specification (e.g., ZigBee™), a Bluetooth™ standard, or similar. In some instances, source device 102 and / or destination device 116 may include their respective system-on-a-chip (SoC) devices.For example, source device 102 may include a SoC device to perform the functionality assigned to media encoder 200 and / or output interface 108, and destination device 116 may include a SoC device to perform the functionality assigned to point cloud decoder 300 and / or input interface 122.

[0049] The techniques of this disclosure can be applied to encode and decode in support of any of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors and processing devices, such as local or remote servers, geographic mapping or other applications.

[0050] The input interface 122 of the destination device 116 receives an encoded bitstream from the computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, or similar). The encoded bitstream may include signaling information defined by the point cloud encoder 200, which is also used by the point cloud decoder 300, as syntax elements having values ​​that describe the characteristics and / or processing of encoded units (e.g., slices, images, image groups, sequences, or similar). The data consumer 118 uses the decoded data. For example, the data consumer 118 may use the decoded data to determine the locations of physical objects. In some examples, the data consumer 118 may understand a display for Petition 870250085531, dated 09 / 22 / 2025, pages 275 / 359 16 / 67 present images based on a point cloud.

[0051] Each of the 200 point cloud encoder and 300 point cloud decoder can be implemented as any one of a variety of suitable encoder and / or decoder circuit sets, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, a device may store instructions for the software in a suitable non-transient, computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure.Each of the 200 point cloud encoder and 300 point cloud decoder can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective device. A device that includes the 200 point cloud encoder and / or the 300 point cloud decoder may comprise one or more integrated circuits, microprocessors, and / or other types of devices.

[0052] The 200 point cloud encoder and the 300 point cloud decoder can operate according to an encoding standard, such as a video point cloud compression standard (V-PCC) or a geometry point cloud compression standard (G-PCC). This disclosure may refer generally to the process of encoding (e.g., encoding and decoding) images to include the process of encoding or decoding data. An encoded bitstream generally includes a series of values ​​for syntax elements representing decisions of the encoding process (e.g., encoding modes).

[0053] This disclosure may generally refer to the signaling of certain Petition 870250085531, dated 09 / 22 / 2025, pp. 276 / 359 17 / 67 information, such as syntax elements. The term signaling can generally refer to the communication of values ​​for syntax elements and / or other data used to decode the encoded data. That is, the point cloud encoder 200 can signal values ​​for syntax elements in the bitstream. In general, signaling refers to the generation of a value in the bitstream. As mentioned above, the source device 102 can transport the bitstream to the destination device 116 substantially in real time or in non-real time, as may occur during the storage of syntax elements in the storage device 112 for later retrieval by the destination device 116.

[0054] ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) is studying the possible need for standardization of point cloud encoding technology with a compression capability that significantly exceeds that of current approaches and aims to create the standard.The group is working together on this exploration activity in a collaborative effort known as the 3DG (3-Dimensional Graphics Team) to evaluate compression technology designs proposed by their experts in this area.

[0055] Point cloud compression activities are categorized into two different approaches. The first approach is video point cloud compression (V-PCC), which segments the three-dimensional object and projects the segments onto multiple two-dimensional planes (which are represented as patches in the two-dimensional frame), which are additionally encoded by a legacy two-dimensional video codec, such as a High Efficiency Video Coding (HEVC) codec (ITU-T H.265). The second approach is geometry-based point cloud compression (G-PCC), which directly compresses the three-dimensional geometry, i.e., the position of a set of points in three-dimensional space, and the associated attribute values ​​(for each point associated with the three-dimensional geometry). G-PCC addresses point cloud compression in category 1 (static point clouds) and category 3 (dynamically acquired point clouds).A recent preliminary version of the G-PCC standard is available in G. Petition 870250085531, dated 09 / 22 / 2025, pp. 277 / 359 18 / 67 PCC DIS, ISO / IEC JTC1 / SC29 / WG11 w19088, Brussels, Belgium, January 2020, and a description of the codec is available in G-PCC Codec Description v6, ISO / IEC JTC1 / SC29 / WG11 w19091, Brussels, Belgium, January 2020.

[0056] A point cloud contains a set of points in three-dimensional space and may have attributes associated with the point. Attributes may be color information, such as R, G, B or Y, Cb, Cr, or reflectance information, or other attributes. Point clouds can be captured by various cameras or sensors, such as LIDAR sensors and three-dimensional scanning devices, and can also be computer-generated. Point cloud data is used in various applications including, but not limited to, construction (modeling), graphics (three-dimensional models for visualization and animation), and the automotive industry (LIDAR sensors used to aid navigation).

[0057] The three-dimensional space occupied by point cloud data can be enclosed by a virtual bounding box. The position of the points in the bounding box can be represented with a certain precision; therefore, the positions of one or more points can be quantified based on precision. At the smallest level, the bounding box is divided into voxels, which are the smallest unit of space represented by a unit cube. A voxel in the bounding box can be associated with zero, one, or more than one point. The bounding box can be divided into multiple cube / cuboid regions, which can be called tiles. Each tile can be encoded into one or more slices. The partitioning of the bounding box into slices and tiles can be based on the number of points in each partition or can be based on other considerations (e.g., a particular region can be encoded as tiles).The slice regions can be further partitioned using splitting decisions similar to those in video codecs.

[0058] Figure 2 is a block diagram illustrating an example 200 point cloud encoder. The modules shown are logical and do not necessarily correspond, one-to-one, to the code implemented in the G-PCC codec reference implementation, i.e., the model software of Petition 870250085531, dated 09 / 22 / 2025, pp. 278 / 359 19 / 67 TMC13 test studied by ISO / IEC MPEG (JTC 1 / SC 29 / WG 11).

[0059] In both the 200 point cloud encoder and the 300 point cloud decoder, point cloud positions are encoded first. Attribute encoding depends on the decoded geometry. Compressed geometry is typically represented as an octree from the root down to a leaf level of individual voxels.

[0060] At each node of an octree, an occupancy is signaled (when not inferred) for one or more of its child nodes (up to eight nodes). Multiple neighborhoods are specified, including (a) nodes that share a face with a current octree node, (b) nodes that share a face, edge, or vertex with the current octree node, etc. In each neighborhood, the occupancy of a node and / or its children can be used to predict the occupancy of the current node or its children. For points that are sparsely populated at certain nodes in the tree, the codec also supports a direct encoding mode, where the three-dimensional position of the point is encoded directly. A flag can be signaled to indicate that a direct mode is signaled. At the lowest level, the number of points associated with the node / leaf node of the octree can also be encoded.

[0061] Once the geometry is encoded, the attributes corresponding to the geometry points are encoded. When there are multiple attribute points corresponding to a reconstructed / decoded geometry point, an attribute value can be derived that is representative of the reconstructed point.

[0062] There are three attribute coding methods using G-PCC: region adaptive hierarchical transform (RAHT) coding, interpolation-based hierarchical nearest neighbor prediction (predictive transform), and interpolation-based hierarchical nearest neighbor prediction with an update / uplift step (uplift transform). RAHT and uplift are typically used for category 1 data, while prediction is typically used for category 3 data. However, any method can be used for any data, and so too Petition 870250085531, dated 09 / 22 / 2025, pp. 279 / 359 With the 20 / 67 geometry codecs in G-PCC, the attribute encoding method used to encode the point cloud is specified in the bitstream.

[0063] Attribute encoding can be conducted at a level of detail (LOD) where, with each level of detail, a finer representation of the point cloud attribute can be obtained. Each level of detail can be specified based on the distance metric from neighboring nodes or based on a sampling distance.

[0064] In the 200 point cloud encoder, the residuals obtained as output from the coding methods for the attributes are quantized. The residuals can be obtained by subtracting the attribute value from a prediction that is derived based on points in the neighborhood of the current point and based on the attribute values ​​of previously coded points. The quantized residuals can be coded using context-adaptive arithmetic coding.

[0065] In the example in Figure 2, the point cloud encoder 200 includes deep learning-based geometry encoder 202, geometry reconstruction unit 206, geometry downscaling unit 210, recoloring unit 212, color transformation unit 204, region-based adaptive hierarchical transform (RAHT) unit 218, LOD generation unit 220, elevation unit 222, coefficient quantization unit 224, and arithmetic encoding unit 226.

[0066] As shown in the example in Figure 2, the point cloud encoder 200 can obtain a set of point positions in the point cloud and a set of attributes. The point cloud encoder 200 can obtain the set of point positions in the point cloud and the set of attributes from data source 104 (Figure 1). Positions can include coordinates of points in a point cloud. Attributes can include information about the points in the point cloud, such as colors, reflectance, intensity, or similar, associated with the points in the point cloud. The deep learning-based geometry encoder 202 of the point cloud encoder 200 can generate Petition 870250085531, dated 09 / 22 / 2025, pages 280 / 359 21 / 67 a geometry bitstream 203 that includes an encoded representation of the positions of the points in the point cloud. The point cloud encoder 200 can also generate an attribute bitstream 205 that includes an encoded representation of the attribute set.

[0067] After encoding the geometry information, the geometry reconstruction unit 206 can decode and reconstruct the geometry information. The geometry downscaling unit 210 can downscale the geometry information. As mentioned above, geometry information can be represented using an octree. A root node of the octree can be partitioned into eight subnodes. For each node of the octree, the node can be further partitioned into eight subnodes, including splitting the node in half along the X, Y, and Z dimensions. Such partitioning can continue until, for example, reaching a node of the smallest size for the octree, called a leaf node of the octree. That is, a leaf node has no child nodes and is unpartitioned.

[0068] In order to downscale the octree, the geometry downscaling unit 210 can determine that a node has eight leaf subnodes, and determine a number of those that are occupied among the eight leaf subnodes. If the number is above a threshold (for example, zero), the geometry downscaling unit 210 can represent the node as an occupied leaf node in a downscaled octree. Otherwise, if the number is less than or equal to the threshold, the geometry downscaling unit 210 can represent the node as an unoccupied leaf node in the downscaled octree. For example, if the threshold is zero, all eight leaf subnodes would need to be unoccupied to represent the node as an unoccupied leaf node in the downscaled octree; otherwise, the node would be represented as an occupied leaf node in the downscaled octree.

[0069] After downscaling the geometry information, recoloring unit 212 can apply the attribute information to the scaled-down octree points. For example, recoloring unit 212 can similarly downscale the original attribute information to the same degree as the Petition 870250085531, dated 09 / 22 / 2025, pp. 281 / 359 22 / 67 geometry information. Such downscaling may include selectively removing attribute data, merging attribute data, or otherwise reducing attribute data so that the downscaled attribute data can be applied to the downscaled geometry.

[0070] The color transformation unit 204 can transform the color information of attributes to a different domain. For example, the color transformation unit 204 can transform color information from an RGB color space to a YCbCr color space.

[0071] Furthermore, the RAHT 218 unit can apply RAHT encoding to the attributes of the reconstructed points. In some examples, under RAHT, the attributes of a 2x2x2 block of point positions are taken and transformed along one direction to obtain four low (L) and four high (H) frequency nodes. Subsequently, the four low frequency (L) nodes are transformed in a second direction to obtain two low (LL) and two high (LH) frequency nodes. The two low frequency (LL) nodes are transformed along a third direction to obtain one low frequency (LLL) and one high frequency (LLH) node. The low frequency LLL node corresponds to the DC coefficients and the high frequency H, LH, and LLH nodes correspond to the AC coefficients. The transformation in each direction can be a 1-D transformation with two coefficient weights.Low-frequency coefficients can be considered as 2x2x2 block coefficients for the next higher level of the RAHT transformation, and the AC coefficients are encoded without changes; such transformations continue up to the upper root node. The tree traversal for encoding is top-down and is used to calculate the weights to be used for the coefficients; the transformation order is bottom-up. The coefficients can then be quantized and encoded.

[0072] Alternatively or additionally, the LOD generation unit 220 and the elevation unit 222 can apply LOD processing and elevation, respectively, to the attributes of the reconstructed points. LOD generation is used to divide the attributes into different levels of refinement. Each level of Petition 870250085531, dated 09 / 22 / 2025, pp. 282 / 359 23 / 67 Refinement provides a refinement to the attributes of the point cloud. The first refinement level provides a rough approximation and contains few points; the subsequent refinement level typically contains more points, and so on. Refinement levels can be constructed using a distance-based metric or can also use one or more other classification criteria (e.g., subsampling from a particular order). In this way, all reconstructed points can be included in a refinement level. Each level of detail can be produced by taking a union of all points up to a particular refinement level: for example, LOD1 is obtained based on refinement level RL1, LOD2 is obtained based on RL1 and RL2, ... LODN is obtained by joining RL1, RL2, ... RLN.In some cases, LOD generation may be followed by a prediction scheme (e.g., prediction transformation) where the attributes associated with each point in the LOD are predicted from a weighted average of the previous points, and the residual part is quantized and entropy-encoded. The elevation scheme is based on the prediction transformation mechanism, where an update operator is used to update the coefficients and an adaptive quantization of the coefficients is performed.

[0073] The RAHT unit 218 and the elevation unit 222 can generate coefficients based on attributes. The coefficient quantization unit 224 can quantize the coefficients generated by the RAHT unit 218 or the elevation unit 222. The arithmetic coding unit 226 can apply arithmetic coding to the syntax elements representing the quantized coefficients. The point cloud encoder 200 can output these syntax elements in the attribute bitstream 205. The attribute bitstream 205 can also include other syntax elements, including non-arithmetically coded syntax elements.

[0074] In some examples, the geometry downscaling unit 210 can downscale geometry data by a certain amount, represented by a particular value, such as a downscaling factor. Point cloud encoder 200 can encode the value as a parameter of Petition 870250085531, dated 09 / 22 / 2025, pp. 283 / 359 24 / 67 a set of parameters, such as a sequence parameter set (SPS) or an attribute parameter set (APS), a slice header, a frame header, or other high-level syntax (HLS). In some examples, the downscaling factor may also indicate an amount by which to downscale the attribute data before recoloring. In some examples, point cloud encoder 200 may encode a second value, separate from the first value, indicating the amount by which to downscale the attribute data.

[0075] Figure 3 is a block diagram illustrating an example 300 point cloud decoder. In the example in Figure 3, the 300 point cloud decoder includes a deep learning-based geometry decoder 302, an attribute arithmetic decoding unit 304, a geometry downscaling unit 306, an inverse quantization unit 308, a RAHT unit 314, a LoD generation unit 316, an inverse upscaling unit 318, an inverse color transformation unit 322, and a point cloud upscaling unit 324.

[0076] The point cloud decoder 300 can receive a geometry bitstream 203 and an attribute bitstream 205. The deep learning-based geometry decoder 302 generally decodes geometry data from the geometry bitstream 203. The attribute arithmetic decoding unit 304 of the decoder 300 can apply arithmetic decoding (e.g., context-adaptive binary arithmetic coding (CABAC) or other type of arithmetic decoding) to syntax elements in the attribute bitstream 205 to decode the attribute bitstream 205.

[0077] In general, attribute bitstream 250 represents a scaled-down version of attribute data relative to the geometry data of geometry bitstream 203. Therefore, geometry downscaling unit 306 can downscale the reproduced geometry data from the deep learning-based geometry decoder 302, for example, from Petition 870250085531, dated 09 / 22 / 2025, pp. 284 / 359 25 / 67 according to a downscaling value. The 300 point cloud decoder can decode the downscaling value from high-level syntax (HLS) data, such as a sequence parameter set (SPS), an attribute parameter set (APS), a slice header, a frame header, or the like. The 306 geometry downscaling unit can downscale geometry data according to the downscaling factor.

[0078] Additionally, the inverse quantization unit 308 can inversely quantize attribute values. The attribute values ​​can be based on syntax elements obtained from the attribute bitstream 205 (for example, including syntax elements decoded by the attribute arithmetic decoding unit 304).

[0079] Depending on how the attribute values ​​are encoded, the RAHT unit 314 can perform RAHT encoding to determine, based on the inversely quantized attribute values, the color values ​​for the point cloud points. RAHT decoding is done from top to bottom of the tree. At each level, the low- and high-frequency coefficients, which are derived from the inverse quantization process, are used to derive the constituent values. At the leaf node, the derived values ​​correspond to the attribute values ​​of the coefficients. The weight derivation process for the points is similar to the process used in the point cloud encoder 200. Alternatively, the LOD generation unit 316 and the inverse elevation unit 318 can determine color values ​​for point cloud points using a detail-based technique level.The LOD 316 generation unit decodes each LOD, resulting in progressively finer representations of the point attribute. Using a prediction transformation, the LOD 316 generation unit derives the point prediction from a weighted sum of points that are in previous LODs or previously reconstructed in the same LOD. The LOD 316 generation unit can add the prediction to the residual (which is obtained after inverse quantization) to obtain the reconstructed attribute value. When the elevation scheme is used, the LOD 316 generation unit... Petition 870250085531, dated 09 / 22 / 2025, pages 285 / 359 26 / 67 can also include an update operator to update the coefficients used to derive the attribute values. The LOD generation unit 316 can, in this case, also apply inverse adaptive quantization.

[0080] In addition, the inverse color transformation unit 322 can apply an inverse color transformation to color values. The inverse color transformation can be an inverse of a color transformation applied by the color transformation unit 204 of the encoder 200. For example, the color transformation unit 204 can transform color information from an RGB color space to a YCbCr color space. Consequently, the inverse color transformation unit 322 can transform color information from the YCbCr color space to the RGB color space.

[0081] After decoding both geometry and attribute information, the 324 point cloud upscaling unit can reconstruct the point cloud. In particular, the 324 point cloud upscaling unit can upscale the attribute information to the scale of the geometry information. The 324 point cloud upscaling unit can upscale the attribute information according to the downscaling factor. Alternatively, the 300 point cloud decoder can decode a separate value representing an amount of upscaling to be applied to the attribute information. In some examples, the 324 point cloud upscaling unit can be a deep learning-based attribute supersampler.Finally, the 324 point cloud upscaling unit can apply the upscaled attribute information to the points of the geometry information to reconstruct the point.

[0082] The various units in Figure 2 and Figure 3 are illustrated to help understand the operations performed by the encoder 200 and the decoder 300. The units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide particular functionality and are predefined in the Petition 870250085531, dated 09 / 22 / 2025, pp. 286 / 359 27 / 67 operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and to provide flexible functionality in the operations that can be performed. For example, programmable circuits can execute software or firmware that causes the programmable circuits to operate in the manner defined by the software or firmware instructions. Fixed-function circuits can execute software instructions (for example, to receive parameters or to emit parameters), but the types of operations that fixed-function circuits perform are generally immutable. In some examples, one or more of the units may be distinct circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0083] As described above, machine learning techniques, such as deep learning, can be used for encoding and decoding point cloud data. The techniques in this disclosure include the use of a deep learning-based lossy point cloud geometry compression scheme for dynamic point cloud compression. A lossy geometry scheme predicts the latent representation of the current frame using the previous frame by employing a prediction network. The example techniques perform P-frame inter-frame point cloud encoding where the current frame is encoded with reference to a previously decoded frame. The architecture can be implemented using a sparse convolutional neural network (CNN) with sparse tensors.The example architecture employs convolution on target coordinates to map the latent representation of the previous frame to the subsampled coordinates of the current frame to predict the embedding of current frame attributes. The encoder transmits the residual of the predicted attributes and the actual attributes by compressing them using a learned probabilistic factored entropy model. Compared to G-PCC and V-PCC, machine learning techniques show better compression performance on dense point clouds with runtime. Petition 870250085531, dated 09 / 22 / 2025, pp. 287 / 359 28 / 67 efficient encoding / decoding.

[0084] There may be problems with using machine learning to encode / decode (e.g., compress / decompress) point cloud data. Often, a point cloud compression codec may perform well in geometry encoding and have deficiencies in attribute encoding, or vice versa. This can be true for deep learning-based point cloud compression schemes, where deep learning-based point cloud compression schemes often perform well in geometry compression but cannot perform or have deficiencies in attribute compression.

[0085] In one or more examples described in this disclosure, to achieve coding efficiencies such as that of deep learning, there may be a benefit in being able to separate geometry compression and attribute compression from the compression frameworks and then being able to combine geometry compression methods from one codec to the attribute compression method of another codec to be able to achieve better compression performance. In one or more examples, this disclosure describes flexible configurations in the compression framework where point cloud encoder 200 and point cloud decoder 300 employ deep learning-based geometry compression with traditional attribute compression methods for point cloud compression.

[0086] By employing downscaling and subsequent upscaling for attribute information, additional flexibility can be provided. Most of the bit rate in previous encoding schemes was consumed by attribute information. Any improvement in the attribute compression scheme can greatly improve the overall encoding efficiency of the structure. Thus, according to the techniques of this disclosure, the 200 point cloud encoder can downscale the attribute information before encoding, and the 300 point cloud decoder can decode and then upscale. Petition 870250085531, dated 09 / 22 / 2025, pages 288 / 359 29 / 67 of the attribute information.

[0087] High-level syntax (HLS) data can be signaled by point cloud encoder 200 and received by point cloud decoder 300. The HLS data may include an attribute encoding type employed to encode the recolored attributes for point cloud decoder 300 in order to reconstruct the attribute values. This encoding method, for example, the G-PCC's region-adaptive hierarchical transform (RAHT), can be signaled in a parameter set, for example, the sequence parameter set (SPS) or the attribute parameter set (APS) as an identifier. The bitstream portion carrying the encoded attribute bits is, for example, a NALU. A list of encoding methods can be specified, each encoding method providing a means to encode point cloud attributes (optionally, they can also encode geometry in a lossless manner).An index for this list can be signaled in the bitstream to indicate the encoding method used to encode the attributes. This index can be signaled in a parameter set (e.g., APS, SPS) or other means. When deep learning mechanisms are used for recoloring or decoding, the parameters / coefficients corresponding to the recoloring or decoding can also be signaled in the bitstream.

[0088] Figure 4 is a block diagram illustrating an example 400 encoding structure. Figure 5 is a block diagram illustrating an example 500 decoding structure. In Figure 4, the 400 encoding structure includes deep learning-based geometry encoder 404, G-PCC recoloring unit 406, and lossless and lossy G-PCC-based attribute geometry encoder 408. In general, the 400 encoding structure can correspond to the 200 point cloud encoder of Figures 1 and 2, where the deep learning-based geometry encoder 404 can correspond to the deep learning-based geometry encoder 202, the G-PCC recoloring unit 406 can correspond to the recoloring unit 212, and the Petition 870250085531, dated 09 / 22 / 2025, pages 289 / 359 The lossless geometry encoder and the lossy attribute encoder by G-PCC 408 can correspond to any or all of the following: color transformation unit 204, RAHT unit 218, LOD generation unit 220, elevation unit 222, coefficient quantization unit 224, and arithmetic encoding unit 226. In this example, the original point cloud 402 is provided to both the deep learning-based geometry encoder 404 and the G-PCC recoloring unit 406. The deep learning-based geometry encoder 404 encodes geometry information from the original point cloud 402 and forms the geometry bitstream 410 including encoded geometry information for the original point cloud 402.

[0089] The deep learning-based geometry encoder 404 also decodes and reconstructs geometry information, for example, an octree including nodes that indicate whether points are present at the nodes (i.e., whether the nodes are occupied). Occupied nodes can be partitioned into eight subnodes, each of which can include indications of being occupied. The G-PCC recoloring unit 406 can receive the reconstructed geometry information from the deep learning-based geometry encoder 404 and use the geometry and attribute information from the original point cloud 402 to recolor the reconstructed geometry. The lossless and lossy G-PCC-based geometry encoder 408 can then encode the recolored reconstructed geometry to form the attribute bitstream 412.

[0090] The 500 decoding structure in Figure 5 includes the deep learning-based geometry decoder 502 (which may correspond to the deep learning-based geometry decoder 302) and the G-PCC decoder 504 (which may correspond to the attribute arithmetic decoding unit 304, inverse quantization unit 308, RAHT unit 314, LOD generation unit 316, inverse elevation unit 318, and inverse color transformation unit 322). In this example, the deep learning-based geometry decoder 502 receives a geometry bitstream 508. The Petition 870250085531, dated 09 / 22 / 2025, pp. 290 / 359 The 31 / 67 deep learning-based geometry decoder 502 decodes the geometry bitstream 508 and reconstructs the geometry information. The G-PCC decoder 504 receives the attribute bitstream 510 and uses the reconstructed geometry information to decode the attribute bitstream 510 and reconstruct the point cloud. For example, the G-PCC decoder 504 can apply the decoded attribute information to the reconstructed geometry information to form the recolored reconstructed point cloud 506.

[0091] In Figures 4 and 5, the 400 encoding structure compresses and the 500 decoding structure decompresses the geometry using a deep learning-based encoder / decoder to obtain a reconstructed point cloud. The reconstructed point cloud differs in geometry from the original point cloud and may not simply be used as the attributes of the original point cloud. In some examples, the 400 encoding structure compresses and the 500 decoding structure may use the G-PCC recoloring scheme to alter the attributes of the original point cloud in order to create newer attributes for the reconstructed point cloud. The recolored point cloud has geometry from the reconstructed point cloud and attributes that were derived from the original point cloud. The recolored point cloud is G-PCC encoded, where the geometry was encoded in a lossless manner and the attributes were encoded in a lossy manner.The reconstructed geometry is combined with reconstructed attributes to obtain the reconstructed point cloud.

[0092] The following changes / flexibility can be added to the structure illustrated in Figures 4 and 5. As an example, although Figures 4 and 5 illustrate deep learning-based encoders and decoders, the geometry encoder or decoder is not necessarily limited to a deep learning-based geometry encoder or decoder, and any geometry encoder or decoder can be used.

[0093] For recoloring, the 400 encoding structure compresses and the 500 decoding structure can employ a recoloring scheme. Petition 870250085531, dated 09 / 22 / 2025, pp. 291 / 359 32 / 67 based on nearest neighbor search based on weighted distance employed in the G-PCC standard: WG 7, MPEG 3D Graphics Coding, G-PCC codec description, document N00271, January 2022. The recoloring scheme can alter attributes to fit the most recent geometry. Any recoloring scheme that can alter attribute values ​​and / or their correspondence with geometry can be employed as a recoloring scheme. The recoloring algorithm need not be limited to recoloring color attributes, e.g., RGB or YCbCr, but can more generically be an algorithm that recomputes attribute values, such as normal vectors, reflectance, etc., from point positions in one geometry to point positions in a second geometry. Deep learning mechanisms can also be applied to perform the recoloring.

[0094] As illustrated in Figure 4, the lossless geometry and lossy attribute encoder using G-PCC is shown. The example techniques are not limited to lossless geometry and lossy attribute encoding using G-PCC, for example, the adaptive hierarchical region-by-region transform (RAHT) of G-PCC. This method can be employed with any encoding scheme, including other deep learning-based encoding or using V-PCC. Furthermore, instead of employing lossless geometry and lossy attribute encoding, the example can use any lossy attribute encoder. Deep learning mechanisms can also be applied to perform the decoding.

[0095] The 400 encoding structure and the 500 decoding structure use a geometry encoder / decoder from one codec and an attribute encoder / decoder from a separate codec to create a complete codec structure that would outperform the two individual codecs. In lossy point cloud compression, the geometry changes after compression, and linking attributes from a separate codec to their corresponding geometry is challenging. Thus, in the 400 encoding structure and the 500 point cloud decoding structure, a recoloring scheme is used to link attributes together. Petition 870250085531, dated 09 / 22 / 2025, pp. 292 / 359 33 / 67 with its corresponding geometry.

[0096] Figure 6 is a block diagram illustrating an example 600 point cloud coding structure according to the techniques of this disclosure. In this example, the 600 point cloud coding structure includes deep learning-based geometry encoder 604 (which may correspond to deep learning-based geometry encoder 202), geometry downscaling unit 606 (which may correspond to geometry downscaling unit 210), recoloring unit 608 (which may correspond to recoloring unit 212), and attribute encoder 610 (which may correspond to color transformation unit 204, RAHT unit 218, LOD generation unit 220, elevation unit 222, coefficient quantization unit 224, and arithmetic coding unit 226).The 600 point cloud encoding structure differs from the 400 encoding structure in that the 600 point cloud encoding structure includes the geometry downscaling unit 606 and the attribute encoder 610 encodes a downscaling version of the attribute information, as discussed in more detail below.

[0097] In general, the deep learning-based geometry encoder Unit 604 receives geometry information from the original point cloud 602, while the recoloring unit 608 receives both geometry and attribute information from the original point cloud 602. The deep learning-based geometry encoder 604 encodes the geometry information to form the geometry bitstream 612. The deep learning-based geometry encoder 604 decodes and reconstructs the geometry information and provides the reconstructed geometry information to the geometry downscaling unit 606. The geometry downscaling unit 606 can then downscale the geometry information and provide downscaled geometry information to the recoloring unit 608. The recoloring unit 608 can then form a downscaled recolored point cloud from the Petition 870250085531, dated 09 / 22 / 2025, pp. 293 / 359 34 / 67 scaled-down geometry information and the original geometry and attribute information of the original point cloud 602. The attribute encoder 610 can then encode attribute information from the scaled-down recolored point cloud to form the attribute bitstream 614.

[0098] Figure 7 is a block diagram illustrating an example 700 point cloud decoding structure according to the techniques of this disclosure. In this example, the 700 point cloud decoding structure includes a deep learning-based geometry decoder 702, a geometry downscaling unit 704, an attribute decoder 706, and a deep learning-based attribute supersampler 708.The deep learning-based geometry decoder 702 can correspond to the deep learning-based geometry decoder 302, the geometry downscaling unit 704 can correspond to the geometry downscaling unit 306, the attribute decoder 706 can correspond to the arithmetic attribute decoding unit 304, inverse quantization unit 308, RAHT unit 314, LOD generation unit 316, inverse upscaling unit 318 and inverse color transformation unit 322, and the deep learning-based attribute supersampler 708 can correspond to the point cloud upscaling unit 324.

[0099] In general, the deep learning-based geometry decoder 702 receives a geometry bitstream 710. The deep learning-based geometry decoder 702 decodes geometry information from the geometry bitstream 710 and reconstructs geometry information from the decoded geometry information. The geometry downscaling unit 704 downscales the geometry information to form downscaled geometry information and provides the downscaled geometry information to the attribute decoder 706. The attribute decoder 706 receives the attribute bitstream 712, including encoded attribute information, and decodes the encoded attribute information. The decoder of Petition 870250085531, dated 09 / 22 / 2025, pp. 294 / 359 35 / 67 attribute 706 can provide the decoded attribute information to the deep learning-based attribute supersampler 708, which can also receive the original reconstructed geometry information and scale the attribute information up to the scale of the reconstructed geometry information. The deep learning-based attribute supersampler 708 can also apply the scaled-up attribute information to the reconstructed geometry information to form a reconstructed point cloud 714.

[0100] According to the techniques of this disclosure, the attribute encoder The 610 encodes attributes at a reduced scale to save attribute bits, which can increase encoding efficiency. Geometry information is compressed using the deep learning-based geometry encoder 604, and then decoded and reconstructed to obtain a reconstructed point cloud. The geometry downscaling unit 606 can downscale the reconstructed geometry information using various downscaling factors (e.g., 1, 2, 4, 8, etc.) as discussed in more detail below. The recoloring unit 608 then recolors the geometry at a reduced scale using the attributes of the original point cloud according to a recoloring scheme, such as the G-PCC recoloring scheme. The attribute encoder 610 can then encode the recolored point cloud attributes with an attribute encoding scheme (lossy or lossless), such as G-PCC RAHT or G-PCC predictive / elevation transformation.

[0101] In the point cloud decoding structure 700, the geometry bitstream 710 is decoded by the deep learning-based geometry decoder 702. The reconstructed geometry is scaled down by the geometry scaling unit 704, and then provided to the attribute decoder 706. The attribute decoder 706 can perform inverse GPCC RAHT or predictive / uplift transformation. The attribute decoder 706 can link the scaled-down decoded attributes to their corresponding Petition 870250085531, dated 09 / 22 / 2025, pages 295 / 359 36 / 67 Reconstructed geometry at a reduced scale to obtain a point cloud at a reduced scale. As an example, the deep learning-based point cloud attribute supersampler is employed to oversample the attributes and map the oversampled attributes to the reconstructed geometry.

[0102] In some examples, explicit downscaling of the reconstructed geometry may not be performed. For example, the geometry decoding process itself may include one or more downscaled versions of geometry that are scaled up / processed to obtain the reconstructed geometry. The one or more downscaled versions of geometry may be passed to the attribute decoder and used instead of the downscaled reconstructed geometry. This may avoid the need to perform the explicit downscaling operation on the reconstructed geometry, thus saving processing time and resources.

[0103] Figures 8A and 8B are conceptual diagrams illustrating examples of voxel downscaling of a point cloud. Figure 8A depicts a sample voxel 800 including subvoxels 802A, 802B, and 802C, each of which is occupied, and other subvoxels are not occupied. In this example, downscaling voxel 800 results in the downscaled voxel 804 which is occupied, as subvoxels 802A, 802B, and 802C are occupied. Figure 8B depicts a sample voxel 810 including all unoccupied subvoxels. So, in this example, downscaling voxel 810 results in the downscaled voxel 812 which is unoccupied.

[0104] In some examples, the geometry downscaling unit 306 or the geometry downscaling unit 704 may employ a voxel grid downscaling of KxKxK, where K is a value that defines a downscaling factor. If K is 2, then the point cloud is divided into a grid of 2x2x2 voxels, where each voxel is a point in three-dimensional space. Then, the 8 voxels within the 2x2x2 voxel grid are integrated into a single voxel. In some examples, the final voxel is considered occupied if any of the 8 voxels in its Petition 870250085531, dated 09 / 22 / 2025, pp. 296 / 359 The 37 / 67 interior is also occupied. In some examples, the final voxel is considered occupied if the number of occupied subvoxels exceeds a threshold value.

[0105] Figure 9 is a block diagram illustrating a set of example stages that can be included in a deep learning-based attribute supersampler, such as the deep learning-based attribute supersampler 708 or the point cloud upscaling unit 324. In this example, the stage set includes deep learning layers 902, supersampling layer 910, and deep learning layers 920. Deep learning layers 902 initially process the geometry and attribute information of the scaled-down point cloud 900, providing results to the supersampling layer 910. The supersampling layer 910 then oversamples the results of the deep learning layers 902 using reconstructed geometry data 904. Deep learning layers 920 then process the oversampled geometry and attribute information to produce the reconstructed point cloud 930.

[0106] Figure 10 is a block diagram illustrating a set of example stages that can be included in a deep learning-based feature supersampler, such as the deep learning-based feature supersampler 708 or the point cloud upscaling unit 324. In this example, the stage set includes deep learning layers 1002 (which may correspond to deep learning layers 902), supersampling layer 1010 (which may correspond to supersampling layer 910), and deep learning layers 1020 (which may correspond to deep learning layers 920). Specifically, the deep learning layers 102 include the 3x3x3 sparse convolutional (SConv) layer 1004, the inception residual block (IRB) layer 1006, and the 3x3x3 SConv layer 1008. In this example, the supersampling layer 1010 is a 5x5x5 convolutional layer in target coordinates.In this example, the layers of deep learning. Petition 870250085531, dated 09 / 22 / 2025, pp. 297 / 359 38 / 67 1020 includes the 3x3x3 SConv layer 1022, the IRB layer 1024, and the SConv layer 1026. These layers process the geometry and attributes of the scaled-down point cloud 1000 sequentially, as well as the reconstructed geometry data 1012, to produce the reconstructed point cloud 1030.

[0107] Sparse convolutional layers with a kernel size of 3x3x3 are used in Figure 10 for example purposes, but other deep learning-based layers can be used additionally or alternatively. Similarly, although a convolution in target coordinates with a kernel size of 5x5x5 is used as an example, other supersampling techniques can be used in place of this layer. The convolution layer in target coordinates can be employed to map the attributes of the scaled-down geometry to the scaled-up geometry. If the scale-down / supersampling factor is large, multiple successive supersampling layers can be employed, or multiple successive attribute supersamplers can be used to oversample the attributes.

[0108] For an attribute supersampler, a mean squared error (MSE) loss can be employed during training instead of a classification loss. The MSE loss can compare a predicted attribute to an original attribute to help train the network. Training is not limited to MSE loss but can employ any loss function that compares and attempts to minimize the distance between the predicted attributes and the original attributes. The network can be trained using an Adam optimizer. The network does not need to be limited to an Adam optimizer and can instead or additionally include root mean squared propagation (RMSProp), stochastic gradient descent (SGD), adaptive gradient algorithm (Adagrad), or any other optimizer like these.

[0109] The attribute supersampler does not necessarily need to be based on deep learning. Examples of attribute supersamplers Petition 870250085531, dated 09 / 22 / 2025, pp. 298 / 359 39 / 67 baseados em aprendizado não profundo incluem aqueles discutidos em Marc Alexa, Johannes Behr, Daniel Cohen-Or, Shachar Fleishman, David Levin e Claudio T. Silva, Computing and rendering point set surfaces, IEEE Trans. Vis. & Comp. Graphics, 9(1):3-15, 2003; Yaron Lipman, Daniel Cohen-Or, David Levin e Hillel TalEzer, Parameterization-free projection for geometry reconstruction, ACM Trans. on Graphics (SIGGRAPH), 26(3):22:1-5, 2007; Hui Huang, Dan Li, Hao Zhang, Uri Ascher e Daniel Cohen-Or, Consolidation of unorganized point clouds for surface reconstruction, ACM Trans. On Graphics (SIGGRAPH Asia), 28(5):176:1-7, 2009; Hui Huang, Shihao Wu, Minglun Gong, Daniel Cohen-Or, Uri Ascher e Hao Zhang, Edge-aware point set resampling, ACM Trans. on Graphics, 32(1):9:1-12, 2013; e Shihao Wu, Hui Huang, Minglun Gong, Matthias Zwicker e Daniel Cohen-Or, Deep points consolidation, ACM Trans. on Graphics (SIGGRAPH Asia), 34(6):176:1-13, 2015.Where a fully convolutional network using sparse convolutions is shown in the example in Figure 10, the same could be achieved using different layers based on deep learning. Similarly, for the supersampling layer, similar results could be achieved using other deep learning layers, such as transposed convolutions, deconvolutions, degrouping, or the like.

[0110] The attributes in the structure are not restricted to color information. For example, these techniques can be applied to encode / oversample surface normal, reflectance, intensity, or similar information. Color space conversion can be applied to individual modules or the entire structure. For example, the YCbCr color space can be employed during recoloring by the attribute encoder, attribute decoder, and / or deep learning-based attribute oversampler. The 200 point cloud encoder can encode syntax elements representing color spaces associated with individual modules or the entire structure in the bitstream, and similarly, the 300 point cloud decoder can decode and use the syntax elements to determine which spaces Petition 870250085531, dated 09 / 22 / 2025, pp. 299 / 359 40 / 67 colors should be used in which modules?

[0111] The 200 point cloud encoder and the 309 point cloud decoder can encode a downscaling factor so that the 300 point cloud decoder determines how much adaptive geometry downscaling should be performed and how much attribute oversampling to perform. If no downscaling is performed in the 200 point cloud encoder, then there is no need for upscaling in the decoder, and the architecture shown in Figures 3 and 4 can be employed. This downscaling factor can be signaled in a parameter set, a slice, or other syntax structure in the bitstream, for example, in the sequence parameter set (SPS) or the attribute parameter set (APS). Although called an undersampling factor (from the encoder's perspective), the 300 point cloud decoder can use this factor for other operations, including oversampling.In some examples, two downscaling factors can be signaled in the bitstream: one for geometry and one for attributes. In some examples, a subsampling factor can be signaled that applies to both geometry and attributes.

[0112] The type of attribute encoding technique employed to encode the recolored attributes can be signaled to the decoder side so that the 300 point cloud decoder reconstructs the attribute values. This encoding method, for example, the G-PCC's region-adaptive hierarchical transform (RAHT), can be signaled in a parameter set, for example, the sequence parameter set (SPS) or the attribute parameter set (APS) as an identifier. The bitstream portion carrying the encoded attribute bits can be, for example, a network abstraction layer unit (NALU).

[0113] A list of encoding techniques can be specified. Each encoding technique can provide a way to encode the attributes of the point cloud (optionally, they can also encode the geometry in a way without Petition 870250085531, dated 09 / 22 / 2025, pages 300 / 359 41 / 67 losses). An index for this list can be signaled in the bitstream to indicate the encoding technique used to encode the attributes. This index can be signaled in a parameter set (e.g., APS, SPS) or in another HLS.

[0114] When deep learning mechanisms are used for recoloring or decoding, the parameters / coefficients corresponding to the recoloring or decoding can also be signaled in the bitstream.

[0115] Figure 11 is a flowchart illustrating an example point cloud data encoding method according to the techniques of this disclosure. The method in Figure 11 is explained in relation to the 200 point cloud encoder of Figure 2. Other point cloud encoding devices, such as those conforming to the 600 point cloud encoding framework of Figure 6, may perform this method or a similar method.

[0116] Initially, point cloud encoder 200 encodes geometry information from a point cloud (1100). For example, point cloud encoder 200 can use a deep learning-based point cloud encoding technique to encode the geometry information. Point cloud encoder 200 can then decode and reconstruct the geometry information (1102). Point cloud encoder 200 can also downscale the geometry information (1104). Point cloud encoder 200 can recolor the downscaled geometry information using the point cloud attribute information (1106). Point cloud encoder 200 can then downscale the attribute information (1108).In some examples, the 200 point cloud encoder can additionally encode HLS information by specifying, for example, a downscaling factor for geometry information and / or an oversampling factor for attribute information.

[0117] In this way, the method in Figure 11 represents an example of a method for encoding point cloud information, including: decoding encoded point cloud geometry data into a point cloud, in order to Petition 870250085531, dated 09 / 22 / 2025, pages 301 / 359 42 / 67 reconstruct point cloud geometry data into the point cloud; downscale the point cloud geometry data to form downscaled point cloud geometry data; and encode attribute data into the point cloud using the downscaled point cloud geometry.

[0118] Figure 12 is a flowchart illustrating an example point cloud data decoding method according to the techniques of this disclosure. The method in Figure 12 is explained in relation to the 300 point cloud decoder of Figure 3. Other point cloud decoding devices, such as those conforming to the 700 point cloud decoding framework of Figure 7, may perform this method or a similar method.

[0119] Initially, point cloud decoder 300 decodes geometry information (1200). Point cloud decoder 300 can also decode a downscaling factor, for example, in high-level syntax (HLS) data. Point cloud decoder 300 can downscale geometry information (1202), for example, according to the downscaling factor. Point cloud decoder 300 can then decode attribute information (1204). Point cloud decoder 300 can then oversample attribute information (1206), for example, using a deep learning-based attribute oversampler and / or a bitstream-included upscaling factor.

[0120] In this way, the method in Figure 12 represents an example of a method for encoding point cloud information, including: decoding encoded point cloud geometry data to a point cloud in order to reconstruct point cloud geometry data for the point cloud; downscaling the point cloud geometry data to form downscaled point cloud geometry data; and encoding attribute data for the point cloud using the downscaled point cloud geometry.

[0121] Figure 13 is a conceptual diagram illustrating a laser package. 1300, such as a LIDAR sensor or other system that includes one or more lasers, Petition 870250085531, dated 09 / 22 / 2025, pages 302 / 359 43 / 67 scanning points in three-dimensional space. The 1300 laser package may correspond to the 380 LIDAR in Figure 7. Data source 104 (Figure 1) may include a 1300 laser package.

[0122] As shown in Figure 13, point clouds can be captured using the 1300 laser package; that is, the sensor scans points in three-dimensional space. It should be understood, however, that some point clouds are not generated by a true LIDAR sensor, but can be encoded as if they were. In the example in Figure 13, the 1300 laser package includes a 1302 LIDAR head that includes multiple 1304A to 1304E lasers (collectively, 1304 lasers) arranged in a vertical plane at different angles relative to an origin point. The 1300 laser package can rotate around a vertical axis 1308. The 1300 laser package can use the returned laser light to determine the distances and positions of the points in the point cloud. The laser beams 1306A to 1306E (collectively, laser beams 1306) emitted by the lasers 1304 of the laser package 1300 can be characterized by a set of parameters.The distances denoted by the arrows 1310, 1312 denote example laser correction values ​​for lasers 1304B and 1304A, respectively.

[0123] Lasers 1300 can be used to obtain both geometry data and attribute data for points from the geometry data. According to the techniques of this disclosure, point cloud geometry data can be scaled down, and then attribute data can be encoded using the scaled-down point cloud geometry data.

[0124] Figure 14 is a conceptual diagram illustrating an example range measurement system 1400 that can be used with one or more techniques of this disclosure. In the example of Figure 14, the range measurement system 1400 includes an illuminator 1402 and a sensor 1404. The illuminator 1402 can emit light 1406. In some examples, the illuminator 1402 can emit light 1406 as one or more laser beams. The light 1406 can be at one or more wavelengths, such as an infrared wavelength or a visible light wavelength. In Petition 870250085531, dated 09 / 22 / 2025, pp. 303 / 359 44 / 67 other examples, light 1406 is not a coherent laser light. When light 1406 encounters an object, such as object 1408, light 1406 creates a return light 1410. The return light 1410 may include backscattered and / or reflected light. The return light 1410 may pass through a lens 1411 that directs the return light 1410 to create an image 1412 of object 1408 on sensor 1404. Sensor 1404 generates signals 1414 based on the image 1412. The image 1412 may comprise a set of points (for example, as represented by points in image 1412 of Figure 14).

[0125] In some examples, illuminator 1402 and sensor 1404 may be mounted on a rotating structure so that illuminator 1402 and sensor 1404 capture a 360-degree view of an environment. In other examples, the range measurement system 1400 may include one or more optical components (e.g., mirrors, collimators, diffraction gratings, etc.) that enable illuminator 1402 and sensor 1404 to detect objects within a specific range (e.g., up to 360 degrees). Although the example in Figure 14 shows only a single illuminator 1402 and sensor 1404, the range measurement system 1400 may include multiple sets of illuminators and sensors.

[0126] In some examples, the illuminator 1402 generates a structured light pattern. In such examples, the range measurement system 1400 may include multiple sensors 1404 on which the respective images of the structured light pattern are formed. The range measurement system 1400 may use disparities between the images of the structured light pattern to determine a distance to an object 1408 from which the structured light pattern backscatters. Structured light-based range measurement systems may have a high level of accuracy (e.g., submillimeter range accuracy) when the object 1408 is relatively close to the sensor 1404 (e.g., from 0.2 meters to 2 meters). This high level of accuracy may be useful in facial recognition applications, such as unlocking mobile devices (e.g., mobile phones, tablet computers, etc.) and for security applications.

[0127] In some examples, the 1400 range measurement system is one. Petition 870250085531, dated 09 / 22 / 2025, pp. 304 / 359 45 / 67 time-of-flight (ToF) based system. In some examples where the range measurement system 1400 is a ToF-based system, the illuminator 1402 generates light pulses. In other words, the illuminator 1402 can modulate the amplitude of the emitted light 1406. In these examples, the sensor 1404 detects the return light 1410 from the light pulses 1406 generated by the illuminator 1402. The range measurement system 1400 can then determine a distance to the object 1408 from which the light 1406 backscatters based on a delay between when the light 1406 was emitted and detected and the known speed of light in air. In some examples, instead of (or in addition to) modulating the amplitude of the emitted light 1406, the illuminator 1402 can modulate the phase of the emitted light 1406.In these examples, sensor 1404 can detect the phase of the return light 1410 coming from object 1408 and determine the distances to points on object 1408 using the speed of light and based on the time differences between when illuminator 1402 generated light 1406 at a specific phase and when sensor 1404 detected the return light 1410 at that specific phase.

[0128] In other examples, a point cloud can be generated without using the illuminator 1402. For example, in some examples, the sensor 1404 of the range measurement system 1400 may include two or more optical cameras. In these examples, the range measurement system 1400 may use the optical cameras to capture stereoscopic images of the environment, including the object 1408. The range measurement system 1400 (e.g., the point cloud generator 1420) can then calculate the disparities between locations in the stereoscopic images. The range measurement system 1400 can then use the disparities to determine the distances to the locations shown in the stereoscopic images. From these distances, the point cloud generator 1420 can generate a point cloud.

[0129] Sensors 1404 can also detect other object attributes 1408, such as color and reflectance information. In the example in Figure 14, a point cloud generator 1420 can generate a point cloud based on signals 1418 Petition 870250085531, dated 09 / 22 / 2025, pages 305 / 359 46 / 67 generated by sensor 1404. The range measurement system 1400 and / or the point cloud generator 1420 may be part of data source 104 (Figure 1).

[0130] Figure 15 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of this disclosure may be used. In the example in Figure 15, a vehicle 1500 includes a laser package 1502, as a LIDAR system. The laser package 1502 can be implemented in the same way as the laser package 600 (Figure 13). Although not shown in the example in Figure 15, the vehicle 1500 may also include a data source, such as data source 104 (Figure 1), and a G-PCC encoder, such as G-PCC encoder 200 (Figure 1). In the example in Figure 15, the laser package 1502 emits laser beams 1504 that are reflected by pedestrians 1506 or other objects on a highway. The vehicle data source 1500 can generate a point cloud based on signals generated by the laser package 1502.The vehicle 1500's G-PCC encoder can encode the point cloud to generate 1508 bitstreams, such as the geometry bitstream in Figure 2 and the attribute bitstream in Figure 2. The 1508 bitstreams can include a much smaller number of bits than the unencoded point cloud obtained by the G-PCC encoder. An output interface of the vehicle 1500 (e.g., output interface 108 (Figure 1)) can transmit 1508 bitstreams to one or more other devices. Thus, the vehicle 1500 may be able to transmit 1508 bitstreams to other devices more quickly than unencoded point cloud data. Furthermore, 1508 bitstreams may require less data storage capacity.

[0131] The techniques in this disclosure can further reduce the number of bits in the 1508 bitstreams. For example, by downscaling the point cloud geometry data and then encoding attribute data for the point cloud using the downscaled point cloud geometry data, the amount of attribute data to be encoded can be significantly reduced, thus reducing the number of bits in the 1508 bitstream. By subsequently recoloring a set of decoded geometry data into Petition 870250085531, dated 09 / 22 / 2025, pp. 306 / 359 47 / 67 real scale using the reduced-scale attribute data, the resulting reconstructed point cloud can, however, represent a high-resolution reproduction of the original point cloud after decoding.

[0132] In the example in Figure 15, vehicle 1500 can transmit bitstreams 1508 to another vehicle 1510. Vehicle 1510 may include a G-PCC decoder, such as the G-PCC decoder 300 (Figure 1). The G-PCC decoder in vehicle 1510 can decode the bitstreams 1508 to reconstruct the point cloud. Vehicle 1510 can use the reconstructed point cloud for various purposes. For example, vehicle 1510 can determine, based on the reconstructed point cloud, that there are pedestrians 1506 on the highway ahead of vehicle 1500 and therefore begin to reduce speed, for example, even before a driver of vehicle 1510 notices that there are pedestrians 1506 on the highway. Thus, in some examples, vehicle 1510 can perform an autonomous navigation operation, generate a notification or warning, or perform another action based on the reconstructed point cloud.

[0133] Additionally or alternatively, vehicle 1500 can transmit bitstreams 1508 to a server system 1512. The server system 1512 can use bitstreams 1508 for various purposes. For example, the server system 1512 can store bitstreams 1508 for subsequent reconstruction of point clouds. In this example, the server system 1512 can use the point clouds along with other data (e.g., vehicle telemetry data generated by vehicle 1500) to train an autonomous driving system. In another example, the server system 1512 can store bitstreams 1508 for subsequent reconstruction for forensic collision investigations (e.g., if vehicle 1500 collides with pedestrians 1506).

[0134] Figure 16 is a conceptual diagram illustrating an example extended reality system in which one or more techniques from this disclosure can be used. Extended reality (XR) is a term used to encompass a range of technologies that includes augmented reality (AR), mixed reality (MR), and virtual reality (VR). In the example in Figure 16, a first user Petition 870250085531, dated 09 / 22 / 2025, pages 307 / 359 48 / 67 1600 is located at a first location 1602. User 1600 uses an XR headset 1604. As an alternative to the XR headset 1604, user 1600 can use a mobile device (e.g., a mobile phone, a tablet computer, etc.). The XR headset 1604 includes a depth-sensing sensor, such as a LIDAR system, which detects the point positions on objects 1606 at location 1602. A data source for the XR headset 1604 can use the signals generated by the depth-sensing sensor to generate a point cloud representation of the objects 1606 at location 1602. The XR headset 1604 may include a G-PCC encoder (e.g., the G-PCC encoder 200 of Figure 1) that is configured to encode the point cloud in order to generate bitstreams 1608.

[0135] The techniques in this disclosure can further reduce the number of bits in the 1608 bitstreams. For example, by downscaling the point cloud geometry data and then encoding attribute data for the point cloud using the downscaled point cloud geometry data, the amount of attribute data to be encoded can be significantly reduced, thus reducing the number of bits in the 1608 bitstream. By subsequently recoloring a full-scale decoded geometry dataset using the downscaled attribute data, the resulting reconstructed point cloud can, however, represent a high-resolution reproduction of the original point cloud after decoding.

[0136] The XR 1604 headset can transmit bitstreams 1608 (e.g., via a network such as the Internet) to an XR 1610 headset worn by a user 1612 at a second location 1614. The XR 1610 headset can decode the bitstreams 1608 to reconstruct the point cloud. The XR 1610 headset can use the point cloud to generate an XR visualization (e.g., an AR, MR, or VR visualization) representing the objects 1606 at location 1602. Thus, in some examples, such as when the XR 1610 headset generates a VR visualization, the user 1612 at location 1614 can have an experience Petition 870250085531, dated 09 / 22 / 2025, pages 308 / 359 49 / 67 Immersive three-dimensional view of location 1602. In some examples, the XR 1610 headset can determine the position of a virtual object based on the reconstructed point cloud. For example, the XR 1610 headset can determine, based on the reconstructed point cloud, that an environment (e.g., location 1602) includes a flat surface and then determine that a virtual object (e.g., a cartoon character) should be positioned on the flat surface. The XR 1610 headset can generate an XR view in which the virtual object is in the determined position. For example, the XR 1610 headset can show the cartoon character on the flat surface.

[0137] Figure 17 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure may be used. In the example in Figure 17, a mobile device 1700, such as a mobile phone or tablet computer, includes a depth-sensing sensor, such as a LIDAR system, that detects point positions on objects 1702 in a mobile device 1700 environment. A data source of the mobile device 1700 may use the signals generated by the depth-sensing sensor to generate a point cloud representation of the objects 1702. The mobile device 1700 may include a G-PCC encoder (for example, the G-PCC encoder 200 of Figure 1) that is configured to encode the point cloud in order to generate bitstreams 1704. In the example in Figure 17, the mobile device 1700 may transmit bitstreams to a remote device 1706, such as a server system or another mobile device.Remote device 1706 can decode bitstreams 1704 to reconstruct the point cloud. Remote device 1706 can use the point cloud for various purposes. For example, remote device 1706 can use the point cloud to generate a map of an environment from mobile device 1700. For example, remote device 1706 can generate a map of a building's interior based on the reconstructed point cloud. In another example, remote device 1706 can generate images (e.g., computer graphics) based on the cloud. Petition 870250085531, dated 09 / 22 / 2025, pages 309 / 359 50 / 67 points. For example, the 1706 remote device can use points from the point cloud as polygon vertices and use the point color attributes as the basis for shading the polygons. In some examples, the 1706 remote device can perform facial recognition using the point cloud.

[0138] Figures 18 and 19 are flow diagrams illustrating example deep learning-based geometry encoder and decoder networks. In particular, Figure 18 depicts an example deep learning-based geometry encoder network, and Figure 19 depicts an example deep learning-based geometry decoder network.In the example in Figure 18, the deep learning-based geometry encoding network includes a processing block 1800 including a 16x33 convolutional layer and a rectified linear unit (ReLU); a processing block 1802 including a 32x33 / 2 convolutional layer, a ReLU, and an inception residual block (IRB); a processing block 1804 including a 32x33 / 2 convolutional layer, a first ReLU, a 64x33 / 2 convolutional layer, a second ReLU, and an IRB; and a processing block 1806 including a 64x33 convolutional layer, a first ReLU, a 32x33 / 2 convolutional layer, a second ReLU, a PRB, and an 8x33 convolutional layer.

[0139] In the example in Figure 19, the deep learning-based geometry decoding network includes a processing block 1900 including a transposed convolutional 64x33 / 2 layer, a first ReLU, a convolutional 64x33 layer, a second ReLU, and an IRB; a processing block 1902 including a transposed convolutional 32x33 / 2 layer, a first ReLU, a convolutional 32x33 layer, a second ReLU, and an IRB; a processing block 1904 including a transposed convolutional 16x33 / 2 layer, a first ReLU, a convolutional 16x33 layer, a second ReLU, and an IRB; and classifiers 1906, 1908, and 1910.

[0140] In the examples in Figures 18 and 19, conv Nx33se refers to a 3x3x3 convolutional filter with N channels (where N can be, for example, 8, 16, 32 or Petition 870250085531, dated 09 / 22 / 2025, pp. 310 / 359 51 / 67 affines). IRB corresponds to residual inception block. T-Conv corresponds to transposed convolutions that scale up by 2. 2f corresponds to a 2-step scale up, while 2; corresponds to a 2-step scale down.

[0141] Deep learning-based networks can be trained using loss functions, for example, three loss functions, one loss function at each scale. The 1906, 1908, and 1910 classifiers can use a binary cross-entropy loss as a classification loss to classify each voxel as occupied or unoccupied. For example, the following loss function can be used: ^BCE = 1Σ-(°logpv +(1 - °v)log(1 - p^) (1) V where Ov is the truth of the field as to whether voxel v is occupied (1) or unoccupied (0).

[0142] Along with a reconstruction loss (binary cross-entropy as explained above), a bit rate loss can be used to optimize rate skewness. The network as a whole can be trained with joint reconstruction and bit rate loss in an end-to-end manner. An Adam optimizer can be used with a decayed learning rate from 0.0008 to 0.00001. In the examples in Figures 18 and 19, if the classification loss is replaced by mean squared error (MSE) loss, such networks would form deep learning-based attribute oversampling / undersampling networks. That is, instead of being trained to determine point positions, such networks would be trained in attribute data processing (undersampling / oversampling).

[0143] The following clauses represent several examples of the techniques of this disclosure.

[0144] Clause 1: A device for encoding point cloud data, wherein the device comprises: a memory configured to store point cloud data; and one or more processors implemented in the set of Petition 870250085531, dated 09 / 22 / 2025, pp. 311 / 359 52 / 67 circuits and configured to: decode encoded point cloud geometry data into a point cloud in order to reconstruct point cloud geometry data into the point cloud; downscale point cloud geometry data to form downscaled point cloud geometry data; and encode attribute data into the point cloud using the downscaled point cloud geometry.

[0145] Clause 2: The device of clause 1, wherein, in order to encode the attribute data for the point cloud, one or more processors are configured to encode the attribute data for the point cloud, and wherein the one or more processors are additionally configured to encode the point cloud geometry data using a deep learning-based geometry encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data.

[0146] Clause 3: The device of clause 2, wherein one or more processors are additionally configured to encode a value representing a downscaling amount to be applied to the point cloud geometry data, wherein to downscale the point cloud geometry data, one or more processors are configured to downscale the point cloud geometry data according to the value representing the downscaling amount.

[0147] Clause 4: The device of clause 1, wherein, in order to encode the attribute data for the point cloud, one or more processors are configured to decode the attribute data for the point cloud in order to form scaled-down point cloud attribute data, and wherein one or more processors are additionally configured to scale up the scaled-down point cloud attribute data.

[0148] Clause 5: The device of clause 4, whereby, to increase the scale of the point cloud attribute data on a reduced scale, one or more Petition 870250085531, dated 09 / 22 / 2025, pp. 312 / 359 53 / 67 processors are configured to upscale point cloud attribute data using a deep learning-based attribute supersampler.

[0149] Clause 6: The device of clause 4, wherein, to increase the scale of downscaled point cloud attribute data, one or more processors are configured to apply a 5x5x5 layer of convolutional target coordinates to downscaled point cloud attribute data.

[0150] Clause 7: The device of clause 4, wherein, to increase the scale of the point cloud attribute data on a reduced scale, one or more processors are configured to reconstruct point cloud attribute data on an increased scale for the point cloud, and wherein one or more processors are additionally configured to apply the point cloud attribute data on an increased scale to the point cloud geometry data to reconstruct the point cloud.

[0151] Clause 8: The device of clause 4, wherein one or more processors are additionally configured to decode a value representing a downscaling amount to be applied to point cloud geometry data, wherein to downscale point cloud geometry data, one or more processors are configured to downscale point cloud geometry data according to the value representing the downscaling amount.

[0152] Clause 9: The device of clause 8, wherein, to increase the scale of point cloud attribute data in a scaled-down manner, one or more processors are configured to increase the scale of point cloud attribute data in a scaled-down manner according to the value representing the amount of scaling down to be applied to the point cloud geometry data.

[0153] Clause 10: The provision of clause 8, in which the value comprises a Petition 870250085531, dated 09 / 22 / 2025, pp. 313 / 359 54 / 67 first value, and wherein one or more processors are additionally configured to decode a second value that represents an amount of upscaling to be applied to the downscaled point cloud attribute data, wherein, to upscale the downscaled point cloud attribute data, one or more processors are configured to upscale the downscaled point cloud attribute data according to the second value that represents the amount of upscaling to be applied to the point cloud attribute data.

[0154] Clause 11: The device of clause 1, wherein the attribute data includes color data in one of a red-green-blue (RGB) format or a luminance, blue hue chrominance and red hue chrominance (YCbCr) format.

[0155] Clause 12: The device of clause 1, wherein, to downscale point cloud geometry data, one or more processors are configured to: for each node of an octree that includes eight leaf subnodes where at least one of the eight leaf subnodes is occupied by a point, redefine the node as an occupied leaf node in a downscaled octree; and for each node of the octree that includes eight leaf subnodes where none of the eight leaf subnodes is occupied by a point, redefine the node as an unoccupied leaf node in the downscaled octree.

[0156] Clause 13: The device of clause 1, wherein, to downscale point cloud geometry data, one or more processors are configured to: for each node of an octree that includes eight leaf subnodes where a number of those that are occupied among the eight leaf subnodes is greater than a threshold, redefine the node as an occupied leaf node in a downscaled octree; and for each node of the octree that includes eight leaf subnodes where a number of those that are occupied among the eight leaf subnodes is less than or equal to the threshold, redefine the node as an unoccupied leaf node in the downscaled octree.

[0157] Clause 14: A method for encoding point cloud data, wherein the method comprises: decoding point cloud geometry data Petition 870250085531, dated 09 / 22 / 2025, pp. 314 / 359 55 / 67 encoded to a point cloud in order to reconstruct point cloud geometry data for the point cloud; downscale the point cloud geometry data to form downscaled point cloud geometry data; and encode attribute data to the point cloud using the downscaled point cloud geometry.

[0158] Clause 15: The method of clause 14, wherein the encoding of attribute data for the point cloud comprises encoding the attribute data for the point cloud, wherein the method further comprises encoding the point cloud geometry data using a deep learning-based geometry encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data.

[0159] Clause 16: The method of clause 15, which further comprises encoding a value representing a scaling amount to be applied to the point cloud geometry data, wherein the scaling of the point cloud geometry data comprises scaling down the point cloud geometry data according to the value representing the scaling amount.

[0160] Clause 17: The method of clause 14, wherein the encoding of attribute data for the point cloud comprises decoding the attribute data for the point cloud in order to form scaled-down point cloud attribute data, wherein the method further comprises scaling up the scaled-down point cloud attribute data.

[0161] Clause 18: The method of clause 17, wherein scale-up point cloud attribute data comprises scaling up point cloud attribute data using a deep learning-based attribute supersampler.

[0162] Clause 19: The method of clause 17, in which scaling up point cloud attribute data to a scale down comprises applying a Petition 870250085531, dated 09 / 22 / 2025, pages 315 / 359 56 / 67 layer of 5x5x5 convolutional target coordinates to scaled-down point cloud attribute data.

[0163] Clause 20: The method of clause 17, wherein the scale-up of point cloud attribute data to a reduced scale comprises applying one or more of a transposed convolutional layer, a deconvolutional layer, or a degrouping layer to the point cloud attribute data to a reduced scale.

[0164] Clause 21: The method in clause 17, wherein the scale-up of point cloud attribute data comprises reconstructing scale-up point cloud attribute data for the point cloud, wherein the method further comprises applying scale-up point cloud attribute data to point cloud geometry data to reconstruct the point cloud.

[0165] Clause 22: The method of clause 17, which further comprises decoding a value representing a scaling amount to be applied to the point cloud geometry data, wherein the scaling of the point cloud geometry data comprises scaling down the point cloud geometry data according to the value representing the scaling amount.

[0166] Clause 23: The method of clause 22, wherein the scale-up of point cloud attribute data at scale-down comprises scaling up point cloud attribute data at scale-down according to the value representing the amount of scale-down to be applied to point cloud geometry data.

[0167] Clause 24: The method of clause 22, wherein the value comprises a first value, wherein the method further comprises decoding a second value representing a scaling amount to be applied to the scaled-down point cloud attribute data, wherein scaling down the scaled-down point cloud attribute data Petition 870250085531, dated 09 / 22 / 2025, pp. 316 / 359 57 / 67 involves increasing the scale of the point cloud attribute data by a smaller scale according to the second value, which represents the amount of scaling to be applied to the point cloud attribute data.

[0168] Clause 25: The method of clause 14, wherein the point cloud geometry data represent points in a three-dimensional space for the point cloud, and wherein the attribute data represent attributes of the points.

[0169] Clause 26: The method of clause 14, wherein the attribute data include color data in one of a red-green-blue (RGB) format or a luminance, blue hue chrominance, and red hue chrominance (YCbCr) format.

[0170] Clause 27: The method of clause 14, wherein the downscaling of point cloud geometry data comprises: for each node of an octree that includes eight leaf subnodes where at least one of the eight leaf subnodes is occupied by a point, redefine the node as an occupied leaf node in a downscaled octree; and for each node of the octree that includes eight leaf subnodes where none of the eight leaf subnodes is occupied by a point, redefine the node as an unoccupied leaf node in the downscaled octree.

[0171] Clause 28: The method of clause 14, in which the downscaling of point cloud geometry data comprises: for each node of an octree that includes eight leaf subnodes where a number of those that are occupied among the eight leaf subnodes is greater than a threshold, redefine the node as an occupied leaf node in a downscaled octree; and for each node of the octree that includes eight leaf subnodes where a number of those that are occupied among the eight leaf subnodes is less than or equal to the threshold, redefine the node as an unoccupied leaf node in the downscaled octree.

[0172] Clause 29: A computer-readable storage medium having stored therein instructions which, when executed, cause a processor to: decode encoded point cloud geometry data into a point cloud in order to reconstruct point cloud geometry data into the point cloud; downscale the point cloud geometry data Petition 870250085531, dated 09 / 22 / 2025, pp. 317 / 359 58 / 67 points to form scaled-down point cloud geometry data; and encode attribute data for the point cloud using scaled-down point cloud geometry.

[0173] Clause 30: A device for encoding point cloud data, wherein the device comprises: means for decoding encoded point cloud geometry data to a point cloud in order to reconstruct point cloud geometry data for the point cloud; means for downscaling point cloud geometry data to form downscaled point cloud geometry data; and means for encoding attribute data to the point cloud using the downscaled point cloud geometry.

[0174] Clause 31: A device for encoding point cloud data, wherein the device comprises: a memory configured to store point cloud data; and one or more processors implemented in the circuit suite and configured to: decode encoded point cloud geometry data for a point cloud in order to reconstruct point cloud geometry data for the point cloud; downscale point cloud geometry data to form downscaled point cloud geometry data; and encode attribute data for the point cloud using the downscaled point cloud geometry.

[0175] Clause 32: The device of clause 31, wherein, in order to encode the attribute data for the point cloud, one or more processors are configured to encode the attribute data for the point cloud, and wherein one or more processors are additionally configured to encode the point cloud geometry data using a deep learning-based geometry encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data.

[0176] Clause 33: The device of clause 32, in which one or more processors are additionally configured to encode a value representing a downscaling amount to be applied to the data of Petition 870250085531, dated 09 / 22 / 2025, pp. 318 / 359 59 / 67 point cloud geometry, where to downscale point cloud geometry data, one or more processors are configured to downscale point cloud geometry data according to the value representing the amount of downscaling.

[0177] Clause 34: The device of clause 31, wherein, in order to encode the attribute data for the point cloud, one or more processors are configured to decode the attribute data for the point cloud in order to form scaled-down point cloud attribute data, and wherein one or more processors are additionally configured to scale up the scaled-down point cloud attribute data.

[0178] Clause 35: The device of clause 34, wherein, to scale down point cloud attribute data, one or more processors are configured to scale down point cloud attribute data using a deep learning-based attribute supersampler.

[0179] Clause 36: The device of either clause 34 and 35, wherein, to increase the scale of downscaled point cloud attribute data, one or more processors are configured to apply a 5x5x5 layer of convolutional target coordinates to downscaled point cloud attribute data.

[0180] Clause 37: The device of any of clauses 34 to 36, wherein, to increase the scale of the scaled-down point cloud attribute data, one or more processors are configured to reconstruct scaled-up point cloud attribute data for the point cloud, and wherein one or more processors are additionally configured to apply the scaled-up point cloud attribute data to the point cloud geometry data to reconstruct the point cloud.

[0181] Clause 38: The device of any of clauses 34 to 37, in which one or more processors are additionally configured to decode Petition 870250085531, dated 09 / 22 / 2025, pp. 319 / 359 60 / 67 is a value that represents a scaling amount to be applied to point cloud geometry data, where to scale down point cloud geometry data, one or more processors are configured to scale down point cloud geometry data according to the value representing the scaling amount.

[0182] Clause 39: The device of clause 38, wherein, to increase the scale of point cloud attribute data in a scale-down manner, one or more processors are configured to increase the scale of point cloud attribute data in a scale-down manner according to the value representing the amount of scale-down to be applied to the point cloud geometry data.

[0183] Clause 40: The device of clause 38, wherein the value comprises a first value, and wherein the one or more processors are additionally configured to decode a second value representing a scaling amount to be applied to the scaled-down point cloud attribute data, wherein, to scale down the scaled-down point cloud attribute data, the one or more processors are configured to scale down the scaled-down point cloud attribute data according to the second value representing the scaling amount to be applied to the point cloud attribute data.

[0184] Clause 41: The device of any of clauses 31 to 40, wherein the attribute data includes color data in one of a red-green-blue (RGB) format or a luminance, blue hue chrominance and red hue chrominance (YCbCr) format.

[0185] Clause 42: The device of any of clauses 31 to 41, wherein, to downscale point cloud geometry data, one or more processors are configured to: for each node of an octree that includes eight leaf subnodes where at least one of the eight leaf subnodes is occupied by a point, redefine the node as an occupied leaf node in a downscaled octree; Petition 870250085531, dated 09 / 22 / 2025, pages 320 / 359 61 / 67 and for each node of the octree that includes eight leaf subnodes where none of the eight leaf subnodes is occupied by a point, redefine the node as an unoccupied leaf node in the scaled-down octree.

[0186] Clause 43: The device of any of clauses 31 to 41, wherein, to downscale point cloud geometry data, one or more processors are configured to: for each node of an octree that includes eight leaf subnodes where a number of those that are occupied among the eight leaf subnodes is greater than a threshold, redefine the node as an occupied leaf node in a downscaled octree; and for each node of the octree that includes eight leaf subnodes where a number of those that are occupied among the eight leaf subnodes is less than or equal to the threshold, redefine the node as an unoccupied leaf node in the downscaled octree.

[0187] Clause 44: A method for encoding point cloud data, wherein the method comprises: decoding encoded point cloud geometry data into a point cloud in order to reconstruct point cloud geometry data for the point cloud; downscaling the point cloud geometry data to form downscaled point cloud geometry data; and encoding attribute data into the point cloud using the downscaled point cloud geometry.

[0188] Clause 45: The method of clause 44, wherein the encoding of attribute data for the point cloud comprises encoding the attribute data for the point cloud, wherein the method further comprises encoding the point cloud geometry data using a deep learning-based geometry encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data.

[0189] Clause 46: The method of clause 45, which further comprises encoding a value representing a scaling amount to be applied to the point cloud geometry data, wherein scaling of the point cloud geometry data comprises scaling down Petition 870250085531, dated 09 / 22 / 2025, pp. 321 / 359 62 / 67 of the point cloud geometry data according to the value that represents the amount of scale reduction.

[0190] Clause 47: The method of clause 44, wherein the encoding of attribute data for the point cloud comprises decoding the attribute data for the point cloud in order to form scaled-down point cloud attribute data, wherein the method further comprises scaling up the scaled-down point cloud attribute data.

[0191] Clause 48: The method of clause 47, wherein scale-up point cloud attribute data comprises scaling up point cloud attribute data using a deep learning-based attribute supersampler.

[0192] Clause 49: The method of either clause 47 and 48, wherein the scale-up of downscaled point cloud attribute data comprises applying a 5x5x5 layer of convolutional target coordinates to downscaled point cloud attribute data.

[0193] Clause 50: The method of any of clauses 47 to 49, wherein the scaling up of downscaled point cloud attribute data comprises applying one or more of a transposed convolutional layer, a deconvolutional layer, or a degrouping layer to downscaled point cloud attribute data.

[0194] Clause 51: The method of any of clauses 47 to 50, wherein scaling up point cloud attribute data at scale down comprises reconstructing point cloud attribute data at scale up for the point cloud, wherein the method further comprises applying point cloud attribute data at scale up to point cloud geometry data to reconstruct the point cloud.

[0195] Clause 52: The method of any of clauses 47 to 51, which additionally comprises decoding a value representing a scaling amount to be applied to point cloud geometry data, Petition 870250085531, dated 09 / 22 / 2025, pp. 322 / 359 63 / 67 where scaling down point cloud geometry data involves scaling down point cloud geometry data according to the value representing the amount of scaling down.

[0196] Clause 53: The method of clause 52, wherein the scale-up of point cloud attribute data on a scale-down scale comprises scaling up point cloud attribute data on a scale-down scale according to the value representing the amount of scale-down to be applied to the point cloud geometry data.

[0197] Clause 54: The method of clause 52, wherein the value comprises a first value, wherein the method further comprises decoding a second value representing a scaling amount to be applied to the scaled-down point cloud attribute data, wherein scaling down the point cloud attribute data comprises scaling down the point cloud attribute data according to the second value representing the scaling amount to be applied to the point cloud attribute data.

[0198] Clause 55: The method of any of clauses 44 to 54, wherein the point cloud geometry data represent points in a three-dimensional space for the point cloud, and wherein the attribute data represent attributes of the points.

[0199] Clause 56: The method of any of clauses 44 to 55, wherein the attribute data include color data in one of a red-green-blue (RGB) format or a luminance, blue hue chrominance, and red hue chrominance (YCbCr) format.

[0200] Clause 57: The method of any of clauses 44 to 56, wherein the downscaling of point cloud geometry data comprises: for each node of an octree that includes eight leaf subnodes where at least one of the eight leaf subnodes is occupied by a point, redefine the node as an occupied leaf node in a downscaled octree; and for each node of the octree that includes eight subnodes Petition 870250085531, dated 09 / 22 / 2025, pp. 323 / 359 64 / 67 of a leaf where none of the eight leaf subnodes are occupied by a point, redefine the node as an unoccupied leaf node in the scaled-down octree.

[0201] Clause 58: The method of any of clauses 44 to 56, wherein the downscaling of point cloud geometry data comprises: for each node of an octree that includes eight leaf subnodes where a number of those that are occupied among the eight leaf subnodes is greater than a threshold, redefine the node as an occupied leaf node in a downscaled octree; and for each node of the octree that includes eight leaf subnodes where a number of those that are occupied among the eight leaf subnodes is less than or equal to the threshold, redefine the node as an unoccupied leaf node in the downscaled octree.

[0202] Clause 59: A computer-readable storage medium having stored therein instructions which, when executed, cause a processor to: decode encoded point cloud geometry data for a point cloud in order to reconstruct point cloud geometry data for the point cloud; downscale point cloud geometry data to form downscaled point cloud geometry data; and encode attribute data for the point cloud using the downscaled point cloud geometry.

[0203] Clause 60: A device for encoding point cloud data, wherein the device comprises: means for decoding encoded point cloud geometry data to a point cloud in order to reconstruct point cloud geometry data for the point cloud; means for downscaling point cloud geometry data to form downscaled point cloud geometry data; and means for encoding attribute data to the point cloud using the downscaled point cloud geometry.

[0204] It should be recognized that, depending on the example, certain actions or events of any of the techniques described in the present invention may be performed in a different sequence, may be added, combined or completely omitted (for example, not all actions or events described Petition 870250085531, dated 09 / 22 / 2025, pp. 324 / 359 65 / 67 are needed for the practice of the techniques). Furthermore, in certain examples, actions or events can be performed simultaneously, for example, through multi-threaded processing, interrupt processing, or multiple processors, instead of sequentially.

[0205] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored or transmitted as one or more instructions or code in a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to a tangible medium, such as data storage media, or communication media, including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. In this way, computer-readable media may generally correspond to (1) tangible computer-readable storage media that are non-transient or (2) a communication medium, such as a signal or a carrier wave.Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0206] By way of example, and not limitation, such computer-readable storage media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory or any other media that Petition 870250085531, dated 09 / 22 / 2025, pages 325 / 359 66 / 67 can be used to store the desired program code in the form of instructions or data structures that can be accessed by a computer. Furthermore, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave will be included in the definition of medium. It should be understood, however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead refer to tangible non-transient storage media.As used in the present invention, disks (disk and disc) include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically by means of lasers. Combinations of the above items should also be included within the scope of computer-readable media.

[0207] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other sets of distinct or equivalent integrated logic circuits. Consequently, the terms processor and processing circuit set, as used in the present invention, may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described in the present invention. Furthermore, in some respects, the functionality described in the present invention may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a codec. Petition 870250085531, dated 09 / 22 / 2025, pages 326 / 359 67 / 67 combined. Furthermore, the techniques can be fully implemented in one or more circuits or logic elements.

[0208] The techniques in this disclosure can be implemented in a wide variety of devices or appliances, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Several components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require implementation by different hardware units. Instead, as described above, several units can be combined into a codec hardware unit or provided by a collection of interoperable hardware units, including one or more processors, as described above, in conjunction with suitable software and / or firmware.

[0209] Several examples have been described. These and other examples are within the scope of the following claims. Petition 870250085531, dated 09 / 22 / 2025, pp. 327 / 359

Claims

1 / 6 CLAIMS 1. Device for encoding point cloud data, the device being characterized by comprising: a memory configured to store point cloud data; and one or more processors implemented in the circuit assembly and configured to: decode encoded point cloud geometry data for a point cloud in order to reconstruct point cloud geometry data for the point cloud; downscale the point cloud geometry data to form downscaled point cloud geometry data; and encode attribute data for the point cloud using the downscaled point cloud geometry data.

2. Device according to claim 1, characterized in that, to encode attribute data for the point cloud, one or more processors are configured to encode attribute data for the point cloud, and wherein the one or more processors are additionally configured to encode point cloud geometry data using a deep learning-based geometry encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data.

3. Device according to claim 2, characterized in that one or more processors are additionally configured to encode a value representing a downscaling amount to be applied to point cloud geometry data, wherein to downscale the point cloud geometry data, the one or more processors are configured to downscale the point cloud geometry data according to the value representing the downscaling amount. Petition 870250085531, dated 09 / 22 / 2025, pp. 354 / 359 2 / 6 4. Device according to claim 1, characterized in that, to encode attribute data for the point cloud, one or more processors are configured to decode the attribute data for the point cloud in order to form scaled-down point cloud attribute data, and wherein one or more processors are additionally configured to scale up the scaled-down point cloud attribute data.

5. Device according to claim 4, characterized in that, to increase the scale of downscaled point cloud attribute data, one or more processors are configured to increase the scale of downscaled point cloud attribute data using a deep learning-based attribute supersampler.

6. Device according to claim 4, characterized in that, to increase the scale of downscaled point cloud attribute data, one or more processors are configured to apply a 5x5x5 layer of convolutional target coordinates to downscaled point cloud attribute data.

7. Device according to claim 4, characterized in that to increase the scale of point cloud attribute data on a reduced scale, one or more processors are configured to reconstruct point cloud attribute data on an increased scale for the point cloud, and wherein one or more processors are additionally configured to apply the point cloud attribute data on an increased scale to the point cloud geometry data to reconstruct the point cloud.

8. Device, according to claim 4, characterized in that one or more processors are additionally configured to decode a value representing a downscaling amount to be applied to point cloud geometry data, wherein to downscale the point cloud geometry data, the one or more processors are configured to downscale the point cloud geometry data according to the value representing the downscaling amount.

9. Device according to claim 8, characterized in that, to increase the scale of point cloud attribute data in a scale-down manner, one or more processors are configured to increase the scale of point cloud attribute data in a scale-down manner according to the value representing the amount of scale reduction to be applied to the point cloud geometry data.

10. Device according to claim 8, characterized in that the value comprises a first value, and wherein one or more processors are additionally configured to decode a second value representing a scaling amount to be applied to the scaled-down point cloud attribute data, wherein, to scale down the scaled-down point cloud attribute data, one or more processors are configured to scale down the scaled-down point cloud attribute data according to the second value representing the scaling amount to be applied to the point cloud attribute data.

11. Device according to claim 1, characterized in that, to downscale point cloud geometry data, one or more processors are configured to: for each node of an octree that includes eight leaf subnodes where at least one of the eight leaf subnodes is occupied by a point, redefine the node as an occupied leaf node in a downscaled octree; and for each node of the octree that includes eight leaf subnodes where none of the eight leaf subnodes is occupied by a point, redefine the node as an unoccupied leaf node in the downscaled octree.

12. Method for encoding point cloud data, the method being characterized by comprising: Petition 870250085531, dated 09 / 22 / 2025, pp. 356 / 359 4 / 6 decoding encoded point cloud geometry data into a point cloud in order to reconstruct point cloud geometry data for the point cloud; downscaling the point cloud geometry data to form downscaled point cloud geometry data; and encoding attribute data into the point cloud using the downscaled point cloud geometry.

13. A method according to claim 12, wherein the encoding of attribute data for the point cloud comprises encoding the attribute data for the point cloud, the method being characterized by further comprising: encoding the point cloud geometry data using a deep learning-based geometry encoder to form the encoded point cloud geometry data before decoding the encoded point cloud geometry data; and encoding a value representing a scaling amount to be applied to the point cloud geometry data, wherein the scaling of the point cloud geometry data comprises scaling down the point cloud geometry data according to the value representing the scaling amount.

14. A method according to claim 12, wherein the encoding of attribute data for the point cloud comprises decoding the attribute data for the point cloud in order to form scaled-down point cloud attribute data, the method being characterized by further comprising: scaling up the scaled-down point cloud attribute data, wherein the scaling up of the scaled-down point cloud attribute data comprises scaling up the scaled-down point cloud attribute data using a deep learning-based attribute supersampler. Petition 870250085531, dated 22 / 09 / 2025, pp. 357 / 359 5 / 6 15. A method according to claim 14, characterized in that the downscaling of point cloud attribute data comprises applying at least one of a 5x5x5 convolutional target coordinate layer, a transposed convolutional layer, a deconvolutional layer, or a degrouping layer to the downscaling point cloud attribute data.

16. A method according to claim 14, characterized in that scaling up point cloud attribute data to scale down comprises reconstructing scale up point cloud attribute data for the point cloud, wherein the method further comprises applying scale up point cloud attribute data to point cloud geometry data to reconstruct the point cloud.

17. A method according to claim 14, characterized by further comprising decoding a value representing a scaling amount to be applied to point cloud geometry data, wherein the scaling of point cloud geometry data comprises scaling down the point cloud geometry data according to the value representing the scaling amount.

18. Method, according to claim 17, characterized by the scaling up of point cloud attribute data at a reduced scale comprising scaling up point cloud attribute data at a reduced scale according to the value representing the amount of scaling down to be applied to the point cloud geometry data.

19. Method according to claim 17, wherein the value comprises a first value, the method being characterized by further comprising decoding a second value representing an amount of scaling to be applied to the scaled-down point cloud attribute data, wherein scaling down the point cloud attribute data comprises scaling down the point cloud attribute data according to the second value representing the amount of scaling to be applied to the point cloud attribute data.

20. A method according to claim 12, characterized by the downscaling of point cloud geometry data, comprises: for each node of an octree that includes eight leaf subnodes where at least one of the eight leaf subnodes is occupied by a point, redefining the node as an occupied leaf node in a downscaled octree; and for each node of the octree that includes eight leaf subnodes where none of the eight leaf subnodes is occupied by a point, redefining the node as an unoccupied leaf node in the downscaled octree.

21. Product, process, system, kit, means or use, characterized by comprising one or more elements described in the descriptive report, claims, drawings, sequence listing, or summary of this application, when applicable. Petition 870250085531, dated 22 / 09 / 2025, pp. 359 / 359