Variable code rate compression of point cloud geometry

By employing variable bitrate compression technology and deep neural networks in point cloud compression, the artifact problem in point cloud reconstruction is solved, achieving high-quality point cloud reconstruction that meets user-defined quality and bandwidth requirements.

CN120917487APending Publication Date: 2025-11-07SONY GROUP CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480024625.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-19
Filing Date
2024-04-04
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing point cloud compression methods are prone to artifacts when reconstructing point cloud geometry and do not support original point encoding modes or region-of-interest encoding modes, resulting in poor reconstruction quality.

Method used

By employing variable bitrate compression technology, switching between different RD operation points, combining locally applied maximum allowable cost thresholds and region-of-interest encoding, and using deep neural networks to encode point clouds, lossless encoding of original points is allowed, thus improving the quality of reconstructed visuals.

Benefits of technology

It effectively reduces reconstruction artifacts, improves the visual quality of point cloud reconstruction, and meets user-defined quality requirements and network bandwidth constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120917487A_ABST
    Figure CN120917487A_ABST
Patent Text Reader

Abstract

An electronic device and method for variable code rate compression of point cloud geometry are provided. An electronic device stores a set of RD operating points and a coding mode associated with the set of RD operating points. An electronic device receives a 3D point cloud geometry and segments the 3D geometry into a set of blocks. After partitioning, the electronic device selects a block and calculates a set of loss values associated with one or more compression indicators. Such loss values correspond to a set of coding modes associated with at least a subset of the set of RD operating points. The electronic device selects a coding mode from the set of coding modes, a loss value of the coding mode in the set of loss values being lower than a loss threshold of the coding mode. Thereafter, the electronic device encodes the block based on the encoding mode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references / incorporation of related applications

[0002] This application claims priority to U.S. Patent Application No. 18 / 303056, filed April 19, 2023, with the U.S. Patent and Trademark Office. All of the above-cited applications are hereby incorporated herein by reference in their entirety. Technical Field

[0003] Various embodiments of this disclosure relate to three-dimensional (3D) point cloud compression (PCC). More specifically, various embodiments of this disclosure relate to variable bitrate compression of point cloud geometry. Background Technology

[0004] Advances in 3D scanning have enabled the creation of high-fidelity 3D geometric representations of 3D objects. 3D point clouds are an example of such 3D geometric representations and have been adopted in various applications, such as free-viewpoint displays for sports or live event broadcasts, geographic information systems, cultural heritage representation, or autonomous vehicle navigation. Typically, point clouds consist of a large number of unstructured 3D points (e.g., each point has X, Y, and Z coordinates) along with associated attributes, such as textures including color or reflectivity. When compressing 3D point clouds, it is best to preserve most of the attributes associated with the 3D point cloud to facilitate efficient reconstruction of the point cloud from the encoded point cloud data. Therefore, an efficient point cloud compression (PCC) method may be required. For geometric compression, conventional PCC methods may only provide a limited number of operation points. When applied to point cloud geometry, this approach can lead to artifacts in the reconstructed point cloud geometry.

[0005] As described in the remainder of this application and with reference to the accompanying drawings, the limitations and disadvantages of conventional methods will become apparent to those skilled in the art by comparing the system with certain aspects of this disclosure. Summary of the Invention

[0006] An electronic device and method for variable bit rate compression of point cloud geometry are provided, as set forth more fully in the claims, and substantially as illustrated in at least one figure and / or described in conjunction with at least one figure.

[0007] These and other features and advantages of this disclosure can be understood by careful study of the following detailed description and the accompanying drawings, in which the same reference numerals always refer to the same parts. Attached Figure Description

[0008] Figure 1 This is a block diagram illustrating an exemplary environment for variable bitrate compression of point cloud geometry according to embodiments of the present disclosure.

[0009] Figure 2is an exemplary encoder and an exemplary decoder for variable bit rate compression of point cloud geometry according to embodiments of the present disclosure. Figure 1 a block diagram of an exemplary electronic device of

[0010] Figure 3 is a block diagram of an exemplary encoder and an exemplary decoder for variable bit rate compression of point cloud geometry according to embodiments of the present disclosure.

[0011] Figure 4 is a diagram illustrating exemplary components of circuitry of Figure 2

[0012] Figure 5 is a diagram illustrating an exemplary processing pipeline for variable bit rate compression of point cloud geometry according to embodiments of the present disclosure.

[0013] Figure 6 is a diagram illustrating an exemplary search approach for modes of RD operating points according to embodiments of the present disclosure.

[0014] Figure 7 is a diagram illustrating an exemplary comparison between lossy and lossless reconstructed outputs of point cloud geometry according to embodiments of the present disclosure.

[0015] Figure 8 is a diagram illustrating exemplary 3D point cloud geometry structure and selection of regions of interest (RoI) in point cloud geometry for point cloud compression according to embodiments of the present disclosure.

[0016] Figure 9 is a flowchart diagram illustrating exemplary operations for variable bit rate compression of point cloud geometry according to embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] ​Implementations described below can be found in electronic devices and methods of variable bitrate compression of disclosed point cloud geometry. Exemplary aspects of the present disclosure provide an electronic device that can include a memory configured to store a set of rate-distortion (RD) operating points and one or more encoding modes associated with each RD operating point of the set of RD operating points. The electronic device can also include circuitry that can be configured to receive three-dimensional (3D) point cloud geometry relating to one or more objects in a 3D space. The electronic device can also be configured to partition the 3D point cloud geometry into a set of blocks and select a first block from the set of blocks. For the selected first block, the electronic device can also be configured to compute a first set of loss values associated with one or more compression metrics. The set of loss values can correspond to a set of encoding modes associated with at least a subset of the set of RD operating points. The electronic device can also be configured to select an encoding mode from the set of encoding modes for which a loss value of the first set of loss values is below a loss threshold for the encoding mode. Thereafter, the electronic device can encode the selected first block based on the selected encoding mode.

[0018] In conventional point cloud compression (PCC) methods, an easy approach is to select a mode to compress point cloud geometry by using a full search algorithm that checks which mode produces the best rate-distortion (RD) performance. One possible implementation of adaptive block-based point cloud geometry compression is based on machine learning. Several models can be trained, each of which is tuned for learned local point cloud features. RD control can be achieved through implicit and explicit quantization. In these conventional methods, unpredictable local artifacts can occur in the locally reconstructed geometry. Such artifacts can result from strong nonlinearities introduced by machine learning-based point cloud compression schemes. Additionally, these conventional methods can not support raw point encoding modes or region-of-interest based encoding modes. In certain existing methods (e.g., adaptive deep learning PCC (ADL-PCC)), given a particular RD operating point, a model (i.e., an encoding mode) can return the minimum cost (e.g., rate-distortion, bit rate, or a combination of both). In the present disclosure, it is generally considered that the minimum cost in all models given an RD point can be too high (above a threshold). Thus, the electronic device of the present disclosure allows the encoder to switch between different RD operating points to search for a model (i.e., a mode) that satisfies a cost constraint. In a more general formulation, a bidirectional search of all modes at RD points can enable the encoder to satisfy any user-defined constraint or quality requirement (e.g., fewer reconstruction artifacts). The present disclosure allows the use of a combination of different RD modes to encode different parts of a point cloud and improves the visual quality of the reconstructed point cloud by locally applying a maximum allowed cost threshold. The present disclosure also proposes a lossless mode to encode raw points that can be enabled if all available modes cannot satisfy the maximum cost criterion. Alternatively, the present disclosure allows region-of-interest based encoding of different slices of point cloud geometry.

[0019] Figure 1 is a block diagram illustrating an exemplary environment for variable bit rate compression of point cloud geometry in accordance with embodiments of the present disclosure. Referring to Figure 1 , a network environment 100 is shown. The network environment 100 can include an electronic device 102, a scanning setup 104, a server 106, a database 108, and a computing device 110. The scanning setup 104 can include one or more image sensors (not shown) and one or more depth sensors (not shown) associated with the one or more image sensors. The electronic device 102 can be communicatively coupled to the scanning setup 104, the server 106, and the computing device 110 via a communication network 112. Also shown in the figure is a three-dimensional (3D) point cloud geometry 114 of a 3D point cloud associated with at least one object (e.g., a person) in a 3D space.

[0020] The electronic device 102 can include suitable logic, circuitry, interfaces, and / or code that can be configured to encode and / or decode 3D point cloud geometry (e.g., 3D point cloud geometry 114). A 3D point cloud can include a collection of data points in a space. These points can represent a 3D shape of an object such that each point position corresponds to a set of Cartesian coordinates. For example, points can be represented as (x, y, z, r, g, b, a), where (x, y, z) represents the 3D coordinates of a point on an object, (r, g, b) represents the red, green, and blue values of the point, and (a) can represent the transparency value of the point.

[0021] In certain embodiments, the electronic device 102 can be configured to generate a 3D point cloud of an object or a plurality of objects (e.g., a 3D scene including objects in the foreground and background). The electronic device 102 can obtain a 3D point cloud geometry 114 of the object (or the plurality of objects) from the 3D point cloud. Examples of the electronic device 102 can include, but are not limited to, a computing device, a video conferencing system, an augmented reality (AR) device, a virtual reality (VR device), a mixed reality (MR) device, a gaming console, a smart wearable device, a mainframe, a server, a computer workstation, and / or a consumer electronics (CE) device.

[0022] The scanning setup 104 can include suitable logic, circuitry, interfaces, and / or code that can be configured to scan a 3D environment including an object to generate a raw 3D scan (also referred to as a raw 3D point cloud). According to an embodiment, the scanning setup 104 can include a single image capture device or a plurality of image capture devices (arranged at a plurality of viewpoints) to capture a plurality of color images. In certain cases, additional depth sensors can be included in the scanning setup 104 to capture depth information of the object. The plurality of color images and the depth information of the object can be captured from different viewpoints. In such cases, a 3D point cloud can be generated based on the captured plurality of color images and the corresponding depth information of the object.

[0023] The scanning setup 104 can be configured to perform a 3D scan of an object in a 3D space and generate a dynamic 3D point cloud (i.e., a sequence of point clouds) that can capture different attributes and geometric deformations of 3D points at different time steps. The scanning setup 104 can be configured to transmit the generated 3D point cloud, the plurality of color images, and / or the corresponding depth information to the electronic device 102 and / or a server via the communication network 112.

[0024] According to one embodiment, the scanning setup 104 can include a combination of multiple sensors, such as a depth sensor, a color sensor (e.g., a red-green-blue (RGB) sensor), and / or a combination of an infrared (IR) projector and an IR sensor. For example, the depth sensor can capture information associated with point cloud geometry (3D positions of points), while the RGB and IR sensors can capture information associated with point cloud attributes (e.g., color and temperature). In one embodiment, the IR projector and the IR sensor can be used to estimate depth information. The combination of the depth sensor, the RGB sensor, and the IR sensor can be used to capture a point cloud frame (a single static point cloud) or multiple point cloud frames (3D video) with associated geometry and attributes.

[0025] According to one embodiment, the scanning setup 104 can include an active 3D scanner that relies on radiation or light to capture 3D structure of an object in a 3D space. Further, the scanning setup 104 can include an image sensor that can capture color information associated with the object. For example, the active 3D scanner can be a time-of-flight (TOF)-based 3D laser scanner, a laser rangefinder, a TOF camera, a handheld laser scanner, a structured light 3D scanner, a modulated light 3D scanner, a CT scanner that outputs point cloud data, an aerial light detection and ranging (LiDAR, laser radar) scanner, a 3D LiDAR, a 3D motion sensor, and the like.

[0026] In Figure 1 the scanning setup 104 is shown as being separate from the electronic device 102. However, in certain embodiments, the scanning setup 104 can be integrated into the electronic device 102. In alternative embodiments, the entire functionality of the scanning setup 104 can be incorporated into the electronic device 102 without departing from the scope of the present disclosure. Examples of the scanning setup 104 can include, but are not limited to, a depth sensor, a RGB sensor, an IR sensor, an image sensor, a light cage with a camera, and / or a motion detector device.

[0027] The server 106 can include suitable logic, circuitry, interfaces, and / or code that can be configured to perform operations, such as data / file storage, 3D rendering or 3D reconstruction operations (e.g., photogrammetry reconstruction operations) to generate a 3D point cloud of an object. As an example and not by way of limitation, the 3D reconstruction operations can be performed by using photogrammetry-based methods (e.g., Structure from Motion (SfM)), methods that require stereo images, or methods that require monocular cues (e.g., Shape from Shading (SfS), photometric stereo, or Shape from Texture (SfT)). For brevity, details of such methods are omitted from the present disclosure. Examples of the server 106 can include, but are not limited to, an application server, a cloud server, a Web server, a database server, a file server, a game server, a mainframe server, or a combination thereof.

[0028] The database 108 can include suitable logic, circuitry, interfaces, and / or code that can be configured to store point cloud geometry. The database 108 can be sourced from a relational or non-relational database, or a set of comma separated value (csv) files in a regular storage or big data storage. The database 108 can be stored or cached on a device such as the server 106. In certain embodiments, the database 108 can be hosted on multiple servers stored at the same or different locations. The operations of the database 108 can be performed using hardware including a processor, microprocessor (e.g., to perform or control performance of one or more operations), field programmable gate array (FPGA), or application specific integrated circuit (ASIC). In other cases, the database 108 can be implemented using software.

[0029] The computing device 110 can include suitable logic, circuitry, interfaces, and / or code that can be configured to communicate with the electronic device 102 and / or the server via the communication network 112. According to one embodiment, the computing device 110 can include a memory configured to store a set of rate-distortion (RD) operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points. The computing device 110 can be configured to receive an encoded 3D point cloud geometry (e.g., as part of multimedia content) from the electronic device 102. The computing device 110 can be configured to decode the encoded 3D point cloud geometry to render a 3D model of an object. Examples of the computing device 110 can include, but are not limited to, a desktop, a personal computer, a notebook, a computer workstation, a tablet computing device, a smartphone, a cellular phone, a mobile phone, a consumer electronic (CE) device with a display, a television (TV), a wearable display, a head-mounted display, a digital signage, a digital mirror (or a smart mirror) capable of storing or rendering multimedia content.

[0030] According to another embodiment, the computing device 110 can be configured to receive input from a user to determine a portion of the 3D point cloud geometry as a region of interest (ROI). The user input can be combined with an object detection operation or a semantic segmentation operation to determine the ROI, i.e., to determine a portion of the 3D point cloud as the ROI.

[0031] In operation, the electronic device 102 can receive a 3D point cloud geometry associated with at least one object in a 3D space (e.g., the 3D point cloud geometry 114). For example, the 3D point cloud data can be obtained from a 3D point cloud (or 3D scan) that includes geometry and various attributes. The 3D point cloud can be a static point cloud or a frame of a dynamic point cloud (i.e., a sequence of point clouds). Generally, the 3D point cloud is a representation of the geometry information (e.g., 3D coordinates of points) and attribute information of an object in a 3D space. The attribute information can include, for example, color information, reflectance information, opacity information, normal vector information, material identifier information, or texture information associated with the object in the 3D space. The texture information can represent a spatial arrangement of colors or intensities in a plurality of color images of the object. The reflectance information can represent information associated with an empirical model of local illumination at points of the 3D point cloud, such as a Phong shading model or a Gouraud shading model. The empirical model of local illumination can correspond to reflectance (rough or shiny surface portions) on the surface of the object. The opacity information can represent transparency of points. The normal vector information can represent a direction perpendicular to a tangent plane at a point in the point cloud.

[0032] After receiving, the electronic device 102 can generate a plurality of voxels from the 3D point cloud geometry 114. The generation of voxels can be referred to as voxelization of the 3D point cloud geometry 114. Conventional techniques for voxelization of a 3D point cloud are known to those of ordinary skill in the art. Therefore, for the sake of brevity, the details of voxelization are omitted from the present disclosure. According to one embodiment, the 3D point cloud geometry 114 can be received as a voxelized point cloud.

[0033] If the 3D point cloud geometry 114 includes a large number of data points (e.g., 10 4 The transmission or reception of such data points can consume a high network bandwidth. Similarly, the data points in an uncompressed state can consume a large amount of storage space. In certain network-based streaming applications, a 3D point cloud or a sequence of 3D point clouds can have to be streamed to more than one media device in near real-time (e.g., free-viewpoint video) for rendering operations. Prior to the rendering operations can be performed, the 3D point cloud or the sequence of 3D point clouds must be encoded at a source device (e.g., the electronic device 102) to achieve a size (e.g., in bytes) that is less than the uncompressed size of the 3D point cloud and the sequence of 3D point clouds. The size can help present a smoother streaming experience when the rendering operations are performed. Therefore, the 3D point cloud geometry 114 can be encoded (i.e., compressed) to minimize the network bandwidth usage and storage space usage for the transmission / reception of the 3D point cloud geometry 114. The encoding process of the 3D point cloud geometry 114 is described herein.

[0034] Electronic device 102 can segment the 3D point cloud geometry 114 into a set of blocks, and can select the first block (e.g., B0) from this set of blocks. The block selection can be performed iteratively. For each selection, it can proceed from RD0 to RD0. N Perform a linear search, where each RD can be searched in a linear manner. i Choose the encoding mode (i.e., from mode) 0 (model 0 ) to mode M To calculate the loss value. For the selected first block, the search yields a pair of modes. j RD i Its loss value can be lower than the mode. j The loss threshold for a given coding pattern can be a static value or dynamically adjusted based on human input or target rate distortion or quality. Table 1 provides examples of twenty-five (25) coding patterns for a set of five RD operation points, as follows:

[0035] Table 1: Exemplary Compression Metrics

[0036] As part of the linear search, the electronic device 102 can compute a first set of loss values ​​for the selected first block. The computed first set of loss values ​​can be associated with more than one compression metric. For example, the metric may include a bitrate metric (e.g., rate distortion) or a mean squared error (MSE) metric. This set of loss values ​​may correspond to a set of encoding modes. Such modes may be associated with at least one subset of the set of RD operation points. For example, Table 1 discloses that for each RD operation point (i.e., RD0 to RD4), there are multiple modes (i.e., Mode0 to Mode4). Each encoding mode may correspond to a deep neural network (called... ( (where j is the index of the pattern and i is the index of the RD point). A deep neural network can be trained to encode the first block of the selected 3D point cloud geometry 114 to generate the encoded first block.

[0037] For any block of point cloud geometry 114, the linear search can terminate at any arbitrary encoding pattern at any rate-distortion point. Therefore, the set of encoding patterns (for which the set of loss values ​​is computed) can correspond only to a subset of the set of RD operation points. In the worst case (i.e., with the worst time complexity), the linear search can cover all encoding patterns and all RD points. In this case, the subset could include all RD operation points in the set of RD operation points.

[0038] From the set of encoding modes, the electronic device 102 can select the encoding modes with loss values in the first set of loss values that are below a loss threshold. The loss threshold can be specific to the encoding mode or can be the same for all encoding modes for a particular RD operating point. Thereafter, the electronic device 102 can encode the selected first block based on the selected encoding mode. Similarly, the electronic device 102 can iteratively select modes for all remaining blocks of the 3D point cloud geometry 114 and can encode the remaining blocks of the 3D point cloud geometry 114. The encoded block data for all blocks of the 3D point cloud geometry 114 can be combined to generate an encoded 3D point cloud geometry.

[0039] In one embodiment, the electronic device 102 can generate supplemental information associated with the encoded 3D point cloud geometry. Examples of the supplemental information can include, but are not limited to, an encoding table, mode selection, index values of geometry information, and quantization parameters. The electronic device 102 can transmit the encoded 3D point cloud geometry to another electronic device that includes a decoder for reconstruction of the 3D point cloud geometry. The supplemental information can be transmitted along with the encoded 3D point cloud geometry.

[0040] In one embodiment, the electronic device 102 can be configured to obtain a calibration point cloud from the computing device 110, the database 108, and / or the server 106. For blocks of the calibration point cloud, the electronic device 102 can compute the first quartile of loss values for respective modes of RD operating points in the set of RD operating points. The electronic device 102 can set the loss threshold to the first quartile loss value corresponding to the encoding mode.

[0041] Figure 2 is a block diagram of an exemplary electronic device that illustrates an embodiment of the present disclosure. Figure 1 is a block diagram of an exemplary electronic device that illustrates an embodiment of the present disclosure. Figure 1 is a block diagram of an exemplary electronic device that illustrates an embodiment of the present disclosure. Figure 2 is a block diagram of an exemplary electronic device that illustrates an embodiment of the present disclosure. Figure 2 Referring to FIG. 2, a block diagram 200 of the electronic device 102 is shown. The electronic device 102 can include circuitry 202. The circuitry can include a processor 204, a classifier model 206, and a codec 208. In certain embodiments, the codec 208 can also include an encoder 208A. In certain embodiments, the codec can include a decoder 208B. The electronic device 102 can also include a memory 210, an input / output (I / O) device 212, and a network interface 214. The I / O device 212 can include a display device 212A that can be used to render multimedia content, such as a 3D point cloud or a 3D graphical model rendered from a 3D point cloud. The circuitry 202 can be communicatively coupled to the memory 210, the I / O device 212, and the network interface 214. The circuitry 202 can be configured to communicate with the server 106, the scanning setup 104, and the computing device 110 by using the network interface 214.

[0042] The processor 204 can include suitable logic, circuitry, and / or interfaces that can be configured to execute instructions associated with encoding of a 3D point cloud of an object. Further, the processor 204 can be configured to execute instructions associated with generation of a 3D point cloud of an object in a 3D space and / or reception of a plurality of color images and corresponding depth information. The processor 204 can be further configured to execute various operations related to sending and / or receiving the 3D point cloud (as multimedia content) to and / or from the computing device 110. An example of the processor 204 can be a graphics processing unit (GPU), a central processing unit (CPU), a tensor processing unit (TPU), a reduced instruction set computer (RISC) processor, an application-specific integrated circuit (ASIC) processor, a complex instruction set computer (CISC) processor, a co-processor, other processors, and / or combinations thereof. According to an embodiment, the processor 204 can be configured to support encoding of the 3D point cloud by the encoder 208A and decoding of the encoded 3D point cloud by the decoder 208B, while implementing other functionalities of the electronic device 102.

[0043] The encoder 208A can include suitable logic, circuitry, and / or interfaces that can be configured to encode 3D point cloud geometry corresponding to an object in a 3D space. In an embodiment, the encoder 208A can encode the 3D point cloud by encoding each 3D block associated with the 3D point cloud geometry. In certain embodiments, the encoder 208A can be configured to manage storage of the encoded 3D point cloud geometry in the memory 210 and / or forwarding of the encoded 3D point cloud geometry to other media devices (e.g., a portable media player) via the communication network 112.

[0044] In certain embodiments, the encoder 208A can be implemented as a deep neural network (in the form of computer executable code) on a GPU, CPU, TPU, RISC processor, ASIC processor, CISC processor, co-processor, other processor, and / or a combination thereof. In other embodiments, the encoder 208A can be implemented as a deep neural network on special-purpose hardware connected to other computing circuitry of the electronic device 102. In such implementations, the encoder 208A can be associated with a particular packaging form on a particular computing circuitry. Examples of particular computing circuitry can include, but are not limited to, a field-programmable gate array (FPGA), a programmable logic device (PLD), an ASIC, a programmable-ASIC (PL-ASIC), an application-specific integrated component (ASSP), and a system-on-chip (SOC) based on a standard microprocessor (MPU) or a digital signal processor (DSP). According to one embodiment, the encoder 208A can also be connected with a GPU to parallelize the operations of the encoder 208A. According to another embodiment, the encoder 208A can be implemented as a combination of programmable instructions stored in the memory 210 and logic units (or programmable logic units) on the hardware circuitry of the electronic device 102.

[0045] The decoder 208B can include suitable logic, circuitry, and / or interfaces that can be configured to decode the encoded information that can represent the geometry information of the object. The encoded information can also include supplemental information, such as encoding tables, weight information, mode information, index values of the geometry information, and quantization parameters, to assist the decoder 208B. For example, the encoded information can include an encoded 3D point cloud geometry. The decoder 208B can be configured to reconstruct the 3D point cloud geometry by decoding the encoded 3D point cloud geometry. According to one embodiment, the decoder 208B can be present on the computing device 110. According to one embodiment, the codec 208 can be integrated as part of an integrated circuit, such as a chip, a system-on-chip (SOC), or the like.

[0046] Memory 210 can include suitable logic, circuitry, and / or interfaces that can be configured to store instructions that can be executed by circuitry 202. Memory 210 can be configured to store an operating system and associated applications. Memory 210 can also be configured to store 3D point clouds corresponding to objects, including 3D point cloud geometry 114. According to one embodiment, memory 210 can be configured to store information related to a plurality of modes and a table mapping the plurality of modes to categories and operating conditions. According to another embodiment, memory 210 can be configured to store a set of rate-distortion (RD) operating points and one or more encoding modes associated with each RD operating point of the set of RD operating points. Examples of implementations of memory 210 can include, but are not limited to, random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), a hard disk drive (HDD), a solid-state drive (SSD), a CPU cache, and / or a secure digital (SD) card.

[0047] I / O device 212 can include suitable logic, circuitry, interfaces, and / or code that can be configured to receive user input. I / O device 212 can also be configured to provide output in response to user input. I / O device 212 can include various input and output devices that can be configured to communicate with circuitry 202. Examples of input devices can include, but are not limited to, a touchscreen, a keyboard, a mouse, a joystick, and / or a microphone. Examples of output devices can include, but are not limited to, display device 212A and / or a speaker.

[0048] Display device 212A can include suitable logic, circuitry, interfaces, and / or code that can be configured to render 3D point clouds onto a display screen of display device 212A. According to one embodiment, display device 212A can be a touch-enabled screen to receive user input. Display device 212A can be implemented through several known technologies, such as, but not limited to, liquid crystal display (LCD) display, light emitting diode (LED) display, plasma display, and / or organic LED (OLED) display technologies, and / or other display technologies. According to one embodiment, display device 212A can refer to a display screen of a smart eyewear device, a 3D display, a see-through display, a projection-based display, an electrochromic display, and / or a transparent display.

[0049] The network interface 214 can include suitable logic, circuitry, interfaces and / or code that can be configured to establish communication between the electronic device 102, the server 106, the scanning setup 104 and the computing device 110 via the communication network 112. The network interface 214 can be implemented by using various known technologies to support wired or wireless communication of the electronic device 102 with the communication network 112. The network interface 214 can include, but not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, and / or a local buffer.

[0050] The network interface 214 can communicate via wireless communications with a network, such as the Internet, an intranet and / or a wireless network, such as a cellular telephone network, a wireless local area network (LAN) and / or a metropolitan area network (MAN). The wireless communication can use any of a plurality of communications standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long Term Evolution (LTE), Fifth Generation (5G) New Radio (NR), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 1002.11a, IEEE 1002.11b, IEEE 1002.11g and / or IEEE 1002.11n), Voice over Internet Protocol (VoIP), Light Fidelity (Li-Fi), Wi-MAX, an email protocol, an instant messaging protocol, and / or Short Message Service (SMS).

[0051] Figure 3 is a block diagram of an exemplary encoder and an exemplary decoder of variable bit rate compression of point cloud geometry according to embodiments of the present disclosure. The elements of Figure 1 and Figure 2 are explained in conjunction with Figure 3 . Referring to Figure 3 , a block diagram 300 is shown that includes an encoder 302A and a decoder 302B. The encoder 302A can be an exemplary implementation of the encoder 208A of Figure 2 , and the decoder 302B can be an exemplary implementation of the decoder 208B of Figure 2 .

[0052] In one embodiment, the encoder 302A and the decoder 302B can be implemented on separate electronic devices. In another embodiment, the encoder 302A and the decoder 302B can both be implemented on the electronic device 102. The decoder 302B can also be implemented on the computing device 110.

[0053] The encoder 302A can include a set of encoders, such as a 1st encoder (e.g., encoder-1 304A),... and an Nth encoder (e.g., encoder-N 304N). The set of encoders of the encoder 302A can each include an associated classifier model, such as a neural network model. For example, the encoder-1 304A can be operatively coupled to a 1st deep neural network (DNN) model, such as a DNN model-1 306A. Further, the encoder-N 304N can be operatively coupled to an Nth DNN model, such as a DNN model-N 306N. The encoder 302A can also include a mode selector 308 that can be communicatively coupled to each of the encoder-1 304A,... and the encoder-N 304N.

[0054] Each deep neural network model (e.g., DNN model-1 306A) can be a neural network model that includes a system of artificial neurons or computational network arranged as nodes in multiple layers. The multiple layers of the neural network model can include an input layer, one or more hidden layers, and an output layer. Each of the multiple layers can include one or more nodes (or artificial neurons, represented, for example, by a circle). The output of all nodes in the input layer can be coupled to the input of at least one node of the hidden layer. Similarly, the output of each hidden layer can be coupled to the input of at least one node in other layers of the neural network model. The output of each hidden layer can be coupled to the input of at least one node in other layers of the neural network model. The nodes in the final layer can receive input from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer can be determined according to hyperparameters of the neural network model. Such hyperparameters can be set before or after the neural network model is trained on a training dataset.

[0055] Each node of the neural network model can correspond to a mathematical function (e.g., a Sigmoid function or a rectified linear unit) with a set of parameters that can be adjusted during network training. The set of parameters can include, for example, a weight parameter, a regularization parameter, and the like. Each node can use the mathematical function to compute an output based on one or more inputs from nodes in other layers (e.g., one or more previous layers) of the neural network model. All or a portion of the nodes of the neural network model can correspond to the same or identical mathematical function.

[0056] In the training of the neural network model, one or more parameters of each node of the neural network can be updated based on whether the output of the final layer for a given input (from the training dataset) matches a correct result based on a loss function of the neural network model. The above process can be repeated for the same or different inputs until a minimum of the loss function can be reached and the training error can be minimized. Several training methods are known in the art, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, metaheuristics, and the like.

[0057] The neural network model may include electronic data, which may be implemented as a software component of an application executable, for example, on an electronic device (e.g., electronic device 102). The neural network model may rely on libraries, external scripts, or other logic / instructions executed by a processing device such as circuit system 202. The neural network model may include code and routines configured to enable a computing device such as circuit system 202 to perform more than one operation to encode or decode 3D blocks associated with 3D point cloud geometry. Additionally or alternatively, classifier models, such as neural network models, may be implemented using hardware including processors, microprocessors (e.g., performance for performing or controlling more than one operation), field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs). Alternatively, in some embodiments, the neural network model may be implemented using a combination of hardware and software.

[0058] Decoder 302B includes a set of decoders, such as a first decoder (e.g., decoder-1 310A), ... and an Nth decoder (e.g., decoder-N 310N). Each decoder in this set of decoders may include an associated neural network model. For example, decoder-1 310A may include a first DNN model, such as DNN model-1 306A. Furthermore, decoder-N 310N may include an Nth DNN model, such as DNN model-N 306N. Figure 3 The block segmenter 312A associated with encoder 302A and binarizer and merger 312B associated with decoder 302B are shown. Also shown are the encoded bitstream and supplementary information 314A, signaling bitstream 314B, input point cloud 316A, reconstructed point cloud 316N, and 3D block set 318.

[0059] Figure 4 This is a diagram illustrating embodiments according to the present disclosure. Figure 2 A diagram illustrating exemplary components of a circuit system. Combined with... Figure 1 , Figure 2 and Figure 3 To explain the elements Figure 4 . Reference Figure 4 The diagram illustrates various components of a circuit system 202 for variable bit rate compression of point cloud geometry. Components 400 of the circuit system 202 may include a block segmenter 402, a classifier model 404, a loss computer 406, a mode selector 408, and an encoder 410.

[0060] The circuitry 202 can be configured to obtain a 3D point cloud of one or more objects (e.g., a person) in a 3D space as an input point cloud. The 3D point cloud can be a representation of geometry information and attribute information of the one or more objects in the 3D space. The geometry information can represent 3D coordinates (e.g., XYZ coordinates) of individual feature points of the 3D point cloud. Without attribute information, the 3D point cloud can be represented as a 3D point cloud geometry (e.g., the 3D point cloud geometry 114) associated with the one or more objects. The attribute information can include, for example, color information, reflectance information, opacity information, normal vector information, material identifier information, and texture information of the one or more objects. According to an embodiment, the 3D point cloud can be received from the scanning setup 104 via the communication network 112, or can be obtained directly from a built-in scanner, which can have the same functionality as the scanning setup 104.

[0061] An individual feature point in the 3D point cloud can be represented as (x, y, z, Y, Cb, Cr, a, a1,... an), where (x, y, z) can be 3D coordinates that can represent geometry information, (Y, Cb, Cr) can be luminance, chrominance blue-difference, and chrominance red-difference components of the feature point (in YCbCr or YUV color space), a can be a transparency value of the feature point, and a1~an can represent one or more attributes such as material identifier and normal vector. n n The one or more attributes can include, for example, material identifier and normal vector. n The Y, Cb, Cr, a, and a1~an can collectively represent attribute information of the individual feature point of the 3D point cloud.

[0062] The block partitioner 402 can receive an input point cloud and can partition the input point cloud into a set of blocks to generate a block stream 412. The block partitioner can perform a voxelization operation on the 3D point cloud. In the voxelization operation, the processor 204 can be configured to generate a plurality of voxels from the 3D point cloud. Each generated voxel can represent a volumetric element of the one or more objects in the 3D space. The volumetric element can represent attribute information and geometry information corresponding to a set of feature points of the 3D point cloud.

[0063] ​A 3D space corresponding to a 3D point cloud can be considered as a cube, which can be recursively partitioned into a plurality of sub-cubes (e.g., octants). A size of each sub-cube can be based on a density of feature points in the 3D point cloud. A plurality of feature points of the 3D point cloud can occupy different sub-cubes. Each sub-cube can correspond to a voxel, and can contain a set of feature points of the 3D point cloud within a particular volume of the corresponding sub-cube. The processor 204 can be configured to compute an average of attribute information associated with the set of feature points of the corresponding voxel. Further, the processor 204 can be configured to compute a center coordinate of each of the plurality of voxels based on geometry information associated with the corresponding set of feature points within the corresponding voxel. The generated plurality of voxels can be represented by the center coordinate and the average of the attribute information associated with the corresponding set of feature points.

[0064] According to an embodiment, the voxelization process of the 3D point cloud can be accomplished using conventional techniques known to those of ordinary skill in the art. Therefore, further details of such conventional techniques are omitted from the present disclosure for the sake of brevity. The plurality of voxels can represent geometry information and attribute information of more than one object in the 3D space. Further, the plurality of voxels can include occupied voxels and unoccupied voxels. The unoccupied voxels can not represent geometry information and attribute information of more than one object in the 3D space. Only the occupied voxels can represent geometry information and attribute information (e.g., color information) of the more than one object. According to an embodiment, the processor 204 can be configured to identify the occupied voxels from the plurality of voxels.

[0065] The block partitioner 402 can be configured to partition the plurality of voxels of the 3D point cloud geometry 114 into a set of blocks (e.g., the block stream 412). As an example and not by way of limitation, the processor 204 can partition the 3D point cloud geometry 114 into blocks, each block can have a predetermined size, such as 64x64x64. In one embodiment, the 3D point cloud geometry 114 can be partitioned into blocks of the same size. In another embodiment, the 3D point cloud geometry 114 can be partitioned into blocks of different sizes. For example, the plurality of voxels can include a first set of voxels that can be densely occupied and a second set of voxels that can be sparsely occupied. A portion of the 3D point cloud geometry 114 including the densely occupied voxels can be partitioned into a first set of blocks of size 32x32x32, while another portion of the 3D point cloud geometry 114 including the sparsely occupied voxels can be partitioned into a second set of blocks of size 64x64x64. According to an embodiment, the processor 204 can select block sizes to partition different portions of the 3D point cloud geometry 114 based on a trade-off between a computational cost associated with the partitioning operation and an occupancy density of the partitioned blocks.

[0066] The classifier model 404 can receive a plurality of voxels as input. The classifier model can be a neural network model, such as a DNN model. As a neural network model, the classifier model 404 can be a system of artificial neurons or computational network arranged in multiple layers. The multiple layers of the neural network model can include an input layer, one or more hidden layers, and an output layer. Each of the multiple layers can include one or more nodes (or artificial neurons, represented, for example, by a circle). The output of all nodes in the input layer can be coupled to the input of at least one node of the hidden layer. Similarly, the input of each hidden layer can be coupled to the output of at least one node in other layers of the neural network model. The output of each hidden layer can be coupled to the input of at least one node in other layers of the neural network model. The nodes in the final layer can receive input from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer can be determined according to the hyperparameters of the neural network model. Such hyperparameters can be set before or after the neural network model is trained on the training dataset.

[0067] Each node of the neural network model can correspond to a mathematical function (e.g., a Sigmoid function or a rectified linear unit) with a set of parameters that can be adjusted during network training. The set of parameters can include, for example, weight parameters, regularization parameters, and the like. Each node can use the mathematical function to compute an output based on one or more inputs from nodes in other layers (e.g., one or more previous layers) of the neural network model. All or a portion of the nodes of the neural network model can correspond to the same or different mathematical functions.

[0068] In the training of the neural network model, one or more parameters of each node of the neural network can be updated based on whether the output of the final layer for a given input (from the training dataset) matches the correct result based on a loss function of the neural network model. The above process can be repeated for the same or different inputs until a minimum of the loss function is reached and the training error is minimized. Several training methods are known in the art, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, metaheuristics, and the like.

[0069] The classifier model 404 can include electronic data that can be implemented as a software component of an application executable on an electronic device (e.g., the electronic device 102), for example. The classifier model 404 can rely on libraries, external scripts, or other logic / instructions executed by a processing device such as the circuitry 202. The classifier model 404 can include code and data that can enable a computing device such as the circuitry 202 to perform one or more operations to encode or decode a block associated with a 3D point cloud geometry. The classifier model 404 can be implemented using hardware including a processor, microprocessor (e.g., to perform one or more operations or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in certain embodiments, the classifier model 404 can be implemented using a combination of hardware and software.

[0070] The loss computer 406 can receive input from the classifier model 404. The loss computer 406 can be configured to compute loss values for blocks of the block stream 412. The processor 204 can select a block or a set of blocks from the block stream 412. For the selected block or set of blocks, a set of loss values associated with one or more compression metrics can be computed. The set of loss values can correspond to a set of encoding modes associated with at least a subset of the set of RD operating points. The first set of loss values can be computed based on the output of the classifier model 404 for the input. The set of loss values can include one or more loss values computed for one or more encoding modes corresponding to a first RD operating point in the set of RD operating points. The loss values can be computed for blocks of the block stream 412.

[0071] The mode selector 408 can receive input from the loss computer 406. Based on the input from the loss computer 406, a mode selection operation can be performed. In the mode selection operation, the processor 204 can be configured to determine a mode for a block in the set of blocks. In alternative embodiments, the mode selection operation can be performed by the encoder 208A. Further, the processor 204 can be configured to select one or more encoding modes (e.g., the selected one or more encoding modes) for the block from a plurality of encoding modes based on a comparison of the computed loss values for the block or the set of blocks to loss thresholds for the encoding modes. Here, each mode in the plurality of modes can correspond to a function that can be used to encode a block.

[0072] In one embodiment, the one or more modes can be selected based on looking up a table or metric that maps modes to categories and operating conditions. In another embodiment, the one or more modes can be selected based on modes used by blocks neighboring the current block in a spatial arrangement of the set of blocks in the 3D point cloud geometry 114.

[0073] When the "one or more modes" includes more than one mode, the processor 204 can determine the rate-distortion cost associated with each of the selected modes and can compare the determined rate-distortion costs with each other. Based on the comparison of the determined rate-distortion costs, the processor 204 can select the mode with the lowest rate-distortion cost from the selected modes as the optimal mode to encode the current 3D block. In another scenario, if the "one or more modes" includes a single mode, the rate-distortion cost of that mode may not be determined. Instead, the single mode itself may be the optimal mode for encoding the current block. Figure 5 The text further describes in detail the determination of the pattern and the selection of more than one pattern.

[0074] Encoder 410 can be configured to encode blocks based on one or more selected modes. For example, encoder 410 can encode blocks based on one or more selected modes to obtain encoded block 414.

[0075] Figure 5 This is a diagram illustrating an exemplary processing pipeline for variable bitrate compression of point cloud geometry according to embodiments of the present disclosure. (In conjunction with...) Figure 1 , Figure 2 , Figure 3 and Figure 4 To explain the elements Figure 5 . Reference Figure 5 The diagram illustrates a processing pipeline 500. Within the processing pipeline 500, a sequence of operations from 502 to 526 is shown. This sequence of operations can be executed by any computing device, such as the circuitry 202 of electronic device 102.

[0076] At 502, the first block 502A (i.e., 3D block) can be selected from the block set 502B (i.e., 3D blocks) of the 3D point cloud geometry. The circuit system 202 can be configured to partition the 3D point cloud geometry into the block set 502B. This selection can be part of an iterative selection process to search for the optimal encoding mode and optimal RD operation point for each block. The search can be a bidirectional search traversing all modes corresponding to the RD operation point set. During the search, the circuit system 202 (e.g., a block encoder) can be allowed to operate on different RD... i Switch between operation points to search for patterns that meet the set cost constraints (i.e., loss thresholds).

[0077] Rate distortion (RD) can be performed in 504. i Operation point selection. In RD i In point selection, you can select from the RD operation point set (e.g., such as...) Figure 1 As described in Table 1, RD0 to RD4 are selected. iAn operating point. For example, as described in Table 1, RD0(i = 0) can be selected from RD0~RD4. Each RD operating point can correspond to a particular rate-distortion between the original point cloud block and the reconstructed point cloud block. The rate-distortion can be based on the point-to-point distance or the plane-to-plane distance between corresponding points in the original point cloud block and the reconstructed point cloud block (or any other objective or subjective distortion metric), and the estimated number of bits required to encode the corresponding block. Each RD operating point in the set of RD operating points can be associated with more than one loss threshold, which corresponds to more than one encoding mode for the RD operating point.

[0078] At 506, a mode selection operation can be performed. In the mode selection operation, the circuitry 202 can be configured to select an encoding mode (e.g., M0) from more than one encoding mode (e.g., M0~M4 of Table 1) associated with the selected RD operating point (RD0). Such encoding modes can correspond to deep neural networks, each of which can be trained to encode the selected 1st block 502A of 3D point cloud geometry to generate an encoded 1st block. The selection of the encoding mode can be performed in a linear fashion. The starting index (j) of the selected encoding (M) can be set to 0, and the index can be incremented in further iterations. It should be noted that the selected encoding mode should not be considered as the final encoding mode to be used in the encoding operation of the 1st block 502A. At 508, the selected encoding mode (M0) should only be considered as a candidate encoding mode for the selected 1st block 502A.

[0079] At 508, a loss computation operation can be performed on the selected 1st block 502A. The circuitry 202 can compute a loss value for the selected 1st block 502A based on the selected encoding mode (e.g., M0). The loss value can be associated with a compression metric such as a code rate metric or an MSE metric. According to one embodiment, the circuitry 202 can input the selected 1st block 502A to a classifier model. In this case, the loss value can be computed based on the output of the classifier model for the input. The classifier model can be, for example, a deep neural network (DNN) model trained for implicit or explicit geometric properties of a test block of at least one point cloud. For example, such properties can include the density of points associated with the point cloud.

[0080] At 510, the computed loss value for the selected 1st block 502A can be compared to the loss threshold of the selected encoding mode. During the comparison, it can be determined whether the computed loss of the selected 1st block 502A is less than the threshold of the selected mode. If the computed loss of the selected 1st block 502A is less than the threshold of the selected mode, control can pass to 512. If the computed loss of the selected 1st block 502A is greater than the threshold of the selected mode, control can pass to 514.

[0081] According to one embodiment, the loss threshold can be different values for each encoding mode of the set of RD operating points. For example, the circuitry 202 can be configured to obtain a calibration point cloud from the computing device 110 and / or the server 106. For a block of the calibration point cloud, the circuitry 202 can calculate a first quartile of loss values corresponding to each mode of the RD operating points of the set of RD operating points. The circuitry 202 can set the first quartile of loss values as the loss threshold corresponding to the encoding mode. Table 2 gives an example of loss thresholds for different RD operating points, as follows:

[0082] Table 2: Loss thresholds for RD points

[0083] In Table 2, various cost / loss thresholds for RDi operating points are shown by "i" from 0 to 4. Among the RD operating points, RD4 is associated with the smallest loss threshold of 0.161, and RD0 is associated with the largest loss threshold of 0.800.

[0084] According to one embodiment, the circuitry 202 can set the loss threshold for an encoding mode based on user input. User input can be provided to adjust the loss threshold for a particular RD operating point. For example, the user input can require the loss threshold for RD2 to change from 0.267 to 0.134. An example of adjustment of loss thresholds is given in Table 3, as follows:

[0085] Table 3: Loss thresholds for RD points

[0086] In Table 3, various cost / loss thresholds for RDi operating points are shown by "i" from 0 to 4. Among the RD operating points, RD4 is associated with the smallest loss threshold of 0.080, and RD1 is associated with the largest loss threshold of 0.400.

[0087] According to one embodiment, the loss threshold can be a fixed value for each encoding mode corresponding to the set of RD operating points. An example of fixed loss thresholds is provided in Table 4, as follows:

[0088] Table 4: Fixed loss thresholds for encoding modes

[0089] At 512, an encoding operation can be performed. As part of this operation, the circuitry 202 can encode the selected first block 502A based on the selected encoding mode. According to one embodiment, the selected encoding mode can correspond to a first deep neural network (e.g., DNN model-1 306A), and the selected first block 502A can be encoded based on applying the first deep neural network to the selected first block 502A. After 512, control can continue to proceed to 524.

[0090] At 514, an operation can be performed to determine whether other modes are available for the selected RD operating point (e.g., RD0). If other modes are available for the selected RD operating point (i.e., the selected mode (Mj) is not the last mode for the RD operating point), then control can proceed to 516. If other modes are not available for the selected RD operating point (i.e., the selected mode (Mj) is the last mode for the RD operating point), then control can proceed to 518.

[0091] At 516, an operation can be performed to switch to the next mode (e.g., M1) associated with the selected RD operating point (e.g., RD0). The switch can be made by increasing the value of the model index (j) by 1. After the switch, the next mode (e.g., M1) can be selected, and the operations from 506 to 510 can be iterated for the next mode until the loss value exceeds the loss threshold for the next mode.

[0092] After traversing all of the modes for the RD operating point, a set of loss values can be obtained for the selected first block 502A. For the selected first block 502A, the circuitry 202 can compute a first set of loss values that can be associated with more than one compression metric. The first set of loss values can correspond to a set of encoding modes associated with a subset of the set of RD operating points. For example, from RD0 to RD2 (i.e., a subset of 5 RD operating points), there can be 12 modes (4 modes per RD point), and the first set of loss values can include 12 loss values. The first set of loss values includes more than one first loss value that can be computed for more than one encoding mode corresponding to a first RD operating point (e.g., RD0) in the set of RD operating points. The number of loss computations for a selected block can vary, and can depend on the values of the loss thresholds that can be set for the modes. i The first set of loss values includes more than one first loss value that can be computed for more than one encoding mode corresponding to a first RD operating point (e.g., RD0) in the set of RD operating points. The number of loss computations for a selected block can vary, and can depend on the values of the loss thresholds that can be set for the modes.

[0093] At 518, an operation can be performed to determine whether there are other RDi operating points available for selection from the set of RD operating points. If there are other RDi operating points available for selection, then control can proceed to 520. If there are no other RDi operating points available for selection (i.e., the RD operating point selected at 504 is the last RD operating point in the set of RD operating points), then control can proceed to 522.

[0094] At 520, an operation can be performed to switch to a next RD operation point (e.g., RD1) in the set of RD operation points (e.g., RD0...RD4). The switch can be made by increasing the value of the RD index (i) by 1. After the switch, the next RD operation point (e.g., RD1) can be selected at 504, and the operations from 506 to 510 can be performed for the next mode iteration until the loss value exceeds the loss threshold for the next mode. For example, the circuitry 202 can switch to a second RD operation point (RD1) in the set of RD operation points based on a determination that the one or more first loss values are greater than the one or more loss thresholds for the one or more encoding modes corresponding to the first RD operation point (e.g., RD0) (i.e., the modes such as M0-M4 of RD0). The set of first loss values (as described in 516) can include one or more second loss values, which can be computed for the one or more encoding modes corresponding to the second RD operation point (e.g., RD1).

[0095] At 522, a lossless encoding operation can be performed. As part of the operation, the original block encoding mode (LLmode) can be turned on to handle the case where, during the local reconstruction at the encoding stage, none of the RD operation points (selected at 504) satisfy the local quality criterion for encoding of the selected first block 502A (i.e., the loss value of all RD operation points is higher than the loss threshold), and an undesirable artifact (e.g., a hole) is detected in the reconstruction of the point cloud geometry. The first block 502A can be stored losslessly. In particular, the circuitry 202 can select a lossless encoding scheme for the first block 502A based on a determination that each loss value in the set of first loss values (as described in 516) is higher than the loss threshold for the corresponding encoding mode in the set of encoding modes. The circuitry 202 can encode the selected first block 502A based on the selected lossless encoding scheme. For example, Figure 9 An example application of the lossless encoding scheme is provided in Section.

[0096] At 524, an operation can be performed to determine whether the selected first block 502A is the last block in the set of blocks 502B for selection. If it is determined that the selected first block 502A is the last block, the circuitry 202 can prepare and send the encoded bitstream 528 associated with the point cloud geometry to the computing device 110 or the server 106. Thereafter, control can pass to the end. If it is determined that the selected first block 502A is not the last block, control can pass to 526.

[0097] At 526, an operation can be performed to select a next block (e.g., a second block) in the set of blocks. After the selection, the operations from 504 to 524 can be repeated to process the next block and subsequent blocks until the last block in the set of blocks is processed.

[0098] Figure 6 This is a diagram illustrating an exemplary search method for the pattern of RD operation points according to an embodiment of this disclosure. (Refer to...) Figure 6 Figure 600 illustrates an exemplary search method for patterns of RD operation points. Combined with... Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 5 To describe the elements Figure 6 Figure 600 shows a table of five RD operation points RD0–RD4 and five corresponding modes mode0–mode4. To search for the optimal encoding mode for a specific RD operation point, a search can be performed (e.g., ...). Figure 5 The search can start from RD0 and mode0, and can traverse RD (as shown by the double arrows). i All Perform a bidirectional search. The number of steps to find the optimal encoding pattern may vary and typically depends on the value of the loss threshold set for these patterns.

[0099] If given an RD operation point (e.g., RD0) ~ If none of the encoding requirements are met (i.e., the block's loss value is higher than the loss threshold for all modes), then the circuit system 202 can switch to the next RD operation point in the set of RD operation points. For the selected block, the search may produce a pair of modes. j RD i Its loss value can be maintained in mode j The loss threshold is below the target value. The loss threshold for a given encoding pattern can be a static value or dynamically adjusted based on human input or target rate distortion or quality.

[0100] Figure 7 This is a diagram illustrating an exemplary comparison between lossy and lossless reconstruction outputs of point cloud geometry according to embodiments of the present disclosure. (In conjunction with...) Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 6 To describe the elements Figure 7 The elements. (Refer to) Figure 7FIG. 7 shows an exemplary comparison 700 between a lossy reconstruction output and a lossless reconstruction output of point cloud geometry. In comparison 700, point cloud geometry 702, point cloud geometry 704, and point cloud geometry 706 are shown. Point cloud geometry 702 represents a reference (original) point cloud of a human head. Point cloud geometry 704 can be a lossy reconstruction of encoded point cloud data that can be obtained from point cloud geometry 702 after applying encoder 208A to a block of point cloud geometry 702. For example, an ML model can be configured to encode a block of a point cloud based on RD points and best modes corresponding to the RD points. For example, Figure 5 The selection of RD points and best modes is described in mode ). The LL mode mode can be used to losslessly encode the block in region 702A to ensure that such undesirable artifacts do not appear in the local reconstruction. Point cloud geometry 706 is shown to include region 706A that is reconstructed from lossless encoded blocks and corresponds to regions 704A and 702A. As shown, there are no visible artifacts such as holes in region 706A.

[0101] Figure 8 FIG. 10 is a diagram illustrating an exemplary 3D point cloud geometry and selection of a region of interest (RoI) in point cloud geometry for point cloud compression according to embodiments of the present disclosure. Referring to FIG. 10, a 3D point cloud geometry 800 is shown to include portion 802, portion 804, portion 806, and portion 808. In conjunction with the elements of Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 elements of Figure 8 are described in conjunction with the elements of

[0102] During operation, the electronic device 102 can determine a portion of the 3D point cloud geometry as a region of interest (ROI). The determination can be made based on input from a user (e.g., a 3D artist) or an automated operation, such as an object detection operation performed on the 3D point cloud geometry 800 or a semantic segmentation operation performed on the 3D point cloud geometry 800. For example, the user can explicitly define a RoI by dividing the 3D point cloud geometry 800 into different slices of the portion 802, the portion 804, the portion 806, and the portion 808. The user can also assign modes or RD points to the respective portions, such as assigning a lossless mode (LL mode ) to the portion 802, assigning RD4 to the portion 804, assigning RD0 to the portion 806, and assigning RD3 to the portion 808. The electronic device 102 can encode the blocks corresponding to the portion 802 of the 3D point cloud geometry 800 based on the lossless encoding scheme. The blocks corresponding to the other portions (e.g., the portion 804, the portion 806, and the portion 808) of the 3D point cloud geometry 800 can be encoded with the best mode associated with the assigned RD points. For example, such modes can be identified using the operations described in Figure 5 .

[0103] Figure 9 is a flowchart illustrating exemplary operations of variable rate compression of point cloud geometry, in accordance with an embodiment of the present disclosure. Referring to FIG. 11, a flowchart 900 is shown. The flowchart 900 is described in conjunction with Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 , and Figure 8 . The operations 902-916 can be implemented on the electronic device 102. The method described in the flowchart 900 can start at 902 and proceed to 904.

[0104] At 904, a set of RD operating points and one or more encoding modes associated with each RD operating point of the set of RD operating points can be stored. The electronic device 102 can include the memory 210, which can be configured to store the set of RD operating points and the one or more encoding modes associated with each RD operating point of the set of RD operating points.

[0105] At 906, a 3D point cloud geometry (e.g., the 3D point cloud geometry 114) can be received. In one embodiment, the circuitry 202 can be configured to receive the 3D point cloud geometry 114. The 3D point cloud geometry 114 can be received from the scanning setup 104 or the server 106 via the communication network 112. For example, The reception of the 3D point cloud geometry is further described in Figure 4 .

[0106] At 908, the 3D point cloud geometry 114 can be partitioned into a set of blocks (e.g., the block stream 412). In one embodiment, the circuitry 202 can be configured to partition the 3D point cloud geometry 114 into the set of blocks. For example, Figure 4 The partitioning of the 3D point cloud geometry is further described in

[0107] At 910, a 1st block 502A can be selected from the set of blocks 502B. In one embodiment, the circuitry 202 can be configured to select the 1st block 502A from the set of blocks 502B.

[0108] At 912, a 1st set of loss values associated with one or more compression metrics can be computed for the selected 1st block 502A. The 1st set of loss values can correspond to a set of encoding modes associated with at least a subset of the set of RD operating points. In one embodiment, the circuitry 202 can be configured to compute the 1st set of loss values for the selected 1st block 502A. For example, Figure 1 and Figure 5 The computation of the loss values is further described in

[0109] At 914, an encoding mode having a loss value below a loss threshold in the 1st set of loss values can be selected for the selected 1st block 502A from the set of encoding modes. In one embodiment, the circuitry 202 can be configured to select the encoding mode from the set of encoding modes.

[0110] At 916, the 1st block 502A can be encoded based on the selected encoding mode. In one embodiment, the circuitry 202 can be configured to encode the 1st block 502A based on the selected encoding mode. For example, Figure 5 The encoding of the 1st block 502A is further described in

[0111] Various embodiments of the present disclosure can provide a non-transitory computer readable medium and / or storage medium having stored thereon instructions executable by a machine and / or computer to operate an electronic device (e.g., the electronic device 100), the method comprising: Figure 1computer-executable instructions of the computer program product. Such instructions can cause the electronic device 102 to perform operations that can include storing a set of RD operating points and one or more encoding modes associated with each RD operating point of the set of RD operating points. The operations can further include receiving 3D point cloud geometry related to one or more objects in a 3D space and segmenting the 3D point cloud geometry into a set of blocks. The operations can further include selecting a first block from the set of blocks and computing a first set of loss values associated with one or more compression metrics for the selected first block. The set of loss values can correspond to a set of encoding modes associated with at least a subset of the set of RD operating points. The operations can further include selecting an encoding mode from the set of encoding modes and encoding the selected first block based on the selected encoding mode, where a loss value of the encoding mode in the first set of loss values is lower than a loss threshold of the encoding mode.

[0112] An exemplary aspect of the present disclosure provides an electronic device (e.g., the electronic device 102) that can include a memory (e.g., the memory 210) that can be configured to store a set of RD operating points and one or more encoding modes associated with each RD operating point of the set of RD operating points. The electronic device can further include circuitry (e.g., the circuitry 202) that can be configured to receive 3D point cloud geometry related to one or more objects in a 3D space. The circuitry 202 can be further configured to segment the 3D point cloud geometry into a set of blocks and select a first block from the set of blocks. For the selected first block, the circuitry 202 can be configured to compute a first set of loss values associated with one or more compression metrics. The set of loss values can correspond to a set of encoding modes associated with at least a subset of the set of RD operating points. The circuitry 202 can be further configured to select an encoding mode from the set of encoding modes, where a loss value of the encoding mode in the first set of loss values is lower than a loss threshold of the encoding mode. Thereafter, the circuitry 202 can encode the selected first block based on the selected encoding mode.

[0113] According to one embodiment, the one or more compression metrics can include a rate-distortion metric or a mean squared error (MSE) metric.

[0114] According to one embodiment, the circuitry 202 can be further configured to input the selected first block to a classifier model. The first set of loss values can be further computed based on an output of the input to the classifier model. The classifier model can be a deep neural network (DNN) model trained for one or more geometric properties of a test block of a point cloud. Such properties can include a density of points associated with the point cloud.

[0115] According to an embodiment, the circuitry 202 can be further configured to determine a portion of the 3D point cloud geometry as a region of interest (ROI) and encode a block corresponding to the determined portion in the 3D point cloud geometry based on a lossless encoding scheme. The portion of the 3D point cloud geometry can be determined as the ROI based on at least one of a user input, an object detection operation, or a semantic segmentation operation.

[0116] According to an embodiment, each RD operating point in the set of RD operating points can be associated with one or more loss thresholds corresponding to one or more encoding modes.

[0117] According to an embodiment, the encoding modes can correspond to deep neural networks, each of which can be trained to encode a selected 1st block of the 3D point cloud geometry to generate an encoded 1st block. The selected 1st block can be encoded based on applying a 1st deep neural network in the deep neural networks to the selected 1st block. The 1st deep neural network can correspond to the selected encoding mode.

[0118] According to an embodiment, the set of 1st loss values can include one or more 1st loss values that can be computed for one or more encoding modes corresponding to a 1st RD operating point in the set of RD operating points. The circuitry 202 can be further configured to switch to a 2nd RD operating point in the set of RD operating points based on a determination that the one or more 1st loss values are greater than one or more loss thresholds for the one or more encoding modes corresponding to the 1st RD operating point. The set of 1st loss values can include one or more 2nd loss values that can be computed for one or more encoding modes corresponding to the 2nd RD operating point.

[0119] According to an embodiment, the circuitry 202 can be further configured to obtain a calibration point cloud and compute, for a block of the calibration point cloud, a first quartile of loss values corresponding to each mode of an RD operating point in the set of RD operating points. The first quartile of loss values corresponding to an encoding mode can be set as the loss threshold for the encoding mode.

[0120] According to an embodiment, the loss threshold can be a fixed value for each encoding mode corresponding to the set of RD operating points.

[0121] According to an embodiment, the circuitry 202 can be further configured to set the loss threshold for an encoding mode based on a user input.

[0122] According to an embodiment, the circuitry 202 can be further configured to select a lossless encoding scheme for the 1st block based on a determination that each loss value in the set of 1st loss values is higher than the loss threshold for a corresponding encoding mode in the set of encoding modes. The circuitry 202 can encode the selected 1st block based on the selected lossless encoding scheme.

[0123] The present disclosure can be realized in hardware, or a combination of hardware and software. The present disclosure can be realized in a centralized fashion in at least one computer system or in a distributed fashion where different elements are spread across several interconnected computer systems. Any kind of computer system or other apparatus adapted for carrying out the methods described herein is suited. A combination of hardware and software can also be realized as specialized or general-purpose hardware, having a computer program that, when being loaded and executed, can control the computer system such that it carries out the methods described herein. The present disclosure can be realized in a hardware component that comprises an integrated circuit

[0124] The present disclosure can also be embedded in a computer program product, which comprises all the features enabling the implementation of the methods described herein, and which - when being loaded in a computer system - is able to carry out these methods. Computer program means or computer program in the present context mean any expression, in any language, code or notation, of a set of instructions intended to cause a system having information processing capability to perform a particular function either directly or after either or both of the following: a) conversion to another language, code or notation; b) reproduction in a different material form.

[0125] While the present disclosure has been described with reference to certain embodiments thereof, a person of ordinary skill in the art will understand that various changes can be made and equivalents substituted for elements thereof without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt a particular situation or material to the teachings of the present disclosure without departing from the scope thereof. Therefore, it is intended that the present disclosure not be limited to the particular embodiment disclosed, but that the present disclosure will include all embodiments falling within the scope of the appended claims.

Claims

1. An electronic device, comprising: a memory configured to store a set of rate-distortion (RD) operating points and one or more encoding modes associated with each RD operating point of the set of RD operating points; and circuitry configured to: receive a three-dimensional (3D) point cloud geometry; segment the 3D point cloud geometry into a set of blocks; select a first block from the set of blocks; for the selected first block, compute a first set of loss values associated with one or more compression metrics, wherein the first set of loss values corresponds to a set of encoding modes associated with at least a subset of the set of RD operating points; select an encoding mode from the set of encoding modes, wherein a loss value of the encoding mode in the first set of loss values is lower than a loss threshold of the encoding mode; and encode the selected first block based on the selected encoding mode.

2. The electronic device of claim 1, wherein the one or more compression metrics comprise a rate metric or a mean squared error (MSE) metric.

3. The electronic device of claim 1, wherein the circuitry is further configured to input the selected first block to a classifier model, and wherein the first set of loss values is further computed based on an output of the classifier model for the input.

4. The electronic device of claim 3, wherein the classifier model is a deep neural network (DNN) model trained for one or more geometric characteristics of a test block of a point cloud.

5. The electronic device of claim 4, wherein the one or more geometric characteristics comprise a density of points associated with the point cloud.

6. The electronic device of claim 1, wherein the circuitry is further configured to: determine a portion of the 3D point cloud geometry as a region of interest (ROI); and encode a block corresponding to the determined portion of the 3D point cloud geometry based on a lossless encoding scheme.

7. The electronic device of claim 6, wherein the portion of the 3D point cloud geometry is determined as the ROI based on at least one of a user input, an object detection operation, or a semantic segmentation operation.

8. The electronic device of claim 1, wherein each RD operating point of the set of RD operating points is associated with one or more loss thresholds corresponding to the one or more encoding modes.

9. The electronic device of claim 1, wherein the encoding modes correspond to deep neural networks, each deep neural network trained to encode the selected first block of the 3D point cloud geometry to generate an encoded first block.

10. The electronic device of claim 9, wherein the selected first block is encoded based on applying a first deep neural network of the deep neural networks to the selected first block, and the first deep neural network corresponds to the selected encoding mode.

11. The electronic device of claim 1, wherein the first set of loss values comprises one or more first loss values computed for one or more encoding modes corresponding to a first RD operating point of the set of RD operating points. ​ 12.The electronic device of claim 11, wherein the circuitry is further configured to, based on a determination that the one or more first loss values are greater than one or more loss thresholds of the one or more coding modes corresponding to the first RD operating point, switch to a second RD operating point in the set of RD operating points, and wherein the first set of loss values comprises one or more second loss values calculated for one or more coding modes corresponding to the second RD operating point. 13.The electronic device of claim 1, wherein the circuitry is further configured to: obtain a calibration point cloud; and calculate, for a block of the calibration point cloud, a first quartile of loss values corresponding to each mode of an RD operating point in the set of RD operating points; and set the first quartile of the loss values corresponding to the coding mode as the loss threshold. 14.The electronic device of claim 1, wherein the loss threshold is a fixed value for each coding mode corresponding to the set of RD operating points. 15.The electronic device of claim 1, wherein the circuitry is further configured to set the loss threshold for the coding mode based on a user input. 16.The electronic device of claim 1, wherein the circuitry is further configured to: select, for a first block, a lossless coding scheme based on a determination that each loss value in the first set of loss values is higher than a loss threshold of a corresponding coding mode in the set of coding modes; and encode the selected first block based on the selected lossless coding scheme. 17.A method comprising: in an electronic device: storing a rate-distortion (RD) operating point set and one or more coding modes associated with each RD operating point in the RD operating point set; receiving a three-dimensional (3D) point cloud geometry; segmenting the 3D point cloud geometry into a block set; selecting a first block from the block set; calculating, for the selected first block, a first set of loss values associated with one or more compression metrics, wherein the first set of loss values corresponds to a set of coding modes associated with at least a subset of the RD operating point set; selecting a coding mode from the set of coding modes, wherein a loss value of the coding mode in the first set of loss values is lower than a loss threshold of the coding mode; and encoding the selected first block based on the selected coding mode. 18.The method of claim 17, further comprising: selecting, for a first block, a lossless coding scheme based on a determination that each loss value in the first set of loss values is higher than a loss threshold of a corresponding coding mode in the set of coding modes; and encoding the selected first block based on the selected lossless coding scheme. 19.A non-transitory computer-readable medium having stored thereon computer-executable instructions that, when executed by an electronic device, cause the electronic device to perform operations comprising: storing a rate-distortion (RD) operating point set and one or more coding modes associated with each RD operating point in the RD operating point set; receiving a three-dimensional (3D) point cloud geometry; ​ ​ segmenting the 3D point cloud geometry into a set of blocks; selecting a first block from the set of blocks; for the selected first block, computing a first set of loss values associated with one or more compression metrics, wherein the first set of loss values corresponds to a set of encoding modes associated with at least a subset of the set of RD operating points; selecting an encoding mode from the set of encoding modes, wherein the loss value of the encoding mode in the first set of loss values is lower than a loss threshold of the encoding mode; and encoding the selected first block based on the selected encoding mode.

20. The non-transitory computer readable medium of claim 19, wherein the operations comprise: based on a determination that each loss value in the first set of loss values is higher than a loss threshold of a corresponding encoding mode in the set of encoding modes, selecting a lossless encoding scheme for the first block; and encoding the selected first block based on the selected lossless encoding scheme.