Variable-rate compression of point cloud geometry

The variable-rate compression method addresses reconstruction artifacts in point clouds by enabling flexible RD operating point switching and region-of-interest encoding, improving visual quality and efficiency in point cloud compression.

JP2026515763APending Publication Date: 2026-05-19SONY GROUP CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2024-04-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Conventional point cloud compression methods often result in artifacts during reconstruction due to the use of full-search algorithms and machine learning-based approaches, which do not support raw point coding modes or region-of-interest-based coding, leading to unpredictable local artifacts and inefficiencies.

Method used

The proposed solution involves an electronic device that employs a variable-rate compression method, allowing switching between different RD operating points and modes to satisfy user-defined constraints, supports reversible encoding for raw points, and enables region-of-interest-based encoding, using a bidirectional search across all modes to improve visual quality and reduce reconstruction artifacts.

Benefits of technology

This approach enhances the visual quality of reconstructed point clouds by minimizing artifacts and optimizing compression efficiency, while supporting diverse encoding strategies tailored to specific regions of interest.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026515763000001_ABST
    Figure 2026515763000001_ABST
Patent Text Reader

Abstract

An electronic device and method for variable-rate compression of point cloud geometry are provided. The electronic device stores a set of RD operating points and encoding modes associated with the set of RD operating points. The electronic device receives a 3D point cloud geometry and divides the 3D geometry into a set of blocks. After division, the electronic device selects a block and calculates a set of loss values ​​associated with one or more compression metrics. Such loss values ​​correspond to a set of encoding modes associated with at least a subset of the set of RD operating points. From the set of encoding modes, the electronic device selects an encoding mode in which the loss value from the set of loss values ​​falls below a loss threshold for that encoding mode. The electronic device then encodes the block based on the encoding mode.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Incorporation through cross-reference / citation to related applications]

[0001] This application claims priority to U.S. Patent Application No. 18 / 303,056, filed with the U.S. Patent and Trademark Office on 19 April 2023. Each of the above applications is incorporated herein by reference in its entirety.

[0002]

[0002] Various embodiments of the present disclosure relate to three-dimensional (3D) point cloud compression (PCC). More specifically, various embodiments of the present disclosure relate to variable-rate compression of point cloud geometry. [Background technology]

[0003]

[0003] Advances in three-dimensional (3D) scanning have made it possible to create high-fidelity 3D geometric representations of 3D objects. 3D point clouds are one example of such 3D geometric representations and have been employed in various applications such as free-viewpoint displays for broadcasting sports or live events, geographic information systems, representations of cultural heritage, or autonomous navigation of vehicles. Typically, a point cloud consists of a large number of unstructured 3D points (e.g., each point has X, Y, and Z coordinates) and associated attributes (e.g., textures including color or reflectivity). When compressing a 3D point cloud, it is desirable that most of the attributes associated with the 3D point cloud be preserved in order to facilitate the efficient reconstruction of the point cloud from the encoded point cloud data. Therefore, it is sometimes desirable to have an efficient point cloud compression (PCC) method. In geometry compression, conventional PCC methods may only provide a limited number of working points. Applying such methods to point cloud geometry may result in artifacts in the reconstructed point cloud geometry.

[0004]

[0004] Those skilled in the art will be able to see the limitations and disadvantages of conventional methods by comparing the described system with some aspects of the disclosure shown with reference to the drawings in the remainder of this application. [Overview of the project] [Problems that the invention aims to solve]

[0005]

[0005] Provided are electronic devices and methods for variable-rate compression of point cloud geometry, substantially shown in at least one figure and / or described in relation to these figures, and more fully shown in the claims.

[0006]

[0006] These and other features and advantages of the Disclosure can be understood by considering the following detailed description of the Disclosure with reference to the accompanying drawings, which indicate the same elements throughout by the same reference numerals. [Brief explanation of the drawing]

[0007] [Figure 1] This block diagram shows an exemplary environment for variable-rate compression of point cloud geometry according to embodiments of the present disclosure. [Figure 2] This is a block diagram of an exemplary electronic device shown in Figure 1, according to an embodiment of the present disclosure. [Figure 3] This is a block diagram of an exemplary encoder and an exemplary decoder for variable-rate compression of point cloud geometry according to an embodiment of the present disclosure. [Figure 4] This figure shows an example of the components of the circuit in Figure 2 according to an embodiment of the present disclosure. [Figure 5] This figure shows an exemplary processing pipeline for variable-rate compression of point cloud geometry according to an embodiment of the present disclosure. [Figure 6] This figure shows an exemplary search pattern for the mode of the RD operating point according to an embodiment of the present disclosure. [Figure 7]A diagram showing an exemplary comparison between an irreversible reconstruction output and a reversible reconstruction output of point cloud geometry according to an embodiment of the present disclosure. [Figure 8] A diagram showing an exemplary 3D point cloud geometry and the selection of a region of interest (ROI) within the point cloud geometry for point cloud compression according to an embodiment of the present disclosure. [Figure 9] A flowchart showing an exemplary operation for variable rate compression of point cloud geometry according to an embodiment of the present disclosure.

MODE FOR CARRYING OUT THE INVENTION

[0008]

[0016] The implementations described below can be found in the disclosed electronic devices and methods for variable rate compression of point cloud geometry. Exemplary aspects of the present disclosure can include an electronic device that can include a memory configured to store a set of rate distortion (RD) operating points and one or more coding modes associated with each RD operating point in the set of RD operating points. The electronic device can further include circuitry configured to receive a three-dimensional (3D) point cloud geometry regarding one or more objects in 3D space. The electronic device can be further configured to divide the 3D point cloud geometry into a set of blocks and select a first block from the set of blocks. The electronic device can be further configured to calculate a first set of loss values associated with one or more compression metrics for the selected first block. The first set of loss values can correspond to a set of coding modes associated with at least a subset of the set of RD operating points. The electronic device can be further configured to select, from the set of coding modes, a coding mode for which a loss value in the first set of loss values is below a loss threshold for the coding mode. Thereafter, the electronic device can encode the selected first block based on the selected coding mode.

[0009]

[0017] Conventional point cloud compression (PCC) methods typically compress point cloud geometry by selecting a mode using a full-search algorithm that investigates which mode yields the best rate-distortion (RD) performance. Possible implementations of adaptive block-based point cloud geometry compression are machine learning-based. Several models can be trained on each model, tuned to specific learned local point cloud characteristics. RD control can be implemented by implicit and explicit quantization. These conventional methods may introduce unpredictable local artifacts into the locally reconstructed geometry. Such artifacts may be due to the strong nonlinearity introduced by machine learning-based point cloud compression schemes. Furthermore, these conventional methods may not support raw point coding modes or region-of-interest-based coding modes. Some existing methods (e.g., Adaptive Deep Learning PCC (ADL-PCC)) may, given a specific RD operating point, return a model (i.e., coding mode) that returns the minimum cost (e.g., rate-distortion, bitrate, or a combination of both). In this disclosure, it is considered that the minimum cost among all models for a given RD point may be too high (above a threshold). Therefore, the electronic devices of this disclosure enable the encoder to switch between different RD operating points to search for a model (i.e., mode) that satisfies cost constraints. In a more general formulation, by bidirectional search across all modes at an RD point, the encoder can satisfy user-defined constraints or quality requirements (e.g., reduction of reconstruction artifacts). This disclosure enables encoding different portions of a point cloud using different combinations of RD modes and improves the visual quality of the reconstructed point cloud by locally imposing a maximum allowable cost threshold. This disclosure further presents a reversible mode for encoding raw points, which can be enabled when none of the available modes can satisfy the maximum cost criterion. Alternatively, this disclosure enables region-of-interest-based encoding of different slices of point cloud geometry.

[0010]

[0018] FIG. 1 is a block diagram showing an exemplary environment for variable rate compression of point cloud geometry according to an embodiment of the present disclosure. Referring to FIG. 1, a network environment 100 is shown. The network environment 100 can include an electronic device 102, a scanning setup 104, a server 106, a database 108, and a computing device 110. The scanning setup 104 can include one or more image sensors (not shown) and one or more depth sensors (not shown) associated with the one or more image sensors. The electronic device 102 can be communicatively coupled to the scanning setup 104, the server 106, and the computing device 110 via a communication network 112. A three-dimensional (3D) point cloud geometry 114 of a 3D point cloud associated with at least one object (e.g., a person) in a 3D space is further shown.

[0011]

[0019] The electronic device 102 can include suitable logic, circuitry, interfaces, and / or code configured to encode and / or decode 3D point cloud geometry (e.g., 3D point cloud geometry 114). A 3D point cloud can include a set of data points in space. These points can represent the 3D shape of an object such that the position of each point corresponds to a set of orthogonal coordinates. As an example, each point can be represented as (x, y, z, r, g, b, α), where (x, y, z) represents the 3D coordinates of the point on the object, (r, g, b) represents the red, green, and blue values of the point, and (α) can represent the transparency value of the point.

[0012]

[0020] In some embodiments, the electronic device 102 can be configured to generate a 3D point cloud of one or more objects (e.g., a 3D scene including objects in the foreground and background). The electronic device 102 can obtain a 3D point cloud geometry 114 of one (or more) objects from the 3D point cloud. Examples of the electronic device 102 include, but are not limited to, computing devices, video conferencing systems, augmented reality (AR) devices, virtual reality (VR) devices, mixed reality (MR) devices, game consoles, smart wearable devices, mainframe machines, servers, computer workstations, and / or consumer electronics (CE) devices.

[0013]

[0021] The scanning setup 104 may include preferred logic, circuitry, interfaces, and / or code that can be configured to scan a 3D environment containing objects to generate a raw 3D scan (also known as a raw 3D point cloud). According to one embodiment, the scanning setup 104 may include a single image capture device or multiple image capture devices (arranged at multiple viewpoints) to capture multiple color images. In certain cases, the scanning setup 104 may include an additional depth sensor to capture depth information of objects. Multiple color images and depth information of objects may be captured from different viewpoints. In such cases, a 3D point cloud can be generated based on the captured multiple color images and corresponding depth information of objects.

[0014]

[0022] The scanning setup 104 can be configured to perform a 3D scan of an object in 3D space and generate a dynamic 3D point cloud (i.e., a point cloud sequence) that can capture changes in different attributes and the geometry of 3D points at different time steps. The scanning setup 104 can be configured to transmit the generated 3D point cloud, multiple color images, and / or corresponding depth information to the electronic device 102 and / or server via the communication network 112.

[0015]

[0023] According to one embodiment, the scanning setup 104 may include multiple sensors, such as a combination of a depth sensor and a color sensor (e.g., a red-green-blue (RGB) sensor), and / or a combination of an infrared (IR) projector and an IR sensor. For example, the depth sensor may capture information associated with the point cloud geometry (3D position of points), and the RGB and IR sensors may capture information associated with the point cloud attributes (e.g., color and temperature). In one embodiment, the IR projector and IR sensor may be used to estimate depth information. A combination of the depth sensor, RGB sensor, and IR sensor may be used to capture a point cloud frame (a single static point cloud) or multiple point cloud frames (3D video) containing the associated geometry and attributes.

[0016]

[0024] According to one embodiment, the scanning setup 104 may include an active 3D scanner that relies on radiation or light to capture the 3D structure of an object in 3D space. The scanning setup 104 may also include an image sensor that can capture color information associated with the object. For example, the active 3D scanner may be a time-of-flight (TOF) based 3D laser scanner, laser rangefinder, TOF camera, handheld laser scanner, structured light 3D scanner, modulated light 3D scanner, CT scanner that outputs point cloud data, airborne light detection and ranging (LiDAR) scanner, 3D LiDAR, 3D motion sensor, etc.

[0017]

[0025] In Figure 1, the scanning setup 104 is shown separately from the electronic device 102. However, in some embodiments, the scanning setup 104 can be integrated into the electronic device 102. In alternative embodiments, the entire functionality of the scanning setup 104 can be incorporated into the electronic device 102 without departing from the scope of this disclosure. Examples of the scanning setup 104 include, but are not limited to, a depth sensor, an RGB sensor, an IR sensor, an image sensor, a light cage including a camera, and / or a motion detection device.

[0018]

[0026] Server 106 may include suitable logic, circuitry, interfaces, and / or code that can be configured to perform operations such as data / file storage, 3D rendering, or 3D reconstruction operations (such as photogrammetry reconstruction operations). For example, but not limited to, 3D reconstruction operations may be performed using photogrammetry-based methods (such as structure from motion (SfM)), methods requiring stereoscopic images, or methods requiring monocular cues (such as shape from shading (SfS), photometric stereo, or shape from texture (SfT)). Details of such methods are omitted from this disclosure for brevity. Examples of Server 106 include, but are not limited to, application servers, cloud servers, web servers, database servers, file servers, game servers, mainframe servers, or combinations thereof.

[0019]

[0027] Database 108 may include suitable logic, interfaces, and / or code that can be configured to store point cloud geometries. Database 108 can retrieve data from relational or non-relational databases, or from a set of comma-separated value (CSV) files in conventional or big data storage. Database 108 can be stored or cached on a device such as server 106. In some embodiments, database 108 can be hosted on multiple servers stored in the same or different locations. The operation of database 108 can be performed using hardware including a processor, a microprocessor (e.g., performing or controlling the execution of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some other cases, database 108 can be implemented using software.

[0020]

[0028] The computing device 110 may include preferred logic, circuitry, interfaces, and / or code that can be configured to communicate with the electronic device 102 and / or a server via a communication network 112. According to one embodiment, the computing device 110 may include memory configured to store a set of rate-distortion (RD) operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points. The computing device 110 may be configured to receive encoded 3D point cloud geometry (e.g., as part of multimedia content) from the electronic device 102. The computing device 110 may be configured to decode the encoded 3D point cloud geometry to render a 3D model of an object. Examples of the computing device 110 include, but are not limited to, desktops, personal computers, laptops, computer workstations, tablet computing devices, smartphones, cellular phones, mobile phones, consumer electronic (CE) devices with displays, televisions (TVs), wearable displays, head-mounted displays, digital signage, and digital mirrors (or smart mirrors) capable of storing or rendering multimedia content.

[0021]

[0029] According to another embodiment, the computing device 110 can be configured to receive user input and determine a portion of the 3D point cloud geometry as a region of interest (ROI). The ROI can be determined using user input in conjunction with object detection or semantic segmentation operations; that is, a portion of the 3D point cloud is determined as the ROI.

[0022]

[0030] During operation, the electronic device 102 can receive 3D point cloud geometry (such as 3D point cloud geometry 114) associated with at least one object in 3D space. For example, 3D point cloud data can be obtained from a 3D point cloud (or 3D scan) that includes geometry and various attributes. The 3D point cloud can be a static point cloud or a frame of a dynamic point cloud (i.e., a point cloud sequence). Generally, a 3D point cloud is a representation of geometric information (e.g., 3D coordinates of points) and attribute information of an object in 3D space. Attribute information may include, for example, color information, reflectivity information, opacity information, normal vector information, material identifier information, or texture information associated with the object in 3D space. Texture information can represent the spatial arrangement of color or intensity in multiple color images of the object. Reflectivity information can represent information associated with an empirical model of local illumination of points in the 3D point cloud (e.g., the Phong shading model or the Gouraud shading model). An empirical model of local illumination can correspond to the reflectivity of an object's surface (rough or glossy surface areas). Opacity information can represent the transparency of a point. Normal vector information can represent the direction perpendicular to the tangent plane at a point in a point cloud.

[0023]

[0031] After receiving, the electronic device 102 can generate multiple voxels from the 3D point cloud geometry 114. The generation of voxels can be referred to as voxelization of the 3D point cloud geometry 114. Prior art for voxelization of 3D point clouds may be well known to those skilled in the art. Therefore, for the sake of brevity, details of voxelization are omitted from this disclosure. According to one embodiment, the 3D point cloud geometry 114 can be received as a voxelized point cloud.

[0024]

[0032] 3D point cloud geometry 114 has a large number of data points (for example, 10 4 If the data points include the above order of magnitude, transmitting or receiving such data points may consume high network bandwidth. Similarly, uncompressed data points may consume a large amount of storage space. Some network-based streaming applications may require streaming a 3D point cloud or 3D point cloud sequence to one or more media devices in near real-time (e.g., free-viewpoint video) for rendering operations. Before performing the rendering operation, the 3D point cloud or 3D point cloud sequence must be encoded on the source device (e.g., electronic device 102) to achieve a size smaller than the uncompressed size of the 3D point cloud or 3D point cloud sequence (e.g., in bytes). This size can help render a smoother streaming experience during the rendering operation. Therefore, the 3D point cloud geometry 114 can be encoded (i.e., compressed) to minimize network bandwidth usage and storage space usage for transmitting / receiving the 3D point cloud geometry 114. This specification describes the encoding process for the 3D point cloud geometry 114.

[0025]

[0033] The electronic device 102 can divide the 3D point cloud geometry 114 into a set of blocks and select a first block (e.g., B0) from the set of blocks. Block selection can be performed iteratively. For each selection, from RD0 to RD N Linear search can be performed up to each RD. i The coding mode for is linear (i.e., mode 0 ~mode M A selection can be made from the first block, and the loss value can be calculated. For the first selected block, the search is performed so that the loss value is calculated in mode j A mode that can fall below the loss threshold for j and RDi This can result in pairs. The loss threshold for a given coding mode can be a static value or can be dynamically adjusted based on human input or target rate distortion or quality. Examples of 25 coding modes for a set of five RD operating points are shown in Table 1 below. Exemplary compression metrics [Table 1]

[0026]

[0034] As part of a linear search, the electronic device 102 can calculate a first set of loss values ​​for a selected first block. The calculated first set of loss values ​​can be associated with one or more compression metrics. For example, the metrics can include a rate metric (e.g., rate distortion) or a mean squared error (MSE) metric. The set of loss values ​​can correspond to a set of coding modes. Such modes can be associated with at least a subset of the set of RD operating points. For example, Table 1 shows that for all RD operating points (i.e., RD0 to RD4), there are multiple modes (i.e., modes 0 to 4). Each coding mode corresponds to a deep neural network ( This is called JPEG2026515763000003.jpg8170, where j is the mode index and i is the RD point index. A deep neural network can be trained to encode a selected first block of 3D point cloud geometry 114 and generate an encoded first block.

[0027]

[0035] For any block of the point cloud geometry 114, the linear search can terminate in any coding mode for any rate-distortion point. Thus, the set of coding modes (for which a set of loss values ​​is calculated) can only correspond to a subset of the set of RD operating points. In the worst case (i.e., at worst-case time complexity), the linear search can cover all coding modes and all RD points. In such a case, the subset can include all RD operating points in the set of RD operating points.

[0028]

[0036] The electronic device 102 can select an encoding mode from a set of encoding modes in which the loss value in a first set of loss values ​​falls below a loss threshold. The loss threshold can be specific to an encoding mode or it can be the same for all encoding modes at a particular RD operating point. The electronic device 102 can then encode the selected first block based on the selected encoding mode. Similarly, the electronic device 102 can iteratively select a mode for all remaining blocks of the 3D point cloud geometry 114 and encode the remaining blocks of the 3D point cloud geometry 114. The encoded block data of all blocks of the 3D point cloud geometry 114 can be combined to generate the encoded 3D point cloud geometry.

[0029]

[0037] In one embodiment, the electronic device 102 can generate supplementary information associated with the encoded 3D point cloud geometry. Examples of supplementary information include, but are not limited to, an encoding table, mode selection, index values ​​for geometric information, and quantization parameters. The electronic device 102 can transmit the encoded 3D point cloud geometry to another electronic device, which includes a decoder for reconstructing the 3D point cloud geometry. The supplementary information can be transmitted together with the encoded 3D point cloud geometry.

[0030]

[0038] In one embodiment, the electronic device 102 can be configured to acquire a calibration point cloud from the computing device 110, the database 108, and / or the server 106. For blocks of the calibration point cloud, the electronic device 102 can calculate the first quartile of loss value corresponding to each mode of the RD operating point in the set of RD operating points. The electronic device 102 can set the first quartile of loss value corresponding to the encoding mode as the loss threshold.

[0031]

[0039] Figure 2 is a block diagram of an exemplary electronic device of Figure 1 according to an embodiment of the present disclosure. The description of Figure 2 will be made in relation to the elements of Figure 1. Referring to Figure 2, a block diagram 200 of the electronic device 102 is shown. The electronic device 102 may include a circuit 202. The circuit may include a processor 204, a classifier model 206, and a codec 208. In some embodiments, the codec 208 may also include an encoder 208A. In some embodiments, the codec may include a decoder 208B. The electronic device 102 may further include a memory 210, an input / output (I / O) device 212, and a network interface 214. The I / O device 212 may include a display device 212A, which can be used to render multimedia content such as a 3D point cloud or a 3D graphic model rendered from a 3D point cloud. The circuit 202 may be communicatively coupled to the memory 210, the I / O device 212, and the network interface 214. The circuit 202 can be configured to communicate with the server 106, the scanning setup 104, and the computing device 110 using the network interface 214.

[0032]

[0040] The processor 204 may include preferred logic, circuitry, and / or interfaces that can be configured to execute instructions associated with encoding a 3D point cloud of an object. The processor 204 may also be configured to execute instructions associated with generating a 3D point cloud of an object in 3D space and / or receiving multiple color images and corresponding depth information. The processor 204 can be further configured to perform various operations related to transmitting and / or receiving a 3D point cloud (as multimedia content) to and from the computing device 110. Examples of the processor 204 may be a graphics processing unit (GPU), a central processing unit (CPU), a tensor processing unit (TPU), a reduced instruction set computer (RISC) processor, an application-specific integrated circuit (ASIC) processor, a composite instruction set computer (CISC) processor, a coprocessor, other processors, and / or combinations thereof. According to one embodiment, the processor 204 may be configured such that an encoder 208A encodes a 3D point cloud, a decoder 208B decodes the encoded 3D point cloud, and other functions of the electronic device 102 are facilitated.

[0033]

[0041] The encoder 208A may include suitable logic, circuitry, and / or interfaces that can be configured to encode 3D point cloud geometry corresponding to objects in 3D space. In one embodiment, the encoder 208A can encode a 3D point cloud by encoding each 3D block associated with the 3D point cloud geometry. In a particular embodiment, the encoder 208A may be configured to manage the storage of encoded 3D point cloud geometry to memory 210 and / or the transfer of encoded 3D point cloud geometry to other media devices (e.g., a portable media player) via a communication network 112.

[0034]

[0042] In some embodiments, the encoder 208A can be implemented as a deep neural network (in the form of computer executable code) on a GPU, CPU, TPU, RISC processor, ASIC processor, CISC processor, coprocessor, other processor, and / or combination thereof. In some other embodiments, the encoder 208A can be implemented as a deep neural network on dedicated hardware interfaced with other computing circuits of the electronic device 102. In such implementations, the encoder 208A can be associated with a specific form factor on a particular computing circuit. Examples of specific computing circuits include, but are not limited to, field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), ASICs, programmable ASICs (PL-ASICs), application-specific integrated circuits (ASSPs), and systems-on-chip (SOCs) based on standard microprocessors (MPUs) or digital signal processors (DSPs). According to one embodiment, the encoder 208A can also be interfaced with a GPU to parallelize the operation of the encoder 208A. According to another embodiment, the encoder 208A can be implemented as a combination of programmable instructions stored in memory 210 and a logic device (or programmable logic unit) on the hardware circuit of the electronic device 102.

[0035]

[0043] Decoder 208B may include preferred logic, circuitry, and / or interfaces that can be configured to decode encoded information that can represent the geometric information of an object. The encoded information may include supplementary information, such as an encoding table, weight information, mode information, index values ​​for the geometric information, and quantization parameters, which may also assist Decoder 208B. As an example, the encoded information may include encoded 3D point cloud geometry. Decoder 208B may be configured to reconstruct the 3D point cloud geometry by decoding the encoded 3D point cloud geometry. According to one embodiment, Decoder 208B may reside on a computing device 110. According to one embodiment, the codec 208 may be integrated as part of an integrated circuit, such as a chip or a system-on-a-chip (SOC).

[0036]

[0044] Memory 210 may include preferred logic, circuitry, and / or interfaces that can be configured to store instructions that circuitry 202 can execute. Memory 210 may be configured to store an operating system and associated applications. Memory 210 may be further configured to store a 3D point cloud (including 3D point cloud geometry 114) corresponding to an object. According to one embodiment, memory 210 may be configured to store information related to a plurality of modes and a table mapping the plurality of modes to classes and operating conditions. According to another embodiment, memory 210 may be configured to store a set of rate-distortion (RD) operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points. Examples of implementations of memory 210 include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), hard disk drives (HDDs), solid-state drives (SSDs), CPU caches, and / or secure digital (SD) cards.

[0037]

[0045] The I / O device 212 may include suitable logic, circuitry, interfaces, and / or code that can be configured to receive user input. The I / O device 212 may be further configured to provide outputs in response to user input. The I / O device 212 may include various input and output devices, which may be configured to communicate with circuitry 202. Examples of input devices include, but are not limited to, a touchscreen, keyboard, mouse, joystick, and / or microphone. Examples of output devices include, but are not limited to, a display device 212A and / or a speaker.

[0038]

[0046] The display device 212A may include suitable logic, circuitry, interfaces, and / or code that can be configured to render a 3D point cloud on the display screen of the display device 212A. According to one embodiment, the display device 212A is a touch-enabled screen that can receive user input. The display device 212A can be implemented through several known technologies, and / or other display technologies, such as liquid crystal display (LCD) displays, light-emitting diode (LED) displays, plasma displays, and / or organic LED (OLED) display technologies, but is not limited to the following. According to one embodiment, the display device 212A may mean a display screen for a smart glasses device, a 3D display, a see-through display, a projection-based display, an electrochromic display, and / or a transparent display.

[0039]

[0047] The network interface 214 may include suitable logic, circuitry, interfaces, and / or code that can be configured to establish communication between the electronic device 102, the server 106, the scanning setup 104, and the computing device 110 via the communication network 112. The network interface 214 may be implemented using various known techniques to support wired or wireless communication between the electronic device 102 and the communication network 112. The network interface 214 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identification module (SIM) card, and / or a local buffer.

[0040]

[0048] The network interface 214 can communicate wirelessly with networks such as the Internet, intranets, and / or cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs). Wireless communication may use any of several communication standards, protocols, and technologies, including Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long-Term Evolution (LTE), 5th Generation (5G) New Radio (NR), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (IEEE 1002.11a, IEEE 1002.11b, IEEE 1002.11g and / or IEEE 1002.11n, etc.), Voice over Internet Protocol (VoIP), Light Fidelity (Li-Fi), Wi-MAX, and protocols for email, instant messaging, and / or Short Message Service (SMS).

[0041]

[0049] Figure 3 is a block diagram of an exemplary encoder and an exemplary decoder for variable-rate compression of point cloud geometry according to an embodiment of the present disclosure. The description of Figure 3 will be made in relation to the elements of Figures 1 and 2. Referring to Figure 3, a block diagram 300 is shown, which includes encoder 302A and decoder 302B. Encoder 302A may be an exemplary implementation of encoder 208A in Figure 2, and decoder 302B may be an exemplary implementation of decoder 208B in Figure 2.

[0042]

[0050] In one embodiment, the encoder 302A and the decoder 302B can be mounted on separate electronic devices. In another embodiment, both the encoder 302A and the decoder 302B can be mounted on the electronic device 102. The decoder 302B can also be mounted on the computing device 110.

[0043]

[0051] Encoder 302A may include a set of encoders, such as a first encoder (e.g., encoder 1 304A), ... and an Nth encoder (e.g., encoder N 304N). Each encoder in the set of encoders of encoder 302A may include an associated classifier model, such as a neural network model. For example, encoder 1 304A can be operably coupled to a first deep neural network (DNN) model, such as DNN model 1 306A. Furthermore, encoder N 304N can be operably coupled to an Nth DNN model, such as DNN model N 306N. Encoder 302A may further include a mode selector 308, which can be communically coupled to each of encoders 1 304A, ... and encoder N 304N.

[0044]

[0052] Each deep neural network model (e.g., DNN model 1 306A) can be a neural network model that includes a system of computational networks or artificial neurons arranged in multiple layers as nodes. The multiple layers of the neural network model can include an input layer, one or more hidden layers, and an output layer. Each of the multiple layers can include one or more nodes (or, for example, artificial neurons represented by circles). The outputs of all nodes in the input layer can be coupled to at least one node in the (one or multiple) hidden layers. Similarly, the input of each hidden layer can be coupled to the output of at least one node in the other layers of the neural network model. The output of each hidden layer can be coupled to the input of at least one node in the other layers of the neural network model. The (one or multiple) nodes in the final layer can receive input from at least one hidden layer and output a result. The number of layers and the number of nodes in each layer can be determined from the hyperparameters of the neural network model. Such hyperparameters can be set before or after training the neural network model with the training dataset.

[0045]

[0053] Each node in a neural network model can correspond to a mathematical function (e.g., a sigmoid function or a normalized linear unit) with a set of parameters that can be adjusted during network training. The parameter set may include, for example, weight parameters and regularization parameters. Each node can use the mathematical function to compute an output based on one or more inputs from nodes in other layers (e.g., previous layers) of the neural network model. All or some nodes in a neural network model can correspond to the same or different mathematical functions.

[0046]

[0054] In training a neural network model, one or more parameters of each node in the neural network can be updated based on whether the output of the final layer for a given input (from the training dataset) matches the correct result based on the loss function for the neural network model. This process can be repeated for the same or different inputs until the minimum value of the loss function can be achieved and the training error can be minimized. Several training methods are known in the art, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, and metaheuristics.

[0047]

[0055] A neural network model may include electronic data that can be implemented, for example, as a software component of an application executable on an electronic device (e.g., electronic device 102). The neural network model may rely on libraries, external scripts, or other logic / instructions executed by a processing device such as circuit 202. The neural network model may include code and routines configured to enable a computing device such as circuit 202 to perform one or more operations for encoding or decoding 3D blocks associated with 3D point cloud geometry. In addition, or alternatively, a classifier model such as a neural network model may also be implemented using hardware including a processor, a microprocessor (e.g., performing or controlling one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the neural network model may be implemented using a combination of hardware and software.

[0048]

[0056] Decoder 302B may include a set of decoders such as a first decoder (e.g., decoder 1 310A), ... and N decoders (e.g., decoder N 310N). Each decoder in the set of decoders may include an associated neural network model. For example, decoder 1 310A may include a first DNN model such as DNN model 1 306A. Furthermore, decoder N 310N may include an Nth DNN model such as DNN model N 306N. Figure 3 shows a block division unit 312A associated with encoder 302A and a binarization / merging unit 312B associated with decoder 302B. Also shown are the encoded bitstream and supplementary information 314A, the signaling bitstream 314B, the input point cloud 316A, the reconstructed point cloud 316N, and a set of 3D blocks 318.

[0049]

[0057] Figure 4 shows exemplary components of the circuit of Figure 2 according to an embodiment of the present disclosure. The description of Figure 4 will be made in relation to the elements of Figures 1, 2 and 3. Referring to Figure 4, various components of the circuit 202 for variable rate compression of point cloud geometry are shown. The components 400 of the circuit 202 may include a block division unit 402, a classifier model 404, a loss calculation unit 406, a mode selector 408, and an encoder 410.

[0050]

[0058] Circuit 202 can be configured to obtain, as an input point cloud, a 3D point cloud of one or more objects (such as a person) in 3D space. The 3D point cloud can be a representation of the geometric information and attribute information of one or more objects in 3D space. The geometric information can indicate the 3D coordinates (such as XYZ coordinates) of the individual feature points of the 3D point cloud. When there is no attribute information, the 3D point cloud can be represented as a 3D point cloud geometry (for example, 3D point cloud geometry 114) associated with one or more objects. The attribute information can include, for example, color information, reflectivity information, opacity information, normal vector information, material identifier information, and texture information of one or more objects. According to an embodiment, the 3D point cloud can be received from the scanning setup 104 via the communication network 112, or can be directly obtained from a built-in scanner having the same function as the scanning setup 104.

[0051]

[0059] Each feature point in the 3D point cloud can be represented as (x, y, z, Y, Cb, Cr, α, a1,... a n ), where (x, y, z) can be 3D coordinates that can represent geometric information, and (Y, Cb, Cr) can be the luminance, chroma - blue difference, and chroma - red difference components of the feature point in the (YCbCr or YUV color space). α can be the transparency value of the feature point, and a1~a n represent one - dimensional or multi - dimensional attributes such as material identifiers and normal vectors. Y, Cb, Cr, α, and a1~a n can jointly represent the attribute information of each feature point of the 3D point cloud.

[0052]

[0060] The block division unit 402 can receive an input point cloud and divide the input point cloud into a set of blocks to generate a block stream 412. The block division unit can perform a voxelization operation on the 3D point cloud. In the voxelization operation, the processor 204 can be configured to generate multiple voxels from the 3D point cloud. Each generated voxel can represent a volume element of one or more objects in 3D space. The volume element can represent attribute information and geometric information corresponding to a group of feature points in the 3D point cloud.

[0053]

[0061] The 3D space corresponding to a 3D point cloud can be thought of as a cube that can be recursively divided into multiple subcubes (such as octants). The size of each subcube can be based on the density of feature points in the 3D point cloud. Multiple feature points in the 3D point cloud can occupy different subcubes. Each subcube can correspond to a voxel, and a set of feature points from the 3D point cloud can be contained within a specific volume of the corresponding subcube. The processor 204 can be configured to calculate the average of the attribute information associated with the set of feature points of the corresponding voxel. The processor 204 can also be configured to calculate the center coordinates for each of the multiple voxels based on the geometric information associated with the set of feature points within the corresponding voxel. Each of the generated multiple voxels can be represented by its center coordinates and the average of the attribute information associated with the set of feature points of the corresponding voxel.

[0054]

[0062] According to one embodiment, the process of voxelizing a 3D point cloud can be carried out using prior art that is well known to those skilled in the art. Therefore, further details of such prior art are omitted from this disclosure for brevity. A plurality of voxels can represent geometric and attribute information of one or more objects in 3D space. A plurality of voxels can also include occupied voxels and unoccupied voxels. Unoccupied voxels cannot represent geometric and attribute information of one or more objects in 3D space. Only occupied voxels can represent geometric and attribute information (such as color information) of one or more objects. According to one embodiment, the processor 204 can be configured to identify occupied voxels from a plurality of voxels.

[0055]

[0063] The block division unit 402 can be configured to divide a plurality of voxels of the 3D point cloud geometry 114 into a set of blocks (e.g., a block stream 412). For example, the processor 204 can divide the 3D point cloud geometry 114 into blocks, each of which may have a predetermined size, such as 64 × 64 × 64. In one embodiment, the 3D point cloud geometry 114 can be divided into blocks of the same size. In another embodiment, the 3D point cloud geometry 114 can be divided into blocks of different sizes. For example, a plurality of voxels may include a first set of voxels that may have a high density of occupancy and a second set of voxels that may have a low density of occupancy. A portion of the 3D point cloud geometry 114 containing densely occupied voxels can be divided into a first set of blocks of size 32 × 32 × 32, while another portion of the 3D point cloud geometry 114 containing sparsely occupied voxels can be divided into a second set of blocks of size 64 × 64 × 64. According to one embodiment, the processor 204 can select block sizes to divide different parts of the 3D point cloud geometry 114 based on a trade-off between the computational cost associated with the division operation and the occupancy density of the divided blocks.

[0056]

[0064] The classifier model 404 can accept multiple voxels as input. The classifier model can be a neural network model, such as a DNN model. As a neural network model, the classifier model 404 can be a system of computational networks or artificial neurons that can be arranged in multiple layers. The multiple layers of the neural network model can include an input layer, one or more hidden layers, and an output layer. Each of the multiple layers can include one or more nodes (or artificial neurons represented by circles, for example). The outputs of all nodes in the input layer can be coupled to at least one node in the (one or multiple) hidden layers. Similarly, the input of each hidden layer can be coupled to the output of at least one node in the other layers of the neural network model. The output of each hidden layer can be coupled to the input of at least one node in the other layers of the neural network model. The (one or multiple) nodes in the final layer can receive input from at least one hidden layer and output a result. The number of layers and the number of nodes in each layer can be determined from the hyperparameters of the neural network model. These hyperparameters can be set before or after training the neural network model on the training dataset.

[0057]

[0065] Each node in a neural network model can correspond to a mathematical function (e.g., a sigmoid function or a normalized linear unit) with a set of parameters that can be adjusted during network training. The parameter set may include, for example, weight parameters and regularization parameters. Each node can use the mathematical function to compute an output based on one or more inputs from nodes in other layers (e.g., previous layers) of the neural network model. All or some nodes in a neural network model can correspond to the same or different mathematical functions.

[0058]

[0066] In training a neural network model, one or more parameters of each node in the neural network can be updated based on whether the output of the final layer for a given input (from the training dataset) matches the correct result based on the loss function for the neural network model. The above process can be repeated for the same or different inputs until the minimum value of the loss function is achieved and the training error is minimized. Several training methods are known in the art, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, and metaheuristics.

[0059]

[0067] The classifier model 404 may include electronic data that can be implemented, for example, as a software component of an application executable on an electronic device (e.g., electronic device 102). The classifier model 404 may rely on libraries, external scripts, or other logic / instructions executed by a processing device such as circuit 202. The classifier model 404 may include code and data that enable a computing device such as circuit 202 to perform one or more operations for encoding or decoding blocks associated with 3D point cloud geometry. The classifier model 404 can be implemented using hardware including a processor, a microprocessor (e.g., performing or controlling one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the classifier model 404 can be implemented using a combination of hardware and software.

[0060]

[0068] The loss calculation unit 406 can receive input from the classifier model 404. The loss calculation unit 406 can be configured to calculate loss values ​​for blocks in the blockstream 412. The processor 204 can select a block or a set of blocks from the blockstream 412. For the selected block or set of blocks, a set of loss values ​​associated with one or more compression metrics can be calculated. The set of loss values ​​can correspond to a set of coding modes associated with at least a subset of the set of RD operating points. The first set of loss values ​​can be calculated based on the output of the classifier model 404 for the input. The set of loss values ​​can include one or more loss values ​​calculated for one or more coding modes corresponding to a first RD operating point in the set of RD operating points. Loss values ​​can be calculated for each block in the blockstream 412.

[0061]

[0069] The mode selector 408 can receive input from the loss calculation unit 406. Based on the input from the loss calculation unit 406, a mode selection operation can be performed. In the mode selection operation, the processor 204 may be configured to determine a mode for a block in a set of blocks. In an alternative embodiment, the mode selection operation can be performed by the encoder 208A. Furthermore, the processor 204 may be configured to select one or more encoding modes (e.g., one or more selected encoding modes) for a block from a plurality of encoding modes, based on a comparison of the loss value calculated for the block or set of blocks with a loss threshold for the encoding mode. In this specification, each of the plurality of modes may correspond to a function that can be used to encode a block.

[0062]

[0070] In one embodiment, one or more modes can be selected based on a search from a table or metric that can map modes to classes and operating conditions. In another embodiment, one or more modes can be selected in the spatial arrangement of a set of blocks in the 3D point cloud geometry 114 based on the modes used by blocks adjacent to the current block.

[0063]

[0071] If one or more modes include more than one mode, the processor 204 can determine the rate distortion cost associated with each of the selected modes and compare the determined rate distortion costs with one another. Based on the comparison of the determined rate distortion costs, the processor 204 can select the mode with the smallest rate distortion cost from the selected modes as the optimal mode to encode the current 3D block. In another scenario, if one or more modes include a single mode, the rate distortion cost of the mode cannot be determined. Instead, the single mode itself may be the optimal mode for encoding the current block. The determination of modes and the selection of one or more modes are explained in more detail in Figure 5.

[0064]

[0072] The encoder 410 can be configured to encode a block based on one or more selected modes. For example, the encoder 410 can encode a block based on one or more selected modes to obtain an encoded block 414.

[0065]

[0073] Figure 5 shows an exemplary processing pipeline for variable-rate compression of point cloud geometry according to an embodiment of the present disclosure. The description of Figure 5 will be made in relation to the elements of Figures 1, 2, 3 and 4. Referring to Figure 5, the processing pipeline 500 is shown. The processing pipeline 500 shows a sequence of operations 502-526. The sequence of operations can be executed by any computing device, such as circuit 202 of electronic device 102.

[0066]

[0074] In 502, a first block 502A (i.e., a 3D block) can be selected from a set of blocks 502B (i.e., 3D blocks) of 3D point cloud geometry. Circuit 202 can be configured to divide the 3D point cloud geometry into a set of blocks 502B. This selection can be part of an iterative selection process to find the optimal encoding mode and optimal RD operating point for each block. This search can be a bidirectional search across all modes corresponding to the set of RD operating points. During the search, circuit 202 (e.g., a block encoder) can select different RDs i By switching between operating points, it is possible to search for a mode that satisfies the set cost constraint (i.e., loss threshold).

[0067]

[0075] In 504, rate distortion (RD) i ) The operating point can be selected. RD i In selecting points, RD iThe operating point can be selected from a set of RD operating points (e.g., RD0 to RD4, as shown in Table 1). For example, RD0 (when i=0) can be selected from RD0 to RD4, as shown in Table 1. Each RD operating point can correspond to a specific rate distortion between the original point cloud block and the reconstructed point cloud block. The rate distortion can be based on the inter-point distance or inter-planar distance (or any other arbitrary objective or subjective distortion metric) between corresponding points in the original and reconstructed point cloud blocks, and the estimated number of bits required to encode the corresponding block. Each RD operating point in the set of RD operating points can be associated with one or more loss thresholds corresponding to one or more encoding modes for the RD operating point.

[0068]

[0076] In 506, a mode selection operation can be performed. In the mode selection operation, circuit 202 can be configured to select an encoding mode (M0, etc.) from one or more encoding modes (M0 to M4, etc. in Table 1) associated with the selected RD operating point (RD0). Such encoding modes can correspond to deep neural networks, each of which can be trained to encode a selected first block 502A of 3D point cloud geometry and generate an encoded first block. The selection of an encoding mode can be performed linearly. The starting index (j) of the selected encoding mode (M) can be set to 0, and the index can be incremented in further iterations. Note that the selected encoding mode should not be considered the final encoding mode used in the encoding operation of the first block 502A. In 508, the selected encoding mode (M0) should only be considered as a candidate encoding mode for the selected first block 502A.

[0069]

[0077] In 508, a loss calculation operation can be performed for the selected first block 502A. Circuit 202 can calculate a loss value for the selected first block 502A based on the selected encoding mode (e.g., M0). The loss value can be associated with a compression metric such as a rate metric or an MSE metric. According to one embodiment, circuit 202 can input the selected first block 502A to a classifier model. In such a case, the loss value can be calculated based on the output of the classifier model for the input. The classifier model can be, for example, a deep neural network (DNN) model trained on implicit or explicit geometric properties of at least one point cloud test block. Such properties can include, for example, the density of points associated with the point cloud.

[0070]

[0078] In step 510, the loss value calculated for the selected first block 502A can be compared with the loss threshold for the selected coding mode. During the comparison, it can be determined whether the loss calculated for the selected first block 502A is less than the threshold for the selected mode. If the loss calculated for the selected first block 502A is less than the threshold for the selected mode, control can proceed to step 512. If the loss calculated for the selected first block 502A is greater than the threshold for the selected mode, control can proceed to step 514.

[0071]

[0079] According to one embodiment, the loss threshold can be a different value for each coding mode in the set of RD operating points. For example, circuit 202 may be configured to acquire a calibration point cloud from computing device 110 and / or server 106. For a block of calibration point cloud, circuit 202 can calculate the first quartile of the loss value corresponding to each mode of the RD operating point in the set of RD operating points. Circuit 202 can set the first quartile of the loss value as the loss threshold corresponding to the coding mode. Examples of loss thresholds for different RD operating points are shown in Table 2 below. Loss threshold for RD point [Table 2] Table 2 shows RD i The various cost / loss thresholds for the operating point are indicated by "i" from 0 to 4. Of the RD operating points, RD4 is associated with a minimum loss threshold of 0.161, and RD0 is associated with a maximum loss threshold of 0.800.

[0072]

[0080] According to one embodiment, the circuit 202 can set a loss threshold for an encoding mode based on user input. User input can be provided to adjust the loss threshold for a specific RD operating point. For example, user input may request a change in the loss threshold for RD2 from 0.267 to 0.134. An example of loss threshold adjustment is shown in Table 3 below. Loss threshold for RD point [Table 3] Table 3 shows RD i The various cost / loss thresholds for the operating point are indicated by "i" from 0 to 4. Of the RD operating points, RD4 is associated with a minimum loss threshold of 0.080, and RD0 is associated with a maximum loss threshold of 0.400.

[0073]

[0081] According to one embodiment, the loss threshold can be a fixed value for each encoding mode corresponding to a set of RD operating points. Examples of fixed loss thresholds are shown in Table 4 below. Fixed loss threshold for encoding mode [Table 4]

[0074]

[0082] At 512, an encoding operation can be performed. As part of this operation, circuit 202 can encode a selected first block 502A based on a selected encoding mode. According to one embodiment, the selected encoding mode may correspond to a first deep neural network (e.g., DNN Model 1306A), and the selected first block 502A can be encoded based on applying the first deep neural network to the selected first block 502A. After 512, control can proceed to 524.

[0075]

[0083] In 514, an operation can be performed to determine whether other modes are available for the selected RD operating point (e.g., RD0). If other modes are available for the selected RD operating point (i.e., the selected mode (M j If (M) is not the last mode for the RD operating point, control can proceed to 516. If no other modes are available for the selected RD operating point (i.e., the selected mode (M) j If () is the last mode of the RD operating point, control can proceed to 518.

[0076]

[0084] In step 516, an operation can be performed to switch to the next mode (e.g., M1) associated with the selected RD operating point (e.g., RD0). This switch can be performed by incrementing the model index (j) value by 1. After the switch, the next mode (e.g., M1) can be selected, and operations 506-510 can be iteratively performed for the next mode until the loss value exceeds the loss threshold for the next mode.

[0077]

[0085] After iterating through all modes of the (single and multiple) RD operating points, a set of loss values ​​can be obtained for the selected first block 502A. Circuit 202 can calculate a first set of loss values ​​for the selected first block 502A that can be associated with one or more compression metrics. The first set of loss values ​​is RD i It can accommodate a set of coding modes associated with at least a subset of the set of operating points. For example, there may be 12 modes (4 modes per RD point) from RD0 to RD2 (i.e., a subset of 5 RD operating points), and the first loss value set may contain 12 loss values. The first loss value set includes one or more first loss values ​​that can be calculated for one or more coding modes corresponding to a first RD operating point (e.g., RD0) in the set of RD operating points. The number of loss calculations for a selected block may vary and may depend on the value of the loss threshold that can be set for the mode.

[0078]

[0086] In 518, from the set of RD operating points, other RD i It is possible to perform an operation to determine whether the operating point is selectable. i If the operating point is selectable, control can proceed to 520. Other RD i If no operating point is selectable (i.e., the RD operating point selected in 504 is the last RD operating point in the set of RD operating points), control can proceed to 522.

[0079]

[0087] In 520, an operation can be performed to switch to the next RD operating point (e.g., RD1) in the set of RD operating points (e.g., RD0...RD4). This switch can be performed by incrementing the value of RD index (i) by 1. After the switch, in 504, the next RD operating point (e.g., RD1) can be selected, and operations 506-510 can be iteratively performed for the next mode until the loss value exceeds the loss threshold for the next mode. For example, circuit 202 can switch to the second RD operating point (RD1) in the set of RD operating points based on the determination that one or more first loss values ​​are greater than one or more loss thresholds for one or more encoding modes corresponding to the first RD operating point (e.g., modes such as M0-M4 of RD0). The first set of loss values ​​(described in 516) may include one or more second loss values ​​that can be calculated for one or more encoding modes corresponding to the second RD operating point (e.g., RD1).

[0080]

[0088] In 522, a reversible coding operation can be performed. As part of this operation, a raw block coding mode (LL mode) can be turned on to address cases where none of the RD operating points (selected in 504) can meet the local quality criteria for coding the selected first block 502A (i.e., the loss values ​​for all RD operating points exceed the loss threshold), and where undesirable artifacts (such as holes) are detected in the reconstruction of the point cloud geometry during local reconstruction at the coding stage. The first block 502A can be stored reversibly. Specifically, circuit 202 can select a reversible coding scheme for the first block 502A based on the determination that each loss value in a first set of loss values ​​(described in 516) exceeds the loss threshold for the corresponding coding mode in a set of coding modes. Based on the selected reversible coding scheme, circuit 202 can code the selected first block 502A. An example of the application of a reversible coding scheme is shown, for example, in Figure 9.

[0081]

[0089] In step 524, an operation can be performed to determine whether the selected first block 502A is the last block in the set of blocks 502B. If it is determined that the selected first block 502A is the last block, the circuit 202 can prepare an encoded bitstream 528 associated with the point cloud geometry and send it to the computing device 110 or the server 106. Control can then proceed to termination. If it is determined that the selected first block 502A is not the last block, control can proceed to step 526.

[0082]

[0090] In step 526, the operation to select the next block (for example, the second block) from the set of blocks can be performed. After the selection, operations 504-524 are repeated until the last block in the set of blocks is processed, allowing the next block and subsequent blocks to be processed.

[0083]

[0091] Figure 6 is a diagram illustrating an exemplary search pattern for modes of RD operating points according to embodiments of the present disclosure. Referring to Figure 6, an exemplary search pattern for modes of RD operating points is shown in Figure 600. The description of Figure 6 will be made in relation to the elements of Figures 1, 2, 3, 4 and 5. Figure 600 shows a table of five RD operating points, RD0 to RD4, and five corresponding modes, Mode 0 to Mode 4. A search can be performed (as described in Figure 5) to find the optimal coding mode at a particular RD operating point. The search can start from RD0, Mode 0, and (as indicated by the double-headed arrow) RD i All within This can proceed as a bidirectional search across JPEG2026515763000007.jpg8170. The number of steps to find the optimal encoding mode can vary and can typically depend on the value of the loss threshold that can be set for the mode.

[0084]

[0092] For a given RD operating point (e.g., RD0) JPEG2026515763000008.jpg8170~ If none of the JPEG2026515763000009.jpg8170 encoding requirements are met (i.e., the loss value for the block exceeds the loss threshold for all modes), circuit 202 can switch to the next RD operating point in the set of RD operating points. For the selected block, the search is performed if the loss value for all modes j A mode that can remain below the loss threshold for j and RD i This can result in a pair of values. The loss threshold for a given encoding mode can be a static value or can be dynamically adjusted based on human input or target rate distortion or quality.

[0085]

[0093] Figure 7 is a diagram illustrating an exemplary comparison between irreversible and reversible reconstruction outputs of a point cloud geometry according to embodiments of the present disclosure. The elements of Figure 7 are described in relation to the elements of Figures 1, 2, 3, 4, 5, and 6. Referring to Figure 7, an exemplary comparison 700 between irreversible and reversible reconstruction outputs of a point cloud geometry is shown. Comparison 700 shows point cloud geometry 702, point cloud geometry 704, and point cloud geometry 706. Point cloud geometry 702 represents a reference (original) point cloud of a human head. Point cloud geometry 704 can be an irreversible reconstruction of encoded point cloud data that can be obtained from point cloud geometry 702 after applying encoder 208A to a block of point cloud geometry 702. For example, the ML model can be configured to encode a block of point cloud based on RD points and the optimal mode corresponding to the RD points. The selection of RD points and optimal modes is described, for example, in Figure 5. Point cloud geometry 704 shows region 704A, which contains a hole artifact corresponding to region 702A of point cloud geometry 702. Hole artifacts can occur due to irreversible reconstruction. If none of the RD points meet the local quality criteria for a particular block (such as the block in region 702A), and an undesirable artifact (as shown in 704A) is detected during local reconstruction at the encoding stage, the "raw block" encoding mode (LL mode) can be turned on. Using LL mode, the block in region 702A can be reversibly encoded to ensure that such undesirable artifacts are not present in local reconstruction. Point cloud geometry 706 shows region 706A, which is reconstructed from the reversibly encoded block, corresponding to regions 704A and 702A. As shown in the figure, region 706A does not contain any visible artifacts such as holes.

[0086]

[0094] Figure 8 shows an exemplary 3D point cloud geometry and the selection of a region of interest (ROI) within the point cloud geometry for point cloud compression, according to an embodiment of the present disclosure. Referring to Figure 8, a 3D point cloud geometry 800 is shown, including parts 802, 804, 806, and 808. The elements of Figure 8 will be described in relation to the elements of Figures 1, 2, 3, 4, 5, 6, and 7.

[0087]

[0095] During operation, the electronic device 102 can determine a portion of the 3D point cloud geometry as a region of interest (ROI). This determination can be made based on input from a user (e.g., a 3D artist) or on an automated operation such as an object detection operation or a semantic segmentation operation performed on the 3D point cloud geometry 800. For example, the user can explicitly define the ROI through different slices that divide the 3D point cloud geometry 800 into portions 802, 804, 806, and 808. The user can further assign a mode or RD point to each portion, such as reversible mode (LL mode) for portion 802, RD4 for portion 804, RD0 for portion 806, and RD3 for portion 808. The electronic device 102 can encode the block corresponding to portion 802 of the 3D point cloud geometry 800 based on a reversible encoding scheme. Blocks corresponding to other parts of the 3D point cloud geometry 800 (such as parts 804, 806, and 808) can be encoded in the optimal mode associated with the assigned RD points. Such modes can be identified using the operation described in Figure 5, for example.

[0088]

[0096] Figure 9 is a flowchart illustrating an exemplary operation for variable-rate compression of point cloud geometry according to an embodiment of the present disclosure. Referring to Figure 9, flowchart 900 is shown. The description of flowchart 900 will be made in relation to the elements of Figures 1, 2, 3, 4, 5, 6, 7, and 8. Operations 902-916 can be implemented on the electronic device 102. The method described in flowchart 900 can be started from 902 and proceeded to 904.

[0089]

[0097] In 904, a set of RD operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points can be stored. The electronic device 102 may include a memory 210 that can be configured to store a set of RD operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points.

[0090]

[0098] At 906, a 3D point cloud geometry (e.g., 3D point cloud geometry 114) can be received. In one embodiment, circuit 202 can be configured to receive the 3D point cloud geometry 114. The 3D point cloud geometry 114 can be received from the scanning setup 104 or server 106 via the communication network 112. The reception of the 3D point cloud geometry is further illustrated, for example, in Figure 4.

[0091]

[0099] In 908, the 3D point cloud geometry 114 can be divided into a set of blocks (e.g., a block stream 412). In one embodiment, the circuit 202 can be configured to divide the 3D point cloud geometry 114 into a set of blocks. The division of the 3D point cloud geometry is further illustrated, for example, in Figure 4.

[0092]

[0100] In 910, a first block 502A can be selected from the set of blocks 502B. In one embodiment, the circuit 202 can be configured to select a first block 502A from the set of blocks 502B.

[0093]

[0101] In 912, a first set of loss values ​​associated with one or more compression metrics can be calculated for the selected first block 502A. The first set of loss values ​​may correspond to a set of coding modes associated with at least a subset of the set of RD operating points. In one embodiment, the circuit 202 may be configured to calculate the first set of loss values ​​for the selected first block 502A. The calculation of loss values ​​is further illustrated, for example, in Figures 1 and 5.

[0094]

[0102] In 914, for the selected first block 502A, an encoding mode can be selected from a set of encoding modes in which the loss value from the first set of loss values ​​is below the loss threshold. In one embodiment, the circuit 202 can be configured to select an encoding mode from a set of encoding modes.

[0095]

[0103] In 916, the first block 502A can be encoded based on the selected encoding mode. In one embodiment, circuit 202 can be configured to encode the first block 502A based on the selected encoding mode. The encoding of the first block 502A is further illustrated, for example, in Figure 5. Control can then proceed to termination.

[0096]

[0104] Various embodiments of this disclosure can provide a non-temporary computer-readable medium and / or storage medium storing computer-executable instructions that can be executed by a machine and / or computer to operate an electronic device (e.g., electronic device 102 in Figure 1). Such instructions can cause the electronic device 102 to perform an operation which may include storing a set of RD operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points. The operation may further include receiving a 3D point cloud geometry relating to one or more objects in 3D space and dividing the 3D point cloud geometry into a set of blocks. The operation may further include selecting a first block from the set of blocks and calculating a first set of loss values ​​associated with one or more compression metrics for the selected first block. The first set of loss values ​​may correspond to a set of encoding modes associated with at least a subset of the set of RD operating points. The operation may further include selecting an encoding mode from a set of encoding modes in which the loss value among a first set of loss values ​​is below a loss threshold for the encoding mode, and encoding the selected first block based on the selected encoding mode.

[0097]

[0105] An exemplary aspect of the present disclosure provides an electronic device (such as electronic device 102) that can include a memory (such as memory 210) which can be configured to store a set of RD operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points. The electronic device may further include a circuit (such as circuit 202) which can be configured to receive 3D point cloud geometry relating to one or more objects in 3D space. Circuit 202 may be further configured to divide the 3D point cloud geometry into a set of blocks and select a first block from the set of blocks. For the selected first block, Circuit 202 may be configured to calculate a first set of loss values ​​associated with one or more compression metrics. The first set of loss values ​​may correspond to a set of encoding modes associated with at least a subset of the set of RD operating points. Circuit 202 may be further configured to select from the set of encoding modes an encoding mode whose loss value in the first set of loss values ​​is below a loss threshold for the encoding mode. Circuit 202 may then encode the selected first block based on the selected encoding mode.

[0098]

[0106] According to one embodiment, one or more compression metrics may include a rate metric or a mean squared error (MSE) metric.

[0099]

[0107] According to one embodiment, the circuit 202 can be further configured to input a selected first block to a classifier model. A first set of loss values ​​can be further calculated based on the output of the classifier model for the input. The classifier model can be a deep neural network (DNN) model trained on one or more geometric properties of the test block of the point cloud. Such properties may include the density of points associated with the point cloud.

[0100]

[0108] According to one embodiment, the circuit 202 can be further configured to determine a portion of the 3D point cloud geometry as a region of interest (ROI) and to encode blocks corresponding to the determined portion of the 3D point cloud geometry based on a reversible coding scheme. The portion of the 3D point cloud geometry can be determined as an ROI based on at least one of user input, object detection operations, or semantic segmentation operations.

[0101]

[0109] According to one embodiment, each RD operating point in a set of RD operating points can be associated with one or more loss thresholds corresponding to one or more encoding modes.

[0102]

[0110] According to one embodiment, the encoding mode can correspond to a deep neural network, each deep neural network can be trained to encode a selected first block of 3D point cloud geometry and generate an encoded first block. The selected first block can be encoded based on applying a first deep neural network of deep neural networks to the selected first block. The first deep neural network can correspond to the selected encoding mode.

[0103]

[0111] According to one embodiment, the first loss value set may include one or more first loss values ​​that can be calculated for one or more coding modes corresponding to a first RD operating point in the set of RD operating points. The circuit 202 may be further configured to switch to a second RD operating point in the set of RD operating points based on a determination that one or more first loss values ​​are greater than one or more loss thresholds for one or more coding modes corresponding to the first RD operating point. The first loss value set may include one or more second loss values ​​that can be calculated for one or more coding modes corresponding to a second RD operating point.

[0104]

[0112] According to one embodiment, the circuit 202 may be further configured to acquire a calibration point cloud and, for blocks of the calibration point cloud, calculate a first quartile of the loss value corresponding to each mode of the RD operating point in the set of RD operating points. The first quartile of the loss value corresponding to the coding mode may be set as the loss threshold.

[0105]

[0113] According to one embodiment, the loss threshold can be a fixed value for each coding mode corresponding to a set of RD operating points.

[0106]

[0114] According to one embodiment, the circuit 202 can be further configured to set a loss threshold for the encoding mode based on user input.

[0107]

[0115] According to one embodiment, the circuit 202 may be further configured to select a reversible coding scheme for a first block based on a determination that each loss value in a first set of loss values ​​exceeds a loss threshold for a corresponding coding mode in a set of coding modes. The circuit 202 can then encode the selected first block based on the selected reversible coding scheme.

[0108]

[0116] This disclosure can be implemented in hardware form or in a combination of hardware and software. This disclosure can be implemented centrally within at least one computer system or in a distributed manner, where different elements are distributed across multiple interconnected computer systems. A computer system or other device adapted to perform the methods described herein may be suitable. The hardware and software combination may be a general-purpose computer system including a computer program that, when loaded and executed, can control the computer system to perform the methods described herein. This disclosure can also be implemented in hardware form, including part of an integrated circuit that also performs other functions.

[0109]

[0117] This disclosure includes all features that enable the implementation of the methods described herein and can be incorporated into a computer program product that can perform these methods when loaded onto a computer system. In this context, a computer program means any expression in any language, code, or notation of an instruction set intended to be executed by a system having information processing capabilities, either directly or after either a) conversion to another language, code, or notation, or b) reproduction in a different form of content.

[0110]

[0118] While this disclosure has been described with reference to specific embodiments, those skilled in the art will understand that various modifications can be made and equivalents can be substituted without departing from the scope of this disclosure. Furthermore, many modifications can be made to adapt the teachings of this disclosure to specific circumstances or content without departing from the scope of this disclosure. Therefore, this disclosure is not limited to the embodiments disclosed, but is intended to include all embodiments that fall within the claims. [Explanation of symbols]

[0111] 100 Network Environment 102 Electronic Devices 104 Scanning Setup 106 servers 108 Databases 110 Computing Devices 112 Communication Networks 114 3D Point Cloud Geometry 200 Block Diagram 202 circuits 204 Processors 206 Classifier Models 208 codecs 208A Encoder 208B Decoder 210 memory 212 Input / Output (I / O) Devices 212A Display Device 214 Network Interfaces 300 Block Diagram 302A Encoder 302B Decoder 304A Encoder 1 304N Encoder N 306A DNN Model 1 306N DNN Model N 308 Mode Selector 310A Decoder 1 310N Decoder N 312A Block division section 312B Binarization and Merging Section 314A encoded bitstream and supplementary information 314B signaling bitstream 316A Input Point Cloud 316N Reconstructed Point Cloud Set of 318 3D blocks 400 components 402 Block division section 404 Classifier Models 406 Loss calculation section 408 Mode Selector 410 encoder 412 Blockstream 414 encoded blocks 500 processing pipelines Select block 502 502A Block 1 502B Block Set 504 (RD i ) Selection 506 RD i mode (M j Select ) 508 Calculate the loss value 510 Is the loss a threshold for the mode? 512 encoding 514 Selected M j RD i Is this the final mode? 516 Next j+1 518 Selected RD i Is that the last RD point? 520 Next i+1 522 lossless encoding 524 Is the selected block the last block? 526 Select the next block 528 encoded bitstream Figure 600 700 comparison 702 Point Cloud Geometry 702A area 704 Point Cloud Geometry 704A area 706 Point Cloud Geometry 706A area 800 3D point cloud geometries 802 parts 804 parts 806 parts 808 parts 900 Flowcharts 902 start 904 Stores a set of rate distortion (RD) operating points and one or more encoding modes associated with each RD operating point in the set. Receive 906 3D point cloud geometries Divide 908 3D point cloud geometries into sets of blocks. Select the first block from the set of 910 blocks. 912 For the selected first block, calculate a first set of loss values ​​associated with one or more compression metrics. From the set of 914 coding modes, select a coding mode whose loss value in the first set of loss values ​​falls below the loss threshold for that coding mode. 916 Encode the selected first block based on the selected encoding mode.

Claims

1. It is an electronic device, A memory configured to store a set of rate-distortion (RD) operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points, Receiving 3D point cloud geometry, Dividing the aforementioned 3D point cloud geometry into a set of blocks, Selecting a first block from the aforementioned set of blocks, For the selected first block, the calculation of a first set of loss values ​​associated with one or more compression metrics, The first set of loss values ​​corresponds to a set of coding modes associated with at least a subset of the set of RD operating points, From the set of encoding modes, select an encoding mode in which the loss value among the first set of loss values ​​is below the loss threshold for the encoding mode. Encoding the selected first block based on the selected encoding mode, A circuit configured to perform the following: An electronic device characterized by including

2. The electronic device according to claim 1, characterized in that the one or more compression metrics include a rate metric or a mean squared error (MSE) metric.

3. The electronic device according to claim 1, wherein the circuit is further configured to input the selected first block to a classifier model, and the first loss value set is further calculated based on the output of the classifier model for the input.

4. The electronic device according to claim 3, characterized in that the classifier model is a deep neural network (DNN) model trained on one or more geometric properties of a point cloud test block.

5. The electronic device according to claim 4, characterized in that the one or more geometric properties include the density of points associated with the point cloud.

6. The aforementioned circuit is A portion of the aforementioned 3D point cloud geometry is determined as a region of interest (ROI), Based on a lossless coding scheme, the blocks corresponding to the determined portion of the 3D point cloud geometry are coded. The electronic device according to claim 1, further characterized by being configured in such a way.

7. The electronic device according to claim 6, characterized in that the portion of the 3D point cloud geometry is determined as the ROI based on at least one of user input, object detection operations, or semantic segmentation operations.

8. The electronic device according to claim 1, characterized in that each RD operating point in the set of RD operating points is associated with one or more loss thresholds corresponding to one or more encoding modes.

9. The electronic device according to claim 1, characterized in that the encoding mode corresponds to a deep neural network, and each deep neural network is trained to encode the selected first block of the 3D point cloud geometry to generate an encoded first block.

10. The electronic device according to claim 9, characterized in that the selected first block is encoded by applying a first deep neural network from the deep neural networks to the selected first block, and the first deep neural network corresponds to the selected encoding mode.

11. The electronic device according to claim 1, characterized in that the first loss value set includes one or more first loss values ​​calculated for one or more encoding modes corresponding to a first RD operating point among the set of RD operating points.

12. The circuit is further configured to switch to a second RD operating point in the set of RD operating points based on a determination that the one or more first loss values ​​are greater than one or more loss thresholds for the one or more encoding modes corresponding to the first RD operating point. The electronic device according to claim 11, characterized in that the first loss value set includes one or more second loss values ​​calculated for one or more encoding modes corresponding to the second RD operating point.

13. The aforementioned circuit is Obtain a calibration point cloud, For each block of the calibration point cloud, calculate the first quartile of the loss value corresponding to each mode of the RD operating point in the set of RD operating points. The first quartile of the loss value corresponding to the encoding mode is set as the loss threshold. The electronic device according to claim 1, further characterized by being configured in such a way.

14. The electronic device according to claim 1, characterized in that the loss threshold is a fixed value for each encoding mode corresponding to the set of RD operating points.

15. The electronic device according to claim 1, wherein the circuit is further configured to set the loss threshold for the encoding mode based on user input.

16. The aforementioned circuit is Based on the determination that each loss value in the first set of loss values ​​exceeds the loss threshold for the corresponding encoding mode in the set of encoding modes, a lossless encoding scheme is selected for the first block. The selected first block is encoded based on the selected reversible encoding scheme. The electronic device according to claim 1, further characterized by being configured in such a way.

17. It is a method, In electronic devices, A step of storing a set of rate distortion (RD) operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points, Steps include receiving a 3D point cloud geometry, The steps include dividing the 3D point cloud geometry into a set of blocks, The steps include selecting a first block from the set of blocks, A step of calculating a first set of loss values ​​associated with one or more compression metrics for the selected first block, The first set of loss values ​​corresponds to a set of coding modes associated with at least a subset of the set of RD operating points, and includes steps The steps include selecting an encoding mode from the set of encoding modes in which the loss value among the first set of loss values ​​is below the loss threshold for the encoding mode, The steps include encoding the selected first block based on the selected encoding mode, A method characterized by including the following.

18. A step of selecting a lossless encoding scheme for a first block based on the determination that each loss value in the first set of loss values ​​exceeds the loss threshold for the corresponding encoding mode in the set of encoding modes, The steps include encoding the selected first block based on the selected reversible encoding scheme, The method according to claim 17, further comprising:

19. A non-temporary computer-readable medium in which computer-executable instructions causing an electronic device to perform an action when executed by the electronic device are stored, wherein the action is: The system stores a set of rate-distortion (RD) operating points and one or more encoding modes associated with each RD operating point in the set of RD operating points. Receiving 3D point cloud geometry, Dividing the aforementioned 3D point cloud geometry into a set of blocks, Selecting a first block from the aforementioned set of blocks, For the selected first block, the calculation of a first set of loss values ​​associated with one or more compression metrics, The first set of loss values ​​corresponds to a set of coding modes associated with at least a subset of the set of RD operating points, From the set of encoding modes, select an encoding mode in which the loss value among the first set of loss values ​​is below the loss threshold for the encoding mode. Encoding the selected first block based on the selected encoding mode, A non-temporary computer-readable medium characterized by including [a specific element].

20. The aforementioned operation is, A lossless encoding scheme for a first block is selected based on the determination that each loss value in the first set of loss values ​​exceeds the loss threshold for the corresponding encoding mode in the set of encoding modes. Encoding the selected first block based on the selected reversible encoding scheme, A non-temporary computer-readable medium according to claim 19, characterized by including the following: