Motion Compensation-Based Neighbor Construction for TriSoup Centroid Information

The point cloud encoding system addresses the challenge of large data size through context-based encoding and dynamic occupancy trees, achieving efficient compression for diverse applications with reduced bitrate and maintained visual quality.

JP2026502870APending Publication Date: 2026-01-27COMCAST CABLE COMM LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025536668
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-22
Filing Date
2023-12-22
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

The large data size of point clouds necessitates efficient compression techniques for storage and transmission, as raw representations require significant bandwidth and storage resources, and existing methods often compromise visual quality or necessitate lossless compression for critical applications.

Method used

A point cloud encoding system utilizing a context-based encoding of centroid residual values and a dynamic occupancy tree structure, combined with a dynamic OBUF scheme, to achieve efficient lossy or lossless compression.

Benefits of technology

The system achieves high compression ratios with reduced bitrate, maintaining visual quality and enabling practical deployment in applications like AR, VR, and autonomous driving, while supporting both lossy and lossless compression scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502870000001_ABST
    Figure 2026502870000001_ABST
Patent Text Reader

Abstract

One or more methods, devices, computer-readable storage media, and systems are disclosed for entropy encoding edge vertex information in a voxelized space of a point cloud. The symbols of the current edge's neighborhood may be determined based on one or more previously encoded edges. The previously encoded edges may be selected from the edge's spatial topology. The use of motion-compensated point clouds to encode centroid residual values ​​may enhance inter-frame correlation used to determine context or probability models. This increased correlation may improve coder selection and result in enhanced compression of the centroid residual values.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 434,526, filed December 22, 2022. The above-referenced application is incorporated herein by reference in its entirety. [Background technology]

[0002] An object or scene can be described using volumetric visual data consisting of a series of points. The points can be stored in a point cloud format, which includes a collection of points in three-dimensional space. Because point clouds can be very large in data size, transmitting and processing point cloud data can require data compression schemes specifically designed for the unique characteristics of point cloud data. Summary of the Invention

[0003] The following summary provides a simplified overview of certain features. It is not an extensive overview and is not intended to identify key or critical elements.

[0004] Coding (e.g., encoding, decoding) may be used to compress and decompress point cloud frames or sequences for efficient storage and transmission. A point cloud coding system may include a source device that may encode a point cloud sequence into a bitstream. The point cloud coding system may include a transmission medium that transmits the encoded bitstream. The point cloud coding system may include a destination device that may obtain a decoded point cloud sequence based on the encoded bitstream. Visual data (e.g., centroid residual values) may be coded (e.g., encoded or decoded) based on a context (or probability model). The context (or probability model) may be selected, for example, based on a lookup table for coding the centroid residual values. A second centroid vertex of the TriSoup may be determined based on the coded centroid residual values ​​and the first centroid vertex of the TriSoup. Using the second centroid vertex as a pivot vertex may reduce reconstruction error and / or distortion.

[0005] These and other features and advantages are described in more detail below. [Brief explanation of the drawings]

[0006] Certain features are illustrated by way of example, and not by way of limitation, in the accompanying drawings in which like numerals refer to like elements and in which:

[0007] [Figure 1] 1 illustrates an exemplary point cloud encoding system. [Figure 2] The Morton order of eight sub-cuboids divided from a cuboid is shown. [Figure 3] 1 shows an example of an occupancy tree scan order. [Figure 4] 10 illustrates an exemplary neighborhood of cuboids for entropy coding the occupancy of a child cuboid. [Figure 5] An example of a dynamic shrink function (DR) that can be used in a dynamic OBUF is shown below. [Figure 6]10 illustrates an exemplary method for encoding cuboid occupancy using dynamic OBUFs. [Figure 7] An example of an occupied cuboid corresponding to a TriSoup node in the occupied tree is shown below. [Figure 8A] 1 shows an exemplary cuboid corresponding to a TriSoup node. [Figure 8B] 1 shows an exemplary refinement to the TriSoup model. [Figure 9] An example of voxelization is shown below. [Figure 10A] Denotes a rectangular parallelepiped whose volume intersects with the current TriSoup edge being entropy encoded. [Figure 10B] Denotes a rectangular parallelepiped whose volume intersects with the current TriSoup edge being entropy encoded. [Figure 11A] Indicates a TriSoup edge that can be used to entropy encode the current TriSoup edge. [Figure 11B] Indicates a TriSoup edge that can be used to entropy encode the current TriSoup edge. [Figure 11C] Indicates a TriSoup edge that can be used to entropy encode the current TriSoup edge. [Figure 12] 1 illustrates an exemplary encoding method. [Figure 13] An example of encoding centroid residual values ​​is shown below. [Figure 14A] 1 shows an example of a TriSoup node and compensated points belonging to a motion compensated point cloud. [Figure 14B] 1 shows an example of a TriSoup node and compensated points belonging to a motion compensated point cloud. [Figure 15] 1 shows an exemplary TriSoup node and compensated points belonging to a motion compensated point cloud. [Figure 16A] 10 illustrates an exemplary method for encoding centroid residual values ​​for a TriSoup node. [Figure 16B] 10 illustrates an exemplary method for decoding centroid residual values ​​of a TriSoup node. [Figure 17] 1 illustrates an exemplary computer system that may use any of the embodiments described herein. [Figure 18] 1 illustrates exemplary elements of a computing device that may be used to implement any of the various devices described herein. DETAILED DESCRIPTION OF THE INVENTION

[0008] The accompanying drawings and description provide examples. It should be understood that the embodiments shown in the drawings and / or description are non-exclusive and that the features shown and described may be practiced in other embodiments. Examples are provided for the operation of a point cloud or point cloud sequence encoding or decoding system. More specifically, the techniques disclosed herein may relate to point cloud compression for use in encoding and / or decoding devices and / or systems.

[0009] Visual data may describe an object or scene using a series of points. Each point may include a two-dimensional (x and y) position and one or more optional attributes, such as color. Volumetric visual data may add another positional dimension to this visual data. Volumetric visual data may describe an object or scene using a series of points, each including a three-dimensional (x, y, and z) position and one or more optional attributes, such as color, reflectance, timestamp, etc. Volumetric visual data may, for example, provide a more immersive way to experience visual data than traditional visual data.

[0010] For example, an object or scene described by volumetric visual data can be viewed from any angle (or multiple angles), whereas traditional visual data is generally only viewable from the angle at which it was captured or rendered. Volumetric visual data can be used in many applications, including augmented reality (AR), virtual reality (VR), and mixed reality (MR). Scattered volumetric visual data can be used in the automotive industry for the representation of three-dimensional (3D) maps (e.g., cartography) or as input to advanced driver assistance systems. For advanced driver assistance systems, volumetric visual data can typically be input into driving decision algorithms. Volumetric visual data can be used to store valuable objects in digital form. In applications for preserving cultural heritage, the goal can be to preserve representations of objects that may be threatened by natural disasters. For example, statues, vases, and temples can be scanned in their entirety and stored as volumetric visual data with billions of samples. This use case for volumetric visual data can be particularly relevant for valuable objects in locations where earthquakes, tsunamis, and typhoons occur frequently. Volumetric visual data can take the form of a volumetric frame. A volumetric frame may describe an object or scene captured at a particular time instance. Volumetric visual data may take the form of a sequence of volumetric frames (called a volume sequence or volumetric video). A sequence of volumetric frames may describe an object or scene captured at multiple different time instances.

[0011] A point cloud is a format for storing volumetric visual data. A point cloud may include a collection of points in 3D space. Each point in the point cloud may include geometric shape information that indicates the point's location in 3D space. The geometric shape information may indicate the point's location in 3D space, for example, using three Cartesian coordinates (x, y, and z) or using spherical coordinates (r, phi, theta) (e.g., when acquired by a rotational sensor). The positions of points in a point cloud may be quantified according to spatial precision. The spatial precision may be the same or different in each dimension. The quantization process may generate a grid in 3D space. One or more points residing within each subgrid volume may be mapped to subgrid center coordinates called voxels. A voxel may be considered a 3D extension of a pixel corresponding to a 2D image grid coordinate. Points in a point cloud may further include one or more types of attribute information. The attribute information may indicate characteristics of the point's visual appearance. The attribute information may indicate, for example, the texture (e.g., color) of the point, the material type of the point, transparency information of the point, reflectance information of the point, a normal vector to the surface of the point, the velocity of the point, the acceleration at the point, a timestamp indicating when the point was captured, or a modality (e.g., running, walking, or flying) indicating how the point was captured. Points in the point cloud may include light field data in the form of multiple view-dependent texture information. The light field data may be another type of arbitrary attribute information.

[0012] Points in a point cloud may describe an object or scene. The points in a point cloud may, for example, describe the exterior surface and / or interior structure of an object or scene. The object or scene may be synthetically generated by a computer. The object or scene may be generated from capturing a real-world object or scene. Geometry information of a real-world object or scene may be obtained by 3D scanning and / or photogrammetry. 3D scanning may include different types of scanning, such as laser scanning, structured light scanning, and / or modulated light scanning. 3D scanning may obtain the geometry information. 3D scanning may obtain the geometry information, for example, by moving one or more laser heads, structured light cameras, and / or modulated light cameras relative to the object or scene being scanned. Photogrammetry may obtain the geometry information. Photogrammetry may obtain the geometry information, for example, by triangulating the same features or points in different spatially shifted 2D photographs. Point cloud data may take the form of a point cloud frame. A point cloud frame may describe an object or scene captured at a particular time instance. Point cloud data may take the form of a sequence of point cloud frames, which may be referred to as a point cloud sequence or a point cloud video, which may describe an object or scene captured at multiple different instances of time.

[0013] The data size of a point cloud frame or point cloud sequence may be too large for storage and / or transmission in many applications. A single point cloud may, for example, contain more than one million points, or even more than one billion points. Each point may include geometric information and one or more types of attribute information. The geometric information for each point may include, for example, three Cartesian coordinates (x, y, and z) or spherical coordinates (r, phi, and theta), each represented using at least 10 bits per component or a total of 30 bits. The attribute information for each point may include texture corresponding to three color components (e.g., R, G, and B color components). Each color component may be represented using, for example, 8 to 10 bits per component or a total of 24 to 30 bits. Thus, a single point may contain at least 54 bits of information, with at least 30 bits of geometric information and at least 24 bits of texture, in this example. If a point cloud frame contains 1 million such points, each point cloud frame may require 54 million bits or 54 megabits to represent. For a dynamic point cloud that changes over time, a data rate of 1.32 gigabits per second may be required to transmit (e.g., transmit) the points of a point cloud sequence at a frame rate of 30 frames per second. Thus, a raw representation of a point cloud may require a large amount of data, and practical deployment of point cloud-based technologies may require compression techniques that enable the storage and distribution of point clouds at a reasonable cost.

[0014] Encoding may be used to compress and / or reduce the data size of a point cloud frame or sequence to provide more efficient storage and / or transmission. Decoding may be used to decompress a compressed point cloud frame or sequence for display and / or other forms of consumption (e.g., by a machine learning-based device, a neural network-based device, an artificial intelligence-based device, or other types of machine-based processing algorithms and / or devices). Compression of the point cloud may be lossy (introducing differences to the original data) for distribution to and visualization by an end user, for example, on AR or VR glasses or any other 3D-enabled device. Lossy compression may enable high compression ratios but may imply a trade-off between compression and visual quality perceived by the end user. Other frameworks, such as those for medical applications or autonomous driving, may require lossless compression to avoid altering the transmission (e.g., transmission) and results of decisions made based on analysis of the decompressed point cloud frames.

[0015] FIG. 1 illustrates an exemplary point cloud encoding (e.g., encoding and / or decoding) system (e.g., point cloud encoding system 100). The point cloud encoding system 100 may include a source device 102, a transmission medium 104, and a destination device 106. The source device 102 may encode a point cloud sequence 108 into a bitstream 110 for more efficient storage and / or transmission. The source device 102 may store and / or transmit (e.g., transmit) the bitstream 110 to the destination device 106 via the transmission medium 104. The destination device 106 may decode the bitstream 110 to display the point cloud sequence 108 or for other forms of consumption (e.g., further analysis, storage, etc.). The destination device 106 may receive the bitstream 110 from the source device 102 via the storage medium or transmission medium 104. The source device 102 and the destination device 106 may include any amount (e.g., number) of different devices. The source device 102 and the destination device 106 may include, for example, a cluster of interconnected computer systems acting as a seamless pool of resources (also called a cloud of computers or cloud computing), a server, a desktop computer, a laptop computer, a tablet computer, a smartphone, a wearable device, a television, a camera, a video game console, a set-top box, a video streaming device, a vehicle (e.g., an autonomous vehicle), or a head-mounted display. The head-mounted display may allow a user to view a VR, AR, or MR scene and adjust the view of the scene based on the user's head movements. The head-mounted display may be tethered to a processing device (e.g., a server, desktop computer, set-top box, or video game console) or may be completely self-contained.

[0016] The source device 102 may include a point cloud source 112, an encoder 114, and an output interface 116. To encode the point cloud sequence 108 into a bitstream 110, the source device 102 may include the point cloud source 112, the encoder 114, and the output interface 116. The point cloud source 112 may provide or generate the point cloud sequence 108 from the capture of natural and / or synthetically generated scenes. The synthetically generated scenes may be scenes including computer-generated graphics. The point cloud source 112 may include one or more point cloud capture devices, a point cloud archive containing previously captured natural and / or synthetically generated scenes, a point cloud feed interface for receiving captured natural and / or synthetically generated scenes from a point cloud content provider, and / or a processor for generating the synthesized point cloud scenes. The point cloud capture devices may include, for example, one or more laser scanning devices, structured light scanning devices, modulated light scanning devices, and / or passive scanning devices.

[0017] As shown in FIG. 1 , a point cloud sequence 108 may include a series of point cloud frames 124. A point cloud frame may describe an object or scene captured at a particular time instance. The point cloud sequence 108 may achieve the impression of motion by sequentially presenting the point cloud frames 124 of the point cloud sequence 108 using a constant or variable time. A point cloud frame may include a collection of points (e.g., voxels) 126 in 3D space. Each point 126 may include geometric shape information that indicates the point's location in 3D space. The geometric shape information may indicate the point's location in 3D space using, for example, three Cartesian coordinates (x, y, and z). One or more of the points 126 may further include one or more types of attribute information. The attribute information may indicate characteristics of the point's visual appearance. The attribute information may indicate, for example, the texture (e.g., color) of the point, the material type of the point, transparency information of the point, reflectance information of the point, a normal vector relative to the surface of the point, the velocity of the point, the acceleration at the point, a timestamp indicating when the point was captured, and a modality (e.g., running, walking, or flying) indicating how the point was captured. One or more of the points 126 may include light field data, for example, in the form of multiple view-dependent texture information. The light field data may be any other type of attribute information. The color attribute information of one or more of the points 126 may include a luminance value and two color difference values. The luminance value may represent the luminance (e.g., luma component, Y) of the point. The color difference values ​​may represent the blue and red components (e.g., chroma components, Cb and Cr) of the point, respectively, separate from its brightness. The other color attribute values ​​may be represented based on a different color scheme (e.g., RGB or monochrome color scheme).

[0018] The encoder 114 may encode the point cloud sequence 108 into a bitstream 110. To encode the point cloud sequence 108, the encoder 114 may use one or more lossless or lossy compression techniques to reduce redundant information in the point cloud sequence 108. To encode the point cloud sequence 108, the encoder 114 may use one or more prediction techniques to reduce redundant information in the point cloud sequence 108. Redundant information is information that can be predicted at the decoder 120 and thus may not need to be sent (e.g., transmitted) to the decoder 120 for accurate decoding of the point cloud sequence 108. For example, the Motion Picture Expert Group (MPEG) introduced the Geometry-Based Point Cloud Compression (G-PCC) standard (ISO / IEC Standard 23090-9: Geometry-Based Point Cloud Compression). G-PCC specifies encoded bitstream syntax and semantics for transmission and / or storage of compressed point cloud frames, as well as decoder operations for reconstructing the compressed point cloud frames from the bitstream. During the standardization of G-PCC, reference software (ISO / IEC Standard 23090-21: Reference Software for G-PCC) was developed to encode the geometric shape and attribute information of a point cloud frame. To encode the geometric shape information of a point cloud frame, the G-PCC reference software encoder may perform voxelization. The G-PCC reference software encoder may perform voxelization, for example, by quantifying the positions of points within a point cloud. Quantifying the positions of points within a point cloud may generate a grid in 3D space. The G-PCC reference software encoder may map points to the center coordinates of sub-grid volumes (e.g., voxels) within which their quantized positions reside. The G-PCC reference software encoder may perform geometric shape analysis using an occupancy tree to compress the geometric shape information. The G-PCC reference software encoder may entropy encode the results of the geometric shape analysis to further compress the geometric shape information.To encode the point cloud attribute information, the G-PCC reference software encoder may use transform tools such as a region-adaptive hierarchical transform (RAHT), a predictive transform, and / or a lifting transform. The lifting transform may be built on top of the predictive transform. The lifting transform may include an additional update / lifting step. The lifting transform and the predictive transform may also be referred to as a predictive / lifting transform or a "pred lift." The encoder 114 may operate in the same or similar manner as the encoder provided by the G-PCC reference software.

[0019] The output interface 116 may be configured to write and / or store the bitstream 110 on the transmission medium 104. The bitstream 110 may be sent (e.g., transmitted) to the destination device 106. Additionally or alternatively, the output interface 116 may be configured to send (e.g., transmit), upload, and / or stream the bitstream 110 to the destination device 106 via the transmission medium 104. The output interface 116 may include a wired and / or wireless transmitter configured to send (e.g., transmit), upload, and / or stream the bitstream 110 according to one or more proprietary and / or standardized communication protocols. The one or more proprietary and / or standardized communication protocols may include, for example, the Digital Video Broadcasting (DVB) standard, the Advanced Television Systems Committee (ATSC) standard, the Integrated Services Digital Broadcasting (ISDB) standard, the Data Over Cable Service Interface Specification (DOCSIS) standard, the 3rd Generation Partnership Project (3GPP®) standard, the Institute of Electrical and Electronics Engineers (IEEE) standard, the Internet Protocol (IP) standard, and the Wireless Application Protocol (WAP) standard.

[0020] Transmission medium 104 may include wireless, wired, and / or computer-readable media. For example, transmission medium 104 may include one or more wires, cables, air interfaces, optical disks, flash memory, and / or magnetic memory. Additionally or alternatively, transmission medium 104 may include one or more networks (e.g., the Internet) or file servers configured to store and / or transmit (e.g., transmit) encoded video data (e.g., bitstream 110).

[0021] The destination device 106 may include an input interface 118, a decoder 120, and a point cloud display 122. To decode the bitstream 110 into a point cloud sequence 108 for display or other consumption, the destination device 106 may be equipped with the input interface 118, the decoder 120, and the point cloud display 122. The input interface 118 may be configured to read the bitstream 110 stored on the transmission medium 104. The bitstream 110 may be stored on the transmission medium 104 by the source device 102. Additionally or alternatively, the input interface 118 may be configured to receive, download, and / or stream the bitstream 110 from the source device 102 over the transmission medium 104. The input interface 118 may comprise a wired and / or wireless receiver configured to receive, download, and / or stream the bitstream 110 according to one or more proprietary and / or standardized communication protocols. For example, the DVB (Digital Video Broadcasting) standard, the ATSC (Advanced Television Systems Committee) standard, the ISDB (Integrated Services Digital Broadcasting) standard, the DOCSIS (Data Over Cable Service Interface Specification) standard, the 3GPP (registered trademark) (3rd Generation Partnership Project) standard, the IEEE (Institute of Electrical and Electronics Engineers) standard, the IP (Internet Protocol) standard, and the WAP (Wireless Application Protocol) standard.

[0022] The decoder 120 may decode the point cloud sequence 108 from the encoded bitstream 110. The decoder 120 may operate in the same or similar manner as, for example, the decoder provided by the G-PCC reference software. The decoder 120 may decode a point cloud sequence that approximates the point cloud sequence 108. The decoder 120 may decode a point cloud sequence that approximates the point cloud sequence 108 due to, for example, lossy compression of the point cloud sequence 108 by the encoder 114 and / or errors introduced into the encoded bitstream 110 when transmission to, for example, the destination device 106 occurred.

[0023] The point cloud display 122 may display the point cloud sequence 108 to a user. The point cloud display 122 may include, for example, a cathode ray tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light emitting diode (LED) display, a 3D display, a holographic display, a head-mounted display, or any other display device suitable for displaying the point cloud sequence 108.

[0024] The point cloud encoding / decoding system 100 is presented by way of example and not limitation. In the embodiment of FIG. 1, the point cloud encoding / decoding system 100 may have other components and / or arrangements. The point cloud source 112 may be, for example, external to the source device 102. The point cloud display device 122 may be, for example, external to the destination device 106, or may be omitted entirely if the point cloud sequence is intended for consumption by a machine and / or a storage device. The source device 102 may further comprise, for example, a point cloud decoder. The destination device 104 may comprise, for example, a point cloud encoder. The source device 102 may further be configured to receive an encoded bitstream from the destination device 106. Receiving the encoded bitstream from the destination device 106 may support bidirectional point cloud transmission between the devices.

[0025] As described herein, the encoder may quantify the location of points within a point cloud according to a spatial precision, which may be the same or different in each dimension of the points. The quantization process may generate a grid in 3D space. The encoder may map any point that resides within each subgrid volume to a subgrid center coordinate called a voxel. A voxel may be considered a 3D extension of a pixel that corresponds to a 2D image grid coordinate.

[0026] The encoder may represent or encode the voxelized point cloud. The encoder may represent or encode the voxelized point cloud using, for example, an occupation tree. The encoder may divide an initial volume or cuboid containing the voxelized point cloud into sub-cuboids. The initial volume or cuboid may be referred to as a bounding box. The cuboid may be, for example, a cube. The encoder may recursively divide each sub-cuboid that contains at least one point of the point cloud. The encoder may not further divide a sub-cuboid that does not contain at least one point of the point cloud. A sub-cuboid that contains at least one point of the point cloud may be referred to as an occupied sub-cuboid. A sub-cuboid that does not contain at least one point of the point cloud may be referred to as an unoccupied sub-cuboid. The encoder may divide an occupied cuboid into, for example, two sub-cuboids (to form a binary tree), four sub-cuboids (to form a quadtree), or eight sub-cuboids (to form an octree). The encoder may divide the occupied cuboid to obtain sub-cuboids, which may have the same size and shape at a given depth level of the occupancy tree. The sub-cuboids may have the same size and shape at a given depth level of the occupancy tree, for example, if the encoder divides the occupied cuboid along a plane that passes through the center of the cuboid's edges.

[0027] The initial volume or cuboid containing the voxelized point cloud may correspond to the root node of the occupancy tree. Each occupied subcuboid split from the initial volume may correspond to a node (of the root node) at a second level of the occupancy tree. Each occupied subcuboid split from an occupied subcuboid at the second level may correspond to a node at a third level of the occupancy tree (the node off the occupied subcuboid at the second level from which it was split). The occupancy tree structure may continue to be formed in this manner for each recursive splitting iteration, for example, until some maximum depth level of the occupancy tree is reached or until each occupied subcuboid has a volume corresponding to one voxel.

[0028] Each non-leaf node of the occupancy tree may include or be associated with an occupancy word that represents the occupancy state of the cuboid corresponding to the node. A node of the occupancy tree corresponding to a cuboid divided into eight sub-cuboids may include or be associated with a one-byte occupancy word. Each bit of the one-byte occupancy word (called an occupancy bit) may represent or indicate the occupancy of a different one of the eight sub-cuboids. Each occupied sub-cuboid may be represented or indicated by a binary "1" in the one-byte occupancy word. Each unoccupied sub-cuboid may be represented or indicated by a binary "0" in the one-byte occupancy word. Occupied and unoccupied sub-cuboids may be represented or indicated by opposite one-bit binary values ​​in the one-byte occupancy word (e.g., a binary "0" representing or indicating an occupied sub-cuboid and a binary "1" representing or indicating an unoccupied sub-cuboid).

[0029] Each bit of the occupancy word may represent or indicate the occupancy of a different one of the eight sub-rectangles. Each bit of the occupancy word may represent or indicate the occupancy of a different one of the eight sub-rectangles, for example, according to the so-called Morton order. The least significant bit of the occupancy word may represent or indicate, for example, the occupancy of a first sub-rectangle of the eight sub-rectangles, for example, according to Morton order. The second least significant bit of the occupancy word may represent or indicate, for example, the occupancy of a second sub-rectangle of the eight sub-rectangles, for example, according to Morton order, etc.

[0030] 2 shows the Morton order of eight sub-rectangles (e.g., sub-rectangles 202-216) divided from a rectangle (e.g., rectangle 200). Sub-rectangles 202-216 are labeled based on their Morton order, with child node 202 coming first and child node 216 coming last. The Morton order for sub-rectangles 202-216 is a local lexicographic order in xyz.

[0031] The voxelized point cloud geometry can be represented by and determined from the initial volumes and occupancy words of the nodes in the occupancy tree. The encoder may send (e.g., transmit) the initial volumes and occupancy words of the nodes in the occupancy tree in a bitstream to a decoder to reconstruct the point cloud. The encoder may entropy encode the occupancy words. For example, the encoder may entropy encode the occupancy words before sending the initial volumes and occupancy words of the nodes in the occupancy tree. The encoder may encode occupancy bits of the occupancy word of the node corresponding to the cuboid. For example, the encoder may encode occupancy bits of the occupancy word of the node corresponding to the cuboid based on one or more occupancy bits of the occupancy words of other nodes corresponding to cuboids that are adjacent to or spatially close to the cuboid of the occupancy bit being encoded.

[0032] The encoder and / or decoder may encode the occupied bits of consecutive occupied words in scan order. Scan order may also be referred to as scanning order. The encoder and / or decoder may scan the occupation tree in breadth-first order. All occupied words of nodes at a given depth (e.g., level) in the occupation tree may be scanned. All occupied words of nodes at a given depth (e.g., level) in the occupation tree may be scanned before scanning the occupied words of nodes at the next depth (e.g., level). Within a given depth, the encoder and / or decoder may scan the occupied words of the nodes in Morton order. Within a given node, the encoder and / or decoder may also scan the occupied bits of the occupied words of the node in Morton order.

[0033] 3 illustrates an example of a scan order (e.g., breadth-first order as described herein) of an occupancy tree (e.g., occupancy tree 300). FIG. 3 illustrates a scan order of the first three exemplary levels of occupancy tree 300. In FIG. 3, a cuboid 302 corresponding to the root node of occupancy tree 300 may be divided into eight sub-cuboids. Two sub-cuboids 304 and 306 of the eight sub-cuboids may be occupied. The other six sub-cuboids of the eight sub-cuboids may be unoccupied. According to Morton order, the first 8-bit occupancy word occW 1,1 is constructed to represent the occupied word of the root node. The first 8-bit occupied word occW 1,1 The least significant occupancy bit of the first 8-bit occupancy word occW represents or indicates the occupancy of the first of the eight sub-cuboids in Morton order. 1,1 The second least significant occupied bit of represents or indicates the occupation of the second of the eight sub-cuboids in Morton order, etc.

[0034] Each of the two occupied sub-cuboids 304 and 306 corresponds to a node from the root node of the second-level occupancy tree 300. Each of the two occupied sub-cuboids 304 and 306 is further divided into eight sub-cuboids. Of the eight sub-cuboids divided from sub-cuboid 304, one of sub-cuboid 308 may be occupied. The other seven of the eight sub-cuboids divided from sub-cuboid 304 may be unoccupied. Of the eight sub-cuboids divided from sub-cuboid 306, three of sub-cuboids 310, 312, and 314 may be occupied. The other five of the eight sub-cuboids divided from sub-cuboid 306 may be unoccupied. Two second 8-bit occupancy words occW 2,1 and occW 2,2 are constructed in this order to represent the occupancy words of the nodes corresponding to sub-cuboid 304 and sub-cuboid 306, respectively.

[0035] Each of the four occupied sub-cuboids 308, 310, 312, and 314 corresponds to a node in the third level occupancy tree 300. Each of the four occupied sub-cuboids 308, 310, 312, and 314 is further divided into a total of eight sub-cuboids or 32 sub-cuboids. Four third level 8-bit occupied words occW 3,1 , occW 3,2 , occW 3,3 , and occW 3,4 are constructed in this order to represent the occupancy words of the nodes corresponding to sub-cuboid 308, the occupancy words of the nodes corresponding to sub-cuboid 310, the occupancy words of the nodes corresponding to sub-cuboid 312, and the occupancy words of the nodes corresponding to sub-cuboid 314, respectively.

[0036] The occupied words of the occupancy tree 300 are sorted, for example, according to a scan order (e.g., breadth-first order) described herein, into seven occupied words occW 1,1 ~occW 3,4The occupied words of the current child node may be entropy encoded (e.g., entropy encoded by an encoder and entropy decoded by a decoder) as a succession of the following: As a result of the breadth-first scan order, the occupied words of all nodes having the same depth (e.g., level) as the current parent node may already be entropy encoded, for example, if the occupied words of the current child node belonging to the current parent node have been entropy encoded. The occupied words of all nodes having the same depth (e.g., level) as the current child node and having a lower Morton order than the current child node may also already be entropy encoded, for example, if the occupied word of the current child node has been entropy encoded. A portion of the already encoded occupied words can be used to entropy encode the occupied word of the current child node. The already encoded occupied words of adjacent parent and / or child nodes may be used, for example, to entropy encode the occupied word of the current child node. The occupied bits of occupied words having a lower Morton order than a particular occupied bit of the occupied word of the current child node may also already be entropy encoded. An occupied bit of an occupied word having a lower Morton order than a particular occupied bit may be used to encode the occupied bit of an occupied word of a current child node, for example, when the particular occupied bit is being encoded.

[0037] 4 illustrates exemplary neighborhoods of cuboids for entropy coding the occupancy of a child cuboid. Neighbors of cuboids with already-coded occupancy bits can be used to entropy code the occupancy bits of a current child cuboid 400. Neighbors of cuboids with already-coded occupancy bits can be determined. Neighbors of cuboids with already-coded occupancy bits can be determined, for example, based on a scan order of an occupancy tree representing the geometry of the cuboids of FIG. 4 described herein. For a current child cuboid, the cuboid neighbors may include one or more of: a cuboid close to the current child cuboid, a cuboid sharing a vertex with the current child cuboid, a cuboid sharing an edge with the current child cuboid, a cuboid sharing a face with the current child cuboid, a parent cuboid close to the current child cuboid, a parent cuboid sharing a vertex with the current child cuboid, a parent cuboid sharing an edge with the current child cuboid, a parent cuboid sharing a face with the current child cuboid, a parent cuboid close to the current parent cuboid, a parent cuboid sharing a vertex with the current parent cuboid, a parent cuboid sharing an edge with the current parent cuboid, a parent cuboid sharing a face with the current parent cuboid, etc. As shown in FIG. 4 , a current child cuboid 400 may belong to a current parent cuboid 402. According to the scan order of the occupancy words and occupancy bits of the nodes of the occupancy tree, the occupancy bits of the four child cuboids 404, 406, 408, and 410 belonging to the same current parent cuboid 402 have already been coded. The occupancy bits of the child cuboid 412 of the preceding parent cuboid have already been coded. The occupancy bits of the parent cuboid 414, whose child occupancy bits have not yet been coded, have already been coded. Therefore, the occupancy bits of the current child cuboid 400 can be coded using the occupancy bits already coded of the cuboids 404, 406, 408, 410, 412, and 414.

[0038] The number (e.g., quantity) of possible occupancy configurations (e.g., sets of one or more occupancy words and / or occupancy bits) for the neighbors of the current child cuboid is 2 Nwhere N is the number (e.g., quantity) of cuboids in the neighborhood of the current child cuboid that have already coded occupancy bits. The neighborhood of the current child cuboid may include tens of cuboids. The neighborhood of the current child cuboid may include 26 neighboring parent cuboids that share faces, edges, and / or vertices with the parent cuboid of the current child cuboid, and also several neighboring child cuboids that share faces, edges, and / or vertices with the current child cuboid. The occupancy configuration of the neighborhood of the current child cuboid may be limited to a subset of neighboring cuboids or may have billions of possible occupancy configurations, making its direct use impractical. The encoder and / or decoder may use the occupancy configurations for the neighborhood of the current child cuboid to select a context (e.g., a probability model) from a set of contexts of a binary entropy coder (e.g., a binary arithmetic coder) that codes the occupancy bits of the current child cuboid. Context-based binary entropy coding may be similar to the context-adaptive binary arithmetic coder (CABAC) used in MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC)).

[0039] The encoder and / or decoder may use several methods to reduce the adjacent occupancy configuration of the current child cuboid to be encoded to a practical number (e.g., quantity) of reduced occupancy configurations. 6 That is, 64 occupancy configurations can be reduced to 9 occupancy configurations. Occupancy configurations can be reduced by using geometric invariants. The occupancy score of the current child cuboid is 2x the occupancy score of the 26 neighboring parent cuboids. 26 A score can be obtained from the occupancy configurations. The score can be further reduced to a ternary occupancy prediction (e.g., "predicted-occupied," "uncertain," or "predicted-unoccupied") by using a score threshold. The number (e.g., quantity) of nearby occupied child cuboids and the number (e.g., quantity) of nearby unoccupied child cuboids can be used instead of the individual occupancies of these child cuboids.

[0040] The encoder and / or decoder may reduce the number (e.g., quantity) of possible occupancy configurations for the neighbors of the current child cuboid to a more manageable number (e.g., several thousand). It has been observed that instead of directly associating the reduction in the number (e.g., quantity) of contexts (e.g., probability models) with the reduction in occupancy configurations, another mechanism, namely, OBUF (Optimal Binary Coders with Update on the Fly), may be used. The encoder and / or decoder may implement OBUF to limit the number (e.g., quantity) of contexts to a lower number (e.g., 32 contexts).

[0041] The OBUF may use a limited number (e.g., 32) of contexts (e.g., probability models). The number (e.g., quantity) of contexts in the OBUF may be a fixed number (e.g., a fixed quantity). The contexts used by the OBUF may be ordered and referenced by a context index (e.g., a context index ranging from 0 to 31), and a "1" may be encoded from the lowest to the highest virtual probability. A context index lookup table (LUT) may be initialized at the beginning of the point cloud encoding process. The LUT may initially point to a context having a median virtual probability for encoding a "1" for all inputs. The LUT may initially point to a context having a median virtual probability for encoding a "1" among the limited number (e.g., quantity) of contexts for all inputs. This LUT may take as input the occupancy configuration of neighbors of the current child cuboid and output a context index associated with the occupancy configuration. The LUT may have the same number of entries as the reduced occupancy configuration (e.g., approximately several thousand entries). Encoding the occupancy bit of the current child cuboid may include determining a reduced occupancy configuration of the current child node, obtaining a context index by using the reduced occupancy configuration as an entry into a LUT, encoding the occupancy bit of the current child cuboid by using the context pointed to (e.g., indicated) by the context index, and updating the LUT entry corresponding to the reduced occupancy configuration. The LUT entry may be updated, for example, based on the value of the encoded occupancy bit of the current child cuboid. For encoding a binary "0" (e.g., indicating that the current child cuboid is unoccupied), the LUT entry may be decreased to a lower context index value (e.g., associated with a lower virtual probability). For encoding a binary "1" (e.g., indicating that the current child cuboid is occupied), the LUT entry may be increased to a higher context index value (e.g., associated with a higher virtual probability).The context index update process may be based on a theoretical model of optimal distribution for hypothetical probabilities associated with a limited number (e.g., quantity) of contexts. The hypothetical probabilities may be fixed by the model. The hypothetical probabilities may differ from the internal probabilities of the contexts that evolve during the encoding of bits of data. The evolution of the internal contexts may follow a well-known process similar to that of CABAC.

[0042] The encoder and / or decoder may implement a "dynamic OBUF" scheme. The "dynamic OBUF" scheme can handle a much larger number (e.g., quantity) of occupancy configurations for the neighbors of the current child cuboid than a typical OBUF. Using a larger number (e.g., quantity) of occupancy configurations for the neighbors of the current child cuboid may result in improved compression performance. Using a larger number (e.g., quantity) of occupancy configurations for the neighbors of the current child cuboid may also keep the complexity within a reasonable range. By using an occupancy tree compressed by OBUF, the encoder and / or decoder may achieve lossless compression performance as good as 1 bit per point (bpp) for encoding dense point cloud geometry. The encoder and / or decoder may implement a dynamic OBUF to further reduce the bitrate, potentially by more than 25%, down to 0.7 bpp.

[0043] The OBUF cannot take as input a wide variety of reduced occupancy configurations for the neighbors of the current child cuboid. This can potentially lead to a loss of useful correlations. The OBUF can increase the size of the context index LUT to process a larger variety of occupancy configurations as input than for the neighbors of the current child cuboid. Doing so can dilute statistics and worsen compression performance. For example, if the LUT has millions of entries and the point cloud has hundreds of thousands of points, most entries will never be visited (e.g., looked up, accessed, etc.). In some instances, many entries may be visited only a few times, and their associated context indexes may not be updated enough times to reflect any meaningful correlation between occupancy configuration values ​​and the occupancy probability of the current child cuboid. A dynamic OBUF can be implemented to mitigate the dilution of statistics due to an increase in the number (e.g., quantity) of occupancy configurations for the neighbors of the current child cuboid. This mitigation is achieved by “dynamic shrinking” of occupancy configurations in the dynamic OBUF.

[0044] The dynamic OBUF may add an additional step to reduce the occupancy configuration of the neighbors of the current child cuboid. The dynamic OBUF may add an additional step to shrink the occupancy configuration of the neighbors of the current child cuboid, for example, before using the LUT of the context index. This step may be called dynamic shrinking because it evolves based on the progress of the point cloud encoding, or more precisely, based on the occupancy configurations already visited (e.g., looked up in the LUT).

[0045] As described herein, many possible occupancy configurations for the neighbors of the current child cuboid are potentially involved, but only a subset may be visited, for example, when encoding a point cloud. This subset of visited occupancy configurations may characterize the type of point cloud. For example, most of the visited occupancy configurations may represent occupied neighboring cuboids of the current child cuboid, for example, when an AR or VR dense point cloud is encoded. On the other hand, most of the visited occupancy configurations may represent only a few occupied neighboring cuboids of the current child cuboid, for example, when a sparse point cloud acquired by a sensor is encoded. The role of dynamic shrinking may be, for example, to obtain a more accurate correlation based on the most frequently visited occupancy configurations by refraining from (e.g., actively reducing) other occupancy configurations that are much less frequently visited. Dynamic shrinking may be updated on the fly. Dynamic shrinking may be updated on the fly. Dynamic shrinking may be updated on the fly, for example, after each visit of an occupancy configuration (e.g., lookup in a LUT). A visit to a proprietary configuration (eg, a lookup of a LUT) may occur, for example, when encoding of proprietary data occurs.

[0046] 5 shows an example of a dynamic shrinking function (DR) that can be used in a dynamic OBUF. The dynamic shrinking function (DR) is a function of bit β j can be obtained by masking β=β1...β K It consists of K bits. The size of the mask can be reduced, for example, if the occupancy configuration is visited (e.g., looked up in a LUT) a certain number of times (e.g., a certain number). The initial dynamic reduction function DR 0 is a constant function DR for all occupancy configurations β 0 All bits may be masked for all occupancy configurations so that (β)=0. The dynamic reduction function is the function DR n Updated function from DR n+1 The dynamic shrinking function may evolve into, for example, a function DR n Updated function from DR n+1 The function can evolve to β'=DR n (β)=β1...β kn(β) where k n (β) 510 is the number of unmasked bits (e.g., quantity). 0 The initialization of k(β) may correspond to k0(β)=0, and the natural evolution of the shrinkage function for finer statistics is the increase in the number of unmasked bits (e.g., quantity) k n (β)≦k n+1 (β). The dynamic shrinkage function is k for all occupancy configurations β. n can be completely determined by the value of

[0047] A visit to an occupied configuration (e.g., an instance of a lookup in the LUT) is performed for all dynamically reduced occupied configurations β'=DR n (β) can be tracked by a variable NV(β'). The corresponding number (e.g., quantity) of visits NV(β V ') can be increased by 1. The corresponding number (e.g., quantity) of visits NV(β V ') is, for example, the occupied configuration β V After each instance of encoding the occupancy bits based on , the number (e.g., quantity) of visits NV(β V ') is the threshold th V If it is greater than NV(β V ')>th V Next, the number of unmasked bits (e.g., quantity) k n (β) is β V’ This means that the dynamically reduced occupancy configuration β V ' into a dynamically reduced new two-occupancy configuration β defined by the following equation: 0 ' and β 1 ' corresponds to replacing it with '.

number

[0048] In other words, the number (e.g., amount) of unmasked bits is n (β)=β V’ For all occupancy configurations β, k n+1 (β)=k n (β)+1, which is an increase of 1. The number of visits (e.g., quantity) of the new dynamically reduced two-occupancy configuration may be initialized to zero. NV(β 0 ')=NV(β 1 ')=0 (I)

[0049] At the beginning of encoding, an initial dynamic reduction function DR 0 The initial number (e.g., quantity) of visits may be set as follows: NV(DR 0 (β))=NV(0)=0 The evolution of NVs in dynamically reduced occupancy configurations can be fully defined.

[0050] The corresponding LUT entry LUT[β V '] is β V Two new entries LUT[β 0 '] and LUT[β 1 ']. The corresponding LUT entry LUT[β V '] is, for example, the dynamically reduced occupancy configuration β V ' is dynamically reduced to a new two-occupancy configuration β 0 ' and β 1 ', then β V Two new entries LUT[β 0 '] and LUT[β 1 '], LUT[β 0 ']=LUT[β 1 ']=LUT[β V '] (II) They are then evolved separately. The evolution of the coder index LUT on the dynamically reduced occupancy configuration can be fully defined.

[0051] Reduction Function DR n is the occupancy configuration β'=DR in which the leaf node 530 is reduced. n (β) is a set of growing binary trees T n 520. The initial tree can be modeled by 0=DR 0 There can be a single root node associated with (β). 0 ' and β 1 ' by β V The replacement of the dynamically reduced β V ' from the leaf node associated with the tree T n This corresponds to growing β 0 ' and β 1 ' by β V The replacement of the dynamically reduced β 0 ' and β 1 ' by attaching two new nodes related to β V ' from the leaf node associated with tree T n It corresponds to growing a tree. n+1 can be obtained by this growth. The number of visits (e.g., quantity), NV, and LUT of context indexes are defined on the leaf nodes and can evolve with the growth of the tree through equations (I) and (II).

[0052] A practical implementation of the dynamic OBUF is to use the arrays NV[β'] and LUT[β'] of context indices and the tree T n 520. An alternative to storing the tree is to store an array k of the number of unmasked bits (e.g., quantity). n [β] 510 can be stored.

[0053] A limitation for implementing a dynamic OBUF may be its memory footprint. In some instances, millions of occupied configurations may be processed in practice, resulting in approximately 20 bits βi that configure the input configuration β to the reduction function DR. Each bit β imay correspond to the occupancy of the neighboring cubes of the current child cube, or to the set of neighboring cubes of the current child cube.

[0054] The higher (e.g., more significant) bit β i (e.g., β0, β1, etc.) may be the first unmasked bit. i (e.g., β0, β1, etc.) may be, for example, the first bit that is not masked during the evolution of the dynamic shrinkage function DR. i The order of neighbor-based information placed in the β i , for example, from higher weight to lower weight. The priority may be, from most important to least important, the occupancy of a set of close adjacent child cuboids, then the occupancy of close adjacent child cuboids, then the occupancy of close adjacent parent cuboids, then the occupancy of non-close adjacent child nodes, and finally the occupancy of non-close adjacent parent nodes. Neighboring nodes that share a face with the current child node may also have a higher priority than neighboring nodes that share an edge (but not a face) with the current child node. Neighboring nodes that share an edge with the current child node may have a higher priority than neighboring nodes that share only vertices with the current child node.

[0055] FIG. 6 illustrates an exemplary method for encoding the occupancy of a cuboid using a dynamic OBUF. More specifically, FIG. 6 illustrates a flowchart of an exemplary method for encoding the occupancy (e.g., indicated by a single bit) of a current child cuboid using a dynamic OBUF. More specifically, FIG. 6 illustrates a flowchart of exemplary method steps for encoding the occupancy of a current child cuboid using a dynamic OBUF. The exemplary method, or one or more operations of the method, may be performed by one or more computing devices or entities. For example, all or part of the flowchart may be implemented by a coder (e.g., encoder 114 of FIG. 1 and / or decoder 120 of FIG. 1), exemplary computer system 1700 of FIG. 17, and / or exemplary computing device 1830 of FIG. 18.

[0056] In step 602, the encoder and / or decoder may determine the occupancy configuration β of the current child cuboid. For example, the encoder and / or decoder may determine the occupancy configuration β of the current child cuboid based on the occupancy bits of already-encoded cuboids adjacent to the current child cuboid. In step 604, the encoder and / or decoder may determine the occupancy configuration β of the current child cuboid based on the occupancy bits of already-encoded cuboids adjacent to the current child cuboid. n (For example, β' = DR n (β)) may be used to dynamically reduce the occupancy configuration β to a reduced occupancy configuration. At step 606, the encoder and / or decoder may look up a context index LUT[β'] in the LUT of the dynamic OBUF. At step 608, the encoder and / or decoder may select a context (e.g., a probability model) pointed to by the context index. At step 610, the encoder and / or decoder may entropy code (e.g., arithmetic code) the occupancy bits of the current child cuboid based on the context.

[0057] Although not shown in FIG. 6, the encoder and / or decoder may use a reduction function DR based on the occupied bits of the current child cuboid. n DR n+1and / or update the context index LUT[β']. The method of Figure 6 may be repeated for additional or all child cuboids of a parent cuboid corresponding to a node in the occupancy tree in a scan order, such as the scan order described herein with respect to Figure 3.

[0058] Occupancy trees are lossless compression techniques. Occupancy trees can be adapted to provide lossy compression, for example, by modifying the point cloud on the encoder side (e.g., downsampling, removing points, moving points, etc.), but lossy compression may have weak compression performance. This can be a useful lossless compression technique for dense point clouds.

[0059] An approach to lossy compression for point cloud geometry may be to set the maximum depth of the occupancy tree so that it does not reach a minimum volume size of one voxel. Instead, the maximum depth of the occupancy tree may be set to stop at a larger volume size (e.g., an N×N×N cuboid, where N>1). The geometry of the points belonging to each occupied leaf node associated with the larger volume may then be modeled. This approach may be particularly suitable for dense, smooth point clouds that can be locally modeled by a smooth function, e.g., a plane or a polynomial. The encoding cost may be the cost of the occupancy tree plus the cost of a local model of each occupied leaf node.

[0060] A scheme for modeling the geometry of points belonging to each occupied leaf node associated with a volume size larger than one voxel may use a set of triangles as a local model. The scheme may be called a "TriSoup" scheme. TriSoup is an abbreviation for "triangle soup" because the connections between triangles may not be part of the model. An occupied leaf node in the occupation tree corresponding to a cuboid with a volume greater than one voxel may be referred to as a TriSoup node. An edge belonging to at least one cuboid corresponding to a TriSoup node may be referred to as a TriSoup edge. A TriSoup node stores an existence flag (s) for each TriSoup edge of its corresponding occupied cuboid. k ) TriSoup edge existence flag (s k ) is a TriSoup vertex (V k ) on a TriSoup edge. At most one TriSoup vertex (V k ) can exist on the TriSoup edge. Each vertex (V k ), the TriSoup node corresponding to the occupied cuboid is the vertex along the TriSoup edge (V k ) position (p k ) may further include.

[0061] In addition to the occupancy word of the occupancy tree, the encoder may entropy encode the TriSoup vertex presence flag and the position of each TriSoup edge belonging to a TriSoup node in the occupancy tree. The decoder may similarly entropy decode the occupancy word of the occupancy tree, as well as the TriSoup vertex presence flag and the position of each TriSoup edge belonging to a TriSoup node in the occupancy tree.

[0062] FIG. 7 shows an example of an occupation cuboid (e.g., occupation cuboid 700) corresponding to a TriSoup node in an occupation tree. Occupation cuboid 700 may be of size N×N×N, where N>1. Occupation cuboid 700 may include TriSoup edges 710-721. The TriSoup node corresponding to occupation cuboid 700 stores an existence flag (s k ) The existence flag for TriSoup edge 714 may indicate that TriSoup vertex V1 is on TriSoup edge 714. The existence flag for TriSoup edge 715 may indicate that TriSoup vertex V2 is on TriSoup edge 715. The existence flag for TriSoup edge 716 may indicate that TriSoup vertex V3 is on TriSoup edge 716. The existence flag for TriSoup edge 717 may indicate that TriSoup vertex V4 is on TriSoup edge 717. The existence flags for the remaining TriSoup edges may each indicate that the TriSoup vertex is not on the corresponding TriSoup edge. The TriSoup node corresponding to occupied cuboid 700 may further include the location of each TriSoup vertex that is along one of its TriSoup edges 710-721. More specifically, the TriSoup node corresponding to occupied cuboid 700 may further include a position p1 of TriSoup vertex V1, a position p2 of TriSoup vertex V2, a position p3 of TriSoup vertex V3, and a position p4 of TriSoup vertex V4.

[0063] Figure 8A shows an example of a cuboid corresponding to a TriSoup node. The cuboid 800 is a cuboid that is connected to the TriSoup vertex V k Within the cuboid 800, the TriSoup triangles may correspond to TriSoup nodes with a quantity (e.g., number) K of TriSoup vertices V k A TriSoup triangle can be constructed from, for example, TriSoup vertices V if there are at least three (K≧3) TriSoup vertices on the TriSoup edges of the rectangular solid 800. kIn the example of FIG. 8A, there are four TriSoup vertices, and a TriSoup triangle is constructed. The TriSoup triangle can be constructed around a centroid vertex C. The centroid vertex C is located at the center of the TriSoup vertex V. k The main direction can be determined and the vertex V k can be ordered by rotating around this direction, and the following K TriSoup triangles can be constructed: V1V2C, V2V3C, ..., V K V1C. The principal direction may be chosen from among three directions each parallel to an axis in 3D space, for example, to increase or maximize the 2D surface of the triangle when the triangle is projected along the principal direction. The principal direction may be somewhat perpendicular to the local surface defined by the points of the point cloud belonging to the TriSoup node.

[0064] FIG. 8B shows an example fine-tuning for a TriSoup model. The TriSoup model can be fine-tuned by encoding the centroid residual value. The centroid residual value C res can be coded into the bitstream. res For example, use C+C instead of C as the pivot vertex of the triangle. res C+C as the pivot vertices of the triangle. res By using res may be closer to the points in the point cloud than the centroid C, lowering the reconstruction error and thereby res The lower distortion can be achieved at the cost of a small increase in the bit rate required to encode the image.

[0065] FIG. 9 shows an example of voxelization. Voxelization may refer to the reconstruction of a decoded point cloud from a set of TriSoup triangles. Voxelization may be performed by ray tracing for each triangle individually. Voxelization may be performed by ray tracing for each triangle individually, for example, before removing overlapping points between voxelized triangles. As shown in FIG. 9, a ray 900 may be shot parallel to one of three axes in 3D space. The ray 900 may be projected along integer coordinate P start The intersection point P of the ray 900 with the TriSoup triangle 901 belonging to the rectangular parallelepiped 902 corresponding to the TriSoup node int (if any) may be rounded to obtain the decoded point. This intersection point P int can be found using, for example, the Moller-Trumbore algorithm.

[0066] Existence flag (s k ) and existence flag (s k ) indicates the existence of a vertex, the current TriSoup edge position (p k ) can be entropy coded. k ) and position (p k ) may be individually or collectively referred to as vertex information. k ), and existence flag (s k ) indicates the existence of a vertex, the current TriSoup edge position (p k ) can be entropy coded, for example, based on the already coded existence flags and the positions of the TriSoup edges adjacent to the current TriSoup edge. k ) and existence flag (s k ) indicates the existence of a vertex, the current TriSoup edge position (p k ) may additionally or alternatively be entropy coded. The existence flag of the current TriSoup edge (s k ) and position (p k) can additionally or alternatively be entropy coded, for example, based on the occupancy of the cuboids adjacent to the current TriSoup edge. Similar to the entropy coding of the occupancy bits of the occupancy tree, the neighbors of the current TriSoup edge (neighborhood configuration β TS (also called) configuration β TS Obtain the reduced configuration β TS '=DR n (β TS ) can be dynamically reduced to the current TriSoup edge neighbor configuration β TS and obtain the reduced configuration β by using, for example, TriSoup's dynamic OBUF scheme. TS '=DR n (β TS ) can be dynamically reduced to the context index LUT[β TS '] may be obtained from the OBUF LUT. At least a portion of the vertex information of the current TriSoup edge may be entropy coded using the context (e.g., a probability model) pointed to by the context index.

[0067] The TriSoup vertex position (p k ) (if present) can be binarized. TriSoup vertex positions (p k ) (if present) may be binarized to entropy encode at least a portion of the vertex information of the current TriSoup edge, for example, using a binary entropy coder. b The number (e.g., quantity) of TriSoup vertices along a TriSoup edge of length N (p k ) can be set to quantify the length of a TriSoup edge. Nb The quantization interval can be divided evenly. By doing so, the TriSoup vertex position (p k ) can be individually encoded by a dynamic OBUF scheme. b Bit(p k j 、 j=1,...,N b ), and existence flags (s k ) can be represented by the bits corresponding to the neighbor configuration βTS , OBUF reduction function DR n , and therefore the context index depends on the nature of the coded bits (e.g., presence flag (s k ), the highest bit (p k 1 ), the second highest bit (p k 2 There are several possible dynamic OBUF schemes, each of which depends on a specific bit of information in the vertex information (e.g., existence flag (s k ) or position bit (p k j)) is exclusive to this product.

[0068] 10A and 10B show rectangular parallelepipeds (e.g., cubes 1000-1003 in FIG. 10A, cubes 1010-1013, and cubes 1020-1023 in FIG. 10B) having volumes that intersect with the current TriSoup edge (e.g., current TriSoup E) being entropy coded. Current TriSoup edge E is an edge of rectangular parallelepipeds 1000-1003. The start point of current TriSoup edge E intersects with rectangular parallelepipeds 1010-1013. The end point of current TriSoup edge E intersects with rectangular parallelepipeds 1020-1023. The neighbor configuration β of current TriSoup edge E is calculated using the occupied bits of one or more of the 12 rectangular parallelepipeds 1000-1003, 1010-1013, and 1020-1023. TS can be determined.

[0069] TriSoup edges may be oriented from start to end according to the orientation of one of the three axes in 3D space to which they are parallel. The overall ordering of TriSoup edges may be defined as a lexicographical order over the set (e.g., start, end). Vertex information associated with TriSoup edges may be encoded according to the TriSoup edge order. The causal neighbors of the current TriSoup edge may be obtained from already encoded TriSoup edges that are adjacent to the current TriSoup edge.

[0070] 11A, 11B, and 11C show TriSoup edges (E' and E'') that may be used to entropy encode the current edge E. In some instances, up to five TriSoup edges (E' and E'') may be used to entropy encode the current TriSoup edge E. The five TriSoup edges are: - an edge E' that is parallel to the current TriSoup edge E and has an end point equal to the start point of the current TriSoup edge E; - four edges E'' that are perpendicular to the current TriSoup edge E and have a start point or end point equal to the start point of the current TriSoup edge E. Depending on the direction of the current TriSoup edge E, any two (in FIG. 11C, direction z), three (in FIG. 11B, direction y), or four (in FIG. 11A, direction x) of the four perpendicular TriSoup edges may already be encoded, and their vertex information may be used to create the neighbor configuration β for the current TriSoup edge E. TS A TriSoup edge E' may already be encoded for each direction of the current TriSoup edge E, and its vertex information may be used to construct an adjacency configuration β of the current TriSoup edge E, independent of its direction. TS can be constructed.

[0071] As described herein, the neighbor configuration β of the current TriSoup edge E TS can be obtained from one or more occupied bits of the cuboid and from the vertex information of adjacent already encoded TriSoup edges. TS can be obtained from one or more of the 12 occupied bits of the 12 cuboids shown in Figures 10A and 10B, and from the vertex information of up to five adjacent already-encoded TriSoup edges (E' and E'') shown in Figures 11A, 11B, and 11C.

[0072] Performance can be improved by using inter-frame prediction, for example, in video compression. The bit rate required to compress inter-frames can typically be one to two orders of magnitude lower than the bit rate within a frame without inter-frame prediction, by definition. Point cloud data may behave differently because 3D geometry is coded, unlike video coding, where typically only attributes (e.g., color) are coded after projecting the 3D geometry onto a 2D plane (e.g., a camera sensor). Even if the 2D projected attributes are expected to have higher temporal correlation than the underlying 3D geometry, it can be expected that inter-frame prediction between 3D point clouds can provide improved compression capabilities over intra-frame prediction within point clouds alone. Octrees can benefit from inter-frame prediction and geometry compression gains.

[0073] FIG. 12 shows an exemplary encoding method. One or more steps of FIG. 12 may be performed by an encoder and / or decoder (e.g., the encoder 114 and / or decoder 120 of FIG. 1), the exemplary computer system 1700 of FIG. 17, and / or the exemplary computing device 1830 of FIG. 18. The general framework of inter-frame prediction of 3D point clouds may be similar to one of the video compression encoding (e.g., encoding) processes described herein with respect to FIG. 12. A current frame 1200 (e.g., an image or a point cloud) may be encoded based on an already-encoded reference frame 1210 (e.g., an image or a point cloud). A motion search 1220 may be performed, for example, from the already-encoded reference frame 1210 toward the current frame 1200 to obtain a motion vector 1221. The motion vector 1221 may represent the flow of motion between the already-encoded reference frame 1210 and the current frame 1200.

[0074] A motion vector may be a two-component (e.g., 2D) vector that may represent, for example, movement from a reference block of pixels to a current block of pixels in at least some video compression. A motion vector may be a three-component (e.g., 3D) vector that may represent, for example, movement from a reference set of 3D points to a current set of 3D points in at least some point cloud compression. The motion vector 1221 may be entropy coded into the bitstream 1250 (as shown in FIG. 12 , step 1225). The reference frame 1210 may be motion compensated (as shown in FIG. 12 , step 1230), for example, to obtain a motion-compensated frame 1231. Motion compensation may involve, for example, moving pixels of the reference image (respectively, a point cloud) according to a 2D motion vector and / or moving points of the reference point cloud according to a 3D motion vector. The obtained motion-compensated frame 1231 may be “closer” to the current frame 1200 than the reference frame 1210. The obtained motion-compensated frame 1231 may be closer to the current frame 1200 than the reference frame 1210, for example, in that the color difference (and / or point distance) between the motion-compensated frame 1231 and the current frame 1200 may be smaller than the color difference between the reference frame 1210 and the current frame 1200. The obtained motion-compensated frame 1231 may be closer to the current frame 1200 than the reference frame 1210, for example, the color difference and / or point distance between the motion-compensated frame 1231 and the current frame 1200 may be smaller, on average, than the color difference and / or point distance between the reference frame 1210 and the current frame 1200. In step 1240, inter-frame prediction may be performed, for example, to obtain an inter-residual 1241. The inter-residual 1241 may be entropy coded into a bitstream 1250 (at step 1245, as shown in FIG. 12 ). The residual 1241 may contain more compressible information than the current frame 1200 or the current frame that has undergone an intra-prediction process.Entropy encoding 1245 may be more efficient in obtaining a bitstream 1250 that is smaller in size compared to the bitstream obtained by encoding the current frame 1200, which may not have benefited from inter-frame prediction.

[0075] The inter-residual (e.g., inter-residual 1241) may be constructed, for example, in video coding, as a pixel-by-pixel color difference between a current block of pixels belonging to a current frame (e.g., image) and a co-located compensated block of pixels belonging to a motion-compensated frame (e.g., image). The inter-residual (e.g., inter-residual 1241) may be an array of color differences, which may have a small size and therefore can be efficiently compressed.

[0076] There may be no concept such as the difference between two sets of points. The concept of residual difference may not be simply generalized to point clouds, for example, in point cloud compression. For prediction of octrees that may represent point clouds, the concept of residual difference may be replaced by conditional entropy coding, and the conditional information for performing the conditional entropy coding may be constructed, for example, based on a motion-compensated point cloud. This approach may be extended to the framework of Optimal Binary Coders with Update on the Fly (OBUF).

[0077] As described herein, the current occupancy bits of the octree may be coded by an entropy coder. The entropy coder may be selected by the output of a dynamic OBUF lookup table (LUT) of coder indexes, which may use the neighbor configuration β as input. The neighbor configuration β may be constructed, for example, based on already coded occupancy bits associated with volumes adjacent to the current volume. The current volume may be associated with a current node, whose occupancy may be signaled by a current occupancy bit. The construction of the neighbor configuration β may be extended, for example, by using inter-frame information. An inter-predictor occupancy bit may be defined relative to the current occupancy bit as a bit representing the presence of at least one point of the motion-compensated point cloud within the current volume. Because the current compensated point cloud and the motion-compensated point cloud may be close to each other, a strong correlation between the current occupancy bits and the inter-predictor occupancy bits may exist, for example, when motion compensation is efficient. Using the inter-predictor occupancy bits as bits of the neighbor configuration β may result in improved compression performance of the octree (e.g., by dividing the size of the octree bitstream by a factor of 2).

[0078] The inter-octree motion field may include a 3D motion vector associated with a 3D prediction unit. The 3D prediction unit may have a volume embedded in a volume (e.g., a cube) associated with a node of the octree. Motion compensation may be performed volume-by-volume based on the 3D motion vector, for example, to obtain a motion-compensated point cloud in one or more current volumes. An inter-predictor occupancy bit may be obtained, for example, based on the presence of at least one point in the motion-compensated point cloud.

[0079] As discussed herein, FIG. 8A illustrates a TriSoup vertex V k 8. Within cube 800, a TriSoup triangle is a TriSoup vertex V if there are at least three (K≧3) TriSoup vertices on a TriSoup edge of cube 800. kIn the example of FIG. 8A, there are four (4) TriSoup vertices, and therefore a TriSoup triangle may be constructed. The TriSoup triangle may be constructed around centroid vertex C. Centroid vertex C is located at the center of TriSoup vertex V. k To do so, the main direction can first be determined, and then the vertex V k can be ordered by rotating around this direction, and the following K TriSoup triangles can be constructed: V1V2C, V2V3C, ..., V K V1C. The main direction may be chosen from among three directions parallel to the axes of 3D space, for example, to increase or maximize the 2D surface of the triangle when projected along the main direction. In doing so, the main direction may be perpendicular (or partially perpendicular) to the local surface defined by the points of the point cloud belonging to the TriSoup node.

[0080] Furthermore, as mentioned above, FIG. 8B shows the triangle with C+C instead of C as the pivot vertex. res To use the centroid residual value C res We demonstrate a fine-tuning to the TriSoup model by encoding into the bitstream: res can be closer to the points of the point cloud than the centroid C, resulting in a reduction in the reconstruction error and the encoding C res Considering the slight increase in bit rate required, distortion is lower.

[0081] 13 shows an example of encoding centroid residual values. The encoder may encode, for example, the centroid residual value C res can be encoded into the bitstream. res can be used instead of C, for example, as the pivot vertex of a triangle.

number

number

number

number

number

number

number

[0082] The encoder calculates the residual value α res The encoder may determine the residual value α res the current point cloud and line

number

[0083] Residual value α res can be quantified. The residual value α res For example, TriSoup vertex V k The quantization error can be quantified by a uniform quantification function that may have a quantification step similar to the quantification accuracy of V. The quantization error is then calculated for all vertices V such that the local surface can be uniformly approximated. k and / or C+C res The residual value α res may be binarized and / or coded (e.g., entropy coded) into a bitstream. res may be binarized and / or encoded (e.g., entropy coded) into a bitstream, for example, by using a typical unary-based scheme. For example, the residual value α res A flag f0 may be encoded to indicate whether α is equal to zero. res If |α| is not equal to zero, the sign bit may be coded and / or the magnitude of the residual |α res |-1 is the magnitude of the residual |α res continuation flag f indicating whether | is equal to 'i' i (i≧1), the residual value α res is a flag f that can be entropy coded by a coder (e.g., a binary entropy coder). i (i≧0) and / or sign bits.

[0084] Residual value α res The compression of the residual value α res The compression of can be improved by determining the boundaries as shown in FIG.

number

[0085] The coder (e.g., a binary entropy coder) used to encode the binarized residual value α res can be a context-adaptive binary arithmetic coder (CABAC). The probability (or context) used to encode at least 1 bit (f res or the sign bit) of the binarized residual value α i can change, for example, according to the value of at least 1 bit (f res or the sign bit) of the binarized residual value α i The context / probability of the coder (e.g., a binary entropy coder) can be determined, for example, based on context information. The context information can include the values of the boundaries m and M, the location of the vertex V k , or the size of the TriSoup node. The selection of the coder (or the selection of the coder's probability or context) can be implemented by a dynamic OBUF scheme that can use the context information described herein as an entry.

[0086] For lossy coding of dense point clouds, the TriSoup method may be more efficient than octree-only approaches, which may be primarily lossless methods. Even using inter-frame prediction as described herein, octrees may not be competitive with octrees for lossy coding of point clouds. TriSoup is an enhancement of incomplete octrees; for example, because TriSoup may be an enhancement of incomplete octrees, inter-octree prediction may benefit the overall TriSoup scheme by reducing the bitrate of the octrees that TriSoup enhances.

[0087] The TriSoup method, for example, calculates the centroid residual value C res or α res may not be coded based on any reference frame and therefore may not fully benefit from inter-frame correlation. Examples of the present disclosure provide a centroid residual value C that may be coded based on, for example, a motion compensated point cloud. res or α res The motion compensated point cloud may be determined, for example, when the coding of the underlying octree occurs, as discussed herein.

[0088] The centroid residual value C is calculated using the motion compensated point cloud. res or α res Coding σ can add inter-frame correlation to the intra-frame correlation on which probability / context decisions can be based. The addition of inter-frame correlation leads to a better choice of entropy coder (equivalent to its associated probability model) and a better centroid residual value (C res or α res ), resulting in a reduction in the overall amount (e.g., number) of bits required to represent the geometry of the compressed point cloud.

[0089] The motion compensated point cloud is, for example, a centroid residual value C res or α resThe context information may be used to select the probability / context of a coder (e.g., an entropy coder) that may encode at least a portion (or bits) of the ensemble. The selection of the probability model / context may be performed, for example, by a dynamic OBUF scheme. The context information may be constructed, for example, based on a motion-compensated point cloud. The context information may be used, for example, as an entry in (or input to) an OBUF LUT, which may output a coder index to select an entropy coder (equivalent to a probability model).

[0090] 14A and 14B show exemplary TriSoup nodes and compensated points belonging to a motion-compensated point cloud. More specifically, FIG. 14A shows a current TriSoup node 1400 and a compensated point 1410 belonging to the motion-compensated point cloud. The motion-compensated point cloud may be determined, for example, based on a reference point cloud (as described herein) to better match (e.g., be closer to) the point cloud of the current TriSoup node 1400. The occupancy of a point in the motion-compensated point cloud may be determined, for example, based on a centroid residual value C associated with the current TriSoup node 1400. res and / or α res The occupancy of points in the motion compensated point cloud that may be within some distance to the TriSoup node 1400 may be determined, for example, by the centroid residual value C associated with the current TriSoup node 1400. res and / or α res The individual occupancies of the point locations in the motion compensated point cloud may be used directly as bits to determine one or more probability models / contexts for encoding (e.g., entropy encoding, arithmetic encoding) the centroid residual value C associated with the current TriSoup node 1400. res and / or α resThe occupied bits may constitute context information used to determine one or more probability models / contexts for encoding (e.g., entropy encoding, arithmetic encoding) the ensemble of ensemble-of-occupied bits. The amount (e.g., number) of such occupied bits may be large and / or may result in an impractical amount (e.g., number) of probability models / contexts to be stored and / or tracked by the coder. The occupied bits may be combined and / or reduced.

[0091] The centroid residual value C associated with the current TriSoup node 1400 res and / or α res The amount (e.g., number) of compensated points of a motion-compensated point cloud used to determine one or more probability models / contexts for entropy encoding may be reduced, for example, by using compensated points of a motion-compensated point cloud that belong to (and / or are contained within) the current TriSoup node 1400. This may be advantageous because determining whether a compensated point of a motion-compensated point cloud belongs to (and / or can be contained within) the current TriSoup node 1400 may be performed efficiently, for example, when building the underlying octree occurs and / or when searching for compensated points neighboring the TriSoup node is not necessary.

[0092] The centroid residual value C associated with the current TriSoup node 1400 res and / or α res The amount (e.g., number) of compensated points of the motion compensated point cloud used to determine one or more probability models / contexts for encoding σ may be reduced, for example, by using compensated points 1430 of the motion compensated point cloud that belong to (and / or are included within) the current TriSoup node 1400. Additionally or alternatively, the centroid residual value C associated with the current TriSoup node 1400 may be reduced. res and / or α resThe amount (e.g., number) of compensated points in the motion compensated point cloud used to determine one or more probability models / contexts for encoding the line

number

number

number

[0093] 15 illustrates an example TriSoup node and compensated points belonging to a motion-compensated point cloud. As described herein with respect to FIG. 15, the compensated intersection point C comp 1540 is the compensated point 1530 and the line

number

number

number

number

number

[0094] The compensated residual value α comp is the centroid residual value α res The centroid residual value α res The encoding of the compensated residual value α comp The second centroid residual α' can be calculated based on the res (For example, α' res =α res -α comp ) can be determined and / or coded into the bitstream. res The amplitude of the centroid residual value α res The amplitude of the second centroid residual α' may be smaller than that of the res is, for example, the second centroid residual α' res The amplitude of the centroid residual value α res Since the amplitude of res The second centroid residual α' can be compressed more easily than res The above combinations m' and M' can be determined. The second centroid residual α' res may be encoded using methods discussed herein (e.g., binarization, truncated non-aryl encoding, etc.).

[0095] The compensated residual value α compFor example, the centroid residual value α res may be used to select one or more context / probability models for a coder (e.g., entropy coder, arithmetic coder) that can encode the compensated residual value α comp is Q(α comp ) The compensated residual value α comp The quantization function Q(.) of the centroid residual value α res The quantization function of the quantized value Q(α comp ) is, for example, the centroid residual value α res may be used to select one or more context / probability models for a coder (e.g., entropy coder, arithmetic coder) that may encode {circumflex over (x)}.

[0096] Centroid residual value α res The probability / context selection of a coder (e.g., an entropy coder) that may encode a flag f0 that may indicate whether α is equal to zero may be determined, for example, by the probability / context selection of the quantized compensated residual value Q(α comp ) the magnitude of |Q(α comp The probability / context selection can be performed based on, for example, whether the flag f0 being true is a function of the amplitude |Q(α comp )| is small, which can be advantageous. res The choice of probability / context of a coder (e.g., an entropy coder) that can encode the code bits of the quantized compensated residual value Q(α comp ) can be selected based on the sign and magnitude of Q(α comp ) can be advantageous because it can be correlated with the sign of Q(α comp ) may indicate a stronger correlation.

[0097] FIG. 16A illustrates an exemplary method for encoding a centroid residual value of a TriSoup node. More specifically, FIG. 16A illustrates a flowchart 1600 of exemplary method steps for encoding a centroid residual value of a TriSoup node. One or more steps of flowchart 1600 may be implemented by an encoder, such as encoder 114 shown in FIG. 1. In step 1602, the encoder may determine a centroid residual value of the TriSoup node. The TriSoup node may be similar to TriSoup node 1400 shown in and described herein with respect to FIGS. 14A and 14B. For example, the centroid residual value may include three components. A centroid residual value may be a single component (e.g., a linear component, as shown in FIGS. 14A and 14B).

number

[0098] In step 1604, the encoder may select a context / probability model for encoding the centroid residual value. The encoder may select the context / probability model for encoding the centroid residual value, for example, based on a motion-compensated point cloud. The motion-compensated point cloud may be determined, for example, based on a reference point cloud (e.g., as described herein), to better match (e.g., be "closer") to the point cloud of the TriSoup node. The encoder may select the context / probability model, for example, based on a compensated point within the motion-compensated point cloud. For example, the compensated point may be within the TriSoup node and / or may be along a line (e.g., line C as shown in FIGS. 14A and 14B) that may pass through the centroid point of the TriSoup node (e.g., point C as shown in FIGS. 14A and 14B).

number

[0099] The encoder may, for example, generate a compensated intersection point (e.g., a compensated intersection point C as shown in FIG. 15) based on the compensated point. comp ) can be determined. The compensated intersection point can be, for example, a line (e.g., line ∇ ...

number

number

[0100] In step 1606, the encoder may encode (e.g., entropy encode, arithmetically encode) the centroid residual value. The encoder may encode (e.g., entropy encode, arithmetically encode) the centroid residual value, for example, based on a context / probability model. The encoder may encode (e.g., entropy encode, arithmetically encode) the centroid residual value, for example, based on a second centroid residual value α′ discussed herein. resThe encoder may encode the centroid residual value by encoding (e.g., entropy encoding)

[0043] . The encoder may select a context / probability model for encoding the centroid residual value. The encoder may select a context / probability model for encoding the centroid residual value based on, for example, a lookup table that may map a neighboring configuration to a context / probability model. The neighboring configuration may be determined based on, for example, a motion-compensated point cloud.

[0101] The encoder may select a context / probability model for encoding the centroid residual value based on, for example, a lookup table that may map a subset of symbols of a neighboring configuration (e.g., only a subset of symbols of a neighboring configuration) to a context / probability model. The subset of symbols of a neighboring configuration may be determined based on, for example, a motion-compensated point cloud. The amount (e.g., number) of symbols in the subset may increase based on, for example, the amount (e.g., number) of encoded centroid residual values ​​having neighboring information that includes the same subset of symbols. The encoder may update the lookup table to map the subset of symbols of a neighboring configuration to a different context / probability model. The encoder may update the lookup table to map the subset of symbols of a neighboring configuration to a different context / probability model based on, for example, the centroid residual value.

[0102] Figure 16B shows the centroid residual value of the TriSoup node (e.g., C in Figure 8B). res ) More specifically, FIG. 16B shows a flowchart 1650 of example method steps for decoding centroid residual values ​​of a TriSoup node. One or more steps of flowchart 1650 may be implemented by a decoder, such as decoder 120 as shown in FIG. 1.

[0103] In step 1654, the decoder calculates the centroid residual value (e.g., C as shown in FIG. 8B). res) The decoder may select the context / probability model for decoding the centroid residual value based on, for example, a motion-compensated point cloud. The motion-compensated point cloud may be determined, for example, based on a reference point cloud (e.g., as described herein), to better match (or be "closer") to the point cloud of the TriSoup node. The decoder may select the context / probability model based, for example, on a compensated point within the motion-compensated point cloud. For example, the compensated point may be within a TriSoup node and / or may be along a line (e.g., line C as shown in FIGS. 14A and 14B) that may pass through the centroid point of the TriSoup node (e.g., point C as shown in FIGS. 14A and 14B).

number

number

[0104] The decoder may determine the first centroid vertex of the TriSoup (e.g., point C as shown in Figures 8A and 8B). As described herein with respect to Figures 8A and 8B, the decoder may determine the TriSoup vertex V k The decoder may determine the first centroid vertex based on, for example, the TriSoup vertex V of the TriSoup node. k The first centroid vertex may be determined as the average of

[0105] In step 1656, the decoder calculates the centroid residual value (e.g., C as described herein with respect to Figures 8B, 14A, 14B, 15, 16A, and 16B). resThe decoder may decode (e.g., entropy decode, arithmetically decode) the centroid residual value α′, e.g., based on a context / probability model. The decoder may decode (e.g., entropy decode, arithmetically decode) the second centroid residual value α′, e.g., as described herein. res The decoder may decode the centroid residual value by decoding the centroid residual value. The decoder may select a context / probability model for decoding the centroid residual value. The decoder may select the context / probability model for decoding the centroid residual value based on, for example, a lookup table that may map a neighboring configuration to a context / probability model. The neighboring configuration may be determined based on, for example, a motion-compensated point cloud.

[0106] The decoder may select a context / probability model for decoding the centroid residual value, for example, based on a lookup table that may map a subset of symbols of a neighboring configuration (e.g., only a subset of symbols of a neighboring configuration) to a context / probability model. The subset of symbols of a neighboring configuration may be determined, for example, based on a motion-compensated point cloud. The amount (e.g., number) of symbols in the subset may increase, for example, based on the amount (e.g., number) of decoded centroid residual values ​​that have neighboring information that includes the same subset of symbols. The decoder may update the lookup table to map the subset of symbols of a neighboring configuration to a different context / probability model. The decoder may update the lookup table to map the subset of symbols of a neighboring configuration to a different context / probability model, for example, based on the centroid residual value.

[0107] The decoder may determine a second centroid vertex, which may be a line (e.g., line ∂ ...

number

number

number

[0108] The decoder may, for example, map one or more of the compensated points to a line (e.g., line ∇ ...

number

[0109] Figure 17 illustrates an exemplary computer system that may use any of the embodiments described herein. For example, as shown in Figure 32, an exemplary computer system 1700 may implement one or more methods described herein. For example, various devices and / or systems described herein (e.g., Figures 1, 2, and 3) may be implemented in the form of one or more computer systems 1700. Furthermore, each of the steps of the flowcharts illustrated in this disclosure may be implemented on one or more computer systems 1700.

[0110] Computer system 1700 may include one or more processors, such as processor 1704. Processor 1704 may be a special purpose processor, a general purpose processor, a microprocessor, and / or a digital signal processor. Processor 1704 may be connected to a communications infrastructure 1702 (e.g., a bus or network). Computer system 1700 may also include main memory 1706 (e.g., random access memory (RAM) and / or secondary memory 1708).

[0111] The secondary memory 1708 may include a hard disk drive 1710 and / or a removable storage drive 1712 (e.g., a magnetic tape drive, optical disk drive, or the like). The removable storage drive 1712 may read from and / or write to a removable storage unit 1716. The removable storage unit 1716 may include a magnetic tape, optical disk, and / or the like. The removable storage unit 1716 may be read by and / or written to the removable storage drive 1712. The removable storage unit 1716 may include a computer-usable storage medium having computer software and / or data stored therein.

[0112] Secondary memory 1708 may include other similar means for allowing computer programs or other instructions to be loaded into computer system 1700. Such means may include removable storage units 1718 and / or interfaces 1714. Examples of such means may include program cartridges and / or cartridge interfaces (such as for video game devices), removable memory chips (e.g., erasable programmable read-only memory (EPROM) or programmable read-only memory (PROM)), and associated sockets, thumb drives, and universal serial bus (USB) ports, and / or other removable storage units 1718 and interfaces 1714 that may allow software and / or data to be transferred from removable storage units 1718 to computer system 1700.

[0113] Computer system 1700 may include a communications interface 1720. Communications interface 1720 may allow software and / or data to be transferred between computer system 1700 and external devices. Examples of communications interface 1720 may include a modem, a network interface (e.g., an Ethernet card), a communications port, etc. The software and / or data transferred via communications interface 1720 may be in the form of signals, which may be electronic, electromagnetic, optical, and / or other signals that can be received by communications interface 1720. The signals may be provided to communications interface 1720 via communications path 1722. Communications path 1722 may transmit signals and / or may be implemented using wire or cable, fiber optics, a telephone line, a cellular phone link, a radio frequency (RF) link, and / or other communications channels.

[0114] The computer system 1700 may also include one or more sensors 1724. The sensors 1724 may measure and / or detect one or more physical quantities. The sensors 1724 may convert the measured or detected physical quantities into electrical signals in digital and / or analog form. For example, the sensors 1724 may include an eye-tracking sensor for tracking a user's eye movements. The display of the point cloud may be updated based on the user's eye movements. In another example, the sensors 1724 may include a head-tracking sensor (e.g., a gyroscope) for tracking a user's head movements. The display of the point cloud may be updated based on the user's head movements. In yet another example, the sensors 1724 may include a camera sensor for taking photographs and / or a 3D scanning device (e.g., a laser scanning, structured light scanning, and / or modulated light scanning device). The 3D scanning device may acquire geometric shape information by moving one or more laser heads, structured light, and / or modulated light cameras relative to the object or scene being scanned. The geometric shape information may be used to construct a point cloud.

[0115] Computer program medium and / or computer-readable medium may be used to refer to tangible (e.g., non-transitory) storage media, such as removable storage units 1716 and 1718, or a hard disk installed in hard disk drive 1710. These computer program products may be means for providing software to computer system 1700. Computer programs (also called computer control logic) may be stored in main memory 1706 and / or secondary memory 1708. Computer programs may be received via communications interface 1720. When executed, these computer programs may enable computer system 1700 to implement one or more exemplary embodiments of the present disclosure, as discussed herein. In particular, when executed, the computer programs may enable processor 1704 to perform processes of the present disclosure, such as any of the methods described herein. Thus, these computer programs may represent controllers of computer system 1700.

[0116] 18 shows exemplary elements of a computing device that may be used to implement any of the various devices described herein. For example, a source device (e.g., 102), an encoder (e.g., 114), a destination device (e.g., 106), a decoder (e.g., 120), and / or any computing device may be described herein. The computing device 1830 may include one or more processors 1831 that may execute instructions stored on random access memory (RAM) 2233, removable media 1834 (e.g., a universal serial bus (USB) drive, a compact disc (CD) or digital versatile disc (DVD), or a floppy disk drive), or any other desired storage medium. Instructions may also be stored on an attached (or internal) hard drive 1835. Computing device 1830 may also include a security processor (not shown) that may execute instructions of one or more computer programs to monitor processes running on processor 1831 and any processes requesting access to any hardware and / or software components of computing device 1830 (e.g., ROM 1832, RAM 1833, removable media 1834, hard drive 1835, device controller 1837, network interface 1839, GPS 1841, Bluetooth interface 1842, WiFi interface 1843, etc.). Computing device 1830 may include one or more output devices such as a display 1836 (e.g., a screen, display device, monitor, television, etc.) and may include one or more output device controllers 1837 such as a video processor. There may also be one or more user input devices 1838 such as a remote control, keyboard, mouse, touchscreen, microphone, etc. Computing device 1830 may also include one or more network interfaces such as network interface 1839, which may be a wired interface, a wireless interface, or a combination of the two.The network interface 1839 may provide an interface through which the computing device 1830 communicates with a network 1840 (e.g., a RAN, or some other network). The network interface 1839 may include a modem (e.g., a cable modem), and the external network 1840 may include a communications link, an external network, a home network, a provider's wireless, coaxial, fiber, or hybrid fiber / coaxial distribution system (e.g., a DOCSIS network), or any other desired network. Additionally, the computing device 1830 may include a location detection device such as a global positioning system (GPS) microprocessor 1841, which may be configured to receive and process global positioning signals and, with possible assistance from external servers and antennas, determine the geographic location of the computing device 1830.

[0117] While the example of FIG. 18 may be a hardware configuration, the components shown may be implemented as software. Changes may be made, as desired, to add, remove, combine, divide, etc., components of computing device 1830. Additionally, components may be implemented using basic computing devices and components, and the same components (e.g., processor 1831, ROM storage 1832, display 1836, etc.) may be used to implement any of the other computing devices and components described herein. For example, the various components described herein may be implemented using a computing device having components such as a processor that executes computer-executable instructions stored on a computer-readable medium, as shown in FIG. 18. Some or all of the entities described herein may be software-based and coexist on a common physical platform (e.g., a requesting entity may be a separate software process and program from a dependent entity, both of which may run as software on a common computing device).

[0118] Various features are highlighted below in sets of numbered clauses or paragraphs. These features are not to be construed as limiting the invention or inventive concept, but are provided merely as highlighting some of the features described herein, without implying the importance or relevance of any particular order of such features.

[0119] Clause 1. A method comprising determining a first centroid vertex of a TriSoup node.

[0120] Clause 2. The method of clause 1, further comprising selecting a context for decoding the centroid residual value based on the motion compensated point cloud.

[0121] Clause 3. The method of clause 1 or 2, further comprising decoding the centroid residual value based on the selected context.

[0122] Clause 4. The method of any one of clauses 1 to 3, further comprising determining a second centroid vertex based on the first centroid vertex and the decoded centroid residual value.

[0123] Clause 5. The method of any one of clauses 1 to 4, wherein the centroid residual value comprises at least a first component.

[0124] Clause 6. The method of any one of clauses 1 to 5, wherein selecting a context further comprises selecting a context based on a compensated point in the motion compensated point cloud.

[0125] Clause 7. The method of any one of clauses 1 to 6, wherein the compensated point is within at least one of a TriSoup node or a distance to a line intersecting the centroid value of the TriSoup node.

[0126] Clause 8. The method of any one of clauses 1 to 7, further comprising determining a compensated intersection point based on a compensated point in the motion compensated point cloud.

[0127] Clause 9. The method of any one of clauses 1 to 8, wherein selecting further comprises selecting a context based on the compensated intersection.

[0128] Clause 10. The method of any one of clauses 1 to 9, wherein the second centroid vertex belongs to a line that passes through the first centroid vertex.

[0129] Clause 11. The method of any one of clauses 1 to 10, wherein the motion compensated point cloud is associated with a video frame or a point cloud frame.

[0130] Clause 12. The method of any one of clauses 1 to 11, wherein selecting further comprises selecting the context based on an association between the context and a neighboring configuration associated with the motion compensated point cloud.

[0131] Clause 13. The method of any one of clauses 1 to 12, wherein the centroid residual value comprises three components.

[0132] Clause 14. The method of any one of clauses 1 to 13, further comprising selecting a context for decoding the centroid residual value based on an association between a subset of symbols of a neighboring configuration and the context, the association being determined based on the motion-compensated point cloud.

[0133] Clause 15. The method of any one of clauses 1 to 14, wherein the number of symbols in a subset of symbols is increased based on the number of decoded centroid residual values ​​having neighboring configuration information that includes the same subset of symbols.

[0134] Clause 16. A computing device comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of clauses 1 to 15.

[0135] Clause 17. A system comprising: a first computing device configured to perform the method of any one of clauses 1 to 15; and a second computing device configured to encode point cloud frames or video frames.

[0136] Clause 18. A computer-readable medium storing instructions that, when executed, cause performance of the method of any one of clauses 1 to 15.

[0137] Clause 19. A method comprising determining a first centroid vertex of a TriSoup node.

[0138] Clause 20. The method of clause 19, further comprising selecting a context for decoding the centroid residual value based on an association between the context and a subset of symbols of the adjacent configuration, the association being determined based on the motion compensated point cloud.

[0139] Clause 21. The method of clause 19 or 20, further comprising decoding the centroid residual value based on the selected context.

[0140] Clause 22. The method of any one of clauses 19 to 21, further comprising determining a second centroid vertex based on the first centroid vertex and the decoded residual value.

[0141] Clause 23. The method of any one of clauses 19 to 22, wherein the number of symbols in a subset of symbols is increased based on the number of decoded centroid residual values ​​having neighboring configuration information that includes the same subset of symbols.

[0142] Clause 24. The method of any one of clauses 19 to 23, further comprising updating the associations to associate subsets of symbols of a neighboring configuration with different contexts based on centroid residual values.

[0143] Clause 25. The method of any one of clauses 19 to 24, further comprising determining a compensated intersection point based on a compensated point in the motion compensated point cloud, wherein the compensated intersection point belongs to a line passing through the first centroid vertex.

[0144] Clause 26. The method of any one of clauses 19 to 25, wherein the motion compensated point cloud is associated with a video frame or a point cloud frame.

[0145] Clause 27. The method of any one of clauses 19 to 26, wherein the compensated point is within at least one of a TriSoup node or a distance to a line intersecting the centroid value of the TriSoup node.

[0146] Clause 28. The method of any one of clauses 19 to 27, wherein selecting a context further comprises selecting a context based on quantization of a compensated centroid residual value associated with the compensated intersection point.

[0147] Clause 29. A computing device comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of clauses 19 to 28.

[0148] Clause 30. A system comprising: a first computing device configured to perform the method of any one of clauses 19 to 28; and a second computing device configured to encode point cloud frames or video frames.

[0149] Clause 31. A computer-readable medium storing instructions that, when executed, cause performance of the method of any one of clauses 19 to 28.

[0150] Clause 32. A method comprising determining a centroid residual value for a TriSoup node.

[0151] Clause 33. The method of clause 32, further comprising selecting a context for encoding the centroid residual value based on the motion compensated point cloud.

[0152] Clause 34. The method of clause 32 or 33, further comprising encoding the centroid residual value based on the selected context.

[0153] Clause 35. The method of any one of clauses 32 to 34, wherein selecting further comprises selecting a context for encoding the centroid residual value based on an association between a neighboring configuration and the context associated with the motion compensated point cloud.

[0154] Clause 36. The method of any one of clauses 32 to 35, wherein the centroid residual value comprises at least a first component.

[0155] Clause 37. The method of any one of clauses 32 to 36, wherein selecting a context further comprises selecting a context based on a compensated point in the motion compensated point cloud.

[0156] Clause 38. The method of any one of clauses 32 to 37, wherein the compensated intersection point belongs to a line passing through the centroid point of the TriSoup node.

[0157] Clause 39. The method of any one of clauses 32 to 37, wherein determining the compensated intersection points further comprises projecting one or more of the compensated points onto a line passing through the centroid point of the TriSoup node.

[0158] Clause 40. The method of any one of clauses 32 to 37, wherein determining the compensated intersection points further comprises averaging projections of one or more of the compensated points on the line.

[0159] Clause 41. The method of any one of clauses 32 to 37, wherein the number of symbols in a subset is increased based on the number of coded centroid residual values ​​having neighboring information that includes the same subset of symbols.

[0160] Clause 42. The method of any one of clauses 32-37, further comprising updating a lookup table to map subsets of symbols of adjacent configurations to different contexts based on centroid residual values.

[0161] Clause 43. A computing device comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of clauses 32 to 42.

[0162] Clause 44. A system comprising: a first computing device configured to perform the method of any one of clauses 32 to 42; and a second computing device configured to encode point cloud frames or video frames.

[0163] Clause 45. A computer-readable medium storing instructions that, when executed, cause performance of the method of any one of clauses 32 to 42.

[0164] A computing device may perform a method including a plurality of operations. The computing device may determine a first centroid vertex of a TriSoup node. The computing device may select a context for decoding a centroid residual value, for example, based on a motion-compensated point cloud. The computing device may decode the centroid residual value, for example, based on the selected context. The computing device may determine a second centroid vertex, for example, based on the first centroid vertex and the decoded centroid residual value. The centroid residual value may include at least a first component. The computing device may select the context, for example, based on a compensated point in the motion-compensated point cloud. The compensated point may be within at least one of a distance to the TriSoup node or a line intersecting the centroid value of the TriSoup node. The computing device may determine a compensated intersection point, for example, based on the compensated point in the motion-compensated point cloud. The computing device may select the context, for example, based on the compensated intersection point. The second centroid vertex may belong to a line passing through the first centroid vertex. The motion-compensated point cloud may be associated with a video frame or a point cloud frame. The computing device may select a context based on, for example, an association between a neighboring configuration and the context associated with the motion-compensated point cloud. The centroid residual value may include three components. The computing device may select a context for decoding the centroid residual value based on, for example, an association between a subset of symbols of the neighboring configuration and the context, determined based on the motion-compensated point cloud. The number of symbols in the subset of symbols may increase based on, for example, the number of decoded centroid residual values ​​having neighboring configuration information that includes the same subset of symbols. The computing device may include one or more processors and memory storing instructions that, when executed by the one or more processors, cause the computing device to perform the described methods, additional operations, and / or include additional elements.The system may include a first computing device configured to perform the described methods, additional operations, and / or include additional elements, and a second computing device configured to encode (e.g., encode or decode) the video frames, point cloud frames, or point cloud sequences. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.

[0165] A computing device may perform a method including multiple operations. The computing device may determine a first centroid vertex of a TriSoup node. The computing device may select a context for decoding a centroid residual value based on an association between a subset of symbols of a neighboring configuration and the context, for example, determined based on the motion-compensated point cloud. The computing device may decode the centroid residual value based on the selected context. The computing device may determine a second centroid vertex based on the first centroid vertex and the decoded residual value. The number of symbols in the subset of symbols may increase based on, for example, the number of decoded centroid residual values ​​having neighboring configuration information including the same subset of symbols. The computing device may update the association based on, for example, the centroid residual value to associate the subset of symbols of a neighboring configuration with a different context. The computing device may determine a compensated intersection based on, for example, a compensated point in the motion-compensated point cloud. The compensated intersection may belong to a line passing through the first centroid vertex. The motion-compensated point cloud may be associated with a video frame or a point cloud frame. The compensated point may be within at least one of a distance to a TriSoup node or a line intersecting with a centroid value of the TriSoup node. The computing device may select the context based on, for example, a quantization of a compensated centroid residual value associated with the compensated intersection point. The computing device may include one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform the described methods, additional operations, and / or include additional elements. The system may include a first computing device configured to perform the described methods, additional operations, and / or include additional elements, and a second computing device configured to code (e.g., encode or decode) a video frame, a point cloud frame, or a point cloud sequence.The computer-readable medium may store instructions that, when executed, perform the methods described, cause additional actions, and / or include additional elements.

[0166] A computing device may perform a method including multiple operations. The computing device may determine a centroid residual value of a TriSoup node. The computing device may select a context for encoding the centroid residual value, for example, based on a motion-compensated point cloud. The computing device may encode the centroid residual value, for example, based on the selected context. The computing device may select a context for encoding the centroid residual value, for example, based on an association between a neighboring configuration associated with the motion-compensated point cloud and a context. The centroid residual value may include at least a first component. The computing device may select the context, for example, based on a compensated point in the motion-compensated point cloud. The compensated intersection point may belong to a line that passes through the centroid point of the TriSoup node. The computing device may determine the compensated intersection point, for example, based on projecting one or more of the compensated points onto a line that passes through the centroid point of the TriSoup node. The computing device may determine the compensated intersection point, for example, based on averaging projections of one or more of the compensated points on the line. The number of symbols in a subset may increase based on, for example, the number of coded (e.g., encoded or decoded) centroid residual values ​​having neighboring information that includes the same subset of symbols. The computing device may update the lookup table to map neighboring subsets of symbols to different contexts, for example, based on the centroid residual values. The computing device may include one or more processors and memory that stores instructions that, when executed by the one or more processors, cause the computing device to perform the described methods, additional operations, and / or include additional elements.The system may include a first computing device configured to perform the described methods, additional operations, and / or include additional elements, and a second computing device configured to encode (e.g., encode or decode) the video frames, point cloud frames, or point cloud sequences. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.

[0167] A computing device may perform a method including multiple operations. The computing device may determine a centroid residual value of a TriSoup node. The computing device may select a context / probability model for encoding the centroid residual value, for example, based on a motion-compensated point cloud. The computing device may entropy encode the centroid residual value, for example, based on the context / probability model. The centroid residual value may include three components. The centroid residual value may include a single component. The computing device may select the context / probability model, for example, based on a compensated point in the motion-compensated point cloud. The compensated point may be within the TriSoup node. The compensated point may be within a distance to a line that intersects with the centroid value of the TriSoup node. The computing device may determine a compensated intersection point, for example, based on the compensated point. The compensated intersection point may belong to a line that passes through the centroid point of the TriSoup node. The computing device may determine the compensated intersection point, for example, based on projecting one or more of the compensated points onto a line that passes through the centroid point of the TriSoup node. The computing device may determine the compensated intersection point, for example, based on averaging projections of one or more of the compensated points onto the line. The computing device may select a context / probability model, for example, based on the compensated intersection point. The computing device may determine a second residual, for example, based on a difference between a centroid residual value and a compensated centroid residual value determined based on the compensated intersection point. The computing device may entropy encode the second residual. The computing device may select a context / probability model, for example, based on the compensated centroid residual value. The computing device may select a context / probability model, for example, based on quantization of the compensated centroid residual value.The computing device may select a context / probability model for encoding the centroid residual value, for example, based on a lookup table that maps a neighboring configuration determined based on a motion-compensated point cloud to a context / probability model. The computing device may select a context / probability model for encoding the centroid residual value, for example, based on a lookup table that maps only a subset of symbols of a neighboring configuration determined based on a motion-compensated point cloud to a context / probability model. The number of symbols in the subset may increase based on, for example, the number of coded centroid residual values ​​with neighboring information that may include the same subset of symbols. The computing device may update the lookup table to map a subset of symbols of a neighboring configuration to a different context / probability model, for example, based on the centroid residual value. The computing device may include one or more processors and memory that stores instructions that, when executed by the one or more processors, cause the computing device to perform the described methods, additional operations, and / or include additional elements. The system may include a first computing device configured to perform the described methods, additional operations, and / or include additional elements, and a second computing device configured to encode (e.g., encode or decode) the video frames, point cloud frames, or point cloud sequences. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.

[0168] One or more examples herein may be described as a process, which may be depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, and / or a block diagram. While a flowchart may describe operations as a sequential process, one or more operations may be performed in parallel or concurrently. The order of operations shown may be rearranged. A process may terminate when its operations are completed, but may have additional steps not shown in the figures. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.

[0169] The operations described herein may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program product) to perform the necessary tasks may be stored on a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Features of the present disclosure may be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and / or gate arrays. Implementation of hardware state machines to perform the functions described herein will also be apparent to those skilled in the art.

[0170] One or more features described herein may be implemented in computer-usable data and / or computer-executable instructions, such as one or more program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types when executed by a processor within a computer or other data processing device. Computer-executable instructions may be stored on one or more computer-readable media, such as hard disks, optical disks, removable storage media, solid-state memory, RAM, etc. The functionality of the program modules may be combined or distributed as desired. Functionality may be implemented in whole or in part in firmware or hardware equivalents, such as integrated circuits, field programmable gate arrays (FPGAs), etc. Particular data structures may be used to more effectively implement one or more features described herein, and such data structures are contemplated within the scope of the computer-executable instructions and computer-usable data described herein. Computer-readable media may include, but are not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media on which data may be stored and which do not include carrier waves and / or transitory electronic signals propagated via wireless or wired connections. Examples of non-transitory media include, but are not limited to, magnetic disks or tapes, optical storage media such as compact disks (CDs) or digital versatile disks (DVDs), flash memory, memory or memory devices. Computer-readable media may store code and / or machine-executable instructions, which may represent procedures, functions, subprograms, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements.A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0171] A non-transitory tangible computer-readable medium may include instructions executable by one or more processors configured to cause the operations described herein. An article of manufacture may include a non-transitory tangible computer-readable machine-accessible medium encoded with instructions to enable programmable hardware that causes a device (e.g., an encoder, decoder, transmitter, receiver, etc.) to perform the operations described herein. A device, or one or more devices, such as in a system, may include one or more processors, memory, interfaces, and / or the like.

[0172] Communications described herein may be determined, generated, sent, and / or received using any amount of messages, information elements, fields, parameters, values, indications, information, bits, and / or the like. While one or more examples may be described herein using any of the terms / phrases message, information element, field, parameter, value, indication, information, bit, and / or the like, those skilled in the art will understand that such communications may be implemented using any one or more of these terms, including other such terms. For example, one or more parameters, fields, and / or information elements (IEs) may include one or more information objects, values, and / or any other information. An information object may include one or more other objects. At least some (or all) parameters, fields, IEs, and / or the like may be used and may be interchangeable depending on the context. Where meanings or definitions are given, such meanings or definitions are controlling.

[0173] One or more elements of the examples described herein may be implemented as a module. A module may be an element that performs a defined function and / or has a defined interface to other elements. A module may be implemented in hardware, software in combination with hardware, firmware, wetware (e.g., hardware with biological components), or a combination thereof, all of which may be behaviorally equivalent. For example, a module may be implemented as a software routine written in a computer language configured to run on a hardware machine (e.g., C, C++, Fortran, Java, Basic, Matlab, etc.) or Simulink, Stateflow, GNU Octave, or LabVIEW MathScript. Additionally or alternatively, it may be possible to implement a module using physical hardware incorporating discrete or programmable analog, digital, and / or quantum hardware. Examples of programmable hardware include computers, microcontrollers, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and / or complex programmable logic devices (CPLDs). Computers, microcontrollers, and / or microprocessors may be programmed using languages ​​such as assembly, C, C++, etc. FPGAs, ASICs, and CPLDs are often programmed using hardware description languages ​​(HDLs) such as Verilog or Verilog Hardware Description Language (VHDL), which allow for the construction of connections between the less functional internal hardware modules of the programmable device. The techniques described above can be used in combination to achieve a functionally modular result.

[0174] One or more of the operations described herein may be conditional. For example, one or more operations may be performed if certain criteria are met, such as by a computing device, a communication device, an encoder, a decoder, a network, a combination of the above, and / or the like. Exemplary criteria may be based on one or more conditions of device configuration, traffic load, initial system setup, packet size, traffic characteristics, a combination of the above, and / or the like. Various examples may be used when one or more criteria are met. It may be possible to implement any part of the examples described herein in any order and based on any condition.

[0175] While examples are described above, features and / or steps of these examples may be combined, divided, omitted, rearranged, revised, and / or extended in any desired manner. Various changes, modifications, and improvements will readily occur to those skilled in the art. Such changes, modifications, and improvements, although not expressly described herein, are intended to be part of this specification and are intended to be within the spirit and scope of the description herein. Accordingly, the foregoing description is illustrative only and not limiting.

Claims

1. 1. A method comprising: determining a first centroid vertex of the TriSoup node; selecting a context for decoding the centroid residual value based on the motion compensated point cloud; decoding the centroid residual value based on the selected context; determining a second centroid vertex based on the first centroid vertex and the decoded centroid residual value.

2. The method of claim 1 , wherein the centroid residual value includes at least a first component.

3. selecting the context, The method of claim 1 or 2, further comprising selecting the context based on a compensated point in the motion compensated point cloud.

4. The compensated point is the TriSoup node, or The method according to any one of claims 1 to 3, wherein the distance to the line of intersection with the centroid value of the TriSoup node is within at least one of the following:

5. The method of any one of claims 1 to 4, further comprising determining a compensated intersection point based on a compensated point in the motion compensated point cloud.

6. The selecting The method of any one of claims 1 to 5, further comprising selecting the context based on the compensated intersection.

7. The method according to any one of claims 1 to 6, wherein the second centroid vertex belongs to a line that passes through the first centroid vertex.

8. The method of any one of claims 1 to 7, wherein the motion compensated point cloud is associated with a video frame.

9. The selecting The method of any one of claims 1 to 8, further comprising selecting the context based on an association between the context and a neighboring configuration associated with the motion compensated point cloud.

10. The method of any one of claims 1 to 9, wherein the centroid residual value comprises three components.

11. 11. The method of claim 1, further comprising selecting a context for decoding a centroid residual value based on an association between the context and a subset of symbols of a neighboring configuration, the association being determined based on a motion compensated point cloud.

12. The method of claim 11 , wherein the number of symbols in the subset of symbols is increased based on the number of decoded centroid residual values ​​having neighboring configuration information that includes the same subset of symbols.

13. 1. A computing device comprising: one or more processors; a memory storing instructions that, when executed by said one or more processors, cause said computing device to perform a method according to any one of claims 1 to 12.

14. 1. A system comprising: a first computing device configured to perform the method of any one of claims 1 to 12; a second computing device configured to encode the point cloud frames or the video frames.

15. A computer readable medium storing instructions that, when executed, cause the performance of a method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

    WO2022169176A1

  • Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

    WO2022186675A1