Centroid localization for voxelization of triangles in point cloud coding
By employing voxelization and occupancy tree coding techniques, the problem of large data volumes in point cloud data is solved, achieving efficient compression and low-distortion data transmission, which is suitable for applications such as extended reality, virtual reality, and mixed reality.
Patent Information
- Application Number
- CN202480039404.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-13
- Filing Date
- 2024-04-12
- Publication Date
- 2026-02-24
AI Technical Summary
The large size of point cloud data leads to low transmission and processing efficiency, and existing compression technologies struggle to maintain a high compression ratio while avoiding rendering distortion.
By converting point cloud data into cuboids, encoding is performed using occupancy tree and Morton order, combined with entropy coding to optimize data transmission, reduce redundant information, and adopt the G-PCC standard for encoding and decoding.
It achieves efficient compression and transmission of point cloud data, reduces rendering distortion, is suitable for various application scenarios such as augmented reality, virtual reality and mixed reality, and improves data processing efficiency.
Smart Images

Figure CN121569486A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 459,242, filed April 13, 2023. The entire contents of the application cited above are incorporated herein by reference. Background Technology
[0003] Objects or scenes can be described using volumetric visual data containing a series of points. Points can be stored in point cloud format, which consists of a set of points in three-dimensional space. Because point cloud data can be quite large, transmitting and processing point cloud data may require data compression methods specifically designed for the unique characteristics of point cloud data. Summary of the Invention
[0004] The following summary presents a simplified overview of certain features. This summary is not a comprehensive overview and is not intended to identify any important or key elements.
[0005] A cuboid may contain a centroid and multiple vertices. The centroid and vertices can form triangles, which can be used to represent portions of a point cloud, representing objects or scenes within the content. Changes in the position of the vertices can cause a significant shift in the position of the centroid, potentially leading to distortion during encoding (e.g., encoding / decoding). Assigning weights to each vertex can adjust the position of the centroid to make it more equidistant from the vertices. These weights can be determined, for example, based on the area of the triangle formed between the centroid and the vertices or the distance between the vertices. By determining the optimal position of the centroid, distortion in rendering can be minimized.
[0006] These and other features and advantages are described in more detail below. Attached Figure Description
[0007] Examples of several embodiments of the various embodiments of this disclosure are described herein with reference to the accompanying drawings.
[0008] Figure 1 An example point cloud encoding system is shown.
[0009] Figure 2 An example of Morton order is shown.
[0010] Figure 3 An example scan order is shown.
[0011] Figure 4 An example neighborhood of a cuboid with already encoded occupant bits is shown.
[0012] Figure 5 An example of the dynamically decreasing function DR in the optimal binary encoder (OBUF) that can be used to dynamically support instant updates is shown.
[0013] Figure 6 An example method for encoding the occupancy of a cuboid using dynamic OBUF is shown.
[0014] Figure 7 An example of an occupied cuboid is shown.
[0015] Figure 8A An example cuboid corresponding to a TriSoup node is shown.
[0016] Figure 8B An example refinement of the TriSoup model is shown.
[0017] Figure 9 An example of voxelization is shown.
[0018] Figure 10 An example of the coordinates of a point relative to the centroid of the TriSoup triangle is shown.
[0019] Figure 11 An example of encoding centroid residual values is shown.
[0020] Figure 12A An example of the centroid vertex determined from the TriSoup vertices of a cuboid is shown.
[0021] Figure 12B An example of the centroid vertex determined from the TriSoup vertices of a cuboid is shown.
[0022] Figure 13 An example of a TriSoup triangle in a cuboid is shown.
[0023] Figure 14A An example of determining the centroid vertex is shown.
[0024] Figure 14B An example of determining the weights of TriSoup vertices is shown.
[0025] Figure 15 An example method for determining the centroid vertex is shown.
[0026] Figure 16 An example method for obtaining a portion of the point cloud in a cuboid is shown.
[0027] Figure 17A An example method for encoding the centroid residual of a centroid vertex is shown.
[0028] Figure 17B An example method for decoding the centroid residual of a centroid vertex is shown.
[0029] Figure 18A block diagram of an example computer system in which instances of the present disclosure may be implemented is shown.
[0030] Figure 19 Example elements of a computing device are shown that can be used to implement any of the various devices described herein. Detailed Implementation
[0031] The accompanying drawings and description provide examples. It should be understood that the examples shown and / or described in the drawings are non-exclusive, and the features shown and described can be practiced in other examples. Examples of operation for point cloud or point cloud sequence encoding or decoding systems are provided. More specifically, the techniques disclosed herein can relate to point cloud compression, such as that used in encoding and / or decoding apparatuses and / or systems.
[0032] At least some visual data can use a series of points to describe objects or scenes in content and / or media. Each point can contain a position in two-dimensional (x and y) form and one or more optional attributes, such as color. Volumetric visual data can add another positional dimension to these visual data. For example, volumetric visual data can use a series of points to describe objects or scenes in content and / or media, each point can contain a position in three-dimensional (x, y, and z) form and one or more optional attributes, such as color, reflectivity, timestamp, etc. For example, volumetric visual data can provide a more immersive way to experience visual data than the aforementioned at least some visual data. For example, objects or scenes described by volumetric visual data can be viewed from any (or more) angles, while objects or scenes described by the aforementioned at least some visual data can typically only be viewed from the angle from which the object or scene is captured or rendered. As a representation format for visual data (e.g., volumetric visual data, 3D video data, etc.), point clouds are universal because they can represent all types of three-dimensional (3D) objects, scenes, and visual content. Point clouds are well-suited for a wide range of applications, including but not limited to: film post-production, real-time 3D immersive media or telepresence, extended reality, free-view video, geographic information systems, autonomous driving, 3D mapping, visualization, medicine, multi-view replay, and real-time light detection and ranging (LiDAR) data acquisition.
[0033] As explained in this article, volumetric visual data can be used in many applications, including extended reality (XR). XR encompasses various types of immersive technologies, including augmented reality (AR), virtual reality (VR), and mixed reality (MR). Sparse volumetric visual data can be used in the automotive industry to represent three-dimensional (3D) maps (e.g., cartography) or as input to driver assistance systems. In the case of driver assistance systems, volumetric visual data can often be fed into driving decision-making algorithms. Volumetric visual data can be used to digitally store valuable objects. In applications for cultural heritage preservation, the goal can be to maintain a representation of objects that may be threatened by natural disasters. For example, statues, vases, and temples can be fully scanned and stored as volumetric visual data with billions of samples. This use case for volumetric visual data may be particularly relevant to valuable objects in locations prone to earthquakes, tsunamis, and typhoons. Volumetric visual data can take the form of volumetric frames. A volumetric frame can describe an object or scene captured at a specific time instance. Volumetric visual data can also take the form of a sequence of volumetric frames (referred to as a volumetric sequence or volumetric video). A volumetric frame sequence can describe an object or scene captured at multiple different time instances.
[0034] Volumetric visual data can be stored in various formats. One format for storing volumetric visual data can be a point cloud. A point cloud can contain a collection of points in 3D space. Such points can be used to create meshes containing vertices and polygons, or other forms of visual content. As described herein, point cloud data can take the form of point cloud frames that describe objects or scenes in content captured in a specific time instance. Point cloud data can also take the form of a sequence of point cloud frames (e.g., point cloud video). As further described herein, point cloud data can be generated from a source device (e.g., as described herein regarding...). Figure 1 The source device 102 encodes the point cloud data, and outputs a bitstream containing the encoded point cloud data. The source device may encode the point cloud data based on point cloud compression coding, such as geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding, or next-generation coding. The destination device (e.g., as described herein regarding...) Figure 1The destination device 106 receives a bitstream containing point cloud data and decodes the bitstream containing point cloud data. The destination device can decode the point cloud data by performing point cloud decompression encoding. Decompression encoding can be the reverse process of point cloud compression encoding. Point cloud decompression encoding can include, for example, G-PCC encoding. Decoding can be used to decompress the point cloud data for display and / or other forms of consumption (e.g., further analysis, storage, etc.). The destination device (or different devices) can include, for example, a renderer for rendering the decoded point cloud data. The renderer can output content, for example, by rendering the point cloud data. The renderer can output content, for example, by rendering the point cloud data along with other data (e.g., audio data).
[0035] A point cloud can contain a collection of points in 3D space. Each point in a point cloud can contain geometric information that indicates the point's location in 3D space. For example, the geometric information can indicate the point's location in 3D space using, for example, three Cartesian coordinates (x, y, and z) and / or spherical coordinates (r, φ, θ) (e.g., if acquired by a rotation sensor). The location of points in a point cloud can be quantized according to spatial precision. Spatial precision can be the same or different in each dimension. The quantization process can create a grid in 3D space. One or more points residing within each sub-grid volume can be mapped to the coordinates of the sub-grid center, referred to as voxels. A voxel can be viewed as a 3D extension of a pixel corresponding to a 2D image grid coordinate. Points in a point cloud can contain one or more types of attribute information. Attribute information can indicate the properties of the point's visual appearance. For example, attribute information can indicate the point's texture (e.g., color), the point's material type, the point's transparency information, the point's reflectivity information, the surface normal vector of the point, the velocity at the point, the acceleration at the point, a timestamp indicating when the point was captured, or an indication of how the point's modality was captured (e.g., running, walking, or flying). Points in a point cloud can contain light field data in the form of multi-view related texture information. The light field data can also be another type of optional attribute information.
[0036] Points in a point cloud can describe objects or scenes. For example, points in a point cloud can describe the external surfaces and / or internal structures of an object or scene. Objects or scenes can be generated synthetically by computer. Objects or scenes can be generated from captures of real-world objects or scenes. Geometric information of real-world objects or scenes can be obtained through 3D scanning and / or photogrammetry. 3D scanning can include different types of scanning, such as laser scanning, structured light scanning, and / or modulated light scanning. 3D scanning can obtain geometric information. 3D scanning can obtain geometric information, for example, by moving one or more laser heads, structured light cameras, and / or modulated light cameras relative to the scanned object or scene. Photogrammetry can obtain geometric information. Photogrammetry can obtain geometric information, for example, by triangulating the same features or points in 2D photographs at different spatial displacements. Point cloud data can be in the form of point cloud frames. Point cloud frames can describe objects or scenes captured at a specific time instance. Point cloud data can be in the form of a sequence of point cloud frames. A sequence of point cloud frames can be referred to as a point cloud sequence or point cloud video. Point cloud frame sequences can describe objects or scenes captured at multiple different time instances.
[0037] In many applications, the data size of a point cloud frame or sequence of point clouds may be too large for storage and / or transmission (e.g., too big). For example, a single point cloud may contain, for example, more than one million points or even billions of points. Each point may contain geometric information as well as one or more optional types of attribute information. The geometric information of each point may contain three Cartesian coordinates (x, y, and z) and / or spherical coordinates (r, φ, θ), each Cartesian and / or spherical coordinate may be represented, for example, using at least 10 bits per component or 30 bits in total. The attribute information of each point may contain textures corresponding to multiple (e.g., three) color components (e.g., R, G, and B color components). Each color component may be represented, for example, using 8-10 bits per component or 24-30 bits in total. For example, a single point may contain at least 54 bits of information, with at least 30 bits of geometric information and at least 24 bits of texture. If a point cloud frame includes one million such points, each point cloud frame may require 54 million bits or 54 megabits to represent. For dynamic point clouds that change over time, at a frame rate of 30 frames per second, a data rate of 1.32 gigabits per second may be required to send (e.g., transmit) the points of a point cloud sequence. The raw representation of the point cloud may require a large amount of data, and the practical deployment of point cloud-based technologies may require compression techniques that enable the storage and distribution of point clouds at a reasonable cost.
[0038] Encoding can be used to compress and / or reduce the data size of point cloud frames or sequences to provide more efficient storage and / or transmission. Decoding can be used to decompress compressed point cloud frames or sequences for display and / or other forms of consumption (e.g., other forms of consumption by machine learning-based devices, neural network-based devices, artificial intelligence-based devices, or other types of consumption by other types of machine-based processing algorithms and / or devices). For example, distribution to and visualization by end users on AR or VR glasses or any other 3D-enabled devices, point cloud compression may be lossy (introducing differences relative to the original data). Lossy compression can allow high compression ratios but may imply a trade-off between compression and visual quality perceived by the end user. Other frameworks, such as those used in medical applications or autonomous driving, may require lossless compression to avoid altering decisions obtained, for example, based on analysis of sent (e.g., transmitted) and decompressed point cloud frames.
[0039] Figure 1 An example point cloud encoding (e.g., encoding / decoding) system 100 is illustrated. The point cloud encoding system 100 may include a source device 102, a transmission medium 104, and a destination device 106. The source device 102 may encode a point cloud sequence 108 into a bit stream 110 for more efficient storage and / or transmission. The source device 102 may store the bit stream 110 and / or send (e.g., transmit) the bit stream to the destination device 106 via the transmission medium 104. The destination device 106 may decode the bit stream 110 to display the point cloud sequence 108 or for other forms of consumption (e.g., further analysis, storage, etc.). The destination device 106 may receive the bit stream 110 from the source device 102 via the storage medium or the transmission medium 104. The source device 102 and the destination device 106 may include any number of different devices. Source device 102 and destination device 106 may include, for example, interconnected clusters of computer systems, servers, desktop computers, laptop computers, tablet computers, smartphones, wearable devices, televisions, cameras, video game consoles, set-top boxes, video streaming devices, vehicles (e.g., autonomous vehicles), or head-mounted displays that act as seamless resource pools (also known as computer clouds or cloud computing). Head-mounted displays may allow users to view VR, AR, or MR scenes and adjust the view of the scene, for example, based on the movement of the user's head. Head-mounted displays may be connected (e.g., tethered) to processing devices (e.g., servers, desktop computers, set-top boxes, or video game consoles) or may be completely independent.
[0040] Source device 102 may include point cloud source 112, encoder 114, and output interface 116. For example, to encode point cloud sequence 108 into bitstream 110, source device 102 may include point cloud source 112, encoder 114, and output interface 116. For example, point cloud source 112 may provide (e.g., generate) point cloud sequence 108 from captures of natural scenes and / or synthetically generated scenes. Synthetically generated scenes may be scenes containing computer-generated graphics. Point cloud source 112 may include one or more point cloud capture devices, a point cloud archive containing previously captured natural scenes and / or synthetically generated scenes, a point cloud feed interface for receiving captured natural scenes and / or synthetically generated scenes from a point cloud content provider, and / or a processor for generating synthetic point cloud scenes. Point cloud capture devices may include, for example, one or more laser scanning devices, structured light scanning devices, modulated light scanning devices, and / or passive scanning devices.
[0041] Point cloud sequence 108 may contain a series of point cloud frames 124 (e.g., Figure 1 (Example shown). A point cloud frame can describe an object or scene captured at a specific time instance. A point cloud sequence 108 can achieve the impression of motion by continuously presenting point cloud frames 124 of the point cloud sequence 108 using constant or variable time. A point cloud frame can contain a set of points (e.g., voxels) 126 in 3D space. Each point 126 can contain geometric information that can indicate the position of the point in 3D space. The geometric information can use three Cartesian coordinates (x, y, and z) to indicate, for example, the position of the point in 3D space. One or more points 126 can contain one or more types of attribute information. Attribute information can indicate the nature of the visual appearance of the point. For example, attribute information can indicate, for example, the texture of the point (e.g., color), the material type of the point, the transparency information of the point, the reflectivity information of the point, the surface normal of the point, the velocity at the point, the acceleration at the point, a timestamp indicating when the point was captured, a modality indicating how the point was captured (e.g., running, walking, or flying), etc. One or more points 126 can contain light field data, for example, in the form of multi-view related texture information. The light field data can be another type of optional attribute information. The color attribute information of one or more points 126 can include a lightness value and two chromaticity values. The lightness value can represent the brightness of the point (e.g., the lightness component Y). The chromaticity values can represent the blue and red components of the point, separate from the lightness (e.g., chromaticity components Cb and Cr). Other color attribute values can be represented, for example, based on different color schemes (e.g., RGB or monochrome color schemes).
[0042] Encoder 114 can encode point cloud sequence 108 into bitstream 110. To encode point cloud sequence 108, encoder 114 can use one or more lossless or lossy compression techniques to reduce redundant information in point cloud sequence 108. To encode point cloud sequence 108, encoder 114 can use one or more prediction techniques to reduce redundant information in point cloud sequence 108. Redundant information is information that can be predicted at decoder 120 and may not need to be sent (e.g., transmitted) to decoder 120 for accurate decoding of point cloud sequence 108. For example, the Movie Experts Group (MPEG) introduced the Geometry-Based Point Cloud Compression (G-PCC) standard (ISO / IEC Standard 23090-9: Geometry-Based Point Cloud Compression). G-PCC specifies the syntax and semantics of the encoded bitstream for transmission and / or storage of compressed point cloud frames, and the decoder operations for reconstructing compressed point cloud frames from the bitstream. During the standardization of G-PCC, reference software (ISO / IEC Standard 23090-21: Reference Software for G-PCC) was developed to encode the geometric and attribute information of point cloud frames. To encode the geometric information of point cloud frames, the G-PCC reference software encoder can perform voxelization. The G-PCC reference software encoder can perform voxelization, for example, by quantizing the positions of points in the point cloud. Quantizing the positions of points in the point cloud can create a mesh in 3D space. The G-PCC reference software encoder can map points to the center coordinates of the sub-mesh volume (e.g., voxel) where their quantized positions are located. The G-PCC reference software encoder can use occupancy trees to perform geometric analysis to compress the geometric information. The G-PCC reference software encoder can entropy encode the results of the geometric analysis to further compress the geometric information. To encode the attribute information of the point cloud, the G-PCC reference software encoder can use transformation tools such as Region Adaptive Hierarchical Transformation (RAHT), predictive transformation, and / or lifting transformation. Lifting transformation can be built on top of predictive transformation. Lifting transformation can include additional update / lifting steps. The lift transform and the predictive transform can be referred to as predictive / lift transform or predictive lift. Encoder 114 can operate in the same or similar manner as the encoder provided in the G-PCC reference software.
[0043] Output interface 116 can be configured to write and / or store bit stream 110 onto transmission medium 104. Bit stream 110 can be sent (e.g., transmitted) to destination device 106. Alternatively or concurrently, output interface 116 can be configured to send (e.g., transmit), upload, and / or stream bit stream 110 to destination device 106 via transmission medium 104. Output interface 116 may include wired and / or wireless transmitters configured to send (e.g., transmit), upload, and / or stream bit stream 110 according to one or more proprietary, open-source, and / or standardized communication protocols. One or more proprietary, open-source, and / or standardized communication protocols may include, for example, the Digital Video Broadcasting (DVB) standard, the Advanced Television Systems Committee (ATSC) standard, the Integrated Services Digital Broadcasting (ISDB) standard, the Cable Data Service Interface Specification (DOCSIS) standard, the 3rd Generation Partnership Project (3GPP) standard, the Institute of Electrical and Electronics Engineers (IEEE) standard, the Internet Protocol (IP) standard, the Wireless Application Protocol (WAP) standard, and / or any other communication protocol.
[0044] The transmission medium 104 may comprise wireless, wired, and / or computer-readable media. For example, the transmission medium 104 may comprise one or more wires, cables, air interfaces, optical discs, flash memory, and / or magnetic storage. Alternatively or additionally, the transmission medium 104 may comprise one or more networks (e.g., the Internet) or file servers configured to store and / or transmit (e.g., transfer) encoded video data.
[0045] Destination device 106 can decode bitstream 110 into point cloud sequence 108 for display or other forms of consumption. Destination device 106 may include one or more of input interface 118, decoder 120, and / or point cloud display 122. Input interface 118 may be configured to read bitstream 110 stored on transmission medium 104. Bitstream 110 may be stored on transmission medium 104 by source device 102. Alternatively, input interface 118 may be configured to receive, download, and / or stream bitstream 110 from source device 102 via transmission medium 104. Input interface 118 may include wired and / or wireless receivers configured to receive, download, and / or stream bitstream 110 according to one or more proprietary, open-source, standardized communication protocols and / or any other communication protocols. Examples of protocols include the Digital Video Broadcasting (DVB) standard, the Advanced Television Systems Committee (ATSC) standard, the Integrated Services Digital Broadcasting (ISDB) standard, the Cable Data Services Interface Specification (DOCSIS) standard, the 3rd Generation Partnership Project (3GPP) standard, the Institute of Electrical and Electronics Engineers (IEEE) standard, the Internet Protocol (IP) standard, and the Wireless Application Protocol (WAP) standard.
[0046] Decoder 120 can decode the point cloud sequence 108 from the encoded bit stream 110. For example, decoder 120 can operate in the same or similar manner as the decoder provided in the G-PCC reference software. Decoder 120 can decode a point cloud sequence that approximates the point cloud sequence 108. Decoder 120 can decode a point cloud sequence that approximates the point cloud sequence 108 due to, for example, lossy compression of the point cloud sequence 108 by encoder 114 and / or errors introduced into the encoded bit stream 110, for example, in the event of transmission to destination device 106.
[0047] The point cloud display 122 can display the point cloud sequence 108 to a user. The point cloud display 122 may include, for example, a cathode rate tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light-emitting diode (LED) display, a 3D display, a holographic display, a head-mounted display, or any other display device suitable for displaying the point cloud sequence 108.
[0048] Point cloud encoding (e.g., encoding / decoding) system 100 is presented by way of example and not limitation. Point cloud encoding systems different from and / or modified versions of point cloud encoding system 100 may perform the methods and processes described herein. For example, point cloud encoding system 100 may include other components and / or arrangements. Point cloud source 112 may be, for example, external to source device 102. Point cloud display device 122 may be, for example, external to destination device 106 or omitted entirely (e.g., if point cloud sequence 108 is intended to be consumed by a machine and / or storage device). Source device 102 may further include, for example, a point cloud decoder. Destination device 106 may include, for example, a point cloud encoder. For example, source device 102 may be configured to further receive an encoded bit stream from destination device 106. Receiving an encoded bit stream from destination device 106 can support bidirectional point cloud transfer between devices.
[0049] As described in this paper, an encoder can quantize the position of points in a point cloud with spatial precision, which can be the same or different in each dimension of the point. The quantization process can create a grid in 3D space. The encoder can map any point residing within each sub-grid volume to the coordinates of the sub-grid center, referred to as a voxel or volume pixel. A voxel can be viewed as a 3D extension of the pixels corresponding to the 2D image grid coordinates.
[0050] An encoder can represent or encode a voxelized point cloud. The encoder can, for example, use an occupied tree to represent or encode a voxelized point cloud. For example, the encoder can split an initial volume or cuboid containing a voxelized point cloud into sub-cuboids. The initial volume or cuboid can be referred to as a bounding box. The cuboid can be, for example, a rectangular prism. The encoder can recursively split each sub-cuboid containing at least one point of the point cloud. The encoder can choose not to further split sub-cuboids that do not contain at least one point of the point cloud. A sub-cuboid containing at least one point of the point cloud can be referred to as an occupied sub-cuboid. A sub-cuboid that does not contain at least one point of the point cloud can be referred to as an unoccupied sub-cuboid. The encoder can split an occupied sub-cuboid into, for example, two sub-cuboids (to form a binary tree), four sub-cuboids (to form a quadtree), or eight sub-cuboids (to form an octree). The encoder can split an occupied sub-cuboid to obtain additional sub-cuboids. Subcubes can have the same size and shape at a given depth level in the occupancy tree. For example, if the encoder splits the occupied subcube along a plane passing through the middle of the subcube's edge, the subcube can have the same size and shape at a given depth level in the occupancy tree.
[0051] An initial volume or cuboid containing a voxelized point cloud can correspond to the root node of the occupancy tree. Each occupied sub-cuboid split from the initial volume can correspond to a node (of the root node) in the second level of the occupancy tree. Each occupied sub-cuboid split from the occupied sub-cuboids in the second level can correspond to a node in the third level of the occupancy tree (outside the occupied sub-cuboids in the second level from which it splits). For each recursive splitting iteration, the occupancy tree structure can continue to form in this manner until, for example, a maximum depth level of the occupancy tree is reached or each occupied sub-cuboid has a volume corresponding to a voxel.
[0052] Each non-leaf node of the occupancy tree may contain an occupancy word or be associated with an occupancy word representing the occupancy status of the cuboid corresponding to the node. For example, a node in the occupancy tree corresponding to a cuboid split into eight sub-cubicles may contain a 1-byte occupancy word or be associated with a 1-byte occupancy word. Each bit of the 1-byte occupancy word (referred to as an occupancy bit) may represent or indicate the occupancy of a different sub-cubicle among the eight sub-cubicles. Occupied sub-cubicles may each be represented or indicated by a binary "1" in the 1-byte occupancy word. Unoccupied sub-cubicles may each be represented or indicated by a binary "0" in the 1-byte occupancy word. Occupied and unoccupied sub-cubicles may be represented or indicated by the opposite 1-bit binary value in the 1-byte occupancy word (e.g., a binary "0" representing or indicating an occupied sub-cubicle and a binary "1" representing or indicating an unoccupied sub-cubicle).
[0053] Each bit of the occupancy word can represent or indicate the occupancy of a different sub-cube among the eight sub-cubes. For example, the least significant bit of the occupancy word can represent or indicate the occupancy of the first sub-cube among the eight sub-cubes following the so-called Merton order. The second least significant bit of the occupancy word can represent or indicate the occupancy of the second sub-cube among the eight sub-cubes following the Merton order, and so on.
[0054] Figure 2 An example of Morton's order is shown. More specifically, Figure 2 The Morton order of the eight sub-cubes 202-216 split from cuboid 200 is shown. Sub-cubes 202-216 can be labeled, for example, based on their Morton order, where child node 202 is the first in the Morton order and child node 216 is the last. The Morton order of sub-cubes 202-216 can be a local lexicographical order in xyz.
[0055] The geometry of the voxelized point cloud can be represented by the initial volume and occupancy word of the nodes in the occupancy tree, and can be determined from the initial volume and the occupancy word. The encoder can send (e.g., transmit) the initial volume and occupancy word of the nodes in the occupancy tree to the decoder in a bitstream for reconstructing the point cloud. The encoder can entropy encode the occupancy word. The encoder can entropy encode the occupancy word, for example, before sending (e.g., transmitting) the initial volume and occupancy word of the nodes in the occupancy tree. The encoder can encode the occupancy bit of the occupancy word of the node corresponding to the cuboid. The encoder can encode the occupancy bit of the occupancy word of the node corresponding to the cuboid that is adjacent to or spatially close to the cuboid whose occupancy bit is being encoded, for example, based on one or more occupancy bits of the occupancy word of another node corresponding to the cuboid that is adjacent to or spatially close to the cuboid whose occupancy bit is being encoded.
[0056] The encoder and / or decoder can encode (e.g., encode / decode) the occupants of occupants in scan order. Scan order can also be referred to as scanning order. For example, the encoder and / or decoder can scan the occupant tree in breadth-first order. All occupants of nodes at a given depth (e.g., level) within the occupant tree can be scanned. All occupants of nodes at a given depth (e.g., level) within the occupant tree can be scanned, for example, before scanning the occupants of nodes at the next depth (e.g., level). Within a given depth, the encoder and / or decoder can scan the occupants of nodes in Morton order. Within a given node, the encoder and / or decoder can further scan the occupants of the node's occupants in Morton order.
[0057] Figure 3 An example scan order is shown. Figure 3 An example scan order (e.g., breadth-first order as described herein) is shown for an occupied tree of 300. More specifically, Figure 3 The scan order for the first three example levels of the occupies tree 300 is shown. Figure 3 In the diagram, the cuboid (e.g., cuboid 302) corresponding to the root node of the occupying tree 300 can be divided into eight sub-cubes (e.g., sub-cubes). Two of the eight sub-cubes, 304 and 306, may be occupied. The other six sub-cubes may be unoccupied. Following Merton order, the first eight occupying words (e.g., occW) are... 1,1 The first eight bits of the placeholder can be constructed to represent the root node. 1,1 Each occupant bit can represent or indicate the occupancy of a sub-cuboid in eight sub-cuboids in Morton order. For example, the first eight-bit occupant word occW 1,1 The least significant occupant can represent or indicate the occupancy of the first sub-cube in the eight sub-cubes in Morton order. The first eight-bit occupancy word is occW. 1,1 The second least effective occupant can indicate or indicate the occupancy of the second sub-cube in the eight sub-cubes in Morton order, and so on.
[0058] Each of the occupied subcubes (e.g., the two occupied subcubes 304 and 306) can correspond to a node other than the root node in the second level of the occupancy tree 300. Each of the occupied subcubes (e.g., the two occupied subcubes 304 and 306) can each be further subdivided into eight subcubes. For example, one of the eight subcubes subdivided from subcube 304, subcube 308, may be occupied, and the other seven subcubes may be unoccupied. Three of the eight subcubes subdivided from subcube 306, subcubes 310, 312, and 314, may be occupied, and the other five subcubes may be unoccupied. Two second octet occWs can be constructed in this order. 2,1 and occW 2,2 , to represent the occupancy word corresponding to the node of subcube 304 and the occupancy word corresponding to the node of subcube 306, respectively.
[0059] Each of the occupied subcubes (e.g., four occupied subcubes 308, 310, 312, and 314) can correspond to a node in the third level of the occupancy tree 300. Each of the occupied subcubes (e.g., four occupied subcubes 308, 310, 312, and 314) can be further subdivided into eight subcubes, or a total of 32 subcubes. For example, four third-level occupancy words (occW) can be constructed in this order. 3,1 occW 3,2 occW 3,3 and occW 3,4 , respectively representing the occupancy word corresponding to the node of sub-cube 308, the occupancy word corresponding to the node of sub-cube 310, the occupancy word corresponding to the node of sub-cube 312 and the occupancy word corresponding to the node of sub-cube 314.
[0060] The occupants of the example occupant tree 300 can be entropy encoded (e.g., entropy encoded by the encoder and / or entropy decoded by the decoder) following, for example, the scan order discussed herein (e.g., Morton order). The occupants of the example occupant tree 300 can be entropy encoded (e.g., entropy encoded by the encoder and / or entropy decoded by the decoder) into a sequence of seven occupants, for example, following the scan order discussed herein. 1,1 to occW 3,4 The scanning order discussed in this paper can be a breadth-first scanning order. For example, if the occupancy word of the current child node belonging to the current parent node is being entropy-encoded, then the occupancy words of all nodes with the same depth (e.g., level) as the current parent node may have already been entropy-encoded. For example, the occupancy words of all nodes with the same depth (e.g., level) as the current child node and with a lower Morton order than the current child node may also have already been entropy-encoded. A portion of the already encoded occupancy words can be used to entropy-encode the occupancy word of the current child node. The already encoded occupancy words of adjacent parent and child nodes can be used, for example, to entropy-encode the occupancy word of the current child node. For example, if a specific occupancy bit of the occupancy word of the current child node is being encoded (e.g., entropy-encoded), then the occupancy bits of occupancy words with a lower Morton order than that specific occupancy bit may have already been entropy-encoded and can be used to encode the occupancy bits of the occupancy word of the current child node.
[0061] Figure 4 An example neighborhood of a cuboid is shown for entropy encoding of the occupancy of a sub-cuboid. More specifically, Figure 4 An example neighborhood of a cuboid with already encoded occupant bits is shown. The neighborhood of a cuboid with already encoded occupant bits can be used for entropy encoding of the occupant bits of the current child cuboid 400. This can be based, for example, on the representation as discussed herein. Figure 4The scanning order of the occupancy tree of the cuboid geometry determines the neighborhood of the cuboid with the encoded occupant positions. The neighborhood of a cuboid, i.e., the neighborhood of the current child cuboid, can include one or more of the following: cuboids adjacent to the current child cuboid, cuboids sharing vertices with the current child cuboid, cuboids sharing edges with the current child cuboid, cuboids sharing faces with the current child cuboid, parent cuboids adjacent to the current child cuboid, parent cuboids sharing vertices with the current child cuboid, parent cuboids sharing edges with the current child cuboid, parent cuboids sharing faces with the current child cuboid, parent cuboids adjacent to the current parent cuboid, parent cuboids sharing vertices with the current parent cuboid, parent cuboids sharing edges with the current parent cuboid, parent cuboids sharing faces with the current parent cuboid, etc. Figure 4 As shown, the current child cuboid 400 can belong to the current parent cuboid 402. Following the scanning order of the occupancy words and occupancy bits of the occupancy tree nodes, the occupancy bits of the four child cuboids 404, 406, 408, and 410 belonging to the same current parent cuboid 402 may have already been encoded. The occupancy bits of the previous parent cuboid's child cuboid 412 may have already been encoded. The occupancy bits of the parent cuboid 414 may have already been encoded, while the occupancy bits of its child cuboids have not yet been encoded. The encoded occupancy bits of cuboids 404, 406, 408, 410, 412, and 414 can be used to encode the occupancy bits of the current child cuboid 400.
[0062] The number (e.g., quantity) of possible occupancy configurations (e.g., a set of one or more occupancy words and / or occupancy bits) in the neighborhood of the current sub-cuboid can be 2. N , where N is the number (e.g., quantity) of cuboids with encoded occupants in the neighborhood of the current child cuboid. The neighborhood of the current child cuboid can contain dozens of cuboids. The neighborhood of the current child cuboid (e.g., dozens of cuboids) can contain 26 neighboring parent cuboids that share faces, edges, and / or vertices with the parent cuboid of the current child cuboid, and several neighboring child cuboids that share faces, edges, and / or vertices with the current child cuboid and have encoded occupants. The occupancy configuration of the neighborhood of the current child cuboid can have billions of possible occupancy configurations, or even be limited to a subset of neighboring cuboids, making its direct use impractical. The encoder and / or decoder can use the occupancy configuration of the neighborhood of the current child cuboid to select a context (e.g., a probabilistic model) from the context set for a binary entropy encoder (e.g., a binary arithmetic encoder) that can encode the occupants of the current child cuboid. Context-based binary entropy coding can be similar to the context-adaptive binary arithmetic encoder (CABAC) used in MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC)).
[0063] The encoder and / or decoder can use several methods to reduce the occupancy configuration of the neighborhood of the current child cuboid being encoded to an actual number (e.g., quantity) of reduced occupancy configurations. This includes reducing the occupancy configuration to the actual number of six neighboring parent cuboids sharing a face with the current child cuboid. 6 Alternatively, 64 occupied configurations can be reduced to 9 occupied configurations. This reduction can be achieved by using geometric invariants. It can be reduced from 26 neighboring parent cuboids. 26 Each occupancy configuration yields the occupancy score of the current sub-cube. The score can be further reduced to a ternary occupancy prediction (e.g., "predicted occupancy", "uncertain", or "predicted unoccupancy") by using a score threshold. The individual occupancy of these sub-cubes can be replaced by the number of occupied neighboring sub-cubes (e.g., quantity) and the number of unoccupied neighboring sub-cubes (e.g., quantity).
[0064] Using / employing one or more of the methods described herein, the encoder and / or decoder can reduce the number (e.g., quantity) of possible occupancy configurations of the current subcuboid's neighborhood to a more manageable number (e.g., thousands). It has been observed that instead of directly associating the reduced number (e.g., quantity) of contexts (e.g., probabilistic models) with the reduced occupancy configuration, another mechanism, namely the Optimal Binary Encoder (OBUF) that supports on-the-fly updates, can be used. The encoder and / or decoder can implement the OBUF to limit the number (e.g., quantity) of contexts to a lower number (e.g., 32 contexts).
[0065] OBUF can use a finite number (e.g., 32) of contexts (e.g., probabilistic models). The number of contexts in OBUF (e.g., quantity) can be a fixed number (e.g., fixed quantity). Contexts used by OBUF can be sorted, indexed by context indices (e.g., context indices in the range of 0 to 31), and associated from the lowest virtual probability to the highest virtual probability to encode "1". A lookup table (LUT) for context indices can be initialized at the start of the point cloud encoding process. For example, the LUT can initially point to contexts with median virtual probabilities (e.g., context index 15) to encode "1" for all inputs. The LUT can initially point to contexts with median virtual probabilities to encode "1" for all inputs in a finite number (e.g., quantity) of contexts. This LUT can take the occupancy configuration of the neighborhood of the current subcube as input and output the context index associated with the occupancy configuration. The LUT can have as many entries as the reduced occupancy configuration (e.g., approximately several thousand entries). Encoding the occupancy bits of the current child cuboid can include the following steps: determining the reduced occupancy configuration of the current child node; obtaining a context index by using the reduced occupancy configuration as an entry in the LUT; encoding the occupancy bits of the current child cuboid using the context pointed to (e.g., indicated by) the context index; and updating the LUT entry corresponding to the reduced occupancy configuration, for example, based on the value of the encoded occupancy bit of the current child cuboid. For example, if a binary "0" (e.g., indicating that the current child cuboid is not occupied) is being encoded, the LUT entry can be reduced to a lower context index value. For example, if a binary "1" (e.g., indicating that the current child cuboid is occupied) is being encoded, the LUT entry can be increased to a higher context index value. The context index update process can, for example, be based on a theoretical model of the optimal distribution of virtual probabilities associated with a finite number (e.g., quantity) of contexts. This virtual probability can be fixed by the model and can differ from the internal probabilities of the context that may evolve, for example, when the encoding of data bits occurs. The evolution of the internal context can follow a well-known process similar to that in CABAC.
[0066] The encoder and / or decoder can implement a “dynamic OBUF” scheme. For example, compared to a general OBUF, a “dynamic OBUF” scheme allows the encoder and / or decoder to handle a much larger number (e.g., quantity) of occupancy configurations in the neighborhood of the current sub-cube. The use of a larger number (e.g., quantity) of occupancy configurations in the neighborhood of the current sub-cube can result in improved compression capabilities while keeping complexity within reasonable limits. By using an occupancy tree compressed by OBUF, the encoder and / or decoder can achieve lossless compression performance as good as 1 bit / point (bpp) for encoding the geometry of dense point clouds. The encoder and / or decoder can implement dynamic OBUF to potentially further reduce the bit rate by more than 25%, to 0.7 bpp.
[0067] OBUF may not take into account the various reduced occupancy configurations of the current child cuboid's neighborhood as input, and may potentially lead to a loss of useful relevance. With OBUF, the size of the LUT with the context index can be increased to handle more diverse occupancy configurations of the current child cuboid's neighborhood as input. Due to this increase, statistics can be diluted, and compression performance may deteriorate. For example, if the LUT has millions of entries and the point cloud has hundreds of thousands of points, most entries may never be accessed (e.g., lookup, access, etc.). Many entries may only be accessed a few times, and their associated context indexes may not be updated enough to reflect any meaningful correlation between the current child cuboid's occupancy configuration value and occupancy probability. Dynamic OBUF can be implemented to mitigate the dilution of statistics due to the increase in the number (e.g., quantity) of occupancy configurations of the current child cuboid's neighborhood. This mitigation can be performed through a "dynamic reduction" of the occupancy configuration in dynamic OBUF.
[0068] Dynamic OBUF can, for example, add an extra step to reduce the occupancy configuration of the current child cuboid's neighborhood before using a context-indexed LUT. This step can be called dynamic reduction because it evolves, for example, based on the progress of the point cloud encoding, or more precisely, based on the occupancy configuration that has already been visited (e.g., looked up in the LUT).
[0069] As discussed in this paper, there may be many possible occupancy configurations potentially involving the neighborhood of the current sub-cube, but if point cloud encoding occurs, only a subset can be accessed. This subset can characterize the type of point cloud. For example, most accessed occupancy configurations may represent the occupied neighboring cuboids of the current sub-cube, e.g., if an AR or VR dense point cloud is being encoded. On the other hand, for example, if a sparse point cloud acquired by a sensor is being encoded, most accessed occupancy configurations may only represent a few occupied neighboring cuboids of the current sub-cube. The effect of dynamic reduction can be, for example, to obtain a more accurate correlation by simultaneously shelving (e.g., actively reducing) other occupancy configurations accessed much less frequently based on the most accessed occupancy configuration. Dynamic reduction can be updated on an ad-hoc basis. For example, if occupancy data encoding occurs, dynamic reduction can be updated on an ad-hoc basis, e.g., after each access to an occupancy configuration (e.g., a lookup in a LUT).
[0070] Figure 5 An example of the dynamic decrease function DR that can be used in a dynamic OBUF is shown. The dynamic decrease function DR can be used by masking the bit β of configuration 500. j To obtain,
[0071] β = β1 … β K
[0072] The occupancy configuration consists of K bits. For example, if the occupancy configuration is accessed (e.g., looked up in the LUT) a certain number of times (e.g., a quantity), the mask size can be reduced. The initial dynamic reduction function DR... 0 It can mask all bits of all occupied configurations, such that for all occupied configurations β, it is a constant function DR. 0 (β) = 0. The dynamically decreasing function can be derived from the function DR. n Evolved to the update function DR n+1 The dynamic reduction function can, for example, be applied to the DR function after each encoding of the occupant bit. n Evolved to the update function DR n+1 A function can be defined as follows:
[0073] β' = DR n (β) = β1 … β kn(β)
[0074] Where k n (β) 510 is the number of unmasked bits (e.g., quantity). DR 0 The initialization can correspond to k0(β)=0, and the natural evolution of the decreasing function toward finer statistics can cause an increase in the number (e.g., quantity) of unmasked bits, k n (β) ≤ k n+1(β). The dynamic reduction function can be completely determined by all k occupying the configuration β. n The value is determined.
[0075] For all dynamically decreasing occupancy configurations β' = DR n Access to the occupied configuration (β), such as an instance lookup in a LUT, can be tracked by the variable NV(β'). For example, in a LUT based on the occupied configuration β... V After each instance where the placeholder is encoded, the corresponding access count (e.g., number) NV(β) V ') can be increased by one. If this number of visits (e.g., quantity) NV(β) V ') greater than the threshold th V ,
[0076] NV(β V ') > th V
[0077] Then for dynamically decreasing to β V 'All occupied configurations β, the number (e.g., quantity) of unmasked bits k n (β) can be increased by one. This corresponds to using two new dynamically decreasing occupancy configurations β. 0 'and β 1 'Replace the dynamically reduced occupancy configuration β' V The two new dynamically reduced occupancy configurations are defined as follows:
[0078] β 0 ' = β V '0 = β V 1 … β V kn(β) 0 and β 1 ' = β V '1 = β V 1 … β V kn(β) 1.
[0079] In other words, for all occupied configurations β, the number (e.g., quantity) of unmasked bits has increased by one, k n+1 (β) =k n (β) + 1, such that DR n (β) = β V The access count (e.g., quantity) of two new dynamically decreasing occupied configurations can be initialized to zero.
[0080] NV(β 0 ') = NV(β 1 ') = 0. (I)
[0081] At the start of encoding, the initial dynamic decrease function DR 0 The initial number of visits (e.g., quantity) can be set to
[0082] NV(DR 0 (β)) = NV(0) = 0,
[0083] Furthermore, the evolution of NV with dynamically decreasing occupancy configuration can be fully defined.
[0084] The corresponding LUT entry LUT[β] V '] can be derived from β V 'Two new entries LUT[β] initialized with the associated encoder index 0 '] and LUT[β] 1 Replacement. For example, if the dynamically decreasing occupancy configuration β V 'By two new dynamically reduced occupied configurations β 0 'and β 1 'Replace, then the corresponding LUT entry LUT[β] V '] can be derived from β V 'Two new entries LUT[β] initialized with the associated encoder index 0 '] and LUT[β] 1 ']replace,
[0085] LUT[β 0 '] = LUT[β 1 '] = LUT[β V '],(II)
[0086] And then it evolves independently. The evolution of the LUT with a dynamically decreasing occupancy configuration encoder index can be fully defined.
[0087] Decrease function DR n It can be formed by a series of growing binary trees T n Model 520, whose leaf node 530 is a reduced-occupancy configuration β' = DR n (β). The initial tree can be 0 = DR 0 (β) The associated single root node. This will be dynamically reduced to β. V Replace with β 0 'and β 1 'Can correspond to from β V 'Associated leaf node growth tree T' n For example, by using β 0 'and β 1 Two new nodes associated with each other are attached to the leaf node. Tree T n+1This can be obtained through such growth. The LUTs for the number of visits (e.g., quantity) NV and the context index can be defined on the leaf nodes and evolve as the tree grows via equations (I) and (II).
[0088] The practical implementation of dynamic OBUF can be achieved by storing an array NV[β'] and a LUT[β'] of context indexes, as well as a tree T. n 520 is used for this. An alternative to storing the tree could be an array k storing the number (e.g., quantity) of the unmasked bits. n [β]510.
[0089] One limitation of implementing dynamic OBUF is its memory footprint. In some applications, millions of occupied configurations can be handled, resulting in approximately 20-bit beta. i The entries β constitute the configuration for the decrease function DR. Each bit β i It can correspond to the occupancy state of the adjacent cuboids of the current sub-cube or the set of adjacent cuboids of the current sub-cube.
[0090] The higher (e.g., higher effective) bit β i (For example, β0, β1, etc.) can be the first bit without masking. Higher (e.g., higher-active) bits β i (For example, β0, β1, etc.) could be, for example, the first unmasked bit during the evolution of the dynamically decreasing function DR. Place bit β... i The order of neighbor-based information in compression can affect compression performance. Neighbor information can be ordered from highest (e.g., highest) priority to lowest priority, and then placed into bit β in this order from highest weight to lowest weight. i In this context, priority is ranked from most important to least important: occupancy of adjacent child cuboids, followed by adjacent child cuboids, then adjacent parent cuboids, then non-adjacent child nodes, and finally non-adjacent parent nodes. A neighboring node sharing a face with the current child node can have a higher priority than a neighboring node sharing an edge (but not a face) with the current child node. Similarly, a neighboring node sharing an edge with the current child node can have a higher priority than a neighboring node sharing only a vertex with the current child node.
[0091] Figure 6 An example method for encoding the occupancy of a cuboid using dynamic OBUF is shown. More specifically, Figure 6 An example method for encoding the occupants of the current child cuboid using dynamic OBUF is shown. Figure 6 One or more steps can be performed by an encoder and / or decoder (e.g., Figure 1The encoder 114 and / or decoder 120 in the flowchart are executed. All or part of the flowchart can be executed by the encoder (e.g., ...). Figure 1 The example computer system 2000 in FIG20 and / or the example computing device 2130 in FIG21 are implemented by encoder 114 and / or decoder 120.
[0092] At step 602, the occupancy configuration of the current sub-cuboid (e.g., occupancy configuration β) can be determined. The occupancy configuration of the current sub-cuboid (e.g., occupancy configuration β) can be determined, for example, based on the occupancy bits of already encoded cuboids in the neighborhood of the current sub-cuboid. At step 604, the occupancy configuration (e.g., occupancy configuration β) can be dynamically reduced. For example, a dynamic reduction function DR can be used. n This allows for dynamic reduction of the occupancy configuration. For example, the occupancy configuration β can be dynamically reduced to a reduced occupancy configuration β' = DR. n (β). At step 606, the context index can be looked up in a lookup table (LUT), for example. For example, the encoder and / or decoder can look up the context index LUT[β'] in the LUT of the dynamic OBUF. At step 608, a context (e.g., a probabilistic model) can be selected. For example, the context pointed to by the context index (e.g., a probabilistic model) can be selected. At step 610, the occupancy of the current child cuboid can be entropy encoded. For example, the occupancy of the current child cuboid can be entropy encoded (e.g., arithmetic encoding) based on the context. The occupancy of the current child cuboid can be encoded based on the occupancy of the already encoded cuboids adjacent to the current child cuboid.
[0093] although Figure 6 Not shown, but the encoder and / or decoder can update the reduction function and / or update the context index. For example, the encoder and / or decoder can update the reduction function DR. n Updated to DR n+1 And / or, for example, update the context index LUT[β'] based on the current child cuboid's occupancy. Figure 6 The method can be based on the scanning order, as discussed in this article. Figure 3 The scanning order discussed repeats for the additional or all child cuboids of the parent cuboid corresponding to the node occupying the tree.
[0094] Occupation tree is typically a lossless compression technique. Occupation tree can be adapted to provide lossy compression, for example, by modifying the point cloud on the encoder side (e.g., downsampling, removing points, moving points, etc.). Lossy compression performance can be weak. For dense point clouds, lossy compression can be a useful lossless compression technique.
[0095] One approach to lossy compression of point cloud geometry could be to set the maximum depth of the occupancy tree to stop at a larger volume size (e.g., an NxNxN cuboid, where N > 1) instead of reaching the minimum volume size of a voxel. The geometry of points belonging to each occupied leaf node associated with this larger volume can then be modeled. This approach may be particularly well-suited for dense and smooth point clouds that can be locally modeled using smoothing functions such as planes or polynomials. The encoding cost can be reduced to the cost of the occupancy tree plus the cost of the local model within each occupied leaf node.
[0096] A scheme for modeling the geometry of points belonging to each occupied leaf node associated with a volume larger than one voxel can use a set of triangles as a local model. This scheme can be called a "TriSoup" scheme. TriSoup is short for "triangle soup" because the connections between triangles may not be part of the model. An occupied leaf node corresponding to a cuboid with a volume larger than one voxel in the tree can be called a TriSoup node. An edge belonging to at least one cuboid corresponding to a TriSoup node can be called a TriSoup edge. A TriSoup node can contain an existence flag (s) for each TriSoup edge of its corresponding occupied cuboid. k The existence flag of a TriSoup edge (s) k ) can indicate the TriSoup vertex (V k Does a vertex exist on a TriSoup edge? At most one TriSoup vertex (V) exists. k A vertex (V) can exist on a TriSoup edge. For each vertex (V) existing on a TriSoup edge of an occupied cuboid... k The TriSoup node corresponding to the occupied cuboid can further contain vertices (V). k ) along the position of TriSoup edge (p k ).
[0097] In addition to the occupation word of the occupation tree, the encoder can also entropy encode the TriSoup vertex presence flag and position for each TriSoup edge belonging to the TriSoup node of the occupation tree. Similarly, in addition to the occupation word of the occupation tree, the decoder can also entropy decode the TriSoup vertex presence flag and position for each TriSoup edge and the vertices along the corresponding TriSoup edges belonging to the TriSoup nodes of the occupation tree.
[0098] Figure 7 An example of an occupied cuboid (e.g., a rectangular prism) 700 is shown. More specifically, Figure 7An example of an occupied cuboid (e.g., a rectangular prism) 700 of size NxNxN (where N > 1) corresponding to a TriSoup node in the occupied tree is shown. The occupied cuboid 700 may contain edges (e.g., TriSoup edges 710-721). The TriSoup node corresponding to the occupied cuboid 700 may contain an existence flag (s) for each edge (e.g., each TriSoup edge in TriSoup edges 710-721). k For example, the presence flag of TriSoup edge 714 can indicate that TriSoup vertex V1 exists on TriSoup edge 714. The presence flag of TriSoup edge 715 can indicate that TriSoup vertex V2 exists on TriSoup edge 715. The presence flag of TriSoup edge 716 can indicate that TriSoup vertex V3 exists on TriSoup edge 716. The presence flag of TriSoup edge 717 can indicate that TriSoup vertex V4 exists on TriSoup edge 717. The presence flags of the remaining TriSoup edges can each indicate that no TriSoup vertex exists on its corresponding TriSoup edge. A TriSoup node corresponding to the occupied cuboid 700 can contain the position of each TriSoup vertex that exists along one of its TriSoup edges 710-721. More specifically, a TriSoup node corresponding to the occupied cuboid 700 can contain the position p1 of TriSoup vertex V1, the position p2 of TriSoup vertex V2, the position p3 of TriSoup vertex V3, and the position p4 of TriSoup vertex V4. TriSoup vertices can be shared among TriSoup nodes along common TriSoup edges.
[0099] The existence of the current TriSoup edge can be flagged (s) k ) and (in the presence of signs (s k (The location (p) can indicate the existence of a vertex) k Entropy encoding is performed. Existence flag (s) k ) and location (p k The information can be referred to individually or collectively as vertex information or TriSoup vertex information. For example, the existence flag (s) of the current TriSoup edge can be determined based on the encoded existence flags and positions of the existing TriSoup vertices of the TriSoup edges adjacent to the current TriSoup edge. k ) and (in the presence of signs (s k (indicating the presence of a vertex) position (p) k Entropy encoding is performed on the current TriSoup edge. Alternatively, the presence flag (s) of the current TriSoup edge can be used as an alternative. k ) and (in the presence of signs (sk (The location (p) can indicate the existence of a vertex) k Entropy encoding is performed on the edges (e.g., indicating the position of the vertices along which the edges are located). The presence flag of the current TriSoup edge (s) k ) and location (p k Entropy encoding can be performed alternatively, for example, based on the occupancy of cuboids adjacent to the current TriSoup edge. Similar to entropy encoding of occupancy bits in an occupancy tree, the configuration β of the neighborhood of the current TriSoup edge can be obtained. TS (Also known as neighborhood configuration β) TS ), and for example, by using TriSoup's dynamic OBUF scheme to dynamically reduce it to a reduced configuration β. TS ' = DR n (β TS ). Context index LUT[β TS The information can be obtained from the OBUF LUT. At least a portion of the vertex information for the current TriSoup edge can be entropy encoded using the context pointed to by the context index (e.g., a probabilistic model).
[0100] The position of the TriSoup vertex along its TriSoup edge (p k (If it exists) can be binarized. The position of the TriSoup vertex along its TriSoup edge (p k (If present) can be binarized, for example, by using a binary entropy encoder to entropy encode at least a portion of the vertex information of the current TriSoup edge. The number of bits (e.g., quantity) N can be set. b To quantize the TriSoup vertex positions (p) along a TriSoup edge of length N. k A TriSoup edge of length N can be uniformly divided into 2. Nb Quantization interval. By doing so, the TriSoup vertex position (p k ) can be individually encoded by an N that can be generated using a dynamic OBUF scheme. b p k j , j=1, …, N b ) and corresponding to the existence flag (s) k The bit representation of ). Neighborhood configuration β TS、 OBUF decrease function DR n The context index can depend on encoded bits (e.g., presence flags). k ), highest position (p) k1 ), second highest position (p) k2Encoded bits (e.g., presence flags) k ), highest position (p) k 1 ), second highest position (p) k 2 Properties, characteristics, and / or attributes of vertex information. In reality, there may be several dynamic OBUF schemes, each dedicated to specific bits of vertex information (e.g., presence flags). k ) or position (p) k j )).
[0101] Figure 8(a) shows an example cuboid (e.g., rectangular prism) 800 corresponding to a TriSoup node. Cuboid 800 can correspond to a TriSoup node with a number of K vertices V. k The TriSoup node. Within the cuboid 800, the TriSoup triangle can be formed by the TriSoup vertex V. k Construction. For example, if there are at least three (K≥3) TriSoup vertices on the TriSoup edges of a cuboid 800, then a TriSoup triangle can be constructed from TriSoup vertices V. k Construction. For example, regarding Figure 8(a), there can be four TriSoup vertices, and a TriSoup triangle can be constructed. The TriSoup triangle can be constructed around the centroid vertex C, which is defined as the TriSoup vertex V. k The mean of the values. The primary direction can be determined, and then the vertex V can be adjusted by rotating around this direction. k Sort the data and construct the following K TriSoup triangles: V1V2C, V2V3C, ..., V K V1C. For example, if a triangle is projected along a principal direction, the principal direction can be selected from three directions that are respectively parallel to the axes of 3D space to increase or maximize the 2D surface of the triangle. By doing so, the principal direction can be slightly perpendicular to the local surface defined by the points of the point cloud belonging to the TriSoup node.
[0102] Figure 8B An example refinement of the TriSoup model is shown. The TriSoup model can be derived by adjusting the centroid residual value C. res Encoding is used to refine the value. Centroid residual C res It can be encoded into a bitstream. Centroid residual value C res It can be encoded into a bitstream, for example, to use C+C++. res Instead of using C as the pivot vertex of the triangle, it is done by using C+C. res As the pivot vertex of the triangle, vertex C+Cres Points can be located closer to the point cloud than the centroid C, which can reduce reconstruction error and thus reduce distortion. However, the cost is that C is encoded more efficiently. res The required bit rate will increase slightly.
[0103] Figure 9 An example of voxelization is shown. Voxelization can refer to reconstructing a decoded point cloud from a set of TriSoup triangles. Voxelization can be performed by ray tracing each triangle individually. For example, voxelization can be performed by ray tracing each triangle individually before removing duplicate points between voxelized triangles. Figure 9 As shown, ray 900 can be emitted parallel to one of the three axes in 3D space. Ray 900 can originate from integer coordinates P. 开始 The emission begins at 905 (e.g., the origin). The intersection point P of ray 900 and the TriSoup triangle 901 belonging to the cuboid (e.g., cuboid) 902 corresponding to the TriSoup node can be rounded. int 904 (if it exists) to obtain the decoded point. This intersection point P can be found, for example, using the Möller-Trumbore algorithm. int .
[0104] TriSoup vertices of TriSoup nodes may need to be quantized to certain acceptable vertex positions to ensure continuity in triangle-based modeling between TriSoup nodes. Therefore, TriSoup modeling approximating occupied voxels within a TriSoup node may not match occupied voxels determined to be located within TriSoup triangles with quantized TriSoup vertices. For example, if voxelization of TriSoup triangles occurs, some voxels may be missed. As described in this paper, techniques including the “halo” method and the “fine ray emission” method have been introduced to enhance the voxelization process, thereby attempting to improve voxel reconnection between triangles.
[0105] Both the halo method and the fine ray emission method attempt to increase the possible intersections between the emitted rays and the triangle in order to “recapture” missed voxels resulting from quantizing the vertices of the TriSoup triangle. While the halo method does not significantly increase complexity and processing cost, the fine ray emission method can significantly increase processing cost because it additionally emits multiple rays at non-integer coordinates (called fine rays) for each ray at integer coordinates to increase the possible intersections between the rays and the TriSoup triangle. Implementing the fine ray emission method not only significantly increases processing cost, but it also overlaps with the halo method and may recapture some of the same missed voxels, which may reduce its effectiveness. Examples of this disclosure include enhancing the voxelization process by adding one or more additional points near each defined point (e.g., intersection) of the TriSoup triangle and quantizing or voxelizing said one or more additional points, as well as the defined points of the TriSoup triangle. The one or more additional points may extend from the defined points, for example, in a direction not aligned with the plane of the TriSoup triangle. The computational intensity of this quantization process is lower than that of the fine ray emission method, and the quantization of these one or more additional points can recapture voxels that may not have been recognized by the halo method, and thus can be implemented using the halo method to enhance voxelization.
[0106] Figure 10 An example of the coordinates of a point relative to the centroid of the TriSoup triangle is shown. More specifically, Figure 10 An example is shown of the barycentric coordinates (u, v, w) of point 1002 (e.g., P) relative to a TriSoup triangle 1000 in 3D space with vertices labeled A, B, and C. For example, point 1002 can be identified as the intersection of a ray with a plane of the TriSoup triangle 1000 (e.g., containing or passing through the three vertices A, B, and C of the TriSoup triangle 1000). For example, a ray can be emitted parallel to one of the three coordinate axes in 3D space. This intersection point 1002 can be uniquely represented as the sum of the three vertices of the TriSoup triangle 1000:
[0107] P = uA + vB + wC
[0108] Under the condition u + v + w = 1, any point P in the plane (containing TriSoup triangle 1000) can have unique coordinates (u, v, w) in the barycentric coordinate system. Points with barycentric coordinates (u, v, w) can consist of ordered triples of u, v, and w. Points whose sum of barycentric coordinates (u, v, w) is 1 (i.e., u + v + w = 1) can be called homogeneous barycentric coordinates or normalized barycentric coordinates. The barycentric coordinates of the intersection point relative to TriSoup triangle 1000 can be determined using, for example, the well-known Möller-Trumbore algorithm.
[0109] The three vertices A, B, and C of TriSoup triangle 1000 can be represented by their corresponding barycentric coordinates A (1,0,0), B (0,1,0), and C (0,0,1) by converting points in 3D space with Cartesian coordinates to homogeneous barycentric coordinates. The convex hull of the three vertices A, B, and C (i.e., TriSoup triangle 1000) is equal to the set of all points P such that the barycentric coordinates u, v, and w are each greater than or equal to zero.
[0110] 0 ≤ u, v, w
[0111] An intersection point can be determined to belong to TriSoup triangle 1000, for example, based on the fact that its ordered triplet values are each greater than or equal to zero. If at least one of the centroid coordinates (e.g., one of u, v, or w) is negative or less than 0, then the intersection point can be determined not to belong to TriSoup triangle 1000, because the intersection point can be on a plane, but not on an edge or inside TriSoup triangle 1000. A point determined to belong to TriSoup triangle 1000 can be a ray that intersects TriSoup triangle 1000 (e.g., inside or at an edge of TriSoup triangle 1000).
[0112] In the Möller-Trumbore algorithm, the intersection point of a ray and the plane containing the TriSoup triangle can be determined, for example, based on calculating the barycentric coordinates u, v, and w for the intersection point. The intersection point can be determined, for example, by verifying that each of the barycentric coordinates u, v, and w is greater than or equal to 0 (e.g., 0 ≤ u, v, w) is within the TriSoup triangle (e.g., on its side or inside). The intersection point can also be determined, for example, to be outside the TriSoup triangle.
[0113] Figure 11 An example of encoding centroid residual values is shown. More specifically, Figure 11The centroid residual value C in / from the potential stream is shown. res A more detailed example of coding is provided, enabling the use of adjusted centroid C+C++. res Instead of the centroid C, generate TriSoup triangles for cuboids 1100 (corresponding to TriSoup nodes) that represent parts of the point cloud. This can be, for example, based on an adjusted centroid C+C. res The TriSoup triangles are generated by selecting neighboring vertex pairs from vertices V1-V4, as determined herein with respect to Figure 8(a). As described herein, the TriSoup triangles of cuboid 1100 can be voxelized at the decoder, for example, to generate voxels representing (or modeling) the portion of the point cloud corresponding to cuboid 1100. Unit vector (For example, also known as a normalized vector) can be determined as triangles (V1V2C, V2V3C, ..., V1V2C) constructed by pivoting around the centroid C and the vertex pairs of the cuboid 1100. K The normalized mean vector of the normal vector of V1C (e.g., as described herein with respect to Figure 8(a)). Unit vector This can be based, for example, on the mean of the cross product representing the area of a triangle. And thus determined as a normalized vector. For example, a unit vector. This can be determined by dividing the mean vector (n) by the norm (or length) of the mean vector (e.g., ).
[0114] The value produced by each cross product is equal to the area of the parallelogram formed by the two vectors in the cross product. This value can also represent the area of the triangle formed by the two vectors, for example, since the area of the triangle is equal to half this value. Vector It can indicate the direction orthogonal to the local surface representing the part of the point cloud, for example, because the vector This can indicate the orientation of triangles (e.g., TriSoup triangles) representing portions of a point cloud. It can also indicate the orientation along the line (C, The single-component residual α of 1110 res Instead of encoding 3D residuals, for example, to maximize the effect of centroid residuals and / or minimize their encoding cost.
[0115]
[0116] residual value α res For example, the encoder can determine the current point cloud and line (C, The intersection points between the normalized vectors can be along the normalized vector. In the same direction. For example, the set of points in a portion of the point cloud that are closest to the line (e.g., within a threshold distance, threshold number of points / count) can be determined by the encoder. The set of points can be projected onto the line, and the residual value α res The mean component can be determined by the encoder as the mean of the line along the projection points. This mean can be determined, for example, as a weighted mean, where the weights can depend on the distance of the set of points from the line. For example, points in the set closer to the line can have a higher weight than another point in the set farther from the line.
[0117] The residual value α can be quantified. res For example, the residual value α res It can be formed by having a vertex V similar to TriSoup. k Quantization is performed using a uniform quantization function with quantization step size and quantization precision similar to TriSoup. k The quantization precision, quantization step size, uniform quantization function, and quantization residual value α res It can maintain the quantization error at all vertices V k and C+C res The surface is uniform, which allows for a uniform approximation of the local surface.
[0118] residual value α res It can be binarized and / or entropy encoded into the bitstream, for example, by using a unary-based coding scheme. The residual value α res Encoding can be performed using a set of flags, for example. For instance, the flag f0 can be encoded to indicate the residual value α. res Is it equal to zero? If the flag f0 indicates the residual value α res If the value is zero, then no additional syntax element may be needed. If the flag f0 indicates the residual value α... res If the value is not zero, then the sign bit of the indicator can be encoded, and / or entropy coding can be used to encode the residual value |α. res |-1 is used for encoding. For example, a unary coding scheme can be used to encode the residual value, which can encode the indicator residual value |α. res | Whether it is equal to the consecutive flag f of 'i' i (i≥1) are encoded. The encoder (e.g., a binary entropy encoder) can encode the residual value α. res Binary formation marker f i (i≥0), and the binary residual value and the sign bit are encoded (e.g., entropy coding).
[0119] residual value α res Compression can be achieved, for example, by determining such... Figure 11 Improvements are made based on the boundaries shown in this article. Figure 11As described, the line (C, ) 1110 can intersect the current cuboid 1100 (corresponding to the TriSoup node) at two boundary points 1120 and 1121, and the encoder can enforce that the adjusted centroid vertex C + C res can be located between the two boundary points 1120 and 1121. These boundary points 1120 and 1121 can limit the residual value α res (which can be quantized), ensuring that it belongs to the integral interval [m, M], where m ≤ 0 ≤ M. By doing so, some bits of the binary residual value α res can be inferred. For example, if m = M = 0, the residual value α res must be equal to zero. For example, if m = 0 < M, the sign bit must be positive. If the residual value α res is not equal to zero and its sign is known, its magnitude |α res | can be determined to be bounded by |m| or M, such that the magnitude can be encoded by a truncated unary coding scheme, which can infer the value of the last flag in the sequence of consecutive flags f i (i ≥ 1).
[0120] The binary entropy encoder that can be used to encode the binary residual value α res [[ID=
[0122] Figure 12A An example of the centroid vertex determined from the TriSoup vertices of a cuboid is shown. More specifically, Figure 12A An example of a centroid vertex 1240 determined from TriSoup vertices 1220, 1221, 1222, and 1223 of a cuboid 1200 indicated by TriSoup nodes corresponding to a portion of a point cloud is shown. For example, the local spatial distribution of points in the portion of the point cloud corresponding to cuboid 1200 can be represented (or modeled) by a local surface 1210 (e.g., plane 1210), which can, for example, intersect cuboid 1200 at four points located on the four corresponding edges of cuboid 1200 associated with the TriSoup nodes. These four points can be quantized as corresponding TriSoup vertices 1220, 1221, 1222, and 1223. For example, other methods can be used to model the local distribution of points as a set of TriSoup vertices on the edges of cuboid 1200 (e.g., three or more edges, or four or more edges, etc.). For example, instead of modeling points as local surfaces (e.g., as... Figure 12A The plane 1210 shown can be used to average the positions of points in the point cloud portion near each edge (e.g., within a distance or threshold) to determine whether to encode the TriSoup vertices on the edge, and if so, to encode the positions of the TriSoup vertices on the edge.
[0123] The centroid vertex 1240 of cuboid 1200 can be determined as a simple average of TriSoup vertices 1220, 1221, 1222, and 1223. For example, each coordinate of centroid vertex 1240 can be determined as the average of the corresponding coordinates of TriSoup vertices 1220, 1221, 1222, and 1223. Small changes in the local spatial distribution of points in the point cloud may result in small changes to the local surface modeled by local surface 1210, but will lead to changes in the number / number and / or position of TriSoup vertices 1220, 1221, 1222, and 1223, which are the intersections between local surface 1210 and the edges of cuboid 1200 associated with TriSoup nodes.
[0124] Figure 12B An example of the centroid vertex determined from the TriSoup vertices of a cuboid is shown. For example, as discussed in this paper... Figure 12B The centroid vertex 1250 can be determined as a result of a small change in the local spatial distribution of points in the point cloud, which causes the local surface (e.g., illustratively representing the local spatial distribution) to change from the local surface 1210 (e.g. Figure 12A As shown) becomes local surface 1211 (as shown) Figure 12B (As shown).
[0125] As described in this article, TriSoup vertices 1220, 1221, 1222, and 1223 can be directly determined from partial points in the point cloud. (See also: Regarding...) Figure 12A and Figure 12B As described, a small change in the distribution of these points may cause the four TriSoup vertices 1220, 1221, 1222, and 1223 to become five TriSoup vertices 1220, 1221, 1222, 1232, and 1234. A small change in the point distribution may cause discontinuities in the number / scale of TriSoup vertices. The centroid vertex can be determined as a simple average of the TriSoup vertices. The centroid vertex can change discontinuously (e.g., from...). Figure 12A The centroid vertex 1240 in the middle becomes Figure 12B (Centroid vertex 1250 in the image). For example, if inter-frame prediction is used between two frames of a dynamic point cloud, the discontinuity in determining the centroid vertex from the points in the point cloud can lead to inefficient predictions between the two point clouds. This discontinuity can result in suboptimal compression performance.
[0126] Figure 13 An example of a TriSoup triangle within a cuboid is shown. More specifically, Figure 13 Examples of TriSoup triangles 1330, 1331, 1332, 1333, and 1334, formed by TriSoup vertices 1310, 1311, 1312, 1313, and 1314 and centroid vertex 1320, are shown in cuboid 1300 corresponding to TriSoup nodes. The centroid vertex 1320 may not be located "in the middle" of a portion of the local surface representing the part of the point cloud covered by cuboid 1300. Two of the five TriSoup vertices 1310, 1311, 1312, 1313, and 1314, namely TriSoup vertices 1313 and 1314, may be located on the edges of cuboid 1300. The distance between two TriSoup vertices 1313 and 1314 may be closer to each other than to the other vertices. The simple average of the TriSoup vertices 1310, 1311, 1312, 1313, and 1314 results in centroid vertex 1320, which is positioned closer to the two vertices 1313 and 1314 than to the other three vertices 1310, 1311, and 1312. TriSoup triangles 1330, 1331, 1332, 1333, and 1334, which can be formed by pairs of neighboring vertices and centroid vertex 1320, can have significantly different sizes. For example, as regarding... Figure 13 As mentioned, TriSoup triangle 1333 is much smaller than TriSoup triangle 1331.
[0127] Differences in the size of TriSoup triangles can lead to suboptimal local interpolation of the point cloud by the TriSoup model's triangles and / or cause greater distortion and / or suboptimal compression performance. It may be necessary to determine the centroid vertex continuously, for example, based on the local distribution of points within the point cloud. Additionally, it may be necessary to adjust the position of the centroid vertex to be more equidistant from the TriSoup vertices of the cuboid.
[0128] As described in this paper, the centroid vertex of a cuboid can be determined based on a weighted sum of its TriSoup vertices. The TriSoup vertices of the cuboid can correspond to portions of the point cloud. Each vertex in the TriSoup can have a weight. The weight of each vertex can be determined, for example, based on its neighboring vertices. Triangles (e.g., TriSoup triangles) can be determined, for example, based on the centroid vertex and pairs of vertices. Triangles can be voxelized to determine voxels representing portions of the point cloud (e.g., as described in this paper regarding...). Figure 9 and / or Figure 10 (as described).
[0129] For each vertex, the area of the triangle formed by the centroid vertex, the vertex itself, and each of the vertex's adjacent vertices can be determined. The weight of the corresponding vertex can be determined, for example, based on the area of the triangle. Alternatively, the weight of each vertex can be determined based on the distance between the vertex and each of its adjacent vertices. For example, the distance can be based on the L1 norm (e.g., also known as the Manhattan norm) or the L2 norm (e.g., also known as the Euclidean norm). For example, the weight can be determined to be proportional to the sum of the areas of the triangles, such as being the same as or half the sum of the areas of the triangles. Alternatively, the weight can be determined to be proportional to the sum of the distances between the vertex and each of its adjacent vertices, such as being the same as or half the sum of the distances between the vertex and each of its adjacent vertices.
[0130] Figure 14A An example of determining the centroid vertex of a cuboid is shown. More specifically, Figure 14A An example is shown where the centroid vertex of a cuboid is determined as, for example, a weighted sum of the TriSoup vertices of the cuboid corresponding to a TriSoup node. (See the section on...) Figure 12A , Figure 12B and Figure 13 The cuboid can correspond to (e.g., cover) a portion of the point cloud. For example... Figure 14A As shown in two-dimensional form, the centroid vertex C can be determined, for example, based on a weighted sum (or weighted average):
[0131]
[0132] In the TriSoup vertex V, which can belong to the cuboid whose centroid vertex will be targeted. i Above. The centroid vertex C can be determined as a weighted average (or weighted mean), in which case the weights w i The sum of all values equals one. This method ensures that the centroid vertex C belongs to (e.g., is contained within) the cuboid associated with the TriSoup node.
[0133]
[0134] weight w i The sum is one. For example, if the centroid vertex C is determined as the weighted average, then the weight w i Each of them can be between 0 and 1.
[0135] For example, if the centroid vertex C is determined based on a weighted sum (and the weights w) i (The sum may not always be one), then the centroid vertex C can be determined as the sum of the weighted sum divided by the weights. This formula is equivalent to dividing the weights w by the weights w. i Each weight in the equation is defined as weight w. i The sum of the fractions. As described in this article, it may be necessary to determine or approximate the weights w. i For example, to ensure continuity and / or to cause the position of the centroid vertex to be aligned with the TriSoup vertex V of the cuboid. i Equidistant (or nearly equidistant).
[0136] Figure 14B An example of determining the weights of TriSoup vertices is shown. More specifically, Figure 14B A triangle (e.g.) is shown and The area of (e.g.) and ) can be used to determine TriSoup vertices (e.g., The weights of (e.g.) Centroid vertex (e.g., centroid vertex C, such as...) Figure 14A (As shown) can be based, for example, on TriSoup vertices (e.g., The weights of (e.g.) To determine.
[0137] As this article is about Figure 14B As stated above, it is known from physics that points exist that satisfy these conditions: points with vertices C and V. i and V i+1 triangle Surfaces formed by composites (or combinations or unions) The center of gravity G. If the surface If a particle has a uniform density ρ, then its centroid G is defined as follows:
[0138]
[0139] Where M is the surface Let dm be a point on the surface at point M, and ds be a local infinitesimal mass at point M. By introducing an arbitrary origin O, the barycenter G can be alternatively defined as follows:
[0140]
[0141] Then
[0142]
[0143] Where S is... The defined total area. (This is achieved by considering the surface area.) Divided into sections with corresponding areas S i+1 / 2 triangle As in this article about Figure 14B As stated, the following will be obtained:
[0144]
[0145] in It is a triangle known only by the following definition Center of gravity:
[0146]
[0147] Removing the origin O from the notation and directly using the point coordinates, the centroid G satisfies the following equation:
[0148]
[0149] Then:
[0150]
[0151] For example, to determine whether the centroid vertex C is close to the barycenter G, we can use G=C to simplify the equation as described in this article to obtain:
[0152]
[0153] Furthermore, the centroid vertex C is effectively obtained as a weighted sum of the TriSoup vertices, and the weights satisfy the following equation:
[0154]
[0155] The sum of the weights equals one (e.g., 1).
[0156]
[0157] The centroid vertex C can be determined, for example, based on the weighted average of the TriSoup vertices of the cuboid belonging to centroid vertex C. Vertices (e.g., V...) i Each weight of ) (e.g., ) can be determined, for example, by including vertex V i The TriSoup triangle formed by the triples and area The average value. The centroid vertex C can be determined, for example, as the barycenter, as described in this paper. Causal problems may arise, for example, because it may be necessary to know the centroid vertex C to determine the TriSoup triangle (e.g., and ), its area (e.g., and ) and weight The centroid vertex C is calculated as a weighted sum. Centroid vertex C is located at both the beginning and end of the causal chain. Vertex V i Each weight (e.g., It can be based, for example, on a combination of vertices V. i The TriSoup triangle formed by the triples (e.g., and The area of (e.g.) and The representative value is used to determine this.
[0158] The initial centroid vertex C0 can be determined from the TriSoup vertices that can form a TriSoup triangle, for example, to avoid causal problems. The initial centroid vertex C0 can be any approximation of the centroid vertex C; for example, the initial centroid vertex C0 can be obtained from a simple (unweighted) average of the TriSoup vertices (e.g., as in the prior art). Weights of the vertices are generated using the initial centroid vertex C0, from which a weighted sum (or average) of the vertices is calculated, for example, to determine the centroid vertex C. The resulting centroid vertex (e.g., centroid vertex C) can approximate the highest centroid vertex.
[0159] For example, the area of a triangle can be approximated even without knowing its exact location (e.g., and To avoid causal problems, the area of a triangle can be calculated without any prior knowledge of the centroid vertex C and / or its approximation C0 (e.g., and The approximate value of ).
[0160] Figure 15 An example method for determining the centroid vertex is shown. More specifically, Figure 15A flowchart 1500 illustrates example method steps for determining the centroid vertex from vertices (e.g., TriSoup vertices) of a cuboid (e.g., associated with a TriSoup node) corresponding to a portion of the point cloud. One or more steps of the example flowchart 1500 can be determined by a decoder (e.g., such as...). Figure 1 The decoder 120 shown) and / or the encoder (e.g., such as Figure 1 The encoder 114 shown is used for implementation.
[0161] At step 1502, the order of the vertices (Vi) on the boundary of the cuboid can be determined. The boundary may include edges or faces of the cuboid and / or cuboid. (See the section on...) Figure 11 , Figure 12A , Figure 12B , Figure 13 , Figure 14A and / or Figure 14B As described herein, vertices can be located on the edges of the cuboid. For example, each vertex can be located on a different edge of the cuboid. There can be at least three vertices or at least four vertices, which can be as described herein. Figure 8A , Figure 8B , Figure 12A and / or Figure 12B Determined by the aforementioned. Vertex V i For example, sorting can be based on a principal direction, which is selected from one of three axes parallel to the 3D space, and is based on vertex V. i Specific (e.g., as about) Figure 8A (as described).
[0162] At step 1504, the weights of the vertices can be determined. The weights of the corresponding vertices (e.g., V) are... i Each weight of ) (e.g., It can be based on the representation of vertex V i The area of the triangle formed by the centroid vertex C to be determined and a triplet of a neighboring vertex in the sorted order (e.g., as shown in the figure). Figure 8A , Figure 8B and Figure 13 (As shown).
[0163] From TriSoup vertex V i Determine the initial centroid vertex C0, for example, to determine the centroid vertex C. For instance, the initial centroid vertex C0 can be determined as, for example, the simple mean of the TriSoup vertices.
[0164] Where N is the vertex V of TriSoup i Quantity / Number.
[0165] The initial centroid vertex C0 can be determined as the center of the cuboid associated with the TriSoup node.
[0166] If the initial centroid vertex C0 is determined, then the triangle can be formed by the initial centroid vertex C0 and following the vertex V. i The order is formed by triplets of neighboring vertices. The value can be calculated. The value approximates (or represents) a triangle. area It can be based, for example, on approximate values. Determine weights For example, with vertex V i Associated weights It can be determined as two triangles and Two approximations of the area The average value of the two triangles is that the two triangles have a TriSoup vertex V as a common vertex. i ,like Figure 14B As shown. For example, weights It can be calculated as:
[0167]
[0168] Where A is the sum of all area approximations. The sum of the weights equals one.
[0169] If the initial centroid vertex C0 is determined, then the area approximation is... It can be computed as a function of vertex V i V i+1 and the triangle defined by C0 The area or an approximation of the area. For example, this area can be represented by a triangle. The area is obtained by the cross product of the two sides. (Approximate area value) It can be determined (e.g., calculated) as:
[0170]
[0171] The norm It can be the Euclidean norm (e.g., also known as the L2 norm, which provides the true area) or any other norm, such as the Manhattan norm (e.g., also known as the L1 norm), to simplify and reduce calculations.
[0172] The value representing the area of the triangle at step 1504 can be determined even without knowing the initial (or any other) centroid vertex C0. For example, it can be determined by using edges... The length is based on the triangle that is subsequently determined (or rendered). The area approximation is determined by the side length. This is unrelated to the centroid vertex C, which will be determined later. Area or its approximation Can be with the edge The length is proportional, for example, because the future centroid vertex C, which is to be determined, can be located relative to the TriSoup vertex V. i Roughly equidistant. For example, it can be based on the representation of vertex V. i Its neighboring vertex V i+1 The value of the distance between them is used to determine (e.g., to calculate) an approximate value of the area of the triangle. As shown below:
[0173]
[0174] The norm It can be the Euclidean norm (e.g., also known as the L2 norm, which provides the true area) or any other norm, such as the Manhattan norm (e.g., also known as the L1 norm), to simplify and reduce calculations.
[0175] The L1 norm can be used as a distance metric between vertices and their neighboring vertices in a sorting process, as shown below:
[0176]
[0177] To offer the best compromise between compressed results and rapid implementation. Area approximation. It can be a vector The sum of the absolute values of the three coordinates.
[0178] At step 1506, the centroid vertex C can be determined, for example, based on the sum of vertices weighted by their respective weights as determined at step 1504. For instance, the centroid vertex C can be determined as a weighted sum of the TriSoup vertices (e.g., ).
[0179] Figure 16 An example method for obtaining a portion of a point cloud within a cuboid is shown. A portion of the point cloud can, for example, be represented by at least the determined centroid vertex. More specifically, Figure 16 A flowchart 1600 illustrates example method steps for determining the centroid vertex from vertices (e.g., TriSoup vertices) of a cuboid (e.g., associated with a TriSoup node) corresponding to a portion of the point cloud. One or more steps of example flowchart 1600 may be derived by a decoder (e.g., as described herein regarding...). Figure 1 The decoder 120 is described above, or is generated by an encoder (e.g., as described herein). Figure 1 The encoder 114 is used for implementation. For example, unless otherwise explicitly stated, the decoder and encoder can perform reciprocal operations.
[0180] At step 1602, multiple vertices of a cuboid corresponding to a portion of the point cloud can be determined. For example, the multiple vertices may include at least three or at least four vertices. The multiple vertices may be on the boundary of the cuboid. For example, the boundary may include the edges of the cuboid. For example, the boundary may include the edges and / or faces of the cuboid. For example, the multiple vertices may be on three or more edges of the cuboid.
[0181] At the decoder, determining multiple vertices may include, for each of the multiple vertices, decoding from the bitstream (e.g., entropy decoding): the presence of the vertex on an edge of the cuboid, and the position of the vertex along the edge. At the encoder, multiple vertices of the cuboid may be determined, for example, based on the spatial positions of points in a portion of the point cloud relative to the edges of the cuboid, as described in this paper. Figure 12A and Figure 12B The encoder can encode an indication of the presence of a vertex on an edge (e.g., entropy coding), and if present, an indication of the vertex's position along the edge.
[0182] Each of the adjacent vertices can be adjacent to a vertex in the ordering of multiple vertices. Pairs of vertices (e.g., referenced at step 1602) can be adjacent vertices in the ordering. The initial centroid vertex and / or the ordering of multiple vertices of the cuboid can be determined, for example, by pivoting about a principal direction selected (or determined) from one of three directions corresponding to three axes in 3D space. For example, the principal direction can be determined, for example, based on the initial centroid vertex and multiple vertices, as described herein. Figure 8A The initial centroid vertex can be determined as the average of multiple vertices.
[0183] The sorting can be cyclic ordering (e.g., called circular ordering). For example, multiple vertices can be sorted using, for instance, a circular buffer or an array with index management, such that the first and last vertices of the array can be considered to be adjacent and / or neighboring to each other. For example, multiple neighboring vertices of a vertex can include: a first neighboring vertex having a first index equal to the remainder of the vertex's index minus one divided by the number of vertices; and a second neighboring vertex having a second index equal to the remainder of the vertex's index plus one divided by the number of vertices.
[0184] At step 1604, the centroid vertex of the cuboid can be determined. The centroid vertex can be determined, for example, based on a weighted sum of multiple vertices. Each vertex in the plurality of vertices can have a weight, for example, based on the vertex's neighboring vertices from the plurality of vertices. The weights can be determined, for example, based on a value representing the distance between the vertex and each of its neighboring vertices. For example, the value representing the distance can include an L1 norm (e.g., called the Manhattan norm) or an L2 norm (e.g., called the Euclidean norm). The weights can be determined as, for example, the sum of the stated values. The centroid vertex can be determined, for example, based on a weighted sum divided by the sum of the weights corresponding to the plurality of vertices. The weights can be determined as, for example, the average of the stated values (e.g., called the mean). The centroid vertex can be determined, for example, based on a weighted average divided by the average of the weights corresponding to the plurality of vertices. The weights can be determined as, for example, the sum of the stated values divided by the sum of the weights corresponding to the plurality of vertices. The values can include: a first value representing a first distance between the vertex and a first neighboring vertex among its neighboring vertices, and / or a second value representing a second distance between the vertex and a second neighboring vertex among its neighboring vertices.
[0185] The weighted sum can be determined, for example, by dividing a linear combination of vertices by the sum of their weights, where the coefficients of the linear combination correspond to the weights of the vertices. The weights can be determined, for example, based on the cross product of the edges formed between a vertex and each of its adjacent vertices. The weight of each vertex can be proportional to the mean (or sum) of the areas of the two triangles formed by said vertex in a triangle. For example, the weight of each vertex can represent the mean (or sum) of the areas of the two triangles formed by said vertex in a triangle. The weights can be determined, for example, based on the triangle formed by the initial centroid vertex described herein. For example, the weight of each vertex can be determined, for example, based on: a first value representing the first area of a first triangle formed by the vertex, the initial centroid vertex, and the first adjacent vertex among the adjacent vertices; and / or a second value representing the second area of a second triangle formed by the vertex, the initial centroid vertex, and the second adjacent vertex among the adjacent vertices of the vertex.
[0186] The first value can be determined, for example, as the first cross product of the two sides of a first triangle or half of the first cross product. The second value can be determined, for example, as the second cross product of the two sides of a second triangle or half of the second cross product.
[0187] The weight of a vertex can be determined, for example, as proportional to the sum of the first and second values. The weight can be determined, for example, as the sum of the first and second values. The weight can be determined, for example, as half the sum of the first and second values.
[0188] The weight of a vertex can be determined, for example, proportional to the mean of the first and second values. The weight can be determined, for example, the mean of the first and second values. The weight can be determined, for example, half the mean of the first and second values.
[0189] At step 1606, triangles can be determined. Triangles can be determined, for example, based on the centroid vertex and pairs of vertices. The number of vertices can equal the number of triangles. Each triangle can be determined by the centroid vertex and the order of the vertices (e.g., as described herein). Figure 15 And in Figure 16 In step 1604, the sorting of multiple vertices (as described in the previous section) forms different neighboring vertices. The consistency and quality of triangle voxelization can be improved by adjusting the centroid vertices. For example, the consistency and quality of triangle voxelization can be improved by adjusting the centroid vertices based on the centroid residual, as discussed herein. Figure 11 As stated above.
[0190] At the decoder, determining the triangle may include, for example, decoding the centroid residual from the bitstream (e.g., entropy decoding). The decoder may determine a second centroid vertex, for example, based on the centroid vertex and the centroid residual. Each triangle comprising a set of three vertices may contain a second centroid vertex and two other vertices that may be adjacent in an order of multiple vertices. The centroid residual may be decoded from the bitstream by the decoder, which may decode a first indication (e.g., entropy decoding) to determine whether the centroid residual is equal to zero. Based on the first indication that the centroid residual is non-zero, the decoder may decode a second indication of the sign of the centroid residual and / or a third indication associated with the magnitude of the centroid residual (e.g., entropy decoding). The third indication may indicate a value equal to the magnitude of the centroid residual minus one. The decoder may determine the centroid residual as said value (e.g., indicated by the third indication and / or decoded from the third indication) plus one.
[0191] At the encoder, the encoder can determine the centroid residual, for example, based on the centroid vertex and a portion of the point cloud. Each triangle, comprising a set of three vertices, can contain a second centroid vertex and two other vertices that may be adjacent in an ordering of multiple vertices. The encoder can encode the centroid residual in the bitstream. The encoder can determine the triangle (e.g., a TriSoup triangle), for example, based on the centroid vertex and pairs of multiple vertices. The encoder can compute a normalized vector, for example, based on the mean normal of the triangle, along which the centroid residual is located.
[0192] The encoder can determine the second centroid vertex along the normalized vector, for example, based on the average of the set of points within a threshold distance of lines extending from both ends of the normalized vector in a portion of the point cloud. The centroid residual can be equal to the difference between the second centroid vertex and the centroid vertex. At step 1608, the triangle can be voxelized to determine the voxels representing the portion of the point cloud, as described herein. Figure 9 and Figure 10 As stated above.
[0193] Figure 17A An example method for encoding the centroid residual of a centroid vertex is shown. More specifically, Figure 17A A flowchart 1700 illustrates example method steps for encoding (e.g., entropy coding) the centroid residuals of centroid vertices determined from vertices (e.g., TriSoup vertices) of a cuboid (e.g., associated with a TriSoup node) corresponding to a portion of the point cloud. One or more steps of example flowchart 1700 may be generated by an encoder (e.g., as described herein). Figure 1 The encoder 114 is implemented.
[0194] At step 1702, the encoder can determine the vertices of a cuboid (e.g., TriSoup vertices) corresponding to the portion of the point cloud. For example, the cuboid can be indicated by TriSoup nodes. Step 1702 can correspond to, as described herein, […]. Figure 16 Step 1602.
[0195] At step 1704, the encoder can determine the centroid vertex of the cuboid, for example, based on a weighted sum of the first vertices. Each vertex in the first set of vertices can have a weight, for example, based on a weight of the vertex's neighboring vertices. Step 1704 can correspond to, as described herein, […]. Figure 16 Step 1604.
[0196] At step 1706, the encoder can determine the centroid residual, for example, based on the centroid vertices and portions of the point cloud. For instance, the centroid residual can be determined from the centroid vertices, as described herein. Figure 11 As stated above.
[0197] The encoder can, for example, determine the triangle based on the centroid vertex and pairs of vertices (e.g., the TriSoup triangle), as discussed in this paper. Figure 8A , Figure 8B , Figure 13 , Figure 14A and Figure 14BThe encoder can, for example, determine (e.g., compute) a normalized vector (e.g., called a unit vector) based on the average normal of a triangle. The centroid residual can be determined along the normalized vector. The encoder can, for example, determine the second centroid vertex along the normalized vector based on the average of the set of points within a threshold distance of a line extending from both ends of the normalized vector in a portion of the point cloud, as described herein. Figure 11 The centroid residual can be determined, for example, as the difference between the second centroid vertex and the centroid vertex.
[0198] At step 1708, the encoder can encode (e.g., entropy coding) the centroid residual in the bit stream, as described herein. Figure 11 The encoder can encode the centroid residual, for example, using a set of syntax elements (e.g., referred to as indicators or flags). For example, the encoder can encode a first indicator of whether the centroid residual is equal to zero (e.g., entropy coding). For example, based on the first indicator that the centroid residual is non-zero, the encoder can encode a second indicator of the sign of the centroid residual and / or a third indicator associated with the magnitude of the centroid residual (e.g., entropy coding). For example, the third indicator can be encoded as a unary codeword (e.g., using a truncated unary coding scheme). If the first indicator indicates that the centroid residual is non-zero, the third indicator can indicate a value equal to the magnitude of the centroid residual minus one. The centroid residual can be equal to said value plus one.
[0199] The encoder can, for example, determine the triangle based on the centroid vertex and pairs of vertices (e.g., the TriSoup triangle), as discussed in this paper. Figure 8A , Figure 8B , Figure 13 , Figure 14A and / or Figure 14B The encoder can, for example, determine (e.g., compute) a normalized vector (e.g., called a unit vector) based on the average normal of a triangle. The centroid residual can be determined along the normalized vector. The encoder can, for example, determine the second centroid vertex along the normalized vector based on the average of the set of points within a threshold distance of a line extending from both ends of the normalized vector in a portion of the point cloud, as described herein. Figure 11 The centroid residual can be determined, for example, as the difference between the second centroid vertex and the centroid vertex.
[0200] Figure 17B An example method for decoding the centroid residual of a centroid vertex is shown. More specifically, Figure 17BA flowchart 1750 illustrates example method steps for decoding (e.g., entropy decoding) the centroid residuals of centroid vertices determined from vertices (e.g., TriSoup vertices) of a cuboid (e.g., associated with a TriSoup node) corresponding to a portion of the point cloud. One or more steps of example flowchart 1750 can be performed by a decoder (e.g., as described herein regarding...). Figure 1 The decoder 120 is implemented.
[0201] At step 1752, the decoder can determine the vertices of the cuboid (e.g., TriSoup vertices) corresponding to the portion of the point cloud. For example, the cuboid can be indicated by TriSoup nodes. Step 1752 can correspond to, as described herein, […]. Figure 16 Step 1602. At step 1754, the decoder can determine the centroid vertex of the cuboid, for example, based on a weighted sum of the first vertices. Each vertex in the first set of vertices can have a weight, for example, based on a weight of its neighboring vertices. Step 1754 can correspond to, as described herein, regarding... Figure 16 Step 1604.
[0202] At step 1756, the decoder can decode (e.g., entropy decoding) the centroid residual (e.g., as described herein) from the bitstream. Figure 11 The decoder can, for example, use a set of syntax elements (e.g., referred to as indicators or flags) to decode the centroid residual. For example, the decoder can decode a first indicator to determine whether the centroid residual is equal to zero. For example, based on a first indicator that the centroid residual is non-zero, the decoder can decode a second indicator of the sign of the centroid residual and / or a third indicator associated with the magnitude of the centroid residual. For example, the third indicator can be decoded as a unary coded codeword (e.g., using a truncated unary coding scheme). If the first indicator indicates that the centroid residual is non-zero, the third indicator can indicate a value equal to the magnitude of the centroid residual minus one. The centroid residual can be equal to said value plus one.
[0203] At step 1758, the decoder can determine the second centroid residual (e.g., C+C). res Or the centroid adjusted via the centroid residual). The decoder can determine the second centroid residual, for example, based on the centroid vertex and the centroid residual decoded from the bitstream (e.g., C+C). res Or the centroid adjusted through the centroid residual, as discussed in this paper. Figure 8B (as described).
[0204] Figure 18 An example computer system in which instances of this disclosure may be implemented is shown. For example, Figure 18The example computer system 1800 shown can implement one or more of the methods described herein. For example, various apparatuses and / or systems described herein (e.g., in...) Figure 1 , 2 (3) can be implemented in the form of one or more computer systems 1800. Furthermore, each step in the flowcharts depicted in this disclosure can be implemented on one or more computer systems 1800.
[0205] Computer system 1800 may include one or more processors, such as processor 1804. Processor 1804 may be a dedicated processor, a general-purpose processor, a microprocessor, and / or a digital signal processor. Processor 1804 may be connected to communication infrastructure 1802 (e.g., a bus or network). Computer system 1800 may also include main memory 1806 (e.g., random access memory (RAM)) and / or secondary memory 1808.
[0206] Secondary memory 1808 may include hard disk drive 1810 and / or removable storage drive 1812 (e.g., magnetic tape drive, optical disc drive, etc.). Removable storage drive 1812 can read from and / or write to removable storage unit 1816. Removable storage unit 1816 may include magnetic tape, optical disc, etc. Removable storage unit 1816 can read from and / or write to removable storage drive 1812. Removable storage unit 1816 may contain computer-usable storage media with computer software and / or data stored therein.
[0207] Secondary memory 1808 may include other similar means for allowing computer programs or other instructions to be loaded into computer system 1800. Such means may include removable memory cell 1818 and / or interface 1814. Examples of such means may include program cartridges and / or cartridge interfaces (e.g., in video game devices) that allow software and / or data to be transferred from removable memory cell 1818 to computer system 1800, removable memory chips (e.g., erasable programmable read-only memory (EPROM) or programmable read-only memory (PROM)) and associated sockets, thumb drives and USB interfaces, and / or other removable memory cells 1818 and interfaces 1814.
[0208] Computer system 1800 may also include communication interface 1820. Communication interface 1820 allows software and data to be transferred between computer system 1800 and external devices. Examples of communication interface 1820 may include modems, network interfaces (e.g., Ethernet cards), communication ports, etc. Software and / or data transmitted via communication interface 1820 may be in the form of signals, which may be electronic, electromagnetic, optical, and / or other signals that can be received by communication interface 1820. Signals may be provided to communication interface 1820 via communication path 1822. Communication path 1822 may carry signals and may be implemented using wires or cables, optical fibers, telephone lines, cellular telephone links, RF links, and / or any other communication channels.
[0209] Computer program media and / or computer-readable media can be used to refer to tangible storage media, such as removable storage units 1816 and 1818 or a hard disk installed in hard disk drive 1810. The computer program product can be a means for providing software to computer system 1800. The computer program (which may also be referred to as computer control logic) can be stored in main memory 1806 and / or secondary memory 1808. The computer program can be received via communication interface 1820. Such a computer program, when executed, can enable computer system 1800 to implement the present disclosure as discussed herein. Specifically, when executed, the computer program can enable processor 1804 to implement the processes of the present disclosure, such as any of the methods described herein. Therefore, such a computer program can represent a controller of computer system 1800.
[0210] The features of this disclosure can be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. Implementing a hardware state machine to perform the functions described herein will also be apparent to those skilled in the art.
[0211] Figure 19Example elements of a computing device are shown that can be used to implement any of the various devices described herein, including, for example, a source device (e.g., 102), an encoder (e.g., 200), a destination device (e.g., 106), a decoder (e.g., 300), and / or any computing device described herein. The computing device 1930 may include one or more processors 1931 that can execute instructions stored in random access memory (RAM) 1933, removable media 1934 (e.g., a Universal Serial Bus (USB) drive, a compact disc (CD) or digital versatile disc (DVD) or floppy disk drive), or any other desired storage medium. Instructions may also be stored in an attached (or internal) hard disk drive 1935. The computing device 1930 may also include a security processor (not shown) that can execute instructions of one or more computer programs to monitor processes executing on the processor 1931 and any processes requesting access to any hardware and / or software components of the computing device 1930 (e.g., ROM 1932, RAM 1933, removable media 1934, hard disk drive 1935, device controller 1937, network interface 1939, GPS 1941, Bluetooth interface 1942, WiFi interface 1943, etc.). The computing device 1930 may include one or more output devices, such as a display 1936 (e.g., screen, display device, monitor, television, etc.), and may include one or more output device controllers 1937, such as a video processor. One or more user input devices 1938 may also be present, such as a remote control, keyboard, mouse, touchscreen, microphone, etc. The computing device 1930 may also include one or more network interfaces (e.g., network interface 1939), which may be wired, wireless, or a combination of both. Network interface 1939 can provide computing device 1930 with an interface to communicate with network 1940 (e.g., RAN or any other network). Network interface 1939 may include a modem (e.g., a cable modem), and external network 1940 may include a communication link, external network, home network, provider wireless, coaxial cable, fiber optic, or hybrid fiber / coaxial cable distribution system (e.g., DOCSIS network), or any other desired network. Additionally, computing device 1930 may include a location detection device, such as a Global Positioning System (GPS) microprocessor 1941, which can be configured to receive and process GPS signals and determine the geographic location of computing device 1930 with possible assistance from external servers and antennas.
[0212] Figure 19The examples shown can be hardware configurations, but the components illustrated can also be implemented as software. Modifications can be made to add, remove, combine, divide, etc., components of computing device 1930 as needed. Furthermore, basic computing devices and components can be used to implement components, and the same components (e.g., processor 1931, ROM storage device 1932, display 1936, etc.) can be used to implement any other computing devices and components described herein. For example, the various components described herein can be implemented using computing devices having components such as processors that execute computer-executable instructions stored on a computer-readable medium, such as… Figure 19 As shown. Some or all of the entities described herein may be software-based and may coexist on a common physical platform (e.g., the requesting entity may be a separate software process and program from the relevant entity, both of which may be executed as software on a common computing device).
[0213] In the following text, various features will be highlighted in a set of numbered clauses or paragraphs. These features should not be construed as limitations on the invention or inventive concept, but are merely highlights of certain features described herein, without implying a particular order of importance or relevance of such features.
[0214] Clause 1. A method comprising: determining a plurality of vertices of a cuboid associated with a portion of a point cloud associated with content.
[0215] Clause 2. The method according to Clause 1 further comprises determining the centroid vertex of the cuboid based on a weighted sum of the plurality of vertices.
[0216] Clause 3. The method according to any one of Clauses 1 to 2, wherein each of the plurality of vertices has a weight based on the vertex’s neighboring vertices among the plurality of vertices.
[0217] Clause 4. The method according to any one of Clauses 1 to 3, further comprising determining a triangle based on the centroid vertex and pairs of the plurality of vertices; and
[0218] Clause 5. The method according to any one of Clauses 1 to 4 further comprises voxelizing the triangle to determine voxels representing the portion of the point cloud.
[0219] Clause 6. The method according to any one of Clauses 1 to 5, wherein determining the triangle further comprises decoding the centroid residual from the bitstream.
[0220] Clause 7. The method according to any one of Clauses 1 to 6 further comprises determining a second centroid vertex based on the centroid vertex and the centroid residual.
[0221] Clause 8. The method according to any one of Clauses 1 to 7, wherein each of the triangles comprises a triple of vertices of the plurality of vertices, the triple of vertices comprising the second centroid vertex and neighboring vertices that are different in the ordering of the plurality of vertices.
[0222] Clause 9. The method according to any one of Clauses 1 to 8, wherein determining the triangle further comprises decoding the centroid residual from the bitstream.
[0223] Clause 10. The method according to any one of Clauses 1 to 9, wherein the decoding comprises decoding the first indication to determine whether the centroid residual is equal to zero.
[0224] Clause 11. The method according to any one of Clauses 1 to 10, further comprising rendering a point cloud frame associated with said portion of the point cloud based on at least one of the voxels.
[0225] Clause 12. The method according to any one of Clauses 1 to 11, wherein the weight is based on the distance between the vertex and each of the adjacent vertices.
[0226] Clause 13. The method according to any one of Clauses 1 to 12, wherein the weight of each vertex is proportional to the mean or sum of the areas of the two triangles associated with the vertex in the triangle.
[0227] Clause 14. The method according to any one of Clauses 1 to 13, wherein determining the plurality of vertices comprises: for each of the plurality of vertices, decoding from a bitstream: the presence of the vertex on an edge of the cuboid; and the position of the vertex along the edge.
[0228] Clause 15. The method according to any one of Clauses 1 to 14, wherein the plurality of vertices are on the boundary of the cuboid.
[0229] Clause 16. The method according to any one of Clauses 1 to 15, wherein the plurality of vertices comprises at least three vertices or at least four vertices.
[0230] Clause 17. The method according to any one of Clauses 1 to 16, wherein the pair of vertices are neighboring vertices in the sorting.
[0231] Clause 18. The method according to any one of Clauses 1 to 17, wherein the sorting is a cyclic sorting.
[0232] Clause 19. The method according to any one of Clauses 1 to 18 further comprises: determining the initial centroid vertex of the cuboid.
[0233] Clause 20. The method according to any one of Clauses 1 to 19 further comprises determining the order of the plurality of vertices based on pivoting about the determined direction.
[0234] Clause 21. The method according to any one of Clauses 1 to 20, wherein the adjacent vertex comprises: a first adjacent vertex having a first index equal to the remainder of the vertex's index minus one divided by the number of the plurality of vertices; and a second adjacent vertex having a second index equal to the remainder of the vertex's index plus one divided by the number of the plurality of vertices.
[0235] Clause 22. The method according to any one of Clauses 1 to 21 further comprises: determining the initial centroid vertex of the cuboid.
[0236] Clause 23. The method according to any one of Clauses 1 to 22 further comprises determining the ordering of the plurality of vertices by pivoting about a principal direction, the principal direction being determined based on the initial centroid vertex and the plurality of vertices being determined as one of three directions corresponding to three axes in 3D space.
[0237] Clause 24. The method according to any one of Clauses 1 to 23, wherein the initial centroid vertex is determined as the mean of the plurality of vertices.
[0238] Clause 25. The method according to any one of Clauses 1 to 24, wherein the value representing the distance contains the L1 norm (also known as the Manhattan norm) or the L2 norm (also known as the Euclidean norm).
[0239] Clause 26. The method according to any one of Clauses 1 to 25, wherein the value comprises: a first value representing a first distance between the vertex and a first adjacent vertex among the adjacent vertices; and a second value representing a second distance between the vertex and a second adjacent vertex among the adjacent vertices.
[0240] Clause 27. The method according to any one of Clauses 1 to 26, wherein the weights are determined as the sum or mean of values.
[0241] Clause 28. The method according to any one of Clauses 1 to 27, wherein the centroid vertex is further based on the weighted sum divided by the sum or mean of the weights corresponding to the plurality of vertices.
[0242] Clause 29. The method according to any one of Clauses 1 to 28, wherein the weight is determined as the sum of the values divided by the sum of the weights corresponding to the plurality of vertices.
[0243] Clause 30. The method according to any one of Clauses 1 to 29, wherein the weight is based on the cross product of the edges formed between the vertex and each of the adjacent vertices.
[0244] Clause 31. The method according to any one of Clauses 1 to 30, wherein the weight of each vertex is further determined based on: a first value representing a first area of a first triangle formed by the vertex, the initial centroid vertex and a first adjacent vertex among the adjacent vertices; and a second value representing a second area of a second triangle formed by the vertex, the initial centroid vertex and a second adjacent vertex among the adjacent vertices of the vertex.
[0245] Clause 32. The method according to any one of Clauses 1 to 31, wherein the first value is determined as a first cross product of two sides of the first triangle or half of the first cross product, and wherein the second value is determined as a second cross product of two sides of the second triangle or half of the second cross product.
[0246] Clause 33. The method according to any one of Clauses 1 to 32, wherein the weight is determined as: the sum of the first value and the second value; or the mean of the first value and the second value.
[0247] Clause 34. The method according to any one of Clauses 1 to 33, wherein the weight of the vertex is proportional to the mean of the first value and the second value.
[0248] Clause 35. The method according to any one of Clauses 1 to 34, wherein the boundary comprises the edges and faces of the cuboid, and wherein the plurality of vertices lie on three or more edges of the cuboid.
[0249] Clause 36. The method according to any one of Clauses 1 to 35, wherein each of the triangles is formed by the centroid vertex and neighboring vertices that are different in the order of the plurality of vertices.
[0250] Clause 37. The method according to any one of Clauses 1 to 36, wherein the decoding comprises decoding based on a first indication that the centroid residual is non-zero: a second indication of the sign of the centroid residual; and a third indication associated with the magnitude of the centroid residual.
[0251] Clause 38. The method according to any one of Clauses 1 to 37, wherein the third indication indicates a value equal to the magnitude of the centroid residual minus one, and wherein the centroid residual is determined to be the value plus one.
[0252] Clause 39. A computing device comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of Clauses 1 to 38.
[0253] Clause 40. A system comprising: a first computing device configured to perform a method according to any one of Clauses 1 to 38; and a second computation configured to encode the point cloud.
[0254] Clause 41. A computer-readable medium storing instructions that, when executed, cause to perform the method according to any one of Clauses 1 to 38.
[0255] Clause 42. A method comprising: determining a plurality of vertices of a cuboid associated with a portion of a point cloud associated with content.
[0256] Clause 43. The method according to Clause 42 further comprises determining the centroid vertex of the cuboid based on a weighted sum of the plurality of vertices.
[0257] Clause 44. The method according to Clauses 42 to 43, wherein each of the plurality of vertices has a weight based on the vertex’s neighboring vertices among the plurality of vertices.
[0258] Clause 45. The method according to any one of Clauses 42 to 44 further comprises determining the centroid residual based on the centroid vertex and the portion of the point cloud.
[0259] Clause 46. The method according to any one of Clauses 42 to 45, wherein each of the triangles comprises a triple of vertices of the plurality of vertices, the triple of vertices comprising a second centroid vertex and distinct neighboring vertices.
[0260] Clause 47. The method according to any one of Clauses 42 to 46 further includes encoding the centroid residual in the bit stream.
[0261] Clause 48. The method according to any one of Clauses 42 to 47 further comprises: determining a triangle based on the centroid vertex and pairs of the plurality of vertices.
[0262] Clause 49. The method according to any one of Clauses 42 to 48 further comprises calculating a normalized vector based on the average normal of the triangle, wherein the centroid residual is along the normalized vector.
[0263] Clause 50. The method according to any one of Clauses 42 to 49, further comprising: determining a second centroid vertex along the normalized vector based on the average of a set of points in the portion of the point cloud within a threshold distance of a line extending from both ends of the normalized vector, wherein the centroid residual comprises the difference between the second centroid vertex and the centroid vertex.
[0264] Clause 51. The method according to any one of Clauses 42 to 50, wherein the weighted sum is equal to a linear combination of the plurality of vertices divided by the sum of the weights, wherein the coefficients of the linear combination correspond to the weights of the plurality of vertices.
[0265] Clause 52. The method according to any one of Clauses 42 to 51, wherein the weight of each vertex represents the mean or sum of the areas of the two triangles associated with the vertex in the triangle.
[0266] Clause 53. The method according to any one of Clauses 42 to 52, wherein the number of the plurality of vertices is equal to the number of the triangles.
[0267] Clause 54. The method according to any one of Clauses 42 to 53, wherein the encoding includes encoding a first indication of whether the centroid residual is equal to zero.
[0268] Clause 55. The method according to any one of Clauses 42 to 54 further comprises: determining a triangle based on the centroid vertex and pairs of the plurality of vertices.
[0269] Clause 56. The method according to any one of Clauses 42 to 55 further comprises: calculating a normalized vector based on the average normal of the triangle, wherein the centroid residual is along the normalized vector.
[0270] Clause 57. The method according to any one of Clauses 42 to 56, further comprising: determining a second centroid vertex along the normalized vector based on the average of the set of points in the portion of the point cloud within a threshold distance of a line extending from both ends of the normalized vector.
[0271] Clause 58. The method according to any one of Clauses 42 to 57, wherein the centroid residual comprises the difference between the second centroid vertex and the centroid vertex.
[0272] Clause 59. A computing device comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of Clauses 42 to 58.
[0273] Clause 60. A system comprising: a first computing device configured to perform the method according to any one of Clauses 42 to 58; and a second calculation configured to decode the centroid residual.
[0274] Clause 61. A computer-readable medium storing instructions that, when executed, cause to perform the method according to any one of Clauses 42 to 58.
[0275] Clause 62. A method comprising: determining a plurality of vertices of a cuboid associated with a portion of a point cloud associated with content.
[0276] Clause 63. The method according to Clause 62 further comprises determining the centroid vertex of the cuboid based on a weighted sum of the plurality of vertices, wherein each of the plurality of vertices has a plurality of weights based on a corresponding weight of the vertex's adjacent vertices among the plurality of vertices;
[0277] Clause 64. The method according to any one of Clauses 62 to 63 further comprises determining the centroid residual based on the centroid vertex and the portion of the point cloud.
[0278] Clause 65. The method according to any one of Clauses 62 to 64 further includes encoding the centroid residual in the bit stream.
[0279] Clause 66. The method according to any one of Clauses 62 to 65 further comprises: determining a plurality of triangles based on the centroid vertex and pairs of the plurality of vertices.
[0280] Clause 67. The method according to any one of Clauses 62 to 66 further comprises: calculating a normalized vector based on the average normal of the plurality of triangles, wherein the centroid residuals are along the normalized vector.
[0281] Clause 68. The method according to any one of Clauses 62 to 67, wherein the encoding includes encoding an indication of whether the centroid residual is equal to zero.
[0282] Clause 69. The method according to any one of Clauses 62 to 68, wherein the encoding comprises encoding the centroid residual in the bit stream based on at least one of the following: a first indication indicating that the centroid residual is non-zero; a second indication indicating the sign of the centroid residual; and a third indication associated with the magnitude of the centroid residual.
[0283] Clause 70. A computing device comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of Clauses 62 to 69.
[0284] Clause 71. A system comprising: a first computing device configured to perform a method according to any one of Clauses 62 to 69; and a second computation configured to encode or decode the point cloud.
[0285] Clause 72. A computer-readable medium storing instructions that, when executed, cause to perform the method according to any one of Clauses 62 to 69.
[0286] Clause 73. A method comprising: determining a plurality of vertices of a cuboid corresponding to a portion of a point cloud.
[0287] Clause 74. The method according to Clause 73 further comprises determining the centroid vertex of the cuboid based on a weighted sum of the plurality of vertices, wherein each of the plurality of vertices has a weight in which the vertex has a corresponding weight based on the vertex’s neighboring vertices from the plurality of vertices.
[0288] Clause 75. The method according to any one of Clauses 73 to 74 further comprises determining the centroid residual based on the centroid vertex and the portion of the point cloud.
[0289] Clause 76. The method according to any one of Clauses 73 to 75 further includes encoding the centroid residual in the bit stream.
[0290] Clause 77. The method according to any one of Clauses 73 to 76 further comprises determining a triangle based on the centroid vertex and pairs of the plurality of vertices.
[0291] Clause 78. The method according to any one of Clauses 73 to 77 further comprises calculating a normalized vector based on the average normal of the triangle, wherein the centroid residual is along the normalized vector.
[0292] Clause 79. The method according to any one of Clauses 73 to 78, further comprising determining a second centroid vertex along the normalized vector based on the average of the set of points in said portion of the point cloud within a threshold distance of a line extending from both ends of said normalized vector; and
[0293] Clause 80. The method according to any one of Clauses 73 to 79, wherein the centroid residual comprises the difference between the second centroid vertex and the centroid vertex.
[0294] Clause 81. The method according to any one of Clauses 73 to 80, wherein the encoding includes encoding a first indication of whether the centroid residual is equal to zero.
[0295] Clause 82. The method according to any one of Clauses 73 to 81, wherein the encoding comprises encoding based on a first indication that the centroid residual is non-zero: a second indication of the sign of the centroid residual; and a third indication associated with the magnitude of the centroid residual.
[0296] Clause 83. The method according to any one of Clauses 73 to 82, wherein the third indication indicates a value equal to the magnitude of the centroid residual minus one, and wherein the centroid residual is determined to be the value plus one.
[0297] Clause 84. A computing device comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of Clauses 73 to 83.
[0298] Clause 85. A system comprising: a first computing device configured to perform the method according to any one of Clauses 73 to 83; and a second calculation configured to decode the centroid residual.
[0299] Clause 86. A computer-readable medium storing instructions that, when executed, cause to perform the method according to any one of Clauses 73 to 83.
[0300] The computing device can perform a method comprising multiple operations. The computing device can determine multiple vertices of a cuboid, which can be associated with portions of a point cloud associated with content. The computing device can determine the centroid vertex of the cuboid, for example, based on a weighted sum of the multiple vertices. Each vertex in the multiple vertices can have a weight, for example, based on the weight of its neighboring vertices in the multiple vertices. The computing device can determine a triangle, for example, based on the centroid vertex and pairs of multiple vertices. The computing device can voxelize the triangle to determine voxels representing portions of the point cloud. The computing device can decode (e.g., entropy decoding) the centroid residual from the bitstream. The computing device can determine a second centroid vertex based on the centroid vertex and the centroid residual. Each of the triangles can contain a triple of vertices in the multiple vertices, the triple containing the second centroid vertex and neighboring vertices that are different in the ordering of the multiple vertices. The computing device can decode the centroid residual from the bitstream. The computing device can decode a first indication to determine whether the centroid residual is equal to zero. The computing device can render a point cloud frame associated with a portion of the point cloud, for example, based on at least one voxel in the voxels. Weights can be based on the distance between a vertex and each of its neighboring vertices. The weight of each vertex can be proportional to the mean or sum of the areas of the two triangles associated with that vertex. Determining multiple vertices can involve, for each vertex, decoding from the bitstream: the vertex's presence on an edge of the cuboid; and the vertex's position along the edge. Multiple vertices can be on the boundaries of the cuboid. Multiple vertices can contain at least three vertices or at least four vertices. Pairs of multiple vertices are neighboring vertices in an order. The order can be a cyclic order. The computing device can determine the initial centroid vertex of the cuboid. The computing device can determine the order of multiple vertices based on pivoting about a determined direction. Neighboring vertices can contain: a first neighboring vertex having a first index equal to the remainder of the vertex's index minus one divided by the number of vertices; and a second neighboring vertex having a second index equal to the remainder of the vertex's index plus one divided by the number of vertices. The computing device can determine the initial centroid vertex of the cuboid. The computing device can determine the ordering of multiple vertices by pivoting about a principal direction, which is based on an initial centroid vertex and the multiple vertices are determined as one of three directions corresponding to three axes in 3D space. The initial centroid vertex can be determined as the mean of the multiple vertices. Values representing distances can include the L1 norm (also known as the Manhattan norm) or the L2 norm (also known as the Euclidean norm). The values can include: a first value representing a first distance between a vertex and its first neighboring vertex; and a second value representing a second distance between a vertex and its second neighboring vertex. Weights can be determined as the sum or mean of the values. The centroid vertex can be further based on a weighted sum divided by the sum or mean of the weights corresponding to the multiple vertices.Weights can be determined as the sum of values divided by the sum of weights corresponding to multiple vertices. Weights can be based on the cross product of the edges formed between a vertex and each of its neighboring vertices. The weight of each vertex can be further determined based on: a first value representing the first area of a first triangle formed by the vertex, the initial centroid vertex, and the first adjacent vertex; and a second value representing the second area of a second triangle formed by the vertex, the initial centroid vertex, and the second adjacent vertex. The first value can be determined as the first cross product of the two sides of the first triangle or half of the first cross product, and the second value can be determined as the second cross product of the two sides of the second triangle or half of the second cross product. Weights can be determined as: the sum of the first and second values; or the average of the first and second values. The weight of a vertex can be proportional to the average of the first and second values. The boundary can contain the edges and faces of a cuboid, and multiple vertices lie on three or more edges of the cuboid. Each of the triangles can be formed by the centroid vertex and its neighboring vertices, which are different in the order of the multiple vertices. Decoding may include, for example, decoding based on a first indication that the centroid residual is non-zero; a second indication of the sign of the centroid residual; and a third indication associated with the magnitude of the centroid residual. The third indication may indicate a value equal to the magnitude of the centroid residual minus one, and wherein the centroid residual is determined to be the value plus one. The computing device may include one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform the described methods, additional operations, and / or include additional elements. The system may include: a first computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to encode a point cloud. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0301] The computing device can perform a method comprising multiple operations. The computing device can determine multiple vertices of a cuboid, which may be associated with portions of a point cloud associated with content. The computing device can determine the centroid vertex of the cuboid, for example, based on a weighted sum of the multiple vertices. Each vertex may have a weight based on its neighboring vertices among the multiple vertices. The computing device can determine the centroid residual, for example, based on the centroid vertex and portions of the point cloud. Each triangle may contain a triple of vertices among the multiple vertices, the triple containing a second centroid vertex and distinct neighboring vertices. The computing device can encode (e.g., entropy encoding) the centroid residual in a bitstream. The computing device can determine triangles based on the centroid vertex and pairs of multiple vertices. The computing device can compute a normalized vector based on the average normal of the triangle, where the centroid residual is along the normalized vector. The computing device can determine a second centroid vertex along the normalized vector, for example, based on the average of the set of points in the portion of the point cloud within a threshold distance of lines extending from both ends of the normalized vector. The centroid residual contains the difference between the second centroid vertex and the centroid vertex. The weighted sum can be equal to a linear combination of multiple vertices divided by the sum of their weights, where the coefficients of the linear combination can correspond to the weights of the multiple vertices. The weight of each vertex can represent the mean or sum of the areas of the two triangles associated with that vertex. The number of vertices can be equal to the number of triangles. The encoding can include encoding a first indication of whether the centroid residual is equal to zero. The computing device can determine the triangle based on the centroid vertex and pairs of multiple vertices. The computing device can compute a normalized vector based on the mean normal of the triangle, where the centroid residual is along the normalized vector. The computing device can determine a second centroid vertex along the normalized vector based on the average of the set of points in a portion of the point cloud within a threshold distance of lines extending from both ends of the normalized vector. The centroid residual can include the difference between the second centroid vertex and the centroid vertex. The computing device can include one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform the described method, additional operations, and / or include additional elements. The system may include: a first computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to decode the centroid residual. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0302] The computing device can perform a method comprising multiple operations. The computing device can determine multiple vertices of a cuboid, which may be associated with portions of a point cloud associated with content. The computing device can determine the centroid vertex of the cuboid, for example, based on a weighted sum of the multiple vertices. Each vertex may have a weight among multiple weights based on its adjacent vertices. The computing device can determine the centroid residual, for example, based on the centroid vertex and portions of the point cloud. The computing device can encode the centroid residual in a bit stream. The computing device can determine multiple triangles, for example, based on the centroid vertex and pairs of multiple vertices. The computing device can compute a normalized vector, for example, based on the average normals of the multiple triangles. The centroid residual can be along the normalized vector. Encoding may include encoding an indication of whether the centroid residual is equal to zero. Encoding may include encoding the centroid residual in the bit stream based on at least one of the following: a first indication indicating that the centroid residual is non-zero; a second indication indicating the sign of the centroid residual; and a third indication associated with the magnitude of the centroid residual. A computing device may include one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform the described methods, additional operations, and / or include additional elements. A system may include: a first computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to encode or decode a point cloud. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0303] The computing device can perform a method comprising multiple operations. The computing device can determine the vertices of a cuboid corresponding to a portion of the point cloud. The computing device can determine the centroid vertex of the cuboid based on a weighted sum of first vertices. Each vertex in the first set of vertices can have a weight based on a weight of its neighboring vertices. The computing device can determine a centroid residual based on the centroid vertex and the portion of the point cloud. The computing device can encode (e.g., entropy encoding) the centroid residual in a bitstream. The computing device can determine a triangle based on the centroid vertex and pairs of vertices. The computing device can compute a normalized vector based on the average normal of the triangle. The centroid residual can be along the normalized vector. The computing device can determine a second centroid vertex along the normalized vector based on the average of the set of points in the portion of the point cloud within a threshold distance of lines extending from both ends of the normalized vector. The centroid residual can contain the difference between the second centroid vertex and the centroid vertex. Encoding can include encoding a first indication of whether the centroid residual is equal to zero. The encoding may include encoding based on a first indication that the centroid residual is non-zero; a second indication of the sign of the centroid residual; and a third indication associated with the magnitude of the centroid residual. The third indication indicates a value equal to the magnitude of the centroid residual minus one, and wherein the centroid residual is determined to be the value plus one. The computing device may include one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform the described methods, additional operations, and / or include additional elements. The system may include: a first computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to decode the centroid residual. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0304] One or more instances in this document can be described as processes that can be depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, and / or block diagrams. Although a flowchart can describe operations as a continuous process, one or more of the operations can be executed in parallel or simultaneously. The order of the operations shown can be rearranged. A process can be terminated when its operations are completed, but may have additional steps not shown in the diagram. A process can correspond to a method, function, program, subroutine, subroutines, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.
[0305] The operations described herein can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments for performing necessary tasks (e.g., computer program products) can be stored on a computer-readable or machine-readable medium. A processor can perform the necessary tasks. The features of this disclosure can be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. Implementing a hardware state machine to perform the functions described herein will also be apparent to those skilled in the art.
[0306] One or more features described herein may be implemented in computer-usable data and / or computer-executable instructions, executable by one or more computers or other devices, such as in one or more program modules. Generally, a program module includes routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type when executed by a processor or other data processing device in a computer. Computer-executable instructions may be stored on one or more computer-readable media, such as hard disks, optical disks, removable storage media, solid-state storage, RAM, etc. The functionality of a program module may be combined or distributed as needed. Functionality may be implemented wholly or partially as firmware or hardware equivalents, such as integrated circuits, field-programmable gate arrays (FPGAs), etc. One or more features described herein may be implemented more efficiently using specific data structures, and such data structures are contemplated within the scope of the computer-executable instructions and computer-usable data described herein. Computer-readable media may include, but are not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored but do not include carrier waves and / or transient electronic signals propagated wirelessly or via wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as CDs or DVDs, flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions that can represent any combination of programs, functions, subroutines, routines, subroutines, modules, software packages, classes or instructions, data structures, or program statements. Code segments can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0307] Non-transitory tangible computer-readable media may contain instructions executable by one or more processors configured to cause the operations described herein. Articles of manufacture may contain non-transitory tangible computer-readable machine-accessible media having instructions encoded thereon for causing programmable hardware to enable devices (e.g., encoders, decoders, transmitters, receivers, etc.) to perform the operations described herein. Devices, or one or more devices such as in a system, may include one or more processors, memories, interfaces, etc.
[0308] The communication described herein can be determined, generated, sent, and / or received using any number of messages, information elements, fields, parameters, values, indications, information, bits, etc. While this document may use any of the terms / phrases message, information element, field, parameter, value, indication, information, bit, etc., to describe one or more instances, those skilled in the art will understand that any one or more of these terms, including other such terms, can be used to perform such communication. For example, one or more parameters, fields, and / or information elements (IEs) may contain one or more information objects, values, and / or any other information. An information object may contain one or more other objects. At least some (or all) of the parameters, fields, IEs, etc., may be used and may be interchangeable depending on the context. Where a meaning or definition is given, such meaning or definition shall prevail.
[0309] One or more elements in the examples described herein can be implemented as modules. A module can be an element that performs a defined function and / or has a defined interface to other elements. Modules can be implemented as hardware, software combined with hardware, firmware, wet hardware (e.g., hardware with biological elements), or a combination thereof, all of which can be behaviorally equivalent. For example, a module can be implemented as software routines written in a computer language configured to be executed by a hardware machine (such as C, C++, Fortran, Java, Basic, Matlab, etc.) or a modeling / simulation program (such as Simulink, Stateflow, GNU Octave, or LabVIEW MathScript). Alternatively or concurrently, modules can be implemented using physical hardware that incorporates discrete or programmable analog, digital, and / or quantum hardware. Examples of programmable hardware can include: computers, microcontrollers, microprocessors, application-specific integrated circuits (ASICs); field-programmable gate arrays (FPGAs); and / or complex programmable logic devices (CPLDs). Computers, microcontrollers, and / or microprocessors can be programmed using languages such as assembly, C, C++, etc. FPGAs, ASICs, and CPLDs are typically programmed using hardware description languages (HDLs), such as VHSIC Hardware Description Language (VHDL) or Verilog. These languages configure connections between internal hardware modules with limited functionality on a programmable device. The techniques mentioned above can be combined to achieve the desired functional modules.
[0310] One or more operations described herein may be conditional. For example, one or more operations may be performed if certain criteria are met in a computing device, communication device, encoder, decoder, network, or a combination thereof. Example criteria may be based on one or more conditions, such as device configuration, traffic load, initial system settings, packet size, service characteristics, or a combination thereof. Various instances may be used if the one or more criteria are met. Any part of the instances described herein may be implemented in any order and based on any conditions.
[0311] Although examples have been described above, features and / or steps of those examples can be combined, divided, omitted, rearranged, modified, and / or expanded in any desired manner. Various changes, modifications, and improvements will readily occur to those skilled in the art. While not expressly stated herein, such changes, modifications, and improvements are intended to be part of this specification and are intended to be within the spirit and scope of the description herein. Therefore, the above description is illustrative only and not restrictive.
Claims
1. A method comprising: Identify multiple vertices of a cuboid associated with a portion of the point cloud related to the content; The centroid vertex of the cuboid is determined based on a weighted sum of the plurality of vertices, wherein each vertex of the plurality of vertices has a weight based on the vertex’s neighboring vertices among the plurality of vertices; The triangle is determined based on the centroid vertex and the pairs of vertices. as well as The triangle is voxelized to determine the voxels representing the portion of the point cloud.
2. The method of claim 1, wherein determining the triangle further comprises: Decoding the centroid residual from the bitstream; and The second centroid vertex is determined based on the centroid vertex and the centroid residual, wherein each of the triangles contains a triple of vertices from the plurality of vertices, the triple of vertices containing the second centroid vertex and neighboring vertices that are different in the order of the plurality of vertices.
3. The method according to any one of claims 1 to 2, wherein determining the triangle further comprises: Decoding the centroid residual from the bitstream, wherein the decoding includes decoding a first indication to determine whether the centroid residual is equal to zero.
4. The method according to any one of claims 1 to 3, further comprising: A point cloud frame associated with the portion of the point cloud is rendered based on at least one of the voxels.
5. The method according to any one of claims 1 to 4, wherein the weight is based on the distance between the vertex and each of the adjacent vertices.
6. The method according to any one of claims 1 to 5, wherein the weight of each vertex is proportional to the mean or sum of the areas of the two triangles associated with the vertex in the triangle.
7. The method according to any one of claims 1 to 6, wherein determining the vertex comprises: For each of the vertices, decode from the bitstream: The existence of the vertex on the edge of the cuboid; and The position of the vertex along the edge.
8. The method according to any one of claims 1 to 7, wherein the plurality of vertices are on the boundary of the cuboid.
9. The method according to any one of claims 1 to 8, wherein the plurality of vertices comprises at least three vertices or at least four vertices.
10. The method of claim 2, wherein the paired plurality of vertices are neighboring vertices in the sorting.
11. The method according to any one of claims 2 or 10, wherein the sorting is a cyclic sorting.
12. The method according to any one of claims 1 to 11, further comprising: Determine the initial centroid vertex of the cuboid; and The order of the plurality of vertices is determined based on pivoting around the initial centroid vertex and the plurality of vertices in a direction determined by the initial centroid vertex and the plurality of vertices.
13. A computing device comprising: One or more processors; and A memory for storing instructions, which, when executed, cause the wireless device to perform the method according to any one of claims 1 to 12.
14. A system comprising: A first computing device, the first computing device being configured to perform the method according to any one of claims 1 to 12; and A second computing device is configured to encode the point cloud.
15. A computer-readable medium storing instructions that, when executed, cause the method according to any one of claims 1 to 12 to be performed.