Encoding point cloud attributes
By combining the G-PCC and V-PCC standards with occupancy tree and predictive transformation techniques, the problem of excessively large point cloud data size is solved, achieving efficient storage and transmission, and making it suitable for fields such as extended reality, autonomous driving, and medicine.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- COMCAST CABLE COMM LLC
- Filing Date
- 2024-07-12
- Publication Date
- 2026-06-05
AI Technical Summary
The large size of point cloud data leads to low transmission and storage efficiency, and existing technologies struggle to effectively compress and decode it.
The geometry-based point cloud compression (G-PCC) standard and the video-based point cloud compression (V-PCC) standard are adopted. The point cloud frames are encoded by an encoder, and occupancy tree and predictive transform techniques are used to reduce redundant information. Entropy coding and lift transform are combined to further compress the point cloud data.
It achieves efficient storage and transmission of point cloud data, reduces bit rate and distortion, and is suitable for various application scenarios such as extended reality, autonomous driving and medical fields.
Smart Images

Figure CN122162380A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefits of U.S. Provisional Application No. 63 / 526,537, filed July 13, 2023; U.S. Provisional Application No. 63 / 543,770, filed October 12, 2023; and U.S. Provisional Application No. 63 / 620,452, filed January 12, 2024. All of the aforementioned applications are hereby incorporated by reference in their entirety. Background Technology
[0003] Objects or scenes can be described using volumetric visual data consisting of a series of points. Points can be stored in point cloud format, which includes a set of points in three-dimensional space. Because point cloud data can be quite large, transmitting and processing point cloud data may require data compression schemes specifically designed for the unique characteristics of point cloud data. Summary of the Invention
[0004] The following summary presents a simplified overview of certain features. This summary is neither a comprehensive overview nor intended to identify important or key elements.
[0005] Point cloud frames associated with content can include geometric and attribute information (e.g., the color or texture of the geometry). Attribute information can be encoded separately from geometric information. A reference point cloud frame can be selected to predict the attributes of the current point cloud frame. One or more attribute predictors can be determined, for example, based on projecting the attributes of the reference point cloud frame onto the geometry of the current point cloud frame. The encoder can encode residual attributes that indicate the difference between the (original or target) attribute and the attribute predictor. The decoder can obtain attribute information by generating the attribute predictor and decoding the received residual attributes. By encoding the residual attributes, for example, without encoding the (original or target) attribute and / or the attribute predictor, the coding cost (e.g., bit rate) and / or distortion of inter-frame prediction can be reduced.
[0006] These and other features and advantages are described in more detail below. Attached Figure Description
[0007] The accompanying drawings illustrate some features by way of example rather than limitation. The same numbers in the drawings indicate similar elements.
[0008] Figure 1 An example point cloud coding system is shown.
[0009] Figure 2 An example of Morton order is shown.
[0010] Figure 3 An example scan order is shown.
[0011] Figure 4 An example neighborhood of a cuboid is shown for entropy writing coding of the occupancy of a sub-cuboid.
[0012] Figure 5 An example of the dynamically decreasing function DR that can be used in dynamic OBUF is shown.
[0013] Figure 6 An example method for writing code to the occupancy of a cuboid using dynamic OBUF is shown.
[0014] Figure 7 An example of an occupied cuboid is shown.
[0015] Figure 8A An example cuboid corresponding to a TriSoup node is shown.
[0016] Figure 8B An example refinement of the TriSoup model is shown.
[0017] Figure 9 An example of voxelization is shown.
[0018] Figure 10 An example encoding method using inter-frame prediction is shown.
[0019] Figure 11 An example method for encoding point cloud attributes based on predictive transformations is shown.
[0020] Figure 12 An example method for decoding point cloud attributes based on predictive transformation is shown.
[0021] Figure 13 An example method for encoding point cloud attributes based on prediction boosting transformations is shown.
[0022] Figure 14 An example method for decoding point cloud properties based on prediction boosting transformations is shown.
[0023] Figure 15 An example of the Region Adaptive Hierarchical Transformation (RAHT) transformation for an octree is shown.
[0024] Figure 16 Another example of the RAHT transformation for an octree is shown.
[0025] Figure 17 An example method for encoding point cloud frames is shown.
[0026] Figure 18 An example method for decoding point cloud frames is shown.
[0027] Figure 19 An example method for encoding the attributes of a point cloud frame is shown.
[0028] Figure 20 Another example method for decoding the properties of a point cloud frame is shown.
[0029] Figure 21 Another example method for encoding the properties of point cloud frames is shown.
[0030] Figure 22 Another example method for decoding the properties of a point cloud frame is shown.
[0031] Figure 23 An example method for encoding residual properties is shown.
[0032] Figure 24 An example method for decoding residual properties is shown.
[0033] Figure 25 Another example method for encoding the properties of point cloud frames is shown.
[0034] Figure 26 Another example method for decoding the properties of a point cloud frame is shown.
[0035] Figure 27 An example method for encoding the attributes of a point cloud frame is shown.
[0036] Figure 28 An example method for decoding the properties of a point cloud frame is shown.
[0037] Figure 29 An example computer system in which the examples of this disclosure may be implemented is shown.
[0038] Figure 30 Example elements of a computing device that can be used to implement any of the various devices described herein are shown. Detailed Implementation
[0039] The accompanying figures and description provide examples. It should be understood that the examples shown and / or described in the figures are non-exclusive, and the features shown and described can be practiced in other examples. Examples of operation for point cloud or point cloud sequence encoding or decoding systems are provided. More specifically, the techniques disclosed herein can relate to point cloud compression, such as that used in encoding and / or decoding apparatuses and / or systems.
[0040] At least some visual data can use a series of points to describe objects or scenes in content and / or media. Each point may include a position in two-dimensional (x and y) form and one or more optional attributes, such as color. Volumetric visual data can add another positional dimension to these visual data. For example, volumetric visual data can use a series of points to describe objects or scenes in content and / or media, each point may include a position in three-dimensional (x, y, and z) form and one or more optional attributes, such as color, reflectivity, timestamp, etc. For example, volumetric visual data can provide a more immersive way to experience visual data than the aforementioned at least some visual data. For example, an object or scene described by volumetric visual data can be viewed from any (or more) angles, while an object or scene described by the aforementioned at least some visual data can typically only be viewed from the angle from which the object or scene is captured or rendered. As a representation format for visual data (e.g., volumetric visual data, 3D video data, etc.), point clouds are universal because they can represent all types of three-dimensional (3D) objects, scenes, and visual content. Point clouds are well-suited for a wide range of applications, including but not limited to: film post-production, real-time 3D immersive media or telepresence, extended reality, free-view video, geographic information systems, autonomous driving, 3D mapping, visualization, medicine, multi-view replay, and real-time light detection and ranging (LiDAR) data acquisition.
[0041] As explained in this article, volumetric visual data can be used in many applications, including extended reality (XR). XR encompasses various types of immersive technologies, including augmented reality (AR), virtual reality (VR), and mixed reality (MR). Sparse volumetric visual data can be used in the automotive industry to represent three-dimensional (3D) maps (e.g., cartography) or as input to driver assistance systems. In the case of driver assistance systems, volumetric visual data can often be fed into driving decision-making algorithms. Volumetric visual data can be used to digitally store valuable objects. In applications for the protection of cultural heritage, the goal can be to maintain a representation of objects that may be threatened by natural disasters. For example, statues, vases, and temples can be fully scanned and stored as volumetric visual data with billions of samples. This use case for volumetric visual data may be particularly relevant to valuable objects in locations prone to earthquakes, tsunamis, and typhoons. Volumetric visual data can take the form of volumetric frames. A volumetric frame can describe an object or scene captured at a specific time instance. Volumetric visual data can also take the form of a sequence of volumetric frames (referred to as a volumetric sequence or volumetric video). A sequence of volumetric frames can describe an object or scene captured at multiple different time instances.
[0042] Volumetric visual data can be stored in various formats. Point clouds can include a collection of points in 3D space. Such points can be used to create meshes including vertices and polygons, or other forms of visual content. As described herein, point cloud data can take the form of point cloud frames that describe objects or scenes in content captured in a specific time instance. Point cloud data can also take the form of a sequence of point cloud frames (e.g., point cloud video). As further described herein, point cloud data can be generated from a source device (e.g., as described herein regarding...). Figure 1 The source device 102 encodes the point cloud data, outputting a bitstream containing the encoded point cloud data. The source device can encode the point cloud data based on point cloud compression coding, for example, geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding, or next-generation coding. The destination device (e.g., as described herein regarding...) Figure 1 The destination device 106 receives a bitstream containing point cloud data and decodes the bitstream containing point cloud data. The destination device can decode the point cloud data by performing point cloud decompression coding. Decompression coding can be the reverse process of point cloud compression coding. Point cloud decompression coding can include, for example, G-PCC coding. Decoding can be used to decompress the point cloud data for display and / or other forms of consumption (e.g., further analysis, storage, etc.). The destination device (or different devices) can include, for example, a renderer for rendering the decoded point cloud data. The renderer can output content, for example, by rendering the point cloud data. The renderer can output content, for example, by rendering the point cloud data along with other data (e.g., audio data).
[0043] One format for storing volumetric visual data can be a point cloud. A point cloud can comprise a collection of points in 3D space. Each point in a point cloud can include geometric information that can indicate the point's location in 3D space. For example, the geometric information can indicate the point's location in 3D space using, for example, three Cartesian coordinates (x, y, z) and / or spherical coordinates (r, φ, θ) (e.g., if acquired by a rotation sensor). The locations of points in a point cloud can be quantized according to spatial precision. Spatial precision can be the same or different in each dimension. The quantization process can create a grid in 3D space. One or more points residing within each sub-grid volume can be mapped to the coordinates of the sub-grid center, referred to as voxels. A voxel can be viewed as a 3D extension of a pixel corresponding to a 2D image grid coordinate. For example, similar to how a pixel is the smallest unit in an example of dividing 2D space (or a 2D image) into discrete, uniform (e.g., equal-sized) regions, a voxel can be the smallest volume unit in an example of dividing 3D space into discrete, uniform regions. Points in a point cloud can include one or more types of attribute information. Attribute information can indicate the properties of a point's visual appearance. For example, attribute information can indicate the point's texture (e.g., color), material type, transparency, reflectivity, surface normal, velocity, acceleration, timestamp indicating when the point was captured, or modality (e.g., running, walking, or flying). Points in a point cloud can include light field data in the form of multi-view related texture information. Light field data can be another type of optional attribute information.
[0044] Points in a point cloud can describe objects or scenes. For example, points in a point cloud can describe the external surfaces and / or internal structures of an object or scene. Objects or scenes can be generated synthetically by computer. Objects or scenes can be generated from captures of real-world objects or scenes. Geometric information of real-world objects or scenes can be obtained through 3D scanning and / or photogrammetry. 3D scanning can include different types of scanning, such as laser scanning, structured light scanning, and / or modulated light scanning. 3D scanning can obtain geometric information. 3D scanning can obtain geometric information, for example, by moving one or more laser heads, structured light cameras, and / or modulated light cameras relative to the scanned object or scene. Photogrammetry can obtain geometric information. Photogrammetry can obtain geometric information, for example, by triangulating the same features or points in 2D photographs at different spatial displacements. Point cloud data can be in the form of point cloud frames. Point cloud frames can describe objects or scenes captured at a specific time instance. Point cloud data can be in the form of point cloud frame sequences. Point cloud frame sequences can be referred to as point cloud sequences or point cloud videos. Point cloud frame sequences can describe objects or scenes captured at multiple different time instances.
[0045] In many applications, the data size of a point cloud frame or sequence of point clouds may be too large for storage and / or transmission (e.g., too big). For example, a single point cloud may include, for example, more than one million points or even billions of points. Each point may include geometric information and one or more optional types of attribute information. The geometric information of each point may include three Cartesian coordinates (x, y, z) and / or spherical coordinates (r, φ, θ), each Cartesian and / or spherical coordinate may be represented, for example, using at least 10 bits per component or 30 bits in total. The attribute information of each point may include a texture corresponding to multiple (e.g., three) color components (e.g., R, G, and B color components). Each color component may be represented, for example, using 8-10 bits per component or 24-30 bits in total. For example, a single point may include at least 54 bits of information, with at least 30 bits of geometric information and at least 24 bits of texture. If a point cloud frame includes one million such points, each point cloud frame may require 54 million bits or 54 megabits to represent. For dynamic point clouds that change over time, at a frame rate of 30 frames per second, a data rate of 1.32 gigabits per second may be required to send (e.g., transmit) the points of a point cloud sequence. The raw representation of the point cloud may require a large amount of data, and the practical deployment of point cloud-based technologies may require compression techniques that enable the storage and distribution of point clouds at a reasonable cost.
[0046] Encoding can be used to compress and / or reduce the data size of point cloud frames or sequences to provide more efficient storage and / or transmission. Decoding can be used to decompress compressed point cloud frames or sequences for display and / or other forms of consumption (e.g., other forms of consumption by machine learning-based devices, neural network-based devices, artificial intelligence-based devices, or other types of consumption by other types of machine-based processing algorithms and / or devices). For example, distribution to and visualization by end users on AR or VR glasses or any other 3D-enabled devices, point cloud compression may be lossy (introducing differences relative to the original data). Lossy compression can allow high compression ratios but may imply a trade-off between compression and visual quality perceived by the end user. Other frameworks, such as those used in medical applications or autonomous driving, may require lossless compression to avoid altering decisions obtained, for example, based on analysis of sent (e.g., transmitted) and decompressed point cloud frames.
[0047] Figure 1An example point cloud coding (e.g., encoding and / or decoding) system 100 is illustrated. The point cloud coding system 100 may include a source device 102, a transmission medium 104, and a destination device 106. The source device 102 may encode a point cloud sequence 108 into a bit stream 110 for more efficient storage and / or transmission. The source device 102 may store the bit stream 110 and / or send (e.g., transmit) the bit stream to the destination device 106 via the transmission medium 104. The destination device 106 may decode the bit stream 110 to display the point cloud sequence 108 or for other forms of consumption (e.g., further analysis, storage, etc.). The destination device 106 may receive the bit stream 110 from the source device 102 via the storage medium or the transmission medium 104. The source device 102 and the destination device 106 may include any number of different devices. Source device 102 and destination device 106 may include, for example, interconnected clusters of computer systems, servers, desktop computers, laptop computers, tablet computers, smartphones, wearable devices, televisions, cameras, video game consoles, set-top boxes, video streaming devices, vehicles (e.g., autonomous vehicles), or head-mounted displays that act as seamless resource pools (also known as computer clouds or cloud computing). Head-mounted displays may allow users to view VR, AR, or MR scenes and adjust the view of the scene, for example, based on the movement of the user's head. Head-mounted displays may be connected (e.g., tethered) to processing devices (e.g., servers, desktop computers, set-top boxes, or video game consoles) or may be completely independent.
[0048] Source device 102 may include point cloud source 112, encoder 114, and output interface 116. For example, to encode point cloud sequence 108 into bitstream 110, source device 102 may include point cloud source 112, encoder 114, and output interface 116. For example, point cloud source 112 may provide (e.g., generate) point cloud sequence 108 from captures of natural scenes and / or synthetically generated scenes. Synthetically generated scenes may be scenes including computer-generated graphics. Point cloud source 112 may include one or more point cloud capture devices, a point cloud archive including previously captured natural scenes and / or synthetically generated scenes, a point cloud feed interface for receiving captured natural scenes and / or synthetically generated scenes from a point cloud content provider, and / or a processor for generating synthetic point cloud scenes. Point cloud capture devices may include, for example, one or more laser scanning devices, structured light scanning devices, modulated light scanning devices, and / or passive scanning devices.
[0049] Point cloud sequence 108 may include a series of point cloud frames 124 (e.g., Figure 1(Example shown). Point cloud frames can describe objects or scenes captured at a specific time instance. Point cloud sequence 108 can achieve the impression of motion by continuously presenting point cloud frames 124 of point cloud sequence 108 using constant or variable time. Point cloud frames can include a set of points (e.g., voxels) 126 in 3D space. Each point 126 can include geometric information that can indicate the position of the point in 3D space. Geometric information can indicate, for example, the position of the point in 3D space using three Cartesian coordinates (x, y, z). One or more points 126 can include one or more types of attribute information. Attribute information can indicate the nature of the visual appearance of the point. For example, attribute information can indicate, for example, the texture (e.g., color) of the point, the material type of the point, the transparency information of the point, the reflectivity information of the point, the surface normal of the point, the velocity at the point, the acceleration at the point, a timestamp indicating when the point was captured, a modality indicating how the point was captured (e.g., running, walking, or flying), etc. One or more points 126 can include light field data, for example, in the form of multi-view related texture information. Light field data can be another type of optional attribute information. The color attribute information for one or more points 126 may include a lightness value and two chromaticity values. The lightness value may represent the brightness of the point (e.g., the lightness component Y). The chromaticity values may represent the blue and red components of the point, separate from the lightness (e.g., chromaticity components Cb and Cr). Other color attribute values may be represented, for example, based on different color schemes (e.g., RGB or a monochrome color scheme).
[0050] Encoder 114 can encode point cloud sequence 108 into bitstream 110. To encode point cloud sequence 108, encoder 114 can use one or more lossless or lossy compression techniques to reduce redundant information in point cloud sequence 108. To encode point cloud sequence 108, encoder 114 can use one or more prediction techniques to reduce redundant information in point cloud sequence 108. Redundant information is information that can be predicted at decoder 120 and may not need to be sent (e.g., transmitted) to decoder 120 for accurate decoding of point cloud sequence 108. For example, the Movie Experts Group (MPEG) introduced the Geometry-Based Point Cloud Compression (G-PCC) standard (ISO / IEC Standard 23090-9: Geometry-Based Point Cloud Compression). G-PCC specifies the encoded bitstream syntax and semantics for transmitting and / or storing compressed point cloud frames, and the decoder operations for reconstructing compressed point cloud frames from the bitstream. During the standardization of G-PCC, reference software (ISO / IEC Standard 23090-21: Reference Software for G-PCC) was developed to encode the geometric and attribute information of point cloud frames. To encode the geometric information of point cloud frames, the G-PCC reference software encoder can perform voxelization. The G-PCC reference software encoder can perform voxelization, for example, by quantizing the positions of points in the point cloud. Quantizing the positions of points in the point cloud can create a mesh in 3D space. The G-PCC reference software encoder can map points to the center coordinates of the sub-mesh volume (e.g., voxel) where their quantized positions are located. The G-PCC reference software encoder can use occupancy trees to perform geometric analysis to compress the geometric information. The G-PCC reference software encoder can entropy encode the results of the geometric analysis to further compress the geometric information. To encode the attribute information of the point cloud, the G-PCC reference software encoder can use transformation tools such as Region Adaptive Hierarchical Transformation (RAHT), predictive transformation, and / or lifting transformation. Lifting transformation can be constructed on top of predictive transformation. Lifting transformation can include additional update / lifting steps. The lift transform and the predictive transform can be referred to as predictive / lift transform or predictive lift. Encoder 114 can operate in the same or similar manner as the encoder provided in the G-PCC reference software.
[0051] Output interface 116 can be configured to write and / or store bit stream 110 onto transmission medium 104. Bit stream 110 can be sent (e.g., transmitted) to destination device 106. Alternatively or additionally, output interface 116 can be configured to send (e.g., transmit), upload, and / or stream bit stream 110 to destination device 106 via transmission medium 104. Output interface 116 may include wired and / or wireless transmitters configured to send (e.g., transmit), upload, and / or stream bit stream 110 according to one or more proprietary, open-source, and / or standardized communication protocols. One or more proprietary, open-source, and / or standardized communication protocols may include, for example, the Digital Video Broadcasting (DVB) standard, the Advanced Television Systems Committee (ATSC) standard, the Integrated Services Digital Broadcasting (ISDB) standard, the Cable Data Service Interface Specification (DOCSIS) standard, the 3rd Generation Partnership Project (3GPP) standard, the Institute of Electrical and Electronics Engineers (IEEE) standard, the Internet Protocol (IP) standard, the Wireless Application Protocol (WAP) standard, and / or any other communication protocol.
[0052] The transmission medium 104 may include wireless, wired, and / or computer-readable media. For example, the transmission medium 104 may include one or more wires, cables, air interfaces, optical discs, flash memory, and / or magnetic storage. Alternatively or additionally, the transmission medium 104 may include one or more networks (e.g., the Internet) or file servers configured to store and / or transmit (e.g., transfer) encoded video data.
[0053] Destination device 106 can decode bitstream 110 into point cloud sequence 108 for display or other forms of consumption. Destination device 106 may include one or more of input interface 118, decoder 120, and / or point cloud display 122. Input interface 118 may be configured to read bitstream 110 stored on transmission medium 104. Bitstream 110 may be stored on transmission medium 104 by source device 102. Alternatively, input interface 118 may be configured to receive, download, and / or stream bitstream 110 from source device 102 via transmission medium 104. Input interface 118 may include a wired and / or wireless receiver configured to receive, download, and / or stream bitstream 110 according to one or more proprietary, open-source, standardized communication protocols and / or any other communication protocols. Examples of protocols include the Digital Video Broadcasting (DVB) standard, the Advanced Television Systems Committee (ATSC) standard, the Integrated Services Digital Broadcasting (ISDB) standard, the Cable Data Services Interface Specification (DOCSIS) standard, the 3rd Generation Partnership Project (3GPP) standard, the Institute of Electrical and Electronics Engineers (IEEE) standard, the Internet Protocol (IP) standard, and the Wireless Application Protocol (WAP) standard.
[0054] Decoder 120 can decode the point cloud sequence 108 from the encoded bit stream 110. For example, decoder 120 can operate in the same or similar manner as the decoder provided in the G-PCC reference software. Decoder 120 can decode a point cloud sequence that approximates the point cloud sequence 108. Decoder 120 can decode a point cloud sequence that approximates the point cloud sequence 108 due to, for example, lossy compression of the point cloud sequence 108 by encoder 114 and / or errors introduced into the encoded bit stream 110, for example, in the event of transmission to destination device 106.
[0055] The point cloud display 122 can display the point cloud sequence 108 to a user. The point cloud display 122 may include, for example, a cathode rate tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light-emitting diode (LED) display, a 3D display, a holographic display, a head-mounted display, or any other display device suitable for displaying the point cloud sequence 108.
[0056] Point cloud coding (e.g., encoding / decoding) system 100 is presented by way of example and not limitation. Point cloud coding systems different from and / or modified versions of point cloud coding system 100 may perform the methods and processes described herein. For example, point cloud coding system 100 may include other components and / or arrangements. Point cloud source 112 may be, for example, external to source device 102. Point cloud display device 122 may be, for example, external to destination device 106 or omitted entirely (e.g., if point cloud sequence 108 is intended to be consumed by a machine and / or storage device). Source device 102 may further include, for example, a point cloud decoder. Destination device 106 may include, for example, a point cloud encoder. For example, source device 102 may be configured to further receive an encoded bit stream from destination device 106. Receiving an encoded bit stream from destination device 106 may support bidirectional point cloud transfer between devices.
[0057] As described in this paper, an encoder can quantize the position of points in a point cloud with spatial precision, which can be the same or different in each dimension of the point. The quantization process can create a grid in 3D space. The encoder can map any point residing within each sub-grid volume to the coordinates of the sub-grid center, referred to as a voxel or volume pixel. A voxel can be viewed as a 3D extension of the pixels corresponding to the 2D image grid coordinates.
[0058] An encoder can represent a point cloud (e.g., a voxelized point cloud) or write codes to a point cloud. An encoder can, for example, use an occupancy tree to represent or write codes to a point cloud. For example, an encoder can subdivide an initial volume or cuboid containing a point cloud into sub-cuboids. The initial volume or cuboid can be referred to as a bounding box. The cuboid can be, for example, a cube. The encoder can recursively subdivide each sub-cuboid containing at least one point of the point cloud. The encoder can choose not to further subdivide sub-cuboids that do not contain at least one point of the point cloud. A sub-cuboid containing at least one point of the point cloud can be referred to as an occupied sub-cuboid. A sub-cuboid that does not contain at least one point of the point cloud can be referred to as an unoccupied sub-cuboid. The encoder can subdivide an occupied sub-cuboid into, for example, two sub-cuboids (to form a binary tree), four sub-cuboids (to form a quadtree), or eight sub-cuboids (to form an octree). The encoder can subdivide occupied sub-cuboids to obtain additional sub-cuboids. Subcubes can have the same size and shape at a given depth level in the occupancy tree. For example, if the encoder splits an occupied subcube along a plane passing through the middle of the subcube's edge, the subcubes can have the same size and shape at a given depth level in the occupancy tree.
[0059] An initial volume or cuboid containing point clouds can correspond to the root node of the occupancy tree. Each occupied sub-cuboid split from the initial volume can correspond to a node in the second level of the occupancy tree (of the root node). Each occupied sub-cuboid split from the occupied sub-cuboids in the second level can correspond to a node in the third level of the occupancy tree (outside the occupied sub-cuboids in the second level from which it splits). For each recursive splitting iteration, the occupancy tree structure can continue to form in this way until, for example, a maximum depth level of the occupancy tree is reached or each occupied sub-cuboid has a volume corresponding to a voxel.
[0060] Each non-leaf node of the occupancy tree may include an occupancy word or be associated with an occupancy word representing the occupancy status of the cuboid corresponding to the node. For example, a node in the occupancy tree corresponding to a cuboid split into eight sub-cubes may include a 1-byte occupancy word or be associated with a 1-byte occupancy word. Each bit of the 1-byte occupancy word (referred to as an occupancy bit) may represent or indicate the occupancy of a different sub-cube among the eight sub-cubes. Occupied sub-cubes may be represented or indicated by a binary "1" in the 1-byte occupancy word. Unoccupied sub-cubes may be represented or indicated by a binary "0" in the 1-byte occupancy word. Occupied and unoccupied sub-cubes may be represented or indicated by the opposite 1-bit binary value in the 1-byte occupancy word (e.g., a binary "0" indicating or indicating an occupied sub-cube and a binary "1" indicating or indicating an unoccupied sub-cube).
[0061] Each bit of the occupancy word can represent or indicate the occupancy of a different sub-cube among the eight sub-cubes. For example, the least significant bit of the occupancy word can represent or indicate the occupancy of the first sub-cube among the eight sub-cubes following the so-called Merton order. The second least significant bit of the occupancy word can represent or indicate the occupancy of the second sub-cube among the eight sub-cubes following the Merton order, and so on.
[0062] Figure 2 An example of the Morton order is shown. More specifically, Figure 2 The Morton order of the eight sub-cubes 202-216, split from cuboid 200, is shown. Sub-cubes 202-216 can be labeled, for example, based on their Morton order, where child node 202 is the first in the Morton order and child node 216 is the last. The Morton order of sub-cubes 202-216 can be a local lexicographical order in xyz.
[0063] The geometry of a point cloud can be represented by the initial volume and occupancy word of a node in an occupancy tree, and can be determined from the initial volume and the occupancy word. An encoder can send (e.g., transmit) the initial volume and occupancy word of a node in the occupancy tree to a decoder in a bitstream for point cloud reconstruction. The encoder can entropy encode the occupancy word. The encoder can entropy encode the occupancy word, for example, before sending (e.g., transmitting) the initial volume and occupancy word of a node in the occupancy tree. The encoder can encode the occupancy bits of the occupancy word of a node corresponding to a cuboid. The encoder can encode the occupancy bits of the occupancy word of a node corresponding to a cuboid that is adjacent to or spatially close to the cuboid whose occupancy bit is being encoded, for example, based on one or more occupancy bits of the occupancy word of another node corresponding to a cuboid that is adjacent to or spatially close to the cuboid whose occupancy bit is being encoded.
[0064] The encoder and / or decoder can encode (e.g., encode and / or decode) the occupancy bits of occupancy words in scan order. Scan order can also be referred to as scanning order. For example, the encoder and / or decoder can scan the occupancy tree in breadth-first order. All occupancy words of nodes at a given depth (e.g., level) within the occupancy tree can be scanned. All occupancy words of nodes at a given depth (e.g., level) within the occupancy tree can be scanned, for example, before scanning the occupancy words of nodes at the next depth (e.g., level). Within a given depth, the encoder and / or decoder can scan the occupancy words of nodes in Morton order. Within a given node, the encoder and / or decoder can further scan the occupancy bits of the node's occupancy words in Morton order.
[0065] Figure 3An example scan order is shown. Figure 3 An example scan order (e.g., breadth-first order as described herein) is shown for occupied tree 300. More specifically, Figure 3 The scan order for the first three example levels of occupancy tree 300 is shown. Figure 3 In the given tree, the cuboid (e.g., cube) 302 corresponding to the root node of occupancy tree 300 can be divided into eight sub-cuboids (e.g., sub-cuboids). Two of the eight sub-cuboids, 304 and 306, may be occupied. The other six sub-cuboids may be unoccupied. Following Merton order, the first eight occupancy words (e.g., occW) are... 1,1 The occupancy word can be constructed to represent the root node. The first eight bits of the occupancy word (e.g., occW) 1,1 Each occupancy bit can represent or indicate the occupancy of a sub-cube in eight sub-cubes ordered by Morton. For example, the first eight-bit occupancy word occW 1,1 The least significant occupancy bit can represent or indicate the occupancy of the first sub-cube in the eight sub-cubes ordered by Morton. The first eight-bit occupancy word is occW. 1,1 The second least significant occupancy bit can represent or indicate the occupancy of the second sub-cube in the eight sub-cubes in Morton order, etc.
[0066] Each of the occupied subcubes (e.g., the two occupied subcubes 304 and 306) can correspond to a node other than the root node in the second level of the occupancy tree 300. Each of the occupied subcubes (e.g., the two occupied subcubes 304 and 306) can be further subdivided into eight subcubes. For example, one of the eight subcubes subdivided from subcube 304, subcube 308, may be occupied, and the other seven may be unoccupied. Three of the eight subcubes subdivided from subcube 306, subcubes 310, 312, and 314, may be occupied, and the other five may be unoccupied. Two second octet occWs can be constructed in this order. 2,1 and occW 2,2 , to represent the occupancy word corresponding to the node of subcube 304 and the occupancy word corresponding to the node of subcube 306, respectively.
[0067] Each of the occupied sub-cubicles (e.g., four occupied sub-cubicles 308, 310, 312, and 314) can correspond to a node in the third level of the occupancy tree 300. Each of the occupied sub-cubicles (e.g., four occupied sub-cubicles 308, 310, 312, and 314) can be further subdivided into eight sub-cubicles each, or a total of 32 sub-cubicles. For example, four third-level eight-bit occupancy words (occW) can be constructed in this order. 3,1 occW 3,2 occW 3,3 and occW 3,4 , respectively representing the occupancy word corresponding to the node of sub-cube 308, the occupancy word corresponding to the node of sub-cube 310, the occupancy word corresponding to the node of sub-cube 312, and the occupancy word corresponding to the node of sub-cube 314.
[0068] The occupancy words of the example occupancy tree 300 can be entropy-written coded (e.g., entropy-encoded by an encoder and / or entropy-decoded by a decoder) following, for example, the scan order discussed herein (e.g., Morton's order). The occupancy words of the example occupancy tree 300 can be entropy-written coded (e.g., entropy-encoded by an encoder and / or entropy-decoded by a decoder) into, for example, a sequence of seven occupancy words occW following the scan order discussed herein. 1,1 to occW 3,4 The scanning order discussed in this paper can be a breadth-first scanning order. For example, if the occupancy word of the current child node belonging to the current parent node is being entropy-written, then the occupancy words of all nodes with the same depth (or level) as the current parent node may have already been entropy-written. For example, the occupancy words of all nodes with the same depth (e.g., level) as the current child node and with a lower Morton order than the current child node may also have already been entropy-written. A portion of the written occupancy words can be used to entropy-write the occupancy word of the current child node. The written occupancy words of adjacent parent and child nodes can be used, for example, to entropy-write the occupancy word of the current child node. For example, if a specific occupancy bit of the occupancy word of the current child node is being written (e.g., entropy-written), then the occupancy bits of occupancy words with a lower Morton order than that specific occupancy bit may also have already been entropy-written and can be used to write the occupancy bits of the occupancy word of the current child node.
[0069] Figure 4 An example neighborhood of a cuboid is shown for entropy-writing coding of the occupancy of a sub-cuboid. More specifically, Figure 4 An example neighborhood of a cuboid with written-code occupant bits is shown. The neighborhood of a cuboid with written-code occupant bits can be used for entropy writing of the occupant bits of the current sub-cuboid 400. This can be based, for example, on the representation as discussed herein. Figure 4The scanning order of the occupancy tree of the cuboid geometry determines the neighborhood of the cuboid with the written code occupant bit. The neighborhood of a cuboid, i.e., the neighborhood of the current child cuboid, can include one or more of the following: cuboids adjacent to the current child cuboid, cuboids sharing vertices with the current child cuboid, cuboids sharing edges with the current child cuboid, cuboids sharing faces with the current child cuboid, parent cuboids adjacent to the current child cuboid, parent cuboids sharing vertices with the current child cuboid, parent cuboids sharing edges with the current child cuboid, parent cuboids sharing faces with the current child cuboid, parent cuboids adjacent to the current parent cuboid, parent cuboids sharing vertices with the current parent cuboid, parent cuboids sharing edges with the current parent cuboid, parent cuboids sharing faces with the current parent cuboid, etc. For example... Figure 4 As shown, the current child cuboid 400 can belong to the current parent cuboid 402. Following the scanning order of the occupancy words and occupancy bits of the occupancy tree nodes, the occupancy bits of the four child cuboids 404, 406, 408, and 410 belonging to the same current parent cuboid 402 may have already been written. The occupancy bits of the previous parent cuboid's child cuboid 412 may have already been written. The occupancy bits of the parent cuboid 414 may have already been written, while the occupancy bits of its child cuboids have not yet been written. The written occupancy bits of cuboids 404, 406, 408, 410, 412, and 414 can be used to write the occupancy bits of the current child cuboid 400.
[0070] The number (e.g., quantity) of possible occupancy configurations (e.g., a set of one or more occupancy words and / or occupancy bits) in the neighborhood of the current sub-cuboid can be 2. N , where N is the number (e.g., quantity) of cuboids with written occupancy bits in the neighborhood of the current child cuboid. The neighborhood of the current child cuboid can include dozens of cuboids. The neighborhood of the current child cuboid (e.g., dozens of cuboids) can include 26 neighboring parent cuboids that share faces, edges, and / or vertices with the parent cuboid of the current child cuboid, and several neighboring child cuboids that share faces, edges, and / or vertices with written occupancy bits. The occupancy configuration of the neighborhood of the current child cuboid can have billions of possible occupancy configurations, or even be limited to a subset of neighboring cuboids, making it impractical to use directly. Encoders and / or decoders can use the occupancy configuration of the neighborhood of the current child cuboid to select a context (e.g., a probabilistic model) from the context set for a binary entropy writer (e.g., a binary arithmetic writer) that can write occupancy bits of the current child cuboid. Context-based binary entropy coding can be similar to the context-adaptive binary arithmetic coder (CABAC) used in MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC)).
[0071] The encoder and / or decoder can use several methods to reduce the occupancy configuration of the neighborhood of the current child cuboid being coded to an actual number (e.g., quantity) of reduced occupancy configurations. This includes the six neighboring parent cuboids sharing a face with the current child cuboid. 6 Alternatively, 64 occupied configurations can be reduced to 9 occupied configurations. This reduction can be achieved by using geometric invariants. It can be reduced from 26 neighboring parent cuboids. 26 An occupancy configuration yields the current occupancy score for a sub-cube. The score can be further reduced to a ternary occupancy prediction (e.g., "predicted occupied," "uncertain," or "predicted unoccupied") by using a score threshold. Individual occupancy of these sub-cubes can be replaced by the number (e.g., quantity) of occupied neighboring sub-cubes and the number (e.g., quantity) of unoccupied neighboring sub-cubes.
[0072] Using / employing one or more of the methods described herein, the encoder and / or decoder can reduce the number (e.g., quantity) of possible occupancy configurations in the neighborhood of the current subcube to a more manageable number (e.g., thousands). It has been observed that instead of directly associating the reduced number (e.g., quantity) of contexts (e.g., probabilistic models) with the reduced occupancy configuration, another mechanism, namely the Optimal Binary Writer (OBUF) that supports on-the-fly updates, can be used. The encoder and / or decoder can implement the OBUF to limit the number (e.g., quantity) of contexts to a lower number (e.g., 32 contexts).
[0073] OBUF can use a finite number (e.g., 32) of contexts (e.g., probabilistic models). The number of contexts in OBUF (e.g., quantity) can be a fixed number (e.g., fixed quantity). Contexts used by OBUF can be sorted, indexed by context indices (e.g., context indices in the range of 0 to 31), and associated from the lowest virtual probability to the highest virtual probability to encode "1". A lookup table (LUT) for context indices can be initialized at the start of the point cloud coding process. For example, the LUT can initially point to a context with a median virtual probability (e.g., with context index 15) to encode "1" for all inputs. The LUT can initially point to a context with a median virtual probability to encode "1" for all inputs in a finite number (e.g., quantity) of contexts. This LUT can take the occupancy configuration of the neighborhood of the current subcube as input and output the context index associated with the occupancy configuration. The LUT can have as many entries as the reduced occupancy configuration (e.g., approximately several thousand entries). The write-code of the current child cuboid's occupancy bits may include the following steps: determining the reduced occupancy configuration of the current child node; obtaining a context index by using the reduced occupancy configuration as an entry in the LUT; writing the current child cuboid's occupancy bits using the context pointed to (e.g., indicated by) the context index; and updating the LUT entry corresponding to the reduced occupancy configuration, for example, based on the value of the written-code occupancy bits of the current child cuboid. For example, if a binary "0" (e.g., indicating that the current child cuboid is not occupied) is being written, the LUT entry may be reduced to a lower context index value. For example, if a binary "1" (e.g., indicating that the current child cuboid is occupied) is being written, the LUT entry may be increased to a higher context index value. The context index update process may, for example, be based on a theoretical model of the optimal distribution of virtual probabilities associated with a finite number (e.g., quantity) of contexts. This virtual probability may be fixed by the model and may differ from the internal probabilities of the context that may evolve, for example, when the write-code of a data bit occurs. The evolution of the internal context may follow a well-known process similar to that in CABAC.
[0074] The encoder and / or decoder can implement a “dynamic OBUF” scheme. For example, compared to a general OBUF, a “dynamic OBUF” scheme allows the encoder and / or decoder to handle a much larger number (e.g., quantity) of occupancy configurations in the neighborhood of the current sub-cube. The use of a larger number (e.g., quantity) of occupancy configurations in the neighborhood of the current sub-cube can result in improved compression capabilities while keeping complexity within reasonable limits. By using an occupancy tree compressed by OBUF, the encoder and / or decoder can achieve lossless compression performance as good as 1 bit / point (bpp) for writing geometry to dense point clouds. The encoder and / or decoder can implement dynamic OBUF to potentially further reduce the bit rate by more than 25%, to 0.7 bpp.
[0075] OBUF may not take into account the various reduced occupancy configurations of the current child cuboid's neighborhood as input, and may potentially lead to a loss of useful relevance. With OBUF, the size of the LUT with the context index can be increased to handle more diverse occupancy configurations of the current child cuboid's neighborhood as input. Due to this increase, statistics can be diluted, and compression performance may deteriorate. For example, if the LUT has millions of entries and the point cloud has hundreds of thousands of points, most entries may never be accessed (e.g., lookup, access, etc.). Many entries may only be accessed a few times, and their associated context index may not be updated enough to reflect any meaningful correlation between the current child cuboid's occupancy configuration value and occupancy probability. Dynamic OBUF can be implemented to mitigate the dilution of statistics caused by an increase in the number (e.g., quantity) of occupancy configurations in the current child cuboid's neighborhood. This mitigation can be performed through a "dynamic reduction" of the occupancy configuration in dynamic OBUF.
[0076] Dynamic OBUF can, for example, add an extra step to reduce the occupancy configuration of the current child cuboid's neighborhood before using a context-indexed LUT. This step can be called dynamic reduction because it evolves, for example, based on the progress of writing code to the point cloud, or more precisely, based on the occupancy configuration that has already been visited (e.g., looked up in the LUT).
[0077] As discussed in this paper, there may be many possible occupancy configurations potentially involving the neighborhood of the current sub-cube, but only a subset can be accessed if a write-code of the point cloud occurs. This subset can characterize the type of point cloud. For example, most accessed occupancy configurations may show the occupied neighboring cuboids of the current sub-cube, e.g., if an AR or VR dense point cloud is being written-coded. On the other hand, for example, if a sparse point cloud acquired by a sensor is being written-coded, most accessed occupancy configurations may only show a few occupied neighboring cuboids of the current sub-cube. The effect of dynamic reduction can be, for example, to obtain a more accurate correlation by simultaneously shelving (e.g., actively reducing) other occupancy configurations accessed much less frequently based on the most accessed occupancy configuration. Dynamic reduction can be updated on an in real-time basis. For example, if a write-code of occupancy data occurs, dynamic reduction can be updated on an in real-time basis, e.g., after each access to an occupancy configuration (e.g., a lookup in a LUT).
[0078] Figure 5 An example of a dynamically decreasing function DR that can be used in a dynamic OBUF is shown. The dynamically decreasing function DR can be implemented by masking the occupancy of bit β in configuration 500. j To obtain,
[0079] β = β1 … β K
[0080] The occupancy configuration consists of K bits. For example, if the occupancy configuration is accessed (e.g., looked up in the LUT) a certain number of times (e.g., a quantity), the mask size can be reduced. The initial dynamic reduction function DR... 0 It can mask all bits of all occupied configurations, making it a constant function DR for all occupied configurations β. 0 (β) = 0. The dynamically decreasing function can be derived from the function DR. n Evolved to the update function DR n+1 The dynamic reduction function can, for example, be applied to the DR function after each write operation of the occupied bits. n Evolved to the update function DR n+1 A function can be defined as follows:
[0081] β' = DR n (β) = β1 … β kn(β)
[0082] Where k n (β) 510 is the number of unmasked bits (e.g., quantity). DR 0 The initialization can correspond to k0(β) = 0, and the natural evolution of the decreasing function toward finer statistics can cause an increase in the number (e.g., quantity) of unmasked bits, k n (β) ≤ k n+1(β). The dynamic reduction function can be completely determined by all k occupying the configuration β. n The value is determined.
[0083] For all dynamically decreasing occupancy configurations β' = DR n Access to the occupancy configuration (β), such as an instance lookup in a LUT, can be tracked by the variable NV(β'). For example, in a LUT based on the occupancy configuration β... V After each instance of writing the occupant bit, the corresponding access count (e.g., quantity) NV(β) V ') can be increased by one. If this number of visits (e.g., quantity) NV(β) V ') greater than the threshold th V ,
[0084] NV(β V ') > th V
[0085] Then the number (e.g., quantity) of unmasked bits k n (β) for dynamically decreasing to β V For all occupancy configurations β, one can be added. This corresponds to the occupancy configuration β that will dynamically decrease. V 'Use two new dynamically reduced occupancy configurations β 0 'and β 1 The two new dynamically reduced occupancy configurations are defined as follows:
[0086] β 0 ' = β V '0 = β V 1 … β V kn(β) 0 and β 1 ' = β V '1 = β V 1 … β V kn(β) 1.
[0087] In other words, for all occupancy configurations β, the number (e.g., quantity) of unmasked bits has increased by one, k n+1 (β) =k n (β) + 1, making DR n (β) = β V The access count (e.g., quantity) of two new dynamically decreasing occupancy configurations can be initialized to zero.
[0088] NV(β 0 ') = NV(β 1 ') = 0. (I)
[0089] At the start of writing code, the initial dynamic decrease function DR 0 The initial number of visits (e.g., quantity) can be set to
[0090] NV(DR 0 (β)) = NV(0) = 0,
[0091] Furthermore, the evolution of NV with dynamically decreasing occupancy configuration can be fully defined.
[0092] The corresponding LUT entry LUT[β] V '] can be two new entries LUT[β] 0 '] and LUT[β] 1 The two new entries are replaced by '], and the two new entries are replaced by β. V 'Associated writer index initialization. For example, the corresponding LUT entry LUT[β] V '] can be two new entries LUT[β] 0 '] and LUT[β] 1 The two new entries are replaced by '], and the two new entries are replaced by β. V 'Associated writer index initialization, if dynamically reduced occupancy configuration β V 'By two new dynamically reduced occupancy configurations β 0 'and β 1 'replace,
[0093] LUT[β 0 '] = LUT[β 1 '] = LUT[β V '], (II)
[0094] And then it evolves independently. The evolution of the LUT with a dynamically decreasing occupancy configuration for the writer index can be fully defined.
[0095] Decrease function DR n It can be formed by a series of growing binary trees T n Model 520, whose leaf node 530 is a reduced occupancy configuration β' = DR n (β). The initial tree can be 0 = DR 0 (β) The associated single root node. This will be dynamically reduced to β. V Replace with β 0 'and β 1 'Can correspond to growth tree T' n (from β) V 'Associated leaf nodes', for example, by connecting with β 0 'and β 1 Two new nodes associated with each other are attached to the leaf node. Tree T n+1This can be obtained through such growth. The LUTs for the number of visits (e.g., quantity) NV and the context index can be defined on the leaf nodes and evolve as the tree grows via equations (I) and (II).
[0096] The practical implementation of dynamic OBUF can be achieved by storing an array NV[β'] and a LUT[β'] of context indexes, as well as a tree T. n 520 is used for this. An alternative to storing the tree could be an array k storing the number (e.g., quantity) of the unmasked bits. n [β]510.
[0097] One limitation of implementing dynamic OBUF is its memory footprint. In some applications, millions of occupied configurations can be processed, resulting in approximately 20 bits of beta. i The entries β constitute the configuration for the decrease function DR. Each bit β i It can correspond to the occupancy status of the adjacent cuboids of the current child cuboid or the set of adjacent cuboids of the current child cuboid.
[0098] The higher (e.g., higher effective) bit β i (For example, β0, β1, etc.) can be the first bit without masking. Higher (e.g., higher-active) bits β i (For example, β0, β1, etc.) could be, for example, the first unmasked bit during the evolution of the dynamically decreasing function DR. Place bit β... i The order of neighbor-based information in compression can affect compression performance. Neighbor information can be ordered from highest (e.g., highest) priority to lowest priority, and then placed into bit β in this order from highest weight to lowest weight. i In this context, priority is ranked from most important to least important: occupancy of adjacent child cuboids, followed by adjacent child cuboids, then adjacent parent cuboids, then non-adjacent child nodes, and finally non-adjacent parent nodes. Neighboring nodes sharing a face with the current child node can have a higher priority than neighboring nodes sharing an edge (but not a face) with the current child node. Neighboring nodes sharing an edge with the current child node can have a higher priority than neighboring nodes sharing only a vertex with the current child node.
[0099] Figure 6 An example method for writing code to the occupancy of a cuboid using dynamic OBUF is shown. More specifically, Figure 6 An example method for writing the occupancy bits of the current subcube using dynamic OBUF is shown. Figure 6 One or more steps can be performed by an encoder and / or decoder (e.g., Figure 1The encoder 114 and / or decoder 120 in the flowchart can be executed in whole or in part by a programmer (e.g., Figure 1 (encoder 114 and / or decoder 120 in the middle) Figure 27 Example computer system 2700 and / or Figure 28 Example computing device 2830 is implemented.
[0100] At step 602, the occupancy configuration of the current sub-cuboid (e.g., occupancy configuration β) can be determined. The current occupancy configuration (e.g., occupancy configuration β) can be determined, for example, based on the occupancy bits of the written-code cuboids in the neighborhood of the current sub-cuboid. At step 604, the occupancy configuration (e.g., occupancy configuration β) can be dynamically reduced. For example, a dynamic reduction function DR can be used. n This allows for dynamic reduction of the occupied configuration. For example, the occupied configuration β can be dynamically reduced to a reduced occupied configuration β' = DR. n (β). At step 606, the context index can be looked up in a lookup table (LUT), for example. For example, the encoder and / or decoder can look up the context index LUT[β'] in the LUT of the dynamic OBUF. At step 608, a context (e.g., a probabilistic model) can be selected. For example, the context pointed to by the context index (e.g., a probabilistic model) can be selected. At step 610, the occupancy of the current sub-cube can be entropy-coded. For example, the occupancy bits of the current sub-cube can be entropy-coded (e.g., arithmetic coding) based on the context. The occupancy bits of the current sub-cube can be coded based on the occupancy bits of the adjacent coded cuboids.
[0101] although Figure 6 Not shown, but the encoder and / or decoder can update the reduction function and / or update the context index. For example, the encoder and / or decoder can update the reduction function DR. n Updated to DR n+1 And / or, for example, update the context index LUT[β'] based on the occupancy bits of the current child cuboid. Figure 6 The method can be based on the scanning order, as discussed in this article. Figure 3 The scanning order discussed repeats for the additional or all child cuboids of the parent cuboid corresponding to the node occupying the tree.
[0102] Occupation tree is typically a lossless compression technique. Occupation tree can be adapted to provide lossy compression, for example, by modifying the point cloud on the encoder side (e.g., downsampling, removing points, moving points, etc.). Lossy compression performance can be weak. For dense point clouds, lossy compression can be a useful lossless compression technique.
[0103] One approach to lossy compression of point cloud geometry could be to set the maximum depth of the occupancy tree to stop at a larger volume size (e.g., an NxNxN cuboid (e.g., a cube), where N > 1) instead of reaching the minimum volume size of a single voxel. The geometry of points belonging to each occupied leaf node associated with this larger volume can then be modeled. This approach may be particularly well-suited for dense and smooth point clouds that can be locally modeled using smoothing functions such as planes or polynomials. The coding cost can be reduced to the cost of the occupancy tree plus the cost of the local model within each occupied leaf node.
[0104] A scheme for modeling the geometry of points belonging to each occupied leaf node associated with a volume larger than one voxel can use a set of triangles as a local model. This scheme can be called a "TriSoup" scheme. TriSoup is short for "Triangle Soup" because the connections between triangles may not be part of the model. An occupied leaf node of the occupancy tree corresponding to a cuboid with a volume larger than one voxel can be called a TriSoup node. An edge belonging to at least one cuboid corresponding to a TriSoup node can be called a TriSoup edge. A TriSoup node can include an existence flag (s) for each TriSoup edge of its corresponding occupied cuboid. k The existence flag of a TriSoup edge (s) k ) can indicate the TriSoup vertex (V k Does a vertex exist on a TriSoup edge? At most one TriSoup vertex (V) exists. k A vertex (V) can exist on a TriSoup edge. For each vertex (V) existing on a TriSoup edge of an occupied cuboid... k The TriSoup node corresponding to the occupied cuboid can further include vertices (V). k ) along the position of TriSoup edge (p k ).
[0105] In addition to the occupancy word of the occupancy tree, the encoder can also entropy encode the TriSoup vertex presence flag and position for each TriSoup edge belonging to the TriSoup node of the occupancy tree. Similarly, in addition to the occupancy word of the occupancy tree, the decoder can also entropy decode the TriSoup vertex presence flag and position for each TriSoup edge and the vertices along the corresponding TriSoup edges belonging to the TriSoup nodes of the occupancy tree.
[0106] Figure 7An example of an occupied cuboid (e.g., a cube) 700 is shown. More specifically, Figure 7 An example of an occupied cuboid (e.g., a cube) 700 of size NxNxN (where N > 1) corresponding to a TriSoup node in the occupied tree is shown. The occupied cuboid 700 may include edges (e.g., TriSoup edges 710-721). The TriSoup node corresponding to the occupied cuboid 700 may include an presence flag (s) for each edge (e.g., each TriSoup edge in TriSoup edges 710-721). k For example, the presence flag of TriSoup edge 714 can indicate that TriSoup vertex V1 exists on TriSoup edge 714. The presence flag of TriSoup edge 715 can indicate that TriSoup vertex V2 exists on TriSoup edge 715. The presence flag of TriSoup edge 716 can indicate that TriSoup vertex V3 exists on TriSoup edge 716. The presence flag of TriSoup edge 717 can indicate that TriSoup vertex V4 exists on TriSoup edge 717. The presence flags of the remaining TriSoup edges can each indicate that no TriSoup vertex exists on its corresponding TriSoup edge. The TriSoup node corresponding to the occupied cuboid 700 can include the position of each TriSoup vertex that exists along one of its TriSoup edges 710-721. More specifically, the TriSoup node corresponding to the occupied cuboid 700 can include the position p1 of TriSoup vertex V1, the position p2 of TriSoup vertex V2, the position p3 of TriSoup vertex V3, and the position p4 of TriSoup vertex V4. TriSoup vertices can be shared among TriSoup nodes along common TriSoup edges.
[0107] The existence of the current TriSoup edge can be flagged (s) k ) and (in the presence of signs (s k (The location (p) can indicate the existence of a vertex) k Entropy coding is performed. Existence flag (s) k ) and location (p k The information can be referred to individually or collectively as vertex information or TriSoup vertex information. For example, the existence flag (s) of the current TriSoup edge can be determined based on the existence flags and positions of the existing TriSoup vertices of the TriSoup edges adjacent to the current TriSoup edge. k ) and (in the presence of signs (s k (indicating the presence of a vertex) position (p) kEntropy coding is performed on the current TriSoup edge. Alternatively or alternatively, the existence flag (s) of the current TriSoup edge can be added. k ) and (in the presence of signs (s k (The location (p) can indicate the existence of a vertex) k (e.g., indicating the position of the vertex along which the edge is located) is entropy-coded. The presence flag of the current TriSoup edge (s) k ) and location (p k Entropy coding can be performed alternatively or alternatively, for example, based on the occupancy of cuboids adjacent to the current TriSoup edge. Similar to the entropy coding of occupancy bits in an occupancy tree, the configuration β of the neighborhood of the current TriSoup edge can be obtained. TS (Also known as neighborhood configuration β) TS ), and, for example, dynamically reduce it to a reduced configuration β by using TriSoup's dynamic OBUF scheme. TS ' = DR n (β) TS ). Context index LUT[β TS The information can be obtained from the OBUF LUT. At least a portion of the vertex information for the current TriSoup edge can be entropy-coded using the context pointed to by the context index (e.g., a probabilistic model).
[0108] The position of the TriSoup vertex along its TriSoup edge (p k (If it exists) can be binarized. The position of the TriSoup vertex along its TriSoup edge (p k (If present) can be binarized, for example, by using a binary entropy writer to entropy-code at least a portion of the vertex information of the current TriSoup edge. The number of bits (e.g., quantity) N can be set. b To quantize the TriSoup vertex positions (p) along a TriSoup edge of length N. k A TriSoup edge of length N can be uniformly divided into 2. Nb Each quantization interval. By doing so, the TriSoup vertex position (p k (This can be written separately by N using a dynamic OBUF scheme) b p k j 、 j=1、……、N b ) and corresponding to the existence flag (s) k The bit representation of ). Neighborhood configuration β TS、 OBUF decrease function DR n The context index can depend on the bits of the write-code (e.g., the presence flag).k ), highest position (p) k1 ), second highest position (p) k2 The bits of the written code (e.g., the presence of a flag (s)) k ), highest position (p) k 1 ), second highest position (p) k 2 Properties, characteristics, and / or attributes of vertex information. In reality, there may be several dynamic OBUF schemes, each dedicated to specific bits of vertex information (e.g., presence flags). k ) or position (p) k j )).
[0109] Figure 8A An example cuboid 800 (e.g., a cube) corresponding to a TriSoup node is shown. The cuboid 800 can correspond to a TriSoup node with a number of K vertices V. k The TriSoup node. Within the cuboid 800, the TriSoup triangle can be formed by the TriSoup vertex V. k Construction. For example, if there are at least three (K≥3) TriSoup vertices on the TriSoup edges of a cuboid 800, then a TriSoup triangle can be constructed from TriSoup vertex V. k Construction. For example, regarding... Figure 8A A TriSoup can have four vertices and can construct a TriSoup triangle. The TriSoup triangle can be constructed around the centroid vertex C, which is defined as the TriSoup vertex V. k The mean of the values. The primary direction can be determined, and then the vertex V can be adjusted by rotating around this direction. k Sort the data and construct the following K TriSoup triangles (list the triplets with vertices as vertices): V1V2C, V2V3C, ..., V K V1C. For example, if a triangle is projected along a principal direction, the principal direction can be selected from three directions that are respectively parallel to the axes of 3D space to increase or maximize the 2D surface of the triangle. By doing so, the principal direction can be slightly perpendicular to the local surface defined by the points of the point cloud belonging to the TriSoup node.
[0110] Figure 8B A detailed example of the TriSoup model is shown. The TriSoup model can be improved by writing the centroid residual values. Centroid residual value C res It can be written into the bitstream. Centroid residual value C resIt can be written to a bitstream, for example, to use C+C++. res Instead of C as the pivot vertex of the triangle, by using C + C res As the pivot vertex of the triangle, vertex C + C res Points can be located closer to the point cloud than the centroid C, which can reduce reconstruction error and thus reduce distortion, but at the cost of writing code C. res The required bit rate will increase slightly.
[0111] Reconstructing the decoded point cloud from a set of TriSoup triangles can be called “voxing”, and can be performed individually for each triangle by ray tracing or rasterization, for example, before removing duplicate voxels from the voxed triangles.
[0112] Figure 9 An example of voxelization is shown. For example, ray 900 can be drawn from integer coordinates P. start Begin firing parallel to one of the three coordinate axes in 3D space. This can be done at the intersection point P with the TriSoup triangle 901 belonging to cube 902. int (If present) Rounding is performed to determine the decoded point. Cube 902 may correspond to a TriSoup node. This intersection can be found (e.g., determined) using the Möller-Trumbore algorithm, for example.
[0113] The existence flag (s) of vertices along the current TriSoup edge can be determined based on, for example, the existence flags and positions of the written codes of the TriSoup edges (existing TriSoup vertices) adjacent to the current TriSoup edge. k ) and (in the presence of signs (s k (indicating the presence of a vertex) position (p) k Entropy coding is performed. Existence flag (s) k ) and location (p k This can be referred to individually or collectively as vertex information. It can also be additionally or alternatively, for example, based on the occupancy of cuboids adjacent to the current TriSoup edge, to indicate the presence of the current TriSoup edge. k ) and (in the presence of signs (s k (indicating the presence of a vertex) position (p) k Entropy writing codes are performed (e.g., indicating the position of a vertex along the edge). Similar to the entropy writing codes of occupancy bits in an occupancy tree, the configuration β of the neighborhood of the current TriSoup edge can be determined. TS (Also known as neighborhood configuration β) TS ), and / or, for example, dynamically reduce it to a reduced configuration β by using TriSoup's dynamic OBUF scheme. TS ' = DRn (β) TS ). Context index LUT[β TS The information can be determined from the OBUF LUT. For example, the context (or probabilistic model) pointed to by the context index can be used to entropy-code at least a portion of the vertex information of the current TriSoup edge.
[0114] The position of the TriSoup vertex along its TriSoup edge (p k (If present) can be binarized. A binary entropy writer can entropy-code at least a portion of the vertex information of the current TriSoup edge. The number of bits N can be set. b To quantize the TriSoup vertex positions (p) along a TriSoup edge of length N. k The edge is uniformly divided into 2 Nb One quantization interval. For example, if the number of bits N b Set to be used for quantizing the TriSoup vertex positions (p) along a TriSoup edge of length N. k ), then the position of the TriSoup vertex (p k ) can be made by N b Units digit (p) k j j=1,…,N b The edge is uniformly divided into 2 Nb Quantization intervals. N b The unit digit and the corresponding existence marker (s) k The bits of ) can be individually coded using the dynamic OBUF scheme. Neighborhood configuration β TS OBUF decrease function DR n And / or context indexes can depend on the nature / characteristics / attributes of the written code points (e.g., presence flags). k ), highest position (p) k 1 ) or second highest position (p k 2 Several dynamic OBUF schemes can be implemented, where each dynamic OBUF scheme is dedicated to a specific bit of vertex information (e.g., presence flag). k ) or position (p) k j )).
[0115] In video compression, performance can be improved by using inter-frame prediction. The bit rate used to compress inter-frames can be one to two orders of magnitude lower than the bit rate of intra-frames that do not use inter-frame prediction. Point cloud data may behave differently than, for example, 2D video data. For point cloud data, 3D geometry can be encoded using 3D point locations. Each point location in the 3D point locations can be associated with an attribute (e.g., color). The geometry and / or attributes may vary between frames. Encoding can be performed, for example, for each frame on different 3D point locations and / or attributes associated with the corresponding 3D point locations. 2D video data can be obtained, for example, by projecting the 3D geometry and / or attributes onto a 2D plane with a fixed geometry (e.g., a camera sensor). For video encoding, attributes can be encoded, but geometry may not be encoded (and may not need to be encoded). Even if, for example, the properties of the expected 2D projection have a higher temporal relevance than the underlying 3D geometry, it can be anticipated that inter-frame prediction between 3D point clouds can provide improved compression capabilities compared to intra-frame prediction within the point cloud (e.g., individual intra-frame prediction). Octrees can benefit from inter-frame prediction and / or geometric compression gains. The general framework for inter-frame prediction of 3D point clouds can be analogous to one of the methods used in video compression.
[0116] Figure 10 Example encoding method 1000 is shown. Figure 10One or more steps can be performed by an encoder (e.g., encoder 114). Encoding method 1000 can use inter-frame prediction between different point cloud frames. The current frame 1001 (e.g., image or point cloud) can be written to relative to a written reference frame 1010 (e.g., image or point cloud). At step 1020, a motion search can be performed from the written reference frame 1010 toward the current frame 1001 to determine a motion vector 1021 representing the motion flow between the two frames 1010 and 1001. In video compression, the motion vector can be a 2-component (or 2D) vector representing the motion from a reference pixel block to the current pixel block. In point cloud compression, the motion vector can be a 3-component (or 3D) vector representing the motion from a reference 3D point set (e.g., in a reference point cloud) to the current 3D point set (e.g., in the current point cloud). At step 1025, the motion vector 1021 can be entropy-coded into bitstream 1050. At step 1030, motion compensation can be performed on reference frame 1010 to determine motion-compensated frame 1031. Motion compensation may involve moving pixels of the reference image based on 2D motion vectors, and / or moving points of the reference point cloud based on 3D motion vectors. The determined motion-compensated frame may be "closer" to the current frame than the reference frame; for example, the chromatic difference (or dot distance) between motion-compensated frame 1031 and the current frame 1001 may be on average smaller than the chromatic difference (or dot distance) between reference frame 1010 and the current frame 1001. At step 1040, inter-frame prediction can be performed to determine inter-frame residual 1041. At step 1045, the inter-frame residual 1040 can be entropy-coded into bitstream 1050. The inter-frame residual may carry more compressible information than the current frame, which may or may not have undergone intra-frame prediction. Therefore, the entropy write coding performed at step 1045 may be more efficient in determining a bitstream 1050 with a reduced size compared to the bitstream determined by writing coding to the current frame 1001 which does not benefit from inter-frame prediction.
[0117] In video coding, inter-frame residuals can be constructed as the color, pixel-per-pixel difference between the current pixel block belonging to the current frame (e.g., an image) and the co-located compensated pixel block belonging to a motion-compensated frame (e.g., an image). Inter-frame residuals can be color difference arrays, which typically have small values and can be effectively compressed.
[0118] In point cloud compression, there may be no "difference" between two sets of points because there may not be a one-to-one mapping between them. The concept of inter-frame residuals may not be directly applicable (e.g., generalization) relative to point clouds. To predict octrees representing point clouds, the concept of inter-frame residuals can be replaced with conditional entropy write-code, where conditional information for performing conditional entropy write-code can be constructed based on motion-compensated point clouds. This can be extended to the framework of dynamic OBUF.
[0119] As described herein, the occupancy of the current volume (e.g., the current volume associated with the current node of the octree) can be provided by a number of occupancy bits, for example, 8 occupancy bits. The current occupancy bits of the octree can be written by an entropy writer selected by the output of a dynamic OBUF LUT using a writer index with a neighborhood configuration β as input. The neighborhood configuration β can be constructed based on the written occupancy bits. The written occupancy bits can be associated with neighboring volumes (e.g., with the neighboring nodes of the current node). The construction of the neighborhood configuration β can be extended using inter-frame information. Inter-frame predictor occupancy bits can be indicated by the current occupancy bits (e.g., defined for the current occupancy bits) as bits representing the presence of at least one point in the motion-compensated point cloud within the current volume. For example, if this motion compensation is valid, there may be a strong correlation between the current occupancy bits and the inter-frame predictor occupancy bits. This could be because the current point cloud and the motion-compensated point cloud may be close to each other. Using the inter-frame predictor occupancy bits as the bits of the neighborhood configuration β can result in better compression performance of the octree (e.g., by dividing the size of the octree bitstream by a factor of two).
[0120] The motion field between octrees may include 3D motion vectors associated with a 3D prediction unit (PU). A PU may have a volume that may include at least a portion of one or more volumes (e.g., cuboids) associated with nodes of the octree. Motion compensation for each volume (e.g., each cuboid's cuboid) may be performed based on the 3D motion vectors to determine a motion-compensated point cloud in one or more current volumes. Inter-frame predictor occupancy bits may be determined, for example, based on the presence of at least one point in this motion-compensated point cloud.
[0121] The TriSoup scheme can benefit from motion-compensated frames determined, for example, during octree coding prior to TriSoup coding. Predictors for the presence and / or location of TriSoup vertices can be determined based on the motion-compensated point cloud. For example, a predictor can be determined based on the intersection of edges between the compensated point cloud and TriSoup nodes. Predictors for centroid residual values can also be determined.
[0122] Entropy coding of TriSoup vertex and / or centroid residual values can be performed, for example, by using these inter-frame predictors. The inter-frame predictor can constitute the context information β of a dynamic OBUF instance that codes out TriSoup syntax elements. inter Part of the input. Alternatively or additionally, the context can be selected based on an inter-frame predictor. The selected context can be used by an entropy writer (e.g., CABAC) to determine the probability of arithmetic entropy writing codes for TriSoup syntax elements.
[0123] For example, after writing to the underlying geometry (e.g., the location of points in 3D space), the attributes associated with the points in the point cloud can be written to. If the geometry writing is lossless (e.g., using an octree scheme), the encoder can directly access the attribute values associated with the written-coded points. If the geometry writing is lossy (e.g., using a TriSoup scheme), the written-coded geometry can differ from the original geometry. The original attributes can be mapped from the original geometry to the written-coded geometry by the encoder to determine the mapped attributes on the written-coded geometry.
[0124] As described in this paper, attributes can indicate the nature of a point's visual appearance (e.g., texture, color, material, transparency, reflectivity, timestamp, or velocity). For color attributes, the attribute mapping performed by the encoder can be referred to as a recoloring process. This is likely because the color of the original geometry can be used to color (e.g., recolor) the writable geometry (e.g., reconstructed geometry).
[0125] The write-coded geometry associated with the mapping attributes of the write-coded geometry may include a point cloud representing the original point cloud in the geometry and attributes. There may be more than one (e.g., two) attribute writing schemes that can be used and / or selected for writing the attributes associated with the write-coded geometry. Attribute writing schemes may include, for example, a prediction-lift transformation (“pred-lift”) scheme and / or a region adaptive hierarchical transformation (“RAHT”) scheme. Attribute writing schemes may be used, for example, in G-PCC.
[0126] Predictive boosting schemes can begin by decomposing the writable geometry into levels of detail (also known as LoD). For a set (S) of points (e.g., all points) of the writable geometry, this set can be decomposed into disjoint subsets Si. i , making L levels of detail can be defined as the point cloud geometry tower.
[0127]
[0128] For example, if the set can be decomposed into disjoint subsets S i So that Then the set of points It can be the first (e.g., the coarsest) level of detail, and a set of points. It can be the Lth (e.g., the finest) level of detail.
[0129] Attribute a j Points s in the set S of all points whose geometry can be written can be used. j Related. Considering a subset a of the attributes.0 ... a L-1 subset a i Attribute a in i j Can be with subset S i point s in i j Related. The i-th level of detail ( It can have a subset a 0 ... a i-1 The cascading related attributes in the set. A set 'a' of attributes (e.g., all attributes) can be partitioned into subsets 'a'. 0 ... a L-1 .
[0130] Figure 11 , 12 Figures 13 and 14 illustrate examples of writing (e.g., encoding or decoding) point cloud attributes. This writing can be based on an intra-frame transform scheme, such as a predictive transform scheme or a predictive boosting scheme. The predictive transform scheme can be a variant of the predictive boosting scheme without an update operation. The intra-frame transform scheme can be an example of wavelet transform, which can transform (e.g., convert) attribute values into wavelet coefficients that can be more effectively compressed compared to the original attribute values. The wavelet coefficients can represent the values of residual attributes, which can be smaller and more effectively compressed compared to the original attribute values. The residual attributes can be referred to as and / or include wavelet coefficients (or transformed / transformed coefficients). The wavelet coefficients can be generated by applying a predictive transform scheme and / or a predictive boosting scheme.
[0131] As described herein, prediction transformation schemes and / or prediction enhancement schemes can operate using predictions within and / or between attribute levels of detail. At the encoder, attributes at higher (e.g., finer) levels of detail can be predicted based on attributes at lower (e.g., coarser) levels of detail. For example, each level of detail starting from the highest level can be predicted sequentially based on lower levels of detail. The decoder can perform the inverse operation, such that attributes at lower levels of detail can be predicted and / or reconstructed, for example, based on residual attributes at higher levels of detail. For example, each level of detail starting from the lowest level can be predicted and reconstructed sequentially based on higher levels of detail. Although... Figure 11-14 The example shown illustrates three levels of detail (LoD), but it should be understood that different numbers of LoDs can be used, for example, by extending (e.g., iteratively) the decomposition scheme.
[0132] Figure 11 An example method for encoding point cloud attributes is shown. The encoding can be based on a predictive transformation scheme. Figure 11 One or more steps can be accomplished by an encoder (e.g., Figure 1The encoder 114 in the middle is executed.
[0133] Predictions can be made, for example, within and between one or more (e.g., three) levels of detail, starting from the highest (e.g., the first prediction) level of detail (e.g., associated with the highest frequency, such as res a). 2 ) down to the lowest (e.g., third) level of detail (e.g., associated with the lowest frequency, such as a 0 The encoder writes the attribute set 'a'. At step 1110, the encoder can split the attribute set 'a' into a subset a. 2 The first set of attributes in the property and includes two subsets a 0 and a 1 The second set of attributes in the set. At step 1120, the encoder can obtain the second set of attributes (a 0 and a 1 The attributes in ) determine the first attribute set (a 2 The predicted values of the attributes in ) are then determined. At step 1130, the encoder can determine the first residual value 'res a'. 2 The encoder can, for example, obtain the first set of attributes (a). 2 The first residual value 'resa' is determined by subtracting the predicted value from the attribute in the property. 2 At step 1170, the encoder can output the first residual value 'res a'. 2 'Encode into the bitstream. The operations at steps 1120-1130 can be iteratively applied (e.g., applied to) each consecutive lower (e.g., coarser) LoD.
[0134] At step 1140, the encoder can transmit the second attribute set (a 0 and a 1 The set is split into a third attribute set and a fourth attribute set. The third attribute set may include a subset a. 1 The attributes in. The fourth attribute set can include a subset a. 0 The attributes in. At step 1150, the encoder can obtain the attributes from the fourth attribute set (a 0 The attributes in ) determine the third attribute set (a 1 The predicted values of the attributes in ) are then used. At step 1160, the encoder can determine the second residual value 'res a'. 1 The encoder can, for example, be obtained from a third set (a 1 The second residual value 'res a' is determined by subtracting the predicted value from the attribute in the property. 1 At step 1170, the encoder can output the second residual value 'res a'. 1 'Encoded into the bitstream. The encoder can convert the fourth attribute set (a 0The attributes in the stream are encoded into the bitstream.
[0135] The bitstream may include a representation of the first residual 'res a' 2 '、Second residual'res a 1 'and / or subset a 0 (For example, the data of attributes in the fourth attribute set). At step 1170, the residual values can be entropy-coded. The encoder can then encode the subset a. 0 The attributes in the code are encoded into the bitstream. Alternatively, the encoder can encode a subset a of the code to be written. 0 The current attribute a in 0 j Perform intra-frame prediction. The encoder can, for example, be based on a subset a. 0 The attribute of the written code in the data is used to perform the operation on subset a. 0 The current attribute a in 0 j Intra-frame prediction. This can improve subset a. 0 The compression efficiency of attributes in the data.
[0136] For example, if lossy attribute write codes are allowed, the encoder can quantize a subset a. 0 The attribute in, the first residual value 'res a 2 'and / or the second residual'res a 1 The encoder can convert subset a 0 The attribute (or subset a) in 0 Quantified attributes in the data), first residual value 'res a 2 '(or the first quantized residual value'res a) 2 ') and / or the second residual value 'resa 1 '(or the quantified second residual value'res a) 1 The encoding (e.g., entropy encoding) is added to the bitstream.
[0137] Figure 12 An example method for decoding point cloud properties is shown. Decoding can be based on a predictive transformation scheme. Figure 12 One or more steps can be accomplished by a decoder (e.g., Figure 1 The decoder (120) in the code executes. It can decode the attribute set 'a'. It can also encode the attribute set 'a', for example, as described in this article... Figure 11 As described. Decoding can use predictions between one or more (e.g., three) levels of detail, from the lowest (e.g., the third) level of detail to the highest (e.g., the first) level of detail (e.g., as described in this paper regarding...). Figure 11 The encoder described is in reverse order.
[0138] At step 1210, the decoder can obtain the first residual value 'res a' from the bitstream pair. 2 '、Second residual'res a 1 'and / or the fourth set of attributes (a 0 The decoder can decode from the attributes in the fourth attribute set (a). The decoder can use (e.g., apply) dequantization (e.g., for lossy compression). At step 1220, the decoder can decode from the fourth attribute set (a 0 The decoded attributes determine the third attribute set (a) 1 The predicted value of China's attributes. The decoder can, for example, be used with information about... Figure 11 The predicted value is determined in a similar manner to that described in step 1150. At step 1230, the decoder can determine the predicted value, for example, by adding the predicted value to the decoded first residual value 'res a'. 1 'To determine the third attribute set (a) 1 The decoded attributes in ) . The operation at steps 1120-1230 can be iteratively applied (e.g., applied to) each successively higher (e.g., finer) LoD.
[0139] At step 1240, the decoder can determine the second attribute set (a 0 and a 1 The decoder can, for example, merge a third set of attributes (a 1 ) and the fourth attribute set (a 0 ) to determine the second attribute set (a 0 and a 1 Step 1240 can be the inverse operation of step 1140. At step 1250, the decoder can obtain the second attribute set (a 0 and a 1 The attributes in ) determine the first attribute set (a 2 The decoder can determine the predicted values of the attributes in the () and the decoder can determine the predicted values in a manner similar to that described with respect to step 1120. At step 1260, the decoder can, for example, by adding the predicted values to the decoded second residual value 'res a 2 'Determine the first set of attributes (a) 2 The decoded attributes in ) . At step 1270, the decoder can, for example, merge the first attribute set (a 2 ) and the second attribute set (a 0 and a 1 This determines the set of decoded attributes 'a' of the write-coded geometry S (e.g., the entire write-coded geometry S). Step 1270 can be the inverse operation of step 1110.
[0140] Figure 13An example method for encoding point cloud attributes is shown. The encoding can be based on a prediction boosting transformation scheme with prediction and updates. Figure 13 One or more steps can be accomplished by an encoder (e.g., Figure 1 The encoder 114 in the code performs the operation. Predictions can be made using one or more (e.g., three) levels of detail, starting from the highest (e.g., first) level of detail (e.g., associated with the highest frequency, such as resembling a). 2 down to the lowest (e.g., third) level of detail (e.g., associated with the lowest frequency, such as up up a) 0 Encode the attribute set 'a'.
[0141] At step 1310, the encoder can split the attribute set 'a' into a subset a. 2 The first set of attributes in the property and includes two subsets a 0 and a 1 The second set of attributes in the property. Step 1310 can be similar to the section on the second set of attributes in this paper. Figure 11 The process is performed as described in step 1110. At step 1320, the encoder can access the second attribute set (a... 0 and a 1 The attributes in ) determine the first attribute set (a 2 The predicted values of the attributes in the first attribute set (a). Step 1320 can be performed similarly to that described herein with respect to step 1120. At step 1330, the encoder can, for example, obtain the predicted values of the attributes from the first attribute set (a 2 The first residual value 'res a' is determined by subtracting the predicted value from the attribute in the property. 2 Step 1320 can be performed similarly to that described herein with respect to step 1130. At step 1370, the encoder can convert the first residual value 'res a' 2 'Encode into the bitstream. Step 1370 can be performed similarly to the description in this document regarding step 1170.'
[0142] At step 1375, the encoder can obtain the first residual value 'res a' 2 'Determine the updated attribute value. The encoder can determine the updated attribute value, for example, based on a first residual value. The updated attribute value can be determined as the first residual value multiplied by a scaling factor (e.g., ½, ¼, 1 / 8), which can be predetermined or signaled in the bit stream.' At step 1380, the encoder can add the updated attribute value to a second attribute set (a 0 and a 1 The attribute values of ) are used to update the second attribute set (a) 0 and a 1 The attribute value 'up a' 0 'and'up a1 The operations at steps 1320, 1330, 1375, and 1380 can be iteratively applied (e.g., applied to) each successively lower (e.g., coarser) LoD.
[0143] At step 1340, the encoder can transmit the second attribute set (a 0 and a 1 The set is split into a third attribute set and a fourth attribute set. The third attribute set may include a subset a. 1 The updated attribute value 'up a 1 The fourth attribute set can include a subset a. 0 The updated attribute value 'up a 0 At step 1350, the encoder can obtain the fourth attribute set (a 0 The updated attribute value 'up a' 0 'Determine the set of third attributes (a)' 1 The updated attribute value 'up a' 1 The predicted value of '. At step 1360, the encoder can, for example, obtain the predicted value from the third attribute set (a 1 The updated attribute value 'up a' 1 'Subtract the predicted value from the middle to determine the third residual' 1 At step 1370, the encoder can update the third residual value. 1 'Encode into the bitstream. At step 1385, the encoder can retrieve the third residual value from the bitstream.' 1 'Determine the updated attribute value. At step 1390, the encoder can, for example, add the updated attribute value to the fourth attribute set (a 0 The updated attribute value 'up a' 0 'Determine the fourth attribute set (a) 0 The further updated attribute value 'up up a' 0 At step 1370, the encoder can convert the fourth attribute set (a 0 The further updated attribute value 'up up a' 0 'Encoded into bitstream.'
[0144] The bit stream may include data representing transformed attributes (e.g., representing the first residual value 'res a'). 2 '、Third residual'res up a 1 'and / or the further updated attribute values of the fourth attribute set' up up a 0 The encoder can further update the attribute values of the fourth attribute set.0 'Encoding is in the bitstream. The encoder can, for example, perform further updates to the current attribute value of the code to be written based on the attributes of the written codes in the fourth attribute set.'up upa 0 j Intra-frame prediction. This can improve the further updated attribute values of the fourth attribute set. 0 Compression efficiency.
[0145] For example, if lossy attribute write coding is allowed, the encoder can quantize the further updated attribute values of the fourth attribute set. 0 '、First residual'res a 2 'and / or third residual' res up a 1 The encoder can further update the attribute values of the fourth attribute set. 0 '、First residual value'res a 2 '(or the first quantized residual value'res a) 2 ') and / or the third residual value'res up a 1 '(or the quantified third residual value'res up a 1 ') Entropy encoding is incorporated into the bitstream.
[0146] Figure 14 An example method for decoding point cloud attributes is shown. Decoding can be based on a prediction-enhancing transformation scheme with predictions and updates. Figure 14 One or more steps can be accomplished by a decoder (e.g., Figure 1 The decoder (120) in the document is executed. This can be done as described in this article. Figure 13 The description describes encoding a set of attributes 'a'. The set of attributes 'a' (e.g., the encoded set of attributes 'a') can be decoded. Decoding can use predictions between one or more (e.g., three) levels of detail, from the lowest (e.g., the third) level of detail to the highest (e.g., the first) level of detail (as described in this paper regarding...). Figure 13 The encoder described is in reverse order.
[0147] At step 1410, the decoder can obtain the first residual 'res a' from the bitstream pair. 2 '、Third residual value' res up a 1 'and / or the fourth set of attributes (a 0 The further updated attribute value 'up up a' 0 Decoding is then performed. The decoder can use (e.g., by application) optional dequantization for lossy compression. At step 1475, the decoder can obtain the decoded third residual value from the decrypted third residual value. 1'Determine the updated attribute value. Step 1475 can be performed in a manner similar to that described herein with respect to step 1385. The decoder can determine the updated attribute value based on the third residual value. The updated attribute value can be determined as the third residual value multiplied by a scaling factor (e.g., ½, ¼, 1 / 8), which can be predetermined or signaled in the bit stream.'
[0148] At step 1480, the decoder can, for example, obtain the fourth attribute set (a 0 The decoded and further updated attribute value 'up up a 0 Subtract the updated attribute value from the middle to determine the fourth attribute set (a) 0 The updated attribute value 'upa' 0 At step 1420, the decoder can obtain the fourth attribute set (a 0 The updated attribute value 'up a' 0 'Determine the set of third attributes (a)' 1 The updated attribute value 'up a' 1 The predicted value of '. Step 1420 can be performed in a similar manner to that described herein with respect to step 1350. At step 1430, the decoder can, for example, add the predicted value to the decoded third residual value 'res up a 1 'To determine the third attribute set (a) 1 The updated attribute value 'up a' 1 At step 1440, the decoder can, for example, merge the third attribute set (a 1 ) and the fourth attribute set (a 0 ) to determine the second attribute set (a 0 and a 1 Step 1440 can be the inverse operation of step 1340. The second attribute set (a) 0 and a 1 This can include the updated attribute value 'up a'. 0 'and updated attribute values'up a 1 At step 1485, the decoder can obtain the first residual value 'resa' from the decoded value. 2 'Determine the updated attribute value. Step 1485 can be performed in a similar manner to that described herein with respect to step 1375. The operations at steps 1420, 1430, 1440, 1475, and 1480 can be iteratively applied (e.g., applied to) each successively higher (e.g., finer) LoD.
[0149] At step 1490, the decoder can, for example, obtain the second attribute set (a 0 and a 1 The updated attribute value 'up a' 0 'Neutralize the updated attribute values from the second attribute set'up a 1 Subtract the updated attribute value from the middle to determine the second attribute set (a) 0 and a 1 The attribute value 'a' 0 'and attribute value'a 1 At step 1450, the decoder can obtain the second attribute set (a 0 and a 1 The attributes in ) determine the first attribute set (a 2 The predicted values of the attributes in ) are then used. Step 1450 can be performed in a similar manner to that described herein with respect to step 1320. At step 1460, the decoder can, for example, add the predicted values to the decoded second residual value 'res a'. 2 'Determine the first set of attributes (a) 2 The decoded attributes in ) . At step 1470, the decoder can, for example, merge the first attribute set (a 2 ) and the second attribute set (a 0 and a 1 This determines the set of decoded attributes 'a' of the decoded geometry S (e.g., the entire decoded geometry). Step 1470 can be the inverse operation of step 1310.
[0150] Predictive lifting schemes can be similar to lifting schemes used for (e.g., applied to) wavelet writing. Predictive lifting schemes may include update steps not present in the predictive transform scheme (e.g., in addition to the prediction step in the predictive transform scheme). Update steps can provide better compression performance (e.g., in conjunction with the prediction step). This allows energy to be compressed at the lowest level of detail, which can reduce distortion in lossy writing.
[0151] The RAHT scheme can be used to encode attributes. The RAHT scheme can be used iteratively based on a two-point transform. In point cloud attribute encoding, the two-point RAHT transform can be used (e.g., applied to) two attribute sets (A1 and A2). Each of A1 and A2 can have w1 and w2 numbers of attributes, respectively. Each of A1 and A2 can have a corresponding correlation coefficient c. A1 and cA2。 c A1 and c A2 Each value in the set represents the sum of the attribute values in the corresponding set divided by the square root of the number of attributes.
[0152] . ( (*)
[0153] A two-point RAHT transformation can depend on weights w1 and w2. The two-point RAHT transformation can be defined by the following 2×2 matrix.
[0154] .
[0155] Two new coefficients, DC and AC, can be determined, for example, when used (e.g., applied to) two coefficients c. A1 and c A2 In this case.
[0156]
[0157] As described in this article, the above properties (*) regarding coefficients hold for DC coefficients.
[0158]
[0159] A two-point RAHT transformation can be iteratively applied (e.g., applied to) the DC coefficients. This can be called the RAHT iterative method. For example, once determined, the AC coefficients may not undergo further transformations. At the start of the RAHT iterative method, there may be an initial set A of properties as many as the points present in the written code geometry S. i Each initial attribute set A i It can contain one attribute (w) i =1). Coefficient c Ai It can be equal to the value of a property, and / or can satisfy property (*). By induction, property (*) can hold for, for example, subsequent DC coefficients (e.g., all subsequent DC coefficients) determined after iterative application of the two-point RAHT transformation.
[0160] At a specific stage of the RAHT iterative method, the determined coefficients can be the union of the set of DC coefficients and the set of AC coefficients satisfying property (*). The RAHT iterative method can continue until all DC coefficients are exhausted, and may leave only one DC coefficient. A DC coefficient can be equal to C. A , where A can be a set of attributes (e.g., all attributes) associated with the writable code geometry S (e.g., the complete writable code geometry S). The RAHT iterative method can follow the order in the DC coefficient pairs.
[0161] The two-point inverse Raht transform can be defined by the following 2×2 matrix.
[0162]
[0163] The two-point inverse RAHT transform can be used (e.g., applied to) DC and AC coefficients to recover the two coefficients c. A1and cA2。
[0164]
[0165] The inverse iterative RAHT method obtains the DC and AC coefficients by applying the inverse two-point RAHT in reverse order (e.g., applying it to) the DC and AC coefficients, relative to obtaining them through the iterative RAHT method. At the end of the inverse iterative RAHT method, the initial property set A can be obtained. i The associated coefficient c Ai These coefficients c Ai It can be equal to the initial set A i The value of an associated attribute.
[0166] For lossy RAHT compression of attributes, the coefficients can be further compressed, for example, before the coefficients are encoded in the bitstream, based on quantization applied to (e.g., applied to) the DC and AC coefficients. The decoder can, for example, use (e.g., apply) dequantization after decoding the quantized DC and AC coefficients from the bitstream.
[0167] The RAHT iterative method can, for example, follow an octree in a specific iterative order within G-PCC. One or more (e.g., up to eight) DC coefficients associated with one or more (e.g., up to eight) occupied child nodes of a parent node in the octree can undergo a cascade of two-point RAHT transformations until one DC coefficient remains, along with the remaining (e.g., up to seven) AC coefficients. This DC coefficient can be pushed at the parent node level. The method can be repeated, for example, at higher octree depths until the root node is reached.
[0168] Figure 15 An example RAHT transformation is shown. The RAHT transformation can be applied to the child nodes of an octree parent node along three consecutive directions. The parent node 1500 can have multiple (e.g., five) occupied child nodes, each with a corresponding association coefficient c. i and weight w i The first RAHT transformation 1510 can be performed along the first direction 1511. If there are two adjacent occupied child nodes 1513 along the first direction, these two adjacent occupied child nodes 1513 can undergo a two-point RAHT transformation. New DC coefficients 1514 and AC coefficients 1515 can be determined and pushed into the AC coefficient set 1550. For example, if there is only one occupied child node 1516 along the first direction, the node can remain as is, and its DC coefficients can remain 1517. In this example, the child node can be folded along the first direction to determine a new set of nodes 1519 (e.g., a set of three nodes), and associated new DC coefficients can be determined.
[0169] A second RAHT transformation 1520 can be performed, for example, after the first RAHT transformation. The second RAHT transformation can be performed along a second direction 1521. The second RAHT transformation can be performed similarly to the first RAHT transformation. Child node 1522 can be determined. Child node 1522 may have been folded along the first two directions 1511 and 1521. AC coefficient 1523 can be pushed to AC coefficient set 1550. A third RAHT transformation 1530 can be performed, for example, after the second RAHT transformation. The third RAHT transformation can be performed along a third direction 1531. The third RAHT transformation can be performed similarly to the first and / or second RAHT transformation. Child node 1532 can be determined (e.g., uniquely). Child node 1532 can be generated by folding along all three directions. AC coefficient 1533 can be pushed to AC coefficient set 1550. Child node 1532 may have associated DC coefficients that are pushed to the parent node (e.g., as...). Figure 16 (as shown in the image).
[0170] Figure 16 An example RAHT transformation is illustrated. The RAHT transformation can be used (e.g., applied to) octree nodes at depth 'd' (e.g., all octree nodes) to determine the DC coefficients and AC coefficients at depth d-1. Occupied nodes 1600 of the octree can be at depth d. Occupied node 1600 can undergo the RAHT transformation along three directions. The DC coefficients of each node can be pushed to the corresponding occupied parent node 1610 belonging to the octree at depth d-1. The three DC coefficients of child node 1601 can, for example, undergo the RAHT transformation along the three directions. A unique DC coefficient associated with parent node 1611 can be determined, and two AC coefficients 1621 can be pushed to AC coefficient set 1620. By performing this method on (e.g., all) occupied nodes 1600 of the octree at depth d, the DC coefficients associated with the occupied nodes of the octree at depth d can be transformed into the DC coefficients associated with the occupied node 1610 of the octree at depth d-1 and the AC coefficient set 1620.
[0171] This bottom-up approach can, for example, be repeated depth-by-depth until a minimum depth (root node) is reached. The result of the RAHT transformation on an octree (e.g., a complete octree) can be a set of coefficients including unique DC coefficients and a set (many) of AC coefficients. The RAHT transformation method can start from the highest depth of the unique point (voxel) of the occupied child node corresponding to the write-coded point cloud S. The unique point can be associated with a unique attribute in the attribute set 'a'. The DC coefficient at the highest depth can be set to the value of the unique attribute associated with each occupied node. The weight 'w' can be set to 1.
[0172] The inverse RAHT method on an octree can be a top-down approach, descending from the root node to the final depth comprised of leaf nodes, each containing a point in the point cloud and an associated attribute. For example, by applying the inverse two-point RAHT transformation to (e.g., to) the DC coefficients of each occupied node at depth d-1 of the octree and using the associated AC coefficients from the AC coefficient set 1620, the DC coefficients of occupied node 1610 at depth d-1 of the octree can be inversely transformed to the DC coefficients of occupied node 1600 at depth d of the octree. The inverse two-point RAHT transformation can be applied in reverse order along three directions to reverse the relationships described herein. Figure 15 The described node transformation allows for the acquisition of DC coefficients for leaf nodes, with each DC coefficient corresponding to a unique attribute associated with each leaf node. Similar to the geometric write-coding of point clouds, the write-coding of attributes associated with points in the current point cloud can benefit from inter-frame prediction using a motion-compensated point cloud. A motion-compensated point cloud can inherit attributes from a motion-compensated reference point cloud. For example, if motion occurs, a point can retain its associated attributes. Motion-compensated attributes (e.g., attributes associated with points in the motion-compensated point cloud) can be used. This improves the compression of attributes of the write-coded geometry of the current point cloud.
[0173] Inter-frame prediction upscaling can use motion-compensated attributes or residual attributes based on the difference between the attribute and the motion-compensated attribute. Using residual attributes can increase compression. Inter-frame prediction upscaling can be used in prediction upscaling, for example, by inserting it into prediction steps 1120, 1150, 1220, 1250, 1320, 1350, 1420, and / or 1450. This can be achieved, for example, through lower levels of detail. subset a 0 ... a i-1 The attribute (or residual attribute) is used to predict the point set S. i The attribute (or residual attribute) a i Alternatively, it can be achieved by comparing with the motion-compensated point cloud S. inter The set a associated with the points inter The motion-compensated properties (or residual properties) are used to predict the point set S. i The attribute (or residual attribute) a i For example, it can be based on subset a. 0 ... a i-1 and enhanced lower level of detail set a inter The prediction step is performed using the properties (or residual properties) of the data.
[0174] The encoder and / or decoder can obtain data from the motion-compensated point cloud S. inter The set a associated with the points inter The set of fourth attributes (or residual attributes) is determined by the motion-compensated attributes (or residual attributes) (coarsest level of detail). subset a 0 The predicted values of the attributes (or residual attributes) in the fourth attribute (or residual attribute) set. The encoder and / or decoder can subtract the predicted values from the attributes (or residual attributes) in the fourth attribute (or residual attribute) set to determine the residual value 'res a'. 0 The encoder can output the residual value 'res a'. 0 The attribute (or residual of the residual attribute) is encoded into the bitstream instead of the attribute (or residual attribute) in the fourth attribute set.
[0175] Inter-frame RAHT schemes can use inter-frame prediction to predict the values of DC and AC coefficients determined by the RAHT iterative method. For example, maintaining the current point cloud S for the code to be written. coded And the motion-compensated point cloud S inter The shared property of an octree structure may be beneficial, as the generation of DC and AC coefficients follows an octree. A common bounding box enclosing the two point clouds can be determined. For both point clouds, octree partitioning can be performed from the root node associated with the common bounding box. For example, if the point clouds are not equal, this may result in the two octree partitions potentially being different. The two octrees can have a common subtree starting from the root node. On the common subtree, the topology of the occupied nodes can be the same, and / or a common set of DC and AC coefficients can be determined for both point clouds. The coefficients can be determined from the motion-compensated point cloud S. inter The DC and AC coefficients, determined by the attributes, are used to predict nodes associated with common subtrees and from the current point cloud S. coded The attributes determine a subset of the DC and AC coefficients. The encoder and / or decoder can, for example, be derived from nodes associated with a common subtree and from the current point cloud S. coded Subtract the DC and AC coefficients determined by the motion-compensated point cloud S from the attribute-determined values. inter The DC and AC coefficients, determined by the attributes, determine the coefficient residual values. The encoder can encode the coefficient residual values into the bitstream, and / or can choose not to associate them with nodes in the common subtree and from the current point cloud S. coded The DC and AC coefficients, whose attributes are determined, are encoded into the bitstream.
[0176] The DC and AC coefficients that are not associated with nodes in the common subtree may not be predicted, and / or the DC and AC coefficients may be directly written in a manner similar to that performed without inter-frame prediction. Alternatively, the current point cloud S may be determined at a certain depth. codedThe predicted DC coefficients (e.g., without predicting AC coefficients at the same depth). This can be based on the assumption that the current point cloud S... coded The octree and the motion-compensated point cloud S inter The octrees all have the same node occupancy rate at this depth. This can be seen from the motion-compensated point cloud S. inter The predicted DC coefficients are determined by the corresponding co-located DC coefficients. This can be achieved, for example, by using the current point cloud S. coded The DC residual value is determined by subtracting the predicted DC coefficients from the DC coefficients. The RAHT transform can, for example, begin by increasing the DC residual value in the octree from the DC coefficients of the current point cloud after the predicted DC coefficients have been determined.
[0177] A RAHT scheme that does not use information from a reference frame different from the current frame can be called an intra-frame RAHT scheme. Intra-frame prediction can be performed between the DC and AC coefficients of an intra-frame RAHT scheme. Intra-frame inter-depth prediction within the current frame may have been integrated into a RAHT scheme such as GGC. The inter-frame depth prediction mechanism can, for example, predict the DC coefficients associated with nodes at depth d in an octree by interpolating the DC coefficients associated with nodes at a lower depth d-1 in the octree.
[0178] To compress colored, dynamically dense point clouds (e.g., in an AR / VR context), attribute-related data can constitute a large portion of the bitstream. Effectively compressing attribute-related data can improve the storage and / or transmission of such point clouds. Geometry can be lossily compressed (e.g., using the TriSoup scheme), and / or color can be lossily compressed (e.g., using the RAHT transform or prediction boosting scheme). For point cloud compression, the RAHT transform can be chosen, for example, based on its better performance compared to prediction boosting schemes.
[0179] In at least some systems, inter-frame RAHT may have already achieved compression gain compared to intra-frame RAHT. Inter-frame RAHT can be based on coefficient predictions from motion-compensated colored point clouds. In regions with complex motion, the gain from inter-frame RAHT may (e.g., significantly) decrease. For example, if (motion) compensation cannot (e.g., is unlikely) provide adequate compression for the current point cloud S. coded And the motion-compensated point cloud S inter If the two octrees of a partition provide a sufficiently large common portion, descent may occur. The number of DC and AC coefficients that can be predicted inter-frame may be limited, and using inter-frame RAHT may be ineffective. In at least some systems, inter-frame RAHT may be ineffective for the current point cloud S. coded With motion-compensated point cloud S interThe geometric differences between them are sensitive, and the property correlations between the two point clouds may be difficult to exploit.
[0180] As described herein, projection-based compression can be used. Projection-based compression can be used to compress the properties of point clouds. An encoder can determine the properties of the reconstructed geometry of a point cloud frame. This determination can be based on the properties of the geometry of the point cloud frame. The encoder can, for example, map the properties of the geometry of the point cloud frame to the reconstructed geometry in lossy compression. The properties of the reconstructed geometry can be determined based on the mapping. For example, in lossless compression, the properties of the geometry can be the same as the properties of the reconstructed geometry. The encoder can determine an attribute predictor for the properties of the reconstructed geometry. The attribute predictor can be determined, for example, based on projecting the properties of a reference point cloud frame for the properties onto the reconstructed geometry. The properties of the reconstructed geometry can be encoded based on the attribute predictor. The properties can be encoded as attribute residuals. The attribute residuals, indicating the difference between the properties of the reconstructed geometry and the attribute predictor, can be encoded into the bitstream associated with the point cloud frame. The attribute residuals can, for example, be used by a decoder to determine the properties of the reconstructed geometry of the point cloud frame.
[0181] The decoder can perform inverse operations to reconstruct (e.g., decode) the attributes of a point cloud frame. The decoder can determine the reconstructed geometry of the point cloud frame, for example, by decoding the geometric information of the point cloud frame. The reconstructed geometry determined by the decoder can be the same as the reconstructed geometry determined by the encoder. An attribute predictor for the attributes of the reconstructed geometry can be determined by the decoder, for example. The decoder can determine the attribute predictor in a manner similar to that of the encoder. The decoder can determine the attribute predictor, for example, based on projecting the attributes of a reference point cloud frame for the attribute onto the reconstructed geometry. The attribute predictor can be determined by the encoder and decoder in an iterative manner (e.g., independently and / or identically). The attribute predictor may not be indicated in the bitstream (e.g., not encoded into the bitstream or not signaled in the bitstream). The decoder can decode the attributes of the reconstructed geometry, for example, based on the attribute predictor. The decoder can receive / determine (e.g., decoded) attribute residuals from the bitstream. For example, attributes can be determined based on adding the attribute predictor to the attribute residuals.
[0182] The reference point cloud frame for an attribute can be a coded reference frame. A coded reference frame can be a reference frame used to encode the geometric information of the point cloud frame. Alternatively, a coded reference frame can be a different reference frame from the one used to encode the geometry of the point cloud frame. The reference point cloud frame for an attribute can be a motion-compensated point cloud frame. A motion-compensated point cloud frame can be, for example, based on a coded reference frame (e.g., as described herein). Figure 10The attribute residuals can be determined (e.g., generated) using the described method. Intra-frame transform schemes (e.g., predictive lift transform or (intra-frame) RAHT schemes) can be used to further compress (e.g., encode and / or decode) the attribute residuals. Projected attributes and writable attributes can belong to the same geometry of the point cloud frame. Predicting writable attributes based on projected attributes may be more efficient than (e.g., directly) predicting from motion-compensated point cloud frames, for example, because the geometric difference between the reconstructed geometry and the motion-compensated point cloud geometry may have been reduced or eliminated.
[0183] As described herein, the decoded geometry of a point cloud frame can correspond to the reconstructed geometry of the point cloud frame. An encoder can determine the reconstructed geometry, for example, by encoding the geometry of the point cloud frame and decoding the encoded geometry. The reconstructed geometry at the encoder can be the same as the geometry decoded at the decoder. The decoder can reconstruct the same reconstructed geometry as the reconstructed geometry at the encoder.
[0184] Figure 17 An example method for encoding point cloud frames is shown. The geometry and / or attributes of the current point cloud frame (e.g., current point cloud frame 1710) can be encoded. Current point cloud frame 1710 can be a frame in a sequence of point cloud frames (e.g., a dynamic point cloud). Figure 17 One or more steps of the example method (e.g., method 1700) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 29 Example computer system 2900 and / or Figure 30 The example computing device 3030 is used to perform and / or implement this. In some examples, steps 1720, 1730, 1740, and 1750 may represent components within the encoder. Figure 17 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0185] At step 1720, an encoder (e.g., a geometry encoder) may encode the geometry information 1721 of the current point cloud frame 1710. The encoded geometry information 1721 may be sent to bitstream 1760. The encoder may determine the decoded geometry 1722 (e.g., reconstructed geometry) of the current point cloud frame. The encoder may determine the decoded geometry 1722, for example, based on decoding the encoded geometry information 1721. The decoded geometry 1722 of the current point cloud frame 1710 and the geometry (e.g., the original geometry) may be different. For example, if the geometry compression is lossy, the decoded geometry 1722 and the geometry (e.g., the original geometry) may be significantly different.
[0186] At step 1730, the encoder (e.g., an attribute determiner) may determine a mapping attribute 1731 associated with the decoded geometry 1722 of the current point cloud frame. The encoder may determine the mapping attribute 1731, for example, by mapping attributes associated with the geometry (e.g., the original) of the current point cloud frame 1710 to the decoded geometry 1722 of the current point cloud frame.
[0187] Attributes may include color. Mapped attributes can be determined, for example, based on recoloring the attributes. Mapped attributes can also be determined, for example, based on a k-nearest neighbor (KNN) search algorithm that determines the nearest point from the geometry of the current point cloud frame 1710 to the decoded geometry 1722. The k-nearest neighbor (KNN) search algorithm may include, for example, using spatial partitioning algorithms such as KD-tree search, ball / metric tree search, brute-force search, etc. The mapped attributes of the points in the decoded geometry 1722 may be, for example, average attribute values associated with the nearest points in the current point cloud frame 1710 relative to the points in the decoded geometry. For example, if geometry compression is lossless, the decoded geometry 1722 (e.g., the reconstructed geometry) of the current point cloud frame's geometry may be identical to the geometry of the current point cloud frame (e.g., the original geometry). Attribute mapping can associate the attributes of each point in the current point cloud frame's geometry with the corresponding points in the decoded geometry 1722.
[0188] At step 1740, the encoder (e.g., an attribute projector) may determine the projection attribute 1742, for example, by projecting the attribute of the reference point cloud frame 1741 onto the decoded geometry 1722. At step 1750, the encoder (e.g., an attribute encoder) may encode (e.g., generate) the attribute information 1751 and transmit / signal the attribute information 1751 to the bit stream 1760 / the bit stream. The attribute information 1751 may represent a mapping attribute 1731 associated with the decoded geometry 1722. The attribute information 1751 may be based on attribute predictions determined from the projection attribute 1742. Attribute predictions may include, for example, from... Figure 19 and Figure 25 The attribute predictor defined by the projected attribute 1742.
[0189] The reference point cloud frame for the attribute can be a coded reference point cloud frame. This coded reference point cloud frame may have been previously selected to encode the geometry of the current point cloud frame (e.g., at step 1720). The coded reference point cloud frame used to determine the projection attribute 1742 can be different from the coded reference point cloud frame used at step 1720. The coded reference point cloud frame can be motion-compensated, for example, by using a motion vector MV field. The coded reference point cloud frame and / or the motion vector MV field can be encoded in bitstream 1760.
[0190] The reference point cloud frame 1741 for an attribute can be determined from a coded reference point cloud frame. The coded reference point cloud frame can be motion-compensated by a motion vector MV field, which can be encoded in a bit stream 1760. The attribute of the coded reference point cloud frame can move with the point, such that, for example, if motion compensation occurs, the reference point cloud frame 1741 for the attribute possesses the attribute.
[0191] Figure 18 An example method for decoding a point cloud frame is shown. The geometry and / or attributes of the current point cloud frame can be decoded. The current point cloud frame can be a frame in a sequence of point cloud frames (e.g., a dynamic point cloud). The current point cloud frame can be as described in this document regarding... Figure 17 The description is encoded. Figure 18 One or more steps of the example method (e.g., method 1800) can be decoded by a decoder (e.g., Figure 1 decoder 120) Figure 29 Example computer system 2900 and / or Figure 30 The example computing device 3030 is used to perform and / or implement this. In some examples, steps 1810, 1820, and 1830 may represent components within the decoder. Figure 18 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0192] At step 1810, the decoder (e.g., a geometry decoder) can provide (e.g., determine) the decoded geometry 1812 of the current point cloud frame (e.g., the reconstructed geometry at the encoder). The decoded geometry 1812 can be provided / determined by decoding the geometry information 1811 from the bitstream 1840. The decoded geometry 1812 can be as described herein. Figure 17 The described decoded geometry is 1722.
[0193] At step 1820, the decoder (e.g., an attribute projector) can determine projection attribute 1822. This projection attribute can be determined, for example, by projecting the attribute of a reference point cloud frame 1821 onto the decoded geometry 1812. Projection attribute 1822 can be related to, as described herein, […]. Figure 17 The projection attribute 1742 described is the same. Similarly, the reference point cloud frame 1821 for the attribute can be the same as described in this paper. Figure 17 The reference point cloud frame 1741 described for the attribute is the same.
[0194] At step 1830, the decoder (e.g., an attribute decoder) can determine the decoded attribute 1832 associated with the decoded geometry 1812. The decoded attribute 1832 can be determined, for example, by decoding attribute information 1831 from bitstream 1840 based on attribute prediction. The attribute prediction and / or decoded attribute 1832 can be determined from the projected attribute 1822.
[0195] The reference point cloud frame 1821 for the attribute can be a coded reference point cloud frame. The coded reference point cloud frame may have been previously determined to decode the geometry of the current point cloud frame at step 1810. Motion vectors indicating the coded reference point cloud frame can be decoded from the geometric information 1811. The coded reference point cloud frame used to determine the projection attribute 1822 can be the same as or different from the coded reference point cloud frame used at step 1810. The coded reference point cloud frame can be motion-compensated, for example, by using a motion vector MV field. The coded reference point cloud frame and / or the motion vector MV field can be decoded from the bit stream 1840.
[0196] The reference point cloud frame 1821 for the attribute can be a motion-compensated point cloud frame determined from a coded reference point cloud frame. The motion-compensated point cloud frame can be motion-compensated by a motion vector (MV) field. The MV field can be decoded from the bit stream 1840. The attribute of the coded reference point cloud frame can move with the point, such that, for example, if motion compensation occurs, the reference point cloud frame 1821 for the attribute can possess the attribute. The reference point cloud frame 1821 for the attribute can be compared with, as described herein... Figure 17 The reference point cloud frame 1741 described for the attribute is the same as (or different from) the reference point cloud frame 1741.
[0197] Projection attribute 1742, projection attribute 1822, mapping attribute 1731, and / or decoded attribute 1832 may belong to the same decoded geometry 1722 (and / or decoded geometry 1812). An attribute predictor for mapping attribute 1731 (e.g., corresponding to the current frame 1710) can be determined from projection attribute 1742 and / or projection attribute 1822. The attribute predictor for mapping attribute 1731 (e.g., corresponding to the current frame 1710) can be determined from projection attribute 1742 and / or projection attribute 1822 in a manner similar to that described herein with respect to steps 1750 and 1830. Determining attribute predictions (e.g., attribute predictors) from projection attributes can be performed iteratively (e.g., independently and / or identically) at the encoder and decoder. This can reduce the size of compressed attribute information 1751 in bitstream 1760 (and / or reduce the size of compressed attribute information 1831 in bitstream 1840).
[0198] It is possible to improve the scheme based on prediction (e.g., as discussed in this article). Figure 11 and / or Figure 13 The attribute predictor is determined by (as described) the intra-frame transform used for (e.g., applied to) residual attributes, which is the residual attribute predictor. This can be based, for example, on information from a lower level of detail. Attribute a 0 ... a i-1 and / or the set of points S predicted from projection attribute 1742 i Attribute a i To determine the attribute predictor. It should be understood that input a i It can represent residual attributes. Input 'a' i You can choose to represent an attribute value or not.
[0199] It is possible to improve the scheme based on prediction (e.g., as discussed in this article). Figure 12 and / or Figure 14 The attribute predictor is determined (as described). This can be done, for example, from a lower level of detail. Attribute a 0 ... a i-1 And / or predict the set of points S from the projection attribute 1822 i Attribute a i Attribute predictions can be determined, for example, from projection attribute 1742 (and / or projection attribute 1822). Attribute predictions can be based on (e.g., using) an inter-frame RATH scheme. The RATH scheme can use the same octree partitioning for mapping attribute 1731, decoded attribute 1832, projection attribute 1742, and / or projection attribute 1822. For example, if both the mapping attribute and the projection attribute are associated with the same decoded geometry 1722 (and / or decoded geometry 1812), the same octree partitioning can be used.
[0200] Each DC and / or AC coefficient determined to be applied to the output of the inter-frame RAHT scheme on mapping attribute 1731 can be predicted by co-located DC and / or AC coefficients obtained as the output of the inter-frame RAHT scheme applied to projection attribute 1742. Each DC and / or AC coefficient determined to be applied to the output of the inter-frame RAHT scheme on decoded attribute 1832 can be predicted by co-located DC and / or AC coefficients obtained as the output of the inter-frame RAHT scheme applied to projection attribute 1822. For example, if the mapping attribute, projection attribute, and / or decoded attribute are associated with the same decoded geometry 1722 (and / or the same decoded geometry 1812), the problem of the inter-frame RAHT scheme's sensitivity to geometric differences between the current point cloud and the motion-compensated point cloud can be addressed.
[0201] Figure 19An example method for encoding the mapping attributes of a point cloud frame is shown. The mapping attributes can be encoded, for example, based on attribute predictions determined from the projection attributes (e.g., an attribute predictor). Figure 19 One or more steps of the example method (e.g., method 1900) can be accomplished by an encoder (e.g., Figure 1 encoder 114 in Figure 29 Example computer system 2900 and / or Figure 30 Example computing device 3030 is used to perform and / or implement. Figure 19 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0202] Method 1900 can correspond to Figure 17 Step 1750. At step 1910, the encoder can determine the predicted attribute (e.g., also referred to as the attribute predictor) from (e.g., based on) the projection attribute 1742. The encoder can determine the residual attribute 1911, for example, by subtracting the predicted attribute from the mapping attribute 1731. At step 1920, the encoder can determine the quantized residual attribute 1921, for example, by quantizing the residual attribute 1911. The encoder can determine the quantized residual attribute 1921, for example, for lossy compression. For example, if lossy compression is allowed, quantization can be used. At step 1930, the encoder can encode the residual attribute 1911 as attribute information 1751 into the bitstream 1760. Alternatively or concurrently, the encoder can encode the quantized residual attribute 1921 as attribute information 1751 into the bitstream 1760.
[0203] Figure 20 An example method for decoding attribute information of a point cloud frame is shown. The attribute information can be decoded, for example, based on attribute prediction (e.g., an attribute predictor). Attribute predictions can be determined from projection attributes. Figure 20 One or more steps of the example method (e.g., method 2000) can be decoded by a decoder (e.g., Figure 1 decoder 120 in Figure 29 Example computer system 2900 and / or Figure 30 Example computing device 3030 is used to perform and / or implement. Figure 10 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0204] Method 2000 can correspond to Figure 18Step 1830. At step 2010, the decoder can determine the residual attribute (e.g., quantized residual attribute 2011) by decoding the attribute information 1831 from the bitstream 1840, for example. For example, if the residual attribute 2011 is quantized, then at step 2020, the decoder can dequantize the residual attribute 2011. The decoder can determine the decoded residual attribute, for example, by dequantizing the quantized residual attribute 2011. At step 2030, the decoder can determine the prediction attribute (e.g., also referred to as the attribute predictor) from the projection attribute 1822. The decoder can determine the decoded attribute 1832, for example, by adding the decoded residual attribute 2021 to the prediction attribute.
[0205] Figure 21 An example method for encoding the mapping attributes of a point cloud frame is shown. The mapping attributes can be encoded, for example, based on attribute predictions (e.g., attribute predictors). Attribute predictions can be determined from the projection attributes. Figure 21 One or more steps of the example method (e.g., method 2100) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 29 Example computer system 2900 and / or Figure 30 Example computing device 3030 is used to perform and / or implement. Figure 21 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0206] Method 2100 can correspond to Figure 17 Step 1750. At step 2140, the encoder may determine the smoothed projection attribute 2141, for example, by smoothing the projection attribute 1742. The smoothed projection attribute 1742 may include removing pseudo-high frequencies that may impair compression performance. The smoothed value of the projection attribute associated with a point of the decoded geometry may be obtained, for example, by averaging the values of the projection attributes associated with points of the decoded geometry that belong to the neighborhood of that point (e.g., before smoothing). This may be referred to as smoothing on a 3D kernel. The 3D kernel used for smoothing may define and / or be used to determine the shape of the function used to take the average of neighboring points. The 3D kernel may define and / or be used to determine the neighborhood as points where voxels intersect with the voxel of that point.
[0207] At step 2110, the encoder can determine the predicted attribute from the smoothed projection attribute 2141. The encoder can determine the residual attribute 2111 (e.g., the residual value of the attribute) by subtracting the predicted attribute from the mapped attribute 1731, for example. At step 2120, the encoder can determine the quantized residual attribute 2121 by quantizing the residual attribute 2111. For example, quantization can be used if lossy compression is permissible. At step 2130, the encoder can encode the residual attribute 2111 as attribute information 1751 into bitstream 1760. Alternatively, the encoder can encode the quantized residual attribute 2121 (e.g., the quantized residual value) as attribute information 1751 into bitstream 1760.
[0208] Figure 22 An example method for decoding attribute information of a point cloud frame is shown. The attribute information can be decoded, for example, based on attribute prediction (e.g., an attribute predictor). Attribute predictions can be determined from projection attributes. Figure 22 One or more steps of the example method (e.g., method 2200) can be decoded by a decoder (e.g., Figure 1 decoder 120) Figure 29 Example computer system 2900 and / or Figure 30 Example computing device 3030 is used to perform and / or implement. Figure 22 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0209] Method 2200 can correspond to Figure 18 Step 1830. At step 2210, the decoder may determine the residual attribute (e.g., quantized residual attribute 2211) by decoding the attribute information 1831 from the bitstream 1840, for example. For example, if the residual attribute is quantized at step 2210, the decoder may dequantize the residual attribute 2211 at step 2220. The decoder may determine the decoded residual attribute 2221 by dequantizing the quantized residual attribute 2211. At step 2240, the decoder may determine the smoothed projection attribute 2241 (e.g., attribute value) by smoothing the projection attribute 1822, for example. Smoothing the projection attribute 1822 may include removing pseudo-high frequencies that may impair compression performance.
[0210] At step 2230, the decoder can determine the prediction attribute from the smoothed projection attribute 2241. The decoder can determine the decoded attribute 1832, for example, by adding the decoded residual attribute 2221 to the prediction attribute. For example, the residual attribute can be encoded (e.g., at steps 1930, 2130) and decoded (e.g., at steps 2010, 2210) by any intra-frame attribute writing scheme, since inter-frame correlation can be determined by constructing the residual attribute. The residual attribute can be encoded and / or decoded by a prediction boosting scheme, where S can be the set of residual attributes to be written (e.g., all residual attributes).
[0211] Figure 23 An example method for encoding residual properties is shown. Step 2300 can correspond to... Figure 19 Step 1930 or Figure 21 Step 2130. Figure 23 One or more steps of the example method (e.g., method 2300) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 29 Example computer system 2900 and / or Figure 30 Example computing device 3030 is used to perform and / or implement. Figure 23 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0212] At step 2310, the encoder can determine the transformed coefficients 2311. The transformed coefficients 2311 can be determined, for example, by applying an intra-frame transform to (e.g., applying to) residual attributes. At step 2320, the encoder can determine the quantized transformed coefficients 2321, for example, by quantizing the transformed coefficients 2321. For example, quantization can be used if lossy compression is permissible. At step 2330, the encoder can encode (e.g., entropy encoding) the transformed coefficients 2311 into attribute information 1751. Alternatively or concurrently, the encoder can encode (e.g., entropy encoding) the quantized transformed coefficients 2321 into attribute information 1751.
[0213] Figure 24 An example method for decoding residual properties is shown. Figure 17 One or more steps of the example method (e.g., method 1700) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 29 Example computer system 2900 and / or Figure 30 Example computing device 3030 is used to perform and / or implement. Figure 17The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0214] Method 2400 can correspond to Figure 20 Step 2010 or Figure 22 Step 2210. At step 2410, the decoder can determine the transformed coefficients (e.g., quantized transformed coefficients 2411) for example by entropy decoding of the attribute information 1831. For example, if the transformed coefficients 2411 at step 2410 are quantized, then at step 2420, the decoder can dequantize the transformed coefficients 2411. The decoder can determine the transformed coefficients 2421 by dequantizing the quantized transformed coefficients 2411. At step 2430, the decoder can determine the residual attributes by applying (e.g., applying to) the residual attributes using the inverse intra-frame transform. The operation of the inverse intra-frame transform can correspond to the inverse operation of the intra-frame transform as described herein with respect to step 2310. The intra-frame transform can be an adaptive DCT. The inverse intra-frame transform can be an inverse adaptive DCT (A-DCT). The intra-frame transform can be a RAHT transform. The inverse intra-frame transform can be the inverse RAHT transform of the RAHT scheme.
[0215] For example, if lossless attribute write coding is performed, quantization can be omitted, and the residual attribute can be entropy-coded (e.g., without transformation) (e.g., at steps 2310 and / or 2430). For example, if lossless attribute write coding is performed, the intra-frame transform can be a Haar transform. For example, if lossless attribute write coding is performed, the inverse intra-frame transform can be an inverse Haar transform.
[0216] Depending on the local prediction quality of the projection attributes (e.g., steps 1742, 1822, 2241), the attributes of the current point cloud frame may not need to be encoded using the projection attributes. For example, in some regions of the decoded geometry of the current point cloud frame, the attributes of the current point cloud frame may not need to be encoded using the projection attributes. For example, if the correlation between the attributes of the current point cloud frame and the projection attributes is low, the attributes of the current point cloud frame may not need to be encoded using the projection attributes.
[0217] The activation of writing attributes associated with the decoded geometry of the current point cloud frame can be signaled in the bitstream using projection attributes. This signal can be generated, for example, by an inter-frame residual activation flag indicating whether the projection attribute is locally used. Either the encoder or the decoder can determine whether the inter-frame residual activation flag indicates the use of the projection attribute. For example, if the inter-frame residual activation flag indicates the use of the projection attribute, the residual attribute can be encoded in the bitstream. For example, if the inter-frame residual activation flag indicates that the projection attribute is not used, the residual attribute can not be encoded in the bitstream, and / or the attribute associated with the decoded geometry of the current point cloud can be encoded using an alternative method.
[0218] Inter-frame residual activation flags can be associated with a spatial region of the current point cloud frame. A spatial region can include spatial blocks (as defined in the high-level syntax of the GGC codec). A spatial block can be a partition of 3D space surrounding point clouds, each with its own write-code parameters, which can be written independently of each other. A spatial region can include nodes or groups of nodes of a RAHT octree. Dedicated partitions of the space can be performed. Each partition portion can be associated with an inter-frame residual activation flag.
[0219] The inter-frame residual activation flag can be inferred from the prediction quality of the mapping attributes associated with the reconstructed (e.g., decoded) geometry. The inter-frame residual activation flag can be sent (e.g., explicitly). The determination of whether the attribute predictor is used to encode attributes of the reconstructed geometry can be based on the prediction quality of the attributes associated with the reconstructed geometry. The prediction quality can be determined by the encoder. Similarly, at the decoder, the determination of whether the attribute predictor is used to decode attributes of the reconstructed geometry can be based on the prediction quality of the attributes associated with the reconstructed geometry. The prediction quality can be determined by the decoder.
[0220] Predicted quality can be determined (e.g., evaluated) based on a quality metric. The quality metric can be based on the difference between the mapped attributes and the attributes associated with the reconstructed geometry. Projection quality can be determined, for example, based on projection distance. For example, a larger projection distance may result in lower quality. Projection distance can be based on the difference between the vertices associated with the attributes of the motion-compensated point cloud frame and the vertices associated with the mapped attributes of the decoded geometry. An inter-frame residual activation flag can be used (e.g., treated as being sent as a signal) based on quality exceeding a threshold, for example. Projection attributes can be determined based on quality exceeding a threshold, for example. Distance can be evaluated in steps 1740 and / or 1820.
[0221] Alternatively or additionally, the inter-frame residual activation flag can be written into the bitstream by the encoder and / or decoded from the bitstream by the decoder. The inter-frame residual activation flag can be entropy-coded and / or decoded according to the prediction quality. The prediction quality can be evaluated, for example, by projection distance as described herein.
[0222] Decoding information (e.g., geometric or attribute information) can instruct the information to be decoded from: at least one single bit (e.g., a flag), each comprising at least one codeword of more than one bit, or a combination of at least one flag and at least one codeword. Decoding information from a bitstream can instruct the parsing of a bitstream based on a specific syntax and / or reading from the bitstream: at least one single bit (e.g., a flag), each comprising at least one codeword of more than one bit, or a combination of at least one flag and at least one word representing the information. Transmitting information with signals can include encoding the information into a bitstream and / or decoding the information from the bitstream.
[0223] Figure 25 An example method for encoding the mapping attributes of a point cloud frame is shown. The encoding can be based on transform coefficient predictions (e.g., a transform attribute coefficient predictor). The transform coefficient predictions can be determined, for example, from the projection attributes. Figure 25 One or more steps of the example method (e.g., method 2500) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 29 Example computer system 2900 and / or Figure 30 Example computing device 3030 is used to perform and / or implement. Figure 25 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0224] Method 2500 can correspond to Figure 17 Step 1750. Regarding... Figure 25 The description can refer to Figure 17 In step 2510, the encoder can determine the transformed coefficients 2511. The transformed coefficients 2511 can be determined, for example, by applying an intra-frame transform to (e.g., applying to) mapping attribute 1731. In step 2520, the encoder can determine the predicted transformed coefficients (e.g., also referred to as the transformed attribute coefficient predictor) from projection attribute 1742, for example. A second intra-frame transform can be applied to (e.g., applying to) projection attribute 1742, for example, to determine the predicted transformed coefficients. The second intra-frame transform can be the same intra-frame transform as in step 2510, a partial intra-frame transform of the intra-frame transform in step 2510, or a different intra-frame transform. The encoder can determine the transformed coefficient residual 2521, for example, by subtracting the predicted transformed coefficients from the transformed coefficients 2511.
[0225] At step 2530, the encoder can determine the quantized transform coefficient residual 2531, for example, by quantizing the transform coefficient residual 2521. For example, if lossy compression is performed, the encoder can determine the quantized transform coefficient residual 2531. The transform coefficient residual 2521 can be quantized, for example, based on quantization parameters (e.g., quantization indicators). Quantization parameters can be used to determine quantization scaling values / factors. For example, if lossy compression is allowed or enabled, quantization can be used.
[0226] For example, if lossless compression is performed, the operation of step 2530 can be bypassed (e.g., omitted or skipped). The transformed coefficient residual 2521 can be used at step 2540. In some cases, the quantized transformed coefficient residual 2531 can be equal to the transformed coefficient residual 2521. At step 2540, the encoder can encode (e.g., entropy encoding) the transformed coefficient residual 2521 (e.g., if step 2520 is performed) or the quantized transformed coefficient residual 2531 (if step 2530 is performed) into bitstream 1760. Attribute information 1751 may include the encoded transformed coefficient residual 2521 or the quantized transformed coefficient residual 2531.
[0227] Figure 26 An example method for decoding attribute information of a point cloud frame is shown. Decoding can be based on transform coefficient prediction (e.g., a transform attribute coefficient predictor). The transform coefficient prediction can be determined, for example, from the projection attributes. Figure 26 One or more steps of the example method (e.g., method 2600) can be achieved through an encoder (e.g., Figure 1 encoder 120) Figure 29 Example computer system 2900 and / or Figure 30 Example computing device 3030 is used to perform and / or implement. Figure 26 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0228] Method 2600 can correspond to Figure 18 Step 1830. Regarding... Figure 26 The description can refer to the description in this article. Figure 18 The part. Step 2600 can be reversed (e.g., inverted), for example, as described in this article regarding... Figure 25 The operation of step 2500 performed by the encoder is described. At step 2610, the decoder may determine the quantized transformed coefficient residual 2611 (or transformed coefficient residual) by decoding the attribute information 1831 from the bit stream 1840 (e.g., entropy decoding).
[0229] At step 2620, the decoder can determine the transformed coefficient residual 2621, for example, by dequantizing the quantized residual attribute 2611 (e.g., if enabled or implemented). The transformed coefficient residual 2621 can be dequantized (e.g., dequantized) based, for example, on quantization parameters (e.g., quantization indicators) used to determine the inverse quantization scaling value / factor. The inverse quantization scaling value / factor can be, for example, as described herein. Figure 25 Step 2530 describes the inverse of the quantization scaling value / factor determined by the encoder.
[0230] As described herein with respect to step 2530, quantization may be disabled or not performed, for example, for lossless compression. In this case, dequantization may be performed, and the operation in step 2620 may be bypassed (e.g., omitted or skipped). The transformed coefficient residual 2621 may be used at step 2630. In some cases, the quantized transformed coefficient residual 2611 may be equal to the transformed coefficient residual 2621.
[0231] At step 2630, the decoder may, for example, determine the predicted transformed coefficients from projection attribute 1822 (e.g., also referred to as the transformed attribute coefficient predictor). A second intra-frame transform may be used (e.g., applied to) projection attribute 1742, for example, to determine the predicted transformed coefficients. The second intra-frame transform may be the same intra-frame transform as described herein with respect to step 2510, a partial intra-frame transform as described herein with respect to step 2510, or a different intra-frame transform. The decoder may, for example, determine the transformed coefficients 2631 by adding the transformed coefficient residuals 2621 to the predicted transformed coefficients.
[0232] At step 2640, the decoder can determine the decoded attribute 1832, for example, by applying (e.g., to) the transform coefficients using an inverse intra-frame transform. The operation of the inverse intra-frame transform can correspond to the intra-frame transform (e.g., as described herein regarding...). Figure 25 The inverse operation described in step 2510. This paper describes Figure 25 Intra-frame transformation and / or mentioned in Figure 26 Examples of corresponding inverse intra-frame transforms mentioned in the text. The intra-frame transform can be an adaptive DCT and / or the inverse intra-frame transform can be an inverse adaptive DCT (A-DCT). The intra-frame transform can be a RAHT transform and / or the inverse intra-frame transform can be the inverse RAHT transform of a RAHT scheme. For example, if lossless attribute write coding is performed (e.g., lossless compression), the intra-frame transform can be an integer Haar transform and / or the inverse intra-frame transform can be an inverse integer Haar transform.
[0233] In steps 2520 and / or 2630, the predicted transformed coefficients can be obtained, for example, by applying (e.g., to) the intra-frame transform to the projection attributes. The predicted transformed coefficients can also be obtained, for example, by applying (e.g., to) a partial intra-frame transform. A partial intra-frame transform can be used to obtain the predicted transformed coefficients, for example, such that a subset of the coefficients can be used for prediction. Coefficients not used for prediction can be set to zero, such that a portion (e.g., a subset) of the coefficients can be used for prediction. In some cases, a smaller than the entire intra-frame transform can be computed. In some cases, a subset of coefficients can be computed. The subset of coefficients can be fixed. The subset of coefficients can exclude the number of highest frequency coefficients (e.g., a predetermined number or based on a threshold).
[0234] A subset of coefficients can be signaled in the bit stream because the number of decomposition levels to be predicted is signaled. The decoder can use the number of decoded decomposition levels to determine the subset of coefficients. The subset of coefficients can be signaled, for example, by following a RAHT decomposition tree. For example, for a given node, a bit can be signaled in the bit stream to achieve a prediction for that node (e.g., included in the subset of coefficients associated with that node). For example, for a given node, a bit can be signaled in the bit stream to prune its subtree (e.g., excluding coefficients associated with that node and all coefficients associated with its child nodes from the subset).
[0235] Figure 27 An example method for encoding attributes of point cloud frames is shown. More specifically, Figure 27 A flowchart 2700 illustrates an example method for encoding attributes of a point cloud frame. The point cloud frame can be the current point cloud frame. Encoding can be based on an attribute predictor. The method in flowchart 2700 can be implemented using an encoder (e.g., Figure 1 encoder 114 in Figure 29 Example computer system 2900 and / or Figure 30 The example computing device 3030 is used to perform and / or implement the method. Method 2700 may correspond to the method described herein. Figure 17 Method 1700. Figure 27 The steps of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0236] At step 2702, the encoder can determine the attributes of the reconstructed geometry of the point cloud frame. The attributes of the reconstructed geometry can be determined, for example, based on the attributes of the point cloud frame's geometry. Determining the attributes of the reconstructed geometry can include mapping the attributes of the point cloud frame's geometry to the reconstructed geometry. An attribute predictor can be further determined based on the mapped attributes. The attribute can be color. The mapped attributes can be determined, for example, based on recoloring. The mapping attributes of each point in the reconstructed geometry can be determined, for example, based on a nearest neighbor search of the nearest points from the point cloud frame's geometry to the points in the reconstructed geometry.
[0237] At step 2704, the encoder can determine an attribute predictor for the properties of the reconstructed geometry. The encoder can determine the attribute predictor for the properties of the reconstructed geometry, for example, based on projecting the properties of a reference point cloud frame onto the reconstructed geometry. The reference point cloud frame can be an attribute-specific reference point cloud frame. An attribute-specific reference point cloud frame can be a coded reference point cloud frame or a motion-compensated point cloud frame. The motion-compensated point cloud frame can, for example, be derived from a coded reference point cloud frame (e.g., as described herein regarding...). Figure 17 (as described) to determine.
[0238] At step 2706, the encoder can encode the attributes of the reconstructed geometry. The encoder can encode the attributes of the reconstructed geometry, for example, based on an attribute predictor. The encoder can determine residual attributes, for example, based on the difference between the attributes of the reconstructed geometry and the attribute predictor. The encoder can encode the residual attributes in a bitstream.
[0239] Encoding residual attributes can include, for example, determining transform coefficients. Transform coefficients can be determined, for example, based on applying an intra-frame transform to (e.g., applying) the residual attributes. Encoding residual attributes can include encoding the transform coefficients corresponding to (e.g., representing or indicating) the residual attributes in the bitstream. Encoding residual attributes can also include quantizing the transform coefficients and entropy coding the quantized transform coefficients.
[0240] Figure 28 An example method for decoding the attributes of a point cloud frame is shown. More specifically, Figure 28 A flowchart 2800 illustrates an example method for decoding attributes of a point cloud frame. The point cloud frame can be the current point cloud frame. Decoding can be based on an attribute predictor. The method in flowchart 2800 can be implemented using a decoder (e.g., Figure 1 decoder 120 in Figure 29 Example computer system 2900 and / or Figure 30 The example computing device 3030 is used to perform and / or implement the method. Method 2800 may correspond to the method described herein. Figure 18Method 1800. The decoder may include a geometry decoder, an attribute decoder, and / or an attribute projector / determiner, as described herein. Figure 18 As described.
[0241] At step 2802, the decoder (e.g., a geometry decoder) can decode the geometry of the point cloud frame to determine the reconstructed geometry of the point cloud frame. The decoder can decode geometric information of, for example, the point cloud from the bit stream. The decoded geometry can correspond to the geometry of the point cloud frame reconstructed (e.g., encoded, then decoded) at the encoder. To decode attributes associated with the reconstructed geometry, the decoder can decode residual attributes from the bit stream. The residual attributes can indicate the difference between the attributes of the reconstructed geometry and the attribute predictor. The decoded attributes can be determined, for example, by adding the attribute predictor to the decoded residual attributes.
[0242] Decoding residual attributes can include decoding (e.g., entropy decoding) the transformed coefficients corresponding to (e.g., representing or indicating) the residual attributes from the bitstream. The residual attributes can be determined, for example, based on applying (e.g., applying) an inverse intra-frame transform to the decoded transformed coefficients. Decoding residual attributes can also include dequantizing the transformed coefficients. The residual attributes can be determined by applying (e.g., applying) an inverse intra-frame transform to the dequantized transformed coefficients.
[0243] At step 2804, the decoder may, for example, determine an attribute predictor of the attributes of the reconstructed geometry based on the projection of the attributes of a reference point cloud frame for the attributes onto the reconstructed geometry. The reference point cloud frame for the attributes may be a coded reference point cloud frame or a motion-compensated point cloud frame. The motion-compensated point cloud frame may, for example, be derived from a coded reference point cloud frame (e.g., as described herein regarding...). Figure 18 (as described) to determine.
[0244] At step 2806, the decoder may decode the attributes of the reconstructed geometry, for example, based on an attribute predictor. As described herein, the attribute predictor may be determined iteratively (e.g., independently and / or identically) at the encoder and decoder. The attribute predictor may be determined, for example, based on the attributes of the motion-compensated point cloud frame projected onto the reconstructed geometry of the point cloud frame. The attribute predictor may include a corresponding attribute predictor for each corresponding point in the points of the reconstructed geometry. The corresponding attribute predictor may be based on the projection attributes corresponding to that point in the projection attributes. Determining the attribute predictor may include smoothing the projection attributes. The predicted attributes may be determined from the smoothed projection attributes.
[0245] Motion-compensated point cloud frames can be determined from coded reference point cloud frames. Motion-compensated point cloud frames can be determined, for example, based on motion compensation of the coded reference point cloud frame. Motion compensation can be based on motion vectors (MVs). Motion vectors can be determined, for example, based on the difference between the reconstructed geometry and the geometry of the coded reference point cloud frame adjusted by the motion vectors. Distortion-reducing motion vectors can be selected as the determined motion vectors. The motion vectors and / or indications (e.g., indexes or IDs) of the coded reference point cloud frame can be transmitted in the bitstream as signals (e.g., encoded by an encoder and / or decoded by a decoder).
[0246] As described herein, an attribute predictor can be used to determine residual attributes (e.g., at the encoder) and / or combine them with decoded residual attributes (e.g., at the decoder) to decode (e.g., reconstruct) the attributes at the decoder. The residual attributes can be encoded and / or decoded, for example, based on an intra-frame transform scheme (e.g., a prediction-boosting (prediction-enhanced) transform scheme, an adaptive DCT and its corresponding inverse A-DCT, a RAHT transform and its corresponding inverse RAHT transform, a Haar transform and / or its corresponding inverse Haar transform).
[0247] Figure 29 An example computer system in which the present disclosure can be implemented is shown. For example, Figure 29 The example computer system 2900 shown can implement one or more methods described herein. For example, various devices and / or systems described herein (e.g., in...) Figure 1 , 2 (3) can be implemented in the form of one or more computer systems 2900. Furthermore, each step of the flowchart depicted in this disclosure can be implemented on one or more computer systems 2900.
[0248] Computer system 2900 may include one or more processors, such as processor 2904. Processor 2904 may be a dedicated processor, a general-purpose processor, a microprocessor, and / or a digital signal processor. Processor 2904 may be connected to communication infrastructure 2902 (e.g., a bus or network). Computer system 2900 may also include main memory 2906 (e.g., random access memory (RAM)) and / or secondary memory 2908.
[0249] Secondary storage 2908 may include hard disk drive 2910 and / or removable storage drive 2912 (e.g., magnetic tape drive, optical disc drive, etc.). Removable storage drive 2912 may read from and / or write to removable storage unit 2916. Removable storage unit 2916 may include magnetic tape, optical disc, etc. Removable storage unit 2916 may be read from and / or write-to by removable storage drive 2912. Removable storage unit 2916 may include computer-usable storage media having computer software and / or data stored therein.
[0250] Secondary memory 2908 may include other similar components for allowing computer programs or other instructions to be loaded into computer system 2900. Such components may include removable storage unit 2918 and / or interface 2914. Examples of such components may include program boxes and / or box interfaces (such as in video game devices) that allow software and / or data to be transferred from removable storage unit 2918 to computer system 2900, removable memory chips (such as erasable programmable read-only memory (EPROM) or programmable read-only memory (PROM)) and associated sockets, flash drives and USB ports, and / or other removable storage units 2918 and interfaces 2914.
[0251] Computer system 2900 may also include communication interface 2920. Communication interface 2920 allows software and data to be transferred between computer system 2900 and external devices. Examples of communication interface 2920 may include a modem, network interface (e.g., Ethernet card), communication port, etc. Software and / or data transferred via communication interface 2920 may be in the form of signals, which may be electronic, electromagnetic, optical, and / or other signals that can be received by communication interface 2920. Signals may be provided to communication interface 2920 via communication path 2922. Communication path 2922 may carry signals and may be implemented using wires or cables, optical fibers, telephone lines, cellular telephone links, RF links, and / or any other communication channels.
[0252] Computer program media and / or computer-readable media can be used to refer to tangible storage media, such as removable storage units 2916 and 2918 or a hard disk installed in hard disk drive 2910. A computer program product can be a component for providing software to computer system 2900. A computer program (which may also be referred to as computer control logic) can be stored in main memory 2906 and / or secondary memory 2908. A computer program can be received via communication interface 2920. When executed, such a computer program can enable computer system 2900 to implement the present disclosure as discussed herein. Specifically, when executed, a computer program can enable processor 2904 to implement the processes of the present disclosure, such as any of the methods described herein. Therefore, such a computer program can represent a controller of computer system 2900.
[0253] The features of this disclosure can be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. Implementing a hardware state machine to perform the functions described herein will also be apparent to those skilled in the art.
[0254] Figure 30Example elements of a computing device are shown that can be used to implement any of the various devices described herein, including, for example, a source device (e.g., 102), an encoder (e.g., 114), a destination device (e.g., 106), a decoder (e.g., 120), and / or any computing device described herein. The computing device 3030 may include one or more processors 3031 that can execute instructions stored in random access memory (RAM) 3033, removable media 3034 (e.g., a Universal Serial Bus (USB) drive, an optical disc (CD) or digital versatile optical disc (DVD), or a floppy disk drive), or any other desired storage medium. Instructions may also be stored in an attached (or internal) hard disk drive 3035. The computing device 3030 may also include a security processor (not shown) that can execute instructions of one or more computer programs to monitor processes executing on the processor 3031 and any processes requesting access to any hardware and / or software components of the computing device 3030 (e.g., ROM 3032, RAM 3033, removable media 3034, hard disk drive 3035, device controller 3037, network interface 3039, GPS 3041, Bluetooth interface 3042, WiFi interface 3043, etc.). The computing device 3030 may include one or more output devices, such as a display 3036 (e.g., screen, display device, monitor, television, etc.), and may include one or more output device controllers 3037, such as a video processor. One or more user input devices 3038 may also be present, such as a remote control, keyboard, mouse, touchscreen, microphone, etc. The computing device 3030 may also include one or more network interfaces, such as a network interface 3039, which may be a wired interface, a wireless interface, or a combination of both. Network interface 3039 can provide the computing device 3030 with an interface to communicate with network 3040 (e.g., RAN or any other network). Network interface 3039 may include a modem (e.g., a cable modem), and external network 3040 may include a communication link, external network, home network, provider wireless, coaxial cable, fiber optic, or hybrid fiber / coaxial cable distribution system (e.g., DOCSIS network), or any other desired network. Additionally, computing device 3030 may include a location detection device, such as a Global Positioning System (GPS) microprocessor 3041, which can be configured to receive and process GPS signals and determine the geographic location of computing device 3030 with possible assistance from external servers and antennas.
[0255] Figure 30The examples shown may be hardware configurations, but the components illustrated can also be implemented as software. Modifications can be made to add, remove, combine, partition, etc., components of computing device 3030 as needed. Alternatively, basic computing devices and components can be used to implement components, and the same components (e.g., processor 3031, ROM storage device 3032, display 3036, etc.) can be used to implement any other computing devices and components described herein. For example, the various components described herein can be implemented using a computing device having components such as a processor that executes computer-executable instructions stored on a computer-readable medium, such as... Figure 30 As shown in the figure. Some or all of the entities described herein may be software-based and may coexist on a common physical platform (e.g., the requesting entity may be a separate software process and program from the relevant entity, both of which may be executed as software on a common computing device).
[0256] A computing device can perform a method including multiple operations. The computing device may include a decoder. The computing device can determine a reconstructed geometry of a point cloud frame based on decoding geometric information of a point cloud frame associated with content. The computing device can determine an attribute predictor associated with the reconstructed geometry based on projecting attributes of a reference point cloud frame onto the reconstructed geometry. The computing device can determine the attribute predictor associated with the reconstructed geometry based on projecting attributes of a reference point cloud frame onto the reconstructed geometry. The computing device can decode attribute information of the reconstructed geometry by decoding residual attributes from a bitstream that indicate the difference between the attribute information of the reconstructed geometry and the attribute predictor. The computing device can determine attribute information of the reconstructed geometry based on the attribute predictor and the residual attributes. The computing device can decode the residual attributes by decoding transformed coefficients corresponding to the residual attributes from a bitstream. The computing device can apply an inverse intra-frame transform to the decoded transformed coefficients. The computing device can also decode transformed coefficients corresponding to the residual attributes from a bitstream. The computing device can dequantize the decoded transformed coefficients. The computing device can decode the attribute information of the reconstructed geometry by: determining a transformed attribute predictor by applying an intra-frame transform to an attribute predictor; determining a transformed residual attribute indicating the difference between the transformed attribute of the reconstructed geometry and the transformed attribute predictor; determining the transformed attribute of the reconstructed geometry based on the transformed attribute predictor and the decoded transformed residual attribute; and applying an inverse intra-frame transform to the transformed attribute of the reconstructed geometry. The computing device can determine the transformed residual attribute by: decoding the transformed coefficients corresponding to the transformed residual attribute; and dequantizing the transformed coefficients. The intra-frame transform can include at least one of the following: adaptive DCT; RAHT transform; or Haar transform. A coded reference point cloud frame associated with a reference point cloud frame can be used to decode the geometric information of the point cloud frame. Determining the attribute predictor can include smoothing the projection attributes of the reference point cloud frame. The computing device can include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to encode point cloud frames. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0257] A computing device can perform a method including multiple operations. The computing device may include an encoder. The computing device can determine a reconstructed geometry of a point cloud frame associated with content. The computing device can determine an attribute predictor associated with the reconstructed geometry based on projecting attributes of a reference point cloud frame onto the reconstructed geometry. The computing device can encode attribute information associated with the reconstructed geometry based on the attribute predictor. The computing device can encode the attribute information by: determining residual attributes based on the difference between the attribute information of the reconstructed geometry and the attribute predictor; and encoding the residual attributes into a bitstream associated with the point cloud frame. The computing device can encode the residual attributes by: determining transformed coefficients based on applying an intra-frame transform to the residual attributes; and entropy encoding the transformed coefficients corresponding to the residual attributes in the bitstream. The computing device can quantize the transformed coefficients before entropy encoding. The computing device can further determine the attribute predictor based on mapping attribute information of the geometry to the reconstructed geometry. The attribute information of the reconstructed geometry may include color. A computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to decode point cloud frames. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0258] A computing device can perform a method comprising multiple operations. The computing device may include a decoder. The computing device may determine an attribute predictor associated with a geometry based on projecting attributes of a reference point cloud frame onto the geometry of the point cloud frame. The computing device may receive residual attributes from a bitstream, indicating a difference between the attribute predictor and attribute information of the geometry. The computing device may use the attribute predictor to decode the residual attributes to determine attributes associated with the geometry of the point cloud frame. The computing device may receive the residual attributes by: receiving transform coefficients corresponding to the residual attributes; and applying an inverse intra-frame transform to the decoded transform coefficients. The intra-frame transform may include at least one of: adaptive DCT; RAHT transform; or Haar transform. A coded reference point cloud frame associated with the reference point cloud frame is used to decode the geometric information of the point cloud frame. The computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to encode point cloud frames. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0259] A computing device can perform a method comprising multiple operations. The computing device may include an encoder. The computing device may determine an attribute predictor associated with a geometry based on projecting attributes of a reference point cloud frame onto the geometry of the point cloud frame. The computing device may encode residual attributes, indicating the difference between the attribute predictor and attribute information of the geometry, in a bit stream. The computing device may encode the residual attributes by: determining transformed coefficients based on applying an inverse intra-frame transform to the residual attributes; and encoding the entropy of the transformed coefficients corresponding to the residual attributes in the bit stream. The inverse intra-frame transform may include at least one of: inverse adaptive DCT; inverse RAHT transform; or inverse Haar transform. The computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to decode a point cloud frame. A computer-readable medium may store instructions that, when executed, enable the performance of the described methods, additional operations, and / or include additional elements.
[0260] A computing device can perform a method comprising multiple operations. The computing device can be an encoder. The computing device can determine attribute information of a reconstructed geometry of a point cloud frame based on attribute information of the geometry of the point cloud frame. The computing device can determine an attribute predictor of the attribute information of the reconstructed geometry based on projecting the attributes of a reference point cloud frame for the attributes onto the reconstructed geometry. The computing device can encode the attribute information of the reconstructed geometry based on the attribute predictor. The computing device can determine the attribute information of the reconstructed geometry by mapping the attributes of the geometry of the point cloud frame to the reconstructed geometry, wherein the attribute predictor can further determine the attribute based on the mapped attributes. The attribute can be color, and the mapped attribute can be determined based on recoloring. The mapping attribute of each point of the reconstructed geometry can be determined based on a nearest neighbor search from the geometry of the point cloud frame to the nearest point of the reconstructed geometry. The computing device can encode the attribute by determining a residual attribute based on the difference between the attributes of the reconstructed geometry and the attribute predictor; and encoding the residual attribute in a bit stream. The computing device can encode residual attributes by: determining transformed coefficients based on applying an intra-frame transform to the residual attributes; and entropy encoding the transformed coefficients corresponding to the residual attributes in the bitstream. The computing device can quantize the transformed coefficients, which are then entropy encoded. Residual attributes can be encoded and decoded based on a prediction-enhanced (prediction-boosted) transform scheme. A reference point cloud frame for the attribute can be determined from a coded reference point cloud frame. The coded reference point cloud frame can be used to encode or decode the geometry of the point cloud frame. A motion-compensated point cloud frame can be determined based on motion compensation of the coded reference point cloud frame to determine the motion-compensated reference point cloud frame. The coded reference point cloud frame can be motion-compensated by motion vectors. The computing device can encode the motion vectors and / or the indications of the coded reference point cloud frame. The computing device can determine the motion vectors based on the difference between the reconstructed geometry and the geometry adjusted by the motion vectors of the coded reference point cloud frame. The computing device can determine an attribute predictor by smoothing the projection attributes, the predicted attributes being determined from the smoothed projection attributes. The attribute predictor can include a corresponding attribute predictor for each corresponding point in the reconstructed geometry, based on the projection attributes corresponding to that point in the projection attributes. The computing device can encode (or signal) inter-frame residual activation flags indicating whether attributes are encoded based on the attribute predictor in the bit stream. The inter-frame residual activation flags can be encoded based on the projection quality of the attributes associated with the reconstructed geometry. Attribute information of the reconstructed geometry can be encoded using the attribute predictor based on the prediction quality of the attributes associated with the reconstructed geometry. The projection quality can be determined based on the projection distance of the attributes associated with the decoded geometry. The inter-frame residual activation flags can be associated with a spatial region of a point cloud frame.A computing device can encode the geometry of a point cloud frame and determine a reconstructed geometry based on decoding the encoded geometry. The computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to decode a point cloud frame. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0261] A computing device can perform a method including multiple operations. The computing device may include a decoder. The computing device can decode the geometry of a point cloud frame to determine a reconstructed geometry of the point cloud frame. The computing device can determine an attribute predictor of the attributes of the reconstructed geometry based on projecting the attributes of a reference point cloud frame for the attribute onto the reconstructed geometry. The computing device can decode attribute information of the reconstructed geometry based on the attribute predictor. The computing device can decode attributes associated with the reconstructed geometry by: decoding residual attributes from a bitstream that indicate the difference between the attributes of the reconstructed geometry and the attribute predictor; and determining the decoded attribute based on adding the attribute predictor to the decoded residual attribute. The computing device can decode the residual attribute by: entropy decoding of transform coefficients corresponding to the residual attribute from a bitstream; and determining the residual attribute based on applying an inverse intra-frame transform to the decoded transform coefficients. The computing device can dequantize the transform coefficients, and the residual attribute is determined by applying an inverse intra-frame transform to the dequantized transform coefficients. The reference point cloud frame for the attribute can be determined from a coded reference point cloud frame. A coded reference point cloud frame can be used to encode or decode the geometry of the point cloud frame. A reference point cloud frame for an attribute can be determined based on motion compensation performed on the coded reference point cloud frame to determine a motion-compensated point cloud frame. The coded reference point cloud frame is motion-compensated by motion vectors. The computing device can encode the motion vectors and / or indications of the coded reference point cloud frame. The computing device can determine the motion vectors based on the difference between the reconstructed geometry and the geometry adjusted by the motion vectors of the coded reference point cloud frame. The computing device can determine an attribute predictor, including smoothed projection attributes, from which the predicted attributes are determined. The attribute predictor can include a corresponding attribute predictor for each corresponding point in the reconstructed geometry, based on the projection attributes corresponding to said point in the projection attributes. The computing device can receive from the bit stream an inter-frame residual activation flag indicating whether the attribute is coded based on the attribute predictor, wherein determining the attribute predictor can be based on the inter-frame residual activation flag. The attribute predictor can be used to decode attribute information of the reconstructed geometry based on the prediction quality of the attributes associated with the reconstructed geometry. Projection quality can be determined based on the projection distance of attributes associated with the decoded geometry. Inter-frame residual activation flags can be associated with spatial regions of point cloud frames. The computing device can decode attributes associated with the reconstructed geometry by: applying an intra-frame transform to an attribute predictor; decoding transformed residual attributes from the bitstream that indicate the difference between the transformed attributes of the reconstructed geometry and the transformed attribute predictor; determining transformed attributes based on adding the transformed attribute predictor to the decoded transformed residual attributes; and determining decoded attributes based on applying an inverse intra-frame transform to the determined transformed attributes.A computing device can decode transformed residual properties by: entropy decoding of transformed coefficients corresponding to the transformed residual properties from a bitstream; and dequantizing the transformed coefficients to determine the transformed residual properties. The computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to encode point cloud frames. A computer-readable medium may store instructions that, when executed, enable the performance of the described methods, additional operations, and / or include additional elements.
[0262] In the following text, various features will be highlighted in a set of numbered clauses or paragraphs. These features should not be construed as limitations on the invention or inventive concept, but are merely highlights of certain features described herein, without implying a particular order of importance or relevance of such features.
[0263] Clause 1. A method comprising determining the reconstructed geometry of a point cloud frame by a decoder based on geometric information of the point cloud frame associated with content.
[0264] Clause 2. The method of claim 1 further comprises determining an attribute predictor associated with the reconstructed geometry based on projecting the attributes of a reference point cloud frame onto the reconstructed geometry.
[0265] Clause 3. The method according to any one of Clauses 1 to 2 further includes decoding the attribute information of the reconstructed geometry based on the attribute predictor.
[0266] Clause 4. The method according to any one of Clauses 1 to 3, wherein decoding the attribute information of the reconstructed geometry comprises: decoding a residual attribute from a bitstream that indicates a difference between the attribute information of the reconstructed geometry and the attribute predictor; and determining the attribute information of the reconstructed geometry based on the attribute predictor and the residual attribute.
[0267] Clause 5. The method according to any one of Clauses 1 to 4, wherein the decoding of the residual attribute comprises: decoding transform coefficients corresponding to the residual attribute from the bitstream; applying an inverse intra-frame transform to the decoded transform coefficients.
[0268] Clause 6. The method according to any one of Clauses 1 to 5 further comprises: decoding the transformed coefficients corresponding to the residual property from the bit stream; and dequantizing the decoded transformed coefficients.
[0269] Clause 7. The method according to any one of Clauses 1 to 6, wherein decoding the attribute information of the reconstructed geometry comprises: determining a transformed attribute predictor by applying an intra-frame transform to the attribute predictor; determining a transformed residual attribute indicating the difference between the transformed attribute of the reconstructed geometry and the transformed attribute predictor; determining the transformed attribute of the reconstructed geometry based on the transformed attribute predictor and the decoded transformed residual attribute; and applying an inverse intra-frame transform to the transformed attribute of the reconstructed geometry.
[0270] Clause 7. The method according to any one of Clauses 1 to 6, wherein determining the transformed residual property comprises: decoding the transformed coefficients corresponding to the transformed residual property; and dequantizing the transformed coefficients.
[0271] Clause 8. The method according to any one of Clauses 1 to 7, wherein the intra-frame transform comprises at least one of: adaptive DCT; RAHT transform; or Haar transform.
[0272] Clause 9. The method according to any one of Clauses 1 to 8, wherein a coded reference point cloud frame associated with the reference point cloud frame is used to decode the geometric information of the point cloud frame.
[0273] Clause 10. The method according to any one of Clauses 1 to 9, wherein determining the attribute predictor includes smoothing the projection attributes of the reference point cloud frame.
[0274] Clause 11. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of Clauses 1 to 10.
[0275] Clause 12. A system comprising: a computing device configured to perform the method according to any one of Clauses 1 to 10; and a second computing device configured to encode the point cloud frame.
[0276] Clause 13. A computer-readable medium storing instructions that, when executed, cause to perform the method according to any one of Clauses 1 to 10.
[0277] Clause 14. A method comprising determining, by an encoder, the reconstructed geometry of a point cloud frame associated with content.
[0278] Clause 15. The method according to Clause 14 further includes determining an attribute predictor associated with the reconstructed geometry based on projecting the attributes of a reference point cloud frame onto the reconstructed geometry.
[0279] Clause 16. The method according to any one of Clauses 14 to 15 further includes encoding attribute information associated with the reconstructed geometry based on the attribute predictor.
[0280] Clause 17. The method according to any one of Clauses 14 to 16, wherein encoding the attribute information comprises: determining residual attributes based on the difference between the attributes of the reconstructed geometry and the attribute predictor; and encoding the residual attributes into a bitstream associated with the point cloud frame.
[0281] Clause 18. The method according to any one of Clauses 14 to 17, wherein encoding the residual attribute comprises: determining transformed coefficients based on applying an intra-frame transform to the residual attribute; and entropy encoding the transformed coefficients corresponding to the residual attribute in the bitstream.
[0282] Clause 19. The method according to any one of Clauses 14 to 18 further includes quantizing the transformed coefficients prior to the entropy encoding.
[0283] Clause 20. The method according to any one of Clauses 14 to 19, wherein the determination of the attribute predictor is further based on mapping the attributes of the geometry to the reconstructed geometry.
[0284] Clause 21. The method according to any one of Clauses 14 to 20, wherein the attribute information of the reconstructed geometry includes color.
[0285] Clause 22. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform: the method according to any one of claims 14 to 21.
[0286] Clause 23. A system comprising: a computing device configured to perform the method according to any one of claims 14 to 21; and a second computing device configured to decode the point cloud frame.
[0287] Clause 24. A computer-readable medium storing instructions that, when executed, cause the method according to any one of claims 14 to 21 to be performed.
[0288] Clause 25. A method comprising a decoder determining an attribute predictor associated with a geometry based on projecting attributes of a reference point cloud frame onto the geometry of the point cloud frame.
[0289] Clause 26. The method according to Clause 25 further includes receiving from the bit stream a residual attribute indicating the difference between the attribute predictor and the attribute of the geometry.
[0290] Clause 27. The method according to any one of Clauses 25 to 26 further includes using the attribute predictor to decode the residual attribute to determine an attribute associated with the geometry of the point cloud frame.
[0291] Clause 28. The method according to any one of Clauses 25 to 27, wherein receiving the residual attribute comprises: receiving transformed coefficients corresponding to the residual attribute; and applying an inverse intra-frame transform to the decoded transformed coefficients.
[0292] Clause 29. The method according to any one of Clauses 25 to 28, wherein the intra-frame transform comprises at least one of: adaptive DCT; RAHT transform; or Haar transform.
[0293] Clause 30. The method according to any one of Clauses 25 to 29, wherein a coded reference point cloud frame associated with the reference point cloud frame is used to decode the geometric information of the point cloud frame.
[0294] Clause 31. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of Clauses 25 to 30.
[0295] Clause 32. A system comprising: a computing device configured to perform the method according to any one of Clauses 25 to 30; and a second computing device configured to encode the point cloud frame.
[0296] Clause 33. A computer-readable medium storing instructions that, when executed, cause the method according to any one of Clauses 25 to 30 to be performed.
[0297] Clause 34. A method comprising an encoder determining an attribute predictor associated with a geometry based on projecting attributes of a reference point cloud frame onto the geometry of the point cloud frame.
[0298] Clause 35. The method according to Clause 34 further includes encoding residual attributes in a bitstream that indicate the difference between the attribute predictor and attribute information of the geometry.
[0299] Clause 36. The method according to any one of Clauses 34 to 35, wherein encoding the residual attribute comprises: determining transformed coefficients based on applying an inverse intra-frame transform to the residual attribute; and entropy encoding the transformed coefficients corresponding to the residual attribute in the bitstream.
[0300] Clause 37. The method according to any one of Clauses 34 to 36, wherein the inverse intra-frame transform comprises at least one of: inverse adaptive DCT; inverse RAHT transform; or inverse Haar transform.
[0301] Clause 38. The method according to any one of Clauses 34 to 37, wherein a coded reference point cloud frame associated with the reference point cloud frame is used to encode the geometric information of the current point cloud frame.
[0302] Clause 39. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform the method according to any one of claims 34 to 38.
[0303] Clause 40. A system comprising: a computing device configured to perform the method according to any one of claims 34 to 38; and a second computing device configured to decode the point cloud frame.
[0304] Clause 41. A computer-readable medium storing instructions that, when executed, cause the method according to any one of claims 34 to 38 to be performed.
[0305] Clause 42. A method comprising determining attribute information of a reconstructed geometry of a point cloud frame based on attribute information of the geometry of the point cloud frame.
[0306] Clause 43. The method of claim 42 further includes an attribute predictor that determines attribute information of the reconstructed geometry based on projecting attributes of a reference point cloud frame for the attribute onto the reconstructed geometry.
[0307] Clause 44. The method according to any one of Clauses 42 to 43 further includes encoding the attribute information of the reconstructed geometry based on the attribute predictor.
[0308] Clause 45. The method according to any one of Clauses 42 to 44 further includes determining the attribute information of the reconstructed geometry by: mapping the attributes of the geometry of the point cloud frame to the reconstructed geometry, and wherein the attribute predictor further determines the attribute information based on the mapped attributes.
[0309] Clause 46. The method according to any one of Clauses 42 to 45, wherein the attribute is color and the mapping attribute is determined based on recoloring.
[0310] Clause 47. The method according to any one of Clauses 42 to 46, wherein the mapping attribute of each point of the reconstructed geometry is determined based on a nearest neighbor search from the geometry of the point cloud frame to the nearest point of the reconstructed geometry.
[0311] Clause 48. The method according to any one of Clauses 42 to 47, wherein encoding the attribute comprises: determining a residual attribute based on the difference between the attribute of the reconstructed geometry and the attribute predictor; and encoding the residual attribute in a bitstream.
[0312] Clause 49. The method according to any one of Clauses 42 to 48, wherein encoding the residual attribute comprises: determining transformed coefficients based on applying an intra-frame transform to the residual attribute; and entropy encoding the transformed coefficients corresponding to the residual attribute in the bitstream.
[0313] Clause 50. The method according to any one of Clauses 42 to 49 further includes quantizing the transformed coefficients, wherein the quantized transformed coefficients are entropy encoded.
[0314] Clause 51. The method according to any one of Clauses 42 to 50, wherein the residual property is encoded and decoded based on a prediction-enhanced (predictive boost) transformation scheme.
[0315] Clause 52. The method according to any one of Clauses 42 to 51, wherein the reference point cloud frame for the attribute is determined from a coded reference point cloud frame.
[0316] Clause 53. The method according to any one of Clauses 42 to 52, wherein the coded reference point cloud frame is used to encode or decode the geometry of the point cloud frame.
[0317] Clause 54. The method according to any one of Clauses 42 to 53, wherein the reference point cloud frame for an attribute is determined based on motion compensation of the coded reference point cloud frame to determine a motion-compensated point cloud frame.
[0318] Clause 55. The method according to any one of Clauses 42 to 54, wherein the coded reference point cloud frame is motion-compensated by motion vectors.
[0319] Clause 56. The method according to any one of Clauses 42 to 55 further includes encoding the motion vector and / or the indication of the coded reference point cloud frame.
[0320] Clause 57. The method according to any one of Clauses 42 to 56 further includes determining the motion vector based on the difference between the reconstructed geometry and the geometry adjusted by the motion vector of the coded reference point cloud frame.
[0321] Clause 58. The method according to any one of Clauses 42 to 57, wherein determining the attribute predictor comprises: smoothing the projected attribute, the predicted attribute being determined from the smoothed projected attribute.
[0322] Clause 59. The method according to any one of Clauses 42 to 58, wherein the attribute predictor includes a corresponding attribute predictor for each corresponding point in the points of the reconstructed geometry, based on the projection attributes corresponding to the point in the projection attributes.
[0323] Clause 60. The method according to any one of Clauses 42 to 59 further includes encoding an inter-frame residual activation flag indicating whether the attribute is encoded based on the attribute predictor in the bit stream (or signaling the inter-frame residual activation flag in the bit stream).
[0324] Clause 61. The method according to any one of Clauses 42 to 60, wherein the inter-frame residual activation flag is encoded based on the projection quality of the attribute associated with the reconstructed geometry.
[0325] Clause 62. The method according to any one of Clauses 42 to 61, wherein the attribute information of the reconstructed geometry is encoded using the attribute predictor based on the prediction quality of the attribute associated with the reconstructed geometry.
[0326] Clause 63. The method according to any one of Clauses 42 to 62, wherein the projection quality is determined based on the projection distance of the attribute associated with the decoded geometry.
[0327] Clause 64. The method according to any one of Clauses 42 to 63, wherein the inter-frame residual activation flag is associated with a spatial region of the point cloud frame.
[0328] Clause 65. The method according to any one of Clauses 42 to 64 further comprises: encoding the geometry of the point cloud frame; and determining the reconstructed geometry based on decoding the encoded geometry.
[0329] Clause 66. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform the method according to any one of claims 42 to 65.
[0330] Clause 67. A system comprising: a computing device configured to perform the method according to any one of claims 42 to 65; and a second computing device configured to decode the point cloud frame.
[0331] Clause 68. A computer-readable medium storing instructions that, when executed, cause the method according to any one of claims 42 to 65 to be performed.
[0332] Clause 69. A method comprising decoding the geometry of a point cloud frame to determine a reconstructed geometry of the point cloud frame.
[0333] Clause 70. The method according to Clause 69 further includes an attribute predictor that determines the attributes of the reconstructed geometry based on projecting the attributes of a reference point cloud frame for the attributes onto the reconstructed geometry.
[0334] Clause 71. The method according to any one of Clauses 69 to 70 further includes decoding the attribute information of the reconstructed geometry based on the attribute predictor.
[0335] Clause 72. The method according to any one of Clauses 69 to 71, wherein decoding the attribute associated with the reconstructed geometry comprises: decoding a residual attribute from a bitstream that indicates a difference between the attribute of the reconstructed geometry and the attribute predictor; and determining the decoded attribute based on adding the attribute predictor to the decoded residual attribute.
[0336] Clause 73. The method according to any one of Clauses 69 to 72, wherein the decoding of the residual attribute comprises: entropy decoding of transform coefficients corresponding to the residual attribute from the bitstream; and determining the residual attribute based on applying an inverse intra-frame transform to the decoded transform coefficients.
[0337] Clause 74. The method according to any one of Clauses 69 to 73 further includes dequantizing the transformed coefficients, wherein the residual property is determined by applying the inverse intra-frame transform to the dequantized transformed coefficients.
[0338] Clause 75. The method according to any one of Clauses 69 to 74, wherein the reference point cloud frame for the attribute is determined from a coded reference point cloud frame.
[0339] Clause 76. The method according to any one of Clauses 69 to 75, wherein the coded reference point cloud frame is used to encode or decode the geometry of the point cloud frame.
[0340] Clause 77. The method according to any one of Clauses 69 to 76, wherein the reference point cloud frame for an attribute is determined based on motion compensation of the coded reference point cloud frame to determine a motion-compensated point cloud frame.
[0341] Clause 78. The method according to any one of Clauses 69 to 77, wherein the coded reference point cloud frame is motion-compensated by motion vectors.
[0342] Clause 79. The method according to any one of Clauses 69 to 78 further includes encoding the motion vector and / or the indication of the coded reference point cloud frame.
[0343] Clause 80. The method according to any one of Clauses 69 to 79 further includes determining the motion vector based on the difference between the reconstructed geometry and the geometry adjusted by the motion vector of the coded reference point cloud frame.
[0344] Clause 81. The method according to any one of Clauses 69 to 80, wherein determining the attribute predictor includes smoothing the projected attribute, and the predicted attribute is determined from the smoothed projected attribute.
[0345] Clause 82. The method according to any one of Clauses 69 to 81, wherein the attribute predictor includes a corresponding attribute predictor for each corresponding point in the points of the reconstructed geometry, based on the projection attributes corresponding to the point in the projection attributes.
[0346] Clause 83. The method according to any one of Clauses 69 to 82 further includes receiving from the bit stream an inter-frame residual activation flag indicating whether the attribute is written based on the attribute predictor, wherein the determination that the attribute predictor can be based on the inter-frame residual activation flag.
[0347] Clause 84. The method according to any one of Clauses 69 to 83, wherein the attribute information of the reconstructed geometry is decoded using the attribute predictor based on the prediction quality of the attribute associated with the reconstructed geometry.
[0348] Clause 85. The method according to any one of Clauses 69 to 84, wherein the projection quality is determined based on the projection distance of the attribute associated with the decoded geometry.
[0349] Clause 86. The method according to any one of Clauses 69 to 85, wherein the inter-frame residual activation flag is associated with a spatial region of the point cloud frame.
[0350] Clause 87. The method according to any one of Clauses 69 to 86, wherein decoding the attribute associated with the reconstructed geometry comprises: applying an intra-frame transform to the attribute predictor; decoding a transformed residual attribute from a bitstream that indicates the difference between the transformed attribute of the reconstructed geometry and the transformed attribute predictor; determining the transformed attribute based on adding the transformed attribute predictor to the decoded transformed residual attribute; and determining the decoded attribute based on applying an inverse intra-frame transform to the determined transformed attribute.
[0351] Clause 88. The method according to any one of Clauses 69 to 87, wherein the decoding of the transformed residual attribute comprises: entropy decoding of transformed coefficients corresponding to the transformed residual attribute from the bit stream; and dequantizing the transformed coefficients to determine the transformed residual attribute.
[0352] Clause 89. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by said one or more processors, cause the computing device to perform the method according to any one of claims 69 to 88.
[0353] Clause 90. A system comprising: a computing device configured to perform the method according to any one of claims 69 to 88; and a second computing device configured to encode the point cloud frame.
[0354] Clause 91. A computer-readable medium storing instructions that, when executed, cause the method according to any one of claims 69 to 88 to be performed.
[0355] One or more examples in this document can be described as processes that can be depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, and / or block diagrams. Although a flowchart can describe operations as a continuous process, one or more of the operations can be executed in parallel or simultaneously. The order of the operations shown can be rearranged. A process can be terminated when its operations are completed, but may have additional steps not shown in the diagram. A process can correspond to a method, function, program, subroutine, subroutines, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.
[0356] The operations described herein can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments for performing necessary tasks (e.g., computer program products) can be stored on a computer-readable or machine-readable medium. A processor can perform the necessary tasks. The features of this disclosure can be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. Implementing a hardware state machine to perform the functions described herein will also be apparent to those skilled in the art.
[0357] One or more features described herein may be implemented in computer-usable data and / or computer-executable instructions, as in one or more program modules, which are executed by one or more computers or other devices. Typically, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type when executed by a processor in a computer or other data processing device. Computer-executable instructions may be stored on one or more computer-readable media, such as hard disks, optical disks, removable storage media, solid-state drives, RAM, etc. The functionality of program modules may be combined or distributed as needed. The functionality may be implemented, in whole or in part, in firmware or hardware equivalents, such as integrated circuits, field-programmable gate arrays (FPGAs), etc. One or more features described herein may be implemented more efficiently using specific data structures, and such data structures are contemplated within the scope of the computer-executable instructions and computer-usable data described herein. Computer-readable media may include, but are not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored but do not include carrier waves and / or transient electronic signals propagated wirelessly or via wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as CDs or DVDs, flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions that can represent any combination of programs, functions, subroutines, routines, subroutines, modules, software packages, classes or instructions, data structures, or program statements. Code segments can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0358] Non-transitory tangible computer-readable media may include instructions executable by one or more processors configured to cause the operations described herein. Articles of manufacture may include non-transitory tangible computer-readable machine-accessible media having instructions encoded thereon for enabling programmable hardware to allow devices (e.g., encoders, decoders, transmitters, receivers, etc.) to perform the operations described herein. Devices, or one or more devices such as in a system, may include one or more processors, memories, interfaces, etc.
[0359] The communication described herein can be determined, generated, sent, and / or received using any number of messages, information elements, fields, parameters, values, indications, information, bits, etc. While this document may use any of the terms / phrases message, information element, field, parameter, value, indication, information, bit, etc., to describe one or more examples, those skilled in the art will understand that any one or more of these terms, including other such terms, can be used to perform such communication. For example, one or more parameters, fields, and / or information elements (IEs) may include one or more information objects, values, and / or any other information. An information object may include one or more other objects. At least some (or all) parameters, fields, IEs, etc., may be used and may be interchangeable depending on the context. Where a meaning or definition is given, such meaning or definition shall prevail.
[0360] One or more elements in the examples described herein can be implemented as modules. A module can be an element that performs a defined function and / or has a defined interface to other elements. Modules can be implemented as hardware, software combined with hardware, firmware, wet hardware (e.g., hardware with biological elements), or a combination thereof, all of which can be behaviorally equivalent. For example, a module can be implemented as software routines written in a computer language configured to be executed by a hardware machine (such as C, C++, Fortran, Java, Basic, Matlab, etc.) or a modeling / simulation program (such as Simulink, Stateflow, GNU Octave, or LabVIEW MathScript). Alternatively or alternatively, modules can be implemented using physical hardware that incorporates discrete or programmable analog, digital, and / or quantum hardware. Examples of programmable hardware can include: computers, microcontrollers, microprocessors, application-specific integrated circuits (ASICs); field-programmable gate arrays (FPGAs); and / or complex programmable logic devices (CPLDs). Computers, microcontrollers, and / or microprocessors can be programmed using languages such as assembly, C, C++, etc. Hardware description languages (HDLs) such as VHSIC (VHDL) or Verilog are typically used to program FPGAs, ASICs, and CPLDs. These HDLs can configure connections between internal hardware modules with limited functionality on a programmable device. The techniques mentioned above can be combined to achieve the desired functional modules.
[0361] One or more operations described herein may be conditional. For example, one or more operations may be performed if certain criteria are met in a computing device, communication device, encoder, decoder, network, or a combination thereof. Example criteria may be based on one or more conditions, such as device configuration, traffic load, initial system settings, packet size, service characteristics, or a combination thereof. Various examples may be used if the one or more criteria are met. Any part of the examples described herein may be implemented in any order and based on any conditions.
[0362] Although examples have been described above, features and / or steps of those examples can be combined, divided, omitted, rearranged, modified, and / or expanded in any desired manner. Various changes, modifications, and improvements will readily occur to those skilled in the art. While not explicitly stated herein, such changes, modifications, and improvements are intended to be part of this specification and are intended to be within the spirit and scope of this specification. Therefore, the above description is illustrative only and not restrictive.
Claims
1. A method comprising: The decoder determines the reconstructed geometry of the point cloud frame based on the geometric information of the point cloud frame associated with the content. An attribute predictor associated with the reconstructed geometry is determined by projecting the attributes of a reference point cloud frame onto the reconstructed geometry. as well as The attribute information of the reconstructed geometry is decoded based on the attribute predictor.
2. The method according to claim 1, wherein decoding the attribute information of the reconstructed geometry includes: Decode the residual attributes from the bitstream that indicate the difference between the attribute information of the reconstructed geometry and the attribute predictor; as well as The attribute information of the reconstructed geometry is determined based on the attribute predictor and the residual attributes.
3. The method of claim 2, wherein the decoding of the residual attribute comprises: Decode the transformed coefficients corresponding to the residual properties from the bit stream; as well as The inverse intra-frame transform is applied to the decoded transform coefficients.
4. The method according to any one of claims 2 to 3, further comprising: Decode the transformed coefficients corresponding to the residual properties from the bit stream; as well as The decoded transform coefficients are dequantized.
5. The method according to any one of claims 1 to 4, wherein decoding the attribute information of the reconstructed geometry comprises: The transformed attribute predictor is determined by applying intra-frame transformation to the attribute predictor. Determine the transformed residual attribute that indicates the difference between the transformed attribute information of the reconstructed geometry and the transformed attribute predictor; The transformed attribute information of the reconstructed geometry is determined based on the transformed attribute predictor and the decoded transformed residual attributes; as well as The inverse intra-frame transform is applied to the transformed attribute information of the reconstructed geometry.
6. The method of claim 5, wherein determining the transformed residual property comprises: Decode the transformed coefficients corresponding to the transformed residual properties; as well as The transformed coefficients are dequantized.
7. The method according to any one of claims 5 to 6, wherein the intra-frame transformation comprises at least one of the following: Adaptive DCT; RAHT transform; or Haar transform.
8. The method according to any one of claims 1 to 7, wherein a coded reference point cloud frame associated with the reference point cloud frame is used to decode the geometric information of the point cloud frame.
9. The method according to any one of claims 1 to 8, wherein determining the attribute predictor includes smoothing the projection attributes of the reference point cloud frame.
10. A method comprising: The decoder determines the attribute predictor associated with the geometry based on the properties of the reference point cloud frame projected onto the geometry of the point cloud frame; Receive residual attributes from the bitstream that indicate the difference between the attributes of the attribute predictor and the attributes of the geometry; as well as The attribute predictor is used to decode the residual attributes to determine the attributes associated with the geometry of the point cloud frame.
11. The method of claim 10, wherein receiving the residual property comprises: Receive the transformed coefficients corresponding to the residual properties; as well as The inverse intra-frame transform is applied to the transformed coefficients.
12. The method of claim 11, wherein the intra-frame transformation comprises at least one of the following: Adaptive DCT; RAHT transform; or Haar transform.
13. A computing device, comprising: One or more processors; And a memory that stores instructions, which, when executed by the one or more processors, cause the computing device to perform the method according to any one of claims 1 to 12.
14. A system comprising: A computing device configured to perform the method according to any one of claims 1 to 12; as well as A second computing device is configured to encode the point cloud frame.
15. A computer-readable medium storing instructions that, when executed, cause the method according to any one of claims 1 to 12 to be performed.