Dual motion field for coding geometry and attributes of point clouds
By combining G-PCC and V-PCC with RAHT transformation, efficient encoding and decoding of point cloud data is achieved, solving the problem of excessively large point cloud data size, improving storage and transmission efficiency, and ensuring data reliability and quality in AR/VR and autonomous driving.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- COMCAST CABLE COMM LLC
- Filing Date
- 2024-07-12
- Publication Date
- 2026-06-05
Smart Images

Figure CN122162381A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefits of U.S. Provisional Application No. 63 / 526,761, filed July 14, 2023, and U.S. Provisional Application No. 63 / 526,537, filed July 13, 2023. The aforementioned applications are hereby incorporated in their entirety by reference. Background Technology
[0003] Objects or scenes can be described using volumetric visual data consisting of a series of points. Points can be stored in point cloud format, which includes a set of points in three-dimensional space. Because point cloud data can be quite large, transmitting and processing point cloud data may require data compression schemes specifically designed for the unique characteristics of point cloud data. Summary of the Invention
[0004] The following summary presents a simplified overview of certain features. This summary is neither a comprehensive overview nor intended to identify important or key elements.
[0005] Point cloud information associated with content can include geometric information and attribute information (e.g., the color or texture of the geometry). Attribute information can be encoded separately from geometric information. A reference point cloud frame can be selected to predict the attributes of the current point cloud frame. The geometric information associated with the point cloud frame and the attributes associated with the reconstructed geometry of the point cloud frame can be encoded. The reconstructed geometry of the point cloud frame can be determined by decoding the geometry associated with the point cloud frame. Attributes associated with the reconstructed geometry can be decoded based on attribute motion vectors. Encoding and / or decoding attributes can reduce the coding cost (e.g., bit rate) and / or distortion of inter-frame prediction.
[0006] These and other features and advantages are described in more detail below. Attached Figure Description
[0007] Examples of several embodiments of the various embodiments of this disclosure are described herein with reference to the accompanying drawings.
[0008] Figure 1 An example point cloud coding system is shown.
[0009] Figure 2 An example of Morton order is shown.
[0010] Figure 3 An example scan order is shown.
[0011] Figure 4 An example neighborhood of a cuboid is shown for entropy writing coding of the occupancy of a sub-cuboid.
[0012] Figure 5 An example of the dynamically decreasing function DR that can be used in dynamic OBUF is shown.
[0013] Figure 6 An example method for writing code to the occupancy of a cuboid using dynamic OBUF is shown.
[0014] Figure 7 An example of an occupied cuboid is shown.
[0015] Figure 8A An example cuboid corresponding to a TriSoup node is shown.
[0016] Figure 8B An example refinement of the TriSoup model is shown.
[0017] Figure 9 An example of voxelization is shown.
[0018] Figure 10 An example encoding method using inter-frame prediction is shown.
[0019] Figure 11 An example method for encoding point cloud attributes based on predictive transformations is shown.
[0020] Figure 12 An example method for decoding point cloud attributes based on predictive transformation is shown.
[0021] Figure 13 An example method for encoding point cloud attributes based on prediction boosting transformations is shown.
[0022] Figure 14 An example method for decoding point cloud properties based on prediction boosting transformations is shown.
[0023] Figure 15 An example of the Region Adaptive Hierarchical Transformation (RAHT) transformation for an octree is shown.
[0024] Figure 16 Another example of the RAHT transformation for an octree is shown.
[0025] Figure 17 An example method for encoding the current point cloud frame is shown.
[0026] Figure 18 An example method for decoding the current point cloud frame is shown.
[0027] Figure 19 An example method for encoding the current point cloud frame is shown.
[0028] Figure 20An example method for decoding the current point cloud frame is shown.
[0029] Figure 21 An example method for encoding the attributes of the current point cloud frame is shown.
[0030] Figure 22 An example method for decoding the properties of the current point cloud frame is shown.
[0031] Figure 23 An example method for encoding the attributes of a point cloud frame is shown.
[0032] Figure 24 An example method for decoding attribute information is shown.
[0033] Figure 25 An example method for encoding the attributes of a point cloud frame is shown.
[0034] Figure 26 An example method for decoding attribute information is shown.
[0035] Figure 27 An example method for encoding residual properties is shown.
[0036] Figure 28 An example method for decoding residual properties is shown.
[0037] Figure 29 An example method for encoding the current point cloud frame is shown.
[0038] Figure 30 An example method for decoding the current point cloud frame is shown.
[0039] Figure 31 An example method for encoding attribute motion vectors is shown.
[0040] Figure 32 An example method for decoding attribute motion vectors is shown.
[0041] Figure 33 An example method for encoding point cloud frames is shown.
[0042] Figure 34 An example method for decoding point cloud frames is shown.
[0043] Figure 35 An example computer system in which this disclosure can be implemented is shown.
[0044] Figure 36Example elements of a computing device that can be used to implement any of the various devices described herein are shown. Detailed Implementation
[0045] The accompanying figures and description provide examples. It should be understood that the examples shown and / or described in the figures are non-exclusive, and the features shown and described can be practiced in other examples. Examples of operation for point cloud or point cloud sequence encoding or decoding systems are provided. More specifically, the techniques disclosed herein can relate to point cloud compression, such as that used in encoding and / or decoding apparatuses and / or systems.
[0046] At least some visual data can use a series of points to describe objects or scenes in content and / or media. Each point may include a position in two-dimensional (x and y) form and one or more optional attributes, such as color. Volumetric visual data can add another positional dimension to these visual data. For example, volumetric visual data can use a series of points to describe objects or scenes in content and / or media, each point may include a position in three-dimensional (x, y, and z) form and one or more optional attributes, such as color, reflectivity, timestamp, etc. For example, volumetric visual data can provide a more immersive way to experience visual data than the aforementioned at least some visual data. For example, an object or scene described by volumetric visual data can be viewed from any (or more) angles, while an object or scene described by the aforementioned at least some visual data can typically only be viewed from the angle from which the object or scene is captured or rendered. As a representation format for visual data (e.g., volumetric visual data, 3D video data, etc.), point clouds are universal because they can represent all types of three-dimensional (3D) objects, scenes, and visual content. Point clouds are well-suited for a wide range of applications, including but not limited to: film post-production, real-time 3D immersive media or telepresence, extended reality, free-view video, geographic information systems, autonomous driving, 3D mapping, visualization, medicine, multi-view replay, and real-time light detection and ranging (LiDAR) data acquisition.
[0047] As explained in this article, volumetric visual data can be used in many applications, including extended reality (XR). XR encompasses various types of immersive technologies, including augmented reality (AR), virtual reality (VR), and mixed reality (MR). Sparse volumetric visual data can be used in the automotive industry to represent three-dimensional (3D) maps (e.g., cartography) or as input to driver assistance systems. In the case of driver assistance systems, volumetric visual data can often be fed into driving decision-making algorithms. Volumetric visual data can be used to digitally store valuable objects. In applications for the protection of cultural heritage, the goal can be to maintain a representation of objects that may be threatened by natural disasters. For example, statues, vases, and temples can be fully scanned and stored as volumetric visual data with billions of samples. This use case for volumetric visual data may be particularly relevant to valuable objects in locations prone to earthquakes, tsunamis, and typhoons. Volumetric visual data can take the form of volumetric frames. A volumetric frame can describe an object or scene captured at a specific time instance. Volumetric visual data can also take the form of a sequence of volumetric frames (referred to as a volumetric sequence or volumetric video). A sequence of volumetric frames can describe an object or scene captured at multiple different time instances.
[0048] Volumetric visual data can be stored in various formats. Point clouds can include a collection of points in 3D space. Such points can be used to create meshes including vertices and polygons, or other forms of visual content. As described herein, point cloud data can take the form of point cloud frames that describe objects or scenes in content captured in a specific time instance. Point cloud data can also take the form of a sequence of point cloud frames (e.g., point cloud video). As further described herein, point cloud data can be generated from a source device (e.g., as described herein regarding...). Figure 1 The source device 102 encodes the point cloud data, outputting a bitstream containing the encoded point cloud data. The source device can encode the point cloud data based on point cloud compression coding, for example, geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding, or next-generation coding. The destination device (e.g., as described herein regarding...) Figure 1 The destination device 106 receives a bitstream containing point cloud data and decodes the bitstream containing point cloud data. The destination device can decode the point cloud data by performing point cloud decompression coding. Decompression coding can be the reverse process of point cloud compression coding. Point cloud decompression coding can include, for example, G-PCC coding. Decoding can be used to decompress the point cloud data for display and / or other forms of consumption (e.g., further analysis, storage, etc.). The destination device (or different devices) can include, for example, a renderer for rendering the decoded point cloud data. The renderer can output content, for example, by rendering the point cloud data. The renderer can output content, for example, by rendering the point cloud data along with other data (e.g., audio data).
[0049] One format for storing volumetric visual data can be a point cloud. A point cloud can comprise a collection of points in 3D space. Each point in a point cloud can include geometric information that can indicate the point's location in 3D space. For example, the geometric information can indicate the point's location in 3D space using, for example, three Cartesian coordinates (x, y, z) and / or spherical coordinates (r, φ, θ) (e.g., if acquired by a rotation sensor). The locations of points in a point cloud can be quantized according to spatial precision. Spatial precision can be the same or different in each dimension. The quantization process can create a grid in 3D space. One or more points residing within each sub-grid volume can be mapped to the coordinates of the sub-grid center, referred to as voxels. A voxel can be viewed as a 3D extension of a pixel corresponding to a 2D image grid coordinate. For example, similar to how a pixel is the smallest unit in an example of dividing 2D space (or a 2D image) into discrete, uniform (e.g., equal-sized) regions, a voxel can be the smallest volume unit in an example of dividing 3D space into discrete, uniform regions. Points in a point cloud can include one or more types of attribute information. Attribute information can indicate the properties of a point's visual appearance. For example, attribute information can indicate the point's texture (e.g., color), material type, transparency, reflectivity, surface normal, velocity, acceleration, timestamp indicating when the point was captured, or modality (e.g., running, walking, or flying). Points in a point cloud can include light field data in the form of multi-view related texture information. Light field data can be another type of optional attribute information.
[0050] Points in a point cloud can describe objects or scenes. For example, points in a point cloud can describe the external surfaces and / or internal structures of an object or scene. Objects or scenes can be generated synthetically by computer. Objects or scenes can be generated from captures of real-world objects or scenes. Geometric information of real-world objects or scenes can be obtained through 3D scanning and / or photogrammetry. 3D scanning can include different types of scanning, such as laser scanning, structured light scanning, and / or modulated light scanning. 3D scanning can obtain geometric information. 3D scanning can obtain geometric information, for example, by moving one or more laser heads, structured light cameras, and / or modulated light cameras relative to the scanned object or scene. Photogrammetry can obtain geometric information. Photogrammetry can obtain geometric information, for example, by triangulating the same features or points in 2D photographs at different spatial displacements. Point cloud data can be in the form of point cloud frames. Point cloud frames can describe objects or scenes captured at a specific time instance. Point cloud data can be in the form of point cloud frame sequences. Point cloud frame sequences can be referred to as point cloud sequences or point cloud videos. Point cloud frame sequences can describe objects or scenes captured at multiple different time instances.
[0051] In many applications, the data size of a point cloud frame or sequence of point clouds may be too large for storage and / or transmission (e.g., too big). For example, a single point cloud may include, for example, more than one million points or even billions of points. Each point may include geometric information and one or more optional types of attribute information. The geometric information of each point may include three Cartesian coordinates (x, y, z) and / or spherical coordinates (r, φ, θ), each Cartesian and / or spherical coordinate may be represented, for example, using at least 10 bits per component or 30 bits in total. The attribute information of each point may include a texture corresponding to multiple (e.g., three) color components (e.g., R, G, and B color components). Each color component may be represented, for example, using 8-10 bits per component or 24-30 bits in total. For example, a single point may include at least 54 bits of information, with at least 30 bits of geometric information and at least 24 bits of texture. If a point cloud frame includes one million such points, each point cloud frame may require 54 million bits or 54 megabits to represent. For dynamic point clouds that change over time, at a frame rate of 30 frames per second, a data rate of 1.32 gigabits per second may be required to send (e.g., transmit) the points of a point cloud sequence. The raw representation of the point cloud may require a large amount of data, and the practical deployment of point cloud-based technologies may require compression techniques that enable the storage and distribution of point clouds at a reasonable cost.
[0052] Encoding can be used to compress and / or reduce the data size of point cloud frames or sequences to provide more efficient storage and / or transmission. Decoding can be used to decompress compressed point cloud frames or sequences for display and / or other forms of consumption (e.g., other forms of consumption by machine learning-based devices, neural network-based devices, artificial intelligence-based devices, or other types of consumption by other types of machine-based processing algorithms and / or devices). For example, distribution to and visualization by end users on AR or VR glasses or any other 3D-enabled devices, point cloud compression may be lossy (introducing differences relative to the original data). Lossy compression can allow high compression ratios but may imply a trade-off between compression and visual quality perceived by the end user. Other frameworks, such as those used in medical applications or autonomous driving, may require lossless compression to avoid altering decisions obtained, for example, based on analysis of sent (e.g., transmitted) and decompressed point cloud frames.
[0053] Figure 1An example point cloud coding (e.g., encoding and / or decoding) system 100 is illustrated. The point cloud coding system 100 may include a source device 102, a transmission medium 104, and a destination device 106. The source device 102 may encode a point cloud sequence 108 into a bit stream 110 for more efficient storage and / or transmission. The source device 102 may store the bit stream 110 and / or send (e.g., transmit) the bit stream to the destination device 106 via the transmission medium 104. The destination device 106 may decode the bit stream 110 to display the point cloud sequence 108 or for other forms of consumption (e.g., further analysis, storage, etc.). The destination device 106 may receive the bit stream 110 from the source device 102 via the storage medium or the transmission medium 104. The source device 102 and the destination device 106 may include any number of different devices. Source device 102 and destination device 106 may include, for example, interconnected clusters of computer systems, servers, desktop computers, laptop computers, tablet computers, smartphones, wearable devices, televisions, cameras, video game consoles, set-top boxes, video streaming devices, vehicles (e.g., autonomous vehicles), or head-mounted displays that act as seamless resource pools (also known as computer clouds or cloud computing). Head-mounted displays may allow users to view VR, AR, or MR scenes and adjust the view of the scene, for example, based on the movement of the user's head. Head-mounted displays may be connected (e.g., tethered) to processing devices (e.g., servers, desktop computers, set-top boxes, or video game consoles) or may be completely independent.
[0054] Source device 102 may include point cloud source 112, encoder 114, and output interface 116. For example, to encode point cloud sequence 108 into bitstream 110, source device 102 may include point cloud source 112, encoder 114, and output interface 116. For example, point cloud source 112 may provide (e.g., generate) point cloud sequence 108 from captures of natural scenes and / or synthetically generated scenes. Synthetically generated scenes may be scenes including computer-generated graphics. Point cloud source 112 may include one or more point cloud capture devices, a point cloud archive including previously captured natural scenes and / or synthetically generated scenes, a point cloud feed interface for receiving captured natural scenes and / or synthetically generated scenes from a point cloud content provider, and / or a processor for generating synthetic point cloud scenes. Point cloud capture devices may include, for example, one or more laser scanning devices, structured light scanning devices, modulated light scanning devices, and / or passive scanning devices.
[0055] Point cloud sequence 108 may include a series of point cloud frames 124 (e.g., Figure 1(Example shown). Point cloud frames can describe objects or scenes captured at a specific time instance. Point cloud sequence 108 can achieve the impression of motion by continuously presenting point cloud frames 124 of point cloud sequence 108 using constant or variable time. Point cloud frames can include a set of points (e.g., voxels) 126 in 3D space. Each point 126 can include geometric information that can indicate the position of the point in 3D space. Geometric information can indicate, for example, the position of the point in 3D space using three Cartesian coordinates (x, y, z). One or more points 126 can include one or more types of attribute information. Attribute information can indicate the nature of the visual appearance of the point. For example, attribute information can indicate, for example, the texture (e.g., color) of the point, the material type of the point, the transparency information of the point, the reflectivity information of the point, the surface normal of the point, the velocity at the point, the acceleration at the point, a timestamp indicating when the point was captured, a modality indicating how the point was captured (e.g., running, walking, or flying), etc. One or more points 126 can include light field data, for example, in the form of multi-view related texture information. Light field data can be another type of optional attribute information. The color attribute information for one or more points 126 may include a lightness value and two chromaticity values. The lightness value may represent the brightness of the point (e.g., the lightness component Y). The chromaticity values may represent the blue and red components of the point, separate from the lightness (e.g., chromaticity components Cb and Cr). Other color attribute values may be represented, for example, based on different color schemes (e.g., RGB or a monochrome color scheme).
[0056] Encoder 114 can encode point cloud sequence 108 into bitstream 110. To encode point cloud sequence 108, encoder 114 can use one or more lossless or lossy compression techniques to reduce redundant information in point cloud sequence 108. To encode point cloud sequence 108, encoder 114 can use one or more prediction techniques to reduce redundant information in point cloud sequence 108. Redundant information is information that can be predicted at decoder 120 and may not need to be sent (e.g., transmitted) to decoder 120 for accurate decoding of point cloud sequence 108. For example, the Movie Experts Group (MPEG) introduced the Geometry-Based Point Cloud Compression (G-PCC) standard (ISO / IEC Standard 23090-9: Geometry-Based Point Cloud Compression). G-PCC specifies the encoded bitstream syntax and semantics for transmitting and / or storing compressed point cloud frames, and the decoder operations for reconstructing compressed point cloud frames from the bitstream. During the standardization of G-PCC, reference software (ISO / IEC Standard 23090-21: Reference Software for G-PCC) was developed to encode the geometric and attribute information of point cloud frames. To encode the geometric information of point cloud frames, the G-PCC reference software encoder can perform voxelization. The G-PCC reference software encoder can perform voxelization, for example, by quantizing the positions of points in the point cloud. Quantizing the positions of points in the point cloud can create a mesh in 3D space. The G-PCC reference software encoder can map points to the center coordinates of the sub-mesh volume (e.g., voxel) where their quantized positions are located. The G-PCC reference software encoder can use occupancy trees to perform geometric analysis to compress the geometric information. The G-PCC reference software encoder can entropy encode the results of the geometric analysis to further compress the geometric information. To encode the attribute information of the point cloud, the G-PCC reference software encoder can use transformation tools such as Region Adaptive Hierarchical Transformation (RAHT), predictive transformation, and / or lifting transformation. Lifting transformation can be constructed on top of predictive transformation. Lifting transformation can include additional update / lifting steps. The lift transform and the predictive transform can be referred to as predictive / lift transform or predictive lift. Encoder 114 can operate in the same or similar manner as the encoder provided in the G-PCC reference software.
[0057] Output interface 116 can be configured to write and / or store bit stream 110 onto transmission medium 104. Bit stream 110 can be sent (e.g., transmitted) to destination device 106. Alternatively or additionally, output interface 116 can be configured to send (e.g., transmit), upload, and / or stream bit stream 110 to destination device 106 via transmission medium 104. Output interface 116 may include wired and / or wireless transmitters configured to send (e.g., transmit), upload, and / or stream bit stream 110 according to one or more proprietary, open-source, and / or standardized communication protocols. One or more proprietary, open-source, and / or standardized communication protocols may include, for example, the Digital Video Broadcasting (DVB) standard, the Advanced Television Systems Committee (ATSC) standard, the Integrated Services Digital Broadcasting (ISDB) standard, the Cable Data Service Interface Specification (DOCSIS) standard, the 3rd Generation Partnership Project (3GPP) standard, the Institute of Electrical and Electronics Engineers (IEEE) standard, the Internet Protocol (IP) standard, the Wireless Application Protocol (WAP) standard, and / or any other communication protocol.
[0058] The transmission medium 104 may include wireless, wired, and / or computer-readable media. For example, the transmission medium 104 may include one or more wires, cables, air interfaces, optical discs, flash memory, and / or magnetic storage. Alternatively or additionally, the transmission medium 104 may include one or more networks (e.g., the Internet) or file servers configured to store and / or transmit (e.g., transfer) encoded video data.
[0059] Destination device 106 can decode bitstream 110 into point cloud sequence 108 for display or other forms of consumption. Destination device 106 may include one or more of input interface 118, decoder 120, and / or point cloud display 122. Input interface 118 may be configured to read bitstream 110 stored on transmission medium 104. Bitstream 110 may be stored on transmission medium 104 by source device 102. Alternatively, input interface 118 may be configured to receive, download, and / or stream bitstream 110 from source device 102 via transmission medium 104. Input interface 118 may include a wired and / or wireless receiver configured to receive, download, and / or stream bitstream 110 according to one or more proprietary, open-source, standardized communication protocols and / or any other communication protocols. Examples of protocols include the Digital Video Broadcasting (DVB) standard, the Advanced Television Systems Committee (ATSC) standard, the Integrated Services Digital Broadcasting (ISDB) standard, the Cable Data Services Interface Specification (DOCSIS) standard, the 3rd Generation Partnership Project (3GPP) standard, the Institute of Electrical and Electronics Engineers (IEEE) standard, the Internet Protocol (IP) standard, and the Wireless Application Protocol (WAP) standard.
[0060] Decoder 120 can decode the point cloud sequence 108 from the encoded bit stream 110. For example, decoder 120 can operate in the same or similar manner as the decoder provided in the G-PCC reference software. Decoder 120 can decode a point cloud sequence that approximates the point cloud sequence 108. Decoder 120 can decode a point cloud sequence that approximates the point cloud sequence 108 due to, for example, lossy compression of the point cloud sequence 108 by encoder 114 and / or errors introduced into the encoded bit stream 110, for example, in the event of transmission to destination device 106.
[0061] The point cloud display 122 can display the point cloud sequence 108 to a user. The point cloud display 122 may include, for example, a cathode rate tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light-emitting diode (LED) display, a 3D display, a holographic display, a head-mounted display, or any other display device suitable for displaying the point cloud sequence 108.
[0062] Point cloud coding (e.g., encoding / decoding) system 100 is presented by way of example and not limitation. Point cloud coding systems different from and / or modified versions of point cloud coding system 100 may perform the methods and processes described herein. For example, point cloud coding system 100 may include other components and / or arrangements. Point cloud source 112 may be, for example, external to source device 102. Point cloud display device 122 may be, for example, external to destination device 106 or omitted entirely (e.g., if point cloud sequence 108 is intended to be consumed by a machine and / or storage device). Source device 102 may further include, for example, a point cloud decoder. Destination device 106 may include, for example, a point cloud encoder. For example, source device 102 may be configured to further receive an encoded bit stream from destination device 106. Receiving an encoded bit stream from destination device 106 may support bidirectional point cloud transfer between devices.
[0063] As described in this paper, an encoder can quantize the position of points in a point cloud with spatial precision, which can be the same or different in each dimension of the point. The quantization process can create a grid in 3D space. The encoder can map any point residing within each sub-grid volume to the coordinates of the sub-grid center, referred to as a voxel or volume pixel. A voxel can be viewed as a 3D extension of the pixels corresponding to the 2D image grid coordinates.
[0064] An encoder can represent a point cloud (e.g., a voxelized point cloud) or write codes to a point cloud. An encoder can, for example, use an occupancy tree to represent or write codes to a point cloud. For example, an encoder can subdivide an initial volume or cuboid containing a point cloud into sub-cuboids. The initial volume or cuboid can be referred to as a bounding box. The cuboid can be, for example, a cube. The encoder can recursively subdivide each sub-cuboid containing at least one point of the point cloud. The encoder can choose not to further subdivide sub-cuboids that do not contain at least one point of the point cloud. A sub-cuboid containing at least one point of the point cloud can be referred to as an occupied sub-cuboid. A sub-cuboid that does not contain at least one point of the point cloud can be referred to as an unoccupied sub-cuboid. The encoder can subdivide an occupied sub-cuboid into, for example, two sub-cuboids (to form a binary tree), four sub-cuboids (to form a quadtree), or eight sub-cuboids (to form an octree). The encoder can subdivide occupied sub-cuboids to obtain additional sub-cuboids. Subcubes can have the same size and shape at a given depth level in the occupancy tree. For example, if the encoder splits an occupied subcube along a plane passing through the middle of the subcube's edge, the subcubes can have the same size and shape at a given depth level in the occupancy tree.
[0065] An initial volume or cuboid containing point clouds can correspond to the root node of the occupancy tree. Each occupied sub-cuboid split from the initial volume can correspond to a node in the second level of the occupancy tree (of the root node). Each occupied sub-cuboid split from the occupied sub-cuboids in the second level can correspond to a node in the third level of the occupancy tree (outside the occupied sub-cuboids in the second level from which it splits). For each recursive splitting iteration, the occupancy tree structure can continue to form in this way until, for example, a maximum depth level of the occupancy tree is reached or each occupied sub-cuboid has a volume corresponding to a voxel.
[0066] Each non-leaf node of the occupancy tree may include an occupancy word or be associated with an occupancy word representing the occupancy status of the cuboid corresponding to the node. For example, a node in the occupancy tree corresponding to a cuboid split into eight sub-cubes may include a 1-byte occupancy word or be associated with a 1-byte occupancy word. Each bit of the 1-byte occupancy word (referred to as an occupancy bit) may represent or indicate the occupancy of a different sub-cube among the eight sub-cubes. Occupied sub-cubes may be represented or indicated by a binary "1" in the 1-byte occupancy word. Unoccupied sub-cubes may be represented or indicated by a binary "0" in the 1-byte occupancy word. Occupied and unoccupied sub-cubes may be represented or indicated by the opposite 1-bit binary value in the 1-byte occupancy word (e.g., a binary "0" indicating or indicating an occupied sub-cube and a binary "1" indicating or indicating an unoccupied sub-cube).
[0067] Each bit of the occupancy word can represent or indicate the occupancy of a different sub-cube among the eight sub-cubes. For example, the least significant bit of the occupancy word can represent or indicate the occupancy of the first sub-cube among the eight sub-cubes following the so-called Merton order. The second least significant bit of the occupancy word can represent or indicate the occupancy of the second sub-cube among the eight sub-cubes following the Merton order, and so on.
[0068] Figure 2 An example of the Morton order is shown. More specifically, Figure 2 The Morton order of the eight sub-cubes 202-216, split from cuboid 200, is shown. Sub-cubes 202-216 can be labeled, for example, based on their Morton order, where child node 202 is the first in the Morton order and child node 216 is the last. The Morton order of sub-cubes 202-216 can be a local lexicographical order in xyz.
[0069] The geometry of a point cloud can be represented by the initial volume and occupancy word of a node in an occupancy tree, and can be determined from the initial volume and the occupancy word. An encoder can send (e.g., transmit) the initial volume and occupancy word of a node in the occupancy tree to a decoder in a bitstream for point cloud reconstruction. The encoder can entropy encode the occupancy word. The encoder can entropy encode the occupancy word, for example, before sending (e.g., transmitting) the initial volume and occupancy word of a node in the occupancy tree. The encoder can encode the occupancy bits of the occupancy word of a node corresponding to a cuboid. The encoder can encode the occupancy bits of the occupancy word of a node corresponding to a cuboid that is adjacent to or spatially close to the cuboid whose occupancy bit is being encoded, for example, based on one or more occupancy bits of the occupancy word of another node corresponding to a cuboid that is adjacent to or spatially close to the cuboid whose occupancy bit is being encoded.
[0070] The encoder and / or decoder can encode (e.g., encode and / or decode) the occupancy bits of occupancy words in scan order. Scan order can also be referred to as scanning order. For example, the encoder and / or decoder can scan the occupancy tree in breadth-first order. All occupancy words of nodes at a given depth (e.g., level) within the occupancy tree can be scanned. All occupancy words of nodes at a given depth (e.g., level) within the occupancy tree can be scanned, for example, before scanning the occupancy words of nodes at the next depth (e.g., level). Within a given depth, the encoder and / or decoder can scan the occupancy words of nodes in Morton order. Within a given node, the encoder and / or decoder can further scan the occupancy bits of the node's occupancy words in Morton order.
[0071] Figure 3An example scan order is shown. Figure 3 An example scan order (e.g., breadth-first order as described herein) is shown for occupied tree 300. More specifically, Figure 3 The scan order for the first three example levels of occupancy tree 300 is shown. Figure 3 In the given tree, the cuboid (e.g., cube) 302 corresponding to the root node of occupancy tree 300 can be divided into eight sub-cuboids (e.g., sub-cuboids). Two of the eight sub-cuboids, 304 and 306, may be occupied. The other six sub-cuboids may be unoccupied. Following Merton order, the first eight occupancy words (e.g., occW) are... 1,1 The occupancy word can be constructed to represent the root node. The first eight bits of the occupancy word (e.g., occW) 1,1 Each occupancy bit can represent or indicate the occupancy of a sub-cube in eight sub-cubes ordered by Morton. For example, the first eight-bit occupancy word occW 1,1 The least significant occupancy bit can represent or indicate the occupancy of the first sub-cube in the eight sub-cubes ordered by Morton. The first eight-bit occupancy word is occW. 1,1 The second least significant occupancy bit can represent or indicate the occupancy of the second sub-cube in the eight sub-cubes in Morton order, etc.
[0072] Each of the occupied subcubes (e.g., the two occupied subcubes 304 and 306) can correspond to a node other than the root node in the second level of the occupancy tree 300. Each of the occupied subcubes (e.g., the two occupied subcubes 304 and 306) can be further subdivided into eight subcubes. For example, one of the eight subcubes subdivided from subcube 304, subcube 308, may be occupied, and the other seven may be unoccupied. Three of the eight subcubes subdivided from subcube 306, subcubes 310, 312, and 314, may be occupied, and the other five may be unoccupied. Two second octet occWs can be constructed in this order. 2,1 and occW 2,2 , to represent the occupancy word corresponding to the node of subcube 304 and the occupancy word corresponding to the node of subcube 306, respectively.
[0073] Each of the occupied sub-cubicles (e.g., four occupied sub-cubicles 308, 310, 312, and 314) can correspond to a node in the third level of the occupancy tree 300. Each of the occupied sub-cubicles (e.g., four occupied sub-cubicles 308, 310, 312, and 314) can be further subdivided into eight sub-cubicles each, or a total of 32 sub-cubicles. For example, four third-level eight-bit occupancy words (occW) can be constructed in this order. 3,1 occW 3,2 occW 3,3 and occW 3,4 , respectively representing the occupancy word corresponding to the node of sub-cube 308, the occupancy word corresponding to the node of sub-cube 310, the occupancy word corresponding to the node of sub-cube 312, and the occupancy word corresponding to the node of sub-cube 314.
[0074] The occupancy words of the example occupancy tree 300 can be entropy-written coded (e.g., entropy-encoded by an encoder and / or entropy-decoded by a decoder) following, for example, the scan order discussed herein (e.g., Morton's order). The occupancy words of the example occupancy tree 300 can be entropy-written coded (e.g., entropy-encoded by an encoder and / or entropy-decoded by a decoder) into, for example, a sequence of seven occupancy words occW following the scan order discussed herein. 1,1 to occW 3,4 The scanning order discussed in this paper can be a breadth-first scanning order. For example, if the occupancy word of the current child node belonging to the current parent node is being entropy-written, then the occupancy words of all nodes with the same depth (or level) as the current parent node may have already been entropy-written. For example, the occupancy words of all nodes with the same depth (e.g., level) as the current child node and with a lower Morton order than the current child node may also have already been entropy-written. A portion of the written occupancy words can be used to entropy-write the occupancy word of the current child node. The written occupancy words of adjacent parent and child nodes can be used, for example, to entropy-write the occupancy word of the current child node. For example, if a specific occupancy bit of the occupancy word of the current child node is being written (e.g., entropy-written), then the occupancy bits of occupancy words with a lower Morton order than that specific occupancy bit may also have already been entropy-written and can be used to write the occupancy bits of the occupancy word of the current child node.
[0075] Figure 4 An example neighborhood of a cuboid is shown for entropy-writing coding of the occupancy of a sub-cuboid. More specifically, Figure 4 An example neighborhood of a cuboid with written-code occupant bits is shown. The neighborhood of a cuboid with written-code occupant bits can be used for entropy writing of the occupant bits of the current sub-cuboid 400. This can be based, for example, on the representation as discussed herein. Figure 4The scanning order of the occupancy tree of the cuboid geometry determines the neighborhood of the cuboid with the written code occupant bit. The neighborhood of a cuboid, i.e., the neighborhood of the current child cuboid, can include one or more of the following: cuboids adjacent to the current child cuboid, cuboids sharing vertices with the current child cuboid, cuboids sharing edges with the current child cuboid, cuboids sharing faces with the current child cuboid, parent cuboids adjacent to the current child cuboid, parent cuboids sharing vertices with the current child cuboid, parent cuboids sharing edges with the current child cuboid, parent cuboids sharing faces with the current child cuboid, parent cuboids adjacent to the current parent cuboid, parent cuboids sharing vertices with the current parent cuboid, parent cuboids sharing edges with the current parent cuboid, parent cuboids sharing faces with the current parent cuboid, etc. For example... Figure 4 As shown, the current child cuboid 400 can belong to the current parent cuboid 402. Following the scanning order of the occupancy words and occupancy bits of the occupancy tree nodes, the occupancy bits of the four child cuboids 404, 406, 408, and 410 belonging to the same current parent cuboid 402 may have already been written. The occupancy bits of the previous parent cuboid's child cuboid 412 may have already been written. The occupancy bits of the parent cuboid 414 may have already been written, while the occupancy bits of its child cuboids have not yet been written. The written occupancy bits of cuboids 404, 406, 408, 410, 412, and 414 can be used to write the occupancy bits of the current child cuboid 400.
[0076] The number (e.g., quantity) of possible occupancy configurations (e.g., a set of one or more occupancy words and / or occupancy bits) in the neighborhood of the current sub-cuboid can be 2. N , where N is the number (e.g., quantity) of cuboids with written occupancy bits in the neighborhood of the current child cuboid. The neighborhood of the current child cuboid can include dozens of cuboids. The neighborhood of the current child cuboid (e.g., dozens of cuboids) can include 26 neighboring parent cuboids that share faces, edges, and / or vertices with the parent cuboid of the current child cuboid, and several neighboring child cuboids that share faces, edges, and / or vertices with written occupancy bits. The occupancy configuration of the neighborhood of the current child cuboid can have billions of possible occupancy configurations, or even be limited to a subset of neighboring cuboids, making it impractical to use directly. Encoders and / or decoders can use the occupancy configuration of the neighborhood of the current child cuboid to select a context (e.g., a probabilistic model) from the context set for a binary entropy writer (e.g., a binary arithmetic writer) that can write occupancy bits of the current child cuboid. Context-based binary entropy coding can be similar to the context-adaptive binary arithmetic coder (CABAC) used in MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC)).
[0077] The encoder and / or decoder can use several methods to reduce the occupancy configuration of the neighborhood of the current child cuboid being coded to an actual number (e.g., quantity) of reduced occupancy configurations. This includes the six neighboring parent cuboids sharing a face with the current child cuboid. 6 Alternatively, 64 occupied configurations can be reduced to 9 occupied configurations. This reduction can be achieved by using geometric invariants. It can be reduced from 26 neighboring parent cuboids. 26 An occupancy configuration yields the current occupancy score for a sub-cube. The score can be further reduced to a ternary occupancy prediction (e.g., "predicted occupied," "uncertain," or "predicted unoccupied") by using a score threshold. Individual occupancy of these sub-cubes can be replaced by the number (e.g., quantity) of occupied neighboring sub-cubes and the number (e.g., quantity) of unoccupied neighboring sub-cubes.
[0078] Using / employing one or more of the methods described herein, the encoder and / or decoder can reduce the number (e.g., quantity) of possible occupancy configurations in the neighborhood of the current subcube to a more manageable number (e.g., thousands). It has been observed that instead of directly associating the reduced number (e.g., quantity) of contexts (e.g., probabilistic models) with the reduced occupancy configuration, another mechanism, namely the Optimal Binary Writer (OBUF) that supports on-the-fly updates, can be used. The encoder and / or decoder can implement the OBUF to limit the number (e.g., quantity) of contexts to a lower number (e.g., 32 contexts).
[0079] OBUF can use a finite number (e.g., 32) of contexts (e.g., probabilistic models). The number of contexts in OBUF (e.g., quantity) can be a fixed number (e.g., fixed quantity). Contexts used by OBUF can be sorted, indexed by context indices (e.g., context indices in the range of 0 to 31), and associated from the lowest virtual probability to the highest virtual probability to encode "1". A lookup table (LUT) for context indices can be initialized at the start of the point cloud coding process. For example, the LUT can initially point to a context with a median virtual probability (e.g., with context index 15) to encode "1" for all inputs. The LUT can initially point to a context with a median virtual probability to encode "1" for all inputs in a finite number (e.g., quantity) of contexts. This LUT can take the occupancy configuration of the neighborhood of the current subcube as input and output the context index associated with the occupancy configuration. The LUT can have as many entries as the reduced occupancy configuration (e.g., approximately several thousand entries). The write-code of the current child cuboid's occupancy bits may include the following steps: determining the reduced occupancy configuration of the current child node; obtaining a context index by using the reduced occupancy configuration as an entry in the LUT; writing the current child cuboid's occupancy bits using the context pointed to (e.g., indicated by) the context index; and updating the LUT entry corresponding to the reduced occupancy configuration, for example, based on the value of the written-code occupancy bits of the current child cuboid. For example, if a binary "0" (e.g., indicating that the current child cuboid is not occupied) is being written, the LUT entry may be reduced to a lower context index value. For example, if a binary "1" (e.g., indicating that the current child cuboid is occupied) is being written, the LUT entry may be increased to a higher context index value. The context index update process may, for example, be based on a theoretical model of the optimal distribution of virtual probabilities associated with a finite number (e.g., quantity) of contexts. This virtual probability may be fixed by the model and may differ from the internal probabilities of the context that may evolve, for example, when the write-code of a data bit occurs. The evolution of the internal context may follow a well-known process similar to that in CABAC.
[0080] The encoder and / or decoder can implement a “dynamic OBUF” scheme. For example, compared to a general OBUF, a “dynamic OBUF” scheme allows the encoder and / or decoder to handle a much larger number (e.g., quantity) of occupancy configurations in the neighborhood of the current sub-cube. The use of a larger number (e.g., quantity) of occupancy configurations in the neighborhood of the current sub-cube can result in improved compression capabilities while keeping complexity within reasonable limits. By using an occupancy tree compressed by OBUF, the encoder and / or decoder can achieve lossless compression performance as good as 1 bit / point (bpp) for writing geometry to dense point clouds. The encoder and / or decoder can implement dynamic OBUF to potentially further reduce the bit rate by more than 25%, to 0.7 bpp.
[0081] OBUF may not take into account the various reduced occupancy configurations of the current child cuboid's neighborhood as input, and may potentially lead to a loss of useful relevance. With OBUF, the size of the LUT with the context index can be increased to handle more diverse occupancy configurations of the current child cuboid's neighborhood as input. Due to this increase, statistics can be diluted, and compression performance may deteriorate. For example, if the LUT has millions of entries and the point cloud has hundreds of thousands of points, most entries may never be accessed (e.g., lookup, access, etc.). Many entries may only be accessed a few times, and their associated context index may not be updated enough to reflect any meaningful correlation between the current child cuboid's occupancy configuration value and occupancy probability. Dynamic OBUF can be implemented to mitigate the dilution of statistics caused by an increase in the number (e.g., quantity) of occupancy configurations in the current child cuboid's neighborhood. This mitigation can be performed through a "dynamic reduction" of the occupancy configuration in dynamic OBUF.
[0082] Dynamic OBUF can, for example, add an extra step to reduce the occupancy configuration of the current child cuboid's neighborhood before using a context-indexed LUT. This step can be called dynamic reduction because it evolves, for example, based on the progress of writing code to the point cloud, or more precisely, based on the occupancy configuration that has already been visited (e.g., looked up in the LUT).
[0083] As discussed in this paper, there may be many possible occupancy configurations potentially involving the neighborhood of the current sub-cube, but only a subset can be accessed if a write-code of the point cloud occurs. This subset can characterize the type of point cloud. For example, most accessed occupancy configurations may show the occupied neighboring cuboids of the current sub-cube, e.g., if an AR or VR dense point cloud is being written-coded. On the other hand, for example, if a sparse point cloud acquired by a sensor is being written-coded, most accessed occupancy configurations may only show a few occupied neighboring cuboids of the current sub-cube. The effect of dynamic reduction can be, for example, to obtain a more accurate correlation by simultaneously shelving (e.g., actively reducing) other occupancy configurations accessed much less frequently based on the most accessed occupancy configuration. Dynamic reduction can be updated on an in real-time basis. For example, if a write-code of occupancy data occurs, dynamic reduction can be updated on an in real-time basis, e.g., after each access to an occupancy configuration (e.g., a lookup in a LUT).
[0084] Figure 5 An example of a dynamically decreasing function DR that can be used in a dynamic OBUF is shown. The dynamically decreasing function DR can be implemented by masking the occupancy of bit β in configuration 500. j To obtain,
[0085] β = β1 … β K
[0086] The occupancy configuration consists of K bits. For example, if the occupancy configuration is accessed (e.g., looked up in the LUT) a certain number of times (e.g., a quantity), the mask size can be reduced. The initial dynamic reduction function DR... 0 It can mask all bits of all occupied configurations, making it a constant function DR for all occupied configurations β. 0 (β) = 0. The dynamically decreasing function can be derived from the function DR. n Evolved to the update function DR n+1 The dynamic reduction function can, for example, be applied to the DR function after each write operation of the occupied bits. n Evolved to the update function DR n+1 A function can be defined as follows:
[0087] β' = DR n (β) = β1 … β kn(β)
[0088] Where k n (β) 510 is the number of unmasked bits (e.g., quantity). DR 0 The initialization can correspond to k0(β) = 0, and the natural evolution of the decreasing function toward finer statistics can cause an increase in the number (e.g., quantity) of unmasked bits, k n (β) ≤ k n+1(β). The dynamic reduction function can be completely determined by all k occupying the configuration β. n The value is determined.
[0089] For all dynamically decreasing occupancy configurations β' = DR n Access to the occupancy configuration (β), such as an instance lookup in a LUT, can be tracked by the variable NV(β'). For example, in a LUT based on the occupancy configuration β... V After each instance of writing the occupant bit, the corresponding access count (e.g., quantity) NV(β) V ') can be increased by one. If this number of visits (e.g., quantity) NV(β) V ') greater than the threshold th V ,
[0090] NV(β V ') > th V
[0091] Then the number (e.g., quantity) of unmasked bits k n (β) for dynamically decreasing to β V For all occupancy configurations β, one can be added. This corresponds to the occupancy configuration β that will dynamically decrease. V 'Use two new dynamically reduced occupancy configurations β 0 'and β 1 The two new dynamically reduced occupancy configurations are defined as follows:
[0092] β 0 ' = β V '0 = β V 1 … β V kn(β) 0 and β 1 ' = β V '1 = β V 1 … β V kn(β) 1.
[0093] In other words, for all occupancy configurations β, the number (e.g., quantity) of unmasked bits has increased by one, k n+1 (β) =k n (β) + 1, making DR n (β) = β V The access count (e.g., quantity) of two new dynamically decreasing occupancy configurations can be initialized to zero.
[0094] NV(β 0 ') = NV(β 1 ') = 0. (I)
[0095] At the start of writing code, the initial dynamic decrease function DR 0 The initial number of visits (e.g., quantity) can be set to
[0096] NV(DR 0 (β)) = NV(0) = 0,
[0097] Furthermore, the evolution of NV with dynamically decreasing occupancy configuration can be fully defined.
[0098] The corresponding LUT entry LUT[β] V '] can be two new entries LUT[β] 0 '] and LUT[β] 1 The two new entries are replaced by '], and the two new entries are replaced by β. V 'Associated writer index initialization. For example, the corresponding LUT entry LUT[β] V '] can be two new entries LUT[β] 0 '] and LUT[β] 1 The two new entries are replaced by '], and the two new entries are replaced by β. V 'Associated writer index initialization, if dynamically reduced occupancy configuration β V 'By two new dynamically reduced occupancy configurations β 0 'and β 1 'replace,
[0099] LUT[β 0 '] = LUT[β 1 '] = LUT[β V '], (II)
[0100] And then it evolves independently. The evolution of the LUT with a dynamically decreasing occupancy configuration for the writer index can be fully defined.
[0101] Decrease function DR n It can be formed by a series of growing binary trees T n Model 520, whose leaf node 530 is a reduced occupancy configuration β' = DR n (β). The initial tree can be 0 = DR 0 (β) The associated single root node. This will be dynamically reduced to β. V Replace with β 0 'and β 1 'Can correspond to growth tree T' n (from β) V 'Associated leaf node', for example by using β 0 'and β 1 Two new nodes associated with each other are attached to the leaf node. Tree T n+1This can be obtained through such growth. The LUTs for the number of visits (e.g., quantity) NV and the context index can be defined on the leaf nodes and evolve as the tree grows via equations (I) and (II).
[0102] The practical implementation of dynamic OBUF can be achieved by storing an array NV[β'] and a LUT[β'] of context indexes, as well as a tree T. n 520 is used for this. An alternative to storing the tree could be an array k storing the number (e.g., quantity) of the unmasked bits. n [β]510.
[0103] One limitation of implementing dynamic OBUF is its memory footprint. In some applications, millions of occupied configurations can be processed, resulting in approximately 20 bits of beta. i The entries β constitute the configuration for the decrease function DR. Each bit β i It can correspond to the occupancy status of the adjacent cuboids of the current child cuboid or the set of adjacent cuboids of the current child cuboid.
[0104] The higher (e.g., higher effective) bit β i (For example, β0, β1, etc.) can be the first bit without masking. Higher (e.g., higher-active) bits β i (For example, β0, β1, etc.) could be, for example, the first unmasked bit during the evolution of the dynamically decreasing function DR. Place bit β... i The order of neighbor-based information in compression can affect compression performance. Neighbor information can be ordered from highest (e.g., highest) priority to lowest priority, and then placed into bit β in this order from highest weight to lowest weight. i In this context, priority is ranked from most important to least important: occupancy of adjacent child cuboids, followed by adjacent child cuboids, then adjacent parent cuboids, then non-adjacent child nodes, and finally non-adjacent parent nodes. Neighboring nodes sharing a face with the current child node can have a higher priority than neighboring nodes sharing an edge (but not a face) with the current child node. Neighboring nodes sharing an edge with the current child node can have a higher priority than neighboring nodes sharing only a vertex with the current child node.
[0105] Figure 6 An example method for writing code to the occupancy of a cuboid using dynamic OBUF is shown. More specifically, Figure 6 An example method for writing the occupancy bits of the current subcube using dynamic OBUF is shown. Figure 6 One or more steps can be performed by an encoder and / or decoder (e.g., Figure 1The encoder 114 and / or decoder 120 in the flowchart can be executed in whole or in part by a programmer (e.g., Figure 1 (encoder 114 and / or decoder 120 in the middle) Figure 27 Example computer system 2700 and / or Figure 28 Example computing device 2830 is implemented.
[0106] At step 602, the occupancy configuration of the current sub-cuboid (e.g., occupancy configuration β) can be determined. The current occupancy configuration (e.g., occupancy configuration β) can be determined, for example, based on the occupancy bits of the written-code cuboids in the neighborhood of the current sub-cuboid. At step 604, the occupancy configuration (e.g., occupancy configuration β) can be dynamically reduced. For example, a dynamic reduction function DR can be used. n This allows for dynamic reduction of the occupied configuration. For example, the occupied configuration β can be dynamically reduced to a reduced occupied configuration β' = DR. n (β). At step 606, the context index can be looked up in a lookup table (LUT), for example. For example, the encoder and / or decoder can look up the context index LUT[β'] in the LUT of the dynamic OBUF. At step 608, a context (e.g., a probabilistic model) can be selected. For example, the context pointed to by the context index (e.g., a probabilistic model) can be selected. At step 610, the occupancy of the current sub-cube can be entropy-coded. For example, the occupancy bits of the current sub-cube can be entropy-coded (e.g., arithmetic coding) based on the context. The occupancy bits of the current sub-cube can be coded based on the occupancy bits of the adjacent coded cuboids.
[0107] although Figure 6 Not shown, but the encoder and / or decoder can update the reduction function and / or update the context index. For example, the encoder and / or decoder can update the reduction function DR. n Updated to DR n+1 And / or, for example, update the context index LUT[β'] based on the occupancy bits of the current child cuboid. Figure 6 The method can be based on the scanning order, as discussed in this article. Figure 3 The scanning order discussed repeats for the additional or all child cuboids of the parent cuboid corresponding to the node occupying the tree.
[0108] Occupation tree is typically a lossless compression technique. Occupation tree can be adapted to provide lossy compression, for example, by modifying the point cloud on the encoder side (e.g., downsampling, removing points, moving points, etc.). Lossy compression performance can be weak. For dense point clouds, lossy compression can be a useful lossless compression technique.
[0109] One approach to lossy compression of point cloud geometry could be to set the maximum depth of the occupancy tree to stop at a larger volume size (e.g., an NxNxN cuboid (e.g., a cube), where N > 1) instead of reaching the minimum volume size of a single voxel. The geometry of points belonging to each occupied leaf node associated with this larger volume can then be modeled. This approach may be particularly well-suited for dense and smooth point clouds that can be locally modeled using smoothing functions such as planes or polynomials. The coding cost can be reduced to the cost of the occupancy tree plus the cost of the local model within each occupied leaf node.
[0110] A scheme for modeling the geometry of points belonging to each occupied leaf node associated with a volume larger than one voxel can use a set of triangles as a local model. This scheme can be called a "TriSoup" scheme. TriSoup is short for "Triangle Soup" because the connections between triangles may not be part of the model. An occupied leaf node of the occupancy tree corresponding to a cuboid with a volume larger than one voxel can be called a TriSoup node. An edge belonging to at least one cuboid corresponding to a TriSoup node can be called a TriSoup edge. A TriSoup node can include an existence flag (s) for each TriSoup edge of its corresponding occupied cuboid. k The existence flag of a TriSoup edge (s) k ) can indicate the TriSoup vertex (V k Does a vertex exist on a TriSoup edge? At most one TriSoup vertex (V) exists. k A vertex (V) can exist on a TriSoup edge. For each vertex (V) existing on a TriSoup edge of an occupied cuboid... k The TriSoup node corresponding to the occupied cuboid can further include vertices (V). k ) along the position of TriSoup edge (p k ).
[0111] In addition to the occupancy word of the occupancy tree, the encoder can also entropy encode the TriSoup vertex presence flag and position for each TriSoup edge belonging to the TriSoup node of the occupancy tree. Similarly, in addition to the occupancy word of the occupancy tree, the decoder can also entropy decode the TriSoup vertex presence flag and position for each TriSoup edge and the vertices along the corresponding TriSoup edges belonging to the TriSoup nodes of the occupancy tree.
[0112] Figure 7An example of an occupied cuboid (e.g., a cube) 700 is shown. More specifically, Figure 7 An example of an occupied cuboid (e.g., a cube) 700 of size NxNxN (where N > 1) corresponding to a TriSoup node in the occupied tree is shown. The occupied cuboid 700 may include edges (e.g., TriSoup edges 710-721). The TriSoup node corresponding to the occupied cuboid 700 may include an presence flag (s) for each edge (e.g., each TriSoup edge in TriSoup edges 710-721). k For example, the presence flag of TriSoup edge 714 can indicate that TriSoup vertex V1 exists on TriSoup edge 714. The presence flag of TriSoup edge 715 can indicate that TriSoup vertex V2 exists on TriSoup edge 715. The presence flag of TriSoup edge 716 can indicate that TriSoup vertex V3 exists on TriSoup edge 716. The presence flag of TriSoup edge 717 can indicate that TriSoup vertex V4 exists on TriSoup edge 717. The presence flags of the remaining TriSoup edges can each indicate that no TriSoup vertex exists on its corresponding TriSoup edge. The TriSoup node corresponding to the occupied cuboid 700 can include the position of each TriSoup vertex that exists along one of its TriSoup edges 710-721. More specifically, the TriSoup node corresponding to the occupied cuboid 700 can include the position p1 of TriSoup vertex V1, the position p2 of TriSoup vertex V2, the position p3 of TriSoup vertex V3, and the position p4 of TriSoup vertex V4. TriSoup vertices can be shared among TriSoup nodes along common TriSoup edges.
[0113] The existence of the current TriSoup edge can be flagged (s) k ) and (in the presence of signs (s k (The location (p) can indicate the existence of a vertex) k Entropy coding is performed. Existence flag (s) k ) and location (p k The information can be referred to individually or collectively as vertex information or TriSoup vertex information. For example, the existence flag (s) of the current TriSoup edge can be determined based on the existence flags and positions of the existing TriSoup vertices of the TriSoup edges adjacent to the current TriSoup edge. k ) and (in the presence of signs (s k (indicating the presence of a vertex) position (p) kEntropy coding is performed on the current TriSoup edge. Alternatively or alternatively, the existence flag (s) of the current TriSoup edge can be added. k ) and (in the presence of signs (s k (The location (p) can indicate the existence of a vertex) k (e.g., indicating the position of the vertex along which the edge is located) is entropy-coded. The presence flag of the current TriSoup edge (s) k ) and location (p k Entropy coding can be performed alternatively or alternatively, for example, based on the occupancy of cuboids adjacent to the current TriSoup edge. Similar to the entropy coding of occupancy bits in an occupancy tree, the configuration β of the neighborhood of the current TriSoup edge can be obtained. TS (Also known as neighborhood configuration β) TS ), and, for example, dynamically reduce it to a reduced configuration β by using TriSoup's dynamic OBUF scheme. TS ' = DR n (β) TS ). Context index LUT[β TS The information can be obtained from the OBUF LUT. At least a portion of the vertex information for the current TriSoup edge can be entropy-coded using the context pointed to by the context index (e.g., a probabilistic model).
[0114] The position of the TriSoup vertex along its TriSoup edge (p k (If it exists) can be binarized. The position of the TriSoup vertex along its TriSoup edge (p k (If present) can be binarized, for example, by using a binary entropy writer to entropy encode at least a portion of the vertex information of the current TriSoup edge. The number of bits (e.g., quantity) N can be set. b To quantize the TriSoup vertex positions (p) along a TriSoup edge of length N. k A TriSoup edge of length N can be uniformly divided into 2. Nb Each quantization interval. By doing so, the TriSoup vertex position (p k (This can be written separately by N using a dynamic OBUF scheme) b Units digit (p) k j j=1, …, N b ) and corresponding to the existence flag (s) k The bit representation of ). Neighborhood configuration β TS、 OBUF decreases function DR n The context index can depend on the bits of the write-code (e.g., the presence flag). k), highest position (p) k1 ), second highest position (p) k2 The bits of the written code (e.g., the presence of a flag (s)) k ), highest position (p) k 1 ), second highest position (p) k 2 Properties, characteristics, and / or attributes of vertex information. In reality, there may be several dynamic OBUF schemes, each dedicated to specific bits of vertex information (e.g., presence flags). k ) or position (p) k j )).
[0115] Figure 8A An example cuboid 800 (e.g., a cube) corresponding to a TriSoup node is shown. The cuboid 800 can correspond to a TriSoup node with a number of K vertices V. k The TriSoup node. Within the cuboid 800, the TriSoup triangle can be formed by the TriSoup vertex V. k Construction. For example, if there are at least three (K≥3) TriSoup vertices on the TriSoup edges of a cuboid 800, then a TriSoup triangle can be constructed from TriSoup vertex V. k Construction. For example, regarding... Figure 8A A TriSoup can have four vertices and can construct a TriSoup triangle. The TriSoup triangle can be constructed around the centroid vertex C, which is defined as the TriSoup vertex V. k The mean of the values. The primary direction can be determined, and then the vertex V can be adjusted by rotating around this direction. k Sort the data and construct the following K TriSoup triangles (list the triplets with vertices as vertices): V1V2C, V2V3C, ..., V K V1C. For example, if a triangle is projected along a principal direction, the principal direction can be selected from three directions that are respectively parallel to the axes of 3D space to increase or maximize the 2D surface of the triangle. By doing so, the principal direction can be slightly perpendicular to the local surface defined by the points of the point cloud belonging to the TriSoup node.
[0116] Figure 8B A detailed example of the TriSoup model is shown. The TriSoup model can be improved by writing the centroid residual values. Centroid residual value C res It can be written into the bitstream. Centroid residual value C res It can be written to a bitstream, for example, to use C+C++.res Instead of C as the pivot vertex of the triangle, by using C + C res As the pivot vertex of the triangle, vertex C + C res Points can be located closer to the point cloud than the centroid C, which can reduce reconstruction error and thus reduce distortion, but at the cost of writing code C. res The required bit rate will increase slightly.
[0117] Reconstructing the decoded point cloud from a set of TriSoup triangles can be called “voxing”, and can be performed individually for each triangle by ray tracing or rasterization, for example, before removing duplicate voxels from the voxed triangles.
[0118] Figure 9 An example of voxelization is shown. For example, ray 900 can be drawn from integer coordinates P. start Begin firing parallel to one of the three coordinate axes in 3D space. This can be done at the intersection point P with the TriSoup triangle 901 belonging to cube 902. int (If present) Rounding is performed to determine the decoded point. Cube 902 may correspond to a TriSoup node. This intersection can be found (e.g., determined) using the Möller-Trumbore algorithm, for example.
[0119] The existence flag (s) of vertices along the current TriSoup edge can be determined based on, for example, the existence flags and positions of the written codes of the TriSoup edges (existing TriSoup vertices) adjacent to the current TriSoup edge. k ) and (in the presence of signs (s k (indicating the presence of a vertex) position (p) k Entropy coding is performed. Existence flag (s) k ) and location (p k This can be referred to individually or collectively as vertex information. It can also be additionally or alternatively, for example, based on the occupancy of cuboids adjacent to the current TriSoup edge, to indicate the presence of the current TriSoup edge. k ) and (in the presence of signs (s k (indicating the presence of a vertex) position (p) k Entropy writing codes are performed (e.g., indicating the position of a vertex along the edge). Similar to the entropy writing codes of occupancy bits in an occupancy tree, the configuration β of the neighborhood of the current TriSoup edge can be determined. TS (Also known as neighborhood configuration β) TS ), and / or, for example, dynamically reduce it to a reduced configuration β by using TriSoup's dynamic OBUF scheme. TS ' = DR n (β)TS ). Context index LUT[β TS The information can be determined from the OBUF LUT. For example, the context (or probabilistic model) pointed to by the context index can be used to entropy-code at least a portion of the vertex information of the current TriSoup edge.
[0120] The position of the TriSoup vertex along its TriSoup edge (p k (If present) can be binarized. A binary entropy writer can entropy-code at least a portion of the vertex information of the current TriSoup edge. The number of bits N can be set. b To quantize the TriSoup vertex positions (p) along a TriSoup edge of length N. k The edge is uniformly divided into 2 Nb One quantization interval. For example, if the number of bits N b Set to be used for quantizing the TriSoup vertex positions (p) along a TriSoup edge of length N. k ), then the position of the TriSoup vertex (p k ) can be made by N b Units digit (p) k j j=1,…,N b The edge is uniformly divided into 2 Nb Quantization intervals. N b The unit digit and the corresponding existence marker (s) k The bits of ) can be individually coded using the dynamic OBUF scheme. Neighborhood configuration β TS OBUF decrease function DR n And / or context indexes can depend on the nature / characteristics / attributes of the written code points (e.g., presence flags). k ), highest position (p) k 1 ) or second highest position (p k 2 Several dynamic OBUF schemes can be implemented, where each dynamic OBUF scheme is dedicated to a specific bit of vertex information (e.g., presence flag). k ) or position (p) k j )).
[0121] In video compression, performance can be improved by using inter-frame prediction. The bit rate used to compress inter-frames can be one to two orders of magnitude lower than the bit rate of intra-frames that do not use inter-frame prediction. Point cloud data may behave differently than, for example, 2D video data. For point cloud data, 3D geometry can be encoded using 3D point locations. Each point location in the 3D point locations can be associated with an attribute (e.g., color). The geometry and / or attributes may vary between frames. Encoding can be performed, for example, for each frame on different 3D point locations and / or attributes associated with the corresponding 3D point locations. 2D video data can be obtained, for example, by projecting the 3D geometry and / or attributes onto a 2D plane with a fixed geometry (e.g., a camera sensor). For video encoding, attributes can be encoded, but geometry may not be encoded (and may not need to be encoded). Even if, for example, the properties of the expected 2D projection have a higher temporal relevance than the underlying 3D geometry, it can be anticipated that inter-frame prediction between 3D point clouds can provide improved compression capabilities compared to intra-frame prediction within the point cloud (e.g., individual intra-frame prediction). Octrees can benefit from inter-frame prediction and / or geometric compression gains. The general framework for inter-frame prediction of 3D point clouds can be analogous to one of the methods used in video compression.
[0122] Figure 10 Example encoding method 1000 is shown. Figure 10One or more steps can be performed by an encoder (e.g., encoder 114). Encoding method 1000 can use inter-frame prediction between different point cloud frames. The current frame 1001 (e.g., image or point cloud) can be written to relative to a written reference frame 1010 (e.g., image or point cloud). At step 1020, a motion search can be performed from the written reference frame 1010 toward the current frame 1001 to determine a motion vector 1021 representing the motion flow between the two frames 1010 and 1001. In video compression, the motion vector can be a 2-component (or 2D) vector representing the motion from a reference pixel block to the current pixel block. In point cloud compression, the motion vector can be a 3-component (or 3D) vector representing the motion from a reference 3D point set (e.g., in a reference point cloud) to the current 3D point set (e.g., in the current point cloud). At step 1025, the motion vector 1021 can be entropy-coded into bitstream 1050. At step 1030, motion compensation can be performed on reference frame 1010 to determine motion-compensated frame 1031. Motion compensation may involve moving pixels of the reference image based on 2D motion vectors, and / or moving points of the reference point cloud based on 3D motion vectors. The determined motion-compensated frame may be "closer" to the current frame than the reference frame; for example, the chromatic difference (or dot distance) between motion-compensated frame 1031 and the current frame 1001 may be on average smaller than the chromatic difference (or dot distance) between reference frame 1010 and the current frame 1001. At step 1040, inter-frame prediction can be performed to determine inter-frame residual 1041. At step 1045, the inter-frame residual 1040 can be entropy-coded into bitstream 1050. The inter-frame residual may carry more compressible information than the current frame, which may or may not have undergone intra-frame prediction. Therefore, the entropy write coding performed at step 1045 may be more efficient in determining a bitstream 1050 with a reduced size compared to the bitstream determined by writing coding to the current frame 1001 which does not benefit from inter-frame prediction.
[0123] In video coding, inter-frame residuals can be constructed as the color, pixel-per-pixel difference between the current pixel block belonging to the current frame (e.g., an image) and the co-located compensated pixel block belonging to a motion-compensated frame (e.g., an image). Inter-frame residuals can be color difference arrays, which typically have small values and can be effectively compressed.
[0124] In point cloud compression, there may be no "difference" between two sets of points because there may not be a one-to-one mapping between them. The concept of inter-frame residuals may not be directly applicable (e.g., generalization) relative to point clouds. To predict octrees representing point clouds, the concept of inter-frame residuals can be replaced with conditional entropy write-code, where conditional information for performing conditional entropy write-code can be constructed based on motion-compensated point clouds. This can be extended to the framework of dynamic OBUF.
[0125] As described herein, the occupancy of the current volume (e.g., the current volume associated with the current node of the octree) can be provided by a number of occupancy bits, for example, 8 occupancy bits. The current occupancy bits of the octree can be written by an entropy writer selected by the output of a dynamic OBUF LUT using a writer index with a neighborhood configuration β as input. The neighborhood configuration β can be constructed based on the written occupancy bits. The written occupancy bits can be associated with neighboring volumes (e.g., with the neighboring nodes of the current node). The construction of the neighborhood configuration β can be extended using inter-frame information. Inter-frame predictor occupancy bits can be indicated by the current occupancy bits (e.g., defined for the current occupancy bits) as bits representing the presence of at least one point in the motion-compensated point cloud within the current volume. For example, if this motion compensation is valid, there may be a strong correlation between the current occupancy bits and the inter-frame predictor occupancy bits. This could be because the current point cloud and the motion-compensated point cloud may be close to each other. Using the inter-frame predictor occupancy bits as the bits of the neighborhood configuration β can result in better compression performance of the octree (e.g., by dividing the size of the octree bitstream by a factor of two).
[0126] The motion field between octrees may include 3D motion vectors associated with a 3D prediction unit (PU). A PU may have a volume that may include at least a portion of one or more volumes (e.g., cuboids) associated with nodes of the octree. Motion compensation for each volume (e.g., each cuboid's cuboid) may be performed based on the 3D motion vectors to determine a motion-compensated point cloud in one or more current volumes. Inter-frame predictor occupancy bits may be determined, for example, based on the presence of at least one point in this motion-compensated point cloud.
[0127] The TriSoup scheme can benefit from motion-compensated frames determined, for example, during octree coding prior to TriSoup coding. Predictors for the presence and / or location of TriSoup vertices can be determined based on the motion-compensated point cloud. For example, a predictor can be determined based on the intersection of edges between the compensated point cloud and TriSoup nodes. Predictors for centroid residual values can also be determined.
[0128] Entropy coding of TriSoup vertex and / or centroid residual values can be performed, for example, by using these inter-frame predictors. The inter-frame predictor can constitute the context information β of a dynamic OBUF instance that codes out TriSoup syntax elements. inter Part of the input. Alternatively or additionally, the context can be selected based on an inter-frame predictor. The selected context can be used by an entropy writer (e.g., CABAC) to determine the probability of arithmetic entropy writing codes for TriSoup syntax elements.
[0129] For example, after writing to the underlying geometry (e.g., the location of points in 3D space), the attributes associated with the points in the point cloud can be written to. If the geometry writing is lossless (e.g., using an octree scheme), the encoder can directly access the attribute values associated with the written-coded points. If the geometry writing is lossy (e.g., using a TriSoup scheme), the written-coded geometry can differ from the original geometry. The original attributes can be mapped from the original geometry to the written-coded geometry by the encoder to determine the mapped attributes on the written-coded geometry.
[0130] As described in this paper, attributes can indicate the nature of a point's visual appearance (e.g., texture, color, material, transparency, reflectivity, timestamp, or velocity). For color attributes, the attribute mapping performed by the encoder can be referred to as a recoloring process. This is likely because the color of the original geometry can be used to color (e.g., recolor) the writable geometry (e.g., reconstructed geometry).
[0131] The write-coded geometry associated with the mapping attributes of the write-coded geometry may include a point cloud representing the original point cloud in the geometry and attributes. There may be more than one (e.g., two) attribute writing schemes that can be used and / or selected for writing the attributes associated with the write-coded geometry. Attribute writing schemes may include, for example, a prediction-lift transformation (“pred-lift”) scheme and / or a region adaptive hierarchical transformation (“RAHT”) scheme. Attribute writing schemes may be used, for example, in G-PCC.
[0132] Predictive boosting schemes can begin by decomposing the writable geometry into levels of detail (also known as LoD). For a set (S) of points (e.g., all points) of the writable geometry, this set can be decomposed into disjoint subsets Si. i , making L levels of detail can be defined as the point cloud geometry tower.
[0133]
[0134] For example, if the set can be decomposed into disjoint subsets S i So that Then the set of points It can be the first (e.g., the coarsest) level of detail, and a set of points. It can be the Lth (e.g., the finest) level of detail.
[0135] Attribute a j Points s in the set S of all points whose geometry can be written can be used. j Related. Considering a subset a of the attributes.0 ... a L-1 subset a i Attribute a in i j Can be with subset S i point s in i j Related. The i-th level of detail ( It can have a subset a 0 ... a i-1 The cascading related attributes in the set. A set 'a' of attributes (e.g., all attributes) can be partitioned into subsets 'a'. 0 ... a L-1 .
[0136] Figure 11 , 12 Figures 13 and 14 illustrate examples of writing (e.g., encoding or decoding) point cloud attributes. This writing can be based on an intra-frame transform scheme, such as a predictive transform scheme or a predictive boosting scheme. The predictive transform scheme can be a variant of the predictive boosting scheme without an update operation. The intra-frame transform scheme can be an example of wavelet transform, which can transform (e.g., convert) attribute values into wavelet coefficients that can be more effectively compressed compared to the original attribute values. The wavelet coefficients can represent the values of residual attributes, which can be smaller and more effectively compressed compared to the original attribute values. The residual attributes can be referred to as and / or include wavelet coefficients (or transformed / transformed coefficients). The wavelet coefficients can be generated by applying a predictive transform scheme and / or a predictive boosting scheme.
[0137] As described herein, prediction transformation schemes and / or prediction enhancement schemes can operate using predictions within and / or between attribute levels of detail. At the encoder, attributes at higher (e.g., finer) levels of detail can be predicted based on attributes at lower (e.g., coarser) levels of detail. For example, each level of detail starting from the highest level can be predicted sequentially based on lower levels of detail. The decoder can perform the inverse operation, such that attributes at lower levels of detail can be predicted and / or reconstructed, for example, based on residual attributes at higher levels of detail. For example, each level of detail starting from the lowest level can be predicted and reconstructed sequentially based on higher levels of detail. Although... Figure 11-14 The example shown illustrates three levels of detail (LoD), but it should be understood that different numbers of LoDs can be used, for example, by extending (e.g., iteratively) the decomposition scheme.
[0138] Figure 11 An example method for encoding point cloud attributes is shown. The encoding can be based on a predictive transformation scheme. Figure 11 One or more steps can be accomplished by an encoder (e.g., Figure 1The encoder 114 in the middle is executed.
[0139] Predictions can be made, for example, using one or more (e.g., three) levels of detail within and between, starting from the highest (e.g., first prediction) level of detail (e.g., associated with the highest frequency, such as res a). 2 From the lowest (e.g., third) level of detail (e.g., associated with the lowest frequency, such as a) 0 The encoder writes the attribute set 'a'. At step 1110, the encoder can split the attribute set 'a' into a subset a. 2 The first set of attributes in the property and includes two subsets a 0 and a 1 The second set of attributes in the set. At step 1120, the encoder can obtain the second set of attributes (a 0 and a 1 The attributes in ) determine the first attribute set (a 2 The predicted values of the attributes in ) are then determined. At step 1130, the encoder can determine the first residual value 'res a'. 2 The encoder can, for example, obtain the first set of attributes (a). 2 The first residual value 'resa' is determined by subtracting the predicted value from the attribute in the property. 2 At step 1170, the encoder can output the first residual value 'res a'. 2 'Encode into the bitstream. The operations at steps 1120-1130 can be iteratively applied (e.g., applied to) each consecutive lower (e.g., coarser) LoD.
[0140] At step 1140, the encoder can transmit the second attribute set (a 0 and a 1 The set is split into a third attribute set and a fourth attribute set. The third attribute set may include a subset a. 1 The attributes in. The fourth attribute set can include a subset a. 0 The attributes in. At step 1150, the encoder can obtain the attributes from the fourth attribute set (a 0 The attributes in ) determine the third attribute set (a 1 The predicted values of the attributes in ) are then used. At step 1160, the encoder can determine the second residual value 'res a'. 1 The encoder can, for example, be obtained from a third set (a 1 The second residual value 'res a' is determined by subtracting the predicted value from the attribute in the property. 1 At step 1170, the encoder can output the second residual value 'res a'. 1 'Encoded into the bitstream. The encoder can convert the fourth attribute set (a 0The attributes in the stream are encoded into the bitstream.
[0141] The bitstream may include a representation of the first residual 'res a' 2 '、Second residual'res a 1 'and / or subset a 0 (For example, the data of attributes in the fourth attribute set). At step 1170, the residual values can be entropy-coded. The encoder can then encode the subset a. 0 The attributes in the code are encoded into the bitstream. Alternatively, the encoder can encode a subset a of the code to be written. 0 The current attribute a in 0 j Perform intra-frame prediction. The encoder can, for example, be based on a subset a. 0 The attribute of the written code in the data is used to perform the operation on subset a. 0 The current attribute a in 0 j Intra-frame prediction. This can improve subset a. 0 The compression efficiency of attributes in the data.
[0142] For example, if lossy attribute write codes are allowed, the encoder can quantize a subset a. 0 The attribute in, the first residual value 'res a 2 'and / or the second residual'res a 1 The encoder can convert subset a 0 The attribute (or subset a) in 0 Quantified attributes in the data), first residual value 'res a 2 '(or the first quantized residual value'res a) 2 ') and / or the second residual value 'resa 1 '(or the quantified second residual value'res a) 1 The encoding (e.g., entropy encoding) is added to the bitstream.
[0143] Figure 12 An example method for decoding point cloud properties is shown. Decoding can be based on a predictive transformation scheme. Figure 12 One or more steps can be accomplished by a decoder (e.g., Figure 1 The decoder (120) in the code executes. It can decode the attribute set 'a'. It can also encode the attribute set 'a', for example, as described in this article... Figure 11 As described. Decoding can use predictions between one or more (e.g., three) levels of detail, from the lowest (e.g., the third) level of detail to the highest (e.g., the first) level of detail (e.g., as described in this paper regarding...). Figure 11 The encoder described is in reverse order.
[0144] At step 1210, the decoder can obtain the first residual value 'res a' from the bitstream pair. 2 '、Second residual'res a 1 'and / or the fourth set of attributes (a 0 The decoder can decode from the attributes in the fourth attribute set (a). The decoder can use (e.g., apply) dequantization (e.g., for lossy compression). At step 1220, the decoder can decode from the fourth attribute set (a 0 The decoded attributes determine the third attribute set (a) 1 The predicted value of China's attributes. The decoder can, for example, be used with information about... Figure 11 The predicted value is determined in a similar manner to that described in step 1150. At step 1230, the decoder can determine the predicted value, for example, by adding the predicted value to the decoded first residual value 'res a'. 1 'To determine the third attribute set (a) 1 The decoded attributes in ) . The operation at steps 1120-1230 can be iteratively applied (e.g., applied to) each successively higher (e.g., finer) LoD.
[0145] At step 1240, the decoder can determine the second attribute set (a 0 and a 1 The decoder can, for example, merge a third set of attributes (a 1 ) and the fourth attribute set (a 0 ) to determine the second attribute set (a 0 and a 1 Step 1240 can be the inverse operation of step 1140. At step 1250, the decoder can obtain the second attribute set (a 0 and a 1 The attributes in ) determine the first attribute set (a 2 The decoder can determine the predicted values of the attributes in the () and the decoder can determine the predicted values in a manner similar to that described with respect to step 1120. At step 1260, the decoder can, for example, by adding the predicted values to the decoded second residual value 'res a 2 'Determine the first set of attributes (a) 2 The decoded attributes in ) . At step 1270, the decoder can, for example, merge the first attribute set (a 2 ) and the second attribute set (a 0 and a 1 This determines the set of decoded attributes 'a' of the write-coded geometry S (e.g., the entire write-coded geometry S). Step 1270 can be the inverse operation of step 1110.
[0146] Figure 13An example method for encoding point cloud attributes is shown. The encoding can be based on a prediction boosting transformation scheme with prediction and updates. Figure 13 One or more steps can be accomplished by an encoder (e.g., Figure 1 The encoder 114 in the code performs the operation. Predictions can be made using one or more (e.g., three) levels of detail, starting from the highest (e.g., first) level of detail (e.g., associated with the highest frequency, such as resembling a). 2 down to the lowest (e.g., third) level of detail (e.g., associated with the lowest frequency, such as up up a) 0 Encode the attribute set 'a'.
[0147] At step 1310, the encoder can split the attribute set 'a' into a subset a. 2 The first set of attributes in the property and includes two subsets a 0 and a 1 The second set of attributes in the property. Step 1310 can be similar to the section on the second set of attributes in this paper. Figure 11 The process is performed as described in step 1110. At step 1320, the encoder can access the second attribute set (a... 0 and a 1 The attributes in ) determine the first attribute set (a 2 The predicted values of the attributes in the first set of attributes (a). Step 1320 can be performed similarly to that described herein with respect to step 1120. At step 1330, the encoder can, for example, obtain the predicted values of the attributes from the first set of attributes (a 2 The first residual value 'res a' is determined by subtracting the predicted value from the attribute in the property. 2 Step 1320 can be performed similarly to that described herein with respect to step 1130. At step 1370, the encoder can convert the first residual value 'res a' 2 'Encode into the bitstream. Step 1370 can be performed similarly to the description in this document regarding step 1170.'
[0148] At step 1375, the encoder can obtain the first residual value 'res a' 2 'Determine the updated attribute value. The encoder can determine the updated attribute value, for example, based on a first residual value. The updated attribute value can be determined as the first residual value multiplied by a scaling factor (e.g., ½, ¼, 1 / 8), which can be predetermined or signaled in the bit stream.' At step 1380, the encoder can add the updated attribute value to a second attribute set (a 0 and a 1 The attribute values of ) are used to update the second attribute set (a) 0 and a 1 The attribute value 'up a' 0 'and'up a1 The operations at steps 1320, 1330, 1375, and 1380 can be iteratively applied (e.g., applied to) each successively lower (e.g., coarser) LoD.
[0149] At step 1340, the encoder can transmit the second attribute set (a 0 and a 1 The set is split into a third attribute set and a fourth attribute set. The third attribute set may include a subset a. 1 The updated attribute value 'up a 1 The fourth attribute set can include a subset a. 0 The updated attribute value 'up a 0 At step 1350, the encoder can obtain the fourth attribute set (a 0 The updated attribute value 'up a' 0 'Determine the set of third attributes (a)' 1 The updated attribute value 'up a' 1 The predicted value of '. At step 1360, the encoder can, for example, obtain the predicted value from the third attribute set (a 1 The updated attribute value 'up a' 1 'Subtract the predicted value from the middle to determine the third residual' 1 At step 1370, the encoder can update the third residual value. 1 'Encode into the bitstream. At step 1385, the encoder can retrieve the third residual value from the bitstream.' 1 'Determine the updated attribute value. At step 1390, the encoder can, for example, add the updated attribute value to the fourth attribute set (a 0 The updated attribute value 'up a' 0 'Determine the fourth attribute set (a) 0 The further updated attribute value 'up up a' 0 At step 1370, the encoder can convert the fourth attribute set (a 0 The further updated attribute value 'up up a' 0 'Encoded into bitstream.'
[0150] The bit stream may include data representing transformed attributes (e.g., representing the first residual value 'res a'). 2 '、Third residual'res up a 1 'and / or the further updated attribute values of the fourth attribute set' up up a 0 The encoder can further update the attribute values of the fourth attribute set.0 'Encoding is in the bitstream. The encoder can, for example, perform further updates to the current attribute value of the code to be written based on the attributes of the written codes in the fourth attribute set.'up upa 0 j Intra-frame prediction. This can improve the further updated attribute values of the fourth attribute set. 0 Compression efficiency.
[0151] For example, if lossy attribute write coding is allowed, the encoder can quantize the further updated attribute values of the fourth attribute set. 0 '、First residual value'res a 2 'and / or third residual' res up a 1 The encoder can further update the attribute values of the fourth attribute set. 0 '、First residual value'res a 2 '(or the first quantized residual value'res a) 2 ') and / or the third residual value 'res up a 1 '(or the quantified third residual value'res up a 1 ') Entropy encoding is incorporated into the bitstream.
[0152] Figure 14 An example method for decoding point cloud attributes is shown. Decoding can be based on a prediction-enhancing transformation scheme with predictions and updates. Figure 14 One or more steps can be accomplished by a decoder (e.g., Figure 1 The decoder (120) in the document is executed. This can be done as described in this article. Figure 13 The description describes encoding a set of attributes 'a'. The set of attributes 'a' (e.g., the encoded set of attributes 'a') can be decoded. Decoding can use predictions between one or more (e.g., three) levels of detail, from the lowest (e.g., the third) level of detail to the highest (e.g., the first) level of detail (as described in this paper regarding...). Figure 13 The encoder described is in reverse order.
[0153] At step 1410, the decoder can obtain the first residual 'res a' from the bitstream pair. 2 '、Third residual value' res up a 1 'and / or the fourth set of attributes (a 0 The further updated attribute value 'up up a' 0 Decoding is then performed. The decoder can use (e.g., by application) optional dequantization for lossy compression. At step 1475, the decoder can obtain the decoded third residual value from the decrypted third residual value. 1'Determine the updated attribute value. Step 1475 can be performed in a manner similar to that described herein with respect to step 1385. The decoder can determine the updated attribute value based on the third residual value. The updated attribute value can be determined as the third residual value multiplied by a scaling factor (e.g., ½, ¼, 1 / 8), which can be predetermined or signaled in the bit stream.'
[0154] At step 1480, the decoder can, for example, obtain the fourth attribute set (a 0 The decoded and further updated attribute value 'up up a 0 Subtract the updated attribute value from the middle to determine the fourth attribute set (a) 0 The updated attribute value 'upa' 0 At step 1420, the decoder can obtain the fourth attribute set (a 0 The updated attribute value 'up a' 0 'Determine the set of third attributes (a)' 1 The updated attribute value 'up a' 1 The predicted value of '. Step 1420 can be performed in a similar manner to that described herein with respect to step 1350. At step 1430, the decoder can, for example, add the predicted value to the decoded third residual value 'res up a 1 'To determine the third attribute set (a) 1 The updated attribute value 'up a' 1 At step 1440, the decoder can, for example, merge the third attribute set (a 1 ) and the fourth attribute set (a 0 ) to determine the second attribute set (a 0 and a 1 Step 1440 can be the inverse operation of step 1340. The second attribute set (a) 0 and a 1 This can include the updated attribute value 'up a'. 0 'and updated attribute values'up a 1 At step 1485, the decoder can obtain the first residual value 'resa' from the decoded value. 2 'Determine the updated attribute value. Step 1485 can be performed in a similar manner to that described herein with respect to step 1375. The operations at steps 1420, 1430, 1440, 1475, and 1480 can be iteratively applied (e.g., applied to) each successively higher (e.g., finer) LoD.
[0155] At step 1490, the decoder can, for example, obtain the second attribute set (a 0 and a 1 The updated attribute value 'up a' 0 'Neutralize the updated attribute values from the second attribute set'up a 1 Subtract the updated attribute value from the middle to determine the second attribute set (a) 0 and a 1 The attribute value 'a' 0 'and attribute value'a 1 At step 1450, the decoder can obtain the second attribute set (a 0 and a 1 The attributes in ) determine the first attribute set (a 2 The predicted values of the attributes in ) are then used. Step 1450 can be performed in a similar manner to that described herein with respect to step 1320. At step 1460, the decoder can, for example, add the predicted values to the decoded second residual value 'res a'. 2 'Determine the first set of attributes (a) 2 The decoded attributes in ) . At step 1470, the decoder can, for example, merge the first attribute set (a 2 ) and the second attribute set (a 0 and a 1 This determines the set of decoded attributes 'a' of the decoded geometry S (e.g., the entire decoded geometry). Step 1470 can be the inverse operation of step 1310.
[0156] Predictive lifting schemes can be similar to lifting schemes used for (e.g., applied to) wavelet writing. Predictive lifting schemes may include update steps not present in the predictive transform scheme (e.g., in addition to the prediction step in the predictive transform scheme). Update steps can provide better compression performance (e.g., in conjunction with the prediction step). This allows energy to be compressed at the lowest level of detail, which can reduce distortion in lossy writing.
[0157] The RAHT scheme can be used to encode attributes. The RAHT scheme can be used iteratively based on a two-point transform. In point cloud attribute encoding, the two-point RAHT transform can be used (e.g., applied to) two attribute sets (A1 and A2). Each of A1 and A2 can have w1 and w2 numbers of attributes, respectively. Each of A1 and A2 can have a corresponding correlation coefficient c. A1 and cA2。 c A1 and c A2 Each value in the set represents the sum of the attribute values in the corresponding set divided by the square root of the number of attributes.
[0158] . ( (*)
[0159] A two-point RAHT transformation can depend on weights w1 and w2. The two-point RAHT transformation can be defined by the following 2×2 matrix.
[0160] .
[0161] Two new coefficients, DC and AC, can be determined, for example, when used (e.g., applied to) two coefficients c. A1 and c A2 In this case.
[0162]
[0163] As described in this article, the above properties (*) regarding coefficients hold for DC coefficients.
[0164]
[0165] A two-point RAHT transformation can be iteratively applied (e.g., applied to) the DC coefficients. This can be called the RAHT iterative method. For example, once determined, the AC coefficients may not undergo further transformations. At the start of the RAHT iterative method, there may be an initial set A of properties as many as the points present in the written code geometry S. i Each initial attribute set A i It can contain one attribute (w) i =1). Coefficient c Ai It can be equal to the value of a property, and / or can satisfy property (*). By induction, property (*) can hold for, for example, subsequent DC coefficients (e.g., all subsequent DC coefficients) determined after iterative application of the two-point RAHT transformation.
[0166] At a specific stage of the RAHT iterative method, the determined coefficients can be the union of the set of DC coefficients and the set of AC coefficients satisfying property (*). The RAHT iterative method can continue until all DC coefficients are exhausted, and may leave only one DC coefficient. A DC coefficient can be equal to C. A , where A can be a set of attributes (e.g., all attributes) associated with the writable code geometry S (e.g., the complete writable code geometry S). The RAHT iterative method can follow the order in the DC coefficient pairs.
[0167] The two-point inverse Raht transform can be defined by the following 2×2 matrix.
[0168]
[0169] The two-point inverse RAHT transform can be used (e.g., applied to) DC and AC coefficients to recover the two coefficients c. A1and cA2。
[0170]
[0171] The inverse iterative RAHT method obtains the DC and AC coefficients by applying the inverse two-point RAHT in reverse order (e.g., applying it to) the DC and AC coefficients, relative to obtaining them through the iterative RAHT method. At the end of the inverse iterative RAHT method, the initial property set A can be obtained. i The associated coefficient c Ai These coefficients c Ai It can be equal to the initial set A i The value of an associated attribute.
[0172] For lossy RAHT compression of attributes, the coefficients can be further compressed, for example, before the coefficients are encoded in the bitstream, based on quantization applied to (e.g., applied to) the DC and AC coefficients. The decoder can, for example, use (e.g., apply) dequantization after decoding the quantized DC and AC coefficients from the bitstream.
[0173] The RAHT iterative method can, for example, follow an octree in a specific iterative order within G-PCC. One or more (e.g., up to eight) DC coefficients associated with one or more (e.g., up to eight) occupied child nodes of a parent node in the octree can undergo a cascade of two-point RAHT transformations until one DC coefficient remains, along with the remaining (e.g., up to seven) AC coefficients. This DC coefficient can be pushed at the parent node level. The method can be repeated, for example, at higher octree depths until the root node is reached.
[0174] Figure 15 An example RAHT transformation is shown. The RAHT transformation can be applied to the child nodes of an octree parent node along three consecutive directions. The parent node 1500 can have multiple (e.g., five) occupied child nodes, each with a corresponding association coefficient c. i and weight w i The first RAHT transformation 1510 can be performed along the first direction 1511. If there are two adjacent occupied child nodes 1513 along the first direction, these two adjacent occupied child nodes 1513 can undergo a two-point RAHT transformation. New DC coefficients 1514 and AC coefficients 1515 can be determined and pushed into the AC coefficient set 1550. For example, if there is only one occupied child node 1516 along the first direction, the node can remain as is, and its DC coefficients can remain 1517. In this example, the child node can be folded along the first direction to determine a new set of nodes 1519 (e.g., a set of three nodes), and associated new DC coefficients can be determined.
[0175] A second RAHT transformation 1520 can be performed, for example, after the first RAHT transformation. The second RAHT transformation can be performed along a second direction 1521. The second RAHT transformation can be performed similarly to the first RAHT transformation. Child node 1522 can be determined. Child node 1522 may have been folded along the first two directions 1511 and 1521. AC coefficient 1523 can be pushed to AC coefficient set 1550. A third RAHT transformation 1530 can be performed, for example, after the second RAHT transformation. The third RAHT transformation can be performed along a third direction 1531. The third RAHT transformation can be performed similarly to the first and / or second RAHT transformation. Child node 1532 can be determined (e.g., uniquely). Child node 1532 can be generated by folding along all three directions. AC coefficient 1533 can be pushed to AC coefficient set 1550. Child node 1532 may have associated DC coefficients that are pushed to the parent node (e.g., as...). Figure 16 As shown in the image).
[0176] Figure 16 An example RAHT transformation is illustrated. The RAHT transformation can be used (e.g., applied to) octree nodes at depth 'd' (e.g., all octree nodes) to determine the DC coefficients and AC coefficients at depth d-1. Occupied nodes 1600 of the octree can be at depth d. Occupied node 1600 can undergo the RAHT transformation along three directions. The DC coefficients of each node can be pushed to the corresponding occupied parent node 1610 belonging to the octree at depth d-1. The three DC coefficients of child node 1601 can, for example, undergo the RAHT transformation along three directions. A unique DC coefficient associated with parent node 1611 can be determined, and two AC coefficients 1621 can be pushed to AC coefficient set 1620. By performing this method on (e.g., all) occupied nodes 1600 of the octree at depth d, the DC coefficients associated with the occupied nodes of the octree at depth d can be transformed into DC coefficients associated with the occupied node 1610 of the octree at depth d-1 and the AC coefficient set 1620.
[0177] This bottom-up approach can, for example, be repeated depth-by-depth until a minimum depth (root node) is reached. The result of the RAHT transformation on an octree (e.g., a complete octree) can be a set of coefficients including unique DC coefficients and a set (many) of AC coefficients. The RAHT transformation method can start from the highest depth of the unique point (voxel) of the occupied child node corresponding to the write-coded point cloud S. The unique point can be associated with a unique attribute in the attribute set 'a'. The DC coefficient at the highest depth can be set to the value of the unique attribute associated with each occupied node. The weight 'w' can be set to 1.
[0178] The inverse RAHT method on an octree can be a top-down approach, descending from the root node to the final depth comprised of leaf nodes, each containing a point in the point cloud and an associated attribute. For example, by applying the inverse two-point RAHT transformation to (e.g., to) the DC coefficients of each occupied node at depth d-1 of the octree and using the associated AC coefficients from the AC coefficient set 1620, the DC coefficients of occupied node 1610 at depth d-1 of the octree can be inversely transformed to the DC coefficients of occupied node 1600 at depth d of the octree. The inverse two-point RAHT transformation can be applied in reverse order along three directions to reverse the relationships described herein. Figure 15 The described node transformation allows for the acquisition of DC coefficients for leaf nodes, with each DC coefficient corresponding to a unique attribute associated with each leaf node. Similar to the geometric write-coding of point clouds, the write-coding of attributes associated with points in the current point cloud can benefit from inter-frame prediction using a motion-compensated point cloud. A motion-compensated point cloud can inherit attributes from a motion-compensated reference point cloud. For example, if motion occurs, a point can retain its associated attributes. Motion-compensated attributes (e.g., attributes associated with points in the motion-compensated point cloud) can be used. This improves the compression of attributes of the write-coded geometry of the current point cloud.
[0179] Inter-frame prediction upscaling can use motion-compensated attributes or residual attributes based on the difference between the attribute and the motion-compensated attribute. Using residual attributes can increase compression. Inter-frame prediction upscaling can be used in prediction upscaling, for example, by inserting it into prediction steps 1120, 1150, 1220, 1250, 1320, 1350, 1420, and / or 1450. This can be achieved, for example, through lower levels of detail. subset a 0 ... a i-1 The attribute (or residual attribute) is used to predict the point set S. i The attribute (or residual attribute) a i Alternatively, it can be achieved by comparing with the motion-compensated point cloud S. inter The set a associated with the points inter The motion-compensated properties (or residual properties) are used to predict the point set S. i The attribute (or residual attribute) a i For example, it can be based on subset a. 0 ... a i-1 and enhanced lower level of detail set a inter The prediction step is performed using the properties (or residual properties) of the data.
[0180] The encoder and / or decoder can obtain data from the motion-compensated point cloud S. inter The set a associated with the points inter The set of fourth attributes (or residual attributes) is determined by the motion-compensated attributes (or residual attributes) (coarsest level of detail). subset a 0 The predicted values of the attributes (or residual attributes) in the fourth attribute (or residual attribute) set. The encoder and / or decoder can subtract the predicted values from the attributes (or residual attributes) in the fourth attribute (or residual attribute) set to determine the residual value 'res a'. 0 The encoder can output the residual value 'res a'. 0 The attribute (or residual of the residual attribute) is encoded into the bitstream instead of the attribute (or residual attribute) in the fourth attribute set.
[0181] Inter-frame RAHT schemes can use inter-frame prediction to predict the values of DC and AC coefficients determined by the RAHT iterative method. For example, maintaining the current point cloud S for the code to be written. coded And the motion-compensated point cloud S inter The shared property of an octree structure may be beneficial, as the generation of DC and AC coefficients follows an octree. A common bounding box enclosing the two point clouds can be determined. For both point clouds, octree partitioning can be performed from the root node associated with the common bounding box. For example, if the point clouds are not equal, this may result in the two octree partitions potentially being different. The two octrees can have a common subtree starting from the root node. On the common subtree, the topology of the occupied nodes can be the same, and / or a common set of DC and AC coefficients can be determined for both point clouds. The coefficients can be determined from the motion-compensated point cloud S. inter The DC and AC coefficients, determined by the attributes, are used to predict nodes associated with common subtrees and from the current point cloud S. coded The attributes determine a subset of the DC and AC coefficients. The encoder and / or decoder can, for example, be derived from nodes associated with a common subtree and from the current point cloud S. coded Subtract the DC and AC coefficients determined by the motion-compensated point cloud S from the attribute-determined values. inter The DC and AC coefficients, determined by the attributes, determine the coefficient residual values. The encoder can encode the coefficient residual values into the bitstream, and / or can choose not to associate them with nodes in the common subtree and from the current point cloud S. coded The DC and AC coefficients, whose attributes are determined, are encoded into the bitstream.
[0182] The DC and AC coefficients that are not associated with nodes in the common subtree may not be predicted, and / or the DC and AC coefficients may be directly written in a manner similar to that performed without inter-frame prediction. Alternatively, the current point cloud S may be determined at a certain depth. codedThe predicted DC coefficients (e.g., without predicting AC coefficients at the same depth). This can be based on the assumption that the current point cloud S... coded The octree and the motion-compensated point cloud S inter The octrees all have the same node occupancy rate at this depth. This can be seen from the motion-compensated point cloud S. inter The predicted DC coefficients are determined by the corresponding co-located DC coefficients. This can be achieved, for example, by using the current point cloud S. coded The DC residual value is determined by subtracting the predicted DC coefficients from the DC coefficients. The RAHT transform can, for example, begin by increasing the DC residual value in the octree from the DC coefficients of the current point cloud after the predicted DC coefficients have been determined.
[0183] A RAHT scheme that does not use information from a reference frame different from the current frame can be called an intra-frame RAHT scheme. Intra-frame prediction can be performed between the DC and AC coefficients of an intra-frame RAHT scheme. Intra-frame inter-depth prediction within the current frame may have been integrated into a RAHT scheme such as GGC. The inter-frame depth prediction mechanism can, for example, predict the DC coefficients associated with nodes at depth d in an octree by interpolating the DC coefficients associated with nodes at a lower depth d-1 in the octree.
[0184] Figure 17 An example method for encoding the current point cloud frame is shown. More specifically, Figure 17 An example method for encoding the geometry and properties of a current point cloud frame is shown. The encoding can be performed using a motion-compensated point cloud frame, which is determined from a motion field selected for geometry coding (e.g., to optimize or reduce geometric discrepancies). Figure 17 One or more steps of the example method (e.g., method 1700) can be accomplished by an encoder (e.g., Figure 1 The encoder 114) is used to perform and / or implement this.
[0185] The encoder can, for example, obtain the current point cloud frame 1710 in a sequence of point cloud frames of a dynamic point cloud. The encoder can also, for example, obtain a coded reference point cloud frame 1705. The coded reference point cloud frame can be obtained, for example, from a plurality of coded reference point cloud frames. At step 1720, the encoder can determine a geometric motion vector (MV) 1721. The geometric motion vector (MV) 1721 can be determined, for example, by performing a geometric motion search. For example, the geometric MV 1721 can approximate a 3D motion field of geometry from the reference point cloud frame 1705 to the current point cloud frame 1710. The geometric MV 1721 can be selected to best approximate the 3D motion field. The geometric MV 1721 can be selected, for example, based on a cost minimization function to best approximate the 3D motion field. The cost minimization function can minimize the difference between the geometry of the reference point cloud frame and the geometry of the current point cloud frame. The geometry of a reference point cloud frame can be adjusted using the geometric MV (e.g., using geometric MV 1721) of a motion-compensated point cloud frame. The encoder can encode the geometric MV information 1722 into a bitstream 1770, for example, as a representation of geometric MV 1721. The geometric MV information can indicate a coded reference point cloud frame 1705 among multiple coded reference point cloud frames.
[0186] At step 1730, the encoder can determine the geometrically motion-compensated point cloud frame 1731. The geometrically motion-compensated point cloud frame 1731 can be determined, for example, by performing motion compensation on the reference point cloud frame 1705 based on geometric MV 1721. At step 1740, the encoder can encode the geometry of the current point cloud frame 1710. The geometry of the current point cloud frame 1710 can be encoded, for example, based on geometric MV 1721. The geometry of the current point cloud frame 1710 can be encoded, for example, based on the geometrically motion-compensated point cloud frame 1731. The geometry of the current point cloud frame 1710 can be encoded as geometric information 1742 into the bitstream 1770 based on the geometrically motion-compensated point cloud frame 1731. The geometry of the current point cloud frame 1710 can be encoded as geometric information 1742 into the bitstream 1770.
[0187] The encoder can provide a decoded geometry 1741 (e.g., a reconstructed geometry) of the current point cloud frame 1710. The encoder can, for example, write (e.g., decode) the encoded geometry to determine a reconstructed geometry corresponding to the decoded geometry at the decoder. For example, if the geometry compression is lossy, the decoded geometry 1741 and the geometry of the current point cloud frame 1710 can be significantly different. At step 1750, the encoder can determine a mapping attribute 1751 associated with the decoded geometry 1741. The mapping attribute can be determined, for example, by mapping attributes associated with the geometry of the current point cloud frame 1710 to the decoded geometry 1741. Attributes can include color. The mapping attribute can be determined, for example, based on recoloring the attribute.
[0188] The mapping attributes can be determined, for example, based on a k-nearest neighbor (KNN) search algorithm that determines the nearest point from the geometry of the current point cloud frame 1710 to the decoded geometry 1741. The k-nearest neighbor (KNN) search algorithm can include, for example, spatial partitioning algorithms such as KD-tree search, ball / metric tree search, brute-force search, etc. The mapping attributes of the points in the decoded geometry 1741 can be, for example, average attribute values. The average attribute values can be associated with the nearest points in the current point cloud frame 1710 relative to the points in the decoded geometry.
[0189] For example, if geometry compression is lossless, the decoded geometry 1741 (e.g., reconstructed geometry) and the geometry of the current point cloud frame 1710 can be the same. Attribute mapping can associate the attributes of each point of the geometry of the current point cloud frame 1710 with each point of the decoded geometry 1741. The mapped attributes can be attributes of the current point cloud frame 1710. At step 1760, the encoder can encode the mapped attributes 1751. The mapped attributes 1751 can be encoded as attribute information 1762 into bitstream 1770. The mapped attributes 1751 can be encoded as attribute information 1762, for example, based on the attributes of the geometry motion-compensated point cloud frame 1731.
[0190] Figure 18 An example method for decoding the current point cloud frame is shown. More specifically, Figure 18 An example method for decoding the geometry and attributes of the current point cloud frame is shown. This can be seen as described in this article regarding... Figure 17 The described method encodes geometry and attributes. Decoding can be performed using motion-compensated point cloud frames, determined from a motion field selected (e.g., optimized to reduce geometric discrepancies) for coding the geometry. Method 1800 illustrates an operation corresponding to but inverse of the operation in method 1700 as described herein. Method 1800 can be the inverse of method 1700. Figure 18One or more steps of the example method (e.g., method 1800) can be decoded by a decoder (e.g., Figure 1 The decoder 120 is used to perform and / or implement the process.
[0191] The decoder can obtain the decoded reference point cloud frame 1805. The decoder can obtain the decoded reference point cloud frame 1805, for example, from multiple decoded reference point cloud frames. At step 1810, the decoder can determine the geometric MV 1811. The geometric MV 1811 can be determined, for example, by decoding the geometric MV information 1812 from the bit stream 1850. As per this document regarding... Figure 17 The described geometry MV 1721 can be geometry MV 1811. At step 1820, the decoder can determine the geometry motion-compensated point cloud frame 1821. The geometry motion-compensated point cloud frame 1821 can be determined, for example, by performing motion compensation on a reference point cloud frame 1805. Performing motion compensation on the reference point cloud frame 1805 to determine the geometry motion-compensated point cloud frame 1821 can be based on geometry MV 1811. At step 1830, the decoder can determine the decoded geometry 1831 of the current point cloud frame. The decoded geometry 1831 of the current point cloud frame can be determined, for example, by decoding the geometry information 1832 from the bit stream 1850 based on the geometry motion-compensated point cloud frame 1821. At step 1840, the decoder can decode the attribute information 1842 from the bit stream 1850. The decoder can determine the decoded attributes 1841 associated with the decoded geometry 1831. The decoder can, for example, determine the decoded attribute 1841 associated with the decoded geometry 1831 based on the decoded attribute information 1842 and the attributes associated with the geometrically motion-compensated point cloud frame 1821.
[0192] In at least some systems, inter-frame attribute write codes (e.g., as discussed in this paper) Figure 17 Step 1760 and decoding (e.g., as described herein with respect to step 1840) can be based on geometrically motion-compensated point cloud frames (e.g., geometrically motion-compensated point cloud frames 1731, 1821) determined from geometric MVs optimized for geometric writing (e.g., geometric MVs 1721, 1811). However, the choice of geometry-based MVs may not be optimal for attribute writing.
[0193] Geometric motion search (e.g., as discussed in this paper) Figure 17The geometric motion search described in step 1720 can be an iterative method that tests multiple candidate geometric motion vectors. The iterative method can also select candidate geometric motion vectors from the candidate geometric motion vectors that minimize geometric (e.g., spatial) distortion between the geometry of the current point cloud frame 1710 and the geometric vector determined by motion compensation using the candidate geometric motion vectors from the reference point cloud frame 1705. The selected geometric motion vector may not provide a geometrically motion-compensated point cloud frame (e.g., geometrically motion-compensated point cloud frames 1731, 1821) that is effective in compressing mapping properties (e.g., compressing mapping property 1751 as described herein with respect to step 1760).
[0194] For example, for some types of point clouds, it may be impossible to determine a motion-compensated point cloud frame that is favorable for predicting the geometry and attributes of the current point cloud frame. For instance, a static object (e.g., a house) may have a static geometry but attributes that move over time (e.g., non-static or changing) (e.g., the shadow of a moving object projected onto the walls of the house). The optimal motion vector field for coding the geometry of the object (e.g., the house) may be determined as a zero motion field (i.e., the motion vectors are equal to zero). However, the optimal motion vector field for coding the attributes of the object (e.g., the color of the walls) may be determined as a non-zero motion field (e.g., the motion of the shadow on the walls approximates the motion of the shadow). As another example, a moving object (e.g., a moving semi-trailer truck) may have a static local color (e.g., the semi-trailer may be entirely white but cast with the shadow of a static object). The optimal motion field for geometry coding may be determined as non-zero. The optimal motion field for attribute coding may be determined as zero. For example, if the motion field is unfavorable for geometry coding or attribute coding, suboptimal compression may be obtained because the size of the bitstream increases.
[0195] As described in this paper, the optimal geometric motion field may not always be optimal for attribute coding. The improvements described in this paper include advantages such as determining the attribute motion field consisting of attribute motion vectors selected (e.g., optimized) for attribute coding. Attribute motion vectors can be used to determine the attribute motion-compensated point cloud frame for coding (e.g., encoding and / or decoding) the attributes of the current point cloud frame.
[0196] Dual motion fields can be used to encode and / or decode the geometry and attributes of the current point cloud frame. Geometric motion fields can be used for geometric encoding and / or decoding. Attribute motion fields can be used for attribute encoding and / or decoding. Attribute motion fields can include attribute motion vectors. Attribute motion vectors can be used to perform motion compensation on an attribute reference point cloud frame, for example, to determine an attribute motion-compensated point cloud frame for encoding and / or decoding the attributes of the current point cloud frame.
[0197] Figure 19 An example method for encoding the current point cloud frame is shown. More specifically, Figure 19 An example method for encoding the geometry and attributes of the current point cloud frame is shown. Figure 19 One or more steps of the example method (e.g., method 1900) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 35 Example computer system 3500 and / or Figure 36 Example computing device 3630 is used to perform and / or implement. Figure 19 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0198] As this article is about Figure 17 Steps 1720, 1730, 1740, and 1750, described for encoding the geometry of the current point cloud frame 1710, can be respectively compared with those described herein. Figure 19 The steps 1920, 1930, 1940, and 1950 described for encoding the geometry of the current point cloud frame 1910 are the same. (See also the section on...) Figure 17 The described geometric MV1721, geometrically motion-compensated point cloud frame 1731, decoded geometry 1741, and mapping attribute 1751 can be respectively compared with those described in this paper. Figure 19 The described geometric MV 1921, geometrically motion-compensated point cloud frame 1931, decoded geometry 1941, and mapping attribute 1951 are the same.
[0199] The encoder can, for example, obtain the current point cloud frame 1910 in a sequence of point cloud frames of a dynamic point cloud. The encoder can, for example, obtain a coded reference point cloud frame 1905 and a coded attribute reference point cloud frame 1981 from a plurality of decoded reference point cloud frames. The coded reference point cloud frame 1905 can be selected from a first plurality of coded reference point cloud frames. The attribute reference point cloud frame 1981 can be selected from a second plurality of coded point cloud frames. The first plurality of coded reference point cloud frames and the second plurality of coded reference point cloud frames can be the same. The first plurality of coded reference point cloud frames and the second plurality of coded reference point cloud frames can be different. At step 1980, the encoder can, for example, determine the attribute motion vector (MV) 1983 by performing an attribute motion search. The encoder can perform the attribute motion search such that the attribute MV 1983 best approximates the 3D attribute motion field from reference point cloud frame 1981 to the current point cloud frame 1910. The encoder can encode attribute MV information 1982 into bitstream 1970 as a representation of attribute MV 1983. Attribute MV information can indicate the coded reference point cloud frame 1981 in multiple coded reference point cloud frames.
[0200] Attribute motion search (e.g., as described herein with respect to step 1980) can be an iterative method of locally testing multiple candidate attribute motion vectors. Attribute motion search can select candidate attribute motion vectors that minimize attribute distortion from the candidate attribute motion vectors. Attribute distortion may occur between attributes of the current point cloud frame 1910 and attributes determined by motion compensation (e.g., adjustment) of attributes of attribute reference point cloud frame 1981. Attributes can be determined, for example, by performing motion compensation of attributes of attribute reference point cloud frame 1981 using candidate attribute motion vectors. Attribute distortion can be based on the difference between attributes of the current point cloud frame 1910 and attributes determined by motion compensation of attributes of attribute reference point cloud frame 1981. Attribute distortion can be based, for example, on the sum of absolute differences (SAD), sum of absolute transformed differences (SATD), sum of squared errors (SSE), etc., between attributes of the current point cloud frame 1910 and attributes determined by motion compensation of attributes of attribute reference point cloud frame 1981. Attributes can be determined, for example, by performing motion compensation of attributes of attribute reference point cloud frame 1981 using candidate attribute motion vectors.
[0201] At step 1990, the encoder can determine the attribute motion-compensated point cloud frame 1991. The encoder can determine the attribute motion-compensated point cloud frame 1991, for example, by performing motion compensation on the attribute reference point cloud frame 1981 based on attribute MV 1983. The attribute distortion of a point for attribute motion search can be determined, for example, by comparing the attribute value of a point in the current point cloud frame 1910 with the attribute value of one of the nearest neighbors of that point in the attribute motion-compensated point cloud frame 1991.
[0202] At step 1960, the encoder can encode the mapped attribute 1951 as attribute information 1962 into the bitstream 1970. The encoder can encode the mapped attribute 1951 as attribute information 1962 into the bitstream 1970, for example, based on the attributes of the attribute motion-compensated point cloud frame 1991. The encoder can encode the mapped attribute 1951 as attribute information 1962 based on the attributes of the attribute motion-compensated point cloud frame 1991 instead of the geometrically motion-compensated point cloud frame 1931. The encoder can encode the mapped attribute 1951 as attribute information 1962 based on the attribute motion vector 1983.
[0203] For example, if a geometric motion search is performed, a motion field can be determined. This determined motion field can minimize geometric distortion. For example, the motion field determined by performing a geometric motion search (as described herein with respect to step 1920) can minimize geometric distortion. For example, if an attribute motion search is performed, a motion field can be determined. This determined motion can minimize attribute distortion. For example, the motion field determined by performing an attribute motion search as described herein with respect to step 1980 can minimize attribute distortion. Geometrically motion-compensated point cloud frames and attribute-motion-compensated point cloud frames (e.g., geometrically motion-compensated point cloud frame 1931 and attribute-motion-compensated point cloud frame 1991) can effectively predict the geometry and attributes of the current point cloud frame 1910, respectively. The compression capability of the point cloud encoder can be improved, for example, by using geometrically motion-compensated point cloud frames and attribute-motion-compensated point cloud frames to predict the geometry and attributes of the current point cloud frame 1910, respectively.
[0204] Figure 20 An example method for decoding the current point cloud frame is shown. More specifically, Figure 20 An example method for decoding the geometry and attributes of the current point cloud frame is shown. Figure 20 One or more steps of the example method (e.g., method 2000) can be decoded by a decoder (e.g., Figure 1 decoder 120) Figure 35 Example computer system 3500 and / or Figure 36 Example computing device 3630 is used to perform and / or implement. Figure 20 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0205] Decoding methods can be used to decode data as described in this article. Figure 19 The point cloud frames encoded using the described encoding method are then decoded, as described in this paper. Figure 18 Steps 1810, 1820, and 1830 of the described decoding method can be respectively compared with... Figure 20 The steps in 2010, 2020, and 2030 are the same. (See the section on...) Figure 18 The described geometric MV 1811, geometrically motion-compensated point cloud frame 1821, and decoded geometry 1831 can be identical to geometric MV 2011, geometrically motion-compensated point cloud frame 2021, and decoded geometry 2031, respectively. The decoder can, for example, obtain a coded reference point cloud frame 2005 and a decoded attribute reference point cloud frame 2071 from multiple decoded reference point cloud frames. The coded reference point cloud frame 1905 and the coded reference point cloud frame 2005 can be identical. The decoded attribute reference point cloud frame 1981 and the decoded attribute reference point cloud frame 2071 can be identical.
[0206] At step 2060, the decoder can determine attribute motion vector (MV) 2061. The decoder can determine attribute motion vector (MV) 2061, for example, by decoding attribute motion vector information 2062 from bitstream 2050. Attribute MV information 2062 and attribute motion vector (MV) information 1982 can be the same. Attribute MV 1983 and attribute MV 2061 can be the same. At step 2070, the decoder can determine attribute motion-compensated point cloud frame 2072. The decoder can determine attribute motion-compensated point cloud frame 2072, for example, by performing motion compensation on attribute reference point cloud frame 2071. The motion compensation of attribute reference point cloud frame 2071 can be based on attribute MV 2061. At step 2040, the decoder can decode attribute information 2042 from bitstream 2050. The decoder can determine the decoded attribute 2041 of the current point cloud frame, for example, based on the attribute associated with attribute motion-compensated point cloud frame 2072.
[0207] As this article is about Figure 20 The described decoding method can generate decoded attribute 2041 for the current point cloud frame. (See the section on...) Figure 20 The described decoding method can, for example, generate decoded attributes 2041 for the current point cloud frame by decoding attribute information 2042. Attribute information 2042 can represent attributes associated with decoded geometry 2031. Decoding attribute information 2042 can be based on attribute motion-compensated point cloud frame 2072, rather than geometry motion-compensated point cloud frame 2021.
[0208] Inter-frame attribute encoding (e.g., as described herein with respect to step 1960) and decoding (e.g., as described herein with respect to step 2040) can be performed by any inter-frame attribute writer. Inter-frame attribute encoding (e.g., as described herein with respect to step 1960) and decoding (e.g., as described herein with respect to step 2040) can be a prediction boosting scheme. The prediction boosting scheme can use points from attribute motion-compensated point cloud frames (e.g., 1991, 2072) in its prediction step. Point set S i Attribute a i It can be from a lower level of detail subset a 0 ... a i-1 The prediction is based on the attributes in the point cloud frame and is made by the attributes associated with the points in the attribute-motion compensated point cloud frame (e.g., attribute-motion compensated point cloud frame 1991, attribute-motion compensated point cloud frame 2072).
[0209] Inter-frame attribute encoding (e.g., as described herein with respect to 1960) and decoding (e.g., as described herein with respect to 2040) can be inter-frame RAHT schemes. Inter-frame RAHT schemes can predict the RAHT coefficients (e.g., AC and DC coefficients) of the current point cloud frame. Inter-frame RAHT schemes can predict the RAHT coefficients (e.g., AC and DC coefficients) of the current point cloud frame, for example, based on the RAHT coefficients (e.g., AC and DC coefficients) of attribute motion-compensated point cloud frames (e.g., attribute motion-compensated point cloud frame 1991, attribute motion-compensated point cloud frame 2072).
[0210] Figure 21 An example method for encoding attributes is shown. More specifically, Figure 21 An example method for encoding the attributes of the current point cloud frame is shown. Figure 21 One or more steps of the example method (e.g., method 2100) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 35 Example computer system 3500 and / or Figure 36 Example computing device 3630 is used to perform and / or implement. Figure 21 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0211] At step 2110, the encoder can determine the projection attribute 2111. The encoder can then generate the attribute-motion-compensated point cloud frame 1991 (as described in this document regarding...). Figure 19 The properties described are projected onto the decoded geometry 1941 (as described in this paper). Figure 19The encoder can determine the projection attribute 2111, for example, based on the attribute motion vector (e.g., attribute motion vector 1983).
[0212] At step 2120, the encoder can encode the attribute information 2121 in bit stream 1970 (as described herein). Figure 19 In the description), attribute information 2121 can be represented by the decoded geometry 1941 (as described in this document). Figure 19 (As described) Related to (as discussed in this article) Figure 19 The mapping attribute 1951 described in the diagram. The encoder can encode the attribute information 2121 based on, for example, attribute predictions determined from the projection attribute 2111 (e.g., an attribute predictor).
[0213] The attribute predictor may include a corresponding attribute predictor for each corresponding point in the points of the decoded geometry 1941. The corresponding attribute predictor may be based on the projection attribute corresponding to that point in the projection attributes. The projection attribute and the mapping attribute may belong to the same geometry of the point cloud frame (e.g., decoded geometry 1941). For example, if geometric differences have been eliminated, prediction of the mapping attribute 1951 based on the projection attribute 2111 may be more efficient.
[0214] Figure 22 An example method for decoding the properties of the current point cloud frame is shown. Figure 22 One or more steps of the example method (e.g., method 2200) can be decoded by a decoder (e.g., Figure 1 decoder 120) Figure 35 Example computer system 3500 and / or Figure 36 Example computing device 3630 is used to perform and / or implement. Figure 22 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0215] At step 2210, the decoder can determine projection attributes 2211. This can be achieved, for example, by projecting the attributes of the attribute-motion-compensated point cloud frame 2072 onto the decoded geometry 2031. Figure 20 The decoder determines the projection attribute 2211, for example, based on the attribute motion vector (e.g., attribute motion vector 2061). At step 2220, the decoder can determine the decoded attribute 2041 of the current point cloud frame. The decoded attribute 2041 can be determined, for example, based on an attribute prediction (e.g., an attribute predictor) determined from the projection attribute 2211. The decoded attribute 2041 can be determined, for example, by decoding the attribute information 2042. Figure 20The attribute prediction (e.g., an attribute predictor) is determined from the projection attribute 2211. The attribute predictor may include a corresponding attribute predictor for each corresponding point in the points of the decoded geometry 2031. The corresponding attribute predictor for each corresponding point in the points of the decoded geometry 2031 may be based on the projection attribute corresponding to that point in the projection attribute.
[0216] Figure 23 An example method for encoding mapped attributes is shown. More specifically, Figure 23 An example method for encoding mapped attributes based on attribute predictions (e.g., attribute predictors) is shown. Attribute predictions (e.g., attribute predictors) can be determined, for example, from projected attributes. Figure 23 One or more steps of the example method (e.g., method 2300) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 35 Example computer system 3500 and / or Figure 36 Example computing device 3630 is used to perform and / or implement. Figure 23 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added. Method 2300 may correspond to the method described herein. Figure 21 Step 2120 as described.
[0217] At step 2310, the encoder can determine the predicted attribute. The encoder can, for example, determine the predicted attribute from the projected attribute 2111 (as described herein). Figure 21 The encoder can determine the predicted attribute, for example, based on the predicted attribute (e.g., an attribute predictor) and the mapped attribute 1951 (as described herein). Figure 19 The encoder determines the residual attribute 2311 by the difference between the residual attribute 2311 and the quantized residual attribute 2321. The encoder can determine the quantized residual attribute 2321, for example, by quantizing the residual attribute 2311. For example, if lossy compression is permitted and / or performed, the encoder can determine the quantized residual attribute 2321 by quantizing the residual attribute 2311. At step 2330, the encoder can encode the residual attribute or the quantized residual attribute 2321 into attribute information 2121 (as described herein). Figure 21 (As described).
[0218] Figure 24 An example method for decoding attribute information is shown. More specifically, Figure 24 An example method for decoding attribute information based on attribute predictions determined from projected attributes is shown. Figure 24 One or more steps of the example method (e.g., method 2400) can be decoded by a decoder (e.g., Figure 1 decoder 120) Figure 35 Example computer system 3500 and / or Figure 36 The example computing device 3630 is used to perform and / or implement this. Figure 24 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added. Method 2400 may correspond to the method described herein. Figure 22 Step 2220 as described.
[0219] At step 2410, the decoder can determine the quantized residual attribute 2411 (or residual attribute). The decoder can, for example, determine the attribute information 2042 (as described herein). Figure 20 Decoding is performed to determine the quantized residual attribute 2411 (or residual attribute) as described herein. At step 2420, the decoder can determine the decoded residual attribute 2421. The decoder can determine the decoded residual attribute 2421, for example, by dequantizing the quantized residual attribute 2411. At step 2430, the decoder can determine the decoded residual attribute 2421 from the projection attribute 2211 (as described herein). Figure 22 The decoder can determine the predicted attribute (e.g., the attribute predictor) as described herein. Figure 20 (As described). The decoder can determine the decoded attribute 2041, for example, based on adding the decoded residual attribute 2421 to the prediction attribute 221 (e.g., an attribute predictor).
[0220] Figure 25 An example method for encoding mapped attributes is shown. More specifically, Figure 25 An example method for encoding mapped attributes based on attribute predictions determined from projected attributes is shown. Figure 25 One or more steps of the example method (e.g., method 2500) can be achieved through an encoder (e.g., Figure 1 encoder 114) Figure 35 Example computer system 3500 and / or Figure 36 The example computing device 3630 is used to perform and / or implement this. Figure 25 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0221] At step 2510, the encoder can determine the smoothed projection attribute 2511. The encoder can, for example, by smoothing the projection attribute 2111 (as described herein). Figure 21(As described) to determine the smoothing projection property 2511. The smoothing projection property 2111 can remove pseudo-high frequencies that may impair compression performance.
[0222] At step 2520, the encoder can determine the predicted attribute (e.g., attribute predictor) from the smoothed projection attribute 2511. The encoder can determine the residual attribute 2311 (as described herein). Figure 23 The encoder can, for example, be based on predicted attributes (e.g., attribute predictors) and mapped attributes 1951 (as described herein). Figure 19 The difference between (as described) is used to determine the residual property 2311.
[0223] Figure 26 An example method for decoding attribute information is shown. More specifically, Figure 26 An example method for decoding attribute information based on attribute predictions determined from projected attributes is shown. Figure 26 One or more steps of the example method (e.g., method 1700) can be decoded by a decoder (e.g., Figure 1 decoder 120) Figure 35 Example computer system 3500 and / or Figure 36 The example computing device 3630 is used to perform and / or implement this. Figure 26 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0224] At step 2610, the decoder can determine the smoothed projection attribute 2611 (e.g., the smoothed value). The decoder can, for example, determine the smoothed projection attribute 2211 (as described herein). Figure 22 The smoothed projection attribute 2611 is determined by (as described herein). The smoothed projection attribute 2221 can remove pseudo-high frequencies that may impair compression performance. At step 2620, the decoder can determine the prediction attribute from the smoothed projection attribute 2611. The decoder can, for example, determine the prediction attribute by using the decoded residual attribute 2421 (as described herein). Figure 24 As described above, the predicted attribute is added to determine the decoded attribute 2041 (as described in this article). Figure 20 (As described).
[0225] Residual attributes can be constructed to utilize (e.g., leverage) inter-frame correlations (e.g., temporal correlations). Residual attributes can be encoded and / or decoded using intra-frame attribute writing schemes to reduce intra-frame correlations and improve compression performance. Residual attributes can be encoded and / or decoded based on prediction with lifting (e.g., pred-lift) schemes. A prediction lifting scheme can include a set S of all residual attributes (or quantized residual attributes) to be encoded and / or decoded.
[0226] Figure 27 An example method for encoding residual properties is shown. Figure 27 One or more steps of the example method (e.g., method 2700) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 35 Example computer system 3500 and / or Figure 36 The example computing device 3630 is used to perform and / or implement this. Figure 27 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added. Method 2700 may correspond to, as described herein, Figure 23 Step 2330 as described.
[0227] At step 2710, the encoder can determine the transformed coefficients 2711. The encoder can, for example, base this on applying the intra-frame transform to (e.g., applying to) the (quantized) residual property 2321 (as described herein). Figure 23 The encoder determines the transformed coefficients 2711 (as described herein). At step 2720, the encoder can determine the quantized transformed coefficients 2721. The encoder can determine the quantized transformed coefficients 2721, for example, by quantizing the transformed coefficients 2711. For example, quantization can be used if lossy compression is permitted and / or performed. At step 2730, the encoder can encode (e.g., entropy encoding) the transformed coefficients or the quantized transformed coefficients 2721 into attribute information 2121 (as described herein). Figure 21 (As described).
[0228] Figure 28 An example method for decoding residual properties is shown. Figure 28 One or more steps of the example method (e.g., method 2800) can be decoded by a decoder (e.g., Figure 1 decoder 114) Figure 35 Example computer system 3500 and / or Figure 36 The example computing device 3630 is used to perform and / or implement this. Figure 28The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added. Method 2800 may correspond to the method described herein. Figure 24 Step 2410 as described.
[0229] At step 2810, the decoder can determine the quantized transformed coefficients 2811 (or transformed coefficients). The decoder can, for example, determine the attribute information 2042 (as described herein). Figure 20 The decoder performs decoding (e.g., entropy decoding) to determine the quantized transform coefficients 2811 (or transform coefficients). At step 2820, the decoder can determine the transform coefficients 2821. The decoder can determine the transform coefficients 2821, for example, by dequantizing the quantized transform coefficients 2811. At step 2830, the decoder can determine the (quantized) residual property 2411. The decoder can determine the (quantized) residual property 2411, for example, based on applying an inverse intra-frame transform to (e.g., applying) the transform coefficients 2821. The operation of the inverse intra-frame transform can correspond to, as described herein, the transformation coefficients 2811, and the transform coefficients 2821. Figure 27 The inverse operation of the intra-frame transform used (e.g., applied) in step 2710 is described.
[0230] Intra-frame transforms may include adaptive DCTs, and inverse intra-frame transforms may include inverse adaptive DCTs (A-DCTs). Intra-frame transforms may include RAHT transforms, and inverse intra-frame transforms may include inverse RAHT transforms of RAHT schemes. Quantization may not be performed, and residual properties may be encoded (e.g., entropy coding) without transformation (e.g., without applying the intra-frame transform as described herein with respect to step 2710 and the inverse intra-frame transform as described herein with respect to step 2830). For example, if lossless property write coding is performed, quantization may not be performed, and residual properties may be encoded without transformation. Intra-frame transforms may include Haar transforms, and inverse intra-frame transforms may include inverse Haar transforms. For example, if lossless property write coding is performed, intra-frame transforms may include Haar transforms, and inverse intra-frame transforms may include inverse Haar transforms.
[0231] Figure 29 An example method for encoding the current point cloud frame is shown. More specifically, Figure 29 An example method for encoding the geometry and attributes of the current point cloud frame is shown. Figure 29 One or more steps of the example method (e.g., method 2900) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 35 Example computer system 3500 and / or Figure 36The example computing device 3630 is used to perform and / or implement this. Figure 29 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0232] At step 2910, the encoder can select an attribute reference point cloud frame from a plurality of coded reference point cloud frames (e.g., as described herein regarding...). Figure 19 The described attribute reference point cloud frame 1981). The encoder can encode the attribute reference point cloud frame information 2911 representing the selected attribute reference point cloud frame (e.g., attribute reference point cloud frame 1981) into a bitstream 1970. Figure 19 In the context of the map, the reference point cloud frame 1981 can be selected independently of the reference point cloud frame 1905 used to encode the mapping attribute 1951.
[0233] Attribute reference point cloud frame information 2911 can indicate the global selection of attribute reference point cloud frames. A single attribute reference point cloud frame 1981 can be selected to encode the mapped attribute 1951. A single attribute reference point cloud frame 1981 and reference point cloud frame 1905 can be different. A single attribute reference point cloud frame 1981 and reference point cloud frame 1905 can be the same. Attribute reference point cloud frame information 2911 can indicate whether attribute reference point cloud frames 1981 and 1905 are the same. Attribute reference point cloud frame 1981 and geometrically motion-compensated point cloud frame 1931 can be different. Attribute reference point cloud frame 1981 and geometrically motion-compensated point cloud frame 1931 can be the same.
[0234] Geometric MV 1921 can provide a good approximation of the motion field between the current point cloud frame 1910 and the attribute reference point cloud frame 1981. For example, if geometric MV 1921 is a good first approximation of attribute MV 1983, then geometric MV 1921 can provide a good approximation of the motion field between the current point cloud frame 1910 and the attribute reference point cloud frame 1981.
[0235] The attribute reference point cloud frame information 2911 can indicate whether the attribute reference point cloud frame 1981 and the geometrically motion-compensated cloud frame 1931 are the same. The attribute reference point cloud frame information 2911 can indicate an index pointing to an index table that references the attribute reference point cloud frame 1981. Each index in the index table can reference a specific reference point cloud frame among multiple coded reference point cloud frames.
[0236] The encoder can determine spatial regions from the decoded geometry 1941. Spatial regions, such as spatial blocks, can be determined from the decoded geometry 1941. A spatial block can be a partition of the 3D space surrounding the point cloud, and each partition can have its own coding parameters, which can be coded independently of each other or independently of the set of tree nodes associated with the decoded geometry 1941. The encoder can select an attribute reference point cloud frame 1981. The encoder can select an attribute reference point cloud frame 1981 for each spatial region to encode the attributes of points belonging to that spatial region in the point cloud frame 1910. Attribute reference point cloud frame information 2911 can indicate the local selection of attribute reference point cloud frame 1981. Attribute reference point cloud frame information 2911 can indicate the definition of the spatial region. The definition of the spatial region can include boundary information, such as a bounding box.
[0237] Attribute reference point cloud frame information 2911 can indicate whether attribute reference point cloud frame 1981 is selected (e.g., global selection) or whether attribute reference point cloud frame 1981 is selected for each spatial region (e.g., local selection). Attribute reference point cloud frame information 2911 can indicate the definition of a spatial region. Attribute reference point cloud frame information 2911 can indicate a reference point cloud frame 1905 for each spatial region. The definition of a spatial region may include boundary information, such as a bounding box. The attribute reference point cloud frame selected for a spatial region (e.g., as described herein with respect to step 2910) may not be equal to the reference point cloud frame 1905 used to encode the geometry of points belonging to that spatial region. The attribute reference point cloud frame selected for a spatial region (e.g., as described herein with respect to step 2910) may be equal to the reference point cloud frame 1905 used to encode the geometry of points belonging to that spatial region.
[0238] The attribute reference point cloud frame information 2911 can indicate whether the attribute reference point cloud frame 1981 selected for a spatial region can be equal to the reference point cloud frame 1905 used to encode the geometry of points belonging to that spatial region. The attribute reference point cloud frame 1981 selected for a spatial region can be equal to the geometrically motion-compensated point cloud frame 1931 used to encode the geometry of points belonging to that spatial region. The attribute reference point cloud frame 1981 selected for a spatial region may not be equal to the geometrically motion-compensated point cloud frame 1931 used to encode the geometry of points belonging to that spatial region.
[0239] The attribute reference point cloud frame information 2911 can indicate whether the attribute reference point cloud frame 1981 selected for a spatial region is equal to the geometrically motion-compensated point cloud frame 1931 used to encode the geometry of points belonging to that spatial region. The attribute reference point cloud frame information 2911 can indicate an index in an index table that references the attribute reference point cloud frame 1981 of the spatial region. Each index in the index table can reference a specific reference point cloud frame among multiple coded reference point cloud frames.
[0240] Figure 30 An example method for decoding the current point cloud frame is shown. More specifically, Figure 30 An example method for decoding the geometry and attributes of the current point cloud frame is shown. Figure 30 One or more steps of the example method (e.g., method 3000) can be decoded by a decoder (e.g., Figure 1 decoder 120) Figure 35 Example computer system 3500 and / or Figure 36 The example computing device 3630 is used to perform and / or implement this. Figure 30 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0241] At step 3010, the decoder can access bitstream 2050 (as described in this document regarding...). Figure 20 The decoder decodes attribute reference point cloud frame information 3011 (as described in this document). The decoder can select attribute reference point cloud frame 2071 from among multiple decoded reference point cloud frames (as described in this document). Figure 20 (As described). The decoder can, for example, select an attribute reference point cloud frame 2071 from multiple decoded reference point cloud frames based on the decoded attribute reference point cloud frame information 3011.
[0242] Attribute reference point cloud frame 2071 can be independent of reference point cloud frame 2005 used to determine decoded attributes 2041 (as described in this article). Figure 20 The attribute reference point cloud frame 2071 can be selected to indicate the global selection of the attribute reference point cloud frame. The attribute reference point cloud frame 2071 can be selected to determine the decoded attribute 2041. Attribute reference point cloud frame 2071 and reference point cloud frame 2005 can be different. Attribute reference point cloud frame 2071 and reference point cloud frame 2005 can be the same. Attribute reference point cloud frame information (e.g., attribute reference point cloud frame information 3011) can indicate whether attribute reference point cloud frame 2071 and reference point cloud frame 2005 are the same.
[0243] The attribute reference point cloud frame 2071 and the geometrically motion-compensated point cloud frame 2021 can be different. A single attribute reference point cloud frame 2071 and the geometrically motion-compensated point cloud frame 2021 can be the same.
[0244] The geometric MV 2011 can provide a good approximation of the motion field between the current point cloud frame 1910 and the attribute reference point cloud frame 1981. For example, if the geometric MV 2011 is a good first approximation of the attribute MV 2061, then the geometric MV 2011 can provide a good approximation of the motion field between the current point cloud frame 1910 and the attribute reference point cloud frame 1981. The attribute reference point cloud frame information 3011 can indicate whether the attribute reference point cloud frame 2071 and the geometrically motion-compensated cloud frame 2021 are the same. The attribute reference point cloud frame information 3011 can indicate an index pointing to an index table that references a single attribute reference point cloud frame 2071. Each index in the index table can reference a specific reference point cloud frame among multiple decoded reference point cloud frames.
[0245] The decoder can determine spatial regions from the decoded geometry 2031. Spatial regions, such as spatial blocks, can be determined from the decoded geometry 2031. A spatial block can be a partition of the 3D space surrounding the point cloud, and each partition can have its own coding parameters, which can be coded independently of each other or independently of the set of tree nodes associated with the decoded geometry 2031. The decoder can select an attribute reference point cloud frame 2071. The decoder can select an attribute reference point cloud frame 2071 for each spatial region to determine decoded attributes 2041 belonging to that spatial region. Attribute reference point cloud frame information 3011 can indicate the local selection of attribute reference point cloud frame 2071. Attribute reference point cloud frame information 3011 can indicate the definition of the spatial region. The definition of the spatial region can include boundary information, such as a bounding box.
[0246] The attribute reference point cloud frame information 3011 can indicate whether to select attribute reference point cloud frame 2071 (e.g., global selection) or whether to select attribute reference point cloud frame 2071 for each spatial region (e.g., local selection). The attribute reference point cloud frame information 3011 can indicate the definition of the spatial region. The definition of the spatial region may include boundary information, such as a bounding box. The attribute reference point cloud frame information 3011 can indicate the reference point cloud frame 2005 for each spatial region. The definition of the spatial region may include boundary information, such as a bounding box. The attribute reference point cloud frame selected for a spatial region (e.g., as described herein with respect to step 3010) may not be equal to the reference point cloud frame 2005 used to decode the geometry of points belonging to that spatial region. The attribute reference point cloud frame selected for a spatial region (e.g., as described herein with respect to step 3010) may be equal to the reference point cloud frame 2005 used to decode the geometry of points belonging to that spatial region. The attribute reference point cloud frame information 3011 can indicate whether the attribute reference point cloud frame 2071 selected for the spatial region is equal to the reference point cloud frame 2005 used to decode the geometry of points belonging to that spatial region.
[0247] The attribute reference point cloud frame 2071 selected for a spatial region can be equal to the geometrically motion-compensated point cloud frame 2021 used to decode the geometry of points belonging to that spatial region. The attribute reference point cloud frame 2071 selected for a spatial region may not be equal to the geometrically motion-compensated point cloud frame 2021 used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frame information 3011 can indicate whether the attribute reference point cloud frame 2071 selected for a spatial region can be equal to the geometrically motion-compensated point cloud frame 2021 used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frame information 3011 can indicate an index in an index table referencing the attribute reference point cloud frame 2071 of the spatial region. Each index in the index table can reference a specific reference point cloud frame among multiple decoded reference point cloud frames.
[0248] Attribute MV information 1982 (as in this article about Figure 19 (as described) and attribute MV information 2062 (as described in this article) Figure 20 The description can indicate that attribute motion vectors (e.g., attribute motion vector 1983, attribute motion vector 2061) and geometric motion vectors (e.g., geometric motion vector 1921, geometric motion vector 2011) are associated with the same motion field structure.
[0249] The motion field structure can be constructed by partitioning the 3D space surrounding the decoded geometry (e.g., decoded geometry 1941, decoded geometry 2031) into motion units (MUs). Geometric motion vectors can be associated with each of the motion units (e.g., corresponding to each of the motion units). Motion compensation for reference point cloud frame 1905 (as described herein with respect to step 1920) and motion compensation for reference point cloud frame 2005 (as described herein with respect to step 2020) can be performed, for example, within each motion unit. Motion compensation for reference point cloud frame 1905 and motion compensation for reference point cloud frame 2005 can be performed, for example, based on the associated geometric motion vectors within each motion unit. Attribute motion vectors can be associated with each of the motion units (e.g., corresponding to each of the motion units). Motion compensation for attribute reference point cloud frame 1981 (as described herein with respect to step 1990) and motion compensation for attribute reference point cloud frame 2071 (as described herein with respect to step 2070) can be performed, for example, within each motion unit. Motion compensation for attribute reference point cloud frame 1981 and attribute reference point cloud frame 2071 can be performed, for example, based on the associated attribute motion vector within each motion unit.
[0250] The motion unit (MU) can be a non-intersecting cuboid. Attribute MV information (e.g., attribute MV information 1982, attribute MV information 2062) can define and / or indicate the motion field structure. The motion field structure can be associated with attribute MV (e.g., attribute MV 1983, attribute MV 2061) and geometric MV (e.g., geometric MV 1921, geometric MV 2011). Attribute MV information (e.g., attribute MV information 1982, attribute MV information 2062) can define and / or indicate the motion field structure, for example, via partitions of 3D space associated with attribute MV (e.g., attribute MV 1983, attribute MV 2061) and geometric MV (e.g., geometric MV 1921, geometric MV 2011). Attribute MV information (e.g., attribute MV information 1982, attribute MV information 2062) can indicate that attribute motion vectors (e.g., attribute motion vector 1983, attribute motion vector 2061) and geometric motion vectors (e.g., geometric motion vector 1921, geometric motion vector 2011) are associated with different motion field structures.
[0251] The attribute motion field structure associated with attribute MV (e.g., attribute MV 1983, attribute MV 2061) can be determined. The attribute motion field structure associated with attribute MV (e.g., attribute MV 1983, attribute MV 2061) can be determined, for example, from a first partition that divides the 3D space surrounding the decoded geometry (e.g., 1941, 2031) into attribute motion units. The geometric motion field structure associated with geometric motion vectors (e.g., geometric motion vector 1921, geometric motion vector 2011) can be determined, for example, from a second partition that divides the 3D space surrounding the decoded geometry (e.g., decoded geometry 1941, decoded geometry 2031) into geometric motion units. The motion field structure can be defined by the shape, size, and / or location of motion units in a partition of 3D space that surrounds the decoded geometry (e.g., decoded geometry 1941, decoded geometry 2031).
[0252] Attribute MV information (e.g., attribute MV information 1982, attribute MV information 2062) may indicate boundaries (e.g., thresholds, range limits, etc.) on the maximum value of geometric motion vectors (e.g., geometric motion vector 1921, geometric motion vector 2011). Alternatively, attribute MV information (e.g., attribute MV information 1982, attribute MV information 2062) may indicate boundaries (e.g., thresholds, range limits, etc.) on attribute motion vectors (e.g., attribute motion vector 1983, attribute motion vector 2061).
[0253] Figure 31 An example method for encoding attribute motion vectors into a bitstream is shown. Figure 31 One or more steps of the example method (e.g., method 3100) can be accomplished by an encoder (e.g., Figure 1 encoder 114) Figure 35 Example computer system 3500 and / or Figure 36 Example computing device 3630 is used to perform and / or implement. Figure 31 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0254] At step 3110, the encoder can determine attribute MV (e.g., attribute MV 1983). The encoder can determine attribute MV 1983, for example, by performing an attribute motion search. The attribute motion search can be performed such that attribute MV 1983 best approximates the 3D attribute motion field from attribute reference point cloud frame 1981 to the current point cloud frame 1910. At step 3120, the encoder can determine the attribute motion field structure associated with attribute MV 1983. The encoder can determine the attribute motion field structure associated with attribute MV 1983, for example, from the geometric motion field structure associated with geometry MV 1921.
[0255] At step 3130, the encoder can determine a motion vector residual (e.g., motion vector residual 3131). The encoder can determine the motion vector residual 3131, for example, based on the difference between the geometric MV 1921 of the geometric motion field structure and the attribute MV 1983 of the attribute motion field structure. The geometric MV 1921 can be subtracted from the attribute MV 1983. At step 3140, the encoder can determine a quantized motion vector residual 3141. The encoder can determine the quantized motion vector residual 3141, for example, by quantizing the motion vector residual 3131. At step 3150, the encoder can encode the (quantized) motion vector residual 3141 (e.g., entropy coding). The encoder can determine the attribute MV information 1982 as representing the encoded motion residual vector. The number of bits required to encode the attribute MV information 1982 in the bit stream 1970 can be reduced, for example, by encoding the (quantized) motion vector residual 3142 and determining the attribute MV information 1982 as a representation of the encoded motion residual vector.
[0256] The geometric motion field structure can be determined (e.g., defined) by partitioning the 3D space surrounding the decoded geometry 1941 into geometric MUs. A geometric MV 1921 can be associated with (e.g., corresponding to) each of the geometric MUs. An attribute motion field structure can be determined (e.g., defined) for example by partitioning the geometric motion field structure into attribute MUs. Attribute MUs can be determined, for example, by partitioning the geometric MUs. The encoder can determine the motion vector residual, for example, by subtracting the geometric MV 1921 associated with the geometric MU from the attribute MV 1983 and attribute MV 2061 associated with the attribute MU. Attribute MV 1983 and attribute MV 2061 can be associated with the attribute MUs partitioned from the geometric MUs.
[0257] Figure 32 An example method for decoding attribute MV information from a bitstream is shown. Figure 32 One or more steps of the example method (e.g., method 3200) can be decoded by a decoder (e.g., Figure 1 decoder 120) Figure 35 Example computer system 3500 and / or Figure 36 Example computing device 3630 is used to perform and / or implement. Figure 32 The steps (e.g., boxes) of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0258] At step 3210, the decoder can determine the (quantized) motion vector residual 3211. The decoder can determine the (quantized) motion vector residual 3211, for example, by decoding the attribute MV information 2062 (e.g., entropy decoding). At step 3220, the decoder can determine the motion vector residual 3221. The decoder can determine the motion vector residual 3221, for example, by inverse quantization of the quantized motion vector residual 3211. At step 3230, the decoder can determine the attribute MV 2061 of the attribute motion field structure. The decoder can determine the attribute MV 2061 of the attribute motion field structure, for example, by adding the motion vector residual 3211 and the geometric MV 2011 of the geometric motion field structure.
[0259] The geometric motion field structure can be determined (e.g., defined) for example by partitioning the 3D space surrounding the decoded geometry 2031 into geometric MUs. The attribute motion field structure can be determined (e.g., defined) for example by partitioning the geometric motion field structure into attribute MUs. The attribute MUs can be determined, for example, by partitioning the geometric MUs. The decoder can determine the attribute MV 2061 associated with the attribute MUs (e.g., one attribute MV per attribute MU) for example by adding the motion vector residual 3221 to the geometric MV 2011 associated with the geometric MU containing the attribute MUs.
[0260] References to encoding information (e.g., geometric information, attribute information, geometric MV information, attribute MV information) in this specification may indicate encoding the information as at least one single bit (e.g., a flag), or as at least one word each comprising more than one bit, or as a combination of at least one flag and at least one word. Encoding information into a bitstream may indicate writing at least one single bit (e.g., a flag) representing the information, or at least one word each comprising more than one bit, or a combination of at least one flag and at least one word, into the bitstream according to a specific syntax.
[0261] References to decoding information (e.g., geometric information, attribute information, geometric MV information, attribute MV information) in this specification may indicate decoding the information from at least one single bit (e.g., a flag), or from at least one word each comprising more than one bit, or from a combination of at least one flag and at least one word. Decoding information from a bitstream may indicate parsing the bitstream according to a specific syntax and reading from the bitstream at least one single bit (e.g., a flag), or at least one word each comprising more than one bit, or a combination of at least one flag and at least one word.
[0262] Figure 33 An example method for encoding point cloud frames is shown. More specifically, Figure 33 A flowchart 3300 illustrates an example method for encoding the geometry and attributes of a point cloud frame. The point cloud frame can be the current point cloud frame. The method in flowchart 3300 can be implemented using an encoder (e.g., Figure 1 encoder 114 in Figure 35 Example computer system 3500 and / or Figure 36 Example computing device 3630 is used to perform and / or implement. Figure 33 The steps of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0263] At step 3305, the encoder can determine a geometric motion vector (MV). The encoder can determine the geometric MV, for example, by performing a geometric motion search. The geometric MV can approximate a 3D motion field of geometry from a reference point cloud frame to the current point cloud frame. The reference point cloud frame can be determined (e.g., selected) from a first plurality of coded point cloud frames. The encoder can encode the geometric MV information into a bitstream. The encoder can encode the geometric MV information into a bitstream, for example, as a representation of one or more geometric MVs.
[0264] At step 3310, the encoder may perform motion compensation on the reference point cloud frame. The encoder may perform motion compensation on the reference point cloud frame discussed herein with respect to step 3305. The encoder may perform motion compensation on the reference point cloud frame of the current point cloud frame, for example, based on geometric motion vectors. The encoder may perform motion compensation on the reference point cloud frame of the current point cloud frame based on geometric motion vectors, for example, to determine a geometrically motion-compensated point cloud frame.
[0265] At step 3315, the encoder can determine the attribute motion vector (MV). The encoder can determine the attribute MV, for example, by performing an attribute motion search. The encoder can perform an attribute motion search such that the attribute MV best approximates the 3D attribute motion field from a reference point cloud frame (e.g., an attribute reference point cloud frame) to the current point cloud frame. The reference point cloud frame (e.g., an attribute reference point cloud frame) can be determined (e.g., selected) from a first plurality of coded point cloud frames. The encoder can encode the attribute MV information into a bitstream. The encoder can encode the attribute MV information into a bitstream, for example, as a representation of one or more attribute MVs.
[0266] At step 3320, the encoder may perform motion compensation on the attribute reference point cloud frame. The encoder may perform motion compensation on the attribute reference point cloud frame discussed herein with respect to step 3315. The encoder may perform motion compensation on the attribute reference point cloud frame of the point cloud frame, for example, based on attribute motion vectors. The encoder may perform motion compensation on the attribute reference point cloud frame of the point cloud frame based on attribute motion vectors, for example, to determine an attribute-motion-compensated point cloud frame. The attribute reference point cloud frame may be determined (e.g., selected) from a second plurality of coded point cloud frames. The first plurality of coded point cloud frames and the second plurality of coded point cloud frames may be the same.
[0267] At step 3330, the encoder may encode the geometry associated with the point cloud frame. The encoder may encode the geometry associated with the point cloud frame, for example, based on the geometry MV. The encoder may encode the geometry associated with the point cloud frame, for example, based on a geometry motion-compensated point cloud frame. At step 3340, the encoder may encode the attributes associated with the reconstructed geometry of the point cloud frame. The encoder may encode the attributes associated with the reconstructed geometry of the point cloud frame, for example, based on the attribute MV. The encoder may encode the attributes associated with the reconstructed geometry of the point cloud frame, for example, based on the attributes of a point cloud frame with attribute motion compensation.
[0268] Figure 34 An example method for decoding point cloud frames is shown. More specifically, Figure 34 A flowchart 3400 illustrates an example method for decoding the geometry and attributes of a point cloud frame. The point cloud frame can be the current point cloud frame. The method in flowchart 3400 can be implemented using a decoder (e.g., Figure 1 The decoder 120 in the middle is used to perform and / or implement. Figure 24 The steps of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0269] At step 3405, the decoder can determine the geometric motion vector (MV). The decoder can determine the geometric MV, for example, by decoding the geometric MV information from the bitstream. At step 3410, the decoder can perform motion compensation on a reference point cloud frame. The decoder can perform motion compensation on a reference point cloud frame of the current point cloud frame, for example, based on the geometric motion vector. The decoder can perform motion compensation on a reference point cloud frame of the current point cloud frame based on the geometric motion vector, for example, to determine a geometrically motion-compensated point cloud frame. The reference point cloud frame can be determined, for example, from a first plurality of coded point cloud frames (e.g., selected).
[0270] At step 3415, the decoder can determine the attribute motion vector (MV). The decoder can determine the attribute MV, for example, through attribute geometry MV information from the bitstream. At step 3420, the decoder can perform motion compensation on the attribute reference point cloud frame. The decoder can perform motion compensation on the attribute reference point cloud frame of the point cloud frame, for example, based on the attribute MV. The decoder can perform motion compensation on the attribute reference point cloud frame of the point cloud frame based on the attribute MV, for example, to determine the attribute motion-compensated point cloud frame. The attribute reference point cloud frame can be determined (e.g., selected), for example, from a second plurality of coded point cloud frames. The first plurality of coded point cloud frames and the second plurality of coded point cloud frames can be the same.
[0271] At step 3430, the decoder may decode the geometry associated with the point cloud frame. The decoder may decode the geometry associated with the point cloud frame, for example, based on the geometry MV. The decoder may decode the geometry associated with the point cloud frame based on the geometry MV, for example, to determine the reconstructed geometry of the point cloud frame. The decoder may decode the geometry associated with the point cloud frame, for example, based on a geometry motion-compensated point cloud frame. The decoder may decode the geometry associated with the point cloud frame based on a geometry motion-compensated point cloud frame, for example, to determine the reconstructed geometry of the point cloud frame. At step 3440, the decoder may decode the attributes associated with the reconstructed geometry. The decoder may decode the attributes associated with the reconstructed geometry, for example, based on the attribute MV. The decoder may decode the attributes associated with the reconstructed geometry, for example, based on the attributes of a point cloud frame with attribute motion compensation.
[0272] Figure 35 An example computer system in which this disclosure can be implemented is shown. Figure 35 The example computer system 3500 shown can implement one or more methods described herein. Various apparatuses and / or systems described herein (e.g., in...) Figure 1 , 2(3) can be implemented in the form of one or more computer systems 3500. Furthermore, each step of the flowchart depicted in this disclosure can be implemented on one or more computer systems 3500.
[0273] Computer system 3500 may include one or more processors, such as processor 3504. Processor 3504 may be a dedicated processor, a general-purpose processor, a microprocessor, and / or a digital signal processor. Processor 3504 may be connected to communication infrastructure 3502 (e.g., a bus or network). Computer system 3500 may also include main memory 3506 (e.g., random access memory (RAM)) and / or secondary memory 3508.
[0274] Secondary storage 3508 may include hard disk drive 3510 and / or removable storage drive 3512 (e.g., magnetic tape drive, optical disc drive, etc.). Removable storage drive 3512 may read from and / or write to removable storage unit 3516. Removable storage unit 3516 may include magnetic tape, optical disc, etc. Removable storage unit 3516 may be read from and / or write-to by removable storage drive 3512. Removable storage unit 3516 may include computer-usable storage media having computer software and / or data stored therein.
[0275] Secondary memory 3508 may include other similar components for allowing computer programs or other instructions to be loaded into computer system 3500. Such components may include removable storage unit 3518 and / or interface 3514. Examples of such components may include program boxes and / or box interfaces (such as in video game devices) that allow software and / or data to be transferred from removable storage unit 3518 to computer system 3500, removable memory chips (such as erasable programmable read-only memory (EPROM) or programmable read-only memory (PROM)) and associated sockets, flash drives and USB ports, and / or other removable storage units 3518 and interfaces 3514.
[0276] Computer system 3500 may also include communication interface 3520. Communication interface 3520 allows software and data to be transferred between computer system 3500 and external devices. Examples of communication interface 3520 may include a modem, network interface (e.g., Ethernet card), communication port, etc. Software and / or data transmitted via communication interface 3520 may be in the form of signals, which may be electronic, electromagnetic, optical, and / or other signals that can be received by communication interface 3520. Signals may be provided to communication interface 3520 via communication path 3522. Communication path 3522 may carry signals and may be implemented using wires or cables, optical fibers, telephone lines, cellular telephone links, RF links, and / or any other communication channels.
[0277] Computer program media and / or computer-readable media can be used to refer to tangible storage media, such as removable storage units 3516 and 3518 or a hard disk installed in hard disk drive 3510. A computer program product can be a component for providing software to computer system 3500. A computer program (which may also be referred to as computer control logic) can be stored in main memory 3506 and / or secondary memory 3508. A computer program can be received via communication interface 3520. When executed, such a computer program can enable computer system 3500 to implement the present disclosure as discussed herein. Specifically, when executed, the computer program can enable processor 3504 to implement the processes of the present disclosure, such as any of the methods described herein. Therefore, such a computer program can represent a controller of computer system 3500.
[0278] The features of this disclosure can be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. Implementing a hardware state machine to perform the functions described herein will also be apparent to those skilled in the art.
[0279] Figure 36Example elements of a computing device are shown that can be used to implement any of the various devices described herein, including, for example, a source device (e.g., 102), an encoder (e.g., 114), a destination device (e.g., 106), a decoder (e.g., 120), and / or any computing device described herein. The computing device 3630 may include one or more processors 3631 that can execute instructions stored in random access memory (RAM) 3633, removable media 3634 (e.g., a Universal Serial Bus (USB) drive, an optical disc (CD) or digital versatile optical disc (DVD), or a floppy disk drive), or any other desired storage medium. Instructions may also be stored in an attached (or internal) hard disk drive 3635. The computing device 3630 may also include a security processor (not shown) that can execute instructions of one or more computer programs to monitor processes executing on the processor 3631 and any processes requesting access to any hardware and / or software components of the computing device 3630 (e.g., ROM 3632, RAM 3633, removable media 3634, hard disk drive 3635, device controller 3637, network interface 3639, GPS 3641, Bluetooth interface 3642, WiFi interface 3643, etc.). The computing device 3630 may include one or more output devices, such as a display 3636 (e.g., screen, display device, monitor, television, etc.), and may include one or more output device controllers 3637, such as a video processor. One or more user input devices 3638 may also be present, such as a remote control, keyboard, mouse, touchscreen, microphone, etc. The computing device 3630 may also include one or more network interfaces, such as a network interface 3639, which may be a wired interface, a wireless interface, or a combination of both. Network interface 3639 can provide the computing device 3630 with an interface to communicate with network 3640 (e.g., RAN or any other network). Network interface 3639 may include a modem (e.g., a cable modem), and external network 3640 may include a communication link, external network, home network, provider's wireless, coaxial cable, fiber optic, or hybrid fiber / coaxial cable distribution system (e.g., DOCSIS network), or any other desired network. Additionally, computing device 3630 may include a location detection device, such as a Global Positioning System (GPS) microprocessor 3641, which can be configured to receive and process GPS signals and determine the geographic location of computing device 3630 with possible assistance from external servers and antennas.
[0280] Figure 36The examples shown can be hardware configurations, but the components illustrated can also be implemented as software. Modifications can be made to add, remove, combine, partition, etc., components of computing device 3630 as needed. Alternatively, basic computing devices and components can be used to implement components, and the same components (e.g., processor 3631, ROM storage device 3632, display 3636, etc.) can be used to implement any other computing devices and components described herein. For example, the various components described herein can be implemented using a computing device having components such as a processor that executes computer-executable instructions stored on a computer-readable medium, such as... Figure 36 As shown in the figure. Some or all of the entities described herein may be software-based and may coexist on a common physical platform (e.g., the requesting entity may be a separate software process and program from the relevant entity, both of which may be executed as software on a common computing device).
[0281] A computing device can perform a method including multiple operations. The computing device may include a decoder. The computing device can determine a geometric motion vector from a bit stream. The computing device can determine an attribute motion vector from the bit stream. The computing device can perform motion compensation on a reference point cloud frame of a current point cloud frame based on the geometric motion vector to determine a geometrically motion-compensated point cloud frame. The computing device can perform motion compensation on an attribute reference point cloud frame of a point cloud frame based on the attribute motion vector to determine an attribute-motion-compensated point cloud frame. The computing device can determine the reconstructed geometry of the point cloud frame. The point cloud frame may be associated with content. The computing device can determine the reconstructed geometry of the point cloud frame by decoding the geometry associated with the point cloud frame. Motion compensation on a reference point cloud frame of the point cloud frame can be performed based on the geometric motion vector to determine the reconstructed geometry of the point cloud frame. The geometry associated with the point cloud frame can be decoded based on the geometric motion vector. The computing device can decode geometric motion vector information indicating geometric motion vectors, wherein the geometric motion vectors are associated with the same motion field structure. The computing device can decode attribute motion vector information indicating attribute motion vectors. The computing device can decode attribute motion vector information by determining motion vector residuals based on the decoding of attribute motion vector information. The computing device can also decode attribute motion vector information by adding the motion vector residuals to the geometric motion vectors of the geometric motion field structure to determine the attribute motion vectors of the attribute motion field structure. The computing device can decode attributes associated with the reconstructed geometry. The computing device can decode attributes associated with the reconstructed geometry by decoding residual attributes that indicate the difference between the attributes of the reconstructed geometry and the attribute predictor associated with the reconstructed geometry. The computing device can decode attributes associated with the reconstructed geometry by adding the attribute predictor to the residual attributes to determine the attributes of the reconstructed geometry. Attributes associated with the reconstructed geometry can be decoded based on attribute motion vectors. The computing device can determine projection attributes based on attribute motion vectors. The computing device can determine projection attributes based on motion compensation of a coded reference point cloud frame, wherein the coded reference point cloud frame is motion compensated by attribute motion vectors. The computing device can determine projection attributes by determining a point cloud frame with attribute motion compensation. The computing device can determine projection attributes by projecting the attributes of a motion-compensated point cloud frame onto the reconstructed geometry. The computing device can determine an attribute predictor based on the projection attributes to determine the attributes associated with the reconstructed geometry. The computing device can decode the attributes associated with the reconstructed geometry based on the attribute predictor. The computing device can decode attribute reference point cloud frame information indicating an attribute reference point cloud frame. The computing device can select an attribute reference point cloud frame from multiple decoded reference point cloud frames based on the attribute reference point cloud frame information.A computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to encode point cloud frames. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0282] A computing device can perform a method including multiple operations. The computing device may include an encoder. The computing device can determine geometric motion vectors based on a point cloud frame associated with content. The computing device can determine attribute motion vectors based on the point cloud frame. The computing device can perform motion compensation on a reference point cloud frame of the current point cloud frame based on the geometric motion vectors to determine a geometrically motion-compensated point cloud frame. The computing device can perform motion compensation on an attribute reference point cloud frame of the point cloud frame based on the attribute motion vectors to determine an attribute-motion-compensated point cloud frame. The computing device can encode the geometry associated with the point cloud frame based on the geometric motion vectors. The computing device can encode attributes associated with the reconstructed geometry of the point cloud frame based on the attribute motion vectors. The computing device can encode attributes associated with the reconstructed geometry by: determining projection attributes based on the attribute motion vectors; determining an attribute predictor for attributes associated with the reconstructed geometry based on the projection attributes; and encoding attributes associated with the reconstructed geometry based on the attribute predictor. The computing device can determine projection attributes based on motion compensation of a coded reference point cloud frame, wherein the coded reference point cloud frame is motion-compensated by attribute motion vectors. The computing device can determine the attribute motion vectors based on the difference between: attributes associated with the reconstructed geometry of the point cloud frame; and attributes of the geometry of the coded reference point cloud frame, which are adjusted by the attribute motion vectors. The computing device can determine attributes associated with the reconstructed geometry of the point cloud frame based on mapping the attributes of the geometry associated with the point cloud frame to the reconstructed geometry. Attributes associated with the reconstructed geometry of the point cloud frame include color. The computing device can determine the mapping attributes of the geometry associated with the point cloud frame based on recoloring. The computing device can perform motion compensation on a reference point cloud frame of the point cloud frame based on geometric motion vectors. The computing device can perform motion compensation on a reference point cloud frame based on geometric motion vectors to determine the reconstructed geometry of the point cloud frame. The computing device can encode geometric motion vector information indicating geometric motion vectors, wherein the geometric motion vectors are associated with the same motion field structure. The computing device can encode attribute motion vector information indicating attribute motion vectors. A computing device can encode attribute motion vector information by: determining a motion vector residual based on the difference between the attribute motion vector and the geometric motion vector; and encoding the motion vector residual into attribute motion vector information. The computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to decode point cloud frames. A computer-readable medium may store instructions that, when executed, enable the performance of the described methods, additional operations, and / or include additional elements.
[0283] A computing device can perform a method comprising multiple operations. The computing device may include a decoder. The computing device can determine attribute motion vectors and geometric motion vectors from a bitstream. The computing device can determine a reconstructed geometry of a point cloud frame associated with content by decoding the geometry associated with the geometric motion vectors. The computing device can determine projection attributes based on the attribute motion vectors. The computing device can determine an attribute predictor for attributes associated with the reconstructed geometry based on the projection attributes. The computing device can decode residual attributes from the bitstream indicating the difference between the attributes associated with the reconstructed geometry and the attribute predictor. The computing device can decode the residual attributes by: decoding transform coefficients indicating the residual attributes from the bitstream; and determining the residual attributes based on applying an inverse intra-frame transform to the decoded transform coefficients. The residual attributes can be decoded based on a prediction-enhanced transform scheme. The computing device can determine the attribute predictor by: smoothing the projection attributes; and determining the attribute predictor based on the smoothed projection attributes. The computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to encode point cloud frames. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0284] A computing device can perform a method including multiple operations. The computing device may include an encoder. The computing device can perform motion compensation on a reference point cloud frame of a current point cloud frame based on geometric motion vectors to determine a geometrically motion-compensated point cloud frame. The computing device can perform motion compensation on an attribute reference point cloud frame of the point cloud frame based on attribute motion vectors to determine an attribute-motion-compensated point cloud frame. The computing device can encode the geometry associated with the point cloud frame based on the geometrically motion-compensated point cloud frame. The computing device can encode attributes associated with the reconstructed geometry of the point cloud frame based on the attributes of the attribute-motion-compensated point cloud frame. The geometrically motion-compensated point cloud frame can be determined from a coded reference point cloud frame. The geometrically motion-compensated point cloud frame can be determined based on motion compensation of a coded reference point cloud frame. The coded reference point cloud frame can be motion-compensated by geometric motion vectors. The attribute-motion-compensated point cloud frame can be determined from a coded reference point cloud frame. The attribute-motion-compensated point cloud frame can be determined based on motion compensation of a coded reference point cloud frame. The coded reference point cloud frame can be motion-compensated by attribute motion vectors. The computing device can encode the attribute motion vectors and / or the indication (e.g., index or ID) of the coded reference point cloud frame. The computing device can determine the attribute motion vectors based on the difference between the attributes associated with the reconstructed geometry and the attributes of the coded reference point cloud frame adjusted by the attribute motion vectors. The computing device can encode the geometric motion vectors and / or the indication (e.g., index or ID) of the coded reference point cloud frame. The computing device can determine the geometric motion vectors based on the difference between the reconstructed geometry and the geometry adjusted by the geometric motion vectors of the coded reference point cloud frame. The computing device can encode attributes by: determining residual attributes based on the difference between the attributes of the reconstructed geometry and the attribute predictor; and encoding the residual attributes. The computing device can determine the attribute predictor by smoothing the projected attributes, the attribute predictor being determined from the smoothed projected attributes. The attribute predictor can include a corresponding attribute predictor for each corresponding vertex in the vertices of the reconstructed geometry, based on the projected attributes corresponding to said vertex in the projected attributes. The computing device can encode residual attributes by: determining transformed coefficients based on applying an intra-frame transform to the residual attributes; and entropy encoding the transformed coefficients corresponding to (e.g., representing or indicating) the residual attributes in the bitstream. The computing device can quantize the transformed coefficients, which are then entropy encoded. Residual attributes can be encoded based on a prediction-enhanced (prediction-boosted) transform scheme. Intra-frame transforms can include adaptive DCT, RAHT transform, or Haar transform. Inverse intra-frame transforms include inverse adaptive DCT (A-DCT), inverse RAHT transform of the RAHT scheme, or inverse Haar transform. The computing device can select an attribute reference point cloud frame from multiple pre-coded reference point cloud frames.The computing device can encode attribute reference point cloud frame information representing a selected attribute reference point cloud frame. The attribute reference point cloud frame is selected independently of the reference point cloud frame. The attribute reference point cloud frame information can indicate the selection of a single attribute reference point cloud frame. A single attribute reference point cloud frame and the reference point cloud frame can be the same or different. The attribute reference point cloud frame information can indicate whether a single attribute reference point cloud frame and the reference point cloud frame are the same. A single attribute reference point cloud frame and a geometrically motion-compensated point cloud frame can be the same or different. The attribute reference point cloud frame information can indicate whether a single attribute reference point cloud frame and a geometrically motion-compensated point cloud frame are the same. The attribute reference point cloud frame information can indicate an index pointing to an index table referencing a single attribute reference point cloud frame, each index referencing a specific reference point cloud frame among multiple decoded reference point cloud frames. The computing device can determine a spatial region from the decoded geometry; and select an attribute reference point cloud frame for each spatial region. The attribute reference point cloud frame information can indicate whether to select a single attribute reference point cloud frame or whether to select an attribute reference point cloud frame for each spatial region. Attribute reference point cloud frame information can indicate the definition of a spatial region. Attribute reference point cloud frame information can indicate the reference point cloud frame for each spatial region. The attribute reference point cloud frame selected for a spatial region may or may not be equal to the reference point cloud frame used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frame information can indicate whether the attribute reference point cloud frame selected for a spatial region is equal to the reference point cloud frame used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frames can be selected for a spatial region, which may or may not be equal to the geometrically motion-compensated point cloud frame used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frame information can indicate whether the attribute reference point cloud frame selected for a spatial region is equal to the geometrically motion-compensated point cloud frame used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frame information can indicate an index pointing to an index table of attribute reference point cloud frames referencing a spatial region, each index referencing a specific reference point cloud frame from multiple decoded reference point cloud frames. The computing device can encode geometric motion vector information representing geometric motion vectors and attribute motion vector information representing attribute motion vectors. The attribute motion vector information can indicate that the attribute motion vectors and geometric motion vectors are associated with the same motion field structure. The motion field structure can include partitions that divide the 3D space surrounding the reconstructed geometry into motion units, wherein geometric motion vectors are associated with each of the motion units, and wherein motion compensation of a reference point cloud frame is performed within each motion unit based on the associated geometric motion vectors, wherein attribute motion vectors are associated with each of the motion units, and wherein motion compensation of an attribute reference point cloud frame is performed within each motion unit based on the associated attribute motion vectors. A motion unit can include a non-intersecting cuboid. The attribute motion vector information can indicate the motion field structure.Attribute motion information can indicate that attribute motion vectors and geometric motion vectors are associated with different motion field structures. The attribute motion field structure associated with the attribute motion vector is determined from a first partition that divides the 3D space surrounding the decoded geometry associated with the point cloud frame into attribute motion units, and the geometric motion field structure associated with the geometric motion vector is determined from a second partition that divides the 3D space surrounding the decoded geometry associated with the point cloud frame into geometric motion units. A computing device can encode the attribute motion vector information by: determining the attribute motion field structure associated with the attribute motion vector from the geometric motion field structure associated with the geometric motion vector; determining a motion vector residual based on the difference between: the attribute motion vector of the attribute motion field structure; and the geometric motion vector of the geometric motion field structure; and encoding the motion vector residual as attribute motion information. The attribute motion information can indicate boundaries (e.g., thresholds, ranges, limits, etc.) on the maximum magnitude of the geometric motion vector and / or boundaries on the attribute motion vector. The computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to decode a current point cloud frame. A computer-readable medium may store instructions that, when executed, cause the described methods, additional operations, and / or include additional elements.
[0285] A computing device can perform a method including multiple operations. The computing device may include a decoder. The computing device can perform motion compensation on a reference point cloud frame of a current point cloud frame based on geometric motion vectors to determine a geometrically motion-compensated point cloud frame. The computing device can perform motion compensation on an attribute reference point cloud frame of the point cloud frame based on attribute motion vectors to determine an attribute-motion-compensated point cloud frame. The computing device can decode the geometry associated with the point cloud frame based on the geometrically motion-compensated point cloud frame to determine a reconstructed geometry of the point cloud frame. The computing device can decode the attributes associated with the reconstructed geometry based on the attributes of the attribute-motion-compensated point cloud frame. The computing device can decode the attributes associated with the reconstructed geometry by: determining an attribute predictor based on projecting the attributes of the attribute-motion-compensated point cloud frame onto the reconstructed geometry; and decoding the attributes of the reconstructed geometry based on the attribute predictor. The computing device can decode attributes associated with the reconstructed geometry by: decoding residual attributes from the bitstream that indicate the difference between the reconstructed geometry attribute and the attribute predictor; and determining the decoded attribute based on adding the attribute predictor to the decoded residual attribute. The computing device can also decode residual attributes by: decoding transform coefficients corresponding to (e.g., representing or indicating) the residual attribute from the bitstream; and determining the residual attribute based on applying an inverse intra-frame transform to the decoded transform coefficients. The computing device can dequantize the transform coefficients, and the residual attribute is determined by applying an inverse intra-frame transform to the dequantized transform coefficients. The residual attribute can be decoded based on a prediction-boosting (prediction-enhanced) transform scheme. Intra-frame transforms can include adaptive DCT, RAHT transform, or Haar transform. Inverse intra-frame transforms include inverse adaptive DCT (A-DCT), inverse RAHT transform of the RAHT scheme, or inverse Haar transform. The computing device can select an attribute reference point cloud frame from multiple coded reference point cloud frames. The computing device can decode attribute reference point cloud frame information representing a selected attribute reference point cloud frame; and select an attribute reference point cloud frame from a plurality of decoded reference point cloud frames based on the decoded attribute reference point cloud frame information. The attribute reference point cloud frame is selected independently of the reference point cloud frame. The attribute reference point cloud frame information can indicate the selection of a single attribute reference point cloud frame. A single attribute reference point cloud frame and a reference point cloud frame can be the same or different. The attribute reference point cloud frame information can indicate whether a single attribute reference point cloud frame and a reference point cloud frame are the same. A single attribute reference point cloud frame and a geometrically motion-compensated point cloud frame can be the same or different. The attribute reference point cloud frame information can indicate whether a single attribute reference point cloud frame and a geometrically motion-compensated point cloud frame are the same. The attribute reference point cloud frame information can indicate an index pointing to an index table referencing a single attribute reference point cloud frame, each index referencing a specific reference point cloud frame among a plurality of decoded reference point cloud frames.The computing device can determine spatial regions from the decoded geometry; and select attribute reference point cloud frames for each spatial region. Attribute reference point cloud frame information can indicate whether a single attribute reference point cloud frame is selected or whether attribute reference point cloud frames are selected for each spatial region. Attribute reference point cloud frame information can indicate the definition of the spatial region. Attribute reference point cloud frame information can indicate the reference point cloud frame for each spatial region. The attribute reference point cloud frame selected for a spatial region may or may not be equal to the reference point cloud frame used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frame information can indicate whether the attribute reference point cloud frame selected for a spatial region is equal to the reference point cloud frame used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frames can be selected for a spatial region, which may or may not be equal to the geometrically motion-compensated point cloud frame used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frame information can indicate whether the attribute reference point cloud frame selected for a spatial region is equal to the geometrically motion-compensated point cloud frame used to decode the geometry of points belonging to that spatial region. Attribute reference point cloud frame information can indicate an index of an index table pointing to attribute reference point cloud frames in a reference space region, each index referencing a specific reference point cloud frame from multiple decoded reference point cloud frames. The computing device can decode geometric motion vector information representing geometric motion vectors and attribute motion vector information representing attribute motion vectors. Attribute motion vector information can indicate that attribute motion vectors and geometric motion vectors are associated with the same motion field structure. The motion field structure can include partitions that divide the 3D space surrounding the reconstructed geometry into motion units, wherein geometric motion vectors are associated with each of the motion units, and wherein motion compensation of reference point cloud frames is performed within each motion unit based on the associated geometric motion vectors, wherein attribute motion vectors are associated with each of the motion units, and wherein motion compensation of attribute reference point cloud frames is performed within each motion unit based on the associated attribute motion vectors. Motion units can include non-intersecting cuboids. Attribute motion vector information can indicate motion field structures. Attribute motion information can indicate that attribute motion vectors and geometric motion vectors are associated with different motion field structures. The attribute motion field structure associated with the attribute motion vector is determined by dividing the 3D space surrounding the decoded geometry associated with the point cloud frame into attribute motion units, and the geometric motion field structure associated with the geometric motion vector is determined by dividing the 3D space surrounding the decoded geometry associated with the point cloud frame into geometric motion units.A computing device can decode attribute motion vector information by: determining an attribute motion field structure associated with an attribute motion vector from a geometric motion field structure associated with a geometric motion vector; determining a motion vector residual based on the difference between: the attribute motion vector of the attribute motion field structure; and the geometric motion vector of the geometric motion field structure; and encoding the motion vector residual into attribute motion information. The attribute motion information may indicate boundaries (e.g., thresholds, ranges, limits, etc.) on the maximum magnitude of the geometric motion vector and / or boundaries on the attribute motion vector. The computing device may include: one or more processors; and a memory storing instructions that, when executed by the one or more processors, perform the methods described herein. A system may include: a computing device configured to perform the described methods, additional operations, and / or include additional elements; and a second computing device configured to encode a current point cloud frame. A computer-readable medium may store instructions that, when executed, enable the performance of the described methods, additional operations, and / or include additional elements.
[0286] In the following text, various features will be highlighted in a set of numbered clauses or paragraphs. These features should not be construed as limitations on the invention or inventive concept, but are merely highlights of certain features described herein, without implying a particular order of importance or relevance of such features.
[0287] Clause 1. A method comprising determining a geometric motion vector from a bitstream by a decoder.
[0288] Clause 2. The method described in Clause 1 further includes determining a geometric motion vector from the bitstream by a decoder.
[0289] Clause 3. The method according to any one of Clauses 1 to 2 further includes determining the reconstructed geometry of the point cloud frame by decoding the geometry associated with the point cloud frame associated with the content based on the geometric motion vector.
[0290] Clause 4. The method according to any one of Clauses 1 to 3 further includes decoding the attributes associated with the reconstructed geometry based on the attribute motion vector.
[0291] Clause 5. The method according to Clause 4, wherein decoding the attribute associated with the reconstructed geometry comprises: determining a projection attribute based on the attribute motion vector; determining an attribute predictor for the attribute associated with the reconstructed geometry based on the projection attribute; and decoding the attribute associated with the reconstructed geometry based on the attribute predictor.
[0292] Clause 6. The method according to Clause 5, wherein determining the projection attribute further comprises: determining an attribute motion-compensated point cloud frame; and determining the projection attribute by projecting the attributes of the attribute motion-compensated point cloud frame onto the reconstructed geometry.
[0293] Clause 7. The method according to any one of Clauses 4 to 6, wherein decoding the attribute associated with the reconstructed geometry comprises: decoding a residual attribute indicating the difference between the attribute of the reconstructed geometry and the attribute predictor associated with the reconstructed geometry; and determining the attribute of the reconstructed geometry based on adding the attribute predictor to the residual attribute.
[0294] Clause 8. The method according to any one of Clauses 3 to 7, wherein motion compensation is performed on a reference point cloud frame of the point cloud frame based on the geometric motion vector to determine the reconstructed geometry of the point cloud frame.
[0295] Clause 9. The method according to Clause 8 further comprises: decoding geometric motion vector information indicating the geometric motion vector, wherein the geometric motion vector is associated with the same motion field structure; and decoding attribute motion vector information indicating the attribute motion vector.
[0296] Clause 10. The method according to Clause 9, wherein decoding the attribute motion vector information comprises: determining a motion vector residual based on decoding the attribute motion vector information; and determining the attribute motion vector of the attribute motion field structure based on adding the motion vector residual to the geometric motion vector of the geometric motion field structure.
[0297] Clause 11. The method according to any one of Clauses 3 to 11 further includes: decoding attribute reference point cloud frame information indicating the attribute reference point cloud frame; and selecting the attribute reference point cloud frame from a plurality of decoded reference point cloud frames based on the attribute reference point cloud frame information.
[0298] Clause 12. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by said one or more processors, cause the computing device to perform a method according to any one of Clauses 1 to 11.
[0299] Clause 13. A system comprising: a computing device configured to perform the method according to any one of Clauses 1 to 11; and a second computing device configured to encode the point cloud frame.
[0300] Clause 14. A computer-readable medium storing instructions that, when executed, cause the execution of the method according to any one of Clauses 1 to 11.
[0301] Clause 15. A method comprising determining geometric motion vectors by an encoder based on point cloud frames associated with content.
[0302] Clause 16. The method according to Clause 15 further includes: determining attribute motion vectors based on the point cloud frame.
[0303] Clause 17. The method according to any one of Clauses 15 to 16 further comprises: encoding the geometry associated with the point cloud frame based on the geometric motion vector.
[0304] Clause 18. The method according to any one of Clauses 15 to 17 further comprises: encoding attributes associated with the reconstructed geometry of the point cloud frame based on the attribute motion vector.
[0305] Clause 19. The method according to Clause 18, wherein encoding the attribute associated with the reconstructed geometry of the point cloud frame comprises: determining a projection attribute based on the attribute motion vector; determining an attribute predictor for the attribute associated with the reconstructed geometry based on the projection attribute; and encoding the attribute associated with the reconstructed geometry based on the attribute predictor.
[0306] Clause 20. The method according to any one of Clauses 18 to 19 further comprises: determining projection attributes based on motion compensation of a coded reference point cloud frame, wherein the coded reference point cloud frame is motion compensated by the attribute motion vector.
[0307] Clause 21. The method according to Clause 20 further comprises: determining the attribute motion vector based on the difference between: the attribute associated with the reconstructed geometry of the point cloud frame; and the attribute of the geometry of the coded reference point cloud frame, which is adjusted by the attribute motion vector.
[0308] Clause 22. The method according to any one of Clauses 18 to 21 further comprises: determining the attribute associated with the reconstructed geometry of the point cloud frame based on mapping the attribute of the geometry associated with the point cloud frame to the reconstructed geometry.
[0309] Clause 23. The method according to Clause 22, wherein the attribute associated with the reconstructed geometry of the point cloud frame includes color.
[0310] Clause 24. The method according to Clause 23 further includes: determining the mapping properties of the geometry associated with the point cloud frame based on recoloring.
[0311] Clause 25. The method according to any one of Clauses 18 to 24, wherein motion compensation is performed on a reference point cloud frame of the point cloud frame based on a geometric motion vector to determine the reconstructed geometry of the point cloud frame.
[0312] Clause 26. The method according to Clause 25 further comprises: encoding geometric motion vector information indicating the geometric motion vector, wherein the geometric motion vector is associated with the same motion field structure; and encoding attribute motion vector information indicating the attribute motion vector.
[0313] Clause 27. The method according to Clause 26, wherein encoding the attribute motion vector information comprises: determining a motion vector residual based on the difference between the attribute motion vector and the geometric motion vector; and encoding the motion vector residual as the attribute motion vector information.
[0314] Clause 28. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by said one or more processors, cause the computing device to perform a method according to any one of Clauses 15 to 27.
[0315] Clause 29. A system comprising: a computing device configured to perform the method according to any one of Clauses 15 to 27; and a second computing device configured to decode the point cloud frame.
[0316] Clause 30. A computer-readable medium storing instructions that, when executed, cause to perform the method according to any one of Clauses 15 to 27.
[0317] Clause 31. A method comprising determining attribute motion vectors and geometric motion vectors from a bitstream by a decoder.
[0318] Clause 32. The method according to Clause 31 further includes determining the reconstructed geometry of the point cloud frame by decoding the geometry associated with the point cloud frame associated with the content based on the geometric motion vector.
[0319] Clause 33. The method according to any one of Clauses 31 to 32 further includes: determining the projection attribute based on the attribute motion vector.
[0320] Clause 34. The method according to Clause 33 further includes: an attribute predictor that determines attributes associated with the reconstructed geometry based on the projection attributes.
[0321] Clause 35. The method according to Clause 34 further comprises: decoding from the bitstream residual attributes indicating the difference between the attribute associated with the reconstructed geometry and the attribute predictor.
[0322] Clause 36. The method according to Clause 35, wherein decoding the residual attribute comprises: decoding transform coefficients indicating the residual attribute from the bitstream; and determining the residual attribute based on applying an inverse intra-frame transform to the decoded transform coefficients.
[0323] Clause 37. The method according to any one of Clauses 35 to 36, wherein the residual property is decoded based on a transformation scheme for prediction enhancement.
[0324] Clause 38. The method according to any one of Clauses 35 to 37, wherein determining the attribute predictor further comprises: smoothing the projected attribute; and determining the attribute predictor based on the smoothed projected attribute.
[0325] Clause 39. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of Clauses 31 to 38.
[0326] Clause 40. A system comprising: a computing device configured to perform the method according to any one of Clauses 31 to 38; and a second computing device configured to encode the point cloud frame.
[0327] Clause 41. A computer-readable medium storing instructions that, when executed, cause to perform the method according to any one of Clauses 31 to 38.
[0328] Clause 42. A method comprising performing motion compensation on a reference point cloud frame of a current point cloud frame based on geometric motion vectors to determine a geometrically motion-compensated point cloud frame.
[0329] Clause 43. The method according to Clause 42 further comprises: performing motion compensation on an attribute reference point cloud frame of the point cloud frame based on an attribute motion vector to determine an attribute motion-compensated point cloud frame.
[0330] Clause 44. The method according to any one of Clauses 42 to 43 further includes encoding the geometry associated with the point cloud frame based on the geometrically motion-compensated point cloud frame.
[0331] Clause 45. The method according to any one of Clauses 43 to 44 further includes encoding attributes associated with the reconstructed geometry of the point cloud frame based on the attributes of the attribute-motion-compensated point cloud frame.
[0332] Clause 46. The method according to Clause 45, wherein encoding the attribute associated with the reconstructed geometry of the point cloud frame comprises: determining the attribute associated with the reconstructed geometry based on projecting the attribute of the attribute-motion-compensated point cloud frame onto the reconstructed geometry; and encoding the attribute of the reconstructed geometry based on the attribute predictor.
[0333] Clause 47. The method according to any one of Clauses 44 to 46 further comprises: determining the attribute associated with the reconstructed geometry of the point cloud frame based on mapping the attributes of the geometry of the point cloud frame to the reconstructed geometry.
[0334] Clause 48. The method according to Clause 47 further comprises: encoding the geometry of the point cloud frame; and determining the reconstructed geometry based on decoding the encoded geometry.
[0335] Clause 47. The method according to any one of Clauses 47 to 48, wherein the attribute is color and the mapping attribute is determined based on recoloring.
[0336] Clause 48. The method according to any one of Clauses 47 to 48, wherein the mapping attribute of each point of the reconstructed geometry is determined based on a nearest neighbor search from the geometry of the point cloud frame to the nearest point of the point in the reconstructed geometry.
[0337] Clause 49. The method according to any one of Clauses 46 to 48, wherein the encoding of the attribute comprises: determining a residual attribute based on the difference between the attribute of the reconstructed geometry and the attribute predictor; and encoding the residual attribute.
[0338] Clause 50. The method according to Clause 49, wherein the encoding of the residual attribute comprises: determining transformed coefficients based on applying an intra-frame transform to the residual attribute; and entropy encoding the transformed coefficients corresponding to (e.g., representing or indicating) the residual attribute in the bitstream.
[0339] Clause 51. The method according to Clause 50 further includes quantizing the transformed coefficients, wherein the quantized transformed coefficients are entropy encoded.
[0340] Clause 52. A computing device comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of Clauses 42 to 51.
[0341] Clause 53. A system comprising: a computing device configured to perform the method according to any one of Clauses 42 to 51; and a second computing device configured to decode the current point cloud frame.
[0342] Clause 54. A computer-readable medium storing instructions that, when executed, cause to perform the method according to any one of Clauses 42 to 51.
[0343] Clause 55. A method comprising performing motion compensation on a reference point cloud frame of a current point cloud frame based on geometric motion vectors to determine a geometrically motion-compensated point cloud frame.
[0344] Clause 56. The method according to Clause 55 further comprises: performing motion compensation on an attribute reference point cloud frame of the point cloud frame based on attribute motion vectors to determine an attribute motion-compensated point cloud frame.
[0345] Clause 57. The method according to any one of Clauses 55 to 56 further includes decoding the geometry associated with the point cloud frame based on the geometrically motion-compensated point cloud frame to determine the reconstructed geometry of the point cloud frame.
[0346] Clause 58. The method according to any one of Clauses 56 to 57 further includes decoding attributes associated with the reconstructed geometry based on the attributes of the attribute-motion-compensated point cloud frame.
[0347] Clause 59. The method according to Clause 58, wherein decoding the attribute associated with the reconstructed geometry comprises: determining the attribute associated with the reconstructed geometry based on projecting the attribute of the attribute-motion-compensated point cloud frame onto the reconstructed geometry; and decoding the attribute of the reconstructed geometry based on the attribute predictor.
[0348] Clause 60. The method according to Clause 59, wherein the decoding of the attribute associated with the reconstructed geometry comprises: decoding a residual attribute from a bitstream that indicates a difference between the attribute of the reconstructed geometry and the attribute predictor; and determining the decoded attribute based on adding the attribute predictor to the decoded residual attribute.
[0349] Clause 61. The method according to Clause 60, wherein the decoding of the residual attribute comprises: entropy decoding from the bitstream of transform coefficients corresponding to (e.g., representing or indicating) the residual attribute; and determining the residual attribute based on applying an inverse intra-frame transform to the decoded transform coefficients.
[0350] Clause 62. The method according to Clause 61 further comprises: dequantizing the transform coefficients, wherein the residual property is determined by applying the inverse intra-frame transform to the dequantized transformed coefficients.
[0351] Clause 63. The method according to any one of Clauses 60 to 62, wherein the residual property is decoded based on a prediction-enhanced (prediction-enhanced) transformation scheme.
[0352] Clause 64. The method according to any one of Clauses 60 to 63, wherein the intra-frame transform comprises adaptive DCT, RAHT transform, or Haar transform.
[0353] Clause 65. The method according to any one of Clauses 60 to 63, wherein the inverse intra-frame transform comprises inverse adaptive DCT (A-DCT), inverse RAHT transform of the RAHT scheme, or inverse Haar transform.
[0354] Clause 66. The method according to any one of Clauses 56 to 65, wherein the geometrically motion-compensated point cloud frame is determined from a coded reference point cloud frame.
[0355] Clause 67. The method according to Clause 66, wherein the geometrically motion-compensated point cloud frame is determined based on motion compensation of the coded reference point cloud frame.
[0356] Clause 68. The method according to any one of Clauses 66 to 67, wherein the geometrically motion-compensated point cloud frame is determined based on motion compensation of the coded reference point cloud frame.
[0357] Clause 69. The method according to any one of Clauses 66 to 68 further includes encoding the geometric motion vector and / or the indication (e.g., index or ID) of the coded reference point cloud frame.
[0358] Clause 70. The method according to any one of Clauses 66 to 69 further includes determining the geometric motion vector based on the difference between the reconstructed geometry and the geometry adjusted by the geometric motion vector of the coded reference point cloud frame.
[0359] Clause 71. The method according to any one of Clauses 56 to 70, wherein the attribute-motion-compensated point cloud frame is determined from a coded reference point cloud frame.
[0360] Clause 72. The method according to Clause 71, wherein the attribute-motion-compensated point cloud frame is determined based on motion compensation of the coded reference point cloud frame.
[0361] Clause 73. The method according to any one of Clauses 71 to 72, wherein the coded reference point cloud frame is motion-compensated by attribute motion vectors.
[0362] Clause 74. The method according to any one of Clauses 71 to 73 further includes encoding the attribute motion vector and / or the indication (e.g., index or ID) of the coded reference point cloud frame.
[0363] Clause 75. The method according to any one of Clauses 71 to 74 further includes determining the attribute motion vector based on the difference between the attribute associated with the reconstructed geometry and the attribute adjusted by the attribute motion vector of the geometry of the coded reference point cloud frame.
[0364] Clause 76. The method according to any one of Clauses 59 to 75, wherein determining the attribute predictor includes smoothing the projected attribute, the attribute predictor being determined from the smoothed projected attribute.
[0365] Clause 77. The method according to any one of Clauses 59 to 76, wherein the attribute predictor includes a corresponding attribute predictor for each corresponding vertex in the vertices of the reconstructed geometry, based on the projection attributes corresponding to the vertex in the projection attributes.
[0366] Clause 78. The method according to any one of Clauses 58 to 77 further includes: decoding attribute reference point cloud frame information representing a selected attribute reference point cloud frame; and selecting the attribute reference point cloud frame from a plurality of decoded reference point cloud frames based on the decoded attribute reference point cloud frame information.
[0367] Clause 79. The method according to Clause 78, wherein the attribute reference point cloud frame is selected independently of the reference point cloud frame.
[0368] Clause 80. The method according to any one of Clauses 78 to 79, wherein the attribute reference point cloud frame information indicates the selection of a single attribute reference point cloud frame.
[0369] Clause 81. The method according to Clause 80, wherein the single attribute reference point cloud frame is different from the reference point cloud frame.
[0370] Clause 82. The method according to Clause 80, wherein the single attribute reference point cloud frame is the same as the reference point cloud frame.
[0371] Clause 83. The method according to any one of Clauses 80 to 82, wherein the attribute reference point cloud frame information indicates whether the individual attribute reference point cloud frame and the reference point cloud frame are the same.
[0372] Clause 84. The method according to Clause 80, wherein the single attribute reference point cloud frame is different from the geometrically motion-compensated point cloud frame.
[0373] Clause 85. The method according to Clause 80, wherein the single attribute reference point cloud frame and the geometrically motion-compensated point cloud frame are the same.
[0374] Clause 86. The method according to Clause 80, wherein the attribute reference point cloud frame information indicates whether the individual attribute reference point cloud frame and the geometrically motion-compensated cloud frame are the same.
[0375] Clause 87. The method according to any one of Clauses 80 to 86, wherein the attribute reference point cloud frame information indicates an index pointing to an index table referencing the single attribute reference point cloud frame, each index referencing a specific reference point cloud frame among a plurality of decoded reference point cloud frames.
[0376] Clause 88. The method according to any one of Clauses 80 to 87 further includes: determining a spatial region from the decoded geometry; and selecting an attribute reference point cloud frame for each spatial region.
[0377] Clause 89. The method according to Clause 88, wherein the attribute reference point cloud frame information indicates whether a single attribute reference point cloud frame or an attribute reference point cloud frame is selected per spatial region.
[0378] Clause 90. The method according to any one of Clauses 88 to 89, wherein the attribute reference point cloud frame information indicates the definition of the spatial region.
[0379] Clause 91. The method according to any one of Clauses 88 to 90, wherein the attribute reference point cloud frame information indicates a reference point cloud frame for each spatial region.
[0380] Clause 92. The method according to Clause 91, wherein the attribute reference point cloud frame selected for the spatial region is not equal to the reference point cloud frame used to decode the geometry of the point belonging to the spatial region.
[0381] Clause 93. The method according to any one of Clauses 88 to 92, wherein the attribute reference point cloud frame selected for the spatial region is equal to the reference point cloud frame used to decode the geometry of the point belonging to the spatial region.
[0382] Clause 94. The method according to any one of Clauses 88 to 93, wherein the attribute reference point cloud frame information indicates whether the attribute reference point cloud frame selected for the spatial region is equal to the reference point cloud frame used to decode the geometry of the point belonging to the spatial region.
[0383] Clause 95. The method according to any one of Clauses 88 to 94, wherein the attribute reference point cloud frame selected for the spatial region is equal to the geometrically motion-compensated point cloud frame used to decode the geometry of the points belonging to the spatial region.
[0384] Clause 96. The method according to any one of Clauses 88 to 94, wherein the attribute reference point cloud frame selected for the spatial region is not equal to the geometrically motion-compensated point cloud frame used to decode the geometry of the points belonging to the spatial region.
[0385] Clause 97. The method according to any one of Clauses 88 to 96, wherein the attribute reference point cloud frame information indicates whether the attribute reference point cloud frame selected for the spatial region is equal to the geometrically motion-compensated point cloud frame used to decode the geometry of the points belonging to the spatial region.
[0386] Clause 98. The method according to any one of Clauses 88 to 97, wherein the attribute reference point cloud frame information indicates an index of an index table of attribute reference point cloud frames pointing to a reference space region, each index referencing a specific reference point cloud frame from a plurality of decoded reference point cloud frames.
[0387] Clause 99. The method according to any one of Clauses 58 to 98 further includes: decoding geometric motion vector information representing the geometric motion vector; and decoding attribute motion vector information representing the attribute motion vector.
[0388] Clause 100. The method according to Clause 99, wherein the attribute motion vector information indicates that the attribute motion vector and the geometric motion vector are associated with the same motion field structure.
[0389] Clause 101. The method according to Clause 100, wherein the motion field structure includes partitioning a 3D space surrounding the reconstructed geometry into motion units, wherein a geometric motion vector is associated with each of the motion units, and wherein the motion compensation of the reference point cloud frame is performed within each motion unit based on the associated geometric motion vector, wherein an attribute motion vector is associated with each of the motion units, and wherein the motion compensation of the attribute reference point cloud frame is performed within each motion unit based on the associated attribute motion vector.
[0390] Clause 102. The method according to Clause 101, wherein the motion unit comprises a non-intersecting cuboid.
[0391] Clause 103. The method according to any one of Clauses 101 to 102, wherein the attribute motion vector information indicates the motion field structure.
[0392] Clause 104. The method according to any one of Clauses 99 to 103, wherein the attribute motion information indicates that the attribute motion vector and the geometric motion vector are associated with different motion field structures.
[0393] Clause 105. The method according to Clause 104, wherein an attribute motion field structure associated with the attribute motion vector is determined from a first partition of the 3D space surrounding the decoded geometry associated with the point cloud frame, which is divided into attribute motion units, and wherein a geometric motion field structure associated with the geometric motion vector is determined from a second partition of the 3D space surrounding the decoded geometry associated with the point cloud frame, which is divided into geometric motion units.
[0394] Clause 106. The method according to Clause 105, wherein decoding the attribute motion vector information comprises: determining a motion vector residual based on decoding the attribute motion information; and determining the attribute motion vector of the attribute motion field structure based on adding the motion vector residual to the geometric motion vector of the geometric motion field structure.
[0395] Clause 107. The method according to any one of Clauses 99 to 106, wherein the attribute motion information indicates a boundary (e.g., threshold, range, limit, etc.) on the maximum value of the geometric motion vector and / or a boundary on the attribute motion vector.
[0396] Clause 108. A computing device comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform a method according to any one of Clauses 55 to 107.
[0397] Clause 109. A system comprising: a computing device configured to perform the method according to any one of Clauses 55 to 107; and a second computing device configured to encode the current point cloud frame.
[0398] Clause 110. A computer-readable medium storing instructions that, when executed, cause to perform the method according to any one of Clauses 55 to 107.
[0399] One or more examples in this document can be described as processes that can be depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, and / or block diagrams. Although a flowchart can describe operations as a continuous process, one or more of the operations can be executed in parallel or simultaneously. The order of the operations shown can be rearranged. A process can be terminated when its operations are completed, but may have additional steps not shown in the diagram. A process can correspond to a method, function, program, subroutine, subroutines, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.
[0400] The operations described herein can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments for performing necessary tasks (e.g., computer program products) can be stored on a computer-readable or machine-readable medium. A processor can perform the necessary tasks. The features of this disclosure can be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. Implementing a hardware state machine to perform the functions described herein will also be apparent to those skilled in the art.
[0401] One or more features described herein may be implemented in computer-usable data and / or computer-executable instructions, as in one or more program modules, which are executed by one or more computers or other devices. Typically, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type when executed by a processor in a computer or other data processing device. Computer-executable instructions may be stored on one or more computer-readable media, such as hard disks, optical disks, removable storage media, solid-state drives, RAM, etc. The functionality of program modules may be combined or distributed as needed. The functionality may be implemented, in whole or in part, in firmware or hardware equivalents, such as integrated circuits, field-programmable gate arrays (FPGAs), etc. One or more features described herein may be implemented more efficiently using specific data structures, and such data structures are contemplated within the scope of the computer-executable instructions and computer-usable data described herein. Computer-readable media may include, but are not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored but do not include carrier waves and / or transient electronic signals propagated wirelessly or via wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as CDs or DVDs, flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions that can represent any combination of programs, functions, subroutines, routines, subroutines, modules, software packages, classes or instructions, data structures, or program statements. Code segments can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0402] Non-transitory tangible computer-readable media may include instructions executable by one or more processors configured to cause the operations described herein. Articles of manufacture may include non-transitory tangible computer-readable machine-accessible media having instructions encoded thereon for enabling programmable hardware to allow devices (e.g., encoders, decoders, transmitters, receivers, etc.) to perform the operations described herein. Devices, or one or more devices such as in a system, may include one or more processors, memories, interfaces, etc.
[0403] The communication described herein can be determined, generated, sent, and / or received using any number of messages, information elements, fields, parameters, values, indications, information, bits, etc. While this document may use any of the terms / phrases message, information element, field, parameter, value, indication, information, bit, etc., to describe one or more examples, those skilled in the art will understand that any one or more of these terms, including other such terms, can be used to perform such communication. For example, one or more parameters, fields, and / or information elements (IEs) may include one or more information objects, values, and / or any other information. An information object may include one or more other objects. At least some (or all) parameters, fields, IEs, etc., may be used and may be interchangeable depending on the context. Where a meaning or definition is given, such meaning or definition shall prevail.
[0404] One or more elements in the examples described herein can be implemented as modules. A module can be an element that performs a defined function and / or has a defined interface to other elements. Modules can be implemented as hardware, software combined with hardware, firmware, wet hardware (e.g., hardware with biological elements), or a combination thereof, all of which can be behaviorally equivalent. For example, a module can be implemented as software routines written in a computer language configured to be executed by a hardware machine (such as C, C++, Fortran, Java, Basic, Matlab, etc.) or a modeling / simulation program (such as Simulink, Stateflow, GNU Octave, or LabVIEW MathScript). Alternatively or alternatively, modules can be implemented using physical hardware that incorporates discrete or programmable analog, digital, and / or quantum hardware. Examples of programmable hardware can include: computers, microcontrollers, microprocessors, application-specific integrated circuits (ASICs); field-programmable gate arrays (FPGAs); and / or complex programmable logic devices (CPLDs). Computers, microcontrollers, and / or microprocessors can be programmed using languages such as assembly, C, C++, etc. Hardware description languages (HDLs) such as VHSIC (VHDL) or Verilog are typically used to program FPGAs, ASICs, and CPLDs. These HDLs can configure connections between internal hardware modules with limited functionality on a programmable device. The techniques mentioned above can be combined to achieve the desired functional modules.
[0405] One or more operations described herein may be conditional. For example, one or more operations may be performed if certain criteria are met in a computing device, communication device, encoder, decoder, network, or a combination thereof. Example criteria may be based on one or more conditions, such as device configuration, traffic load, initial system settings, packet size, service characteristics, or a combination thereof. Various examples may be used if the one or more criteria are met. Any part of the examples described herein may be implemented in any order and based on any conditions.
[0406] Although examples have been described above, features and / or steps of those examples can be combined, divided, omitted, rearranged, modified, and / or expanded in any desired manner. Various changes, modifications, and improvements will readily occur to those skilled in the art. While not explicitly stated herein, such changes, modifications, and improvements are intended to be part of this specification and are intended to be within the spirit and scope of this specification. Therefore, the above description is illustrative only and not restrictive.
Claims
1. A method comprising: The geometric motion vector is determined from the bit stream by the decoder; Determine the attribute motion vector from the bit stream; The reconstructed geometry of the point cloud frame is determined by decoding the geometry associated with the point cloud frame related to the content based on the geometric motion vector. as well as The attributes associated with the reconstructed geometry are decoded based on the attribute motion vector.
2. The method of claim 1, wherein decoding the attribute associated with the reconstructed geometry comprises: The projection attributes are determined based on the motion vector of the aforementioned attributes; An attribute predictor that determines the attributes associated with the reconstructed geometry based on the projection attributes; as well as The attribute predictor decodes the attribute associated with the reconstructed geometry.
3. The method according to claim 2, wherein determining the projection attribute includes: The projection attribute is determined based on motion compensation of the coded reference point cloud frame, wherein the coded reference point cloud frame is motion compensated by the attribute motion vector.
4. The method according to claim 3, wherein determining the projection attribute further comprises: Identify the point cloud frame after attribute motion compensation; as well as The projection attributes are determined by projecting the attributes of the attribute-motion-compensated point cloud frame onto the reconstructed geometry.
5. The method according to any one of claims 1 to 4, wherein decoding the attribute associated with the reconstructed geometry comprises: Decode the residual attributes that indicate the difference between the attribute of the reconstructed geometry and the attribute predictor associated with the reconstructed geometry; as well as The attributes of the reconstructed geometry are determined by adding the attribute predictor to the residual attributes.
6. The method according to any one of claims 1 to 5, wherein motion compensation is performed on a reference point cloud frame of the point cloud frame based on the geometric motion vector to determine the reconstructed geometry of the point cloud frame, the method further comprising: Decode the geometric motion vector information that indicates the geometric motion vector, wherein the geometric motion vector is associated with the same motion field structure; as well as The attribute motion vector information that indicates the attribute motion vector is decoded.
7. The method according to claim 6, wherein decoding the attribute motion vector information comprises: The motion vector residual is determined based on the decoding of the attribute motion vector information; as well as The attribute motion vector of the attribute motion field structure is determined by adding the residual motion vector to the geometric motion vector of the geometric motion field structure.
8. The method according to any one of claims 1 to 7, further comprising: Decode the attribute reference point cloud frame information of the indicator attribute reference point cloud frame; as well as Based on the attribute reference point cloud frame information, the attribute reference point cloud frame is selected from multiple decoded reference point cloud frames.
9. A method comprising: The decoder determines the attribute motion vector and geometric motion vector from the bit stream; The reconstructed geometry of the point cloud frame is determined by decoding the geometry associated with the point cloud frame related to the content based on the geometric motion vector. The projection attributes are determined based on the motion vector of the aforementioned attributes; An attribute predictor that determines the attributes associated with the reconstructed geometry based on the projection attributes; as well as The residual attributes, which indicate the difference between the attributes associated with the reconstructed geometry and the attribute predictor, are decoded from the bitstream.
10. The method of claim 9, wherein decoding the residual property comprises: Decode the transformed coefficients indicating the residual properties from the bit stream; as well as The residual properties are determined by applying the inverse intra-frame transform to the decoded transform coefficients.
11. The method according to any one of claims 9 to 10, wherein the residual property is decoded based on a prediction-enhanced transformation scheme.
12. The method according to any one of claims 9 to 11, wherein determining the attribute predictor further comprises: Smooth the projection properties; as well as The attribute predictor is determined based on the smoothed projection properties.
13. A computing device, comprising: One or more processors; and a memory that stores instructions, which, when executed by the one or more processors, cause the computing device to perform: The method according to any one of claims 1 to 8; or The method according to any one of claims 9 to 12.
14. A system comprising: A computing device configured to perform: The method according to any one of claims 1 to 8; or The method according to any one of claims 9 to 12; and A second computing device is configured to encode the point cloud frame.
15. A computer-readable medium storing instructions that, when executed, cause the following to be performed: The method according to any one of claims 1 to 8; or The method according to any one of claims 9 to 12.