Method, device and medium for point cloud coding and decoding
By determining the azimuth angle of key points and their predicted values in point cloud encoding and decoding, and establishing an entropy encoding and decoding context, the problem of insufficient point cloud encoding and decoding efficiency in the prior art is solved, especially when processing sparse point cloud data, more efficient encoding and decoding performance is achieved.
Patent Information
- Application Number
- CN202380074132.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-10-19
- Publication Date
- 2025-06-06
AI Technical Summary
The existing point cloud encoding and decoding technology has shortcomings in encoding and decoding efficiency, especially when processing sparse point cloud data, it is difficult to effectively utilize the motion information captured by LIDAR.
By determining multiple azimuth values of key points associated with the current frame node and their predicted values, an entropy codec context for the plane position of at least one axis is established, thereby improving the efficiency of point cloud codec.
It improves the efficiency of point cloud encoding and decoding, especially when processing sparse point cloud data, it can more effectively utilize the motion information captured by LIDAR, and improves the encoding and decoding performance.
Smart Images

Figure CN120112949A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video coding techniques, and more particularly to entropy coding of plane information for point cloud coding. Background Art
[0002] A point cloud is a collection of individual data points in a three-dimensional (3D) plane, where each point has set coordinates on the X, Y, and Z axes. Point clouds can therefore be used to represent the physical content of a three-dimensional space. Point clouds have proven to be a promising way to represent 3D visual data for a variety of immersive applications, from augmented reality to self-driving cars.
[0003] Point cloud codec standards have evolved primarily through the development of the well-known MPEG organization. MPEG stands for Moving Picture Experts Group, which is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Codec Group (3DG) released a Call for Proposals (CFP) document to begin the development of a point cloud codec standard. The final standard will encompass two categories of solutions. Video-based point cloud compression (V-PCC or VPCC) is suitable for point sets with relatively uniform point distribution. Geometry-based point cloud compression (G-PCC or GPCC) is suitable for more sparse distributions. However, it is generally expected to further improve the codec efficiency of conventional point cloud codec techniques. Summary of the invention
[0004] The embodiments of the present disclosure provide a solution for point cloud encoding and decoding.
[0005] In a first aspect, a method for point cloud encoding and decoding is proposed. The method includes: during the conversion between a current frame of a point cloud sequence and a bitstream of the point cloud sequence, determining a plurality of azimuth angle measurements of a plurality of key points associated with a node of the current frame, the node representing a spatial segmentation of the current frame; determining a prediction of the azimuth angle measurement of the node; based on the plurality of azimuth angle measurements and the prediction of the azimuth angle measurement, determining a context for entropy encoding and decoding of a planar position of at least one axis associated with the node; and performing the conversion based on the context for entropy encoding and decoding. The method according to the first aspect of the present disclosure determines a context for entropy encoding and decoding of a planar position of at least one axis based on the azimuth angles and predicted azimuth angles of several key points, and thus can improve the efficiency of point cloud encoding and decoding.
[0006] In a second aspect, an apparatus for processing a point cloud sequence is provided. The apparatus for processing a point cloud sequence comprises a processor and a non-volatile memory having instructions thereon. These instructions, when executed by the processor, cause the processor to perform a method according to the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect, a non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bit stream of a point cloud sequence generated by a method executed by a point cloud sequence processing device. The method includes: determining multiple azimuth angle values of multiple key points associated with a node of a current frame of the point cloud sequence, the node representing a spatial domain segmentation of the current frame; determining a prediction of the azimuth angle value of the node; based on the multiple azimuth angle values and the prediction of the azimuth angle value, determining a context for entropy coding and decoding of a planar position of at least one axis associated with the node; and generating a bit stream based on the entropy coding and decoding context.
[0009] In a fifth aspect, a method for storing a bitstream of a point cloud sequence is proposed. The method includes: determining multiple azimuth angle values of multiple key points associated with a node of a current frame of the point cloud sequence, the node representing a spatial segmentation of the current frame; determining a prediction of the azimuth angle value of the node; determining a context for entropy coding and decoding of a plane position of at least one axis associated with the node based on the multiple azimuth angle values and the prediction of the azimuth angle value; generating a bitstream based on the entropy coding and decoding context; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] Figure 1 A block diagram illustrating an example point cloud encoding and decoding system according to some embodiments of the present disclosure is shown;
[0013] Figure 2 shows a block diagram illustrating an example of a GPCC encoder according to some embodiments of the present disclosure;
[0014] Figure 3 shows a block diagram illustrating an example of a GPCC decoder according to some embodiments of the present disclosure;
[0015] Figure 4An example of an encoding and decoding process for improving point cloud geometry encoding and decoding using LIDAR characteristics is shown;
[0016] Figure 5 A flowchart of a method for point cloud encoding and decoding according to some embodiments of the present disclosure is shown; and
[0017] Figure 6 A block diagram of a computing device is shown in which various embodiments of the present disclosure may be implemented.
[0018] Same or similar reference numbers generally refer to same or similar elements throughout the drawings. DETAILED DESCRIPTION
[0019] The principle of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, without implying any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0020] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0021] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art to affect correlation with other embodiments.
[0022] It should be understood that, although the terms "first" and "second" etc. may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another element. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0023] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "include", "comprises", "has", "has", "includes" and / or "comprising" are used herein to indicate the presence of the features, elements and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0024] Example Environment
[0025] Figure 1 is a block diagram illustrating an example point cloud codec system 100 that may utilize the techniques of the present disclosure. As shown, the point cloud codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a point cloud codec device, and the destination device 120 may also be referred to as a point cloud decoding device. In operation, the source device 110 may be configured to generate encoded point cloud data, and the destination device 120 may be configured to decode the encoded point cloud data generated by the source device 110. The techniques of the present disclosure are generally directed to encoding and decoding (encoding and / or decoding) point cloud data, i.e., supporting point cloud compression. The codec may be effective in compressing and / or decompressing point cloud data.
[0026] Source device 100 and destination device 120 may include any of a variety of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets (such as smart phones and mobile phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, vehicles (e.g., land or sea vehicles, spacecraft, aircraft, etc.), robots, LIDAR devices, satellites, extended reality devices, etc. In some cases, source device 100 and destination device 120 may be equipped for wireless communication.
[0027] The source device 100 may include a data source 112, a memory 114, a GPCC encoder 116, and an input / output (I / O) interface 118. The destination device 120 may include an input / output (I / O) interface 128, a GPCC decoder 126, a memory 124, and a data consumer 122. According to the present disclosure, the GPCC encoder 116 of the source device 100 and the GPCC decoder 126 of the destination device 120 may be configured to apply the techniques related to point cloud encoding and decoding of the present disclosure. Therefore, the source device 100 represents an example of an encoding device, and the destination device 120 represents an example of a decoding device. In other examples, the source device 100 and the destination device 120 may include other components or arrangements. For example, the source device 100 may receive data (e.g., point cloud data) from an internal source or an external source. Similarly, the destination device 120 may be connected to an external data consumer interface instead of including the data consumer in the same device.
[0028] In general, the data source 112 represents a source of point cloud data (i.e., raw, unencoded point cloud data) and can provide a continuous series of "frames" of point cloud data to the GPCC encoder 116, which encodes the point cloud data for the frames. In some examples, the data source 112 generates point cloud data. The data source 112 of the source device 100 may include a point cloud acquisition device, such as any of a variety of cameras or sensors, such as one or more cameras, an archive containing previously acquired point cloud data, a 3D scanner or a light detection and ranging (LIDAR) device, and / or a data feed interface for receiving point cloud data from a data content provider. Therefore, in some examples, the data source 112 can generate point cloud data based on a signal from a LIDAR device. Alternatively or additionally, the point cloud data can be generated by a computer from a scanner, camera, sensor, or other data. For example, the data source 112 can generate point cloud data, or produce a combination of real-time point cloud data, archived point cloud data, and computer-generated point cloud data. In each case, the GPCC encoder 116 encodes the acquired, pre-acquired, or computer-generated point cloud data. The GPCC encoder 116 can rearrange the frames of the point cloud data from the receive order (sometimes referred to as the "display order") to the codec order for encoding and decoding. The GPCC encoder 116 can generate one or more bit streams including the encoded point cloud data. The source device 100 can then output the encoded point cloud data via the I / O interface 118 for reception and / or retrieval by, for example, the I / O interface 128 of the destination device 120. The encoded point cloud data can be transmitted directly to the destination device 120 via the network 130A via the I / O interface 118. The encoded point cloud data can also be stored on the storage medium / server 130B for access by the destination device 120.
[0029] The memory 114 of the source device 100 and the memory 124 of the destination device 120 may represent general purpose memory. In some examples, the memory 114 and the memory 124 may store raw point cloud data, for example, raw point cloud data from the data source 112 and raw, decoded point cloud data from the GPCC decoder 126. Additionally or alternatively, the memory 114 and the memory 124 may store software instructions executable by, for example, the GPCC encoder 116 and the GPCC decoder 126, respectively. Although the memory 114 and the memory 124 are shown separately from the GPCC encoder 116 and the GPCC decoder 126 in this example, it should be understood that the GPCC encoder 116 and the GPCC decoder 126 may also include internal memory for functionally similar or equivalent purposes. In addition, the memory 114 and the memory 124 may store encoded point cloud data, for example, encoded point cloud data output from the GPCC encoder 116 and input to the GPCC decoder 126. In some examples, portions of memory 114 and memory 124 may be allocated as one or more caches, for example, to store raw point cloud data, decoded point cloud data, and / or encoded point cloud data. For example, memory 114 and memory 124 may store point cloud data.
[0030] I / O interface 118 and I / O interface 128 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component that operates according to any of a variety of IEEE 802.11 standards, or other physical components. In examples where I / O interface 118 and I / O interface 128 include wireless components, I / O interface 118 and I / O interface 128 may be configured to transmit data, such as encoded point cloud data, according to a cellular communication standard (such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc.). In some examples where I / O interface 118 includes a wireless transmitter, I / O interface 118 and I / O interface 128 may be configured to transmit data, such as encoded point cloud data, according to other wireless standards (such as IEEE 802.11 specifications). In some examples, source device 100 and / or destination device 120 may include corresponding system-on-chip (SoC) devices. For example, source device 100 may include a SoC device for performing the functions attributed to GPCC encoder 116 and / or I / O interface 118 , and destination device 120 may include a SoC device for performing the functions attributed to GPCC decoder 126 and / or I / O interface 128 .
[0031] The techniques disclosed herein may be applied to encoding and decoding to support any of a variety of applications, such as communications between autonomous vehicles, communications between scanners, cameras, sensors and processing devices (e.g., local servers or remote servers), geographic mapping, or other applications.
[0032] The I / O interface 128 of the destination device 120 receives the encoded bitstream from the source device 110. The encoded bitstream may include signaling information defined by the GPCC encoder 116, which is also used by the GPCC decoder 126, such as a syntax element having a value representing a point cloud. The data consumer 122 uses the decoded data. For example, the data consumer 122 may use the decoded point cloud data to determine the location of a physical object. In some examples, the data consumer 122 may include a display for presenting an image based on the point cloud data.
[0033] The GPCC encoder 116 and the GPCC decoder 126 can each be implemented as any of a variety of suitable encoder circuit systems and / or decoder circuit systems, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the technology of the present disclosure. Each of the GPCC encoder 116 and the GPCC decoder 126 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device. The device including the GPCC encoder 116 and / or the GPCC decoder 126 may include one or more integrated circuits, microprocessors, and / or other types of devices.
[0034] The GPCC encoder 116 and the GPCC decoder 126 may operate in accordance with a codec standard, such as the Video Point Cloud Compression (VPCC) standard or the Geometry Point Cloud Compression (GPCC) standard. In general, the present disclosure may refer to the encoding and decoding of frames (e.g., encoding and decoding) to include the process of encoding data or decoding data. The encoded bitstream typically includes a series of values for syntax elements that represent codec decisions (e.g., codec modes).
[0035] A point cloud may contain a set of points in 3D space and may have attributes associated with the points. The attributes may be color information, such as R, G, B or Y, Cb, Cr or reflectivity information, or other attributes. Point clouds may be collected by various cameras or sensors (such as LIDAR sensors and 3D scanners), and may also be computer generated. Point cloud data is used in a variety of applications, including but not limited to architecture (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors to aid navigation).
[0036] Figure 2 is a block diagram showing an example of a GPCC encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of a GPCC encoder 116 in the system 100 is shown. Figure 3 is a block diagram showing an example of a GPCC decoder 300 according to some embodiments of the present disclosure. The GPCC decoder 300 may be Figure 1 An example of a GPCC decoder 126 in the system 100 is shown.
[0037] In both the GPCC encoder 200 and the GPCC decoder 300, the point cloud position is first encoded and decoded. The attribute encoding and decoding depends on the decoded geometry. Figure 2 and Figure 3 , the region adaptive hierarchical transform (RAHT) unit 218, the surface approximation analysis unit 212, the RAHT unit 314, and the surface approximation synthesis unit 310 are options that are commonly used for category 1 data. The level of detail (LOD) generation unit 220, the lifting unit 222, the LOD generation unit 316, and the delifting unit 318 are options that are commonly used for category 3 data. All other units are common between category 1 and category 3.
[0038] For category 3 data, the compressed geometry is typically represented as an octree from the root down to the leaf level for each voxel. For category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to the leaf level for blocks larger than a voxel) plus a model for approximating the surface within each leaf node of the pruned octree. In this way, both category 1 and category 3 data share the octree codec mechanism, while category 1 data can additionally utilize a surface model to approximate the voxels within each leaf node. The surface model used is a triangulation of 1 to 10 triangles per block, resulting in a triangle soup. Therefore, category 1 geometry codecs are called triangle soup geometry codecs, while category 3 geometry codecs are called octree geometry codecs.
[0039] exist Figure 2 In the example, the GPCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometric reconstruction unit 216, a RAHT unit 218, an LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224 and an arithmetic coding unit 226.
[0040] like Figure 2 As shown in the example of , the GPCC encoder 200 can receive a set of positions and a set of attributes. The positions can include the coordinates of a point in the point cloud. The attributes can include information about the point in the point cloud, such as the color associated with the point in the point cloud.
[0041] The coordinate transformation unit 202 may apply a transformation to the coordinates of the point to transform the coordinates from the initial domain to the transformed domain. The present disclosure may refer to the transformed coordinates as transformed coordinates. The color transformation unit 204 may apply a transformation to transform the color information of the attribute to a different domain. For example, the color transformation unit 204 may transform the color information from the RGB color space to the YCbCr color space.
[0042] In addition, Figure 2 In the example of , the voxelization unit 206 may voxelize the transformed coordinates. Voxelization of the transformed coordinates may include quantizing and removing some points of the point cloud. In other words, multiple points of the point cloud may be grouped into a single "voxel", which may thereafter be considered as a point in some aspects. In addition, the octree analysis unit 210 may generate an octree based on the voxelized transformed coordinates. Additionally, in Figure 2 In the example of , the surface approximation analysis unit 212 can analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 can perform arithmetic coding on syntax elements representing information of the octree and / or information of the surface determined by the surface approximation analysis unit 212. The GPCC encoder 200 can output these syntax elements in a geometry bitstream.
[0043] The geometric reconstruction unit 216 may reconstruct the transformed coordinates of the points in the point cloud based on the octree, the data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transformed coordinates reconstructed by the geometric reconstruction unit 216 may be different from the original number of points in the point cloud. The present disclosure may refer to the resulting points as reconstructed points. The attribute transfer unit 208 may transfer the attributes of the original points of the point cloud to the reconstructed points of the point cloud data.
[0044] In addition, the RAHT unit 218 may apply RAHT coding to the attributes of the reconstruction point. Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 may apply LOD processing and lifting, respectively, to the attributes of the reconstruction point. The RAHT unit 218 and the lifting unit 222 may generate coefficients based on the attributes. The coefficient quantization unit 224 may quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 may apply arithmetic coding to the syntax elements representing the quantized coefficients. The GPCC encoder 200 may output these syntax elements in the attribute bitstream.
[0045] exist Figure 3 In the example, the GPCC decoder 300 may include a geometric arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometric reconstruction unit 312, a RAHT unit 314, an LOD generation unit 316, an inverse lifting unit 318, a coordinate inverse transformation unit 320 and a color inverse transformation unit 322.
[0046] The GPCC decoder 300 may obtain a geometry bitstream and an attribute bitstream. The geometry arithmetic decoding unit 302 of the decoder 300 may apply arithmetic decoding (e.g., CABAC or other types of arithmetic decoding) to syntax elements in the geometry bitstream. Similarly, the attribute arithmetic decoding unit 304 may apply arithmetic decoding to syntax elements in the attribute bitstream.
[0047] The octree synthesis unit 306 may synthesize the octree based on the syntax elements parsed from the geometry bitstream. In the case where surface approximation is used in the geometry bitstream, the surface approximation synthesis unit 310 may determine the surface model based on the syntax elements parsed from the geometry bitstream and based on the octree.
[0048] Furthermore, the geometry reconstruction unit 312 may perform reconstruction to determine the coordinates of the points in the point cloud. The coordinate inverse transformation unit 320 may apply an inverse transformation to the reconstructed coordinates to convert the reconstructed coordinates (positions) of the points in the point cloud from the transformed domain back to the original domain.
[0049] Additionally, in Figure 3 In the example of , the inverse quantization unit 308 may inverse quantize the property value. The property value may be based on a syntax element obtained from the property bitstream (eg, including a syntax element decoded by the property arithmetic decoding unit 304).
[0050] Depending on how the attribute values are encoded, the RAHT unit 314 may perform RAHT decoding to determine color values for points in the point cloud based on the dequantized attribute values. Alternatively, the LOD generation unit 316 and the de-lifting unit 318 may use a level of detail based technique to determine color values for points in the point cloud.
[0051] In addition, Figure 3 In the example of , the color inverse transform unit 322 can apply an inverse color transform to the color values. The inverse color transform can be the inverse of the color transform applied by the color transform unit 204 of the encoder 200. For example, the color transform unit 204 can transform the color information from the RGB color space to the YCbCr color space. Accordingly, the color inverse transform unit 322 can transform the color information from the YCbCr color space to the RGB color space.
[0052] Figure 2 and Figure 3 Various units are shown to help understand the operations performed by the encoder 200 and the decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functions and are preset with respect to the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions in the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by fixed-function circuits are generally immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.
[0053] Some exemplary embodiments of the present disclosure are described in detail below. It should be understood that the section titles used in this document are for ease of understanding, and do not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to GPCC or other specific point cloud codecs, the disclosed techniques are also applicable to other point cloud coding and decoding technologies. In addition, although some embodiments describe the point cloud coding and decoding steps in detail, it should be understood that the corresponding decoding steps of the de-coding will be implemented by the decoder.
[0054] 1. Brief Overview
[0055] The present disclosure relates to point cloud coding and decoding technology. In particular, it relates to point cloud geometry coding and decoding using LIDAR characteristics. These concepts can be applied to any point cloud coding and decoding standard or non-standard point cloud codec, such as the geometry-based point cloud compression (G-PCC) under development, either alone or in various combinations.
[0056] 2. Abbreviations
[0057] G-PCC Geometry-based Point Cloud Compression
[0058] MPEG Moving Picture Experts Group
[0059] 3DG 3D Graphics Codec Group
[0060] CFP Request for Proposals
[0061] V-PCC video-based point cloud compression
[0062] DCM direct encoding and decoding mode
[0063] IDCM inference direct codec mode
[0064] 3. Introduction
[0065] MPEG is short for Moving Picture Experts Group, which is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Codec Group (3DG) released a call for proposals (CFP) document to start the development of a point cloud codec standard. The final standard will include two categories of solutions. Video-based point cloud compression (V-PCC) is suitable for point sets with relatively uniform point distribution. Geometry-based point cloud compression (G-PCC) is suitable for more sparse distributions. Both V-PCC and G-PCC support encoding and decoding for single point clouds and point cloud sequences.
[0066] In a point cloud, there can be geometric information and attribute information. Geometric information is used to describe the geometric position of data points. Attribute information is used to record some details of data points, such as texture, normal vector, reflection, etc. One of the important applications of point clouds is autonomous driving. In autonomous driving, point cloud data is mainly captured by LIDAR. Therefore, some important characteristics of LIDAR can be used to compress point clouds. For example, for standard spin-type LIDARs, they always consist of multiple laser diodes aligned vertically, resulting in an effective vertical (elevation) field of view. The entire unit can then rotate only around its vertical axis at a fixed speed to provide a full 360-degree azimuth field of view. The elevation and azimuth angles of the laser beam can be used to compress point cloud geometric information.
[0067] Point cloud codecs can handle various information in different ways. Usually, there are many optional tools in the codec to support the encoding and decoding of geometry information and attribute information respectively. Among the geometry codec tools in G-PCC, the following tools have an important impact on the performance of point cloud geometry codec.
[0068] 3.1 Octree Geometry Compression
[0069] In G-PCC, one of the important point cloud geometry codec tools is octree geometry compression, which utilizes the spatial correlation of point cloud geometry. If the geometry codec tool is enabled, the cube axis-aligned bounding box associated with the octree root node will be determined based on the point cloud geometry information. The bounding box will then be subdivided into 8 sub-cubes, which are associated with the 8 child nodes of the root node (cubes are equivalent to nodes below). An 8-bit code is then generated in a specific order to individually indicate whether the 8 child nodes contain points, with one bit associated with a node. The 8-bit code is named an occupancy code and will be transmitted by signaling based on the occupancy information of neighboring nodes. Then, nodes containing only points will be further subdivided into 8 child nodes. This process will be recursively performed until the node size is 1. Therefore, the point cloud geometry information is converted into an occupancy code sequence.
[0070] At the decoder side, the occupancy code sequence will be decoded and the point cloud geometry information can be reconstructed based on the occupancy code sequence.
[0071] 3.2 Plane Mode
[0072] Planar mode is a tool used to improve the occtree node's occupancy code more efficiently. Before encoding or decoding the node's occupancy code, the node is judged to be suitable for planar mode according to certain suitability conditions in three dimensions.
[0073] Take the z-axis as an example. If plane mode is applicable on the z-axis, a binary flag zIsPlanar is decoded to signal whether its occupied child nodes belong to the same horizontal plane. If zIsPlaner is true, an additional bit zPlanePosition is signaled to indicate whether the plane is the lower plane or the upper plane, and the empty plane occupancy code can be ignored. Otherwise, the node will continue the normal tree encoding and decoding process. Applicability is based on tracking the probability that the previously decoded node is a plane, as follows.
[0074] If and only if p planar ≥T and d local >3, the node is applicable, where T is a user-defined probability threshold and d localIt is the local density that can be derived based on the information of neighboring nodes.
[0075] When node occupancy is encoded (decoded) or / and node plane information is encoded (decoded), the following changes are made:
[0076] The new probability p planar :
[0077] p planar =(L×p planar +δ) / (L+1)
[0078] Where L=255, and δ is 1 if the encoded node is a plane, and 0 otherwise.
[0079] The flag zIsPlaner is encoded and decoded using the 3 contexts based on the axis information by using the binary arithmetic codec. If zIsPlaner is true, zPlanePosition is encoded and decoded by using the binary arithmetic codec.
[0080] 3.3 Inferred Direct Coding Mode (IDCM)
[0081] The octree representation (or more generally any tree representation) is efficient at representing points with spatial correlations because the tree tends to decompose the high-order bits of the point coordinates. For the octree, each depth level refines the point coordinates within the child nodes by one bit per component, at a cost of 8 bits per refinement. Further compression is obtained by entropy encoding and decoding the partitioning information associated with each tree node.
[0082] However, if a node of the octree contains isolated points, directly encoding and decoding its relative coordinates in the node is better than the octree representation. Because no other points exist in the node, spatial correlation cannot be used. Directly encoding and decoding the point coordinates in the node / child node is called direct coding mode (DCM). On the other hand, using DCM will reduce the time complexity because the octree recursive partitioning process cannot be performed.
[0083] In G-PCC, each node is judged to be suitable for DCM according to certain suitability conditions, which is called Inferred Direct Codec Mode (IDCM). If the node is suitable for DCM, a binary flag is encoded to signal to the node whether DCM is applied (flag = 1) or not (flag = 0). If the flag is equal to 1, the points belonging to the associated volume are directly encoded and decoded using DCM. Otherwise (flag is equal to 0), the tree encoding and decoding process continues for the current node.
[0084] Currently, there are two applicability conditions for IDCM.
[0085] Based on the suitability of the parent. There is only one occupied child (= current node) at the parent node level, and
[0086] A grandparent node has at most two occupied children (= parent node + possibly one other node).
[0087] 6N applicability. There is only one occupied child (= current node) at the parent node level, and no occupied neighbors N (among the six neighbors that share faces with the current cube associated with the current node).
[0088] 3.4 Angle Mode
[0089] In G-PCC, angle mode is introduced to improve the compression of isolated point relative coordinates in IDCM and planar positions in planes. It can be used only for real-time LIDAR captured point cloud data. For standard spindle LIDAR, each laser has a fixed elevation angle and a fixed maximum number of points is captured per rotation. Angle mode uses the previous fixed elevation angle of each laser. It uses the subnode elevation distance to the laser elevation angle to improve the compression of binary occupancy codec through the prediction of the planar position of the plane mode and the prediction of the z-coordinate bit in the DCM node.
[0090] Angular mode is applied to nodes that meet elevation suitability, i.e., if the elevation dimension is below the minimum elevation increment between two adjacent lasers. If a node is suitable, it is only passed by one laser in the elevation direction. Then, the laser passing the node elevation will be found and several keypoint elevations of the node will be calculated. Based on the relationship of several keypoint elevations to the laser passing the node elevation, a context will be determined to help encode and decode the z coordinate bit in the DCM and the planar position of the z axis in the planar mode.
[0091] 3.5 Azimuth Mode
[0092] Similar to the angle mode, the azimuth mode is introduced to improve the compression of the relative coordinates of isolated points in the IDCM and the planar positions in the plane. It can also be used for real-time LIDAR to capture point cloud data. The azimuth mode uses the prior information of a fixed maximum number of points captured per laser per rotation. It uses the azimuth of the already encoded node to improve the compression of the binary occupancy codec through the prediction of the x or y planar position of the planar mode and the prediction of the x or y coordinate bits in the DCM node.
[0093] In the current G-PCC, if a node is applicable to angle mode, it is applicable to azimuth mode. If a node is applicable to azimuth mode, the index of the laser passing through the node will be found. Based on the laser information and the azimuth of the already encoded and decoded node with the same laser as the current node, the predicted azimuth is determined. Then the azimuths of several key points of the node are calculated. Based on the positional relationship between the azimuths of several key points and the predicted azimuth, the context will be determined to help encode and decode the x-coordinate or y-coordinate bits in the DCM and the plane position of the x or y axis in the plane mode.
[0094] 4. Question
[0095] The existing point cloud geometry encoding and decoding design has the following problems:
[0096] 1. In the current G-PCC, the occupancy information of the parent node and the neighboring nodes is used for the IDCM applicability condition. In other words, whether the point in the current node is an isolated point is only derived from the occupancy information of the parent node or the neighboring node. However, for LIDAR captured point cloud data, there is some prior information that can be used for isolated point judgment. For example, if a node is only passed by one laser beam, it is most likely to contain only one point, which means that the point is most likely to be an isolated point.
[0097] 2. In the current G-PCC, a node is applicable to the azimuth mode if and only if it is applicable to the angle mode. However, the applicability condition for the angle mode can only ensure that it is only passed by one laser beam in the elevation direction. It is unclear whether the node is only passed by one laser beam in the azimuth direction.
[0098] Therefore, the applicability conditions in the azimuth direction should be considered. 5. Specific implementation methods
[0100] In order to solve the above problems and some unmentioned problems, the following method is disclosed. The embodiments should be considered as examples to explain the general concept and should not be interpreted in a narrow way. In addition, these embodiments can be applied individually or in combination in any way.
[0101] 1) For a node, a capture laser for capturing the node can be determined.
[0102] a. In one example, the acquisition laser of the current node may be determined based on at least one representative node position and / or at least one laser elevation angle.
[0103] i. In one example, the elevation angle of a node may be calculated based on the node position.
[0104] ii. In one example, the laser with the closest elevation angle to the elevation angle of the node can be considered as the capture laser.
[0105] iii. Alternatively, the laser having the smallest elevation angle among the lasers having elevation angles greater than that of the node may be regarded as the capture laser.
[0106] iv. Alternatively, the laser having the largest elevation angle among the lasers having elevation angles smaller than that of the node may be regarded as the capture laser.
[0107] v. In one example, a predefined point position of a node may be used as a representative position of the node, such as a midpoint position, a vertex position, an original point position, etc.
[0108] 1. Alternatively, the representative position may be signaled from the encoder to the decoder.
[0109] vi. In one example, a function value of a node's elevation angle may be used to represent its elevation angle.
[0110] 1. The function can be tangent, cotangent, sine, cosine, etc.
[0111] vii. In one example, a function value of the elevation angle of a laser can be used to represent its elevation angle.
[0112] 1. The function can be tangent, cotangent, sine, cosine, etc.
[0113] b. Alternatively, the acquisition laser of the current node may be determined based on the acquisition lasers of other nodes.
[0114] i. In one example, the acquisition laser of the current node may be the acquisition laser of its parent node.
[0115] c. Alternatively, in addition, the acquisition laser of the current node can be used to determine whether the node is passed by only one laser beam in the elevation direction or the azimuth direction.
[0116] 2) Determine during the encoding / decoding process whether the node is restricted to be passed by only one laser beam in the elevation direction.
[0117] a. In one example, whether a node is passed by only one laser beam in the elevation direction may depend on the relationship between the elevation size covered by the node and the effective elevation scan size of its acquisition laser.
[0118] i. In one example, if the elevation dimension covered by a node is smaller than the effective elevation scan dimension of its acquisition laser, the node may be considered to be traversed by only one laser beam in the elevation direction.
[0119] ii. Alternatively, if the elevation size of the coverage by a node is less than or equal to the effective elevation scan size of its acquisition laser, then the node may be considered to be traversed by only one laser beam in the elevation direction.
[0120] b. In one example, the effective elevation scan size of its acquisition laser may be equal to the minimum absolute value of the elevation angle differences of all adjacent lasers.
[0121] c. Alternatively, the effective elevation scan size of its acquisition laser may be equal to half of the absolute value of the difference in elevation angles of two adjacent lasers of its acquisition laser.
[0122] d. In one example, the elevation dimension covered by a node may be determined by the elevation angle of at least one key point of the node.
[0123] i. In one example, a predetermined point position of a node may be used as a key point position of the node, such as a midpoint position, a vertex position, an original point position, etc.
[0124] 1. Alternatively, the keypoint locations can be signaled from the encoder to the decoder.
[0125] ii. In one example, the elevation dimension covered by a node may be equal to the absolute value of the elevation difference between the midpoint of the upper surface and the midpoint of the lower surface of the node along the z-axis.
[0126] iii. Alternatively, the elevation angle size covered by the node may be equal to the absolute value of the difference between the maximum elevation angle and the minimum elevation angle of the eight vertices of the node.
[0127] iv. In one example, the function value of the elevation angle of a point can be used to represent its elevation angle.
[0128] 1. The function can be tangent, cotangent, sine, cosine, etc.
[0129] v. In one example, a function value of the elevation angle of a laser can be used to represent its elevation angle.
[0130] 1. The function can be tangent, cotangent, sine, cosine, etc.
[0131] e. In one example, the elevation dimension covered by a node may be replaced by the length of a particular line segment covered by the node.
[0132] i. In one example, the specific line segment may be a line segment between a midpoint of an upper surface and a midpoint of a lower surface of the node along the z-axis.
[0133] f. In one example, the effective elevation scan size of a node acquisition laser may be replaced by the length of a particular line segment covered by the effective elevation scan range of its acquisition laser.
[0134] i. In one example, the specific line segment may be a line segment that passes through the midpoint of the node and is parallel to a line segment between the midpoint of the upper surface and the midpoint of the lower surface of the node along the z-axis.
[0135] 3) Determine during the encoding / decoding process whether the node is constrained to be passed by only one laser beam in the azimuth direction.
[0136] a. In one example, whether a node is only passed by one laser beam in the azimuth direction may depend on the relationship between the azimuth size covered by the node and the effective azimuth scan size of its capturing laser beam.
[0137] i. In one example, a node may be considered to be traversed by only one laser beam in the azimuth direction if the azimuth dimension covered by the node is smaller than the effective azimuth scan dimension of its capturing laser beam.
[0138] ii. Alternatively, a node may be considered to be traversed by only one laser beam in the azimuthal direction if the azimuthal size covered by a node is less than or equal to the effective azimuthal scan size of its capturing laser beam.
[0139] b. In one example, the effective azimuth scan size of its capture laser beam can be determined by the scan parameters of its capture laser.
[0140] i. In one example, the effective azimuth scan size of its capture laser beam can be determined by the scan range and frequency of its capture laser.
[0141] c. In one example, the azimuth dimension covered by a node may be determined by the azimuth of at least one key point of the node.
[0142] i. In one example, a predetermined point position of a node may be used as a key point position of the node, such as a midpoint position, a vertex position, an original point position, etc.
[0143] 1. Alternatively, the keypoint locations can be signaled from the encoder to the decoder.
[0144] ii. In one example, the azimuthal dimension covered by a node may be equal to the absolute value of the difference in azimuthal angle between the midpoint of the upper surface and the midpoint of the lower surface of the node along the x-axis or y-axis.
[0145] iii. Alternatively, the azimuth size covered by a node may be equal to the absolute value of the difference between the maximum azimuth and the minimum azimuth of the eight vertices of the node.
[0146] d. In one example, the angular dimension covered by a node may be replaced by the length of a particular line segment covered by the node.
[0147] i. In one example, the specific line segment may be a line segment between a midpoint of an upper surface and a midpoint of a lower surface of a node along an x-axis or a y-axis.
[0148] e. In one example, the effective azimuth scan size of a node's capture laser beam may be replaced by the length of a particular line segment covered by the effective azimuth scan range of its capture laser beam.
[0149] i. In one example, the specific line segment may be a line segment that passes through the midpoint of the node and is parallel to a line segment between the midpoint of the upper surface and the midpoint of the lower surface of the node along the x-axis or the y-axis.
[0150] 4) If a node is only passed by one laser beam during the encoding / decoding process, a specific encoding / decoding mode can be applied.
[0151] a. In one example, if a node is only passed by one laser beam, the node can be signaled by the DCM.
[0152] b. In one example, if a node is only passed by one laser in the elevation direction, the node can be considered to be passed by only one laser beam.
[0153] c. Alternatively, if a node is only passed by one laser beam in the azimuth direction, the node can be considered to be passed by only one laser beam.
[0154] d. Alternatively, if a node is passed by only one laser beam in both the elevation direction and the azimuth direction, the node may be considered to be passed by only one laser beam.
[0155] e. In one example, the condition that a node is only passed by one laser beam may be the only condition for a particular codec mode.
[0156] f. Alternatively, the condition that a node is only passed by one laser beam can be combined with other conditions for a specific codec mode.
[0157] 5) An indicator (eg, a binary value) may be signaled to indicate whether angle information is used for the applicability condition of a particular codec mode.
[0158] a. In one example, if the value of the indicator is equal to X (e.g., X=1), the angle information will be used for the applicability condition of the specific codec mode. Otherwise (if it is equal to (1-X)), the angle information will not be used for the applicability condition of the specific codec mode.
[0159] b. In one example, the specific codec mode may be DCM.
[0160] c. In one example, the indication may be encoded and decoded using fixed length codec, unary codec, truncated unary codec, etc.
[0161] d. In one example, the indication may be encoded in a predictive manner.
[0162] 6) Line density can be used to help determine the applicability of angle information for a particular codec mode.
[0163] a. In one example, line density can be equal to the average number of spots along a laser beam.
[0164] b. Alternatively, the line density can be equal to the maximum number of spots along one laser beam.
[0165] c. In one example, the angle information is used for the applicability condition of a specific codec mode only when the line density is less than a density threshold.
[0166] d. Alternatively, the angle information is used for the applicability condition of a specific codec mode only when the line density is less than or equal to a density threshold.
[0167] e. In one example, the specific codec mode may be DCM.
[0168] f. In one example, line density may be used to derive a value for an indicator indicating whether angle information is used for a particular codec mode's applicability condition.
[0169] 7) The context for entropy coding and decoding of the plane position of at least one axis can be determined by several keypoint azimuths and predicted azimuths.
[0170] a. In one example, the axis can be the x-axis or the y-axis.
[0171] b. In one example, a key point may be a vertex, a center point, or a middle point of a side of a rectangle, where the rectangle is the projection of the node on the xy plane.
[0172] c. In one example, the number of keypoints may be two.
[0173] i. In one example, one of the key points can be the center of a rectangle that is the node in
[0174] Projection onto the xy-plane.
[0175] ii. In one example, one of the key points may be determined by comparing the absolute values of the x-coordinate and the y-coordinate.
[0176] 1. In one example, at least one of the x-coordinate value and the y-coordinate value may be a coordinate value of a center of a rectangle that is a projection of the node on the xy plane.
[0177] 2. In one example, if the absolute value of the x-coordinate is less than the absolute value of the y-coordinate, the key point may be the midpoint of the edge in the negative direction of the x-axis, otherwise, the key point may be the midpoint of the edge in the negative direction of the y-axis, otherwise.
[0178] 3. In one example, if the absolute value of the x-coordinate is less than the absolute value of the y-coordinate, the key point may be the midpoint of the edge in the negative direction of the y-axis, otherwise, the key point may be the midpoint of the edge in the negative direction of the x-axis, otherwise.
[0179] d. In one example, the context may be determined by the difference between the predicted azimuth and the keypoint azimuth.
[0180] e. In one example, the azimuth angle can be replaced by the value of the corresponding tangent angle.
[0181] 6. Examples
[0182] Figure 4 An example of a codec flow 400 for improving point cloud geometry codec using LIDAR features is depicted in FIG. At box 410, a line density is calculated. At box 420, a determination is made as to whether the line density is less than a density threshold. If it is determined at box 420 that the line density is greater than or equal to the density threshold, then at box 470, the current node may be signaled via a conventional codec mode. If it is determined at box 420 that the line density is less than the density threshold, then at box 430, a determination is made as to whether the current node is only passed by one laser beam in the elevation direction. If it is determined at box 430 that the current node is only passed by one laser beam in the elevation direction, then at box 460, the current node may be signaled by an octree. If it is determined at box 430 that the current node is only passed by one laser beam in the elevation direction, then at box 440, a determination is made as to whether the current node is only passed by one laser beam in the azimuth direction. If it is determined at block 440 that the current node is passed by more than one laser beam in the azimuth direction, the current node may be signaled by the octree at block 460. If it is determined at block 440 that the current node is passed by only one laser beam in the azimuth direction, the current node may be signaled by the DCM at block 450.
[0183] Embodiments of the present disclosure relate to motion information encoding and decoding for point cloud encoding and decoding. As used herein, the term "point cloud sequence" may refer to a sequence of one or more point clouds. The term "frame" may refer to a point cloud in a point cloud sequence. The term "point cloud" may refer to a frame in a point cloud sequence.
[0184] Figure 51 shows a flow chart of a method 500 for point cloud encoding and decoding according to some embodiments of the present disclosure. The method 500 may be implemented during the conversion between the current frame of the point cloud sequence and the bit stream of the point cloud sequence. Figure 5 As shown, method 500 begins at block 502, where a plurality of orientation angle values of a plurality of key points associated with nodes of a current frame are determined. The nodes represent a spatial segmentation of the current frame.
[0185] At block 504, a prediction of an azimuth angle magnitude of the node is determined. The prediction may include a predicted value of the azimuth angle or a predicted tangent value of the azimuth angle. At block 506, a context for entropy coding and decoding of a plane position of at least one axis associated with the node is determined based on a plurality of azimuth angle magnitudes and a prediction of the azimuth angle magnitudes. In some example embodiments, the context for entropy coding and decoding of a plane position may indicate a position above or below a plane. The plane may be an xy plane. By determining a context for entropy coding and decoding of a plane position based on a plurality of azimuth angle magnitudes and a prediction of the azimuth angle magnitudes, coding and decoding efficiency may be improved.
[0186] At block 508, conversion is performed based on the context of entropy coding and decoding. In some embodiments, conversion may include encoding the current frame into a bitstream. Alternatively or additionally, conversion may include decoding the current frame from a bitstream.
[0187] In some example embodiments, at least one axis includes an x-axis or a y-axis. In some example embodiments, at least one axis may include an x-axis and a y-axis.
[0188] In some example embodiments, the azimuth angle magnitude may include the value of the azimuth angle. Alternatively, in some example embodiments, the azimuth angle magnitude may include the tangent value of the azimuth angle. It should be understood that other metric values, such as the tangent value, sine value, or cosine value of the azimuth angle, may also be applied. The scope of the present application is not limited in this regard.
[0189] In some example embodiments, method 500 may further include determining a rectangular projection of the node on a plane associated with at least one axis. As an example, the plane associated with at least one axis may include an xy plane. Method 500 may further include determining a plurality of key points based on the rectangular projection.
[0190] In some example embodiments, the plurality of key points include at least one of: a vertex of a rectangular projection, a center point of a rectangular projection, or a midpoint of an edge of a rectangular projection. That is, the key point may be a vertex, a center point, or a midpoint of an edge of a rectangle, where the rectangle is a projection of the node on the xy plane.
[0191] In some example embodiments, a first coordinate value associated with a first axis of the node and a second coordinate value associated with a second axis of the node may be determined. A plurality of key points may be determined by comparing an absolute value of the first coordinate value and an absolute value of the second coordinate value.
[0192] In some example embodiments, at least one of the first coordinate value and the second coordinate value may be determined based on the coordinate value of the center point of the rectangular projection. In some example embodiments, the first axis includes an x-axis, and the second axis includes a y-axis. As an example, at least one of the x-coordinate value and the y-coordinate value may be a coordinate value of the center of the rectangular projection of the node on the xy plane.
[0193] In some example embodiments, if the first absolute value of the first coordinate value is less than the second absolute value of the second coordinate value, the midpoint of the first side of the rectangle projected in the negative direction along the first axis may be determined as one of the multiple key points. Alternatively or additionally, in some example embodiments, if the first absolute value of the first coordinate value is greater than or equal to the second absolute value of the second coordinate value, the midpoint of the second side of the rectangle projected in the negative direction along the second axis may be determined as one of the multiple key points.
[0194] In some example embodiments, if the first absolute value of the first coordinate value is less than the second absolute value of the second coordinate value, the midpoint of the second side of the rectangle projected in the negative direction along the second axis may be determined as one of the multiple key points. Alternatively or additionally, in some example embodiments, if the first absolute value of the first coordinate value is greater than or equal to the second absolute value of the second coordinate value, the midpoint of the first side of the rectangle projected in the negative direction along the first axis may be determined as one of the multiple key points.
[0195] In some example embodiments, the plurality of key points include two key points or more than two key points. As an example, one of the key points may be the center point of a rectangular projection of the node on the xy plane. One of the key points may be determined by comparing the absolute value of the x coordinate with the absolute value of the y coordinate.
[0196] In some example embodiments, a plurality of differences between a plurality of azimuth angle magnitudes for a plurality of key points and predictions of the azimuth angle magnitudes may be determined. At block 506, a context for entropy coding and decoding for the plane position may be determined based on the plurality of differences. For example, the context may be determined by the difference between the predicted azimuth and the azimuth of the key point.
[0197] According to an embodiment of the present disclosure, a non-transitory computer-readable recording medium is proposed. A bit stream of a point cloud sequence is stored in a non-transitory computer-readable recording medium. The bit stream of the point cloud sequence is generated by a method executed by a point cloud sequence processing device. According to the method, multiple azimuth angle values of multiple key points associated with a node of a current frame of the point cloud sequence are determined. The node represents a spatial domain segmentation of the current frame. A prediction of the azimuth angle value of the node is determined. Based on the multiple azimuth angle values and the prediction of the azimuth angle value, a context for entropy coding and decoding of a plane position of at least one axis associated with the node is determined. A bit stream is generated based on the entropy coding and decoding context.
[0198] According to an embodiment of the present disclosure, a method for storing a bitstream of a point cloud sequence is proposed. In the method, multiple azimuth angle values of multiple key points associated with a node of a current frame of the point cloud sequence are determined. The node represents a spatial segmentation of the current frame. A prediction of the azimuth angle value of the node is determined. Based on the multiple azimuth angle values and the prediction of the azimuth angle value, a context for entropy coding and decoding of a plane position of at least one axis associated with the node is determined. A bitstream is generated based on the entropy coding and decoding context. The bitstream is stored in a non-transitory computer-readable recording medium.
[0199] Implementations of the present disclosure may be described according to the following items, features of which may be combined in any reasonable manner.
[0200] Item 1. A method for point cloud encoding and decoding, comprising: during conversion between a current frame of a point cloud sequence and a bitstream of the point cloud sequence, determining a plurality of azimuth angle measurements of a plurality of key points associated with a node of the current frame, the node representing a spatial segmentation of the current frame; determining a prediction of the azimuth angle measurement of the node; determining a context for entropy encoding and decoding of a planar position of at least one axis associated with the node based on the plurality of azimuth angle measurements and the prediction of the azimuth angle measurement; and performing the conversion based on the entropy encoding and decoding context.
[0201] Item 2. A method according to Item 1, wherein the at least one axis includes an x-axis or a y-axis.
[0202] Item 3. A method according to Item 1 or Item 2, wherein the multiple key points include two key points.
[0203] Item 4. The method according to any one of Items 1 to 3 further includes: determining a rectangular projection of the node on a plane associated with the at least one axis; and determining the multiple key points based on the rectangular projection.
[0204] Item 5. A method according to Item 4, wherein the plane associated with the at least one axis comprises an xy plane.
[0205] Item 6. A method according to Item 4 or Item 5, wherein the multiple key points include at least one of the following: a vertex of the rectangular projection, a center point of the rectangular projection, or a midpoint of an edge of the rectangular projection.
[0206] Item 7. A method according to any one of Items 4 to 6, wherein determining the multiple key points includes: determining a first coordinate value associated with a first axis of the node and a second coordinate value associated with a second axis of the node; and determining the multiple key points by comparing the absolute value of the first coordinate value and the absolute value of the second coordinate value.
[0207] Item 8. A method according to Item 7, wherein determining the first coordinate value and the second coordinate value includes: determining at least one of the first coordinate value and the second coordinate value based on the coordinate value of the center point of the rectangular projection.
[0208] Item 9. A method according to Item 7 or Item 8, wherein determining the multiple key points by comparing the absolute value of the first coordinate value and the absolute value of the second coordinate value includes: if the first absolute value of the first coordinate value is less than the second absolute value of the second coordinate value, determining the midpoint of the first side of the rectangle projected in the negative direction along the first axis as one of the multiple key points; and if the first absolute value of the first coordinate value is greater than or equal to the second absolute value of the second coordinate value, determining the midpoint of the second side of the rectangle projected in the negative direction along the second axis as one of the multiple key points.
[0209] Item 10. A method according to Item 7 or Item 8, wherein determining the multiple key points by comparing the absolute value of the first coordinate value and the absolute value of the second coordinate value includes: if the first absolute value of the first coordinate value is less than the second absolute value of the second coordinate value, determining the midpoint of the second side of the rectangle projected in the negative direction along the second axis as one of the multiple key points; and if the first absolute value of the first coordinate value is greater than or equal to the second absolute value of the second coordinate value, determining the midpoint of the first side of the rectangle projected in the negative direction along the first axis as one of the multiple key points.
[0210] Item 11. A method according to any one of Items 7 to 10, wherein the first axis comprises an x-axis and the second axis comprises a y-axis.
[0211] Item 12. A method according to any one of Items 1 to 11, wherein determining the context for entropy coding and decoding for the planar position comprises: determining a plurality of differences between the plurality of azimuth angle values of the plurality of key points and the predictions of the azimuth angle values; and determining the context for entropy coding and decoding for the planar position based on the plurality of differences.
[0212] Item 13. A method according to any one of Items 1 to 12, wherein the azimuth angle magnitude comprises one of: the value of the azimuth angle, or the tangent value of the azimuth angle.
[0213] Item 14. A method according to any one of Items 1 to 13, wherein the converting comprises encoding the current frame into the bitstream.
[0214] Item 15. A method according to any one of Items 1 to 13, wherein the converting comprises decoding the current frame from the bitstream.
[0215] Item 16. An apparatus for processing point cloud data, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Items 1 to 15.
[0216] Item 17. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 15.
[0217] Item 18. A non-transitory computer-readable recording medium storing a bit stream of a point cloud sequence, the bit stream being generated by a method executed by a point cloud processing device, wherein the method comprises: determining a plurality of azimuth angle values of a plurality of key points associated with a node of a current frame of the point cloud sequence, the node representing a spatial segmentation of the current frame; determining a prediction of the azimuth angle value of the node; determining a context for entropy encoding and decoding of a planar position of at least one axis associated with the node based on the plurality of azimuth angle values and the prediction of the azimuth angle values; and generating the bit stream based on the context of entropy encoding and decoding.
[0218] Item 19. A method for storing a bitstream of a point cloud sequence, comprising: determining multiple azimuth angle measurements of multiple key points associated with a node of a current frame of the point cloud sequence, the node representing a spatial segmentation of the current frame; determining a prediction of the azimuth angle measurement of the node; determining a context for entropy coding and decoding of a planar position of at least one axis associated with the node based on the multiple azimuth angle measurements and the prediction of the azimuth angle measurement; generating the bitstream based on the context of entropy coding and decoding; and storing the bitstream in a non-transitory computer-readable recording medium.
[0219] Example Device
[0220] Figure 6 A block diagram of a computing device 600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 600 may be implemented as a source device 110 (or a GPCC encoder 116 or 200) or a destination device 120 (or a GPCC decoder 126 or 300), or may be included in a source device 110 (or a GPCC encoder 116 or 200) or a destination device 120 (or a GPCC decoder 126 or 300).
[0221] It should be understood that Figure 6 The computing device 600 shown in FIG. 6 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.
[0222] like Figure 6 As shown, computing device 600 comprises a general computing device 600. Computing device 600 may include at least one or more processors or processing units 610, memory 620, storage unit 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660.
[0223] In some embodiments, the computing device 600 can be implemented as any user terminal or server terminal with computing power. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination thereof, including the accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 600 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0224] The processing unit 610 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 620. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to increase the parallel processing capability of the computing device 600. The processing unit 610 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0225] The computing device 600 typically includes various computer storage media. Such media can be any media accessible by the computing device 600, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 620 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (such as a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 630 can be any removable or non-removable medium, and can include machine-readable media, such as a memory, a flash drive, a disk, or other media that can be used to store information and / or data and can be accessed in the computing device 600.
[0226] The computing device 600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 6 Although not shown in the figure, a disk drive for reading from and / or writing to a removable nonvolatile disk and an optical drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to the bus (not shown) via one or more data medium interfaces.
[0227] The communication unit 640 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 600 can be implemented by a single computing cluster or multiple computing machines, which can communicate via a communication connection. Therefore, the computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general network nodes.
[0228] The input device 650 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. The output device 660 may be one or more of various output devices, such as a display, a speaker, a printer, etc. With the aid of the communication unit 640, the computing device 600 may also communicate with one or more external devices (not shown), such as storage devices and display devices, and the computing device 600 may also communicate with one or more devices that enable a user to interact with the computing device 600, or, if necessary, the computing device 600 may also communicate with any device (e.g., a network card, a modem, etc.) that enables the computing device 600 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).
[0229] In some embodiments, some or all components of the computing device 600 may also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage services, which will not require the end user to know the physical location or configuration of the system or hardware that provides these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using a suitable protocol. For example, a cloud computing provider provides an application via a wide area network, which can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. The computing resources in a cloud computing environment may be merged or distributed at the location of a remote data center. Cloud computing infrastructure can provide services through a shared data center, although they appear as a single access point to the user. Therefore, the cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server, or may be installed on a client device directly or otherwise.
[0230] In an embodiment of the present disclosure, the computing device 600 may be used to implement point cloud encoding / decoding. The memory 620 may include one or more point cloud encoding / decoding modules 625 having one or more program instructions. These modules are accessible and executable by the processing unit 610 to perform the functions of the various embodiments described herein.
[0231] In an example embodiment of performing point cloud encoding, input device 650 may receive point cloud data as input to be encoded 670. The point cloud data may be processed, for example, by point cloud encoding and decoding module 625 to generate an encoded bitstream. The encoded bitstream may be provided as output 680 via output device 660.
[0232] In an example embodiment of performing point cloud decoding, the input device 650 may receive an encoded bitstream as input 670. The encoded bitstream may be processed, for example, by the point cloud codec module 625 to generate decoded point cloud data. The decoded point cloud data may be provided as output 680 via the output device 660.
[0233] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be appreciated by those skilled in the art that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These modifications are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for point cloud encoding and decoding, include: During conversion between a current frame of a point cloud sequence and a bitstream of the point cloud sequence, determining a plurality of orientation angle values of a plurality of key points associated with nodes of the current frame, the nodes representing a spatial segmentation of the current frame; determining a prediction of an angular magnitude of the node; determining a context for entropy coding of a plane position of at least one axis associated with the node based on the plurality of azimuth angle magnitudes and the prediction of the azimuth angle magnitudes; as well as The conversion is performed based on the context of the entropy coding and decoding. The method of claim 1 , wherein the at least one axis comprises an x-axis or a y-axis.
3. The method according to claim 1 or claim 2, wherein the plurality of key points comprises two key points.
4. The method according to any one of claims 1 to 3, further comprising: include: determining a rectangular projection of the node on a plane associated with the at least one axis; as well as Based on the rectangular projection, the plurality of key points are determined. The method of claim 4 , wherein the plane associated with the at least one axis comprises an xy plane.
6. The method according to claim 4 or claim 5, wherein the plurality of key points comprises at least one of the following: The vertices of the rectangular projection, the center point of the rectangular projection, or The midpoint of the side of the rectangle's projection.
7. The method according to any one of claims 4 to 6, wherein determining the plurality of key points include: determining a first coordinate value associated with a first axis of the node and a second coordinate value associated with a second axis of the node; as well as The plurality of key points are determined by comparing the absolute value of the first coordinate value with the absolute value of the second coordinate value.
8. The method according to claim 7, wherein determining the first coordinate value and the second coordinate value include: At least one of the first coordinate value and the second coordinate value is determined based on the coordinate value of the center point of the rectangular projection.
9. The method according to claim 7 or claim 8, wherein the plurality of key points are determined by comparing the absolute value of the first coordinate value with the absolute value of the second coordinate value. include: If a first absolute value of the first coordinate value is less than a second absolute value of the second coordinate value, determining a midpoint of a first side of the rectangle projected in the negative direction along the first axis as one of the multiple key points; as well as If the first absolute value of the first coordinate value is greater than or equal to the second absolute value of the second coordinate value, a midpoint of a second side of the rectangle projected in the negative direction along the second axis is determined as one of the multiple key points.
10. The method according to claim 7 or claim 8, wherein the plurality of key points are determined by comparing the absolute value of the first coordinate value with the absolute value of the second coordinate value. include: If a first absolute value of the first coordinate value is less than a second absolute value of the second coordinate value, determining a midpoint of a second side of the rectangle projected in the negative direction along the second axis as one of the multiple key points; as well as If the first absolute value of the first coordinate value is greater than or equal to the second absolute value of the second coordinate value, a midpoint of a first side of the rectangle projected in the negative direction along the first axis is determined as one of the multiple key points.
11. The method of any one of claims 7 to 10, wherein the first axis comprises an x-axis and the second axis comprises a y-axis.
12. The method according to any one of claims 1 to 11, wherein the context for entropy coding and decoding of the plane position is determined include: determining a plurality of differences between the plurality of azimuth angle magnitudes for the plurality of key points and the predictions of azimuth angle magnitudes; as well as Based on the plurality of differences, the context for entropy coding of the plane position is determined.
13. The method according to any one of claims 1 to 12, wherein the azimuth angle value comprises one of the following: the value of the azimuth, or The tangent of the azimuth angle.
14. The method according to any one of claims 1 to 13, wherein the converting comprises encoding the current frame into the bitstream.
15. The method according to any one of claims 1 to 13, wherein the converting comprises decoding the current frame from the bitstream.
16. An apparatus for processing point cloud data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 15. 17 . A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to execute the method according to any one of claims 1 to 15.
18. A non-transitory computer-readable recording medium storing a bit stream of a point cloud sequence, the bit stream being generated by a method executed by a point cloud processing device, wherein the method include: determining a plurality of orientation angle values of a plurality of key points associated with a node of a current frame of the point cloud sequence, the node representing a spatial segmentation of the current frame; determining a prediction of an angular magnitude of the node; determining a context for entropy coding of a plane position of at least one axis associated with the node based on the plurality of azimuth angle magnitudes and the prediction of the azimuth angle magnitudes; as well as The bitstream is generated based on the entropy-encoded context.
19. A method for storing a bit stream of a point cloud sequence, include: determining a plurality of orientation angle values of a plurality of key points associated with a node of a current frame of the point cloud sequence, the node representing a spatial segmentation of the current frame; determining a prediction of an angular magnitude of the node; determining a context for entropy coding of a plane position of at least one axis associated with the node based on the plurality of azimuth angle magnitudes and the prediction of the azimuth angle magnitudes; Generate the bitstream based on the context of entropy coding and decoding; as well as The bit stream is stored in a non-transitory computer-readable recording medium.