Method, apparatus, and medium for point cloud coding
The method addresses inefficiencies in conventional point cloud coding by optimizing nearest neighbor search across multiple LODs and using motion compensation, resulting in improved accuracy and efficiency for point cloud coding techniques.
Patent Information
- Application Number
- JP2024531412
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-26
- Filing Date
- 2022-11-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-25
AI Technical Summary
Conventional point cloud coding techniques suffer from inefficiencies in accuracy and complexity in nearest neighbor search and attribute interpolation prediction, particularly in geometry-based point cloud compression (G-PCC), due to limitations in search centers, search ranges, and geometric distances used in inter-prediction methods.
The proposed method improves nearest neighbor search accuracy and efficiency by allowing search within multiple levels of detail (LODs) of a reference point cloud sample, using different geometric distances and motion compensation, and storing search results in multiple lists for attribute interpolation prediction.
Enhances the accuracy and efficiency of point cloud coding by improving nearest neighbor search and attribute interpolation prediction, reducing computational complexity and enhancing coding performance.
Smart Images

Figure 0007708511000005 
Figure 0007708511000006 
Figure 0007708511000007
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to point cloud coding techniques, and more particularly, to optimized inter-prediction for point cloud attribute coding based on nearest neighbor search.
Background Art
[0002] A point cloud is a set of individual data points in a three-dimensional (3D) plane, where each point has set coordinates on the X, Y, and Z axes. Thus, point clouds can be used to represent the physical content of 3D space. Point clouds have been shown to be a promising way to represent 3D visual data for a wide range of immersive applications, from augmented reality to autonomous vehicles.
[0003] The point cloud coding standard has mainly evolved through the development of the well-known MPEG organization. MPEG, short for Moving Picture Experts Group, is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Coding Group (3DG) issued a Call for Proposals (CFP) document to initiate the development of a point cloud coding standard. The final standard consists of two classes of solutions. Video-based point cloud compression (V-PCC or VPCC) is suitable for point sets where points are relatively uniformly distributed. Geometry-based point cloud compression (G-PCC or GPCC) is suitable for sparser distributions. However, the coding efficiency of conventional point cloud coding techniques is generally expected to be further improved.
Summary of the Invention
[0004] Embodiments of the present disclosure provide a solution for point cloud coding.
[0005] In a first aspect, a method for point cloud coding is proposed. The method includes, during conversion between a current point cloud (PC) sample and a bitstream of a point cloud sequence, determining, for a current point in the current PC sample of the point cloud sequence, at least one neighboring point from a set of points in a reference PC sample of the current PC sample, wherein the set of points is in a group of levels of detail (LOD) of the reference PC sample, and performing the conversion based on the at least one neighboring point.
[0006] Based on the method according to the first aspect of the present disclosure, points within a group of LODs of a reference PC sample are searched to obtain neighboring points of a current point. Compared with a conventional solution that only searches for points at the same LOD as the current point, the proposed method can advantageously improve the accuracy of nearest neighbor search and attribute interpolation prediction.
[0007] In a second aspect, another method for point cloud coding is proposed. The method includes, during conversion between a current point cloud (PC) sample and a bitstream of a point cloud sequence, determining, for a current point in the current PC sample of the point cloud sequence, a first plurality of neighboring points from points in a second plurality of PC samples of the point cloud sequence, and performing the conversion based on the first plurality of neighboring points, wherein the first plurality of neighboring points are stored in a third plurality of lists.
[0008] Based on the method according to the second aspect of the present disclosure, a plurality of neighboring points of a current point are stored in a plurality of lists. Compared with a conventional solution that stores neighboring points in only one list, the proposed method can advantageously improve the efficiency of nearest neighbor search and attribute interpolation prediction.
[0009] In a third aspect, an apparatus for processing point cloud data is proposed. The apparatus for processing point cloud data includes a processor and a non-transitory memory storing instructions. When executed by the processor, the instructions cause the processor to perform the method according to the first or second aspect of the present disclosure.
[0010] In a fourth aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first or second aspect of the present disclosure.
[0011] In a fifth aspect, a non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a point cloud sequence generated by a method executed by a point cloud processing device. The method includes determining, for a current point in a current point cloud (PC) sample of the point cloud sequence, at least one neighboring point from a set of points in a reference PC sample of the current PC sample, the set of points being in a group of detail levels (LODs) of the reference PC sample, and generating the bitstream based on the at least one neighboring point.
[0012] In a sixth aspect, a method of storing a bitstream of a point cloud sequence is proposed. The method includes determining, for a current point in a current point cloud (PC) sample of the point cloud sequence, at least one neighboring point from a set of points in a reference PC sample of the current PC sample, the set of points being in a group of detail levels (LODs) of the reference PC sample, generating the bitstream based on the at least one neighboring point, and storing the bitstream in a non-transitory computer-readable recording medium.
[0013] In a seventh aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a point cloud sequence generated by a method executed by a point cloud processing device. The method includes determining, for a current point in a current point cloud (PC) sample of the point cloud sequence, a first plurality of neighboring points from points in a second plurality of PC samples of the point cloud sequence, and generating the bitstream based on the first plurality of neighboring points, where the first plurality of neighboring points are stored in a third plurality of lists.
[0014] In an eighth aspect, a method for storing a bitstream of a point cloud sequence is proposed. The method includes determining, for a current point in a current point cloud (PC) sample of the point cloud sequence, a first plurality of neighboring points from points in a second plurality of PC samples of the point cloud sequence, generating the bitstream based on the first plurality of neighboring points, and storing the bitstream in a non-transitory computer-readable recording medium, where the first plurality of neighboring points are stored in a third plurality of lists.
[0015] The content of this invention is provided to introduce, in a simplified form, a selection of concepts that will be further described in the following detailed description. The content of this invention is not intended to identify the main features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
Brief Description of the Drawings
[0016] The above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, usually, the same reference numerals refer to the same components.
[0017]
Figure 1
[0018]
Figure 2
[0019]
Figure 3
[0020]
Figure 4
[0021]
Figure 5
[0022]
Figure 6
[0023] Throughout the drawings, the same or similar reference numerals generally refer to the same or similar elements.
DETAILED DESCRIPTION OF THE INVENTION
[0024] Next, the principles of the present disclosure will be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and are intended to assist those skilled in the art in understanding and implementing the present disclosure, and do not imply any limitation with respect to the scope of the present disclosure. The disclosure described herein can be implemented in various ways other than the methods described below.
[0025] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0026] References to "one embodiment", "an embodiment", "exemplary embodiment", etc. in this disclosure indicate that the described embodiment may include a particular feature, structure, or characteristic, but not necessarily all embodiments include the particular feature, structure, or characteristic. Also, such phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in relation to an exemplary embodiment, it is pointed out that it is within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in relation to other embodiments, whether or not explicitly described.
[0027] Terms such as "first" and "second" may be used herein to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiment, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0028] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an", and "the" are to be construed to include the plural forms as well, unless the context clearly dictates otherwise. It will be further understood that the terms "comprises", "comprising", "includes", "including", "has", "having", and / or "contains", when used herein, specify the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0029] [Exemplary Environment] FIG. 1 is a block diagram showing an exemplary point cloud coding system 100 that can utilize the technology of the present disclosure. As shown, the point cloud coding system 100 may include a source device 110 and a destination device 120. The source device 110 is also called a point cloud coding device, and the destination device 120 may also be called a point cloud decoding device. During operation, the source device 110 may be configured to generate encoded point cloud data, and the destination device 120 may be configured to decode the encoded point cloud data generated by the source device 110. The technology of the present disclosure generally targets the coding (encoding and / or decoding) of point cloud data, i.e., the support for point cloud compression. The coding may be effective for compressing and / or decompressing point cloud data.
[0030] The source device 100 and the destination device 120 may include any of a wide range of devices including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephones such as smartphones and mobile phones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, vehicles (e.g., land vehicles or marine vessels, spacecraft, aircraft, etc.), robots, LIDAR devices, satellites, augmented reality devices, etc. In some cases, the source device 100 and the destination device 120 may be equipped for wireless communication.
[0031] The source device 100 may include a data source 112, a memory 114, a GPCC encoder 116, and an input / output (I / O) interface 118. The destination device 120 may include an input / output (I / O) interface 128, a GPCC decoder 126, a memory 124, and a data consumer 122. According to the present disclosure, the GPCC encoder 116 of the source device 100 and the GPCC decoder 126 of the destination device 120 may be configured to apply the techniques of the present disclosure related to point cloud coding. Thus, the source device 100 represents an example of an encoding device, and the destination device 120 represents an example of a decoding device. In other examples, the source device 100 and the destination device 120 may include other components or arrangements. For example, the source device 100 may receive data (e.g., point cloud data) from an internal or external source. Similarly, the destination device 120 may interface with an external data consumer rather than including a data consumer in the same device.
[0032] Generally, data source 112 represents a source of point cloud data (i.e., raw, unencoded point cloud data), and can provide a continuous series of "frames" of point cloud data to a GPCC encoder 116 that encodes the point cloud data of the frames. In some examples, data source 112 generates point cloud data. The data source 112 of source device 100 can include any of a variety of cameras or sensors, such as one or more video cameras, an archive containing previously captured point cloud data, a 3D scanner or a light detection and ranging (LIDAR) device, etc., which are point cloud capture devices, and / or a data feed interface that receives point cloud data from a data content provider. Thus, in some examples, data source 112 can generate point cloud data based on signals from a LIDAR device. Alternatively or additionally, the point cloud data can be computer-generated from scanners, cameras, sensors or other data. For example, data source 112 can generate point cloud data or a combination of live point cloud data, archived point cloud data, and computer-generated point cloud data. In any case, GPCC encoder 116 encodes the captured, pre-captured, or computer-generated point cloud data. GPCC encoder 116 can rearrange the frames of point cloud data from the order received (which may also be called the "display order") to a coding order for coding. GPCC encoder 116 can generate one or more bitstreams containing the encoded point cloud data. Next, source device 100 can output the encoded point cloud data via I / O interface 118 for reception and / or retrieval, for example, by I / O interface 128 of destination device 120. The encoded point cloud data can be transmitted directly to destination device 120 through network 130A via I / O interface 118. The encoded point cloud data may be stored in storage medium / server 130B for access by destination device 120.
[0033] The memory 114 of the source device 100 and the memory 124 of the destination device 120 may represent general-purpose memory. In some examples, the memory 114 and the memory 124 may store raw point cloud data, e.g., raw point cloud data from the data source 112, and raw decoded point cloud data from the GPCC decoder 126. Additionally or alternatively, the memory 114 and the memory 124 may store software instructions executable, e.g., by the GPCC encoder 116 and the GPCC decoder 126 respectively. In this example, although the memory 114 and the memory 124 are shown separately from the GPCC encoder 116 and the GPCC decoder 126, it should be understood that the GPCC encoder 116 and the GPCC decoder 126 may also include internal memory for functionally similar or equivalent purposes. Further, the memory 114 and the memory 124 may store, e.g., encoded point cloud data output from the GPCC encoder 116 and input to the GPCC decoder 126. In some examples, a portion of the memory 114 and the memory 124 may be allocated as one or more buffers to store raw, decoded, and / or encoded point cloud data. For example, the memory 114 and the memory 124 may store point cloud data.
[0034] I / O interfaces 118 and 128 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of various IEEE 802.11 standards, or other physical components. In examples where I / O interfaces 118 and 128 include wireless components, I / O interfaces 118 and 128 may be configured to transfer data such as encoded point cloud data according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), LTE-Advanced, 5G, etc. In some examples where I / O interface 118 includes a wireless transmitter, I / O interfaces 118 and 128 may be configured to transfer data such as encoded point cloud data according to other wireless standards such as the IEEE 802.11 standard. In some examples, source device 100 and / or destination device 120 may each include a respective system-on-chip (SoC) device. For example, source device 100 may include an SoC device that performs functions attributable to GPCC encoder 116 and / or I / O interface 118, and destination device 120 may include an SoC device that performs functions attributable to GPCC decoder 126 and / or I / O interface 128.
[0035] The techniques of the present disclosure may be applied to encoding and decoding that supports any of various applications such as communication between autonomous vehicles, communication between processing device devices such as scanners, cameras, sensors, and local or remote servers, geographic mapping, or other applications.
[0036] The I / O interface 128 of the destination device 120 receives the encoded bitstream from the source device 110. The encoded bitstream may include signaling information defined by the GPCC encoder 116, which is also used by the GPCC decoder 126, such as syntax elements having values representing point clouds. The data consumer 122 uses the decoded data. For example, the data consumer 122 may use the decoded point cloud data to determine the position of a physical object. In some examples, the data consumer 122 may include a display that presents an image based on the point cloud data.
[0037] The GPCC encoder 116 and the GPCC decoder 126 may each be embodied as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or combinations thereof. Where the technology is embodied partially in software, the device may store software instructions on a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the techniques of this disclosure. Each of the GPCC encoder 116 and the GPCC decoder 126 may be included in one or more encoders or decoders, but any of them may be integrated as part of a combined encoder / decoder (CODEC) in their respective devices. Devices including the GPCC encoder 116 and / or the GPCC decoder 126 may include one or more integrated circuits, microprocessors, and / or other types of devices.
[0038] The GPCC encoder 116 and the GPCC decoder 126 can operate according to a coding standard such as the Video Point Cloud Compression (VPCC) standard or the Geometry Point Cloud Compression (GPCC) standard. The present disclosure can generally refer to the coding (e.g., encoding and decoding) of a frame including a process of encoding or decoding data. The encoded bitstream generally includes a series of values of syntax elements representing coding decisions (e.g., coding modes).
[0039] A point cloud can include a set of points in 3D space and can have attributes associated with those points. The attributes can be color information such as R, G, B, Y, Cb, Cr, or reflectance information, or other attributes. The point cloud may be captured by various cameras and sensors such as LIDAR sensors and 3D scanners, or may be generated by a computer. Point cloud data is used in various applications including, but not limited to, construction (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors used for navigation).
[0040] FIG. 2 is a block diagram showing an example of a GPCC encoder 200 that can be an example of the GPCC encoder 116 in the system 100 shown in FIG. 1 according to some embodiments of the present disclosure. FIG. 3 is a block diagram showing an example of a GPCC decoder 300 that can be an example of the GPCC decoder 126 in the system 100 shown in FIG. 1 according to some embodiments of the present disclosure.
[0041] In both the GPCC encoder 200 and the GPCC decoder 300, the point cloud positions are first encoded. The attribute coding depends on the encoded geometry. In FIGS. 2 and 3, the region adaptive hierarchical transform (RAHT) unit 218, the surface approximation analysis unit 212, the RAHT unit 314, and the surface approximation synthesis unit 310 are options typically used for category 1 data. The level of detail (LOD) generation unit 220, the lifting unit 222, the LOD generation unit 316, and the inverse lifting unit 318 are options typically used for category 3 data. All other units are common to categories 1 and 3.
[0042] For category 3 data, the compressed geometry is typically represented as an octree that goes from the root to the leaf level of individual voxels. For category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root to the leaf level of blocks larger than voxels) and a model that approximates the surface within each leaf of the pruned octree. In this way, both category 1 data and category 3 data share the octree coding mechanism, but category 1 data can further approximate the voxels within each leaf with a surface model. The surface model used is a triangulation with 1 to 10 triangles per block, resulting in a triangle soup. Thus, the category 1 geometry codec is known as the Trisoup geometry codec, and the category 3 geometry codec is known as the octree geometry codec.
[0043] In the example of FIG. 2, the GPCC encoder 200 may include a coordinate conversion unit 202, a color conversion unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometry reconstruction unit 216, a RAHT unit 218, a LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.
[0044] As shown in the example of FIG. 2, the GPCC encoder 200 may receive a set of positions and a set of attributes. The positions may include the coordinates of the points within the point cloud. The attributes may include information regarding the points within the point cloud, such as the color associated with the points within the point cloud.
[0045] The coordinate conversion unit 202 may apply a conversion to the coordinates of the points to convert the coordinates from an initial region to a conversion region. In the present disclosure, the converted coordinates may be referred to as conversion coordinates. The color conversion unit 204 may apply a conversion for converting the color information of the attributes to a different domain. For example, the color conversion unit 204 may convert the color information from an RGB color space to a YCbCr color space.
[0046] Furthermore, in the example of FIG. 2, the voxelization unit 206 may voxelize the conversion coordinates. The voxelization of the conversion coordinates may include quantizing and deleting some of the points of the point cloud. In other words, a plurality of points of the point cloud may be included within a single "voxel" and may then be treated as one point in some respects. Furthermore, the octree analysis unit 210 may generate an octree based on the voxelized conversion coordinates. Additionally, in the example of FIG. 2, the surface approximation analysis unit 212 may analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 may perform arithmetic coding on the syntax elements representing the octree and / or surface information determined by the surface approximation analysis unit 212. The GPCC encoder 200 may output these syntax elements in a geometry bitstream.
[0047] The geometry reconstruction unit 216 can reconstruct the transformed coordinates of the points in the point cloud based on the octree, the data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. The number of transformed coordinates reconstructed by the geometry reconstruction unit 216 can be different from the number of original points in the point cloud for voxelization and surface approximation. In the present disclosure, the resulting points can be referred to as reconstructed points. The attribute transfer unit 208 can transfer the attributes of the original points in the point cloud to the reconstructed points of the point cloud data.
[0048] Furthermore, the RAHT unit 218 can apply RAHT coding to the attributes of the reconstructed points. Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 can apply LOD processing and lifting to the attributes of the reconstructed points, respectively. The RAHT unit 218 and the lifting unit 222 can generate coefficients based on the attributes. The coefficient quantization unit 224 can quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 can apply arithmetic coding to the syntax elements representing the quantized coefficients. The GPCC encoder 200 can output these syntax elements in the attribute bitstream.
[0049] In the example of FIG. 3, the GPCC decoder 300 can include a geometry arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometry reconstruction unit 312, a RAHT unit 314, an LOD generation unit 316, an inverse lifting unit 318, a coordinate inverse transformation unit 320, and a color inverse transformation unit 322.
[0050] The GPCC decoder 300 can acquire a geometry bitstream and an attribute bitstream. The geometry arithmetic decoding unit 302 of the decoder 300 can apply arithmetic decoding (e.g., CABAC or other types of arithmetic decoding) to syntax elements in the geometry bitstream. Similarly, the attribute arithmetic decoding unit 304 can apply arithmetic decoding to syntax elements in the attribute bitstream.
[0051] The octree synthesis unit 306 can synthesize an octree based on syntax elements parsed from the geometry bitstream. When surface approximation is used in the geometry bitstream, the surface approximation synthesis unit 310 can determine a surface model based on the syntax elements parsed from the geometry bitstream and the octree.
[0052] Furthermore, the geometry reconstruction unit 312 can perform reconstruction to determine the coordinates of points in the point cloud. The coordinate inverse transformation unit 320 can apply an inverse transformation to the reconstructed coordinates to convert the reconstructed coordinates (positions) of points in the point cloud back from the transformed region to the initial region.
[0053] Additionally, in the example of FIG. 3, the inverse quantization unit 308 can inverse-quantize the attribute values. The attribute values may be based on syntax elements obtained from the attribute bitstream (e.g., including syntax elements decoded by the attribute arithmetic decoding unit 304).
[0054] Depending on how the attribute values are encoded, the RAHT unit 314 can perform RAHT coding to determine the color values of the points in the point cloud based on the inverse-quantized attribute values. Alternatively, the LOD generation unit 316 and the inverse lifting unit 318 can use a level-of-detail-based technique to determine the color values of the points in the point cloud.
[0055] Furthermore, in the example of FIG. 3, the color inverse conversion unit 322 may apply an inverse color conversion to the color values. The inverse color conversion may be the inverse of the color conversion applied by the color conversion unit 204 of the encoder 200. For example, the color conversion unit 204 may convert color information from the RGB color space to the YCbCr color space. Accordingly, the color inverse conversion unit 322 may convert the color information from the YCbCr color space to the RGB color space.
[0056] The various units in FIGS. 2 and 3 are shown to assist in understanding the operations performed by the encoder 200 and the decoder 300. The units may be implemented as fixed function circuits, programmable circuits, or combinations thereof. A fixed function circuit refers to a circuit that provides a specific function and is pre-set with respect to executable operations. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides a flexible function with respect to executable operations. For example, a programmable circuit may execute software or firmware that operates the programmable circuit in a manner defined by software or firmware instructions. A fixed function circuit may execute software instructions (e.g., receive or output parameters), but the type of operations performed by the fixed function circuit is generally immutable. In some examples, one or more units may be separate circuit blocks (fixed function or programmable), and in some examples, one or more units may be integrated circuits.
[0057] Some exemplary embodiments of the present disclosure will be described in detail below. It should be understood that section headings are used herein for ease of understanding, and the embodiments disclosed in a section are not limited to that section only. Furthermore, although specific embodiments are described with reference to GPCC or other specific point cloud coders, the disclosed technology is applicable to other point cloud coding techniques as well. Additionally, although some embodiments describe the point cloud encoding steps in detail, it will be understood that the corresponding decoding steps for reversing the coding are implemented by the decoder. [1. Overview] This disclosure relates to point cloud coding techniques. Specifically, it relates to point cloud attribute prediction in inter prediction. This idea can be applied to point cloud coding standards or non-standard point cloud coders (e.g., the developing Geometry-based Point Cloud Compression (G-PCC)) individually or in various combinations. [2. Abbreviations] G-PCC: Geometry based Point Cloud Compression (Geometry-based Point Cloud Compression) MPEG: Moving Picture Experts Group (Moving Picture Experts Group) 3DG: 3D Graphics Coding Group (3D Graphics Coding Group) CFP: Call For Proposal (Call For Proposal) V-PCC: Video-based Point Cloud Compression (Video-based Point Cloud Compression) LOD: Level of Detail (Level of Detail) CE: Core Experiment (Core Experiment) EE: Exploration Experiment (Exploration Experiment) inter-EM: Inter Exploration Model (Inter Exploration Model) PC: Point Cloud (Point Cloud) RDO: Rate-distortion Optimization (Rate-distortion Optimization) [3. Background] The point cloud coding standard has mainly evolved through the development of the well-known MPEG organization. MPEG is short for Moving Picture Experts Group and is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Coding Group (3DG) issued a Call for Proposals (CFP) document to initiate the development of the point cloud coding standard. The final standard consists of two classes of solutions. Video-based Point Cloud Compression (V-PCC) is suitable for point sets where points are relatively uniformly distributed. Geometry-based Point Cloud Compression (G-PCC) is suitable for sparser distributions. To explore future point cloud coding techniques in G-PCC, Core Experiment (CE) 13.5 and Exploration Experiment (EE) 13.2 were established to develop inter prediction techniques in G-PCC. Since then, many new inter prediction methods have been adopted by MPEG and incorporated into the reference software named inter Exploration Model (inter-EM). In a point cloud, there can be geometry information and attribute information. Geometry information is used to describe the geometric positions of data points. Attribute information is used to record details of data points such as texture, normal vectors, and reflections. The point cloud codec can process various information in various ways. Usually, the codec has many optional tools that support the coding and decoding of geometry information and attribute information respectively. [3.1 Attribute Intra Prediction] In G-PCC, two attribute coding methods for performing attribute intra prediction using geometry information have been proposed. [3.1.1 Prediction of Transform] Transform prediction is an interpolation-based hierarchical nearest neighbor prediction method and is usually used for sparse point cloud content. First, a Level of Detail (LOD) structure is generated. Next, the nearest neighbor is searched based on the LOD structure. Then, attribute prediction is performed based on the search result. [3.1.1.1 Generation of LOD] In the LOD generation process, the geometry information is utilized to construct a hierarchical structure of point clouds that defines a set of "levels of detail". The hierarchical structure is used to efficiently predict attributes. It also becomes possible to provide advanced functions such as progressive transmission and scalable rendering. In the LOD generation process, according to the user-defined parameter L indicating the LOD number, the point cloud points are reorganized into refinement levels (point sets) R0, R1, …, R L-1 Next, the attributes of the point cloud points are encoded from R0 to R L-1 up to. The level of detail l, LOD l can be obtained by taking the union of the refinement levels R0, R1, …, R l : [Number] [3.1.1.2 Nearest Neighbor Search Considering Point Distribution] In G-PCC, two neighbor lists, list1 and list2, are constructed to search for three approximate nearest neighbors of the current point. List1 contains three approximate nearest neighbors obtained by a LOD-based approximate nearest neighbor search algorithm. List2 contains the three points that dropped out when list1 was updated. Considering the point distribution information, the concepts of strict opposition and loose opposition are defined. According to the relative position to the current point (x, y, z), all the nearest neighbor points (x n , y n , z n ) are assigned to the direction index dirIdx. According to the direction index dirIdx, strict opposition and loose opposition are defined as shown in Table 3-1. [Table 1] It should be noted that since there are not enough neighbors generated by updating list1 using the points in list2 with strict opposition eligibility check and loose opposition eligibility check, and the neighbor pruning process is executed, the number of points in the final list1 can be less than 3. [3.1.1.3 Prediction of Attributes] After obtaining list1, a plurality of predictor candidates are created based on list1. Each predictor candidate is assigned one index. Next, the variation of the attributes of the points in list1 is calculated. If the variation is smaller than the threshold, the attribute of the current point is predicted using the weighted average value. Otherwise, the best predictor is selected by applying the rate distortion optimization (RDO) procedure. [3.1.2 Lifting Transform] The lifting transform is usually used for dense point cloud content and is built on top of the prediction transform method. The main differences between the lifting transform and the prediction transform are the update operator and the adaptive quantization strategy. In the lifting transform, each point is associated with an influence weight value. Points with a low level of detail (LOD) are used more frequently and are assigned higher weight values. The influence weights are used in the quantization process. [3.2 Attribute Inter-Prediction] In inter-EM, several inter-prediction tools have been proposed to perform attribute inter-coding. There is one list1 that stores the nearest neighbors in the current frame and the previous one frame. The attributes of the points in the list are used to generate predictor candidates and obtain the predicted value of the current point in the same way as in intra-frame coding. First, the points in the current frame and the reference frame are sorted based on the Morton code. Each point is associated with one Morton index indicating the Morton order. Next, for each point, a nearest neighbor search is performed in the current frame and the reference frame. There is one parameter Search_Range that controls the search range. a) In the current frame, the previous Search_Range points of the current point in Morton order are traversed. b) In the reference frame, the search center is the point with the same Morton index in the reference frame. The previous Search_Range points before the search center, the next Search_Range points after the search center, and the search center point are searched. The nearest neighbor search is based on the Euclidean distance from the searched point to the current point. Three nearest neighbor points are selected and stored in list1. It should be noted that the weight of the points from the reference frame should be lower than that of the points from the current frame. Finally, a plurality of predictor candidates are created based on list1, and predicted attribute values are generated in the same manner as in intra-coding. [4. Problems] The existing design of point cloud attribute inter-prediction has the following problems. 1. In the current inter-EM, the search center in the reference frame is the point with the same Morton index. However, there is no strict correspondence between the permuted points of the current frame and the reference frame. In some cases, the geometric positions of the points with the same Morton index may be very different, resulting in inaccurate search results and prediction results. 2. In the current inter-EM, the nearest neighbor search is performed based on the Euclidean distance. The calculation of the Euclidean distance is very complex and affects the overall complexity of encoding and decoding. 3. In the current inter-EM, the search ranges of the current frame and the reference frame are the same. However, the points in the current frame and the points in the difference frame have different effects on the current point. Using the same search range limits the prediction efficiency. [5. Detailed Solutions] To solve the above problems and several other problems not mentioned, a method as summarized below is disclosed. The solutions should be considered as examples for explaining general concepts and should not be interpreted narrowly. Furthermore, these solutions can be applied individually or in any combination. In the following description, list1 can be a list for storing the nearest neighbors. 1) It is proposed that at least one search center can be derived for the nearest neighbor search in attribute inter-prediction. a. In one example, the nearest neighbor search can be performed within a given frame to be searched. i. In one example, a given frame to be searched may be the current frame. ii. In one example, a given frame to be searched may be another frame. iii. In one example, a given frame to be searched may be a reference frame of the current frame. b. In one example, the points to be searched within a given frame may be reordered prior to the nearest neighbor search. i. In one example, the reordering may be performed based on the Morton code, Hilbert code, or other conversion code of the points. ii. In one example, the reordering may be performed based on the polar coordinates of the points. iii. In one example, the reordering may be performed based on the spherical coordinates of the points. iv. In one example, the reordering may be performed based on the cylindrical coordinates of the points. v. In one example, the reordering may be performed based on the scanning order of the radar. vi. In one example, the search is performed according to the order (which may be reordered) of the points. c. In one example, there may be one search center for a given frame to be searched. i. In one example, the previous points before the search center and the search center in the reordered order may be searched. ii. In one example, the next points after the search center and the search center in the reordered order may be searched. iii. In one example, the previous points before the search center, the search center, and the next points after the search center in the reordered order may be searched. d. In one example, the search center may be an approximate nearest neighbor point at a geometric location within the frame to be searched. i. In one example, the search center may be selected from all or some of the points within the frame to be searched. ii. In one example, the search center may be the point having the closest distance from the current point. (1) The distance may be Euclidean distance, Manhattan distance, Chebyshev distance, etc. iii. In one example, the search center can be a point having an approximate nearest distance from the current point. (1) The search center can be selected from partial points within the frame to be searched. (2) The distance can be a Euclidean distance, a Manhattan distance, a Chebyshev distance, etc. iv. In one example, the search center can be a point having the nearest distance from the current point on the conversion code. (1) The distance can be the difference of the conversion codes. (2) The conversion code can be a Morton code, a Hilbert code, etc. v. In one example, the search center can be a point having an approximate nearest distance from the current point on the conversion code. (1) The search center can be selected from partial points within the frame to be searched. (2) In one example, the partial points can be points whose conversion code is greater than the conversion code of the current point. (3) In one example, the partial points can be points whose conversion code is smaller than the conversion code of the current point. (4) The distance can be the difference of the conversion codes. (5) The conversion code can be a Morton code, a Hilbert code, etc. e. In one example, a plurality (such as N) of search centers can be derived. i. Alternatively, further, the search can be performed from one or more search centers. ii. In one example, the N search centers can be N points having N nearest distances from the current point. iii. In one example, the N search centers can be N points having N nearest distances from the current point on the conversion code. 2) It is proposed to use different search ranges in different directions and / or different frames. a. In one example, there can be at least one search range for one frame to be searched. i. In one example, there can be one search range for indicating the number of points before the search center that needs to be searched. ii. In one example, there may be one search range for indicating the number of points after the search center to be searched. iii. In one example, there may be one search range for indicating both the number of points before the search center to be searched and the number of points after the search center. b. In one example, different search ranges may exist for the current frame and the reference frame. 3) It is proposed to signal the search range to the decoder by coding an indication indicating the search range. a. In one example, at least one indication for indicating the search range may be signaled to the decoder. i. In one example, when the codec performs a nearest neighbor search for all points, the indication may be a pre - defined signal. ii. In one example, when the search range is selected from several pre - defined search ranges, the indication may be selected from several pre - defined signals. iii. In one example, the indication may be a value of the search range. iv. In one example, the indication may be a pre - defined mathematical transformation (such as logarithm, square root, etc.) of the search range. b. In one example, the indication may be coded using fixed - length coding, unary coding, truncated unary coding, etc. c. In one example, the indication may be coded in a predictive manner. 4) Different geometric distances may be used for nearest neighbor search and generation of neighborhood weights in attribute inter - prediction. a. In one example, by performing a nearest neighbor search in the current frame and the reference frame, at least one point may be stored in list1. i. In one example, the selected point may be a point having the closest geometric distance (such as Euclidean distance, Manhattan distance, Chebyshev distance, etc.) from the current point. ii. In one example, the selected point may be from the searched points defined by the search center and the search range. b. In one example, the geometric distance of each point in list1 may be used for generation of neighborhood weights. c. In one example, the processes of nearest neighbor search and neighborhood weight generation may use different geometric distances. i. In one example, the Manhattan distance of each searched point may be used for nearest neighbor search, and the Euclidean distance of each point in list1 may be used for neighborhood weight generation. 5) It is proposed to apply motion compensation to the reference frame before attribute interpolation prediction. a. In one example, there may be motion compensation for the reference frame. b. In one example, motion compensation may be applied to the reference frame before attribute interpolation prediction. c. In one example, a reference frame with motion compensation may be used in attribute interpolation prediction. d. In one example, a reference frame without motion compensation may be used in attribute interpolation prediction. e. In one example, an indication indicating whether motion compensation is applied may be signaled to the decoder. i. In one example, the indication may be coded using fixed-length coding, unary coding, truncated unary coding, etc. ii. In one example, the indication may be coded in a predictive manner. 6) Points at different LODs of the reference frame may be searched in attribute interpolation prediction. a. In one example, the points in the reference frame may be divided into one or more LODs. b. In one example, the points in the current frame may be divided into one or more LODs. c. In one example, there may be one LOD level that refers to the LOD of each point. d. In one example, for the current point, points having the same LOD level in the reference frame may be searched to perform nearest neighbor search. e. In one example, for the current point, points having a lower LOD level in the reference frame may be searched to perform nearest neighbor search. f. In one example, for the current point, points having a higher LOD level in the reference frame may be searched to perform nearest neighbor search. g. In one example, an indication indicating whether points at all LODs are to be searched can be signaled to the decoder. i. In one example, the indication can be coded using, for example, fixed-length coding, unary coding, truncated unary coding, etc. ii. In one example, the indication can be coded in a predictive manner. h. In one example, an indication indicating whether only points at the same LOD are to be searched can be signaled to the decoder. i. In one example, the indication can be signaled conditionally, for example, according to whether points at all LODs are to be searched. ii. In one example, the indication can be coded using, for example, fixed-length coding, unary coding, truncated unary coding, etc. iii. In one example, the indication can be coded in a predictive manner. 7) It is proposed to save search results in different frames using multiple lists and combine all the lists to generate a predictor list. a. In one example, there may be a first list (such as list1) for saving search results in the current frame. b. In one example, there may be a second list (such as list2) for saving search results in each reference frame. c. In one example, the nearest neighbor search in the current frame may only modify list1. d. In one example, the nearest neighbor search in one reference frame may only modify the corresponding list. e. In one example, a predictor list can be generated using information of points in all the lists. 8) The foregoing "frame" can be replaced by other processing units, such as sub-regions within the frame. 9) The above method is applicable to other coding modules of G-PCC or other search methods in addition to the nearest neighbor search method. [6. Embodiment] 1) In this embodiment, an example of a method for performing nearest neighbor search using the Manhattan distance in attribute inter-prediction will be described. In this example, the search center within the reference frame is set to the point with the closest Morton code. The search ranges of both the current frame and the reference frame are set to 128. For each frame, the reference frame is the previous one frame, and the attribute inter-prediction is performed by an encoder and a decoder. First, the points of the current frame and the reference frame are reordered. The Morton code of each point is calculated, and the points within one frame are reordered based on the Morton code order. Second, for each point in the current frame, three approximate nearest neighbors in the current frame and the reference frame are searched and stored in list1. There are three flags indicating whether the nearest neighbor is from the current frame. a. The search center of the current frame is the current point. The previous 128 points before the search center in Morton code order are traversed. Among the traversed points, up to three points with the closest Manhattan distance are selected. The position, flag, and index of each point in list1 are recorded. The Manhattan distance d between two points (x1, y1, z1) and (x2, y2, z2) is calculated by the following formula.
Equation
Equation
[0058] Further details of embodiments of the present disclosure related to optimized inter-prediction for point cloud attribute coding based on nearest neighbor search are described below.
[0059] As used herein, the term "point cloud sequence" may refer to a sequence of one or more point clouds. The term "frame" may refer to a point cloud within a point cloud sequence. The term "PC sample" may refer to a unit for performing coding in point cloud sequence coding, such as a point cloud frame, a sub-region within a point cloud frame, a picture, a slice, a tile, a sub-picture, a node, a point, or other unit including one or more nodes or points.
[0060] FIG. 4 shows a flowchart of a method 400 for point cloud coding according to some embodiments of the present disclosure. Method 400 may be implemented during the conversion between the current PC sample of the point cloud sequence and the bitstream of the point cloud sequence. As shown in FIG. 4, method 400 starts at 402, where for the current point in the current PC sample, at least one neighboring point is determined from the set of points in the reference PC sample of the current PC sample. The set of points is in a group of the detail levels (LODs) of the reference PC sample. As an example, the at least one neighboring point may be at least one of the nearest neighbors of the current point. To obtain the at least one nearest neighbor, a nearest neighbor search may be performed for points at the same LOD as the current point, and a nearest neighbor search may be performed for points at an LOD lower than the current point.
[0061] At 404, a conversion is performed based on the at least one neighboring point. For example, the attribute value of the current point may be predicted by calculating a weighted average of the attribute values of the at least one neighboring point. The conversion may be performed based on the predicted attribute value. In some embodiments, the conversion may include encoding the current PC sample into a bitstream. Additionally or alternatively, the conversion may include decoding the current PC sample from the bitstream. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.
[0062] Based on the above, points within the group of LODs of the reference PC sample are searched to obtain the neighboring points of the current point. Compared with the conventional solution that only searches for points at the same LOD as the current point, the proposed method can advantageously improve the accuracy of the nearest neighbor search and attribute interpolation prediction.
[0063] In some embodiments, the points in the reference PC sample may be divided into multiple LODs, and the multiple LODs may include a group of LODs. Additionally or alternatively, the points in the current PC sample may be divided into one or more LODs.
[0064] In some embodiments, at least one neighboring point can be determined at 402 by performing a nearest neighbor search on a set of points. In some embodiments, for each point in the reference PC sample and / or the current PC sample, there can be one LOD level indicating a refinement list for one of the plurality of LODs. In one example, the set of points can include points at the same LOD level as the current point. In another example, the set of points can include points at an LOD level lower than the current point. In a further example, the set of points can include points at an LOD level higher than the current point. It should be understood that the above description is provided for illustrative purposes only. The scope of the present disclosure is not limited in this regard.
[0065] In some embodiments, a first indication indicating whether a nearest neighbor search is to be performed for points at all LODs of the reference PC sample can be included in the bitstream. For example, the first indication can be signaled to the decoder. In one example, the first indication can be coded with a fixed-length coding. Alternatively, the first indication can be coded with a unary coding. In a further example, the first indication can be coded with a truncated unary coding. In yet another example, the first indication can be coded in a predictive manner.
[0066] Additionally or alternatively, a second indication indicating whether a nearest neighbor search is to be performed for points at the LOD of the same reference PC sample as the current point can be included in the bitstream. The second indication can be signaled conditionally. In one example, the second indication can be included in the bitstream based on whether a nearest neighbor search can be performed for points at all LODs of the reference PC sample. In one example, the second indication can be coded with a fixed-length coding. Alternatively, the second indication can be coded with a unary coding. In a further example, the second indication can be coded with a truncated unary coding. In yet another example, the second indication can be coded in a predictive manner.
[0067] In some embodiments, at least one neighboring point can be determined at 402 based on a first geometric distance. At 404, at least one weight associated with the at least one neighboring point can be determined based on a second geometric distance, and the transformation can be performed based on the at least one weight and the at least one neighboring point. The second geometric distance is different from the first geometric distance. In some examples, the first geometric distance can be one of a Euclidean distance, a Manhattan distance, or a Chebyshev distance. In one example, the first geometric distance can be a Manhattan distance, and the second geometric distance can be a Euclidean distance. Alternatively, the second geometric distance can be determined based on the first geometric distance. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.
[0068] In some embodiments, at least one neighboring point can be stored in a list. In some embodiments, the at least one neighboring point can include a target point. The first geometric distance between the target point and the current point may be the smallest among the first geometric distances between the current point and each point in the set of points. In some embodiments, the set of points can be defined by a search center and a search range.
[0069] In some embodiments, at 404, a compensated reference frame can be obtained by applying motion compensation to the reference frame. Further, a predicted attribute value of the current point can be determined based on the at least one neighboring point and the compensated reference frame, and the transformation can be performed based on the predicted attribute value. That is, motion compensation can be applied to the reference frame before attribute inter-prediction.
[0070] Alternatively, at 404, a predicted attribute value of the current point can be determined based on the at least one neighboring point and the reference frame, and the transformation can be performed based on the predicted attribute value. That is, a reference frame without motion compensation can be used for attribute inter-prediction.
[0071] In some embodiments, a third indication indicating whether motion compensation is applied to a reference frame may be included in the bitstream. In one example, the third indication may be coded with fixed-length coding. Alternatively, the third indication may be coded with unary coding. In a further example, the third indication may be coded with truncated unary coding. In yet another example, the third indication may be coded in a predictive manner.
[0072] According to an embodiment of the present disclosure, a non-transitory computer-readable recording medium is proposed. A bitstream of a point cloud sequence is stored in the non-transitory computer-readable recording medium. The bitstream may be generated by a method executed by a point cloud processing device. According to this method, for a current point in a current PC sample, at least one neighboring point is determined from a set of points in a reference PC sample of the current PC sample. The set of points is in a group of detail levels (LOD) of the reference PC sample. Further, the bitstream is generated based on the at least one neighboring point.
[0073] According to an embodiment of the present disclosure, a method of storing a bitstream of a point cloud sequence is proposed. In this method, for a current point in a current PC sample, at least one neighboring point is determined from a set of points in a reference PC sample of the current PC sample. The set of points is in a group of detail levels (LOD) of the reference PC sample. Further, the bitstream is generated based on the at least one neighboring point. The bitstream is stored in a non-transitory computer-readable recording medium.
[0074] FIG. 5 shows a flowchart of another method 500 for point cloud coding according to some embodiments of the present disclosure. Method 500 may be implemented during the conversion between the current PC sample of the point cloud sequence and the bitstream of the point cloud sequence. As shown in FIG. 5, method 500 starts at 502 where, for the current point in the current PC sample, a first plurality of neighboring points are determined from the points in a second plurality of PC samples of the point cloud sequence. The first plurality of neighboring points are stored in a third plurality of lists. As an example, the first plurality of neighboring points may include at least one nearest neighbor of the current point. The nearest neighbor search may be performed on the points of the current PC sample, and the search results may be stored in the first list. Further, the nearest neighbor search may be performed on the points of the reference PC sample of the current PC sample, and the search results may be stored in a second list different from the second list.
[0075] At 404, a conversion is performed based on the first plurality of neighboring points. For example, the attribute value of the current point may be predicted by calculating a weighted average of the attribute values of the first plurality of neighboring points. The conversion may be performed based on the predicted attribute value. In some embodiments, the conversion may include encoding the current PC sample into a bitstream. Additionally or alternatively, the conversion may include decoding the current PC sample from the bitstream. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.
[0076] Based on the above, a plurality of neighboring points of the current point are stored in a plurality of lists. Compared with the conventional solution where the neighboring points are stored in only one list, the proposed method can advantageously improve the efficiency of nearest neighbor search and attribute inter-prediction.
[0077] In some embodiments, a first set of neighboring points among the first plurality of neighboring points may be stored in a first list among the third plurality of lists, where the first set of neighboring points is determined from points in the current PC sample. That is, there is a first list in the current PC sample for storing search results. Additionally or alternatively, a second set of neighboring points among the first plurality of neighboring points may be stored in a second list among the third plurality of lists, where the second set of neighboring points is determined from points in the reference PC sample of the current PC sample. The second list may be different from the first list. That is, there is a second list in the reference PC sample for storing search results.
[0078] In some embodiments, during the determination at 502, a first neighboring point among the first plurality of neighboring points may be determined from points in the current PC sample, and the first neighboring point may be stored in the first list. For example, in the nearest neighbor search of the current PC sample, only the first list may be modified. Additionally or alternatively, during the determination at 502, a second neighboring point among the first plurality of neighboring points may be determined from points in the reference PC sample, and the second neighboring point may be stored in the second list. For example, in the nearest neighbor search within one reference frame, only the corresponding list may be modified.
[0079] In some embodiments, at 504, at least one predicted attribute value of the current point may be determined based on neighboring points stored in the third plurality of lists, and the transformation may be performed based on the at least one predicted attribute value. That is, the information of the points stored in all the lists may be used to predict the attribute value of the current point.
[0080] In some embodiments, methods 400 and 500 may be applicable to other coding processes in geometry-based point cloud compression (GPCC) or search processes other than the nearest neighbor search. For example, methods 400 and 500 may be implemented in other coding modules in GPCC.
[0081] According to an embodiment of the present disclosure, a non-transitory computer-readable recording medium is proposed. A bitstream of a point cloud sequence is stored in the non-transitory computer-readable recording medium. The bitstream can be generated by a method executed by a point cloud processing device. According to this method, for a current point in a current PC sample, a first plurality of neighboring points are determined from points in a second plurality of PC samples of the point cloud sequence. The first plurality of neighboring points are stored in a third plurality of lists. Further, the bitstream is generated based on the first plurality of neighboring points.
[0082] According to an embodiment of the present disclosure, a method of storing a bitstream of a point cloud sequence is proposed. In this method, for a current point in a current PC sample, a first plurality of neighboring points are determined from points in a second plurality of PC samples of the point cloud sequence. The first plurality of neighboring points are stored in a third plurality of lists. Further, the bitstream is generated based on the first plurality of neighboring points. The bitstream is stored in a non-transitory computer-readable recording medium.
[0083] Embodiments of the present disclosure can be described in consideration of the following clauses, and their features can be combined in any reasonable manner.
[0084] Clause 1. A method for point cloud coding, comprising: during conversion between a current point cloud (PC) sample and a bitstream of a point cloud sequence, determining at least one neighboring point for a current point in the current PC sample of the point cloud sequence from a set of points in a reference PC sample of the current PC sample, wherein the set of points is in a group of levels of detail (LOD) of the reference PC sample; and performing the conversion based on the at least one neighboring point.
[0085] Clause 2. The method according to Clause 1, wherein the points in the reference PC sample are divided into a plurality of LODs, and the plurality of LODs include the group of LODs.
[0086] The method according to any one of clauses 1 to 2, wherein the points in the current PC sample are divided into one or more LODs.
[0087] The method according to any one of clauses 2 to 3, wherein for each point in the reference PC sample, there is one LOD level indicating one refinement list among the plurality of LODs.
[0088] The method according to clause 4, wherein the step of determining the at least one neighboring point includes the step of determining the at least one neighboring point by performing a nearest neighbor search on the set of points.
[0089] The method according to clause 5, wherein the set of points includes points at the same LOD level as the current point.
[0090] The method according to clause 5, wherein the set of points includes points at a lower LOD level than the current point.
[0091] The method according to clause 5, wherein the set of points includes points at a higher LOD level than the current point.
[0092] The method according to clause 5, wherein the bitstream includes a first indication indicating whether the nearest neighbor search is performed for points at all LODs of the reference PC sample.
[0093] The method according to clause 9, wherein the first indication is encoded in one of fixed-length coding, unary coding, or truncated unary coding.
[0094] The method according to clause 9, wherein the first indication is encoded in a predictive manner.
[0095] Clause 12. The method according to clause 5, wherein a second instruction indicating whether the nearest neighbor search is performed for a point at the LOD of the same reference PC sample as the current point is included in the bit stream.
[0096] Clause 13. The method according to clause 12, wherein the second instruction is included in the bit stream based on whether the nearest neighbor search is performed for points at all LODs of the reference PC sample.
[0097] Clause 14. The method according to clause 12, wherein the second instruction is encoded by one of fixed - length coding, unary coding, or truncated unary coding.
[0098] Clause 15. The method according to clause 12, wherein the second instruction is encoded in a predictive manner.
[0099] Clause 16. The method according to any one of clauses 1 to 15, wherein the at least one neighboring point is determined based on a first geometric distance, and the step of performing the transformation includes determining at least one weight associated with the at least one neighboring point based on a second geometric distance, wherein the second geometric distance is different from the first geometric distance, and performing the transformation based on the at least one weight and the at least one neighboring point.
[0100] Clause 17. The method according to clause 16, wherein the at least one neighboring point is stored in a list.
[0101] Clause 18. The method according to any one of clauses 16 to 17, wherein the at least one neighboring point includes a target point, and the first geometric distance between the target point and the current point is the minimum among the first geometric distances between the current point and each point in the set of points.
[0102] Clause 19. The method according to any one of Clauses 16 to 18, wherein the first geometric distance is one of a Euclidean distance, a Manhattan distance, or a Chebyshev distance.
[0103] Clause 20. The method according to any one of Clauses 16 to 19, wherein the set of points is defined by a search center and a search range.
[0104] Clause 21. The method according to any one of Clauses 16 to 20, wherein the second geometric distance is determined based on the first geometric distance.
[0105] Clause 22. The method according to any one of Clauses 16 to 20, wherein the first geometric distance is a Manhattan distance and the second geometric distance is a Euclidean distance.
[0106] Clause 23. The method according to any one of Clauses 1 to 15, wherein the step of performing the transformation includes obtaining a compensated reference frame by applying motion compensation to the reference frame, determining a predicted attribute value of the current point based on the at least one neighboring point and the compensated reference frame, and performing the transformation based on the predicted attribute value.
[0107] Clause 24. The method according to any one of Clauses 1 to 15, wherein the step of performing the transformation includes determining a predicted attribute value of the current point based on the at least one neighboring point and the reference frame, and performing the transformation based on the predicted attribute value.
[0108] Clause 25. The method according to any one of Clauses 1 to 15, wherein a third indication indicating whether motion compensation is applied to the reference frame is included in the bitstream.
[0109] Clause 26. The method according to Clause 25, wherein the third indication is coded in one of a fixed-length coding, a unary coding, or a truncated unary coding.
[0110] Clause 27. The third instruction is the method according to Clause 25, which is encoded in a predictive manner.
[0111] Clause 28. The at least one neighboring point is the method according to any one of Clauses 1 to 27, including at least one nearest neighbor of the current point.
[0112] Clause 29. A method for point cloud coding, comprising, during the conversion between a current point cloud (PC) sample and a bitstream of a point cloud sequence, determining, for a current point in the current PC sample of the point cloud sequence, a first plurality of neighboring points from points in a second plurality of PC samples of the point cloud sequence, and performing the conversion based on the first plurality of neighboring points, wherein the first plurality of neighboring points are stored in a third plurality of lists.
[0113] Clause 30. A first set of neighboring points among the first plurality of neighboring points are stored in a first list among the third plurality of lists, and the first set of neighboring points are determined from points in the current PC sample, according to the method described in Clause 29.
[0114] Clause 31. A second set of neighboring points among the first plurality of neighboring points are stored in a second list among the third plurality of lists, and the second set of neighboring points are determined from points in a reference PC sample of the current PC sample, and the second list is different from the first list, according to the method described in Clause 30.
[0115] Clause 32. The step of determining the first plurality of neighboring points includes determining a first neighboring point among the first plurality of neighboring points from points in the current PC sample, and storing the first neighboring point in the first list, according to any one of Clauses 30 to 31.
[0116] Clause 33. The step of determining the first plurality of neighboring points includes a step of determining a second neighboring point among the first plurality of neighboring points from a point in the reference PC sample, and a step of storing the second neighboring point in the second list, the method according to Clause 31.
[0117] Clause 34. The step of performing the conversion includes a step of determining at least one predicted attribute value of the current point based on the neighboring points stored in the third plurality of lists, and a step of performing the conversion based on the at least one predicted attribute value, the method according to any one of Clauses 29 to 33.
[0118] Clause 35. The first plurality of neighboring points includes at least one nearest neighbor of the current point, the method according to any one of Clauses 29 to 34.
[0119] Clause 36. The current PC sample is a point cloud frame in the point cloud sequence or a sub-region within the point cloud frame in the point cloud sequence, the method according to any one of Clauses 1 to 35.
[0120] Clause 37. The method is applicable to other coding processes in geometry-based point cloud compression (GPCC) or search processes other than nearest neighbor search, the method according to any one of Clauses 1 to 36.
[0121] Clause 38. The conversion includes a step of encoding the current PC sample into the bitstream, the method according to any one of Clauses 1 to 37.
[0122] Clause 39. The conversion includes a step of decoding the current PC sample from the bitstream, the method according to any one of Clauses 1 to 37.
[0123] Clause 40. An apparatus for processing point cloud data including a processor and a non-transitory memory storing instructions, which when executed by the processor, cause the processor to execute the method according to any one of Clauses 1 to 39.
[0124] Clause 41. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of Clauses 1 to 39.
[0125] Clause 42. A non-transitory computer-readable recording medium storing a bitstream of a point cloud sequence generated by a method executed by a point cloud processing apparatus, the method including: determining at least one neighboring point from a set of points in a reference point cloud (PC) sample of the current PC sample for a current point in the current PC sample of the point cloud sequence, the set of points being in a group of a detailed level of detail (LOD) of the reference PC sample; and generating the bitstream based on the at least one neighboring point.
[0126] Clause 43. A method for storing a bitstream of a point cloud sequence, the method including: determining at least one neighboring point from a set of points in a reference point cloud (PC) sample of the current PC sample for a current point in the current PC sample of the point cloud sequence, the set of points being in a group of a detailed level of detail (LOD) of the reference PC sample; generating the bitstream based on the at least one neighboring point; and storing the bitstream in a non-transitory computer-readable recording medium.
[0127] A non - transitory computer - readable recording medium storing a bitstream of a point - cloud sequence generated by a method executed by a point - cloud processing device, wherein the method includes: determining a first plurality of neighboring points from points in a second plurality of point - cloud (PC) samples of the point - cloud sequence for a current point in a current PC sample of the point - cloud sequence; and generating the bitstream based on the first plurality of neighboring points, wherein the first plurality of neighboring points are stored in a third plurality of lists.
[0128] A method for storing a bitstream of a point - cloud sequence, the method including: determining a first plurality of neighboring points from points in a second plurality of point - cloud (PC) samples of the point - cloud sequence for a current point in a current PC sample of the point - cloud sequence; generating the bitstream based on the first plurality of neighboring points; and storing the bitstream in a non - transitory computer - readable recording medium, wherein the first plurality of neighboring points are stored in a third plurality of lists.
[0129] [Exemplary Device] FIG. 6 shows a block diagram of a computing device 600 that can embody various embodiments of the present disclosure. The computing device 600 can be embodied as or included in a source device 110 (or GPCC encoder 116 or 200) or a destination device 120 (or GPCC decoder 126 or 300).
[0130] It will be understood that the computing device 600 shown in FIG. 6 is for illustrative purposes only and is not intended to limit in any way the functions and scope of the embodiments of the present disclosure.
[0131] As shown in FIG. 6, computing device 600 includes a general-purpose computing device 600. Computing device 600 may include at least one or a plurality of processors or processing units 610, a memory 620, a storage unit 630, one or a plurality of communication units 640, one or a plurality of input devices 650, and one or a plurality of output devices 660.
[0132] In some embodiments, computing device 600 may be embodied as any user terminal or server terminal having computing capabilities. The server terminal may be, for example, a server provided by a service provider or a large-scale computing device. The user terminal may be, for example, a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination thereof (including accessories and peripherals of these devices, or any combination thereof), and may be any type of mobile terminal, fixed terminal, or portable terminal. It is contemplated that computing device 600 can support any type of interface for the user (such as a "wearable" circuit).
[0133] The processing unit 610 can be a physical or virtual processor and can implement various processes based on a program stored in the memory 620. In a multiprocessor system, in order to improve the parallel processing ability of the computing device 600, a plurality of processing units execute computer-executable instructions in parallel. The processing unit 610 can be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0134] The computing device 600 typically includes various computer storage media. Such media can be any media accessible by the computing device 600, including but not limited to volatile and non-volatile media, or removable and non-removable media. The memory 620 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or any combination thereof. The storage unit 630 can be any removable or non-removable media and can be used to store information and / or data, and can include machine-readable media such as memory, flash memory drives, magnetic disks, or other media accessible by the computing device 600.
[0135] The computing device 600 can further include additional removable / non-removable, volatile / non-volatile memory media. Although not shown in FIG. 6, it is possible to provide a magnetic disk drive for reading and writing a removable non-volatile magnetic disk and an optical disk drive for reading and writing a removable non-volatile optical disk. In such a case, each drive can be connected to a bus (not shown) via one or more data media interfaces.
[0136] The communication unit 640 communicates with additional computing devices via a communication medium. Further, the functionality of components within the computing device 600 may be embodied by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or additional general network nodes.
[0137] The input device 650 can be one or more of various input devices such as a mouse, keyboard, trackball, voice input device, etc. The output device 660 can be one or more of various output devices such as a display, loudspeaker, printer, etc. The communication unit 640 enables the computing device 600 to further communicate with one or more external devices (not shown) such as a storage device and a display device, and enables a user to interact with the computing device 600 via one or more devices, or, if necessary, enables the computing device 600 to communicate with one or more other computing devices via any device (such as a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0138] In some embodiments, rather than being integrated into a single device, some or all components of computing device 600 may be disposed in a cloud computing architecture. In a cloud computing architecture, components are provided remotely and may cooperate to perform the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, and an end user need not be aware of the physical location or configuration of the system or hardware that provides these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides applications via a wide area network that can be accessed through a web browser or other computing component. Software or components and corresponding data of a cloud computing architecture may be stored on a server at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at the location of a remote data center. A cloud computing infrastructure operates as a single access point for a user but may provide services through a shared data center. Thus, a cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.
[0139] Computing device 600 may be used to implement point cloud encoding / decoding in embodiments of this disclosure. Memory 620 may include one or more point cloud coding modules 625 having one or more program instructions. These modules are accessible and executable by processing unit 610 to perform the functions of the various embodiments described herein.
[0140] In an exemplary embodiment for performing point cloud symbolization, the input device 650 may receive point cloud data to be symbolized as input 670. The point cloud data may be processed, for example, by a point cloud coding module 625 to generate an encoded bitstream. The encoded bitstream may be provided as output 680 via an output device 660.
[0141] In an exemplary embodiment for performing point cloud decoding, the input device 650 may receive an encoded bitstream as input 670. The encoded bitstream may be processed, for example, by a point cloud coding module 625 to generate decoded point cloud data. The decoded point cloud data may be provided as output 680 via an output device 660.
[0142] Although the present disclosure has been particularly illustrated and described with reference to its preferred embodiments, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the application as defined by the appended claims. Such modifications are intended to be included within the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for point cloud coding, comprising: During the conversion between a current point cloud (PC) sample and a bitstream of a point cloud sequence, for a current point in the current PC sample of the point cloud sequence, based on a first geometric distance, determining at least one neighboring point from a set of points in a reference PC sample of the current PC sample; Determining at least one weight associated with the at least one neighboring point based on a second geometric distance, wherein the second geometric distance is different from the first geometric distance; Executing the conversion based on the at least one weight and the at least one neighboring point.
2. The method according to claim 1, wherein the at least one neighboring point is stored in a predictor list.
3. The method according to claim 1, wherein the at least one neighboring point includes a target point, and the first geometric distance between the target point and the current point is the smallest among the first geometric distances between the current point and each point in the set of points.
4. The first geometric distance is Determined based on the Manhattan distance, according to the method of claim 1.
5. The set of points is defined by a search center and a search range, according to the method of claim 1.
6. The set of points is included in a group of the detailed level (LOD) of the reference PC sample, according to the method of claim 1.
7. The step of determining the at least one neighboring point Includes determining the at least one neighboring point by performing a nearest neighbor search on the set of points, according to the method of claim 6.
8. The set of points includes at least one of points at the same LOD level as the current point, Points at a lower LOD level than the current point, or Points at a higher LOD level than the current point, according to the method of claim 7.
9. The step of executing the conversion Includes obtaining a compensated reference PC sample by applying motion compensation to the reference PC sample; Determining a predicted attribute value of the current point based on the at least one neighboring point and the compensated reference PC sample; Executing the conversion based on the predicted attribute value, according to the method of claim 1.
10. The step of performing the conversion comprises: determining a predicted attribute value of the current point based on the at least one neighboring point and the reference PC sample; performing the conversion based on the predicted attribute value, according to the method of claim 1. **Claim 11** The method of claim 1, wherein a third indication indicating whether motion compensation is applied to the reference PC sample is included in the bitstream. **Claim 12** The method of claim 2, wherein the predictor list is generated by combining a plurality of lists for storing neighboring points of the current point determined from different PC samples. **Claim 13** The method of claim 1, wherein the current PC sample is a slice within a point cloud frame in the point cloud sequence. **Claim 14** The method of claim 1, wherein the conversion includes encoding the current PC sample into the bitstream. **Claim 15** The method of claim 1, wherein the conversion includes decoding the current PC sample from the bitstream. **Claim 16** An apparatus for processing point cloud data, comprising a processor and a non-transitory memory storing instructions, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 15. **Claim 17** A non-transitory computer-readable storage medium storing instructions for causing a processor to perform the method according to any one of claims 1 to 15. **Claim 18** A method for storing a bitstream of a point cloud sequence, comprising: determining at least one neighboring point from a set of points in a reference PC sample of the current PC sample, based on a first geometric distance, for a current point in the current point cloud (PC) sample of the point cloud sequence; determining at least one weight associated with the at least one neighboring point, based on a second geometric distance, wherein the second geometric distance is different from the first geometric distance; generating the bitstream based on the at least one weight and the at least one neighboring point; and storing the bitstream in a non-transitory computer-readable recording medium.
Citation Information
Patent Citations
Method and apparatus for video coding
US20200021844A1
Dynamic Point Cloud Compression Using Inter-Prediction
US20210099711A1
Trimming Search Space For Nearest Neighbor Determinations in Point Cloud Compression
US20210103780A1
Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
US20210320960A1
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
WO2020175708A1