Method, device and medium for point cloud coding and decoding
By introducing multi-reference inter prediction and hierarchical codec priority into point cloud sequences, the limitations of inter prediction design in the prior art are solved, and the quality and efficiency of point cloud codec are improved.
Patent Information
- Application Number
- CN202380072328.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-04-14
- Publication Date
- 2025-05-30
AI Technical Summary
In the existing point cloud compression technology, there are multiple limitations in inter-frame prediction design, including using only one reference frame for prediction. The reference frame can only be frames with an earlier time stamp, and the encoding and decoding priority is not flexible enough, which affects the encoding and decoding performance.
A multi-reference inter prediction method is proposed, allowing inter prediction using multiple reference PC samples in a point cloud sequence, and by indicating whether multi-reference inter prediction is enabled. In addition, the hierarchical reference relationship and hierarchical codec accuracy are used to dynamically adjust the codec parameters according to the codec priority.
Through multi-reference inter prediction and dynamic codec parameter adjustment, the encoding and codec quality and efficiency of point cloud codec are improved, and the prediction accuracy and codec performance are enhanced.
Smart Images

Figure CN120077662A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to point cloud encoding and decoding techniques, and more particularly to multi-reference inter-frame prediction for point cloud encoding and decoding. Background Art
[0002] A point cloud is a collection of individual data points in a three-dimensional (3D) plane, where each point has set coordinates on the X-axis, Y-axis, and Z-axis. Thus, point clouds can be used to represent the physical content of 3D space. For various immersive applications ranging from augmented reality to autonomous vehicles, point clouds have proven to be a promising way to represent 3D visual data.
[0003] Point cloud encoding and decoding standards have mainly evolved through the well-known MPEG organization. MPEG is short for the Moving Picture Experts Group, which is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Encoding and Decoding Group (3DG) released a Call for Proposals (CFP) document to initiate the development of point cloud encoding and decoding standards. The final standard will encompass two categories of solutions. Video-based Point Cloud Compression (V-PCC or VPCC) is applicable to point sets with relatively uniform point distributions. Geometry-based Point Cloud Compression (G-PCC or GPCC) is applicable to more sparse distributions. However, there is an overall expectation to further improve the encoding and decoding efficiency of conventional point cloud encoding and decoding techniques. Summary of the Invention
[0004] Embodiments of the present disclosure provide a solution for point cloud encoding and decoding.
[0005] In a first aspect, a method for point cloud encoding and decoding is proposed. The method includes: obtaining a first indication indicating whether multi-reference inter-frame prediction is enabled for a point cloud sequence for the conversion between a current point cloud (PC) sample of the point cloud sequence and a bitstream, where multiple reference PC samples are used in the multi-reference inter-frame prediction; and performing the conversion based on the first indication.
[0006] Based on the method according to the first aspect of the present disclosure, the conversion between the point cloud sequence and the bitstream is performed based on an indication indicating whether multi-reference inter-frame prediction is enabled for the point cloud sequence. In this way, the proposed method can advantageously facilitate the application of multi-reference inter-frame prediction, and thus can improve the encoding and decoding quality of point cloud encoding and decoding.
[0007] In a second aspect, a method for point cloud encoding and decoding is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0008] In a third aspect, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.
[0009] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a point cloud encoding / decoding device for a point cloud sequence. The method includes: obtaining a first indication indicating whether multi-reference frame inter prediction is enabled for the point cloud sequence, where multiple reference PC samples are used in the multi-reference frame inter prediction; and generating a bitstream based on the first indication.
[0010] In a fifth aspect, a method for storing a bitstream of a point cloud sequence is provided. The method includes: obtaining a first indication indicating whether multi-reference frame inter prediction is enabled for the point cloud sequence, where multiple reference PC samples are used in the multi-reference frame inter prediction; generating a bitstream based on the first indication; and storing the bitstream in a non-transitory computer-readable recording medium.
[0011] The present invention content is provided to introduce a selection of concepts in a simplified form, which will be further described in the following detailed implementation. The present invention content is not intended to identify the key features or basic features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Through the following detailed description with reference to the accompanying drawings, the above and other objectives, features, and advantages of the exemplary embodiments of the present disclosure will become more clear. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0013] Figure 1 is a block diagram showing an exemplary point cloud encoding / decoding system that can utilize the technology of the present disclosure;
[0014] Figure 2 shows a block diagram of an exemplary point cloud encoder for showing some embodiments according to the present disclosure;
[0015] Figure 3 shows a block diagram of an exemplary point cloud decoder for showing some embodiments according to the present disclosure;
[0016] Figure 4 shows a schematic diagram for showing an example of inter prediction for predictive geometry encoding / decoding;
[0017] Figure 5 shows a schematic diagram for showing an example of a group of pictures (GOF) structure with a GOF size of 8;
[0018] Figure 6A schematic diagram showing an example of a hierarchical reference relationship for showing a GOF is shown;
[0019] Figure 7 A schematic diagram showing another example of a hierarchical reference relationship for showing a GOF is shown;
[0020] Figure 8 A schematic diagram showing an example of deriving the prediction direction of a child node is shown;
[0021] Figure 9 A schematic diagram showing an example of a reference relationship of an IPPP GOF structure is shown;
[0022] Figure 10 A flowchart showing a method for point cloud encoding and decoding according to some embodiments of the present disclosure; and
[0023] Figure 11 A block diagram of a computing device in which various embodiments of the present disclosure can be implemented is shown.
[0024] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description
[0025] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is only for the purpose of illustration and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways other than those described below.
[0026] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.
[0027] As used in the present disclosure, the terms "one embodiment", "embodiment", "example embodiment", etc. indicate that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment must include that specific feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a specific feature, structure, or characteristic is described in connection with an example embodiment, it is claimed that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.
[0028] It should be understood that although terms such as "first" and "second" may be used to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0029] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes" and / or "including" when used herein specify the presence of the stated features, elements and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example environment
[0030] Figure 1 is a block diagram showing an exemplary point cloud encoding and decoding system 100 that can utilize the techniques of the present disclosure. As shown, the point cloud encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a point cloud encoding device, and the destination device 120 may also be referred to as a point cloud decoding device. In operation, the source device 110 may be configured to generate encoded point cloud data, and the destination device 120 may be configured to decode the encoded point cloud data generated by the source device 110. The techniques of the present disclosure generally aim at encoding and decoding (encoding and / or decoding) point cloud data, that is, supporting point cloud compression. Encoding and decoding may be effective in compressing and / or decompressing point cloud data.
[0031] The source device 100 and the destination device 120 may include any of a variety of devices, including desktop computers, notebooks (i.e., laptops) computers, tablet computers, set-top boxes, telephone handsets (such as smart phones and mobile phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, vehicles (e.g., land or sea vehicles, spacecraft, aircraft, etc.), robots, LIDAR devices, satellites, extended reality devices, and so on. In some cases, the source device 100 and the destination device 120 may be equipped for wireless communication.
[0032] The source device 100 may include a data source 112, a memory 114, a GPCC encoder 116, and an input / output (I / O) interface 118. The destination device 120 may include an input / output (I / O) interface 128, a GPCC decoder 126, a memory 124, and a data consumer 122. According to the present disclosure, the GPCC encoder 116 of the source device 100 and the GPCC decoder 126 of the destination device 120 may be configured to apply the techniques related to point cloud encoding and decoding of the present disclosure. Thus, the source device 100 represents an example of an encoding device, and the destination device 120 represents an example of a decoding device. In other examples, the source device 100 and the destination device 120 may include other components or arrangements. For example, the source device 100 may receive data (e.g., point cloud data) from an internal source or an external source. Similarly, the destination device 120 may interface with an external data consumer instead of including a data consumer in the same device.
[0033] Generally, the data source 112 represents a source of point cloud data (i.e., raw, unencoded point cloud data) and may provide a continuous series of “frames” of point cloud data to the GPCC encoder 116, which encodes the point cloud data for the frames. In some examples, the data source 112 generates the point cloud data. The data source 112 of the source device 100 may include a point cloud acquisition device, such as any of various cameras or sensors, e.g., one or more cameras, an archive containing previously acquired point cloud data, a 3D scanner, or a light detection and ranging (LIDAR) device, and / or a data feed interface that receives point cloud data from a data content provider. Thus, in some examples, the data source 112 may generate point cloud data based on signals from a LIDAR device. Alternatively or additionally, the point cloud data may be generated from a scanner, a camera, a sensor, or other data by a computer. For example, the data source 112 may generate point cloud data, or a combination of real-time point cloud data, archived point cloud data, and computer-generated point cloud data. In each case, the GPCC encoder 116 encodes the acquired, pre-acquired, or computer-generated point cloud data. The GPCC encoder 116 may rearrange the frames of the point cloud data from the received order (sometimes referred to as “display order”) to an encoding / decoding order for encoding and decoding. The GPCC encoder 116 may generate one or more bitstreams including the encoded point cloud data. The source device 100 may then output the encoded point cloud data via the I / O interface 118 for reception and / or retrieval by, for example, the I / O interface 128 of the destination device 120. The encoded point cloud data may be directly transmitted to the destination device 120 via the I / O interface 118 through the network 130A. The encoded point cloud data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0034] The memory 114 of the source device 100 and the memory 124 of the destination device 120 may represent general memories. In some examples, the memories 114 and 124 may store raw point cloud data, e.g., raw point cloud data from the data source 112 and raw, decoded point cloud data from the GPCC decoder 126. Additionally or alternatively, the memories 114 and 124 may store software instructions executable by, e.g., the GPCC encoder 116 and the GPCC decoder 126, respectively. Although the memories 114 and 124 are shown separately from the GPCC encoder 116 and the GPCC decoder 126 in this example, it should be understood that the GPCC encoder 116 and the GPCC decoder 126 may also include internal memories for functionally similar or equivalent purposes. Further, the memories 114 and 124 may store encoded point cloud data, e.g., encoded point cloud data output from the GPCC encoder 116 and input to the GPCC decoder 126. In some examples, portions of the memories 114 and 124 may be allocated as one or more caches, e.g., for storing raw point cloud data, decoded point cloud data, and / or encoded point cloud data. For example, the memories 114 and 124 may store point cloud data.
[0035] The I / O interfaces 118 and 128 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where the I / O interfaces 118 and 128 include wireless components, the I / O interfaces 118 and 128 may be configured to transmit data, such as encoded point cloud data, according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc. In some examples where the I / O interface 118 includes a wireless transmitter, the I / O interfaces 118 and 128 may be configured to transmit data, such as encoded point cloud data, according to other wireless standards such as the IEEE 802.11 specifications. In some examples, the source device 100 and / or the destination device 120 may include respective system-on-a-chip (SoC) devices. For example, the source device 100 may include an SoC device for performing functions attributed to the GPCC encoder 116 and / or the I / O interface 118, and the destination device 120 may include an SoC device for performing functions attributed to the GPCC decoder 126 and / or the I / O interface 128.
[0036] The techniques of the present disclosure can be applied to encoding and decoding to support any one of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors, and processing devices (e.g., local or remote servers), geographical mapping, or other applications.
[0037] The I / O interface 128 of the destination device 120 receives the encoded bitstream from the source device 110. The encoded bitstream may include signaling information defined by the GPCC encoder 116, which is also used by the GPCC decoder 126, such as syntax elements having values representing point clouds. The data consumer 122 uses the decoded data. For example, the data consumer 122 may use the decoded point cloud data to determine the position of a physical object. In some examples, the data consumer 122 may include a display for presenting an image based on the point cloud data.
[0038] Each of the GPCC encoder 116 and the GPCC decoder 126 can be implemented as any one of a variety of suitable encoder circuit systems and / or decoder circuit systems, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the GPCC encoder 116 and the GPCC decoder 126 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device. Devices including the GPCC encoder 116 and / or the GPCC decoder 126 may include one or more integrated circuits, microprocessors, and / or other types of devices.
[0039] The GPCC encoder 116 and the GPCC decoder 126 can operate according to encoding and decoding standards, such as the Video Point Cloud Compression (VPCC) standard or the Geometric Point Cloud Compression (GPCC) standard. Generally, the present disclosure may refer to the encoding and decoding (e.g., encoding and decoding) of frames to include the process of encoding data or decoding data. The encoded bitstream typically includes a series of values for syntax elements representing encoding and decoding decisions (e.g., encoding and decoding modes).
[0040] A point cloud can include a set of points in 3D space and can have attributes associated with the points. The attributes can be color information, such as R, G, B or Y, Cb, Cr or reflectivity information, or other attributes. The point cloud can be acquired by various cameras or sensors (such as LIDAR sensors and 3D scanners) and can also be computer-generated. Point cloud data is used in various applications, including but not limited to construction (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors for assisting navigation).
[0041] Figure 2 is a block diagram showing an example of a GPCC encoder 200 according to some embodiments of the present disclosure. The GPCC encoder 200 can be Figure 1 an example of the GPCC encoder 116 in the system 100 shown. Figure 3 is a block diagram showing an example of a GPCC decoder 300 according to some embodiments of the present disclosure. The GPCC decoder 300 can be Figure 1 an example of the GPCC decoder 126 in the system 100 shown.
[0042] In both the GPCC encoder 200 and the GPCC decoder 300, the point cloud positions are first encoded and decoded. Attribute encoding and decoding depend on the decoded geometry. In Figure 2 and Figure 3 , the Region Adaptive Hierarchical Transform (RAHT) unit 218, the Surface Approximation Analysis unit 212, the RAHT unit 314, and the Surface Approximation Synthesis unit 310 are options typically used for Category 1 data. The Level of Detail (LOD) Generation unit 220, the Boosting unit 222, the LOD Generation unit 316, and the Inverse Boosting unit 318 are options typically used for Category 3 data. All other units are common between Category 1 and Category 3.
[0043] For Category 3 data, the compressed geometry is typically represented as an octree from the root all the way down to the leaf level of individual voxels. For Category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to the leaf level of blocks larger than voxels) plus a model for approximating the surface within each leaf node of the pruned octree. In this way, both Category 1 data and Category 3 data share the octree encoding and decoding mechanism, while Category 1 data can additionally utilize the surface model to approximate the voxels within each leaf node. The surface model used is a triangulation including 1 to 10 triangles per block, resulting in a triangle soup. Therefore, the Category 1 geometry codec is referred to as the Triangle Soup (Trisoup) geometry codec, while the Category 3 geometry codec is referred to as the octree geometry codec.
[0044] InFigure 2 In the example of, the GPCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometry reconstruction unit 216, a RAHT unit 218, a LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.
[0045] As Figure 2 shown in the example of, the GPCC encoder 200 may receive a set of positions and a set of attributes. The positions may include the coordinates of the points in the point cloud. The attributes may include information about the points in the point cloud, such as the color associated with the points in the point cloud.
[0046] The coordinate transformation unit 202 may apply a transformation to the coordinates of the points to transform the coordinates from an initial domain to a transformed domain. The present disclosure may refer to the transformed coordinates as transformed coordinates. The color transformation unit 204 may apply a transformation to convert the color information of the attributes to a different domain. For example, the color transformation unit 204 may convert the color information from the RGB color space to the YCbCr color space.
[0047] In addition, in Figure 2 the example of, the voxelization unit 206 may voxelize the transformed coordinates. The voxelization of the transformed coordinates may include quantizing and removing some of the points of the point cloud. In other words, multiple points of the point cloud may be grouped into a single "voxel", which may be considered a point in some aspects thereafter. Additionally, the octree analysis unit 210 may generate an octree based on the voxelized transformed coordinates. Additionally, in Figure 2 the example of, the surface approximation analysis unit 212 may analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 may perform arithmetic coding on the syntax elements representing the information of the octree and / or the information of the surface determined by the surface approximation analysis unit 212. The GPCC encoder 200 may output these syntax elements in a geometry bitstream.
[0048] The geometry reconstruction unit 216 may reconstruct the transformed coordinates of the points in the point cloud based on the octree, the data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transformed coordinates reconstructed by the geometry reconstruction unit 216 may be different from the original number of points in the point cloud. The present disclosure may refer to the resulting points as reconstructed points. The attribute transfer unit 208 may transfer the attributes of the original points of the point cloud to the reconstructed points of the point cloud data.
[0049] In addition, the RAHT unit 218 may apply RAHT coding to the attributes of the reconstructed points. Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 may apply LOD processing and lifting to the attributes of the reconstructed points, respectively. The RAHT unit 218 and the lifting unit 222 may generate coefficients based on the attributes. The coefficient quantization unit 224 may quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 may apply arithmetic coding to the syntax elements representing the quantized coefficients. The GPCC encoder 200 may output these syntax elements in the attribute bitstream.
[0050] In Figure 3 the example of, the GPCC decoder 300 may include a geometric arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometric reconstruction unit 312, a RAHT unit 314, an LOD generation unit 316, an inverse lifting unit 318, a coordinate inverse transformation unit 320, and a color inverse transformation unit 322.
[0051] The GPCC decoder 300 may obtain a geometric bitstream and an attribute bitstream. The geometric arithmetic decoding unit 302 of the decoder 300 may apply arithmetic decoding (e.g., CABAC or other types of arithmetic decoding) to the syntax elements in the geometric bitstream. Similarly, the attribute arithmetic decoding unit 304 may apply arithmetic decoding to the syntax elements in the attribute bitstream.
[0052] The octree synthesis unit 306 may synthesize an octree based on the syntax elements parsed from the geometric bitstream. In the case where surface approximation is used in the geometric bitstream, the surface approximation synthesis unit 310 may determine a surface model based on the syntax elements parsed from the geometric bitstream and based on the octree.
[0053] In addition, the geometric reconstruction unit 312 may perform reconstruction to determine the coordinates of the points in the point cloud. The coordinate inverse transformation unit 320 may apply an inverse transformation to the reconstructed coordinates to transform the reconstructed coordinates (positions) of the points in the point cloud from the transform domain back to the initial domain.
[0054] Additionally, in Figure 3 the example of, the inverse quantization unit 308 may perform inverse quantization on the attribute values. The attribute values may be based on the syntax elements obtained from the attribute bitstream (e.g., including the syntax elements decoded by the attribute arithmetic decoding unit 304).
[0055] Depending on how the attribute values are encoded, the RAHT unit 314 may perform RAHT decoding to determine color values for points in the point cloud based on the dequantized attribute values. Alternatively, the LOD generation unit 316 and the inverse lifting unit 318 may use level-of-detail based techniques to determine color values for points in the point cloud.
[0056] In addition, in Figure 3 the example of, the color inverse transform unit 322 may apply an inverse color transform to the color values. The inverse color transform may be the inverse of the color transform applied by the color transform unit 204 of the encoder 200. For example, the color transform unit 204 may transform color information from the RGB color space to the YCbCr color space. Correspondingly, the color inverse transform unit 322 may transform color information from the YCbCr color space to the RGB color space.
[0057] Figure 2 and Figure 3 The various units of and are shown to assist in understanding the operations performed by the encoder 200 and the decoder 300. These units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function and is preset with respect to the operations that can be performed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations that can be performed. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit may execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuit are generally immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.
[0058] Some exemplary embodiments of the present disclosure will be described in detail below. It should be understood that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the sections to that section. In addition, although some embodiments are described with reference to GPCC or other specific point cloud codecs, the disclosed techniques are also applicable to other point cloud coding and decoding techniques. In addition, although some embodiments describe the point cloud coding and decoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. 1. Overview The present disclosure relates to point cloud coding and decoding techniques. Specifically, the present disclosure relates to the coding, decoding, and encapsulation of coding and decoding parameters in point cloud coding and decoding. These concepts can be applied, alone or in various combinations, to any point cloud coding standard or non-standard point cloud codec, such as the geometry-based point cloud compression (G-PCC) being developed. 2. Abbreviations Geometry-based Point Cloud Compression (G-PCC) Moving Picture Experts Group (MPEG) 3D Graphics Coding and Decoding Group (3DG) Call for Proposals (CFP) Video-based Point Cloud Compression (V-PCC) Core Experiment (CE) Exploration Experiment (EE) Inter-frame Exploration Model (inter-EM) Group of Frames (GOF) Rate-Distortion Optimization (RDO) Global Motion (GM) Quantization Parameter (QP) Random Access (RA) First In First Out (FIFO) Occupancy Code (OC) Picture Order Count (POC) Point Cloud (PC) 3. Introduction The point cloud coding and decoding standard has mainly evolved through the well-known Moving Picture Experts Group (MPEG). MPEG, short for Moving Picture Experts Group, is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Coding and Decoding Group (3DG) released a Call for Proposals (CFP) document to initiate the development of the point cloud coding and decoding standard. The final standard will encompass two categories of solutions. Video-based Point Cloud Compression (V-PCC) is applicable to point sets with relatively uniform point distributions. Geometry-based Point Cloud Compression (G-PCC) is applicable to sparser distributions. To explore future point cloud coding and decoding techniques in G-PCC, Core Experiment (CE) 13.5 and Exploration Experiment (EE) 13.2 were formed to develop inter-frame prediction techniques in G-PCC. Subsequently, many new inter-frame prediction methods were adopted by MPEG and incorporated into the reference software called the Inter-frame Exploration Model (inter-EM). In a point cloud frame, there are many data points to describe a 3D object or scene. For each data point, there can be corresponding geometric information and attribute information. Geometric information is used to record the spatial position of the data point. Attribute information is used to record more details of the data point, such as texture, normal vector, and reflection. In inter-EM, there are some optional tools that respectively support the inter-frame predictive coding and decoding of geometric information and attribute information. For the attribute information, the codec performs inter-frame prediction for each point in the current frame using the attribute information of the reference points. The reference points are selected from the data points in the current frame and the reference frame based on the geometric distance of the points. Each reference point corresponds to a weight value, which is based on the geometric distance from the current point. The predicted attribute value can be the weighted average of the attribute values of the reference points or one of the attribute values of the reference points. The decision of the predicted attribute value is based on the rate-distortion optimization (RDO) method. For the geometric information, there are two main methods for performing inter-frame prediction coding, namely the octree-based method and the prediction-tree-based method. In the first method, the geometric information is represented by an octree structure and the occupancy code (OC) of each node. For each node in the octree of the current frame, the codec will decide whether to perform octagonal partitioning based on the number of points in the current node. The same partitioning is performed on the corresponding reference node in the reference frame. At the same time, the occupancy codes of the current node and the reference node are calculated. The codec will use the occupancy code of the reference node to perform prediction coding for the occupancy code of the current node. In the second method, the points in the point cloud are sorted to form a prediction tree. As Figure 4 shown, for each point, the previously decoded point will be selected as point A. Then the point in the reference frame with the same scaled azimuth angle and laser ID as point A will be selected as point B. Finally, the point in the reference frame as follows will be selected as point C, which is the first point with a scaled azimuth angle larger than that of point B. The codec will use the geometric information of point C to perform prediction coding for the geometric information of the current point. In the current inter-EM, the IPPP structure is applied, which means that if inter-frame prediction is applied to the current frame, the reference frame of the current frame is the previous frame. At the same time, the inter-EM uses the quantization parameter (QP) to control the bitrate point, and all frames share the same QP value. 4. Problems The existing designs for inter-frame prediction of point cloud compression have the following problems: 1. In the current inter-EM, there is only one reference frame for each frame to perform inter-frame prediction. In theory, the more reference information, the more accurate the prediction result. Using only one reference frame will limit the prediction accuracy and affect the coding and decoding efficiency. 2. In the current inter-EM, the reference frame can only be the frame with an earlier timestamp (i.e., a smaller POC value). The purpose of inter-frame prediction is to eliminate the redundant information between consecutive frames. However, the redundant information exists not only between the previous frame and the current frame, but also between the current frame and the subsequent frames. Using only the frame with an earlier timestamp will limit the coding and decoding performance. 3. In the current inter-EM, the QP value for each frame is the same. However, some frames are reference frames for other frames, which means their encoding and decoding priorities should be higher. In the case of limited transmission resources, lower QP values should be assigned to them to ensure they can be transmitted more accurately. When the transmission resources are very limited, applying the same encoding and decoding accuracy to all frames will affect the encoding and decoding performance. 5. Detailed solutions To solve the above problems and some other problems not mentioned, the following summarized methods are disclosed. This solution should be regarded as an example to explain the general concept and should not be interpreted in a narrow sense. In addition, these solutions can be applied individually or in any combined way. In the following discussion, the term "PC sample" refers to the unit for predictive encoding and decoding in the point cloud sequence encoding and decoding, such as a frame / picture / slice / tile / sub-picture / node / point / other unit containing one or more nodes or points. 1) It is proposed to divide frames into one or more groups of frames (GOF) in a point cloud sequence to perform point cloud compression. a. In one example, N consecutive frames in timestamp order can be clustered into a GOF. i. In one example, each frame can belong to one GOF. ii. In one example, N can be equal to the GOF size. b. In one example, the first frame of the GOF in decoding order can be an I frame. i. In one example, there can be only intra-frame prediction for the I frame. c. In one example, the first frame of the GOF in decoding order can not be an I frame. i. In one example, the first frame of the GOF in decoding order can be a P frame. ii. In one example, the first frame of the GOF in decoding order can be a P frame or a B frame, where all reference frames are before the current frame in timestamp order. d. Whether to encode and decode the first frame of the GOF in decoding order as an I frame can depend on the intra-frame period / random access period. e. In one example, the GOF size can be equal to the intra-frame period / random access period. f. In one example, the GOF size can be less than the intra-frame period / random access period. g. In one example, the indication of the GOF size and / or the encoding and decoding structure within the GOF can be signaled. 2) It is proposed to use one or more reference PC samples to perform inter-frame prediction for the current PC sample. a. In one example, for a current PC sample, there can be one or more reference PC samples. b. In one example, multiple reference PC samples can come from different reference slices / frames. i. Alternatively, multiple reference PC samples can come from the same reference slice / frame. c. In one example, a reference PC sample can be derived from at least one PC reconstruction sample. i. In one example, a reference PC sample can be a PC reconstruction sample. ii. In one example, a reference PC sample can be the result of a process applied to at least one PC reconstruction sample. For example, the process can be sampling or upsampling. iii. In one example, a reference PC sample can be a combined PC sample from multiple PC samples. (1) In one example, the combined PC sample of multiple PC samples can be a clustering of all points in the PC samples. (2) Alternatively, the combined PC sample of multiple PC samples can be a clustering of some points in the PC samples. a) In one example, some points are generated by a downsampling process. iv. In one example, a reference PC sample can be the result of a process applied to at least one combined sample from multiple PC reconstruction samples. For example, the process can be such as sampling or upsampling. d. In one example, a reference PC sample can come from the same slice / frame as the current PC sample. e. Alternatively, in addition, an indication of whether to use multiple reference PC samples can be signaled to the decoder. f. In one example, reference information for the current PC sample (e.g., where the reference PC sample comes from and / or which reference PC sample to use) can be derived at the decoder. g. In one example, reference information for the current PC sample (e.g., where the reference PC sample comes from and / or which reference PC sample to use) can be signaled to the decoder. i. In one example, the reference direction can be signaled. (1) In one example, the reference direction can include: a) The reference direction can be a unidirectional prediction from a reference frame in a first reference list (denoted as L0). b) The reference direction can be a unidirectional prediction from a reference frame in a second reference list (denoted as L1). c) The reference direction can be bidirectional prediction (the first reference frame in L0 and the second reference frame in L1). (2) In one example, the relative positions of the reference frames in the reference list are fixed for a specific frame within a GOF. (4) a) In one example, the N previously decoded frames in display order (e.g., N = 2) can be used as reference frames. i. Alternatively, in addition, the indication of N can be signaled. ii. Alternatively, in addition, the N frames are consecutive previously decoded frames before a specific current frame. (4) b) In one example, the relative positions of the reference frames in the reference list can be adaptive for a specific frame in the GOF. For example, the position can be derived based on the GOF size. (3) In one example, the reference direction can be signaled conditionally, for example, according to the reference picture list information. ii. In one example, the indication of the reference frame from which the reference PC sample is derived can be signaled. (1) The indication of the reference frame can be signaled as a reference list index (L0 or L1) and a reference frame index in the reference list. a) Alternatively, it can be signaled by the reference direction and the reference frame index for each direction. (2) The reference list index can be signaled conditionally. a) If there is only one reference list, the signaling of the reference list index can be skipped. (3) The reference frame index for the reference list can be signaled conditionally. a) If there is only one reference frame in the reference list, the signaling of the reference frame index can be skipped. iii. Alternatively, the indication of the number of reference PC samples can be signaled to the decoder. iv. In addition, for the samples, at least one indication of at least one reference PC sample can be signaled to the decoder to indicate the reference relationship. (1) For example, depending on whether other samples rather than the previous sample are used as the reference PC sample, the indication can be signaled conditionally. (2) The indication can be represented by some indices (e.g., sample id) that indicate the associated samples to be used as the reference PC samples. (3) The indication can be encoded and decoded using fixed - length coding, unary coding, truncated unary coding, etc. (4) The indication can be decoded in a predictive manner. v. In one example, there is at least one reference sample. h. In one example, the geometric information of the reference PC sample can be used to perform geometric inter prediction for the current PC sample. i. In one example, the geometric information of the reference PC sample can be used to derive the predicted geometric value of the current PC sample. (1) In one example, the predicted geometric value can be selected from some candidate predictors. a) The candidate predictors can be derived from one or more geometric values of the reference sample. b) The candidate predictors can be derived as a function of one or more geometric values of the reference PC sample. c) The candidate predictors can be derived from one or more predicted geometric values of the current PC sample or previously decoded samples. d) The candidate predictors can be derived as a function of one or more predicted geometric values of the current PC sample or previously decoded samples. (2) In one example, the candidate predictors can include but are not limited to the average value, weighted average value, one of the geometric information of the reference PC sample, etc. (3) In one example, the selection of the predicted value can be based on rate optimization methods, distortion optimization methods, RDO methods, etc. (4) In one example, the selection can be derived at the decoder. (5) In one example, for each sample, an indication related to the selected predicted value can be signaled to the decoder. a) The indication can be encoded and decoded using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding. b) The indication can be encoded and decoded in a predictive manner. (6) In one example, the residual between the predicted geometric information and the actual geometric information can be derived and signaled to the decoder. a) The residual can be encoded and decoded using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. b) The residual can be encoded and decoded in a predictive manner. ii. In one example, the geometric information of the reference PC sample can be used as context information for predictive coding and decoding of the geometric information of the current node. i. In one example, the attribute information of the reference PC sample can be used to perform attribute inter prediction for the current PC sample. i. In one example, the attribute information of the reference PC sample can be used to derive the predicted attribute value of the current PC sample. (1) In one example, a predicted attribute value can be selected from a number of candidate prediction values. a) The candidate prediction values can be derived by referring to one or more attribute values of a PC sample. b) The candidate prediction values can be derived as a function of one or more attribute values of a reference PC sample. c) The candidate prediction values can be derived by one or more predicted attribute values of the current PC sample or a previously decoded sample. d) The candidate prediction values can be derived as a function of one or more predicted attribute values of the current PC sample or a previously decoded sample. (2) In one example, the candidate prediction values can include, but are not limited to, an average value, a weighted average value, one of the attribute information of a reference PC sample, etc. (3) In one example, the selection of the prediction value can be based on a rate optimization method, a distortion optimization method, an RDO method, etc. (4) In one example, the selection can be derived at the decoder. (5) In one example, for each sample, an indication related to the selected prediction value can be signaled to the decoder. a) The indication can be encoded and decoded using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. b) The indication can be encoded and decoded in a predictive manner. (6) In one example, the residual between the predicted attribute information and the actual attribute information can be derived and signaled to the decoder. a) The residual can be encoded and decoded using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. b) The residual can be encoded and decoded in a predictive manner. ii. In one example, the attribute information of a reference PC sample can be used as context information for predictive coding and decoding of the attribute information of the current node. 3) It is proposed to use at least one indication to indicate whether to enable the method of using multiple reference PC samples for a point cloud sequence. a. In one example, the indication can be signaled to the decoder. i. The indication can be encoded and decoded using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. ii. The indication can be encoded and decoded in a predictive manner. b. In one example, if the method of using multiple reference PC samples is disabled for a point cloud sequence, at most one reference PC sample can be used for inter-frame prediction of a PC sample. 4) It is proposed to use at least one indication for each PC sample to indicate whether multiple reference PC samples are used for inter - frame prediction of the current PC sample. a. In one example, the indication can be derived at the encoder. b. In one example, the indication can be derived at the decoder. c. In one example, the indication can be signaled to the decoder. i. The indication can be encoded and decoded using fixed - length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. ii. The indication can be encoded and decoded in a predictive manner. 5) It is proposed to use at least one GOF structure in a point - cloud sequence. a. In one example, frames in different GOF structures can have different reference relationships. i. In one example, except for the first frame, frames in the IPPP GOF structure can have only the previous frame as the reference frame. ii. In one example, except for the first frame, frames in the IBBB GOF structure can have two reference frames. b. In one example, one GOF structure can be applied to all GOFs in a point - cloud sequence. c. In one example, multiple GOF structures can be applied to GOFs in a point - cloud sequence. d. In one example, there can be at least one indication for indicating whether only one GOF structure is applied to all GOFs in a point - cloud sequence. i. In one example, the indication can be signaled to the decoder. (1) The indication can be encoded and decoded using fixed - length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. (2) The indication can be encoded and decoded in a predictive manner. e. In one example, if only one GOF structure is applied to all GOFs in a point - cloud sequence, there is an indication for indicating which GOF structure is applied. i. In one example, the indication can be signaled to the decoder. (1) The indication can be encoded and decoded using fixed - length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. (2) The indication can be encoded and decoded in a predictive manner. f. In one example, GOF motion information can be used to determine which GOF structure is applied to a GOF. i. In one example, GOF motion information can be derived at the encoder. (1) In one example, the GOF motion information can be the motion information between the first frame in a GOF and the first frame in the next GOF. (2) Alternatively, the GOF motion information can be the motion information between the first frame and the last frame in a GOF. (3) Alternatively, the GOF motion information can be the motion information between the first I-frame and the next I-frame in a GOF. ii. In one example, it is determined that the IBBB GOF structure is applied to a GOF only when the GOF motion information satisfies the GOF constraint condition. Otherwise, it is determined that the IPPP GOF structure is applied to the GOF. (1) In one example, the GOF motion condition can be that the GOF motion information is less than at least one threshold. a) In one example, the threshold can be derived at the encoder. b) In one example, the threshold can be predefined. iii. In one example, the decision can be made at the encoder. iv. In one example, the decision can be made at the decoder. g. In one example, if multiple GOF structures are applied to the GOFs in a point cloud sequence, there can be at least one indication for indicating which GOF structure is applied to a GOF. i. In one example, the indication can be signaled to the decoder. (1) The indication can be encoded and decoded by using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. (2) The indication can be encoded and decoded in a predictive manner. h. In the above description, a frame can be replaced by a slice / block or other processing unit. 6) In one example, information on how to manage the decoded frames can be signaled for the frames in point cloud coding. a. In one example, the decoded frames can be identified by an index counted in the display order. b. In one example, the decoded frames can be identified by an index counted in the coding / decoding order. c. In one example, it can be signaled which of the decoded frames should be kept in the frame buffer. d. In one example, it can be signaled which of the decoded frames should be removed from the frame buffer. e. In one example, it can be signaled which of the decoded frames should be used as a reference frame for a specific frame. f. In one example, it can be signaled which of the decoded frames should be put into which reference list. g. In one example, the order of reference frames can be signaled. h. In one example, information associated with a frame can be signaled. i. In one example, information independent of a frame can be signaled. 7) Samples with later timestamps in a frame can be used as reference PC samples for the current PC sample. a. In one example, there is timestamp information for each frame in a timed point cloud sequence. b. In one example, the timestamp order can be the same as the display order. c. In one example, the timestamp order can be the same as the rendering order. d. In one example, the timestamp of each sample is equal to the timestamp of the frame to which it belongs. e. In one example, samples with earlier timestamps can be used as reference PC samples for the current PC sample. f. In one example, samples with the same timestamp can be used as reference PC samples for the current PC sample. g. In one example, samples with later timestamps can be used as reference PC samples for the current PC sample. h. Alternatively, in addition, an indication of whether samples with later timestamps are allowed to be used as reference PC samples can be signaled to the decoder. 8) In one example, there can be a low-latency mode for point cloud compression. a. Alternatively, in addition, with the low-latency mode, the timestamp order and the decoding order must be the same. b. Alternatively, in addition, an indication of whether to use the low-latency mode can be signaled to the decoder. c. Alternatively, in addition, multiple reference frames can be used for the low-latency mode. 9) It is proposed to use samples with earlier timestamps or the same timestamp in a frame as reference samples in the low-latency mode. a. In one example, samples with earlier timestamps can be used as reference PC samples for the current PC sample in the low-latency mode. b. In one example, samples with the same timestamp can be used as reference PC samples for the current PC sample in the low-latency mode. 10) It is proposed to perform the encoding and decoding processes based on the reference relationship rather than the timestamp order of the samples. a. In one example, the reference PC sample can be encoded before the current PC sample. b. In one example, the reference PC sample can be decoded before the current PC sample. 11) In one example, the timestamp order of each PC sample can be signaled to the decoder. a. In one example, the timestamp order of each PC sample can be different from the encoding / decoding order of each PC sample. b. In one example, the timestamp order can be in the form of consecutive increasing integers. c. In one example, the timestamp order can be directly signaled to the decoder. i. The timestamp order can be encoded / decoded using fixed-length coding / decoding, unary coding / decoding, truncated unary coding / decoding, etc. ii. The timestamp order can be encoded / decoded in a predictive manner. d. In one example, the timestamp order can be indirectly signaled to the decoder. i. In one example, the relative timestamp order can be derived at the encoder. (1) In one example, the relative timestamp order of a PC sample can be derived based on the timestamp order of the current PC sample and the timestamp order of another specific PC sample. (2) In one example, the specific PC sample can be before the current PC sample in the encoding / decoding order. (3) In one example, the specific PC sample can be the previous PC sample in the encoding / decoding order. (4) In one example, the specific PC sample can be a previous PC sample in the encoding / decoding order that satisfies certain characteristics. a) In one example, the specific PC sample can be a previous PC sample for which inter-frame prediction is disabled. b) In one example, the specific PC sample can be a previous PC sample that uses only one reference PC sample in inter-frame prediction. ii. In one example, the relative timestamp order can be signaled to the decoder. (1) The relative timestamp order can be encoded / decoded using fixed-length coding / decoding, unary coding / decoding, truncated unary coding / decoding, etc. (2) The relative timestamp order can be encoded / decoded in a predictive manner. iii. In one example, the timestamp order can be derived at the decoder based on the relative timestamp order. 12) For inter-frame coded slices / frames for which inter-frame prediction is enabled, the information of the reference frame can be signaled. a. The information of the reference frame may include, i. The number of reference frames. ii. The number of reference lists. iii. The number of reference frames in each reference list. iv. The reference frames in each reference list. (1) The reference frame can be indicated by its timestamp or POC or other means. b. This information can be shared by multiple frames, for example, transmitted by signaling in a higher-level syntax structure (e.g., in SPS / PPS). 13) It is proposed to encode and decode PC samples in a different order rather than in a fixed order, and / or to encode and decode PC samples with different encoding and decoding precisions. a. In one example, it is proposed to apply hierarchical encoding and decoding precision based on the encoding and decoding priorities of PC samples. i. In one example, PC samples in a point cloud sequence can have different encoding and decoding priorities. b. In one example, the encoding and decoding priority of the reference PC sample should be higher than that of the current PC sample. c. In one example, the encoding and decoding precision of samples with higher encoding and decoding priorities should be higher than that of samples with lower encoding and decoding priorities. d. In one example, the encoding and decoding precision can be controlled by the QP value / quantization step size in the point cloud sequence encoding and decoding. e. In one example, the QP / quantization step size value for the reference PC sample can be lower / smaller than that of the current PC sample. f. In one example, the difference value of the QP / quantization step size value for the reference PC sample can be fixed. g. In one example, the difference value of the QP / quantization step size value for the reference PC sample can be derived at the decoder. i. In one example, the difference value of the QP / quantization step size value for the reference PC sample can be derived based on the GOF size. ii. In one example, the difference value of the QP / quantization step size value for the reference PC sample can be derived based on the intra period / random access period. iii. In one example, the difference value of the QP / quantization step size value for the reference PC sample can be derived based on the indicator of the lossless encoding and decoding mode. iv. In one example, the difference value of the QP / quantization step size value for the reference PC sample can be derived based on the indicator of the low-latency encoding and decoding mode. h. In one example, the difference value for the QP / quantization step size value with respect to the reference PC sample can be signaled to the decoder. i. In the above example, the "reference PC sample" can be replaced by the "current PC sample". j. Alternatively, in addition, an indication of whether to use hierarchical QP values and / or QP values / quantization steps is signaled to the decoder. k. In one example, the QP value for each sample can be derived at the decoder. 14) In one example, the QP value for each frame / block / cube / slice is signaled to the decoder. a. The QP value can be encoded and decoded using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. b. The QP value can be encoded and decoded in a predictive manner. 15) It is proposed that when using octree geometry coding and decoding, the occupancy information of multiple reference nodes is used to perform inter-frame prediction on the current node. a. In one example, when using octree geometry coding and decoding, the geometric information can be represented by an octree structure and the occupancy information of octree nodes (such as occupancy codes). b. In one example, for each frame, there can be multiple reference frames. c. Alternatively, in addition, an indication of whether to use multiple reference frames can be signaled to the decoder. d. In one example, for each node, there can be at least one corresponding reference node in each reference frame. e. In one example, for each node, the reference occupancy code can be selected from some candidate values. i. The candidate values can be derived from the occupancy information of one or more reference nodes. ii. The candidate values can be derived as a function of the occupancy information of one or more reference nodes. iii. In one example, the candidate values can include, but are not limited to, the XOR of the occupancy information of reference nodes, the occupancy information of the same reference node, or one of the occupancy information of reference nodes. iv. In one example, the selection of the candidate values can be based on rate optimization methods, distortion optimization methods, RDO methods, etc. v. In one example, the selection can be derived at the decoder. vi. In one example, an indication related to the selected candidate values is signaled to the decoder. (1) The indication can be encoded and decoded using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. (2) This indication can be encoded and decoded in a predictive manner. f. In one example, for each node, a reference occupancy code can be used as predictive occupancy information. i. Additionally, the residual between the predictive occupancy information and the actual occupancy information can be derived and signaled to the decoder. (1) The residual can be encoded and decoded using fixed-length encoding and decoding, unary encoding, truncated unary encoding, etc. (2) The residual can be encoded and decoded in a predictive manner. g. In one example, for each node, the reference occupancy information can be used as context information for predictive encoding and decoding of the occupancy information of the current node. 16) It is proposed to derive a selection of reference occupancy information for child nodes based on the current node and the reference node of the current node when using octree geometry encoding and decoding. a. In one example, when using octree geometry encoding and decoding, the geometry information can be represented by the octree structure and the occupancy information (e.g., occupancy code) of the octree nodes. b. In one example, for each node, there can be an occupancy code, which is an 8-bit binary number. Each bit corresponds to a child node. c. In one example, for each node, there can be multiple reference nodes and corresponding occupancy codes. d. In one example, for each node, there can be a reference occupancy code. e. In one example, for each node, the reference occupancy code can be selected from one of the occupancy codes of the reference nodes. f. In one example, for each node, the selection of the reference occupancy code of the child node can be derived based on the occupancy codes of the current node and the reference node of the current node. i. In one example, for each bit in the occupancy codes of the current node and the reference node of the current node: (1) If the bit values at the same bit position for the current node and a reference node are the same, then the occupancy code of the child node of that reference node can be selected as the reference occupancy code of the child node of the current node. The child node corresponds to that bit position. ii. In one example, calculate the number of bits that do not match between the occupancy code of the current node and the occupancy code of the reference node of the current node: (1) If the number of bits that do not match is the same for all reference nodes, then the selection of the child node can inherit the selection of the current node. (2) If the number of bits that do not match is not the same for all reference nodes, then the occupancy code of the child node of the reference node with the fewest number of mismatches can be selected as the reference occupancy code of the child node. 17) It is proposed to derive the cumulative global motion based on at least one externally estimated global frame for at least one frame. a. In one example, the global motion can be estimated between an externally estimated frame and its subsequent frame. i. In one example, the global motion can be estimated as a pre - processing before point cloud compression. ii. In one example, the global motion can be part of the original data. b. In one example, the externally estimated global motion can be used in the global motion estimation. c. In one example, when the frame distance between the current frame and the reference frame is greater than a threshold (such as 1), the cumulative global motion can be used instead of the externally estimated global motion. d. In one example, the cumulative global motion can be derived at the encoder. e. In one example, the cumulative global motion between the reference frame and the current frame can be derived based on the reference frame and the externally estimated global motions of consecutive frames before the current frame in timestamp order. f. In one example, the cumulative global motion between the reference frame and the current frame can be derived based on the current frame and the externally estimated global motions of one or more consecutive frames before the reference frame in timestamp order. g. In one example, the cumulative global motion can be signaled to the decoder. i. The cumulative global motion can be encoded and decoded using fixed - length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. ii. The cumulative global motion can be encoded and decoded in a predictive manner. iii. The cumulative global motion can be encoded and decoded through context - based coding. iv. The cumulative global motion can be encoded and decoded through bypass coding. 18) It is proposed to derive the inter - attribute frame threshold based on the frame distance of at least one reference frame. a. In one example, for a reference frame, there can be at least one inter - attribute frame threshold to determine whether inter - attribute frame prediction is applied to the reference frame. b. In one example, for a reference frame, the inter - attribute frame threshold can be derived based on the original inter - attribute frame threshold and the frame distance of the reference frame. c. In one example, the requirements for the inter - attribute frame threshold of a reference frame with a larger frame distance can be more stringent than those of a reference frame with a smaller frame distance. i. In one example, the inter - attribute frame threshold can be the original inter - attribute frame threshold divided by the frame distance. d. The threshold can be derived at the decoder or can be signaled from the encoder to the decoder. 19) It is proposed that the search range for attribute inter - frame prediction can be based on the reference relationship. a. The search range for a sample with multiple reference samples can be smaller than that for a sample with one reference sample. i. The search range for a sample with one reference sample can be indicated by an integer (e.g., N); the search range for a sample with M reference samples can be indicated by a smaller integer (e.g., N / M). b. Alternatively, the search range for a sample with multiple reference samples can be larger than that for a sample with one reference sample. c. Alternatively, the search range for a sample with multiple reference samples can be equal to that for a sample with one reference sample. d. Alternatively, the search range can be signaled from the encoder to the decoder. 20) It is proposed to perform geometric inter - frame prediction using one or more reference frames on a set of layers of an octree structure. a. In one example, the set is the first N layers of the octree structure. b. In one example, the set is the last N layers of the octree structure. c. In one example, geometric encoding and decoding can be performed in an octree structure with multiple layers. d. In one example, geometric intra - prediction encoding and decoding can be performed on all layers of the octree structure. e. In one example, geometric inter - frame prediction encoding and decoding with one reference frame can be performed on all layers of the octree structure. f. In one example, geometric inter - frame prediction encoding and decoding with multiple reference frames can be performed on the first N layers of the octree structure. N can be a non - negative integer. i. In one example, N can be a predefined value. ii. In one example, N can be derived at the encoder. (1) N can be derived based on the node size of each layer. (2) N can be derived based on the motion block size. iii. In one example, N can be derived at the decoder. (1) N can be derived based on the node size of each layer. (2) N can be derived based on the motion block size. iv. In one example, N can be signaled to the decoder. (1)N can be encoded and decoded using fixed - length encoding and decoding, unary encoding and decoding, truncated unary encoding and decoding, etc. (2)N can be encoded and decoded in a predictive manner. 21) It is proposed that if the reconstructed sample of a PC is a reference PC sample for other PC samples, the reconstructed sample of the PC is temporarily recorded. a. In one example, a PC sample can be reconstructed at the encoder and / or decoder. b. In one example, a reconstructed sample of a PC can be a reference PC sample for other PC samples. c. In one example, if the reconstructed sample of a PC is a reference PC sample for other PC samples, some memory can be used to record the reconstructed sample of a PC when processing other PC samples. d. In one example, if the reconstructed sample of a PC is not a reference PC sample for any other PC samples, the memory for recording the reconstructed sample of a PC can be released. e. In one example, for each PC sample, there can be at least one indication to indicate whether the PC sample is a reference PC sample for other PC samples. i. In one example, the indication can be a flag for indicating whether the PC sample is a reference PC sample for other PC samples. (1) In one example, the flag can be derived at the encoder. (2) In one example, the flag can be derived at the decoder. (3) Alternatively, the flag can be signaled to the decoder. ii. In one example, the indication can be the number of PC samples using the current PC sample as a reference PC sample. (1) In one example, the number can be derived at the encoder. (2) In one example, when other PC samples are encoded and decoded, the number can change. For example, after a PC sample is referenced by another PC sample, the number can be reduced by 1. If the number becomes zero, the memory for recording the corresponding PC sample can be released. (3) In one example, the number can be derived at the decoder. (4) In one example, the number can be signaled to the decoder. 22) It is proposed to select a reference point from different reference PC samples for the current PC point to perform attribute inter - frame prediction. a. In one example, for each point, there can be multiple reference points to perform attribute inter - frame prediction. b. In one example, a reference point can be selected from multiple samples based on the geometric distance between the reference point and the current point. c. In one example, the attribute information of the reference point can be used to derive the predicted attribute value of the current point. d. In one example, the predicted attribute value can be selected from some candidate predicted values. i. The candidate predicted values can be derived from the attribute information of the reference points from one or more reference PC samples. ii. The candidate predicted values can be derived as a function of the attribute information of the reference points from one or more reference PC samples. iii. In one example, the candidate predicted values can include, but are not limited to, the average value, the weighted average value, one of the attribute information of the reference points, etc. (1) In one example, the weight of each reference point can be the geometric distance from the current point. iv. In one example, the selection of the predicted value can be based on rate optimization methods, distortion optimization methods, RDO methods, etc. v. In one example, for each point, an indication related to the selected predicted value can be signaled to the decoder. (1) The indication can be encoded and decoded using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. (2) The indication can be encoded and decoded in a predictive manner. e. In one example, the residual between the predicted attribute value and the actual attribute value can be derived and signaled to the decoder. i. The residual can be encoded and decoded using fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. ii. The residual can be encoded and decoded in a predictive manner. f. In one example, the predicted attribute value can be used as context information for predictive coding and decoding of the attribute information of the current point. 6. Embodiment 1) This embodiment describes how to use two reference frames and perform inter-frame prediction on the current frame with the frame having a later timestamp as the reference frame. In this example, the point cloud frames in the point cloud sequence are divided into multiple GOFs, and the GOF size is set to 8. As Figure 5 shown, for each GOF, there are 8 consecutive frames in timestamp order. Figure 5 The numbers in The first frame is a random access (RA) point, which means that only intra-frame prediction coding and decoding exist but inter-frame prediction coding and decoding do not exist. For the remaining 7 frames, both intra-frame prediction coding and decoding and inter-frame prediction coding and decoding will be performed based on the reference relationship. As Figure 6 shown, a hierarchical reference relationship is applied for each group of pictures (GOP). In Figure 6 , frame "8" is the first frame of the next GOP. For frames "1" to "7", each frame has two reference frames. One reference frame has a timestamp earlier than the current frame, while the other reference frame has a timestamp later than the current frame. For each frame, the reference frames are shown in Table 1. Table 1 Reference frames for each frame in a GOP Frame timestamp 0 1 2 3 4 5 6 7 Reference frame timestamp None 0,2 0,4 2,4 0,8 4,6 4,8 6,8 To ensure that the reference frames are encoded and decoded before the current frame, the encoding and decoding order for frames "0" to "8" is {0, 8, 4, 2, 1, 3, 6, 5, 7}. It should be noted that frame "8" is the first frame of the next GOP, but it should be processed before frames "1" to "7". And if frame "8" is encoded or decoded in the current GOP, then the processing of frame "8" (which is also frame "0" in the next GOP) should be skipped in the next GOP. 2) This embodiment describes how to apply hierarchical coding precision to attribute inter-frame prediction based on the coding and decoding priority of samples. In the example, the point cloud frames in the point cloud sequence are divided into multiple GOPs, and the GOP size is set to 8. As Figure 6 shown, a hierarchical reference relationship is applied. The hierarchical coding and decoding priority is calculated based on the principle that the reference frame has a higher coding and decoding priority than the current frame. The coding and decoding priority results are shown in Table 2. Table 2 Coding and decoding priorities for each frame in a GOP Frame timestamp 0 1 2 3 4 5 6 7 8 Codec priority 4 1 2 1 3 1 2 1 4 In the example, the quantization parameter (QP) value is used to control the coding and decoding precision. The lower the QP value, the higher the coding and decoding precision. Therefore, a hierarchical QP value structure is applied to the frames so that the coding and decoding precision can be changed based on the coding and decoding priority. For each frame, the QP value is calculated as follows: QP real = QP original + QP_shift The QP_shift value for each frame is shown in Table 3. Table 3 QP_shift values for each frame in a GOP Frame timestamp 0 1 2 3 4 5 6 7 8 QP_shift 0 +3step +2step +3step +step +3step +2step +3step 0 The parameter step is a non - negative number, which is used to control the change ratio of the hierarchical QP value. During testing, it can be 2 / 3 / 4, etc. For example, when step is set to 3, the QP_shift values for each frame are shown in Table 4. Table 4 QP_shift values for each frame in a GOF when step = 3 Frame timestamp 0 1 2 3 4 5 6 7 8 QP_shift 0 +9 +6 +9 +3 +9 +6 +9 0 3) This embodiment describes an example of how to perform geometric inter - prediction using two reference frames when using octree - based geometric coding and decoding. In this example, the geometric information is represented by an octree structure and the occupancy code of each node. As Figure 6 shown, a hierarchical reference relationship is applied. For each frame, there are two reference frames used for geometric inter - prediction. At the encoder, the same octree partitioning is performed on the current frame and the reference frames. Therefore, the octree structures of the current frame and the reference frames are the same. A FIFO queue is used to store the nodes to be processed. For each node, a boolean flag predicted_forward is used to indicate the source of the reference occupancy code: a. If the reference occupancy code is the occupancy code of a node in the reference frame with an earlier timestamp, then predicted_forward is set to 1. b. If the reference occupancy code is the occupancy code of a node in the reference frame with a later timestamp, then predicted_forward is set to 0. The parameter mismatched_count_parent_node is used to indicate the number of bits by which the occupancy code of the parent node does not match the reference occupancy code of the parent node. First, the root node of the octree of the current node is generated and pushed into the queue. The predicted_forward value of the root node is set to 1. The mismatched_count_parent_node value of the root node is set to 0. Second, the following process is performed until the queue is empty: a. Obtain the node at the head of the queue and its corresponding forward reference node and backward reference node. The forward reference node and the backward reference node are the nodes in the two reference frames that share the same position in the octree structure as the current node. The first reference node is in the reference frame with an earlier timestamp, and the other reference node is in the reference frame with a later timestamp. b. Calculate the occupancy codes OC for the current node, the forward reference node, and the backward reference nodecurrent 、 OC forward 、 OC backward 。 c. Count the number of bits that do not match between the current OC current and OC forward as mismatched_count_forward. Count the number of bits that do not match between the current OC current and OC backward as mismatched_count_backward. d. If predicted_forward of the current node is 1, then the reference occupancy code OC reference is set to OC forward ; otherwise, OC reference is set to OC backward 。 e. If mismatched_count_parent_node of the current node is greater than 4, then OC reference is set to all zeros. f. Use OC reference and the intra-frame prediction result as part of the context to perform predictive coding on OC current 。 g. If the current node can be divided into 8 child nodes, then for each bit in OC current and its corresponding child node: i. If there is no point in the child node, skip it. Otherwise, go to the next step. ii. Obtain the corresponding bits in OC forward and OC backward as bit_0 and bit_1 respectively. iii. If bit_0 = 1 and bit_1 = 1, then the predicted_forward value of the child node is set to 0, and the mismatched_count_parent_node value of the child node is set to mismatched_count_backward. iv. If bit_0 = 1 and bit_1 = 0, then the predicted_forward value of the child node is set to 1, and the mismatched_count_parent_node value of the child node is set to mismatched_count_forward. v. If bit_0 = 0 and bit_1 = 0 or bit_0 = 1 and bit_1 = 1: (1) If mismatched_count_backward < mismatched_count_for- ward, the predicted_forward value of the child node is set to 0, and the mismatched_count_parent_node value of the child node is set to mismatched_count_backward. (2) If mismatched_count_backward > mismatched_count_for- ward, the predicted_forward value of the child node is set to 1, and the mismatched_count_parent_node value of the child node is set to mismatched_count_forward. (3) If mismatched_count_backward = mismatched_count_for- ward, the predicted_forward value of the child node is set to the predicted_forward value of the current node and the mismatched_count_par- ent_node value of the child node is set to mismatched_count_forward. vi. Push the child node into the queue. h. Pop the current node from the queue. At the decoder, the same processing is performed on the current frame and the reference frame. Thus, the reference occupancy code can be derived for each node. The occupancy code can be decoded based on the reference occupancy code. 4) This embodiment describes an example of how to perform geometric inter-frame prediction using two reference frames when using octree geometry encoding and decoding. In this example, the geometric information is represented by an octree structure and the occupancy code of each node. As Figure 6 shown, a hierarchical reference relationship is applied. For each frame, there are two reference frames for geometric inter-frame prediction. At the encoder, the same octree partitioning is performed on the current frame and the reference frame. Thus, the octree structure is the same for the current frame and the reference frame. A FIFO queue is used to store the nodes to be processed. A boolean flag predicted_forward is used for each node to indicate the source of the reference occupancy code: c. If the reference occupancy code is the occupancy code of a node in the reference frame with an earlier timestamp, then The predicted_forward is set to 1. d. If the reference occupancy code is the occupancy code of a node with a later timestamp in the reference frame, then set predicted_forward to 0. The parameter mismatched_count_parent_node is used to indicate the number of bits that do not match between the occupancy code of the parent node and the reference occupancy code of the parent node. First, generate the root node of the octree for the current node and push it onto the queue. The predicted_forward value of the root node is set to 1. The mismatched_count_parent_node value of the root node is set to 0. Second, perform the following process until the queue is empty: i. Obtain the node at the head of the queue and its corresponding forward reference node and backward reference node. The forward reference node and the backward reference node are nodes that share the same position in the octree structure with the current node in two reference frames. The first reference node is in the reference frame with an earlier timestamp, and the other reference node is in the reference frame with a later timestamp. j. Calculate the occupancy codes OC current 、OC forward 、OC backward for the current node, the forward reference node, and the backward reference node. k. Count the number of bits that do not match between the current OC current and OC forward as mismatched_count_forward. Count the number of bits that do not match between the current OC current and OC backward as mismatched_count_backward. l. If the predicted_forward of the current node is 1, then the reference occupancy code OC reference is set to OC forward ; otherwise, OC reference is set to OC backward . m. If mismatched_count_parent_node is greater than 4, then OC reference is set to all zeros. n. Use OC reference and the intra-frame prediction result as part of the context to perform predictive coding and decoding on OC current . o. If the current node can be divided into 8 child nodes, then for each bit in OC current and its corresponding child node: i. If there is no point in the child node, skip it. Otherwise, go to the next step. ii. Obtain the corresponding bits in OC forward and OC backward which are bit_0 and bit_1 respectively. iii. If bit_0 = 1 and bit_1 = 1, the predicted_forward value of the child node is set to 0, and the mismatched_count_parent_node value of the child node is set to mismatched_count_backward. iv. If bit_0 = 0 and bit_1 = 0, the predicted_forward value of the child node is set to 1, and the mismatched_count_parent_node value of the child node is set to mismatched_count_forward. v. If bit_0 = 0 and bit_1 = 0 or bit_0 = 1 and bit_1 = 1: (1) If mismatched_count_backward < mismatched_count_for- ward, the predicted_forward value of the child node is set to 0, and the mismatched_count_parent_node value of the child node is set to mismatched_count_backward. (2) If mismatched_count_backward > mismatched_count_for- ward, the predicted_forward value of the child node is set to 1, and the mismatched_count_parent_node value of the child node is set to mismatched_count_forward. (3) If mismatched_count_backward = mismatched_count_for- ward, the predicted_forward value of the child node is set to the predicted_forward value of the current node, and the mismatched_count_par- The value of ent_node is set to mismatched_count_forward. vi. Push the child node into the queue. p. Pop the current node from the queue. At the decoder, the same processing is performed on the current frame and the reference frame. Therefore, the reference occupancy code can be derived for each node. The occupancy code can be decoded based on the reference occupancy code. 5) This embodiment describes an example of how to perform attribute inter-frame prediction using two reference frames. In this example, the attribute information is represented by the reflection value of each point. As Figure 6 shown, a hierarchical reference relationship is applied. For each frame, there are two reference frames for attribute inter-frame prediction. At the encoder, three reference points {point 0, point 1, point 2} will be selected from the current frame and the reference frame. The predicted attribute value will be calculated based on the attribute values of the reference points. Then, the residual between the predicted attribute value and the current attribute value will be calculated and signaled to the decoder. For each point, the array neighbors is used to record the weight values of the selected reference points. The weight value of each reference point is the distance between the reference point and the current point. First, the points in the current frame and the reference frame are reordered in the order of the motion code. Second, for each point, the encoder will search for the three reference points closest to the current point. The search results and their weight values will be stored in neighbors: a. Perform reference point search on the current frame: i. Scan the warp-decoded points within the same motion code level. ii. Scan the warp-decoded points within the search range. The search range is defined by a parameter. b. The reference point search is performed on the reference frame: iii. Scan the warp-decoded points within the same motion code level. iv. Scan the warp-decoded points within the search range. The search range is defined by a parameter. Third, recalculate the weight values of the reference points. The reference points from the current frame should have higher weight values. Fourth, the predicted attribute value will be selected from the candidate list: a. The weighted average of the attribute values of the reference points. b. The attribute value of reference point 0. c. The attribute value of reference point 1. d. The attribute value of reference point 2. For each candidate value, an encoding / decoding score will be calculated based on the compressed bits and the prediction residual. Then the encoder will select the candidate value with the highest encoding / decoding score. An indication related to the selected candidate will be signaled to the decoder. Finally, the residual between the attribute value and the predicted attribute value will be calculated and signaled to the decoder. At the decoder, the reference points for each point will be searched for by the same method as in the encoding process. The candidate list will be calculated in the same way, and the indication will be decoded for each point to obtain the predicted attribute value. Based on this, the prediction residual will be decoded and the actual attribute value will be generated. 6) This embodiment describes an example of how to perform inter-frame prediction for both geometric encoding / decoding and attribute encoding / decoding using two reference frames. A hierarchical GOF structure is proposed to perform inter-frame prediction for geometric encoding / decoding and attribute encoding / decoding. In the hierarchical GOF structure, the first frame in each GOF is an I frame. The other frames in the GOF are B frames, which means the frames will use two reference frames in both the forward and backward directions. As Figure 7 shown, frames "0" to "7" are the frames in one GOF, and frame "8" is the first frame of the next GOF. For frames "0" to "8", the reference frames are shown in Table 5. Table 5 Reference frames for each frame in one GOF Frame timestamp 0 1 2 3 4 5 6 7 8 Reference frame timestamp None 0,2 0,4 2,4 0,8 4,6 4,8 6,8 None The encoding and decoding order for frames "0" to "8" is {0, 8, 4, 2, 1, 3, 6, 5, 7}. For geometric encoding / decoding, the same octree partitioning is performed on the current frame and the two reference frames. For each node in the octree, the occupancy codes of the current node and the reference node are calculated. As Figure 8 shown, the prediction direction of the child nodes of the current node is derived based on the occupancy codes of the current node and the reference node. For each child node of the current node, the corresponding bit values in the occupancy code of the reference node are denoted as bit_pre and bit_follow: If bit_pre = 1 and bit_follow = 0, the prediction direction of the child node is set to use the previous reference (forward) node to perform inter-frame prediction. If bit_pre = 0 and bit_follow = 1, the prediction direction of the child node is set to use the subsequent reference (backward) node to perform inter-frame prediction. If bit_pre = bit_follow, calculate the number of mismatched bits between the occupancy code of the current node and the occupancy code of the reference node. If the number of mismatched bits is different, set the prediction direction of the child node to the prediction direction with fewer mismatches. Otherwise, set the prediction direction of the child node to the prediction direction of the current node. When encoding / decoding an attribute, select three reference points {point 0, point 1, point 2} from the current frame and two reference frames. The predicted attribute value will be calculated based on the attribute values of the reference points, which is similar to Inter-EM. In addition, a hierarchical QP structure is applied to perform attribute encoding / decoding. Based on the reference relationship, there is a QP shift value for each frame. The QP shift value for the reference frame should be lower than the QP shift value of the current frame. For each frame, the true attribute QP value is set to: QP original +QP shift The quantization process is performed based on the true attribute QP value. 7) This embodiment describes an example of how to perform inter-frame prediction for both geometric encoding / decoding and attribute encoding / decoding by merging two reference frames. A hierarchical GOF structure is proposed to perform inter-frame prediction for geometric encoding / decoding and attribute encoding / decoding. In the hierarchical GOF structure, the first frame in each GOF is an I frame. The other frames in the GOF are B frames, which means the frames will use two reference frames in both the forward and backward directions. As Figure 7 shown, frames "0" to "7" are the frames in one GOF, and frame "8" is the first frame of the next GOF. For frames "0" to "8", the reference frames are shown in Table 5. The encoding and decoding order of frames "0" to "8" is {0, 8, 4, 2, 1, 3, 6, 5, 7}. For each frame, first apply the global motion to the two reference frames. Then merge all the points in the two reference frames into a new merged reference frame. For geometric encoding / decoding, perform the same octree partitioning on the current frame and the merged reference frame. Then perform geometric inter-frame prediction on the current frame and the merged reference frame. When encoding / decoding an attribute, select three reference points {point 0, point 1, point 2} from the current frame and the merged reference frame. The predicted attribute value will be calculated based on the attribute values of the reference points, which is similar to Inter-EM. In addition, a hierarchical QP structure is applied to perform attribute encoding and decoding. Based on the reference relationship, there is a QP shift value for each frame. The QP shift value for the reference frame should be lower than the QP shift value of the current frame. For each frame, the true attribute QP value is set to: QP original + QP shift The quantization process is performed based on the true attribute QP value. 8) This embodiment describes an example of how to determine which GOF structure to apply to a GOF. Figure 7 An example of an IBBB GOF structure is shown in Figure 7 . For each frame except the first frame in a GOF, there are two reference frames. Figure 9 An example of an IPPP GOF structure is shown in Figure 9 . For each frame except the first frame in a GOF, there is one reference frame. Frame 8 is the first frame in the next GOF. At the encoder, for each GOF, frame 0 is processed first. If the GOF is the first GOF in a point cloud sequence, frame 0 is encoded or decoded. Otherwise, frame 0 is skipped because it has been encoded or decoded when processing the previous GOF. Then, frame 8 is processed, and the motion information between frame 0 and frame 8 is derived. The rotation degrees (Rx, Ry, Rz) and the translation vector (Sx, Sy, Sz) are derived based on the motion information. If the motion information satisfies two conditions, the IBBB GOF structure is applied to the GOF. Otherwise, the IPPP GOF structure is applied to the GOF. (1) All rotation degrees are less than thr1: thr1 = 0.1 * random_access_period where random_access_period is a parameter used to indicate the minimum frame distance between two I-frames. (2) The translation vector is less than thr2: thr2 = 0.005 * (2 * quantization_bits + 1) * slice_size where quantization_bits is a parameter used to indicate the geometric quantization scale, and slice_size is a parameter used to indicate the bounding box size of frame 8. There is a signal change_GOF_structure for indicating the GOF structure selection result, and this signal is transmitted to the decoder through signal transmission. If the IBBB GOF structure is applied, change_GOF_structure is set to 0. Otherwise, change_GOF_structure is set to 1. At the decoder, first the IBBB GOF structure is applied to each GOF. Only when change_GOF_structure is equal to 1, the IPPP GOF structure is applied during the decoding process of the GOF.
[0059] More details of embodiments of the present disclosure related to multi-reference inter-frame prediction for point cloud coding and decoding will be described below. Embodiments of the present disclosure should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. In addition, these embodiments can be applied individually or in combination in any way.
[0060] As used herein, the term "point cloud sequence" may refer to a sequence of one or more point clouds. The term "point cloud frame" or "frame" may refer to a point cloud in a point cloud sequence. The term "point cloud (PC) sample" may refer to a frame, picture, slice, sheet, sub-picture, node, point, or unit containing one or more nodes or points.
[0061] Figure 10 A flowchart of a method 1000 for point cloud coding and decoding according to some embodiments of the present disclosure is shown. Method 1000 can be implemented during the conversion between the current PC sample of a point cloud sequence and the bitstream of the point cloud sequence. As Figure 10 shown, method 1000 starts at 1002, where a first indication indicating whether multi-reference inter-frame prediction is enabled for a point cloud sequence is obtained, and multiple reference PC samples are used in the multi-reference inter-frame prediction. By way of example and not limitation, if multi-reference inter-frame prediction is used for a PC sample, multiple reference PC samples can be used to encode and decode the PC sample. Refer Figure 7 , frame 1 and frame 8 are used as reference frames for frame 4.
[0062] In some embodiments, the first indication may be determined at the encoder and included in the bitstream. At the decoder, the first indication may be obtained from the bitstream. By way of example and not limitation, the first indication may be a syntax element, an index, a flag, etc. It should be noted that the first indication may be implemented as a single indication, multiple indications, or a combination of multiple indications. In one example, the first indication may be encoded and decoded using fixed-length coding and decoding. In another example, the first indication may be encoded and decoded using unary coding. In another example, the first indication may be encoded and decoded using truncated unary coding. Alternatively, the first indication may be encoded and decoded in a predictive manner. It should be understood that the above diagrams and / or examples are described only for purposes of illustration. The scope of the present disclosure is not limited in this regard.
[0063] At 1004, a transformation is performed based on the first indication. In some embodiments, the transformation may include encoding the current PC sample into the bitstream. Alternatively or additionally, the transformation may include decoding the current PC sample from the bitstream.
[0064] In view of the foregoing, the transformation between the point cloud sequence and the bitstream is performed based on an indication indicating whether multi-reference inter prediction is enabled for the point cloud sequence. In this way, the proposed method can advantageously better support the use of multi-reference inter prediction and facilitate the application of multi-reference inter prediction, and thus can improve the coding and decoding quality of point cloud coding and decoding.
[0065] In some embodiments, if the first indication indicates that multi-reference inter prediction is disabled for the point cloud sequence, a single reference PC sample may be allowed to be used for performing inter prediction on the current PC sample. That is, at most one reference PC sample is allowed to be used for performing inter prediction on the current PC sample. Alternatively, the current PC sample is encoded and decoded based on any other prediction process other than inter prediction (e.g., intra prediction).
[0066] In some embodiments, at 1004, a second indication indicating whether multi-reference inter prediction is used for the current PC sample may be obtained. Further, a transformation is performed based on the first indication and the second indication. By way of example and not limitation, if the first indication indicates that multi-reference inter prediction is enabled for the point cloud sequence, and the second indication indicates that multi-reference inter prediction is used for the current PC sample, the current PC sample may be encoded and decoded based on multi-reference inter prediction by using multiple reference PC samples.
[0067] In some embodiments, the second indication may be determined at the encoder and included in the bitstream. At the decoder, the second indication may be obtained from the bitstream. By way of example and not limitation, the second indication may be a syntax element, an index, a flag, etc. It should be noted that the second indication may be implemented as a single indication, multiple indications, or a combination of multiple indications. In one example, the second indication may be encoded and decoded using fixed-length coding and decoding. In another example, the second indication may be encoded and decoded using unary coding. In another example, the second indication may be encoded and decoded using truncated unary coding. Alternatively, the second indication may be encoded and decoded in a predictive manner. It should be understood that the above diagrams and / or examples are described only for purposes of illustration. The scope of the present disclosure is not limited in this regard.
[0068] In some alternative embodiments, the second indication may be determined at the decoder. In one example, the second indication may be determined based on global motion information, reference structure, etc.
[0069] In some embodiments, the point cloud sequence may include a plurality of PC samples. The position of the current PC sample in the timestamp order of the plurality of PC samples is included in the bitstream. By way of example and not limitation, the timestamp order may be in the form of continuously increasing integers. In one example, the position may be encoded and decoded using fixed-length coding and decoding. In one example, the position may be encoded and decoded using unary coding. In another example, the position may be encoded and decoded using truncated unary coding. Alternatively, the position may be encoded and decoded in a predictive manner. It should be understood that the above diagrams and / or examples are described only for purposes of illustration. The scope of the present disclosure is not limited in this regard.
[0070] In some alternative embodiments, a third indication indicating the position of the current PC sample in the timestamp order of the plurality of PC samples may be included in the bitstream. For example, the position may be indirectly signaled to the decoder. By way of example and not limitation, the plurality of PC samples may include another PC sample different from the current PC sample, and the third indication may include an offset that depends on the position of the current PC sample and the position of the other PC sample in the timestamp order.
[0071] In some embodiments, another PC sample may be before the current PC sample in the decoding order of the plurality of PC samples. The decoding order may be different from the timestamp order. Alternatively, the other PC sample may be immediately before the current PC sample in the decoding order.
[0072] In some alternative embodiments, another PC sample may be immediately before the current PC sample in a set of PC samples of a point cloud sequence that meets one or more specific conditions. In one example, each sample in a set of PC samples meets one of the following conditions: inter-frame prediction is disabled for the corresponding PC sample, or the corresponding PC sample is encoded and decoded based on inter-frame prediction using a single reference PC sample. In another example, inter-frame prediction is disabled for each sample in a set of PC samples. In another example, each PC sample in a set of PC samples is encoded and decoded based on inter-frame prediction using a single reference PC sample.
[0073] It should be noted that the current PC sample may be included in or excluded from a set of PC samples. Another PC sample may be one of the set of PC samples that is immediately before the current PC sample in the encoding and decoding order. It should be understood that the above diagrams and / or examples are described only for the purpose of description. The scope of the present disclosure is not limited in this regard.
[0074] In some embodiments, the offset is determined at the encoder based on the position of the current PC sample and the position of another PC sample. By way of example and not limitation, the offset may be determined as the difference between the position of the current PC sample and the position of another PC sample. Thus, the position of the current PC sample may be determined at the decoder based on the offset.
[0075] By way of example and not limitation, the third indication may be a syntax element, an index, a flag, etc. It should be noted that the third indication may be implemented as a single indication, multiple indications, or a combination of multiple indications. In one example, the third indication may be encoded and decoded using fixed-length encoding and decoding. In another example, the third indication may be encoded and decoded using unary encoding and decoding. In yet another example, the third indication may be encoded and decoded using truncated unary encoding and decoding. Alternatively, the third indication is encoded and decoded in a predictive manner. It should be understood that the above diagrams and / or examples are described only for the purpose of description. The scope of the present disclosure is not limited in this regard.
[0076] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for point cloud encoding and decoding of a point cloud sequence. In the method, a first indication indicating whether multi-reference inter-frame prediction is enabled for the point cloud sequence is obtained, and multiple reference PC samples are used in the multi-reference inter-frame prediction. In addition, a bitstream is generated based on the first indication.
[0077] According to some further embodiments of the present disclosure, a method for storing a bitstream of a point cloud sequence is provided. According to the method, a first indication indicating whether multi-reference inter-frame prediction is enabled for the point cloud sequence is obtained, and multiple reference PC samples are used in the multi-reference inter-frame prediction. In addition, the bitstream is generated based on the first indication, and the bitstream is stored in a non-transitory computer-readable recording medium.
[0078] The implementation of the present disclosure can be described according to the following items, and the features can be combined in any reasonable manner.
[0079] Item 1. A method for point cloud encoding and decoding, including: obtaining a first indication indicating whether multi-reference inter-frame prediction is enabled for the point cloud sequence for the conversion between the current point cloud (PC) sample of the point cloud sequence and the bitstream of the point cloud sequence, and multiple reference PC samples are used in the multi-reference inter-frame prediction; and performing the conversion based on the first indication.
[0080] Item 2. The method according to Item 1, wherein the first indication is included in the bitstream.
[0081] Item 3. The method according to any one of Items 1 to 2, wherein the first indication is encoded and decoded by using one of the following: fixed-length encoding and decoding, unary encoding and decoding, or truncated unary encoding and decoding.
[0082] Item 4. The method according to any one of Items 1 to 2, wherein the first indication is encoded and decoded in a predictive manner.
[0083] Item 5. The method according to any one of Items 1 to 4, wherein if the first indication indicates that the multi-reference inter-frame prediction is disabled for the point cloud sequence, a single reference PC sample is allowed to be used for performing inter-frame prediction on the current PC sample.
[0084] Item 6. The method according to any one of Items 1 to 5, wherein performing the conversion includes: obtaining a second indication indicating whether the multi-reference inter-frame prediction is used for the current PC sample; and performing the conversion based on the first indication and the second indication.
[0085] Item 7. The method according to Item 6, wherein the second indication is determined at the encoder and is included in the bitstream.
[0086] Item 8. The method according to any one of Items 6 to 7, wherein the second indication is encoded and decoded by using one of the following: fixed-length encoding and decoding, unary encoding and decoding, or truncated unary encoding and decoding.
[0087] Item 9. The method according to any one of Items 6 to 7, wherein the second indication is encoded and decoded in a predictive manner.
[0088] Item 10. The method according to Item 6, wherein the second indication is determined at the decoder.
[0089] Item 11. The method according to any one of Items 1 to 10, wherein the point cloud sequence includes a plurality of PC samples, and the position of the current PC sample in the timestamp order of the plurality of PC samples is included in the bitstream.
[0090] Item 12. The method according to Item 11, wherein the position is encoded and decoded using one of the following: fixed-length encoding and decoding, unary encoding, or truncated unary encoding.
[0091] Item 13. The method according to Item 11, wherein the position is encoded and decoded in a predictive manner.
[0092] Item 14. The method according to any one of Items 1 to 10, wherein the point cloud sequence includes a plurality of PC samples, and a third indication indicating the position of the current PC sample in the timestamp order of the plurality of PC samples is included in the bitstream.
[0093] Item 15. The method according to Item 14, wherein the plurality of PC samples includes another PC sample different from the current PC sample, and the third indication includes an offset that depends on the position of the current PC sample and the position of the other PC sample in the timestamp order.
[0094] Item 16. The method according to Item 15, wherein the other PC sample is before the current PC sample in the encoding and decoding order of the plurality of PC samples.
[0095] Item 17. The method according to Item 15, wherein the other PC sample is immediately before the current PC sample in the encoding and decoding order of the plurality of PC samples.
[0096] Item 18. The method according to Item 15, wherein the other PC sample is immediately before the current PC sample in a set of PC samples of the point cloud sequence, and each PC sample in the set of PC samples satisfies one of the following conditions: inter-frame prediction is disabled for the corresponding PC sample, or the corresponding PC sample is encoded based on inter-frame prediction using a single reference PC sample.
[0097] Item 19. The method according to Item 15, wherein the other PC sample immediately precedes the current PC sample in a set of PC samples of the point cloud sequence, and inter-frame prediction is disabled for each PC sample in the set of PC samples.
[0098] Item 20. The method according to Item 15, wherein the other PC sample immediately precedes the current PC sample in a set of PC samples of the point cloud sequence, and each PC sample in the set of PC samples is encoded and decoded based on inter-frame prediction using a single reference PC sample.
[0099] Item 21. The method according to any one of Items 15 to 20, wherein the offset is determined at the encoder based on the position of the current PC sample and the position of the other PC sample.
[0100] Item 22. The method according to any one of Items 15 to 21, wherein the position of the current PC sample is determined at the decoder based on the offset.
[0101] Item 23. The method according to any one of Items 14 to 22, wherein the third indication is encoded and decoded using one of: fixed-length coding, unary coding, or truncated unary coding.
[0102] Item 24. The method according to any one of Items 14 to 22, wherein the third indication is encoded in a predictive manner.
[0103] Item 25. The method according to any one of Items 11 to 24, wherein the timestamp order is different from the coding order of the plurality of PC samples.
[0104] Item 26. The method according to any one of Items 11 to 25, wherein the timestamp order is in the form of consecutive increasing integers.
[0105] Item 27. The method according to any one of Items 1 to 26, wherein the PC sample is one of: a frame, a picture, a slice, a tile, a sub-picture, a node, a point, or a unit containing one or more nodes or points.
[0106] Item 28. The method according to any one of Items 1 to 27, wherein the transformation includes encoding the current PC sample into the bitstream.
[0107] Item 29. The method according to any one of Items 1 to 27, wherein the transformation includes decoding the current PC sample from the bitstream.
[0108] Item 30. An apparatus for point cloud encoding and decoding, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of Items 1 to 29.
[0109] Item 31. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of Items 1 to 29.
[0110] Item 32. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for point cloud encoding and decoding for a point cloud sequence, wherein the method comprises: obtaining a first indication indicating whether multi-reference frame inter prediction is enabled for the point cloud sequence, in which a plurality of reference PC samples are used in the multi-reference frame inter prediction; and generating the bitstream based on the first indication.
[0111] Item 33. A method for storing a bitstream of a point cloud sequence, comprising: obtaining a first indication indicating whether multi-reference frame inter prediction is enabled for the point cloud sequence, in which a plurality of reference PC samples are used in the multi-reference frame inter prediction; generating the bitstream based on the first indication; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0112] Figure 11 A block diagram of a computing device 1100 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1100 may be implemented as the source device 110 (or the GPCC encoder 116 or 200) or the destination device 120 (or the GPCC decoder 126 or 300), or may be included in the source device 110 (or the GPCC encoder 116 or 200) or the destination device 120 (or the GPCC decoder 126 or 300).
[0113] It should be understood that Figure 11 the computing device 1100 shown is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.
[0114] As Figure 11 shown, the computing device 1100 includes a general-purpose computing device 1100. The computing device 1100 may include at least one or more processors or processing units 1110, a memory 1120, a storage unit 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160.
[0115] In some embodiments, computing device 1100 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that computing device 1100 may support any type of interface to the user (such as "wearable" circuitry, etc.).
[0116] Processing unit 1110 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in memory 1120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of computing device 1100. Processing unit 1110 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0117] Computing device 1100 generally includes various computer storage media. Such media may be any media accessible to computing device 1100, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. Memory 1120 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. Storage unit 1130 may be any removable or non-removable media, and may include machine-readable media, such as memory, flash drive, disk, or other media that can be used to store information and / or data and can be accessed in computing device 1100.
[0118] Computing device 1100 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 11 A disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0119] The communication unit 1140 communicates with another computing device via a communication medium. Additionally, the functionality of the components in the computing device 1100 can be implemented by a single computing cluster or multiple computing machines, which can communicate via a communication connection. Thus, the computing device 1100 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.
[0120] The input device 1150 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and so on. The output device 1160 can be one or more of various output devices, such as a display, speaker, printer, and so on. With the aid of the communication unit 1140, the computing device 1100 can also communicate with one or more external devices (not shown), such as storage devices and display devices, the computing device 1100 can also communicate with one or more devices that enable a user to interact with the computing device 1100, or if needed, the computing device 1100 can also communicate with any device (such as a network card, modem, etc.) that enables the computing device 1100 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0121] In some embodiments, some or all of the components of the computing device 1100 can also be arranged in a cloud computing architecture instead of being integrated in a single device. In a cloud computing architecture, the components can be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require the end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses a suitable protocol to provide services via a wide area network (such as the Internet). For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure can provide services through a shared data center, although to the user, they appear as a single access point. Thus, a cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0122] In an embodiment of the present disclosure, the computing device 1100 can be used to implement point cloud encoding / decoding. The memory 1120 may include one or more point cloud encoding / decoding modules 1125 having one or more program instructions. These modules are accessible and executable by the processing unit 1110 to perform the functions of the various embodiments described herein.
[0123] In an example embodiment of performing point cloud encoding, the input device 1150 may receive point cloud data as an input 1170 to be encoded. The point cloud data may be processed, for example, by the point cloud encoding / decoding module 1125 to generate an encoded bitstream. The encoded bitstream may be provided as an output 1180 via the output device 1160.
[0124] In an example embodiment of performing point cloud decoding, the input device 1150 may receive the encoded bitstream as an input 1170. The encoded bitstream may be processed, for example, by the point cloud encoding / decoding module 1125 to generate decoded point cloud data. The decoded point cloud data may be provided as an output 1180 via the output device 1160.
[0125] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for point cloud encoding and decoding, comprising: obtaining a first indication indicating whether multi-reference frame inter prediction is enabled for the point cloud sequence, for the conversion between a current point cloud (PC) sample of the point cloud sequence and the bitstream of the point cloud sequence, where multiple reference PC samples are used in the multi-reference frame inter prediction; and performing the conversion based on the first indication.
2. The method according to claim 1, wherein the first indication is included in the bitstream.
3. The method according to any one of claims 1 to 2, wherein the first indication is encoded and decoded using one of the following: fixed length encoding and decoding, unary encoding and decoding, or truncated unary encoding and decoding.
4. The method according to any one of claims 1 to 2, wherein the first indication is encoded and decoded in a predictive manner.
5. The method according to any one of claims 1 to 4, wherein if the first indication indicates that the multi-reference frame inter prediction is disabled for the point cloud sequence, a single reference PC sample is allowed to be used for performing inter prediction on the current PC sample.
6. The method according to any one of claims 1 to 5, wherein performing the conversion comprises: obtaining a second indication indicating whether the multi-reference frame inter prediction is used for the current PC sample; and performing the conversion based on the first indication and the second indication.
7. The method according to claim 6, wherein the second indication is determined at the encoder and is included in the bitstream.
8. The method according to any one of claims 6 to 7, wherein the second indication is encoded and decoded using one of the following: fixed length encoding and decoding, unary encoding and decoding, or truncated unary encoding and decoding.
9. The method according to any one of claims 6 to 7, wherein the second indication is encoded and decoded in a predictive manner.
10. The method according to claim 6, wherein the second indication is determined at the decoder.
11. The method according to any one of claims 1 to 10, wherein the point cloud sequence includes multiple PC samples, and the position of the current PC sample in the timestamp order of the multiple PC samples is included in the bitstream.
12. The method according to claim 11, wherein the position is encoded and decoded using one of the following: fixed length encoding and decoding, unary encoding and decoding, or truncated unary encoding and decoding.
13. The method according to claim 11, wherein the position is encoded and decoded in a predictive manner.
14. The method according to any one of claims 1 to 10, wherein the point cloud sequence includes multiple PC samples, and a third indication indicating the position of the current PC sample in the timestamp order of the multiple PC samples is included in the bitstream.
15. The method according to claim 14, wherein the multiple PC samples include another PC sample different from the current PC sample, and the third indication includes an offset that depends on the position of the current PC sample and the position of the other PC sample in the timestamp order.
16. The method according to claim 15, wherein the other PC sample is before the current PC sample in the encoding and decoding order of the plurality of PC samples.
17. The method according to claim 15, wherein the other PC sample is immediately before the current PC sample in the encoding and decoding order of the plurality of PC samples.
18. The method according to claim 15, wherein the other PC sample is immediately before the current PC sample in a set of PC samples of the point cloud sequence, and each PC sample in the set of PC samples satisfies one of the following conditions: Inter-frame prediction is disabled for the corresponding PC sample, or The corresponding PC sample is encoded and decoded based on inter-frame prediction using a single reference PC sample.
19. The method according to claim 15, wherein the other PC sample is immediately before the current PC sample in a set of PC samples of the point cloud sequence, and inter-frame prediction is disabled for each PC sample in the set of PC samples.
20. The method according to claim 15, wherein the other PC sample is immediately before the current PC sample in a set of PC samples of the point cloud sequence, and each PC sample in the set of PC samples is encoded and decoded based on inter-frame prediction using a single reference PC sample.
21. The method according to any one of claims 15 to 20, wherein the offset is determined at the encoder based on the position of the current PC sample and the position of the other PC sample.
22. The method according to any one of claims 15 to 21, wherein the position of the current PC sample is determined at the decoder based on the offset.
23. The method according to any one of claims 14 to 22, wherein the third indication is encoded using one of the following: Fixed-length encoding and decoding, Unary encoding, or Truncated unary encoding.
24. The method according to any one of claims 14 to 22, wherein the third indication is encoded in a predictive manner.
25. The method according to any one of claims 11 to 24, wherein the timestamp order is different from the encoding and decoding order of the plurality of PC samples.
26. The method according to any one of claims 11 to 25, wherein the timestamp order is in the form of continuously increasing integers.
27. The method according to any one of claims 1 to 26, wherein the PC sample is one of the following: Frame, Picture, Slice, Tile, Sub-picture, Node, Point, or A unit containing one or more nodes or points.
28. The method according to any one of claims 1 to 27, wherein the conversion includes encoding the current PC sample into the bitstream.
29. The method according to any one of claims 1 to 27, wherein the conversion includes decoding the current PC sample from the bitstream.
30. A device for point cloud encoding and decoding, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of claims 1 to 29.
31. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1 to 29.
32. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by a device for point cloud encoding and decoding for a point cloud sequence, the method comprising: obtaining a first indication indicating whether multi-reference frame inter prediction is enabled for the point cloud sequence, in which multiple reference PC samples are used; and generating the bitstream based on the first indication.
33. A method for storing a bitstream of a point cloud sequence, comprising: obtaining a first indication indicating whether multi-reference frame inter prediction is enabled for the point cloud sequence, in which multiple reference PC samples are used; generating the bitstream based on the first indication; and storing the bitstream in a non-transitory computer-readable recording medium.