Inter prediction candidate selection in point cloud compression
By allowing the selection of the previous frame with global motion compensation or the use of a resampled previous reference frame for point cloud inter-frame prediction, the problem of poor inter-frame prediction performance in the prior art is solved, and the point cloud reconstruction quality and bit transmission efficiency are improved.
Patent Information
- Application Number
- CN202480024740.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-21
- Filing Date
- 2024-04-09
- Publication Date
- 2025-11-07
AI Technical Summary
In existing point cloud compression techniques, the selection of inter-frame prediction candidates is limited by the previous frame or the global motion compensation frame, resulting in poor prediction performance, insufficient reconstruction quality, and low bit transmission efficiency.
Inter-frame prediction can be performed by selecting the previous frame after global motion compensation or by using a resampled previous reference frame, and the use of the previous frame is indicated by an additional syntax element. The value of the syntax element is then used to perform inter-frame prediction of the point cloud data.
This improved the quality of point cloud reconstruction and reduced bit transmission, resulting in better prediction performance and higher coding efficiency.
Smart Images

Figure CN120917752A_ABST
Abstract
Description
Cross Reference to Related Applications
[0001] This application claims priority to U.S. Patent Application No. 18 / 612,724, filed March 21, 2024, and U.S. Provisional Patent Application No. 63 / 496,669, filed April 17, 2023, the entire contents of each of which are incorporated by reference. U.S. Patent Application No. 18 / 612,724, filed March 21, 2024, claims the benefit of U.S. Provisional Patent Application No. 63 / 496,669, filed April 17, 2023. TECHNICAL FIELD
[0002] This disclosure relates to point cloud encoding and decoding. BACKGROUND
[0003] A point cloud is a collection of points in a three-dimensional space. The points can correspond to points on objects within the three-dimensional space. Thus, a point cloud can be used to represent the physical content of a three-dimensional space. Point clouds can have utility in a wide variety of situations. For example, a point cloud can be used in the context of an autonomous vehicle to represent the location of objects on a roadway. In another example, a point cloud can be used in the context of representing the physical content of an environment in order to position virtual objects in an augmented reality (AR) or mixed reality (MR) application. Point cloud compression is a process for encoding and decoding point clouds. Encoding a point cloud can reduce the amount of data needed to store and transmit the point cloud. SUMMARY
[0004] In general, this disclosure describes techniques for inter prediction candidate selection for point cloud compression. Such techniques can include selecting a reference frame for inter prediction and / or determining a list of inter prediction factor candidates.
[0005] In some example implementations, the selection of candidates and reference frames in point cloud compression is limited to selecting between so-called pre-pre frames or previously globally motion compensated frames. Such techniques do not allow the use of a globally motion compensated pre-pre frame or the use of a resampled previous reference frame. However, in some cases, a globally motion compensated pre-pre frame or a resampled previous reference frame can provide better prediction and thus provide a better quality reconstruction of a point cloud by a point cloud decoder and / or fewer bits for transmitting residuals between a point cloud encoder and a point cloud decoder. Thus, it can be desirable to allow inter prediction of point cloud data using a globally motion compensated pre-pre frame or using a resampled previous reference frame. Techniques of this disclosure can include using an additional syntax element whose value can indicate whether a pre-pre frame is allowed for inter prediction of point cloud data to which the additional syntax element applies.
[0006] In one example, the disclosure describes a method of processing point cloud data, the method comprising: processing a syntax element indicating whether a bi-previous frame can be used for inter prediction of one or more points of the point cloud data, wherein the bi-previous frame comprises a reference frame of a previous reference frame; and coding the one or more points based on a determination of whether the bi-previous frame can be used for inter prediction of the one or more points.
[0007] In another example, the disclosure describes an apparatus for processing point cloud data, the apparatus comprising: one or more memories configured to store point cloud data; and one or more processors implemented in circuitry and communicatively coupled to the one or more memories, the one or more processors configured to: process a syntax element indicating whether a bi-previous frame can be used for inter prediction of one or more points of the point cloud data, wherein the bi-previous frame comprises a reference frame of a previous reference frame; and code the one or more points based on a determination of whether the bi-previous frame can be used for inter prediction of the one or more points.
[0008] In another example, the disclosure describes a non-transitory computer- readable storage medium comprising instructions that, when executed, cause one or more processors to: process a syntax element indicating whether a bi-previous frame can be used for inter prediction of one or more points of the point cloud data, wherein the bi-previous frame comprises a reference frame of a previous reference frame; and code the one or more points based on a determination of whether the bi-previous frame can be used for inter prediction of the one or more points.
[0009] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 is a block diagram illustrating an example encoding and decoding system that can perform the techniques of this disclosure.
[0011] Figure 2 is a block diagram illustrating an example geometry point cloud compression (G-PCC) encoder.
[0012] Figure 3 is a block diagram illustrating an example G-PCC decoder.
[0013] Figure 4 is an example octree split for geometry coding according to the techniques of this disclosure.
[0014] Figure 5 is a conceptual diagram illustrating an example of a prediction tree according to one or more techniques of this disclosure.
[0015] Figure 6Aand Figure 6B is a conceptual diagram illustrating an example of a rotating light detection and ranging (LIDAR) acquisition model in accordance with one or more techniques of this disclosure.
[0016] Figure 7 is a conceptual diagram illustrating an example of inter-prediction of a current point from a point in a reference frame in accordance with one or more techniques of this disclosure.
[0017] Figure 8 is a flowchart illustrating an operation of a G-PCC decoder in accordance with one or more techniques of this disclosure.
[0018] Figure 9 is a conceptual diagram illustrating an example of a further inter-predictor point obtained from a first point having an azimuth angle greater than an inter-predictor point in accordance with one or more techniques of this disclosure.
[0019] Figure 10 is a flowchart illustrating an example inter-prediction candidate selection technique in accordance with one or more aspects of this disclosure.
[0020] Figure 11 is a conceptual diagram illustrating an example ranging system that can be used with one or more techniques of this disclosure.
[0021] Figure 12 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of this disclosure can be used.
[0022] Figure 13 is a conceptual diagram illustrating an example extended reality system in which one or more techniques of this disclosure can be used.
[0023] Figure 14 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure can be used. DETAILED DESCRIPTION
[0024] In some example implementations, the selection of candidates and reference frames in point cloud compression is limited to selecting between so-called pre-pre frames or previously globally motion compensated frames. Such techniques do not allow for the use of a globally motion compensated pre-pre frame or the use of a resampled previous reference frame. However, in some cases, a globally motion compensated pre-pre frame or a resampled previous reference frame can provide better prediction and thus provide a better quality reconstruction of a point cloud by a point cloud coder and / or fewer bits for transmitting residuals between a point cloud encoder and a point cloud decoder. Thus, it can be desirable to allow for inter-prediction of point cloud data using a globally motion compensated pre-pre frame or using a resampled previous reference frame.
[0025] Figure 1This is a block diagram illustrating an example encoding and decoding system 100 capable of implementing the techniques of this disclosure. The techniques of this disclosure generally relate to decoding (encoding and / or decoding) point cloud data, i.e., supporting point cloud compression. Generally, point cloud data includes any data used for processing point clouds. This decoding can be effective in compressing and / or decompressing point cloud data.
[0026] like Figure 1 As shown, system 100 includes a source device 102 and a destination device 116. The source device 102 provides encoded point cloud data for decoding by the destination device 116. Specifically, in Figure 1 In this example, source device 102 provides point cloud data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can include any of a wide range of devices, including desktop computers, laptops, tablets, set-top boxes, mobile phones (such as smartphones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, land or sea vehicles, spacecraft, aircraft, robots, LiDAR devices, satellites, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication.
[0027] exist Figure 1 In the example, source device 102 includes a data source 104, a memory 106, a G-PCC encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a G-PCC decoder 300, a memory 120, and a data consumer 118. According to this disclosure, the G-PCC encoder 200 of source device 102 and the G-PCC decoder 300 of destination device 116 can be configured to apply techniques related to inter-frame prediction candidate selection in point cloud compression as described in this disclosure. Therefore, source device 102 represents an example of an encoding device, while destination device 116 represents an example of a decoding device. In other examples, source device 102 and destination device 116 may include other components or arrangements. For example, source device 102 may receive data (e.g., point cloud data) from an internal or external source. Similarly, destination device 116 may interface with an external data consumer without including the data consumer in the same device.
[0028] like Figure 1The illustrated system 100 is merely one example. In general, other digital encoding and / or decoding devices can perform the techniques of this disclosure related to inter prediction candidate selection in point cloud compression. Source device 102 and destination device 116 are merely examples of such devices in which source device 102 generates coded data for transmission to destination device 116. This disclosure refers to “coding” devices as devices that perform coding (e.g., encoding and / or decoding) of data. Thus, G-PCC encoder 200 and G-PCC decoder 300 represent examples of coding devices, in particular, encoders and decoders, respectively. In some examples, source device 102 and destination device 116 can operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes encoding and decoding components. Hence, system 100 can support one-way or two-way transmission between source device 102 and destination device 116, e.g., for streaming, playback, broadcast, telephony, navigation, and other applications.
[0029] In general, data source 104 represents a source of data (i.e., raw, unencoded point cloud data) and can provide a series of sequential “frames” of data to G-PCC encoder 200, which encodes the data of the frames. Data source 104 of source device 102 can include a point cloud capture device such as any of a variety of cameras or sensors (e.g., a 3D scanner or a light detection and ranging (LIDAR) device, one or more video cameras), an archive containing previously captured data, and / or a data feed interface to receive data from a data content provider. Alternatively or additionally, point cloud data can be computer generated from a scanner, camera, sensor, or other data. For example, data source 104 can generate computer graphics-based data as source data, or a combination of live data, archived data, and computer generated data. In each case, G-PCC encoder 200 encodes the captured data, pre-captured data, or computer generated data. G-PCC encoder 200 can rearrange the frames from the received order (sometimes referred to as “display order”) into a coding order for coding. G-PCC encoder 200 can generate one or more bitstreams including encoded data. Source device 102 can then output the encoded data via output interface 108 onto computer-readable medium 110 for reception and / or retrieval by input interface 122 of, e.g., destination device 116.
[0030] The memories 106 of the source device 102 and 120 of the destination device 116 can represent general-purpose memories. In some examples, the memories 106 and 120 can store raw data, e.g., raw data from the data source 104 and raw decoded data from the G-PCC decoder 300. Additionally or alternatively, the memories 106 and 120 can store software instructions capable of being executed by, e.g., the G-PCC encoder 200 and the G-PCC decoder 300, respectively. Although the memories 106 and 120 are shown separate from the G-PCC encoder 200 and the G-PCC decoder 300 in this example, it should be understood that the G-PCC encoder 200 and the G-PCC decoder 300 can also include internal memories for functionally similar or equivalent purposes. Furthermore, the memories 106 and 120 can store encoded data, e.g., output from the G-PCC encoder 200 and input to the G-PCC decoder 300. In some examples, portions of the memories 106 and 120 can be allocated as one or more buffers, e.g., for storing raw decoded data and / or encoded data. For example, the memories 106 and 120 can store data representing point clouds.
[0031] The computer-readable medium 110 can represent any type of medium or device capable of storing the encoded data from the source device 102 and communicating that encoded data to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium to enable the source device 102 to send encoded data directly to the destination device 116 in real-time, e.g., via a radio frequency network or computer-based network. The output interface 108 can modulate the transmission signal including the encoded data, and the input interface 122 can demodulate the received transmission signal according to a communication standard, such as a wireless communication protocol. Communication media can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication media can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication media can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from the source device 102 to the destination device 116.
[0032] In some examples, the source device 102 can output encoded data from the output interface 108 to a storage device 112. Similarly, the destination device 116 can access encoded data from the storage device 112 via the input interface 122. The storage device 112 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded data.
[0033] In some examples, source device 102 can output encoded data to a file server 114 or another intermediate storage device that can store encoded data generated by source device 102. Destination device 116 can access stored data from file server 114 via streaming or download. File server 114 can be any type of server device that is capable of storing encoded data and transmitting that encoded data to destination device 116. File server 114 can represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 can access encoded data from file server 114 through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of the two, suitable for accessing encoded data stored on file server 114. File server 114 and input interface 122 can be configured to operate according to a streaming protocol, a download delivery protocol, or a combination thereof.
[0034] Output interface 108 and input interface 122 can represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of a variety of IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 comprise wireless components, output interface 108 and input interface 122 can be configured to transfer data, such as encoded data, according to a cellular communication standard, such as 4G, 4G-LTE (Long-Term Evolution), LTE Advanced, 5G, or the like. In some examples where output interface 108 comprises a wireless transmitter, output interface 108 and input interface 122 can be configured to transfer data, such as encoded data, according to other wireless standards, such as IEEE 802.11 specifications, IEEE 802.15 specifications (e.g., ZigBee ™ ™ standards, and the like. In some examples, source device 102 and / or destination device 116 can comprise respective system on a chip (SoC) devices. For example, source device 102 can comprise a SoC device to perform functions attributed to G-PCC encoder 200 and / or output interface 108, and destination device 116 can comprise a SoC device to perform functions attributed to G-PCC decoder 300 and / or input interface 122.
[0035] The techniques of this disclosure can be applicable to encoding and decoding to support any of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors, and processing devices such as local or remote servers, geographic mapping, or other applications.
[0036] Input interface 122 of destination device 116 receives an encoded bitstream from computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, or the like). The encoded bitstream can include signaling information defined by G-PCC encoder 200 that is also used by G-PCC decoder 300, such as syntax elements having values that describe characteristics and / or processing of coded units (e.g., slices, pictures, groups of pictures, sequences, or the like). Data consumer 118 uses the decoded data. For example, data consumer 118 can use the decoded data to determine a location of a physical object. In some examples, data consumer 118 can include a display for presenting images based on point clouds.
[0037] G-PCC encoder 200 and G-PCC decoder 300 each can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware or any combinations thereof. When the techniques are implemented partially in software, a device can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of G-PCC encoder 200 and G-PCC decoder 300 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective device. A device that includes G-PCC encoder 200 and / or G-PCC decoder 300 can comprise one or more integrated circuits (ICs), microprocessors, and / or other types of devices.
[0038] G-PCC encoder 200 and G-PCC decoder 300 can operate according to a coding standard, such as the Versatile Video Coding (V-PCC) standard or the Geometric Point Cloud Compression (G-PCC) standard. This disclosure can generally relate to coding (e.g., encoding and decoding) of pictures, including processes of encoding or decoding data. An encoded bitstream generally includes a series of values for syntax elements that represent coding decisions (e.g., coding modes).
[0039] The disclosure can generally relate to “signaling” certain information, such as syntax elements. The term “signaling” can generally refer to a communication of values for syntax elements and / or other data used to decode encoded data. That is, the G-PCC encoder 200 can signal values for syntax elements in a bitstream. Generally, signaling refers to generating values in a bitstream. As noted above, the source device 102 can transmit the bitstream to the destination device 116 in substantially real-time or not in real-time, such as can occur when syntax elements are stored to a storage device 112 for later retrieval by the destination device 116.
[0040] ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) and recently ISO / IEC MPEG 3DG (JTC 1 / SC 29 / WG 7) are investigating the potential need for standardization of point cloud coding techniques with compression capabilities significantly exceeding current approaches. MPEG is conducting this exploratory activity jointly in a collaboration called the Three-Dimensional Graphics Team (3DG) to evaluate compression technology designs proposed by experts in the field.
[0041] Point cloud compression activities are categorized into two different approaches. The first approach is “Video Point Cloud Compression” (V-PCC), which segments 3D objects and projects these segments to multiple 2D planes, which are represented as “patches” in 2D frames, which are further coded by legacy 2D video codecs such as High Efficiency Video Coding (HEVC) (ITU-T H.265) codecs. The second approach is “Geometry-based Point Cloud Compression” (G-PCC), which directly compresses 3D geometry (i.e., positions of point sets in 3D space) and associated attribute values (for each point associated with 3D geometry). G-PCC addresses compression of point clouds in both Category 1 (static point clouds) and Category 3 (dynamically acquired point clouds). A recent draft of the G-PCC standard can be found in ISO / IEC FDIS 23090-9 Geometry-based Point Cloud Compression (ISO / IEC JTC 1 / SC 29 / WG 7 M55637, Telecommunication Standardization Sector of ITU, October 2020), and a description of the codec can be found in G-PCC Codec Description (ISO / IEC JTC 1 / SC 29 / WG 7 MDS 20983, Telecommunication Standardization Sector of ITU, October 2021).
[0042] A point cloud contains a collection of points in 3D space and can have attributes associated with the points. The attributes can be color information such as R, G, B, or Y, Cb, Cr, or reflectance information, or other attributes. Point clouds can be captured by various cameras or sensors such as LIDAR sensors and 3D scanners, and can also be computer generated. Point cloud data is used for a variety of applications including, but not limited to, architecture (modeling), graphics (3D models for visualization and animation), and automotive industry (LIDAR sensors to help with navigation).
[0043] The 3D space occupied by the point cloud data can be enclosed by a virtual bounding box. The location of points in the bounding box can be represented with a certain precision; thus, the location of one or more points can be quantized based on the precision. At the smallest level, the bounding box is partitioned into voxels, which are the smallest unit of space represented by a unit cube. A voxel in the bounding box can be associated with zero, one, or more than one point. The bounding box can be partitioned into multiple cuboid regions, which can be referred to as tiles. Each tile can be coded as one or more slices. The division of the bounding box into slices and tiles can be based on the number of points in each partition, or based on other considerations (e.g., certain regions can be coded as tiles). The slice regions can be further divided using a similar split decision as in video codecs.
[0044] Figure 2 An overview of the G-PCC encoder 200 is provided. Figure 3 An overview of the G-PCC decoder 300 is provided. The illustrated modules are logical and do not necessarily correspond one-to-one with the implemented code in the reference implementation of the G-PCC codec, i.e., the TMC13 test model software studied by ISO / IEC MPEG (JTC 1 / SC 29 / WG 11).
[0045] In both the G-PCC encoder 200 and the G-PCC decoder 300, the point cloud positions are coded first. The attribute coding depends on the decoded geometry. In Figure 2 In the G-PCC encoder 200 and the G-PCC decoder 300, the point cloud positions are coded first. The attribute coding depends on the decoded geometry. In Figure 3 In the G-PCC encoder 200 and the G-PCC decoder 300, the point cloud positions are coded first. The attribute coding depends on the decoded geometry. In
[0046] For geometry, octree and / or prediction tree coding can be utilized. For class 3 data, the compressed geometry is typically represented as an octree from the root all the way to the leaf level of individual voxels. For class 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root to the leaf level of blocks larger than a voxel) plus a model of the surface inside each leaf of the approximated pruned octree. In this way, both class 1 and class 3 data share the octree coding mechanism, while class 1 data can additionally approximate the voxels inside each leaf with a surface model. The surface model used is a triangulation that includes 1-10 triangles per block, forming a trisoup. Thus, the class 1 geometry codec is referred to as the Trisoup geometry codec, while the class 3 geometry codec is referred to as the Octree geometry codec.
[0047] At each node of the octree, the occupancy of one or more of its children (up to eight nodes) is signaled (when not inferred). Multiple neighborhoods are specified, including (a) nodes that share a face with the current octree node, (b) nodes that share a face, edge, or vertex with the current octree node, etc. Within each neighborhood, the occupancy of the current node or its children can be predicted using the occupancy of the nodes and / or their children. For points that are sparsely populated in certain nodes of the octree, the codec also supports a direct coding mode in which the 3D position of the point is directly encoded. A flag can be signaled to indicate that the direct mode is signaled. At the lowest level, the number of points associated with an octree node / leaf node can also be coded.
[0048] Figure 4 An example octree split for geometry coding according to the techniques of this disclosure.
[0049] Once the geometry is coded, the attributes corresponding to the geometry points are coded. When there are multiple attribute points corresponding to one reconstructed / decoded geometry point, the attribute value representative of the reconstructed point can be derived.
[0050] There are three attribute coding methods in G-PCC: region adaptive hierarchical transform (RAHT) coding, interpolation-based hierarchical nearest neighbor prediction (predictive transform), and interpolation-based hierarchical nearest neighbor prediction with an update / boost step (boosted transform). RAHT and boost are typically used for class 1 data, while predict is typically used for class 3 data. However, either method can be used for any data, and like the geometry codecs in G-PCC, the attribute coding method used to code a point cloud is specified in the bitstream.
[0051] The coding of attributes can be done in levels of detail (LOD), where a finer representation of the point cloud attributes is available through each level of detail. Each level of detail can be specified based on a distance metric to neighboring nodes or based on a sampling distance.
[0052] At the G-PCC encoder 200, the residual obtained as the output of the attribute decoding method can be quantized. The residual can be obtained by subtracting the attribute value from the prediction derived from the attribute values of the points in the neighborhood of the current point and based on the attribute values of the previously encoded points. The quantized residual can be decoded using context-adaptive arithmetic decoding.
[0053] exist Figure 2 In the example, the G-PCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometric reconstruction unit 216, a RAHT unit 218, a LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.
[0054] like Figure 2 As shown in the example, the G-PCC encoder 200 can obtain the set of locations and attribute sets of points in a point cloud. The G-PCC encoder 200 can obtain data from data source 104 ( Figure 1 The G-PCC encoder 200 obtains a set of locations and a set of attributes for points in the point cloud. These locations may include the coordinates of the points in the point cloud. Attributes may include information about the points in the point cloud, such as the color associated with a point in the point cloud. The G-PCC encoder 200 can generate a geometric bitstream 203 that includes an encoded representation of the locations of the points in the point cloud. The G-PCC encoder 200 can also generate an attribute bitstream 205 that includes an encoded representation of the attribute set.
[0055] The coordinate transformation unit 202 can apply transformations to the coordinates of a point to transform the coordinates from the initial domain to the transformation domain. The transformed coordinates may be referred to as transformed coordinates. The color transformation unit 204 can apply transformations to transform the color information of an attribute to a different domain. For example, the color transformation unit 204 can transform color information from the RGB color space to the YCbCr color space.
[0056] In addition, Figure 2 In the example, voxelization unit 206 can voxelize the transformed coordinates. Voxelization of the transformed coordinates can include quantization and removal of some points in the point cloud. In other words, multiple points in the point cloud can be grouped into a single "voxel," which can then be treated as a single point in some respects. Furthermore, octree analysis unit 210 can generate an octree based on the voxelized transformed coordinates. Additionally, in Figure 2In the example of FIG. 2, surface approximation analysis unit 212 can analyze the points to potentially determine a surface representation of the set of points. Arithmetic encoding unit 214 can entropy encode syntax elements representing information of the octree and / or the surface determined by surface approximation analysis unit 212. G-PCC encoder 200 can output the syntax elements in geometry bitstream 203. Geometry bitstream 203 can also include other syntax elements, including syntax elements that are not arithmetically encoded.
[0057] Geometric reconstruction unit 216 can reconstruct the transformed coordinates of the points in the point cloud based on the octree, data indicating the surface determined by surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transformed coordinates reconstructed by geometric reconstruction unit 216 can be different from the number of original points in the point cloud. This disclosure can refer to the resulting points as reconstructed points. Attribute transfer unit 208 can transfer attributes of the original points in the point cloud to the reconstructed points in the point cloud.
[0058] Further, RAHT unit 218 can apply RAHT coding to the attributes of the reconstructed points. In some examples, according to RAHT, the attributes of a block of 2x2x2 point positions are fetched and transformed along one direction to obtain four low frequency nodes (L) and four high frequency nodes (H). Subsequently, the four low frequency nodes (L) are transformed in a second direction to obtain two low frequency nodes (LL) and two high frequency nodes (LH). The two low frequency nodes (LL) are transformed along a third direction to obtain one low frequency node (LLL) and one high frequency node (LLH). The low frequency node LLL corresponds to a DC coefficient, and the high frequency nodes H, LH, and LLH correspond to AC coefficients. The transformation in each direction can be a 1-D transform with two coefficient weights. The low frequency coefficients can be fetched as coefficients of a 2x2x2 block at the next higher level for the RAHT transform, and the AC coefficients are encoded without change; such transformations continue until the top root node. The tree traversal for encoding is used top-down to compute the weights to be used for the coefficients; the transformation order is bottom-up. The coefficients can then be quantized and coded.
[0059] Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 can apply LOD processing and lifting, respectively, to the attributes of the reconstructed points. LOD generation is used to split the attributes into different levels of detail. Each level of detail provides a refinement of the attributes of the point cloud. The first level of detail provides a coarse approximation and contains few points; subsequent levels of detail typically contain more points, and so on. The levels of detail can be constructed using a distance-based metric, or one or more other classification criteria can also be used (e.g., subsampling from a particular order). Thus, all reconstructed points can be included in the levels of detail. Each level of detail is produced by taking the union of all points up to a particular level of detail: e.g., LOD1 based on level of detail RL1, LOD2 based on RL1 and RL2,..., LODN by the union of RL1, RL2,..., RLN. In some cases, the LOD generation can be followed by a prediction scheme (e.g., a predictive transform) in which attributes associated with each point in the LOD are predicted from a weighted average of previous points and a residual is quantized and entropy coded. The lifting scheme builds on the predictive transform mechanism in which the coefficients are updated using an update operator and adaptive quantization of the coefficients is performed.
[0060] The RAHT unit 218 and the lifting unit 222 can generate coefficients based on the attributes. The coefficient quantization unit 224 can quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic encoding unit 226 can apply arithmetic coding to syntax elements representing the quantized coefficients. The G-PCC encoder 200 can output these syntax elements in the attribute bitstream 205. The attribute bitstream 205 can also include other syntax elements, including syntax elements that are not arithmetically encoded.
[0061] In the example of FIG. 3, Figure 3 In the example of FIG. 3,
[0062] The G-PCC decoder 300 can obtain the geometry bitstream 203 and the attribute bitstream 205. The geometry arithmetic decoding unit 302 of the G-PCC decoder 300 can apply arithmetic decoding (e.g., context adaptive binary arithmetic coding (CABAC) or other types of arithmetic decoding) to the syntax elements in the geometry bitstream 203. Similarly, the attribute arithmetic decoding unit 304 can apply arithmetic decoding to the syntax elements in the attribute bitstream 205.
[0063] Octree synthesis unit 306 can synthesize the octrees based on syntax elements parsed from geometry bitstream 203. Starting from the root node of the octree, the occupancy of each of the eight child nodes at each octree level is signaled in the bitstream. When signaling indicates that a child node at a particular octree level is occupied, the occupancy of the child nodes of that child node is signaled. The signaling of the nodes at each octree level is signaled before proceeding to subsequent octree levels. At the last level of the octree, each node corresponds to a voxel position; when a leaf node is occupied, one or more points can be designated as being occupied at the voxel position. In some instances, due to quantization, some branches of the octree can terminate early of the final level. In such cases, the leaf node is considered to be an occupied node without child nodes. In instances where surface approximation is used in geometry bitstream 203, surface approximation synthesis unit 310 can determine a surface model based on syntax elements parsed from geometry bitstream 203 and based on the octrees.
[0064] Further, geometry reconstruction unit 312 can perform reconstruction to determine coordinates of points in the point cloud. For each position at a leaf node of the octree, geometry reconstruction unit 312 can reconstruct the node position by using the binary representation of the leaf node in the octree. At each respective leaf node, the number of points at the respective leaf node is signaled; this indicates the number of duplicate points at the same voxel position. When geometry quantization is used, the point positions are scaled to determine reconstructed point position values.
[0065] Inverse transform coordinate unit 320 can apply an inverse transform to the reconstructed coordinates to convert the reconstructed coordinates (positions) of points in the point cloud from the transformed domain back to the original domain. The positions of points in the point cloud can be in a floating point domain, but the point positions in the G-PCC codec are coded in an integer domain. The inverse transform can be used to convert these positions back to the original domain.
[0066] Additionally, in Figure 3 In an example, inverse quantization unit 308 can inverse quantize attribute values. The attribute values can be based on syntax elements obtained from attribute bitstream 205 (e.g., including syntax elements decoded by attribute arithmetic decoding unit 304).
[0067] Depending on how the attribute values are encoded, RAHT unit 314 can perform RAHT decoding to determine the color values of points in the point cloud based on the inverse-quantized attribute values. RAHT decoding is performed from the top to the bottom of the tree. At each level, composition values are derived using low-frequency and high-frequency coefficients derived from the inverse-quantization process. At leaf nodes, the derived values correspond to the attribute values of the coefficients. The point weight derivation process is similar to that used at the G-PCC encoder 200. Alternatively, LOD generation unit 316 and inverse lifting unit 318 can use level-of-detail techniques to determine the color values of points in the point cloud. LOD generation unit 316 decodes each LOD, giving a progressively finer representation of the point's attributes. In the case of prediction transform, LOD generation unit 316 derives the prediction of the point from a weighted sum of points previously reconstructed in the same LOD or earlier. LOD generation unit 316 can add the prediction to the residual (obtained after inverse-quantization) to obtain the reconstructed values of the attributes. When using an enhancement scheme, the LOD generation unit 316 may also include update operators to update the coefficients used to derive attribute values. In this case, the LOD generation unit 316 may also apply inverse adaptive quantization.
[0068] In addition, Figure 3 In the example, the inverse color transformation unit 322 can apply an inverse color transformation to the color value. The inverse color transformation can be the reverse of the color transformation applied by the color transformation unit 204 of the G-PCC encoder 200. For example, the color transformation unit 204 can transform color information from the RGB color space to the YCbCr color space. Therefore, the inverse color transformation unit 322 can transform color information from the YCbCr color space to the RGB color space.
[0069] Figure 2 and Figure 3 The various units are explained to aid in understanding the operations performed by the G-PCC encoder 200 and the G-PCC decoder 300. Units can be implemented as fixed-function circuits, programmable circuits, or combinations thereof. A fixed-function circuit is a circuit that provides specific functionality and is pre-defined for the operations that can be performed. A programmable circuit is a circuit that can be programmed to perform various tasks and provides flexible functionality for the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more units within the unit may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units within the unit may be integrated circuits.
[0070] Predictive geometry coding (see, e.g., G-PCC 2nd edition codec description, ISO / IEC JTC 1 / SC 29 / WG 7 MDS21684, Telecommunications Assembly, July 2022) is introduced as an alternative to octree geometry coding, where nodes are arranged in a tree structure (which defines a prediction structure) and various prediction strategies are used to predict the coordinates of each node in the tree relative to its predictors. Figure 5 An example of a prediction tree is shown, which is a directed graph with arrows pointing in the direction of prediction. The horizontally hashed node is the root vertex and has no predictors; the cross-hatched nodes have two children; the diagonally hashed nodes have 3 children; the non-hashed nodes have one child, while the vertically hashed nodes are leaf nodes and have no children. Each node except the root node has only one parent.
[0071] Figure 5 A conceptual diagram illustrating an example of a prediction tree. Node 500 is the root vertex and has no predictors. Nodes 502 and 504 have two children. Node 506 has 3 children. Nodes 508, 510, 512, 514, and 516 are leaf nodes and have no children. Each of the remaining nodes has one child. Each node except the root node 500 has only one parent.
[0072] Four prediction strategies are specified for each node based on its parent (p0), grandparent (p1), and great-grandparent (p2):
[0073] No prediction / zero prediction (0)
[0074] Incremental prediction (p0)
[0075] Linear prediction (2*p0 - p1)
[0076] Parallelogram prediction (p0 + p1 - p2)
[0077] The G-PCC encoder 200 can employ any algorithm to generate the prediction tree; the algorithm used is determined based on the application / use case and several strategies can be used. Some strategies are described in the G-PCC 2nd edition codec description (ISO / IEC JTC 1 / SC 29 / WG 7 MDS21684, Telecommunications Assembly, July 2022).
[0078] For each node, the residual coordinate values are coded in the bitstream in a depth-first manner, starting from the root node. For example, the G-PCC encoder 200 can code the residual coordinate values in the bitstream.
[0079] For each node, residual coordinate values are coded in the bitstream in a depth-first manner starting from the root node. For example, the G-PCC encoder 200 can code the residual coordinate values in the bitstream.
[0080] Predictive geometry coding is mainly used for category 3 (LIDAR acquired) point cloud data, e.g., for low latency applications.
[0081] Figure 6A and Figure 6B is a conceptual diagram illustrating an example of a rotating LIDAR acquisition model. Now the angular mode for predictive geometry coding is described. The angular mode can be used in predictive geometry coding, where the characteristics of the LIDAR sensor can be used to more efficiently code the prediction tree. The coordinates of a position are converted to (radius, azimuth, and laser index) domain 600, and prediction is performed in this domain 600 (e.g., residuals are coded in the domain). Due to errors in the rounding, the coding in is not lossless, so a second set of residuals corresponding to the Cartesian coordinates can be coded. The description of the encoding and decoding strategies for the angular mode for predictive geometry coding is reproduced from the G-PCC codec description below.
[0082] The angular mode technique can focus on point clouds acquired using a rotating LIDAR model. Here, the LIDAR 602 has N lasers (e.g., N = 16, 32, 64) rotating around the Z axis with an azimuth angle and a height angle. Each laser can have a different elevation angle and height . For example, the laser may hit a point Figure 6A to Figure 6B with Cartesian integer coordinates defined according to the coordinate system shown in
[0083] The technique models the position of M with three parameters in the following way:
[0084]
[0085]
[0086] ,
[0087] More precisely, the technique uses a quantized version of , denoted as , where the three integers and are computed as follows:
[0088]
[0089]
[0090]
[0091] where (a, b) = (a + b, a - b) are quantization parameters controlling the precision of (a, b) and (a, b) respectively. is a function that returns 1 if t is positive and -1 otherwise. is the absolute value of . is a function that returns 1 if t is positive and -1 otherwise. is the absolute value of .
[0092] To avoid reconstruction mismatch due to the use of floating point operations, the values of and are pre-computed and quantized as follows:
[0093]
[0094] is a function that returns 1 if t is positive and -1 otherwise.
[0095] where (a, b) = (a + b, a - b) are quantization parameters controlling the precision of (a, b) and (a, b) respectively. The reconstructed Cartesian coordinates are obtained as follows:
[0096]
[0097]
[0098]
[0099] ,
[0100] where and are approximations of and . The computations can be performed using fixed point representation, look-up tables and / or linear interpolation.
[0101] Note that due to various reasons such as quantization, approximation, model inaccuracy or model parameter inaccuracy, may be different from .
[0102] The reconstruction residual can be defined as follows:
[0103]
[0104]
[0105]
[0106] With this technique, the G-PCC encoder 200 can proceed as follows:
[0107] 1) Encode the model parameters and the quantization parameters
[0108] 2) Apply the geometry predictive scheme described in the text of ISO / IEC FDIS 23090-9 Geometry-based Point Cloud Compression (ISO / IEC JTC 1 / SC29 / WG 7 m55637, Telecommunications Assembly, October 2020) to the representation
[0109] A new predictor exploiting the characteristics of LIDAR can be introduced. For example, the rotational speed of a LIDAR scanner around the z-axis is usually constant. Thus, the G-PCC encoder 200 can predict the current :
[0110]
[0111] where
[0112] is a set of possible velocities that the G-PCC encoder 200 can use. The index can be explicitly written in the bitstream or can be inferred from the context based on a deterministic policy applied by both the G-PCC encoder 200 and the G-PCC decoder 300, and is the number of points that are skipped, which can be explicitly written in the bitstream or can be inferred from the context based on a deterministic policy applied by both the G-PCC encoder 200 and the G-PCC decoder 300. Also referred to herein as the “phi multiplier”. Note that the phi multiplier is currently only used with the delta predictor.
[0113] 3) Encode the residual reconstructed with each node pair
[0114] The G-PCC decoder 300 can proceed as follows:
[0115] 1) Decode the model parameters and the quantization parameters
[0116] 2) Decode the parameters associated with the node according to the geometry predictive scheme described in the text of ISO / IEC FDIS 23090-9 Geometry-based Point Cloud Compression (ISO / IEC JTC 1 / SC 29 / WG 7 m55637, Telecommunications Assembly, October 2020).
[0117] 3) Compute the reconstructed coordinates as described above.
[0118] 4) Decode the residual .
[0119] As discussed in the next section, lossy compression can be supported by quantizing the reconstructed residual .
[0120] 5) Compute the original coordinates as follows:
[0121]
[0122]
[0123]
[0124] Lossy compression can be achieved by applying quantization to the reconstructed residual or by discarding points.
[0125] The quantized reconstructed residual can be computed as follows:
[0126]
[0127]
[0128]
[0129] where , are quantization parameters that control the accuracy of , and respectively. For example, the G-PCC encoder 200 or the G-PCC decoder 300 can compute the quantized residual.
[0130] The G-PCC encoder 200 or the G-PCC decoder 300 can use lattice quantization to further improve the RD (Rate Distortion) performance results.
[0131] The quantization parameters can vary at the sequence / frame / slice / block level to enable region adaptive quality and / or for rate control purposes.
[0132] Inter prediction in G-PCC predictive geometry coding is now discussed. Information on G-PCC predictive geometry coding can be found in the techniques being considered in G-PCC, ISO / IEC JTC 1 / SC 29 / WG 7 MDS21256, Telecommunications Assemblies, October 2021; A. K. Ramasubramonian, B. Ray, L. Pham Van, G. Van der Auwera, M. Karczewicz, [G-PCC] Inter prediction with predictive geometry coding [new], ISO / IEC JTC 1 / SC 29 / WG 7 m56117, January 2021; and A. K. Ramasubramonian, L. Pham Van, G. Van der Auwera, M. Karczewicz, [G-PCC] Report on Inter prediction, Test 2, ISO / IEC JTC 1 / SC 29 / WG 7 m56839, April 2021.
[0133] Predictive geometry coding uses a prediction tree structure to predict the position of a point. When angular coding is enabled, the x, y, z coordinates are transformed into radius, azimuth, and laser ID, and the residual can be signaled in these three coordinates as well as in x, y, z dimensions. Intra prediction for radius, azimuth, and laser ID can be one of four modes, and the prediction factor is a node in the prediction tree classified as a parent, grandparent, and great-grandparent node relative to the current node. Predictive geometry coding as designed in G-PCC version 1 is an intra coding tool, as it only uses points in the same frame for prediction. Additionally, using points from previously decoded frames can provide better prediction and thus better compression performance.
[0134] For inter prediction, the proposal is to predict based on the radius of the reference frame prediction point as initially proposed in the following documents: G-PCC, ISO / IEC JTC 1 / SC 29 / WG 7 MDS20999, Telecommunications Congress, October 2021; and A. K. Ramasubramonian, L. Pham Van, G. Van der Auwera, M. Karczewicz, [G-PCC] Report on Inter Prediction for Predictive Geometry, Test 2, ISO / IEC JTC 1 / SC 29 / WG 7 m56839, April 2021. For each point in the prediction tree, the G-PCC decoder 300 can determine whether the point is inter predicted or intra predicted (e.g., the G-PCC encoder 200 can indicate this inter prediction or intra prediction by a flag that the G-PCC encoder 200 can signal in the bitstream). When intra prediction is performed, the intra prediction mode using predictive geometry coding is used. When inter prediction is used, the azimuth angle and laserID are still predicted with intra prediction, while the radius is predicted from a point in the reference frame that has the same laserID as the current point and the azimuth angle closest to the current azimuth angle. This technique is further modified in A. K. Ramasubramonian, G. Van der Auwera, L. Pham Van, M. Karczewicz, [G-PCC] Additional Results on Inter Prediction for Predictive Geometry [EE13.2-Related], ISO / IEC JTC 1 / SC 29 / WG 7 m56841, April 2021, where, in addition to the radius prediction, inter prediction of the azimuth angle and laserID is also implemented. When inter coding is applied, the radius, azimuth angle, and laserID of the current point are predicted based on points located in the vicinity of the azimuth angle position of the previously decoded points in the reference frame. Additionally, separate context sets are used for inter prediction and intra prediction.
[0135] Figure 7 is a conceptual diagram illustrating an example of inter prediction of a current point from a point in a reference frame. This concept can be found in A. K. Ramasubramonian, G. Van der Auwera, L. Pham Van, M. Karczewicz, [G-PCC] Additional Results on Inter Prediction for Predictive Geometry [EE13.2-Related], ISO / IEC JTC 1 / SC 29 / WG 7 m56841, April 2021. Extending inter prediction to azimuth angle, radius, and laserID includes, for example, the following steps that can be performed by the G-PCC decoder 300:
[0136] 1) For a given point (e.g., current point curPoint 700 in current frame 704), select a previously decoded point (prevDecP0 702).
[0137] 2) In reference frame 70, select a location with the same scaled azimuth and laser ID as prevDecP0 702 (e.g., refFrameP0 706).
[0138] 3) In reference frame 708, find the first point with an azimuth greater than the azimuth of refFrameP0 706 (interPredPt 710). interPredPt can also be referred to as the "Next" inter predictor.
[0139] Figure 8 is a flowchart illustrating the operation of a G-PCC decoder. Figure 8 A decoding process associated with the "inter_flag" signaled for each point is illustrated. This technique is available in InterEM-v3.0.
[0140] For example, G-PCC decoder 300 can determine whether the inter flag is true (e.g., equal to 1) (800). If the inter flag is true ("Yes" path from block 800), G-PCC decoder 300 can select a previously decoded point in decoding order using the radius, azimuth, and laser ID (802). G-PCC decoder 300 can derive the quantized phi, Q(phi) (e.g., the quantized value of the azimuth) of the selected previously decoded point (e.g., prevDecP0 702) (804). G-PCC decoder 300 can check points in a reference frame (e.g., reference frame 708 of FIG. 7) where the quantized phi of such points is greater than Q(phi), which can result in interPredPt 710 (806). G-PCC decoder 300 can then use interPredPt 710 as the inter predictor for the current point curPoint 700 (808). G-PCC decoder 300 can then add the delta phi multiplier (e.g., delta phi multiplier as discussed above Figure 7 ) to the primary residual (810).
[0141] If the inter flag is false (e.g., equal to 0) ("No" path from block 800), G-PCC decoder 300 can select an intra prediction candidate (812) and apply intra prediction. G-PCC decoder 300 can then add the delta phi multiplier to produce the primary residual (810).
[0142] Figure 9 This is a conceptual diagram illustrating an example of an additional inter-frame predictor point obtained from a first point having an azimuth angle greater than that of the inter-frame predictor point. Further predictor candidates are now discussed. Information on these candidates can be found in KL Loi, T. Nishi, and T. Sugio's [G-PCC][New] Inter-frame prediction for improved quantization of azimuth angle in predictive geometry decoding (ISO / IEC JTC1 / SC29 / WG7 m57351, July 2021). In the aforementioned inter-frame prediction technique for predictive geometry, when inter-frame decoding is applied, for example by a G-PCC decoder 300, the radius, azimuth angle, and laser ID of the current point are predicted based on points near the juxtaposed azimuth angle position in the reference frame using the following steps:
[0143] 1) For a given point (e.g., the current point, current point 900), select the previously decoded point (e.g., the previously decoded point 902).
[0144] 2) In reference frame 908, select a location (e.g., reference point 906) with the same azimuth and laser ID as the previously decoded point (e.g., previously decoded point 902).
[0145] 3) Select a position in the reference frame 908 (inter-frame prediction point 910) from a first point that has a larger azimuth angle than the position in the reference frame 908 that has the same scaling as the previously decoded point (e.g., the previously decoded point 902) and a laser ID, to use as an inter-frame prediction point.
[0146] This technology adds the ability to find, for example, [the following]. Figure 9 The additional inter-frame predictor point 912 is obtained by taking the first point with a larger azimuth angle than the inter-frame predictor point shown (e.g., inter-frame predictor point 910). Additional signaling is used to indicate which predictor to select after inter-frame decoding has been applied by the G-PCC encoder 200. For example, the G-PCC encoder 200 may signal to the G-PCC decoder 300 which predictor has been selected. The additional inter-frame predictor point may also be referred to as the "NextNext" inter-frame predictor.
[0147] Improved inter prediction flag coding is now discussed. Information on improved inter prediction flag coding can be found in A. K. Ramasubramonian, L. Pham Van, G. Van der Auwera, M. Karczewicz, [G-PCC] [New Proposal] Improvement of Inter Prediction using Predictive Geometry Coding (ISO / IEC JTC1 / SC29 / WG7 m57299, July 2021). An improved context selection algorithm can be applied to code the inter prediction flag. The G-PCC encoder 200 can use the inter prediction flag values of five previous coded points to select the context of the inter prediction flag in predictive geometry coding.
[0148] Reference frames for inter prediction are now described. Information on reference frames for inter prediction can be found in the following documents: K. L. Loi, T. Nishi, T. Sugio, [G-PCC] [EE13.2] Report on Inter Prediction Test 8 on Predictive Geometry, ISO / IEC JTC1 / SC29 / WG7 m61586, January 2023; and A. K. Ramasubramonian, G. Van der Auwera, M. Karczewicz, Thoughts on EE13.2 Test 8 - Inter Prediction in Predictive Geometry Coding, ISO / IEC JTC1 / SC29 / WG7 m62218, January 2023.
[0149] The following names that can be used in this disclosure can be defined as set forth below:
[0150] Prev - refers to the reference frame used for inter prediction when only one reference frame is used; when more than one reference frame is used, prev can indicate the reference frame that is closer in output order (e.g., frame distance) to the current frame.
[0151] X-Zero - the uncompensated reference frame X; also referred to as X.
[0152] X-Glob - the reference frame X after global motion compensation.
[0153] X-Resam - the resampled reference frame X, which can be obtained from X-Glob and X-Zero.
[0154] PrevPrev - refers to the reference frame of the previous reference frame (the pre-pre reference frame).
[0155] X / Next - refers to the next inter prediction candidate from reference frame X.
[0156] X / NextNext - refers to the next-next inter prediction candidate from reference frame X.
[0157] For predictive geometry, the reference frame selection and hence the inter prediction candidates can be selected as follows:
[0158] When global motion is disabled:
[0159] If resampling is enabled, the resampled reference frame Prev-Resam can be selected. The inter prediction list can be as follows:
[0160] Prev-Resam / Next
[0161] Prev-Resam / NextNext
[0162] If resampling is disabled, the uncompensated reference frame Prev-Zero can be selected. The inter prediction list can be as follows:
[0163] Prev-Zero / Next
[0164] Prev-Zero / NextNext
[0165] When global motion is enabled:
[0166] If resampling is enabled, the resampled reference frame Prev-Resam can be selected. The inter prediction list can be as follows:
[0167] Prev-Resam / Next,
[0168] Prev-Resam / NextNext
[0169] Prev-Glob / Next
[0170] Prev-Glob / NextNext
[0171] If resampling is disabled, the uncompensated reference frame Prev-Zero can be selected. The inter prediction list can be as follows:
[0172] Prev-Zero / Next
[0173] Prev-Zero / NextNext
[0174] Prev-Glob / Next
[0175] Prev-Glob / NextNext
[0176] In K. L. Loi, T. Nishi, T. Sugio, [G-PCC][EE13.2] Report on Inter- Prediction Test 8 on Predictive Geometry (ISO / IEC JTC1 / SC29 / WG7 m61586, January 2023), it was further proposed to check the moving status of the point cloud frame; if the point cloud frame is “moving”, e.g. there is considerable motion between the reference frame and the current frame, then the globally motion compensated candidate can be used as described above; otherwise, it can be considered as static state and the PrevPrev reference frame can be used instead of the globally motion compensated candidate. In the bitstream, effectively, a flag gmForCurrFrame can be used to indicate whether global motion is applied to a particular frame.
[0177] When global motion is disabled:
[0178] If resampling is enabled, a resampled reference frame Prev-Resam can be selected. The inter prediction list can be as follows:
[0179] Prev-Resam / Next
[0180] Prev-Resam / NextNext
[0181] If resampling is disabled, an uncompensated reference frame Prev-Zero can be selected. The inter prediction list can be as follows:
[0182] Prev-Zero / Next
[0183] Prev-Zero / NextNext
[0184] When global motion is enabled:
[0185] If resampling is enabled, a resampled reference frame Prev-Resam can be selected. The inter prediction list can be as follows:
[0186] Prev-Resam / Next,
[0187] Prev-Resam / NextNext
[0188] If gmFrameCurrFrame = 1, Prev-Glob / Next, otherwise PrevPrev / Next
[0189] If gmFrameCurrFrame = 1, Prev-Glob / NextNext, otherwise PrevPrev / NextNext
[0190] If resampling is disabled, the uncompensated reference frame Prev-Zero can be selected. The inter prediction list can be as follows:
[0191] Prev-Zero / Next
[0192] Prev-Zero / NextNext
[0193] If gmFrameCurrFrame = 1, Prev-Glob / Next, else PrevPrev / Next
[0194] If gmFrameCurrFrame = 1, Prev-Glob / NextNext, else PrevPrev / NextNext
[0195] The selection of reference frames as discussed above is limited to selecting between PrevPrev and Prev-Glob. Such techniques do not allow the use of a globally motion compensated PrevPrev frame or the use of a resampled previous reference frame. However, in some cases, a globally motion compensated PrevPrev frame or a resampled previous reference frame can provide better prediction and thus provide better quality reconstruction of the point cloud by the G-PCC decoder 300 and / or fewer bits for transmitting residuals from the G-PCC encoder 200 to the G-PCC decoder 300.
[0196] In some examples, a syntax element can be signaled to indicate whether a PrevPrev frame can be used for inter prediction. For example, the G-PCC encoder 200 can signal a flag, such as usePrevPrevRefFrameFlag, to indicate whether a previous previous reference frame is used. The G-PCC decoder 300 can parse such a flag to determine whether a previous previous reference frame is used. When the syntax element takes one value (e.g., 1), the PrevPrev frame can be used for inter prediction; when the syntax element takes a second value (e.g., 0), the PrevPrev frame is not used for inter prediction.
[0197] When the PrevPrev frame can be used for inter prediction:
[0198] If resampling is not applied, the PrevPrev frame can be applied global motion parameters signaled with the current frame. If resampling is applied, the PrevPrev frame can be applied a second set of global motion parameters signaled with the current frame. For example, the G-PCC encoder 200 can signal the global motion parameters and / or the second set of global motion parameters to the G-PCC decoder 300 in the bitstream. It should be understood that when resampling is enabled for coding the current frame (in which case resampling can be applied to the current frame), the reference frames can be compensated and resampled. However, the G-PCC decoder 300 can determine that resampling is applied to the current frame when processing the current frame.
[0199] In one example, resampling is enabled only when global motion is enabled. When resampling is disabled, the uncompensated reference frame can be used instead of the resampled reference frame. For example, the G-PCC encoder 200 or the G-PCC decoder 300 can use the uncompensated reference frame instead of the resampled reference frame when resampling is disabled.
[0200] In one example, for some reference frames, the prevprev candidate is not selected; instead, the next candidate from another reference frame is selected. For example, the G-PCC encoder 200 or the G-PCC decoder 300 can select the next candidate from another reference frame for inter prediction instead of the prevprev candidate.
[0201] In some examples, when global motion is enabled (for the entire point cloud sequence), a flag can be signaled at the frame / slice level to indicate whether global motion compensation will be applied to the current frame. For example, the G-PCC encoder 200 can signal the flag to the G-PCC decoder 300. In such cases, the last two entries in the inter prediction factor candidate list can be rejected when global motion compensation will not be applied to the current frame.
[0202] In another example, the last two entries in the inter prediction factor candidate list can be rejected when global motion compensation will not be applied to the current frame, and only the current previous frame is not used for inter prediction. For example, if the prevprev frame is not used for inter prediction, and global motion compensation is not applied to the current frame, the G-PCC encoder 200 or the G-PCC decoder 300 can not use the last two entries in the inter prediction factor candidate list.
[0203] In some cases, when global motion is enabled (overall), but global motion is not applied to the current frame, the inter prediction list can be derived similarly to when global motion is disabled (overall).
[0204] In a first example, G-PCC encoder 200 or G-PCC decoder 300 can use the following to select reference frames and inter prediction list candidates for inter prediction:
[0205] In this first example, when global motion is disabled:
[0206] If resampling is enabled, the resampled reference frame Prev-Resam is selected. The inter prediction list is as follows:
[0207] Prev-Resam / Next
[0208] Prev-Resam / NextNext
[0209] If usePrevPrevRefFrameFlag = 1, PrevPrev / Next, else none
[0210] If usePrevPrevRefFrameFlag = 1, PrevPrev / NextNext, else none
[0211] If resampling is disabled, the uncompensated reference frame Prev-Zero is selected. The inter prediction list is as follows:
[0212] Prev-Zero / Next
[0213] Prev-Zero / NextNext
[0214] If usePrevPrevRefFrameFlag = 1, PrevPrev / Next, else none
[0215] If usePrevPrevRefFrameFlag = 1, PrevPrev / NextNext, else none
[0216] In this first example, when global motion is enabled:
[0217] When usePrevPrevRefFrameFlag = 1, PrevPrev is selected as X, else Prev is selected
[0218] When usePrevPrevRefFrameFlag = 1, PrevPrev is selected as Y, else none
[0219] If resampling is enabled, the resampled reference frame Prev-Resam is selected. The inter prediction list is as follows:
[0220] Prev-Resam / Next,
[0221] Prev-Resam / Next Next
[0222] If gmFrameCurrFrame = 1, X-Glob / Next, else Y / Next
[0223] If gmFrameCurrFrame = 1, X-Glob / Next Next, else Y / Next Next
[0224] If resampling is disabled, select uncompensated reference frame Prev-Zero. Inter prediction list is as follows:
[0225] Prev-Zero / Next
[0226] Prev-Zero / Next Next
[0227] If gmFrameCurrFrame = 1, X-Glob / Next, else Y / Next
[0228] If gmFrameCurrFrame = 1, X-Glob / Next Next, else Y / Next Next
[0229] In another example, G-PCC encoder 200 or G-PCC decoder 300 can use the following to select reference frames and inter prediction list candidates for inter prediction. In this second example, when a PrevPrev frame is used, it replaces the third and fourth candidates Prev in the list.
[0230] In this second example, when global motion is disabled:
[0231] If resampling is enabled, select resampled reference frame Prev-Resam. Inter prediction list is as follows:
[0232] Prev-Resam / Next
[0233] Prev-Resam / Next Next
[0234] If usePrevPrevRefFrameFlag = 1, PrevPrev-Resam / Next, else none
[0235] If usePrevPrevRefFrameFlag = 1, PrevPrev-Resam / Next Next, else none
[0236] If resampling is disabled, the uncompensated reference frame Prev-Zero is selected. The inter prediction list is as follows:
[0237] Prev-Zero / Next
[0238] Prev-Zero / NextNext
[0239] If usePrevPrevRefFrameFlag = 1, PrevPrev / Next, else none
[0240] If usePrevPrevRefFrameFlag = 1, PrevPrev / NextNext, else none
[0241] In this second example, when global motion is enabled:
[0242] If resampling is enabled, the resampled reference frame Prev-Resam is selected. The inter prediction list is as follows:
[0243] Prev-Resam / Next
[0244] Prev-Resam / NextNext
[0245] If usePrevPrevRefFrameFlag = 1, PrevPrev-Resam / Next, else none
[0246] If usePrevPrevRefFrameFlag = 1, PrevPrev-Resam / NextNext, else none
[0247] If resampling is disabled, the uncompensated reference frame Prev-Zero is selected. The inter prediction list is as follows:
[0248] Prev-Zero / Next
[0249] Prev-Zero / NextNext
[0250] If usePrevPrevRefFrameFlag = 1, PrevPrev / Next, else none
[0251] If usePrevPrevRefFrameFlag = 1, PrevPrev / NextNext, else none
[0252] Examples in various aspects of the present disclosure can be used individually or in any combination.
[0253] Figure 10 is a flowchart illustrating an example inter prediction candidate selection technique in accordance with one or more aspects of the present disclosure. G-PCC encoder 200 or G-PCC decoder 300 can process a syntax element indicating whether a pre-pre frame can be used for inter prediction of one or more points of point cloud data, where the pre-pre frame comprises a reference frame of a previous reference frame (1000). For example, G-PCC encoder 200 can determine whether a pre-pre frame can be used for inter prediction of one or more points and encode and signal a syntax element indicating whether the pre-pre frame can be used for inter prediction of the one or more points. G-PCC decoder 300 can receive the syntax element in a bitstream and parse the syntax element to determine a value of the syntax element. The value of the syntax element can indicate whether the pre-pre frame can be used for inter prediction of the one or more points.
[0254] G-PCC encoder 200 or G-PCC decoder 300 can code the one or more points based on the determination of whether the pre-pre frame can be used for inter prediction of the one or more points (1002). For example, G-PCC encoder 200 can encode the one or more points based on whether the pre-pre frame can be used for inter prediction of the one or more points and G-PCC decoder 300 can decode the one or more points based on whether the pre-pre frame can be used for inter prediction of the one or more points. For example, if the pre-pre frame can be used for inter prediction of the one or more points, G-PCC encoder 200 and G-PCC decoder 300 can use the pre-pre frame to code the point cloud data for inter prediction of the one or more points. If the pre-pre frame cannot be used for inter prediction, G-PCC encoder 200 and G-PCC decoder 300 can code the point cloud data without using the pre-pre frame for inter prediction of the one or more points.
[0255] In some examples, the pre-pre frame can be used for inter prediction. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine not to apply resampling to a current frame of point cloud data. In such examples, G-PCC encoder 200 or G-PCC decoder 300 can apply a first set of one or more global motion parameters signaled with the current frame to the pre-pre frame.
[0256] In some examples, the pre-pre-frame can be used for inter prediction. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine to apply resampling to the current frame of point cloud data. In such examples, G-PCC encoder 200 or G-PCC decoder 300 can process (e.g., signal or parse) a second set of one or more global motion parameters and apply the second set of one or more global motion parameters to the pre-pre-frame. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can also process (e.g., signal or parse) a first set of global motion parameters and apply the first set of global motion parameters to the previous frame.
[0257] In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine that global motion is not enabled. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can disable resampling based on global motion not being enabled. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can use an uncompensated reference frame instead of a resampled reference frame.
[0258] In some examples, G-PCC encoder 200 or G-PCC decoder 300 can use a next candidate from the reference frames instead of a next-next candidate, which is a candidate after the first and second candidates in the inter prediction factor candidate list.
[0259] In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine that global motion compensation is enabled and process a flag at a frame level or slice level that indicates whether to apply global motion compensation to the current frame of point cloud data. In some examples, global motion compensation will not be applied to the current frame. In such examples, G-PCC encoder 200 or G-PCC decoder 300 can reject the last two entries of the inter prediction factor candidate list. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine that the pre-pre-frame is not used for inter prediction and reject the last two entries of the inter prediction factor candidate list based on the pre-pre-frame not being used for inter prediction.
[0260] In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine that global motion compensation is disabled. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine that resampling is enabled. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can select a previous resampled reference frame as a reference frame for inter prediction based on global motion compensation being disabled and resampling being enabled.
[0261] In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine that global motion compensation is disabled. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine that resampling is disabled. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can select, based on global motion compensation being disabled and resampling being disabled, a previously non-resampled reference frame as the reference frame for inter prediction.
[0262] In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine that global motion compensation is enabled. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine that resampling is enabled. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can select, based on global motion compensation being enabled and resampling being enabled, a previously resampled reference frame as the reference frame for inter prediction.
[0263] In some examples, coding includes decoding, and processing the syntax element includes parsing the syntax element to determine a value of the syntax element, the value indicating whether a pre-preceding frame can be used for inter prediction. In some examples, coding includes encoding, and processing the syntax element includes encoding the syntax element to include a value indicating whether a pre-preceding frame can be used for inter prediction.
[0264] Figure 11 is a conceptual diagram illustrating an example ranging system 1100 that can be used with one or more techniques of this disclosure. In Figure 11 In examples of the ranging system 1100, the ranging system 1100 includes an illuminator 1102 and a sensor 1104. The illuminator 1102 can emit light 1106. In some examples, the illuminator 1102 can emit the light 1106 as one or more laser beams. The light 1106 can be at one or more wavelengths, such as infrared wavelengths or visible light wavelengths. In other examples, the light 1106 is not coherent laser light. When the light 1106 encounters an object, such as an object 1108, the light 1106 produces return light 1110. The return light 1110 can include backscatter light and / or reflected light. The return light 1110 can pass through a lens 1111 that directs the return light 1110 to produce an image 1112 of the object 1108 on the sensor 1104. The sensor 1104 generates a signal 1114 based on the image 1112. The image 1112 can include a set of points (e.g., points as represented by the small dots of the image 1112). Figure 11 In examples of the ranging system 1100, the ranging system 1100 includes an illuminator 1102 and a sensor 1104. The illuminator 1102 can emit light 1106. In some examples, the illuminator 1102 can emit the light 1106 as one or more laser beams. The light 1106 can be at one or more wavelengths, such as infrared wavelengths or visible light wavelengths. In other examples, the light 1106 is not coherent laser light. When the light 1106 encounters an object, such as an object 1108, the light 1106 produces return light 1110. The return light 1110 can include backscatter light and / or reflected light. The return light 1110 can pass through a lens 1111 that directs the return light 1110 to produce an image 1112 of the object 1108 on the sensor 1104. The sensor 1104 generates a signal 1114 based on the image 1112. The image 1112 can include a set of points (e.g., points as represented by the small dots of the image 1112).
[0265] In some examples, the illuminator 1102 and the sensor 1104 can be mounted on a rotating structure such that the illuminator 1102 and the sensor 1104 capture a 360-degree view of the environment (e.g., a rotating LIDAR sensor). In other examples, the ranging system 1100 can include one or more optical components (e.g., mirrors, collimators, diffraction gratings, etc.) that enable the illuminator 1102 and the sensor 1104 to detect objects within a particular range (e.g., a maximum of 360 degrees). Although Figure 11 Examples of the ranging system 1100 show only a single illuminator 1102 and sensor 1104, but the ranging system 1100 can include multiple sets of illuminators and sensors.
[0266] In some examples, the illuminator 1102 generates a structured light pattern. In such examples, the ranging system 1100 can include multiple sensors 1104 on which respective images of the structured light pattern are formed. The ranging system 1100 can use differences between the images of the structured light pattern to determine distances to objects 1108 from which the structured light pattern backscatters. Structured light-based ranging systems can have a high level of accuracy (e.g., accuracy in the sub-millimeter range) when the objects 1108 are relatively close to the sensors 1104 (e.g., 0.2 meters to 2 meters). This high level of accuracy can be useful in facial recognition applications such as unlocking a mobile device (e.g., a mobile phone, a tablet computer, etc.) and for security applications.
[0267] In some examples, the ranging system 1100 is a time-of-flight (ToF)-based system. In some examples in which the ranging system 1100 is a ToF-based system, the illuminator 1102 generates pulses of light. In other words, the illuminator 1102 can modulate the amplitude of the emitted light 1106. In such examples, the sensor 1104 detects the return light 1110 from the pulses of light 1106 generated by the illuminator 1102. The ranging system 1100 can then determine a distance to the object 1108 from which the light 1106 backscatters based on the delay between when the light 1106 is emitted and when it is detected and the known speed of light in air. In some examples, the illuminator 1102 can modulate the phase of the emitted light 1106 instead of (or in addition to) modulating the amplitude of the emitted light 1106. In such examples, the sensor 1104 can detect the phase of the return light 1110 from the object 1108 and determine a distance to a point on the object 1108 using the speed of light and based on the time difference between when the illuminator 1102 generates the light 1106 at a particular phase and when the sensor 1104 detects the return light 1110 at that particular phase.
[0268] In other examples, a point cloud can be generated without using illuminator 1102. For example, in some examples, sensors 1104 of ranging system 1100 can include two or more optical cameras. In such examples, ranging system 1100 can use the optical cameras to capture stereo images of an environment that includes object 1108. Ranging system 1100 can include a point cloud generator 1116 that can compute differences between locations in the stereo images. Ranging system 1100 can then use these differences to determine distances to locations shown in the stereo images. From these distances, point cloud generator 1116 can generate a point cloud.
[0269] Sensor 1104 can also detect other properties of object 1108, such as color and reflectivity information. In Figure 11 examples, point cloud generator 1116 can generate a point cloud based on signals 1114 generated by sensor 1104. Ranging system 1100 and / or point cloud generator 1116 can form part of data source 104 Figure 1 ). Thus, a point cloud generated by ranging system 1100 can be encoded and / or decoded according to any of the techniques in this disclosure.
[0270] Figure 12 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of this disclosure can be used. In Figure 12 examples, vehicle 1200 includes a ranging system 1202. Ranging system 1202 can be implemented in the manner discussed above with Figure 11 reference to ranging system 1100. Although Figure 12 not shown in the example of vehicle 1200, vehicle 1200 can also include a data source (such as data source 104 Figure 1 ) and a G-PCC encoder (such as G-PCC encoder 200 Figure 1 ). In Figure 12 examples, ranging system 1202 emits laser beams 1204 that reflect off of pedestrian 1206 or other objects on the road. A data source of vehicle 1200 can generate a point cloud based on signals generated by ranging system 1202. A G-PCC encoder of vehicle 1200 can encode the point cloud to generate a bitstream 1208, such as a geometry bitstream Figure 2 ) and an attribute bitstream Figure 2 . Inter prediction and residual prediction as described in this disclosure can reduce the size of the geometry bitstream. Bitstream 1208 can include many fewer bits than the unencoded point cloud obtained by the G-PCC encoder.
[0271] An output interface of vehicle 1200 (e.g., output interface 108 Figure 1The bitstream 1208 can be transmitted to one or more other devices. The bitstream 1208 can include many fewer bits than the unencoded point cloud obtained by the G-PCC encoder. Thus, the vehicle 1200 can be able to transmit the bitstream 1208 to other devices faster than the unencoded point cloud data. Additionally, the bitstream 1208 can require less data storage capacity on the devices.
[0272] In the example of FIG. 12, the vehicle 1200 can transmit the bitstream 1208 to another vehicle 1210. The vehicle 1210 can include a G-PCC decoder, such as the G-PCC decoder 300 Figure 12 ). The G-PCC decoder of the vehicle 1210 can decode the bitstream 1208 to reconstruct the point cloud. The vehicle 1210 can use the reconstructed point cloud for various purposes. For example, the vehicle 1210 can determine, based on the reconstructed point cloud, that the pedestrian 1206 is on the road in front of the vehicle 1200 and thus begin to slow down, e.g., even before a driver of the vehicle 1210 is aware of the pedestrian 1206 on the road. Thus, in some examples, the vehicle 1210 can perform autonomous navigation operations based on the reconstructed point cloud. Figure 1
[0273] Additionally or alternatively, the vehicle 1200 can transmit the bitstream 1208 to a server system 1212. The server system 1212 can use the bitstream 1208 for various purposes. For example, the server system 1212 can store the bitstream 1208 for subsequent reconstruction of the point cloud. In this example, the server system 1212 can use the point cloud, along with other data (e.g., vehicle telemetry data generated by the vehicle 1200), to train an autonomous driving system. In other examples, the server system 1212 can store the bitstream 1208 for subsequent reconstruction for accident forensics investigations.
[0274] Figure 13 is a conceptual diagram illustrating an example extended reality system in which one or more techniques of this disclosure can be used. Extended reality (XR) is a term used to cover a range of technologies including augmented reality (AR), mixed reality (MR), and virtual reality (VR). In Figure 13 In the example of FIG. 13, a user 1300 is located in a first location 1302. The user 1300 wears an XR headset 1304. As an alternative to the XR headset 1304, the user 1300 can use a mobile device (e.g., a mobile phone, a tablet computer, etc.). The XR headset 1304 includes a depth detection sensor, such as a ranging system, that detects locations of points on an object 1306 at the location 1302. A data source of the XR headset 1304 can generate a point cloud representation of the object 1306 at the location 1302 using signals generated by the depth detection sensor. The XR headset 1304 can include a G-PCC encoder (e.g., the G-PCC encoder 200 of Figure 1 FIG. 13) configured to encode the point cloud to generate a bitstream 1308. Inter- frame prediction and residual prediction as described in this disclosure can reduce the size of the bitstream 1308.
[0275] The XR headset 1304 can transmit the bitstream 1308 to an XR headset 1310 worn by a user 1312 at a second location 1314 (e.g., via a network, such as the Internet). The XR headset 1310 can decode the bitstream 1308 to reconstruct the point cloud. The XR headset 1310 can use the point cloud to generate an XR visualization (e.g., an AR visualization, an MR visualization, a VR visualization) representing the object 1306 at the location 1302. Thus, in some examples, such as when the XR headset 1310 generates a VR visualization, the user 1312 can have a 3D immersive experience of the location 1302. In some examples, the XR headset 1310 can determine a location of a virtual object based on the reconstructed point cloud. For example, the XR headset 1310 can determine, based on the reconstructed point cloud, that the environment (e.g., the location 1302) includes a flat surface, and then determine that a virtual object (e.g., a cartoon character) is to be positioned on the flat surface. The XR headset 1310 can generate an XR visualization in which the virtual object is located at the determined location. For example, the XR headset 1310 can show the cartoon character sitting on the flat surface.
[0276] Figure 14 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure can be used. In the example of FIG. 14, a mobile device 1400 (e.g., a wireless communication device), such as a mobile phone or a tablet computer, includes a ranging system (such as a LIDAR system) that detects locations of points on an object 1402 in an environment of the mobile device 1400. A data source of the mobile device 1400 can generate a point cloud representation of the object 1402 using signals generated by the depth detection sensor. The mobile device 1400 can include a G-PCC encoder (e.g., the G-PCC encoder 200 of Figure 14 In the example of FIG. 13, a user 1300 is located in a first location 1302. The user 1300 wears an XR headset 1304. As an alternative to the XR headset 1304, the user 1300 can use a mobile device (e.g., a mobile phone, a tablet computer, etc.). The XR headset 1304 includes a depth detection sensor, such as a ranging system, that detects locations of points on an object 1306 at the location 1302. A data source of the XR headset 1304 can generate a point cloud representation of the object 1306 at the location 1302 using signals generated by the depth detection sensor. The XR headset 1304 can include a G-PCC encoder (e.g., the G-PCC encoder 200 of Figure 1 FIG. 13) configured to encode the point cloud to generate a bitstream 1308. Inter- frame prediction and residual prediction as described in this disclosure can reduce the size of the bitstream 1308.
[0275] The XR headset 1304 can transmit the bitstream 1308 to an XR headset 1310 worn by a user 1312 at a second location 1314 (e.g., via a network, such as the Internet). The XR headset 1310 can decode the bitstream 1308 to reconstruct the point cloud. The XR headset 1310 can use the point cloud to generate an XR visualization (e.g., an AR visualization, an MR visualization, a VR visualization) representing the object 1306 at the location 1302. Thus, in some examples, such as when the XR headset 1310 generates a VR visualization, the user 1312 can have a 3D immersive experience of the location 1302. In some examples, the XR headset 1310 can determine a location of a virtual object based on the reconstructed point cloud. For example, the XR headset 1310 can determine, based on the reconstructed point cloud, that the environment (e.g., the location 1302) includes a flat surface, and then determine that a virtual object (e.g., a cartoon character) is to be positioned on the flat surface. The XR headset 1310 can generate an XR visualization in which the virtual object is located at the determined location. For example, the XR headset 1310 can show the cartoon character sitting on the flat surface.
[0276] Figure 14 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure can be used. In the example of FIG. 14, a mobile device 1400 (e.g., a wireless communication device), such as a mobile phone or a tablet computer, includes a ranging system (such as a LIDAR system) that detects locations of points on an object 1402 in an environment of the mobile device 1400. A data source of the mobile device 1400 can generate a point cloud representation of the object 1402 using signals generated by the depth detection sensor. The mobile device 1400 can include a G-PCC encoder (e.g., the G-PCC encoder 200 of Figure 14 In the example of FIG. 13, a user 1300 is located in a first location 1302. The user 1300 wears an XR headset 1304. As an alternative to the XR headset 1304, the user 1300 can use a mobile device (e.g., a mobile phone, a tablet computer, etc.). The XR headset 1304 includes a depth detection sensor, such as a ranging system, that detects locations of points on an object 1306 at the location 1302. A data source of the XR headset 1304 can generate a point cloud representation of the object 1306 at the location 1302 using signals generated by the depth detection sensor. The XR headset 1304 can include a G-PCC encoder (e.g., the G-PCC encoder 200 ofFigure 1 G-PCC encoder 200) configured to encode a point cloud to generate a bitstream 1404. In Figure 14 In examples of the mobile device 1400, the mobile device 1400 can transmit the bitstream to a remote device 1406, such as a server system or other mobile device. Inter- frame prediction and residual prediction as described in this disclosure can reduce the size of the bitstream 1404. The remote device 1406 can decode the bitstream 1404 to reconstruct the point cloud. The remote device 1406 can use the point cloud for various purposes. For example, the remote device 1406 can use the point cloud to generate an environmental map of the mobile device 1400. For example, the remote device 1406 can generate a map of an interior of a building based on the reconstructed point cloud. In another example, the remote device 1406 can generate an image (e.g., computer graphics) based on the point cloud. For example, the remote device 1406 can use points in the point cloud as vertices of polygons and use color attributes of the points as a basis for shading the polygons. In some examples, the remote device 1406 can use the reconstructed point cloud for facial recognition or other security applications.
[0277] This disclosure includes the following non-limiting clauses.
[0278] Clause 1A. A method of processing point cloud data, the method comprising: signaling or parsing a syntax element indicating whether a pre-pre frame can be used for inter- frame prediction, wherein the pre-pre frame comprises a reference frame of a previous reference frame; and processing the point cloud data based on the syntax element.
[0279] Clause 2A. The method of clause 1A, wherein the pre-pre frame can be used for inter- frame prediction, and the method further comprises: determining that no resampling is applied to a current frame of the point cloud data; and applying global motion parameters signaled with the current frame to the pre-pre frame.
[0280] Clause 3A. The method of clause 1A, wherein the pre-pre frame can be used for inter- frame prediction, and the method further comprises: determining that resampling is applied to a current frame of the point cloud data; signaling or parsing a second set of global motion parameters; and applying the second set of global motion parameters to the pre-pre frame.
[0281] Clause 4A. The method of clause 1A, the method further comprising: determining that global motion is not enabled; and based on global motion not being enabled, disabling resampling.
[0282] Clause 5A. The method of clause 4A, the method further comprising using an uncompensated reference frame instead of a resampled reference frame.
[0283] Clause 6A. The method of any of clauses 1A-5A, further comprising using a next candidate from a reference frame in place of a next-next candidate, the next-next candidate being a candidate after a first candidate and a second candidate in an inter prediction factor candidate list.
[0284] Clause 7A. The method of any of clauses 1A-3A or 6A, further comprising: determining that global motion compensation is enabled; and signaling or parsing, at a frame level or a slice level, a flag indicating whether to apply global motion compensation to a current frame of the point cloud data.
[0285] Clause 8A. The method of clause 7A, wherein global motion compensation is not to be applied to the current frame, and wherein the method further comprises rejecting last two entries of an inter prediction factor candidate list.
[0286] Clause 9A. The method of clause 8A, further comprising determining that the previous-previous frame is not used for inter prediction, and wherein the last two entries of the inter prediction factor candidate list are rejected based on the previous-previous frame not being used for inter prediction.
[0287] Clause 10A. The method of clause 1A, further comprising: determining that global motion compensation is disabled; determining that resampling is enabled; based on global motion compensation being disabled and resampling being enabled; selecting a previously resampled reference frame as a reference frame for inter prediction.
[0288] Clause 11A. The method of clause 1A, further comprising: determining that global motion compensation is disabled; determining that resampling is disabled; based on global motion compensation being disabled and resampling being disabled; selecting a previously non-resampled reference frame as a reference frame for inter prediction.
[0289] Clause 12A. The method of clause 1A, further comprising: determining that global motion compensation is enabled; determining that resampling is enabled; based on global motion compensation being enabled and resampling being enabled, selecting a previously resampled reference frame as a reference frame for inter prediction.
[0290] Clause 13A. The method of any of clauses 1A-12A, further comprising generating the point cloud.
[0291] Clause 14A. An apparatus for processing a point cloud, the apparatus comprising one or more means for performing the method of any of clauses 1A-13A.
[0292] Clause 15A. The apparatus of clause 14A, wherein the one or more means comprise one or more processors implemented in circuitry.
[0293] Clause 16A. The device of any of clauses 14A or 15A, further comprising a memory to store the data representative of the point cloud.
[0294] Clause 17A. The device of any of clauses 14A-16A, wherein the device comprises a decoder.
[0295] Clause 18A. The device of any of clauses 14A-17A, wherein the device comprises an encoder.
[0296] Clause 19A. The device of any of clauses 14A-18A, further comprising a device to generate the point cloud.
[0297] Clause 20A. The device of any of clauses 14A-19A, further comprising a display to present an image based on the point cloud.
[0298] Clause 21A. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to perform the method of any of clauses 1A-13A.
[0299] Clause 1B. A method of processing point cloud data, the method comprising: processing a syntax element that indicates whether a triple previous frame can be used for inter prediction of one or more points of the point cloud data, wherein the triple previous frame comprises a reference frame of a previous reference frame; and based on a determination of whether the triple previous frame can be used for inter prediction of the one or more points, coding the one or more points.
[0300] Clause 2B. The method of clause 1B, wherein the triple previous frame can be used for inter prediction, and the method further comprises: determining not to apply resampling to a current frame of the point cloud data; and applying a first set of one or more global motion parameters signaled with the current frame to the triple previous frame.
[0301] Clause 3B. The method of clause 1B, wherein the triple previous frame can be used for inter prediction, and the method further comprises: determining to apply resampling to a current frame of the point cloud data; processing a second set of one or more global motion parameters; and applying the second set of one or more global motion parameters to the triple previous frame.
[0302] Clause 4B. The method of clause 1B, further comprising: determining that global motion is not enabled; and based on global motion not being enabled, disabling resampling.
[0303] Clause 5B. The method of clause 4B, further comprising using an uncompensated reference frame instead of a resampled reference frame.
[0304] Clause 6B. The method of any of clauses 1B-5B, further comprising using a next candidate from a reference frame instead of a next-next candidate, the next-next candidate being a candidate after a first candidate and a second candidate in an inter prediction factor candidate list.
[0305] Clause 7B. The method of any of clauses 1B-3B or 6B, further comprising: determining that global motion compensation is enabled; and processing a flag at a frame level or a slice level that indicates whether to apply global motion compensation to a current frame of the point cloud data.
[0306] Clause 8B. The method of clause 7B, wherein global motion compensation is not to be applied to the current frame, and wherein the method further comprises rejecting last two entries of an inter prediction factor candidate list.
[0307] Clause 9B. The method of clause 8B, further comprising determining that the previous-previous frame is not used for inter prediction, and wherein the last two entries of the inter prediction factor candidate list are rejected based on the previous-previous frame not being used for inter prediction.
[0308] Clause 10B. The method of clause 1B, further comprising: determining that global motion compensation is disabled; determining that resampling is enabled; and based on global motion compensation being disabled and resampling being enabled, selecting a previously resampled reference frame as a reference frame for inter prediction.
[0309] Clause 11B. The method of clause 1B, further comprising: determining that global motion compensation is disabled; determining that resampling is disabled; and based on global motion compensation being disabled and resampling being disabled, selecting a previously uncompensated reference frame as a reference frame for inter prediction.
[0310] Clause 12B. The method of clause 1B, further comprising: determining that global motion compensation is enabled; determining that resampling is enabled; and based on global motion compensation being enabled and resampling being enabled, selecting a previously resampled reference frame as a reference frame for inter prediction.
[0311] Clause 13B. The method of any of clauses 1B-12B, wherein coding comprises decoding, and processing the syntax element comprises parsing the syntax element to determine a value of the syntax element, the value indicating whether the previous-previous frame can be used for inter prediction.
[0312] Clause 14B. The method of any of clauses 1B-12B, wherein coding includes encoding and processing the syntax element includes encoding the syntax element to include a value indicating whether the pre-pre frame can be used for inter prediction.
[0313] Clause 15B. A device for processing a point cloud, the device comprising: one or more memories configured to store point cloud data; and one or more processors implemented in circuitry and communicatively coupled to the one or more memories, the one or more processors configured to: process a syntax element indicating whether a pre-pre frame can be used for inter prediction of one or more points of the point cloud data, wherein the pre-pre frame comprises a reference frame of a previous reference frame; and based on a determination of whether the pre-pre frame can be used for inter prediction of the one or more points, code the one or more points.
[0314] Clause 16B. The device of clause 15B, wherein the pre-pre frame can be used for inter prediction and the one or more processors are further configured to: determine not to apply resampling to a current frame of the point cloud data; and apply a first set of one or more global motion parameters signaled with the current frame to the pre-pre frame.
[0315] Clause 17B. The device of clause 15B, wherein the pre-pre frame can be used for inter prediction and the one or more processors are further configured to: determine to apply resampling to a current frame of the point cloud data; process a second set of one or more global motion parameters; and apply the second set of one or more global motion parameters to the pre-pre frame.
[0316] Clause 18B. The device of any of clauses 15B-17B, wherein the device comprises a decoder configured to parse the syntax element, and wherein as part of parsing the syntax element, the one or more processors are configured to determine a value of the syntax element, the value indicating whether the pre-pre frame can be used for inter prediction.
[0317] Clause 19B. The device of any of clauses 15B-17B, wherein the device comprises an encoder configured to signal the syntax element, and wherein the syntax element has a value indicating whether the pre-pre frame can be used for inter prediction.
[0318] Clause 20B. A non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to: process a syntax element that indicates whether a pre-pre frame can be used for inter prediction of one or more points of point cloud data, wherein the pre-pre frame comprises a reference frame of a previous reference frame; and based on a determination of whether the pre-pre frame can be used for inter prediction of the one or more points, code the one or more points.
[0319] It is recognized that, in accordance with examples, certain acts or events of any of the techniques described herein can be performed in a different sequence, can be added, merged, or omitted altogether (for example, not all described acts or events are necessary for implementation of that technique). Moreover, in certain examples, acts or events can be performed concurrently, for example, through multi-threaded processing, interrupt processing, or multiple processors, rather than sequentially.
[0320] In one or more examples, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer- readable media generally can correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product can include a computer-readable medium.
[0321] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any
[0322] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the terms "processor" and "processing circuitry," as used herein can refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0323] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware components. Rather, as described above, various
[0324] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A method of processing point cloud data, the method comprising: processing a syntax element that indicates whether a previous previous frame can be used for inter prediction of one or more points of the point cloud data, wherein the previous previous frame comprises a reference frame of a previous reference frame; and based on a determination of whether the previous previous frame can be used for inter prediction of the one or more points, coding the one or more points.
2. The method of claim 1, wherein, the previous previous frame can be used for inter prediction, and the method further comprising: determining that no resampling is applied to a current frame of the point cloud data; and applying a first set of one or more global motion parameters signaled with the current frame to the previous previous frame.
3. The method of claim 1, wherein, the previous previous frame can be used for inter prediction, and the method further comprising: determining that resampling is applied to a current frame of the point cloud data; processing a second set of one or more global motion parameters; and applying the second set of one or more global motion parameters to the previous previous frame.
4. The method of claim 1, the method further comprising: determining that global motion is not enabled; and based on global motion not being enabled, disabling resampling.
5. The method of claim 4, the method further comprising using an uncompensated reference frame instead of a resampled reference frame.
6. The method of claim 1, the method further comprising using a next candidate from a reference frame instead of a next next candidate, the next next candidate being a candidate after a first candidate and a second candidate in an inter prediction factor candidate list.
7. The method of claim 1, the method further comprising: determining that global motion compensation is enabled; and processing, at a frame level or a slice level, a flag that indicates whether global motion compensation is to be applied to a current frame of the point cloud data.
8. The method of claim 7, wherein, global motion compensation is not to be applied to the current frame, and wherein the method further comprises rejecting last two entries of an inter prediction factor candidate list.
9. The method of claim 8, the method further comprising determining that the pre- preceding frame is not used for inter prediction, and wherein, rejecting the last two entries of the inter prediction factor candidate list based on the previous previous frame not being used for inter prediction.
10. The method of claim 1, the method further comprising: determining that global motion compensation is disabled; determining that resampling is enabled; and based on global motion compensation being disabled and resampling being enabled, selecting a previously resampled reference frame as a reference frame for inter prediction.
11. The method of claim 1, the method further comprising: determining that global motion compensation is disabled; determining that resampling is disabled; and based on global motion compensation being disabled and resampling being disabled, selecting a previously uncompensated reference frame as a reference frame for inter prediction.
12. The method of claim 1, the method further comprising: determining that global motion compensation is enabled; determining that resampling is enabled; and based on global motion compensation being enabled and resampling being enabled, selecting a previously resampled reference frame as a reference frame for inter prediction.
13. The method of claim 1, wherein, coding comprises decoding, and processing the syntax element comprises parsing the syntax element to determine a value of the syntax element, the value indicating whether the previous previous frame can be used for inter prediction.
14. The method of claim 1, wherein, Decoding includes encoding, and processing the syntax element includes encoding the syntax element to include a value indicating whether the pre-pre frame can be used for inter prediction.
15. A device for processing a point cloud, the device comprising: one or more memories configured to store point cloud data; and one or more processors implemented in circuitry and communicatively coupled to the one or more memories, the one or more processors configured to: process a syntax element indicating whether a pre-pre frame can be used for inter prediction of one or more points of the point cloud data, wherein the pre-pre frame comprises a reference frame of a previous reference frame; and based on a determination of whether the pre-pre frame can be used for inter prediction of the one or more points, code the one or more points.
16. The apparatus of claim 15, wherein, The pre-pre frame can be used for inter prediction, and the one or more processors are further configured to: determine not to apply resampling to a current frame of the point cloud data; and apply the pre-pre frame with a first set of one or more global motion parameters signaled with the current frame.
17. The apparatus of claim 15, wherein, The pre-pre frame can be used for inter prediction, and the one or more processors are further configured to: determine to apply resampling to a current frame of the point cloud data; process a second set of one or more global motion parameters; and apply the pre-pre frame with the second set of one or more global motion parameters.
18. The apparatus of claim 15, wherein, The device comprises a decoder configured to parse the syntax element, and wherein, as part of parsing the syntax element, the one or more processors are configured to determine a value of the syntax element, the value indicating whether the pre-pre frame can be used for inter prediction.
19. The apparatus of claim 15, wherein, The device comprises an encoder configured to signal the syntax element, and wherein the syntax element has a value indicating whether the pre-pre frame can be used for inter prediction.
20. A non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to: processing a syntax element that indicates whether a previous frame can be used for inter prediction of one or more points of the point cloud data, wherein The pre-pre frame comprises a reference frame of a previous reference frame; and based on a determination of whether the pre-pre frame can be used for inter prediction of the one or more points, code the one or more points.