Predictive geometry coding of point clouds

By generating a prediction tree and determining branch termination and new creation based on the azimuth difference between point and cloud points, the problem of inefficient prediction geometry encoding and decoding in the prior art is solved, and more efficient point cloud data encoding and decoding is achieved.

CN120019411APending Publication Date: 2025-05-16QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072016.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-13
Filing Date
2023-10-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the prior art uses prediction geometry to code point clouds, the suboptimal structure of the tree leads to inefficient encoding and decoding efficiency, and when the LIDAR system does not rotate 360 ​​degrees, the prediction relationship between point clouds and points is complex and difficult to effectively utilize.

Method used

By generating a prediction tree, the prediction geometry codec can utilize the inherent dependence between points in the LIDAR system, avoiding the point at which the next scan begins by using the point at which the saccade ends. The specific method includes determining the azimuth difference between consecutive points of the point cloud data, and determining whether to terminate the prediction tree branch and start a new branch based on the difference.

Benefits of technology

The efficiency of point cloud encoding and decoding is improved, and the point cloud data of the LIDAR system can be used more effectively, especially when the LIDAR system does not rotate completely 360 degrees.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019411A_ABST
    Figure CN120019411A_ABST
Patent Text Reader

Abstract

An example apparatus includes a memory configured to store point cloud data and one or more processors configured to determine a first point of the point cloud data as a first node of a first prediction tree branch. The one or more processors are configured to determine that a first azimuth difference between a first point and a second point of the point cloud data does not meet a first azimuth threshold, and based on the determination, determine the second point as a second node of the first prediction tree branch. The one or more processors are configured to determine that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies a first azimuth threshold, and based on the determination, determine the fourth point as a first node of a second prediction tree branch.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Patent Application Serial No. 18 / 486,541, filed on October 13, 2023, and U.S. Provisional Patent Application Serial No. 63 / 379,847, filed on October 17, 2022, the entire contents of which are incorporated herein by reference. U.S. Patent Application Serial No. 18 / 486,541, filed on October 13, 2023, claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 379,847, filed on October 17, 2022. Technical Field

[0002] The present disclosure relates to point cloud encoding and decoding. Background Art

[0003] A point cloud is a collection of points in a three-dimensional space. A point may correspond to a point on an object within the three-dimensional space. Therefore, a point cloud can be used to represent the physical contents of a three-dimensional space. Point clouds can have utility in a wide variety of situations. For example, a point cloud can be used in the context of an autonomous vehicle to represent the location of an object on a road. In another example, a point cloud can be used in the context of representing the physical contents of an environment for the purpose of locating virtual objects in augmented reality (AR) or mixed reality (MR) applications. Point cloud compression is a process used to encode and decode a point cloud. Encoding a point cloud can reduce the amount of data required to store and transmit the point cloud. Summary of the invention

[0004] Generally, this disclosure describes techniques for predictive geometry coding of point cloud data, and in particular, relates to the generation of prediction trees for predictive geometry coding.

[0005] A LIDAR system may include one or more laser sources and detectors / sensors, and each sensor may capture points (e.g., points of a point cloud) in a predetermined order. When encoding and decoding points using a point cloud compression codec, points belonging to different lasers / sensors are typically encoded and decoded together. When encoding and decoding a point cloud using predicted geometry, a prediction tree may be generated. Suboptimal construction of the tree may result in inefficient encoding and decoding. Therefore, it is desirable to generate a tree such that the predicted geometry codec can exploit the inherent dependencies between the individual points of the sensor.

[0006] In one example, the present disclosure describes a device for encoding and decoding point cloud data, including: one or more memories configured to store point cloud data; and one or more processors implemented as circuits and communicatively coupled to the one or more memories, the one or more processors configured to: determine a first point of the point cloud data as a first node of a first prediction tree branch of a prediction tree; determine that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point include sequentially consecutive points; determine the second point as a first node of a first prediction tree branch of a prediction tree based on the first azimuth difference not satisfying the first azimuth threshold; A second node of a prediction tree branch; determining that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies a first azimuth threshold, wherein the third point and the fourth point include consecutive points in sequence, and wherein the third point includes a third node of a first prediction tree branch; based on the second azimuth difference satisfying the first azimuth threshold, terminating the first prediction tree branch at the third point, so that the third point includes a leaf node of the first prediction tree branch, and determining the fourth point as a first node of a second prediction tree branch; connecting the first prediction tree branch and the second prediction tree branch in the prediction tree; and encoding and decoding the point cloud data based on the prediction tree.

[0007] In another example, the present disclosure describes a method for encoding and decoding point cloud data, the method comprising: determining a first point of the point cloud data as a first node of a first prediction tree branch of a prediction tree; determining that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point include sequentially consecutive points; based on the first azimuth difference not satisfying the first azimuth threshold, determining the second point as a second node of the first prediction tree branch; determining that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point include sequentially consecutive points, and wherein the third point includes a third node of the first prediction tree branch; based on the second azimuth difference satisfying the first azimuth threshold, terminating the first prediction tree branch at the third point, so that the third point includes a leaf node of the first prediction tree branch, and determining the fourth point as a first node of the second prediction tree branch; connecting the first prediction tree branch and the second prediction tree branch in the prediction tree; and encoding and decoding the point cloud data based on the prediction tree.

[0008] In another example, the present disclosure describes a device for encoding and decoding point cloud data, the device comprising: a component for determining a first point of the point cloud data as a first node of a first prediction tree branch of a prediction tree; a component for determining that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point include sequentially consecutive points; a component for determining the second point as a second node of the first prediction tree branch based on the first azimuth difference not satisfying the first azimuth threshold; a component for determining that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point include sequentially consecutive points, and wherein the third point includes a third node of the first prediction tree branch; a component for terminating the first prediction tree branch at the third point based on the second azimuth difference satisfying the first azimuth threshold, so that the third point includes a leaf node of the first prediction tree branch and the fourth point is determined as a first node of the second prediction tree branch; a component for connecting the first prediction tree branch and the second prediction tree branch in the prediction tree; and a component for encoding and decoding the point cloud data based on the prediction tree.

[0009] In yet another example, the present disclosure describes a non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to: determine a first point of point cloud data as a first node of a first prediction tree branch of a prediction tree; determine that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point include sequentially consecutive points; based on the first azimuth difference not satisfying the first azimuth threshold, determine the second point as a second node of the first prediction tree branch; determine that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point include sequentially consecutive points, and wherein the third point includes a third node of the first prediction tree branch; based on the second azimuth difference satisfying the first azimuth threshold, terminate the first prediction tree branch at the third point, so that the third point includes a leaf node of the first prediction tree branch, and determine the fourth point as a first node of the second prediction tree branch; connect the first prediction tree branch and the second prediction tree branch in the prediction tree; and encode and decode the point cloud data based on the prediction tree.

[0010] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a block diagram illustrating an example encoding and decoding system that may perform the techniques of this disclosure.

[0012] Figure 2is a block diagram illustrating an example Geometric Point Cloud Compression (G-PCC) encoder.

[0013] Figure 3 is a block diagram illustrating an example G-PCC decoder.

[0014] Figure 4 is a conceptual diagram showing an example octree partitioning for geometry coding.

[0015] Figure 5 is a conceptual diagram illustrating an example of a prediction tree.

[0016] Fig. 6A and 6B is a conceptual diagram illustrating an example of a rotating light detection and ranging (LIDAR) acquisition model.

[0017] Figure 7 is a conceptual diagram of an example prediction tree according to one or more aspects of the present disclosure.

[0018] Figure 8 is a conceptual diagram of another example prediction tree according to one or more aspects of the present disclosure.

[0019] Fig. 9 is a flow chart illustrating an example technique for generating a prediction tree according to one or more aspects of the present disclosure.

[0020] Fig.10 is a flow chart illustrating an example prediction tree generation technique in accordance with one or more aspects of the present disclosure.

[0021] Fig.11 is a conceptual diagram illustrating an example ranging system 1100 that may be used with one or more techniques of this disclosure for coordinate transformation in G-PCC.

[0022] Fig.12 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more of the techniques of this disclosure for coordinate transformation in G-PCC may be used.

[0023] Fig.13 is a conceptual diagram illustrating an example extended reality system in which one or more techniques of the present disclosure for coordinate transformation in G-PCC may be used.

[0024] Fig.14 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure for coordinate conversion in G-PCC may be used. DETAILED DESCRIPTION

[0025] Not all LIDAR systems rotate a full 360 degrees. For example, some LIDAR systems may sweep over a more limited range (e.g., 120 degrees) and then return to the starting angle to perform the next sweep. As a result, consecutive points in a scan or capture sequence, for example, may sometimes be in very different locations, and one point may not provide an accurate prediction of the next consecutive point.

[0026] When encoding and decoding a point cloud using predicted geometry, a prediction tree may be generated. Suboptimal construction of the prediction tree may result in inefficient encoding and decoding. Therefore, it may be desirable to generate a prediction tree so that the predicted geometry encoding and decoding can exploit the inherent dependencies between the various points sensed by the LIDAR system and avoid using the point at the end of a sweep to predict the point at the beginning of the next sweep, such as when the LIDAR system does not rotate a full 360 degrees.

[0027] Figure 1 1 is a block diagram illustrating an example encoding and decoding system 100 that can perform the techniques of the present disclosure. The techniques of the present disclosure generally relate to encoding and decoding (encoding and / or decoding) point cloud data, i.e., supporting point cloud compression. In general, point cloud data includes any data used to process point clouds. The encoding and decoding can effectively compress and / or decompress point cloud data.

[0028] like Figure 1 As shown, system 100 includes a source device 102 and a destination device 116. Source device 102 provides encoded point cloud data to be decoded by destination device 116. Figure 1 In an example of , source device 102 provides point cloud data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 may include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as smart phones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, ground or sea vehicles, spacecraft, aircraft, robots, LIDAR devices, satellites, or the like. In some cases, source device 102 and destination device 116 may be equipped for wireless communication.

[0029] exist Figure 1In the example of , the source device 102 includes a data source 104, one or more memories 106, a G-PCC encoder 200, and an output interface 108. The destination device 116 includes an input interface 122, a G-PCC decoder 300, one or more memories 120, and a data consumer 118. According to the present disclosure, the G-PCC encoder 200 of the source device 102 and the G-PCC decoder 300 of the destination device 116 can be configured to apply the technology related to the generation of prediction trees for predictive geometry coding and decoding of the present disclosure. Therefore, the source device 102 represents an example of an encoding device, and the destination device 116 represents an example of a decoding device. In other examples, the source device 102 and the destination device 116 may include other components or arrangements. For example, the source device 102 can receive data (e.g., point cloud data) from an internal or external source. Similarly, the destination device 116 can interface with an external data consumer instead of including the data consumer in the same device.

[0030] like Figure 1 The system 100 shown is only an example. In general, other digital encoding and / or decoding devices can perform the technology related to the generation of prediction trees for predicting geometric coding and decoding disclosed in the present invention. The source device 102 and the destination device 116 are only examples of such devices in which the source device 102 generates coded data for sending to the destination device 116. The present disclosure refers to a "codec" device as a device that performs the coding and decoding (coding and / or decoding) of data. Therefore, the G-PCC encoder 200 and the G-PCC decoder 300 represent examples of codec devices, specifically, encoders and decoders, respectively. In some examples, the source device 102 and the destination device 116 can operate in a substantially symmetrical manner so that each of the source device 102 and the destination device 116 includes coding and decoding components. Therefore, the system 100 can support one-way or two-way transmission between the source device 102 and the destination device 116, such as for streaming, playback, broadcasting, telephone, navigation and other applications.

[0031] Typically, the data source 104 represents the source of the data (i.e., the original, unencoded point cloud data), and a series of continuous "frames" of data can be provided to the G-PCC encoder 200, and the G-PCC encoder 200 encodes the data of the frame. The data source 104 of the source device 102 may include a point cloud capture device, such as any of a variety of cameras or sensors, such as a 3D scanner or a light detection and ranging (LIDAR) device, one or more cameras, an archive containing previously captured data, and / or a data feed interface for receiving data from a data content provider. Alternatively or additionally, the point cloud data may be generated by a scanner, camera, sensor, or other data computer. For example, the data source 104 may generate computer graphics-based data as source data, or generate a combination of real-time data, archived data, and computer-generated data. In each case, the G-PCC encoder 200 encodes the captured, pre-captured, or computer-generated data. The G-PCC encoder 200 may rearrange the frames from the order in which they are received (sometimes referred to as the "display order") into a codec order for encoding and decoding. G-PCC encoder 200 may generate one or more bitstreams including the encoded data. Source device 102 may then output the encoded data onto computer-readable medium 110 via output interface 108 for receipt and / or retrieval by input interface 122 of destination device 116, for example.

[0032] The memory 106 of the source device 102 and the memory 120 of the destination device 116 can represent one or more general purpose memories. In some examples, the memory 106 and the memory 120 can store original data, for example, original data from the data source 104 and original, decoded data from the G-PCC decoder 300. Additionally or alternatively, the memory 106 and the memory 120 can store software instructions that can be executed by, for example, the G-PCC encoder 200 and the G-PCC decoder 300, respectively. Although the memory 106 and the memory 120 are shown separately from the G-PCC encoder 200 and the G-PCC decoder 300 in this example, it should be understood that the G-PCC encoder 200 and the G-PCC decoder 300 can also include internal memories for functionally similar or equivalent purposes. In addition, the memory 106 and the memory 120 can store encoded data, for example, output from the G-PCC encoder 200 and input to the G-PCC decoder 300. In some examples, portions of memory 106 and memory 120 may be allocated as one or more buffers, for example, to store raw, decoded, and / or encoded data. For example, memory 106 and memory 120 may store data representing a point cloud.

[0033] The computer-readable medium 110 may represent one or more of any type of medium or device capable of transmitting the encoded data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium that enables the source device 102 to send the encoded data directly to the destination device 116 in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard (e.g., a wireless communication protocol), the output interface 108 can modulate a transmission signal including the encoded data, and the input interface 122 can demodulate the received transmission signal. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include a router, a switch, a base station, or any other equipment that can be used to facilitate communication from the source device 102 to the destination device 116.

[0034] In some examples, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded data.

[0035] In some examples, source device 102 may output the encoded data to file server 114 or another intermediate storage device that may store the encoded data generated by source device 102. Destination device 116 may access the stored data from file server 114 via streaming or downloading. File server 114 may be any type of server device capable of storing the encoded data and sending the encoded data to destination device 116. File server 114 may represent a network server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 may access the encoded data from file server 114 via any standard data connection (including an Internet connection). This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both that is suitable for accessing the encoded data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.

[0036] Output interface 108 and input interface 122 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components that operate according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data (e.g., encoded data) according to a cellular communication standard (e.g., 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, or the like). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to transmit data (e.g., encoded data) according to other wireless standards (e.g., IEEE 802.11 specifications, IEEE 802.15 specifications (e.g., ZigBee TM ),Bluetooth TM Standard or the like). In some examples, source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include a SoC device for performing functions attributed to G-PCC encoder 200 and / or output interface 108, and destination device 116 may include a SoC device for performing functions attributed to G-PCC decoder 300 and / or input interface 122.

[0037] The techniques disclosed herein may be applied to encoding and decoding to support any of a variety of applications, such as communications between autonomous vehicles, communications between scanners, cameras, sensors and processing devices (such as local or remote servers), geographic mapping, or other applications.

[0038] The input interface 122 of the destination device 116 receives the encoded bitstream from the computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, and / or the like). The encoded bitstream may include signaling information defined by the G-PCC encoder 200, which is also used by the G-PCC decoder 300, such as syntax elements with values ​​describing characteristics and / or processing of units of the codec (e.g., slices, pictures, groups of pictures, sequences, and the like). The data consumer 118 uses the decoded data. For example, the data consumer 118 may use the decoded data to determine the location of a physical object. In some examples, the data consumer 118 may include a display for presenting an image based on a point cloud.

[0039] The G-PCC encoder 200 and the G-PCC decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device can store instructions for the software in a suitable non-temporary computer-readable medium, and use one or more processors to execute the instructions in hardware to perform the technology disclosed herein. Each of the G-PCC encoder 200 and the G-PCC decoder 300 can be included in one or more encoders or decoders, and any of the encoders or decoders can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device. The device including the G-PCC encoder 200 and / or the G-PCC decoder 300 can include one or more integrated circuits, microprocessors, and / or other types of devices.

[0040] The G-PCC encoder 200 and the G-PCC decoder 300 may operate according to a codec standard such as the Video Point Cloud Compression (V-PCC) standard or the Geometric Point Cloud Compression (G-PCC) standard. The present disclosure may generally refer to the coding and decoding (e.g., encoding and decoding) of a picture to include the process of encoding or decoding data. The encoded bitstream typically includes a series of values ​​of syntax elements representing coding decisions (e.g., coding modes).

[0041] The present disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to the communication of values ​​of syntax elements and / or other data used to decode encoded data. That is, the G-PCC encoder 200 may signal the values ​​of syntax elements in a bitstream. Typically, signaling refers to generating values ​​in a bitstream. As mentioned above, the source device 102 may transmit the bitstream to the destination device 116 in substantially real time or non-real time, such as may occur when storing syntax elements to the storage device 112 for later retrieval by the destination device 116.

[0042] ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) and more recently ISO / IEC MPEG 3DG (JTC 1 / SC29 / WG7) are investigating the potential need for standardization of point cloud codecs with compression capabilities significantly exceeding those of current methods, with the goal of creating a standard. MPEG is conducting this exploratory activity in a collaborative effort called the 3D Graphics Group (3DG) to evaluate compression technology designs proposed by their experts in the field.

[0043] Point cloud compression activities are categorized in two different approaches. The first approach is “Video Point Cloud Compression” (V-PCC), which segments 3D objects and projects the segments in multiple 2D planes (which are represented as “patches” in 2D frames), which are further encoded and decoded by a traditional 2D video codec such as the High Efficiency Video Codec (HEVC) (ITU-T H.265) codec. The second approach is “Geometry-based Point Cloud Compression” (G-PCC), which directly compresses the 3D geometry, i.e., the positions of a collection of points in 3D space, and the associated attribute values ​​(for each point associated with the 3D geometry). G-PCC addresses the compression of point clouds in category 1 (static point clouds) and category 3 (dynamically acquired point clouds). The latest draft of the G-PCC standard is available in ISO / IEC FDIS23090-9 Geometry-based Point Cloud Compression, ISO / IEC JTC1 / SC29 / WG7 m55637, Teleconference, October 2020, and a description of the G-PCC codec is available in G-PCC CodecDescription, ISO / IEC JTC 1 / SC29 / WG7 MDS20983, Teleconference, October 2021 (hereinafter referred to as "G-PCC codec description").

[0044] A point cloud comprises a collection of points in 3D space and may have attributes associated with the point. The attributes may be color information such as R, G, B or Y, Cb, Cr or reflectivity information or other attributes. Point clouds may be captured by various cameras or sensors (such as LIDAR sensors and 3D scanners) and may also be computer generated. Point cloud data is used in a variety of applications including, but not limited to, construction (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors to aid navigation).

[0045] The 3D space occupied by the point cloud data can be surrounded by a virtual bounding box. The position of the points in the bounding box can be represented by a specific precision; therefore, the position of one or more points can be quantified based on the precision. At the smallest level, the bounding box is divided into voxels, which are the smallest spatial units represented by unit cubes. The voxels in the bounding box can be associated with zero, one, or more than one point. The bounding box can be divided into multiple cube / cuboid areas, which can be referred to as tiles. Each tile can be encoded and decoded into one or more slices. The bounding box is divided into slices and tiles based on the number of points in each partition, or based on other considerations (for example, a specific area can be encoded and decoded as a tile). The slice area can be further segmented using a partition decision similar to the partition decision in the video codec.

[0046] Figure 2 An overview of the G-PCC encoder 200 is provided. Figure 3 An overview of a G-PCC decoder 300 is provided. The modules shown are logical and do not necessarily correspond one-to-one with implementation code in the reference implementation of the G-PCC codec, ie, the TMC13 test model software developed by ISO / IEC MPEG (JTC 1 / SC29 / WG 11).

[0047] In the G-PCC encoder 200 and the G-PCC decoder 300, the point cloud positions are first encoded and decoded. The attribute encoding and decoding depends on the decoded geometry. Figure 2 In the surface approximation analysis unit 212 and the RAHT unit 218 and Figure 3 In , the surface approximation analysis unit 310 and the RAHT unit 314 are options that are typically used for category 1 data. The diagonal cross hatch module is an option that is typically used for category 3 data. All other modules are common between category 1 and category 3. See ISO / IEC FDIS23090-9 Geometry-based Point Cloud Compression, ISO / IEC JTC1 / SC29 / WG7 m55637, Teleconference, October 2020.

[0048] For geometry coding, there are two different types of coding techniques: octree and prediction tree coding. Now let's discuss octree coding. For category 3 data, the compressed geometry is typically represented as an octree from the root down to the leaf level of a single voxel. For category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to the leaf level of blocks larger than a voxel) plus a model that approximates the surface within each leaf of the pruned octree. In this way, category 1 and category 3 data share the octree coding mechanism, while category 1 data can additionally utilize a surface model to approximate the voxels within each leaf. The surface model used is a triangulation that includes 1-10 triangles per block, resulting in a triangle soup. Therefore, the category 1 geometry codec is called a triangle soup geometry codec, while the category 3 geometry codec is called an octree geometry codec.

[0049] At each node of the octree, occupancy is signaled (when not inferred) for one or more of its children (up to eight nodes). Multiple neighborhoods are specified, including (a) nodes that share faces with the current octree node, (b) nodes that share faces, edges, or vertices with the current octree node, etc. Within each neighborhood, the occupancy of the node and / or its children can be used to predict the occupancy of the current node or its children. For points that are sparsely populated in certain nodes of the octree, the codec also supports a direct encoding and decoding mode in which the 3D position of the point is encoded directly. A flag can be signaled to indicate that direct mode is signaled. At the lowest level, the number of points associated with the octree node / leaf node can also be encoded and decoded.

[0050] Figure 4 4 is a conceptual diagram showing an example octree partition for geometric coding. At each node of the octree 400, the G-PCC encoder 200 can send the occupancy of one or more sub-nodes (e.g., up to eight nodes) of the sub-nodes of the node to the G-PCC decoder 300 with a signal (when the G-PCC decoder 300 does not infer the occupancy). Specify multiple neighborhoods, including (a) nodes that share a face with the current octree node, (b) nodes that share a face, edge or vertex with the current octree node, etc. In each neighborhood, the occupancy of the node and / or the sub-nodes of the node can be used to predict the occupancy of the current node or the sub-nodes of the node. For points that are sparsely populated in certain nodes of the octree, the codec also supports a direct coding mode in which the 3D position of the point is directly encoded. The G-PCC encoder 200 can send a flag with a signal to indicate that a direct mode is sent with a signal. At the lowest level, the number of points associated with the octree node / leaf node can also be encoded and decoded.

[0051] Once the geometry is encoded and decoded, the attributes corresponding to the geometry points are encoded and decoded. When there are multiple attribute points corresponding to one reconstructed / decoded geometry point, the attribute values ​​representing the reconstructed points can be derived.

[0052] There are three attribute codec methods in G-PCC: Region Adaptive Hierarchical Transform (RAHT) codec, interpolation-based hierarchical nearest neighbor prediction (prediction transform), and interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform). RAHT and lifting are typically used for category 1 data, while prediction is typically used for category 3 data. However, either method can be used for any data, and just like the geometry codec in G-PCC, the attribute codec method used to encode and decode point clouds is specified in the bitstream.

[0053] The encoding and decoding of attributes can be done in levels of detail (LOD), where with each level of detail, a finer representation of the point cloud attributes can be obtained. Each level of detail can be specified based on a distance metric to neighboring nodes or based on a sampling distance.

[0054] At the G-PCC encoder 200, the residual obtained as an output of the coding method of the attribute is quantized. The residual can be obtained by subtracting the attribute value from a prediction derived based on points in the neighborhood of the current point and based on the attribute value of previously encoded points. The quantized residual can be encoded and decoded using context adaptive arithmetic coding.

[0055] exist Figure 2 In the example, the G-PCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometric reconstruction unit 216, a RAHT unit 218, an LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224 and an arithmetic coding unit 226.

[0056] like Figure 2 As shown in the example of , the G-PCC encoder 200 can obtain a set of positions and a set of attributes of points in the point cloud. The G-PCC encoder 200 can obtain the position set and attribute set of points in the point cloud from the data source 104 ( Figure 1 ) obtains a set of positions and a set of attributes for points in a point cloud. The positions may include the coordinates of the points in the point cloud. The attributes may include information about the points in the point cloud, such as a color associated with the points in the point cloud. The G-PCC encoder 200 may generate a geometry bitstream 203 including representations of the positions of the encoded points in the point cloud. The G-PCC encoder 200 may also generate an attribute bitstream 205 including representations of the encoded attribute sets.

[0057] The coordinate transformation unit 202 may apply a transformation to the coordinates of the point to transform the coordinates from the initial domain to the transformed domain. The present disclosure may refer to the transformed coordinates as transformed coordinates. The color transformation unit 204 may apply a transformation to transform the color information of the attribute to a different domain. For example, the color transformation unit 204 may transform the color information from the RGB color space to the YCbCr color space.

[0058] In addition, Figure 2 In the example of , the voxelization unit 206 can voxelize the transformed coordinates. Voxelization of the transformed coordinates can include quantizing and removing some points of the point cloud. In other words, multiple points of the point cloud can be included in a single "voxel", which can then be treated as a point in some aspects. In addition, the octree analysis unit 210 can generate an octree based on the voxelized transformed coordinates. In addition, in Figure 2 In the example of FIG. 1 , the surface approximation analysis unit 212 may analyze the points to potentially determine a surface representation of a set of points. The arithmetic coding unit 214 may entropy encode syntax elements representing information of the octree and / or surface determined by the surface approximation analysis unit 212. The G-PCC encoder 200 may output these syntax elements in a geometry bitstream 203. The geometry bitstream 203 may also include other syntax elements, including syntax elements that are not arithmetically coded.

[0059] The geometric reconstruction unit 216 may reconstruct the transformed coordinates of the points in the point cloud based on the octree, the data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transformed coordinates reconstructed by the geometric reconstruction unit 216 may be different from the original number of points of the point cloud. The present disclosure may refer to the resulting points as reconstructed points. The attribute transfer unit 208 may transfer the attributes of the original points of the point cloud to the reconstructed points of the point cloud.

[0060] In addition, the RAHT unit 218 can apply RAHT coding to the attributes of the reconstructed points. In some examples, under RAHT, the attributes of the block of 2×2×2 point positions are taken and transformed in one direction to obtain four low (L) and four high (H) frequency nodes. Subsequently, the four low-frequency nodes (L) are transformed in the second direction to obtain two low (LL) and two high (LH) frequency nodes. The two low-frequency nodes (LL) are transformed along the third direction to obtain one low-frequency (LLL) and one high-frequency (LLH) node. The low-frequency node LLL corresponds to the DC coefficient, and the high-frequency nodes H, LH and LLH correspond to the AC coefficient. The transformation in each direction can be a 1-D transformation with two coefficient weights. The low-frequency coefficients can be regarded as coefficients of the 2×2×2 block for the next higher-level RAHT transform, and the AC coefficients are encoded without changes; this transformation continues to the top root node. The tree traversal for encoding is used from top to bottom to calculate the weights to be used for the coefficients; the transformation order is from bottom to top. The coefficients can then be quantized and encoded.

[0061] Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 may apply LOD processing and lifting to the attributes of the reconstructed points, respectively. LOD generation is used to divide the attributes into different refinement levels. Each refinement level provides a refinement of the attributes of the point cloud. The first refinement level provides a rough approximation and contains few points; the subsequent refinement levels typically contain more points, and so on. The refinement level can be constructed using a distance-based metric, or one or more other classification criteria (e.g., subsampling from a specific order) can also be used. Therefore, all reconstructed points can be included in the refinement level. Each detail level is generated by lifting the union of all points to a specific refinement level: for example, LOD1 is obtained based on the refinement level RL1, LOD2 is obtained based on RL1 and RL2, ... LODN is obtained by the union of RL1, RL2, ... RLN. In some cases, LOD generation may be followed by a prediction scheme (e.g., a predictive transform), in which the attributes associated with each point in the LOD are predicted based on a weighted average of previous points, and the residual is quantized and entropy encoded. The lifting scheme is built on top of the predictive transform mechanism, where an update operator is used to update the coefficients and perform adaptive quantization of the coefficients.

[0062] The RAHT unit 218 and the lifting unit 222 may generate coefficients based on the attributes. The coefficient quantization unit 224 may quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 may apply arithmetic coding to syntax elements representing the quantized coefficients. The G-PCC encoder 200 may output these syntax elements in the attribute bitstream 205. The attribute bitstream 205 may also include other syntax elements, including non-arithmetically coded syntax elements.

[0063] exist Figure 3 In the example, the G-PCC decoder 300 may include a geometric arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometric reconstruction unit 312, a RAHT unit 314, a LoD generation unit 316, an inverse lifting unit 318, an inverse transform coordinate unit 320 and an inverse transform color unit 322.

[0064] The G-PCC decoder 300 may obtain the geometry bitstream 203 and the attribute bitstream 205. The geometry arithmetic decoding unit 302 of the G-PCC decoder 300 may apply arithmetic decoding (e.g., context adaptive binary arithmetic coding and decoding (CABAC) or other types of arithmetic decoding) to the syntax elements in the geometry bitstream 203. Similarly, the attribute arithmetic decoding unit 304 may apply arithmetic decoding to the syntax elements in the attribute bitstream 205.

[0065] The octree synthesis unit 306 can synthesize an octree based on the syntax elements parsed from the geometry bitstream 203. Starting from the root node of the octree, the occupancy of each of the eight child nodes at each octree level is signaled in the bitstream. When the signaling indicates that the child node at a specific octree level is occupied, the occupancy of the child nodes of the child node is signaled. Before proceeding to the subsequent octree level, the signaling of the nodes at each octree level is signaled. At the final level of the octree, each node corresponds to a voxel position; when a leaf node is occupied, one or more points can be specified to be occupied at the voxel position. In some cases, due to quantization, some branches of the octree may terminate earlier than the final level. In this case, the leaf node is considered to be an occupied node without a child node. In the case of using surface approximation in the geometry bitstream 203, the surface approximation synthesis unit 310 can determine the surface model based on the syntax elements parsed from the geometry bitstream 203 and based on the octree.

[0066] In addition, the geometric reconstruction unit 312 can perform reconstruction to determine the coordinates of the points in the point cloud. For each position at a leaf node of the octree, the geometric reconstruction unit 312 can reconstruct the node position by using the binary representation of the leaf node in the octree. At each corresponding leaf node, the number of points at the corresponding leaf node is signaled; this indicates the number of repeated points at the same voxel position. When geometric quantization is used, the point position is scaled to determine the reconstructed point position value.

[0067] The inverse transform coordinate unit 320 can apply an inverse transform to the reconstructed coordinates to convert the reconstructed coordinates (positions) of the points in the point cloud from the transform domain back to the original domain. The positions of the points in the point cloud can be in the floating point domain, but the point positions in the G-PCC codec are encoded and decoded in the integer domain. The inverse transform can be used to convert the positions back to the original domain.

[0068] In addition, Figure 3 In the example of , the inverse quantization unit 308 may inverse quantize the property value. The property value may be based on syntax elements obtained from the property bitstream 205 (eg, including syntax elements decoded by the property arithmetic decoding unit 304).

[0069] Depending on how the attribute value is encoded, the RAHT unit 314 may perform RAHT encoding and decoding to determine the color value of the point of the point cloud based on the inverse quantized attribute value. RAHT decoding is performed from the top to the bottom of the tree. At each level, the low-frequency coefficients and high-frequency coefficients derived from the inverse quantization process are used to derive the component values. At the leaf node, the derived value corresponds to the attribute value of the coefficient. The weight derivation process for the point is similar to the process used at the G-PCC encoder 200. Alternatively, the LOD generation unit 316 and the inverse lifting unit 318 may use a detail level-based technique to determine the color value of the point of the point cloud. The LOD generation unit 316 decodes each LOD, thereby giving a gradually finer representation of the attribute of the point. Using the prediction transform, the LOD generation unit 316 derives the prediction of the point from the weighted sum of the points reconstructed in the previous LOD or in the same LOD. The LOD generation unit 316 can add the prediction to the residual (which is obtained after inverse quantization) to obtain the reconstructed value of the attribute. When using the lifting scheme, the LOD generation unit 316 may also include an update operator to update the coefficients used to derive the attribute value. In this case, the LOD generation unit 316 may also apply inverse adaptive quantization.

[0070] In addition, Figure 3 In the example of , the inverse color transform unit 322 can apply an inverse color transform to the color value. The inverse color transform can be the inverse of the color transform applied by the color transform unit 204 of the G-PCC encoder 200. For example, the color transform unit 204 can transform the color information from the RGB color space to the YCbCr color space. Therefore, the inverse color transform unit 322 can transform the color information from the YCbCr color space to the RGB color space.

[0071] Figure 2 and Figure 3Various units are shown to help understand the operations performed by the G-PCC encoder 200 and the G-PCC decoder 300. The units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functionality and are preset on executable operations. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions in executable operations. For example, a programmable circuit may execute software or firmware that enables the programmable circuit to operate in a manner defined by instructions of software or firmware. Fixed-function circuits may execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by fixed-function circuits are generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0072] Predictive geometry codecs (see, e.g., G-PCC Codec Description) were introduced as an alternative to octree geometry codecs, where the nodes are arranged in a tree structure (which defines the prediction structure), and various prediction strategies are used to predict the coordinates of each node in the tree relative to its predictor. Figure 5 An example of a prediction tree is shown, with a directed graph where the arrows point in the direction of the prediction. The horizontal hash node is the root vertex and has no predictors; the cross-hatched nodes have two children; the diagonal hash nodes have 3 children; the non-hashed nodes have one child, and the vertical hash nodes are leaf nodes, and these nodes have no children. Except for the root node, each node has only one parent node.

[0073] Figure 5 5 is a conceptual diagram showing an example of a prediction tree. Node 500 is the root vertex and has no predictor. Nodes 502 and 504 have two child nodes. Node 506 has three child nodes. Nodes 508, 510, 512, 514, and 516 are leaf nodes, and these nodes have no child nodes. The remaining nodes each have one child node. Each node except the root node 500 has only one parent node.

[0074] Four prediction strategies are specified for each node based on its parent node (p0), grand-parent node (p1), and great-grand-parent node (p2):

[0075] No forecast / Zero forecast(0)

[0076] Delta prediction (p0)

[0077] Linear prediction (2*p0–p1)

[0078] Parallelogram prediction (2*p0+p1–p2)

[0079] The G-PCC encoder 200 may employ any algorithm to generate the prediction tree; the algorithm used may be determined based on the application / use case, and several strategies may be used. Some strategies are described in G-PCC Codec Description.

[0080] For each node, the residual coordinate values ​​are encoded and decoded in the bitstream starting from the root node in a depth-first manner. For example, the G-PCC encoder 200 may encode and decode the residual coordinate values ​​in the bitstream.

[0081] Predictive geometry codec is mainly used for category 3 (LIDAR acquired) point cloud data, for example, for low-latency applications.

[0082] Fig. 6A and 6B is a conceptual diagram showing an example of a rotating LIDAR acquisition model. Angular modes for predictive geometry codecs are now described. Angular modes can be used for predictive geometry codecs, where characteristics of the LIDAR sensor can be used to more efficiently encode and decode the prediction tree. The coordinates of the position are converted to the (r, φ, i) (radius, azimuth, and laser index) domain 600, and prediction is performed in this domain 600 (e.g., the residuals are encoded and decoded in the r, φ, i domain). Due to errors in rounding, encoding and decoding in r, φ, i is not lossless, so a second set of residuals corresponding to Cartesian coordinates can be encoded and decoded. A description of the encoding and decoding strategies for angular modes for predictive geometry codecs is reproduced below from G-PCCCodec Description.

[0083] The technique focuses on point clouds acquired using a rotating LIDAR model. Here, the LIDAR 602 has N lasers (e.g., N=16, 32, 64) rotating around the Z axis according to an azimuth angle φ. Each laser can have a different elevation angle θ(i) i=1…N and height Assume that the laser i hits Figure 6A-6B The coordinate system shown in is defined by a point M with Cartesian integer coordinates (x,y,z).

[0084] This technique models the position of M using three parameters (r, φ, i), which are calculated as follows:

[0085]

[0086] φ=atan2(y,x)

[0087]

[0088] More precisely, this technique uses a quantized version of (r,φ,i), expressed as Among the three integers and i are calculated as follows:

[0089]

[0090]

[0091] Among them (q r ,o r ) and (q φ ,o φ ) are control and The precision of the quantization parameter. sign(t) is a function that returns 1 if t is positive, otherwise it returns (-1). |t| is the absolute value of t.

[0092] To avoid reconstruction mismatches due to the use of floating-point operations, andtan(θ(i)) i=1…N The values ​​are precomputed and quantized as follows:

[0093]

[0094] in and (q θ ,o θ ) are respectively controlled and The quantization parameter of the accuracy.

[0095] The reconstructed Cartesian coordinates are obtained as follows:

[0096]

[0097] where app_cos(.) and app_sin(.)a are approximations of cos(.) and sin(.). The calculations can be performed using fixed-point representation, lookup tables, and / or linear interpolation.

[0098] Note that due to various reasons such as quantization, approximation, model inaccuracy, model parameter inaccuracy, etc. May be different from (x,y,z).

[0099] Assume (r x ,r y ,r z ) is the reconstructed residual defined as follows:

[0100]

[0101] Using this technique, the G-PCC encoder 200 can proceed as follows:

[0102] 1) Model parameters and And the quantization parameter q r , q θ and q φ Encoding

[0103] 2) Apply the geometry prediction scheme described in the text of ISO / IEC FDIS23090-9 Geometry-based Point Cloud Compression, ISO / IEC JTC 1 / SC29 / WG 7m55637, Teleconference, Oct.2020.

[0104] 3) Indicate

[0105] New predictors that take advantage of LIDAR properties can be introduced. For example, the rotation speed of a LIDAR scanner around the z-axis is usually constant. Therefore, we can predict the current

[0106]

[0107] in

[0108] (δ φ (k)) k=1…K is a set of potential speeds that can be used by the G-PCC encoder 200. The index k can be written explicitly into the bitstream, or can be based on the speeds specified by the G-PCC encoder 200 and the G-

[0109] The deterministic strategy applied by both the G-PCC encoder 200 and the G-PCC decoder 300 is inferred from the context, and n(j) is the number of skipped points that can be explicitly written into the bitstream, or can be inferred from the context based on the deterministic strategy applied by both the G-PCC encoder 200 and the G-PCC decoder 300. n(j) is also referred to herein as the "phi multiplier". Note that the phi multiplier is currently only used with the delta predictor.

[0110] 4) Reconstruct the residual (r) using each node pair x ,r y ,r z ) to encode

[0111] The G-PCC decoder 300 may proceed as follows:

[0112] 1) Decoding model parameters and And the quantization parameter q r , q θ and q φ

[0113] 2) The parameters associated with the nodes according to the geometry prediction scheme described in the text of ISO / IEC FDIS23090-9 Geometry-based Point Cloud Compression, ISO / IEC JTC 1 / SC29 / WG 7m55637, Teleconference, Oct.2020 to decode.

[0114] 3) Calculate the reconstructed coordinates as described above

[0115] 4) Decoding residual (r x ,r y ,r z )

[0116] As discussed in the next section, the reconstructed residual (r x ,r y ,r z ) to support lossy compression.

[0117] 5) Calculate the original coordinates (x, y, z) as follows

[0118]

[0119] Lossy compression can be achieved by reconstructing the residual (r x ,r y ,r z ) is achieved by applying quantization or by discarding points.

[0120] The quantized reconstruction residual can be calculated as follows:

[0121]

[0122]

[0123] Among them, (q x ,o x ), (q y ,o y ) and (q z ,o z ) are respectively controlled and For example, the G-PCC encoder 200 or the G-PCC decoder 300 may calculate the quantized residual.

[0124] The G-PCC encoder 200 or the G-PCC decoder 300 may use trellis quantization to further improve RD (rate-distortion) performance results.

[0125] The quantization parameter can be changed at sequence / frame / slice / block level to achieve region-adaptive quality and / or for rate control purposes.

[0126] Different types of sensors that may be used with system 100 are now discussed. Autonomous driving solutions may utilize accurate mapping of the environment to aid navigation. Several LIDAR systems are common, and different technologies are used for LIDAR sensors. An overview of examples of such systems is provided in S. Royo and M. Ballesta-Garcia, An overview of LIDAR imaging systems for autonomous vehicles, Journal of Applied Sciences, 2019, 9(19), 4093; https: / / doi.org / 10.3390 / app9194093. Among these example systems, rotating LIDAR sensors and solid-state sensors are discussed in this document.

[0127] Rotating LIDAR systems are common systems where one or more laser sources and detectors are mounted on a rotating structure. Each laser source is typically directed at a different elevation angle, and the laser scans the surrounding object / area. Several rotations may occur per second, and the detector captures the reflected light. Rotating LIDAR can capture a 360-degree field of view (FOV). However, these sensors are typically large and bulky due to the mechanical systems involved in the rotation. Examples of rotating LIDAR sensors include Velodyne Sensors VLP-16, Alpha Prime, etc.

[0128] Solid-state LIDAR systems can also be used to capture the surrounding environment. These systems work in a different way than rotating LIDAR systems. Solid-state sensors usually have a lower spatial volume because they do not contain large mechanical systems (as in rotating sensors). Due to their smaller size and limited field of view, multiple solid-state LIDARs are often used in several applications. Each sensor can be used to capture one area of ​​the FOV.

[0129] A LIDAR system may include one or more laser sources and detectors / sensors, and each sensor may capture points (e.g., points of a point cloud) in a predetermined order. This order (which may be referred to as a capture order or a scan order) may be proprietary and / or may have some dependency on the mechanics of the laser / sensor system. For example, for a rotating LIDAR system, each sensor may capture points in a rotating order.

[0130] When encoding and decoding points using the G-PCC encoder 200 or the G-PCC decoder 300 or other point cloud compression codecs, points belonging to different lasers and / or sensors are typically encoded and decoded together. When encoding and decoding a point cloud using predicted geometry, a prediction tree may be generated. Suboptimal construction of the prediction tree may result in inefficient encoding and decoding. Therefore, it may be desirable to generate a prediction tree so that the predicted geometry codec can exploit the inherent dependencies between the various points of the sensor.

[0131] Not all LIDAR systems rotate a full 360 degrees. For example, some LIDAR systems may sweep over a more limited range (e.g., 120 degrees) and then return to the starting angle to perform the next sweep. As a result, consecutive points in a scan or capture sequence, for example, may sometimes be in very different locations, and one point may not provide an accurate prediction of the next consecutive point.

[0132] The techniques of this disclosure may be implemented independently or together in any combination.

[0133] The first prediction tree branch may be constructed by connecting points captured successively by the sensor. The first point in a particular prediction tree branch may in turn be referred to as a master node or a root node of the particular prediction tree branch.

[0134] In one example, the prediction tree branches may be constructed from consecutive points presented to the G-PCC encoder 200, for example, in a codec order, which may be the same as or different from the capture order. For example, the G-PCC encoder 200 or the G-PCC decoder 300 may construct the prediction tree branches from consecutive points in a codec order and / or a capture order.

[0135] The first azimuth threshold may be specified to indicate the maximum azimuth difference between adjacent points in the first prediction tree branch. When the azimuth difference between two consecutive points in the first prediction tree branch (e.g., point A and then point B, where point B is after point A) exceeds the first threshold, the first prediction tree branch may terminate at point A. For example, when the azimuth difference between the two consecutive points exceeds the first threshold, the G-PCC encoder 200 or the G-PCC decoder 300 may terminate the first prediction tree branch at the first of the two consecutive points, so that the first of the two consecutive points becomes a leaf node of the first prediction tree branch, and starts the second prediction tree branch with the second of the consecutive points.

[0136] Although various thresholds and comparisons to thresholds are discussed herein, it should be understood that the determination of whether a difference is greater than (e.g., exceeds) a threshold may be replaced by a determination of whether the difference is greater than or equal to the threshold, and vice versa. In addition, the determination of whether a difference is less than a threshold may be replaced by a determination of whether the difference is less than or equal to the threshold, and vice versa.

[0137] When the azimuth difference between two consecutive points (e.g., point A and then point B, where point B is after point A) exceeds (or is greater than or equal to) a first azimuth threshold in sequence, a new prediction tree branch (e.g., a second prediction tree branch) may be started at point B; point B may be the main node of the new prediction branch. For example, when the azimuth difference between two consecutive points exceeds the first azimuth threshold, the G-PCC encoder 200 or the G-PCC decoder 300 may start the second prediction tree branch at point B. In this way, when the azimuth difference between two consecutive points is relatively large, the second of the two consecutive points may be used to start a new prediction tree branch instead of following the first of the two consecutive points in the existing prediction tree branch. For example, the second of the two consecutive points may be closer to a previous point other than the first of the two consecutive points, and it may be more appropriate to use a point other than the first of the two consecutive points to predict the second of the two consecutive points.

[0138] The first azimuth threshold may be signaled in a bitstream (in a sequence parameter set (SPS), a geometry parameter set (GPS), or a slice), or may be predetermined for the G-PCC encoder 200 and the G-PCC decoder 300. In one example, the first threshold may be limited to a non-negative number. In another example, the first threshold may be a negative number. In another alternative, a positive threshold and a negative threshold may be specified (e.g., signaled). For example, T1 and T2 may be a positive threshold and a negative threshold, respectively. In this example, the condition for starting a new prediction tree branch may be specified as follows: If the azimuth difference (e.g., azimuth difference) between consecutive points is greater than T2 and less than T1, then the second point in the consecutive points is added to the current prediction tree branch. If the azimuth difference is neither greater than T2 nor less than T1, then the second point in the consecutive points is added to the new prediction tree branch. For example, the G-PCC encoder 200 or the G-PCC decoder 300 may apply such a condition.

[0139] The G-PCC encoder 200 or the G-PCC decoder 300 can connect the first prediction tree branch and the second prediction tree branch by adding the master node of the first prediction tree branch as a child node of one of the nodes in the second prediction tree branch. In one example, the master node of the first prediction tree branch is added as a child node of the master node of the second prediction tree branch. Similarly, the G-PCC encoder 200 or the G-PCC decoder 300 can connect the first prediction tree branch and the second prediction tree branch by adding the master node of the second prediction tree branch as a child node of a node in the first prediction tree branch. In one example, the master node of the second prediction tree branch is added as a child node of the master node of the first prediction tree branch.

[0140] In another example, the master node M1 of the first prediction tree branch is added as a child node of the first node in the second prediction tree branch. The first node can be selected as the node with the shortest distance to M1 among the nodes of the second prediction tree branch. Similarly, the master node M2 ​​of the second prediction tree branch can be added as a child node of the first node in the first prediction tree branch. The first node can be selected as the node with the shortest distance to M2 among the nodes of the first prediction tree branch.

[0141] In another example, multiple points N1 may be selected. The master node M1 of the first prediction tree branch may be added as a child node of the first node in the second prediction tree branch. For example, the first node may be selected among the first N1 nodes of the second prediction tree branch so that the first node has the shortest distance to M1. Similarly, the master node M2 ​​of the second prediction tree branch may be added as a child node of the first node in the first prediction tree branch. For example, the first node may be selected among the first N1 nodes of the first prediction tree branch so that the first node has the shortest distance to M2.

[0142] When two prediction tree branches are joined using one of the techniques disclosed above, the resulting tree structure may be referred to as a subtree of the first prediction tree branch and the second prediction tree branch. The master node of the second prediction tree branch may be considered as the master node of the subtree.

[0143] Although the above discussion is for combining branches into subtrees, these techniques can also be applied when combining two subtrees, or when branches are combined with subtrees. In these cases, the resulting tree structure can still be called a subtree.

[0144] The G-PCC encoder 200 can specify a scan line identifier (ID) for each point of the point cloud data. The scan line ID can specify a row in a captured raster scan order (or zigzag order). Each point coordinate can be specified with a radius, an azimuth, and a scan line ID and a residual in a Cartesian coordinate. The point coordinates can be encoded and decoded in a Cartesian coordinate domain, such as a radius, an azimuth, and a scan line ID, with optional residuals in x, y, and z coordinates. One or more characteristics associated with the scan line ID can be signaled by the G-PCC encoder 200 in the bitstream, for example, an elevation angle associated with the scan line or a z elevation angle associated with the scan line.

[0145] An example implementation is now described. In this example, an implementation in which a raster scan jump results in a relatively large negative value in the difference between the azimuth value of the next point and the azimuth value of the current point (azimuth (next point) - azimuth (current point)). The points described below include points that belong to a sensor or are captured by a sensor. Start with a first point P(1); point P(1) can be the root node of a prediction tree. A prediction tree branch for a point can start at P(1). The G-PCC encoder 200 or the G-PCC decoder 300 can keep traversing the points (e.g., in the order of capture) and add the points to the prediction tree as linear prediction tree branches. For example, this can continue as long as azimuth (next point) - azimuth (current point) > threshold. The threshold (which can be a first azimuth threshold) can be small to cover any noise. In some examples, if known, the threshold can be derived based on the sampling distance of the LiDAR system. At the end of the traversal of the points, there can be a long chain of points with P(1) as the root.

[0146] Figure 7 is a conceptual diagram of an example prediction tree according to one or more aspects of the present disclosure. Prediction tree 700 may be a prediction tree generated by G-PCC encoder 200 or G-PCC decoder 300. For example, G-PCC encoder 200 or G-PCC decoder 300 may use point 702 as the root node of the first tree branch from point 702 to point 704. For each pair of consecutive points between points 702 and 704, the value of azimuth (next point) - azimuth (current point) > threshold. For example, the value of azimuth (point 703) - azimuth (point 702) > threshold. This chain of points from point 702 to point 704 may constitute the long chain of points described above. In this case, point 702 corresponds to point P(1).

[0147] When azimuth(next point)-azimuth(current point)≤threshold (e.g., if there is a large jump such as in a raster scan jump, the difference may be a large negative value and therefore less than or equal to the threshold), and in the case where the next point is P(2), the G-PCC encoder 200 or the G-PCC decoder 300 may start a new prediction tree branch from P(2) and add P(2) as a child node to P(1). The G-PCC encoder 200 or the G-PCC decoder 300 may then continue traversing the points as described above to fill in the new prediction tree branch starting from P(2). For example, point 706 may be the root point of a second tree branch from point 706 to point 708. For example, azimuth(point 706)-azimuth(point 704)≤threshold, such that a new branch of the prediction tree 700 should be started.

[0148] At the next jump (e.g., the next azimuth (next point) - azimuth (current point) ≤ threshold), point P(n+1) is added as a child node to P(n) and the linear prediction branch starting at P(n+1) is continued to be filled. For example, the G-PCC encoder 200 or the G-PCC decoder 300 can add point 710 as the root node of the third prediction branch because azimuth (point 710) - azimuth (point 708) ≤ threshold.

[0149] It should be noted that in examples where the raster scan jump results in a large positive value for azimuth (next point) - azimuth (current point) rather than a large negative value, the signs used above may be flipped. In other words, when azimuth (next point) - azimuth (current point) < threshold, the G-PCC encoder 200 or G-PCC decoder 300 may add the next point as another point in the current prediction tree branch, and when azimuth (next point) - azimuth (current point) ≥ threshold, the G-PCC encoder 200 or G-PCC decoder 300 may add the next point as the root point of a new prediction tree branch.

[0150] The G-PCC encoder 200 or the G-PCC decoder 300 may repeat these techniques until all points of the sensor are included in the prediction tree. The resulting prediction tree 700 may be similar to the default tree of G-PCC.

[0151] In some examples, the prediction tree generated for each sensor can be combined with one or more prediction trees generated by other sensors (e.g., by adding the root node of one sensor prediction tree as a child node to one of the nodes of another sensor prediction tree), or can be encoded and decoded separately (within the same slice or in different slices).

[0152] In another example implementation, the G-PCC encoder 200 or the G-PCC decoder 300 may start from a first point P(1) and start a new prediction tree branch for the point at P(1). In this example implementation, the raster scan jump results in a relatively large negative value for the difference between the azimuth value of the next point and the azimuth value of the current point (azimuth(next point)-azimuth(current point)). The G-PCC encoder 200 or the G-PCC decoder 300 may keep traversing the points (e.g., in the order in which they were captured) and add the points to the prediction tree as linear prediction tree branches. This may continue as long as azimuth(next point)-azimuth(current point)>threshold. The threshold may be small to cover any noise. In some examples, if known, the threshold may be derived based on the sampling distance of the LiDAR system. At the end of the traversal of the points, there may be a long chain of points with P(1) as the root.

[0153] Figure 8is a conceptual diagram of another example prediction tree according to one or more aspects of the present disclosure. Prediction tree 800 is similar to prediction tree 700, except that the root node of the consecutive branch does not point to the root node of the previous branch, but points to another node of the previous branch. For example, the G-PCC encoder 200 or the G-PCC decoder 300 can use point 802 as the root node of the first tree branch from point 802 to point 804. For each pair of consecutive points between points 802 and 804, the value of azimuth (next point)-azimuth (current point)>threshold. This point chain from point 802 to point 804 can constitute the above-mentioned long point chain. In this case, point 802 corresponds to point P(1).

[0154] When azimuth(next point)-azimuth(current point)≤threshold (e.g., if there is a large jump such as in a raster jump, the difference may be a large negative value and therefore less than or equal to the threshold), and in the case where the next point is P(2), the G-PCC encoder 200 or the G-PCC decoder 300 may start a new prediction tree branch from P(2) and add P(2) as a child node to the node in the prediction tree branch starting at P(1) that is closest to P(2). The G-PCC encoder 200 or the G-PCC decoder 300 may then continue traversing the points as described above to fill in the new prediction tree branch starting from P(2). For example, point 806 may be the root point of a second tree branch from point 806 to point 808. For example, azimuth(point 806)-azimuth(point 804)≤threshold, such that a new branch of the prediction tree 800 should be started. In this example, point 806 may be closest to point 803. Therefore, the G-PCC encoder 200 or the G-PCC decoder 300 may start the second prediction tree branch from the node associated with the point 803 with the point 806 as the root node.

[0155] At the next jump (e.g., the next azimuth (next point) - azimuth (current point) ≤ threshold), the G-PCC encoder 200 or the G-PCC decoder 300 may add point P(n+1) as a child node to the node in the prediction tree branch starting at P(n) that is closest to P(n+1), and continue to fill the linear prediction tree branch starting at P(n+1). For example, the G-PCC encoder 200 or the G-PCC decoder 300 may add point 810 as the root node of the third prediction branch because azimuth (point 810) - azimuth (point 808) ≤ threshold.

[0156] It should be noted that in implementations where the raster scan jump results in large positive values ​​of azimuth (next point) - azimuth (current point) rather than large negative values, the signs used above may be flipped. In other words, when azimuth (next point) - azimuth (current point) < threshold, the G-PCC encoder 200 or G-PCC decoder 300 may add the next point as another point in the current prediction tree branch, and when azimuth (next point) - azimuth (current point) ≥ threshold, the G-PCC encoder 200 or G-PCC decoder 300 may add the next point as the root point of a new prediction tree branch.

[0157] The G-PCC encoder 200 or the G-PCC decoder 300 may repeat these techniques until all points of the sensor are included in the prediction tree. The resulting prediction tree 800 may be similar to the default tree of G-PCC.

[0158] Fig. 9 is a flow chart illustrating an example technique for generating a prediction tree according to one or more aspects of the present disclosure. The G-PCC encoder 200 or the G-PCC decoder 300 may utilize (e.g., use) a first point of the point cloud data to start a first prediction tree branch (900). For example, the first point is used to start the first prediction tree branch and may be the root node of the first prediction tree branch. The G-PCC encoder 200 or the G-PCC decoder 300 may parse (e.g., decode) a second point in the encoding / decoding / capture order (902). For example, the second point is encoded and decoded, wherein the third point (in Fig. 9) is a point that is encoded / captured before the second point (e.g., the third point is encoded / captured before the second point). The third point may also be part of a first prediction tree branch. The G-PCC encoder 200 or the G-PCC decoder 300 may determine whether the azimuth of the second point minus the azimuth of the third point (e.g., the azimuth difference) is greater than a first threshold (e.g., a first azimuth threshold) (904). If the azimuth of the second point minus the azimuth of the third point is greater than the first threshold (the "yes" branch from box 904), the G-PCC encoder 200 or the G-PCC decoder 300 may add the second point to the first prediction tree branch (906). For example, if the azimuth of the second point minus the azimuth of the third point is greater than the first threshold, then the second point may be added to the first prediction tree branch. Otherwise, the second prediction tree branch is started with the second point as the main / root node. For example, if the azimuth angle of the second point minus the azimuth angle of the third point is not greater than the first threshold (the "No" branch from block 904), the G-PCC encoder 200 or the G-PCC decoder 300 starts a second prediction tree branch with the second point as a main node or root node (908). The G-PCC encoder 200 or the G-PCC decoder 300 may add the second point as a child node to one of the nodes in the first prediction tree branch (910).

[0159] Another example implementation is now described. This example may be similar to the above Figure 8 The example described is different in that the G-PCC encoder 200 or the G-PCC decoder 300 can use the absolute value of the azimuth difference instead of the actual azimuth difference. In this case, the way in which the G-PCC encoder 200 or the G-PCC decoder 300 uses the threshold value can be changed as described below.

[0160] The G-PCC encoder 200 or the G-PCC decoder 300 can start a new prediction tree branch for the point at P(1). The G-PCC encoder 200 or the G-PCC decoder 300 can keep traversing the points (e.g., in the order of capture) and add the points to the prediction tree as linear prediction tree branches. This can continue as long as the absolute value of the azimuth (next point)-azimuth (current point) difference is less than (or less than or equal to) a threshold. The threshold can be large enough to cover any noise. In some examples, the threshold can be derived based on the sampling distance of the LiDAR system (if known) or the azimuth jump in the case where the LiDAR system uses raster scanning. At the end of the traversal of the points, the prediction tree branch may include a long chain of points with P(1) as the root.

[0161] When ABS(azimuth(next point)-azimuth(current point))>threshold (for example, if there is a large jump such as in a raster jump), and in the case where the next point is P(2), the G-PCC encoder 200 or the G-PCC decoder 300 may start a new prediction tree branch from P(2) and add P(2) as a child node to the node in the prediction tree branch starting at P(1) that is closest to P(2). The G-PCC encoder 200 or the G-PCC decoder 300 may then continue traversing the points as described above to fill in the new prediction tree branch starting at P(2).

[0162] At the next jump (e.g., the next time ABS(azimuth(next point)-azimuth(current point))>threshold), point P(n+1) is added as a child node to the node in the prediction tree branch starting at P(n) that is closest to P(n+1), and the linear prediction tree branch starting at P(n+1) is continued to be filled.

[0163] The G-PCC encoder 200 or the G-PCC decoder 300 may repeat these techniques until all points of the point cloud sensed by the sensor are included in the prediction tree. The resulting tree may be similar to the default tree of G-PCC.

[0164] Another example implementation is now described. This example may be similar to the above Figure 7 The example described is different in that the G-PCC encoder 200 or the G-PCC decoder 300 can use the absolute value of the azimuth difference instead of the actual azimuth difference. In this case, the way in which the G-PCC encoder 200 or the G-PCC decoder 300 uses the threshold value can be changed as described below.

[0165] The G-PCC encoder 200 or the G-PCC decoder 300 can start a new prediction tree branch for the point at P(1). The G-PCC encoder 200 or the G-PCC decoder 300 can keep traversing the points (e.g., in the order in which they were captured) and add the points to the tree as linear prediction tree branches. This can continue as long as the absolute value of the azimuth (next point)-azimuth (current point) difference is less than (or less than or equal to) a threshold. The threshold can be large enough to cover any noise. In some examples, the threshold can be derived based on the sampling distance of the LiDAR system (if known) or the azimuth jump in the case where the LiDAR system uses raster scanning. At the end of the traversal of the points, the prediction tree branch may include a long chain of points with P(1) as the root.

[0166] When ABS(azimuth(next point)-azimuth(current point))>threshold (for example, if there is a large jump such as in a raster jump), and in the case where the next point is P(2), the G-PCC encoder 200 or the G-PCC decoder 300 may start a new prediction tree branch from P(2) and add P(2) as a child node to P(1). The G-PCC encoder 200 or the G-PCC decoder 300 may then continue traversing the points as described above to fill in the new prediction tree branch starting at P(2).

[0167] At the next jump (e.g., the next ABS(azimuth(next point)-azimuth(current point))>threshold), the G-PCC encoder 200 or the G-PCC decoder 300 may add point P(n+1) as a child node to P(n) and continue to fill the linear prediction tree branch starting at P(n+1).

[0168] The G-PCC encoder 200 or the G-PCC decoder 300 may repeat these techniques until all points of the sensor are included in the tree.

[0169] Fig.10 is a flow chart illustrating an example prediction tree generation technique according to one or more aspects of the present disclosure. The G-PCC encoder 200 or the G-PCC decoder 300 may determine a first point of the point cloud data as a first node of a first prediction tree branch of a prediction tree (1000).

[0170] The G-PCC encoder 200 or the G-PCC decoder 300 may determine that a first azimuth difference between a first point of the point cloud data and a point of the point cloud data immediately preceding in sequence does not satisfy a first azimuth threshold value (1002). For example, the G-PCC encoder 200 or the G-PCC decoder 300 may determine the first azimuth difference as the difference between the azimuth values ​​of the first point (e.g., point 702) and the second point (e.g., point 703). The G-PCC encoder 200 or the G-PCC decoder 300 may compare the first azimuth difference with the first azimuth threshold value to determine that the first azimuth difference does not satisfy the first azimuth threshold value.

[0171] Based on the first azimuth difference not satisfying the first azimuth threshold, the G-PCC encoder 200 or the G-PCC decoder 300 may determine the second point as a second node of the first prediction tree branch (1004). For example, the G-PCC encoder 200 or the G-PCC decoder 300 may add the second point to the first prediction tree branch, for example, from the first point (e.g., immediately adjacent to the first point), so as to further construct the first prediction tree branch.

[0172] The G-PCC encoder 200 or the G-PCC decoder 300 may determine that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies a first azimuth threshold value (1006). The third point (e.g., point 704) and the fourth point (e.g., point 706) include sequentially consecutive points, and the third point includes a third node of a first prediction tree branch. For example, the G-PCC encoder 200 or the G-PCC decoder 300 may determine the second azimuth difference as the difference between the azimuth values ​​of the third point and the fourth point, the third point being immediately before the fourth point in sequence. The G-PCC encoder 200 or the G-PCC decoder 300 may compare the second azimuth difference with the first azimuth threshold value to determine that the second azimuth difference satisfies the first azimuth threshold value.

[0173] Based on the second azimuth difference satisfying the first azimuth threshold, the G-PCC encoder 200 or the G-PCC decoder 300 may terminate the first prediction tree branch at the third point, so that the third point includes a leaf node of the first prediction tree branch, and determine the fourth point as the first node of the second prediction tree branch (1108). For example, the G-PCC encoder 200 or the G-PCC decoder 300 may start the second prediction tree branch with the fourth point as the main node or root node of the second prediction tree branch. In other words, the G-PCC encoder 200 or the G-PCC decoder 300 may start the second prediction tree branch with the fourth point.

[0174] The G-PCC encoder 200 or the G-PCC decoder 300 may connect the first prediction tree branch and the second prediction tree branch in the prediction tree (1010). For example, the G-PCC encoder 200 or the G-PCC decoder 300 may connect the fourth point (e.g., the main node or root node of the second prediction tree branch) to a specific node of the first prediction tree branch.

[0175] The G-PCC encoder 200 or the G-PCC decoder 300 may encode or decode the point cloud data based on the prediction tree (1012). For example, the G-PCC encoder 200 may encode the point cloud data based on the prediction tree, or the G-PCC decoder 300 may decode the point cloud data based on the prediction tree.

[0176] In some examples, as part of connecting the first prediction tree branch and the second prediction tree branch, the G-PCC encoder 200 or the G-PCC decoder 300 is configured to add the first node of the second prediction tree branch as a child node of the first node of the first prediction tree branch. In some examples, the first node of the first prediction tree branch is the root node of the prediction tree.

[0177] In some examples, as part of connecting a first prediction tree branch and a second prediction tree branch, the G-PCC encoder 200 or the G-PCC decoder 300 is configured to add a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, wherein the node of the first prediction tree branch has a corresponding point of point cloud data, and the corresponding point of the point cloud data has the shortest distance to a fourth point among all nodes of the first prediction tree branch.

[0178] In some examples, as part of connecting a first prediction tree branch and a second prediction tree branch, the G-PCC encoder 200 or the G-PCC decoder 300 is configured to add a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, wherein the node of the first prediction tree branch has a corresponding point of point cloud data, and the corresponding point of the point cloud data has the shortest distance to a fourth point among a predetermined number of nodes of the first prediction tree branch.

[0179] In some examples, the order includes at least one of a sensor capture order or a codec order. In some examples, the G-PCC encoder 200 or the G-PCC decoder 300 is further configured to signal or parse the first azimuth angle threshold in the bitstream. In some examples, the first azimuth angle threshold includes one of a non-negative number or a negative number.

[0180] In some examples, the first azimuth threshold comprises a non-negative number, wherein the second azimuth threshold comprises a negative number. In some examples, as part of determining that the second azimuth difference satisfies the first azimuth threshold, the G-PCC encoder 200 or the G-PCC decoder 300 is configured to determine that the second azimuth difference a) is less than the first azimuth threshold or b) is less than or equal to the first azimuth threshold, and is further configured to terminate the first prediction tree branch at a third point based on determining that the second azimuth difference c) is greater than or equal to the second azimuth threshold or d) is greater than the second azimuth threshold.

[0181] In some examples, as part of determining that the second azimuth difference satisfies the first azimuth threshold, G-PCC encoder 200 or G-PCC decoder 300 is configured to determine that the second azimuth difference is a) less than or equal to the first azimuth threshold or b) less than the first azimuth threshold.

[0182] In some examples, as part of determining that the second azimuth difference satisfies the first azimuth threshold, the G-PCC encoder 200 or the G-PCC decoder 300 is configured to determine that an absolute value of the second azimuth difference is a) greater than the first azimuth threshold or b) greater than or equal to the first azimuth threshold.

[0183] In some examples, the G-PCC encoder 200 or the G-PCC decoder 300 is further configured to determine a first scan line ID for the third point, and to determine a second scan line ID for the fourth point. In some examples, the G-PCC encoder 200 or the G-PCC decoder 300 is further configured to signal or parse one or more characteristics associated with at least one of the first scan line ID or the second scan line ID in the bitstream. In some examples, the G-PCC encoder 200 or the G-PCC decoder 300 is further configured to generate a point cloud.

[0184] Fig.11 is a conceptual diagram illustrating an example ranging system 1100 that can be used for coordinate transformation in G-PCC with one or more techniques of the present disclosure. Fig.11 In an example of , ranging system 1100 includes an illuminator 1102 and a sensor 1104. Illuminator 1102 can emit light 1106. In some examples, illuminator 1102 can emit light 1106 as one or more laser beams. Light 1106 can be in one or more wavelengths, such as infrared wavelengths or visible light wavelengths. In other examples, light 1106 is not a coherent laser. When light 1106 encounters an object (such as object 1108), light 1106 creates return light 1110. Return light 1110 may include backscattered light and / or reflected light. Return light 1110 may pass through lens 1111, which guides return light 1110 to create an image 1112 of object 1708 on sensor 1104. Sensor 1104 generates signal 1114 based on image 1112. Image 1112 may include a collection of points (e.g., such as Fig.11 1112).

[0185] In some examples, the illuminator 1102 and the sensor 1104 can be mounted on a rotating structure so that the illuminator 1102 and the sensor 1104 capture a 360-degree view of the environment (e.g., a rotating LIDAR sensor). In other examples, the ranging system 1100 can include one or more optical components (e.g., mirrors, collimators, diffraction gratings, etc.) that enable the illuminator 1102 and the sensor 1104 to detect the range of objects within a certain range (e.g., up to 360 degrees). Although Fig.11 The example shows only a single illuminator 1102 and sensor 1104, but the ranging system 1100 may include a collection of multiple illuminators and sensors.

[0186] In some examples, the illuminator 1102 generates a structured light pattern. In such an example, the ranging system 1100 may include a plurality of sensors 1104 on which respective images of the structured light pattern are formed. The ranging system 1100 may use differences between the images of the structured light pattern to determine the distance to an object 1108 from which the structured light pattern is backscattered. When the object 1108 is relatively close to the sensor 1104 (e.g., 0.2 meters to 2 meters), the ranging system based on structured light can have a high level of accuracy (e.g., accuracy in the sub-millimeter range). This high level of accuracy may be useful in facial recognition applications such as unlocking a mobile device (e.g., a mobile phone, a tablet computer, etc.) and for security applications.

[0187] In some examples, the ranging system 1100 is a time-of-flight (ToF) based system. In some examples where the ranging system 1100 is a ToF based system, the illuminator 1102 generates light pulses. In other words, the illuminator 1102 can modulate the amplitude of the emitted light 1106. In such an example, the sensor 1104 detects the return light 1110 from the light pulse 1106 generated by the illuminator 1102. The ranging system 1100 can then determine the distance to the object 1108 from which the light 1106 is backscattered based on the delay between the time when the light 1106 is emitted and detected and the known speed of light in air. In some examples, instead of modulating the amplitude of the emitted light 1106 (or in addition to modulating the amplitude of the emitted light 1106), the illuminator 1102 can modulate the phase of the emitted light 1106. In such an example, sensor 1104 can detect the phase of return light 1110 from object 1108 and determine the distance to a point on object 1108 using the speed of light and based on the time difference between when illuminator 1102 generates light 1106 at a particular phase and when sensor 1104 detects return light 1110 at the particular phase.

[0188] In other examples, the point cloud can be generated without using the illuminator 1102. For example, in some examples, the sensor 1104 of the ranging system 1100 may include two or more optical cameras. In such an example, the ranging system 1100 can use the optical cameras to capture a stereoscopic image of the environment (including the object 1108). The ranging system 1100 may include a point cloud generator 1116, which can calculate the difference between the positions in the stereoscopic images. The ranging system 1100 can then use the difference to determine the distance to the position shown in the stereoscopic image. Based on these distances, the point cloud generator 1116 can generate a point cloud.

[0189] Sensor 1104 may also detect other properties of object 1108, such as color and reflectivity information. Fig.11In the example of , the point cloud generator 1116 can generate a point cloud based on the signal 1114 generated by the sensor 1104. The ranging system 1100 and / or the point cloud generator 1116 can form a data source 104 ( Figure 1 ). Therefore, the point cloud generated by the ranging system 1100 can be encoded and / or decoded according to any of the techniques of the present disclosure. Inter-frame prediction and residual prediction as described in the present disclosure can reduce the size of the encoded data.

[0190] Fig.12 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more of the techniques for coordinate transformation in G-PCC disclosed herein may be used. Fig.12 In the example of FIG. 1 , the vehicle 1200 includes a ranging system 1202. The ranging system 1202 can be used to measure the distance between the vehicle and the vehicle. Fig.11 Although Fig.12 , but the vehicle 1200 may also include a data source (such as data source 104 ( Figure 1 )) and a G-PCC encoder (such as G-PCC encoder 200 ( Figure 1 )).exist Fig.12 In the example of FIG. 1 , the ranging system 1202 emits a laser beam 1204, which is reflected from a pedestrian 1206 or other object in the road. The data source of the vehicle 1200 can generate a point cloud based on the signal generated by the ranging system 1202. The G-PCC encoder of the vehicle 1200 can encode the point cloud to generate a bitstream 1208, such as a geometry bitstream ( Figure 2 ) and attribute bitstream ( Figure 2 ). Inter-frame prediction and residual prediction as described in the present disclosure can reduce the size of the geometry bitstream. The bitstream 1208 can include fewer bits than the unencoded point cloud obtained by the G-PCC encoder.

[0191] The output interface of the carrier 1200 (eg, the output interface 108 ( Figure 1 )) can send a bitstream 1208 to one or more other devices. The bitstream 1208 may include fewer bits than the unencoded point cloud obtained by the G-PCC encoder. Therefore, the vehicle 1200 can send the bitstream 1208 to other devices faster than the unencoded point cloud data. In addition, the bitstream 1208 may require less data storage capacity on the device.

[0192] exist Fig.12 In the example of FIG. 1 , the carrier 1200 may send the bitstream 1208 to another carrier 1210. The carrier 1210 may include a G-PCC decoder, such as the G-PCC decoder 300 ( Figure 1). The G-PCC decoder of vehicle 1210 can decode bitstream 1208 to reconstruct the point cloud. Vehicle 1210 can use the reconstructed point cloud for various purposes. For example, vehicle 1210 can determine that pedestrian 1206 is in the road in front of vehicle 1200 based on the reconstructed point cloud, and therefore begin to slow down, for example, even before the driver of vehicle 1210 realizes that pedestrian 1206 is in the road. Therefore, in some examples, vehicle 1210 can perform autonomous navigation operations based on the reconstructed point cloud.

[0193] Additionally or alternatively, the vehicle 1200 can send the bitstream 1208 to the server system 1212. The server system 1212 can use the bitstream 1208 for various purposes. For example, the server system 1212 can store the bitstream 1208 for subsequent reconstruction of the point cloud. In this example, the server system 1212 can use the point cloud and other data (e.g., vehicle telemetry data generated by the vehicle 1200) to train an autonomous driving system. In other examples, the server system 1212 can store the bitstream 1208 for subsequent reconstruction for use in a forensic collision investigation.

[0194] Fig.13 is a conceptual diagram illustrating an example extended reality system in which one or more of the techniques for coordinate transformation in G-PCC of the present disclosure may be used. Extended reality (XR) is a term used to encompass a range of technologies including augmented reality (AR), mixed reality (MR), and virtual reality (VR). Fig.13 In the example of FIG. 1 , user 1300 is located at a first location 1302. User 1300 wears an XR headset 1304. As an alternative to XR headset 1304, user 1300 may use a mobile device (e.g., a mobile phone, a tablet computer, etc.). XR headset 1304 includes a depth detection sensor, such as a range finding system, which detects the location of a point on object 1306 at location 1302. A data source for XR headset 1304 may use a signal generated by the depth detection sensor to generate a point cloud representation of object 1306 at location 1302. XR headset 1304 may include a G-PCC encoder (e.g., Figure 1 The G-PCC encoder 200 is configured to encode the point cloud to generate a bitstream 1308. Inter-frame prediction and residual prediction as described in the present disclosure can reduce the size of the bitstream 1308.

[0195] The XR headset 1304 may send a bitstream 1308 (e.g., via a network such as the Internet) to an XR headset 1310 worn by a user 1312 at a second location 1314. The XR headset 1310 may decode the bitstream 1308 to reconstruct the point cloud. The XR headset 1310 may use the point cloud to generate an XR visualization (e.g., AR, MR, VR visualization) representing an object 1306 at the location 1302. Thus, in some examples, such as when the XR headset 1310 generates a VR visualization, the user 1312 may have a 3D immersive experience of the location 1302. In some examples, the XR headset 1310 may determine the location of a virtual object based on the reconstructed point cloud. For example, the XR headset 1310 may determine that the environment (e.g., location 1302) includes a flat surface based on the reconstructed point cloud, and then determine that the virtual object (e.g., a cartoon character) is to be positioned on the flat surface. The XR headset 1310 may generate an XR visualization in which a virtual object is at a determined location. For example, the XR headset 1310 may show a cartoon character sitting on a flat surface.

[0196] Fig.14 is a conceptual diagram illustrating an example mobile device system in which one or more techniques for coordinate conversion in G-PCC of the present disclosure may be used. Fig.14 In an example of the present invention, a mobile device 1400 (e.g., a wireless communication device) (e.g., a mobile phone or a tablet computer) includes a range finding system, such as a LIDAR system, which detects the location of points on an object 1402 in the environment of the mobile device 1400. A data source of the mobile device 1400 can use a signal generated by a depth detection sensor to generate a point cloud representation of the object 1402. The mobile device 1400 can include a G-PCC encoder (e.g., Figure 1 The G-PCC encoder 200 is configured to encode the point cloud to generate a bitstream 1404. Fig.14In an example, the mobile device 1400 may send a bitstream to a remote device 1406, such as a server system or other mobile device. Inter-frame prediction and residual prediction as described in the present disclosure may reduce the size of the bitstream 1404. The remote device 1406 may decode the bitstream 1404 to reconstruct a point cloud. The remote device 1406 may use the point cloud for various purposes. For example, the remote device 1406 may use the point cloud to generate a map of the environment of the mobile device 1400. For example, the remote device 1406 may generate a map of the interior of a building based on the reconstructed point cloud. In another example, the remote device 1406 may generate an image (e.g., computer graphics) based on the point cloud. For example, the remote device 1406 may use the points of the point cloud as vertices of a polygon, and use the color attributes of the points as the basis for coloring the polygon. In some examples, the remote device 1406 may use the reconstructed point cloud for facial recognition or other security applications.

[0197] The examples in various aspects of this disclosure may be used alone or in any combination.

[0198] This disclosure includes the following non-limiting terms.

[0199] Item 1A. A method for encoding and decoding point cloud data, the method comprising: determining a first point of the point cloud data as a root node of a first prediction tree branch of a prediction tree; determining that an azimuth difference between a second point of the point cloud data and the first point satisfies a first azimuth threshold, wherein the first point and the second point comprise consecutive points in sequence; terminating the first prediction tree branch at the first point based on the azimuth difference satisfying the first azimuth threshold; determining the second point as a root node of a second prediction tree branch; and encoding and decoding the point cloud data based on the prediction tree.

[0200] Clause 2A. The method of Clause 1A, wherein the order comprises at least one of a sensor capture order or a codec order.

[0201] Clause 3A. The method of clause 1A or clause 2A, further comprising signaling or parsing the first azimuth angle threshold in a bitstream.

[0202] Clause 4A. The method of any of Clauses 1A-3A, wherein the first azimuth angle threshold comprises one of a non-negative number or a negative number.

[0203] Clause 5A. A method according to any one of clauses 1A-4A, wherein the first azimuth threshold comprises a non-negative number, wherein the second azimuth threshold comprises a negative number, wherein determining that the azimuth difference satisfies the first azimuth threshold comprises determining that the azimuth difference is less than the first azimuth threshold, and wherein terminating the first prediction tree branch at the first point is also based on determining that the azimuth difference is greater than a second azimuth threshold, and the second azimuth threshold comprises a negative number.

[0204] Clause 6A. A method according to any one of Clauses 1A-4A, wherein determining that the azimuth difference between the second point and the first point satisfies a first azimuth threshold includes determining that the azimuth difference between the second point and the first point is less than or equal to the first azimuth threshold.

[0205] Clause 7A. A method according to any one of Clauses 1A-4A, wherein determining that the azimuth difference between the second point and the first point satisfies a first azimuth threshold includes determining that an absolute value of the azimuth difference between the second point and the first point is greater than the first azimuth threshold.

[0206] Clause 8A. The method of any of Clauses 1A to 5A, further comprising connecting the first prediction tree branch and the second prediction tree branch.

[0207] Clause 9A. The method of Clause 8A, wherein connecting the first prediction tree branch and the second prediction tree branch comprises adding a root node of the first prediction tree branch as a child node of a root node of the second prediction tree branch.

[0208] Clause 10A. The method of clause 8A, wherein connecting the first prediction tree branch and the second prediction tree branch comprises adding a root node of the first prediction tree branch as a child node of a node in the second prediction tree branch that has a shortest distance to the root node of the first prediction tree branch.

[0209] Clause 11A. The method of clause 8A, wherein connecting the first prediction tree branch and the second prediction tree branch comprises adding a root node of the first prediction tree branch as a child node of a node of a predetermined number of nodes of the second prediction tree branch that has a shortest distance to the root node of the first prediction tree branch.

[0210] Clause 12A. The method of any of Clauses 1A-11A, further comprising: determining a first scan line ID for the first point; and determining a second scan line ID for the second point.

[0211] Clause 13A. The method of Clause 12A, further comprising signaling or resolving one or more characteristics associated with at least one of the first scan line ID or the second scan line ID.

[0212] Item 14A. A method for encoding and decoding point cloud data, the method comprising: determining a first point of the point cloud data as a root node of a first prediction tree branch of a prediction tree; determining a second point of the point cloud data; determining a third point of the point cloud data as a part of the first prediction tree branch, the third point being at least one of encoded and decoded before the second point or captured before the second point; determining whether an azimuth difference between the second point and the third point is greater than a first azimuth threshold; and adding the second point to the prediction tree based on determining whether the azimuth difference between the second point and the third point is greater than the first azimuth threshold.

[0213] Clause 15A. The method of Clause 14A, wherein the azimuth difference is greater than the first azimuth threshold, and wherein adding the second point to the prediction tree comprises adding the second point to the first prediction tree branch.

[0214] Clause 16A. A method according to Clause 14A, wherein the azimuth difference is not greater than the first azimuth threshold, and wherein adding the second point to the prediction tree includes: adding the second point as a root node of a second prediction tree branch; and adding the second point as a child node to a node of the first prediction tree branch.

[0215] Clause 17A. The method of any of Clauses 1A-16A, further comprising generating the point cloud.

[0216] Clause 18A. An apparatus for processing a point cloud, the apparatus comprising one or more components for performing the method of any of Clauses 1A-17A.

[0217] Clause 19A. The apparatus of Clause 18A, wherein the one or more components include one or more processors implemented in circuitry.

[0218] Clause 20A. The apparatus of any of Clauses 18A or 19A, further comprising a memory for storing data representing the point cloud.

[0219] Clause 21A. The apparatus of any of clauses 18A-20A, wherein the apparatus comprises a decoder.

[0220] Clause 22A. The apparatus of any of clauses 18A-21A, wherein the apparatus comprises an encoder.

[0221] Clause 23A. The apparatus of any of Clauses 18A-22A, further comprising a device for generating the point cloud.

[0222] Clause 24A. The apparatus of any of Clauses 18A-23A, further comprising a display for presenting an image based on the point cloud.

[0223] Clause 25A. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to perform the method of any of clauses 1A-17A.

[0224] Item 1B. A device for encoding and decoding point cloud data, the device comprising: one or more memories configured to store the point cloud data; and one or more processors implemented in circuits and communicatively coupled to the one or more memories, the one or more processors configured to: determine a first point of the point cloud data as a first node of a first prediction tree branch of a prediction tree; determine that a first azimuth difference between the first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point include sequentially consecutive points; determine the second point as a first prediction tree branch based on the first azimuth difference not satisfying the first azimuth threshold; a second node of a branch of the point cloud data; determining that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point comprise consecutive points in sequence, and wherein the third point comprises a third node of the first prediction tree branch; based on the second azimuth difference satisfying the first azimuth threshold, terminating the first prediction tree branch at the third point, so that the third point comprises a leaf node of the first prediction tree branch, and determining the fourth point as a first node of a second prediction tree branch; connecting the first prediction tree branch and the second prediction tree branch in the prediction tree; and encoding and decoding the point cloud data based on the prediction tree.

[0225] Clause 2B. An apparatus as described in Clause 1B, wherein as part of connecting the first prediction tree branch and the second prediction tree branch, the one or more processors are configured to add the first node of the second prediction tree branch as a child node of the first node of the first prediction tree branch.

[0226] Clause 3B. The apparatus of clause 2B, wherein the first node of the first prediction tree branch is a root node of the prediction tree.

[0227] Clause 4B. A device according to any of clauses 1B-3B, wherein as part of connecting a first prediction tree branch and a second prediction tree branch, one or more processors are configured to add a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, wherein the node of the first prediction tree branch has a corresponding point of point cloud data, and the corresponding point of the point cloud data has the shortest distance to a fourth point among all nodes of the first prediction tree branch.

[0228] Clause 5B. A device according to any of clauses 1B-4B, wherein as part of connecting a first prediction tree branch and a second prediction tree branch, one or more processors are configured to add a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, the node of the first prediction tree branch having a corresponding point of point cloud data, and the corresponding point of the point cloud data has the shortest distance to a fourth point among a predetermined number of nodes of the first prediction tree branch.

[0229] Clause 6B. The apparatus of any of clauses 1B-5B, wherein the order comprises at least one of a sensor capture order or a codec order.

[0230] Clause 7B. The apparatus of any of clauses 1B-6B, wherein the one or more processors are further configured to signal or parse the first azimuth angle threshold in a bitstream.

[0231] Clause 8B. The apparatus of any of clauses 1B-7B, wherein the first azimuth angle threshold comprises one of a non-negative number or a negative number.

[0232] Clause 9B. An apparatus according to any of clauses 1B-8B, wherein the first azimuth threshold comprises a non-negative number, wherein the second azimuth threshold comprises a negative number, wherein as part of determining that the second azimuth difference satisfies the first azimuth threshold, the one or more processors are configured to determine that the second azimuth difference a) is less than the first azimuth threshold or b) is less than or equal to the first azimuth threshold, and wherein the one or more processors are configured to terminate the first prediction tree branch at the third point based on determining that the second azimuth difference c) is greater than or equal to the second azimuth threshold or d) is greater than the second azimuth threshold.

[0233] Clause 10B. A device as described in any of clauses 1B-9B, wherein as part of determining that the second azimuth difference satisfies the first azimuth threshold, one or more processors are configured to determine that the second azimuth difference is a) less than or equal to the first azimuth threshold or b) less than the first azimuth threshold.

[0234] Clause 11B. A device as described in any of clauses 1B-9B, wherein as part of determining that the second azimuth difference satisfies the first azimuth threshold, one or more processors are configured to determine that an absolute value of the second azimuth difference is a) greater than the first azimuth threshold or b) greater than or equal to the first azimuth threshold.

[0235] Clause 12B. A device as described in any of clauses 1B-11B, wherein the one or more processors are further configured to: determine a first scan line ID for the third point; determine a second scan line ID for the fourth point; and signal or parse in a bit stream one or more characteristics associated with at least one of the first scan line ID or the second scan line ID.

[0236] Clause 13B. The apparatus of any of clauses 1B-12B, wherein as part of encoding and decoding the point cloud data, the one or more processors are configured to encode the point cloud data.

[0237] Clause 14B. The apparatus of any of clauses 1B-13B, wherein the one or more processors are configured to decode the point cloud data as part of encoding and decoding the point cloud data.

[0238] Clause 15B. The apparatus of any of clauses 1B-14B, wherein the one or more processors are further configured to generate the point cloud.

[0239] Clause 16B. A method for encoding and decoding point cloud data, the method comprising: determining a first point of the point cloud data as a first node of a first prediction tree branch of a prediction tree; determining that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point include sequentially consecutive points; based on the first azimuth difference not satisfying the first azimuth threshold, determining the second point as a second node of the first prediction tree branch; determining that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point include sequentially consecutive points, and wherein the third point includes a third node of the first prediction tree branch; based on the second azimuth difference satisfying the first azimuth threshold, terminating the first prediction tree branch at the third point, so that the third point includes a leaf node of the first prediction tree branch, and determining the fourth point as a first node of the second prediction tree branch; connecting the first prediction tree branch and the second prediction tree branch in the prediction tree; and encoding and decoding the point cloud data based on the prediction tree.

[0240] Clause 17B. The method of Clause 16B, wherein connecting the first prediction tree branch and the second prediction tree branch comprises adding a first node of the second prediction tree branch as a child node of the first node of the first prediction tree branch.

[0241] Clause 18B. The method of Clause 17B, wherein the first node of the first prediction tree branch is a root node of the prediction tree.

[0242] Clause 19B. A method according to any one of clauses 16B-18B, wherein connecting a first prediction tree branch and a second prediction tree branch includes adding a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, the node of the first prediction tree branch having a corresponding point of point cloud data, and the corresponding point of the point cloud data having the shortest distance to a fourth point among all nodes of the first prediction tree branch.

[0243] Clause 20B. A method according to any one of clauses 16B-19B, wherein connecting a first prediction tree branch and a second prediction tree branch includes adding a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, the node of the first prediction tree branch having a corresponding point of point cloud data, and the corresponding point of the point cloud data has the shortest distance to a fourth point among a predetermined number of nodes of the first prediction tree branch.

[0244] Clause 21B. The method of any of clauses 16B to 20B, wherein the order comprises at least one of a sensor capture order or a codec order.

[0245] Clause 22B. The method of any of clauses 16B-21B, further comprising signaling or parsing the first azimuth angle threshold in the bitstream.

[0246] Clause 23B. The method of any of clauses 16B-22B, wherein the first azimuth angle threshold comprises one of a non-negative number or a negative number.

[0247] Clause 24B. A method according to any one of clauses 16B to 23B, wherein the first azimuth threshold comprises a non-negative number, wherein the second azimuth threshold comprises a negative number, wherein determining that the second azimuth difference satisfies the first azimuth threshold comprises determining that the second azimuth difference a) is less than the first azimuth threshold or b) is less than or equal to the first azimuth threshold, and wherein terminating the first prediction tree branch at the third point is also based on determining that the second azimuth difference c) is greater than or equal to the second azimuth threshold or d) is greater than the second azimuth threshold.

[0248] Clause 25B. The method of any of clauses 16B-24B, wherein determining that the second azimuth difference satisfies the first azimuth threshold comprises determining that the second azimuth difference is a) less than or equal to the first azimuth threshold or b) less than the first azimuth threshold.

[0249] Clause 26B. The method of any of clauses 16B-24B, wherein determining that the second azimuth difference satisfies the first azimuth threshold comprises determining that an absolute value of the second azimuth difference is a) greater than the first azimuth threshold or b) greater than or equal to the first azimuth threshold.

[0250] Clause 27B. The method according to any one of clauses 16B to 26B further includes: determining a first scan line ID for the third point; determining a second scan line ID for the fourth point; and signaling or parsing one or more characteristics associated with at least one of the first scan line ID or the second scan line ID in a bit stream.

[0251] Clause 28B. The method of any of Clauses 16B to 27B, further comprising generating the point cloud.

[0252] Clause 29B. A device for encoding and decoding point cloud data, the device comprising: a component for determining a first point of the point cloud data as a first node of a first prediction tree branch of a prediction tree; a component for determining that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point include sequentially consecutive points; a component for determining the second point as a second node of the first prediction tree branch based on the first azimuth difference not satisfying the first azimuth threshold; a component for determining a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data as a second node of the first prediction tree branch; A component in which the azimuth difference satisfies the first azimuth threshold, wherein the third point and the fourth point include consecutive points in sequence, and wherein the third point includes the third node of the first prediction tree branch; a component for terminating the first prediction tree branch at the third point based on the second azimuth difference satisfying the first azimuth threshold, so that the third point includes a leaf node of the first prediction tree branch, and determining the fourth point as the first node of the second prediction tree branch; a component for connecting the first prediction tree branch and the second prediction tree branch in the prediction tree; and a component for encoding and decoding point cloud data based on the prediction tree.

[0253] Item 30B. A non-transitory computer-readable storage medium having instructions stored thereon, which when executed cause one or more processors to: determine a first point of cloud data as a first node of a first prediction tree branch of a prediction tree; determine that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point include sequentially consecutive points; based on the first azimuth difference not satisfying the first azimuth threshold, determine the second point as a second node of the first prediction tree branch; determine that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point include sequentially consecutive points, and wherein the third point includes a third node of the first prediction tree branch; based on the second azimuth difference satisfying the first azimuth threshold, terminate the first prediction tree branch at the third point, so that the third point includes a leaf node of the first prediction tree branch, and determine the fourth point as a first node of the second prediction tree branch; connect the first prediction tree branch and the second prediction tree branch in the prediction tree; and encode and decode the point cloud data based on the prediction tree.

[0254] It should be appreciated that, depending on the example, certain actions or events of any of the techniques described herein may be performed in a different order, may be added, combined, or omitted entirely (e.g., not all described actions or events are required to practice the techniques). In addition, in some examples, actions or events may be performed simultaneously rather than sequentially, such as through multithreading, interrupt processing, or multiple processors.

[0255] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or sent via a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates, for example, the transfer of a computer program from one place to another according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementation in the techniques described in the present disclosure. A computer program product may include a computer-readable medium.

[0256] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program codes in the form of instructions or data structures and can be accessed by a computer. In addition, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are sent from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other temporary media, but are directed to non-temporary tangible storage media. As used herein, disks and optical disks include compact disks (CDs), laser optical disks, optical optical disks, digital versatile disks (DVDs), floppy disks, and Blu-ray disks, wherein disks typically reproduce data magnetically, and optical disks reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0257] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuit" as used herein may refer to any of the aforementioned structures or any other structure suitable for the implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. In addition, these techniques may be fully implemented in one or more circuits or logic elements.

[0258] The techniques of the present disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. Instead, as described above, in conjunction with appropriate software and / or firmware, the various units may be combined in a codec hardware unit, or provided by a set of interoperable hardware units including one or more processors as described above.

[0259] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A device for encoding and decoding point cloud data, the device comprising: One or more memories configured to store the point cloud data; as well as one or more processors implemented in circuitry and communicatively coupled to the one or more memories, the one or more processors being configured to: Determine the first point of the point cloud data as the first node of the first prediction tree branch of the prediction tree; determining that a first azimuth difference between the first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point comprise sequentially consecutive points; Based on the first azimuth difference not satisfying the first azimuth threshold, determining the second point as a second node of the first prediction tree branch; determining that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point comprise sequentially consecutive points, and wherein the third point comprises a third node of the first prediction tree branch; Based on the second azimuth difference satisfying the first azimuth threshold, terminating the first prediction tree branch at the third point, so that the third point includes a leaf node of the first prediction tree branch, and determining the fourth point as a first node of the second prediction tree branch; Connecting a first prediction tree branch and a second prediction tree branch in the prediction tree; as well as The point cloud data is encoded and decoded based on the prediction tree.

2. The apparatus of claim 1 , wherein as part of connecting the first prediction tree branch and the second prediction tree branch, the one or more processors are configured to add a first node of the second prediction tree branch as a child node of the first node of the first prediction tree branch. 3 . The apparatus of claim 2 , wherein the first node of the first prediction tree branch is a root node of the prediction tree.

4. The apparatus of claim 1 , wherein as part of connecting the first prediction tree branch and the second prediction tree branch, the one or more processors are configured to add a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, the node of the first prediction tree branch having a corresponding point of the point cloud data, the corresponding point of the point cloud data having the shortest distance to the fourth point among all nodes of the first prediction tree branch.

5. The apparatus of claim 1 , wherein as part of connecting the first prediction tree branch and the second prediction tree branch, the one or more processors are configured to add a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, the node of the first prediction tree branch having a corresponding point of the point cloud data, the corresponding point of the point cloud data having a shortest distance to the fourth point among a predetermined number of nodes of the first prediction tree branch.

6. The device according to claim 1, wherein: The order includes at least one of a sensor capture order or a codec order.

7. The device of claim 1, wherein the one or more processors are further configured to signal or parse the first azimuth angle threshold in a bitstream.

8. The apparatus of claim 1, wherein the first azimuth angle threshold comprises one of a non-negative number or a negative number.

9. The apparatus of claim 1 , wherein the first azimuth threshold comprises a non-negative number, wherein the second azimuth threshold comprises a negative number, wherein as part of determining that the second azimuth difference satisfies the first azimuth threshold, the one or more processors are configured to determine that the second azimuth difference a) is less than the first azimuth threshold or b) is less than or equal to the first azimuth threshold, and wherein the one or more processors are further configured to terminate the first prediction tree branch at the third point based on determining that the second azimuth difference c) is greater than or equal to the second azimuth threshold or d) is greater than the second azimuth threshold.

10. The device of claim 1, wherein as part of determining that the second azimuth difference satisfies the first azimuth threshold, the one or more processors are configured to determine that the second azimuth difference is a) less than or equal to the first azimuth threshold or b) less than the first azimuth threshold.

11. The device of claim 1 , wherein as part of determining that the second azimuth difference satisfies the first azimuth threshold, the one or more processors are configured to determine that an absolute value of the second azimuth difference is a) greater than the first azimuth threshold or b) greater than or equal to the first azimuth threshold.

12. The device of claim 1, wherein the one or more processors are further configured to: Determine a first scan line ID of the third point; Determine a second scan line ID of the fourth point; and One or more characteristics associated with at least one of the first scanline ID or the second scanline ID are signaled or parsed in a bitstream.

13. The apparatus according to claim 1, wherein: As part of encoding and decoding the point cloud data, the one or more processors are configured to encode the point cloud data.

14. The apparatus according to claim 1, wherein: As part of encoding and decoding the point cloud data, the one or more processors are configured to decode the point cloud data.

15. The apparatus according to claim 1, wherein: The one or more processors are further configured to generate the point cloud.

16. A method for encoding and decoding point cloud data, the method comprising: Determine the first point of the point cloud data as the first node of the first prediction tree branch of the prediction tree; determining that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point comprise sequentially consecutive points; Based on the first azimuth difference not satisfying the first azimuth threshold, determining the second point as a second node of the first prediction tree branch; determining that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point comprise sequentially consecutive points, and wherein the third point comprises a third node of the first prediction tree branch; Based on the second azimuth difference satisfying the first azimuth threshold, terminating the first prediction tree branch at the third point, so that the third point includes a leaf node of the first prediction tree branch, and determining the fourth point as a first node of the second prediction tree branch; Connecting a first prediction tree branch and a second prediction tree branch in the prediction tree; as well as The point cloud data is encoded and decoded based on the prediction tree.

17. The method of claim 16, wherein connecting the first prediction tree branch and the second prediction tree branch comprises adding a first node of the second prediction tree branch as a child node of the first node of the first prediction tree branch.

18. The method of claim 17, wherein the first node of the first prediction tree branch is a root node of the prediction tree.

19. The method of claim 16, wherein connecting the first prediction tree branch and the second prediction tree branch comprises adding a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, the node of the first prediction tree branch having a corresponding point of the point cloud data, and the corresponding point of the point cloud data has a shortest distance to the fourth point among all nodes of the first prediction tree branch.

20. The method of claim 16, wherein connecting the first prediction tree branch and the second prediction tree branch comprises adding a first node of the second prediction tree branch as a child node of a node of the first prediction tree branch, the node of the first prediction tree branch having a corresponding point of the point cloud data, the corresponding point of the point cloud data having the shortest distance to the fourth point among a predetermined number of nodes of the first prediction tree branch.

21. The method according to claim 16, wherein: The order includes at least one of a sensor capture order or a codec order.

22. The method of claim 16, further comprising signaling or parsing the first azimuth angle threshold in a bitstream.

23. The method according to claim 16, wherein: The first azimuth angle threshold comprises one of a non-negative number or a negative number.

24. The method according to claim 16, wherein: The first azimuth angle threshold comprises a non-negative number, wherein the second azimuth angle threshold comprises a negative number, wherein determining that the second azimuth angle difference satisfies the first azimuth angle threshold comprises: determining that the second azimuth angle difference a) is less than the first azimuth angle threshold or b) is less than or equal to the first azimuth angle threshold, and wherein terminating the first prediction tree branch at the third point is also based on determining that the second azimuth angle difference c) is greater than or equal to the second azimuth angle threshold or d) is greater than the second azimuth angle threshold.

25. The method of claim 16, wherein: Determining that the second azimuth angle difference satisfies the first azimuth angle threshold includes: determining that the second azimuth angle is a) less than or equal to the first azimuth angle threshold or b) less than the first azimuth angle threshold.

26. The method of claim 16, wherein: Determining whether the second azimuth angle difference satisfies the first azimuth angle threshold includes determining that an absolute value of the second azimuth angle difference is a) greater than the first azimuth angle threshold or b) greater than or equal to the first azimuth angle threshold.

27. The method of claim 16, further comprising: Determine a first scan line ID of the third point; Determine a second scan line ID of the fourth point; as well as One or more characteristics associated with at least one of the first scanline ID or the second scanline ID are signaled or parsed in a bitstream.

28. The method of claim 16, further comprising generating the point cloud.

29. A device for encoding and decoding point cloud data, the device comprising: A component for determining a first point of the point cloud data as a first node of a first prediction tree branch of a prediction tree; means for determining that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point comprise sequentially consecutive points; means for determining the second point as a second node of the first prediction tree branch based on the first azimuth difference not satisfying the first azimuth threshold; means for determining that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point comprise sequentially consecutive points, and wherein the third point comprises a third node of the first prediction tree branch; means for terminating the first prediction tree branch at the third point based on the second azimuth difference satisfying the first azimuth threshold, such that the third point includes a leaf node of the first prediction tree branch, and determining the fourth point as a first node of a second prediction tree branch; means for connecting a first prediction tree branch and a second prediction tree branch in the prediction tree; and A component for encoding and decoding the point cloud data based on the prediction tree.

30. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: Determine the first point of the point cloud data as the first node of the first prediction tree branch of the prediction tree; determining that a first azimuth difference between a first point and a second point of the point cloud data does not satisfy a first azimuth threshold, wherein the first point and the second point comprise sequentially consecutive points; Based on the first azimuth difference not satisfying the first azimuth threshold, determining the second point as a second node of the first prediction tree branch; determining that a second azimuth difference between a third point of the point cloud data and a fourth point of the point cloud data satisfies the first azimuth threshold, wherein the third point and the fourth point comprise sequentially consecutive points, and wherein the third point comprises a third node of the first prediction tree branch; Based on the second azimuth difference satisfying the first azimuth threshold, terminating the first prediction tree branch at the third point, so that the third point includes a leaf node of the first prediction tree branch, and determining the fourth point as a first node of the second prediction tree branch; Connecting a first prediction tree branch and a second prediction tree branch in the prediction tree; as well as The point cloud data is encoded and decoded based on the prediction tree.