Motion estimation in geometric point cloud compression
By using GPS information for global motion compensation, the problem of inaccurate motion compensation in point cloud encoding is solved, thus improving decoding efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-07
- Publication Date
- 2026-03-27
AI Technical Summary
In point cloud coding, existing techniques may lead to inaccurate motion compensation when estimating rotation matrices and translation vectors based on feature points between a reference frame and the current frame, increasing distortion and reducing decoding efficiency.
A global motion compensation method based on Global Positioning System (GPS) information is adopted. By identifying and determining the first set of global motion parameters, the method is converted into the rotation matrix and translation vector of the current frame, thereby improving the accuracy of motion compensation.
This increases the accuracy of motion-compensated prediction frames and reduces the residual of the current frame encoding, thereby improving decoding efficiency.
Smart Images

Figure CN116648681B_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application No. 17 / 495,428, filed October 6, 2021, U.S. Provisional Patent Application No. 63 / 088,936, filed October 7, 2020, U.S. Provisional Patent Application No. 63 / 090,627, filed October 12, 2020, and U.S. Provisional Patent Application No. 63 / 090,657, filed October 12, 2020, the entire contents of each of which are incorporated herein by reference. U.S. Patent Application No. 17 / 495,428, filed October 6, 2021, claims the benefit of U.S. Provisional Patent Application No. 63 / 088,936, filed October 7, 2020, U.S. Provisional Patent Application No. 63 / 090,627, filed October 12, 2020, and U.S. Provisional Patent Application No. 63 / 090,657, filed October 12, 2020. TECHNICAL FIELD
[0002] The present disclosure relates to point cloud encoding and decoding. BACKGROUND
[0003] A point cloud is a collection of points in a three-dimensional space. The points can correspond to points on objects within the three-dimensional space. Thus, a point cloud can be used to represent the physical content of a three-dimensional space. Point clouds can have utility in various situations. For example, a point cloud can be used in the context of an autonomous vehicle for representing the positioning of objects on a road. In another example, a point cloud can be used in the context of representing the physical content of an environment for the purpose of positioning virtual objects in an augmented reality (AR) or mixed reality (MR) application. Point cloud compression is a process for encoding and decoding point clouds. Encoding a point cloud can reduce the amount of data needed to store and transmit the point cloud. SUMMARY
[0004] In general, this disclosure describes techniques for improving the visualization of point cloud frames that can use a geometry-based point cloud compression (G-PCC) codec developed by the 3D Graphics Coding (3DG) group within MPEG. A G-PCC coder (e.g., a G-PCC encoder or a G-PCC decoder) can be configured to apply motion compensation to a reference frame to generate a motion compensated frame. For example, a G-PCC encoder can apply “global” motion compensation to a reference frame (e.g., a prediction frame) to account for a rotation of the entire reference frame and / or a translation of the entire reference frame. In this example, the G-PCC encoder can apply “local” motion estimation to account for rotations and / or translations at a finer scale than the global motion compensation. For example, the G-PCC encoder can apply local node motion estimation of one or more nodes (e.g., portions) of the global motion compensated frame.
[0005] According to the techniques of this disclosure, a G-PCC coder (e.g., a G-PCC encoder or a G-PCC decoder) can be configured to apply global motion compensation based on global positioning system information (e.g., information from any satellite system, such as, for example, the Global Positioning System (GPS) implemented in the United States). For example, the G-PCC encoder can identify a first set of global motion parameters from the global positioning system information. The first set of global motion parameters can include orientation parameters (e.g., roll, pitch, yaw, or angular velocity) and / or position parameters (e.g., displacement or velocity along an x, y, or z dimension). In this example, the G-PCC encoder can determine a second set of global motion parameters based on the first set of global motion parameters. For example, the G-PCC encoder can convert the orientation parameters and / or the position parameters to a rotation matrix and a translation vector for a current frame. In this way, the G-PCC coder (e.g., a G-PCC encoder or a G-PCC decoder) can use satellite information to apply global motion compensation, which can be more accurate than estimating a rotation matrix and a translation vector for a current frame based on feature points between a reference frame (e.g., a predicted frame) and the current frame. Increasing the accuracy of the motion compensation can increase the accuracy of the motion compensated predicted frame, which can reduce the residual that is encoded for the current frame, thereby improving coding efficiency.
[0006] In one example, the disclosure describes a device for encoding point cloud data, the device comprising a memory to store the point cloud data and one or more processors coupled to the memory and implemented in circuitry. The one or more processors are configured to identify a first set of global motion parameters from global positioning system information. The one or more processors are further configured to determine a second set of global motion parameters to be used for a global motion estimation of a current frame based on the first set of global motion parameters, and apply motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
[0007] In another example, the disclosure describes a method for encoding point cloud data, comprising identifying, with one or more processors, a first set of global motion parameters from global positioning system information, and determining, with the one or more processors, a second set of global motion parameters to be used for a global motion estimation of a current frame based on the first set of global motion parameters. The method further comprises applying, with the one or more processors, motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
[0008] In another example, the disclosure describes a computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to identify a first set of global motion parameters from global positioning system information, and determine a second set of global motion parameters to be used for a global motion estimation of a current frame based on the first set of global motion parameters. The instructions also cause the one or more processors to apply motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame of the current frame.
[0009] In another example, the disclosure describes a device for processing point cloud data, the device comprising at least one means for identifying a first set of global motion parameters from global positioning system information, and a means for determining a second set of global motion parameters to be used for a global motion estimation of a current frame based on the first set of global motion parameters. The device also comprises a means for applying motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame of the current frame.
[0010] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 is a block diagram illustrating an example encoding and decoding system that can perform the techniques of this disclosure.
[0012] Figure 2 is a block diagram illustrating an example geometry point cloud compression (G-PCC) encoder according to the techniques of this disclosure.
[0013] Figure 3 is a block diagram illustrating an example G-PCC decoder according to the techniques of this disclosure.
[0014] Figure 4 is a block diagram illustrating an example motion estimation flowchart according to the techniques of this disclosure.
[0015] Figure 5 is a block diagram illustrating an example algorithm for estimating global motion according to the techniques of this disclosure.
[0016] Figure 6 is a block diagram illustrating an example algorithm for estimating local node motion vectors according to the techniques of this disclosure.
[0017] Figure 7 is a conceptual diagram illustrating geodetic latitude and longitude of points measured on an ellipsoidal approximation of the Earth.
[0018] Figure 8is a conceptual diagram showing an ECEF (Earth-Centered, Earth-Fixed) coordinate system (X, Y, Z axes) relative to the equator and the prime meridian (0 degrees latitude and longitude).
[0019] Figure 9 is a flowchart illustrating an example encoding process according to the techniques of this disclosure.
[0020] Figure 10 is a flowchart illustrating an example decoding process according to the techniques of this disclosure. DETAILED DESCRIPTION
[0021] A geometry-based point cloud compression (G-PCC) coder (e.g., a G-PCC encoder or a G-PCC decoder) can be configured to apply motion compensation to a reference frame to generate a motion compensated frame. For example, a G-PCC encoder can apply “global” motion compensation to a reference frame to account for a rotation of the entire reference frame and / or a translation of the entire reference frame. In this example, the G-PCC encoder can apply “local” motion estimation to account for rotations and / or translations at a finer scale than the global motion compensation. For example, the G-PCC encoder can apply local node motion estimation of one or more nodes (e.g., portions) of the global motion compensated frame.
[0022] Some systems can estimate a rotation matrix and a translation vector for a current frame based on feature points between the current frame and a reference frame (e.g., a predicted frame). For example, a G-PCC encoder can estimate a motion matrix and a translation vector to enforce a “match” between respective feature points of the reference frame and the current frame. However, the detected feature points can not be reliable, which can result in an erroneous estimate of the rotation matrix and / or the translation vector, which can increase distortion rather than compensate for global motion. This additional distortion can reduce the coding efficiency of a G-PCC encoder and a G-PCC decoder that uses the rotation matrix and the translation vector.
[0023] According to the technology disclosed herein, a G-PCC encoder can be configured to apply global motion compensation based on Global Positioning System (GPS) information. For example, the G-PCC encoder can identify a first set of global motion parameters. This first set of global motion parameters may include orientation parameters (e.g., roll, pitch, yaw, or angular velocity) and / or positioning parameters (e.g., displacement or velocity along the x, y, or z dimension). In this example, the G-PCC encoder can determine a second set of global motion parameters based on the first set. For example, the G-PCC encoder can convert the orientation parameters and / or positioning parameters into a rotation matrix and translation vector for the current frame. In this way, a G-PCC decoder (e.g., a G-PCC encoder or a G-PCC decoder) can use satellite information to apply global motion compensation, which can be more accurate than estimating the rotation matrix and translation vector of the current frame based on feature points between a reference frame (e.g., a predicted frame) and the current frame. Increasing the accuracy of motion compensation can increase the accuracy of the motion-compensated predicted frame, which can reduce the residual encoded for the current frame and thus improve decoding efficiency.
[0024] Figure 1 This is a block diagram illustrating an example encoding and decoding system 100 capable of implementing the techniques of this disclosure. The techniques of this disclosure generally relate to decoding (encoding and / or decoding) point cloud data, i.e., supporting point cloud compression. Generally, point cloud data includes any data used for processing point clouds. Decoding can efficiently compress and / or decompress point cloud data.
[0025] like Figure 1 As shown, system 100 includes a source device 102 and a destination device 116. The source device 102 provides coded point cloud data to be decoded by the destination device 116. Specifically, in Figure 1 In this example, source device 102 provides point cloud data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can include any of a variety of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, handsets such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, ground or sea vehicles, spacecraft, aircraft, robots, LiDAR devices, satellites, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication.
[0026] exist Figure 1In the example, source device 102 includes a data source 104, a memory 106, a G-PCC encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a G-PCC decoder 300, a memory 120, and a data consumer 118. According to this disclosure, the G-PCC encoder 200 of source device 102 and the G-PCC decoder 300 of destination device 116 can be configured to apply the techniques of this disclosure related to techniques for improving the visualization of point cloud frames. Therefore, source device 102 represents an example of an encoding device, and destination device 116 represents an example of a decoding device. In other examples, source device 102 and destination device 116 may include other components or arrangements. For example, source device 102 may receive data (e.g., point cloud data) from an internal or external source. Similarly, destination device 116 may interface with an external data consumer, rather than including the data consumer in the same device.
[0027] like Figure 1 The system 100 shown is merely an example. In general, other digital encoding and / or decoding devices can perform techniques for improving the visualization of point cloud frames. Source device 102 and destination device 116 are merely examples of such devices, where source device 102 generates encoded data for transmission to destination device 116. This disclosure refers to a “decoding” device as a device that performs data decoding (encoding and / or decoding). Thus, G-PCC encoder 200 and G-PCC decoder 300 represent examples of decoding devices that are specifically encoders and decoders, respectively. In some examples, source device 102 and destination device 116 can operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes both encoding and decoding components. Therefore, system 100 can support unidirectional or bidirectional transmission between source device 102 and destination device 116, for example, for streaming, playback, broadcasting, telephone, navigation, and other applications.
[0028] In general, data source 104 represents a source of data (i.e., raw, uncoded point cloud data), and can provide a series of“frames” of data to G-PCC encoder 200, which encodes the data for the frames. Data source 104 of source device 102 can include a point cloud capture device such as any of a variety of cameras or sensors, e.g., a 3D scanner or a light detection and ranging (LIDAR) device, one or more cameras, an archive containing previously captured data, and / or a data feed interface to receive data from a data content provider. Alternatively or additionally, the point cloud data can be computer generated from a scanner, camera, sensor, or other data computer. For example, data source 104 can generate computer graphics-based data as source data, or produce a combination of live data, archived data, and computer-generated data. That is, data source 104 can generate point cloud data. In each case, G-PCC encoder 200 encodes the captured, pre-captured, or computer-generated data. G-PCC encoder 200 can rearrange the frames from reception order (sometimes referred to as“presentation order”) into coding order for coding. G-PCC encoder 200 can generate one or more bitstreams that include encoded data. Source device 102 can then output the encoded data via output interface 108 onto computer- readable medium 110 for reception and / or retrieval by input interface 122 of destination device 116, for example.
[0029] Memory 106 of source device 102 and memory 120 of destination device 116 can represent general-purpose memory. In some examples, memory 106 and memory 120 can store raw data, e.g., raw data from data source 104 and raw decoded data from G-PCC decoder 300. Additionally or alternatively, memory 106 and memory 120 can store software instructions executable by, e.g., G-PCC encoder 200 and G-PCC decoder 300, respectively. Although memory 106 and memory 120 are shown separately from G-PCC encoder 200 and G-PCC decoder 300 in this example, it should be understood that G-PCC encoder 200 and G-PCC decoder 300 can also include internal memory for similarly or equivalently purposed. Moreover, memory 106 and memory 120 can store encoded data, e.g., output from G-PCC encoder 200 and input to G-PCC decoder 300. In some examples, portions of memory 106 and memory 120 can be allocated as one or more buffers, e.g., to store raw, decoded, and / or encoded data. For example, memory 106 and memory 120 can store data representing a point cloud.
[0030] The computer-readable medium 110 can represent any type of medium or device capable of transporting encoded data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium to enable the source device 102 to transmit encoded data directly to the destination device 116 in real-time, such as via a radio frequency network or computer-based network. The output interface 108 can modulate a transmission signal including the encoded data, and the input interface 122 can demodulate the received transmission signal, in accordance with a communication standard, such as a wireless communication protocol. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from the source device 102 to the destination device 116.
[0031] In some examples, the source device 102 can output the encoded data from the output interface 108 to a storage device 112. Similarly, the destination device 116 can access the encoded data from the storage device 112 via the input interface 122. The storage device 112 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded data.
[0032] In some examples, the source device 102 can output the encoded data to a file server 114 or another intermediate storage device that can store encoded data generated by the source device 102. The destination device 116 can access the stored data from the file server 114 via streaming or download. The file server 114 can be any type of server device that is capable of storing encoded data and transmitting that encoded data to the destination device 116. The file server 114 can represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. The destination device 116 can access the encoded data from the file server 114 by any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both that is suitable for accessing encoded data stored on the file server 114. The file server 114 and the input interface 122 can be configured to operate according to streaming protocols, download transmission protocols, or a combination thereof.
[0033] Output interface 108 and input interface 122 can represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of a variety of IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 comprise wireless components, output interface 108 and input interface 122 can be configured to transfer data, such as encoded data, according to a cellular communication standard, such as 4G, 4G-LTE (Long-Term Evolution), LTE Advanced, 5G, or the like. In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured to transfer data, such as encoded data, according to other wireless standards, such as an IEEE 802.11 specification, an IEEE 802.15 specification (e.g., ZigBee TM ), a Bluetooth TM standard, or the like. In some examples, source device 102 and / or destination device 116 can include respective system on a chip (SoC) devices. For example, source device 102 can include an SoC device to perform the functionality considered to be G-PCC encoder 200 and / or output interface 108, and destination device 116 can include an SoC device to perform the functionality considered to be G-PCC decoder 300 and / or input interface 122.
[0034] The techniques of this disclosure can be applied to encoding and decoding to support any of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors, and processing devices (such as local or remote servers), geographic mapping, or other applications.
[0035] Input interface 122 of destination device 116 receives an encoded bitstream from computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, or the like). The encoded bitstream can include signaling information defined by G-PCC encoder 200 that is also used by G-PCC decoder 300, such as syntax elements having values that describe characteristics and / or processing of coding units (e.g., slices, pictures, groups of pictures, sequences, or the like). Data consumer 118 uses the decoded data. For example, data consumer 118 can use the decoded data to determine a location of a physical object. In some examples, data consumer 118 can include a display that renders images based on the point cloud.
[0036] The G-PCC encoder 200 and the G-PCC decoder 300 each can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware or any combinations thereof. When the techniques are implemented partially in software, a device can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of the G-PCC encoder 200 and the G-PCC decoder 300 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective device. A device including the G-PCC encoder 200 and / or the G-PCC decoder 300 can comprise one or more integrated circuits (ICs), microprocessors, and / or other types of devices.
[0037] The G-PCC encoder 200 and the G-PCC decoder 300 can operate according to a coding standard, such as the Video Point Cloud Compression (V-PCC) standard or the Geometric Point Cloud Compression (G-PCC) standard. This disclosure can generally refer to coding (e.g., encoding and decoding) of pictures to include processes of encoding or decoding data. An encoded bitstream typically includes a series of values for syntax elements representing coding decisions (e.g., coding modes).
[0038] This disclosure can generally refer to “signaling” certain information, such as syntax elements. The term “signaling” can generally refer to the communication of values for syntax elements and / or other data used to decode encoded data. That is, the G-PCC encoder 200 can signal values for syntax elements in a bitstream. Typically, signaling refers to generating values in a bitstream. As described above, the source device 102 can transmit the bitstream to the destination device 116 in real time or non-real time, such as can occur when storing syntax elements to the storage device 112 for later retrieval by the destination device 116.
[0039] The ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) is investigating the potential need for standardization of point cloud coding techniques that significantly outperform current approaches in compression capability and will target the creation of a standard. This group is collaborating on this exploratory activity, known as the Three-Dimensional Graphics Team (3DG), to evaluate compression technology designs proposed by experts in the field.
[0040] Point cloud compression activities are classified into two different approaches. The first approach is “Video-based Point Cloud Compression” (V-PCC), which segments 3D objects and projects the segments (represented as “patches” in 2D frames) in multiple 2D planes, which are further coded by traditional 2D video codecs, such as High Efficiency Video Coding (HEVC) (ITU-T H.265) codec. The second approach is “Geometry-based Point Cloud Compression” (G-PCC), which directly compresses 3D geometry, i.e., the positioning of a set of points in 3D space and associated attribute values (for each point associated with the 3D geometry). G-PCC addresses the compression of point clouds in Category 1 (static point clouds) and Category 3 (dynamically acquired point clouds). The latest draft of the G-PCC standard is available in G-PCC DIS (ISO / IEC JTC1 / SC29 / WG11 w19088, Brussels, Belgium, Jan. 2020), and the description of the codec is available in G-PCC Codec Specification v6 (ISO / IEC JTC1 / SC29 / WG11 w19091, Brussels, Belgium, Jan. 2020).
[0041] A point cloud contains a set of points in 3D space and can have attributes associated with the points. The attributes can be color information, such as R, G, B, or Y, Cb, Cr, or reflectance information or other attributes. Point clouds can be captured by various cameras or sensors, such as LIDAR sensors and 3D scanners, and can also be computer generated. That is, source device 102 can generate point cloud data based on signals from a LIDAR device (e.g., a LIDAR sensor and / or a LIDAR apparatus). Point cloud data is used for various applications, including, but not limited to, architecture (modeling), graphics (3D models for visualization and animation), and automotive industry (LIDAR sensors to help with navigation).
[0042] The 3D space occupied by the point cloud data can be enclosed by a virtual bounding box. The positions of the points in the bounding box can be represented with a certain precision; thus, the positioning of one or more points can be quantized based on the precision. At a minimum level, the bounding box is partitioned into voxels, which are the smallest unit of space represented by a unit cube. Voxels in the bounding box can be associated with zero, one, or more points. The bounding box can be partitioned into multiple cube / cuboid regions, which can be referred to as tiles. Each tile can be coded into one or more slices. The division of the bounding box into slices and tiles can be based on the number of points in each partition, or based on other considerations (e.g., a particular region can be coded as a tile). The slice regions can be further divided using partitioning decisions similar to those in video codecs.
[0043] According to techniques of this disclosure, G-PCC encoder 200 can be configured to apply global motion compensation based on global positioning system information. For example, G-PCC encoder 200 can identify a first set of global motion parameters. The first set of global motion parameters can include orientation parameters (e.g., roll, pitch, yaw, or angular velocity) and / or position parameters (e.g., displacement or velocity along an x, y, or z dimension). In this example, G-PCC encoder 200 can determine a second set of global motion parameters based on the first set of global motion parameters. For example, G-PCC encoder 200 can convert the orientation parameters and / or the position parameters to a rotation matrix and a translation vector for a current frame. In this way, G-PCC encoder 200 can apply global motion compensation using satellite information, which can be more accurate than estimating a rotation matrix and a translation vector for a current frame based on feature points between a reference frame (e.g., a predicted frame) and the current frame.
[0044] G-PCC encoder 200 can signal the first set of global motion parameters (e.g., the orientation parameters and / or the position parameters). In some examples, G-PCC encoder 200 can signal the second set of global motion parameters (e.g., the rotation matrix and the translation vector for the current frame). While examples describe signaling the rotation matrix and the translation vector for the current frame to signal the second set of global motion parameters, in some examples, G-PCC encoder 200 can signal a portion of and / or an estimate of the rotation matrix and the translation vector for the current frame. For example, G-PCC encoder 200 can signal a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicating an average velocity of the current frame. In some examples, G-PCC encoder 200 can signal the second set of global motion parameters to include only the translation vector indicating the average velocity of the current frame. In some examples, G-PCC encoder 200 can signal the second set of global motion parameters to include only a magnitude of the translation vector indicating the average velocity of the current frame. In this way, G-PCC encoder 200 can apply global motion compensation using satellite information, which can be more accurate than estimating a rotation matrix and a translation vector for a current frame based on feature points between a reference frame (e.g., a predicted frame) and the current frame. Increasing the accuracy of the motion compensation can increase the accuracy of the motion compensated predicted frame, which can reduce the residual encoded for the current frame and thereby improve coding efficiency.
[0045] Similarly, the G-PCC decoder 300 can apply global motion compensation based on global positioning system information. For example, the G-PCC decoder 300 can decode global motion information (e.g., a first set of global motion parameters and / or a second set of global motion parameters) from the bitstream. In this example, the G-PCC decoder 300 can determine global motion for the current frame based on the global motion information. For example, the G-PCC decoder 300 can apply global motion compensation based on a rotation matrix and a translation vector for the current frame.
[0046] The G-PCC decoder 300 can decode the rotation matrix and the translation vector directly from the bitstream or can estimate and / or derive the rotation matrix and the translation vector from global motion information signaled in the bitstream. For example, the G-PCC decoder 300 can decode a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicating an average velocity of the current frame from the bitstream. In this example, the G-PCC decoder 300 can determine the rotation matrix and the translation vector based on the roll difference, the pitch difference, the yaw difference, and the translation vector. In some examples, the G-PCC decoder 300 can decode the translation vector from the bitstream and approximate the rotation matrix as an identity matrix. In some examples, the G-PCC decoder 300 can decode a magnitude of the translation vector indicating an average velocity of the current frame from the bitstream and approximate the translation vector based on the received magnitude and approximate the rotation matrix as an identity matrix of the rotation matrix. In this way, the G-PCC decoder 300 can apply global motion compensation using satellite information, which can be more accurate than estimating a rotation matrix and a translation vector for the current frame based on feature points between a reference frame (e.g., a predicted frame) and the current frame. Increasing the accuracy of the motion compensation can increase the accuracy of the motion compensated predicted frame, which can reduce the residual encoded for the current frame and thus improve coding efficiency.
[0047] Figure 2 An overview of the G-PCC encoder 200 is provided. Figure 3 An overview of the G-PCC decoder 300 is provided. The illustrated modules are logical modules and do not necessarily correspond one-to-one with implementation code in a reference implementation of the G-PCC codec, i.e., the TMC13 test model software under study by ISO / IEC MPEG (JTC 1 / SC 29 / WG 11).
[0048] In both the G-PCC encoder 200 and the G-PCC decoder 300, point cloud positioning is coded first. Attribute coding depends on the decoded geometry. In the G-PCC encoder 200, the geometry is encoded first, and then the attributes are encoded. Figure 2 and Figure 3In particular, the gray-shaded module is an option typically used for Category 1 data. The diagonal cross-hatched module is an option typically used for Category 3 data. All other modules are generic between Category 1 and Category 3.
[0049] For Category 3 data, the compressed geometry is typically represented as an octree from the root all the way down to the leaf level of individual voxels. For Category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to the leaf level of blocks larger than voxels) plus a model that approximates the surface within each leaf of the pruned octree. In this way, Category 1 and Category 3 data share the octree coding mechanism, while Category 1 data can additionally approximate the voxels within each leaf with a surface model. The surface model used is a triangle partition that includes 1-10 triangles per block, forming a triangle soup. Thus, the Category 1 geometry codec is referred to as a Trisoup geometry codec, while the Category 3 geometry codec is referred to as an Octree geometry codec.
[0050] At each node of the octree, occupancy is signaled (when not inferred) for one or more of its child nodes (up to eight nodes). Multiple neighborhoods are specified, including (a) nodes that share a face with the current octree node, (b) nodes that share a face, edge, or vertex with the current octree node, etc. Within each neighborhood, the occupancy of the node and / or its child nodes can be used to predict the occupancy of the current node or its child nodes. For points that are sparsely distributed in certain nodes of the octree, the codec also supports a direct coding mode in which the 3D position of the point is directly encoded. A flag can be signaled to indicate that the direct mode is signaled. At the lowest level, the number of points associated with an octree node / leaf node can also be coded.
[0051] Once the geometry is coded, the attributes corresponding to the geometry points are coded. When there are multiple attribute points corresponding to one reconstructed / decoded geometry point, the attribute value representing the reconstructed point can be derived.
[0052] There are three attribute coding methods in G-PCC: Region Adaptive Hierarchical Transform (RAHT) coding, interpolation-based hierarchical nearest-neighbor prediction (predictive transform), and interpolation-based hierarchical nearest-neighbor prediction with an update / promotion step (promoted transform). RAHT and promoted are typically used for Category 1 data, while predictive is typically used for Category 3 data. However, any of the methods can be used for any data, just like the geometry codecs in G-PCC, the attribute coding method used to code the point cloud is specified in the bitstream.
[0053] Attribute decoding can be performed at the level of detail (LOD), where a finer representation of the point cloud attributes can be obtained at each level of detail. Each level of detail can be specified based on a distance metric to neighboring nodes or based on the sampling distance.
[0054] At the G-PCC encoder 200, quantization is the residual obtained as the output of the attribute-specific decoding method. The residual can be obtained by subtracting the attribute value from a prediction derived from the attribute values of points in the neighborhood of the current point and based on the attribute values of previously encoded points. The quantized residual can be decoded using context-adaptive arithmetic decoding.
[0055] exist Figure 2 In the example, the G-PCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometric reconstruction unit 216, a RAHT unit 218, a LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.
[0056] like Figure 2 As shown in the example, the G-PCC encoder 200 can obtain a set of locations and a set of attributes for points in a point cloud. The G-PCC encoder 200 can obtain data from data source 104 ( Figure 1 The G-PCC encoder 200 obtains a set of locations and a set of attributes for points in a point cloud. Location can include the coordinates of the points in the point cloud. Attributes can include information about the points in the point cloud, such as the color associated with the points. The G-PCC encoder 200 can generate a geometric bitstream 203, which includes an encoded representation of the locations of the points in the point cloud. The G-PCC encoder 200 can also generate an attribute bitstream 205, which includes an encoded representation of a set of attributes.
[0057] The coordinate transformation unit 202 can apply transformations to the coordinates of a point to transform the coordinates from the initial domain to the transformation domain. The transformed coordinates can be referred to as transformed coordinates. The color transformation unit 204 can apply attribute transformations to transform color information to different domains. For example, the color transformation unit 204 can transform color information from the RGB color space to the YCbCr color space.
[0058] In addition, Figure 2 In the example, voxelization unit 206 can voxelize the transformed coordinates. Voxelization of the transformed coordinates can include quantization and removal of some points in the point cloud. In other words, multiple points in the point cloud can be contained in a single "voxel" and then treated as a single point in some respects.
[0059] The octree analysis unit 210 can generate an octree based on the voxelized transformed coordinates. According to the techniques of this disclosure, the octree analysis unit 210 can be configured to apply global motion compensation based on global positioning system information. For example, the octree analysis unit 210 can identify a first set of global motion parameters. The first set of global motion parameters can include orientation parameters (e.g., roll, pitch, yaw, or angular velocity) and / or position parameters (e.g., displacement or velocity along the x, y, or z dimension). In this example, the octree analysis unit 210 can determine a second set of global motion parameters based on the first set of global motion parameters. For example, the octree analysis unit 210 can convert the orientation parameters and / or the position parameters to a rotation matrix and a translation vector for the current frame. In this way, the octree analysis unit 210 can apply global motion compensation using satellite information, which can be more accurate than estimating a rotation matrix and a translation vector for the current frame based on feature points between a reference frame (e.g., a predicted frame) and the current frame.
[0060] Furthermore, in Figure 2 In examples, the surface approximation analysis unit 212 can analyze the points to potentially determine surface representations for groups of points. The arithmetic encoding unit 214 can entropy encode syntax elements representing information for the octrees and / or surfaces determined by the surface approximation analysis unit 212. The G-PCC encoder 200 can output these syntax elements in the geometry bitstream 203. The geometry bitstream 203 can also include other syntax elements, including syntax elements that are not arithmetically encoded.
[0061] The geometry reconstruction unit 216 can reconstruct the transformed coordinates for the points in the point cloud based on the octrees, data indicating the surfaces determined by the surface approximation analysis unit 212, and / or other information. Due to the voxelization and the surface approximation, the number of transformed coordinates reconstructed by the geometry reconstruction unit 216 can be different from the original number of points of the point cloud. This disclosure can refer to the resulting points as reconstructed points. The attribute transfer unit 208 can transfer attributes of the original points of the point cloud to the reconstructed points of the point cloud.
[0062] Further, the RAHT unit 218 can apply RAHT coding to the attributes of the reconstructed points. In some examples, under RAHT, the attributes of a 2x2x2 point positioned block are obtained and transformed along one direction to obtain four low (L) frequency nodes and four high (H) frequency nodes. Subsequently, the four low frequency nodes (L) are transformed in a second direction to obtain two low (LL) frequency nodes and two high (LH) frequency nodes. The two low frequency nodes (LL) are transformed along a third direction to obtain one low (LLL) frequency node and one high (LLH) frequency node. The low frequency node LLL corresponds to the DC coefficient, while the high frequency nodes H, LH, and LLH correspond to AC coefficients. The transformation in each direction can be a one-dimensional (ID) transform with two coefficient weights. The low frequency coefficients can be treated as coefficients of a 2x2x2 block for the next higher level of RAHT transform, and the AC coefficients are encoded without change; such transformation continues up to the top root node. The tree traversal for encoding is from top to bottom for computing the weights for the coefficients; the transformation order is from bottom to top. The coefficients can then be quantized and coded.
[0063] Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 can apply LOD processing and lifting, respectively, to the attributes of the reconstructed points. The lifting unit 222 can be configured to perform an interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform). In some examples, the lifting unit 222 can be configured to perform global motion compensation.
[0064] The LOD generation unit 220 can be used to partition the attributes into different levels of refinement. Each level of refinement refines the attributes of the point cloud. The first level of refinement provides a coarse approximation and contains few points; subsequent levels of refinement typically contain more points, and so on. The levels of refinement can be constructed using a distance-based metric, or one or more other classification criteria can also be used (e.g., subsampling from a particular order). Thus, all reconstructed points can be included in the levels of refinement. Each level of detail is produced by taking the union of all points up to a particular level of refinement: for example, LOD1 is obtained based on refinement level RL1, LOD2 is obtained based on RL1 and RL2,... LODN is obtained by the union of RL1, RL2,... RLN. In some cases, the LOD generation can be followed by a prediction scheme (e.g., a prediction transform), in which the attributes associated with each point in the LOD are predicted from a weighted average of previous points, and the residuals are quantized and entropy encoded. The lifting scheme builds on the prediction transform mechanism, in which an update operator is used to update the coefficients and perform adaptive quantization of the coefficients.
[0065] The RAHT unit 218 and the lifting unit 222 can generate coefficients based on the attributes. The coefficient quantization unit 224 can quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 can apply arithmetic coding to syntax elements representing the quantized coefficients. The G-PCC encoder 200 can output these syntax elements in the attribute bitstream 205. The attribute bitstream 205 can also include other syntax elements, including syntax elements that are not arithmetically coded.
[0066] The arithmetic encoding unit 214 can signal the first set of global motion parameters (e.g., orientation parameters and / or position parameters). In some examples, the arithmetic encoding unit 214 can signal the second set of global motion parameters (e.g., a rotation matrix and a translation vector of the current frame) in the geometry bitstream 203. While examples describe signaling a rotation matrix and a translation vector of the current frame to signal the second set of global motion parameters, in some examples, the arithmetic encoding unit 214 can signal a portion of and / or an estimate of the rotation matrix and the translation vector of the current frame in the geometry bitstream 203. For example, the arithmetic encoding unit 214 can signal a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicating an average velocity of the current frame. In some examples, the arithmetic encoding unit 214 can signal the second set of global motion parameters to include only the translation vector indicating the average velocity of the current frame. In some examples, the arithmetic encoding unit 214 can signal the second set of global motion parameters to include only a magnitude of the translation vector indicating the average velocity of the current frame. In this way, the arithmetic encoding unit 214 can apply global motion compensation using satellite information, which can be more accurate than estimating a rotation matrix and a translation vector of the current frame based on feature points between a reference frame (e.g., a predicted frame) and the current frame.
[0067] In Figure 3 In examples, the G-PCC decoder 300 can include a geometry arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometry reconstruction unit 312, a RAHT unit 314, a LOD generation unit 316, an inverse lifting unit 318, an inverse transform coordinate unit 320, and an inverse transform color unit 322.
[0068] The G-PCC decoder 300 can obtain the geometry bitstream 203 and the attribute bitstream 205. The geometry arithmetic decoding unit 302 of the decoder 300 can apply arithmetic decoding (e.g., context adaptive binary arithmetic coding (CABAC) or other types of arithmetic decoding) to the syntax elements in the geometry bitstream 203. Similarly, the attribute arithmetic decoding unit 304 can apply arithmetic decoding to the syntax elements in the attribute bitstream 205.
[0069] For example, the geometry arithmetic decoding unit 302 can receive global motion information (e.g., a first set of global motion parameters and / or a second set of global motion parameters) from the geometry bitstream 203. For example, the geometry arithmetic decoding unit 302 can receive a second set of global motion parameters (e.g., a rotation matrix and a translation vector for the current frame). While examples describe receiving a rotation matrix and a translation vector for the current frame to signal the second set of global motion parameters, in some examples, the geometry arithmetic decoding unit 302 can receive a rotation matrix and a portion and / or estimate of a rotation vector for the current frame. For example, the geometry arithmetic decoding unit 302 can receive a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicating an average velocity of the current frame. In some examples, the geometry arithmetic decoding unit 302 can receive a second set of global motion parameters including only a translation vector indicating an average velocity of the current frame. In some examples, the geometry arithmetic decoding unit 302 can receive a second set of global motion parameters including only a magnitude of a translation vector indicating an average velocity of the current frame.
[0070] The octree synthesis unit 306 can synthesize an octree based on the syntax elements parsed from the geometry bitstream 203. Starting from the root node of the octree, the occupancy of each of the eight child nodes of each octree level is signaled in the bitstream. When signaling indicates that a child node of a particular octree level is occupied, the occupancy of the child of that node is signaled. The signaling of the nodes of each octree level is signaled before proceeding to a subsequent octree level. At the final level of the octree, each node corresponds to a voxel location; when a leaf node is occupied, one or more points can be specified as being occupied at the voxel location. In some cases, due to quantization, certain branches of the octree can terminate earlier than the final level. In such cases, the leaf node is treated as an occupied node without child nodes. In cases where a surface approximation is used in the geometry bitstream 203, the surface approximation synthesis unit 310 can determine a surface model based on the syntax elements parsed from the geometry bitstream 203 and based on the octree.
[0071] The octree synthesis unit 306 can determine global motion for the current frame based on the global motion information. For example, the octree synthesis unit 306 can apply global motion compensation based on a rotation matrix and a translation vector for the current frame. Likewise, the geometry arithmetic decoding unit 302 can receive the rotation matrix and the translation vector from the geometry bitstream 203. The octree synthesis unit 306 can determine or estimate the second set of global motion parameters (e.g., the rotation matrix and the translation vector) based on the first set of global motion parameters and / or a portion of the second global motion parameters.
[0072] For example, the octree synthesis unit 306 can decode, from the geometry bitstream 203, a roll difference between the current frame and the reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicative of an average velocity of the current frame. In this example, the octree synthesis unit 306 can determine the rotation and the translation vector based on the roll difference, the pitch difference, the yaw difference, and the translation vector.
[0073] In some examples, the octree synthesis unit 306 can decode, from the geometry bitstream 203, the translation vector and approximate the rotation matrix as an identity matrix. In some examples, the octree synthesis unit 306 can decode, from the geometry bitstream 203, a magnitude of the translation vector indicative of an average velocity of the current frame and approximate the translation vector based on the received magnitude and approximate the rotation matrix as an identity matrix of the rotation matrix. In this way, the octree synthesis unit 306 can use satellite information to apply global motion compensation, which can be more accurate than estimating a rotation matrix and a translation vector for the current frame based on feature points between a reference frame (e.g., a predicted frame) and the current frame.
[0074] The octree synthesis unit 306 can convert the orientation parameters and / or the position parameters to a rotation matrix and a translation vector for the current frame. In this way, the octree synthesis unit 306 can use satellite information to apply global motion compensation, which can be more accurate than estimating a rotation matrix and a translation vector for the current frame based on feature points between a reference frame (e.g., a predicted frame) and the current frame.
[0075] Further, the geometry reconstruction unit 312 can perform reconstruction to determine coordinates of points in the point cloud. For each position at a leaf node of the octree, the geometry reconstruction unit 312 can reconstruct the node position by using the binary representation of the leaf node in the octree. At each respective leaf node, a number of points at the respective leaf node is signaled; this indicates a number of duplicate points at the same voxel position. When using geometry quantization, the scaled point positions are used to determine reconstructed point position values.
[0076] Inverse transform coordinate unit 320 can apply an inverse transform to the reconstructed coordinates to convert the reconstructed coordinates (positioning) of points in the point cloud from the transformed domain back to the original domain. The positioning of points in the point cloud can be in a floating point domain, but the point positioning in the G-PCC codec is coded in an integer domain. The inverse transform can be used to convert the positioning back to the original domain.
[0077] Additionally, in Figure 3 In examples of the G-PCC decoder 300, inverse quantization unit 308 can inverse quantize the attribute values. The attribute values can be based on syntax elements obtained from the attribute bitstream 205 (e.g., including syntax elements decoded by the attribute arithmetic decoding unit 304).
[0078] Depending on how the attribute values are encoded, the RAHT unit 314 can perform RAHT decoding to determine the color values of points of the point cloud based on the inverse quantized attribute values. RAHT decoding is performed from the top of the tree to the bottom. At each level, constituent values are derived using low frequency coefficients and high frequency coefficients derived from the inverse quantization process. At leaf nodes, the derived values correspond to the attribute values of the coefficients. The weight derivation process for points is similar to the process used at the G-PCC encoder 200. Alternatively, the LOD generation unit 316 and the inverse lifting unit 318 can use a technique based on levels of detail to determine the color values of points of the point cloud. The LOD generation unit 316 decodes each LOD, giving progressively finer representations of the attributes of points. With a predictive transform, the LOD generation unit 316 can derive a prediction of a point from a weighted sum of points at a previous LOD or previously reconstructed in the same LOD. The LOD generation unit 316 can add the prediction to a residual (which is obtained after inverse quantization) to obtain a reconstructed value of the attribute. When a lifting scheme is used, the LOD generation unit 316 can also include an update operator to update the coefficients used to derive the attribute values. In this case, the LOD generation unit 316 can also apply inverse adaptive quantization.
[0079] Further, in Figure 3 In examples of the G-PCC decoder 300, inverse transform color unit 322 can apply an inverse color transform to the color values. The inverse color transform can be the inverse of the color transform applied by the color transform unit 204 of the G-PCC encoder 200. For example, the color transform unit 204 can transform color information from an RGB color space to a YCbCr color space. Accordingly, the inverse color transform unit 322 can transform the color information from the YCbCr color space to the RGB color space.
[0080] Figure 2 and Figure 3The various units of G-PCC encoder 200 and G-PCC decoder 300 are shown to assist with understanding the operations performed by G-PCC encoder 200 and G-PCC decoder 300. These units can be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Fixed-function circuitry refers to circuitry that provides particular functionality, and is preset on the types of operations that it can perform. Programmable circuitry refers to circuitry that can be programmed to perform various tasks, and provides flexible functionality in the types of operations that it can perform. For example, programmable circuitry can execute software or firmware that cause the programmable circuitry to operate in a manner defined by instructions of the software or firmware. Fixed-function circuitry can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations that the fixed-function circuitry performs are generally immutable. In some examples, one or more of the units can be distinct circuit blocks (fixed-function or programmable), and in some examples, one or more of the units can be integrated circuitry.
[0081] According to the techniques of this disclosure, G-PCC encoder 200 can represent an example of a device that includes memory for storing point cloud data and one or more processors coupled to the memory and implemented in circuitry. The one or more processors are configured to identify a first set of global motion parameters from global positioning system information, determine a second set of global motion parameters for a global motion estimation of a current frame based on the first set of global motion parameters. The one or more processors are further configured to apply motion compensation to a reference frame using the second set of global motion parameters to generate a global motion compensated frame for the current frame.
[0082] G-PCC decoder 300 can represent an example of a device that includes memory for storing point cloud data and one or more processors coupled to the memory and implemented in circuitry. The one or more processors can be configured to decode a bitstream of symbols that indicate global motion information, determine a set of global motion parameters to be used for a global motion estimation of a current frame based on the global motion information. The one or more processors are further configured to apply motion compensation to a reference frame using the set of global motion parameters to generate a global motion compensated frame for the current frame.
[0083] G-PCC techniques involve two types of motion, global motion matrices and local node motion vectors. Global motion parameters can include a rotation matrix and a translation vector. A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can apply global motion compensation to all points in a reference frame (e.g., a predicted frame). A local node motion vector of a node of an octree is a motion vector that is applied only to points within the node in the predicted (reference) frame. For example, a G-PCC coder can apply local node motion compensation only on a portion (e.g., a set of nodes) of a reference frame.
[0084] Figure 4is a block diagram illustrating an example motion estimation flowchart according to the techniques of this disclosure. Given a predicted frame 402 and a current frame 410, the G-PCC encoder 200 can first estimate global motion on a global scale (404). For example, the G-PCC encoder 200 can generate a rotation matrix and a translation vector according to the techniques described herein. After applying the global motion to the predicted frame 402 (406), the G-PCC encoder 200 can apply local node motion estimation (408) to estimate local motion in an octree at a finer scale and node level. The G-PCC encoder 200 can apply the estimated local node motion in motion compensation. For example, the G-PCC encoder 200 can apply motion estimation and encode motion vector information and point information (412).
[0085] Figure 5 is a block diagram illustrating an example algorithm for estimating global motion. Figure 5 The illustrated process can be an example of step 404 of Figure 4 The G-PCC encoder 200 can be configured to define a global motion matrix to match feature points between a predicted frame (reference) and a current frame. The entire global motion estimation algorithm can be divided into three steps: finding feature points (502), sampling pairs of feature points (504), and performing motion estimation using a least mean square (LMS) algorithm (506).
[0086] The G-PCC encoder 200 can perform the LMS algorithm (506) to define points between the predicted frame and the current frame that have large positional changes as feature points. For each point in the current frame, the G-PCC encoder 200 can find the closest point in the predicted frame, and can establish a pair of points between the current frame and the predicted frame (502). If the distance between the paired points is greater than a threshold, the G-PCC encoder 200 can consider the paired points as feature points.
[0087] After finding the feature points, the G-PCC encoder can perform sampling on the feature points to reduce the scale of the problem (e.g., by selecting a subset of the feature points to reduce the complexity of the motion estimation) (504). The G-PCC encoder 200 can then apply a least mean square (LMS) algorithm to derive motion parameters by trying to reduce the error between the corresponding feature points in the predicted frame and the current frame (506). This process can loop for each pair of feature points of step 502.
[0088] Figure 6 is a block diagram illustrating an example algorithm for estimating local node motion vectors. Figure 6 The illustrated process can be an example of step 408 of Figure 4 The G-PCC encoder 200 can be configured to define a global motion matrix to match feature points between a predicted frame (reference) and a current frame. The entire global motion estimation algorithm can be divided into three steps: finding feature points (502), sampling pairs of feature points (504), and performing motion estimation using a least mean square (LMS) algorithm (506). Figure 6In the example of FIG. 6, G-PCC encoder 200 can estimate motion vectors in a recursive manner. G-PCC encoder 200 can use a cost function to select the most suitable motion vector based on rate-distortion cost.
[0089] If the current node is not split into 8 sub-nodes, G-PCC encoder 200 can determine a motion vector that can result in the lowest cost between the current node 602 and the predicted node. If the current node is split into 8 sub-nodes (610), G-PCC encoder 200 can apply a motion estimation algorithm to find the motion for each sub-node (612), and can obtain the total cost under the split condition by adding the estimated cost values of each sub-node (614). G-PCC encoder 200 can decide whether to split or not by comparing the cost between the split and not split (606). If the current node is split, G-PCC encoder 200 can assign a respective motion vector to each sub-node (or can further split to its sub-nodes). If the current node is not split, G-PCC encoder 200 can find a motion vector that achieves the lowest cost (604), and can assign the motion vector to the current node.
[0090] Two parameters that affect the performance of motion vector estimation are BlockSize and MinPUSize. BlockSize defines the upper bound on the size of the node for which motion vector estimation is applied, and MinPUSize defines the lower bound.
[0091] G-PCC encoder 200 can generate a motion matrix and a translation vector that force G-PCC encoder 200 to "match" detected feature points between a predicted (reference) frame and a current frame, respectively. The detected feature points can be unreliable, which can result in incorrect global motion parameter estimation, introducing additional distortion instead of compensating for the global motion. This additional distortion, in turn, can result in lower coding efficiency than if the global motion were properly compensated.
[0092] In some systems, detected feature points in consecutive frames can not match. For example, the relative rotation and / or translation between individual feature points in consecutive frames can be different. If a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) forces individual feature points in consecutive frames to match each other, distortion can occur.
[0093] A problem with the current motion estimation structure is run-time efficiency. The total run-time can be 25 times that of the anchor version used as a reference. Of the additional total run-time, approximately half is typically spent on global motion estimation and approximately half is typically spent on local node motion estimation. Such high run-time can be impractical for certain applications (e.g., real-time coding of point cloud compression) and thereby hinder motion estimation in such applications.
[0094] According to the techniques of this disclosure, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can be configured to apply global motion compensation based on global positioning system information. According to the techniques of this disclosure, a G-PCC coder can be configured to apply motion compensation using one or more of the following techniques.
[0095] 1) G-PCC encoder 200 can identify a first set of global motion parameters from GPS (global positioning system) information. As used herein, global positioning system (also referred to herein simply as “GPS”) can refer to any satellite system such as, for example, the Global Positioning System (GPS) implemented in the United States, the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), Galileo, Quasi-Zenith Satellite System (QZSS), Indian Regional Navigational Satellite System (IRNSS), or other satellite systems.
[0096] a. The first set of global motion parameters can include a set of orientation parameters, e.g., roll, pitch, yaw, angular velocity, etc. That is, G-PCC encoder 200 can identify the first set of global motion parameters to include a set of orientation parameters. In some examples, the set of orientation parameters can include one or more of a roll, a pitch, a yaw, or an angular velocity of the current frame.
[0097] b. In some examples, three parameters of orientation can be specified based on a reference frame. Any such three parameters can be included in the first set of global motion parameters. For example, given a roll, a pitch, and a yaw of a last frame, the orientation parameters can include a roll difference between the current frame and the last frame, a pitch difference between the current frame and the last frame, and a yaw difference between the current frame and the last frame. That is, G-PCC encoder 200 can identify the set of orientation parameters to include one or more of a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, or an angular velocity difference between the current frame and the reference frame.
[0098] c.The first set of global motion parameters can include a set of positional parameters, e.g., x, y, z dimensional displacement or velocity (e.g., velocity-x, velocity-y, velocity-z). That is, the G-PCC encoder 200 can identify the set of orientation parameters to include a set of positional parameters. In some examples, the set of positional parameters can include one or more of a displacement of the current frame or an average velocity of the current frame.
[0099] d.In some examples, three parameters of a position can be specified based on the reference frame. That is, the G-PCC encoder 200 can identify the set of orientation parameters to include the three parameters of a position based on the reference frame. Any such three parameters can be included in the first set of global motion parameters. For example, the positional parameters can be computed using an East-North-Up coordinate system. That is, the set of positional parameters can include one or more of an East velocity, a North velocity, or an Up velocity of the East-North-Up coordinate system.
[0100] 2) The G-PCC coder (e.g., the G-PCC encoder 200 or the G-PCC decoder 300) can derive a second set of global motion parameters from the first set of global motion parameters for global motion estimation. That is, the G-PCC coder can determine, based on the first set of global motion parameters, a second set of global motion parameters to use for global motion estimation of the current frame.
[0101] a.The second set of global motion parameters can include elements of a global motion matrix, which can represent (e.g., describe) a rotation matrix and a translation vector. That is, the G-PCC coder (e.g., the G-PCC encoder 200 or the G-PCC decoder 300) can derive the second set of global motion parameters from the first set of global motion parameters to include a rotation matrix that indicates a yaw, a pitch, and a roll of the current frame and a translation vector that indicates an average velocity of the current frame.
[0102] b.The second set of global motion parameters can be computed exactly or approximately from the first set of global motion parameters.
[0103] 3) The G-PCC coder (e.g., the G-PCC encoder 200 or the G-PCC decoder 300) can use the second set of global motion parameters to apply motion compensation to the reference frame, resulting in a compensated frame. That is, the G-PCC coder can use the second set of global motion parameters to apply motion compensation to the reference frame to generate a globally motion compensated frame of the current frame.
[0104] a.The G-PCC coder can use the compensated frame as a reference for motion estimation of the current frame.
[0105] b. In some examples, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can apply the compensation by first applying the rotation and then the translation.
[0106] c. In some examples, a G-PCC coder can apply the compensation by first applying the translation and then the rotation.
[0107] 4) G-PCC encoder 200 can signal the second set of global motion parameters in the bitstream. G-PCC encoder 200 can use the second set of global motion parameters to enable G-PCC decoder 300 to estimate the global motion and apply the prediction or motion compensation. That is, G-PCC encoder 200 can signal the second set of global motion parameters in the bitstream. For example, G-PCC encoder 200 can signal a rotation matrix indicating the yaw, pitch, and roll of the current frame and a translation vector indicating the average velocity of the current frame. Similarly, G-PCC decoder 300 can decode from the bitstream the rotation matrix indicating the yaw, pitch, and roll of the current frame and the translation vector indicating the average velocity of the current frame.
[0108] a. In some examples, G-PCC encoder 200 can signal the first set of global motion parameters in the bitstream. In this example, G-PCC decoder 300 can derive the second set of global motion parameters. That is, G-PCC encoder 200 can signal the first set of global motion parameters in the bitstream. For example, G-PCC encoder 200 can signal one or more of a set of orientation parameters or a set of translation parameters identified from global positioning system information. Similarly, G-PCC decoder 300 can decode from the bitstream one or more of a set of orientation parameters or a set of translation parameters identified from global positioning system information. As described below, G-PCC encoder 200 can encode a portion of the first set of global motion parameters and / or the second set of global motion parameters and / or the estimation.
[0109] 5) A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can find an appropriate combination of global motion parameters for local node motion vector estimation to try to find a balance tradeoff between run-time and performance. That is, the G-PCC coder can limit the block size of local motion to be equal to the minimum prediction unit size.
[0110] a. A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can limit the block size of local motion to be equal to minPuSize. This can help to ensure that the run-time of local motion vector estimation is minimized while having a limited impact on coding efficiency.
[0111] 6) A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can improve the global motion estimation algorithm by one or more of the following steps:
[0112] a. First, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can estimate an initial translation vector by minimizing the mean squared error between the current frame and the reference frame. That is, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can estimate an initial translation vector by minimizing the mean squared error between the current frame and the reference frame. When estimating the initial translation vector, the G-PCC coder can take into account the label of whether a point is ground. That is, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can determine whether a point is ground. In this example, the G-PCC coder can estimate a rotation matrix for the second current frame based on whether the point is ground.
[0113] b. A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can combine the initial translation vector with an identity matrix and can feed the combined initial translation vector and identity matrix into an iterative closest point scheme or similar scheme to estimate a rotation matrix and a translation vector.
[0114] c. In some examples, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can first estimate a rotation matrix based on a label of whether a point is ground. That is, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can determine whether a point is ground. In this example, the G-PCC coder can estimate a rotation matrix for the second current frame based on whether the point is ground. For example, G-PCC encoder 200 can derive and signal the label to G-PCC decoder 300. That is, G-PCC encoder 200 can signal a set of labels that indicate whether a point is ground. In some examples, G-PCC encoder 200 and G-PCC decoder 300 can each derive the labels. The G-PCC coder can derive the labels based on a ground estimation algorithm; such an algorithm can be based on the height of the point, the density of the point cloud near the point, the relative distance of the point to the LIDAR origin / fix point, etc.
[0115] d. A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can feed an initial rotation matrix with a zero translation vector into an iterative closest point scheme or similar scheme to estimate a rotation matrix and a translation vector.
[0116] In this example, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can derive global motion parameters from GPS information. The GPS information can include values / parameters that can be used to derive roll-pitch-yaw information and velocity for each timestamp (or point of acquisition, or estimate for each point of acquisition) in an East-North-Up coordinate system. For example, the G-PCC coder can directly compute a global motion matrix and translation vector from these provided information items.
[0117] The rotation matrix can define the change of axes from the reference frame to the current frame. Given the roll ref ,pitch ref ,yaw ref ) and the roll, pitch, and yaw of the current frame (roll cur ,pitch cur ,yaw cur ), the G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can use the differences in roll, pitch, and yaw to make the derivation. The G-PCC coder can compute the differences as Δ roll = roll cur - roll ref , Δ pitch = pitch cur - pitch ref , and Δ yaw = yaw cue - yaw ref .
[0118] The rotation matrix for roll is
[0119]
[0120] The rotation matrix for pitch is
[0121]
[0122] The rotation matrix for yaw is
[0123]
[0124] The final rotation matrix is
[0125] R = R yaw R pitch R roll
[0126] In some examples, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can be configured to determine whether a positive direction of one or more of roll, pitch, and yaw has occurred. The G-PCC coder can change the sign of the angle accordingly, as the positive direction can be defined differently.
[0127] For the translation vector, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can use East-North-Up (ENU) velocity, as the ENU coordinate system is well aligned with the coordinate system of the point cloud frame. The G-PCC coder can compute the average velocity of the reference frame and the current frame as V = (v East ,v North ,v Up ). The G-PCC coder can decompose the velocity vector into coordinates in the coordinate system of the reference frame.
[0128] A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can compute a new rotation matrix for the velocity as
[0129] The rotation matrix for the roll velocity is
[0130]
[0131] The rotation matrix for the pitch velocity is
[0132]
[0133] The rotation matrix for the yaw velocity is
[0134]
[0135] The final rotation matrix for the velocity is
[0136]
[0137] A G-PCC can compute the translation vector as
[0138] T = tR v V
[0139] Here, t is the time of one frame. In addition, the positive direction should be considered to derive the correct translation vector.
[0140] In this example, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can signal only the roll-pitch-yaw and translation vectors, instead of all 12 elements for the global motion. This would reduce the number of global motion parameters from 12 to 6. That is, G-PCC encoder 200 can signal a rotation matrix that indicates the yaw, pitch, and roll of the current frame, which includes 9 elements, and a translation vector that includes 3 elements that indicate the average velocity of the current frame (e.g., a total of 12 elements). In some examples, G-PCC encoder 200 can signal the roll of the current frame, the pitch of the current frame, the yaw of the current frame, and a translation vector that indicates the average velocity of the current frame (e.g., a total of 6 elements) without signaling a complete global rotation matrix.
[0141] A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can determine a rotation matrix and a translation vector. The G-PCC coder can use the roll, pitch, yaw (Δ roll , Δ pitch , Δ yaw ) and the difference of the main parameters of the translation vector. Instead of compressing / signaling the entire rotation matrix, G-PCC encoder 200 can signal only the delta and the translation vector, which would reduce the number of parameters signaled from 12 (9 for the rotation matrix, 3 for the translation vector) to 6 (3 for the delta values, 3 for the translation vector). That is, G-PCC encoder 200 can signal the roll difference between the current frame and a reference frame, the pitch difference between the current frame and the reference frame, the yaw difference between the current frame and the reference frame, and a translation vector that indicates the average velocity of the current frame (e.g., 6 elements). Similarly, G-PCC decoder 300 can decode the roll difference between the current frame and a reference frame, the pitch difference between the current frame and the reference frame, the yaw difference between the current frame and the reference frame, and a translation vector that indicates the average velocity of the current frame (e.g., 6 elements).
[0142] Furthermore, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can reduce the number of global motion parameters by approximation. The elements in the rotation matrix are close to 1 and 0, so the G-PCC coder can approximate the rotation matrix as an identity matrix. In this way, for the translation vector, 6 parameters can be further reduced to only 3. That is, G-PCC encoder 200 can signal a translation vector (e.g., 3 elements) that indicates the average velocity of the current frame, and avoid signaling a rotation matrix that indicates the yaw, pitch, and roll of the current frame. Similarly, G-PCC decoder 300 can decode a translation vector (e.g., 3 elements) that indicates the average velocity of the current frame, and avoid decoding a rotation matrix that indicates the yaw, pitch, and roll of the current frame.
[0143] Assuming that the vehicle (e.g., equipped with LIDAR) will move forward most of the time, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can approximate the translation vector as T appr = [0, abs(T), 0], where abs() is an absolute function that calculates the magnitude of the translation vector T. Then, the total number of global motion parameters can be reduced to only 1. In this example, the system is defined such that the vehicle moves in the positive y direction, but similar derivations apply to other systems as well. That is, G-PCC encoder 200 can signal the magnitude of the translation vector that indicates the average velocity of the current frame. Similarly, G-PCC decoder 300 can decode the translation vector that indicates the magnitude of the translation vector that indicates the average velocity of the current frame.
[0144] Try different combinations of the input parameters BlockSize and MinPUSize to find a balanced trade-off between run-time and performance.
[0145] A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can determine a BlockSize that is twice as large as MinPUSize to achieve a balanced trade-off between performance gain and run-time.
[0146] In some examples, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can set BlockSize to be equal to MinPUSize to achieve a balanced trade-off between performance gain and run-time. That is, the G-PCC coder can perform motion vector estimation for the current frame based on global motion compensation, where, to perform the motion vector estimation, the G-PCC coder is configured to limit the block size of local motion to be equal to the minimum prediction unit size.
[0147] The global motion estimation algorithm can be improved by configuring the G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) to first estimate the translation vector by minimizing the mean squared error (MSE) between the current frame and the predicted (reference) frame. That is, the G-PCC coder can estimate the initial translation vector by minimizing the mean squared error between the second current frame and the second reference frame. Then, after applying the estimated translation vector, the G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can determine the rotation matrix.
[0148] The new global motion estimation algorithm can have two steps. The first step is to configure the G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) to compute the initial translation vector T’. The second step is to have the G-PCC coder feed T’ and the identity matrix into the iterative closest point algorithm (e.g., provided by the open3d library) or a similar alternative.
[0149] One step is to configure the G-PCC coder to estimate T’. Assume T’ = [a, b, c]. The translation vector is assumed to minimize the MSE between the current frame and the predicted frame. Here, the G-PCC can represent the MSE by the following loss function:
[0150]
[0151] Here, the points are from the reference frame, and the points are from the current frame. w i is the weight function, which is defined as the distance to the center: The variable N represents the total number of points.
[0152] To minimize L, we set
[0153]
[0154] The computed a will be:
[0155]
[0156] b and c can also be derived:
[0157]
[0158] However, computing a, b, and c only once can not be accurate enough, as the motion between frames is always large. The G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can define the number of loops, / . For the first iteration, given the reference frame f0and the current frame, the G-PCC coder can compute T o′ The G-PCC decoder can then apply T1 to f0 to obtain a new reference frame f1. The G-PCC decoder can compute T1 between f1 and the current frame o ′ to obtain a new reference frame f1. The G-PCC decoder can compute T1 between f1 and the current frame ′ The G-PCC can perform this loop l times to obtain our final translation vector
[0159] Another modification targets the weight function. Ground points can "confuse" the algorithm used by a G-PCC decoder (e.g., G-PCC encoder 200 or G-PCC decoder 300) to estimate the global motion, so the G-PCC decoder can "turn off" the ground points. The weight function can be modified by:
[0160]
[0161] A G-PCC decoder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can determine whether a point is a ground point or otherwise based on one or more characteristics of the point, such as height, distance from the center of the point cloud frame, distance from the point in the frame along a certain plane (e.g., the x-y plane), etc.
[0162] In addition to the position of a point of an object / scene / point cloud relative to a local reference, some applications use positions relative to a fixed location on Earth. There are several coordinate systems used to describe the geographical positioning of a point. A few of the coordinate systems used in several applications are briefly introduced below.
[0163] One such system is the geodetic system. The geodetic system uses a set of three values - latitude, longitude, and altitude.
[0164] Figure 7 is a conceptual diagram showing the geodetic latitude and longitude of a point measured based on an ellipsoidal approximation of the Earth. Figure 7 It is shown how to obtain the geodetic latitude of a point as the angle that the normal to the Earth's surface (ellipsoidal approximation) at that point makes with the equatorial plane. The angle φ represents the north-south positioning of the point relative to the Earth. The longitude is measured as the angle λ (in degrees) between the prime meridian (chosen as a point in Greenwich, England) and the positive angle measured eastward and the negative angle measured westward. The altitude is defined as the distance measured in units in the direction normal to the ellipsoid surface above the ellipsoid.
[0165] Another coordinate system is the Earth-Centered Earth-Fixed (ECEF) coordinate system, which is now discussed. In this system, the center of the Earth is chosen as the reference point, and the positioning of this point on the Earth (or anywhere in space, typically close to the Earth's surface) is described as a displacement in the x, y, and z dimensions from this origin point. The positive z-axis is chosen as the straight line from the origin to the North Pole. The positive x-axis is chosen as the straight line connecting the center to the point on the Earth's surface with latitude 0 degrees and longitude 0 degrees.
[0166] Figure 8 is a conceptual diagram showing the ECEF (Earth-Centered Earth-Fixed) coordinate system (X, Y, Z axes) relative to the equator and the prime meridian (0 degrees latitude and longitude). Figure 8 A diagram showing the ECEF system relative to latitude and longitude is shown in
[0167] A local tangent system (ENU, NED) is now discussed. A local tangent system specifies a local tangent plane on the Earth's surface and describes the positioning of a point in this plane with East, North, and Up (ENU) displacements. An equivalent system can be described using North, East, and Down (NED) displacements. Different systems can be used in different applications. In some examples, the displacements can be described in meters.
[0168] In addition to the ENU / NED displacements, the system uses a reference from which the displacements are measured. The ENU / NED reference can be described by the ECEF or geodetic system or another coordinate system. One advantage of the ENU / NED system can be that the relative displacement values are typically smaller than the absolute positioning in ECEF coordinates.
[0169] The orientation of a point cloud is now discussed. The orientation of objects in a scene can also be important for applications that can wish to visualize a point cloud using information from the surrounding scene. To this end, it must be known which one or more of the x, y, and z axes of the point cloud frame are oriented relative to a fixed reference. This can be described as a rotation of the fixed reference xyz axes relative to the axes used by the point cloud frame. The rotation can be described as a matrix or as a triple of roll-pitch-yaw angles.
[0170] In some examples of G-PCC, the G-PCC encoder 200 and the G-PCC decoder 300 can code the x, y, z positioning of points in a point cloud relative to a slice origin, and the slice origin can in turn be coded relative to an origin specified in a sequence parameter set (SPS), which can be signaled by the G-PCC encoder 200; this (referred to as the SPS origin) is the actual origin of the point cloud frame. The current signaling also includes a scale flag, which indicates whether the coordinate values obtained after applying a non-standard scaling operation on the decoder side have units of meters.
[0171] However, the absolute positioning of the SPS origin is not specified in current G-PCC examples. In some applications (e.g., geospatial data visualization), it can be useful to convey the actual positioning of the frame origin to indicate where to acquire the point cloud or the relationship of the point cloud to geospatial objects. For example, an application can wish to present point cloud information in addition to other attributes acquired at a particular location at the same time. Without the positioning of the SPS origin, G-PCC decoder 300 or the application can have to perform expensive registration and classification algorithms to identify the location of the SPS origin.
[0172] In some systems, orientation information can also be important in addition to the positioning of the point cloud origin. To enable G-PCC decoder 300 to correctly present the point cloud, G-PCC encoder 200 informs G-PCC decoder 300 of the orientation of the x, y, z axes. G-PCC currently does not support the notification of the orientation of the x, y, z axes. In a more specific example, a LIDAR system on a car captures point clouds about the direction of motion of the car (e.g., the vehicle can assume that the direction of motion of the vehicle is the positive y direction), which can change at every frame. The orientation information (as well as the SPS origin) can be frame-specific.
[0173] One or more of the following techniques can be applied independently or in combination.
[0174] Geographic position / GIS projection is now discussed. In some examples, G-PCC encoder 200 can signal the positioning of the SPS origin relative to a fixed reference. G-PCC decoder 300 can parse the signaled positioning. More generally, the positioning of an origin associated with a point cloud frame can be signaled.
[0175] In some examples, G-PCC encoder 200 can use offsets in respective dimensions to specify the positioning of the SPS origin relative to a fixed origin and coordinate axes. G-PCC decoder 300 can parse the specified positioning. In one example, an ECEF system can be used to describe the SPS origin. In one example, an ENU system can be used to describe the SPS origin. In one example, a geodetic system can be used to describe the SPS origin. More generally, any positioning system can be used.
[0176] In some examples, G-PCC encoder 200 can signal a syntax element that indicates a coordinate system used to describe the positioning of the SPS origin. G-PCC decoder 300 can parse the syntax element to determine the coordinate system.
[0177] In some examples, G-PCC encoder 200 can signal one or more syntax elements that indicate a number of bits used to code a position of a SPS origin in an indicated coordinate system. G-PCC decoder 300 can parse the syntax elements to determine the number of bits.
[0178] Now discuss orientation. In some examples, G-PCC encoder 200 can use parameters that describe an x, y, z axis orientation. For example, the parameters can describe a rotation from a fixed axis system (e.g., ECEF XYZ axes). In some examples, the parameters can describe a rotation matrix. In some examples, the parameters can describe roll, pitch, and yaw.
[0179] In some examples, G-PCC encoder 200 can signal a syntax element that is used to indicate a coordinate system used to describe an orientation of a point cloud frame axis. G-PCC decoder 300 can parse the syntax element to determine the coordinate system.
[0180] In some examples, G-PCC encoder 200 can signal one or more syntax elements that indicate a number of bits used to code parameters to describe an orientation of a point cloud frame. G-PCC decoder 300 can parse the syntax elements to determine the number of bits.
[0181] Now discuss restrictions on position and orientation. G-PCC encoder 200 can apply certain restrictions based on certain constraints that apply due to the nature of the application system being used. For example, for LIDAR data captured by a car, G-PCC encoder 200 can restrict the position information by only providing latitude and longitude without specifying altitude. In other examples, for LIDAR captured data, G-PCC encoder 200 can only provide a velocity of the vehicle in ENU directions to enable G-PCC decoder 300 to derive a position, and similarly provide angular velocity to specify an orientation over multiple frames.
[0182] In some examples, other parameters associated with a point cloud can also be signaled, such as angular velocity, angular acceleration, linear velocity, linear acceleration, time associated with capture of a point cloud frame, etc.
[0183] Now discuss frame based signaling vs sequence based signaling. In some examples, a position and / or orientation of a point cloud origin and axis orientation are fixed for an entire sequence or bitstream. In this case, for example, G-PCC encoder 200 can signal the position of the point cloud origin and axis orientation only once per sequence (e.g., in a parameter set such as an SPS). G-PCC decoder 300 can parse the syntax elements in the parameter set to determine the position and orientation.
[0184] In some examples, the position and / or orientation of the point cloud origin can vary from frame to frame, and the G-PCC encoder 200 can signal the values for one or more frames. The G-PCC decoder 300 can parse the signaled values to determine the position and orientation.
[0185] The following example shows how to describe the position and orientation of the point cloud origin on a per-frame basis.
[0186]
[0187] pcoo_update_flag equal to 0 specifies that the syntax structure contains at least one of the position of the SPS origin and the orientation of the axes of the current frame with respect to some fixed system, as indicated by pcoo_origin_coordinate_system_id and pcoo_origin_orientation_system_id, respectively. pcoo_update_flag equal to 1 specifies that the syntax structure contains the position of the SPS origin and the orientation of the axes of the current frame with respect to some fixed system, as indicated by pcoo_origin_coordinate_system_id and pcoo_origin_orientation_system_id, respectively, where pcoo_origin_coordinate_val[ ] and pcoo_origin_orientation_val[ ] indicate delta-coded values from a previous frame with pcoo_update_flag equal to 0.
[0188] In some examples, a restriction can be added such that pcoo_update_flag can be 0 only for pictures that cannot be removed from the bitstream (e.g., the first frame of a sequence or the frames associated with the lowest frame rate).
[0189] In one example, an ID is specified for each syntax structure, and a reference ID referring to the position and orientation is signaled when pcoo_update_flag is equal to 1.
[0190] In some examples, the reference position and orientation is chosen to be the previous point cloud frame with the same or lower temporal ID as the current frame.
[0191] Alternatively, a frame index is signaled when pcoo_update_flag is equal to 1, and the delta-coded position and orientation are measured with respect to the position and orientation of that frame.
[0192] In one example, the delta coding is applied only to the position values and not to the orientation values.
[0193] pcoo_origin_info_present_flag equal to 1 specifies that the absolute positioning information of the SPS origin is signaled. pcoo_origin_info_present_flag equal to 0 specifies that the absolute positioning information of the point cloud frame is not signaled.
[0194] In this case, the default absolute can be selected or sent by external means.
[0195] pcoo_origin_coordinate_system_id specifies the coordinate system used to describe the absolute position of the SPS origin of the current frame. For bitstreams conforming to this version of this Specification, the value of pcoo_origin_coordinate_system_id shall be in the range of 0 to 2. Other values are reserved for future use by ISO / IEC.
[0196] pcoo_origin_coordinate_num_params specifies the number of parameters signaled for the position in the indicated coordinate system.
[0197] In some alternatives, the value of pcoo_origin_coordinate_num_params can be fixed and can be predetermined based on the value of pcoo_origin_coordinate_system_id. In some alternatives, this syntax element can be coded as a_minusN, where a value less than N is not allowed to be signaled.
[0198] pcoo_origin_coordinate_num_bits is used to specify the number of bits used to signal pcoo_origin_coordinate_val[i].
[0199] For i = 0, pcoo_origin_coordinate_val[i]. pcoo_origin_coordinate_num_params - 1 is used to derive the position of the SPS origin of the current frame.
[0200] The interpretation of pcoo_origin_coordinate_val[] is given by the following table:
[0201]
[0202]
[0203] Note that the above precisions and the precisions in the rest of this disclosure are just examples, and the techniques of this disclosure are applicable to any precision in meters, degrees, or other units.
[0204] In one example, for some systems, the number of parameters for the same system can vary based on the value of update_flag. For example, when update_flag is equal to 0, an ENU system can have six points, while when update_fag is equal to 1, the ENU system can have only three points corresponding to local displacement.
[0205] In another example, an option to signal a local reference can be allowed. When a syntax element indicating the presence of a local reference is signaled (e.g., local reference_present_flag), the local reference can be signaled.
[0206] pcoo_orientation_info_present_flag equal to 1 specifies that the orientation of the XYZ axes of the point cloud frame is signaled. pcoo_orientation_info_present_flag equal to 0 specifies that the orientation of the XYZ axes of the point cloud frame is not signaled.
[0207] pcoo_orientation_system_id specifies the orientation system used to describe the orientation of the XYZ axes of the point cloud frame. For bitstreams conforming to this version of this Specification, the value of pcoo_orientation_system_id shall be in the range of 0 to 1. Other values are reserved for future use by ISO / IEC.
[0208] pcoo_orientation_coordinate_num_params specifies the number of parameters signaled for the orientation in the indicated coordinate system.
[0209] In some examples, the value of pcoo_orientation_coordinate_num_params can be fixed and pre-determined based on the value of pcoo_orientation_coordinate_system_id. In some examples, this syntax element can be coded as a_minusN, where signaling of values less than N is not allowed.
[0210] pcoo_orientation_coordinate_num_bits specifies the number of bits used to signal pcoo_orientation_coordinate_val[i].
[0211] For i = 1, pcoo_orientation_coordinate_val[i]. pcoo_orientation_coordinate_num_params is used to determine the orientation of the XYZ axes of the point cloud frame.
[0212] The interpretation of pcoo_orientation_coordinate_val can be obtained from the following table:
[0213]
[0214] In one example, in the syntax structure, at least one of pcoo_orientation_info_present_flag and pcoo_origin_info_present_flag is constrained to be equal to 1.
[0215] In some examples, the parameters of indices 0 to 8 correspond to the row- scanned elements of the rotation matrix (i.e., the i-th index corresponds to the (i / 3)-th row and (i%2)-th column). More generally, any scan pattern of the rotation matrix elements can be chosen to obtain the pcoo_orientation_coordinate_val[i] parameters.
[0216] Figure 9 is a flowchart illustrating an example encoding process in accordance with the techniques of this disclosure. A G-PCC encoder 200 (e.g., octree analysis unit 210) can identify a first set of global motion parameters from global positioning system information (902). For example, the G-PCC encoder 200 can identify a set of orientation parameters and / or a set of position parameters for a current frame from global positioning system information.
[0217] The G-PCC encoder 200 (e.g., octree analysis unit 210) can determine a second set of global motion parameters to be used for global motion estimation of the current frame based on the first set of global motion parameters (904). For example, the G-PCC encoder 200 can determine a rotation matrix indicating a yaw, a pitch, and a roll of the current frame and a translation vector indicating an average velocity of the current frame.
[0218] The G-PCC encoder 200 (e.g., the octree analysis unit 210) can apply motion compensation to the reference frame based on the second set of global motion parameters to generate a globally motion compensated frame for the current frame (906). For example, the G-PCC encoder 200 (e.g., the octree analysis unit 210) can apply global motion compensation to all points in the reference frame (e.g., the prediction frame) based on the second set of global motion parameters (e.g., the rotation matrix and the translation vector). In this way, the G-PCC encoder 200 can apply global motion compensation using satellite information, which can be more accurate than estimating a rotation matrix and a translation vector for the current frame based on feature points between the reference frame (e.g., the prediction frame) and the current frame.
[0219] The G-PCC encoder 200 (e.g., the octree analysis unit 210) can apply local node motion estimation based on the globally motion compensated frame to generate motion vector information and point information for the current frame (908). For example, the G-PCC encoder 200 (e.g., the octree analysis unit 210) can apply local node motion estimation. For example, the G-PCC encoder 200 (e.g., the octree analysis unit 210) can apply a brute force search within the current node. In this example, the G-PCC encoder 200 (e.g., the octree analysis unit 210) can sample a portion of the points in the current node. The G-PCC encoder 200 (e.g., the octree analysis unit 210) can generate compensated points given an initial motion vector and compute a cost between these points and their counterparts in the reference node. The G-PCC encoder 200 (e.g., the octree analysis unit 210) can select the motion vector that minimizes the cost as the final estimated local node motion vector.
[0220] The G-PCC encoder 200 (e.g., the arithmetic encoding unit 214) can signal the global motion information, the motion vector information, and the point information for the current frame (910). For example, the G-PCC encoder 200 (e.g., the arithmetic encoding unit 214) can signal a first set of global motion parameters (e.g., a set of orientation parameters and / or a set of position parameters). In some examples, the G-PCC encoder 200 (e.g., the arithmetic encoding unit 214) can signal a second set of global motion parameters (e.g., a rotation matrix and a translation vector). The G-PCC encoder 200 (e.g., the arithmetic encoding unit 214) can signal a roll difference between the current frame and the reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector (e.g., 6 elements) that indicates an average velocity of the current frame. In some examples, the G-PCC encoder 200 (e.g., the arithmetic encoding unit 214) can signal a translation vector (e.g., 3 elements) that indicates an average velocity of the current frame. The G-PCC encoder 200 (e.g., the arithmetic encoding unit 214) can signal an approximation of the translation vector as T appr = [0, abs(T), 0], where abs() is an absolute function that calculates the magnitude of the translation vector T.
[0221] Figure 10 is a flowchart illustrating an example decoding process according to the techniques of this disclosure. The G-PCC encoder 200 (e.g., the geometric arithmetic decoding unit 302) can decode the symbols of the bitstream that indicate the global motion information, the motion vector information, and the point information for the current frame (1002).
[0222] For example, the G-PCC decoder 300 (e.g., the geometric arithmetic decoding unit 302) can decode a first set of global motion parameters and / or a second set of global motion parameters from the geometry bitstream 203. The G-PCC decoder 300 (e.g., the geometric arithmetic decoding unit 302) can decode a roll difference between the current frame and the reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector that indicates an average velocity of the current frame. In some examples, the G-PCC decoder 300 (e.g., the geometric arithmetic decoding unit 302) can decode a translation vector that indicates an average velocity of the current frame. In some examples, the G-PCC decoder 300 (e.g., the geometric arithmetic decoding unit 302) can decode a magnitude of a translation vector that indicates an average velocity of the current frame.
[0223] G-PCC decoder 300 (e.g., octree synthesis unit 306) can determine a set of global motion parameters to use for global motion estimation of the current frame based on the global motion information (1004). For example, G-PCC decoder 300 (e.g., octree synthesis unit 306) can determine a rotation matrix and a translation vector based on a roll difference, a pitch difference, a yaw difference, and a translation vector. G-PCC decoder 300 (e.g., octree synthesis unit 306) can decode the translation vector from geometry bitstream 203 and approximate the rotation matrix as an identity matrix. In some examples, G-PCC decoder 300 (e.g., octree synthesis unit 306) can decode a magnitude of the translation vector from geometry bitstream 203 that indicates an average velocity of the current frame. In this example, G-PCC decoder 300 (e.g., octree synthesis unit 306) can approximate the translation vector based on the received magnitude and approximate the rotation matrix as an identity matrix.
[0224] G-PCC decoder 300 (e.g., octree synthesis unit 306) can apply global motion compensation to the reference frame based on the set of sets of global motion parameters to generate a global motion compensated frame of the current frame (1006). In this way, G-PCC decoder 300 (e.g., octree synthesis unit 306) can apply global motion compensation using satellite information, which can be more accurate than estimating a rotation matrix and a translation vector of the current frame based on feature points between the reference frame (e.g., predicted frame) and the current frame.
[0225] G-PCC decoder 300 (e.g., octree synthesis unit 306) can apply local node motion estimation based on the global motion compensated frame to generate the current frame (1008). G-PCC decoder 300 can output the current frame (1012). For example, G-PCC decoder 300 can cause a display to output the current frame.
[0226] Examples in various aspects of the disclosure can be used alone or in any combination.
[0227] Clause Al. A device for encoding point cloud data, the device comprising: a memory to store point cloud data; and one or more processors coupled to the memory and implemented in circuitry, the one or more processors configured to: identify a first set of global motion parameters from global positioning system information; determine a second set of global motion parameters to use for global motion estimation of a current frame based on the first set of global motion parameters; and apply motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame of the current frame.
[0228] Clause A2. The device of clause Al, wherein the one or more processors are configured to signal the second set of global motion parameters in a bitstream.
[0229] Clause A3. The device of clause A1, wherein the one or more processors are configured to signal the first set of global motion parameters in a bitstream.
[0230] Clause A4. The device of clause A1, wherein the first set of global motion parameters comprises a set of orientation parameters.
[0231] Clause A5. The device of clause A4, wherein the set of orientation parameters comprises one or more of a roll, a pitch, a yaw, or an angular velocity of the current frame.
[0232] Clause A6. The device of clause A4, wherein the set of orientation parameters comprises one or more of a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, or an angular velocity difference between the current frame and the reference frame.
[0233] Clause A7. The device of clause A1, wherein the first set of global motion parameters comprises a set of position parameters.
[0234] Clause A8. The device of clause A7, wherein the set of position parameters comprises one or more of a displacement of the current frame or an average velocity of the current frame.
[0235] Clause A9. The device of clause A7, wherein the set of position parameters comprises one or more of an east velocity, a north velocity, or an up velocity of an east-north-up coordinate system.
[0236] Clause A10. The device of clause A1, wherein the second set of global motion parameters comprises a rotation matrix indicating a yaw, a pitch, and a roll of the current frame and a translation vector indicating an average velocity of the current frame.
[0237] Clause A11. The device of clause A1, wherein the one or more processors are configured to signal a roll of the current frame, a pitch of the current frame, a yaw of the current frame, and a translation vector indicating an average velocity of the current frame.
[0238] Clause A12. The device of clause A1, wherein the one or more processors are configured to signal a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicating an average velocity of the current frame.
[0239] Clause A13. The device of clause A1, wherein the one or more processors are configured to signal a translation vector indicating an average velocity of the current frame and refrain from signaling a rotation matrix indicating a yaw, a pitch, and a roll of the current frame.
[0240] Clause A14. The device of clause A1, wherein the one or more processors are configured to signal a magnitude of a translation vector that indicates an average velocity of the current frame.
[0241] Clause A15. The device of clause A1, wherein the one or more processors are configured to perform motion vector estimation for the current frame based on the global motion compensated frame, and wherein, to perform the motion vector estimation, the one or more processors are configured to limit a block size of local motion to be equal to a minimum prediction unit size.
[0242] Clause A16. The device of clause A1, wherein the reference frame is a first reference frame and the current frame is a first current frame, and wherein the one or more processors are configured to estimate an initial translation vector by minimizing a mean squared error between a second current frame and a second reference frame.
[0243] Clause A17. The device of clause A1, wherein the current frame is a first current frame, and wherein the one or more processors are configured to: determine whether a point is ground; and estimate a rotation matrix for a second current frame based on whether the point is ground.
[0244] Clause A18. The device of clause A17, wherein the one or more processors are configured to signal a set of labels that indicate whether the point is ground.
[0245] Clause A19. The device of clause A1, wherein the one or more processors are further configured to generate point cloud data.
[0246] Clause A20. The device of clause A19, wherein the one or more processors are configured to generate the point cloud data based on signals from a LIDAR device as part of generating the point cloud data.
[0247] Clause A21. The device of clause A1, wherein the device is one of a mobile phone, a tablet, a vehicle, or an extended reality device.
[0248] Clause A22. The device of clause A1, wherein the device includes an interface configured to transmit the encoded point cloud data.
[0249] Clause A23. A method for encoding point cloud data, the method comprising: identifying, with one or more processors, a first set of global motion parameters from global positioning system information; determining, with the one or more processors, a second set of global motion parameters to be used for global motion estimation of a current frame based on the first set of global motion parameters; and applying, with the one or more processors, motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
[0250] Clause A24. The method of clause A23, further comprising signaling, with the one or more processors, the second set of global motion parameters in a bitstream.
[0251] Clause A25. The method of clause A23, further comprising signaling, with the one or more processors, the first set of global motion parameters in a bitstream.
[0252] Clause A26. The method of clause A23, wherein the first set of global motion parameters comprises a set of orientation parameters.
[0253] Clause A27. The method of clause A26, wherein the set of orientation parameters comprises one or more of a roll, a pitch, a yaw, or an angular velocity of the current frame.
[0254] Clause A28. The method of clause A26, wherein the set of orientation parameters comprises one or more of a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, or an angular velocity difference between the current frame and the reference frame.
[0255] Clause A29. The method of clause A23, wherein the first set of global motion parameters comprises a set of position parameters.
[0256] Clause A30. The method of clause A29, wherein the set of position parameters comprises one or more of a displacement of the current frame or an average velocity of the current frame.
[0257] Clause A31. The method of clause A29, wherein the set of position parameters comprises one or more of an east velocity, a north velocity, or an up velocity of an east-north-up coordinate system.
[0258] Clause A32. The method of clause A23, wherein the second set of global motion parameters comprises a rotation matrix indicating a yaw, a pitch, and a roll of the current frame and a translation vector indicating an average velocity of the current frame.
[0259] Clause A33. The method of clause A23, further comprising signaling a roll of the current frame, a pitch of the current frame, a yaw of the current frame, and a translation vector indicating an average velocity of the current frame.
[0260] Clause A34. The method of clause A23, further comprising signaling a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicating an average velocity of the current frame.
[0261] Clause A35. The method of clause A23, further comprising signaling, with the one or more processors, a translation vector indicative of an average velocity of the current frame, and refraining from signaling a rotation matrix indicative of a yaw, pitch, and roll of the current frame.
[0262] Clause A36. The method of clause A23, further comprising signaling, with the one or more processors, a magnitude of a translation vector indicative of an average velocity of the current frame.
[0263] Clause A37. The method of clause A23, further comprising performing, with the one or more processors, motion vector estimation for the current frame based on the global motion compensated frame, and wherein performing motion vector estimation comprises limiting a block size of local motion to be equal to a minimum prediction unit size.
[0264] Clause A38. The method of clause A23, wherein the reference frame is a first reference frame and the current frame is a first current frame, the method further comprising estimating an initial translation vector by minimizing a mean squared error between a second current frame and a second reference frame.
[0265] Clause A39. The method of clause A23, wherein the current frame is a first current frame, the method further comprising: determining, with the one or more processors, whether a point is ground; and estimating, with the one or more processors, a rotation matrix for a second current frame based on whether the point is ground.
[0266] Clause A40. The method of clause A39, further comprising signaling, with the one or more processors, a set of labels indicative of whether the point is ground.
[0267] Clause A41. The method of clause A23, further comprising generating, with the one or more processors, point cloud data.
[0268] Clause A42. The method of clause A41, further comprising generating, with the one or more processors, the point cloud data based on signals from a LIDAR device.
[0269] Clause A43. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: identify a first set of global motion parameters from global positioning system information; determine a second set of global motion parameters to be used for global motion estimation of a current frame based on the first set of global motion parameters; and apply motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
[0270] Clause B1. A device for encoding point cloud data, the device comprising: a memory for storing point cloud data; and one or more processors coupled to the memory and implemented in circuitry, the one or more processors configured to: identify a first set of global motion parameters from global positioning system information; determine a second set of global motion parameters to be used for a global motion estimation of a current frame based on the first set of global motion parameters; and apply motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
[0271] Clause B2. The device of clause B1, wherein the one or more processors are configured to signal the second set of global motion parameters in a bitstream.
[0272] Clause B3. The device of clause B1, wherein the one or more processors are configured to signal the first set of global motion parameters in a bitstream.
[0273] Clause B4. The device of any of clauses B1-B3, wherein the first set of global motion parameters comprises a set of orientation parameters.
[0274] Clause B5. The device of clause B4, wherein the set of orientation parameters comprises one or more of a roll, a pitch, a yaw, or an angular velocity of the current frame.
[0275] Clause B6. The device of clause B4, wherein the set of orientation parameters comprises one or more of a roll difference between the current frame and the reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, or an angular velocity difference between the current frame and the reference frame.
[0276] Clause B7. The device of any of clauses B1-B6, wherein the first set of global motion parameters comprises a set of position parameters.
[0277] Clause B8. The device of clause B7, wherein the set of position parameters comprises one or more of a displacement of the current frame or an average velocity of the current frame.
[0278] Clause B9. The device of clause B7, wherein the set of position parameters comprises one or more of an east velocity, a north velocity, or an up velocity of an east-north-up coordinate system.
[0279] Clause B10. The device of any of clauses B1-B9, wherein the second set of global motion parameters comprises a rotation matrix indicating a yaw, a pitch, and a roll of the current frame and a translation vector indicating an average velocity of the current frame.
[0280] Clause B11. The device of any of clauses B1, B4-B10, wherein the one or more processors are configured to signal a roll of the current frame, a pitch of the current frame, a yaw of the current frame, and a translation vector indicative of an average velocity of the current frame.
[0281] Clause B12. The device of any of clauses B1, B4-B10, wherein the one or more processors are configured to signal a roll difference between the current frame and the reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicative of an average velocity of the current frame.
[0282] Clause B13. The device of any of clauses B1, B4-B10, wherein the one or more processors are configured to signal a translation vector indicative of an average velocity of the current frame and refrain from signaling a rotation matrix indicative of a yaw, a pitch, and a roll of the current frame.
[0283] Clause B14. The device of any of clauses B1, B4-B10, wherein the one or more processors are configured to signal a magnitude of a translation vector indicative of an average velocity of the current frame.
[0284] Clause B15. The device of any of clauses B1-B14, wherein the one or more processors are configured to perform motion vector estimation for the current frame based on the global motion compensated frame, and wherein, to perform the motion vector estimation, the one or more processors are configured to limit a block size of the local motion to be equal to a minimum prediction unit size.
[0285] Clause B16. The device of any of clauses B1-B15, wherein the reference frame is a first reference frame and the current frame is a first current frame, and wherein the one or more processors are configured to estimate an initial translation vector by minimizing a mean squared error between a second current frame and a second reference frame.
[0286] Clause B17. The device of any of clauses B1-B15, wherein the current frame is a first current frame, and wherein the one or more processors are configured to: determine whether a point is ground; and estimate a rotation matrix for a second current frame based on whether the point is ground.
[0287] Clause B18. The device of clause B17, wherein the one or more processors are configured to signal a set of labels indicative of whether the point is ground.
[0288] Clause B19. The device of any of clauses B1-B18, wherein the one or more processors are further configured to generate point cloud data.
[0289] Clause B20. The device of clause B19, wherein the one or more processors are configured to generate the point cloud data based on signals from the LIDAR device as part of generating the point cloud data.
[0290] Clause B21. The device of any of clauses B1-B20, wherein the device is one of a mobile phone, a tablet, a vehicle, or an extended reality device.
[0291] Clause B22. The device of any of clauses B1-B21, wherein the device comprises an interface configured to transmit the encoded point cloud data.
[0292] Clause B23. A method for encoding point cloud data, the method comprising: identifying, with one or more processors, a first set of global motion parameters from global positioning system information; determining, with the one or more processors, a second set of global motion parameters to be used for a global motion estimation of a current frame based on the first set of global motion parameters; and applying, with the one or more processors, motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
[0293] Clause B24. The method of clause B23, further comprising signaling, with the one or more processors, the second set of global motion parameters in a bitstream.
[0294] Clause B25. The method of clause B23, further comprising signaling, with the one or more processors, the first set of global motion parameters in a bitstream.
[0295] Clause B26. The method of any of clauses B23-B25, wherein the first set of global motion parameters comprises a set of orientation parameters.
[0296] Clause B27. The method of clause B26, wherein the set of orientation parameters comprises one or more of a roll, a pitch, a yaw, or an angular velocity of the current frame.
[0297] Clause B28. The method of clause B26, wherein the set of orientation parameters comprises one or more of a roll difference between the current frame and the reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, or an angular velocity difference between the current frame and the reference frame.
[0298] Clause B29. The method of any of clauses B23-B28, wherein the first set of global motion parameters comprises a set of position parameters.
[0299] Clause B30. The method of clause B29, wherein the set of position parameters comprises one or more of a displacement of the current frame or an average velocity of the current frame.
[0300] Clause B31. The method of clause B29, wherein the set of positioning parameters comprises one or more of an east velocity, a north velocity, or an up velocity of an east-north-up coordinate system.
[0301] Clause B32. The method of any of clauses B23-B31, wherein the second set of global motion parameters comprises a rotation matrix indicative of a yaw, a pitch, and a roll of the current frame and a translation vector indicative of an average velocity of the current frame.
[0302] Clause B33. The method of clause B23, further comprising signaling a roll of the current frame, a pitch of the current frame, a yaw of the current frame, and a translation vector indicative of an average velocity of the current frame.
[0303] Clause B34. The method of any of clauses B23, B26-B33, further comprising signaling a roll difference between the current frame and the reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicative of an average velocity of the current frame.
[0304] Clause B35. The method of any of clauses B23, B26-B33, further comprising signaling, with the one or more processors, a translation vector indicative of an average velocity of the current frame and refraining from signaling a rotation matrix indicative of a yaw, a pitch, and a roll of the current frame.
[0305] Clause B36. The method of any of clauses B23, B26-B33, further comprising signaling, with the one or more processors, a magnitude of a translation vector indicative of an average velocity of the current frame.
[0306] Clause B37. The method of any of clauses B23-B36, further comprising performing, with the one or more processors, motion vector estimation for the current frame based on the global motion compensated frame, and wherein performing the motion vector estimation comprises limiting a block size of the local motion to be equal to a minimum prediction unit size.
[0307] Clause B38. The method of any of clauses B23-B37, wherein the reference frame is a first reference frame and the current frame is a first current frame, the method further comprising estimating an initial translation vector by minimizing a mean squared error between a second current frame and a second reference frame.
[0308] Clause B39. The method of any of clauses B23-B37, wherein the current frame is a first current frame, the method further comprising determining, with the one or more processors, whether a point is ground and estimating, with the one or more processors, a rotation matrix for a second current frame based on whether the point is ground.
[0309] Clause B40. The method of clause B39, further comprising signaling, with the one or more processors, a set of labels indicating whether the points are ground.
[0310] Clause B41. The method of any of clauses B23-B40, further comprising generating, with the one or more processors, point cloud data.
[0311] Clause B42. The method of clause B41, further comprising generating, with the one or more processors, the point cloud data based on signals from a LIDAR device.
[0312] Clause B43. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: identify a first set of global motion parameters from global positioning system information; determine, based on the first set of global motion parameters, a second set of global motion parameters to use for a global motion estimation of a current frame; and apply motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
[0313] It will be recognized that certain actions or events can be performed in a different order, or omitted, in other examples, or in a different sequence, or in parallel, or with additional actions or events, without departing from the scope of the examples described herein. Furthermore, in certain examples, actions or events can be performed concurrently, such as through multi-threaded processing, interrupt processing, or multiple processors, rather than sequentially.
[0314] In one or more examples, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored or transmitted as one or more instructions or code on a computer-readable medium, which can be accessed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer- readable media generally can correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product can include a computer-readable medium.
[0315] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave may be included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient, tangible storage media. As used herein, disks and optical discs include optical discs (CDs), laser discs, optical optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically magnetically copy data, while optical discs optically copy data using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0316] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuitry" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques can be fully implemented in one or more circuit or logic elements.
[0317] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be provided in combination within a codec hardware unit or by a collection of interoperable hardware units, including one or more processors as described above, combined with suitable software and / or firmware.
[0318] Various examples have been described. These examples, as well as others, are within the scope of the following claims.
Claims
1. A device for encoding point cloud data, the device comprising: a memory to store the point cloud data; and one or more processors coupled to the memory and implemented in circuitry, the one or more processors configured to: identify a first set of global motion parameters from global positioning system information; determine a second set of global motion parameters to be used for global motion estimation of a current frame based on the first set of global motion parameters; and apply motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
2. The device of claim 1, wherein the one or more processors are configured to signal the second set of global motion parameters in a bitstream.
3. The device of claim 1, wherein the one or more processors are configured to signal the first set of global motion parameters in a bitstream.
4. The device of claim 1, wherein the first set of global motion parameters comprises a set of orientation parameters.
5. The device of claim 4, wherein the set of orientation parameters comprises one or more of a roll, a pitch, a yaw, or an angular velocity of the current frame.
6. The device of claim 4, wherein the set of orientation parameters comprises one or more of a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the frame, or an angular velocity difference between the current frame and the reference frame.
7. The device of claim 1, wherein the first set of global motion parameters comprises a set of position parameters.
8. The device of claim 7, wherein the set of position parameters comprises one or more of a displacement of the current frame or an average velocity of the current frame.
9. The device of claim 7, wherein the set of position parameters comprises one or more of an east velocity, a north velocity, or an up velocity of an east-north-up coordinate system.
10. The device of claim 1, wherein the second set of global motion parameters comprises a rotation matrix indicating a yaw, a pitch, and a roll of the current frame and a translation vector indicating an average velocity of the current frame.
11. The device of claim 1, wherein the one or more processors are configured to signal a roll of the current frame, a pitch of the current frame, a yaw of the current frame, and a translation vector indicating an average velocity of the current frame.
12. The device of claim 1, wherein the one or more processors are configured to signal a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicating an average velocity of the current frame.
13. The device of claim 1, wherein the one or more processors are configured to signal a translation vector indicating an average velocity of the current frame and refrain from signaling a rotation matrix indicating a yaw, a pitch, and a roll of the current frame. 14. The device of claim 1, wherein the one or more processors are configured to signal a magnitude of a translation vector that indicates an average velocity of the current frame.
15. The device of claim 1, wherein the one or more processors are configured to perform motion vector estimation for the current frame based on the global motion compensated frame, and wherein, To perform the motion vector estimation, the one or more processors are configured to limit a block size of the local motion to be equal to a minimum prediction unit size.
16. The device of claim 1, wherein the reference frame is a first reference frame and the current frame is a first current frame, and wherein the one or more processors are configured to estimate an initial translation vector by minimizing a mean squared error between a second current frame and a second reference frame.
17. The device of claim 1, wherein the current frame is a first current frame, and wherein the one or more processors are configured to: determine whether a point is ground; and estimate a rotation matrix for a second current frame based on whether the point is ground.
18. The device of claim 17, wherein the one or more processors are configured to signal a set of labels that indicate whether the point is ground.
19. The device of claim 1, wherein the one or more processors are further configured to generate the point cloud data.
20. The device of claim 19, wherein the one or more processors are configured to generate the point cloud data based on signals from a LIDAR device as part of generating the point cloud data.
21. The device of claim 1, wherein the device is one of a mobile phone, a tablet, a vehicle, or an extended reality device.
22. The device of claim 1, wherein the device includes an interface configured to transmit encoded point cloud data.
23. A method for encoding point cloud data, the method comprising: identifying, with one or more processors, a first set of global motion parameters from global positioning system information; determining, with one or more processors, a second set of global motion parameters to be used for global motion estimation of a current frame based on the first set of global motion parameters; and applying, with one or more processors, motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
24. The method of claim 23, further comprising signaling, with the one or more processors, the second set of global motion parameters in a bitstream.
25. The method of claim 23, further comprising signaling, with the one or more processors, the first set of global motion parameters in a bitstream.
26. The method of claim 23, wherein the first set of global motion parameters includes a set of orientation parameters.
27. The method of claim 26, wherein the set of orientation parameters includes one or more of a roll, a pitch, a yaw, or an angular velocity of the current frame.
28. The method of claim 26, wherein the set of orientation parameters comprises one or more of a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, or an angular velocity difference between the current frame and the reference frame.
29. The method of claim 23, wherein the first set of global motion parameters comprises a set of position parameters.
30. The method of claim 29, wherein the set of position parameters comprises one or more of a displacement of the current frame or an average velocity of the current frame.
31. The method of claim 29, wherein the set of position parameters comprises one or more of an east velocity, a north velocity, or an up velocity of an east-north-up coordinate system.
32. The method of claim 23, wherein the second set of global motion parameters comprises a rotation matrix indicative of a yaw, a pitch, and a roll of the current frame and a translation vector indicative of an average velocity of the current frame.
33. The method of claim 23, further comprising signaling a roll of the current frame, a pitch of the current frame, a yaw of the current frame, and a translation vector indicative of an average velocity of the current frame.
34. The method of claim 23, further comprising signaling a roll difference between the current frame and a reference frame, a pitch difference between the current frame and the reference frame, a yaw difference between the current frame and the reference frame, and a translation vector indicative of an average velocity of the current frame.
35. The method of claim 23, further comprising signaling, with the one or more processors, a translation vector indicative of an average velocity of the current frame and refraining from signaling a rotation matrix indicative of a yaw, a pitch, and a roll of the current frame.
36. The method of claim 23, further comprising signaling, with the one or more processors, a magnitude of a translation vector indicative of an average velocity of the current frame.
37. The method of claim 23, further comprising performing, with the one or more processors, motion vector estimation for the current frame based on the global motion compensated frame, and wherein performing motion vector estimation comprises limiting a block size of local motion to be equal to a minimum prediction unit size.
38. The method of claim 23, wherein the reference frame is a first reference frame and the current frame is a first current frame, the method further comprising estimating an initial translation vector by minimizing a mean squared error between a second current frame and a second reference frame.
39. The method of claim 23, wherein the current frame is a first current frame, the method further comprising: determining, with the one or more processors, whether a point is ground; and estimating, with the one or more processors, a rotation matrix for a second current frame based on whether the point is ground.
40. The method of claim 39, further comprising signaling, with the one or more processors, a set of labels indicative of whether the point is ground.
41. The method of claim 23, further comprising generating, with the one or more processors, the point cloud data.
42. The method of claim 41, further comprising generating, with the one or more processors, the point cloud data based on signals from a LIDAR device.
43. A computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: identify a first set of global motion parameters from global positioning system information; determine a second set of global motion parameters to use for global motion estimation of a current frame based on the first set of global motion parameters; and apply motion compensation to a reference frame based on the second set of global motion parameters to generate a global motion compensated frame for the current frame.
44. An apparatus for processing point cloud data, the apparatus comprising means for performing the method of any one of claims 23 to 42.
Citation Information
Patent Citations
Image prediction method, system and device based on three-dimensional point cloud model under cloud environment
CN104715496A
Point cloud attribute compression method based on deletion of 0 elements in quantization matrix
CN108833927A