Method, encoder, decoder and bitstream for encoding and decoding 3D point clouds

By receiving frame information from the bitstream, utilizing the rotation regularity and order index difference of the LiDAR system, and combining it with a coarse representation framework, the encoding and decoding process of point clouds is optimized, solving the problems of redundancy and low efficiency in point cloud compression, and improving the accuracy and efficiency of point cloud reconstruction in applications such as autonomous driving.

CN118830250BActive Publication Date: 2025-12-05BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202480001735.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-02
Publication Date
2025-12-05
Estimated Expiration
2044-02-02

AI Technical Summary

Technical Problem

Existing point cloud compression technologies suffer from redundancy and low efficiency during encoding and decoding, especially in dynamic environments, particularly in autonomous driving applications, where existing coding structures fail to effectively utilize prior information from encoded frames.

Method used

By receiving frame information from the bitstream, the points of the point cloud are determined using at least one previously obtained frame information. The point cloud reconstruction process is optimized by using the order index difference and weighting method. In the encoding process, the rotation regularity of the LiDAR system is utilized, combined with the coarse representation framework, to reduce the encoding of redundant information.

Benefits of technology

It improves the temporal accuracy and efficiency of point cloud reconstruction, especially in dynamic environments, reduces the decoding of redundant information, optimizes the processing of different spatial features, and improves coding efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118830250B_ABST
    Figure CN118830250B_ABST
Patent Text Reader

Abstract

The invention relates to a method of encoding and decoding a 3D point cloud, an encoder, a decoder, a bitstream and a computer readable storage medium. The method of decoding a 3D point cloud is particularly suitable for point clouds acquired by a rotating light detection and ranging, Lidar, system with multiple sensors. The process involves receiving a bitstream comprising frame information of a sequence of frames, wherein each frame represents point cloud data at a particular time. For a starting frame in the sequence, the method comprises obtaining the frame information and determining points of the point cloud based on decoding data therefrom. For subsequent frames, the method involves obtaining the frame information and then determining points of the point cloud based on information from at least one previously obtained frame. The invention is particularly suitable for applications requiring accurate and efficient processing of three-dimensional spatial data, such as autonomous vehicle navigation, geographic mapping and environmental monitoring, in which a detailed and dynamic representation of the environment is of paramount importance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to a method of encoding and decoding a 3D point cloud, an encoder, a decoder, a bitstream and a computer readable storage medium. In particular, it relates to point cloud geometry data captured by a spinning sensor head. BACKGROUND

[0002] As a format of 3D data representation, point clouds have recently gained attention as they are versatile in representing all kinds of 3D objects or scenes. However, for all compression schemes, the reconstruction quality of the point cloud points is of utmost importance. SUMMARY

[0003] It is therefore an object of the present invention to provide a method for decoding geometry of a 3D point cloud from a bitstream with improved efficiency as well as encoding a 3D point cloud into a bitstream.

[0004] This problem is solved by the decoding method according to a first aspect, the encoding method according to another aspect, the decoder according to another aspect, the encoder according to another aspect and the software according to another aspect.

[0005] In a first aspect, a method of decoding geometry of a 3D point cloud from a bitstream is provided, the method preferably implemented in a decoder. The method comprises: receiving a bitstream, wherein the bitstream comprises frame information of a sequence of frames, each frame representing point cloud data at a particular time; for a first frame in the sequence, obtaining the frame information and determining points of the point cloud from decoded data of the first frame; for subsequent frames (S16), obtaining the frame information and determining points of the point cloud based on at least one previously obtained frame information. Wherein the frame information comprises data indicative of relative positions of points in the point cloud; and each point of the point cloud is represented in a multi-dimensional coordinate system, the representation and ordering of the points being based on at least one characteristic of the points being captured. Thus, the invention relates to a method for interpreting and reconstructing 3D spatial data from a sequential data stream, preferably generated by a Lidar system mounted on a rotating head. The rotating Lidar head can comprise one or more sensors to capture respective point clouds of physical objects. The system captures a sequence of “frames”, each frame being a snapshot of 3D data points representative of the surrounding environment at a given time instant. The decoding process starts from receiving an encoded data stream (bitstream) comprising detailed information about these frames, including their relative positions. For the initial frame in the sequence, the method decodes the point cloud directly from its encoded form, as at the current stage, the a priori knowledge of the point cloud is unknown. For subsequent frames, the decoding process utilizes information from at least one previous frame (including the relative position data) to accurately reconstruct the point cloud. Each point in the point cloud is represented in a multi-dimensional coordinate system, enabling flexible and efficient reconstruction of the 3D environment captured by the Lidar system. This method allows for efficient and accurate reconstruction of the 3D environment captured by the Lidar system.

[0006] Preferably, the frame information comprises a sequence index difference, the sequence index difference being configured to indicate a difference between samples of the point cloud, wherein each sample comprises one or more points of the point cloud. Thus, the frame information is specified to comprise a “sequence index difference”. The sequence index difference can be used to show how the samples in the point cloud differ from each other. Each sample can consist of a single point or multiple points in the point cloud. The sequence index difference can be interpreted as any type of metric or value that quantifies or describes how different parts (samples or points) of the point cloud differ from each other in terms of sequence or ordering. In a more specific example, the sequence index difference can be a numerical value indicating a difference in ordering in the sample ordering when capturing or processing the samples. For example, if the Lidar captures the point cloud in a linear path, the sequence index difference can reflect the distance or ordering between consecutive samples along the path.

[0007] Preferably, the sequence index is an index associated with each sample according to the decoding order of the point cloud. Thus, the index is assigned to each sample and is based on the decoding order of the point cloud, i.e. the order in which the point cloud data is encoded (or decoded or arranged). The decoding order and the subsequent sequence index can include various ways, such as the time order of data capture, the spatial arrangement of points or the order of data processing and storage. For a specific example, consider a Lidar system scanning an area in a systematic pattern. The decoding order can refer to the order in which different areas are scanned, and the sequence index can be a number identifying the position of each sample in this sequence. For example, if the Lidar scans from left to right, the sequence index can increase from the leftmost to the rightmost sample. The decoding order can also be as described in the detailed description.

[0008] Preferably, the previously obtained frame information is the frame information immediately preceding in the sequence. Thus, by directly using the preceding frame information to decode the subsequent frame, the method ensures continuity and consistency in the data interpretation process. This approach improves the temporal accuracy of point cloud reconstruction, especially in dynamic environments where consecutive frames are closely related. Redundant information no longer needs to be decoded, thus improving decoding efficiency.

[0009] Preferably, the method further comprises determining whether the frame information in the bitstream corresponds to a high or low portion of the point cloud; and continuing to determine the points of the point cloud based on the at least one previously obtained frame information. Thus, the method determines whether the data belongs to a higher segment or a lower segment of the point cloud and adjusts the decoding process accordingly. This distinction is crucial for applications such as autonomous driving, where the lower and higher portions of the point cloud data do not change frequently and are therefore more predictable based on previously decoded frames. Therefore, this embodiment enhances the decoding process by introducing an additional decision step that identifies the specific segment (high or low) of the point cloud. This segmentation can enable the customization of targeted decoding strategies for different portions of the point cloud, optimizing processing for different spatial features, which is particularly beneficial in complex environments with different terrains. The high portion of the point cloud can correspond to data acquired by higher laser beams (e.g. from 8° to 15°, with a preferred range of 10° to 15°), and the low portion of the point cloud can correspond to data acquired by lower laser beams (e.g. from -14° - Δθ to -25°, with a preferred range of -20° to -25°, where Δθ is the angular difference between two laser beams in the vertical direction). The above angles refer to the range of Lidar scanning in the vertical direction.

[0010] Preferably, for a subsequent frame, the acquisition of frame information and the determination of points of the point cloud are based on at least two previously acquired frame information, wherein each of said two previously acquired frame information is associated with a weight, preferably, frame information from frames closer in the sequence to the current frame is assigned a higher weight relative to frame information from frames further in the sequence. Thus, the method assigns weights to previously acquired frame information, giving higher weight to frames closer in the sequence to the current frame being decoded. This weighting method improves the accuracy and relevance of the decoded data, especially in dynamic environments, where more recently generated data is generally more predictive of the current state than earlier data.

[0011] Preferably, the bitstream further contains a flag configured to enable or disable the aforementioned method, wherein the flag is at least one bit. Thus, the decoding method includes in the bitstream a flag allowing to enable or disable certain processes, providing flexibility and control over the decoding algorithm.

[0012] Preferably, each captured point of the point cloud is represented by a point of a two-dimensional discrete angular plane, the sensor index as a coordinate of a first axis, the azimuth angle as a coordinate of a second axis, the points of the point cloud are ordered according to a lexicographic order, first based on said azimuth angle using said lexicographic order, then based on said sensor index using said lexicographic order, or vice versa; the sensor index is an index associated with a sensor that captured at least one point of the point cloud, and the azimuth angle is a capture angle of said sensor. Thus, the method also uses an innovative method to organize point cloud data, representing each point within a two-dimensional discrete angular plane. This arrangement based on sensor index and azimuth angle and ordered lexicographically ensures that the large amount of data typically generated by Lidar systems is processed in an efficient, systematic manner. Reference is also made to Figure 9 The points represented by solid lines are captured points of the point cloud. Those represented by dashed lines can be unoccupied (e.g. the reflected beam is not received by the emitting laser). Thus, for two consecutive sequential captured points, the sequence index difference can be, for example, 1, 2, 3, etc. With the above information, for example, if the first point sequence information is acquired and the sequence difference is known, the other points of the point cloud can be reconstructed. In addition, the sequence difference of a subsequent frame can also be predicted.

[0013] Preferably, the method further comprises determining that the frame information in the bitstream corresponds to an intermediate portion of the point cloud; and not continuing to determine points of the point cloud based on the at least one previously obtained frame information. Thus, the method further involves an additional step of identifying whether the frame information in the bitstream is associated with an intermediate portion of the point cloud. Upon determining that the frame information corresponds to the intermediate portion, the method requires a procedure to decide not to continue to determine points of the point cloud based on the previously obtained frame information. This claim can be seen as a method to optimize the decoding procedure by skipping certain calculations for the intermediate portion of the point cloud. The "intermediate portion" can refer to any central segment of the scanning region, or determined by its geometric position or the order of data capture. Preferably, the intermediate portion corresponds to data acquired by intermediate laser beams (e.g. from -14° to 8° - Δθ, preferably, the range can be from -6° to 3°, where Δθ is the angular difference of two laser beams in the vertical direction). The above angles refer to the range of Lidar scanning in the vertical direction.

[0014] In another aspect of the application, an encoding method is provided, the method comprising: generating frame information for a sequence of frames, each frame representing point cloud data at a particular time; for a first frame in the sequence, encoding points of the point cloud based on three-dimensional data of the points of the point cloud; for subsequent frames, encoding points of the point cloud based on at least one previously encoded frame information. The embodiments described with reference to the decoding method also apply to the corresponding encoding method.

[0015] In another aspect of the application, an encoder for encoding a 3D point cloud into a bitstream is provided. The encoder comprises a memory and a processor, wherein instructions are stored in the memory that, when executed by the processor, implement the steps of the aforementioned encoding method.

[0016] In another aspect of the application, a decoder for decoding a 3D point cloud from a bitstream is provided. The decoder comprises a memory and a processor, wherein instructions are stored in the memory that, when executed by the processor, implement the steps of the aforementioned decoding method.

[0017] In another aspect of the application, a bitstream is provided, wherein the bitstream is encoded by the steps of the aforementioned encoding method.

[0018] In another aspect of the application, a computer readable storage medium is provided, comprising instructions for implementing the steps of the aforementioned method of encoding a 3D point cloud into a bitstream.

[0019] In another aspect of the application, a computer readable storage medium is provided, comprising instructions for implementing the steps of the aforementioned method of decoding a 3D point cloud from a bitstream. BRIEF DESCRIPTION OF DRAWINGS

[0020] In the following, the application is described in more detail with reference to the accompanying drawings.

[0021] Figure 1 A rotating Lidar head comprising several rotating lasers probing the environment is shown;

[0022] Figure 2 An elevation angle Θ of a rotating laser is shown;

[0023] Figure 3 A 2D angular representation of points acquired by a rotating Lidar is shown;

[0024] Figure 4 A discretization of acquired points is shown;

[0025] Figure 5 A 3D xyz coordinate and an angular-based coordinate are shown; or

[0026] Figure 6 An acquisition order in a plane made of azimuth angle and laser index λ is shown;

[0027] Figure 7 A point in a plane made of azimuth angle and laser index λ is shown with real numbers;

[0028] Figure 8 An ordering of points in a plane made of coarse azimuth angle and laser index λ is shown;

[0029] Figure 9 A representation of a point cloud by difference Δnext for a first dictionary order is shown;

[0030] Figure 10 An overview of an encoding method implementing an example is shown;

[0031] Figure 11 An overview of a decoding method implementing an example is shown;

[0032] Figure 12a and Figure 12b A coarse representation of two Lidar point cloud frames with 1 second time difference is shown;

[0033] Figure 13 An overview of a Lidar point cloud decoding method according to the invention is shown;

[0034] Figure 14 A block diagram of a decoding method according to the invention is shown;

[0035] Figure 15 ​​​a block diagram of an encoding method according to the application is shown;

[0036] Figure 16 a decoder / encoder according to the application is shown. DETAILED DESCRIPTION

[0037] As a format to represent 3D data, point clouds have recently gained interest as they are versatile in their ability to represent all types of 3D objects or scenes. As a result, many use cases can be handled by point clouds, among which

[0038] • movie post-production,

[0039] • real-time 3D immersive telepresence or Virtual Reality (VR) / Augmented Reality (AR) applications,

[0040] • free-viewpoint video (e.g. for watching sports),

[0041] • geographic information systems (a.k.a. cartography),

[0042] • cultural heritage (storing scans of rare items in digital form),

[0043] • autonomous driving, including 3D mapping of the environment and real-time radar Lidar data acquisition

[0044] Point clouds are a set of points located in a 3D space, to which additional values can be optionally attached for each point. These additional values are usually referred to as point attributes. Thus, a point cloud is a combination of geometry (3D position of each point) and attributes.

[0045] The attributes can be, for example, three-component colors, material attributes such as reflectance, and / or two-component normal vectors of the surface associated with the point. Point clouds can be captured by various types of devices, like camera arrays, depth sensors, Lidar, scanners, or can be computer-generated (for example, in movie post-production). Depending on the use case, for mapping applications, point clouds can have from thousands to billions of points. The raw representation of a point cloud requires very high number of bits per point, at least twelve bits per spatial component X, Y or Z, and optionally more bits for the attributes, for example, ten bits times three for the color. Practical deployment of point cloud based applications requires compression techniques that enable storing and distributing point clouds with reasonable storage and transmission infrastructure. For distribution to and visualization by end users, for example, on AR / VR glasses or any other 3D-enabled device, compression can be lossy (as in video compression). Other use cases do require lossless compression, for example, medical applications or autonomous driving, to avoid altering the decision outcomes from the analysis of the compressed and transmitted point clouds. Until recently, point cloud compression (aka PCC) has not been addressed by the mass market and no standardized point cloud codec was available. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11, also known as Moving Picture Experts Group or MPEG, started a work item on point cloud compression. This resulted in two standards, namely

[0046] • MPEG-I Part 5 (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC)

[0047] • MPEG-I Part 9 (ISO / IEC 23090-9) or Geometry-based Point Cloud Compression (G-PCC)

[0048] Both V-PCC and G-PCC standards completed their first versions at the end of 2020 and will soon be available on the market. The V-PCC encoding method compresses point clouds by performing multiple projections of the 3D object to obtain 2D patches packed into images (or into videos when dealing with moving point clouds). The obtained images or videos are then compressed using existing image / video codecs, allowing to leverage already deployed image and video solutions. By its nature, V-PCC is only efficient on dense and continuous point clouds, as image / video codecs cannot compress non-smooth patches, for example, obtained from the projection of sparse geometry data captured from a Lidar.

[0049] There are two schemes for geometry compression in G-PCC coding methods. The first scheme is based on an occupancy tree (octree / quadtrees / binary tree) representation of the point cloud geometry. Occupied nodes are split until a certain size is reached and the occupied leaf nodes provide the positions of the points, typically at the center of these nodes. By using a neighbor-based prediction technique, a high level of compression can be achieved for dense point clouds. Sparse point clouds are also addressed by directly encoding the positions of the points within nodes of non-minimal size by stopping the tree construction when only isolated points exist in a node; this technique is called Direct Coding Mode (DCM). The second scheme is based on a prediction tree, where each node represents a 3D position of a point and the relationship between nodes is a spatial prediction from parent to child. This method can only handle sparse point clouds and has the advantage of lower latency and simpler decoding compared to the occupancy tree. However, the compression performance is only slightly better than the first occupancy-based method and is complex to encode, densely searching for the best predictor (among a long list of potential predictors) when building the prediction tree. In both schemes, attribute (solution) encoding can be performed after the geometry (solution) encoding is completed, resulting in two-pass encoding. Therefore, low latency is achieved by using slices that decompose the 3D space into independently encoded sub-volumes without the need for prediction between sub-volumes. When many slices are used, the compression performance can be severely impacted.

[0050] One important use case is the transmission of Lidar data collected by moving vehicles. This typically requires a simple low-latency on-board encoder for the vehicle. The need for simplicity is because the encoder can be deployed on a computing unit that performs other processing in parallel, such as (semi-)autonomous driving, thereby limiting the processing power available to the point cloud encoder. Low latency is also required in order to quickly transmit from the car to the cloud, in order to view local traffic in real-time based on multi-vehicle collection and make decisions based on traffic information quickly enough. While using 5G can make the transmission latency low enough, the encoder itself should not introduce too much latency due to encoding. Furthermore, compression performance is extremely important because the data stream from millions of cars to the cloud is expected to be very large. How to combine simplicity, low latency, and compression performance of the encoder and decoder is still an open problem that has not been fully solved by existing point cloud codecs.

[0051] Certain priors related to Lidar data collection have been exploited in G-PCC and have brought very significant compression gains. The first technique involves rotating the vertical angle (with respect to the horizontal ground) from which the Lidar collects data (as shown in Figure 1 The Lidar head in Figure 1 can be rotated around the vertical axis (dashed line) to capture the geometry data of an object: this angle θ is fixed as shown in Figure 2

[0052] ​In practice, the representation of the point cloud collected by a Lidar is quasi two-dimensional (2D) in the spherical coordinates above, where r 3D is the 3D distance of a point to the center of the Lidar. It is the linear distance directly measured from the Lidar sensor to a point in space, independent of the horizontal or vertical angle. The parameter r is crucial for determining the absolute distance of an object from the Lidar sensor. It is the basis for generating accurate 3D maps and models of the environment, as it provides the direct spatial relationship between the sensor and the objects in its field of view. is the azimuth angle of the rotation of the Lidar head with respect to a reference. Azimuth angle refers to the horizontal angular displacement measured with respect to a defined origin. Specifically, it is the angle between the projection of the point in question on the horizontal plane and a fixed reference direction on that plane. In Lidar systems, is used to determine the horizontal orientation of a scanned point with respect to the Lidar device. This is crucial for calculating the precise horizontal position of objects in the scanned environment. θ is the elevation angle of the sensor of the Lidar head with respect to the horizontal reference plane. This elevation angle represents the vertical angular displacement with respect to a defined horizontal plane. It is the angle between the line from the point to the center of the Lidar and the horizontal plane, measured on the vertical plane containing that line. θ is crucial for determining the vertical position of objects in the Lidar scan. It can calculate the height or altitude of a point with respect to the Lidar device, contributing to the creation of a three-dimensional representation of the scanned area.

[0053] G-PCC further exploits a second technique, exploiting the regularity of the laser sensing as the Lidar head rotates, as shown in Figure 3 For the data collected by a Lidar, a regular distribution is observed as a function of the azimuth angle This regularity can be exploited to obtain a quasi-1D representation of the point cloud, where, due to noise, preferably only the radius r 3D belongs to a continuous range of values, while the angles and θ can only take discrete values. Basically, the point cloud geometry can be represented on a 2D discrete angular plane (see Figure 4 ), as well as the radius value for each point. G-PCC exploits this quasi-1D property in both the occupancy tree and the prediction tree, predicting the position of the current point with respect to the already coded points in the spherical coordinates by exploiting the discreteness of the angles.

[0054] In practice, the occupancy tree can exploit the DCM extensively and entropy encode the direct position of the points within the nodes by using a context adaptive entropy encoder. The context is obtained from a local conversion of the point position to angular (φ, θ) coordinates. The prediction tree directly encodes the angular coordinates where r 2Dis the projected radius on the horizontal xy-plane (see Figure 5 and then converted to (x, y, z) and then the xyz residual is encoded to account for errors in coordinate conversion, approximation of laser angle and noise (the xyz residual refers to the difference between the actual coordinates of a point and its estimated coordinates).

[0055] As mentioned above, there are mainly two types of encoding structures in the art, namely occupancy tree and prediction tree.

[0056] However, according to the existing encoding structures, there is still redundancy in the encoding, and the prior information of the encoded frames is not taken into account. Therefore, a decoding method according to Figure 14 and an encoding method according to Figure 15 are proposed.

[0057] The decoding method comprises the following steps:

[0058] receiving a bitstream, wherein the bitstream contains frame information of a sequence of frames, each frame representing point cloud data at a specific time (S12);

[0059] for the first frame in the sequence of frames, obtaining the frame information and determining the points of the point cloud based on the decoded data therefrom (S14);

[0060] for the subsequent frames, obtaining the frame information and determining the points of the point cloud based on at least one previously obtained frame information, wherein the frame information comprises data indicative of the relative position of the points in the point cloud; and each point of the point cloud is represented in a multi-dimensional coordinate system, the representation and ordering of the points being based on at least one feature capturing the points.

[0061] The encoding method comprises the following steps:

[0062] generating frame information of a sequence of frames, each frame representing point cloud data at a specific time (S22);

[0063] for the first frame in the sequence, encoding the points of the point cloud based on three-dimensional data of the points of the point cloud (S24);

[0064] for the subsequent frames, encoding the points of the point cloud based on at least one previously encoded frame information, wherein the frame information comprises data indicative of the relative position of the points in the point cloud; and each point of the point cloud is represented in a multi-dimensional coordinate system, the representation and ordering of the points being based on at least one feature capturing the points.

[0065] Therefore, according to the proposed method, the encoding efficiency can be improved by taking into account one or more adjacent Lidar data frames.

[0066] The proposed method can also be used in conjunction with coarse representation of point clouds. The invention will be described below within the framework of coarse representation, and embodiments can be freely combined.

[0067] In some embodiments, all lasers (sensors) can be encoded once using the acquisition sequence, such as... Figure 6 of As shown in the plane. Due to the regular rotation of the LiDAR head and the continuous acquisition by each laser at fixed time intervals, the azimuth distance between two points detected by the same laser is the basic azimuth offset. Multiples of.

[0068] Firstly, the coarse representation can be encoded based on prior data collection, rather than directly encoding the point locations. For example, a coarse representation of the point cloud geometry can be used. Sort the points using lexicographical sort (also known as dictionary order), first by coarse azimuth. First, use dictionary sorting, then use dictionary order in the laser index λ (or sensing elevation index), or vice versa.

[0069] Indicatively, it is possible to According to the plane Figure 7 The points are acquired in sequence as shown. Due to the regular rotation of the LiDAR head and the continuous acquisition by each laser at fixed time intervals, the azimuth distance between two points detected by the same laser is the basic azimuth offset. The multiple of. In reality, not all points are collected; that is, the laser beam may not be reflected, resulting in collection noise, and the laser may not be perfectly aligned. Actual data as... Figure 7 As shown.

[0070] Coarse angle It can be easily obtained through the quantization of φ, as shown below:

[0071]

[0072] The ordinal index o(P) of point P can be obtained by the following formula.

[0073]

[0074] Where, N laser λ is the number of lasers (i.e., the maximum value of λ), where λ is the laser index of the capture point P, in [0, N]. laser The codec encodes point P monotonically according to the order o(P) of point P, for example, using ascending order. Therefore, it can be done according to... Figure 8 The points are encoded in the order shown.

[0075] Coarse representation on a plane Can be encoded by:

[0076] • Number of points N points

[0077] • Value of the first captured point

[0078] • N points - 1 continuous difference Dnext, which is the difference between the current point and the next point ordered lexicographically, as shown in: Figure 9

[0079] The coarse representation is composed of continuous differences Dnext, and the compression of the coarse representation is basically based on the compression of continuous positive values Dnext.

[0080] An overview of the encoding method based on the above method can be shown as: Figure 10

[0081] First, according to the knowledge of Lidar sensor settings, the encoder can convert the xyz point position into laser (sensor) index λ, coarse angle and radius r 2D . Determine the difference Dnext, and apply this method to encode Dnext into the bitstream, as well as useful information about the sensor settings, such as and laser elevation angle.

[0082] Second, the reconstructed azimuth angle can be calculated, which can be directly obtained from the dequantization of the coarse angle . Optionally, the residual can be calculated in difference and encoded into the bitstream. Before encoding, the residual can be quantized to In this case, the reconstructed azimuth angle can be obtained by

[0083]

[0084] Where IQ represents the dequantization process.

[0085] The radius r (here r 2D ), optionally after quantization to Q(r), is encoded. After dequantization, the reconstructed radius

[0086] r rec = IQ(Q(r)).

[0087] Third, the reconstructed azimuth angle and the reconstructed radius r​​rec can be converted back to xy to obtain an estimate of the x and y position of the point

[0088]

[0089] the residual x relative to the original point position xy res and y res can be calculated as

[0090] x res = x - x estim , y res = y - y estim ,

[0091] and encoded into the bitstream.

[0092] Fourth and last, the vertical estimate z can be obtained from the laser angle θ(λ) by estim

[0093] z estim = r rec tan(θ(λ))

[0094] and the residual z relative to the original point position z can be calculated as res

[0095] z res = z - z estim

[0096] and encoded into the bitstream.

[0097] An overview of the related decoding method is shown in Figure 11 Once the encoding process is understood, the operations can be performed straightforwardly.

[0098] First, the decoder can decode from the bitstream the useful information about the sensor setup (e.g. the number of lasers N and the laser elevation angle) and the difference Δnext by using the proposed method. Then, the values of the laser index λ and the coarse angle can be obtained from Δnext as described below.

[0099] Second, the reconstructed azimuth angle can be calculated. It can be obtained directly from the dequantization of the coarse angle Alternatively, the azimuth angle residual can be decoded from the bitstream. The decoded residual can be a quantized version of the residual and the reconstructed azimuth angle is obtained by

[0100]

[0101] where IQ stands for the dequantization process.​

[0102] The radius r (here r 2D ) can also be decoded. The decoded radius can be a quantized version Q(r) of the radius. Dequantizing this gives the reconstructed radius

[0103] r rec = IQ(Q(r)).

[0104] Optionally, the radius can be predicted (e.g. by the preceding decoded radius) and a radius residual is decoded in place of the radius.

[0105] Third, the reconstructed azimuth angle and the reconstructed radius r rec are converted back to xy, estimating the x and y positions

[0106]

[0107] The residuals x res and y res can be decoded from the bitstream and the horizontal positions x dec and y dec of the point can be calculated

[0108] x dec = x estim + x res , y dec = y estim + y res。

[0109] Fourth and last, the vertical estimate z estim can be obtained from the laser angle θ(λ) by

[0110] z estim = r rec tan(θ(λ)),

[0111] The residual z res can be decoded from the bitstream and the decoded vertical position z dec of the point can be calculated

[0112] z dec = z estim + z res .

[0113] In practical applications of Lidar, such as autonomous driving, the scanning speed of a Lidar device can reach 25 Hz, so it can generate 25 frames of Lidar point cloud data per second. The geometric information between several consecutive Lidar point cloud frames is similar, because the surrounding environment may not change significantly within 0.1 seconds, especially for the environment more than 100 meters away from the car. For example, in the city, there are always high-rise buildings and roads in front of the car more than 100 meters away, or sometimes the car in front drives in the same way as the car, and these objects in the front environment that are a certain distance away from the car may not change significantly within a few frames. Therefore, there is temporal redundancy in consecutive Lidar data frames, and this redundancy can be used to improve the performance of Lidar data coding. Therefore, the method proposed according to the present application is preferably implemented in this case.

[0114] When encoding / decoding a sequence of Lidar data captured by a Lidar device of an autonomous driving car, if only the coarse representation of consecutive Lidar data frames is used separately on a plane , there will be a lot of redundant information. However, the encoded sequence of Lidar data does not take advantage of the redundant information between consecutive point cloud frames. Therefore, according to the present application, by taking into account the coarse representation information of adjacent Lidar data frames, the compression performance of the sequence of (Lidar data captured in an autonomous driving car) can be improved. In addition, by combining the method according to the present application with the above-mentioned coarse representation framework, the compression efficiency can be further improved.

[0115] It is suggested to reduce the temporal redundancy of the coarse representation information between consecutive Lidar data frames to further improve the compression performance of the sequence of Lidar data.

[0116] In practical applications, the vertical angular range of Lidar scanning can reach more than 25°, usually around 40° (for example from -25° to +15°), the vertical angular range of Lidar laser beams can be divided into three parts, i.e. higher laser beams (for example from 8° to 15°, the preferred range can be from 10° to 15°), middle laser beams (for example from -14° to 8°-Δθ, the preferred range can be from -6° to 3°, where Δθ is the angular difference of two laser beams in the vertical direction) and lower laser beams (for example from -14°-Δθ to -25°, the preferred range can be from -20° to -25°), the laser beams of different parts can detect objects at different / distinct distances from the car. The higher laser beams can detect distant / higher objects, such as buildings or trees; the middle laser beams can detect nearby objects, such as cars, people, bicycles, etc. in front; the lower laser beams can detect objects really close to the car, usually the road in front of the car. When the Lidar device works on a moving car, the objects detected by the higher laser beams and the lower laser beams do not suddenly disappear within two adjacent Lidar scanning frames (0.1 s time), then the order difference Δnext between two adjacent point cloud frames can have a correlation, and the difference ΔΔnext between them will be less than Δnext, so when encoding the coarse representation of the current Lidar data frame, encoding ΔΔnext into the bitstream saves bits compared to directly encoding Δnext into the bitstream.

[0117] As shown in Figure 12a and Figure 12b , they show the coarse representations of Lidar point clouds at time t seconds and time t+1 seconds on the plane The points in the middle part represent the points captured by the middle laser beams, and the points in the upper part and the points in the lower part represent the points captured by the higher laser beams and the lower laser beams respectively. Some middle points on the plane may change slightly within 1 second due to the movement of objects or the operation of the car in front, and the slight changes can be found by comparing Figure 12a and Figure 12b . Therefore, it can be observed that the coarse representations of adjacent frames captured within 0.1 seconds will be much smaller, especially for the points detected by the higher laser beams and the lower laser beams. Therefore, the coarse representation information of the previous adjacent lidar point cloud can be used to predict the coarse representation of the current encoded lidar point cloud frame.

[0118] In the preferred embodiment, in order to encode the current Lidar point cloud frame (i-th frame), the order difference Δnext between two consecutive points P n and P n-1 in the previous Lidar point cloud frame (i-1-th frame) can be introducedi-1,Pn , to predict the order difference Dnext between consecutive points P n and P n-1 in the current frame. i,Pn The predicted order difference residual DADnext i,Pn may be encoded into the bitstream or decoded from the bitstream. The predicted order difference residual of a point P n may be obtained by the following equation

[0119] Dnext i,Pn = Dnext i,Pn - Dnext i-1,Pn .

[0120] In a preferred embodiment, the order difference prediction method can be used to encode / decode the current point cloud frame for each point.

[0121] A preferred embodiment of the proposed decoding method is illustrated in Figure 13 . In detail, the steps to decode a point P n of a Lidar point cloud frame from the bitstream are as follows:

[0122] • Obtain the predicted order difference Dnext n between two consecutive points P n-1 and P i-1,Pn in the previous frame, and obtain the order o n-1 (P i ) of the previously decoded point P n-1 in the current frame;

[0123] • Decode the predicted order difference residual DADnext i,Pn from the bitstream using Context-Adaptive Binary Arithmetic Coding (CABAC);

[0124] • Then, based on the predicted order difference residual DADnext i,Pn and the order difference Dnext n of P i-1,Pn in the previous frame, obtain the order difference Dnext n of P i,Pn in the current frame

[0125] Dnext i,Pn = Dnext i-1,Pn + DADnext i,Pn

[0126] • Then, based on the obtained order difference Dnext n of P i,Pn and the order o i (P n-1 ) of the previous decoded point in the current frame, obtain the order o i of the current decoded point in the current frame by the following equation(P n )

[0127] o i (P n )=o i (P n-1 )+Δnext i,Pn =o i (P n-1 )+Δnext i-1,Pn +ΔΔnext i,Pn

[0128] • Then, by ordering the points P n according to their order, the laser (sensor) index l and the azimuth sampling angle are calculated. The coarse representation of the points P n in the plane is constructed and their calculation formula is

[0129] l n =o i (P n ) mod N laser ,

[0130]

[0131] where N laser is the total number of laser beams (or sensors) of the Lidar device.

[0132] • Then, the same process as outlined above is followed to decode the radius residual, the azimuth angle residual and the residuals x res , y res and z res to reconstruct the geometric information (x, y, z) of the point o i (P n ) in the current frame.

[0133] After the decoding of the geometric information and the attribute information of the points in the current frame is completed, the decoding of the next point of the current frame can then continue until the last point in the current frame.

[0134] In the current low latency, low complexity Lidar coding method, for each frame, the information encoded into the bitstream or decoded from the bitstream can include the order of the first point P0 in the current frame and the order difference Dnext Pn of all points in the frame to make a coarse representation. In the proposed method described above, the order of the first point P0 of each frame can be encoded into the bitstream or decoded from the bitstream when coding the sequence of Lidar data.

[0135] In a preferred embodiment, only the ordering of the first point P0 of the first frame in the group of pictures (GOP) is encoded into / from the bitstream, while the ordering of the first points of the following frames in the GOP is not directly encoded into / from the bitstream, only the first point order difference between the first point P0 in the current frame (i-th frame) and the first point P’0 in the previous frame ((i-1)-th frame) is encoded into / from the bitstream. And the equation to calculate the first point order difference can be described as

[0136]

[0137] A group of pictures (often abbreviated as GOP) is a series of consecutive frames in a compressed video or image sequence. It is a fundamental concept in digital video compression, used to organize the sequence of frames for efficient encoding and decoding. A GOP typically starts with an I-frame (intra-coded frame) that is encoded independently of other frames, followed by a series of P-frames (predictively coded frames) and B-frames (bidirectionally predicted coded frames), which are encoded based on information from previous and / or subsequent frames in the GOP. I-frames serve as reference points for subsequent frames, facilitating error recovery and random access in the video stream. The length of a GOP, defined as the number of frames it contains, as well as the pattern of I, P, and B frames within it, can be adjusted based on the specific requirements of the video compression application, balancing between compression efficiency, image quality, and processing complexity.

[0138] In a preferred embodiment, the prediction order difference depends only on the 1-frame information of the previous encoding. In a variant, if more than 1 previous frame has been encoded in the GOP, the prediction order difference Δnext pred of the point in the current frame can be obtained from the average of Δnext between two consecutive points P n and P n-1 in the more than 1 previous encoded frames, so for example, if the order difference of 2 previous frames is used, the prediction order difference Δnext pred of the point in the current frame can be obtained by

[0139] Δnext pred = (w1*Δnext i-1,Pn +w2*Δnext i-2,Pn ),

[0140] where the weights W1 and W2 can be optionally applied, preferably W1+W2=1, and more preferably W1>W2. The prediction order difference residual of the capturing point is then obtained by

[0141] ΔΔnext i,Pn = Δnext i,Pn - Δnext pred = Δnexti,Pn - (Anext i-1,Pn + Anext i-2,Pn ) / 2.

[0142] In a preferred embodiment, a flag (e.g., coarse_inter_prediction_flag) can be used to enable / disable the proposed coarse representation inter prediction method. In another embodiment, the flag can be an inter prediction flag indicating whether the inter prediction method of Lidar data sequence encoding is enabled. If the flag is true, the proposed coarse representation inter prediction method (e.g., the delta prediction method) can be used to encode the coarse representation of the current frame in the current Lidar frame encoding process; otherwise, the proposed coarse representation inter prediction method (e.g., the delta prediction method) will not be used in the current Lidar frame encoding process.

[0143] Reference is now made to Figure 16 FIG. 3 shows a simplified block diagram of an example embodiment of an encoder or decoder 300. The encoder or decoder 300 includes a processor 301 and a memory storage device 303. The memory storage device 303 can store a computer program or application containing instructions that, when executed, cause the processor 301 to perform operations such as those described herein. For example, the instructions can encode and output an encoded bitstream according to the methods described herein or decode a bitstream and output points of a point cloud. It will be appreciated that the instructions can be stored on a non-transitory computer readable medium, e.g., an optical disc, a flash memory device, a random access memory, a hard drive, etc. When the instructions are executed, the processor 301 performs the operations and functions specified in the instructions in order to operate as a specialized processor that implements the described processes. In some examples, such a processor can be referred to as a “processor circuit” or “processor circuitry.”

[0144] It can be appreciated that decoders and / or encoders according to the present application can be implemented in many computing devices, including but not limited to servers, general purpose computers programmed appropriately, machine vision systems, and mobile devices. The decoders or encoders can be implemented by software containing instructions for configuring one or more processors to perform the functions described herein. The software instructions can be stored on any suitable non-transitory computer readable memory, including CDs, RAM, ROM, flash memory, etc.

[0145] It should be appreciated that the decoder and / or encoder described herein, as well as the modules, routines, processes, threads or other software components implementing the described methods / processes for configuring an encoder or decoder, can be implemented using standard computer programming techniques and languages. The present application is not limited to a particular processor, computer language, computer programming conventions, data structures, other such implementation details. Those skilled in the art will recognize that the described processes can be implemented as a part of computer-executable code stored in volatile or non-volatile memory, as part of an application-specific integrated circuit (ASIC), or the like.

[0146] Certain adjustments and modifications can be made to the described embodiments. Therefore, the above-discussed embodiments are to be considered illustrative rather than restrictive. In particular, the embodiments can be freely combined with each other.

Claims

1. A method for decoding three-dimensional positions of points of a point cloud from a bitstream, the method comprising: receiving a bitstream, wherein the bitstream comprises frame information of a sequence of frames, each frame representing point cloud data at a particular time; obtaining the frame information; for a first frame in the sequence, determining points of the point cloud based on decoded data of the first frame; for a subsequent frame, determining points of the point cloud based on at least one previously obtained frame information, wherein the frame information comprises data indicative of relative positions of points in the point cloud; and each point of the point cloud is represented in a multi-dimensional coordinate system, the representation and ordering of the points being based on at least one characteristic of capturing the points; for a subsequent frame, the frame information further comprises a predicted order difference residual, the predicted order difference residual being used to indicate a predicted residual between an order index difference of two consecutive neighboring points in the subsequent frame and an order index difference of corresponding two consecutive neighboring points in a previous frame, wherein each sample comprises one or more points of the point cloud, the order index being according to a decoding order of the point cloud, and an index associated with each sample.

2. The method of claim 1, the method further comprising: determining that the frame information in the bitstream corresponds to a high portion or a low portion of the point cloud, wherein the high portion of the point cloud corresponds to data acquired by laser beams in a vertical direction of 8° to 15°, and the low portion of the point cloud corresponds to data acquired by laser beams in a vertical direction of -14°-Δθ to -25°, wherein Δθ is an angle difference in the vertical direction of two laser beams, the angle indicating a range of scanning in the vertical direction by a radar; continuing to determine points of the point cloud based on at least one previously obtained frame information.

3. The method of claim 1, for subsequent frames, the obtaining of the frame information and the determining of the points of the point cloud are based on at least two previously obtained frame information, wherein, each of the two previously obtained frame information is associated with a weight.

4. The method of claim 3, wherein, frame information from a frame closer to a current frame in the sequence is assigned a higher weight relative to frame information from a frame further away in the sequence.

5. The method according to any one of claims 1 to 4, the bitstream further comprising a flag configured to enable or disable the step of determining points of the point cloud based on the at least one previously obtained frame information, wherein, the flag is at least one bit.

6. The method of any one of claims 1 to 4, wherein, each captured point of the point cloud is represented by a point of a two-dimensional discrete angular plane, wherein a sensor index is a coordinate of a first axis of the two-dimensional discrete angular plane, and an azimuth angle is a coordinate of a second axis of the two-dimensional discrete angular plane, points of the point cloud are ordered according to a dictionary order, first based on the azimuth angle using the dictionary order, then based on the sensor index using the dictionary order, or first based on the sensor index using the dictionary order, then based on the azimuth angle using the dictionary order; the sensor index is an index associated with a sensor that captures at least one point of the point cloud, and the azimuth angle is a capturing angle of the sensor.

7. The method of any one of claims 1 to 4, the method further comprising: determining that the frame information in the bitstream corresponds to a middle portion of the point cloud, wherein the middle portion of the point cloud corresponds to data acquired by laser beams in a vertical direction of -14°-Δθ to 8°; and not continuing to determine points of the point cloud based on the at least one previously obtained frame information.

8. The method of any one of claims 1 to 4, wherein, the point cloud is captured by a rotating light detection and ranging radar head having a plurality of sensors.

9. A method for encoding a three-dimensional point cloud into a bitstream, the method comprising: generating frame information for a sequence of frames, each frame representing point cloud data at a particular time; for a first frame of the sequence, encoding points of the point cloud based on three- dimensional data of the points of the point cloud; for subsequent frames, encoding points of the point cloud based on at least one previously encoded frame information, wherein the frame information comprises data indicative of relative positions of points in the point cloud; and each point of the point cloud is represented in a multi-dimensional coordinate system, the representation and ordering of the points being based on at least one characteristic of the points being captured; for subsequent frames, the frame information further comprises a prediction order difference residual, the prediction order difference residual being used to indicate a prediction residual between an order index difference of two consecutive neighboring points in a subsequent frame and an order index difference of two consecutive neighboring points in a previous frame, wherein each sample comprises one or more points of the point cloud, the order index being according to an encoding order of the point cloud, and an index associated with each sample.

10. The method of claim 9, further comprising: determining that the frame information corresponds to a high portion or a low portion of the point cloud, wherein the high portion of the point cloud corresponds to data acquired by laser beams in a vertical direction of 8° to 15°, and the low portion of the point cloud corresponds to data acquired by laser beams in a vertical direction of -14° - Δθ to -25°, wherein Δθ is an angular difference of two laser beams in the vertical direction, the angle indicating a range of scanning of the radar in the vertical direction; and continuing to encode points of the point cloud based on at least one previously encoded frame information.

11. The method of claim 9, wherein, for subsequent frames, the determination of the frame information and the points of the point cloud is based on at least two previously encoded frame information, each of the two previously encoded frame information being associated with a weight, wherein frame information from a frame closer to a current frame in the sequence is assigned a higher weight relative to frame information from a frame further away in the sequence.

12. The method of any one of claims 9-11, wherein, the bitstream further comprises a flag configured to enable or disable encoding points of the point cloud based on the at least one previously encoded frame information, wherein the flag is at least one bit.

13. The method of any one of claims 9-11, wherein, each captured point of the point cloud is represented by a point of a two-dimensional discrete angular plane, wherein a sensor index is a coordinate of a first axis of the two-dimensional discrete angular plane, and an azimuth angle is a coordinate of a second axis of the two-dimensional discrete angular plane, the points of the point cloud are ordered according to a dictionary order, first based on the azimuth angle using the dictionary order, then based on the sensor index using the dictionary order, or first based on the sensor index using the dictionary order, then based on the azimuth angle using the dictionary order; the sensor index is an index associated with a sensor capturing at least one point of the point cloud, and the azimuth angle is a capturing angle of the sensor.

14. The method of any one of claims 9 to 11, wherein, the point cloud is captured by a rotating light detection and ranging radar head having a plurality of sensors.

15. A decoder for decoding a 3D point cloud from a bitstream, the decoder comprising at least one processor and a memory, wherein, the memory stores instructions that, when executed by the processor, perform the steps of the method according to any one of claims 1 to 8.

16. An encoder for encoding a 3D point cloud into a bitstream, the encoder comprising at least one processor and a memory, wherein, the processor stores instructions that, when executed by the processor, perform the steps of the method according to any one of claims 9 to 14.

17. A computer-readable storage medium comprising instructions which, when executed by a processor, perform the steps of the method according to any one of claims 1 to 8 or 9 to 14.

Citation Information

Patent Citations

  • Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

    US20230290006A1

  • Method and apparatus of encoding / decoding point cloud geometry data captured by a spinning sensors head

    US20230401754A1

  • Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

    WO2023276820A1

  • Point cloud encoding method and apparatus, point cloud decoding method and apparatus, device and storage medium

    WO2024011381A1