Method for encoding and decoding a 3D point cloud, encoder, decoder
The method adapts precision levels and quantization steps based on the mobile platform's orientation to enhance compression efficiency and data fidelity in Lidar point clouds, addressing the challenges of low latency and precision in existing technologies.
Patent Information
- Application Number
- PCT/CN2024/078414
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-23
- Publication Date
- 2025-08-28
AI Technical Summary
Existing point cloud compression technologies, particularly for Lidar data from moving vehicles, face challenges in achieving efficient compression with low latency and high precision, especially in scenarios where different views of the environment require varying levels of detail for accurate navigation and obstacle detection.
A method for encoding and decoding 3D point clouds that adapts precision levels based on the orientation relative to the mobile platform's movement, prioritizing higher precision for critical views like the front and side while using lower precision for less critical views, such as the back, and employing adaptive quantization steps based on angular ranges.
Enhances compression efficiency by optimizing computational resources and data fidelity, ensuring critical environmental features are accurately represented for navigation and obstacle detection while reducing overall data size.
Smart Images

Figure CN2024078414_28082025_PF_FP_ABST
Abstract
Description
Method for encoding and decoding a 3D point cloud, encoder, decoderTechnical Field
[0001] The present disclosure relates to a method for encoding and decoding a 3D point cloud, encoder, and decoder. In particular, relates to point cloud geometry data captured by a spinning sensor head.Background
[0002] As a format for the representation of 3D data, point clouds have recently gained traction as they are versatile in their capability to represent all types of 3D objects or scenes. However, as for all compression schemes, the quality of reconstruction of the points of the point cloud is essential.SUMMARY
[0003] Thus, it is an object of the present disclosure to provide a method for decoding the geometry of a 3D point cloud from a bitstream as well as encoding a 3D point cloud into a bitstream with increased efficiency.
[0004] The problem is solved by a method for decoding according to claim 1, a method for encoding according to claim 9, a decoder according to claim 17, an encoder according to claim 18, a bitstream according to claim 19 and software according to claim 20.
[0005] In a first aspect a method for decoding the geometry of a 3D point cloud from a bitstream is provided, preferably implemented in a decoder. The method includes receiving the bitstream, wherein the bitstream contains view information configured to indicate an orientation relative to the mobile platform’s moving direction (S12) ; obtaining the view information (S14) ; determining points of the point cloud based on the view information (S16) . Thus, this disclosure introduces a comprehensive method for decoding the 3D positions of points from a bitstream, especially focusing on data from a Lidar system mounted on a mobile platform. This process starts by receiving a bitstream containing specific view information, which is pivotal in understanding how objects are positioned relative to the platform's movement. View information may be derived from analyzing the geometrical and spatial properties of the point cloud relative to the Lidar system's orientation and position, especially when mounted on a mobile platform like a vehicle. In general, view information could be obtained through methods such as: Using the Lidar system's orientation and the mobile platform's navigation data (e.g., GPS, IMU) to calculate the relative positions of points. Or employing algorithms that analyze the point cloud data to classify points based on their location relative to a predefined orientation axis, which corresponds to the vehicle's direction of movement. This process involves determining the angular position of each point in the point cloud relative to the Lidar system's forward direction and categorizing these points into different views based on their angular positions. In practice, for example, in urban mapping, when the lidar is placed on a vechicle, this method allows for the accurate placement of buildings and streets within a city model by interpreting the orientation of points, enhancing the model's compression performance.
[0006] Preferably, the view information is configured to indicate at least one of a first view, a second view and a third view. The present disclosure further specifies an enhancement to the decoding method outlined above, focusing on the utilization of view information within the bitstream to indicate at least one specific view out of potentially three distinct views. This approach allows for a flexible interpretation of point cloud data, enabling the system to adapt its processing based on the presence and identification of one or more designated views relative to the mobile platform's orientation. The method does not mandate the identification or processing of all three views simultaneously. Instead, it requires that the decoding process can identify and utilize at least one specified view (e.g., a first view) for enhanced processing. This could mean prioritizing points that are directly in the path of the mobile platform (apotential interpretation of the first view) and applying a specific precision level (e.g., smaller quantization step) to these points for detailed analysis. For example, if the system identifies points belonging to the first view (potentially the front view) , these points could be processed with a higher precision level, ensuring critical features directly ahead of the platform are decoded with greater detail for immediate navigation decisions. Points not explicitly marked as belonging to any specified view may be processed using a default precision level, ensuring efficient use of computational resources while maintaining adequate environmental awareness.
[0007] In scenarios where only one or two views are relevant, such as a vehicle primarily concerned with frontal and side obstacles while moving through a narrow corridor, the method allows for selective focus on these views, enhancing processing efficiency and relevance. The ability to selectively identify and process at least one of the potential views provides a significant technical advantage by allowing the decoding system to adapt its precision and focus dynamically based on operational needs and environmental context. This flexibility ensures that critical areas of the point cloud receive the attention and computational resources needed for high-precision decoding, thereby enhancing the overall performance and responsiveness of the system in diverse operational scenarios.
[0008] Preferably, the first view is a front view, the second view is a side view and the third view is a back view relative to the mobile platform’s moving direction. Thus, the disclosure may further utilize view information to categorize points based on their orientation: front, side, or back relative to the mobile platform’s moving direction. This classification may be crucial for applications like autonomous navigation, where understanding the direction of objects in relation to movement can prioritize data processing, focusing computational power on objects in the path of movement, thus enhancing operational efficiency and safety.
[0009] Preferably, the front view indicates points located within an angular range from 45 degrees to the left of the direct front to 45 degrees to the right of the direct front of the mobile platform’s moving direction; the side view indicates points located from 45 degrees to 60 degrees to the left of the direct front of the mobile platform’s moving direction and from 45 degrees to 60 degrees to the right of the direct front of the mobile platform’s moving direction if the field of view of Lidar device is 120 degrees, and if the field of view of Lidar device is 360 degrees (rotating lidar) , the side view indicates points located from 45 degrees to N (N can be a value in range 60≤N≤90) degrees to the left of the direct front of the mobile platform’s moving direction and from 45 degrees to N (N can be a value in range 60≤N≤90) degrees to the right of the direct front of the mobile platform’s moving direction; the back view indicates all points not included in the front view or side view angular ranges. The "direct front" refers to the orientation or direction directly ahead of a mobile platform, aligned with its forward motion. This is the primary axis along which the platform advances, and it serves as a reference point for determining the relative positions of objects in the environment. In the context of Lidar point cloud coding, identifying the direct front helps in classifying point data based on their spatial relation to the movement direction, enabling precise encoding and decoding of objects' positions for navigation, obstacle detection, and mapping applications. Thus, the disclosure further defines the angular ranges for the front, side, and back views more precisely, allowing for a nuanced approach to point cloud decoding. This is particularly useful in applications such as obstacle detection systems in vehicles, where accurately distinguishing between objects directly ahead and those to the side can dictate the vehicle's response to potential hazards, significantly improving the system's effectiveness.
[0010] Preferably, the method further comprising: determining the view information is configured to indicate the front view; determining points of the point cloud according to a precision level A; determining the view information is configured to indicate the side view; determining points of the point cloud according to a precision level B; wherein A > B and greater precision level indicating better quality of the reconstruction of the points. Thus, the method introduces the concept of assigning different precision levels to the decoded points based on their view, each point or each group of points may be associated with a respective view, then prioritizing points in the front view with a higher precision level. The precision level may determine the granularity at which the point cloud data is quantized, where a higher precision level may correspond to a smaller quantization step size. This smaller step size allows for a more detailed and accurate representation of the point cloud by minimizing the loss of information during the quantization process. Essentially, as the precision level increases, the quantization process preserves more detail by using finer intervals to code the data, resulting in higher fidelity of the reconstructed point cloud. Higher precision levels may correspond to a finer spatial resolution in the point cloud, allowing for more detailed and nuanced representation of objects and features in the environment. This means that at higher precision levels, the Lidar system can distinguish between closely spaced objects more effectively. Precision level could also refer to the density of data points within a given area of the point cloud. A higher precision level might result in a higher density of points, offering a more comprehensive coverage and reducing gaps in data that could lead to inaccuracies in object detection and mapping. Beyond the geometric positioning of points, precision level could affect the accuracy of associated attributes, such as reflectivity or color. Higher precision levels might enable more accurate capture and representation of these attributes, improving the fidelity of the point cloud for applications requiring detailed attribute analysis. At higher precision levels, the encoding and decoding process might incorporate more sophisticated noise reduction techniques, enhancing the clarity and usability of the point cloud data by filtering out irrelevant or erroneous signals that could obscure important features. In practical terms, this could be applied in robotic vision systems where identifying and interacting with objects directly in the robot's path requires greater detail, while less precision can be afforded to peripheral areas, optimizing processing power and enhancing interaction capabilities.
[0011] Preferably, the method further comprises: determining the view information is configured to indicate the back view; determining points of the point cloud according to a precision level C; wherein A > C. Thus, the method extends the differentiation in precision levels to include the back view, albeit at a lower precision than the front. This could be particularly relevant in surveillance drones, where monitoring areas behind the drone might require less detail than those ahead or to the sides. By assigning a lower precision level to the back view, resources can be efficiently allocated, extending operational duration while maintaining situational awareness.
[0012] Preferably, wherein A > B > C. Thus, the method further sets a hierarchy of precision levels (A> B > C) applied to different views in the point cloud data, ensuring a strategic allocation of computational resources. This hierarchical precision management facilitates more detailed processing of points that are most critical for the application, such as the front view for a vehicle's navigation system, where obstacles need to be identified with the highest clarity. Meanwhile, it allows for less detailed processing of less critical areas, such as the vehicle's rear view, optimizing overall system efficiency by prioritizing computational resources where they are most needed for accurate and safe navigation. This approach ensures that the encoding and decoding processes are tailored to enhance the utility and performance of the Lidar system in real-world applications, such as autonomous driving, where understanding and reacting to the immediate environment is paramount.
[0013] Preferably, at least two points identified within the point cloud differ in their azimuth, elevation, or radial distance from the Lidar sensor, the at least two points are processed with distinct precision levels. Thus, the present disclosure further introduces a nuanced approach to processing point cloud data, focusing on the differentiation of precision levels for at least two points based on their spatial characteristics. Specifically, this claim emphasizes that when two points within the point cloud exhibit significant differences in their azimuth (horizontal angle) , elevation (vertical angle) , or radial distance (distance from the Lidar sensor) , they should be coded with distinct precision levels. The criterion for 'significant difference' is quantified by thresholds –if the distance between two points exceeds a specified threshold, or if their differences in azimuth, elevation, or radial distance surpass respective threshold values (threshold 1 for azimuth, threshold2 for elevation, and threshold3 for radial distance) , then these points are to be processed differently in terms of precision level during the coding. For example, in a complex urban environment, a Lidar sensor mounted on an autonomous vehicle might detect two pedestrians on the sidewalk, one closer to the vehicle than the other. If the difference in their radial distances from the sensor exceeds a certain threshold, the encoding process could assign a higher precision level to the closer pedestrian, ensuring that this nearer, potentially more critical object is represented with greater detail in the point cloud. This method allows for adaptive data processing, focusing higher precision on points that, due to their position or the significant difference from others, might be of more immediate relevance to the platform's operational needs. The technical effect of this approach is the optimized allocation of computational resources and storage space. By varying precision levels based on spatial differences, the system can prioritize data fidelity where it matters most, enhancing the quality of the environmental model used for navigation, obstacle avoidance, and decision-making, while efficiently managing data volume and processing load.
[0014] Preferably, the precision level applied to points within the point cloud data is determined based on their azimuth, elevation, or radial distance from the Lidar sensor. Thus, the present disclosure further refines the method by specifying that the precision level to be applied during the encoding or decoding of point cloud data is directly determined by the points' spatial characteristics –their azimuth, elevation, or radial distance from the Lidar sensor. This claim reinforces the adaptive nature of the processing, ensuring that each point's representation in the point cloud is customized based on its precise location relative to the sensor. This customization ensures that the point cloud not only accurately reflects the scanned environment but does so in a way that intelligently balances detail and computational efficiency. For instance, objects that are farther away and less likely to immediately impact the mobile platform's navigation may be coded with a lower precision level, reducing data size without significantly impacting the quality of actionable insights derived from the data. Conversely, objects that are closer or positioned in critical areas (directly ahead, for example) can be coded with higher precision, ensuring that the data used for making immediate navigational decisions is as detailed and accurate as possible.
[0015] In another aspect of the present disclosure, an encoding method is provided, the method includes: determining view information indicating an orientation relative to the mobile platform’s moving direction (S22) ; encoding points of the point cloud into the bitstream based on the view information (S24) . The embodiments described with reference to the decoding method also apply to the corresponding encoding method.
[0016] In another aspect of the present disclosure, an encoder is provided for encoding a 3D point cloud into a bitstream. The encoder comprises a memory and a processor, wherein instructions are stored in the memory, which when executed by the processor, performs the steps of the method for encoding described before.
[0017] In another aspect of the present disclosure, a decoder is provided for decoding a 3D point cloud from a bitstream. The decoder comprises a memory and a processor, wherein instructions are stored in the memory, which when executed by the processor, performs the steps of the method for decoding described before.
[0018] In another aspect of the present disclosure, a bitstream is provided, wherein the bitstream is encoded by the steps of the method for encoding described before.
[0019] In another aspect of the present disclosure, a computer-readable storage medium is provided comprising instructions to perform the steps of the method for encoding a 3D point cloud into a bitstream as described above.
[0020] In another aspect of the present disclosure, a computer-readable storage medium is provided comprising instructions to perform the steps of the method for decoding a 3D point cloud from a bitstream as described above.
[0021] FIGURES
[0022] In the following the present disclosure is described in more detail with reference to the accompanying figures.
[0023] Figure 1 illustrates a spinning Lidar head that includes several spinning lasers that probe the environment;
[0024] Figure 2 illustrates the elevation angle θ of a spinning laser;
[0025] Figure 3 illustrates 2D angular representation of the points acquired by a spinning Lidar;
[0026] Figure 4 illustrates acquired points on the discrete representation;
[0027] Figure 5 illustrates 3D xyz coordinates and angle-based coordinates or
[0028] Figure 6 illustrates acquisition order in azimuthal angle and laser index λ;
[0029] Figure 7 illustrates points in azimuthal angle and laser index λ by a real;
[0030] Figure 8 illustrates the ordering of points in coarse azimuthal angle and laser index λ;
[0031] Figure 9 illustrates the representation of the point cloud by differences Δnext for a first lexicographic order
[0032] Figure 10 illustrates an overview of an encoding method implementing an example;
[0033] Figure 11 illustrates an overview of a decoding method implementing an example;
[0034] Figures 12a and 12b illustrate a coarse representation of two Lidar point cloud frames with 1-second time differences;
[0035] Figure 13 illustrates an overview of the Lidar point cloud decoding method according to the present disclosure;
[0036] Figure 14 illustrates a block diagram of the decoding method according to the present disclosure;
[0037] Figure 15 illustrates a block diagram of the encoding method according to the present disclosure;
[0038] Figure 16 illustrates a decoder / encoder according to the present disclosure;
[0039] Figure 17 illustrates an encoder for the predictive tree as in G-PCC,
[0040] Figure 18 illustrates a decoder for the predictive tree as in G-PCC;
[0041] Figure 19 illustrates an example of Lidar device whose field of view is 120 degrees.DETAILED DESCRIPTION
[0042] As a format for the representation of 3D data, point clouds have recently gained traction as they are versatile in their capability to represent all types of 3D objects or scenes. Therefore, many use cases can be addressed by point clouds, among which are
[0043] · movie post-production,
[0044] · real-time 3D immersive telepresence or VR / AR applications,
[0045] · free viewpoint video (for instance for sports viewing) ,
[0046] · Geographical Information Systems (aka cartography) ,
[0047] · culture heritage (storage of scans of rare objects into a digital form) ,
[0048] · Autonomous driving, including 3D mapping of the environment and real-time Lidar data acquisition
[0049] A point cloud is a set of points located in a 3D space, optionally with additional values attached to each of the points. These additional values are usually called point attributes. Consequently, a point cloud is combination of a geometry (the 3D position of each point) and attributes. Attributes may be, for example, three-component colors, material properties like reflectance and / or two-component normal vectors to a surface associated with the point. Point clouds may be captured by various types of devices like an array of cameras, depth sensors, Lidars, scanners, or may be computer-generated (in movie post-production for example) . Depending on the use cases, points clouds may have from thousands to up to billions of points for cartography applications. Raw representations of point clouds require a very high number of bits per point, with at least a dozen of bits per spatial component X, Y or Z, and optionally more bits for the attribute (s) , for instance three times 10 bits for the colors. Practical deployment of point-cloud-based applications requires compression technologies that enable the storage and distribution of point clouds with reasonable storage and transmission infrastructures. Compression may be lossy (like in video compression) for the distribution to and visualization by an end-user, for example on AR / VR glasses or any other 3D-capable device. Other use cases do require lossless compression, like medical applications or autonomous driving, to avoid altering the results of a decision obtained from the analysis of the compressed and transmitted point cloud. Until recently, point cloud compression (aka PCC) was not addressed by the mass market and no standardized point cloud codec was available. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11, also known as Moving Picture Experts Group or MPEG, has initiated work items on point cloud compression. This has led to two standards, namely
[0050] · MPEG-I part 5 (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC)
[0051] · MPEG-I part 9 (ISO / IEC 23090-9) or Geometry-based Point Cloud Compression (G-PCC)
[0052] Both V-PCC and G-PCC standards have finalized their first version in late 2020 and will soon be available to the market. The V-PCC coding method compresses a point cloud by performing multiple projections of a 3D object to obtain 2D patches that are packed into an image (or a video when dealing with moving point clouds) . Obtained images or videos are then compressed using already existing image / video codecs, allowing for the leverage of already deployed image and video solutions. By its very nature, V-PCC is efficient only on dense and continuous point clouds because image / video codecs are unable to compress non-smooth patches as would be obtained from the projection of, for example, Lidar-acquired sparse geometry data.
[0053] The G-PCC coding method has two schemes for the compression of the geometry. The first scheme is based on an occupancy tree (octree / quadtree / binary tree) representation of the point cloud geometry. Occupied nodes are split down until a certain size is reached, and occupied leaf nodes provide the location of points, typically at the center of these nodes. By using neighbor-based prediction techniques, high level of compression can be obtained for dense point clouds. Sparse point clouds are also addressed by directly coding the position of point within a node with non-minimal size, by stopping the tree construction when only isolated points are present in a node; this technique is known as Direct Coding Mode (DCM) . The second scheme is based on a predictive tree, each node representing the 3D location of one point and the relation between nodes is spatial prediction from parent to children. This method can only address sparse point clouds and offers the advantage of lower latency and simpler decoding than the occupancy tree. However, compression performance is only marginally better, and the encoding is complex, relatively to the first occupancy-based method, intensively looking for the best predictor (among a long list of potential predictors) when constructing the predictive tree. In both schemes, attribute (de) coding may be performed after complete geometry (de) coding, leading to a two-pass coding. Thus, low latency is obtained by using slices that decompose the 3D space into sub-volumes that are coded independently, without prediction between the sub-volumes. This may heavily impact the compression performance when many slices are used.
[0054] An important use case is the transmission of Lidar data acquired by a moving vehicle. This usually requires a simple low-latency embarked encoder. Simplicity is required because the encoder is likely to be deployed on computing units which perform other processing in parallel, such as (semi-) autonomous driving, thus limiting the processing power available to the point cloud encoder. Low latency is also required to allow for fast transmission from the car to a cloud in order to have a real-time view of the local traffic, based on multiple-vehicle acquisition, and take adequate fast decisions based on the traffic information. While transmission latency can be low enough by using 5G, the encoder itself shall not introduce too much latency due to coding. Also, compression performance is extremely important since the flow of data from millions of cars to the cloud is expected to be extremely heavy. Combining encoder and decoder simplicity, low latency and compression performance is still a problem that has not been satisfactorily solved by existing point cloud codecs.
[0055] Specific priors related to the acquisition of Lidar data have been already exploited in G-PCC and have led to very significant gains of compression. A first technique concerns the vertical angle (relative to the horizontal ground) of acquisition from a spinning Lidar (depicted in Figure 1) The Lidar head of Figure 1 may spin around the vertical axis (dotted line) to capture geometry data of an object: this angle θ is fixed, as shown on Figure 2.
[0056] Practically, the representation of the Lidar-acquired point cloud is not 3D but quasi 2D in the spherical coordinates where r3D is 3D the distance of a point from the Lidar’s center. It is the linear distance measured directly from the Lidar sensor to a point in space, irrespective of the horizontal or vertical angle. The parameter r is critical for determining the absolute distance of objects from the Lidar sensor. It is fundamental in generating accurate three-dimensional maps and models of the environment, as it provides the direct spatial relationship between the sensor and the objects in its field of view. is an azimuthal angle of a Lidar head’s spin relative to a referential. The azimuthal angle refers to the horizontal angular displacement measured from a defined origin. Specifically, it is the angle between the projection of the point in question onto the horizontal plane and a fixed reference direction on that plane. In Lidar systems, is used to determine the horizontal orientation of a scanned point relative to the Lidar device. It is crucial in calculating the precise horizontal positioning of objects in the scanned environment. θ is the elevation angle of a sensor of the Lidar head relative to a horizontal referential plane. The elevation angle signifies the vertical angular displacement from a defined horizontal plane. It is the angle between the line from the point to the Lidar's center and the horizontal plane, measured in the vertical plane containing the line. θ is essential for determining the vertical positioning of objects in Lidar scans. It allows for the calculation of the height or altitude of points in relation to the Lidar device, contributing to the creation of three-dimensional representations of the scanned area.
[0057] G-PCC has gone even further by exploiting a second technique that makes benefit of the regularity of laser sensing while the Lidar is spinning, as depicted in Figure 3. A regular distribution along the azimuthal angle has been observed on Lida r acquired data. This regularity may be used to obtain a quasi 1D representation of the point cloud where, up to noise, preferably only the radius r3D belongs to a continuous range of value while the angles and θ may take only a discrete number of values. Basically, one may represent the point cloud geometry on a 2D discrete angular plane, see Figure 4, together with a radius value for each point. This quasi 1D property has been exploited in G-PCC in both the occupancy tree and the predictive tree by predicting, in the spherical coordinate, the location of a current point relative to an already coded point by using the discrete nature of angles.
[0058] Practically, the occupancy tree may use DCM intensively and entropy codes the direct location of points within a node by using a context-adaptive entropy coder. Contexts may be obtained from the local conversion of the point location into angular coordinates. The predictive tree directly codes the angular coordinates where r2D is the projected radius on the horizontal xy plane (see Figure 5) , before converting into (x, y, z) and then coding a xyz residual to tackle the errors of coordinate conversion, the approximation of laser angle and noise (xyz residual refers to the difference between the actual coordinates of a point and its estimated coordinates) .
[0059] As explained above there are mainly two types of coding structure in the art, namely the occupancy tree and the predictive tree.
[0060] However, according to existing coding structures, there are still redundancies in coding and the compression performance of the point cloud could be further improved. Therefore, a decoding method according to Figure 14 and an encoding method according to Figure 15 are proposed.
[0061] The decoding method comprises the following steps:
[0062] receiving the bitstream, wherein the bitstream contains view information configured to indicate an orientation relative to the mobile platform’s moving direction (S12) ; obtaining the view information (S14) ; determining points of the point cloud based on the view information (S16) .
[0063] The encoding method comprises the following steps:
[0064] determining view information indicating an orientation relative to the mobile platform’s moving direction (S22) ;
[0065] encoding points of the point cloud into the bitstream based on the view information (S24) .
[0066] Nowadays, there are more and more cars mounted with several Lidar devices to capture the environment of cars, however, the massive 3D point cloud data captured by Lidar devices in cars will take much bandwidth / storage when transmitting it / store it, so the compression of 3D point cloud data is needed. In real applications, the captured Lidar point cloud data may not be viewed by human eyes directly in most cases, it may be transmitted for later machine analysis tasks, like objects detection, objects classification, segmentation, etc, which is very different from the applications of video. One example of classification tasks that use 3D point clouds is the classification of tree species, roof types, road types or pedestrians to guide automatic driving. For automatic driving, the detection performance of objects in front of the car may have a greater influence on driving security than that of objects in the back of the car. Thus, to further improve the compression performance of the 3D Lidar point cloud and keep the driving security at the same time, the reconstructed precision of decoded data belonging to the back environment of the car needs to be reduced.
[0067] Thus, it is proposed according to the appended claims to introduce an adaptive quantization method based on different angle areas in FOV to reduce the coded data size. The main idea is to use a larger quantization step for coding points (geometry position and attribute (reflectance) ) not in the front view of a car than the quantization step used for points in the front view.
[0068] Thus, according to the proposed methods, the coding efficiency could be improved by considering adjacent lidar data frame (s) .
[0069] In real applications of Lidar devices, there are mainly two kinds of Lidar devices according to their FOV (Field of view) , one kind of Lidar device can capture 360 degrees in a horizontal direction, like a rotating spinning Lidar, and the other kind of Lidar device can only capture a horizontal view of around 120 degrees, and the Figure 19 shows an example of Lidar device, FOV of 120 degrees is illustrated, and all laser beams in Figure 19 belong to the same laser (θ) .
[0070] When a car is driving along a road, the detection system is more sensitive to objects in the front view (for example ) of the car than objects on the left side view (for example ) and the right side view (for example ) and on the back view (if the laser has an FOV of 360 degrees) of the car. Thus, in one embodiment, the quantization step for coding points on the left, right and back view of a car is larger than that for coding points in the front view of a car. Assume the quantization step for the geometry position of points in the front view are qgf, the quantization step for the geometry position of points in the side view is qgs, and the quantization step for the geometry position of points in the back view are qgb, then their relationship can be qgf<qgs, qgf<qgb.
[0071] Also, in most applications, compared to the back view (if the laser has an FOV of 360 degrees) of a car, the detection system may be more sensitive to objects on the side view (for example ) , thus the quantization step for coding points on the side view of a car may be larger than that for coding points in the back view of a car, then their relationship can be further refined to be qgf<qgs <qgb.
[0072] In another embodiment, the quantization step qaf for attribute (e.g. reflectance) of points in front view, the quantization step qas for attribute of points in side view, and the quantization step qab for attribute of points in back view satisfy the relationship qaf<qas <qab.
[0073] Wherein reflectance can be estimated by the ratio between transmitting power and receiving power of a laser probing a point of an object, and reflectance can identify the characteristics of surface of an object.
[0074] A preferred embodiment of the proposed geometry encoding method follows the steps below:
[0075] Firstly, determine the quantization step for geometry position of points in front, side and back views, which are qgf, qgs, qgb, and they are encoded into bitstream. Then, transform positions of points from cartesian coordinates (x, y, z) to angular coordinates and obtain prediction of points using the method in the prediction tree coding method / using the method in low latency low complexity lidar codec L3C2 coding method, and then get the geometry residuals Then, quantize the coordinate position residuals of each point based on their estimated azimuthal angle information, if their azimuthal angle is estimated within the side views, then the quantization step qgs may be used to quantize geometry residuals; and if their azimuthal angle is estimated within the back views, then the quantization step qgb may be used to quantize geometry residuals; and if their azimuthal angle is estimated within the front views, then the quantization step qgf may be used to quantize geometry residuals. Preferably, the quantization step qgfis set as 1 to losslessly code points in front view. Finally, the quantized geometry residuals may be encoded into bitstream.
[0076] A preferred embodiment of the proposed geometry decoding method follows the steps below:
[0077] Firstly, decode from bitstream the quantization step for the geometry position of points in front, side and back views, which are qgf, qgs, qgb. Then, decode the quantized geometry residuals from bitstream for each point, and then dequantize the coordinate position residuals of each point based on their estimated azimuthal angle range information. If their azimuthal angle is estimated within the side views, then the quantization step qgs may be used to dequantize geometry residuals; and if their azimuthal angle is estimated within the back views, then the quantization step qgb may be used to dequantize geometry residuals; and if their azimuthal angle may be estimated within the front views, then the quantization step qgf may be used to dequantize geometry residuals. Then, obtain a prediction of points using the method in the prediction tree coding method / using the method in L3C2 coding method, and then add the dequantized geometry residuals to get reconstructed positions
[0078] In one embodiment, for example, when using L3C2 to code a Lidar point cloud captured with a non-360 degree Lidar device, the coarse angle obtained from order difference can be used to estimate azimuthal angle range information of a point (the details of coarse representation will be explained in more detail in the sections below) . When the is smaller than a threshold Th1 or larger than a threshold Th2, then the current point’s azimuthal angle may be estimated to be in the side view of the car, and the quantization step qgs may be used to dequantize geometry residuals of the current point, wherein the thresholds Th1 and Th2 are based on FOV of the device in the horizontal direction and front view range FV definition, as is shown in Figure 19.
[0079] Wherein is the parameter of Lidar device, which represents the azimuthal sampling interval of Lidar device, and the front view range FV may be set as 90°.
[0080] In another embodiment, when using the predictive tree method to code Lidar point cloud, there is no coarse angle generated in the coding process, at decoder side, it doesn’t know the azimuthal information of the current decoded point before dequantize its geometry residuals To estimate the azimuthal information of the current decoded point, the previous coded point’s azimuthal angle can be used to estimate the azimuthal range of the current decoded point. If is smaller than a threshold Th3 or larger than a threshold Th4, then the current point’s azimuthal angle may be estimated to be in the side view of the car, and quantization step qgs may be used to dequantize geometry residuals of the current point.
[0081] and front view range FV may be set as 90°
[0082] In lossy coding of the current Lidar point cloud, there may be quantization steps q along each dimension among r, and θ, and in the proposed method, the quantization steps qgf of points in the front view may be set to be quantization steps q in the current Lidar data coding method in G-PCC / L3C2, and qgf may include quantization steps qr , qθ along three directions of r, and θ, respectively, and they are encoded / decoded into / from bitstream during the encoding / decoding process. For quantization steps qgs of points in the side view, at least one scaling parameter (m, p, k) may be introduced to determine quantization steps along three directions of r, and θ. qgs_r=m*qr, m>1 qgs_θ=k*qθ, k>1
[0083] And in one embodiment, the scaling parameters (m, p, k) are the same, m=p=k,
[0084] In one embodiment, the scaling parameters are encoded / decoded into / from bitstream, and in another embodiment, quantization steps qgs of points in the side view are encoded / decoded into / from the bitstream. For quantization steps qgb of points in the back view, at least one scaling parameter (u, w, v) may be introduced to determine quantization steps along three directions of r, and θ. qgb_r=u*qr, u>1 qgb_θ=w*qθ, w>1
[0085] In one embodiment, the scaling parameters (u, w, v) are the same, u=w=v,
[0086] In one embodiment the scaling parameters are encoded / decoded into / from bitstream, and in another embodiment, quantization steps quantization steps qgb of points in the back view are encoded / decoded into / from the bitstream.
[0087] In lossy coding of the current Lidar point cloud, there may be quantization steps qr for reflectance, and in the proposed method, the quantization steps qaf of points in the front view may be set to be quantization steps qr in the current Lidar data coding method in G-PCC / L3C2, and the quantization step qaf may be encoded / decoded into / from the bitstream. And for the quantization step qas of points in the side view, one scaling parameter t may be introduced to determine its quantization steps. qas=t*qr, t>1
[0088] In one embodiment, the scaling parameter t may be encoded / decoded into / from bitstream, and in another embodiment, the quantization step qas may be encoded / decoded into / from the bitstream.
[0089] For the quantization step qab of points in the back view, one scaling parameter n may be introduced to determine its quantization steps, qab=n*qr, n>1
[0090] In one embodiment, the scaling parameter n may be encoded / decoded into / from the bitstream, and in another embodiment, the quantization step qab may be encoded / decoded into / from the bitstream.
[0091] The following details on the angular implementation of the predictive tree in G-PCC will be explained.
[0092] In the predictive tree encoder, see Figure 17, when using angle-based coding for Lidar data, a point (x, y, z) is projected onto angular coordinates where, as defined previously, r2D is the radius of the xy plane projection and is the azimuthal angle; and where θidx is a “laser index” . Practically, θidx may be the index of the closest elevation angle in a list of possible elevation angles. This list may be part of the parameters of the codec and may be signaled in the bitstream in the geometry parameter set.
[0093] In the following description, unless otherwise specified, *r2D*and *r*may be used interchangeably. In G-PCC predictive trees, the radius r2D value may be quantized to fixed point precision Δr2D=2^M *quantization step
[0094] where M is a parameter of the codec that may be signaled in the bitstream in the geometry parameter set. Typically, M may be 0 for lossless coding; and for lossy geometry coding, M is an integer larger than 0.
[0095] In G-PCC predictive tree, the angle value is quantized using a quantization step
[0096] where N is a parameter of the codec that may be signaled in the bitstream in the geometry parameter set. Typically, N may be 17 for lossless coding.
[0097] The transformation from cartesian coordinates (x, y, z) to angular coordinates applies as r2D = round (sqrt (x*x + y*y) / Δr2D ) (1)
[0098] where round () is the rounding operation to the nearest integer value, sqrt () is the square root function and atan2 (y, x) is the arc tangent applied to y / x. Concerning θidx, the encoder may search for the “laser index” i for which the absolute value of tan(θi) *sqrt (x*x + y*y) –zoffset_i –z is minimized, where tan (θi) is the tangent of the θi angle of the i-th laser, and zoffset_i is a (vertical) offset of the i-th later to the origin. Both tan (θi) and zoffset_i may be signaled, using a fixed-point precision, in the geometry parameter set for each “laser index” i.
[0099] The inverse transformation from angular coordinate to cartesian coordinates (xpred, ypred, zpred) applies as: r = r2D *Δr2D (3) zpred = round (tan (θθidx) *r -zoffset_θidx ) (6)
[0100] A prediction of the point may be derived from the known ancestors (parent, and up to great-grand-parent points / nodes) of the point in the prediction tree. A prediction mode index may be signaled in the bitstream to inform the decoder of the selected prediction mode among a few allowed prediction modes. The predictor obtained by using the selected prediction mode can be refined into by additionally signaling in the bitstream a (positive or negative) integer k representing the number of elementary steps to be added to the taken predictor
[0101] The elementary step is coding parameter that may be signaled in the bitstream in the geometry parameter set with some fixed-point precision The encoder may derive from the frequencies at which a lidar head is performing acquisition at the different elevation angles, for example from the number of probing per head turn.
[0102] The residual between the original angular coordinates and the prediction
[0103] May be losslessly coded in the bitstream. Then, the reconstructed point in angular coordinates
[0104] May be inverse transformed into the Cartesian space as
[0105] and a second residual (xres, yres, zres) between the original point and this Cartesian reconstruction is obtained (xres, yres, zres) = (x, y, z) - (xpred, ypred, zpred)
[0106] and encoded into the bitstream. The second residual coding may be lossless, when x, y, z quantization step are equal to the original point precision (typically 1) , or lossy when quantization steps are larger than the original point precision (typically quantization steps larger than 1) .
[0107] Reconstructed cartesian coordinates (xdec, ydec, zdec) , as obtained by a decoder, may be computed as (xdec, ydec, zdec) = (xpred, ypred, zpred) + (xres, yres, zres)
[0108] and may be used by the encoder for example for ordering (decoded) points before attribute coding.
[0109] The decoder is shown in Figure 18 and essentially some steps are already described for the encoder.
[0110] Once the predictor index is decoded, the prediction may be built as in the encoder. Optionally, the number k of elementary steps to be added to the predictor is obtained from the bistream and is updated to as in the encoder. Then, the residual may be decoded from the bitstream and added to the prediction to obtain the reconstructed angular coordinates
[0111] These reconstructed angular coordinates may be inverse transformed to the Cartesian space to obtain (xpred, ypred, zpred) . Finally, the second residual (xres, yres, zres) may be decoded, inverse quantized and added to the Cartezian predictor to obtain the decoded Cartesian coordinates (xdec, ydec, zdec) . (xdec, ydec, zdec) = (xpred, ypred, zpred) + (xres, yres, zres)
[0112] As discussed previously, the proposed method could also be used in combination with the coarse representation of the point cloud. In the following, the present disclosure will be introduced in combination with the coarse representation framework, in particular, some basic concepts of coarse representation will be explained and the embodiments can be freely combined.
[0113] In some embodiments, all lasers (sensors) may be coded at once by using the order of acquisition, as shown in Figure 6 in the plane. Due to the regular rotation of the Lidar’s head and the continuous acquisition with fixed time intervals by each laser, the azimuthal distance between two points probed by the same laser is a multiple of an elementary azimuthal shift
[0114] Instead of coding the point location directly, a coarse representation may be coded first, based on the acquisition priors. For example, one may use a coarse representation of the point cloud geometry to order the point using a lexicographic order (aka dictionary order) first in the coarse azimuthal angle and second in the laser index (or the sensing elevation angle index) λ or inversely.
[0115] Schematically, the points may be acquired in the order already shown in Figure 7 in the plane. Due to the regular rotation of the Lidar’s head and the continuous acquisition with fixed time interval by each laser, the azimuthal distance between two points probed by the same laser is a multiple of an elementary azimuthal shift Practically, not all points are acquired, i.e. the laser beam may not be reflected, there is acquisition noise and the laser may not be all perfectly aligned. Real data looks as shown in Figure 7.
[0116] The coarse angle may be simply obtained by the quantization of as follows
[0117] and the order index o (P) of a point P may be obtained by
[0118] where Nlaser is the number of lasers (i.e., the maximum of λ) and λ is the index of the laser index, in [0, Nlaser -1] that has acquired point P. The codec encodes the points P following their order o (P) monotonously, for example using an ascending order. Thus, points may be coded in the order depicted in Figure 8.
[0119] The coarse representation in the plane may be coded by
[0120] · the number of points Npoints
[0121] · the value of for the first acquired point
[0122] · the Npoints-1 successive differences Δnext between a current point and a next point as sorted by the lexicographic order, as depicted in Figure 9.
[0123] The coarse representation consists in successive differences Δnext and the compression of the coarse representation is essentially based on the compression of the successive positive values Δnext.
[0124] An overview of an encoding method based on the above-described method may be shown in Figure 10.
[0125] Firstly, the encoder may convert the xyz point location into a laser (sensor) index λ, a coarse angle and a radius r2D thanks to the knowledge of the Lidar sensor setup. Differences Δnext are determined and the method is applied to the coding of the Δnext’s into the bitstream, as well as useful information on the sensor setup (like and laser elevation angles for example) .
[0126] Secondly, a reconstructed azimuthal angle may be computed. It can be obtained directly from the dequantization of the coarse angle. Optionally, a residual may be computed as the difference and encoded into the bitstream. The residual may be quantized into before encoding. In this case, the reconstructed azimuthal angle may be obtained by
[0127] where IQ stands for the inverse quantization process.
[0128] The radius r (here r2D) may also be coded, optionally after quantization into Q (r) . It is inverse quantized to obtain a reconstructed radius rrec = IQ (Q (r) ) .
[0129] Thirdly, the reconstructed azimuthal angle and the reconstructed radius rrec may be converted back to xy to obtain an estimation of the x and y location of the point.
[0130] Residuals xres and yres relative to the original point location xy may be computed xres = x -xestim, yres = y -yestim,
[0131] and encoded into the bitstream.
[0132] Fourthly and finally, a vertical estimate zestim may be obtained from the laser angle θ (λ) by zestim = rrec tan (θ (λ) )
[0133] and a residual zres relative to the original point location z may be computed zres = z -zestim
[0134] and encoded into the bitstream.
[0135] An overview of the associated corresponding decoding method is shown in Figure 11. It is straightforward once the encoding process is understood.
[0136] Firstly, the decoder may decode useful information on the sensor setup (like and laser elevation angles for example) and differences Δnext, by using the proposed method, from the bitstream. Then, the values of the laser index λ and the coarse angle may be obtained from Δnext as explained in the following.
[0137] Secondly, the reconstructed azimuthal angle may be computed. It may be obtained directly from the dequantization of the coarse angle. Optionally, an azimuthal residual may be decoded from the bitstream. The decoded residual may be a quantized version of the residual and the reconstructed azimuthal angle is obtained by
[0138] where IQ stands for the inverse quantization process.
[0139] The radius r (here r2D) may also be decoded. The coded radius may be a quantized version Q (r) of the radius. It is inverse quantized to obtain a reconstructed radius rrec = IQ (Q (r) ) .
[0140] Optionally, the radius may be predicted (by a preceding coded radius for example) and a radius residual may be coded instead of the radius.
[0141] Thirdly, the reconstructed azimuthal angle and the reconstructed radius rrec may be converted back to xy to obtain an estimation of the x and y location of the point.
[0142] Residuals xres and yres may be decoded from the bitstream and decoded horizontal location xdec and ydec of the point may be computed xdec = xestim + xres, ydec = yestim + yres.
[0143] Fourthly and finally, a vertical estimate zestim may be obtained from the laser angle θ (λ) by zestim = rrec tan (θ (λ) ) ,
[0144] a residual zres may be decoded from the bitstream and the decoded vertical location zdec of the point is computed zdec = zestim + zres .
[0145] In real applications of Lidar, like automatic driving, Lidar device scanning speed can reach 25HZ, then it can produce 25 frames of Lidar point cloud data per second, the geometry information between several successive Lidar point cloud frames is similar since the surrounding environment may not change dramatically within time period of around 0.1s, especially for the environment more than 100 meters away from the car. For example, front places that are more than 100 meters away from the car always have tall buildings and roads in a city or sometimes front cars run in the same way as the car, these objects that have a distance from the car in the front environment may not change dramatically within the time of several frames. Thus, there is temporal redundancy existing in successive Lidar data frames, and the redundancy can be exploited to improve the performance of Lidar data coding. Thus, why the proposed method according to the present disclosure is preferably implemented in such a situation.
[0146] To encode / decode sequences of Lidar data captured by Lidar devices of auto-driving cars, the coarse representation in the plane of successive lidar data frames when being used alone, will have much redundancy information. However, coding sequences of lidar data didn’t make use of redundancy information between successive point cloud frames.
[0147] Thus, according to the present disclosure, by considering adjacent lidar data frames’ coarse representation information, the compression performance of lidar data sequences (captured in auto-driving cars) can be improved. Additionally, by combining the method according to the present disclosure with the above-described framework of coarse representation, the compression efficiency can be further improved.
[0148] It is proposed to reduce the temporal redundancy of coarse representation information between successive Lidar data frames to further improve the compression performance of Lidar data sequences.
[0149] In real applications, the vertical angle range of Lidar scanning can reach around more than 25°, generally it is around 40° (for example, from -25° to +15°) , and the Lidar laser beams’ vertical angle range can be divided into three parts, which are higher laser beams (for example from 8° to 15°, preferably the range can be from 10° to 15°) , middle laser beams (for example from -14° to 8°-Δθ, preferably the range can be from -6° to 3°, wherein Δθ is angle difference between two laser beams on vertical direction) and lower laser beams (for example from -14°-Δθ to -25°, preferably the range can be from -20° to -25°) , the lasers of different parts may probe objects that have different / distinguished distances from the car. The higher laser beams can probe objects that are farther and taller, like buildings or trees; And the middle laser beams may probe near objects, like front cars, people, bicycles, and etc; The lower laser beams may probe the places that are really close to the car, usually the places are the road in front of the car. When a Lidar device working on a running car, the objects probed by higher laser beams and lower laser beams will not disappear suddenly within two adjacent Lidar scanning frames (within 0.1s time period) , then the order differences Δnext between two adjacent point cloud frames may have correlations, and their differences ΔΔnext will be smaller than Δnext, and then coding ΔΔnext into bitstream will save bits than coding Δnext directly into bitstream when coding coarse representation for current Lidar data frame.
[0150] As is illustrated in Figures 12a and 12b, they show the coarse representation in the plane of the Lidar point cloud at time t s and time t+1 s, and the points in the middle part represent points captured by middle laser beams, and the points in the upper part., the points in the lower part represent points captured by higher laser beams and lower laser beams respectively. Some middle points in the plane may change a little within 1 second because of the objects' movement or the front cars running, and it can be found that the little changes by comparing Figure 12a with Figure 12b. Thus, it can be observed that the coarse representation of adjacent frames captured within 0.1 second will be much smaller, especially for points probed by higher laser beams and lower beams. Thus, the coarse representation information of the previous adjacent lidar point cloud can be used to predict the coarse representation of the current coded lidar point cloud frame.
[0151] In a preferred embodiment, to code the current Lidar point cloud frame (the i-th frame) , the order difference Δnexti-1, Pn between two consecutive points, Pn and Pn-1 in the previous Lidar point cloud frame (the (i-1) -th frame) may be introduced to predict the order difference Δnexti, Pn between corresponding consecutive points in Pn and Pn-1 in the current frame, and a predicted order difference residual ΔΔnexti, Pn may be encoded / decoded into / from the bitstream. The predicted order difference residual for a point Pn may be obtained by ΔΔnexti, Pn=Δnexti, Pn-Δnexti-1, Pn.
[0152] In a preferred embodiment, the order difference prediction method may be used for each point to encode / decode the current point cloud frame.
[0153] A preferred embodiment of the proposed decoding method is shown in Figure 13. To be detailed, to decode a point Pn of a Lidar point cloud frame from a bitstream, it may follow the steps:
[0154] · Obtain the predicted order difference Δnexti-1, Pn between two consecutive points Pn and Pn-1 of the previous frame and obtain the order oi (Pn-1) of previously coded point Pn-1in the current frame;
[0155] · Decode the predicted order difference residual ΔΔnexti, Pn from bitstream by using Context-Adaptive Binary Arithmetic Coding, CABAC;
[0156] · And then obtain the order difference Δnexti, Pn of Pn in the current frame based on predicted order difference residual ΔΔnexti, Pn and the order difference Δnexti-1, Pn of Pn in the previous frame by Δnexti, Pn=Δnexti-1, Pn+ΔΔnexti, Pn
[0157] · Then obtain the order oi (Pn) of the current coded point in the current frame based on the obtained order difference Δnexti, Pn of Pn and order oi (Pn-1) of previously coded point in the current frame by oi (Pn) =oi (Pn-1) +Δnexti, Pn = oi (Pn-1) +Δnexti-1, Pn+ΔΔnexti, Pn
[0158] · Then construct the coarse representation in plane of the point Pn by calculating the laser (sensor) Index λ and azimuthal sampling angle based on the order of the point Pn, and they may be calculated by λn=oi (Pn) mod Nlaser,
[0159] wherein Nlaser is the total number of laser beams (or sensors) of a Lidar device.
[0160] · Then decode the radius residual, azimuthal residual, and residuals xres, yres and zres following the same process as previously outlined to reconstruct the geometry information (x, y, z) of the point oi (Pn) in the current frame.
[0161] After finishing decoding geometry and attribute information of a point in the current frame, then it may proceed to decode the next point of the current frame until it reaches the last point in the current frame.
[0162] In the current low latency low complexity Lidar coding method, for each frame, the information encoded / decoded into / from bitstream may include the order of first point P0 in the current frame and order differences ΔnextPn of all points in the frame for coarse representation. In the proposed method described above, the order of the first point P0 for each frame may be encoded / decoded into / from bitstream when coding Lidar data sequences.
[0163] In a preferred embodiment, only the order of first point P0 of the first frame in a Group of Pictures (GOP) is encoded / decoded into / from bitstream and the first point order in later frames in a GOP will not be encoded / decoded into / from bitstream directly, and only the first point order difference between the first point P0 in the current frame (the i-th frame) and the first point P’0 in the previous frame ( (the (i-1) -th frame) ) is encoded / decoded into / from bitstream. The equation to calculate the first point order difference may be described by
[0164] A Group of Pictures, commonly abbreviated as GOP, is a series of consecutive frames within a compressed video or image sequence. It is a fundamental concept in digital video compression, used to organize the sequence of frames for efficient encoding and decoding. A GOP typically starts with an I-frame (Intra-coded frame) that is encoded independently of other frames, followed by a series of P-frames (Predictive-coded frames) and B-frames (Bi-directionally predictive-coded frames) that are encoded based on information from preceding and / or following frames within the GOP. The I-frame serves as a reference point for subsequent frames, facilitating error recovery and random access in the video stream. The length of a GOP, defined as the number of frames it contains, and the pattern of I, P, and B-frames within it are adjustable based on the specific requirements of the video compression application, balancing between compression efficiency, image quality, and processing complexity.
[0165] In a preferred embodiment, the predicted order difference only depends on previously coded 1 frame information. In a variant, the predicted order differences Δnextpred of points in the current frame can be obtained based on an average value of Δnext between two consecutive points Pn and Pn-1 in more than 1 previously coded frames if more than 1 previously frames have been coded in a GOP, then, for example, if previously 2 frames order differences are used, the predicted order differences Δnextpred of points in the current frame can be obtained by Δnextpred= (w1*Δnexti-1, Pn+w2*Δnexti-2, Pn) ,
[0166] Wherein, weight W1 and W2 might be optionally applied, preferably W1+W2 = 1 and more preferably W1 > W2. Then the predicted order difference residual for a captured point can be obtained by ΔΔnexti, Pn=Δnexti, Pn-Δnextpred=Δnexti, Pn- (Δnexti-1, Pn+Δnexti-2, Pn) / 2.
[0167] In a preferred embodiment, a flag (for example, coarse_inter_prediction_flag) may be used to enable / disable the usage of the proposed inter-prediction method of coarse representation. In another embodiment, the flag can be inter inter-prediction flag indicating if inter inter-prediction method of Lidar data sequences coding is enabled or not. If the flag is true, then the proposed inter-prediction method of coarse representation (for example, the order difference prediction method) may be used to code the coarse representation of the current frame in the current Lidar frame coding process; otherwise, the proposed inter prediction method of coarse representation (for example, the order difference prediction method) will not be used in current Lidar frame coding process.
[0168] Reference is now made to Figure 16, which shows a simplified block diagram of an example embodiment of an encoder or decoder 300. The encoder or decoder 300 includes a processor 301 and a memory storage device 303. The memory storage device 303 may store a computer program or application containing instructions that, when executed, cause the processor 301 to perform operations such as those described herein. For example, the instructions may encode and output bitstreams encoded or decode bitstreams and output points of a point cloud in accordance with the methods described herein. It will be understood that the instructions may be stored on a non-transitory computer-readable medium, such as a compact disc, flash memory device, random access memory, hard drive, etc. When the instructions are executed, the processor 301 carries out the operations and functions specified in the instructions so as to operate as a special-purpose processor that implements the described process (es) . Such a processor may be referred to as a "processor circuit" or "processor circuitry" in some examples.
[0169] It will be appreciated that the decoder and / or encoder according to the present disclosure may be implemented in a number of computing devices, including, without limitation, servers, suitably programmed general purpose computers, machine vision systems, and mobile devices. The decoder or encoder may be implemented by way of software containing instructions for configuring a processor or processors to carry out the functions described herein. The software instructions may be stored on any suitable non-transitory computer-readable memory, including CDs, RAM, ROM, Flash memory, etc.
[0170] It will be understood that the decoder and / or encoder described herein and the module, routine, process, thread, or other software component implementing the described method / process for configuring the encoder or decoder may be realized using standard computer programming techniques and languages. The present application is not limited to particular processors, computer languages, computer programming conventions, data structures, other such implementation details. Those skilled in the art will recognize that the described processes may be implemented as a part of computer-executable code stored in volatile or non-volatile memory, as part of an application-specific integrated chip (ASIC) , etc.
[0171] Certain adaptations and modifications of the described embodiments can be made. Therefore, the above-discussed embodiments are considered to be illustrative and not restrictive. In particular, embodiments can be freely combined with each other.
Claims
1.A method for decoding, from a bitstream, three-dimensional position of points of a point cloud, preferably the point cloud is captured by a spinning Light Detection and Ranging, lidar, head with a plurality of sensors, preferably the lidar is deployed on a mobile platform, comprising:receiving the bitstream, wherein the bitstream contains view information configured to indicate an orientation relative to the mobile platform’s moving direction (S12) ;obtaining the view information (S14) ;determining points of the point cloud based on the view information (S16) .2.The method accoridng to claim 1, wherein the view information is configured to indicate at least one of a first view, a second view and a third view.3.The method according to claim 2, wherein the first view is a front view, the second view is a side view and the third view is a back view relative to the mobile platform’s moving direction.4.The method according claim 3, further comprising: determining the view information is configured to indicate the front view; determining points of the point cloud according to a precision level A; determining the view information is configured to indicate the side view; determining points of the point cloud according to a precision level B; wherein A > B and greater precision level indicating better quality of the reconstruction of the points.5.The method according to claim 4, further comprising: determining the view information is configured to indicate the back view; determining points of the point cloud according to a precision level C; wherein A > C.6.The method according to claim 5, wherein A > B > C.7.The method according to any one of claims 1 –6, wherein at least two points identified within the point cloud differ in their azimuth, elevation, or radial distance from the Lidar sensor, the at least two points are processed with distinct precision levels.8.The method according to claim 7, wherein the precision level applied to points within the point cloud data is determined based on their azimuth, elevation, or radial distance from the Lidar sensor.9.A method for encoding a three-dimensional position of points of a point cloud into a bitstream, preferably the point cloud is captured by a spinning Light Detection and Ranging, lidar, head with a plurality of sensors, and wherein the Lidar is preferably deployed on a mobile platform, comprising:determining view information indicating an orientation relative to the mobile platform’s moving direction (S22) ;encoding points of the point cloud into the bitstream based on the view information (S24) .10.The method accoridng to claim 9, wherein the view information is configured to indicate at least one of a first view, a second view and a third view.11.The method according to claim 10, wherein the first view is a front view, the second view is a side view and the third view is a back view relative to the mobile platform’s moving direction.12.The method according to claim 11, further comprising: determining the view information is configured to indicate the front view; determining points of the point cloud according to a precision level A; determining the view information is configured to indicate the side view; determining points of the point cloud according to a precision level B; wherein A > B and greater precision level indicating better quality of the reconstruction of the points.13.The method according to claim 12, further comprising: determining the view information is configured to indicate the back view; determining points of the point cloud according to a precision level C; wherein A > C.14.The method according to claim 13, wherein A > B > C.15.The method according to any one of claims 9 –14, wherein at least two points identified within the point cloud differ in their azimuth, elevation, or radial distance from the Lidar sensor, the at least two points are processed with distinct precision levels.16.The method according to claim 15, wherein the precision level applied to points within the point cloud data is determined based on their azimuth, elevation, or radial distance from the Lidar sensor.17.A decoder to decode a 3D point cloud from a bitstream comprising at least one processor and a memory, wherein the memory stores instructions when executed by the processor performs the steps of the method according to any one of claims 1 to 8.18.An encoder to encode a 3D point cloud into a bitstream comprising at least one processor and a memory, wherein the memory stores instructions when executed by the processor perform the steps of the method according to any one of claims 9 to 16.19.A bitstream encoded by the method according to any one of claims 9 to 16.20.A computer-readable storage medium comprising instructions when executed by a processor to perform the steps of the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Coding and decoding method, device and system for lossy compression of point cloud
CN113284248A
Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
US20210407142A1
Methods and apparatus for supporting collaborative extended reality (XR)
WO2023081197A1