Method and apparatus for encoding / decoding point cloud geometry data captured by a spin sensor head
Optimizing point cloud encoding and decoding by dynamic prediction of data lists solves the problem of coding complexity and high latency of spin sensor heads to capture sparse geometric data, achieving more efficient point cloud compression and real-time transmission.
Patent Information
- Application Number
- CN202180096722.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-09
- Filing Date
- 2021-10-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-10-13
AI Technical Summary
When existing point cloud compression technology processes sparse geometric data captured by spin sensor heads, there are problems such as complex encoding, high latency and insufficient compression performance, especially when real-time transmission and processing on autonomous vehicles, it is difficult to meet the requirements of simplicity and low latency.
Using the method of dynamic prediction data list, the encoding and decoding process of point clouds is optimized based on residual radius data by selecting and updating the predictor, and dynamically update the predicted data list to track object distances and improve prediction efficiency.
It improves the compression performance and coding efficiency of point cloud data, reduces latency, and meets the needs of real-time transmission and processing of autonomous vehicles.
Smart Images

Figure CN117157982B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to European Patent Application No. EP21305459.6, filed on April 9, 2021, the content of which is incorporated herein by reference in its entirety. Technical field
[0003] This application generally relates to point cloud compression, and particularly to methods and apparatuses for encoding / decoding point cloud geometric data captured by a spin sensor head. Background art
[0004] This section is intended to introduce the reader to aspects of the field that may be related to aspects of at least one exemplary embodiment of the present application described and / or claimed below. This discussion is considered to be helpful in providing background information to the reader to better understand the aspects of the present application.
[0005] As a format for representing 3D data, point clouds have recently received attention because of their generality in representing all types of physical objects or scenes. Point clouds can be used for various purposes, such as cultural heritage / buildings, where objects such as statues or buildings are scanned in 3D in order to share the spatial configuration of the object without sending or accessing it. Additionally, this is a way to ensure the preservation of knowledge of an object in case it may be damaged; for example, a temple that may be damaged by an earthquake. Such point clouds are typically static, colored, and huge.
[0006] Another use case is in topography and cartography, where using 3D representations allows maps to be not limited to a plane and may include relief. Google Maps is now a good example of a 3D map, but uses meshes instead of point clouds. However, point clouds may be a suitable data format for 3D maps, and such point clouds are typically static, colored, and huge.
[0007] Virtual reality (VR), augmented reality (AR), and immersive worlds have recently become hot topics and are foreseen by many as the future of 2D flat video. The basic idea is to immerse the viewer in the surrounding environment, whereas in contrast, a standard TV only allows the viewer to watch the virtual world in front of him / her. Depending on the viewer's degree of freedom in the environment, there are several levels of immersion. Point clouds are good format candidates for distributing VR / AR worlds.
[0008] The automotive industry, especially foreseeable autonomous vehicles, is also an area where point clouds may be intensively used. Autonomous vehicles should be able to "detect" their environment in order to make good driving decisions based on the detected presence and nature of nearby objects in their vicinity and the road configuration.
[0009] A point cloud is a set of points located in three-dimensional (3D) space, optionally with additional values attached to each point. These additional values are commonly referred to as attributes. Attributes can be, for example, the three-component color of the surface associated with the point, material properties (such as reflectivity), and / or the two-component normal vector.
[0010] Thus, a point cloud is a combination of geometry (the position of points in 3D space is typically represented by three-dimensional Cartesian coordinates x, y, and z) and attributes.
[0011] Point clouds can be captured by various types of devices, such as camera arrays, depth sensors, lasers (optical detection and ranging, also known as lidar), radars, or can be generated by a computer (e.g., in post-production of movies). Depending on the usage scenario, point clouds can have thousands to billions of points for mapping applications. The original representation of a point cloud requires a very high number of bits per point, at least a dozen bits for each Cartesian coordinate x, y, or z, and optionally more bits for attributes, e.g., three times 10 bits for color.
[0012] In many applications, it is important to be able to distribute the point cloud to the end user or store it on a server by consuming only a reasonable amount of bitrate or storage space, while maintaining an acceptable (or preferably very good) quality of experience. The efficient compression of these point clouds is the key to making the distribution chain of many immersive worlds practical.
[0013] Compression can be lossy (similar to video compression) for distribution and visualization to the end user (e.g., on AR / VR glasses or any other 3D-capable device). Other use cases do require lossless compression, such as medical applications or autonomous driving, to avoid changing the decision results obtained from subsequent analysis of the compressed and transmitted point cloud.
[0014] Until recently, point cloud compression (aka PCC) has not been addressed by the mass market, and no standardized point cloud codec was available. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11, also known as the Moving Picture Experts Group or MPEG, launched a working item on point cloud compression. This has led to two standards, namely
[0015] · Part 5 of MPEG-I (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC)
[0016] · Part 9 of MPEG-I (ISO / IEC 23090-9) or Geometry-based Point Cloud Compression (G-PCC)
[0017] The V-PCC coding method compresses point clouds by performing multiple projections of a 3D object to obtain 2D patches that are filled into an image (or video when processing dynamic point clouds). Then, the obtained image or video is compressed using an existing image / video codec to allow leveraging of the deployed image and video solutions. By its nature, V-PCC is only effective on dense and continuous point clouds because the image / video codec cannot compress non-smooth patches that would be obtained from projections of sparse geometric data captured, for example, by lidar.
[0018] The G-PCC coding method has two schemes for compressing the captured geometric data.
[0019] The first scheme is based on an occupancy tree representing the point cloud geometry, where the occupancy tree can be any type of tree among octrees, quadtrees, or binary trees. Occupied nodes are split until a certain size is reached, and the occupied leaf nodes provide the 3D positions of the points, typically at the centers of these nodes. The occupancy information is carried by occupancy flags that signal the occupancy status of each child node of the node. By using neighbor-based prediction techniques, a high level of compression of the occupancy flags can be obtained for dense point clouds. Sparse point clouds can also be addressed by stopping the tree construction when only isolated points are present in a node and directly encoding the positions of the points within nodes having non-minimal size; this technique is called the direct coding mode (DCM).
[0020] The second scheme is based on a prediction tree where each node represents the 3D position of a point and the parent / child relationship between nodes represents a spatial prediction from the parent to the child. This method can only handle sparse point clouds and has the advantages of lower latency and simpler decoding compared to the occupancy tree. However, compared to the first occupancy-based method, the compression performance is only slightly better, and the encoding is also complex because the encoder has to densely search for the best predictor among a long list of potential predictors when constructing the prediction tree.
[0021] In both schemes, the attribute (de)coding is performed after the geometric (de)coding is completed, which effectively results in two-pass encoding. Therefore, joint geometric / attribute low latency is obtained by using slices that decompose the 3D space into sub-volumes that are independently encoded without prediction between sub-volumes. When using many slices, this can severely affect the compression performance.
[0022] Combining the requirements regarding the simplicity of the encoder and decoder, regarding low latency, and regarding compression performance remains a problem that has not been satisfactorily solved by existing point cloud codecs.
[0023] An important use case is the transmission of sparse geometric data captured by a spin sensor head (e.g., a spin lidar head) mounted on a moving vehicle. This typically requires a simple and low-latency on-vehicle encoder. Simplicity is needed because the encoder may be deployed on a computing unit that is performing other processing in parallel, such as (semi) autonomous driving, thus limiting the processing power available for the point cloud encoder. Low latency is also needed to allow for fast transmission from the vehicle to the cloud for real-time viewing of local traffic based on multi-vehicle acquisitions and for making decisions quickly enough based on traffic information. While the transmission latency can be made low enough by using 5G, the encoder itself should not introduce too much latency due to encoding. Additionally, compression performance is very important because the data stream from millions of vehicles to the cloud is expected to be very heavy.
[0024] Specific priors related to the sparse geometric data captured by the spin sensor head have been used to obtain very efficient encoding / decoding methods.
[0025] For example, G-PCC utilizes the elevation angle (relative to the horizontal ground) captured by the spin sensor head, as Figure 1 and Figure 2 shown. The spin sensor head 10 includes a set of sensors 11 (e.g., lasers), represented here as five sensors. The spin sensor head 10 can spin about the vertical axis z to capture the geometric data of a physical object, i.e., the 3D positions of the points of the point cloud. The geometric data captured by the spin sensor head is represented in spherical coordinates (r 3D , φ, θ), where r 3D is the distance of the point P from the center of the spin sensor head, φ is the azimuth angle of the spin of the sensor head relative to a reference object, and θ is the elevation angle with respect to the elevation angle index k of the sensors of the spin sensor head relative to a horizontal reference plane (here the y-axis). The elevation angle index k can be, for example, the elevation angle of sensor k, or in the case where a single sensor is at each of the successive elevation angles during successive detections, the k-th sensor position..
[0026] A regular distribution along the azimuth angle is observed in the geometric data captured by the spin sensor head, as Figure 3 shown. This regularity is used in G-PCC to obtain a quasi-1D representation of the point cloud, where, up to noise, only the radius r 3D belongs to a continuous range of values, while the angles φ and θ only take a discrete number of values, from 0 to I-1, where I is the number of azimuth angles used to capture the points, and from 0 to K-1, where K is the number of sensors of the spin sensor head 10.. Basically, G-PCC represents the sparse geometric data captured by the spin sensor head on a 2D discrete angular plane (φ, θ), as Figure 3 shown, and the radius value r 3D for each point.
[0027] This quasi-1D property has been exploited in G-PCC, both in the occupancy tree and in the prediction tree that predicts the position of the current point in the spherical coordinate space based on the discrete nature of the angles and the already encoded points.
[0028] More precisely, the occupancy tree makes dense use of DCM and entropy-encodes the direct position of the points within the nodes by using a context-adaptive entropy coder. Then, the local transformation from the point position to the coordinates (φ,θ) and the context is obtained from the positions of these coordinates relative to the discrete coordinates (φ i ,θ k ) obtained from the already encoded points.
[0029] The prediction tree uses the quasi-1D property (r,φ i ,θ k ) of this coordinate space to directly encode a first version of the position of the current point P in spherical coordinates (r,φ,θ), where r is the projected radius on the horizontal xy plane, as shown in Figure 4 r 2D . Then, the spherical coordinates (r,φ,θ) are transformed into three-dimensional Cartesian coordinates (x, y, z), and the xyz residuals are encoded to address the errors in the coordinate transformation, the approximation of the elevation and azimuth angles, and potential noise.
[0030] Figure 5 Fig. shows a point cloud encoder similar to the encoder of the G-PCC prediction tree.
[0031] First, the Cartesian coordinates (x, y, z) of the points of the point cloud are transformed into spherical coordinates (r,φ,θ) by the (r,φ,θ) = C2A(x,y,z) transformation.
[0032] The transformation function C2A(.) is given in part by:
[0033] r = round(sqrt(x*x + y*y) / ΔIr)
[0034] φ = round(atan2(y,x) / ΔIφ)
[0035] where round() is the rounding operation to the nearest integer value, sqrt() is the square root function, and atan2(y,x) is the arctangent applied to y / x.
[0036] ΔIr and ΔIφ are the internal precisions for the radius and azimuth angle, respectively. They are typically the same as their respective quantization steps, i.e., ΔIφ = Δφ, and ΔIr = Δr, where
[0037]
[0038] and
[0039] Δr = 2 M *Basic quantization step
[0040] where M and N are two parameters of the encoder that can be signaled in the bitstream in the geometric parameter set, and where the basic quantization step is typically equal to 1. Typically, for lossless coding, N can be 17, and M can be 0.
[0041] The encoder can derive Δφ and Δr by minimizing the cost (e.g., number of bits) of encoding the spherical coordinate representation and the xyz residuals in Cartesian space.
[0042] For simplicity, in the following, Δφ = ΔIφ and Δr = ΔIr.
[0043] Also for clarity and simplicity, θ is used as the elevation angle value in the following, e.g., using
[0044]
[0045] where atan(.) is the arctangent function. However, in G-PCC, for example, θ is an integer value representing the elevation angle index k (i.e., the index of the k-th elevation angle), so the operations (prediction, residual (de)coding, etc.) performed on θ as represented below will be applied to the elevation angle index instead. Those familiar with point cloud compression will easily understand the advantage of using the index k, and how to use the elevation angle index k instead of θ. In addition, those skilled in the art of point cloud compression will easily understand that this subtlety does not affect the principle of the proposed invention. k Subsequently, the residual spherical coordinates (r
[0046] , φ n , θ res ) between the spherical coordinates (r, φ, θ) and the predicted spherical coordinates obtained from the predictor PR res are given by: res ) = (r, φ, θ) - (r
[0047] (r res , φ res , θ res ) = (r, φ, θ) - (r pred , φ pred , θ pred ) = (r, φ, θ) - PR n - (0, m * φ step , 0)(1)
[0048] where PR n is a predictor selected from a list of candidate predictors PR0, PR1, PR2, and PR3, and m is the basic azimuth step for the number of integer values to be added to the azimuth prediction of φstep .
[0049] The encoder can derive the basic azimuth step φ based on the frequency and rotational speed when the spin sensor head performs captures at different elevation angles, for example, based on NP, i.e., the number of detections per head rotation. step :[[]]
[0050]
[0051] And signaled in the bitstream in the geometric parameter set. Alternatively, NP is a parameter of the encoder that can be signaled in the bitstream in the geometric parameter set, and φ is derived similarly in both the encoder and the decoder. step .
[0052] The residual spherical coordinates (r res , φ res , θ res ) can be encoded in the bitstream B.
[0053] The residual spherical coordinates (r res , φ res , θ res ) can be quantized (Q) to the quantized residual spherical coordinates Q(r res , φ res , θ res ). The quantized residual spherical coordinates Q(r res , φ res , θ res ) can be encoded in the bitstream B.
[0054] The prediction index n and the number m are signaled in the bitstream B at each node of the prediction tree, while the basic azimuth step φ with a certain fixed-point precision step is shared by all nodes of the same prediction tree.
[0055] The prediction index n points to the selected predictor in the candidate predictor list.
[0056] The candidate predictor PR0 can be equal to (r min , φ0, θ0), where r min is the minimum radius value (provided in the geometric parameter set), and φ0 and θ0 are equal to 0 if the current node (current point P) has no parent node, or is equal to the azimuth and elevation angles of the point associated with the parent node.
[0057] Another candidate predictor PR1 can be equal to (r0, φ0, θ0), where r0, φ0, and θ0 are the radius, azimuth, and elevation angles of the point associated with the parent node of the current node, respectively.
[0058] Another candidate predictor PR2 can be equal to a linear prediction of the radius, azimuth angle, and elevation angle (r0, φ0, θ0) using the radius, azimuth angle, and elevation angle of the point associated with the parent node of the current node, as well as the radius, azimuth angle, and elevation angle (r1, φ1, θ1) of the point associated with the grandparent node.
[0059] For example, PR2 = 2 * (r0, φ0, θ0) - (r1, φ1, θ1)
[0060] Another candidate predictor PR3 can be equal to a linear prediction of the radius, azimuth angle, and elevation angle (r0, φ0, θ0) using the radius, azimuth angle, and elevation angle of the point associated with the parent node of the current node, the radius, azimuth angle, and elevation angle (r1, φ1, θ1) of the point associated with the grandparent node, and the radius, azimuth angle, and elevation angle (r2, φ2, θ2) of the point associated with the great-grandparent node.
[0061] For example, PR3 = (r0, φ0, θ0) + (r1, φ1, θ1) - (r2, φ2, θ2)
[0062] The predicted Cartesian coordinates (x pred , y pred , z pred ) are obtained by performing an inverse transformation on the decoded spherical coordinates (r dec , φ dec , θ dec ) as follows:
[0063] (x pred , y pred , z pred ) = A2C(r dec , φ dec , θ dec ) (2)
[0064] where the spherical coordinates (r dec , φ dec , θ dec ) decoded by the decoder can be given by:
[0065] (r dec , φ dec , θ dec ) = (r res,dec , φ res,dec , θ res,dec ) + PR n + (0, m * φ step , 0), (3)
[0066] where (r res,dec , φ res,dec , θ res,dec ) are the residual spherical coordinates decoded by the decoder.
[0067] The decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ) can be the result of inverse quantization (IQ) of the quantized residual spherical coordinates Q(r res , φ res , θ res ).
[0068] In G-PCC, there is no quantization of the residual spherical coordinates, and the decoded spherical coordinates (r res,dec , φ res,dec , θ res,dec ) are equal to the residual spherical coordinates (r res , φ res , θ res ). Then, the decoded spherical coordinates (r dec , φ dec , θ dec ) are equal to the spherical coordinates (r, φ, θ).
[0069] The inverse transformation of the decoded spherical coordinates (r dec , φ dec , θ dec ) can be given by:
[0070] r = r dec * Δr
[0071] x pred = round(r * cos(φ dec * Δφ))
[0072] y pred = round(r * sin(φ dec * Δφ)
[0073] z pred = round(tan(θ dec ) * r)
[0074] where sin() and cos() are the sine and cosine functions. These two functions can be approximated by performing operations with fixed-point precision. The value tan(θ dec ) can also be stored as a fixed-point precision value. Therefore, floating-point operations are not used in the decoder. Avoiding floating-point operations is generally a strong requirement for simplifying the hardware implementation of the codec.
[0075] The residual Cartesian coordinates (x pred , y pred , z pred ) between the original point and the predicted Cartesian coordinates (x res , y res , zres ) is given by:
[0076] (x res , y res , z res ) = (x, y, z) - (x pred , y pred , z pred )
[0077] The residual Cartesian coordinates (x res , y res , z res ) are quantized (Q), and the quantized residual Cartesian coordinates Q(x res , y res , z res ) are encoded into a bitstream.
[0078] When the quantization step sizes of x, y, z are equal to the origin precision (usually 1), the residual Cartesian coordinates can be losslessly encoded, or when the quantization step size is greater than the origin precision (usually the quantization step size is greater than 1), the residual Cartesian coordinates can be lossily encoded.
[0079] The Cartesian coordinates (x dec , y dec , z dec ) decoded by the decoder are given by:
[0080] (x dec , y dec , z dec ) = (x pred , y pred , z pred ) + IQ(Q(x res , y res , z res )) (4)
[0081] where IQ(Q(x res , y res , z res )) represents the inverse-quantized quantized residual Cartesian coordinates.
[0082] The decoder can use those decoded Cartesian coordinates (x dec , y dec , z dec ), for example, for sorting the (decoded) points before attribute coding.
[0083] Figure 6 A point cloud decoder similar to the decoder for the prediction tree based on the G-PCC prediction tree is shown.
[0084] The prediction index n and the number m are accessed from the bitstream B of each node of the prediction tree, while the basic azimuth step φ step is accessed from the parameter set of the bitstream B and is shared by all nodes of the same prediction tree.
[0085] The decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ) can be obtained by decoding the residual spherical coordinates (r res , φ res , θ res ) from the bitstream B.
[0086] The quantized residual spherical coordinates Q(r res , φ res , θ res ) can be decoded from the bitstream B. The quantized residual spherical coordinates Q(r res , φ res , θ res ) are inverse quantized to obtain the decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ).
[0087] According to Equation (3), the decoded spherical coordinates (r res,dec , φ res,dec , θ res,dec ) are added to the predicted spherical coordinates (r pred , φ pred , θ pred ) to obtain the decoded spherical coordinates (r dec , φ dec , θ dec ).
[0088] According to Equation (2), the predicted Cartesian coordinates (x dec , y dec , z dec ) are obtained by inverse-transforming the decoded spherical coordinates (r pred , φ pred , θ pred ).
[0089] The quantized residual Cartesian coordinates Q(x res , y res , z res ) are decoded from the bitstream B and inverse quantized to obtain the inverse quantized Cartesian coordinates IQ(Q(x res , y res , z res ). The decoded Cartesian coordinates (x dec , ydec , z dec ) is given by equation (4).
[0090] To quickly process the point cloud data captured by the spin sensor head, the data should be processed according to its acquisition order without spending extra time reordering the points for more efficient compression.
[0091] For example, a slice or a prediction tree can be guided for each sensor index, or alternatively, a degenerate tree can be constructed as a comb tree, where each comb tooth is associated with a sensor index and contains a chain of points for this sensor index in its acquisition order.
[0092] In the presence of noise, the candidate predictors PR2 and PR3 are usually ineffective. Therefore, they are rarely selected because the point cloud data captured by the spin sensor head usually has noise due to the nature of the acquisition process.
[0093] The candidate predictor PR1 (also known as the delta predictor) is often selected because it is less sensitive to noise. However, the sensor usually captures objects at different distances from the sensor, and a large jump in the radius between two consecutive points can be observed, resulting in high-amplitude residual radius values to be encoded in the bitstream. None of the candidate predictors PR1, PR2, or PR3 are suitable for predicting these jumps. If the object is close to the sensor (i.e., close to the minimum radius), the candidate predictor PR0 may sometimes be useful, but for objects far from the sensor, the candidate predictor PR0 is also inefficient. In addition, the disadvantage of the candidate predictor PR0 is that it introduces a one-frame delay in the processing of the point cloud data in order to obtain a suitable minimum radius. For real-time processing, a suboptimal value, such as 0, will be used.
[0094] To improve the prediction of the point cloud geometry data captured by the spin sensor head and thus its compression, a better prediction scheme is needed, especially for predicting the radius information, which accounts for the most important part of the bit budget in the bitstream. Summary of the Invention
[0095] The following section presents a simplified overview of at least one exemplary embodiment to provide a basic understanding of some aspects of the present application. This overview is not an extensive overview of the exemplary embodiment. It is not intended to identify the key or critical elements of the embodiment. The following overview only presents some aspects of at least one of the exemplary embodiments in a simplified form as a prelude to the more detailed description provided elsewhere in the document.
[0096] According to a first aspect of the present application, there is provided a method of encoding a point cloud into a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with at least one radius responsive to a distance of the point from a reference object, the method comprising selecting a predictor of at least predicted radius data representing the radius of a point of the point cloud from at least one predictor derived from at least one prediction data; encoding data representing the prediction data in the bitstream, the selected predictor being derived from the prediction data; encoding in the bitstream residual radius data between the data representing the radius of the point and the predicted radius data derived from the selected predictor; and updating the at least one prediction data based on the residual radius data.
[0097] According to a second aspect of the present application, there is provided a method of decoding a point cloud from a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with at least one radius responsive to a distance of the point from a reference object, the method comprising decoding data representing prediction data in a list of at least one prediction data from the bitstream, a predictor of at least one predicted radius data representing the radius of a point of the point cloud being derived from the at least one prediction data; deriving the predicted radius data of the point from the predictor; decoding the residual radius of the point from the bitstream; deriving data representing the radius of the point from the residual radius data and the predicted radius data; and updating (250) the at least one prediction data based on the residual radius data.
[0098] In an exemplary embodiment, updating the at least one prediction data based on the residual radius data is based on a comparison of the residual radius data with a threshold.
[0099] In an exemplary embodiment, updating the at least one prediction data based on the residual radius data comprises: if the residual radius data is greater than the threshold, adding new prediction data derived from at least the residual radius data to the top of the list of prediction data; and if the residual radius data is below the threshold, the prediction data from which the selected predictor is derived, or prediction data associated with data representing prediction data in a list of at least one prediction data from which at least one predicted radius data representing the radius of a point of the point cloud is derived, is updated based at least on the residual radius data, and the updated prediction data is moved to the top of the list of prediction data.
[0100] In an exemplary embodiment, if the residual radius data is greater than the threshold and if the maximum number of prediction data in the list of prediction data is reached, the last prediction data in the list of prediction data is removed from the list of prediction data.
[0101] In an exemplary embodiment, the predicted radius data of the predicted data is the radius of the decoded point of the point cloud, and the predicted data further includes the azimuth angle of the decoded point, and a predictor including the predicted radius data and the azimuth angle.
[0102] In a variant, the azimuth angle of the predictor represents the sum of the azimuth angle of the decoded point and an integer number of basic azimuth angle steps.
[0103] In an exemplary embodiment, the predicted radius data of the predicted data is the radius of the decoded point of the point cloud, and the predicted data further includes a quantized residual azimuth angle associated with the decoded point, and a predictor including the predicted radius data and the quantized residual azimuth angle.
[0104] In an exemplary embodiment, the predicted radius data of the predicted data is the radius of the quantized decoded point, and the predicted data further includes a quantized residual azimuth angle associated with the decoded point, and a predictor including the predicted radius data and the quantized residual azimuth angle
[0105] In an exemplary embodiment, a list of spherical coordinate predictors is associated with each sensor of a spin sensor head for capturing points of a point cloud.
[0106] According to a third aspect of the present application, there is provided a bitstream representing encoded point cloud data of a physical object, each point of the point cloud being associated with at least one radius responsive to the distance of the point from a reference object, wherein the bitstream further includes data representing predicted data in a list of at least one predicted data, and a predictor of at least one radius of a point of the point cloud is derived from the at least one predicted data.
[0107] According to a fourth aspect of the present application, there is provided an apparatus for encoding a point cloud into a bitstream representing encoded point cloud data of a physical object. The apparatus includes one or more processors configured to execute the method according to the first aspect of the present application.
[0108] According to a fifth aspect of the present application, there is provided an apparatus for decoding points of a point cloud representing a physical object from a bitstream. The apparatus includes one or more processors configured to execute the method according to the second aspect of the present application.
[0109] According to a sixth aspect of the present application, there is provided a computer program product including instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to the second aspect of the present application.
[0110] According to a seventh aspect of the present application, there is provided a non-transitory storage medium carrying instructions for program code for executing the method according to the second aspect of the present application.
[0111] The specific nature of at least one of the exemplary embodiments and other objects, advantages, features, and uses of at least one of the exemplary implementations will become apparent from the following description of the examples in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0112] Reference will now be made, by way of example, to the accompanying drawings, which illustrate exemplary embodiments of the present application, and in which:
[0113] Figure 1 A side view of a sensor head and some of its parameters according to the prior art is shown;
[0114] Figure 2 A top view of a sensor head and some of its parameters according to the prior art is shown;
[0115] Figure 3 A regular distribution of data captured by a spin sensor head according to the prior art is shown;
[0116] Figure 4 A representation of points in 3D space according to the prior art is shown;
[0117] Figure 5 A point cloud encoder similar to an encoder based on a G-PCC prediction tree according to the prior art is shown;
[0118] Figure 6 A point cloud decoder similar to a decoder based on a G-PCC prediction tree according to the prior art is shown;
[0119] Figure 7 A block diagram of the steps of a method 100 for encoding a point cloud representing a physical object according to at least one exemplary embodiment is shown;
[0120] Figure 8 A block diagram of the steps of a method 200 for decoding a point cloud representing a physical object according to at least one exemplary embodiment is shown;
[0121] Figure 9 A block diagram of the steps of a method 300 for updating a list of at least one prediction data PD from residual radius data according to at least one exemplary embodiment is shown; and p
[0122] Figure 10 A schematic block diagram of a system example in which various aspects and exemplary embodiments are implemented is shown.
[0123] Like reference numerals may be used in different drawings to represent like components. DETAILED DESCRIPTION
[0124] At least one of the exemplary embodiments is described more fully below with reference to the accompanying drawings, in which examples of at least one of the exemplary embodiments are shown. However, the exemplary embodiments may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Accordingly, it is to be understood that the exemplary embodiments are not intended to be limited to the particular forms disclosed. On the contrary, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present application.
[0125] When a figure is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a figure is presented in the form of a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0126] At least one of the aspects generally relates to point cloud encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream.
[0127] Furthermore, the present aspect is not limited to MPEG standards such as Part 5 or Part 9 of MPEG-I related to point cloud compression, and can be applied to, for example, other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including Part 5 and Part 9 of MPEG-I). Unless otherwise stated or technically excluded, the aspects described in the present application can be used alone or in combination.
[0128] The present invention relates to a method for encoding / decoding a point cloud representing a physical object, wherein each point of the point cloud is associated with spherical coordinates representing at least one radius in response to the distance of the point from a reference object.
[0129] The present invention relates to the technical field of encoding and decoding, and aims to provide a technical solution for encoding / decoding point cloud data. Since a point cloud is a set of massive data, storing a point cloud may consume a large amount of memory, and it is also impossible to directly transmit the point cloud at the network layer without compressing the point cloud. Therefore, it is necessary to compress the point cloud. Accordingly, as the application of point clouds in aspects such as autonomous navigation, real-time detection, geographic information services, cultural heritage / architecture protection, 3D immersive communication and interaction becomes more and more extensive, the present invention can be applied to many application scenarios.
[0130] This encoding / decoding method particularly relates to encoding / decoding point cloud data based on a dynamic list of prediction data to improve the compression performance of the point cloud. For example, the total bit rate for encoding / decoding the point cloud can be improved.
[0131] The present invention determines a dynamic list of at least one prediction data for deriving at least one candidate predictor for encoding geometric data of points of a point cloud. The list of at least one prediction data is dynamic because during the encoding or decoding of a point, the prediction data is updated based on residual radius data representing the residual radius of the decoded point.
[0132] Such a dynamic list of prediction data allows tracking of various object distances and produces more efficient predictions than a static pre-determined list of predictors used in the prediction tree of GPCC, resulting in an improvement in the total bit rate for encoding both the data representing the selected predictor and the prediction residuals.
[0133] The dynamic list of prediction data is particularly (but not only) useful for obtaining / deriving better predictions after a sensor (laser) beam has moved from a first object at a first distance to another object at a different distance, has passed by it and returned to the first object. This can occur, for example, when one object is in front of another (e.g., a car in front of a wall), or when an object has holes (e.g., a wall with open doors and windows, or an entrance wall).
[0134] Figure 7 A step block diagram of a method 100 for encoding a point cloud representing a physical object according to at least one exemplary embodiment is shown.
[0135] The list of prediction data includes at least one prediction data PD p . The prediction data PD p includes prediction radius data R representing the radius of a point of the point cloud p and data D representing the azimuth angle of said point. p .
[0136] Initially, i.e., at the start of the encoding of the point cloud or a G-PCC slice, for example, the list of prediction data includes a single prediction data PD0. The single prediction data PD0 includes data R0 representing an initial radius and data D0 representing an initial azimuth angle.
[0137] The single prediction data PD0 can be a default predictor pre-known to both the encoder and the decoder. For example, the single prediction data PD0 can include an initial radius and an initial azimuth angle, where the initial radius can be a predetermined minimum radius value extracted from all points of the point cloud, and the initial azimuth angle can be the azimuth angle of the root node of the prediction tree or the minimum value extracted from all possible azimuth angles of the spin sensor head.
[0138] As explained in detail below, under certain conditions, new prediction data PD p can be added to the prediction list. Then, the list of prediction data includes at least one prediction data PD p, and for each prediction data PD p Obtain candidate predictor DPR p . Therefore, the candidate predictor DPR is dynamically obtained according to the update of the list of prediction data p .
[0139] In an exemplary embodiment of the method, the list of prediction data is associated with each elevation angle index k of the sensors of the spin sensor head, and the list of candidate predictors DPR p is associated with each list of prediction data. Then, each list of prediction data is updated independently of each other.
[0140] In an exemplary embodiment of the method, the list of candidate predictors includes at least one candidate predictor DPR of at least prediction radius data representing the radius of the points of the point cloud p .
[0141] Candidate predictor DPR p 's list can replace predictors PR0 to PR3 in G-PCC, or at least one of the predictors PR0, PR1, PR2, and / or PR3 defined in G-PCC can compete with candidate predictor DPR p .
[0142] In step 110, a predictor DPR of at least prediction radius data representing the radius of the points of the point cloud is selected from the list of at least one candidate predictor DPR p derived from the list of at least one prediction data PD p .
[0143] For example, the selection can be performed by using an optimization process based on a rate cost function or a rate-distortion cost function (in the case of lossy coding), and the optimization process selects a candidate predictor DPR that minimizes the cost function.
[0144] In step 120, the data DA representing the prediction data from which the selected predictor DPR is derived is encoded in the bitstream B
[0145] In an exemplary embodiment of step 120, the data DA is the index of the prediction data PD in the list of prediction data from which the selected candidate predictor DPR is derived
[0146] In step 130, the residual radius data between the data representing the radius of the points and the predicted radius data derived from the selected predictor is encoded in the bitstream B
[0147] In a variant, the residual radius is quantized before being encoded
[0148] In step 140, based on the residual radius data, and preferably, according to Figure 9 method 300 updates at least one prediction data PD p .
[0149] Figure 8 FIG. shows a block diagram of steps of a method 200 for decoding a point cloud representing a physical object according to at least one exemplary embodiment.
[0150] The list of prediction data includes at least one prediction data PD p . The prediction data PD p includes predicted radius data R representing the radius of decoded points of the point cloud p and data D representing the azimuth angle of the decoded points. p .
[0151] Initially, for example, at the start of decoding of the point cloud or g-PCC slice, the list of prediction data includes a single prediction data PD0, as Figure 7 explained.
[0152] As explained in detail below, under certain conditions, new prediction data PD p can be added to the prediction list. Then, the list of prediction data includes at least one prediction data PD p , and for each prediction data PR p a candidate predictor DPR p is obtained. Thus, the candidate predictor DPR p is obtained dynamically according to the update of the list of prediction data.
[0153] In an exemplary embodiment of the method, the list of prediction data is associated with each elevation angle index k of the sensors of the spin sensor head, and the list of candidate predictors DPR p is associated with each list of prediction data. Then, each list of prediction data is updated independently of each other.
[0154] In an exemplary embodiment of the method, the list of candidate predictors includes at least one candidate predictor DPR of at least one radius of the points of the point cloud p .
[0155] In step 210, data DA is decoded from the bitstream B. The data DA represents the prediction data in the list of at least one prediction data PD p , and from the prediction data PD p a predictor DPR of at least one predicted radius data representing the radius of the points of the point cloud is derived. The predictor DPR belongs to the list of candidate predictors DPR p .
[0156] In step 220, the predicted radius data for the point is derived from the predictor DPR.
[0157] In step 230, the residual radius data decoded from the bitstream B.
[0158] In step 240, data representing the radius of the point is derived from the residual radius data and the predicted radius data.
[0159] In step 250, based on the residual radius data, and preferably, according to Figure 9 method 300, at least one prediction data PD is updated p .
[0160] Figure 9 shows a block diagram of the steps of method 300 for updating at least one prediction data PD p from the residual radius data according to at least one exemplary embodiment.
[0161] Updating at least one prediction data PD p of the list takes as input the data DA representing the selected candidate predictor DPR and the residual radius threshold Th-r.
[0162] In step 310, the residual radius data is compared with the residual radius threshold Th-r.
[0163] In one exemplary embodiment of method 300, the residual radius threshold Th-r can be a fixed value. For example, it can be 1024, 2048 or 4096 scaled according to the bit depth of the input data and / or according to the quantization parameter.
[0164] In a variant, when the comparison (which will be explained in further detail) is based on the quantized or inverse quantized residual radius, the residual radius threshold Th-r can be equal to the fixed value divided by the quantization step.
[0165] In one exemplary embodiment of method 300, the residual radius threshold Th-r can be a parameter signaled / accessed from the bitstream B. For example, it can belong to the set of geometric parameters of the bitstream B.
[0166] In one exemplary embodiment of method 300, the residual radius threshold Th-r can be a dynamically calculated threshold. For example, it can be calculated based on the variance of the residual radius, which can be estimated based on (some of) the previously encoded / decoded residual radii, for example as the average of the absolute values.
[0167] Since the radius represents the distance of the sensed object from the reference of the spin sensor head, as described below, the update of the list of prediction data based on the comparison of the residual radius data with a threshold keeps track of several distances that may correspond to different objects in the sensing scenario.
[0168] If the residual radius data is greater than or equal to the residual radius threshold Th-r, then in step 320, new prediction data PD is added at the top of the list of prediction data p (i.e., inserted as the new first element). This new prediction data PD p is at least derived from the residual radius data.
[0169] A residual radius exceeding the threshold indicates that the associated point may be part of a new object, for example, a car in front of another object (e.g., a wall), or a new object behind another object when the object has holes (e.g., a wall with open doors and windows, or an entrance wall). Adding the new prediction data PD p to the list of prediction data results in dynamically deriving a new candidate predictor that provides a better prediction for encoding other points on this new object compared to the prediction derived from the predictor obtained from the list of prediction data before the update. However, most importantly, retaining the prediction data obtained from points belonging to a previous object in the list of prediction data allows "remembering" (the distance to) the previously seen object, thus deriving a candidate predictor that provides a better prediction for encoding new points that will again belong to the previous object (e.g., when the spin sensor will pass through the new object) and thus for points with a similar distance.
[0170] In a variant, if the residual radius data is greater than the residual radius threshold Th-r and if the maximum number of prediction data in the list of prediction data is reached, the last prediction data in the list of prediction data is removed from the list of prediction data.
[0171] This will have a similar effect to forgetting the prediction data (and the distance to) the oldest object seen.
[0172] This variant limits the size of the list of prediction data in memory and the encoding cost of the data DA.
[0173] Typically, the list of prediction data is limited to 4, 5, 6, 7, or 8 elements. For such small list sizes, it is usually more efficient to implement the list as a buffer / table array of contiguous memory, which is maintained such that the elements are sorted in memory according to their index in the list of prediction data. Then, this memory can stay in the cache of a typical processing unit, and moving or inserting predictors in the buffer / table / array involves fast operations for moving a small portion of the memory in the cache; and accessing the prediction data PD of an element in the list in the cache is fast to obtain the predictor DPR p to obtain the predictor DPR p is fast.
[0174] If the residual radius data is below the residual radius threshold Th-r, then in step 330, the prediction data PD from which the selected predictor DPR is derived (step 110) or the prediction data PD associated with the data DA representing the predictor DPR (step 210) is updated based at least on the residual radius data, and the updated prediction data PD is moved to the top of the list of prediction data (i.e., it is moved before the first element of the list and thus it becomes the first element of the list).
[0175] Residual radius data below the threshold indicates that the associated point of the point cloud is part of the same object as the object from which the prediction data PD is obtained. Thus, since the two points are part of the same object, the prediction data PD is updated according to the residual radius data. Moving the updated prediction data PD to the top of the list of prediction data is a simple way to allow for better compression performance in the encoding of the data DA of the selected predictor, because the predictor obtained / derived from this element has just been used and it is much more likely to be used again for the next point. This also has the advantage that when the last prediction data is removed from the list, the most recently used prediction data (and thus the prediction information that has a greater chance of being used again later) is not "forgotten".
[0176] In one exemplary embodiment, the present invention can be used to improve the current G-PCC prediction scheme given by equations (1), (2), or (3), where the radius r and the azimuth angle φ are adaptively quantized, as described in European Patent Application No. EP20306674. In this exemplary embodiment, regarding Figure 5 and Figure 6 described Q(r res , φ res , θ res ) is set to be equal to (Qr res , Qφ res , θ res ), where Qr res and Qφ resare the radius residual of adaptive quantization and the azimuth residual of adaptive quantization, as described below, and θ res is the non-quantized elevation (index) residual.
[0177] On the encoding side, the residual radius r obtained from Equation (1) is adaptively quantized as follows res :
[0178] Qr res = Q r (r res , φ pred ) = round(r res / Δr(φ pred )) (5)
[0179] where Q r is an adaptive quantizer using Δr(φ pred ) given by:
[0180] Δr(φ pred ) = Δ / (|sin(φ pred )| + |cos(φ pred )|) (6)
[0181] where φ pred is the predicted azimuth of the azimuth φ given by Equation (1).
[0182] The inverse quantized residual radius IQr is obtained by inverse quantizing Qr res as follows res
[0183] IQr res = IQ r (Qr res , φ pred ) = Qr res * Δr(φ pred )) (7)
[0184] where IQ r is an adaptive inverse quantizer based on the predicted azimuth φ pred .
[0185] The decoded radius r is obtained by adding the inverse quantized residual radius IQr res to the predicted radius r given by Equation (1) pred . dec .
[0186] The residual azimuth φ obtained from Equation (1) is adaptively quantized as follows res :
[0187] Qφ res = Qφ (φ res ,r dec ) = round(φ res / Δφ(r dec )) (8)
[0188] where Q φ is an adaptive quantizer using a quantization step Δφ(r dec ) given by:
[0189] Δφ(r dec ) = Δr(φ pred ) / r dec
[0190] The inverse - quantized residual azimuth angle IQφ res is obtained by inverse - quantizing Qφ res as follows:
[0191] IQφ res = IQ φ (Qφ res ,r dec ) = Qφ res *Δφ(r dec ) (9)
[0192] where IQ φ is an adaptive inverse - quantizer based on the decoded radius r dec .
[0193] Finally, the decoded spherical coordinates (r dec ,φ dec ,θ dec ) of the decoded point can be given by:
[0194] (r dec ,φ dec ,θ dec ) = (IQr res ,IQφ res ,θ res ) + PR n +(0,m*φ step ,0) (10)
[0195] The quantized residual radius Qr res and the quantized residual azimuth angle Qφ res are encoded in the bitstream B
[0196] In the exemplary embodiment of the present invention, the predicted radius data R p is the radius r dec of the decoded point, and the data D p is the azimuth angle φdec , and the residual data is the decoded residual radius associated with the current point of the point cloud.
[0197] In steps 110 and 220, the predictor PR in equation (1) or (3) n can be a candidate predictor DPR belonging to a list derived from a list of prediction data p of predictors. The candidate predictor DPR p is given by:
[0198] DPR p = (r dec,p , φ dec,p , θ0) (11)
[0199] where θ0 is equal to 0 if the node associated with the point has no parent node, or θ0 is equal to the elevation angle of the point associated with the parent node or equal to a predetermined minimum elevation angle, and where r dec,p and φ dec,p are the (previous) decoded radius and (previous) decoded azimuth angle represented by the prediction radius data R p and the data D p respectively, and are included in the prediction data PD p .
[0200] This exemplary embodiment results in a very good prediction of the azimuth angle of points with similar (or higher) radii, especially when the points have close azimuth angles and when the quantization of the residual azimuth angle depends on the radius.
[0201] In steps 140 and 250, based on the decoded residual radius r res,dec , φ res,dec , θ res,dec ) obtained from the decoded residual spherical coordinates associated with the decoded current point given by equation (3) res,dec (residual radius data) to update the list of prediction data. The decoded residual radius r res,dec can also be decoded from the bitstream B.
[0202] In step 320, the new prediction data PD p includes the radius r dec (as the prediction radius data R p ) and the azimuth angle φ dec (as the data D p ) of the decoded current point given by equation (3). Thus, the new prediction data PD p enables the r dec and φ dec set equal to the just decoded radius and azimuth angle r dec,p and φdec,p to construct a new candidate predictor DPR p .
[0203] In step 330, the prediction data PD p is updated to include the data R dec as the radius r of the currently decoded point p and the data D dec as the azimuth angle φ of the currently decoded point p . Thus, the updated prediction data PD p enables the construction of a new candidate predictor DPR dec from the r dec and φ dec,p that are updated to be equal to the radius and azimuth angle that have just been decoded dec,p . p .
[0204] In a variant of the exemplary embodiment, the radius is not adaptively quantized (and thus, (de)coded as in a G-PCC prediction tree-based codec), and the quantization step size Δφ(r dec ) of the azimuth angle is defined by:
[0205] Δφ(r dec ) = Δφ arc / r dec ,
[0206] where Δφ arc represents the quantization step size of the azimuth angle arc length (i.e., the quantization step size of the azimuth angle multiplied by the radius). Δφ arc can be chosen to be equal to 1, or 2 1 / 2 , or 8 / 2*π. In the latter case, for example, if the internal precision representing the azimuth angle ΔIφ is (2*π) / 2 N , then the internal representation of Δφ arc will be ΔIφ arc = ΔIφ * Δφ arc = 1 / 2 N-3 , and thus multiplication or division by a factor of Δφ arc can be implemented by using simple shift operations.
[0207] In another variant of the exemplary embodiment, the azimuth angle representation of the candidate predictor DPR p is the sum of the (previously) decoded azimuth angle φ dec,p (predicted azimuth angle) and the integer number m p of the basic azimuth angle steps that separate the decoded azimuth angle and the azimuth angle φ0 of the point associated with the parent node of the current point.
[0208] In steps 110 and 220, the predictor PR in equation (1) or (3) n can be a candidate predictor DPR belonging to a list derived from a list of prediction data p of candidate predictors. The candidate predictor DPR p is then given by:
[0209] DPR p =(r dec,p , φ dec,p + m p * φ step , θ0) (12)
[0210] where m p is derived from the difference between the azimuth angles φ0 and φ dec,p .
[0211] For example:
[0212] m p = 0 if |φ dec,p - φ0| < φ step
[0213] m p = round((φ dec,p - φ0) / φ step ) otherwise
[0214] where round(.) can be any operation that rounds to an integer.
[0215] Figures 7 to 9 The steps of p remain the same, except for step 320, in which the new prediction data PD p enables the use of equation (12) to construct a new candidate predictor DPR
[0216] This variation of the exemplary embodiment is advantageous because it generally provides better predictions than those given by equation (11), and thus, it reduces the cost of encoding the integer value m used in equations (1) and (3) in the bitstream B.
[0217] In one exemplary embodiment, the present invention can be used to improve the prediction scheme of a single-chain encoding / decoding scheme, as described in European Patent Application No. EP20306672.
[0218] In single-chain encoding, the captured 3D position is represented in a 2D coordinate (Cφ, λ) system together with a radius value r (r 2D or r 3D ). The coordinate Cφ (for coarse φ) is the azimuth angle index of the spin of the sensor head, and its discrete values are represented as Cφi ( to I - 1), corresponding to the effective rotation of the sensor head at an angle φ i The coordinate λ is the sensor index, and its discrete values are represented as λ k ( to K - 1). The radius r belongs to a continuous numerical range.
[0219] For each point of the point cloud, the sensor index λ associated with the sensor that captured the point is obtained (λ is one of the sensor indices λ k ( to K - 1)), the azimuth index Cφ representing the capture angle of the sensor (Cφ is one of the discrete angle indices Cφ i ( to I - 1)), and the radius value r of the spherical coordinates of the point.
[0220] The sensor index λ and the azimuth index Cφ are obtained by transforming the 3D Cartesian coordinates (x, y, z) representing the 3D position of the captured point. These 3D Cartesian coordinates (x, y, z) can be the output of the sensor head. For example, assuming that the angle φ step is the basic azimuth step between two consecutive detections of the sensor head for a given sensor index λ, and assuming that the arctangent value of y / x returned by the function atan2(y, x) takes values in the interval [0; 2*π], then Cφ can be obtained as follows:
[0221] Cφ = round(φ / φ step ),
[0222] where φ = atan2(y, x). Then, the rotation angle φ i will be obtained by the following formula:
[0223] φ i = Cφ i * φ step .
[0224] In this case, the set of discrete angles φ i (0 ≤ i < I) is essentially defined by φ i = i * φ step .
[0225] In addition, λ can be determined as the index λ of the sensor that has an elevation angle θ obtained by minimizing the following formula k k
[0226]
[0226] λ = argmin k abs(θ k - θ),
[0227] where abs(.) is a function that returns the absolute value, and
[0228] θ = z / sqrt(r),
[0229] where r = x * x + y * y, and sqrt(.) is a function that returns the square root.
[0230] Next, sort the points of the point cloud based on the azimuth index Cφ and the sensor index λ.
[0231] In one variant, sort the points lexicographically, first based on the azimuth, and then based on the sensor index. The order index o(P) of point P is obtained as follows:
[0232] o(P) = Cφ * K + λ
[0233] In another variant, sort the points lexicographically, first based on the sensor index, and then based on the azimuth. The order index o(P) of point P is obtained as follows:
[0234] o(P) = λ * I + Cφ
[0235] Encoding the ordered points into the bitstream B may include encoding the order index difference Δo n which n each represents the difference between the order indices of two consecutive points P n-1 and P n for n = 2 to N:
[0236] Δo n = o(P n ) - o(P n-1 )
[0237] The order index o(P1) of the first point P1 can be directly encoded into the bitstream B. This is equivalent to arbitrarily setting the order index of the virtual zero point to zero, i.e., o(P0) = 0, and encoding Δo1 = o(P1) - o(P0) = o(P1).
[0238] Given the order index o(P1) of the first point and the order difference Δo n , then the order index o(P n ) of any point P n can be recursively reconstructed as follows:
[0239] o(P n ) = o(P n-1 ) + Δo n
[0240] Then, the one associated with point P nAssociated sensor index λ n and azimuth index Cφ n :
[0241] λ n = o(P n ) modulo K (13)
[0242] Cφ n = o(P n ) / K (14)
[0243] where the division / K is integer division (also known as Euclidean division). Thus, o(P1) and Δo n are alternative representations of λ n and Cφ n .
[0244] On the encoding side, the residual azimuth φ res is given by:
[0245] φ res = φ - φ pred = φ - Cφ n *φ step (15)
[0246] where, φ pred = φ n = Cφ n *φ step is the predicted azimuth.
[0247] Then, encoding the ordered points into the bitstream B can also include encoding the residuals (r res , Qφ res,res ) associated with the ordered points given by:
[0248] (r res , Qφ res,res ) = (r, Qφ res ) - (r pre d, Qφ res,pred ) (16)
[0249] where, r res is the residual radius, Qφ res is the quantized residual azimuth, Qφ res,res is the residual of the quantized residual azimuth, and Qφ res,pred is the quantized predicted residual azimuth. The elevation angle φ n of each point P n of the point cloud is not prediction-encoded and is considered equal to the elevation angle n associated with the sensor index λ which sensor index is obtained using equation (13) from the order o(Pn )Determined.
[0250] To obtain the quantized residual azimuth Qφ res , an adaptive quantizer Q of Equation (8) with a decoding radius r dec can be used to quantize the residual azimuth φ φ (Equation 15): res (Equation 15):
[0251] r dec = IQr res + r pred ,
[0252] where IQr res is obtained from the quantized radius residual Qr as in Equation (7) res , and the quantized radius residual Qr res is generated by the adaptive quantization of the radius residual r as in Equation (5) res , and the predicted angular azimuth φ pred is used, which is equal to φ in both equations n = Cφ n * φ step .
[0253] In a variant, the decoded residual radius IQr res is obtained as follows:
[0254] IQr res = IQ(Qr res ),
[0255] where Qr res = Q(r res ) is the uniformly quantized radius residual, and where Q is the uniform quantizer and IQ is the inverse uniform quantizer.
[0256] The inverse quantized residual azimuth IQφ can be obtained by using the decoding radius r dec , and applying the adaptive inverse quantizer IQ of Equation (9) φ to the quantized residual azimuth. The decoded residual azimuth φ res is equal to IQφ res,dec : res φ
[0257] = IQφ res,dec = IQ res (Qφ φ , r res , r dec )
[0258] The quantized residual radius Qr resEncoded in the bitstream B by an encoder such that it can be decoded by a decoder and inverse quantized to obtain the same decoded residual radius r res,dec The inverse quantized residual radius IQr of res .
[0259] Then the predicted quantized residual azimuth Qφ res,pred is used to predict the quantized residual azimuth Qφ res , to obtain the residual of the quantized residual azimuth Qφ res,res .
[0260] Qφ res,res = Qφ res - Qφ res,pred
[0261] Then the residual Qφ of the quantized residual azimuth is encoded in the bitstream B res,res .
[0262] At the encoding end, the decoded coordinates (r dec , φ dec θ dec ) are given by:
[0263]
[0264] Encoding the ordered points into the bitstream B may also include obtaining the residual Cartesian coordinates (x res , y res , z res ) of the three-dimensional Cartesian coordinates of the ordered points as follows:
[0265] (x res , y res , z res ) = (x, y, z) - (x pred , y pred , z pred )
[0266] where (x, y, z) are the three-dimensional Cartesian coordinates of the ordered points, and (x pred , y pred , z pred ) are the predicted Cartesian coordinates obtained as follows:
[0267]
[0268] The residual Cartesian coordinates (x res , y res , z res ) are quantized (Q), and the quantized residual Cartesian coordinates Q(x res , y res , z res) are encoded into the bitstream.
[0269] When the x, y, and z quantization steps are equal to the origin precision (usually 1), the residual Cartesian coordinates can be losslessly encoded, or when the quantization step is greater than the origin precision (usually the quantization step is greater than 1), the residual Cartesian coordinates can be lossily encoded.
[0270] Decoding the points of the point cloud from the bitstream requires information such as the number of points N of the point cloud, the order index o(P1) of the first point in the 2D coordinate (Cφ,λ) system, and sensor setting parameters such as the basic azimuth step φ associated with each sensor step or the elevation angle θ k .
[0271] Decode at least one order index difference Δo n (n = 2 to N). For the current point P n Decode each order index difference Δo n .
[0272] For the current point P n , the order index o(P n ) is obtained by the following:
[0273] o(P n ) = o(P n-1 ) + Δo n
[0274] The sensor index λ associated with the sensor that captured the current point n and the azimuth angle φ representing the capture angle of the sensor n = Cφ n * φ step is derived from the order index o(P n ) (equations (13) and (14)).
[0275] Decode the quantized residual radius Qr from the bitstream B res .
[0276] The decoded residual radius r res,dec = IQr res is obtained by using equation (7) and using the predicted azimuth angle φ equal to φ n in the equation to inverse-quantize Qr pred . In one variant, the decoded residual radius r res is obtained by applying a uniform inverse quantizer to the quantized residual radius Qr res = IQr res,dec . res .
[0277] Decode the residual Qφ of the quantized residual azimuth angle from the bitstream B res,res 。
[0278] The quantized residual azimuth angle Qφ res is obtained by the following:
[0279] Qφ res = Qφ res,res + Qφ res,pred ,
[0280] And the decoded residual azimuth angle φ res,dec = IQφ res is obtained by inverse quantizing Qφ res (Equation 9).
[0281] The decoded spherical coordinates are given by:
[0282]
[0283] The decoded Cartesian coordinates (x dec , y dec , z dec ) are given by:
[0284] (x dec , y dec , z dec ) = (x pred , y pred , z pred ) + IQ(Q(x res , y res , z res ))
[0285] where IQ(Q(x res , y res , z res ) represents the inverse quantized decoded quantized residual Cartesian coordinates from the bitstream B.
[0286] In the exemplary embodiment of the present invention, according to Figures 7 to 9 the method, a list PD of prediction data is maintained for each sensor index λ p,λ to maintain the radius prediction highly dependent on the sensor index.
[0287] The predicted radius data R p is the predicted radius r dec equal to the radius r of the decoded point dec,p , and the data A p is the predicted Qφ of the quantized residual azimuth angle equal to the quantized residual Qφ res,dec of the decoded point res,pred,p (see step 320 below).
[0288] In steps 110 and 220, then, a candidate predictor derived from a list of prediction data associated with sensor index λ n is given by: as follows:
[0289]
[0290] where r dec,p is the prediction radius, and Qφ res,pred,p is the prediction of the quantized residual azimuth angle given by the predictor .
[0291] Thus, the r in equation (16) pred is equal to r dec,p .
[0292] Then, the residual Qφ of the quantized residual azimuth angle is obtained by using equation (16) res,res , where
[0293] Qφ res,pred = Qφ res,pred,p :
[0294] Qφ res,res = Qφ res - Qφ res,pred,p .
[0295] The residuals (Qr res , Qφ res,res ) include the quantized residual radius, which is obtained using adaptive quantization of the residual radius or in variable uniform quantization (as described above), and the residual of the quantized residual azimuth angle, and the residuals (Qr res , Qφ res,res ) are encoded in the bitstream B on the encoder side and decoded from the bitstream B on the decoder side. In the encoder and decoder, the decoded radius residual lr res,dec = IQr res is obtained by inverse quantization of the quantized radius residual, and the decoded radius r res,dec is obtained by adding r dec,p to the prediction radius r dec :
[0296] r dec = r res,dec + r dec,p .
[0297] The decoded quantized azimuth angle residual Qφ is obtained by adding the residual Qφ of the quantized residual azimuth angle res,res to the prediction Qφ of the quantized residual azimuth angle res,pred,p res,dec (equal to Qφ res ):
[0298] Qφ res,dec = Qφ res,res + Qφ res,pred,p = Qφ res
[0299] And by using the decoding radius r_dec, the adaptive inverse quantizer IQ of equation (8) φ is applied to the decoded quantized azimuth residual Qφ res,dec , obtaining the decoded residual azimuth residual φ res,dec :
[0300] φ res,dec = IQφ res,dec = IQ φ (Qφ res,dec , r dec ) = IQφ res
[0301] Then, the decoded spherical coordinates are given by:
[0302]
[0303] In steps 140 and 250, based on the decoded residual radius r res,dec to update the list of predicted data.
[0304] In step 320, the new predicted data includes the radius r dec (in the predicted radius data R p ) and the quantized residual azimuth Qφ of the decoded current point res,dec (in the data D p ). Thus, the new predicted data PD p can be set equal to the radius and quantized residual azimuth r dec and Qφ res,dec just decoded, and r dec,p and Qφ res,pred,p are used to construct a new candidate predictor
[0305] In one variant, the decoding radius r is obtained in both the encoder and decoder by applying the adaptive inverse quantizer of equation (7) to the quantized radius Qr instead of the quantized radius residual dec , φ pred equal to φ n :
[0306] r dec = IQ r (Qr, φ n ) (19)
[0307] Qr is the quantizer Q using equation (5), r obtained, where φ pred is equal to φ n , so that Q r is applied to the radius r instead of the radius residual:
[0308] Qr = Q r (r, φ n ). (20)
[0309] In this variant, as will be described in detail below, instead of using the predicted radius to perform the prediction of the radius, the predicted quantized radius is used to perform the prediction of the quantized radius. Then, the predicted radius data R p is the predicted quantized radius Qr equal to the radius Qr of the quantized decoded point p , and the data A p remains the quantized residual azimuth Qφ equal to the quantized residual Qφ of the decoded point res,dec res,pred,p .
[0310] In this variant, in steps 110 and 220, the candidate predictor n derived from the list of predicted data associated with the sensor index λ is given by:
[0311]
[0312] where Qr p is the predicted quantized radius, and Qφ res,pred,p is the prediction Qφ of the quantized residual azimuth given by the predictor res .
[0313] In this variant, then, the residual radius r res is equal to the residual Qr of the quantized radius, res given by:
[0314] r res = Qr res = Qr - Qr p
[0315] where Qr is the quantized radius of the point given by equation (20), and Qr p is the predicted radius of the quantized radius.
[0316] Including the residual of the residual (Qr res , Qφ res,res )Is encoded in the bitstream B on the encoder side and decoded from the bitstream B on the decoder side.
[0317] In both the encoder and the decoder, by adding the quantization radius residual Qr res to the predicted quantization radius Qr p the decoded quantization radius Qr dec is obtained:
[0318] Qr dec = Qr res + Qr p = Qr,
[0319] and the decoded radius r is obtained by inverse quantizing the quantized radius residual using equation (19) dec .
[0320] In step 320, the new prediction data includes the quantization radius Qr of the decoded current point dec (in the prediction radius data R p ) and the quantized residual azimuth angle Qφ res,dec (in the data D p ). Thus, the new prediction data PD p enables a new candidate predictor to be constructed from Qr dec and Qφ res,dec which are set equal to the just decoded quantization radius and quantized residual azimuth angle p of Qr res,dec,p and Qφ
[0321] This encoding / decoding method can be used to encode / decode point clouds for various purposes to improve the compression performance of the point clouds, for example, it can improve the total bit rate for encoding / decoding point cloud data.
[0322] For example, in a use case of transmitting sparse geometric data captured by a spin sensor head mounted on a moving vehicle, an encoding / decoding method can be used to encode / decode a point cloud representing an obstacle. Each point of the point cloud is associated with a radius that represents the distance of the point from a reference of the spin sensor head. Each point of the point cloud can correspond to a point on the obstacle. A dynamic list of prediction data for deriving a predictor for encoding the geometric data of the points is determined, the prediction data initially including data representing a predetermined minimum radius value extracted from all points of the point cloud, e.g., the radius value of a point on an obstacle closest to the reference of the spin sensor head in this use case. During encoding, the data representing the prediction data of the predictor and the residual radius data are encoded in a bitstream, where the residual radius data is the data between the data representing the radius of the point and the predicted radius data derived from the predictor. During decoding, the data representing the prediction data and the residual radius data are decoded from the bitstream, then the predicted radius data can be derived from the predictor derived from the data representing the prediction data, and then the data representing the radius of the point can be derived from the residual radius data and the predicted radius data. During the encoding or decoding of a point, the prediction data is updated based on the residual radius data representing the residual radius of the point. In this use case, for example, the bitstream can be transmitted by an encoder mounted on a moving vehicle to the cloud and then obtained by a decoder on a data processing server, or the bitstream can be directly transmitted to the data processing server, such that the geometric data of the obstacle captured by the moving vehicle can be transmitted to the data processing server for subsequent processing, e.g., determining a driving path for the moving vehicle. Further, since the dynamic list of prediction data allows tracking of various object distances and generates more efficient predictions, the total bit rate for encoding the data can be increased.
[0323] Figure 10 FIG. shows a schematic block diagram illustrating an example of a system in which various aspects and exemplary embodiments are implemented.
[0324] System 400 can be embodied as one or more devices, including various components described below. In various embodiments, system 400 can be configured to implement one or more aspects described in the present application.
[0325] Examples of devices that may form all or part of system 400 include personal computers, laptop computers, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output of the video decoder, pre-processors that provide input to the video encoder, network servers, set-top boxes, and any other device or other communication device for processing point clouds, video, or images. The elements of system 400, individually or in combination, may be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 400 may be distributed across multiple ICs and / or discrete components. In various embodiments, system 400 may be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.
[0326] System 400 may include at least one processor 410 configured to execute instructions loaded therein to implement, for example, the various aspects described in the present application. Processor 410 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 400 may include at least one memory 420 (e.g., volatile memory devices and / or non-volatile storage devices). System 400 may include a storage device 440, which may include non-volatile memory and / or volatile storage, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 may include internal storage devices, connected storage devices, and / or network-accessible storage devices.
[0327] System 400 may include an encoder / decoder module 430 configured to, for example, process data to provide encoded / decoded point cloud geometry data, and the encoder / decoder module 430 may include its own processor and memory. Encoder / decoder module 430 may represent a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding and a decoding module. Additionally, encoder / decoder module 430 may be implemented as a separate element of system 400 or may be incorporated within processor 410 as a combination of hardware and software known to those skilled in the art.
[0328] The program code to be loaded onto the processor 410 or the encoder / decoder 430 to execute the various aspects described in this application can be stored in the storage device 440 and then loaded onto the memory 420 to be executed by the processor 410. According to various embodiments, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 can store one or more of the various items during the execution of the processes described in this application. The items so stored can include, but are not limited to, point cloud frames, encoded / decoded geometry / attribute video / images or portions thereof, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.
[0329] In several embodiments, the memory internal to the processor 410 and / or the encoder / decoder module 430 can be used to store instructions and provide working memory for the processing that can be executed during encoding or decoding.
[0330] However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 410 or the encoder / decoder module 430) can be used for one or more of these functions. The external memory can be the memory 320 and / or the storage device 440, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory can be used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM can be used as the working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), or MPEG-I Part 5 or Part 9.
[0331] Inputs to the elements of the system 300 can be provided through various input devices, as shown by block 490. Such input devices include, but are not limited to, (i) an RF portion that can receive, for example, RF signals transmitted over the air by a broadcaster, (ii) composite input terminals, (iii) USB input terminals, and / or (iv) HDMI input terminals.
[0332] In various embodiments, the input device of block 490 may have respective input processing elements known in the art. For example, the RF section may be associated with the following required elements: (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting the signal band to one band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) de-multiplexing to select a desired data packet stream. The RF section of various embodiments may include one or more elements for performing these functions, such as, for example, a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down-converter, a demodulator, an error corrector, and a de-multiplexer. The RF section may include a tuner that performs various functions of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or a baseband.
[0333] In one set-top box embodiment, the RF section and its associated input processing elements may receive an RF signal transmitted through a wired (e.g., cable) medium. Then, the RF section may perform frequency selection by filtering, down-converting, and again filtering to a desired frequency band.
[0334] Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.
[0335] Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section may include an antenna.
[0336] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 400 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, as needed, within, for example, a separate input processing IC or within the processor 410. Similarly, various aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within the processor 410. The demodulated, error-corrected, and de-multiplexed stream may be provided to various processing elements, including, for example, the processor 410 and the encoder / decoder 430, which operate in combination with memory and storage elements to process the data stream as needed for presentation on an output device.
[0337] The various elements of the system 400 may be disposed within an integrated housing. Within the integrated housing, the various elements may be interconnected using a suitable connection arrangement 490 and data may be transmitted between them, such as, for example, internal buses known in the art, including I2C buses, wiring, and printed circuit boards.
[0338] System 400 may include a communication interface 350 that enables communication with other devices via a communication channel 800. The communication interface 450 may include, but is not limited to, a transceiver configured to send and receive data over the communication channel 800. The communication interface 450 may include, but is not limited to, a modem or a network card, and the communication channel 800 may be implemented, for example, within a wired and / or wireless medium.
[0339] In various embodiments, data may be streamed to System 400 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments may be received via a communication channel 800 and a communication interface 450 suitable for Wi-Fi communication. The communication channel 800 of these embodiments may generally be connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other top-level communications.
[0340] Other embodiments may use a set-top box to provide streamed data to System 400, and the set-top box delivers data through the HDMI connection of the input block 490.
[0341] Still other embodiments may use the RF connection of the input block 490 to provide streamed data to System 400.
[0342] The streamed data may be used as a way of signaling information used by System 400. The signaling information may include a bitstream B and / or such as data DA, the number of points in a point cloud, the coordinates of the first point in a 2D coordinate (Cφ,λ) system or the order o(P1) and / or sensor setting parameters (such as the basic azimuth step φ associated with the sensors of the spin sensor head 10 step or the elevation angle θ k ).
[0343] It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. may be used to send information to a corresponding decoder.
[0344] System 400 may provide output signals to various output devices, including a display 400, a speaker 600, and other peripheral devices 700. In various examples of the embodiments, the other peripheral devices 700 may include one or more of an independent DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of System 400.
[0345] In various embodiments, signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention can be used to transfer control signals between system 400 and display 500, speaker 600, or other peripheral devices 700.
[0346] Output devices can be communicatively coupled to system 400 via dedicated connections through respective interfaces 460, 470, and 480.
[0347] Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 800. Display 500 and speaker 600 can be integrated into a single unit with other components of system 400 in an electronic device such as a television.
[0348] In various embodiments, display interface 460 can include a display driver, such as a timing controller (T Con) chip.
[0349] For example, if the RF portion of input 490 is part of a separate set-top box, then display 500 and speaker 600 are alternatively separated from one or more other components. In various embodiments where display 500 and speaker 600 can be external components, output signals can be provided via a dedicated output connection that includes, for example, an HDMI port, a USB port, or a COMP output.
[0350] In Figures 1 - 10 this, various methods are described, and each method includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.
[0351] Some examples are described with respect to block diagrams and / or operation flowcharts. Each block represents a circuit element, module, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in other implementations, the functions noted in the blocks can occur in the order indicated. For example, in fact, two consecutively displayed blocks can be executed substantially simultaneously, or these blocks can sometimes be executed in the reverse order, depending on the functions involved.
[0352] The implementations and aspects described herein can be implemented, for example, in a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if discussed only in the context of a single implementation form (e.g., only as a method), the implementation of the features discussed can be implemented in other forms (e.g., an apparatus or a computer program).
[0353] These methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device.
[0354] In addition, these methods can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product specifically implemented in one or more computer-readable media and having computer-readable program code specifically implemented thereon that is executable by a computer. The computer-readable storage medium used herein can be considered a non-transitory storage medium, given its inherent ability to store information and to provide information retrieval therefrom. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. It should be understood that while the following provides more specific examples of computer-readable storage media to which the present embodiments can be applied, the following is merely an illustrative list readily understandable by a person of ordinary skill in the art and is not an exhaustive list: a portable computer floppy disk; a hard disk; a read-only memory (ROM); an erasable programmable read-only memory (EPROM or flash memory); a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing.
[0355] The instructions can form an application program specifically implemented tangibly on a processor-readable medium.
[0356] For example, the instructions can be hardware, firmware, software, or a combination thereof. The instructions can be found in an operating system, a separate application program, or a combination of both. Thus, a processor can be characterized as, for example, a device configured to execute a process and a device including a processor-readable medium (such as a storage device) having instructions for executing the process. In addition, in addition to or instead of the instructions, the processor-readable medium can store data values generated by the implementation.
[0357] The apparatus can be implemented using, for example, suitable hardware, software, and firmware. Examples of such devices include personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, head-mounted display devices (HMDs, see-through glasses), projectors (beamers), "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from the video decoder, pre-processors that provide input to the video encoder, web servers, set-top boxes, and any other device or other communication device for processing point clouds, video, or images. It should be clear that the device can be mobile and can even be installed in a moving vehicle.
[0358] The computer software can be implemented by the processor 410 or by hardware or by a combination of hardware and software. As a non-limiting example, the embodiments can also be implemented by one or more integrated circuits. The memory 420 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology, such as optical storage devices, magnetic storage devices, semiconductor-based storage devices, fixed memories, and removable memories, as non-limiting examples. The processor 410 can be of any type suitable for the technical environment and can include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.
[0359] As will be apparent to those of ordinary skill in the art, the implementations can generate various signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal can be formatted to carry the bitstream of the described embodiments. For example, such a signal can be formatted as an electromagnetic wave (e.g., using the radio frequency part of the spectrum) or a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over various different wired or wireless links, which are known. The signal can be stored on a processor-readable medium.
[0360] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" may also include the plural forms unless the context clearly dictates otherwise. It will be further understood that when the terms "comprises / comprising" and / or "includes / including" are used in this specification, the presence of the stated features, integers, steps, operations, elements and / or components may be specified, but one or more other features, integers, steps, operations, elements, components and / or groups are not excluded. Further, when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to other elements, no intervening elements are present.
[0361] It should be understood that, for example, in the case of "A / B", "A and / or B" and "at least one of A and B", any one of the symbols / terms " / " and " / or" and "at least one of" may be intended to include only the first-listed option (A), or only the second-listed option (B), or both options (A and B). As a further example, in the case of "A, B and / or C" and "at least one of A, B and C", such wording is intended to include only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or all three options (A, B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to any number of items listed.
[0362] Various numerical values may be used in this application. Specific values may be for illustrative purposes, and the aspects described are not limited to these specific values.
[0363] It should be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of this application, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. No order is implied between the first element and the second element.
[0364] References to "an exemplary embodiment" or "exemplary embodiments" or "an implementation" or "implementations" and other variations thereof are often used to convey that a particular feature, structure, characteristic, etc. (described in connection with the embodiment / implementation) is included in at least one embodiment / implementation. Thus, the phrases "in an exemplary embodiment" or "in exemplary embodiments" or "in an implementation" or "in implementations" and any other variations thereof that appear throughout this application do not necessarily all refer to the same embodiment.
[0365] Similarly, references herein to "according to an exemplary embodiment / example / implementation" or "in an exemplary embodiment / example / implementation" and other variations thereof are often used to convey that a particular feature, structure, or characteristic (described in connection with the exemplary embodiment / example / implementation) may be included in at least one exemplary embodiment / example / implementation. Thus, the expressions "according to an exemplary embodiment / example / implementation" or "in an exemplary embodiment / example / implementation" that appear throughout the specification do not necessarily all refer to the same exemplary embodiment / example / implementation, and individual or alternative exemplary embodiments / examples / implementations are not necessarily mutually exclusive of other exemplary embodiments / examples / implementations.
[0366] The reference numerals appearing in the claims are for illustrative purposes only and do not limit the scope of the claims. Although not explicitly described, the embodiments / examples and variations may be used in any combination or sub - combination.
[0367] When a figure is represented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a figure is presented in the form of a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0368] Although some figures include arrows on communication paths to indicate the main direction of communication, it should be understood that communication can occur in a direction opposite to that of the depicted arrows.
[0369] Various implementations involve decoding. "Decoding" as used in this application can include, for example, all or part of the process performed on a received point cloud frame (which may include a received bitstream that encodes one or more point cloud frames) to produce a final output suitable for display or further processing in the reconstructed point cloud domain. In various embodiments, such a process includes one or more of the processes typically performed by a decoder. In various embodiments, such a process also or alternatively includes the processes performed by the decoders of the various implementations described in this application, e.g.,
[0370] As a further example, in one embodiment, "decoding" may refer to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" specifically refers to a subset of operations or generally refers to a broader decoding process will be clear based on the context of the specific description, and it is believed that those skilled in the art will well understand it.
[0371] Various implementations involve encoding. In a manner similar to the discussion of "decoding" above, "encoding" as used in this application may include, for example, all or part of the processing performed on an input point cloud frame to produce an encoded bitstream. In various embodiments, such processing includes one or more of the processing typically performed by an encoder. In various embodiments, such processing also or alternatively includes the processing performed by the encoders of the various implementations described in this application.
[0372] As a further example, in one embodiment, "encoding" may refer only to entropy encoding, in another embodiment, "encoding" may refer only to differential encoding, and in another embodiment, "encoding" may refer to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" specifically refers to a subset of operations or generally refers to a broader encoding process will be clear based on the context of the specific description, and it is believed that those skilled in the art will well understand it.
[0373] In addition, this application may be related to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.
[0374] In addition, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory or a bitstream), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0375] In addition, the application may refer to "receiving" various information. Like "accessing", receiving is a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory or a bitstream). In addition, during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is generally involved in one way or another.
[0376] In addition, as used herein, the word "signal" refers to indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals specific information, such as data DA, the number of points of a point cloud, or the coordinates of the first point in a 2D coordinate system (Cφ,λ) or the order o(P1), or sensor setting parameters, such as the basic azimuth step φ associated with sensor k step or the elevation angle θ k . In this way, in one embodiment, the same parameters can be used on the encoder side and the decoder side. Thus, for example, the encoder can send (explicit signaling) specific parameters to the decoder such that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling (implicit signaling) can be used without transmission to simply allow the decoder to know and select the specific parameters. By avoiding transmitting any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to send information to the corresponding decoder. Although the foregoing relates to the verb form of the word "signal", "signal" can also be used as a noun herein
[0377] Numerous implementations have been described. However, it should be understood that various modifications can be made. For example, elements of different implementations can be combined, supplemented, modified, or removed to yield other implementations. In addition, those skilled in the art will understand that other structures and processes can replace those disclosed, and the resulting implementations will perform at least substantially the same functions in at least substantially the same way to achieve at least substantially the same results as the disclosed implementations. Accordingly, these and other implementations are contemplated by this application
Claims
1. A method of encoding a point cloud into a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with at least one radius responsive to the distance of the point from a reference object, the method comprising: - Selecting (110) a predictor of at least one predicted radius data representing the radius of a point of the point cloud from at least one predictor derived from a list of at least one predicted data; - Encoding (120) data representing the predicted data into the bitstream, the selected predictor being derived from the predicted data; - Encoding into the bitstream (130) radius residual data between data representing the radius of the point and predicted radius data derived from the selected predictor; and - Updating (140) the at least one predicted data based on the radius residual data; wherein updating the at least one predicted data based on the radius residual data comprises: - If the radius residual data is greater than a threshold, adding new predicted data derived from at least the radius residual data to the top of the list of the at least one predicted data; and - If the radius residual data is below the threshold, the predicted data derived from the predictor from which the selected predictor is derived, or predicted data associated with data representing the predicted data in the list of the at least one predicted data of the predictor representing at least one predicted radius data of the radius of a point of the point cloud is updated at least based on the radius residual data, and the updated predicted data is moved to the top of the list of the at least one predicted data.
2. A method of decoding a point cloud from a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with at least one radius responsive to the distance of the point from a reference object, the method comprising: - Decoding (210) from the bitstream data representing predicted data in a list of at least one predicted data, wherein a predictor of at least one predicted radius data representing the radius of a point of the point cloud is derived from the at least one predicted data; - Deriving (220) the predicted radius data of the point from the predictor; - Decoding (230) the radius residual data of the point from the bitstream; - Deriving (240) data representing the radius of the point from the radius residual data and the predicted radius data; and - Updating (250) the at least one predicted data based on the radius residual data; wherein updating the at least one predicted data based on the radius residual data comprises: - If the radius residual data is greater than a threshold, adding new predicted data derived from at least the radius residual data to the top of the list of the at least one predicted data; and - If the radius residual data is below the threshold, the prediction data derived from the selected predictor, or the prediction data associated with the prediction data in the list of at least one prediction data representing at least one prediction radius data of the radius representing the points of the point cloud is updated at least based on the radius residual data, and the updated prediction data is moved to the top of the list of the at least one prediction data.
3. The method according to claim 1 or 2, wherein If the radius residual data is greater than the threshold, and if the maximum number of prediction data in the list of the at least one prediction data is reached, the last prediction data in the list of the at least one prediction data is removed from the list of the at least one prediction data.
4. The method according to one of claims 1 to 2, wherein The prediction radius data of the prediction data is the radius of the decoded point of the point cloud, and the prediction data further includes the azimuth angle of the decoded point, and the predictor includes the prediction radius data and the azimuth angle.
5. The method according to claim 4, wherein the azimuth angle of the predictor represents the sum of the azimuth angle of the decoded point and an integer number of basic azimuth angle steps.
6. The method according to one of claims 1 to 2, wherein The prediction radius data of the prediction data is the radius of the decoded point of the point cloud, and the prediction data further includes the quantized residual azimuth angle associated with the decoded point, and the predictor includes the prediction radius data and the quantized residual azimuth angle.
7. The method according to one of claims 1 to 2, wherein The prediction radius data of the prediction data is the radius of the quantized decoded point, and the prediction data further includes the quantized residual azimuth angle associated with the decoded point, and the predictor includes the prediction radius data and the quantized residual azimuth angle.
8. The method according to one of claims 1 to 2, wherein, The list of spherical coordinate predictors is associated with each sensor of the spin sensor head for capturing the points of the point cloud.
9. An apparatus for encoding a point cloud into a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with at least one radius responsive to the distance of the point from a reference object, the apparatus including one or more processors configured to: - Select, from at least one predictor derived from at least one prediction data, a predictor representing at least one prediction radius data of the radius representing the points of the point cloud; - Encode the data representing the prediction data into the bitstream, the selected predictor being derived from the prediction data; -Encode the radius residual data between the data representing the radius of the point and the predicted radius data obtained from the selected predictor into the bitstream; and - Update the at least one prediction data based on the radius residual data; wherein updating the at least one prediction data based on the radius residual data includes: - If the radius residual data is greater than a threshold, adding new prediction data derived from at least the radius residual data to the top of the list of the at least one prediction data; and - If the radius residual data is below the threshold, the predicted data derived from the selected predictor, or the predicted data associated with the data of the predicted data in the list of at least one predicted data of the predictor representing at least one predicted radius data of the radius representing the points of the point cloud is updated at least based on the radius residual data, and the updated predicted data is moved to the top of the list of the at least one predicted data.
10. An apparatus for decoding a point cloud from a bitstream representing encoded point cloud data of a physical object, each point of the point cloud being associated with at least one radius responsive to the distance of the point from a reference object, the apparatus including one or more processors configured to: - Decode data representing predicted data in a list of at least one predicted data from the bitstream, wherein, At least one predictor of predicted radius data representing the radius of the points of the point cloud is derived from the at least one predicted data; - Derive predicted radius data of the point from the predictor; - Decode radius residual data of the point from the bitstream; and - Derive data representing the radius of the point from the radius residual data and the predicted radius data; and - Update the at least one predicted data based on the radius residual data; wherein updating the at least one predicted data based on the radius residual data includes: - If the radius residual data is greater than the threshold, add new predicted data derived from at least the radius residual data to the top of the list of the at least one predicted data; and - If the radius residual data is below the threshold, the predicted data derived from the selected predictor, or the predicted data associated with the data of the predicted data in the list of at least one predicted data of the predictor representing at least one predicted radius data of the radius representing the points of the point cloud is updated at least based on the radius residual data, and the updated predicted data is moved to the top of the list of the at least one predicted data.
11. A computer program product including instructions that, when executed by one or more processors, cause the one or more processors to perform a method of decoding a point cloud from a bitstream representing encoded point cloud data of a physical object, each point of the point cloud being associated with at least one radius responsive to the distance of the point from a reference object, the method including: - Decode data of predicted data in a list representing at least one predicted data from the bitstream, wherein at least one predictor of predicted radius data representing the radius of the points of the point cloud is derived from the at least one predicted data; - Derive predicted radius data of the point from the predictor; - Decode the radius residual of the point from the bitstream; - Derive data representing the radius of the point from the radius residual data and the predicted radius data; and - Update the at least one predicted data based on the radius residual data; wherein updating the at least one predicted data based on the radius residual data includes: - If the radius residual data is greater than a threshold, add new prediction data derived from at least the radius residual data to the top of the list of the at least one prediction data; and - If the radius residual data is below the threshold, the prediction data derived from the selected predictor, or the prediction data associated with the prediction data in the list of the at least one prediction data of the predictor representing at least one predicted radius data representing the radius of the point of the point cloud is updated at least based on the radius residual data, and the updated prediction data is moved to the top of the list of the at least one prediction data.
12. A non-transitory storage medium carrying instructions of program code, the processor executing the program code to implement a method for decoding a point cloud from a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with at least one radius responsive to the distance of the point from a reference object, wherein the method comprises: - Decode data representing prediction data in a list of at least one prediction data from the bitstream, wherein at least one predictor representing a predicted radius data of a point of the point cloud is derived from the at least one prediction data; - Derive the predicted radius data of the point from the predictor; - Decode the radius residual data of the point from the bitstream; and - Derive data representing the radius of the point from the radius residual data and the predicted radius data; and update the at least one prediction data based on the radius residual; wherein updating the at least one prediction data based on the radius residual data comprises: - If the radius residual data is greater than a threshold, add new prediction data derived from at least the radius residual data to the top of the list of the at least one prediction data; and - If the radius residual data is below the threshold, the prediction data derived from the selected predictor, or the prediction data associated with the prediction data in the list of the at least one prediction data of the predictor representing at least one predicted radius data representing the radius of the point of the point cloud is updated at least based on the radius residual data, and the updated prediction data is moved to the top of the list of the at least one prediction data.
Citation Information
Patent Citations
Method and apparatus of quantizing spherical coorinates used for encoding / decoding point cloud geometry data
EP4020397A1
Method and apparatus of encoding / decoding point cloud geometry data captured by a spinning sensors head
EP4020816A1
A point cloud encoding method, a point cloud decoding method, and related devices
CN111699683A
Systems and Methods for Compressing, Representing and Processing Point Clouds
US20190116372A1