Point Cloud Data Processing Method, Apparatus, Computer Program Product, and Storage Medium

Through the prediction tree encoding method of dynamically adjusting the azimuth step size and spherical coordinate conversion, point cloud encoding is optimized, solving the encoding complexity and delay problems of sparse geometric data captured by spin sensor heads in autonomous driving cars, and achieving efficient point cloud compression.

CN117136385BActive Publication Date: 2025-07-08BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180096607.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-09
Filing Date
2021-10-13
Publication Date
2025-07-08
Estimated Expiration
2041-10-13

AI Technical Summary

Technical Problem

Existing point cloud compression technologies are difficult to achieve simple, low latency and efficient coding on sparse geometric data captured by spin sensor heads, especially under the demands of real-time transmission and decision-making in autonomous vehicles. The existing methods have problems of coding complexity and excessive latency.

Method used

By dynamically adjusting the azimuth step size, combining spherical coordinate transformation and prediction tree coding, the encoding method of point clouds is optimized, and the basic azimuth step size is dynamically scaled to reduce coding costs and improve compression performance.

Benefits of technology

It realizes efficient encoding of sparse geometric data captured on the spin sensor head, reduces encoding delay, improves the compression performance of point clouds, and is suitable for the real-time transmission and decision-making needs of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117136385B_ABST
    Figure CN117136385B_ABST
Patent Text Reader

Abstract

Methods and apparatus are provided for encoding a point cloud into a bitstream of encoded point cloud data representing a physical object / decoding a point cloud from a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with spherical coordinates of an azimuth angle representing a capture angle of a sensor of a spin sensor head that captured the point and a radius responsive to a distance of the point from a reference point. The method includes scaling a basic azimuth step of a predicted azimuth angle for the point based on a decoded radius of the point of the point cloud. Thus, the azimuth step is dynamically scaled and thus depends on the radius of the points of the point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to European Patent Application No. EP21305461.2, filed on April 9, 2021, the content of which is incorporated herein by reference in its entirety. Technical field

[0003] This application generally relates to point cloud compression and, in particular, to methods and apparatuses for encoding / decoding point cloud geometry data captured by a spin sensor head. Background art

[0004] This section is intended to introduce the reader to aspects of the art that may be related to aspects of at least one exemplary embodiment of the present application described and / or claimed below. This discussion is considered to be helpful in providing background information to facilitate a better understanding of the aspects of the present application.

[0005] As a format for representing 3D data, point clouds have recently gained attention because of their various capabilities in representing all types of physical objects or scenes. Point clouds can be used for various purposes, such as cultural heritage / buildings, where objects such as statues or buildings are scanned in 3D in order to share the spatial configuration of the objects without sending or accessing the objects. Moreover, it is a way to ensure the preservation of knowledge of the object in case the object may be damaged; for example, a temple damaged by an earthquake. Such point clouds are typically static, colored, and huge.

[0006] Another use case is in topology and cartography, where the use of 3D representation allows maps to be not limited to a plane and can include landforms. Google Maps is now a good example of a 3D map, but it uses meshes instead of point clouds. However, point clouds can be a suitable data format for 3D maps, and such point clouds are typically static, colored, and huge.

[0007] Virtual reality (VR), augmented reality (AR), and immersive worlds have recently become a hot topic and are foreseen by many as the future of 2D flat video. The basic idea is to immerse the viewer in the surrounding environment, while a standard TV only allows the viewer to watch the virtual world in front of his / her eyes. Depending on the degree of freedom of the viewer in the environment, there are several levels of immersion. Point clouds are good format candidates for distributing VR / AR worlds.

[0008] The automotive industry, especially foreseeable autonomous vehicles, is also an area where point clouds can be used extensively. Autonomous vehicles should be able to "detect" their environment in order to make good driving decisions based on the presence and nature of their closest nearby objects and the road configuration.

[0009] A point cloud is a set of points located in three-dimensional (3D) space, optionally with additional value attached to each point. These additional values are commonly referred to as attributes. Attributes can be, for example, three-component color, material properties (such as reflectivity), and / or two-component normal vectors of the surface associated with the points.

[0010] Thus, a point cloud is a combination of geometry (the positions of points in 3D space, typically represented by 3D Cartesian coordinates x, y, and z) and attributes.

[0011] Point clouds can be captured by various types of devices, such as arrays of cameras, depth sensors, lasers (light detection and ranging, also known as lidar), radars, or can be generated by computers (e.g., in post-production of movies). Depending on the use case, point clouds can have thousands to billions of points for mapping applications. The original representation of point clouds requires a very large number of bits per point, at least a dozen bits for each Cartesian coordinate x, y, or z, and optionally more bits for the (one or more) attributes, such as triple of 10 bits for color.

[0012] In many applications, it is very important to be able to distribute point clouds to end users or store them in servers while consuming only a reasonable amount of bit rate or storage space and maintaining an acceptable (or preferably very good) quality of experience. The efficient compression of these point clouds is a key point to make the distribution chain of many immersive worlds practical.

[0013] For distribution and visualization by end users, such as on AR / VR glasses or any other 3D-enabled device, the compression can be lossy (as in video compression). Other use cases do require lossless compression, such as medical applications or autonomous driving, to avoid changing the results of decisions obtained from subsequent analysis of the compressed and transmitted point clouds.

[0014] Until recently, the mass market has not addressed the problem of point cloud compression (aka PCC), nor are there available standardized point cloud codecs. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11, also known as the Moving Picture Experts Group or MPEG, initiated a work item on point cloud compression. This has led to two standards, namely

[0015] · MPEG-I Part 5 (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC)

[0016] · MPEG-I Part 9 (ISO / IEC 23090-9) or Geometry-based Point Cloud Compression (G-PCC)

[0017] The V-PCC coding method compresses point clouds by performing multiple projections on 3D objects to obtain 2D tiles that are packed into an image (or video when processing dynamic point clouds). The resulting image or video is then compressed using existing image / video codecs, thus allowing for the full utilization of already deployed image and video solutions. By its nature, V-PCC is only efficient on dense and continuous point clouds because image / video codecs cannot compress non-smooth tiles, such as those obtained from projections of sparse geometric data captured by lidar.

[0018] The G-PCC coding method has two schemes for compressing the captured geometric data.

[0019] The first scheme is based on an occupancy tree, which locally can be any type of tree among octree, quadtree, or binary tree, representing the point cloud geometry. Occupied nodes are split until a certain size is reached, and the occupied leaf nodes provide the 3D positions of the points, typically at the centers of these nodes. The occupancy information is carried by occupancy flags that signal the occupancy status of each child node of a node. By using neighbor-based prediction techniques, a high level of compression of the occupancy flags for dense point clouds can be achieved. Sparse point clouds can also be addressed by directly encoding the positions of points within nodes of non-minimal size, stopping tree construction when only isolated points are present in a node; this technique is called the direct coding mode (DCM).

[0020] The second scheme is based on a prediction tree, where each node represents the 3D position of a point, and the parent / child relationship between nodes represents a spatial prediction from the parent to the child. This method can only address sparse point clouds and offers the advantages of lower latency and simpler decoding compared to the occupancy tree. However, compared to the first occupancy-based method, the compression performance is only slightly better, and the encoding is also complex because the encoder has to search intensively (among a long list of potential predictors) for the best predictor when constructing the prediction tree.

[0021] In both schemes, attribute (decoding) encoding is performed after the geometric (decoding) encoding is completed, effectively resulting in two encodings. Therefore, joint geometric / attribute low latency is obtained by using slices that decompose the 3D space into independently encoded sub-volumes without prediction between sub-volumes. When many slices are used, this severely affects the compression performance.

[0022] Combining the requirements for encoder and decoder simplicity, low latency, and compression performance remains a problem that existing point cloud codecs have not satisfactorily addressed.

[0023] An important use case is to transmit sparse geometric data captured by a spin sensor head (e.g., a spin lidar head) mounted on a moving vehicle. This typically requires a simple and low-latency embedded encoder. The simplicity is required because the encoder may be deployed on a computing unit that is concurrently performing other processing (such as (semi-)autonomous driving), thus limiting the processing power available for the point cloud encoder. Low latency is also required to allow for fast transmission from the vehicle to the cloud, in order to view local traffic in real time based on multi-vehicle acquisitions and make decisions quickly enough based on the traffic information. Although the transmission latency can be made low enough by using 5G, the encoder itself should not introduce too much latency due to encoding. Moreover, compression performance is extremely important because the data stream from millions of vehicles to the cloud is expected to be very large.

[0024] Specific priors related to the sparse geometric data captured by the spin sensor head have been used to obtain very efficient encoding / decoding methods.

[0025] For example, G-PCC utilizes the elevation angle (relative to the horizontal ground) captured by the spin sensor head, as Figure 1 and Figure 2 depicted above. The spin sensor head 10 includes a set of sensors 11 (e.g., lasers), five sensors are shown here. The spin sensor head 10 can rotate about the vertical axis z to capture the geometric data of a physical object, i.e., the 3D positions of the points of the point cloud. Then, the geometric data captured by the spin sensor head is represented in spherical coordinates (r 3D , φ, θ), where r 3D is the distance of the point P from the center of the spin sensor head, φ is the azimuth angle by which the sensor head spins relative to a reference, and θ is the elevation angle for the sensor of the spin sensor head relative to the horizontal reference plane (here the y-axis) for elevation angle index k. The elevation angle index k can be, for example, the elevation angle of sensor k, or the k-th sensor position in the case where a single sensor successively detects each of the successive elevation angles.

[0026] A regular distribution along the azimuth angle is observed in the geometric data captured by the spin sensor head, as Figure 3 depicted above. This regularity is used in G-PCC to obtain a quasi-1D representation of the point cloud, where, up to noise, only the radius r 3D belongs to a continuous value range, while the angles φ and θ only take on a discrete number of values, from 0 to I-1, where I is the number of azimuth angles used to capture the points, from 0 to K-1, where K is the number of sensors of the spin sensor head 10. Basically, G-PCC represents the sparse geometric data captured by the spin sensor head on the 2D discrete angular plane (φ, θ) as Figure 3 depicted above and the radius value r 3D for each point.

[0027] This quasi-1D property has been exploited in the occupancy tree and prediction tree in G-PCC by predicting the position of the current point based on the already encoded points by using the discrete nature of the angles in the spherical coordinate space.

[0028] More precisely, the occupancy tree makes heavy use of DCM and entropy-encodes the direct position of the points within the nodes by using a context-adaptive entropy coder. Then the local transformation from the point position to the coordinates (φ,θ) and the context for these coordinates with respect to the discrete angular coordinates (φ i ,θ k ) obtained from the already encoded points are obtained.

[0029] Using the quasi-1D nature (r,φ i ,θ k ) of this coordinate space, the prediction tree directly encodes the first version of the position of the current point in the spherical coordinates (r,φ,θ), where r is the projected radius on the horizontal xy plane, as Figure 4 depicted above by r 2D . Then, the spherical coordinates (r,φ,θ) are converted into 3D Cartesian coordinates (x,y,z), and the xyz residuals are encoded to address the errors in the coordinate conversion, the approximation of the elevation and azimuth angles, and the potential noise.

[0030] Figure 5 Fig. illustrates a point cloud encoder similar to the encoder based on the G-PCC prediction tree.

[0031] First, the Cartesian coordinates (x,y,z) of the points of the point cloud are converted into spherical coordinates (r,φ,θ), (r,φ,θ) = C2A(x,y,z).

[0032] The transformation function C2A(.) is partially given by:

[0033] r = round(sqrt(x*x + y*y) / ΔIr)

[0034] φ = round(atan2(y,x) / ΔIφ)

[0035] where round() is the operation of rounding to the nearest integer value, sqrt() is the square root function, and atan2(y,x) is the arctangent applied to y / x.

[0036] ΔIr and ΔIφ are the internal precisions of the radius and azimuth angle, respectively. They are typically the same as their respective quantization steps, i.e., ΔIφ = Δφ, and ΔIr = Δr,

[0037]

[0038] And,

[0039] Δr = 2 M * Basic quantization step

[0040] where M and N are two parameters of the encoder that can be signaled in the bitstream (e.g., in the geometric parameter set), and where the basic quantization step is typically equal to 1. Typically, for lossless coding, N can be 17, and M can be 0.

[0041] The encoder can derive Δφ and Δr by minimizing the cost (e.g., number of bits) for encoding the spherical coordinate representation and the xyz residuals in the Δφ Cartesian space.

[0042] For simplicity, hereinafter, Δφ = ΔIφ and Δr = ΔIr.

[0043] Also for clarity and simplicity, θ is used hereinafter as the elevation angle value obtained, for example, using the following formula

[0044]

[0045] where atan(.) is the arctangent function. However, in G-PCC, for example, θ is the integer value of the elevation angle index k representing θ k (i.e., the index of the k-th elevation angle), so the operations (prediction, residual (decoding) coding, etc.) described hereinafter performed on θ will be applied to the elevation angle index. Those skilled in the art of point cloud compression will readily understand the advantages of using the index k and how to use the elevation angle index k instead of θ. Furthermore, those skilled in the art of point cloud compression will readily understand that such nuances do not affect the principles of the proposed invention.

[0046] Then the residual spherical coordinates (r n , φ res , θ res , θ res ) between the spherical coordinates (r, φ, θ) and the predicted spherical coordinates obtained from the predictor PR

[0047] (r res , φ res , θ res ) are given by: pred , φ pred , θ pred ) = (r, φ, θ) - (r

[0048] = (r, φ, θ) - (r n , φ n , θ n ) - (0, m * φ step , 0) (1)

[0049] where (rn , φ n , θ n ) are the predicted radius, predicted azimuth angle, and predicted elevation angle obtained from predictors selected from the list of candidate predictors PR0, PR1, PR2, and PR3, and m is the integer number of basic azimuth steps φ to be added to the predicted azimuth angle step .

[0050] The encoder can derive the basic azimuth step φ based on the frequency at which the spin sensor head performs captures at different elevation angles and the rotational speed step , for example, based on the number of detections per head turn NP:

[0051]

[0052] For example, the basic azimuth step φ step or the number of detections per head turn NP is encoded in the bitstream B in the geometric parameter set. Optionally, NP is a parameter of the encoder, which can be signaled in the bitstream in the geometric parameter set, and φ step is derived similarly in both the encoder and the decoder

[0053] The residual spherical coordinates (r res , φ res , θ res ) can be encoded in the bitstream B

[0054] The residual spherical coordinates (r res , φ res , θ res ) can be quantized (Q) to the quantized residual spherical coordinates Q(r res , φ res , θ res ). The quantized residual spherical coordinates Q(r res , φ res , θ res ) can be encoded in the bitstream B

[0055] For each node of the prediction tree, φ step the prediction index n and the quantity m are signaled in the bitstream B, and the basic azimuth step has a certain fixed-point precision, shared by all nodes of the same prediction tree

[0056] The prediction index n points to the selected predictor in the list of candidate predictors

[0057] The candidate predictor PR0 can be equal to (r min , φ0, θ0), where r min is the minimum radius value (provided in the geometric parameter set), and if the current node (current point P) has no parent node, then φ0 and θ0 are equal to 0, or equal to the azimuth angle and elevation angle of the point associated with the parent node

[0058] Another candidate predictor PR1 can be equal to (r0, φ0, θ0), where r0, φ0, and θ0 are the radius, azimuth angle, and elevation angle of the point associated with the parent node of the current node, respectively.

[0059] Another candidate predictor PR2 can be equal to a linear prediction of the radius, azimuth angle, and elevation angle using the radius, azimuth angle, and elevation angle (r0, φ0, θ0) of the point associated with the parent node of the current node and the radius, azimuth angle, and elevation angle (r1, φ1, θ1) of the point associated with the grandparent node.

[0060] For example, PR2 = 2*(r0, φ0, θ0) - (r1, φ1, θ1)

[0061] Another candidate predictor PR3 can be equal to a linear prediction of the radius, azimuth angle, and elevation angle using the radius, azimuth angle, and elevation angle (r0, φ0, θ0) of the point associated with the parent node of the current node, the radius, azimuth angle, and elevation angle (r1, φ1, θ1) of the point associated with the grandparent node, and the radius, azimuth angle, and elevation angle (r2, φ2, θ2) of the point associated with the great-grandparent node.

[0062] For example, PR3 = (r0, φ0, θ0) + (r1, φ1, θ1) - (r2, φ2, θ2)

[0063] The predicted Cartesian coordinates (x pred , y pred , z pred ) are obtained by inverse-transforming the decoded spherical coordinates (r dec , φ dec , θ dec ) as follows:

[0064] (x pred , y pred , z pred ) = A2C(r dec , φ dec , θ dec ) (3)

[0065] where the spherical coordinates (r dec , φ dec , θ dec ) decoded by the decoder can be given by:

[0066] (r dec , φ dec , θ dec ) = (r res,dec , φ res,dec , θ res,ded ) + (r pred , φpred , θ pred ) =

[0067] (r res,dec , φ res,dec , θ res,dec ) + (r n , φ n , θ n ) + (0, m * φ step , 0) (4)

[0068] where (r res,dec , φ res,dec , θ res,dec ) is the residual spherical coordinate decoded by the decoder.

[0069] The decoded residual spherical coordinate (r res,dec , φ res,dec , θ res,dec ) can be the result of the inverse quantization (IQ) of the quantized residual spherical coordinate Q(r dec , φ res , θ res ).

[0070] In G-PCC, the residual spherical coordinate is not quantized, and the decoded spherical coordinate (r res,dec , φ res,dec , θ res,dec ) is equal to the residual spherical coordinate (r res , φ res , θ res ). Thus, the decoded spherical coordinate (r dec , φ dec , θ dec ) is equal to the spherical coordinate (r, φ, θ).

[0071] The inverse transformation of the decoded spherical coordinate (r dec , φ dec , θ dec ) can be given by the following formula:

[0072] r = r dec * Δr

[0073] x pred = round(r * cos(φ dec * Δφ))

[0074] y pred = round(r * sin(φ dec * Δφ)

[0075] z pred = round(tan(θ dec ) * r)

[0076] where sin() and cos() are the sine and cosine functions. These two functions can be approximated by operations based on fixed-point precision. The value of tan(θ dec ) can also be stored as a fixed-point precision value. Therefore, floating-point operations are not used in the decoder. Avoiding floating-point operations is generally a strong requirement for simplifying the hardware implementation of the codec.

[0077] The residual Cartesian coordinates (x pred , y pred , z pred ) between the original point and the predicted Cartesian coordinates (x res , y res , z res ) are given by:

[0078] (x res , y res , z res ) = (x, y, z) - (x pred , y pred , z pred )

[0079] The residual Cartesian coordinates (x res , y res , z res ) are quantized (Q) and the quantized residual Cartesian coordinates Q(x res , y res , z res ) are encoded into the bitstream.

[0080] When the quantization step sizes of x, y, z are equal to the original point precision (usually 1), the residual Cartesian coordinates can be losslessly encoded, or when the quantization step size is greater than the original point precision (usually the quantization step size is greater than 1), they can be lossily encoded.

[0081] The Cartesian coordinates (x dec , y dec , z dec ) decoded by the decoder are given by:

[0082] (x dec , y dec , z dec ) = (x pred , y pred , z pred ) + IQ(Q(x res , y res , z res ) (5)

[0083] where IQ(Q(x res , y res , z res)) represent the quantized residual Cartesian coordinates of inverse quantization.

[0084] Those decoded Cartesian coordinates (x dec , y dec , z dec ) can be used by the encoder, for example, for sorting (decoding) points before attribute coding.

[0085] Figure 6 Illustrated is a point cloud decoder similar to the prediction tree decoder based on the G-PCC prediction tree.

[0086] For each node of the prediction tree, access the prediction index n and the quantity m from the bitstream B, while the basic azimuth step φ step or the number of turns per probe NP is accessed from the bitstream B (e.g., from the parameter set) and shared by all nodes of the same prediction tree.

[0087] The decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ) can be obtained by decoding the residual spherical coordinates (r res , φ res , θ res ) from the bitstream B.

[0088] The quantized residual spherical coordinates Q(r res , φ res , θ res ) can be decoded from the bitstream B. The quantized residual spherical coordinates Q(r res , φ res , θ res ) are inverse quantized to obtain the decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ).

[0089] The decoded spherical coordinates (r dec , φ dec , θ dec ) are obtained by adding the decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ) to the predicted spherical coordinates (r pred , φ pred , θ pred ) according to Equation (4).

[0090] The predicted Cartesian coordinates (x pred , y pred , z pred ) are obtained by converting the decoded spherical coordinates (r dec , φ dec, θ dec ) is obtained by performing an inverse transform.

[0091] Decode the quantized residual Cartesian coordinates Q(x res , y res , z res ) from the bitstream B and perform inverse quantization to obtain the inverse quantized Cartesian coordinates IQ(Q(x res , y res , z res ). The decoded Cartesian coordinates (x dec , y dec , z dec ) are given by Equation (5).

[0092] In G-PCC, Equations (1) and (4) are used for the predicted azimuth angle φ and to obtain the decoded azimuth angle φ dec , and Equation (5) is used to decode the Cartesian coordinates of the points.

[0093] Then, when the radius is small enough (i.e., when the point is close enough to the LiDAR sensor), using different values of "m" (e.g., m + 1 or m - 1) in Equations (1) and (4) will produce the same (quantized) Cartesian residual values.

[0094] Figure 7 The above figure illustrates this defect of G-PCC. In Figure 7 's example, the points of the decoded point cloud (black circles) belong to a square in the xy Cartesian space. These squares correspond to the regular sampling of the x and y axes in the Cartesian space. In the case of lossless compression, the size of the square depends on the accuracy of the input point cloud, or roughly on the quantization step in the case of lossy compression. The angular sector with angle φ step represents the area covered by the sensors of the spin sensor head 10 between two captures. Here, the point belongs to the square intersected by two laser beams b1 and b2 corresponding to the orientations of the two sensors, for two consecutive detections / captures at the same elevation angle. Thus, the square is covered by three angular sectors s0, s1, and s2. Then, to calculate the predicted azimuth angle (Equation (1) or (4)), the quantity "m" should be selected here between two values m1 (associated with laser beam b1) and m2 (associated with laser beam b2). It is possible to encode several (here 2, m1, and m2) different m values for the same point of the point cloud, which can lead to a suboptimal encoding of the quantity m (e.g., m1) compared to the encoding of another quantity m that would be obtained if φ step were high enough for a single laser beam to pass through the square to which the point belongs. Therefore, with respect to the radius of the point (i.e., its distance from the sensor), by using the basic azimuth step φ stepThe angular precision obtained is too high compared to the Cartesian precision output. Thus, it is possible to encode several different quantities indicating the sub-optimality of the encoding scheme and compression can be improved.

[0095] To improve G-PCC, a better encoding of the number m of basic azimuth steps is needed. Summary of the Invention

[0096] The following section presents a simplified summary of at least one exemplary embodiment to provide a basic understanding of some aspects of the present application. This summary is not an exhaustive overview of the exemplary embodiments. It is not intended to identify key or critical elements of the embodiments. The following summary only presents some aspects of at least one of the exemplary embodiments in a simplified form as a prelude to the more detailed description provided elsewhere herein.

[0097] According to a first aspect of the present application, there is provided a method of encoding a point cloud into a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with spherical coordinates representing an azimuth angle of a sensor of a spin sensor head in response to capturing the point and a radius in response to the distance of the point from a reference point. The method includes obtaining a scaled basic azimuth step associated with a point of the point cloud, the scaled basic azimuth step being equal to a first data when a second data is strictly below a threshold, the first data being greater than the basic azimuth step, and in other cases the scaled basic azimuth step being equal to the basic azimuth step, the basic azimuth step being derived from the frequency and rotational speed of the spin sensor head capturing the point cloud, and the second data depending on the decoded radius of the point obtained by encoding and decoding the radius associated with the point; encoding the number of scaled basic azimuth steps obtained from the azimuth angle of the point, the prediction of the azimuth angle, and the scaled basic azimuth step into the bitstream; and encoding the residual azimuth angle of the point between the azimuth angle of the point and the predicted azimuth angle derived from the number of scaled basic azimuth steps and the scaled basic azimuth step into the bitstream.

[0098] According to a second aspect of the present application, a method for decoding a point cloud from a bitstream representing encoded point cloud data of a physical object is provided, wherein each point of the point cloud is associated with spherical coordinates representing an azimuth angle corresponding to a capture angle of a sensor of a spin sensor head that captured the point and a radius corresponding to the distance of the point from a reference point. The method includes decoding a basic azimuth step size from the bitstream; obtaining a decoded radius of a point in the point cloud from a decoded residual radius decoded from the bitstream; obtaining a scaled basic azimuth step size associated with a point of the point cloud, wherein when a second data depending on the decoded radius of the point is strictly less than a threshold, the scaled basic azimuth step size is equal to a first data, the first data being greater than the basic azimuth step size, and in other cases the scaled basic azimuth step size is equal to the basic azimuth step size; decoding a number of scaled basic azimuth step sizes from the bitstream; decoding a decoded residual azimuth angle from the bitstream; and obtaining a decoded azimuth angle from the decoded residual azimuth angle and a predicted azimuth angle derived from the number of scaled basic azimuth step sizes and the scaled basic azimuth step size.

[0099] In one exemplary embodiment, the first data depends on the decoded radius.

[0100] In one exemplary embodiment, the first data is inversely proportional to the product of the decoded radius and a scaling factor greater than or equal to 1.

[0101] In one exemplary embodiment, the second data is the decoded radius (r dec ).

[0102] In one exemplary embodiment, the second data is obtained by applying a monotonic function to the decoded radius and comparing the second data with a second threshold obtained by applying the same monotonic function to the threshold.

[0103] In one exemplary embodiment, the monotonic function is defined as a function that provides integer bounds for the residual azimuth angle.

[0104] In one exemplary embodiment, the first data is obtained from an approximation of 2π / (r dec *α*ΔIφ), where r dec is the decoded radius, α is a scaling factor greater than or equal to 1, and ΔIφ corresponds to the internal precision of the azimuth angle.

[0105] In one exemplary embodiment, the first data is obtained by iteratively refining the approximation.

[0106] In one exemplary embodiment, the approximation is obtained by finding the highest power of factor 2 of the basic azimuth step size that is less than 2π / (r dec *α*ΔIφ).

[0107] According to a third aspect of the present application, there is provided an apparatus for encoding a point cloud into a bitstream of encoded point cloud data representing a physical object. The apparatus includes one or more processors configured to execute the method according to the first aspect of the present application.

[0108] According to a fourth aspect of the present application, there is provided an apparatus for decoding points of a point cloud representing a physical object from a bitstream. The apparatus includes one or more processors configured to execute the method according to the second aspect of the present application.

[0109] According to a fifth aspect of the present application, there is provided a computer program product including instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to the first aspect of the present application.

[0110] According to a sixth aspect of the present application, there is provided a computer program product including instructions that, when the program is executed by one or more processors, cause the one or more processors to execute the method according to the second aspect of the present application.

[0111] According to a seventh aspect of the present application, there is provided a non-transitory storage medium carrying instructions for program code for executing the method according to the second aspect of the present application.

[0112] The specific nature of at least one of the exemplary embodiments and other objects, advantages, features, and uses of at least one of the exemplary embodiments will become apparent from the following description of the examples in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0113] Now, by way of example, reference will be made to the drawings showing exemplary embodiments of the present application, and in which:

[0114] Figure 1 A side view showing a sensor head according to the prior art and some of its parameters is illustrated;

[0115] Figure 2 A top view showing a sensor head according to the prior art and some of its parameters is illustrated;

[0116] Figure 3 A regular distribution of data captured by a spin sensor head according to the prior art is illustrated;

[0117] Figure 4 A representation of points in a 3D space according to the prior art is illustrated;

[0118] Figure 5 A point cloud encoder similar to a G-PCC prediction tree-based encoder according to the prior art is illustrated;

[0119] Figure 6 illustrates a point cloud decoder similar to that of a G-PCC prediction tree-based decoder according to the prior art;

[0120] Figure 7 illustrates the disadvantages of G-PCC according to the prior art;

[0121] Figure 8 illustrates a block diagram of steps of a method 100 for encoding a point cloud representing a physical object according to at least one exemplary embodiment;

[0122] Figure 9 illustrates a block diagram of steps of a method 200 for decoding a point cloud representing a physical object according to at least one exemplary embodiment;

[0123] Figure 10 illustrates an example of a method 300 for determining an approximation of a scaled basic azimuth step according to at least one exemplary embodiment;

[0124] Figure 11 illustrates an example of a method 400 for determining an approximation of a scaled basic azimuth step according to at least one exemplary embodiment; and

[0125] Figure 12 illustrates a schematic block diagram of an example of a system in which various aspects and exemplary embodiments are implemented.

[0126] Like reference numerals may have been used to denote like components in different figures. Detailed Description

[0127] Exemplary embodiments of at least one will be described more fully hereinafter with reference to the accompanying drawings, in which examples of exemplary embodiments of at least one are illustrated. However, the exemplary embodiments may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Thus, it is to be understood that the exemplary embodiments are not intended to be limited to the particular forms disclosed. On the contrary, this disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of this application.

[0128] When the accompanying drawings are presented in the form of a flowchart, it should be understood that they also provide a block diagram of the corresponding apparatus. Similarly, when the accompanying drawings are presented in the form of a block diagram, it should be understood that they also provide a flowchart of the corresponding method / process.

[0129] At least one of these aspects generally relates to point cloud encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream.

[0130] Moreover, the present aspect is not limited to MPEG standards such as MPEG-I Part 5 or Part 9 related to point cloud compression, and can be applied to, for example, other standards and recommendations, whether pre-existing or future-developed, as well as extensions of any such standards and recommendations (including MPEG-I Part 5 and Part 9). Unless otherwise indicated or technically precluded, the aspects described in this application can be used alone or in combination.

[0131] The present invention relates to the technical fields of encoding and decoding, and aims to provide a technical solution for encoding / decoding point cloud data. Since point cloud is a collection of massive data, storing point cloud consumes a large amount of memory, and it is impossible to directly transmit point cloud at the network layer without compressing it. Therefore, point cloud compression is required. Thus, as point cloud is increasingly widely used in autonomous navigation, real-time inspection, geographic information services, cultural heritage / architecture protection, 3D immersive communication and interaction, etc., the present invention can be used in many application scenarios.

[0132] This encoding / decoding method specifically relates to encoding / decoding point cloud data based on the azimuth step size of dynamic list scaling to improve the compression performance of point cloud.

[0133] The present invention relates to a method for encoding a point cloud into a bitstream of encoded point cloud data representing a physical object / decoding a point cloud from a bitstream of encoded point cloud data representing a physical object, wherein each point of the point cloud is associated with spherical coordinates of an azimuth angle representing the capture angle of a sensor of a spin sensor head that captures the point and a radius representing the distance of the point from a reference point.

[0134] The method includes scaling a basic azimuth step size for predicting the azimuth angle of a point of the point cloud from the decoded radius of the point. Thus, the azimuth step size is dynamically scaled and thus depends on the radius of the points of the point cloud.

[0135] When the radius of the point is too small to efficiently use the basic azimuth step size φ step then, a plurality of quantities m can be selected, such as Figure 7As shown above. However, scaling the basic azimuth step size by a value greater than 1 statistically reduces the number of corner sectors of all squares to which any point belonging to the same radius belongs, thus statistically reducing or even eliminating the use of an unnecessary additional number of basic azimuth step sizes to be encoded for these points, because on average, only a single (or few) corner sectors cover the square to which the point belongs (if the scaling factor is properly selected). The effect of this scaling is to reduce the average number m to be encoded (averaged over all encoded points), and implicitly reduce the average cost of encoding the number m of basic azimuth step sizes (the number of bits averaged over all encoded points), because the encoding of an unnecessary additional number of basic azimuth step sizes is avoided (or at least reduced).

[0136] Figure 8 FIG. illustrates a block diagram of steps of a method 100 for encoding a point cloud representing a physical object according to at least one exemplary embodiment.

[0137] In step 110, the basic azimuth step size φ step can be given by equation (2).

[0138] In step 120, the decoded radius r of a point in the point cloud can be obtained by equation (4) dec :

[0139] r dec = r res,dec + r n (6)

[0140] where r n is the predicted radius given by the predictor PR n .

[0141] In step 130, when the second data D2 depending on the decoded radius r dec is strictly lower than the threshold TH, the scaled basic azimuth step size S(φ step ,r dec ) is equal to the first data D1, which is greater than the basic azimuth step size, otherwise the scaled basic azimuth step size S(φ step ,r dec ) is equal to the basic azimuth step size φ step .

[0142] In step 140, the number m s of scaled basic azimuth step sizes is encoded into the bitstream B. The number m n is obtained from the azimuth angle φ of the point, the predicted azimuth angle φ step and the scaled basic azimuth step size S(φ dec ,r s ):

[0143] m s = round((φ - φ n ) / S(φ step , r dec ))

[0144] In step 150, the predicted azimuth angle φ step , r dec ) and the quantity m s are derived from the scaled basic azimuth step S(φ pred :

[0145] φ pred = φ n + m s * S(φ step , r dec ) (7)

[0146] In step 160, the residual azimuth angle φ res is encoded in the bitstream B. The residual azimuth angle φ res can be calculated between the azimuth angle φ of the point and the predicted azimuth angle φ pred .

[0147] φ res = φ - φ pred = φ - φ n - m s * S(φ step , r dec ) (8)

[0148] Figure 9 FIG. illustrates a block diagram of the steps of a method 200 for decoding a point cloud representing a physical object according to at least one exemplary embodiment.

[0149] In step 210, the basic azimuth step φ step or the number of detections per turn NP can be decoded from the bitstream B, for example, from a geometric parameter set.

[0150] In step 220, the decoded residual radius r res,dec is decoded from the bitstream B. The decoded radius r dec can be obtained from equation (6).

[0151] In step 130, the scaled basic azimuth step (S(φ step , r dec )) associated with the points of the point cloud is obtained.

[0152] In step 230, the number m s of scaled basic azimuth steps is decoded from the bitstream B.

[0153] In step 150, the predicted azimuth angle φ is derived from equation (7). pred .

[0154] In step 240, the decoded residual azimuth angle φ res,dec is decoded from the bitstream B.

[0155] In step 250, the decoded azimuth angle φ res,dec is obtained from the decoded residual azimuth angle φ pred and the predicted azimuth angle φ dec :

[0156] φ dec = φ res,dec + φ pred = φ res,dec +φ n + m s * S(φ step , r dec ) (9)

[0157] In one exemplary embodiment, the present invention can be used in the G-PCC prediction scheme given by equation (1) or (4), where the azimuth angle φ associated with the points of the point cloud is adaptively quantized as described in European Patent Application No. EP20306674.

[0158] Combined Figure 5 and Figure 6 described Q(r res ,φ res ,θ res ) is set equal to (r res ,Qφ res ,θ res ), where r res is the residual radius, Qφ res is the adaptively quantized residual azimuth angle, as described below, and θ res is the non-quantized elevation angle (index) residual.

[0159] On the encoding side, the residual radius r res is obtained from equation (1), and the decoded radius r dec is obtained from equation (4).

[0160] The residual azimuth angle φ res obtained from equation (1) is adaptively quantized by:

[0161] Qφ res = Q φ (φ res ,r dec ) = round(φ res / Δφ(r dec)) (10)

[0162] where Q φ is an adaptive quantizer using a quantization step Δφ(r dec ) given by:

[0163] Δφ(r dec ) = Δφ arc / r dec , (11)

[0164] where Δφ arc is the arc quantization step.

[0165] If the codec uses fixed-point precision arithmetic and the internal precision ΔIφ of the codec's representation for azimuth is given by ΔIφ = 2π / 2 N , then the arc quantization step Δφ arc is set to be equal to 2π / (x * ΔIφ) = 2 N / x, where for example x = 8.

[0166] The residual azimuth IQφ for inverse quantization res is obtained by inverse quantization Qφ res through the following formula:

[0167] IQφ res = IQ φ (Qφ res , r dec ) = Qφ res * Δφ(r dec ) (12)

[0168] where IQ φ is the adaptive inverse quantizer based on the decoded radius r dec .

[0169] The decoded azimuth φ dec can be given by:

[0170] φ dec = IQφ res + φ n + m s * S(φ step , r dec ) (13)

[0171] or by:

[0172] φ dec = Qφ res * Δφ(r dec ) + φ n + m s * S(φ step , r dec)

[0173] = Qφ res * Δφ arc / r 十进制 + φ n + m s * S(φ step , r dec ) (14)

[0174] Quantized residual azimuth Qφ res is encoded in the bitstream B.

[0175] In an exemplary embodiment, the residual azimuth φ between the azimuth of a point of the point cloud and the predicted angle of the azimuth res can be improved by encoding / decoding using a bounding property given by:

[0176] |φ res | ≤ B

[0177] where |φ res | is the absolute value of the residual azimuth φ res and is bounded by an integer bound B, which is given by:

[0178] B = Q φ (φ step / 2, r dec )

[0179] = round(r dec *(φ step / 2) / Δφ arc )) = round(r dec * φ step / (2 * Δφ arc )) (15)

[0180] where Δφ arc = 2π / (x * ΔIφ) = 2 N / x, for example x = 8.

[0181] In a first exemplary embodiment of step 130, the first data D1 depends on the decoding radius r dec and the second data D2 is the decoding radius r dec .

[0182] In this first embodiment, when the decoding radius r dec (D2) is strictly below the threshold TH, the scaled basic azimuth step S(φ step , r dec ) is equal to the first data D1, and the first data D1 is greater than the basic azimuth step φ step, while in other cases the basic azimuth step size S(φ step ,r dec ) is equal to the basic azimuth step size φ step .

[0183] In a variant, the first data D1 is inversely proportional to the product of the decoding radius and a scaling factor α greater than or equal to 1.

[0184] For example,

[0185] D1 = 2π / (r dec *α*ΔIφ) (16)

[0186] When the decoding radius r dec is small enough (i.e., when r dec *α is less than the number of detections per turn NP, or equivalently r dec <NP / α), this defines the basic azimuth step size S(φ step ,r dec ) as greater than the basic azimuth step size φ step . Then, for a "small" decoding radius, the use of an unnecessary extra number of basic azimuth steps is reduced, and the coding cost (i.e., the number of bits required for coding) of the quantity m s is reduced. The optimal threshold to use will be TH = NP / α.

[0187] In a variant, if the internal precision is then the residual azimuth angle is adaptively quantized by equation (10), where Δφ arc = 2 N / x and α is set equal to x, equation (16) can be rewritten as:[[]]

[0188] D1 = 2 N / (r dec *α) (17)

[0189] In a variant, the threshold TH is equal to the threshold Th0 given by:[[]]

[0190] Th0 = Δφ arc / φ step (18)

[0191] When the residual azimuth angle is bounded by an integer bound B, the threshold Th0 can be obtained as follows: When B = 0, it is best to use S(φ step ,r dec ) = D1, while when B > 0, it is best to use S(φ step ,r dec ) = φ step。Then, according to Equation (15), when the following equivalent inequality is satisfied, the bound B is equal to 0, that is, the bound B is strictly less than 1 (B < 1), because the bound B is a positive number rounded to an integer value:

[0192]

[0193] Then, by solving the following equation, the threshold Th0 is derived from the upper bound of the inequality that appears when B = 1:

[0194]

[0195] Thus, we obtain:

[0196] Th0 = (1 - 0.5) * Δφ arc * 2 / φ step = Δφ arc / φ step (19)

[0197] The threshold Th0 given by Equation (18) can be used either in variants where the residual azimuth is adaptively quantized by Equation (10) or not by Equation (10), or in variants where the residual azimuth is bounded by an integer bound B or not by an integer bound B.

[0198] In a variant, if the internal precision the residual azimuth is adaptively quantized by Equation (10), where Δφ arc = 2 N / x and α is set to be equal to x, then the threshold Th0 is given by:

[0199] Th0 = 2 N / (α * φ step ) (20)

[0200] In the second exemplary embodiment of step 130, the second data D2 is obtained by applying a monotonic function m(.) to the decoding radius r dec and the second data D2 is compared with a threshold TH obtained by applying the same monotonic function m(.) to the threshold Th0:

[0201]

[0202] If the monotonic function m(.) is a monotonically increasing function, then when the second data D2 is greater than or equal to the threshold TH, the scaled basic azimuth step S(φ step , r dec ) is equal to the basic azimuth step φ step , and when the second data D2 is strictly lower than the threshold TH, the scaled basic azimuth step S(φ step,r dec ) is equal to the first data D1.

[0203] Using a monotonically increasing or decreasing function is equivalent, and the scope of the present invention extends to a monotonically increasing or decreasing function m(.).

[0204] For example, if the function m(.) is a monotonically decreasing function, then an equivalent monotonically increasing function m'(x) can be constructed from m(x), for example, m'(x) = -m(x), in order to continue using a monotonically increasing function.

[0205] As another example, if the function m(.) is a directly used monotonically decreasing function, then it is obvious that a method equivalent to the method described in conjunction with the drawings can be obtained: for example, when the second data D2 is strictly lower than the threshold TH, the scaled basic azimuth step S(φ step ,r dec ) will be equal to the basic azimuth step φ step , and when the second data D2 is greater than or equal to the threshold TH, the scaled basic azimuth step S(φ step ,r dec ) will be equal to the first D1.

[0206] In a variant, the monotonic function m(.) is scaled by a scaling factor α greater than or equal to one.

[0207] Then, the second data D2 and the threshold TH are given by

[0208]

[0209] In a variant, if the internal precision the residual azimuth angle is adaptively quantized by equation (10) and Δφ arc = 2 N / x, and α is set equal to x, then the second data D2 and the threshold TH are given by

[0210]

[0211] In a variant, the monotonic function m(.) is defined as a function that provides the integer bound B in equation (15).

[0212] Then, the second data D2 and the threshold TH are given by:

[0213]

[0214] In that variant, since the integer bound B (equation 15) is an integer, the condition B > 0 is equivalent to D2 ≥ TH (i.e., D2 ≥ 1), and the condition B = 0 is equivalent to D2 < 1. When B = 0 (D2 < 1), the scaled basic azimuth step S(φstep , r dec ) is equal to or greater than the basic azimuth step φ step of the first data D1 (Equations 16 or 17), and the scaled basic azimuth step S(φ step , r dec ) is equal to the basic azimuth step φ in other cases step .

[0215] If the integer bound B of Equation (15) is used to bound the encoding of the residual azimuth angle φ res , then it is clearly preferable to use a variant of Equation (23), because the second data D2 is equal to the integer bound B that has already been calculated for each point, and the threshold TH is always 1. In other cases, when the integer bound B is not used to bound the residual azimuth angle, the first embodiment or its variant or one of Equations 21 to 22 is preferable, because the threshold TH only needs to be calculated once for each point cloud, and the second data D2 in Equation (23) requires integer division to be calculated for each point of the point cloud.

[0216] When the second data D2 (integer bound B) is given by Equation (23), the integer division can be calculated as a bitwise shift operation. For example, if the internal precision of the codec is used for the azimuth angle and Δφ arc = 2 N / 8 fixed-point representation, then the integer bound B (D2) is given by:

[0217] D2 = B = round(r dec * φ step / (2 * Δφ arc )) = round(r dec * φ step / (2 N-3 ))

[0218] = (r dec * φ step + (1 << (N - 2))) >> (N - 3) (24)

[0219] where the result of (a) << (b) is (a) shifted left by (b) bits (i.e., the result of (a) multiplied by 2 b ), and the result of (a) >> (b) is (a) shifted right by (b) bits (i.e., the result of the integer division of (a) by 2 b ).

[0220] Equation (16) involves integer division for each point of the point cloud. This is something to be avoided if possible, because division is costly in terms of hardware design, execution runtime, and / or power consumption.

[0221] In a preferred variant of the previous embodiment, the division operation is approximated by an operation with lower hardware cost.

[0222] There are many possibilities for approximating integer division. For example, the division u / v can be simply approximated by a right bitwise shift operation of the number of bits of u equal to the number of bits NV taken by v (i.e., NV = floor(log2(v)+1)). In that case, u / v is approximated as u / 2 NV . Another approach could be to use the Newton - Raphson iteration algorithm (https: / / en.wikipedia.org / wiki / Division_algorithm) in a finite number of iterations. Or, since it has been used in G - PCC, an approximation with a fixed - point precision of 1 / v based on a lookup table can also be used, and u is multiplied by the approximated 1 / v, and then the fixed - point result is rounded to an appropriate fixed - point (or integer) precision.

[0223] Compared with true integer division, due to the approximation, a small error is introduced in the result, so equation (16) will become:

[0224] D1 app = divApprox(2 N ,(r dec *α)) = 2 N / (r dec *α)+ε(r dec ) (25)

[0225] where divApprox(u,v) is an approximation function of the integer division u / v, and +ε(r dec ) is the error introduced by the division approximation.

[0226] For example, divApprox(2 N ,(r dec *α)) = 2 N-M , where M = floor(log2(r dec *α)+1) is the number of bits occupied by r dec *α.

[0227] Due to the approximation error ε(r dec ), compared with using D1 (Equation 16), encoding loss is introduced when using D1 app .

[0228] Three different cases occur:

[0229] 1) D1 app > D1, that is, ε(r dec )>0

[0230] Then, since D1app is too high, so the bit rate will be slightly reduced because the number m of scaled basic azimuth steps s will be smaller, but more distortion will be introduced in the predicted (x pred , y pred ) coordinates.

[0231] 2) D1 app <D1, that is, ε(r dec ) < 0

[0232] Then, since D1 app is too small, the number m of scaled basic azimuth steps s can be slightly increased, so the cost of its encoding will also increase.

[0233] 3) D1 app = D1, so the optimal efficiency is obtained.

[0234] In our experiments, it is generally observed that for the same absolute value of the error ε(r dec ), the first case is more critical than the second case. Therefore, it is preferably to slightly increase the bit rate rather than introduce prediction errors in the (x pred , y pred ) coordinates. Therefore, it is preferable to avoid the first case.

[0235] In a variant of equation (16), the first data D1 (scaled basic azimuth step S(φ step , r dec )) obtained by equation (16) is refined by iteratively increasing it and / or optionally by iteratively decreasing it.

[0236] Figure 2 Illustrates an example of a method 300 for determining an approximation of a scaled basic azimuth step S(φ step , r dec ) according to at least one exemplary embodiment.

[0237] Briefly, the method iteratively refines the scaled basic azimuth angle step S(φ step , r dec ) using some simple operations (bitwise shift operations, increment and compare) in each refinement step.

[0238] As previously explained, when the decoding radius r dec (D2) is greater than or equal to the threshold TH, the scaled basic azimuth step S(φ step , r dec ) is set equal to the basic azimuth step φ step . When the decoding radius r decWhen (D2) is strictly less than the threshold TH (or when the integer bound B calculated from the decoding radius r dec and from the basic azimuth step φ step equals 0), in step 310, from the approximation of, from the decoding radius r dec obtain the first basic azimuth step φ step,0 , for example:

[0239] φ step,0 = divApprox(2 N , r dec *α)

[0240] In step 320, obtain the first integer bound B0 value (Equation 15) from the first basic azimuth step φ step,0 and the decoding radius r dec .

[0241] Then, starting from index i = 0, when the integer bound B i is greater than zero, increment index i by one (step 340), and obtain the new basic azimuth step φ step,i-1 by decrementing the previous basic azimuth step φ step,i by one (step 350): φ step,i = φ step,i-1 - 1, and obtain the new integer bound B step,i from the basic azimuth step φ dec and the decoding radius r i (Equation 15) (step 360).

[0242] When the integer bound B i equals 0, the first case above is resolved, and the scaled basic azimuth step S(φ step , r dec ) can be set equal to φ step,i .

[0243] In a variant, to also resolve the second case above, as Figure 10 additionally shown above, when the integer bound B i equals 0, obtain the new basic azimuth step φ step,i by incrementing the current basic azimuth step φ step,i+1 by one (step 370): φ step,i+1 = φ step,i + 1, and obtain the new integer bound B step,i+1 from the basic azimuth step φ dec and the decoding radius r i+1 (Equation 15) (step 380).

[0244] If the integer bound Bi+1 is equal to 0, then it means that φ step,i can still be increased to φ step,i+1 and thus the index i is incremented by one (step 390).

[0245] Then, a new basic azimuth step size is obtained (step 370) and a new integer bound is obtained (step 380).

[0246] This process is iteratively repeated until the integer bound B i+1 is not equal to zero (i.e., it is greater than zero), then the scaled basic azimuth step size S(φ step , r dec ) is set to be equal to the current basic azimuth step size φ step,i .

[0247] In a variant, the addition and / or subtraction operation(s) can use a dynamically determined increment (and / or decrement) step size.

[0248] For example, the increment value is doubled after each iteration of the method.

[0249] Figure 11 Illustrates another example of a method 400 for determining a scaled basic azimuth step size S(φ step , r dec ) according to at least one exemplary embodiment.

[0250] As previously explained, when the decoding radius r dec (D2) is greater than or equal to the threshold TH, the scaled basic azimuth step size S(φ step , r dec ) is set to be equal to the basic azimuth step size φ step . When the decoding radius r dec (D2) is strictly below the threshold TH (or when the integer bound B calculated from the decoding radius r dec and the basic azimuth step size φ step is equal to 0), in step 410, the first basic azimuth step size φ step,0 is set to be equal to the basic azimuth step size φ step , and the associated arc length φ arc,0 is set to be equal to (step 420):

[0251] φ arc,0 = φ step,0 * r dec * α.

[0252] Then, starting from index i = 0, when the arc length φ arc,i is below the threshold Th1, the index i is incremented by 1, and the previous basic azimuth step size φ step,i-1(Step 430) and the previous associated arc length φ arc,i-1 (Step 440) Multiply by two (using a left bitwise shift operation) to obtain a new basic azimuth step φ step,i and the new associated arc length φ arc,i :

[0253]

[0254] When the arc length φ arc,i is no longer less than the threshold Th1, the scaled basic azimuth step S(φ step , r dec ) can be set equal to the basic azimuth step φ step,i .

[0255] Typically, Th1 = π / ΔIφ, and thus the approximate method 400 for determining the scaled basic azimuth step is equivalent to finding the highest power of factor 2 (2 ) of the basic azimuth step φ step below (i.e., satisfying n step and using φ step *2 n = φ step << n as the scaled basic azimuth step S(φ step , r dec ). This can be regarded as a fast approximation of division (Equation 16).

[0256] As an example, using the variant shown in Figure 11 , in the case of Th1 = π / ΔIφ, we are always in the second or third case above, and ~90% of the coding gain obtained when using true division is retained in our experiments. However, if we use Th1 = π / ΔIφ, but allow one more iteration when abs(2*Th1 - 2*φ arc,i ) < abs(2*Th1 - φ arc,i ), then we allow the first case, but we reduce the absolute value of the approximation error ε(r dec ), and only ~72% of the coding gain is retained, even though the error of this division is less than or equal to the previous case. This shows that it is preferable to avoid the third case.

[0257] This encoding / decoding method can be used to encode / decode point clouds, which can be used for various purposes, especially for encoding / decoding point cloud data based on azimuth step scaling of a dynamic list, which improves the compression performance of the point cloud.

[0258] Figure 12A schematic block diagram is shown illustrating an example of a system in which various aspects and exemplary embodiments are implemented.

[0259] System 500 may be embedded as one or more devices, including various components described below. In various embodiments, System 500 may be configured to implement one or more aspects described in the present application.

[0260] Examples of equipment that may constitute all or part of System 500 include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from video decoders, pre-processors that provide input to video encoders, web servers, set-top boxes, and any other devices for processing point clouds, video, or images, or other communication devices. The elements of System 500 may be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 500 may be distributed across multiple ICs and / or discrete components. In various embodiments, System 500 may be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.

[0261] System 500 may include at least one processor 510 that is configured to execute instructions loaded therein for implementing, for example, the various aspects described in the present application. Processor 510 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 500 may include at least one memory 520 (e.g., volatile memory devices and / or non-volatile memory devices). System 500 may include a storage device 540, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 540 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0262] System 500 may include an encoder / decoder module 530, which is configured to process data, for example, to provide encoded / decoded point cloud geometry data, and the encoder / decoder module 530 may include its own processor and memory. The encoder / decoder module 530 may represent one or more modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 530 may be implemented as a separate element of the system 500 or may be incorporated into the processor 510 as a combination of hardware and software known to those skilled in the art.

[0263] Program code to be loaded onto the processor 510 or the encoder / decoder 530 to perform the various aspects described in this application may be stored in the storage device 540 and subsequently loaded onto the memory 520 for execution by the processor 510. According to various embodiments, during the execution of the processes described in this application, one or more of the processor 510, the memory 520, the storage device 540, and the encoder / decoder module 530 may store one or more of various items. Such stored items may include, but are not limited to, point cloud frames, encoded / decoded geometry / attribute video / images or portions thereof, bitstreams, matrices, variables, and intermediate or final results of equations, formulas, operations, and operational logic processing.

[0264] In several embodiments, the memory internal to the processor 510 and / or the encoder / decoder module 530 may be used to store instructions and provide working memory for the processing that may be performed during encoding or decoding.

[0265] However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 510 or the encoder / decoder module 530) is used for one or more of these functions. The external memory may be the memory 520 and / or the storage device 540, for example, dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, fast external dynamic volatile memory such as RAM may be used as the working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), or MPEG-I Part 5 or Part 9.

[0266] As indicated in block 590, inputs to the elements of system 500 can be provided via a variety of input devices. Such input devices include, but are not limited to, (i) an RF portion that can receive RF signals transmitted over the air, for example, by a broadcast device, (ii) composite input terminals, (iii) USB input terminals, and / or (iv) HDMI input terminals.

[0267] In various embodiments, the input devices of block 590 have corresponding input processing elements associated therewith, as is known in the art. For example, the RF portion can be associated with elements necessary for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band), (ii) down-converting the selected signal, (iii) band-limiting the band again to a narrower band to select a signal band that can be referred to as a channel, for example, in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF portion of various embodiments can include one or more elements that perform these functions, for example, a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband.

[0268] In one set-top box embodiment, the RF portion and its associated input processing elements can receive an RF signal transmitted over a wired (e.g., cable) medium. The RF portion can then perform frequency selection by filtering, down-converting, and filtering again to a desired band.

[0269] Various embodiments reorder the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0270] Adding elements can include inserting elements between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.

[0271] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 500 to other electronic devices via the USB and / or HDMI connections. It should be understood that various aspects of the input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processor 510 when necessary. Similarly, various aspects of the USB or HDMI interface processing may be implemented within a separate interface IC or within the processor 510 when necessary. The demodulated, error-corrected, and demultiplexed streams may be provided to various processing elements, including, for example, the processor 510 and the encoder / decoder 530, which operate in conjunction with memory and storage elements to process the data stream as necessary for presentation on an output device.

[0272] The various elements of the system 500 may be provided within an integrated housing. Within the integrated housing, a suitable connection arrangement 590, such as internal buses (including I2C buses), wiring, and printed circuit boards known in the art, may be used to interconnect the various elements and transfer data between them.

[0273] The system 500 may include a communication interface 550 that enables communication with other devices via a communication channel 900. The communication interface 550 may include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 900. The communication interface 550 may include, but is not limited to, a modem or a network card, and the communication channel 900 may be implemented, for example, within a wired and / or wireless medium.

[0274] In various embodiments, data may be streamed to the system 500 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments may be received via the communication channel 900 and the communication interface 550 suitable for Wi-Fi communication. The communication channel 900 of these embodiments may generally be connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-top communications.

[0275] Other embodiments may use a set-top box to provide streamed data to the system 500, which delivers the data via an HDMI connection of the input block 590.

[0276] Still other embodiments may use an RF connection of the input block 590 to provide streamed data to the system 500.

[0277] The streamed data may be used as a way of signaling information used by the system 500. The signaling information may include the bitstream B and / or the number of points such as a point cloud and / or sensor setting parameters (such as the basic azimuth step φ associated with the sensors of the spin sensor head 10)step or the elevation angle θ k ) information.

[0278] It should be recognized that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. can be used to signal information to the corresponding decoder.

[0279] System 500 can provide output signals to various output devices, including display 500, speaker 700, and other peripheral devices 800. In various examples of the embodiments, other peripheral devices 800 can include one or more of a standalone DVR, disc player, stereo system, lighting system, and other devices based on the output-providing function of system 500.

[0280] In various embodiments, control signals can be communicated between system 500 and display 600, speaker 700, or other peripheral devices 800 using signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention.

[0281] The output devices can be communicatively coupled to system 500 via dedicated connections through corresponding interfaces 560, 570, and 580.

[0282] Optionally, the output devices can be connected to system 500 using communication channel 900 via communication interface 550. Display 600 and speaker 700 can be integrated with other components of system 500 in a single unit in an electronic device such as, for example, a television.

[0283] In various embodiments, display interface 560 can include a display driver such as, for example, a timing controller (TCon) chip.

[0284] For example, if the RF portion of input terminal 590 is part of a separate set-top box, then display 600 and speaker 700 can optionally be separate from one or more of the other components. In various embodiments where display 600 and speaker 700 can be external components, output signals can be provided via dedicated output connections including, for example, HDMI ports, USB ports, or COMP outputs.

[0285] In Figure 1-12 , various methods are described herein, and each method includes one or more steps or actions to implement the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined.

[0286] Some examples are described with respect to block diagrams and / or operational flowcharts. Each block represents a circuit element, module, or portion of code that includes one or more executable instructions for implementing the specified logic function(s). It should also be noted that in other embodiments, the function(s) noted in the blocks may not occur in the order indicated. For example, depending on the functions involved, two consecutive blocks shown may actually be executed substantially concurrently, or sometimes the blocks may be executed in the reverse order.

[0287] The embodiments and aspects described herein can be implemented in, for example, a method or process, apparatus, computer program, data stream, bit stream, or signal. Even if discussed only in the context of a single form of embodiment (e.g., only as a method), the embodiments of the features discussed can be implemented in other forms (e.g., an apparatus or a computer program).

[0288] A method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor also includes a communication device.

[0289] Furthermore, a method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the embodiments) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product implemented in one or more computer-readable media and having computer-readable program code executable by a computer implemented thereon. Considering the inherent ability to store information therein and the inherent ability to retrieve information therefrom, a computer-readable storage medium as used herein can be considered a non-transitory storage medium. The computer-readable storage medium can be, but is not limited to, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. It should be recognized that while the following provides more specific examples of computer-readable storage media to which this embodiment can be applied, it is merely illustrative and not an exhaustive list as would be readily recognized by a person of ordinary skill in the art: a portable computer floppy disk; a hard disk; a read-only memory (ROM); an erasable programmable read-only memory (EPROM or flash memory); a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing.

[0290] The instructions can form an application program tangibly implemented on a processor-readable medium.

[0291] For example, the instructions can be in hardware, firmware, software, or a combination thereof. For example, the instructions can be found in an operating system, a separate application, or a combination of both. Thus, a processor can be characterized as, for example, a device configured to execute a process and a device including a processor-readable medium (such as a storage device) having instructions for executing the process. Additionally, in addition to or instead of the instructions, the processor-readable medium can store data values generated by the implementation.

[0292] The apparatus can be implemented in, for example, suitable hardware, software, and firmware. Examples of such apparatus include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, head-mounted display devices (HMDs, see-through glasses), projectors (projectors), "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from the video decoder, pre-processors that provide input to the video encoder, web servers, set-top boxes, and any other device for processing point clouds, video, or images, or other communication devices. It should be clear that the equipment can be mobile and even installed in a moving vehicle.

[0293] The computer software can be implemented by the processor 510 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can also be implemented by one or more integrated circuits. The memory 520 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). The processor 510 can be of any type suitable for the technical environment and can encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.

[0294] As will be apparent to those of ordinary skill in the art, the implementations can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for executing a method or data generated by one of the described implementations. For example, the signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. The formatting can include, for example, encoding the data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted via various different wired or wireless links. The signal can be stored on a processor-readable medium.

[0295] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" may also be intended to include the plural forms, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "include / comprise" and / or "including / comprising" may specify the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. Also, when an element is referred to as being "responsive" or "connected" to another element, it may be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to another element, no intervening elements are present.

[0296] It should be recognized that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any of the symbols / terms " / ", "and / or", and "at least one" may be intended to cover the selection of only the first-listed option (A), or the selection of only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such language is intended to cover the selection of only the first-listed option (A), or the selection of only the second-listed option (B), or the selection of only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.

[0297] Various numerical values may be used in this application. Specific values may be used for illustrative purposes and the aspects described are not limited to these specific values.

[0298] It will be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of this application, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. No order is implied between the first element and the second element.

[0299] References to "an exemplary embodiment" or "exemplary embodiments" or "an embodiment" or "embodiments" and other variations thereof are frequently used to convey that a particular feature, structure, characteristic, etc. (described in connection with the embodiment / embodiment) is included in at least one embodiment / embodiment. Thus, the appearances of the phrases "in an exemplary embodiment" or "in exemplary embodiments" or "in an embodiment" or "in embodiments" and any other variations thereof that occur throughout this application do not necessarily refer to the same embodiment.

[0300] Similarly, references herein to "according to an exemplary embodiment / example / embodiment" or "in an exemplary embodiment / example / embodiment" and other variations thereof are frequently used to convey that a particular feature, structure, or characteristic (described in connection with the exemplary embodiment / example / embodiment) may be included in at least one exemplary embodiment / example / embodiment. Thus, the expressions "according to an exemplary embodiment / example / embodiment" or "in an exemplary embodiment / example / embodiment" that appear throughout the specification do not necessarily refer to the same exemplary embodiment / example / embodiment, nor do the individual or alternative exemplary embodiments / examples / embodiments have to be mutually exclusive of other exemplary embodiments / examples / embodiments.

[0301] The reference numerals that appear in the claims are for illustrative purposes only and do not limit the scope of the claims. Although not explicitly described, the embodiments / examples and variations thereof can be employed in any combination or sub - combination.

[0302] When a figure is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0303] Although some figures include arrows on communication paths to indicate the main direction of communication, it should be understood that communication can occur in a direction opposite to that of the depicted arrows.

[0304] Various embodiments relate to decoding. As used in this application, "decoding" can cover, for example, all or part of a process performed on a received point cloud frame (which may include a received bitstream that encodes one or more point cloud frames) to produce a final output suitable for display or further processing in a reconstructed point cloud domain. In various embodiments, such processes include one or more of the processes typically performed by a decoder. In various embodiments, for example, such processes also or optionally include processes performed by the decoders of the various embodiments described in this application.

[0305] As a further example, in one embodiment "decoding" may refer to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of entropy decoding and differential decoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process, and it is believed to be well understood by those skilled in the art.

[0306] Various embodiments relate to encoding. Similar to the above discussion regarding "decoding", "encoding" as used in this application may cover, for example, all or part of the process of performing on an input point cloud frame to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder. In various embodiments, such processes also include or optionally include the processes performed by the encoders of the various embodiments described in this application.

[0307] As a further example, in one embodiment "encoding" may refer only to entropy encoding, in another embodiment, "encoding" may refer only to differential encoding, and in another embodiment, "encoding" may refer to a combination of differential encoding and entropy encoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process, and it is believed to be well understood by those skilled in the art.

[0308] In addition, this application may refer to "determining" various information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from a memory.

[0309] Additionally, this application may refer to "accessing" various information. Accessing information may include, for example, one or more of the following: receiving information, retrieving information (e.g., from a memory or a bitstream), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0310] Furthermore, this application may refer to "receiving" various information. Like "accessing", receiving is a broad term. Receiving information may include one or more of the following: for example, accessing information or retrieving information (e.g., from a memory or a bitstream). Additionally, in one way or another, during operations such as, for example: storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is typically involved.

[0311] Moreover, as used herein, the term "signal" particularly refers to indicating something to a corresponding decoder etc. For example, in some embodiments, the encoder signals specific information, such as the number of points in a point cloud or sensor setting parameters (such as the basic azimuth step φ step or elevation angle θ k ). In this way, in embodiments, the same parameters can be used on the encoder side and the decoder side. Thus, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder such that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, then signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameters. By avoiding transmitting any actual functionality, bit savings are achieved in various embodiments. It should be recognized that signaling can be done in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the foregoing related to the verb form of the term "signal", the term "signal" can also be used as a noun herein.

[0312] Multiple embodiments have been described. However, it should be understood that various modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to yield other embodiments. Additionally, those of ordinary skill in the art will understand that other structures and processes can substitute the disclosed structures and processes, and the resulting embodiments will perform at least substantially the same (one or more) functions in at least substantially the same (one or more) ways to achieve at least substantially the same (one or more) results as the disclosed embodiments. Accordingly, this application contemplates these and other embodiments.

Claims

1. A method for processing point cloud data, the method being used to encode a point cloud into a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with spherical coordinates representing an azimuth angle responsive to a capture angle of a sensor of a spin sensor head that captures the point and a radius responsive to a distance of the point from a reference point, the method comprising: - Obtain (130) a scaled basic azimuth step (S( , )) associated with the points of the point cloud, the scaled basic azimuth step (S( , )) being greater than the basic azimuth step ( ) when second data (D2) is below a threshold, and the scaled basic azimuth step (S( , )) being equal to the basic azimuth step ( ) when the second data (D2) is greater than or equal to the threshold, the basic azimuth step ( ) being derived from the frequency and rotational speed at which the point cloud is captured by a spin sensor head, and the second data (D2) being the decoded radius ( r dec ) of the points obtained by encoding and decoding (120) the radii associated with the points; -Encode (140) into the bitstream the number (m) of scaled basic azimuth steps obtained from the azimuth of the point, the prediction of the azimuth, and the scaled basic azimuth step s ); and -Encode the residual azimuth angle of the point between the azimuth angle of the point and the predicted azimuth angle derived (150) from the number (m s ) of scaled basic azimuth steps and the scaled basic azimuth step (S( , )) into the bitstream.

2. A method for processing point cloud data, the method being used to decode a point cloud from a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with spherical coordinates representing an azimuth angle responsive to a capture angle of a sensor of a spin sensor head that captures the point and a radius responsive to a distance of the point from a reference point, the method comprising: - Determine a basic azimuth step from the bitstream ( ) - obtaining (220) a decoded radius of a point in the point cloud from a decoded residual radius decoded from the bitstream ( r res,dec ); r dec ​ - Obtain (130) a scaled basic azimuth step (S( , )) associated with a point of the point cloud, where the scaled basic azimuth step (S( r dec )) is greater than the basic azimuth step ( , ) when a second data (D2) which is the decoding radius ( ) of the point is below a threshold (TH), and the scaled basic azimuth step (S( , )) is equal to the basic azimuth step ( ) when the second data (D2) is greater than or equal to the threshold (TH); - Decode (230) the number (m) of scaled basic azimuth steps from the bitstream s ) - Decode (240) the decoded residual azimuth from the bitstream ( ); and - from the decoded residual azimuth ( ), and from the number (m s ) of scaled basic azimuth steps and the scaled basic azimuth step (S( , )) derive (150) a predicted azimuth and obtain (250) the decoded azimuth ( ).

3. The method according to claim 1 or 2, wherein when the second data (D2) is lower than a threshold (TH), the scaled basic azimuth step size (S( , )) is equal to the first data (D1), wherein the first data (D1) depends on the decoding radius ( r dec ), and wherein the first data (D1) is inversely proportional to the product of the decoding radius and a scaling factor ( ) that is greater than or equal to 1.

4. The method according to claim 3, wherein the second data (D2) is obtained by applying a monotonic function to the decoding radius ( r dec ), and the second data (D2) is compared with a second threshold (TH) obtained by applying the same monotonic function to the threshold.

5. The method according to claim 4, wherein the monotonic function is defined as a function that provides an integer bound for the residual azimuth angle.

6. The method according to claim 4, wherein the first data (D1) is obtained from an approximation of 2π / ( r dec *α* ), where r dec is the decoding radius, α is a scaling factor greater than or equal to 1, and corresponds to the internal precision of the azimuth angle.

7. The method according to claim 6, wherein the first data (D1) is obtained by iteratively refining the approximation.

8. The method according to claim 6, wherein the approximation is obtained by finding the highest power of factor 2 of a basic azimuth step size that is less than 2π / ( r dec *α* ).

9. A point cloud data processing apparatus, the apparatus being used to encode a point cloud into a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with spherical coordinates representing an azimuth angle responsive to a capture angle of a sensor of a spin sensor head that captures the point and a radius responsive to a distance of the point from a reference point, the apparatus comprising one or more processors configured to: - Obtain a scaled basic azimuth step associated with a point of the point cloud, the scaled basic azimuth step being greater than the basic azimuth step when second data is below a threshold, and the scaled basic azimuth step being equal to the basic azimuth step when the second data is greater than or equal to the threshold, the basic azimuth step being derived from a frequency and a rotational speed at which the spin sensor head captures the point cloud, and the second data being a decoded radius of the point obtained by encoding and decoding a radius associated with the point; - Encode a number of scaled basic azimuth steps obtained from the azimuth angle of the point, a prediction of the azimuth angle, and the scaled basic azimuth step into the bitstream; and - Encode a residual azimuth angle of the point between the azimuth angle of the point and a predicted azimuth angle derived from the number of scaled basic azimuth steps and the scaled basic azimuth step into the bitstream.

10. A point cloud data processing apparatus, the apparatus being used to decode a point cloud from a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with spherical coordinates representing an azimuth angle responsive to a capture angle of a sensor of a spin sensor head that captures the point and a radius responsive to a distance of the point from a reference point, the apparatus comprising one or more processors configured to: - Determine a basic azimuth step from the bitstream; - Obtain a decoded radius of a point in the point cloud from a decoded residual radius decoded from the bitstream; - Obtain a scaled basic azimuth step associated with a point of the point cloud, where when second data that is the decoding radius of the point is below a threshold, the scaled basic azimuth step is greater than the basic azimuth step, and when the second data is greater than or equal to the threshold, the scaled basic azimuth step is equal to the basic azimuth step; - Decode the number of scaled basic azimuth steps from the bitstream; - Decode a decoded residual azimuth angle from the bitstream; and - Obtain a decoded azimuth angle from the decoded residual azimuth angle and a predicted azimuth angle derived from the number of scaled basic azimuth steps and the scaled basic azimuth step.

11. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method of encoding a point cloud into a bitstream representing encoded point cloud data of a physical object, wherein each point of the point cloud is associated with spherical coordinates representing an azimuth angle that is the capture angle of a sensor of a spin sensor head that captured the point and a radius that is the distance of the point from a reference point, the method comprising: - Obtain a scaled basic azimuth step associated with a point of the point cloud, where when second data is below a threshold, the scaled basic azimuth step is greater than the basic azimuth step, and when the second data is greater than or equal to the threshold, the scaled basic azimuth step is equal to the basic azimuth step, the basic azimuth step being derived from the frequency and rotational speed at which a spin sensor head captured the point cloud, and the second data being the decoding radius of the point obtained by encoding and decoding a radius associated with the point; - Encode the number of scaled basic azimuth steps obtained from the azimuth angle of the point, the prediction of the azimuth angle, and the scaled basic azimuth step into the bitstream; and - Encode the residual azimuth angle of the point between the azimuth angle of the point and a predicted azimuth angle derived from the number of scaled basic azimuth steps and the scaled basic azimuth step into the bitstream.

12. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method of decoding a point cloud from a bitstream representing encoded point cloud data of a physical object, wherein each point of the point cloud is associated with spherical coordinates representing an azimuth angle that is the capture angle of a sensor of a spin sensor head that captured the point and a radius that is the distance of the point from a reference point, the method comprising: - Determine a basic azimuth step from the bitstream; - Obtain the decoding radius of a point in the point cloud from a decoded residual radius decoded from the bitstream; - Obtain a scaled basic azimuth step associated with a point of the point cloud, where when second data that is the decoding radius of the point is below a threshold, the scaled basic azimuth step is greater than the basic azimuth step, and when the second data is greater than or equal to the threshold, the scaled basic azimuth step is equal to the basic azimuth step; - Decode the number of scaled basic azimuth steps from the bitstream; - Decode a decoded residual azimuth angle from the bitstream; and - Obtain the decoded azimuth angle from the decoded residual azimuth angle and the predicted azimuth angle derived from the number of scaled basic azimuth steps and the scaled basic azimuth step.

13. A non-transitory storage medium carrying instructions of program code for performing a method of decoding a point cloud from a bitstream of encoded point cloud data representing a physical object, each point of the point cloud being associated with spherical coordinates of an azimuth angle representing a capture angle of a sensor of a spin sensor head that captured the point and a radius corresponding to a distance of the point from a reference point, the method being implemented when the instructions are executed by one or more processors as follows: - Determine a basic azimuth step from the bitstream; - Obtain the decoded radius of a point in the point cloud from the decoded residual radius decoded from the bitstream; - Obtain a scaled basic azimuth step associated with a point of the point cloud, the scaled basic azimuth step being greater than the basic azimuth step when second data that is the decoded radius of the point is below a threshold, and the scaled basic azimuth step being equal to the basic azimuth step when the second data is greater than or equal to the threshold; - Decode the number of scaled basic azimuth steps from the bitstream; - Decode a decoded residual azimuth angle from the bitstream; and - Obtain the decoded azimuth angle from the decoded residual azimuth angle and the predicted azimuth angle derived from the number of scaled basic azimuth steps and the scaled basic azimuth step.

Citation Information

Patent Citations

  • Method and apparatus of quantizing spherical coorinates used for encoding / decoding point cloud geometry data

    EP4020397A1

  • Apparatus, method, and system for alignment of 3D datasets

    CN110574071A

  • Video encoding or decoding method and device, computer equipment and storage medium

    CN112616058A