Method and apparatus for encoding / decoding point cloud geometric data

By sorting and entropy coding based on two-dimensional spatial dictionary order of point cloud geometric data, the problem of difficult to combine encoding and decoding complexity, delay and compression performance in the prior art is solved, and efficient compression and low-latency coding of point cloud data are achieved.

CN117981321BActive Publication Date: 2025-07-04BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280060375.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-30
Filing Date
2022-06-21
Publication Date
2025-07-04
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

Existing point cloud codecs are difficult to achieve simplicity, low latency and high compression performance in the encoding and decoding process, especially in point cloud compression of sparse geometric data, especially when using spin lidar and flexible sensor heads, which cannot be effectively combined.

Method used

When encoding point cloud geometric data, the ordered coarse points are sorted according to the dictionary order of the two-dimensional space, the order index difference encoding and decoding are used to identify the occupied coarse points, and the entropy coding technology is used to optimize the encoding process to maintain low latency and efficient compression.

Benefits of technology

Efficient compression of point cloud data at low latency is achieved, ensuring that encoding performance does not significantly reduce under ideal use, while simplifying the design of encoder and decoder.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117981321B_ABST
    Figure CN117981321B_ABST
Patent Text Reader

Abstract

Methods and apparatus are provided for encoding / decoding point cloud geometry data represented by an ordered coarse point representation of some of the discrete positions in a set of discrete positions occupying a two-dimensional space, the ordered coarse points being sorted according to a lexicographical order based on the coordinates of the two-dimensional space. The method includes encoding (110) data (S next ) indicating whether an occupied coarse point (P next ) associated with a point of the point cloud is a later occupied coarse point into a bitstream, where the occupied coarse point is considered a later occupied coarse point (P ref ) when the order of the occupied coarse point in the lexicographical order is lower than the order of a reference coarse point (P next ) associated with a previously encoded point of the point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority and benefit of European Patent Application Serial No. 21306360.5, filed on September 30, 2021, the entire content of which is incorporated herein by reference for all purposes. Technical field

[0003] This application generally relates to point cloud compression and, in particular, to methods and apparatuses for encoding / decoding point cloud geometry data sensed by at least one sensor. Background art

[0004] This section is intended to introduce the reader to aspects of the art that may be related to aspects of at least one embodiment of the present application described and / or claimed below. This discussion is considered to be helpful in providing background information to the reader to facilitate a better understanding of the aspects of the present application.

[0005] As a format for representing 3D data, point clouds have recently gained attention because of their multiple capabilities in representing all types of physical objects or scenes. Point clouds can be used for various purposes, such as cultural heritage / buildings, where objects such as statues or buildings are scanned in 3D in order to share the spatial configuration of the objects without sending or accessing the objects. Moreover, it is a way to ensure the preservation of knowledge of objects in case the objects may be damaged; for example, a temple damaged by an earthquake. Such point clouds are typically static, colored, and huge.

[0006] Another use case is in topology and cartography, where the use of 3D representation allows maps to be not limited to a plane and can include landforms. Google Maps is now a good example of a 3D map, but it uses meshes instead of point clouds. However, point clouds can be a suitable data format for 3D maps, and such point clouds are typically static, colored, and huge.

[0007] Virtual reality (VR), augmented reality (AR), and immersive worlds have recently become a hot topic and are foreseen by many as the future of 2D flat video. The basic idea is to immerse the viewer in the surrounding environment, while a standard TV only allows the viewer to watch the virtual world in front of his / her eyes. Depending on the degree of freedom of the viewer in the environment, there are several levels of immersion. Point clouds are good format candidates for distributing VR / AR worlds.

[0008] The automotive industry, especially foreseeable autonomous vehicles, is also an area where point clouds can be used extensively. Autonomous vehicles should be able to "detect" their environment in order to make good driving decisions based on the presence and nature of their closest nearby objects and the road configuration.

[0009] A point cloud is a set of points located in three-dimensional (3D) space, optionally with additional value attached to each point. These additional values are commonly referred to as attributes. Attributes can be, for example, three-component color, material properties (such as reflectivity), and / or two-component normal vectors of the surface associated with the points.

[0010] Thus, a point cloud is a combination of geometric data (the positions of points in 3D space, typically represented by 3D Cartesian coordinates x, y, and z) and attributes.

[0011] Point clouds can be sensed by various types of devices, such as arrays of cameras, depth sensors, lasers (light detection and ranging, also known as lidar), radar, or can be generated by a computer (e.g., in post-production of a movie). Depending on the use case, point clouds can have thousands to billions of points for mapping applications. The original representation of a point cloud requires a very large number of bits per point, at least a dozen bits for each Cartesian coordinate x, y, or z, and optionally more bits for the (one or more) attributes, such as triple 10 bits for color.

[0012] In many applications, it is very important to be able to distribute point clouds to end users or store them on a server while consuming only a reasonable amount of bitrate or storage space and maintaining an acceptable (or preferably very good) quality of experience. The efficient compression of these point clouds is a key point to make the distribution chain of many immersive worlds practical.

[0013] For distribution and visualization by end users, such as on AR / VR glasses or any other 3D-enabled device, the compression can be lossy (as in video compression). Other use cases do require lossless compression, such as medical applications or autonomous driving, to avoid changing the results of decisions obtained from subsequent analysis of the compressed and transmitted point clouds.

[0014] Until recently, the mass market has not addressed the issue of point cloud compression (aka PCC), nor are there available standardized point cloud codecs. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11, also known as the Moving Picture Experts Group or MPEG, initiated a work item on point cloud compression. This has led to two standards, namely

[0015] · Part 5 of MPEG-I (ISO / IEC 23090-5) or Video-based Point Cloud Compression (aka V-PCC)

[0016] · Part 9 of MPEG-I (ISO / IEC 23090-9) or Geometry-based Point Cloud Compression (aka G-PCC)

[0017] The V-PCC encoding method compresses the point cloud by performing multiple projections on the 3D object to obtain 2D tiles that are packed into an image (or video when processing dynamic point clouds). Then, an existing image / video codec is used to compress the obtained image or video, thus allowing for the full utilization of the already deployed image and video solutions. By its nature, V-PCC is only efficient on dense and continuous point clouds because image / video codecs cannot compress non-smooth tiles, such as those obtained from the projection of sparse geometric data sensed by lidar.

[0018] The G-PCC encoding method has two schemes for compressing sensed sparse geometric data.

[0019] The first scheme is based on an occupancy tree, which locally can be any type of tree among octree, quadtree, or binary tree, representing the point cloud geometry. Occupied nodes (i.e., nodes associated with a cube / cuboid that includes at least one point of the point cloud) are split until a certain size is reached, and the occupied leaf nodes provide the 3D positions of the points, typically at the centers of these nodes. The occupancy information is carried by occupancy data (binary data, flags), and the occupancy flags signal the occupancy status of each child node of the node. By using neighbor-based prediction techniques, a high level of compression of the occupancy data for dense point clouds can be achieved. Sparse point clouds can also be addressed by directly encoding the positions of points within nodes of non-minimal size, stopping the tree construction when only isolated points are present in the node; this technique is called the direct coding mode (DCM).

[0020] The second scheme is based on a prediction tree, where each node represents the 3D position of a point, and the parent / child relationship between nodes represents a spatial prediction from the parent to the child. This method can only address sparse point clouds and offers the advantages of lower latency and simpler decoding compared to the occupancy tree. However, compared to the first occupancy-based method, the compression performance is only slightly better, and the encoding is also complex because the encoder has to search intensively (among a long list of potential predictors) for the best predictor when constructing the prediction tree.

[0021] In both schemes, the attribute (decoding) encoding is performed after the geometric (decoding) encoding is completed, effectively resulting in two encodings. Therefore, joint geometric / attribute low latency is obtained by using slices that decompose the 3D space into independently encoded sub-volumes without prediction between sub-volumes. When many slices are used, this severely affects the compression performance.

[0022] Combining the requirements for encoder and decoder simplicity, low latency, and compression performance together remains an issue that existing point cloud codecs have not satisfactorily addressed.

[0023] An important use case is the transmission of sparse geometric data sensed by at least one sensor mounted on a moving vehicle. This typically requires a simple and low-latency embedded encoder. Simplicity is required because the encoder may be deployed on a computing unit that is performing other processing in parallel, such as (semi-)autonomous driving, thus limiting the processing power available for the point cloud encoder. Low latency is also required to allow for fast transmission from the vehicle to the cloud, in order to view local traffic in real-time based on multi-vehicle acquisitions and make decisions quickly enough based on traffic information. Although the transmission latency can be made low enough by using 5G, the encoder itself should not introduce too much latency due to encoding. Moreover, compression performance is extremely important because the data stream from millions of vehicles to the cloud is expected to be very large.

[0024] Specific priors related to the sparse geometric data sensed by spin lidar have been exploited in G-PCC and have led to very significant compression gains.

[0025] First, G-PCC exploits the sensed elevation angle (with respect to the horizontal ground) from the spin lidar head 10, as Figure 1 and Figure 2 depicted. The lidar head 10 includes a set of sensors 11 (e.g., lasers), five sensors are represented here. The spin lidar head 10 can rotate about the vertical axis z to sense the geometric data of physical objects. Then, the geometric data sensed by the lidar is represented in spherical coordinates (r 3D , φ, θ), where r 3D is the distance of the point P from the center of the lidar head, φ is the azimuth angle of the lidar head spinning relative to a reference, and θ is the elevation angle of the sensor k of the spin lidar head 10 relative to the horizontal reference plane.

[0026] A regular distribution along the azimuth angle has been observed in the data sensed by the lidar, as Figure 3 depicted. This regularity is used in G-PCC to obtain a quasi-1D representation of the point cloud, where, up to noise, only the radius r 3D belongs to a continuous range of values, while the angles φ and θ only take a discrete number of values, from 0 to I - 1, where I is the number of azimuth angles used for sensing points, and from 0 to N sensor - 1, where N sensor is the number of sensors of the spin lidar head 10. Basically, G-PCC represents the sparse geometric data sensed by the lidar in a two-dimensional (discrete) angular coordinate space (φ, θ), as Figure 3 depicted, and the radius value r 3D for each point.

[0027] This quasi-1D property has been exploited in both the occupancy tree and the prediction tree in G-PCC by predicting the position of the current point based on the already encoded points using the discrete nature of the angles in the spherical coordinate space.

[0028] More precisely, the occupancy tree makes extensive use of DCM and entropy-encodes the direct positions of the points within the nodes by using a context-adaptive entropy coder. Then, the local transformation from the point position to the angular coordinates (φ, θ) and the context is obtained from the positions of these angular coordinates relative to the discrete angular coordinates (φ i , θ j ) obtained from the previously encoded points. Using the quasi-1D nature of this angular coordinate space (r 2D , φ i , θ j ), the prediction tree directly encodes the first version of the point position in the angular coordinates (r 2D , φ, θ), where r 2D is the projected radius on the horizontal xy plane, as depicted in Figure 4 . Then, the spherical coordinates (r 2D , φ, θ) are converted to 3D Cartesian coordinates (x, y, z), and the xyz residuals are encoded to address the errors in the coordinate transformation, the approximation of the elevation and azimuth angles, and the potential noise.

[0029] G-PCC does use angular priors to better compress the sparse geometric data sensed by lidar, but does not adapt the encoding structure to the sensing order. By its nature, the occupancy tree must be encoded to its final depth before outputting the points. This occupancy data is encoded in the so-called breadth-first order: first, the occupancy data of the root node is encoded, indicating its occupied child nodes; then, the occupancy data of each occupied child node is encoded, indicating the occupied grandchild nodes; and so on, iterating over the tree depth until the leaf nodes can be determined and the corresponding points are provided / output to the application or the (one or more) attribute encoding schemes. Regarding the prediction tree, the encoder is free to choose the order of the points in the tree, but for good compression performance and to optimize the prediction accuracy, G-PCC recommends encoding one tree per sensor. This mainly has the same drawback as using one encoding slice per sensor, namely, non-optimal compression performance, because prediction between sensors is not allowed and low latency is not provided to the encoder. Even worse, each sensor should have an encoding process, and the number of core encoding units should be equal to the number of sensors; this is impractical.

[0030] In short, in the framework of the lidar sensor head for sparse geometric data of sensed point clouds, the prior art does not address the problem of combining the simplicity of encoding and decoding, low latency, and compression performance.

[0031] In addition, sensing sparse geometric data of a point cloud using a spin sensor head has some drawbacks, and other types of sensor heads can be used.

[0032] The mechanical parts that generate the spin (rotation) of the spin sensor head are prone to damage and are costly. Additionally, by construction, the viewing angle must be 2π. This does not allow for sensing a specific region of interest at a high frequency. For example, sensing in front of a vehicle may be more interesting than sensing behind it. In fact, in most cases, when the sensor is attached to a vehicle, most of the 2π viewing angle is blocked by the vehicle itself, and the blocked viewing angle does not need to be sensed.

[0033] Recently, new types of sensors have emerged that allow for a more flexible selection of the regions to be sensed. In most recent designs, the sensors can move more freely and electronically (thus avoiding fragile mechanical parts) to obtain multiple sensing paths in a 3D scene, as Figure 5 depicted in Figure 5 a set of four sensors is shown. Their relative sensing directions (i.e., azimuth and elevation angles) are fixed relative to each other, but they generally sense the scene along a programmable sensing path depicted by a dashed line in a two-dimensional angular coordinate (φ,θ) space. Then, the points of the point cloud can be regularly sensed along the sensing path. As Figure 6 illustrated in Figure 7 when an area of interest R is detected, some sensor heads can also adjust their sensing frequency by increasing their sensing frequency. Such an area of interest R can be associated with, for example, a nearby object, a moving object, or any object (pedestrian, other vehicle, etc.) that was previously segmented in the previous frame or is dynamically segmented during sensing.

[0034] As Figure 8 depicted in Figure 8 a sensor head including a single sensor can also be used to sense multiple positions ( Figure 8A single sensor with different elevation angles is used to mimic the sensing using a collection of multiple sensors. For simplicity, in the following description and claims, "sensor head" may refer to a collection of physical sensors or a collection of sensing elevation angle indices that mimic a collection of sensors. Additionally, those skilled in the art will understand that "sensor" may also refer to a sensor at each sensing elevation angle index position.

[0035] Combining the simplicity, low latency of the encoder and decoder, and the requirements for the compression performance of the point cloud sensed by any type of sensor remains a problem that existing point cloud codecs have not satisfactorily solved.

[0036] At least one embodiment of the present application has been designed in view of the foregoing. Summary of the Invention

[0037] The following section presents a simplified summary of at least one embodiment to provide a basic understanding of some aspects of the present application. This summary is not an exhaustive overview of the embodiment. It is not intended to identify key or important elements of the embodiment. The following summary only presents some aspects of at least one of the embodiments in a simplified form as a prelude to the more detailed description provided elsewhere in the document.

[0038] According to a first aspect of the present application, there is provided a method of encoding point cloud geometry data represented by ordered coarse points of some discrete positions in a discrete position set occupying a two-dimensional space into a bitstream, the ordered coarse points being sorted according to a lexicographical order based on coordinates in the two-dimensional space, wherein the method includes encoding data indicating whether an occupied coarse point associated with a point of the point cloud is a later occupied coarse point into the bitstream, and when the order index of the occupied coarse point in the lexicographical order is lower than the order index of a first reference coarse point associated with a previously encoded point of the point cloud, the occupied coarse point is considered a later occupied coarse point; if the data indicates that the occupied coarse point is a later occupied coarse point, obtaining a post-point order index difference between the order index of the later occupied coarse point and the order index of a second reference coarse point; and encoding the magnitude of the post-point order index difference into the bitstream.

[0039] According to a second aspect of the present application, there is provided a method for decoding point cloud geometric data represented by an ordered set of coarse points that occupy some discrete positions in a discrete position set in a two-dimensional space from a bitstream, where the ordered coarse points are sorted according to the lexicographical order based on the coordinates of the two-dimensional space. The method includes decoding from the bitstream data indicating whether an occupied coarse point associated with a point of the point cloud is a later occupied coarse point. When the order of the occupied coarse point in the lexicographical order is lower than the order of a first reference coarse point associated with a previously decoded point of the point cloud, the occupied coarse point is considered a later occupied coarse point; and if the data indicates that the occupied coarse point is a later occupied coarse point, decoding from the bitstream the magnitude of the post-point order index difference between the order index of the later occupied coarse point and the order index of a second reference coarse point.

[0040] In some embodiments, the first reference coarse point may be the last encoded or decoded occupied coarse point or the previously encoded or decoded occupied coarse point having the highest order in the lexicographical order, and the second reference coarse point is the first reference coarse point or a coarse point associated with an order equal to the order of the first reference coarse point shifted by an offset.

[0041] In some embodiments, the data may be binary data indicating whether an occupied coarse point associated with a point of the point cloud is a later occupied coarse point, and the data may be entropy encoded or decoded based on at least one other binary data indicating whether at least one previously encoded or decoded occupied coarse point is a later occupied coarse point.

[0042] In some embodiments, the magnitude of the post-point order index difference may be encoded by encoding two positive offsets associated with the coordinates of the two-dimensional space into the bitstream, or the magnitude of the post-point order index difference may be decoded by decoding two positive offsets associated with the coordinates of the two-dimensional space from the bitstream.

[0043] In some embodiments, encoding or decoding one of the positive offsets may include encoding or decoding a first positive value and a second positive value whose sum is equal to the positive offset.

[0044] In some embodiments, encoding or decoding the first positive value or the second positive value may include entropy encoding or decoding a sequence of binary data representing the first positive value or the second positive value.

[0045] In some embodiments, encoding or decoding the other positive offset may include entropy encoding or decoding a sequence of binary data representing the other positive offset.

[0046] According to a third aspect of the present application, there is provided an apparatus for encoding point cloud geometry data represented by an ordered set of rough points of some discrete positions in a discrete position set occupying a two-dimensional space into a bitstream. The apparatus includes one or more processors configured to execute the method according to the first aspect of the present application.

[0047] According to a fourth aspect of the present application, there is provided an apparatus for decoding point cloud geometry data represented by an ordered set of rough points of some discrete positions in a discrete position set occupying a two-dimensional space from a bitstream. The apparatus includes one or more processors configured to execute the method according to the second aspect of the present application.

[0048] According to a fifth aspect of the present application, there is provided a bitstream representing encoded point cloud data of a point cloud geometry represented by an ordered set of rough points of some discrete positions in a discrete position set occupying a two-dimensional space, the ordered rough points being sorted according to a lexicographical order based on coordinates in the two-dimensional space, wherein the bitstream further includes data indicating whether an occupied rough point associated with a point of the point cloud is a later occupied rough point, and when the order of the occupied rough point in the lexicographical order is lower than the order of a reference rough point associated with a previously encoded point of the point cloud, the occupied rough point is considered a later occupied rough point.

[0049] According to a sixth aspect of the present application, there is provided a computer program product including instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to the first aspect of the present application.

[0050] According to a seventh aspect of the present application, there is provided a non-transitory storage medium carrying instructions for program code for executing the method according to the first aspect of the present application.

[0051] According to an eighth aspect of the present application, there is provided a computer program product including instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to the second aspect of the present application.

[0052] According to a ninth aspect of the present application, there is provided a non-transitory storage medium carrying instructions for program code for executing the method according to the second aspect of the present application.

[0053] The specific nature of at least one of the embodiments and other objects, advantages, features, and uses of at least one of the embodiments will become apparent from the following description of the examples in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Embodiments of the present application will now be described by way of example with reference to the drawings, in which:

[0055] Figure 1 A side view schematically showing a sensor head according to the prior art and some of its parameters;

[0056] Figure 2 A top view schematically showing a sensor head according to the prior art and some of its parameters;

[0057] Figure 3 A schematic illustration of a regular distribution of data sensed by a spin sensor head according to the prior art;

[0058] Figure 4 A schematic illustration of the representation of points of a point cloud in 3D space according to the prior art;

[0059] Figure 5 A schematic illustration of an example of a sensor head capable of sensing a real scene along a programmable sensing path according to the prior art;

[0060] Figure 6 A schematic illustration of an example of a sensor head capable of sensing a real scene along a programmable sensing path according to different sensing frequencies according to the prior art;

[0061] Figure 7 A schematic illustration of an example of a sensor head capable of sensing a real scene along a programmable zigzag sensing path according to different sensing frequencies according to the prior art;

[0062] Figure 8 A schematic illustration of a single sensor head capable of sensing a real scene along a programmable zigzag sensing path according to different sensing frequencies;

[0063] Figure 9 A schematic illustration of an example of ordered coarse points in a rough representation according to at least one embodiment.

[0064] Figure 10 A schematic illustration of an example of ordered coarse points in a rough representation sensed by a dual sensor head according to at least one embodiment;

[0065] Figure 11 A schematic illustration of the representation of ordered coarse points in a two-dimensional coordinate (s, λ) space;

[0066] Figure 12 A schematic illustration of ordered coarse points in a rough representation according to at least one embodiment.

[0067] Figure 13 A schematic block diagram showing the steps of method 100 for encoding point cloud geometric data into a bitstream of encoded point cloud data according to at least one embodiment;

[0068] Figure 14 A schematic block diagram showing the steps of a method 200 for decoding point cloud geometry data from a bitstream of encoded point cloud data according to at least one embodiment;

[0069] Figure 15 A schematic block diagram showing the steps of a context - adaptive binary arithmetic encoder according to at least one embodiment;

[0070] Figure 16 A schematic block diagram showing step 140 of method 100 and step 230 of method 200 according to at least one embodiment;

[0071] Figure 17 A schematic block diagram showing step 140 of method 100 and step 230 of method 200 according to at least one embodiment;

[0072] Figure 18 An example of determining a first positive offset and a positive value when the next occupied coarse point and the reference occupied coarse point have different sample indices according to at least one embodiment;

[0073] Figure 19 An example of determining a first positive offset and a positive value when the next occupied coarse point and the reference occupied coarse point have the same sample index according to at least one embodiment;

[0074] Figure 20 An example of determining a first positive offset and a positive value when the next occupied coarse point and the reference occupied coarse point have different sample indices according to at least one embodiment; and

[0075] Figure 21 A schematic block diagram illustrating an example of a system in which various aspects and embodiments are implemented.

[0076] Similar reference numerals may have been used in different figures to denote similar components. Detailed Description

[0077] Embodiments of at least one will be described more fully hereinafter with reference to the accompanying drawings, in which examples of embodiments of at least one are illustrated. However, the embodiments may be implemented in many alternative forms and should not be construed as limited to the examples set forth herein. Thus, it should be understood that there is no intention to limit the embodiments to the particular forms disclosed. On the contrary, this disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of this application.

[0078] At least one aspect generally relates to point cloud encoding and decoding, another aspect generally relates to transmitting a generated or encoded bitstream, and another aspect relates to receiving / accessing a decoded bitstream.

[0079] Furthermore, the present aspects are not limited to MPEG standards such as MPEG-I Part 5 or Part 9 related to point cloud compression, and may be applied, for example, to other standards and recommendations, whether pre-existing or developed in the future, as well as extensions of any such standards and recommendations, including MPEG-I Part 5 and Part 9. Unless otherwise indicated, or technically excluded, the aspects described in this application may be used alone or in combination.

[0080] The present invention relates to encoding / decoding point cloud geometric data represented by ordered coarse points that are coarse representations of some discrete positions in a set of discrete positions occupying a two-dimensional space.

[0081] For example, in the working group ISO / IEC JTC 1 / SC 29 / WG 7 on MPEG 3D Graphics Coding, a new codec called L3C2 (Low Latency Low Complexity Codec) is considered to improve the coding efficiency of lidar-sensed point clouds relative to the G-PCC codec. Codec L3C2 provides an example of a two-dimensional representation (i.e. a coarse representation) of the points of a point cloud. A description of the code can be found in the working group's output document N00167, ISO / IEC JTC 1 / SC29 / WG 7, MPEG 3D Graphics Coding, "Technologies under Consideration in G-PCC", dated August 31, 2021.

[0082] Basically, for each sensed point P of the point cloud n , the sensed point P is represented by the conversion n The 3D Cartesian coordinates (x n ,y n ,z n ) to obtain the corresponding sensing point P n The sensor index λ associated with the sensor n and the azimuth angle φ representing the sensing angle of the sensor n Then, based on the azimuth angle φ n and sensor index λ n The points of the point cloud are sorted, for example, according to a lexicographic order based first on the azimuth angle and then on the sensor index. Then, point P n The order index o(P n ) is obtained by the following formula:

[0083] o(P n )=φ n *K+λ n (1)

[0084] Where K is the number of sensors.

[0085] Figure 9 Schematically shows roughly represented ordered rough points. Five points of the point cloud have been sensed. Each of these five points is roughly represented by a rough point (black dot) in the rough representation: two rough points P n and P n+1 represent two points in the point cloud sensed at time t1 at an angle φ c (in φ i 's), and three rough points represent three points of the point cloud sensed at time t2 at an angle φ c +Δφ. The rough points representing the sensed points of the point cloud are the occupied rough points, and the rough points not representing the sensed points of the point cloud are the unoccupied rough points. Since the points of the point cloud are represented by the occupied rough points in the rough representation, the order index associated with the points of the point cloud is also the order index associated with the occupied rough points.

[0086] Then, a rough representation of the point cloud geometric data can be defined in the two-dimensional coordinate (φ,λ) space.

[0087] A rough representation can also be defined for any type of sensor head, which includes rotating (spinning) or non-rotating sensor heads. The definition is based on a sensing path defined according to the characteristics of the sensor in the two-dimensional angular coordinate (φ,θ) space, which includes the azimuth angle coordinate φ representing the azimuth angle and the elevation angle coordinate θ representing the elevation angle of the sensor relative to a horizontal reference plane. The azimuth angle represents the sensing angle of the sensor relative to a reference object. The sensing path is used to sense the points of the point cloud according to the ordered rough points representing the potential positions of the sensed points of the point cloud. Each rough point is defined according to a sample index s associated with the sensing moment along the sensing path and a sensor index λ associated with the sensor.

[0088] In Figure 10 , a sensor head including two sensors is used. The sensing paths along which the two sensors are located are represented by dashed lines. For each sample index s (each sensing moment), two rough points are defined. The rough point associated with the first sensor is represented by the black shaded dot on Figure 10 , and the rough point associated with the second sensor is represented by the black hash dot. Each of these two rough points belongs to the sensor sensing path (dashed line) defined by the sensing path SP. Figure 11 Schematically shows the representation of the ordered rough points in the two-dimensional coordinate (s,λ) space. Figure 10 and Figure 11 The arrows on illustrate the link between two consecutive ordered rough points.

[0089] Associate an order index o(P) with each coarse point according to the sorting of the coarse point in the ordered coarse points:

[0090] o(P) = λ + s * K (2)

[0091] where K is the number of sensors in the sensor set or the number of different positions of a single sensor for the same sample index, and λ is the sensor index of the sensor that senses the point P in the point cloud at the sensing moment s.

[0092] Figure 12 The figure shows the ordered coarse points in a coarse representation, showing five occupied coarse points (black circles): two coarse points P n and P n+1 are occupied by two points of the point cloud sensed at the sensing moment t1 (corresponding to the sample index s1), and three coarse points are occupied by three points of the point cloud sensed at the sensing moment t2 (corresponding to the sample index s2).

[0093] Then, a coarse representation of the point cloud geometric data can be defined in the two-dimensional coordinate (s, λ) space.

[0094] Independent of the two-dimensional space in which the coarse representation of the point cloud geometric data is defined, encoding the point cloud geometric data includes encoding the occupancy data of the occupied coarse points in the coarse representation by encoding the order index difference Δo. Each order index difference Δo represents the difference between the order indices of two consecutive occupied coarse points P -1 and P:

[0095] Δo = o(P) - o(P -1 ) (3)

[0096] In Figure 9 , assuming that the coordinates of the first occupied coarse point P n in the two-dimensional coordinate (φ, λ) space are known in advance, the first order index difference Δo n+1 is obtained as the difference between the order index o(P n+1 ) associated with the occupied coarse point P n+1 and the order index o(P n ) associated with the occupied coarse point P n . The second order index difference Δo n+2 is obtained as the difference between the order index o(P n+2 ) associated with another occupied coarse point P n+2 and the order index o(P n+1 ) associated with P n+1 ), and so on.

[0097] In Figure 12Among them, assume that the coordinates of the first occupied coarse point P in the two-dimensional coordinate (s, λ) space n are known in advance, then the first-order index difference Δo n+1 is obtained as the difference between the order index o(P n+1 ) of the occupied coarse point P n+1 and the order index o(P n ) of the occupied coarse point P n . In the example, Δo n+1 = 2 because of the unoccupied coarse points (white circles). The second-order index difference Δo n+2 is obtained as the difference between the order index o(P h+2 ) of another occupied coarse point P n+2 and the order index o(P n+1 ) of the occupied coarse point P n+1 , and so on.

[0098] The order index o(P1) of the first coarse point occupied by the first sensed point P1 of the point cloud can be directly encoded into the bitstream B. This is equivalent to arbitrarily setting the order index of the virtual zero-th point to zero, i.e., o(P0) = 0, and encoding Δo1 = o(P1) - o(P0) = o(P1).

[0099] Given the order index o(P1) and the order difference Δo of the first coarse point occupied by the first sensed point P1 of the point cloud, the order index o(P) of any occupied coarse point occupied by the sensed point P of the point cloud can be recursively reconstructed by the following formula:

[0100] o(P) = o(P - 1) + Δo

[0101] Hereinafter, the present invention is mainly described by considering the coarse representation defined in the two-dimensional coordinate (s, λ) space. However, the same description can also be made for the coarse representation defined in the two-dimensional coordinate (φ, λ) space, because the spin sensor head (such as a lidar head) is a specific coarse representation defined in the two-dimensional coordinate (s, λ) space, where at each sensing moment, the sensor of the sensor head detects an object, and the sensed point corresponds to the occupied coarse point of the representation.

[0102] In fact, the order index difference Δo is usually positive because the sensing order of the points of the point cloud is equal to the lexicographical order of the associated occupied coarse points in the coarse representation.

[0103] However, it may happen that some points of the point cloud sensed according to the sensing order do not exactly follow the lexicographical order of the occupied coarse points associated with these sensed points. These points are called post-occupied coarse points.

[0104] For example, when defining a rough representation in a two-dimensional coordinate (φ, λ) space, the azimuth angle φ of the occupied rough points within the rough representation defined by the two-dimensional coordinates (φ, λ) is obtained c , where φ is the azimuth angle associated with the sensed points of the point cloud associated with the occupied rough points, and Δφ is the azimuth shift. Due to the rounding of the ratio of the actual azimuth angle φ to the azimuth shift Δφ, if the value of φ includes additive noise equal to or less than -Δφ / 2, the azimuth angle φ c may be 1 (or more) lower than the expected value in the rough representation according to the lexicographical order; or if the additive noise is equal to or greater than Δφ / 2, it may be 1 (or more) higher than that expected value.

[0105] In the rough representation defined in the two-dimensional coordinate (s, λ) space, due to the noise on the sensed points, the sequence of the sensed points may have sample index coordinates s that are not in an increasing sequence, so the sensed points of the point cloud according to the sensing order may not exactly follow the lexicographical order of the occupied rough points associated with these sensed points.

[0106] The lexicographical order used to sort the rough points within the rough representation may also not be applicable to an actual system in which a sensor head sends data to one or more operating components (such as an encoder and / or an embedded data analyzer) via a transmission bus or channel that is prone to transmission errors (e.g., due to transmission collisions or electromagnetic disturbances). In such a system, when a transmission error is detected, the sensor data sent through the sensor head is usually retransmitted. However, in the meantime, other sensor data may have been transmitted before the previous said sensor data is retransmitted.

[0107] For these two actual use cases, the order of the sensed points of the point cloud received as the input of the encoder may not always be the expected lexicographical order.

[0108] To address the problem of the later occupied rough points, the sensed points received by the encoder can be buffered. Buffering the sensed points allows these points to be reordered to fix the difference between the sensing order (or receiving order) and the lexicographical order. Buffering the sensed points is not always an acceptable solution because the buffer size may need to be made larger (beyond the actual implementation) to support the correction of the maximum possible difference between the sensing order and the lexicographical order, especially when the interference level (noise and / or retransmission latency) is high.

[0109] Additionally, if the buffered sensed points are sent to the encoder prematurely and new points that should have been pre-encoded arrive later, there is no real solution: the sensed points must be discarded (i.e., the later occupied rough points are lost) because they can no longer be encoded.

[0110] It is completely impossible to pre - estimate the maximum buffer size because it depends on the maximum difference between the sensing order and the lexicographical order of the system used. This is especially true for systems prone to transmission errors, where in the worst - case scenario, some sensed points may need to be re - transmitted multiple times to succeed. Thus, if some sensed points arrive late or exceed the buffer size, these sensed points may be lost. To overcome this problem, the maximum buffer size may be overestimated, but this results in a waste of resources.

[0111] In addition, one of the requirements is to design an encoding / decoding of point - cloud geometric data with extremely low latency. This is not compatible with solutions based on re - ordering sensed points buffered and embedded because then an unacceptable delay is introduced before encoding.

[0112] One of the problems to be solved is to process the encoding / decoding of post - occupied coarse points associated with the sensed points of a point cloud relative to the lexicographical order of the occupied coarse points defined in a coarse representation, while maintaining low - latency encoding of the points of the point cloud and not significantly degrading the encoding performance obtained in the ideal usage scenario when all sensed points are correctly ordered.

[0113] The position of the occupied coarse points associated with the sensed points is represented as discrete positions of a grid in a two - dimensional space. The grid includes rows and columns, and each occupied coarse point belongs to the intersection of a column and a row of the grid. Each sensor index λ occupies a row, and each sample index s occupies a column. In the case where s is associated with an angle, a direct solution to the above problem would consist of adding many empty columns from the current occupied coarse point to recover the post - occupied coarse points (i.e., to recover the angle of almost a full rotation until the angle associated with the post - occupied coarse point) by inserting "false" empty circles. It would work in the sense that the points and the post - occupied coarse points could be encoded / decoded. However, the cost of encoding / decoding the empty columns would endanger the compression efficiency of the encoding / decoding.

[0114] In short, the present invention provides a method for encoding point - cloud geometric data represented by ordered coarse points occupying some discrete positions in a set of discrete positions in a two - dimensional space into a bit - stream. The ordered coarse points are sorted according to the lexicographical order based on the coordinates of the two - dimensional space.

[0115] The point - cloud geometry is encoded into a bit - stream by encoding the occupied coarse points representing the points of the point cloud in the coarse representation defined by the two - dimensional space. By considering a first reference coarse point P associated with a previously encoded point of the point cloud ref , each next occupied coarse point P associated with a point of the point cloud nextare encoded one by one. The order o(P next ) and o(P ref ) are determined by equation (1 or 2), i.e., according to the lexicographical order.

[0116] The order index difference Δo is obtained by:

[0117] Δo = o(P next ) - o(P ref ). (4)

[0118] When the order o(P next ) is greater than the order o(P ref ), then the next occupied coarse point P next is considered not to be the later occupied coarse point, and the order index difference Δo is positive.

[0119] When the order o(P next ) is lower than the order o(P ref ), then the next occupied coarse point P next is considered to be the later occupied coarse point, and the order index difference Δo is negative.

[0120] Data S next indicating whether the next occupied coarse point P next is the later occupied coarse point is encoded into the bitstream.

[0121] Encoding the data S next indicating whether the next occupied coarse point P next is the later occupied coarse point is equivalent to encoding the sign of the order index difference Δo.

[0122] If the data S next indicates that the sign of the order index difference Δo is positive, then the order index difference Δo is encoded into the bitstream.

[0123] If the data S next indicates that the occupied coarse point is the later occupied coarse point P next , then the post-point order index difference Δo next between the order index o(P ref ) and the order index of the second reference coarse point P' late is obtained, and the magnitude of the order index difference is encoded into the bitstream.

[0124] Therefore, this method provides a very simple solution to the problem of encoding / decoding the later occupied coarse points, while maintaining low-latency encoding of the points in the point cloud.

[0125] In addition, as discussed below, when all the sensed points are correctly sorted, for the data Snext Encoding does not significantly degrade the encoding performance obtained under ideal usage conditions.

[0126] Figure 13 FIG. shows a schematic block diagram of steps of a method 100 for encoding point cloud geometry data into a bitstream of encoded point cloud data according to at least one embodiment.

[0127] In step 110, an order index difference Δo is calculated using equation (4), and data S is obtained based on the order index difference Δo next and encoded into bitstream B. Data S next can indicate whether the order index difference Δo is negative, and thus, whether the occupied coarse point P next associated with a point of the point cloud is a later occupied coarse point.

[0128] If data S next indicates that the occupied coarse point P next is not a later occupied coarse point, then in step 120, the positive order index difference Δo is encoded into bitstream B.

[0129] For example, the positive order index difference Δo can be binarized into a sequence of binary data, and each binary data of the sequence is entropy encoded into bitstream B.

[0130] If data S next indicates that the occupied coarse point P next is a later occupied coarse point, then in step 130, a post-point order index difference Δo is obtained between the order index o(P next ) and the order index o(P'ref) of a second reference coarse point P' ref . Next, in step 140, the magnitude |Δo late | of the post-point order index difference Δo late is encoded into bitstream B. late |

[0131] Optionally, if the magnitude |Δo late | is not empty, then in step 150, the sign of the post-point order index difference Δo late is encoded into bitstream B.

[0132] Figure 14 FIG. shows a schematic block diagram of steps of a method 200 for decoding point cloud geometry data from a bitstream of encoded point cloud data according to at least one embodiment.

[0133] The decoding method 200 is directly derived from the encoding method 100.

[0134] In step 210, data S is decoded from bitstream Bnext 。Data S next Indicates the occupied coarse point P associated with the point of the point cloud to be decoded next Whether it is a later occupied coarse point. When the order of the occupied coarse point in lexicographical order is lower than the order of the first reference coarse point P associated with the previously decoded point of the point cloud ref The occupied coarse point is considered to be the later occupied coarse point P next 。

[0135] If the data S next Indicates the occupied coarse point P next is not a later occupied coarse point, then in step 220, a positive order index difference Δo is decoded from the bitstream B. The positive order index difference Δo represents the order index o(P next ) of the occupied coarse point P next and the order index of the reference coarse point P ref ) of o(P ref ) difference.

[0136] In one embodiment of step 220, a sequence of binary data can be entropy decoded from the bitstream B, and the positive order index difference Δo can be obtained from the decoded sequence of binary data.

[0137] If the data S next Indicates the occupied coarse point P next is a later occupied coarse point, then in step 230, the magnitude |Δo| of the post-point order index difference Δo is decoded from the bitstream B late of late |.

[0138] The post-point order index difference Δo late is the difference between the order index o(P next ) and the order index of the second reference coarse point P' ref .

[0139] Optionally, in step 240, if the magnitude |Δo late | is not empty, the sign of the post-point order index difference Δo late can be decoded from the bitstream B.

[0140] In one embodiment of step 150(240), the sign of the post-point order index difference Δo late can be bypass-coded (i.e., it is directly coded as a bit without entropy coding).

[0141] In one embodiment, the first reference coarse point P refcan be the last encoded or decoded occupied coarse point P associated with the last encoded or decoded point of the point cloud in the encoding or decoding order last . Equation (4) becomes Equation (5) given by:

[0142] Δo = o(P next ) - o(P last ). (5)

[0143] When a later occupied coarse point appears, it is possible that the last encoded (decoded) occupied coarse point P last is not the previously encoded (decoded) occupied coarse point with the highest order in the lexicographical order.

[0144] In one embodiment, the first reference coarse point P ref can be the previously encoded or decoded occupied coarse point P with the highest order index in the lexicographical order high . Equation (4) becomes Equation (6) given by:

[0145] Δo = o(P next ) - o(P high ). (6)

[0146] In one embodiment, the second reference coarse point P' ref can be the first reference coarse point P ref , that is, the last encoded or decoded occupied coarse point P last , or the previously encoded or decoded occupied coarse point P high .

[0147] According to this embodiment, as calculated in Equation (4), the post-point order index difference Δo late is equal to the order index Δo. Then the post-point order index difference Δo late is a negative order index difference, and it may not be positive or empty. Then, steps 150 and 240 are omitted, and the sign can be inferred, such as the post-point order index difference Δo late is negative.

[0148] When there is small noise on the coordinates φ or s in the two-dimensional space, it is particularly applicable to the embodiment where the second reference coarse point P' ref is the last encoded or decoded occupied coarse point P last , because this can reduce the cost of the magnitude for encoding (decoding) the post-point order index difference.

[0149] In one embodiment, the second reference coarse point P' refcan be an occupied or unoccupied coarse point associated with an order index that is equal to the first reference coarse point P shifted by an offset ref of the order index.

[0150] In one variant, the offset can be calculated based on the negative order index difference Δo obtained from the previously encoded or decoded points and the order index o(P ref ) of the first reference coarse point P ref ) for all or a subset of it.

[0151] For example, the offset is equal to 2 N times the average of the negative order index differences Δo, where N is equal to 3 or 4, for example 2 N of the last encoded or decoded negative order index differences Δo.

[0152] In one variant, the magnitude |Δo late | of the post-point order index difference Δo minus 1 can be encoded into (decoded from) the bitstream B. late |

[0153] In an embodiment of methods 100 and 200, for example, the data S next can be a syntax element represented as "next_point_is_late_flag". Thus, the data S next can be binary data (a flag).

[0154] In one variant, the binary data (flag) can be signaled, for example, in the geometric parameter set, to indicate that the data S next does not need to be encoded and that it can always be inferred to indicate that the next occupied coarse point P next is not post-occupied. Thus, if it is guaranteed that the point cloud is in an ideal use case, there is no encoding loss at all. This binary data informs whether the dictionary order index always increases during decoding, which can be useful for decoding applications, for example, when using an optimized decoding or rendering pipeline in an ideal use case.

[0155] In an embodiment of step 110 (210), the data S next can be binary data indicating whether the occupied coarse point (P next ) associated with a point of the point cloud is a post-occupied coarse point, and it can be entropy encoded (decoded) based on at least one other binary data S next,i . Each of this at least one binary data S next,i indicates whether the previously encoded (decoded) occupied coarse point P i associated with a previously encoded (decoded) point of the point cloud is a post-occupied coarse point.

[0156] encoding (decoding) each binary data S independently of each other next compared to, based on the at least one other binary data S next,i for the binary data S next performing entropy encoding (decoding) provides a more efficient encoding (decoding) of the binary data S next .

[0157] In Figure 15 in one embodiment shown, context adaptive binary arithmetic coding (CABAC) can be used to perform entropy encoding (decoding) on the binary data S next .

[0158] In a variant, the context ctxIdx can be selected based on the binary data S associated with N previously encoded (decoded) occupied coarse points. For example, an N-bit word W is formed by concatenating N (e.g., N = 3) binary data S next,i . next,i The N-bit word W formed

[0159] Having N ctx The context table Tctx with entries usually stores the probabilities associated with the context, and the probability p ctxIdx is obtained as the ctxIdx-th entry of the context table. The context is selected based on the context index ctxIdx by the following formula

[0160] Ctx = Tctx[ctxIdx].

[0161] For example, the context index ctxIdx can correspond to the N-bit word W. Then, the context Ctx is selected as

[0162] Ctx = Tctx[W]

[0163] Using the probability p ctxIdx entropy encodes the binary data S next into the bitstream B (decodes from the bitstream B).

[0164] The entropy encoder (decoder) is usually an arithmetic encoder (decoder), but can also be any other type of entropy encoder (decoder), such as an asymmetric digital system. In any case, the optimal encoder adds -log2(p ctxIdx ) bits to the bitstream B to encode S next = 1, or adds -log2(1 - p ctxIdx ) bits to the bitstream B to encode S next = 0. When the binary data S next is encoded (decoded), the encoded (decoded) binary data S is updated by using an updaternext and p ctxIdx Update the probability p as an entry ctxIdx ; The updater typically performs by using an update table. The updated probability replaces the ctxIdx-th entry of the context table Tctx. Then, another binary data S next can be encoded (decoded), and so on. The update loop of the context table is a bottleneck in the encoding workflow because another binary data S next can only be encoded (decoded) after the update is performed. For this reason, the memory access to the context table must be as fast as possible, and minimizing the size of the context table helps to simplify its hardware implementation.

[0165] Select an appropriate context, that is, estimate at most the probability p next that the binary data S is equal to 1 ctxldx , which is crucial for obtaining good compression. Therefore, context selection should use at least one binary data S next,i associated with the last occupied coarse point of the N previously encoded (decoded) and the correlation between them to obtain the appropriate context.

[0166] In one variant, the context index ctxIdx can be obtained from the binary data S next,i associated with the last occupied coarse point of the N previously encoded (decoded) with sensor indices belonging to the determined sensor index group.

[0167] In one variant, multiple sets of consecutive sensors can be formed.

[0168] For example, if the sensor head includes 32 sensors (32 sensor indices), each group can include 4 consecutive sensor indices: {0, 1, 2, 3}, {4, 5, 6, 7}, …, {28, 29, 30, 31}. Then, an N-bit buffer can be used for each group, and whenever binary data S next has been encoded (decoded) for the sensors belonging to that group, the buffer of that group is updated.

[0169] In one variant, each group can contain unique sensors.

[0170] In one variant, the context can be specific to each group of sensor indices.

[0171] The context adaptive binary arithmetic decoder performs basically the same operations as the context adaptive binary arithmetic encoder, except that the entropy decoder uses the probability p ctxIdx to decode the encoded binary data S next from the bitstream B.

[0172] In Figure 16 In one embodiment shown, the post - dot order index difference Δo can be encoded (step 140) by encoding (decoding for decoding, step 230) a first positive offset (step 310) associated with the first coordinate of the two - dimensional space and encoding (decoding) a second positive offset (step 320) associated with the second coordinate of the two - dimensional space. late of the magnitude |Δo late | (or in a variant, the magnitude |Δo late | minus 1).

[0173] When the two - dimensional coordinate space is a two - dimensional coordinate (φ, λ) space, the first positive offset is the positive azimuth offset φ offset and the second positive offset is the positive sensor index offset λ offset .

[0174] Then, the post - dot order index difference Δo late can be obtained by the following formula:

[0175] Δo late = - φ offset *N sensor - λ offset - 1 (7)

[0176] where N sensor is the number of sensors of the sensor head.

[0177] In one embodiment of step 310, the positive azimuth offset φ offset can be encoded by an expGolomb code.

[0178] In one embodiment of step 310, entropy encoding (decoding) of the positive azimuth offset φ offset can be performed based on adaptive context selection.

[0179] In a variant, the positive azimuth offset φ offset can be binarized in a sequence of binary data b k1 . In one example, the sequence of binary data b k1 is the unary representation of the positive azimuth offset φ offset , where k1 is in the range from 0 to φ offset (i.e., it is a sequence of φ offset consecutive binary data equal to 1, followed by a sequence of a binary data equal to 0, or alternatively a sequence of φ offset binary data equal to 0, followed by a sequence of a binary data equal to 1). Each binary data b k1 is entropy - encoded using the same context. For the binary data b k1Decode the sequence of and reconstruct the positive azimuth offset φ from the decoded binary data b k1 from the sequence of offset .

[0180] In one variant, an entropy coding context for encoding or decoding each bit b can be selected based on a function of index k1, such as min(k1, 1). k1

[0181] When the two-dimensional coordinate space is the two-dimensional coordinate (s, λ) space, the first positive offset is the positive sample index offset s offset , and the second offset is the positive sensor index offset λ offset .

[0182] Then, the post-point order index difference Δo late is obtained by:

[0183] Δo late = (-s offset - 1) * N sensor + λ offset (8)

[0184] This variant (Equation 8) is advantageous when the forced sensor index offset λ offset is positive. If the points are correctly sorted by sensor index, this results in easier encoding (and better compression).

[0185] In one embodiment of step 310, the positive sample index offset s offset can be encoded by an expGolomb code.

[0186] In one embodiment of step 310, the positive sample index offset s offset can be entropy encoded (decoded) based on adaptive context selection.

[0187] In one variant, the positive sample index offset s offset can be binarized in the sequence of binary data b k2 to obtain, for example, a unary representation of the positive sample index offset s offset (where k2 is in the range from 0 to s offset ). Each binary data b k2 is entropy encoded using the same context. Entropy decode the sequence of binary data b k2 and reconstruct the positive sample index offset s k2 from the sequence of decoded binary data b offset .

[0188] In one embodiment of step 320, the positive sensor index offset λ​offset It can be encoded by the expGolomb code.

[0189] In one embodiment of step 320, the positive sensor index offset λ offset can be entropy encoded (decoded) based on an adaptive context selection.

[0190] In a variant, the positive sensor index offset λ offset can be binarized in the sequence of binary data b k3 to obtain, for example, the unary representation of the positive sensor index offset λ offset where k3 is in the range from 0 to λ offset . Each binary data b k3 is entropy encoded using a context. The sequence of binary data b k3 is entropy decoded, and the positive sensor index offset λ k3 is reconstructed from the decoded sequence of binary data b offset .

[0191] In a variant, the context can be selected based on the binary data index k3 and / or by concatenating N (e.g., N = 1 or N = 2) binary data S next,i associated with the last occupied coarse points of N previously encoded (decoded) sensors having sensor indices belonging to a determined sensor index group to form an N-bit word W1.

[0192] In a variant, the context for encoding the binary data with index k3 can also be selected based on the index of the sensor index group to which the sensor with index (λ ref + k3) modulo N sensor belongs, where λ ref is the sensor index associated with the second reference coarse point P'. ref

[0193] It should be noted that equation (7) and the related embodiments can be easily applied to a two-dimensional coordinate space, which is a reorganized two-dimensional coordinate (s, λ) space using s offset instead of φ offset . Additionally, equation (8) and the related embodiments can be easily applied to a two-dimensional coordinate space, which is a reorganized two-dimensional coordinate (φ, λ) space using φ offset instead of s offset .

[0194] In Figure 17 one embodiment shown, the first positive offset (step 310) can be encoded (decoded by decoding, step 230) and by relative to the sensor index λref Encode two positive values λ offset,1 and λ offset,2 to encode the positive sensor index offset λ offset (step 140) to encode the back point order index difference Δo late of the magnitude |Δo late | (or in a variant the magnitude |Δo late | minus 1). Determine the positive values λ offset,1 and λ offset,2 such that their sum is equal to the positive sensor offset λ offset : λ offset = λ offset,1 + λ offset,2 and such that

[0195] λ offset,1 ≤ N laser * λ ref

[0196] If λ offset,1 < N laser - λ ref then λ offset,2 = 0

[0197] If λ offset,1 = N laser - λ ref then λ offset,2 ≥ 0

[0198] In step 330, the first positive value λ offset,1 can be encoded into the bitstream B (decoded from the bitstream B).

[0199] In step 340, if λ offset,1 < N sensor - λ ref then the encoding (decoding) of the positive sensor index offset λ offset can end (λ offset,2 = 0).

[0200] In step 350, if λ offset,1 = N sensor - λ ref then the second positive value λ offset,2 can be encoded into the bitstream B (decoded from the bitstream B) (step 360).

[0201] In a variant, the first positive offset may not be encoded before the positive sensor index offset λ offset and in step 350, if λ offset,1 = N sensor - λ ref, the first positive offset is encoded into the bitstream B (decoded from the bitstream B) (step 310). If λ offset,1 < N sensor - λ ref , the first positive offset is considered equal to zero.

[0202] In one variant, if λ offset,1 = N sensor - λ ref , after the first positive value λ offset,1 has been encoded but before the second positive value λ offset,2 is encoded, i.e., after step 350 and before step 360, the first positive offset can be encoded into the bitstream B (decoded from the bitstream B) (step 310).

[0203] In one variant, the positive sensor index offset λ offset can be further restricted to be lower than the number of sensors N sensor .

[0204] In one embodiment of step 330, entropy encoding (decoding) of the positive offset value λ offset,1 can be performed based on adaptive context selection.

[0205] In one variant, the positive offset value λ offset,1 can be binarized in a sequence of binary data bk4 to obtain, for example, the unary representation of the positive offset value λ offset,1 (where k4 ranges from 0 to λ offset,1 ). Each binary data b k4 is entropy encoded using context. Entropy decoding of the sequence of binary data b k4 is performed, and the positive offset value λ k4 is reconstructed from the decoded sequence of binary data b offset,1 .

[0206] In one variant, the positive offset value λ offset,1 can be binarized in a sequence of binary data b k4 to obtain the unary representation of the positive offset value λ offset,1 if λ offset,1 = N sensor - λ ref , k4 can only range from 0 to λ offset,1 - 1: the last bit of the unary representation can be omitted. This is because higher values of the positive offset value λ offset,1 are not possible / allowed.

[0207] In one variant, it can be based on the sensor index λ ref+k4 is used to select the context, or in a variant, based on the index of the sensor index group including the sensor index λ ref +k4 to select the context.

[0208] In one variant, the context can be further selected based on the data S indicating whether the occupied coarse point P i previously encoded (decoded) is a subsequently occupied coarse point next,i The occupied coarse point P i is associated with the points of the point cloud previously encoded (decoded), for which the sensor index associated with the order index of the second reference coarse point P’ ref(i) is used to calculate the subsequent point order index difference P’ ref(i) , because the occupied coarse point P i is equal to the sensor index λ ref , or in a variant, belongs to the sensor index group including the sensor index λ ref .

[0209] In one variant, the context can be further selected based on a function of k4.

[0210] For example, the function can be equal to min(k4,2).

[0211] In one embodiment of step 360, the positive offset value λ offset,2 can be entropy encoded (decoded) based on adaptive context selection.

[0212] In one variant, the positive offset value λ offset,2 can be binarized in the sequence of binary data bk5 to obtain, for example, the unary representation of the positive offset value λ offset,2 (where k5 is in the range from 0 to λ offset,2 ). Each binary data b k5 is entropy encoded using the context. The sequence of binary data b k5 is entropy decoded, and the positive offset value λ k5 is reconstructed from the decoded sequence of binary data b offset,2 .

[0213] In one variant, the context can be further selected based on the data S indicating whether the occupied coarse point P i previously encoded (decoded) is a subsequently occupied coarse point next,i The occupied coarse point P i is associated with the points of the point cloud previously encoded (decoded), for which the sensor index associated with the order index of the second reference coarse point P’ ref(i) is used to calculate the subsequent point order index difference Δo late(i), because the occupied coarse point Pi is equal to the sensor index λ ref , or in a variant, belongs to a sensor index group including the sensor index λ ref .

[0214] In one variant, the context can be further selected based on a function of i.

[0215] For example, the function can be equal to min(λ offset,2 + k5, 2).

[0216] In one variant, the context table for encoding λ offset,1 can also be used to encode λ offset,2 .

[0217] Figure 18 Shows the case when the next occupied coarse point P next has a sample index s ref equal to s next - 1, determining the first positive offset s offset and the positive value λ offset,2 example, where s ref is the sample index of the second reference coarse point P' ref and has a sensor index λ greater than the second reference coarse point P' ref sensor index λ ref of next .

[0218] In an illustrative example, λ offset,1 ≤ N sensor - λ ref (step 340). Then, the second positive value λ offset,2 is determined to be equal to 0, and the first positive offset s offset is also determined to be equal to 0.

[0219] In this example, the first positive value λ to be encoded (decoded) offset,1 is greater than 0 (λ next > λ ref ). Equation (8) gives the positive sample index offset s offset = 0 and the order index difference Δo.

[0220] Figure 19 Shows the case when the next occupied coarse point P next and the second reference coarse point P' ref have the same sample index, determining the first positive offset s offset and the positive value λ offset,2 example. The sample index offset s offset is equal to 0.

[0221] In an illustrative example, λ offset,1 = N sensor - λ ref (step 350). Then, the second positive value λ offset,2 is encoded into the bitstream B. The positive sample index offset s offset = 0, and the order index difference Δo is given by Equation (8).

[0222] Figure 20 Illustrates an example where the next occupied coarse point P next has a sample index s ref equal to s next - 2, determining the first positive offset s offset and the positive value λ offset,2 where s ref is the sample index of the second reference coarse point P'. ref

[0223] In an illustrative example, λ offset,1 = N sensor - λ last (step 350). Then, the second positive value λ offset,2 is encoded into the bitstream B. The positive sample index offset s offset is equal to 2, and the order index difference Δo is given by Equation (8).

[0224] Figure 21 Illustrates a schematic block diagram of an example of a system in which various aspects and embodiments are implemented.

[0225] System 400 may be embedded as one or more devices, including the various components described below. In various embodiments, system 400 may be configured to implement one or more aspects described in this application.

[0226] ​Examples of equipment that can form all or part of system 400 include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from the video decoder, pre-processors that provide input to the video encoder, web servers, set-top boxes, and any other device for processing point clouds, video, or images, or other communication devices. The elements of system 400 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 400 can be distributed across multiple ICs and / or discrete components. In various embodiments, system 400 can be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.

[0227] System 400 can include at least one processor 410 that is configured to execute instructions loaded therein for implementing, for example, the various aspects described in the present application. Processor 410 can include embedded memory, input / output interfaces, and various other circuits known in the art. System 400 can include at least one memory 420 (e.g., volatile memory devices and / or non-volatile memory devices). System 400 can include a storage device 440, which can include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 can include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0228] System 400 may include an encoder / decoder module 430, which is configured to process data, for example, to provide encoded / decoded point cloud geometry data, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 may represent one or more modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 430 may be implemented as a separate element of the system 400 or may be incorporated into the processor 410 as a combination of hardware and software known to those skilled in the art.

[0229] The program code to be loaded onto the processor 410 or the encoder / decoder 430 to execute the various aspects described in this application may be stored in the storage device 440 and subsequently loaded onto the memory 420 for execution by the processor 410. According to various embodiments, during the execution of the processes described in this application, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 may store one or more of various items. Such stored items may include, but are not limited to, point cloud frames, encoded / decoded geometry / attribute video / images or portions thereof, bitstreams, matrices, variables, and intermediate or final results of equations, formulas, operations, and operation logic processing.

[0230] In several embodiments, the memory internal to the processor 410 and / or the encoder / decoder module 430 may be used to store instructions and provide working memory for processing that may be performed during encoding or decoding.

[0231] However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory may be the memory 420 and / or the storage device 440, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, fast external dynamic volatile memory such as RAM may be used as working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), or MPEG-I Part 5 or Part 9.

[0232] As indicated in block 490, input to the elements of system 400 can be provided via a variety of input devices. Such input devices include, but are not limited to, (i) an RF section that can receive RF signals, such as those transmitted over the air by a broadcast device, (ii) composite input terminals, (iii) USB input terminals, and / or (iv) HDMI input terminals.

[0233] In various embodiments, the input devices of block 490 have associated respective input processing elements, as is known in the art. For example, the RF section can be associated with elements necessary for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band), (ii) down-converting the selected signal, (iii) band-limiting the band again to a narrower band to select a signal band that can be referred to as a channel, for example, in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) de-multiplexing to select a desired data packet stream. The RF section of various embodiments can include one or more elements that perform these functions, such as, for example, a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down-converter, a demodulator, an error corrector, and a de-multiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband.

[0234] In one set-top box embodiment, the RF section and its associated input processing elements can receive RF signals transmitted over a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and filtering again to a desired band.

[0235] Various embodiments re-order the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0236] Adding elements can include inserting elements between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section can include an antenna.

[0237] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 400 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented, for example, within a separate input processing IC or within the processor 410 when necessary. Similarly, various aspects of USB or HDMI interface processing can be implemented within a separate interface IC or within the processor 410 when necessary. The demodulated, error-corrected, and demultiplexed stream can be provided to various processing elements, including, for example, the processor 410 and the encoder / decoder 430, which operate in conjunction with memory and storage elements to process the data stream for presentation on an output device when necessary.

[0238] The various elements of the system 400 can be provided within an integrated housing. Within the integrated housing, a suitable connection arrangement 490, such as internal buses (including I2C buses), wiring, and printed circuit boards known in the art, can be used to interconnect the various elements and transfer data between them.

[0239] The system 400 may include a communication interface 450 that enables communication with other devices via a communication channel 800. The communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 800. The communication interface 450 may include, but is not limited to, a modem or a network card, and the communication channel 800 may be implemented, for example, within a wired and / or wireless medium.

[0240] In various embodiments, data can be streamed to the system 400 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments can be received via the communication channel 800 and the communication interface 450 suitable for Wi-Fi communication. The communication channel 800 of these embodiments can typically be connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-top communications.

[0241] Other embodiments can use a set-top box to provide streamed data to the system 400, which delivers the data via an HDMI connection of the input block 490.

[0242] Still other embodiments can use an RF connection of the input block 490 to provide streamed data to the system 400.

[0243] The streamed data can be used as a means of signaling information used by the system 400. The signaling information can include the bitstream B and / or information such as the number of points, coordinates, and / or sensor setting parameters of a point cloud.

[0244] It should be recognized that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. can be used to signal information to a corresponding decoder.

[0245] System 400 can provide output signals to various output devices, including display 500, speaker 600, and other peripheral devices 700. In various examples of the embodiments, other peripheral devices 700 can include one or more of a standalone DVR, disc player, stereo system, lighting system, and other devices based on the output providing function of system 400.

[0246] In various embodiments, control signals can be communicated between system 400 and display 500, speaker 600, or other peripheral devices 700 using signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention.

[0247] The output devices can be communicatively coupled to system 400 via dedicated connections through respective interfaces 460, 470, and 480.

[0248] Optionally, the output devices can be connected to system 400 using communication channel 800 via communication interface 450. Display 500 and speaker 600 can be integrated with other components of system 400 in an electronic device (such as, for example, a television) in a single unit.

[0249] In various embodiments, display interface 460 can include a display driver, such as, for example, a timing controller (TCon) chip.

[0250] For example, if the RF portion of input terminal 490 is part of a separate set-top box, then display 500 and speaker 600 can optionally be separate from one or more of the other components. In various embodiments where display 500 and speaker 600 can be external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0251] In Figures 1 to 21 this, various methods are described, and each method includes one or more steps or actions to implement the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined.

[0252] Some examples are described with respect to block diagrams and / or operational flowcharts. Each block represents a circuit element, a module, or a portion of code that includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that in other embodiments, the function(s) noted in the blocks may not occur in the order indicated. For example, depending on the functionality involved, two blocks shown in succession may actually be executed substantially concurrently, or sometimes the blocks may be executed in the reverse order.

[0253] The embodiments and aspects described herein can be implemented in, for example, a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if discussed only in the context of a single form of embodiment (e.g., only as a method), the embodiments of the features discussed can be implemented in other forms (e.g., an apparatus or a computer program).

[0254] A method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device.

[0255] In addition, a method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the embodiments) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product implemented in one or more computer-readable media and having computer-readable program code executable by a computer implemented thereon. Considering the inherent ability to store information therein and the inherent ability to retrieve information therefrom, a computer-readable storage medium as used herein can be considered a non-transitory storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. It should be recognized that the following, while providing more specific examples of computer-readable storage media to which the present embodiment can be applied, is merely illustrative and not an exhaustive list: a portable computer floppy disk; a hard disk; a read-only memory (ROM); an erasable programmable read-only memory (EPROM or flash memory); a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing.

[0256] The instructions can form an application program tangibly implemented on a processor-readable medium.

[0257] For example, the instructions can be in hardware, firmware, software, or a combination. For example, the instructions can be found in an operating system, a separate application, or a combination of both. Thus, a processor can be characterized as, for example, a device configured to execute a process and a device that includes a processor-readable medium (such as a storage device) having instructions for executing the process. Additionally, in addition to or instead of the instructions, the processor-readable medium can store data values generated by an implementation.

[0258] The apparatus can be implemented in, for example, suitable hardware, software, and firmware. Examples of such apparatus include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, set-top boxes, and any other device for processing point clouds, video, or images, or other communication devices. It should be clear that the equipment can be mobile and even installed in a moving vehicle.

[0259] The computer software can be implemented by the processor 410 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can also be implemented by one or more integrated circuits. The memory 420 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). The processor 410 can be of any type suitable for the technical environment and can encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.

[0260] As will be apparent to those of ordinary skill in the art, the implementations can generate various signals that are formatted to carry information such as can be stored or transmitted. The information can include, for example, instructions for executing a method or data generated by one of the described implementations. For example, the signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (such as using the radio frequency portion of the spectrum) or a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted via various different wired or wireless links. The signal can be stored on a processor-readable medium.

[0261] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" may also be intended to include the plural forms, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "include / comprise" and / or "including / comprising" may specify the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. Also, when an element is referred to as being "responsive" or "connected" to another element, it may be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to another element, no intervening elements are present.

[0262] It should be recognized that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any of the symbols / terms " / ", "and / or", and "at least one" may be intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such language is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.

[0263] Various numerical values may be used in this application. Specific values may be used for illustrative purposes and the aspects described are not limited to these specific values.

[0264] It will be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of this application, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. No order is implied between the first element and the second element.

[0265] References to "one embodiment" or "an embodiment" or "some embodiments" or "an implementation" or "an implementation" and other variations thereof are frequently used to convey that a particular feature, structure, characteristic, etc. (described in conjunction with an embodiment / implementation) is included in at least one embodiment / implementation. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in some embodiments" or "in an implementation" or "in an implementation" and any other variations appearing in various places in this application do not necessarily all refer to the same embodiment.

[0266] Similarly, references herein to "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" and other variations thereof are frequently used to convey that a particular feature, structure, or characteristic (described in conjunction with an embodiment / example / implementation) may be included in at least one embodiment / example / implementation. Therefore, the expressions "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" appearing in various places in the specification do not necessarily all refer to the same embodiment / example / implementation, nor do separate or alternative embodiments / examples / implementations necessarily exclude other embodiments / examples / implementations.

[0267] Reference numerals appearing in the claims are for illustration only and have no limiting effect on the scope of the claims.The present embodiments / examples and variants may be employed in any combination or sub-combination although not explicitly described.

[0268] When a figure is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.

[0269] While some diagrams include arrows on communication paths to illustrate a primary direction of communication, it should be understood that communication can occur in the opposite direction to the depicted arrows.

[0270] Various embodiments relate to decoding. As used in this application, "decoding" can encompass, for example, all or part of a process performed on a received point cloud frame (which may include a received bitstream encoding one or more point cloud frames) to produce a final output suitable for display or further processing in a reconstructed point cloud domain. In various embodiments, such processes include one or more of the processes typically performed by a decoder. In various embodiments, for example, such processes also or alternatively include processes performed by a decoder of various embodiments described in this application.

[0271] As a further example, in one embodiment "decoding" may refer only to dequantization, in one embodiment "decoding" may refer to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of dequantization, entropy decoding, and differential decoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process, and it is believed that those skilled in the art will well understand this.

[0272] Various embodiments relate to encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application may cover, for example, all or part of the process performed on an input point cloud frame to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder. In various embodiments, such processes also include or optionally include the processes performed by the encoders of the various embodiments described in this application.

[0273] As a further example, in one embodiment "encoding" may refer only to quantization, in one embodiment "encoding" may refer only to entropy encoding, in another embodiment, "encoding" may refer only to differential encoding, and in another embodiment, "encoding" may refer to a combination of quantization, differential encoding, and entropy encoding. Based on the context of the specific description, it will be clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process, and it is believed that those skilled in the art will well understand this.

[0274] In addition, this application may refer to "obtaining" various information. Obtaining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from a memory.

[0275] Additionally, this application may refer to "accessing" various information. Accessing information may include, for example, one or more of the following: receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0276] Furthermore, this application may refer to "receiving" various information. Like "accessing", receiving is intended to be a broad term. Receiving information may include one or more of the following: for example, accessing information or retrieving information (e.g., from a memory). Additionally, in one way or another, during operations such as, for example: storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is generally involved.

[0277] Moreover, as used herein, the term "signal" particularly refers to indicating something to a corresponding decoder, etc. For example, in some embodiments, the encoder signals specific information, such as the number of points or coordinates of a point cloud or sensor setting parameters. In this way, in an embodiment, the same parameters can be used on the encoder side and the decoder side. Thus, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder such that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameters. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It should be recognized that signaling can be accomplished in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the foregoing related to the verb form of the term "signal", the term "signal" can also be used as a noun herein.

[0278] Multiple embodiments have been described. However, it should be understood that various modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to yield other embodiments. Additionally, those of ordinary skill in the art will understand that other structures and processes can replace the disclosed structures and processes, and the resulting embodiments will perform at least substantially the same (one or more) functions in at least substantially the same (one or more) ways to achieve at least substantially the same (one or more) results as the disclosed embodiments. Thus, the present application contemplates these and other embodiments.

Claims

1. A method for encoding point cloud geometric data represented by an ordered set of coarse points of some discrete positions in a discrete position set occupying a two-dimensional space into a bitstream, the ordered coarse points being sorted according to a lexicographical order based on coordinates of the two-dimensional space, wherein the method includes: -Encode data (S next ) indicating whether an occupied coarse point (P next ) associated with a point of the point cloud is a later occupied coarse point into the bitstream. When the order index of the occupied coarse point in the lexicographical order is lower than the order index of a first reference coarse point (P ref ) associated with a previously encoded point of the point cloud, the occupied coarse point is considered a later occupied coarse point (P next ); next next ref next ​​​​ - If the data (S next ) indicates that the occupied coarse point is a later occupied coarse point (P next ), then Obtain the rough point (P) occupied after (130) next ) and the second reference coarse point (P' ref ) between the order indices of the next points ( ), the second reference coarse point (P' ref ) is the first reference coarse point or a coarse point associated with an order equal to the order of the first reference coarse points shifted by an offset; and Encoding (140) the magnitude of the post-point order index difference into the bitstream, wherein the magnitude of the post-point order index difference is encoded by encoding two positive offsets of the two-dimensional space into the bitstream; the two offsets include a first positive offset associated with a first coordinate of the two-dimensional space and a second positive offset associated with a second coordinate of the two-dimensional space.

2. The method according to claim 1, wherein the first reference coarse point (P ref ) is the last encoded occupied coarse point (P last ) or the previously encoded occupied coarse point (P high ) having the highest order in the lexicographical order.

3. The method according to claim 1, wherein the data (S next ) is binary data indicating whether an occupied coarse point (P next ) associated with a point of the point cloud is a subsequently occupied coarse point, and the data (S next,i ) is entropy encoded based on at least one other binary data (S next ) indicating whether at least one previously encoded occupied coarse point is a subsequently occupied coarse point.

4. The method according to claim 1, wherein encoding the second positive offset ( ) includes encoding a first positive value ( ) and a second positive value ( ) whose sum is equal to the second positive offset ( ).

5. The method according to claim 4, wherein encoding the first positive value or the second positive value includes entropy encoding a sequence of binary data representing the first positive value or the second positive value.

6. The method according to any one of claims 4 or 5, wherein encoding the first positive offset includes entropy encoding a sequence of binary data representing the first positive offset, the first positive offset being or .

7. A method for decoding point cloud geometric data represented by an ordered set of coarse points of some discrete positions in a discrete position set occupying a two-dimensional space from a bitstream, the ordered coarse points being sorted according to a lexicographical order based on coordinates of the two-dimensional space, wherein the method includes: - Decode (210) from the bitstream data indicating whether an occupied coarse point (P next ), which is associated with a point of the point cloud, is a later occupied coarse point (S next ). When the order of the occupied coarse point in the lexicographical order is lower than the order of a first reference coarse point (P ref ) associated with a previously decoded point of the point cloud, the occupied coarse point is considered to be a later occupied coarse point (P next ); and - If the data (S next ) indicates that the occupied coarse point is a later occupied coarse point (P next ), then decode (230) from the bitstream the magnitude of the post-point order index difference ( ) between the order index of the later occupied coarse point (P next ) and the order index of the second reference coarse point (P' ref ), where the second reference coarse point (P' ref ) is the first reference coarse point or a coarse point associated with an order that is one offset of the order shift equal to the first reference coarse point; wherein the magnitude of the post-point order index difference is decoded by decoding two positive offsets of the two-dimensional space from the bitstream; the two offsets include a first positive offset associated with a first coordinate of the two-dimensional space and a second positive offset associated with a second coordinate of the two-dimensional space.

8. The method according to claim 7, wherein the first reference coarse point (P ref ) is the last decoded occupied coarse point (P last ) or the previously decoded occupied coarse point (P high ) having the highest order in the lexicographical order.

9. The method according to claim 7, wherein the data (S next ) is binary data indicating whether an occupied coarse point (P next ) associated with a point of the point cloud is a subsequently occupied coarse point, and the data (S next,i ) is decoded based on at least one other binary data (S next ) indicating whether at least one previously decoded occupied coarse point is a subsequently occupied coarse point.

10. The method according to claim 7, wherein decoding the second positive offset ( ) comprises decoding a first positive value ( ) and a second positive value ( ) whose sum is equal to the second positive offset ( ).

11. The method according to claim 10, wherein decoding the first positive value or the second positive value includes decoding a sequence of binary data representing the first positive value or the second positive value.

12. The method according to any one of claims 10 or 11, wherein decoding the first positive offset includes decoding a sequence of binary data representing the first positive offset, the first positive offset being or .

13. An apparatus for encoding point cloud geometric data represented by an ordered set of coarse points of some discrete positions in a discrete position set occupying a two-dimensional space into a bitstream, the ordered coarse points being sorted according to a lexicographical order based on coordinates of the two-dimensional space, wherein the apparatus includes at least one processor configured to: -Encode data (S next ) indicating whether an occupied coarse point (P next ) associated with a point of the point cloud is a later-occupied coarse point into the bitstream, where the occupied coarse point is considered a later-occupied coarse point (P next ) when the order index of the occupied coarse point in the lexicographical order is lower than the order index of a first reference coarse point (P ref ) associated with a previously-encoded point of the point cloud; next ), next when the order index of the occupied coarse point in the lexicographical order is lower than the order index of a first reference coarse point (P ref ) associated with a previously-encoded point of the point cloud, ref the occupied coarse point is considered a later-occupied coarse point (P next ); next ​ - If the data (S next ) indicates that the occupied coarse point is a later occupied coarse point (P next ), then Get the rough points occupied by the latter (P next ) and the second reference coarse point (P' ref ) between the order indices of the next points ( ), the second reference coarse point (P' ref ) is the first reference coarse point or a coarse point associated with an order equal to the order of the first reference coarse points shifted by an offset; and Encoding the magnitude of the post-point order index difference into the bitstream, wherein the magnitude of the post-point order index difference is encoded by encoding two positive offsets of the two-dimensional space into the bitstream; the two offsets include a first positive offset associated with a first coordinate of the two-dimensional space and a second positive offset associated with a second coordinate of the two-dimensional space.

14. An apparatus for decoding point cloud geometric data represented by an ordered set of coarse points of some discrete positions in a discrete position set occupying a two-dimensional space from a bitstream, the ordered coarse points being sorted according to a lexicographical order based on coordinates of the two-dimensional space, wherein the apparatus includes at least one processor configured to: - Decoding from the bitstream data (S next ) indicating whether an occupied coarse point (P next ) associated with a point of the point cloud is a subsequently occupied coarse point, where the occupied coarse point is considered to be a subsequently occupied coarse point (P next ) when the order of the occupied coarse point in the lexicographical order is lower than the order of a first reference coarse point (P ref ) associated with a previously decoded point of the point cloud; and next ), when the order of the occupied coarse point in the lexicographical order is lower than the order of a first reference coarse point (P ref ) associated with a previously decoded point of the point cloud, the occupied coarse point is considered to be a subsequently occupied coarse point (P next ); and next ), when the order of the occupied coarse point in the lexicographical order is lower than the order of a first reference coarse point (P ref ) associated with a previously decoded point of the point cloud, the occupied coarse point is considered to be a subsequently occupied coarse point (P next ); and ref ), when the order of the occupied coarse point in the lexicographical order is lower than the order of a first reference coarse point (P ref ) associated with a previously decoded point of the point cloud, the occupied coarse point is considered to be a subsequently occupied coarse point (P next ); and next ); and - If the data (S next ) indicates that the occupied coarse point is a later occupied coarse point (P next ), then decode from the bitstream the magnitude of the post-point order index difference ( next ) between the order index of the later occupied coarse point (P ref ) and the order index of a second reference coarse point (P’ ), where the second reference coarse point (P’ ref ) is the first reference coarse point or a coarse point associated with an order that is one offset shifted in order from the first reference coarse point; Among them, The magnitude of the post - point order index difference is decoded by decoding two positive offsets of the two - dimensional space from the bitstream; the two offsets include a first positive offset associated with a first coordinate of the two - dimensional space and a second positive offset associated with a second coordinate of the two - dimensional space.

15. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method of encoding point - cloud geometric data represented by ordered coarse points that occupy some of the discrete positions in a set of discrete positions in a two - dimensional space into a bitstream, the ordered coarse points being sorted according to a lexicographical order based on the coordinates of the two - dimensional space, wherein the method comprises: -Encode data (S next ) indicating whether an occupied coarse point (P next ) associated with a point of the point cloud is a later occupied coarse point into the bitstream, where the occupied coarse point is considered a later occupied coarse point (P next ) when the order index of the occupied coarse point in the lexicographical order is lower than the order index of a first reference coarse point (P ref ) associated with a previously encoded point of the point cloud; next next ref next ​​​​ - If the data (S next ) indicates that the occupied coarse point is a later occupied coarse point (P next ), then Obtain the rough point (P) occupied after (130) next ) and the second reference coarse point (P' ref ) between the order indices of the next points ( ), the second reference coarse point (P' ref ) is the first reference coarse point or a coarse point associated with an order equal to the order of the first reference coarse points shifted by an offset; and Encoding (140) the magnitude of the post - point order index difference into the bitstream; wherein the magnitude of the post - point order index difference is encoded by encoding two positive offsets of the two - dimensional space into the bitstream; the two offsets include a first positive offset associated with a first coordinate of the two - dimensional space and a second positive offset associated with a second coordinate of the two - dimensional space.

16. A non - transitory storage medium carrying instructions of program code that, when executed by a processor, implement a method of encoding point - cloud geometric data represented by ordered coarse points that occupy some of the discrete positions in a set of discrete positions in a two - dimensional space into a bitstream, the ordered coarse points being sorted according to a lexicographical order based on the coordinates of the two - dimensional space, wherein the method comprises: -Encode data (S next ) indicating whether an occupied coarse point (P next ) associated with a point of the point cloud is a later occupied coarse point into the bitstream. When the order index of the occupied coarse point in the lexicographical order is lower than the order index of a first reference coarse point (P ref ) associated with a previously encoded point of the point cloud, the occupied coarse point is considered to be a later occupied coarse point (P next );​​​​​​​​ - If the data (S next ) indicates that the occupied coarse point is a later-occupied coarse point (P next ), then Obtain the rough point (P) occupied after (130) next ) and the second reference coarse point (P' ref ) between the order indices of the next points ( ), the second reference coarse point (P' ref ) is the first reference coarse point or a coarse point associated with an order equal to the order of the first reference coarse points shifted by an offset; and Encoding (140) the magnitude of the post - point order index difference into the bitstream; wherein the magnitude of the post - point order index difference is encoded by encoding two positive offsets of the two - dimensional space into the bitstream; the two offsets include a first positive offset associated with a first coordinate of the two - dimensional space and a second positive offset associated with a second coordinate of the two - dimensional space.

17. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method of decoding point - cloud geometric data represented by ordered coarse points that occupy some of the discrete positions in a set of discrete positions in a two - dimensional space from a bitstream, the ordered coarse points being sorted according to a lexicographical order based on the coordinates of the two - dimensional space, wherein the method comprises: - Decode (210) from the bitstream data indicating whether an occupied coarse point (P next ), which is associated with a point of the point cloud, is a later-occupied coarse point (S next ), where an occupied coarse point is considered to be a later-occupied coarse point (P ref ) when the order of the occupied coarse point in the lexicographical order is lower than the order of a first reference coarse point (P next ) associated with a previously decoded point of the point cloud; and - If the data (S next ) indicates that the occupied coarse point is a later occupied coarse point (P next ), then the magnitude of the post-point order index difference ( ) between the order index of the later occupied coarse point (P next ) decoded (230) from the bitstream and the order index of the second reference coarse point (P' ref ) is decoded, where the second reference coarse point (P' ref ) is the first reference coarse point or a coarse point associated with an order that is one offset shifted from the order of the first reference coarse point; wherein the magnitude of the post - point order index difference is decoded by decoding two positive offsets of the two - dimensional space from the bitstream; the two offsets include a first positive offset associated with a first coordinate of the two - dimensional space and a second positive offset associated with a second coordinate of the two - dimensional space.

18. A non - transitory storage medium carrying instructions of program code that, when executed by a processor, implement a method of decoding point - cloud geometric data represented by ordered coarse points that occupy some of the discrete positions in a set of discrete positions in a two - dimensional space from a bitstream, the ordered coarse points being sorted according to a lexicographical order based on the coordinates of the two - dimensional space, wherein the method comprises: - Decode (210) from the bitstream data indicating whether an occupied coarse point (P next ), which is associated with a point of the point cloud, is a subsequently occupied coarse point (S next ), where the occupied coarse point is considered to be a subsequently occupied coarse point (P ref ) when the order of the occupied coarse point in the lexicographical order is lower than the order of a first reference coarse point (P next ) associated with a previously decoded point of the point cloud; and - If the data (S next ) indicates that the occupied coarse point is a later occupied coarse point (P next ), then decode (230) from the bitstream the magnitude of the post-point order index difference ( ) between the order index of the later occupied coarse point (P next ) and the order index of a second reference coarse point (P' ref ), where the second reference coarse point (P' ref ) is the first reference coarse point or a coarse point associated with an order that is one offset of the order shift equal to the first reference coarse point; Wherein, the magnitude of the post dot order index difference is decoded by decoding two positive offsets of the two-dimensional space from the bitstream; the two offsets include a first positive offset associated with a first coordinate of the two-dimensional space and a second positive offset associated with a second coordinate of the two-dimensional space.

Citation Information

Patent Citations

  • Angular priors for improved prediction in tree-based point cloud coding

    WO2021084295A1