Method and apparatus for encoding / decoding a point cloud captured by a spin sensor head

By transmitting scaling offset information in the bitstream for point clouds captured by spin sensor heads, the method enables efficient encoding and decoding of point clouds, addressing the challenges of low latency and high compression performance in existing technologies.

JP7699245B2Active Publication Date: 2025-06-26BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023580931
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-02
Filing Date
2022-04-23
Publication Date
2025-06-26
Estimated Expiration
2042-04-23

AI Technical Summary

Technical Problem

Existing point cloud compression technologies face challenges in achieving efficient encoding and decoding while maintaining low latency and high compression performance, particularly in applications involving sparse geometric data captured by spin sensor heads.

Method used

The method involves encoding and decoding point clouds using spherical coordinates, where scaling offset information is transmitted in the bitstream to facilitate efficient attribute encoding and decoding, eliminating the need for buffering all coordinates before attribute processing.

Benefits of technology

This approach reduces memory occupancy and maintains decoding efficiency by allowing attribute decoding to start without waiting for complete coordinate decoding, while ensuring scaled spherical coordinates are non-negative and within a specific range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699245000021
    Figure 0007699245000021
  • Figure 0007699245000022
    Figure 0007699245000022
  • Figure 0007699245000023
    Figure 0007699245000023
Patent Text Reader

Abstract

The present invention provides a method and apparatus for encoding / decoding a point cloud into / from a bit stream of encoded point cloud data captured by a spin sensor head. Each point of the point cloud is associated with a spherical coordinate and an attribute. The method includes the steps of transmitting a signal in the bit stream to signal scaling offset information representing a scaling offset, for each current point of the point cloud, encoding / decoding the spherical coordinate of the current point, obtaining the decoded spherical coordinate of the current point from the encoded spherical coordinate, scaling the decoded spherical coordinate of the current point with the scaling offset, and encoding / decoding at least one attribute of the current point based on the scaled decoded spherical coordinate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference to Related Applications This application claims the priority and benefit of European Patent Application No. 21305920.7 filed on July 2, 2021, the entire content of which is incorporated herein by reference.

[0002] This application generally relates to point cloud compression, and specifically to a method and apparatus for encoding / decoding the positions and attributes of points in a point cloud captured by a spin sensor head.

Background Art

[0003] This section is intended to introduce the reader to aspects of the art, which are related to aspects of at least one exemplary embodiment of the application described and / or claimed below. This discussion is recognized to be useful in understanding the aspects of the application by providing the reader with background information.

[0004] As a form of representation of 3D data, point clouds have recently attracted attention because they have various functions in representing all types of physical objects or scenes. Point clouds can be used for various purposes such as cultural heritage and buildings. For example, by scanning an object such as a statue or a building in a 3D manner, the spatial arrangement of the object can be shared without sending or visiting the object. Also, it is a way to preserve knowledge of an object in case the object is destroyed, such as a temple destroyed by an earthquake. Such point clouds are usually static, colored, and huge.

[0005] Another example of use is in topology and cartography. When using 3D representations, maps are not limited to flat surfaces and can include terrain with elevation. Currently, Google Maps is a good example of a 3D map, but it uses meshes instead of point clouds. However, point clouds can also be a suitable data format for 3D maps, and such point clouds are usually static, colored, and huge.

[0006] Virtual reality (VR), augmented reality (AR), and immersive worlds have recently become a topic of discussion, and many people foresee them as the future of 2D flat videos. The basic concept is to immerse the viewer in the surrounding environment, in contrast to a standard TV where the viewer can only see the virtual world in front of their eyes. Depending on the degree of freedom of the viewer in the environment, there are several levels of immersion. Point clouds are a candidate for a suitable format for distributing VR / AR worlds.

[0007] The automotive industry, especially with the predicted self-driving cars, is also an area where point clouds can be used extensively. Self-driving cars need to "detect" the surrounding environment and make good driving decisions based on the presence and nature of the nearest detected objects and the road configuration.

[0008] A point cloud is a set of points in three-dimensional (3D) space, and optionally additional values are added to each point. These additional values are usually called attributes. The attributes may be, for example, a color consisting of three elements, material properties (such as reflectivity), and / or a two-component normal vector of the surface associated with the point.

[0009] Therefore, a point cloud is a combination of geometry (position in 3D space, usually represented by 3D Cartesian coordinates x, y, and z) and at least one attribute.

[0010] Point clouds can be captured by various devices such as an array of cameras, depth sensors, laser devices (light detection and ranging, also called lidar), radar, or can be generated by a computer (e.g., in movie post-production, etc.). Depending on the usage example, point clouds may have thousands to up to billions of points when used in mapping applications. The original representation of a point cloud requires a very high number of bits per point, at least a dozen bits for each of the Cartesian coordinates x, y, or z, and optionally more bits for the (one or more) attributes, for example, three times 10 bits for color.

[0011] In many applications, it is very important to distribute point clouds to end users or store them on servers while consuming only an appropriate amount of bitrate or storage space while maintaining an acceptable (or very good) quality of experience. Efficient compression of these point clouds is an important point for practicalizing the distribution chain in many immersive worlds.

[0012] For distribution and visualization by end users, such as in AR / VR glasses and other 3D-capable devices, compression can be lossy compression (e.g., in video compression). Other use cases, such as medical applications and autonomous driving, definitely require lossless compression so that the results of decisions obtained from subsequent analysis of the compressed and transmitted point clouds are not changed.

[0013] Until recently, the problem of point cloud compression (also known as PCC) has not been addressed in the mass market and there has been no available standardized point cloud codec. In 2017, the ISO / JCT1 / SC29 / WG11, a standardization working group also known as the Moving Picture Experts Group or MPEG, started a work item on point cloud compression. This has resulted in the following two standards. Namely, MPEG-I Part 5 (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC) MPEG-I Part 9 (ISO / IEC 23090-9) or Geometry-based Point Cloud Compression (G-PCC)

[0014] The V-PCC coding method obtains 2D patches packed in an image (or video when dealing with dynamic point clouds) by performing multiple projections on a 3D object to compress the point cloud. Then, using a conventional image / video codec, the acquired image or video is compressed, thereby making the most of the already developed solutions for images and videos. Since the image / video codec cannot usually compress non-smooth patches such as non-smooth patches obtained from projections of sparse geometric data captured by a lidar, essentially, V-PCC is only efficient on high-density and continuous point clouds.

[0015] There are two solutions for the G-PCC coding method to compress the captured geometric data.

[0016] The first solution is based on an occupancy tree, which is locally one of an octree, a quadtree, or a binary tree and represents the geometric shape of the point cloud. Occupied nodes are divided until they reach a certain size, and occupied leaf nodes provide the 3D positions of points, usually at the centers of these nodes. The occupancy information is carried by occupancy flags, which send signals to notify the occupancy status of each child node of the node. By using an adjacent-based prediction technique, high-level compression of the occupancy flags of a high-density point cloud becomes possible. Sparse point clouds can be addressed by directly encoding the positions of points other than the minimum size within the node, and when only outliers exist in the node, the tree structure is stopped, and this technique is called the direct coding mode (DCM).

[0017] The second solution is based on a prediction tree, where each node represents the 3D position of one point, and the parent / child relationship between nodes represents a spatial prediction from the parent to the child. This method can only solve sparse point clouds and has the advantage of providing lower latency and simpler decoding than the occupancy tree. However, compared with the first occupancy-based method, the compression performance is only slightly better, and since the encoder needs to intensively search for the optimal predictor (from a long list of potential predictors) when constructing the prediction tree, the encoding also becomes complex.

[0018] In these two types of solutions, the encoding (decoding) of attributes is performed after the geometric encoding (decoding) is completed. In fact, two encodings occur. Therefore, the low latency of joint geometry / attributes is obtained by using slices that divide the 3D space into independently encoded sub-volumes, and there is no need to predict between sub-volumes. Using a large number of slices will have a great impact on the compression performance.

[0019] Combining the requirements of simplicity, low latency, and compression performance of the encoder and decoder is a problem that conventional point cloud codecs have not fully addressed.

[0020] One important use case is the transmission of sparse geometric data captured by a spin sensor head (e.g., a spin rider head) mounted on a moving vehicle. This usually requires a simple and low-latency embedded encoder. Since the encoder may be placed in a computing unit for parallel execution of other processes (e.g., (semi)autonomous driving), simplicity is required, and thus the available processing power of the point cloud encoder is limited. To enable high-speed transmission from the vehicle to the cloud, low latency is also required, which facilitates real-time checking of local traffic based on the collection of multiple vehicles and making decisions at a sufficient speed based on traffic information. Although using 5G can also reduce the latency sufficiently, the encoder itself should not introduce excessive latency due to encoding. And since the data stream from millions of vehicles to the cloud can be very large, compression performance becomes very important.

[0021] Certain a priori related to the sparse geometric data captured by the spin sensor head are used to obtain a very efficient encoding / decoding method.

[0022] For example, G-PCC utilizes the elevation angles (with respect to the horizontal ground) captured by a spin sensor head, such as those depicted in FIGS. 1 and 2. The spin sensor head 10 includes a set of sensors 11 (e.g., lasers), where five sensors are shown. The spin sensor head 10 can capture the geometric data of a physical object, i.e., the 3D positions of the points in the point cloud, by rotating around the vertical axis z. Subsequently, the geometric data captured by the spin sensor head is represented in spherical coordinates (r 3D , φ, θ), where r 3D is the distance between point P and the center of the spin sensor head, φ is the azimuth angle of spin with respect to the reference of the sensor head, and θ is the elevation angle with respect to the elevation angle index k of the horizontal reference plane (where the y-axis) of the sensors of the spin sensor head. The elevation angle index k may be relative, for example, the elevation angle with respect to sensor k, or the k-th sensor position when a single sensor successively detects each of the consecutive elevation angles.

[0023]

Number

[0024] In spherical coordinate space, using the discrete nature of the angles, the position of the current point is predicted based on the encoded points, and this quasi-1D characteristic has already been utilized in the occupancy tree and prediction tree in G-PCC.

[0025] More precisely, the occupancy tree makes extensive use of DCM and uses a context self-adaptive entropy encoder to perform entropy encoding on the direct positions of the points within a node. Subsequently, a local transformation from the point position to the coordinates (φ, θ) and the context is obtained from the discrete angular coordinates (φ i , θ k ) obtained from the encoded points.

[0026] This quasi-1D property of the coordinate space (r, φ i , θ k) is used, and the prediction tree directly encodes the first version of the position of the current point in spherical coordinates (r, φ, θ), where r is the projection radius in the horizontal xy plane, and r on Figure 4 2D is as depicted by. Then, the spherical coordinates (r, φ, θ) are converted to 3D Cartesian coordinates (x, y, z), and the coordinate transformation error, approximation of the elevation angle and azimuth angle, and potential noise are resolved by encoding the xyz residuals.

[0027] Figure 5 shows a similar point cloud encoder of the encoder based on the G-PCC prediction tree.

[0028] First, the Cartesian coordinates (x, y, z) of the points in the point cloud are converted to spherical coordinates (r, φ, θ), where (r, φ, θ) = C2A(x, y, z).

[0029]

Number

[0030]

Number

[0031]

Number

[0032]

Number

[0033]

Number

[0034] [Number]

[0035] [Number]

[0036] [Number]

[0037] For example, the basic azimuth step size φ step or the number of detections per turn NP is encoded in the bitstream B within the geometric parameter set. Optionally, NP is a parameter of the encoder and can signal in the bitstream within the geometric parameter set, and φ step can be similarly derived from both the encoder and the decoder.

[0038] The residual spherical coordinates (r res , φ res , θ res ) can be encoded in the bitstream B.

[0039] The residual spherical coordinates (r res , φ res , θ res ) can be quantized (Q) to the quantized residual spherical coordinates Q(r res , φ res , θ res ). The quantized residual spherical coordinates Q(r res , φ res , θ res ) can be encoded in the bitstream B.

[0040] For each node of the prediction tree, φ step the prediction index n and the number m are signaled in the bitstream B, the basic azimuth step size has a fixed fixed-point precision, and is shared by all nodes of the same prediction tree.

[0041] The prediction index n refers to the predictor selected from the list of candidate predictors.

[0042] The candidate predictor PR0 may be equal to (r min , φ0, θ0), where r min is the minimum radius value (provided in the geometric parameter set), and if there is no parent node for the current node (current point P), φ0 and θ0 are equal to 0, or equal to the azimuth angle and elevation angle of the point associated with the parent node.

[0043] Another candidate predictor PR1 may be equal to (r0, φ0, θ0), where r0, φ0, and θ0 are the radius, azimuth angle, and elevation angle of the point associated with the parent node of the current node, respectively.

[0044] Another candidate predictor PR2 may be equal to the linear prediction of the radius, azimuth angle, and elevation angle using the radius, azimuth angle, and elevation angle (r0, φ0, θ0) of the point associated with the parent node of the current node and the radius, azimuth angle, and elevation angle (r1, φ1, θ1) of the point associated with the grandparent node.

[0045] For example, PR2 = 2 * (r0, φ0, θ0) - (r1, φ1, θ1)

[0046] Another candidate predictor PR3 may be equal to the linear prediction of the radius, azimuth angle, and elevation angle using the radius, azimuth angle, and elevation angle (r0, φ0, θ0) of the point associated with the parent node of the current node, the radius, azimuth angle, and elevation angle (r1, φ1, θ1) of the point associated with the grandparent node, and the radius, azimuth angle, and elevation angle (r2, φ2, θ2) of the point associated with the great-grandparent node.

[0047] For example, PR3 = (r0, φ0, θ0) + (r1, φ1, θ1) - (r2, φ2, θ2)

[0048]

Number

[0049] The decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ) may be the result of the inverse quantization (IQ) of the quantized residual spherical coordinates Q(r dec , φ res , θ res ).

[0050]

Number

[0051]

Number

[0052]

Number

[0053] If the x, y, z quantization step sizes are equal to the origin point precision (usually 1), the residual Cartesian coordinates may be reversible coding, or if the quantization step size is greater than the origin point precision (usually the quantization step size is greater than 1), it may be non-reversible coding.

[0054]

Number

[0055] The decoded Cartesian coordinates (x dec , y dec , z dec ) are available for use by the encoder, for example, the points can be sorted (decoded) before attribute coding.

[0056] Figure 6 shows a point cloud decoder similar to the prediction tree decoder based on the G-PCC prediction tree.

[0057] For each node of the prediction tree, access the prediction index n and the number m from the bitstream B, and the basic azimuth step size φ step or the number of detections NP per turn is accessed from the bitstream B (e.g., from a parameter set) and is shared by all nodes of the same prediction tree.

[0058] The decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ) can be obtained by decoding the residual spherical coordinates (r res , φ res , θ res ) from the bitstream B.

[0059] The quantized residual spherical coordinates Q(r res , φ res , θ res ) can be decoded from the bitstream B. By performing inverse quantization on the quantized residual spherical coordinates Q(r res , φ res , θ res ), the decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ) are obtained.

[0060] The decoded spherical coordinates (r dec, φ dec , θ dec ) are obtained by adding the decoded residual spherical coordinates (r res,dec , φ res,dec , θ res,dec ) and the predicted spherical coordinates (r pred , φ pred , θ pred ) according to Equation (4).

[0061] The predicted Cartesian coordinates (x pred, y pred , z pred ) is obtained by performing an inverse transformation on the decoded spherical coordinates (r dec , φ dec , θ dec ) based on Equation (3).

[0062] From the bitstream B, by decoding the quantized residual Cartesian coordinates Q(x res , y res , z res ) and performing inverse quantization, the inverse quantized Cartesian coordinates IQ(Q(x res , y res , z res )) are obtained. The decoded Cartesian coordinates (x dec , y dec , z dec ) are given by Equation (5).

[0063] The point attributes may be encoded based on the encoded Cartesian coordinates of the points to help disassociate the attribute information based on the spatial relationship / distance between the points.

[0064] In G-PCC, there are mainly two methods for disassociating and encoding the associations of point attributes. One method is represented as RAHT used in region-adaptive hierarchical transformation, and the other method is represented as LoD prediction.

[0065] RAHT encodes the attributes of points using multi - resolution transformation. RAHT successively encodes the transformed attribute values into the sub - bands of the next resolution from low resolution to maximum resolution. The multi - resolution decomposition is performed based on the encoded coordinates of the points, and each decomposition layer from the last encoded point to the first encoded point is obtained by a two - fold resolution reduction in each dimension, and each resolution is associated with the attribute encoding used for the occupancy positions of one octree geometry level (for details, see G - PCC codec description N0057 on https: / / www.mpegstandards.org / standards / MPEG - I / 9 / , January 2021).

[0066] The LoD prediction method may be used to obtain a multi - resolution representation, but the number of decomposition levels (i.e., the number of levels of detail) is parameterizable. Each level of detail is obtained by deterministically selecting a subset of points based on their encoded coordinates. The LoD prediction method obtains sub - samples of points between each layer. When using the prediction transform dissociation method, the attribute value of the currently decoded point (e.g., 3 - channel / component color, single - channel / component reflectance) is predicted using a weighted prediction of the k - nearest attribute values selected from a specified number of the last points in the same layer decoded from the attribute value and / or points in the parent layer (belonging to the points of the previously decoded layer). The number "k" is indicated in the bit - stream (in the attribute parameter set), and the weights in the weighted prediction are determined by the distance between the coordinates of the current point (in Cartesian or spherical coordinates, depending on the arrangement) and the most adjacent coordinates. To limit the complexity of the LoD prediction method, the most adjacent is limited to those belonging to the search window. More details of the method are given in the G - PCC codec description document N0057 of January 2021 (available from https: / / www.mpegstandards.org / standards / MPEG - I / 9 / ).

[0067] The LoD prediction method can use the same mechanism as the above prediction transformation, and a lifting step can be added between each decomposition layer. Therefore, it may also be represented as a lifting transformation, which has better energy compression in the lowest resolution display (i.e., in the first LoD layer). Thus, the efficiency of irreversible attribute coding is higher.

[0068] Observation shows that for the point cloud captured by the spin sensor head, instead of using Cartesian coordinates, attribute coding may benefit from using (r, φ, θ) directly obtained by predictive tree decoding (i.e., decoded spherical coordinates) or spherical coordinates (r, φ, θ) obtained (i.e., calculated) from the Cartesian coordinates decoded when using octree geometry.

[0069] The spherical coordinates (r, φ, θ) are not always represented with equal magnitudes. The radius r has values for the x, y coordinates and quantization parameters, the azimuth angle φ has values for the azimuth accuracy used in the codec, and the elevation angle θ has values for the index of the k-th elevation angle (which may be a small value compared to the other two coordinates). Therefore, to improve attribute coding, the spherical coordinates (r, φ, θ) may be scaled using the scaling coefficients signaled and notified in the bitstream before attribute coding.

[0070] In G-PCC, one scaling coefficient is encoded for each of the three spherical coordinates c k (c0 = r, c1 = φ, and c2 = θ). Each scaling coefficient is encoded by a 5-bit prefix value and a suffix value. The value obtained by adding 1 (1 to 32) to the prefix unsigned integer value (0 to 31) represents the number of bits of the suffix. The suffix is an unsigned integer value s k and corresponds to the 8-bit fixed-point representation of the scaling coefficient σ k where k = 0, 1, 2. Therefore, the scaling coefficient σ kFor s being 1.0, k the encoded value of k is 256 (the value of the encoded prefix is 7), and for encoded s k being equal to 1 (the encoded prefix is 0), the corresponding scaling coefficient is σ k = 1.0 / 256.

[0071] In G-PCC, most of the attribute encoding settings are designed to process positive coordinates. And for each point having index "i" and spherical coordinates c k,i before multiplying by the scaling coefficient, subtract the scaling offset o k from the spherical coordinates c k respectively, where k = 0, 1, 2.

[0072]

Number

[0073] For a given k (in {0, 1, 2}), the scaling offset o k is equal to the minimum value of the spherical coordinates c k,i of all the encoded points (i.e., for any "i"). Thus, since the difference (c k,i - o k ) in Equation (6) is always non - negative and the scaling coefficient s k is positive, the scaled spherical coordinates sc k,i are non - negative.

[0074] Calculating the scaling offset ok (i.e., obtaining the minimum value) requires accessing the spherical coordinates of all the encoded points in both the encoder and the decoder. Then, it is necessary to buffer the spherical coordinates of all the encoded points before encoding or decoding those attributes. This behavior causes a delay in the execution pipeline, increasing the delay and memory occupancy.

[0075] Before starting attribute encoding (decoding), removing the requirement to wait for the encoding (decoding) of the coordinates of all points in the point cloud, and without reducing the encoding (decoding) and without increasing the complexity is an issue to be solved.

Summary of the Invention

Problems to be Solved by the Invention

[0076] The following section provides a basic understanding of some aspects of the present application by presenting a simplified overview of at least one exemplary embodiment. This overview is not a detailed description of the exemplary embodiment. It is not intended to identify critical or important elements of the embodiment. The following overview only presents, in a simplified form, at least some aspects of at least one exemplary embodiment as a prelude to the more detailed description provided elsewhere in this specification.

Means for Solving the Problems

[0077] According to a first aspect of the present application, a method for encoding a point cloud into a bitstream of encoded point cloud data captured by a spin sensor head is provided. Each point in the point cloud is associated with spherical coordinates and at least one attribute. The spherical coordinates represent an azimuth angle representing the capture angle of the sensor of the spin sensor head that captured the point, an elevation angle with respect to the elevation (also called altitude or height) of the sensor that captured the point, and a radius depending on the distance from the point to a reference point. The method includes: - In the bitstream, transmitting a signal to notify scaling offset information representing a scaling offset; For each current point in the point cloud, - Encoding the spherical coordinates of the current point and adding the encoded spherical coordinates of the current point to the bitstream; - Obtaining decoded spherical coordinates by decoding the encoded spherical coordinates of the current point; - Scaling the decoded spherical coordinates based on the scaling offset; -Encoding at least one attribute of the current point based on the scaled and decoded spherical coordinates, and adding the at least one encoded attribute to the bit stream.

[0078] According to a second aspect of the present application, a method for decoding a point cloud from a bit stream of encoded point cloud data captured by a spin sensor head is provided. The method includes: -Accessing the scaling offset information from the bit stream; For each current point of the point cloud: -Obtaining the decoded spherical coordinates by decoding the encoded spherical coordinates of the current point obtained from the bit stream, where the spherical coordinates of the current point represent the azimuth angle representing the capture angle of the sensor of the spin sensor head that captured the current point, the elevation angle with respect to the elevation of the sensor that captured the current point, and the radius depending on the distance from the current point to the reference point. -Scaling the decoded spherical coordinates based on the scaling offset obtained from the scaling offset information; -Decoding the attributes of the current point based on the scaled and decoded spherical coordinates.

[0079] In an exemplary embodiment, the scaling offset is determined to ensure that the scaled spherical coordinates are non - negative and / or within a specific range.

[0080] In an exemplary embodiment, the scaling offset is determined to be smaller than the minimum value calculated from the decoded spherical coordinates of all points of the point cloud.

[0081] In an exemplary embodiment, the scaling offset is determined based on the sensor and capture characteristics of the sensor spin head.

[0082] In an exemplary embodiment, the scaling offset information includes the scaling offset.

[0083] In an exemplary embodiment, the scaling offset information instructs to determine the scaling offset using a specific method.

[0084] In an exemplary embodiment, the specific method uses clipping or modulo operation.

[0085] According to a third aspect of the present application, there is provided an apparatus for encoding a point cloud into a bitstream of encoded point cloud data captured by a spin sensor head. The apparatus includes one or more processors configured to execute the method according to the first aspect of the present application.

[0086] According to a fourth aspect of the present application, there is provided an apparatus for decoding points of a point cloud captured by a spin sensor head from a bitstream. The apparatus includes one or more processors configured to execute the method according to the second aspect of the present application.

[0087] According to a fifth aspect of the present application , Co computer program mmunity is provided, and when the program is executed by one or more processors, the one or more processors are caused to execute the method according to the first aspect of the present application.

[0088] According to a sixth aspect of the present application , Co computer program mmunity is provided, and when the program is executed by one or more processors, the one or more processors are caused to execute the method according to the second aspect of the present application.

[0089] According to a seventh aspect of the present application, there is provided a non-transitory storage medium carrying program code instructions for executing the method according to the first aspect of the present application.

[0090] According to an eighth aspect of the present application, a non-transitory storage medium is provided, and the non-transitory storage medium carries instructions of program code for executing the method according to the second aspect of the present application.

[0091] According to a ninth aspect of the present application, a bitstream of encoded point cloud data captured by a spin sensor head is provided, and the bitstream includes at least one syntax element carrying scaling offset information, and the scaling offset information represents a scaling offset for scaling the decoded spherical coordinates of the points of the point cloud represented by the encoded point cloud data.

[0092] At least one specific property in the exemplary embodiments and at least one other object, advantage, feature, and use in the exemplary embodiments will become apparent from the description of the examples in combination with the following drawings.

Brief Description of the Drawings

[0093] Currently, reference is made as an example to the drawings of the exemplary embodiments of the present application.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

[0094] In different drawings, it is possible to represent similar components with similar reference numerals.

DETAILED DESCRIPTION OF THE INVENTION

[0095] Hereinafter, at least one of the exemplary embodiments will be described with reference to the drawings, where an example of at least one of the exemplary embodiments is shown. However, the exemplary embodiments can be implemented in many alternative forms and should not be understood as limiting the examples described herein. Therefore, it should be understood that the exemplary embodiments should not be limited to the specific forms disclosed. Rather, the present disclosure aims to cover all modifications, equivalents, and alternative solutions within the spirit and scope of the present application.

[0096] It should be understood that when the drawings are shown in the form of a flowchart, a block diagram of the corresponding device is also provided. Similarly, when the drawings are shown in the form of a block diagram, a flowchart of the corresponding method / process is also provided.

[0097] At least one of these aspects generally relates to point cloud encoding and decoding, and at least one other aspect generally relates to the transmission of a generated or encoded bitstream.

[0098] Moreover, this aspect is applicable not only to MPEG standards such as MPEG-I Part 5 or Part 9 related to point cloud compression, but also to other standards and recommendations, such as those that already exist, those that have not yet been developed, and extensions of any such standards and recommendations (including MPEG-I Part 5 and Part 9). Unless otherwise indicated or technically excluded, the aspects described in this application can be used alone or in combination.

[0099] The present invention relates to a method and apparatus for encoding a point cloud into a bitstream of encoded point cloud data captured by a spin sensor head / decoding a point cloud from a bitstream of encoded point cloud data captured by a spin sensor head. Each point of the point cloud is associated with spherical coordinates and attributes. The method includes transmitting a signal in the bitstream to notify scaling offset information representing a scaling offset, encoding / decoding the spherical coordinates of the current point of the point cloud for each current point, obtaining the decoded spherical coordinates of the current point from the encoded spherical coordinates, scaling the decoded spherical coordinates of the current point using the scaling offset, and encoding / decoding at least one attribute of the current point based on the scaled decoded spherical coordinates.

[0100] By transmitting a signal in the bitstream to notify such scaling offset information, it avoids calculating the scaling offset in the decoder and removes the requirement to wait for the decoding of the coordinates of all points before starting attribute decoding, without affecting the decoding efficiency.

[0101] The present invention further reduces memory occupancy by removing the buffer of all coordinates of the decoded points before decoding the attributes of the decoded points.

[0102] FIG. 7 is a block diagram of steps of a method 100 for encoding the attributes of points of a point cloud captured by a spin sensor head according to at least one exemplary embodiment.

[0103] In step 110, a scaling offset o k is determined.

[0104] In step 120, a signal is transmitted in a bit stream B to notify scaling offset information representing the scaling offset o k .

[0105] For each current point of the point cloud, in step 130, the spherical coordinates of the current point are encoded and added to the bit stream B.

[0106] In step 140, by decoding the encoded spherical coordinates, decoded spherical coordinates are obtained.

[0107] In step 150, based on the scaling offset o k , for example using equation (6), the decoded spherical coordinates of the current point are scaled.

[0108] In step 160, at least one attribute of the current point is encoded based on the scaled decoded spherical coordinates, and the at least one encoded attribute is added to the bit stream B.

[0109] Such a method does not require buffering the spherical coordinates of all points in the point cloud before encoding the attributes of the points. This method does not increase the complexity of encoding without reducing the encoding efficiency.

[0110] The scaling offset information may be any information that can determine the scaling offset o k before complete decoding of the spherical coordinates of all encoded points of the point cloud.

[0111] The scaling offset information can be provided from a part of the bitstream that has been read / accessed / acquired / decoded before the spherical coordinates of the encoded points have been reconstructed / decoded yet.

[0112] In an exemplary embodiment, a signal is transmitted in the attribute parameter set of the G-PCC to notify the scaling offset information.

[0113] In an exemplary embodiment, a signal is transmitted in the geometry parameter set of the G-PCC to notify the scaling offset information.

[0114] In an exemplary embodiment, a signal is transmitted in the attribute block header of the G-PCC to notify the scaling offset information.

[0115] In an exemplary embodiment, a signal is transmitted in the geometry block header of the G-PCC to notify the scaling offset information.

[0116] In one exemplary embodiment, to ensure that the spherical coordinate s of the scaling k,i is non-negative and / or within a specific range, the scaling offset o k is determined.

[0117] In a first variant, the encoder determines, in the encoder, the scaling offset o k as the minimum value m for any k calculated from the decoded spherical coordinates c of all the points of the point cloud k,i k

[0118] This variant is advantageous because it removes the requirement to wait for the decoder to decode the spherical coordinates of all the points of the point cloud on the decoder side before starting the attribute decoding.

[0119] However, the encoder still kIt is necessary to wait for and buffer the coordinates of all points before being able to determine and before being able to start attribute encoding.

[0120] In another variant, for any k, the encoder determines that any scaling offset o k is less than the minimum value m k is determined to be less than.

[0121] In one exemplary embodiment of said another variant, the scaling offset o k is determined by the sensor and the capture characteristics.

[0122] This exemplary embodiment is advantageous because the requirement of waiting to encode / decode the spherical coordinates of all points in the point cloud before starting attribute encoding / decoding on the encoding side and the decoding side is removed.

[0123] For example, the sensor and capture characteristics are the binary values indicating whether the sensor spin head has completed a full spin and / or the sensor characteristics itself in order to capture the frame of the point cloud. For example, geometric encoding settings such as the accuracy of the point coordinates or the quantization step size may be used to determine the scaling offset o k may be used.

[0124] As a first example, when the azimuth φ value is included within the range [-2N - 2; 2N - 2] corresponding to the valid azimuth rotation within the range [-π / 2; π / 2], the lidar sensor is attached to the front of the vehicle and is configured to acquire the previous point, corresponding to the acquisition in the hemisphere (see Equation (1)). Thereafter, the encoder may select o1 (for the case of azimuth coordinate k = 1) to be the lower limit of the azimuth acquisition range, for example, -2 N-2 , or a lower value, for example, the lowest value (or lower than the lowest value) that Φ reaches in the case of non - reversible encoding.

[0125] As a second example, the encoder is configured not to encode any point closer to the sensor than the front, rear, left, and right sides of the vehicle to which the sensor is attached. The minimum radius can be determined by the distance from the sensor to the closest side of the vehicle and is used to determine o0 (when targeting the radius coordinate k = 0). This minimum radius is calculated based on the physical distance from the sensor to the closest side of the vehicle, and the encoder setting information (e.g., the quantization step of the radius) and the spatial accuracy of the point cloud in the physical Cartesian space (e.g., a translation of x = 1, y = 1, and z = 1 in the point cloud space corresponds to a translation of 1 mm in x, 1 mm in y, and 1 mm in z in the real / physical space).

[0126] As a third example, the encoder is configured not to encode the "n" points at the first elevation angle from the lidar sensor (e.g., they are encoded in another G-PCC slice and a signal is sent in the block header to notify the scaling offset). Subsequently, the minimum elevation angle index o2 (when targeting the elevation coordinate k = 2) can be determined as "n".

[0127] In one exemplary embodiment, the scaling offset information includes the scaling offset o k and.

[0128] In one exemplary embodiment, the scaling offset information instructs to determine the scaling offset o k using a specific method.

[0129] In one exemplary embodiment, the scaling offset information includes the geometry encoding setting information.

[0130] In a variant, a signal is sent in the geometry parameter set of the G-PCC to notify the geometry encoding setting information.

[0131]

Number

[0132] However, in the prediction tree of G-PCC, there is nothing that enforces or guarantees that this constraint is checked by a G-PCC compliant bitstream. The G-PCC encoder may be carefully designed to generate only azimuth integer representations within the [-b1;b1] range, but from the perspective of decoding, there is no guarantee that the encoder will consider this constraint when generating a G-PCC compliant bitstream.

[0133] To solve this problem, in the variant, the geometry encoding configuration information indicates that the decoded azimuth c 1,i belongs to the range [-b1,b1] and the scaling offset o1 is determined to be o1 = -b1, so that the scaled decoded azimuth sc 1,i is non-negative.

[0134] This variant is advantageous because it does not require modification to the decoding of the point coordinates.

[0135] However, since there may be errors in the encoder that cannot be accommodated by the bitstream, forcing the generation of the bitstream to check this constraint is not the best solution.

[0136] In the variant, the geometric decoding is modified so that even if the difference (c k,i - o k ) is negative, a non-negatively scaled decoded azimuth sc 1,i is obtained.

[0137] In the variant, the decoded azimuth c 1,i is restricted to check whether the decoded azimuth belongs to the range [-b1;b1]. Before calculating the scaled azimuth, if the decoded azimuth c 1,i is lower than -b1, the decoded azimuth c 1,i is set to be equal to -b1, or if the decoded azimuth c 1,i is higher than b1, the decoded azimuth c1,i is set to be equal to b1.

[0138] In another variant, if the decoded azimuth angle c 1,i is higher than b1, since it does not interfere with the decoding of the attribute, the value of the decoded azimuth angle is not modified / limited.

[0139]

Number

[0140] The negatively scaled decoded azimuth angle sc 1,i is limited to zero, and the scaling offset o1 is determined to be o1 = -b1.

[0141] In a variant, the encoder performs a modulo operation (or an equivalent operation) on the decoded azimuth angle c 1,i to limit the decoded azimuth angle c 1,i to belong to the range [-b1, b1] (or equivalent to the range [-π, π]). If the decoded azimuth angle c 1,i is outside [-b1, b1], the decoded azimuth angle is set to be equal to c 1,i modulo b1. For example, if the decoded azimuth angle c 1,i is lower than -b1, the decoded azimuth angle is set to be equal to c 1,i +(1 + 2*b1), and if c 1,i is higher than b1, c 1,i is set to be equal to c 1,i -(1 + 2*b1). This iterative process is preferred because it is less costly in terms of CPU operation time than using a "true" modulo operation.

[0142]

Number

[0143]

Number

[0144] The only difference is that when using Equation (9), c 1,i is strictly lower than b1.

[0145]

Number

[0146] During the geometric decoding process, a restriction or modulo operation may be performed so that the corrected azimuth angle is taken into account in the prediction of subsequent decoded points.

[0147] Optionally, for example as described above, a restriction or modulo operation may be performed during the scaling of the coordinates, after geometric decoding and before attribute decoding.

[0148] In a variant, a scalable scaling offset (which can be completed in a situation where points are not buffered if necessary, for example executed in the previous frame) adjustable by the encoder may be used to transmit and notify a signal in the attribute parameter set, but (c k,i -o k ) cannot always be guaranteed to be zero or more. Then, instead of using Equation (6), the scaling coefficient is determined in a way that ensures that the scaled coefficient sc k,i is zero or more, for example using Equation (7), or a modulo operation, for example using Equation (10).

[0149] FIG. 8 is a block diagram of the steps of a method 200 for decoding the attributes of points of a point cloud captured by at least one exemplary spin sensor head.

[0150] In step 210, access the scaling offset information from the bit stream B.

[0151] For each of the current points of the point cloud, in step 220, based on the encoded spherical coordinates of the current point obtained from the bitstream B, the decoded spherical coordinates are obtained, and the spherical coordinates of the current point represent the azimuth angle representing the capture angle of the sensor of the spin sensor head that captured the current point, the elevation angle with respect to the elevation of the sensor that captured the current point, and the radius that depends on the distance from the current point to the reference point.

[0152] In step 230, for example, using Equation (6), the scaling offset o k obtained from the scaling offset information is used to scale the decoded spherical coordinates.

[0153] In step 240, based on the scaled decoded spherical coordinates, the (one or more) attributes of the points of the point cloud are decoded.

[0154] The bitstream corresponding to G-PCC should include a plurality of setting information for decoding the encoded point cloud data.

[0155] Some of the current syntax elements of the bitstream B corresponding to G-PCC should have specific values and require some new syntax elements to implement the present invention.

[0156] Therefore, in one exemplary embodiment, the scaling offset information includes geometric encoding and attribute encoding setting information that instructs to determine the scaling offset o k using a specific method.

[0157] For example, regarding the conventional syntax elements that constitute the geometric coding and attribute coding setting information, in the geometry parameter set, by setting the conventional syntax element "geom_tree_type" to 1, it is indicated to perform geometric coding using a prediction tree. When the "geometry_angular_enabled_flag" of the conventional syntax element is set to 1 in the geometry parameter set, it is indicated that the prediction tree uses spherical coordinates in the prediction proposal. In the attribute parameter set, by setting the "attr_coding_type / attr_encoding" of the conventional syntax element to 0, it is indicated to perform point attribute coding using LoD prediction conversion. By setting the conventional syntax elements "lod_scalability_enabled_flag / scalable_lifting_enabled_flag" to 0, it is indicated not to use a scalable representation. By setting the conventional syntax elements "max_num_detail_levels_minus1 / num_detail_levels_minus1" to 0, it is indicated to use single-layer detail. By setting the conventional syntax elements "morton_sort_skip_enabled_flag / canonical_point_order_flag" to 1, it is indicated that the coding order of the point attributes is the same as the order of its geometric coding. By setting the conventional syntax element "aps_coord_conv_flag / spherical_coord_flag" to 1, it is indicated to code the point attributes using spherical coordinates, which provides better compression performance for the point attributes and indicates to transmit signals in the bitstream to notify the scaling coefficients.

[0158] According to the present invention, by adding a new syntax element to the bitstream, signals are transmitted to notify the scaling offset information.

[0159] Transmitting a signal to notify scaling offset information indicates that a specific extension mechanism is used for point attribute coding, and the attribute coding parameters are notified by transmitting a signal in the bitstream (e.g., in the extended part of the attribute parameter set of the G-PCC compliant bitstream).

[0160] In one exemplary embodiment, the new syntax element for transmitting a signal to notify scaling offset information includes binary information (e.g., represented as a flag of "attr_coord_conv_scale_fixed_offset_flag") and a scaling offset. By setting the binary information to 1, it effectively indicates transmitting a signal in the subsequent bitstream to notify the scaling offset o k and information representing the scaling offset follows the binary information in the bitstream.

[0161] The decoding method 200 decodes binary information (e.g., the flag "attr_coord_conv_scale_fixed_offset_flag") from the bitstream. If the binary information is equal to 1, the scaling offset o k is obtained by decoding the information representing the scaling offset o from the bitstream. k If the binary information is equal to 0, the decoding method 200 is not enabled.

[0162] In one exemplary embodiment, the new syntax element for transmitting a signal to notify scaling offset information includes binary information, represented as a flag of "attr_coord_conv_scale_fixed_offset_flag" for example, and by setting it to 1, it indicates that the scaling offsets (o1, o1, o2) are (0, -(2 geom_angular_azimuth_scale_log2_minus11+11-1 ), 0) respectively. Setting the binary information to 1 further indicates scaling the decoded spherical coordinates of the point using Equation (7) instead of Equation (6).

[0163] The decoding method 200 decodes binary information (e.g., "attr_coord_conv_scale_fixed_offset_flag") from a bit stream. When the binary information is equal to 1, the scaling offsets (o0, o1, o2) are set to (0, -(2 geom_angular_azimuth_scale_log2 _minus11+11 -1 ), 0) respectively, and instead of equation (6), equation (7) is used to scale the spherical coordinates of the decoded points. When the binary information is equal to 0, the decoding method 200 is not effective.

[0164] In one variation, the attribute coding configuration information (and optionally the geometry coding configuration information) indicates that the attribute coding / decoding method can support the use of negative geometric coordinates (and thus negative values of the scaled decoded spherical coordinates). For example, using more examples than sending a signal to notify a conventional syntax element, in one exemplary embodiment, a signal is sent to notify that a new syntax element of scaling offset information includes binary information, for example, is indicated in a flag of "attr_coord_conv_scale_fixed_offset_flag", and when it is set to 1, it indicates that the scaling offsets (o1, o1, o2) are equal to (0, 0, 0) respectively.

[0165] FIG. 9 shows an exemplary schematic block diagram of a system of each aspect and an exemplary embodiment.

[0166] The system 300 may be incorporated into one or more devices and includes each component described below. In various embodiments, the system 300 may be configured to implement one or more aspects described in this application.

[0167] Examples of devices that can constitute all or part of system 300 include personal computers, notebook computers, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, set-top boxes, point clouds, any other device for processing video or images, or other communication devices. The elements of system 300 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 300 can be distributed across multiple ICs and / or discrete components. In various embodiments, system 300 can be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or dedicated input and / or output ports.

[0168] System 300 can include at least one processor 310, and the at least one processor 310 is configured to implement each aspect described in this application by executing instructions to be loaded. The processor 310 can include an embedded memory, an input / output interface, and various other circuits known in the art. System 300 can include at least one memory 320 (e.g., a volatile memory device and / or a non-volatile memory device). System 300 can include a storage device 340, which can include non-volatile memory and / or volatile memory, and can include, but is not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drive, and / or optical disk drive. As a non-limiting example, the storage device 340 can include an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0169] System 300 can include an encoder / decoder module 330 configured to provide encoded / decoded point cloud geometry data, for example, by processing data, and the encoder / decoder module 330 includes its own processor and memory. The encoder / decoder module 330 can represent one or more modules that can be included in a device to perform encoding and / or decoding functions. As is known, a device can include either or both of an encoding and a decoding module. Also, the encoder / decoder module 330 can be implemented as an independent element of System 300 or can be coupled within the processor 310 as a combination of hardware and software known to those skilled in the art.

[0170] The program code loaded into the processor 310 or the encoder / decoder 330 to execute each aspect described in the present application can be stored in the storage device 340, and then loaded into the memory 320 and executed by the processor 310. According to various embodiments, during the execution of the process described in the present application, one or more of the processor 310, the memory 320, the storage device 340, and the encoder / decoder module 330 can store one or more of various items. Such stored items can include, but are not limited to, point cloud frames, encoded / decoded geometric shape / attribute videos / images or parts thereof, bitstreams, matrices, variables, and expressions, formulas, intermediate or final results of operations and arithmetic logic processing.

[0171] In some embodiments, the memory inside the processor 310 and / or the encoder / decoder module 330 can be used to store instructions and provide a working memory for the processes executed during encoding or decoding.

[0172] However, in other embodiments, a memory external to the processing device (for example, the processing device may be the processor 310 or the encoder / decoder module 330) is used for one or more of these functions. The external memory may be the memory 320 and / or the storage device 340, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM can be used as a working memory for video encoding and decoding operations, for example, for HEVC (Efficient Video Coding), VVC (Versatile Video Coding), or MPEG-I Part 3 or Part 9 with respect to MPEG-2 Part 2 (also called ITU-T Recommendation H.262 and ISO / IEC 13818-2, and also called MPEG-2 video).

[0173] As directed by block 390, inputs to the elements of system 300 can be provided via various input devices. Such input devices can include, but are not limited to, (i) an RF portion that can receive an RF signal wirelessly transmitted, for example, by a broadcast device, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0174] In various embodiments, as is known in the art, the input devices of block 390 have corresponding input processing elements associated therewith. For example, the RF portion may be associated with the following necessary elements: (i) selecting a desired frequency (also called signal selection or restricting a signal frequency band within a frequency band), (ii) a signal selected by downconversion, (iii) selecting a signal frequency band, which in certain embodiments is called a channel, by restricting the frequency band again to a narrow frequency band (for example), (iv) demodulating the downconverted signal and the signal with the restricted frequency band, (v) performing debugging, and (vi) selecting a desired packet stream by demultiplexing. The RF portion of various embodiments can include one or more elements that perform these functions, such as a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a downconverter, a demodulator, a debugging device, and a demultiplexer. The RF portion can include a tuner that performs each of these functions, and these functions include downconverting the received signal to a lower frequency (for example, an intermediate frequency or a frequency near the baseband) or to the baseband.

[0175] In one set-top box embodiment, the RF portion and its associated input processing elements can receive an RF signal transmitted over a wired (for example, cable) medium. The RF portion can then perform frequency selection by filtering, downconverting, and re-filtering to a desired frequency band.

[0176] Various embodiments rearrange the order of the above (and other) elements, delete some of these elements, and / or add other elements that perform similar or different functions.

[0177] Adding elements may involve inserting elements between conventional elements, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.

[0178] Also, the USB and / or HDMI terminals can include corresponding interface processors and are used to connect the system 300 to other electronic devices via USB and / or HDMI connections. Note that, when necessary, each aspect of the input processing (e.g., Reed-Solomon debugging) can be implemented, for example, within a separate input processing IC or within the processor 310. Similarly, when necessary, each aspect of the USB or HDMI interface processing can be implemented within a separate interface IC or within the processor 310. The demodulated, debugged, and de-multiplexed streams can be provided to various processing elements, for example, including the processor 310 and the encoder / decoder 330, which operate in combination with memory and storage elements, thereby processing the data stream when necessary for display at the output device.

[0179] It is to provide various elements of the system 300 within an integrated housing. Within the integrated housing, an appropriate connection arrangement 390 can be made, for example, connecting various elements and transmitting data between them via internal buses (including I2C buses), wiring, and printed circuit boards known in the art.

[0180] System 300 can include a communication interface 350, so that it can communicate with other devices via a communication channel 700. The communication interface 350 includes, but is not limited to, a transceiver that transmits and receives data on the communication channel 700. The communication interface 350 includes, but is not limited to, a modem or a network card, and the communication channel 700 can be implemented, for example, in a wired and / or wireless medium.

[0181] In various embodiments, data can be streamed to System 300 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals in these embodiments can be received by the communication interface 350 and a communication channel 700 that is compatible with Wi-Fi communication. The communication channel 700 in these embodiments can typically be connected to an access point or a router, and the access point or the router provides access to an external network including the Internet and enables streaming applications and other over-the-top communications.

[0182] Other embodiments can provide stream data to System 300 using a set-top box, and the set-top box transmits data via the HDMI connection of the input block 390.

[0183] Other embodiments can also provide stream data to System 300 using the RF of the input block 390.

[0184] The stream data can be used as a signaling information scheme used by System 300. The signaling information can include information such as a bitstream B and / or scaled offset information.

[0185] Note that various signaling can be realized. For example, in various embodiments, one or more syntax elements, flags, etc. may be used to transmit signaling information to the corresponding decoder.

[0186] System 300 can provide output signals to various output devices including a display 400, a speaker 500, and other peripheral devices 600. In various examples of the embodiments, the other peripheral devices 600 can include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices based on the output providing function of the system 300.

[0187] In various embodiments, the control signal can be used to communicate between the system 300 and the display 400, the speaker 500, or the other peripheral devices 600 using signaling of other communication protocols that enable control between devices, such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or the like, with or without user intervention.

[0188] The output devices are coupled to the system 300 via dedicated connections so as to be communicable through corresponding interfaces 360, 370, and 380.

[0189] Optionally, the output devices can be connected to the system 300 using a communication channel 700 via a communication interface 350. The display 400 and the speaker 500 can be integrated into a single unit together with other components of the system 300 within an electronic device (e.g., a television).

[0190] In various embodiments, the display interface 360 can include a display driver such as, for example, a timing controller (T Con) chip.

[0191] For example, if the RF portion of the input terminal 390 is part of a separate set-top box, the display 400 and the speaker 500 can be optionally separated from one or more of the other components. In various embodiments where the display 400 and the speaker 500 may be external components, output signals can be provided via dedicated output connections (including, for example, an HDMI port, a USB port, or a COMP output terminal).

[0192] In FIGS. 1-9, various methods are described herein, and each method includes one or more steps or operations to implement the described method. The exact operation of the method does not require a specific order of steps or operations, and the order and / or use of specific steps and / or operations can be modified or combined, as long as a specific order of steps or operations is not required.

[0193] Several examples of block diagrams and / or operation flowcharts are described. Each block represents a circuit element, module, or portion of code that includes one or more executable instructions for implementing the specified (one or more) logical functions. It should be understood that in other embodiments, the (one or more) functions marked within a block may be different from the indicated order. For example, depending on the functions involved, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may be executed in reverse order.

[0194] For example, in a method or process, apparatus, computer program, data stream, bit stream, or signal, the embodiments and aspects described herein can be implemented. Even when discussing only in the context of a single form of embodiment (e.g., discussing only as a method), the embodiments of the features discussed can be implemented in other forms (e.g., an apparatus or a computer program).

[0195] The method can be implemented, for example, in a processor, which generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor further includes a communication device.

[0196] Also, the method can be implemented by instructions executable by a processor, and such instructions (and / or data values generated by the embodiments) can be stored in a computer-readable storage medium. The computer-readable storage medium can use the form of a computer-readable program product having computer-readable program code that is implemented in and executable by one or more computer-readable media by a computer. Considering the inherent ability to store information and the inherent ability to provide information retrieval, the computer-readable storage medium used herein can be regarded as a non-transitory storage medium. The computer-readable storage medium may be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Although more specific examples to which the computer-readable storage medium of this embodiment is applied are provided below, it is merely illustrative and not a detailed list as can be easily understood by those skilled in the art: floppy disks, hard disks, read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any combination of the foregoing.

[0197] The instructions can form an application program tangibly implemented on a processor-readable medium.

[0198] For example, instructions can exist in hardware, firmware, software, or a combination. For example, instructions can be found in an operating system, an independent application, or a combination of both. Thus, a processor can be characterized as a device configured to execute a process, for example, and a device having a processor-readable medium (e.g., a storage device) for executing the instructions of the process. Also, in addition to or instead of instructions, the processor-readable medium can store data values generated by an embodiment.

[0199] The apparatus can be realized, for example, in suitable hardware, software, and firmware. Examples of such apparatus include personal computers, notebook computers, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, point clouds, videos, or any other device for processing images, or other communication devices. Note that the apparatus is movable and can also be attached to a moving vehicle.

[0200] Computer software can be implemented by a processor 510, hardware, or a combination of hardware and software. As a non-limiting example, embodiments can be implemented by one or more integrated circuits. Memory 520 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology (including, as non-limiting examples, for example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory). Processor 510 can be of any type suitable for the technical environment and can include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture).

[0201] It is obvious to those skilled in the art that embodiments can generate signals formatted to carry, for example, information that can be stored or transmitted. The information can include, for example, instructions for executing a method or data generated by one of the described embodiments. For example, the signal can be formatted to carry a bitstream of the described example. Such a signal can be formatted, for example, as an electromagnetic wave (using, for example, the radio frequency part of the spectrum) or as a baseband signal. The formatting includes, for example, encoding a data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted over various different wired or wireless links. The signal can be stored on a processor-readable medium.

[0202] The terms used in this specification are for the purpose of describing particular embodiments and are not intended to be limiting. Unless explicitly indicated otherwise in the context, the singular forms "a", "an" and "the" used in this specification also include the plural forms. Further, as used in this specification, the terms "include" and / or "comprise" and / or "including" and / or "comprising" can specify the presence of, for example, the described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. And when an element is referred to as "responding" or "connecting" to another element, it can directly respond to or connect to the other element, or there may be intermediate elements. Conversely, when an element is referred to as "directly responding" or "directly connecting" to another element, there are no intermediate elements.

[0203] Note that, for example, in the cases of "A / B", "A and / or B" and "at least one of A and B", the use of any one of the symbols / terms " / ", "and / or" and "at least one of" is intended to include only the selection of the first alternative (A) listed, or only the selection of the second alternative (B) listed, or the selection of both alternatives (A and B). As a further example, in the case of "A, B and / or C" and "at least one of A and B", such expressions are intended to include only the selection of the first alternative (A) listed, or the selection of the second alternative (B) listed, or only the selection of the third alternative (C) listed, or only the selection of the first and second alternatives (A and B) listed, or only the selection of the first and third alternatives (A and C) listed, or only the selection of the second and third alternatives (B and C) listed, or all selections of the three alternatives (A, B and C). As will be apparent to those skilled in the art, this can be extended to the same number of items as those listed.

[0204] In the present application, various numerical values can be used. The specific values can be used for illustrative purposes, and each of the described aspects is not limited to these specific values.

[0205] In addition, terms such as first, second, etc. can be used in this specification to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, unless departing from the teachings of the present application, the first element can also be called the second element, and similarly, the second element can also be called the first element. No order is implied between the first element and the second element.

[0206] References to "exemplary embodiment" or "exemplary examples" or "one embodiment" or "embodiment" and other variations are frequently used to indicate that certain features, structures, characteristics, etc. (described in accordance with the example / embodiment) may be included in at least one example / embodiment. Therefore, the appearance of the terms "in an exemplary embodiment" or "in an exemplary example" or "in one embodiment" or "in an embodiment" and any other variations that appear throughout the present application do not necessarily refer to the same embodiment.

[0207] Similarly, this specification uses references to "according to an exemplary embodiment / example / embodiment" or "in an exemplary embodiment / example / embodiment" and other variations frequently to indicate that certain features, structures or characteristics (described in combination with the exemplary embodiment / example / embodiment) may be included in at least one exemplary embodiment / example / embodiment. Therefore, the descriptions "according to an exemplary embodiment / example / embodiment" or "in an exemplary embodiment / example / embodiment" that appear throughout the present application do not necessarily refer to the same exemplary embodiment / example / embodiment, and independent or alternative exemplary embodiments / examples / embodiments are not necessarily mutually exclusive of other exemplary embodiments / examples / embodiments.

[0208] The reference signs in the claims are for illustrative purposes only and do not limit the scope of the claims. Although not explicitly stated, any combination or partial combination of the embodiments / examples and variations can be adopted.

[0209] When the drawings are represented as flowcharts, block diagrams of the corresponding devices are provided. Similarly, it should be understood that when the drawings are represented as block diagrams, flowcharts of the corresponding methods / processes are also provided.

[0210] Some of the drawings include arrows on the path to indicate the main direction of communication, but it should be understood that communication may occur in the direction opposite to that depicted.

[0211] The various embodiments are related to decoding. As used in this application, "decoding" can include, for example, all or part of the process performed on a received point cloud frame (which may include a received bitstream that has been encoded for one or more point cloud frames), thereby generating a final output that is suitable for further processing in the displayed or reconstructed point cloud region. In various examples, this type of process includes one or more of the processes typically performed by a decoder. In various examples, for example, this type of process can optionally further include the processes performed by the decoders of the various embodiments described in this application.

[0212] As a further example, in one embodiment, "decoding" can refer to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of entropy decoding and differential decoding. Based on the specifically described context, it is clear whether the term "decoding process" specifically refers to a subset of operations or generally to a broader decoding process, and it is recognized that this is well understood by those skilled in the art.

[0213] All embodiments relate to encoding. Similar to the mechanism of "decoding" described above, "encoding" as used in this application can include all or part of a process that, for example, generates an encoded bitstream by performing on an input point cloud frame. In various examples, this type of process can include one or more of the processes typically performed by an encoder. In various examples, this type of process can include or alternatively include the processes performed by the encoders of the various embodiments described in this application.

[0214] As a further example, in one embodiment, "encoding" can refer only to entropy encoding, in another embodiment, "encoding" can refer only to differential encoding, and in another embodiment, "encoding" can refer to a combination of differential encoding and entropy encoding. Based on the limited context described, it is clearly whether the term "decoding process" specifically refers to a subset of operations or generally refers to a broader decoding process, and it is recognized that this is well understood by those skilled in the art.

[0215] Also, in this application, the "determination" of various information is mentioned. The determination of information can include one or more of, for example, the estimation of information, the calculation of information, the prediction of information, or the retrieval of information from memory.

[0216] Also, in this application, the "access" to various information is mentioned. Access to information can include one or more of the reception of information, the retrieval of information (e.g., from memory or a bitstream), the storage of information, the transfer of information, the copying of information, the calculation of information, the determination of information, the prediction of information, or the estimation of information.

[0217] Also, "receiving" of various information was mentioned. Similar to "access", receiving is a broad term. Receiving of information can include, for example, one or more of access to information or retrieval of information (e.g., from a memory or a bitstream). Also, as one or another form, the operation periods such as storage of information, processing of information, transmission of information, movement of information, copying of information, deletion of information, calculation of information, determination of information, prediction of information, or estimation of information are usually related to "receiving".

[0218] And the term "signal" as used herein specifically refers to a corresponding decoder and indicates something etc. For example, in some embodiments, the encoder transmits a signal to notify specific information such as, for example, scaled offset information. In this way, in the embodiments, the same parameters can be used on the encoder side and the decoder side. Thus, for example, the encoder can transmit specific parameters to the decoder (explicit signaling), whereby the decoder can use the same specific parameters. Conversely, when the decoder has specific parameters and other parameters, signaling (using signaling without the need for implicit signaling) allows the decoder to know and select the specific parameters. By avoiding any actual function transmission, bit savings are achieved in each embodiment. Note that signaling can be completed in many ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to transmit a signal to notify the corresponding decoder of information. Although the verb form of "signal" was mentioned in the foregoing, the term "signal" can also be used as a noun herein.

[0219] Multiple embodiments have already been described. However, it should be understood that various modifications are possible. For example, other embodiments can be generated by combining, supplementing, modifying, or removing elements of different embodiments. Also, those skilled in the art will recognize that the disclosed structures and processes can be replaced with other structures and processes, and the resulting embodiments will perform at least substantially the same (one or more) functions in at least substantially the same (one or more) ways, thereby achieving at least substantially the same (one or more) results as the disclosed embodiments. Accordingly, the present application contemplates these and other embodiments.

Claims

1. A method for encoding a point cloud captured by a spin sensor head into a bitstream of encoded point cloud data, wherein each point of the point cloud is associated with spherical coordinates and at least one attribute, and the spherical coordinates represent an azimuth angle representing the capture angle of a sensor of the spin sensor head that captured the point, an elevation angle with respect to the elevation of the sensor that captured the point, and a radius depending on the distance from the point to a reference point, and the method comprises: - In the bitstream, transmitting a signal to notify scaling offset information representing a scaling offset, the scaling offset being determined based on the sensor of the spin sensor head and capture characteristics; For each current point of the point cloud, - Encoding the spherical coordinates of the current point and adding the encoded spherical coordinates of the current point to the bitstream; - Obtaining decoded spherical coordinates by decoding the encoded spherical coordinates of the current point; - Scaling the decoded spherical coordinates based on the scaling offset; - Encoding at least one attribute of the current point based on the scaled decoded spherical coordinates and adding at least one encoded attribute to the bitstream. A method for encoding a point cloud captured by a spin sensor head into a bitstream of encoded point cloud data.

2. A method for decoding a point cloud from a bitstream of encoded point cloud data captured by a spin sensor head, the method comprising: - Accessing scaling offset information representing a scaling offset from the bitstream, the scaling offset being determined based on the sensor of the spin sensor head and capture characteristics; For each current point of the point cloud, - Obtaining the decoded spherical coordinates by decoding the encoded spherical coordinates of the current point obtained from the bitstream, wherein the spherical coordinates of the current point represent the azimuth angle representing the capture angle of the sensor of the spin sensor head that captured the current point, the elevation angle with respect to the elevation of the sensor that captured the current point, and the radius depending on the distance from the current point to a reference point. - Scaling the decoded spherical coordinates based on the scaling offset obtained from the scaling offset information. - Decoding the attributes of the current point based on the scaled decoded spherical coordinates. A method for decoding a point cloud from a bitstream of encoded point cloud data captured by a spin sensor head.

3. The scaling offset information includes the scaling offset. The method according to claim 1 or 2.

4. The scaling offset information instructs to determine the scaling offset using a specific method. The method according to claim 1 or 2.

5. The specific method uses clipping or modulo operation. The method according to claim 4.

6. An apparatus for encoding a point cloud captured by a spin sensor head into a bitstream of encoded point cloud data, wherein each point of the point cloud is associated with spherical coordinates and at least one attribute, the spherical coordinates representing the azimuth angle representing the capture angle of the sensor of the spin sensor head that captured the point, the elevation angle with respect to the elevation of the sensor that captured the point, and the radius depending on the distance from the point to a reference point, the apparatus including one or more processors, and the one or more processors - In the bitstream, sending a signal to notify scaling offset information representing a scaling offset, the scaling offset being determined based on the sensor and capture characteristics of the spin sensor head. For each current point of the point cloud, - Encoding the spherical coordinates of the current point and adding the encoded spherical coordinates to the bitstream. - Obtaining the decoded spherical coordinates of the current point by decoding the encoded spherical coordinates. - Scale the decoded spherical coordinates based on the scaling offset, - Based on the scaled decoded spherical coordinates, encode at least one attribute of the current point and add at least one encoded attribute to the bitstream, An apparatus for encoding a point cloud captured by a spin sensor head into a bitstream of encoded point cloud data. **Claim 7** An apparatus for decoding a point cloud from a bitstream of encoded point cloud data captured by a spin sensor head, the apparatus including one or more processors, the one or more processors being - Access scaling offset information representing a scaling offset from the bitstream, the scaling offset being determined based on a sensor and capture characteristics of the spin sensor head, For each current point of the point cloud, - Obtain decoded spherical coordinates by decoding the encoded spherical coordinates of the current point obtained from the bitstream, the spherical coordinates of the current point representing an azimuth angle representing a capture angle of the sensor of the spin sensor head that captured the current point, an elevation angle with respect to the elevation of the sensor that captured the current point, and a radius depending on the distance from the current point to a reference point, - Scale the decoded spherical coordinates based on the scaling offset obtained from the scaling offset information, - Based on the scaled decoded spherical coordinates, configured to decode an attribute of the current point, An apparatus for decoding a point cloud from a bitstream of encoded point cloud data captured by a spin sensor head. **Claim 8** A computer program which, when executed by one or more processors, causes the one or more processors to execute a method of encoding a point cloud captured by a spin sensor head into a bitstream of encoded point cloud data, each point of the point cloud being associated with spherical coordinates and at least one attribute, the spherical coordinates representing an azimuth angle representing a capture angle of a sensor of the spin sensor head that captured the point, an elevation angle relative to an elevation of the sensor that captured the point, and a radius depending on a distance from the point to a reference point, the method comprising: - In the bitstream, transmitting a signal to notify scaling offset information representing a scaling offset, the scaling offset being determined based on a sensor of the spin sensor head and capture characteristics; For each current point of the point cloud, - Encoding the spherical coordinates of the current point and adding the encoded spherical coordinates of the current point to the bitstream; - Obtaining decoded spherical coordinates by decoding the encoded spherical coordinates of the current point; - Scaling the decoded spherical coordinates based on the scaling offset; - Encoding at least one attribute of the current point based on the scaled decoded spherical coordinates and adding at least one encoded attribute to the bitstream. A computer program.

9. A computer program which, when executed by one or more processors, causes the one or more processors to execute a method of decoding a point cloud from a bitstream of encoded point cloud data captured by a spin sensor head, the method comprising: - Accessing scaling offset information representing a scaling offset from the bitstream, the scaling offset being determined based on a sensor of the spin sensor head and capture characteristics; For each current point of the point cloud, - Obtaining the decoded spherical coordinates by decoding the encoded spherical coordinates of the current point obtained from the bit stream, wherein the spherical coordinates of the current point represent the azimuth angle representing the capture angle of the sensor of the spin sensor head that captured the current point, the elevation angle with respect to the elevation of the sensor that captured the current point, and the radius depending on the distance from the current point to a reference point. - Scaling the decoded spherical coordinates based on the scaling offset obtained from the scaling offset information. - Decoding the attributes of the current point based on the scaled decoded spherical coordinates. A computer program.

10. A non-transitory storage medium carrying program code instructions for executing a method of encoding a point cloud captured by a spin sensor head into a bit stream of encoded point cloud data, wherein each point of the point cloud is associated with spherical coordinates and at least one attribute, and the spherical coordinates represent the azimuth angle representing the capture angle of the sensor of the spin sensor head that captured the point, the elevation angle with respect to the elevation of the sensor that captured the point, and the radius depending on the distance from the point to a reference point, and the method includes: - In the bit stream, transmitting a signal to notify scaling offset information representing a scaling offset, wherein the scaling offset is determined based on the sensor and capture characteristics of the spin sensor head. For each current point of the point cloud, - Encoding the spherical coordinates of the current point and adding the encoded spherical coordinates of the current point to the bit stream. - Obtaining the decoded spherical coordinates by decoding the encoded spherical coordinates of the current point. - Scaling the decoded spherical coordinates based on the scaling offset. - Encoding at least one attribute of the current point based on the scaled decoded spherical coordinates and adding at least one encoded attribute to the bit stream. A non-transitory storage medium.

11. A non-transitory memory medium carrying instructions of program code for performing a method of decoding a point cloud from a bitstream of encoded point cloud data captured by a spin sensor head, the method comprising: - accessing scaling offset information representing a scaling offset from the bitstream, the scaling offset being determined based on a sensor of the spin sensor head and capture characteristics; for each current point of the point cloud, - obtaining decoded spherical coordinates by decoding the encoded spherical coordinates of the current point obtained from the bitstream, the spherical coordinates of the current point representing an azimuth angle representing a capture angle of a sensor of the spin sensor head that captured the current point, an elevation angle with respect to an elevation of the sensor that captured the current point, and a radius depending on a distance from the current point to a reference point; - scaling the decoded spherical coordinates based on a scaling offset obtained from the scaling offset information; - decoding an attribute of the current point based on the scaled decoded spherical coordinates. A non-transitory memory medium.

Citation Information

Patent Citations

  • Method and apparatus for point cloud compression

    WO2020251888A1