Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
Patent Information
- Application Number
- JP2023566140
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-09
- Filing Date
- 2022-10-24
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2042-10-24
AI Technical Summary
【0011】 本開示は、符号化効率を向上できる三次元データ符号化方法、三次元データ復号方法、三次元データ符号化装置、及び三次元データ復号装置を提供できる。
Smart Images

Figure 0007914135000001 
Figure 0007914135000002 
Figure 0007914135000003
Abstract
Description
[Technical Field]
[0001] This disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. [Background technology]
[0002] In the future, devices and services utilizing three-dimensional data are expected to become widespread in a wide range of fields, including computer vision for autonomous operation of automobiles or robots, map information, monitoring, infrastructure inspection, and video distribution. Three-dimensional data can be acquired in various ways, such as using distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras.
[0003] One method of representing three-dimensional data is called a point cloud, which represents the shape of a three-dimensional structure using a cloud of points in three-dimensional space. In a point cloud, the position and color of the points are stored. Point clouds are expected to become the mainstream method of representing three-dimensional data, but point clouds are extremely large in size. Therefore, in the storage or transmission of three-dimensional data, data compression through encoding is essential, just as with two-dimensional moving images (for example, MPEG-4 AVC or HEVC, which are standardized by MPEG).
[0004] Furthermore, point cloud compression is partially supported by publicly available libraries (such as the Point Cloud Library) that handle point cloud-related processing.
[0005] Furthermore, there is a known technique for searching for and displaying facilities located around a vehicle using three-dimensional map data (see, for example, Patent Document 1). [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] International Publication No. 2014 / 020663 [Overview of the project] [Problems that the invention aims to solve]
[0007] In the encoding and decoding of three-dimensional data, improving encoding efficiency is highly desirable.
[0008] This disclosure aims to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device that can improve encoding efficiency. [Means for solving the problem]
[0009] A three-dimensional data encoding method relating to one aspect of this disclosure is: Motion compensation is performed to generate a first set of reference points by correcting the position information of one or more first three-dimensional points to match the coordinate system of the target three-dimensional point to be encoded. A predicted point of the target three-dimensional point is selected from either the first set of reference points or a second set of reference points including the one or more first three-dimensional points with the position information before correction, and the position information of the target three-dimensional point is encoded by referring to at least a portion of the position information of the predicted point. .
[0010] A three-dimensional data decoding method relating to one aspect of this disclosure is: Motion compensation is performed to generate a first reference point group by correcting the position information of one or more first three-dimensional points to match the coordinate system of the target three-dimensional point to be decoded. A predicted point of the target three-dimensional point is selected from either the first reference point group or a second reference point group including the one or more first three-dimensional points with the position information before correction, and the position information of the target three-dimensional point is decoded by referring to at least a portion of the position information of the predicted point. . [Effects of the Invention]
[0011] This disclosure provides a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device that can improve encoding efficiency. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 is a diagram illustrating a method for encoding or decoding a three-dimensional point represented in polar coordinates using interpretation according to Embodiment 1. [Figure 2] Figure 2 illustrates a method for encoding or decoding a three-dimensional point represented in polar coordinates using interpretation according to Embodiment 1. [Figure 3] Figure 3 is a diagram illustrating a method for encoding or decoding a three-dimensional point represented in polar coordinates using interpretation according to Embodiment 1. [Figure 4]Figure 4 is a diagram illustrating a method for encoding or decoding a three-dimensional point represented in polar coordinates using interpretation according to Embodiment 1. [Figure 5] Figure 5 is a flowchart showing an example of the processing procedure for the interpretation prediction method according to Embodiment 1. [Figure 6] Figure 6 is a diagram illustrating a first example in which the method for deriving the predicted value is switched according to the value of the horizontal angle according to Embodiment 1. [Figure 7] Figure 7 shows the formula for deriving the predicted value dpred, which is determined for each of the four directions defined by the horizontal angle φcur according to Embodiment 1. [Figure 8] Figure 8 is a diagram illustrating a second example in which the method for deriving the predicted value is switched according to the value of the horizontal angle in Embodiment 1. [Figure 9] Figure 9 shows the formula for deriving the predicted value dpred, which is determined for each of the eight directions defined by the horizontal angle φcur according to Embodiment 1. [Figure 10] Figure 10 is a block diagram of a three-dimensional data encoding device according to Embodiment 2. [Figure 11] Figure 11 is a block diagram of a three-dimensional data decoding device according to Embodiment 2. [Figure 12] Figure 12 is a flowchart of the encoding or decoding process including interpretation processing according to Embodiment 2. [Figure 13] Figure 13 is a diagram illustrating an example of an inter prediction method according to Embodiment 2. [Figure 14] Figure 14 is a flowchart of the interpretation prediction process according to Embodiment 2. [Figure 15] Figure 15 is a flowchart of the three-dimensional data encoding process according to Embodiment 2. [Figure 16] Figure 16 is a flowchart of the three-dimensional data decoding process according to Embodiment 2. [Modes for carrying out the invention]
[0013] A three-dimensional data encoding method according to one aspect of the present disclosure generates a first reference point group by correcting the positional information of one or more first three-dimensional points to match the coordinate system of a target three-dimensional point to be encoded; selects either the first reference point group or a second reference point group including the one or more first three-dimensional points before correction as a third reference point group for the target three-dimensional point; determines a predicted point using the third reference point group; and encodes the positional information of the target three-dimensional point by referring to at least a portion of the positional information of the predicted point.
[0014] According to this, the three-dimensional data encoding method encodes target points by selectively using a first corrected set of reference points and a second set of reference points before correction. Therefore, this three-dimensional data encoding method has the potential to determine prediction points that reduce prediction errors. Thus, this three-dimensional data encoding method can improve encoding efficiency. Furthermore, this three-dimensional data encoding method can reduce the amount of data handled in the encoding process.
[0015] For example, in the correction, the positional information of the one or more first three-dimensional points may be adjusted to the coordinate system of the target three-dimensional point based on first information indicating the displacement between the coordinate system of the one or more first three-dimensional points and the coordinate system of the target three-dimensional point.
[0016] For example, in the correction, the positional information of one or more second three-dimensional points included in the first reference point group may be derived by projecting the one or more first three-dimensional points onto the coordinate origin of the target three-dimensional point according to the displacement.
[0017] For example, the first information may include at least one of second information relating to movement parallel to the horizontal plane and third information relating to rotation about a vertical axis.
[0018] For example, the position information includes a distance component, a horizontal angle component, and an elevation angle component, and the correction may correct at least one of the distance component and the horizontal angle component. That is, in this embodiment, the elevation angle component in polar coordinates is not corrected. Therefore, this embodiment is suitable when selectively using a reference point cloud with horizontal position correction and a reference point cloud before correction. For example, this embodiment is suitable for a three-dimensional point cloud obtained from a sensor that repeatedly moves and stops in the horizontal direction.
[0019] For example, the three-dimensional data encoding method may further determine whether or not to perform the correction and generate a bitstream that includes the encoded position information of the target three-dimensional point and a fourth piece of information indicating whether or not to perform the correction.
[0020] According to this, the three-dimensional data encoding method can determine prediction points that reduce prediction errors by switching whether or not to perform correction.
[0021] For example, if the above correction is not performed, the second group of reference points may be selected as the third group of reference points.
[0022] For example, if the one or more first three-dimensional points are included in the first processing unit, and the correction is not performed, one of the second group of reference points and the fourth group of reference points, which consists of one or more third three-dimensional points included in a second processing unit different from the first processing unit and which are the one or more third three-dimensional points before the correction, may be selected as the third group of reference points.
[0023] According to this, the three-dimensional data encoding method can reference two uncorrected processing units if no correction is applied. Therefore, the encoding efficiency of the three-dimensional data encoding method can be improved.
[0024] For example, the first group of reference points may be generated by correcting one or more fourth three-dimensional points, which are part of the one or more first three-dimensional points.
[0025] According to this, the three-dimensional data encoding method can reduce processing load by limiting the three-dimensional points to be corrected. For example, if the relative position of the target three-dimensional point and the origin is approximately equal to the relative position of the predicted points included in the reference point set and the origin, the prediction error can be suppressed by using those predicted points without correction. On the other hand, if the relative position of the target three-dimensional point and the origin is different from the relative position of the predicted points included in the reference point set and the origin, the prediction error can be suppressed by using those corrected predicted points. In this way, the prediction error can be reduced by switching whether or not to perform correction depending on the position of the target three-dimensional point.
[0026] For example, the position information includes a distance component, a horizontal angle component, and an elevation angle component, and the fourth three-dimensional point may be a first three-dimensional point among the one or more first three-dimensional points whose elevation angle component is larger than a predetermined value. In other words, in this embodiment, the three-dimensional points to be corrected are limited to three-dimensional points with a large elevation angle component. Three-dimensional points with a large elevation angle component represent, for example, buildings. Buildings are fixed to the ground. Therefore, when the target three-dimensional point and the predicted point represent a building, the relative positional relationship between the target three-dimensional point and the origin is different from the relative positional relationship between the predicted point included in the reference point group and the origin. In this case, using the corrected predicted point can suppress prediction errors. Therefore, the three-dimensional data encoding method limits the targets to buildings and the like.
[0027] Similarly, the fourth three-dimensional point may be one of the one or more first three-dimensional points whose vertical position is higher than a predetermined position.
[0028] A three-dimensional data decoding method according to one aspect of the present disclosure generates a first reference point group by correcting the positional information of one or more first three-dimensional points to match the coordinate system of a target three-dimensional point to be decoded; selects either the first reference point group or a second reference point group including the one or more first three-dimensional points before correction as a third reference point group for the target three-dimensional point; determines a predicted point using the third reference point group; and decodes the positional information of the target three-dimensional point by referring to at least a portion of the positional information of the predicted point.
[0029] According to this, the three-dimensional data decoding method decodes the target points by selectively using a first corrected set of reference points and a second set of reference points before correction. Therefore, this three-dimensional data decoding method has the potential to determine prediction points that reduce prediction errors. Thus, this three-dimensional data decoding method can reduce the amount of data handled in the decoding process.
[0030] For example, in the correction, the positional information of the one or more first three-dimensional points may be adjusted to the coordinate system of the target three-dimensional point based on first information indicating the displacement between the coordinate system of the one or more first three-dimensional points and the coordinate system of the target three-dimensional point.
[0031] For example, in the correction, the positional information of one or more second three-dimensional points included in the first reference point group may be derived by projecting the one or more first three-dimensional points onto the coordinate origin of the target three-dimensional point according to the displacement.
[0032] For example, the first information may include at least one of second information relating to movement parallel to the horizontal plane and third information relating to rotation about a vertical axis.
[0033] For example, the position information includes a distance component, a horizontal angle component, and an elevation angle component, and the correction may correct at least one of the distance component and the horizontal angle component. That is, in this embodiment, the elevation angle component in polar coordinates is not corrected. Therefore, this embodiment is suitable when selectively using a reference point cloud with horizontal position correction and a reference point cloud before correction. For example, this embodiment is suitable for a three-dimensional point cloud obtained from a sensor that repeatedly moves and stops in the horizontal direction.
[0034] For example, the three-dimensional data decoding method may further acquire fourth information from the bitstream indicating whether or not to perform the correction, and decide whether or not to perform the correction based on the fourth information.
[0035] According to this, the three-dimensional data decoding method can determine prediction points that minimize prediction errors by switching whether or not to perform correction.
[0036] For example, if the above correction is not performed, the second group of reference points may be selected as the third group of reference points.
[0037] For example, if the one or more first three-dimensional points are included in the first processing unit, and the correction is not performed, one of the second group of reference points and the fourth group of reference points, which consists of one or more third three-dimensional points included in a second processing unit different from the first processing unit and which are the one or more third three-dimensional points before the correction, may be selected as the third group of reference points.
[0038] According to this, the three-dimensional data decoding method can reference two uncorrected processing units if no correction is applied. Therefore, the encoding efficiency of the three-dimensional data decoding method can be improved.
[0039] For example, the first group of reference points may be generated by correcting one or more fourth three-dimensional points, which are part of the one or more first three-dimensional points.
[0040] According to this, the three-dimensional data decoding method can reduce processing load by limiting the three-dimensional points to be corrected. For example, if the relative positional relationship between the target three-dimensional point and the origin is approximately equal to the relative positional relationship between the predicted points included in the reference point set and the origin, the prediction error can be suppressed by using the predicted points without correction. On the other hand, if the relative positional relationship between the target three-dimensional point and the origin is different from the relative positional relationship between the predicted points included in the reference point set and the origin, the prediction error can be suppressed by using the corrected predicted points. In this way, the prediction error can be reduced by switching whether or not to perform correction depending on the position of the target three-dimensional point.
[0041] For example, the position information includes a distance component, a horizontal angle component, and an elevation angle component, and the fourth three-dimensional point may be one of the one or more first three-dimensional points whose elevation angle component is larger than a predetermined value. In other words, in this embodiment, the three-dimensional points to be corrected are limited to three-dimensional points with a large elevation angle component. Three-dimensional points with a large elevation angle component represent, for example, buildings. Buildings are fixed to the ground. Therefore, when the target three-dimensional point and the predicted point represent a building, the relative positional relationship between the target three-dimensional point and the origin differs from the relative positional relationship between the predicted point included in the reference point group and the origin. In this case, using the corrected predicted point can suppress prediction errors. Therefore, the three-dimensional data decoding method limits the targets to buildings, etc.
[0042] Similarly, the fourth three-dimensional point may be one of the one or more first three-dimensional points whose vertical position is higher than a predetermined position.
[0043] Furthermore, a three-dimensional data encoding device according to one aspect of the present disclosure comprises a processor and a memory, wherein the processor generates a first reference point group by correcting the positional information of one or more first three-dimensional points to match the coordinate system of the target three-dimensional point to be encoded using the memory, selects either the first reference point group or a second reference point group including the one or more first three-dimensional points before correction as a third reference point group for the target three-dimensional point, determines a predicted point using the third reference point group, and encodes the positional information of the target three-dimensional point by referring to at least a portion of the positional information of the predicted point.
[0044] According to this, the three-dimensional data encoding device selectively uses a first corrected set of reference points and a second set of uncorrected reference points to encode the target points. Therefore, this three-dimensional data encoding device has the potential to determine prediction points that minimize prediction errors. Thus, this three-dimensional data encoding device can improve encoding efficiency. Furthermore, this three-dimensional data encoding device can reduce the amount of data handled in the encoding process.
[0045] Furthermore, a three-dimensional data decoding device according to one aspect of the present disclosure comprises a processor and a memory, wherein the processor generates a first reference point group by correcting the positional information of one or more first three-dimensional points to match the coordinate system of the target three-dimensional point to be decoded using the memory, selects either the first reference point group or a second reference point group including the one or more first three-dimensional points before correction as a third reference point group for the target three-dimensional point, determines a predicted point using the third reference point group, and decodes the positional information of the target three-dimensional point by referring to at least a portion of the positional information of the predicted point.
[0046] According to this, the three-dimensional data decoding device selectively uses a first corrected set of reference points and a second set of uncorrected reference points to decode the target points. Therefore, this three-dimensional data decoding device has the potential to determine prediction points that minimize prediction errors. Thus, this three-dimensional data decoding method can reduce the amount of data handled in the decoding process.
[0047] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.
[0048] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, components in the following embodiments that are not described in an independent claim will be described as optional components.
[0049] (Embodiment 1) This embodiment describes a three-dimensional data encoding method and a three-dimensional data decoding method for interpreting a point cloud containing multiple three-dimensional points whose positional information is represented in polar coordinates. Note that positional information may sometimes be simply referred to as "position." The following mainly describes a method for determining one or more candidate points used to determine the predicted values for interpretation.
[0050] Figures 1 to 3 illustrate how to encode or decode three-dimensional points represented in polar coordinates using interpretation.
[0051] Here, interprediction is a method of predicting and encoding multiple three-dimensional points in the target frame by referencing one or more pre-encoded three-dimensional points in a reference frame (a frame different from the target frame) when encoding a three-dimensional point to be encoded in the target frame, and based on the referenced one or more three-dimensional points. Intraprediction is a method of predicting and encoding multiple three-dimensional points in the target frame by referencing at least one of the other one or more pre-encoded three-dimensional points in the target frame, and based on the referenced one or more three-dimensional points. The target frame may also be referred to as the second frame. The reference frame may also be referred to as the first frame.
[0052] The point cloud data has one or more frames, each of which has one or more three-dimensional points. The one or more frames include an encoding frame and a reference frame. For example, each frame may be generated by measurements taken at multiple different locations by a sensor. Each frame may also be generated by measurements taken by multiple different sensors.
[0053] Multiple first three-dimensional points are obtained by measuring the distance to an object in each of several first directions around a first position in space on a reference plane. The first position is the first origin that serves as the reference for the positional information of the multiple first three-dimensional points, which are measurement results from a sensor placed at a third position. The first position may also be referred to as the first reference position. The first position may or may not coincide with the third position where the sensor is placed. Each of the multiple first three-dimensional points is represented in a first polar coordinate system with the first position as the first origin. Multiple first three-dimensional points are included in a reference frame, for example.
[0054] Furthermore, multiple second three-dimensional points are obtained by measuring the distance to an object in each of multiple second directions around the second position in space on the reference plane. The second position is the second origin that serves as the reference for the positional information of the multiple second three-dimensional points, which are measurement results from a sensor placed at the fourth position. The second position may also be referred to as the second reference position. The second position may or may not coincide with the fourth position where the sensor is placed. Each of the multiple second three-dimensional points is represented in a second polar coordinate system with the second position as the second origin.
[0055] The sensor emits electromagnetic waves and acquires reflected waves from the object, thereby generating measurement results that include one or more three-dimensional points. In this embodiment, the sensor may generate a single frame containing the measurement results obtained in a single measurement. Specifically, the sensor measures the time it takes for the emitted electromagnetic wave to be emitted, reflected from the object, and returned to the sensor, and uses the measured time and the wavelength of the electromagnetic wave to calculate the distance between the sensor and a point on the surface of an object surrounding the sensor. The sensor emits electromagnetic waves in a predetermined radial direction from the sensor's reference position. The sensor is, for example, a LiDAR, and the electromagnetic wave is, for example, laser light.
[0056] Each three-dimensional point has at least positional information. The positional information indicates the location of the three-dimensional point having the positional information and is expressed in polar coordinates. Specifically, the positional information includes the distance from the reference point to the three-dimensional point having the positional information, and two angles indicating the direction from the reference point to the three-dimensional point having the positional information. One of the two angles is, for example, the angle (horizontal angle) between the reference direction perpendicular to the axis and the above direction when viewed from an axis perpendicular to the reference plane, and the other is the angle (elevation angle) between the reference plane and the above direction. The reference plane is a horizontal plane, and can be, for example, a plane perpendicular to a predetermined axis of the sensor, such as the rotation axis of a LiDAR, or the ground, floor, or a plane parallel to these.
[0057] Figures 1 to 3 assume the encoding of a point cloud generated by acquiring the three-dimensional positions of objects around a sensor, centered on the sensor's position, as is the case with LiDAR. Figures 1 to 3 show, for example, the positional relationship between the second reference position 13808 of sensor 13806 when measuring the point cloud of the frame to be encoded, the first reference position 13807 of sensor 13805 when measuring the point cloud of the reference frame, the point to be encoded 13801, and the reference candidate points 13802 and 13803 in interpretation. Figures 1 to 3 are plan views from the axial direction, for example, in the case of a sensor that measures the distance to an object by emitting a laser around a predetermined axis (rotation axis), such as a LiDAR sensor. The point to be encoded 13801 and the reference candidate points 13802 and 13803 indicate the three-dimensional positions on the same plane (e.g., a plane) 13810 of object 13804. The point to be encoded 13801 is included in the point cloud of the frame to be encoded. The point cloud of the frame to be encoded is shown as black-filled diamond-shaped dots in Figures 1 to 3. Reference candidate points 13802 and 13803 are included in the point cloud of the reference frame and are also included in a plurality (e.g., n+1) of three-dimensional points that indicate the three-dimensional position on the surface 13810. The point cloud of the reference frame is shown as white-outlined diamond-shaped dots in Figures 1 to 3. Sensors 13805 and 13806 may be the same sensor or different sensors (i.e., separate sensors). If sensors 13805 and 13806 are the same sensor, for example, this occurs when one sensor moves from the first reference position 13807 to the second reference position 13808, or from the second reference position 13808 to the first reference position 13807. In this case, the time when the frame to be encoded was generated and the time when the reference frame was generated are different. If sensor 13805 and sensor 13806 are different sensors, the time at which the encoded frame was generated and the time at which the reference frame was generated may be different or the same.
[0058] The three-dimensional data encoding device measures the distance d from the second reference position 13808 of the sensor 13806 to the point to be encoded 13801. cur Predicted value d predis determined based on the positional relationship among a second reference position 13808 of the sensor 13806 when measuring the point cloud of a frame to be encoded, a first reference position 13807 of the sensor 13805 when measuring the point cloud of a reference frame, an encoding target point 13801, and reference candidate points 13802 and 13803. Then, the three-dimensional data encoding device obtains the determined prediction value d pred may be used for inter prediction. For example, the three-dimensional data encoding device obtains the prediction value d by executing the following procedures 1 to 3 pred may be determined. Note that the three-dimensional data decoding device obtains the prediction value d by executing the same processing as the three-dimensional data encoding device pred is determined, and therefore only the three-dimensional data encoding device will be described below.
[0059] In procedure 1, as shown in FIG. 2, the three-dimensional data encoding device projects at least one reference candidate point 13802, 13803 onto a second polar coordinate system of the second reference position 13808, and obtains the horizontal angle φ of the i-th reference candidate point as viewed from the second reference position 13808 according to formula Z1 ref2 (i) is derived. Note that φ ref1 (i), d ref1 (i), and m respectively represent the horizontal angle at the first reference position 13807 shown in FIG. 2, the distance from the first reference position 13807 to the i-th reference candidate point, and the distance between the first reference position 13807 and the second reference position 13808 (inter-sensor distance or movement distance). The horizontal angle is an angle of a direction in which the i-th reference candidate point exists with respect to the first reference position 13807, with reference to a reference direction 13809 which is the direction connecting the first reference position 13807 and the second reference position 13808. In addition, n is a natural number that can take i and k. i is any natural number.
[0060] φ ref2 (i)=arctan( d ref1 (i)sin(φ ref1 (i)) / (d ref1 (i)cos(φ ref1 (i)) - m) ) (Formula Z1)
[0061] In step 2, the three-dimensional data encoding device determines at least one horizontal angle φ corresponding to the i-th reference candidate point. ref2 (i) From among these, the horizontal angle φ from the second reference position 13808 to the coding target point 13801 cur The horizontal angle φ in the vicinity ref2 Select (k), and as shown in Figure 3, horizontal angle φ ref2 The k-th reference candidate point 13802, pointed to by (k), is selected as the reference point (inter-reference point) 13811 to be used for inter-prediction. Note that k is the n horizontal angles φ ref2 Among these, the horizontal angle φ points from the second reference position 13808 to the encoding target point 13801. cur This is a natural number indicating that it is a horizontal angle in the neighborhood of . In other words, the k-th reference candidate point is a reference point used to calculate the predicted value from among the n reference candidate points, and is an example of a coded first three-dimensional point.
[0062] In step 3, the three-dimensional data encoding device measures the distance d from the second reference position 13808 to the interreference point 13811. ref2 Derive (k), and distance d ref2 (k) predict value d pred This will be decided.
[0063] d pred =d ref2 (k) = d ref1 (k) sin(φ ref1 (k) / sin(φ ref2 (k)) (Formula Z2)
[0064] The point cloud of the frame to be encoded and the point cloud of the reference frame are represented in polar coordinates of coordinate systems with different reference positions. Therefore, when predictively encoding point 13801 using reference candidate points 13802 and 13803 of a reference frame different from the frame to be encoded, it is necessary to perform a coordinate system transformation to convert the coordinate systems of reference candidate points 13802 and 13803 of the reference frame from the first coordinate system in which the point cloud of the reference frame is represented to the second coordinate system in which the point cloud of the encoded frame is represented.
[0065] The three-dimensional data encoding device identifies a first three-dimensional point whose position is represented in first polar coordinates and has been encoded, by performing steps 1 to 3. The first three-dimensional point is a reference point used for prediction values. The three-dimensional data encoding device then identifies the distance d from the second reference position 13808 to the second three-dimensional point whose position is represented in second polar coordinates and has not been encoded. cur To calculate the predicted value, (i) the distance m between the first reference position 13807 and the second reference position 13808, and (ii) the horizontal angle φ between the first line connecting the first reference position 13807 and the second reference position 13808 and the second line connecting the first reference position 13807 and the reference point 13811. ref1 (k), and (iii) the distance d from the first reference position 13807 of the reference point 13811 in the first polar coordinate system. ref1 (k) is identified. Note that the distance d cur This is an example of a second distance. Distance m is an example of the distance between the first and second positions. Horizontal angle φ ref1 (k) is an example of the first angle, showing the angle between the first line and the second line. Distance d ref1 (k) is an example of the first distance.
[0066] Furthermore, for example, a three-dimensional data encoding device calculates a predicted value of the position of an unencoded second three-dimensional point in a second polar coordinate system, using distance m and horizontal angle φ. ref1 (k), and distance d ref1 Using (k), (iv) the horizontal angle φ between the first line and the third line connecting the second reference position 13808 and the reference point 13811. ref2 (k), and (v) the distance d from the second reference position 13808 of reference point 13811 in the second polar coordinate system. cur Calculate the horizontal angle φ. ref2 (k) is an example of a second angle.
[0067] Also, for example, distance d cur In the calculation, when the first line and the reference line for the horizontal angle in the first polar coordinate system are aligned (i.e., when the first line and the reference line for the horizontal angle are parallel), the distance d ref1 (k) and the horizontal angle φ as the first angle ref1 Using (k), the horizontal angle φref2 Calculate (k). Horizontal angle φ ref1 (k) is an example of the first horizontal angle, and is the horizontal angle component of the polar coordinate components representing the position of reference point 13811. The position of reference point 13811 is represented in the first polar coordinate system. In the examples in Figures 1 to 3, the case where the first line and the reference line for the horizontal angle in the first polar coordinate system are aligned (i.e., the first line and the reference line for the horizontal angle are parallel) is shown. However, if the first line and the reference line for the horizontal angle in the first polar coordinate system are not aligned (i.e., the first line and the reference line for the horizontal angle are not parallel), the three-dimensional data encoding device may calculate the difference between the horizontal angle component of the polar coordinate components representing the position of reference point 13811 and the angle that the first line makes with the reference line as the first angle.
[0068] Furthermore, in identifying the reference point 13811, the three-dimensional data encoding device may identify the first three-dimensional point based on another second three-dimensional point whose position is represented in a second polar coordinate system and has already been encoded.
[0069] Based on the above, the three-dimensional data encoding device measures the distance d from the second reference position 13808 of the frame to be encoded to the point 13801. cur This makes it possible to predict with high accuracy, potentially improving the efficiency of interpredictive coding.
[0070] The travel distance m may be generated based on results measured using sensors such as GPS (Global Positioning System) and odometers, or on results derived using self-localization techniques such as SfM (Structure from Motion) and SLAM (Simultaneous Localization and Mapping). These results may also be included in the header information of predetermined data units such as frames and slices, or in the information of the first node of these data units. In this way, these results may be notified from the three-dimensional data encoding device to the three-dimensional data decoding device.
[0071] Furthermore, in step 3 above, d ref2(k) may be derived.
[0072] d ref2 (k) = (d ref1 (k)cos(φ ref1 (k)-m) / cos(φ ref2 (k)) (Formula Z3)
[0073] Alternatively, the three-dimensional data encoding device may determine the interreference point by projecting all candidate reference points of the reference frame onto the second polar coordinate system of the second reference position 13808 of the frame to be encoded. In other words, the three-dimensional data encoding device may perform the above coordinate system transformation on all candidate reference points of the reference frame and determine the interreference point based on all the transformed candidate points. Or, the three-dimensional data encoding device may determine the interreference point based on the horizontal angle φ of the point to be encoded 13801. cur The reference candidate points may be limited to points that fall within a range of the horizontal angle of the reference frame with respect to the sensor 13805, based on the distance m between the first reference position 13807 and the second reference position 13808, for example. cur If the horizontal angle φ is small, or if the point to be encoded is in the direction connecting the first and second reference positions, the process of determining the inter-reference point is not performed. Also, if the distance m is large, the process is limited to reference candidate points in the direction connecting the first and second reference positions, but in the region ahead of the first reference position. This reduces the processing load required to determine the inter-reference point. In other words, the three-dimensional data encoding device uses a horizontal angle φ cur By limiting the reference candidate points subject to coordinate system transformation to a portion of all reference candidate points based on factors such as distance m, the amount of processing involved in coordinate system transformation can be reduced.
[0074] Furthermore, the arithmetic operations such as trigonometric functions and division in each of the above steps may be simplified by using a table with a finite number of elements. Simplification can improve the efficiency of interpredictive coding while suppressing the amount of processing required.
[0075] Furthermore, the angle between the reference direction 13809 and the same plane (e.g., a flat surface) 13810 of object 13804 in Figures 1 to 3 is not limited. That is, the direction of the first line connecting the first reference position 13807 and the second reference position 13808 can take any angle with respect to the horizontal axis contained in the same plane (e.g., a flat surface) 13810 of object 13804. Even in that case, the predicted value can be calculated using the method described above.
[0076] Next, Figure 4 assumes the encoding of a point cloud generated by acquiring the three-dimensional positions of objects around the sensor, centered on the sensor's position, as in LiDAR. In Figure 4, for example, in steps 1 to 3 explained using Figures 1 to 3, the three-dimensional data encoding device views the point 13821 to be encoded from the sensor 13826 of the frame to be encoded at the elevation angle θ. cur Taking this into consideration, the distance d from sensor 13826 to encoding point 13821 of the frame to be encoded is also considered. cur Predicted value d pred You may decide on that.
[0077] Figure 4 shows, for example, the positional relationship between the second reference position 13828 of sensor 13826 in the frame to be encoded, the first reference position 13827 of sensor 13826 in the reference frame, the point to be encoded 13821, and the reference candidate point 13822 in the interpretation. The point to be encoded 13821 and the reference candidate point 13822 represent the three-dimensional positions on the same surface (e.g., plane) 13824 of object 13823. The point to be encoded 13821 is included in the point cloud of the frame to be encoded. In Figure 4, the point cloud of the frame to be encoded is shown as a black-filled diamond shape. The reference candidate point 13822 is included in the point cloud of the reference frame and is a three-dimensional point that represents the three-dimensional position on the surface 13824 of object 13823. A reference frame contains one or more three-dimensional points, and may contain multiple (e.g., n) three-dimensional points. In Figure 4, the point cloud of the reference frame is shown as a white-outlet diamond shape. Sensors 13825 and 13826 may be the same sensor or different sensors (i.e., different sensors). If sensors 13825 and 13826 are the same sensor, for example, one sensor moves from the first reference position 13827 to the second reference position 13828, or from the second reference position 13828 to the first reference position 13827. In this case, the time when the frame to be encoded is generated and the time when the reference frame is generated are different. If sensors 13825 and 13826 are different sensors, the time when the frame to be encoded is generated and the time when the reference frame is generated may be different or the same.
[0078] The three-dimensional data encoding device predicts the value d by performing steps 11 to 13 below. pred You may decide on that.
[0079] In step 11, the three-dimensional data encoding device projects at least one reference candidate point 13822 onto the second polar coordinate system of the second reference position 13828, and calculates the horizontal angle φ of the i-th reference candidate point as seen from the second reference position 13828 using equations Z4 and Z5. ref2 (i) and elevation angle θ ref2 Derive (i). Note that φ ref2 (i), θ ref2 (i), dref1 (i) and m represent the horizontal angle and elevation angle at the first reference position 13827 shown in Figure 4, the distance from the first reference position to the i-th reference candidate point, and the distance between the first reference position 13827 and the second reference position 13828 (distance between sensors or distance traveled), respectively. The horizontal angle is the angle in the direction in which the i-th reference candidate point exists relative to the first reference position 13827, with reference direction 13829 of the first and second reference positions 13827 and 13828 as the reference plane. The elevation angle is the angle in the direction in which the i-th reference candidate point exists relative to the first reference position 13827, with reference to the horizontal plane. Also, n is a natural number that can take the form of i and k. i is an arbitrary natural number.
[0080] φ ref2 (i) = arctan( d ref1 (i)sin(φ ref1 (i)) / (d ref1 (i)cos(φ ref1 (i)) - m) ) (Formula Z4) θ ref2 (i) = arctan( ( tan(θ ref1 (i))+(h ref1 (i)-h ref2 (i)) / d ref1 (i) ) × sin(φ ref2 (i)) / sin(φ ref1 (i)) ) (Formula Z5)
[0081] In addition, in formula Z5, h ref1 (i) indicates the height of sensor 13825 from the reference plane, h ref2 (i) indicates the height of sensor 13826 from the reference plane.
[0082] In step 12, the three-dimensional data encoding device generates at least one horizontal angle φ corresponding to the reference candidate point 13822. ref2 (i) and elevation angle θ ref2 From the combinations in (i), the horizontal angle φ that points from the second reference position 13828 to the coding target point 13821. cur and elevation angle θ cur The horizontal angle φ in the vicinity of the combination ref2(k) and an elevation angle θ ref2 (k), and selects the k-th reference candidate point 13822 indicated by the combination of the horizontal angle φ ref2 (k) and the elevation angle θ ref2 (k) as a reference point (inter-reference point) used for inter prediction. Here, k is a natural number indicating that among n combinations of the horizontal angle φ ref2 and the elevation angle θ ref2 , it is a combination of the horizontal angle φ cur and the elevation angle θ cur that are in the vicinity of the horizontal angle φ ref2 (k) and the elevation angle θ ref2 (k) indicating the encoding target point 13821 from the second reference position 13828. That is, the k-th reference candidate point is a reference point used for calculating a prediction value among n reference candidate points, and is an example of an encoded first three-dimensional point.
[0083] In step 13, the three-dimensional data encoding device calculates the distance d ref2 (k) from the second reference position 13828 to the reference candidate point 13822 selected as the inter-reference point, and sets the distance d ref2 (k) as the prediction value d pred determined as the prediction value d .
[0084] d pred =d ref2 (k)= d ref1 (k) sin(φ ref1 (k)) / sin(φ ref2 (k)) (Formula Z6)
[0085] As described above, the three-dimensional data encoding device can predict the distance d cur from the sensor 13826 of the encoding target frame to the encoding target point 13821 with high accuracy, which may potentially improve the efficiency of inter prediction encoding.
[0086] Note that in the aforementioned step 13, d ref2 (k) may be derived by Formula Z7.
[0087] d ref2 (k)=(d ref1 (k)cos(φref1 (k)-m) / cos(φ ref2 (k)) (Formula Z7)
[0088] Alternatively, the three-dimensional data encoding device may determine the interreference point by projecting all candidate reference points of the reference frame onto the second polar coordinate system of the second reference position 13828 of the sensor 13826 of the frame to be encoded. In other words, the three-dimensional data encoding device may project all candidate reference points of the reference frame onto the second polar coordinate system of the second reference position and determine the interreference point based on all the transformed candidate points. Or, the three-dimensional data encoding device may determine the interreference point based on the horizontal angle φ of the point to be encoded 13821. cur Or elevation angle θ cur Based on the distance m between the first reference position 13827 and the second reference position 13828, the candidate reference points may be limited to points that fall within a certain range of the horizontal angle and elevation angle of the reference frame relative to the sensor 13825. This reduces the processing load required to determine the interreference points. In other words, the three-dimensional data encoding device can determine the horizontal angle φ cur , elevation angle θ cur By limiting the reference candidate points subject to coordinate system transformation to a portion of all reference candidate points based on factors such as distance m, the amount of processing involved in coordinate system transformation can be reduced.
[0089] Furthermore, the trigonometric functions, division, and other arithmetic operations in each of the above steps may be simplified using a table with a finite number of elements. In addition, in step 11, (h ref1 (i)-h ref2 (i)) / d ref1 Assuming (i) is sufficiently small, θ can be expressed by equation Z8. ref2 (i) can also be calculated. Through the simplification described above, the efficiency of interpredictive coding can be improved while suppressing the amount of processing required.
[0090] θ ref2 (i) = arctan( tan(θ ref1 (i)) sin(φ ref2 (i)) / sin(φ ref1 (i)) ) (Formula Z8)
[0091] Figure 5 is a flowchart showing an example of the processing procedure for the interpretation prediction method.
[0092] The three-dimensional data encoding device determines the intra-prediction point (d) in the frame to be encoded. intra ,φ intra ,θ intra The intra-prediction point is determined to be the reference point for inter-prediction (S13801). In addition, intra-prediction information may be notified and the predicted value from the prediction method determined to be appropriate may be used, or the notification of some or all of the intra-prediction information may be omitted, limited to a specific prediction method. Here, the intra-prediction point determined to be the reference point for inter-prediction may be a three-dimensional point used in the calculation of the predicted value of the three-dimensional point to be encoded in the second polar coordinate system.
[0093] Next, the three-dimensional data encoding device generates intra-prediction points (d intra ,φ intra ,θ intra Project ) onto the first polar coordinate system of the reference frame, and use the angle (φ) that serves as the basis for selecting the reference candidate point in the reference frame. sref ,θ sref Determine the angle (φ) (S13802). sref ,θ sref ) may be determined by formulas Z9 and Z10.
[0094] φ sref =arctan(d intra sin(φ intra ) / (d intra cos(φ intra )+m)) (Formula Z9) θ sref = arctan(tan(θ intra )sin(φ sref ) / sin(φ intra )) (Formula Z10)
[0095] Next, the three-dimensional data encoding device uses an angle (φ) in the reference frame. sref ,θ sref ) angle (φ ref1 (i), θ ref1(i)) one or more three-dimensional points having the above characteristics are selected as reference candidate points by a predetermined method (S13803). The three-dimensional data encoding device uses an elevation angle θ sref Select one or more laser scanning lines with an angle close to θ, and the elevation angle θ sref In order of the laser scanning lines closest to each other, the horizontal angle φ in each laser scanning line sref Alternatively, you can select one or more three-dimensional points that are close to the target and set the order of selection as the index of the reference candidate points.
[0096] Next, the three-dimensional data encoding device uses a reference candidate point (d ref1 (i), φ ref1 (i), θ ref1 (i)) is projected onto the second polar coordinate system of the second reference position of the frame to be encoded, and the angle (φ) in the frame to be encoded is obtained. ref2 (i), θ ref2 (i)) is derived (S13804). Note that the angle (φ ref2 (i), θ ref2 (i)) may be derived by equations Z11 and Z12.
[0097] φ ref2 (i) = arctan(d ref1 (i)sin(φ ref1 (i)) / (d ref1 (i)cos(φ ref1 (i))-m)) (Formula Z11) θ ref2 (i) = arctan(tan(θ) ref1 (i))sin(φ ref2 (i)) / sin(φ ref1 (i))) (Formula Z12)
[0098] Next, the three-dimensional data encoding device uses an angle (φ ref2 (i), θ ref2 (i)) Select the intermediate reference point (φ) by the prescribed method. ref2 (k), θ ref2 Select (k)) (S13805). Note that the three-dimensional data encoding device uses an angle (φ cur ,θ curA reference candidate point having the closest angle to ) may be selected as the interreference point. In this way, the interreference point is identified based on the angular component of the polar coordinate component representing the position of another second three-dimensional point included in the frame to be encoded. The interreference point is one in which, among a plurality of first three-dimensional points whose position is represented in the first polar coordinate system and which include other encoded first three-dimensional points, the angular component of the plurality of projected first three-dimensional points after projection from the first polar coordinate system to the second polar coordinate system is closest to the angular component of the point to be encoded. Alternatively, the index k of the interreference point selected by a predetermined method may be notified to the three-dimensional data decoding device.
[0099] Next, the three-dimensional data encoding device measures the distance d from the first reference position of the frame to be encoded to the interreference point. ref2 Derive (k), and distance d ref2 (k) predict value d pred This is determined (S13806). Note that the distance d ref2 (k) may be derived by either equation Z13 or equation Z14.
[0100] d ref2 (k=d ref1 (k)sin(φ ref1 (k) / sin(φ ref2 (k)) (Formula Z13) d ref2 (k) = (d ref1 (k)cos(φ ref1 (k)-m) / cos(φ ref2 (k)) (Formula Z14)
[0101] Furthermore, the intersection reference point itself is used as the predicted value, φ cur Since the corresponding distance is not calculated, whether you use equation Z13 or equation Z14, the distance d ref2 (k) will have the same value.
[0102] Thus, the three-dimensional data encoding device may project the intra-prediction point of the frame to be encoded onto the first polar coordinate system of the first reference position of the reference frame, and select one or more inter-reference point candidates based on that angle. This reduces the number of reference candidate points used for inter-prediction and reduces the processing load for inter-prediction. Furthermore, the arithmetic operations such as trigonometric functions and division in the above procedure may be simplified using a table with a finite number of elements.
[0103] Therefore, it may be possible to improve the efficiency of interpredictive coding while suppressing the amount of processing required.
[0104] Furthermore, calculations or determinations regarding the elevation angle θ may be omitted, and calculations or determinations may be performed using only the horizontal angle φ. In addition, the interpretation and intrapretation in this embodiment may be switched on a per-node or per-slice basis, or intrapretation or other interpretation may be switched on a per-node or per-slice basis.
[0105] In the interpretation method explained using Figures 2 to 5, the distance d ref2 (k) is derived and the predicted value d pred It was decided as such, but as shown in Figure 6, φ cur The predicted value d depends on the value of pred The derivation method may be switched. Figure 6 is a diagram illustrating a first example of switching the derivation method of the predicted value according to the value of the horizontal angle. Figure 6 is a plan view from the axial direction in the case of a sensor that measures the distance to an object by emitting a laser around a predetermined axis (rotation axis), such as a LiDAR sensor. Figure 7 shows the horizontal angle φ cur Predicted values d are determined for each of the four directions defined by pred This figure shows the formula for deriving the expression. The Index in Figure 7 indicates that the hypothetical planes 13840 to 13843 set up in Figure 6 correspond to index values 0 to 3, respectively.
[0106] As shown in Figure 6, a virtual plane perpendicular to the horizontal plane (reference plane) is set in the four directions (front, back, left, and right) as the object, and the horizontal angle φcur For each of the four ranges divided by the formula shown in Figure 7, the predicted value d pred This is derived. In other words, the three-dimensional data encoding device uses the horizontal angle φ of the point to be encoded. cur The horizontal angle φ shown in Figure 9 is obtained. cur Predicted value d based on the formula corresponding to this pred We derive this. Note that in the example in Figure 6, |φ cur The ranges are defined as |≦π and 0<α<π / 2. α may be a predetermined constant, such as α=π / 6. Furthermore, α may be included in the header information of a predetermined data unit, such as a sequence, frame, or slice. This allows α to be notified from the 3D data encoding device to the 3D data decoding device and made modifiable. Also, the inclusion of each range boundary in the encoding and decoding processes only needs to be consistent and does not necessarily have to be as shown in Figure 7.
[0107] Thus, the three-dimensional data encoding device may have different methods for determining the predicted values in predictive coding for multiple three-dimensional points on a first plane and different methods for determining the predicted values in predictive coding for multiple three-dimensional points on a second plane. Predictive coding is an interprediction method in which the points to be encoded in the frame to be encoded are predicted and coded using reference candidate points of a reference frame different from the frame to be encoded. The first plane is a plane that faces the first and second reference positions in a third direction and is perpendicular to the reference plane. The second plane is a plane that faces the first and second reference positions in a fourth direction and is perpendicular to the reference plane. The third direction and the fourth direction are different directions from each other. Some of the multiple three-dimensional points on the first plane are included in the multiple first three-dimensional points included in the reference frame. Some of the multiple three-dimensional points on the first plane are included in the multiple second three-dimensional points included in the frame to be encoded. Some of the multiple three-dimensional points on the second plane are included in the multiple first three-dimensional points included in the reference frame. Other portions of the multiple three-dimensional points on the second plane are included in the multiple second three-dimensional points included in the second frame.
[0108] Based on the above, the distance d from the second reference position of the frame to be encoded to the point to be encoded iscur Based on the point cloud obtained by the sensor 13845 of the reference frame at the first reference position, which is located at a distance m from the second reference position, prediction can be made with higher accuracy than the interpretation method explained using Figures 2 to 5, and the efficiency of interpretation coding can be further improved.
[0109] Furthermore, the trigonometric functions, division, and other arithmetic operations shown in Figure 7 may be simplified by using a table with a finite number of elements. Simplification can improve the efficiency of interpredictive coding while reducing the amount of processing required.
[0110] The predicted values d in the four directions (front, back, left, and right) explained using Figures 6 and 7. pred In addition to the four derivation methods mentioned above, four more derivation methods for four diagonal directions between the four directions may be added, so that the derivation method can be switched in eight directions, as shown in Figure 8. Figure 8 is a diagram illustrating a second example in which the derivation method of the predicted value is switched according to the value of the horizontal angle. Figure 8 is a plan view from the axial direction in the case of a sensor that measures the distance to an object by emitting a laser around a predetermined axis (rotation axis), such as a LiDAR sensor. Figure 9 shows the horizontal angle φ cur Predicted values d are determined for each of the eight directions defined by pred This figure shows the formula for deriving the expression. The Index in Figure 9 indicates that the hypothetical planes 13840 to 13843 set up in Figure 8 correspond to index values 0 to 7.
[0111] As shown in Figure 8, a virtual plane perpendicular to the horizontal plane (reference plane) is set as the object in eight directions: front, back, left, and right, plus four diagonal directions between these four directions, with a horizontal angle φ cur For each of the eight ranges divided by the formula shown in Figure 9, the predicted value d is calculated using the formula shown in Figure 9. pred This is derived. In other words, the three-dimensional data encoding device uses the horizontal angle φ of the point to be encoded. cur The horizontal angle φ shown in Figure 9 is obtained. cur Predicted value d based on the formula corresponding to this pred We derive this. Note that in the example in Figure 8, |φ curThe ranges are defined as |≦π, 0<α<π / 2, and 0<β<π / 2-α. α and β may be predetermined constants, such as α=π / 6 and β=π / 4. Furthermore, α and β may be included in the header information of predetermined data units such as sequences, frames, and slices. This allows α and β to be notified from the 3D data encoding device to the 3D data decoding device and made modifiable. Also, the inclusion of each range boundary in the encoding and decoding processes only needs to be consistent and does not necessarily have to be as shown in Figure 9.
[0112] Based on the above, the distance d from the second reference position of the frame to be encoded to the point to be encoded is cur This can be predicted with higher accuracy than the interpretation method described using Figures 2 to 7, based on the point cloud obtained by sensor 13845 of the reference frame at a first reference position located m away from sensor 13846, and the efficiency of interpretation coding can be further improved.
[0113] Furthermore, the trigonometric functions, division, and other arithmetic operations shown in Figure 9 may be simplified by using a table with a finite number of elements. Simplification may improve the efficiency of interpredictive coding while reducing the amount of processing required.
[0114] (Embodiment 2) This embodiment describes a method for predicting the positional information of three-dimensional points that contain three-dimensional data.
[0115] Three-dimensional data is, for example, point cloud data. A point cloud is a collection of multiple three-dimensional points that represent the three-dimensional shape of an object. Point cloud data includes positional and attribute information of multiple three-dimensional points. This positional information indicates the three-dimensional position of each point. Positional information is sometimes also called geometry information. For example, positional information can be expressed in a Cartesian coordinate system or a polar coordinate system.
[0116] Attribute information can include, for example, color, reflectance, or normal vector. A single three-dimensional point may have one attribute or multiple attribute pieces of information.
[0117] Note that the three-dimensional data is not limited to point cloud data; it may also be other types of three-dimensional data such as mesh data. Mesh data (also called three-dimensional mesh data) is a data format used in computer graphics (CG), and it represents the three-dimensional shape of an object as a collection of surface information. For example, mesh data includes point cloud information (e.g., vertex information). Therefore, the same methods as those used for point cloud data can be applied to this point cloud information.
[0118] Figure 10 is a block diagram showing the configuration of a three-dimensional data encoding device according to this embodiment. The three-dimensional data encoding device 100 supports interpredictive encoding, which encodes a point cloud to be encoded while referring to an already encoded point cloud. This three-dimensional data encoding device 100 includes an encoding unit 101, a motion compensation unit 102, a first buffer 103, a second buffer 104, a switching unit 105, and an interpredictive unit 106.
[0119] Although Figure 10 only shows the configuration for interpredictive coding of location information, the three-dimensional data coding device 100 may also include other processing units for coding location information (e.g., an intraprediction unit), or an attribute information coding unit for coding attribute information.
[0120] The encoding unit 101 generates a bitstream by encoding the input point cloud, which is the point cloud to be encoded. Specifically, the encoding unit 101 extracts a prediction tree (Predtree), which is a unit of encoding processing, from the target point cloud, and encodes each point in the prediction tree while referring to the interprediction points. The encoding unit 101 also outputs decoded points, which are reconstructed points obtained by decoding the bitstream. These decoded points are used for interpretation of subsequent target point clouds (for example, the target point cloud of a subsequent frame or slice).
[0121] Here, the position information of the target point cloud and the position information of the decoded points are expressed, for example, in polar coordinates. Alternatively, the position information of the target point cloud may be expressed in Cartesian coordinates, and the encoding unit 101 may convert the position information in Cartesian coordinates to position information in polar coordinates and encode the converted polar coordinate position information.
[0122] The motion compensation unit 102 performs motion compensation on the decoded points and stores the motion-compensated reference point group (first reference point group) in the first buffer 103. For example, the motion compensation unit 102 performs motion compensation by projecting the position information of the decoded points onto the polar coordinates of the target point group, using an interpretation method in a polar coordinate system as described in Embodiment 1.
[0123] Furthermore, motion compensation is the process of correcting the positional information of the decoded point to match the coordinate system of the target three-dimensional point being encoded. Specifically, motion compensation may include at least one of the following: aligning the origin of the coordinate system of the decoded point with the origin of the coordinate system of the target three-dimensional point, and aligning each axis of the coordinate system of the decoded point with each axis of the coordinate system of the target three-dimensional point. Motion compensation may also include coordinate calculations using translation vectors and rotation matrices.
[0124] Furthermore, the decoded points (a second group of reference points that have not been motion-compensated) are stored in the second buffer 104. The switching unit 105 selects either the first group of reference points stored in the first buffer 103 or the second group of reference points stored in the second buffer 104 as the inter-reference point (third group of reference points) and outputs the inter-reference point to the inter-prediction unit 106.
[0125] The inter-prediction unit 106 determines the inter-prediction point by referring to at least one inter-reference point stored in the first buffer 103 or the second buffer 104. For example, the inter-prediction unit 106 refers to one or more inter-reference points that are in the same or close position as the target point, from among a plurality of inter-reference points included in a reference frame different from the target frame that includes the target point cloud. Here, an inter-reference point that is in the same or close position as the target point is, for example, a point whose elevation angle index and horizontal angle index are the same or close as the target point (for example, an index value that is 1 greater or 1 less). In other words, as a method for determining the inter-prediction point in the inter-prediction unit 106, one three-dimensional point in the reference point cloud may be selected as the prediction point, or the prediction point may be calculated from a plurality of three-dimensional points in the reference point cloud. For example, the average position of a plurality of three-dimensional points may be calculated as the position of the prediction point.
[0126] Figure 11 is a block diagram showing the configuration of a three-dimensional data decoding device according to this embodiment. The three-dimensional data decoding device 200 supports interpredictive decoding, which decodes the point cloud to be decoded while referring to the decoded point cloud. This three-dimensional data decoding device 200 includes a decoding unit 201, a motion compensation unit 202, a first buffer 203, a second buffer 204, a switching unit 205, and an interpredictive unit 206.
[0127] Although Figure 11 only shows the configuration for interprediction and decoding of location information, the three-dimensional data decoding device 200 may also include other processing units for decoding location information (e.g., an intraprediction unit), or an attribute information decoding unit for decoding attribute information.
[0128] The decoding unit 201 generates a set of decoded points by decoding the input bitstream. Specifically, the decoding unit 201 decodes each point in the prediction tree while referring to the interpredicted points and outputs the obtained decoded points. The operation of the motion compensation unit 202, the first buffer 203, the second buffer 204, the switching unit 205, and the interprediction unit 206 is the same as the operation of the motion compensation unit 102, the first buffer 103, the second buffer 104, the switching unit 105, and the interprediction unit 106 included in the three-dimensional data encoding device 100 shown in Figure 10.
[0129] Thus, by using a motion-compensated first reference point cloud in interpredictive coding, it becomes possible to accurately predict the positional information of structures such as buildings and walls around a movement path, such as a road, when coding point clouds acquired by a sensor such as LiDAR while it is moving. Therefore, there is a possibility of improving the efficiency of interpredictive coding. Furthermore, by making both the motion-compensated first reference point cloud and the motion-uncompensated second reference point cloud accessible in interpredictive coding, it becomes possible to accurately predict not only the positional information of structures around the movement path, but also the positional information of points at a nearly constant distance from the sensor, such as objects moving at the same speed as the sensor or the ground around the sensor. Therefore, there is a possibility of further improving the efficiency of interpredictive coding.
[0130] Next, an example of the interpretation process procedure will be described. Figure 12 is a flowchart showing an example of the process procedure for encoding or decoding frames to which interpretation processing is applied in the three-dimensional data encoding device 100 and three-dimensional data decoding device 200 shown in Figures 10 and 11. Note that the process shown in Figure 12 may be repeated for each frame, or it may be repeated for each processing unit (e.g., slice) into which the frame is divided.
[0131] In this example, first, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 acquire motion information regarding the displacement between the coordinates of the processed point cloud that has been encoded or decoded and the coordinates of the target point cloud to be encoded or decoded (S101). For example, the three-dimensional data encoding device 100 uses alignment techniques such as the ICP (Iterative Closest Point) algorithm to detect displacements such as rotation and / or translation between the coordinates of the processed point cloud and the coordinates of the target point cloud, and determines motion information based on the detected displacement. The three-dimensional data encoding device 100 also stores the motion information in the higher-level syntax of the bitstream (SPS, GPS, or slice header, etc.).
[0132] Note that SPS (Sequence Parameter Set) is metadata (parameter set) common to multiple frames. GPS (Geometry Parameter Set) is metadata (parameter set) related to the encoding of location information. For example, GPS is metadata common to multiple frames.
[0133] Furthermore, motion information includes, for example, information about movement parallel to the horizontal plane and information about rotation around the vertical axis. Specifically, motion information includes, for example, a 3x1 translation matrix. Alternatively, motion information includes the absolute value (|mv| described later) and direction (angle α described later) of a translation vector parallel to the horizontal plane. Alternatively, motion information includes, for example, a 3x3 rotation matrix. Alternatively, motion information indicates the rotation angle of the coordinate axis on the horizontal plane (angle β described later) and rotation around the vertical axis.
[0134] The three-dimensional data decoding device 200 acquires motion information from the bitstream and sets displacements such as rotation and / or translation between the coordinates of the processed point cloud and the coordinates of the target point cloud based on the acquired motion information.
[0135] Next, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 project at least a portion of the first processed point cloud onto the coordinates of the target point cloud according to the motion information, and set the resulting point cloud as the first reference point cloud (S102). As a method for projecting the first processed point cloud onto the coordinates of the target point cloud, the inter-prediction method in the polar coordinate system described with Figures 1 to 9 of Embodiment 1 may be used. Here, the first processed point cloud is a processed point cloud contained in one frame or one slice. The first processed point cloud may also be a processed point cloud contained in multiple frames or multiple slices.
[0136] Furthermore, the above projection may be applied to all of the distance, horizontal angle, and elevation angle components included in the position information, or to only some of them. For example, only the distance and horizontal angle components may be changed to values projected from the first processed point cloud to the coordinates of the target point cloud, while the elevation angle component may remain unchanged from the value of the first processed point cloud. This reduces the amount of processing required.
[0137] Next, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 set a point cloud as the second reference point cloud, using the coordinate information of the second processed point cloud as is for at least a portion of the second processed point cloud (S103). Step S103 may be performed before step S101 or S102.
[0138] The first and second processed point clouds may be contained within the same processing unit (same frame or slice, etc.) or within different processing units. For example, the first and second processed point clouds may be the corrected (projected) and uncorrected point clouds of the same processing unit. For example, the point cloud in this processing unit may be the point cloud whose time is closest to the target point cloud. Alternatively, the second processed point cloud may be the point cloud whose time is closest to the target point cloud, and the first processed point cloud may be a different point cloud with a different time from the second processed point cloud.
[0139] The time points used for the first and second processed point clouds may be fixedly set in advance. For example, it may be fixedly set that the point cloud whose time is closest to the target point cloud is used for both the first and second processed point clouds. Alternatively, the three-dimensional data encoding device 100 may determine which time points to use for the first and second processed point clouds and store information indicating this determination in the higher syntax of the bitstream (SPS, GPS, or slice header, etc.). For example, this information may indicate the time distance between the target point cloud and the point cloud used as the first processed point cloud, and the time distance between the target point cloud and the point cloud used as the second processed point cloud. If a common point cloud is used for both the first and second processed point clouds, this information may indicate that common point cloud. For example, this information may be set for each processing unit (e.g., slice or frame).
[0140] This allows the three-dimensional data encoding device 100 to select processed point clouds that are suitable for the characteristics of the target point cloud, such as its time evolution, potentially improving encoding efficiency.
[0141] Next, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 determine an inter-predicted point cloud for each point, and encode or decode the target point by referring to the inter-predicted points (S104~S107).
[0142] First, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 start loop processing for each target point included in the target point cloud (S104). In other words, one of the multiple points included in the target point cloud is selected as the target point to be processed.
[0143] Next, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 determine the inter-prediction point corresponding to the target point by referring to at least a portion of the reference point group, which includes the first reference point group and the second reference point group (S105). For example, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 refer to one or more inter-reference points that are in the same or close position as the target point, from among a plurality of inter-reference points included in a reference frame different from the target frame that includes the target point group. Here, an inter-reference point that is in the same or close position as the target point is, for example, a point that has the same or close elevation index and horizontal angle index as the target point (for example, an index value that is 1 greater or 1 less).
[0144] For example, the three-dimensional data encoding device 100 compares the code amount (residual) when using the inter-prediction point determined using the first processed point cloud with the code amount (residual) when using the inter-prediction point determined using the second processed point cloud, and selects (references) the one with the smaller code amount. Alternatively, the three-dimensional data encoding device 100 may refer to the first and second processed point clouds to determine the inter-prediction point with the smallest code amount.
[0145] Furthermore, whether to use the first processed point cloud or the second processed point cloud may be determined according to the characteristics of the target point or target point cloud. In addition, the three-dimensional data encoding device 100 may store information indicating whether to use the first processed point cloud or the second processed point cloud, or information indicating the interpretation point, in a bitstream, and the three-dimensional data decoding device 200 may refer to this information to decide whether to use the first processed point cloud or the second processed point cloud, or to determine the interpretation point.
[0146] Next, the three-dimensional data encoding device 100 encodes the target point by referring to the interpretation prediction point (S106). Specifically, the three-dimensional data encoding device 100 calculates the residual (difference) between the position information of the target point and the position information of the interpretation prediction point. The three-dimensional data encoding device 100 generates encoded position information by performing quantization and entropy encoding on the obtained residual. Note that the residuals of some of the multiple components of the position information (e.g., distance, elevation angle, horizontal angle) may be calculated, and the original values of the other components may be quantized and entropy encoded as they are. The three-dimensional data encoding device 100 also generates a bitstream containing the encoded position information.
[0147] In the three-dimensional data decoding device 200, the target point is decoded by referring to this interpretation prediction point. Specifically, the three-dimensional data decoding device 200 obtains the encoded position information of the target point from the bitstream. The three-dimensional data decoding device 200 generates the residual of the target point by performing entropy decoding and inverse quantization on the encoded position information of the target point. The three-dimensional data decoding device 200 generates the position information of the target point by adding the residual of the target point and the position information of the interpretation prediction point.
[0148] Next, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 complete the loop processing for the target points (S107). That is, steps S105 and S106 are performed for each of the multiple points included in the target point group.
[0149] Furthermore, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 do not always need to refer to interprediction points to encode or decode target points. For example, the three-dimensional data encoding device 100 may store switching information in the bitstream for each node or slice indicating whether or not to refer to interprediction points. In this case, the three-dimensional data decoding device 200 can switch whether or not to refer to interprediction points based on this information. If interprediction points are not to be referred to, step S105 may be omitted. By making it possible to switch whether or not to refer to interprediction points, the three-dimensional data encoding device 100 can select an encoding method suitable for the characteristics of the target point cloud, such as its time evolution, and potentially improve encoding efficiency.
[0150] Next, an example of an interpretation method in a polar coordinate system will be described. Figure 13 is a diagram illustrating an example of an interpretation method when performing encoding or decoding using polar coordinates. In the motion compensation unit 102 shown in Figure 10 and the motion compensation unit 202 shown in Figure 11, or in step S102 shown in Figure 12, the interpretation method described below may be used instead of the interpretation method in a polar coordinate system described in Embodiment 1. In other words, Figure 13 shows an example of motion compensation.
[0151] In Figure 13, the motion vector mv is the horizontal displacement from the polar coordinate origin of the target frame (the frame being processed) to the polar coordinate origin of the reference frame. Angle α (alpha) is the angle between the horizontal angle of the target frame and the reference direction of the motion vector mv. Angle β (beta) is the angle between the horizontal angle of the target frame and the reference direction of the horizontal angle of the reference frame.
[0152] In this case, when point n in the reference frame is used for interpretation in the encoding of the target frame, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 predict the horizontal angle φ in polar coordinates of the target frame. ref2 (n) and the predicted distance d ref2(n) may also be determined using the following equations (1) and (2). Note that |mv| is the magnitude of the motion vector mv.
[0153] φ ref2 (n) = arctan(d ref1 (n)sin(φ ref1 (n)-α+β) / (d ref1 (n)cos(φ ref1 (n)-α+β)+|mv|))+α...(Formula 1) d ref2 (n) = d ref1 (n)sin(φ ref1 (n)-α+β) / sin(φ ref2 (n)-α) (Formula 2)
[0154] Furthermore, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 use the following (Equation 3) instead of (Equation 2) d ref2 (n) may be determined. Note that if the denominator of the division in either (Equation 2) or (Equation 3) becomes 0 (zero), the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 will use the other equation to determine d ref2 (n) may be determined.
[0155] d ref2 (n) = (d ref1 (n)cos(φ ref1 (n)-α+β)+|mv|) / cos(φ ref2 (n)-α) (Formula 3)
[0156] As a result, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 can predict the distance from the polar coordinate origin of the target frame to the target point with high accuracy. This may improve the efficiency of interpredictive coding.
[0157] Note that the predicted elevation angle θ in polar coordinates for the target frame is also shown. ref2 (n) may be determined using (Equation Z5) shown in Embodiment 1. This may further improve the efficiency of interpredictive coding. Alternatively, the elevation angle θ in polar coordinates of the reference frame.ref1 (n) or its index information directly predicts the elevation angle θ ref2 (n) or its index information may be used. This may improve the efficiency of interpredictive coding while reducing the amount of processing required.
[0158] Note that instead of arctan in (Equation 1), a method that can derive angles in the second and third quadrants, such as atan2(y,x) in the C language, may be used. For example, instead of (Equation 1), the following (Equation 4) using atan2(y,x) in the C language may be used.
[0159] φ ref2 (n) = atan2(d ref1 (n)sin(φ ref1 (n)-α+β),d ref1 (n)cos(φ ref1 (n)-α+β)+|mv|)+α (Equation 4)
[0160] However, since atan2(0,0) is undefined, a constant such as atan2(0,0)=0 may be set in this case.
[0161] Furthermore, the angle β between the motion vector mv and the reference direction of the horizontal angle of the reference frame with respect to the reference direction of the horizontal angle of the target frame may be determined based on results obtained using at least one of the following: (1) measurements using a sensor such as GPS (Global Positioning System) or an odometer; (2) self-localization techniques such as SfM (Structure from Motion) or SLAM (Simultaneous Localization and Mapping); and (3) alignment techniques such as the ICP (Iterative Closest Point) algorithm.
[0162] Furthermore, the three-dimensional data encoding device 100 may store as motion information the motion vector mv and the angle β between the reference direction of the horizontal angle of the target frame and the reference direction of the horizontal angle of the reference frame, or information regarding these values, in the header information of a unit such as a frame or slice. Alternatively, instead of the motion vector mv, the three-dimensional data encoding device 100 may store the magnitude of mv |mv| and the angle α (alpha) between mv and the reference direction of the horizontal angle of the target frame in the header information.
[0163] Figure 14 is a flowchart of the interpretation prediction process. Figure 14 also shows the predicted value φ of the horizontal angle in polar coordinates of the target frame in the interpretation prediction method described using Figure 13. ref2 (n) and the predicted distance d ref2 This is an example of the procedure for deriving (n). In this example, the horizontal angle takes a value within the range of -π to π. For example, in steps S121 and S129 shown in Figure 14, if the result is less than -π, 2π is added to the result, and if the result is greater than π, 2π is subtracted from the result.
[0164] Furthermore, the operation of the three-dimensional data encoding device 100 will be described below, but the operation of the three-dimensional data decoding device 200 is similar.
[0165] In this procedure, first, the three-dimensional data encoding device 100 performs φ' ref1 (n) φ ref1 Set it to (n)-α+β (S121).
[0166] Next, the three-dimensional data encoding device 100 performs φ' ref2 (n) = φ ref2 (n)-α is the case for φ' ref2 Determine which angle range (n) is in, and use one of the following equations (5) to (8) to determine φ' ref2 Derive (n).
[0167] φ' ref2 (n) = arctan(d ref1 (n)sin(φ' ref1 (n) / (d ref1(n)cos(φ' ref1 (n)) + |mv|)) ···(Equation 5) φ' ref2 (n) = arctan(d ref1 (n)sin(φ' ref1 (n) / (d ref1 (n)cos(φ' ref1 (n))+|mv|))+sign(φ' ref1 (n))×π (Equation 6) φ' ref2 (n) = sign(φ' ref1 (n))×π / 2 (Formula 7) φ' ref2 (n)=0...(Formula 8)
[0168] However, sign(x) is 1 for x > 0, 0 for x == 0, and -1 for x < 0.
[0169] Specifically, the three-dimensional data encoding device 100 is d ref1 (n)cos(φ' ref1 If (n) + |mv| > 0 (True in S122), then -π / 2 < φ' ref2 (n) < π / 2, and by executing (Equation 5), we get φ' ref2 Find (n) (S123).
[0170] On the other hand, d ref1 (n)cos(φ' ref1 (n))+|mv|>0 is not satisfied (False in S122), and d ref1 (n)cos(φ' ref1 If (n) + |mv| < 0 (True in S124), the three-dimensional data encoding device 100 will perform φ' ref2 (n) < -π / 2 or π / 2 < φ' ref2 Determine that (n), and then execute (Equation 6) to get φ' ref2 Find (n) (S125).
[0171] On the other hand, d ref1 (n)cos(φ' ref1(n))+|mv|>0 is not satisfied (No in S122), and d ref1 (n)cos(φ' ref1 (n))+|mv|<0 is not satisfied (No in S124), and sin(φ' ref1 (n))≠0 holds (Yes in S126), the three-dimensional data encoding apparatus 100 sets φ' ref2 (n)=-π / 2 or φ' ref2 (n)=π / 2, and obtains φ' ref2 (n) by executing (Equation 7) (S127).
[0172] On the other hand, when all the above three conditions are No (No in S126), the three-dimensional data encoding apparatus 100 determines that tan is undefined, and sets a constant (e.g., 0) to φ' ref2 (n) by executing (Equation 8) (S128).
[0173] Note that in (Equation 5) and (Equation 6), arctan(x) is defined in the range of -π / 2<x<π / 2. Also, in S123, S125 and S127, instead of (Equation 5) to (Equation 7), (Equation 8) to which φ' ref1 (n)=φ ref1 (n)-α+β and φ' ref2 (n)=φ ref2 (n)-α are applied may be used.
[0174] Next, the three-dimensional data encoding apparatus 100 uses φ' ref2 (n) obtained above to set φ ref2 (n) to φ' ref2 (n)+α (S129).
[0175] Next, the three-dimensional data encoding apparatus 100 determines which angular range φ' ref1 (n) falls within, and derives d ref2 (n) using (Equation 9) or (Equation 10).
[0176] d ref2 (n)=d ref1 (n)sin(φ' ref1(n) / sin(φ' ref2 (n)) (Formula 9) d ref2 (n) = (d ref1 (n)cos(φ' ref1 (n))+|mv|) / cos(φ' ref2 (n))...(Formula 10)
[0177] Specifically, sin(φ' ref2 If (n) is not 0 (True in S130), the three-dimensional data encoding device 100 determines that the reference point is not located on the line containing the motion vector mv, and executes (Equation 9) to determine d ref2 Find (n) (S131). On the other hand, sin(φ' ref2 If (n) is 0 (False in S130), the three-dimensional data encoding device 100 determines that the reference point is located on a straight line containing the motion vector mv, and executes (Equation 10) d ref2 Find (n) (S132).
[0178] The following describes some modifications. In the apparatus, processing, or syntax disclosed using Figures 10 to 14, if both the motion vector mv, which indicates the horizontal displacement from the polar coordinate origin of the target frame (processing target frame) to the polar coordinate origin of the reference frame, and the angle β (beta) between the reference direction of the horizontal angle of the reference frame and the reference direction of the horizontal angle of the target frame are smaller than a predetermined value (i.e., the motion is small), the three-dimensional data encoding device 100 may use only the reference point cloud without motion compensation for interpretation and omit storing accompanying information, such as information regarding the selection of interreference points, in the bitstream. Alternatively, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 may use a point cloud from a different time that does not have motion compensation as the reference point cloud instead of the motion-compensated reference point cloud. This may allow for maintaining or improving encoding efficiency while suppressing the amount of processing required when encoding a scene in which a moving object equipped with a sensor is stationary.
[0179] Furthermore, the three-dimensional data encoding device 100 may store information indicating whether or not to use motion-compensated reference point clouds for interpretation in the higher-level syntax (SPS, GPS, or slice header, etc.). Also, if the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 do not use motion-compensated reference point clouds for interpretation, they may use one or more non-motion-compensated reference point clouds for interpretation. This increases the flexibility of the operation of the three-dimensional data encoding device 100 and the design of the encoding algorithm, potentially improving the operability and encoding efficiency of the device.
[0180] Furthermore, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 may project all inter-reference point candidates included in the reference frame onto the polar coordinate origin of the target frame, or they may project only some points onto the polar coordinate origin of the target frame. For example, these some points may be points whose elevation angle with respect to the polar coordinate origin of the reference frame is above (larger than) a predetermined value, or points whose vertical (z-axis direction) position when the reference frame is represented in Cartesian coordinates is above a predetermined position (e.g., the ground). This makes it possible to narrow down the processing target to structures such as buildings or walls where a significant improvement in encoding efficiency can be expected through projection. Therefore, it is possible to maintain or improve encoding efficiency while suppressing the amount of processing and memory usage. Alternatively, the horizontal angle with respect to the polar coordinate origin of the reference frame may be divided into predetermined angles (quantized), and inter-reference point candidates may be limited to only representative points in each angle interval. This may further reduce the amount of processing and memory usage.
[0181] Furthermore, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 may rotate and / or translate the processed point cloud in the orthogonal coordinate space to obtain coordinates of the processed point cloud with respect to the target point cloud in the orthogonal coordinate space, further convert the obtained coordinates of the processed point cloud into coordinates of the target point cloud in the polar coordinate space, and set the obtained point cloud as a motion-compensated reference point cloud. This allows the motion compensation method to be shared between encoding of point clouds represented in orthogonal coordinates using an octree or a prediction tree and encoding of point clouds represented in polar coordinates. Therefore, the structures of the three-dimensional data encoding device and the three-dimensional data decoding device can be simplified, which potentially makes it possible to reduce the scale of circuits or software.
[0182] Furthermore, arithmetic processing such as trigonometric functions and division in inter prediction may be simplified by using approximate arithmetic based on integer precision processing, or by using a table consisting of a finite number of elements. These simplifications potentially make it possible to improve the efficiency of point cloud encoding while reducing the amount of processing and memory required for projection.
[0183] Furthermore, at least part of the above-described device, processing or syntax may be used for encoding vertex information of a three-dimensional mesh. This allows processing to be shared between point cloud encoding and three-dimensional mesh encoding, which potentially makes it possible to reduce the scale of circuits or software.
[0184] As described above, the three-dimensional data encoding device according to the present embodiment performs the processing illustrated in Fig. 15. The three-dimensional data encoding device generates a first reference point cloud by correcting (motion compensating) position information of one or more first three-dimensional points to match the coordinate system of a target three-dimensional point to be encoded (S201), selects either the first reference point cloud or a second reference point cloud including the one or more first three-dimensional points before correction as a third reference point cloud for the target three-dimensional point (S202), determines a prediction point using the third reference point cloud (S203), and encodes the position information of the target three-dimensional point with reference to at least part of the position information of the prediction point (for example, at least part of a plurality of components included in the position information) (S204).
[0185] Alternatively, the three-dimensional data encoding device may determine the predicted points for the target three-dimensional point from the first and second reference point groups instead of performing steps S202 and S203.
[0186] According to this, the three-dimensional data encoding device uses a corrected first set of reference points and an uncorrected second set of reference points to encode the target points. Therefore, the three-dimensional data encoding device may be able to determine prediction points that minimize prediction errors. Thus, the three-dimensional data encoding device can improve encoding efficiency. In addition, the three-dimensional data encoding device can reduce the amount of data handled in the encoding process.
[0187] For example, in the correction (S201), the three-dimensional data encoding device adjusts the position information of one or more first three-dimensional points to the coordinate system of the target three-dimensional point based on first information (e.g., motion information) indicating the displacement between the coordinate system of one or more first three-dimensional points and the coordinate system of the target three-dimensional point.
[0188] For example, in the correction (S201), the three-dimensional data encoding device derives positional information for one or more second three-dimensional points included in the first reference point group by projecting one or more first three-dimensional points onto the coordinate origin of the target three-dimensional point according to its displacement.
[0189] For example, the first information includes at least one of the second information relating to movement parallel to the horizontal plane and the third information relating to rotation around a vertical axis.
[0190] For example, position information includes a distance component, a horizontal angle component, and an elevation angle component, and the three-dimensional data encoding device corrects at least one of the distance component and the horizontal angle component in the correction (S201). This allows the three-dimensional data encoding device to efficiently correct three-dimensional data obtained from a sensor moving in the horizontal direction. In other words, in this embodiment, the elevation angle component in polar coordinates is not corrected. Therefore, this embodiment is suitable when selectively using a reference point cloud with horizontally corrected position and a reference point cloud before correction. For example, this embodiment is suitable for a three-dimensional point cloud obtained from a sensor that repeatedly moves and stops in the horizontal direction.
[0191] For example, the three-dimensional data encoding device further determines whether or not to perform correction and generates a bitstream containing the location information of the encoded target three-dimensional point and a fourth piece of information indicating whether or not to perform correction. This allows the three-dimensional data encoding device to determine prediction points that minimize prediction errors by switching whether or not to perform correction. For example, the three-dimensional data encoding device can select an appropriate method depending on the characteristics of the point cloud being processed. The fourth piece of information may be a flag or a parameter. Furthermore, the fourth piece of information may be provided for each frame, for each processing unit (e.g., slice) within a frame, or for each point.
[0192] For example, if no correction is performed, the three-dimensional data encoding device will select the second reference point group as the third reference point group.
[0193] For example, if one or more first three-dimensional points are included in a first processing unit (e.g., a frame or slice), the three-dimensional data encoding device, when no correction is performed, selects either a second group of reference points or a fourth group of reference points consisting of one or more third three-dimensional points included in a second processing unit different from the first processing unit, which are the one or more third three-dimensional points before correction, as the third group of reference points. In this case, the three-dimensional data encoding device can reference two uncorrected processing units (e.g., two frames) when no correction is performed. Therefore, the three-dimensional data encoding device can improve encoding efficiency.
[0194] For example, the first set of reference points is generated by correcting one or more fourth three-dimensional points, which are part of one or more first three-dimensional points. This allows the three-dimensional data encoding device to reduce processing load by limiting the number of three-dimensional points to be corrected. For instance, if the relative position of the target three-dimensional point and the origin is approximately equal to the relative position of the predicted points included in the reference point set and the origin, using those predicted points without correction can reduce prediction errors. On the other hand, if the relative position of the target three-dimensional point and the origin differs from the relative position of the predicted points included in the reference point set and the origin, using the corrected predicted points can reduce prediction errors. Thus, by switching whether or not to perform correction depending on the position of the target three-dimensional point, prediction errors can be reduced.
[0195] For example, the position information includes a distance component, a horizontal angle component, and an elevation angle component, and the fourth three-dimensional point is a first three-dimensional point among one or more first three-dimensional points whose elevation angle component is larger than a predetermined value. In other words, in this embodiment, the three-dimensional points to be corrected are limited to three-dimensional points with a large elevation angle component. Three-dimensional points with a large elevation angle component represent, for example, buildings. Buildings are fixed to the ground. Therefore, when the target three-dimensional point and the predicted point represent a building, the relative positional relationship between the target three-dimensional point and the origin differs from the relative positional relationship between the predicted point included in the reference point group and the origin. In this case, using the corrected predicted point can suppress prediction errors. Therefore, the three-dimensional data encoding device limits the targets to buildings, etc.
[0196] Similarly, the fourth three-dimensional point is one or more first three-dimensional points whose vertical position is higher than a predetermined position.
[0197] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0198] Furthermore, the three-dimensional data decoding device according to this embodiment performs the processing shown in Figure 16. The three-dimensional data decoding device generates a first reference point group by correcting (motion compensation) the position information of one or more first three-dimensional points to match the coordinate system of the target three-dimensional point to be decoded (S211), selects either the first reference point group or the second reference point group containing one or more first three-dimensional points before correction as the third reference point group for the target three-dimensional point (S212), determines the predicted point using the third reference point group (S213), and decodes the position information of the target three-dimensional point by referring to at least a part of the position information of the predicted point (for example, at least a part of the multiple components included in the position information) (S214).
[0199] Alternatively, the three-dimensional data decoding device may determine the predicted points for the target three-dimensional point from the first and second reference point groups instead of steps S212 and S213.
[0200] According to this, the three-dimensional data decoding device uses a corrected first set of reference points and an uncorrected second set of reference points to decode the target points. Therefore, the three-dimensional data decoding device may be able to determine prediction points that minimize prediction errors. Thus, the three-dimensional data decoding device can reduce the amount of data handled in the decoding process.
[0201] For example, in the correction (S211), the three-dimensional data decoding device adjusts the position information of one or more first three-dimensional points to the coordinate system of the target three-dimensional point based on first information (e.g., motion information) indicating the displacement between the coordinate system of one or more first three-dimensional points and the coordinate system of the target three-dimensional point.
[0202] For example, in the correction (S211), the three-dimensional data decoding device derives the positional information of one or more second three-dimensional points included in the first reference point group by projecting one or more first three-dimensional points onto the coordinate origin of the target three-dimensional point according to its displacement.
[0203] For example, the first information includes at least one of second information relating to movement parallel to a horizontal plane and third information relating to rotation about a vertical axis. According to this, the three-dimensional data decoding device can efficiently correct three-dimensional data obtained by a sensor that moves in the horizontal direction.
[0204] For example, the position information includes a distance component, a horizontal angle component, and an elevation angle component, and in the correction (S211), the three-dimensional data decoding device corrects at least one of the distance component and the horizontal angle component. That is, in this aspect, the elevation angle component in polar coordinates is not corrected. Therefore, this aspect is suitable for a case where a reference point group whose position has been corrected in the horizontal direction and an uncorrected reference point group are selectively used. For example, this aspect is suitable for a three-dimensional point cloud obtained by a sensor that repeatedly moves and stops in the horizontal direction.
[0205] For example, the three-dimensional data decoding device further acquires fourth information indicating whether to perform correction from the bitstream, and determines whether to perform correction based on the fourth information. According to this, the three-dimensional data decoding device can determine a prediction point that reduces a prediction error by switching whether to perform correction. Note that the fourth information may be a flag or a parameter. Further, the fourth information may be provided for each frame, may be provided for each processing unit (e.g., slice) in a frame, or may be provided for each point.
[0206] For example, when not performing correction, the three-dimensional data decoding device selects a second reference point cloud as a third reference point cloud.
[0207] For example, if one or more first three-dimensional points are included in a first processing unit (e.g., a frame or slice), the three-dimensional data decoder, when no correction is performed, selects either a second group of reference points or a fourth group of reference points consisting of one or more third three-dimensional points included in a second processing unit different from the first processing unit, which are the one or more third three-dimensional points before correction, as the third group of reference points. In this case, the three-dimensional data decoder can reference two uncorrected processing units (e.g., two frames) when no correction is performed. Therefore, the three-dimensional data decoder can improve encoding efficiency.
[0208] For example, the first set of reference points is generated by correcting one or more fourth three-dimensional points, which are part of one or more first three-dimensional points. This allows the three-dimensional data decoding device to reduce processing load by limiting the number of three-dimensional points to be corrected. For instance, if the relative position of the target three-dimensional point and the origin is approximately equal to the relative position of the predicted points included in the reference point set and the origin, using those predicted points without correction can reduce prediction errors. On the other hand, if the relative position of the target three-dimensional point and the origin differs from the relative position of the predicted points included in the reference point set and the origin, using the corrected predicted points can reduce prediction errors. Thus, by switching whether or not to perform correction depending on the position of the target three-dimensional point, prediction errors can be reduced.
[0209] For example, the position information includes distance, horizontal angle, and elevation angle components, and the fourth three-dimensional point is a first three-dimensional point among one or more first three-dimensional points whose elevation angle component is greater than a predetermined value. In other words, in this embodiment, the three-dimensional points to be corrected are limited to three-dimensional points with a large elevation angle component. Three-dimensional points with a large elevation angle component represent, for example, buildings. Buildings are fixed to the ground. Therefore, when the target three-dimensional point and the predicted point represent a building, the relative positional relationship between the target three-dimensional point and the origin differs from the relative positional relationship between the predicted point included in the reference point group and the origin. In this case, using the corrected predicted point can reduce the prediction error. Therefore, the three-dimensional data decoding method limits the targets to buildings, etc.
[0210] Similarly, the fourth three-dimensional point is one or more first three-dimensional points whose vertical position is higher than a predetermined position.
[0211] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0212] Although embodiments and modified examples of the three-dimensional data encoding device and three-dimensional data decoding device, etc., of the present disclosure have been described above, the present disclosure is not limited to these embodiments.
[0213] Furthermore, each processing unit included in the three-dimensional data encoding device and the three-dimensional data decoding device, etc., according to the above embodiment is typically implemented as an integrated circuit (LSI). These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip.
[0214] Furthermore, integrated circuit implementation is not limited to LSIs; it may also be achieved using dedicated circuits or general-purpose processors. Field-Programmable Gate Arrays (FPGAs), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow for the reconfiguration of the connections and settings of circuit cells within the LSI, may also be used.
[0215] Furthermore, in each of the above embodiments, each component may be implemented by being composed of dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0216] Furthermore, this disclosure may be implemented as a three-dimensional data encoding method or a three-dimensional data decoding method, etc., performed by a three-dimensional data encoding device and a three-dimensional data decoding device, etc.
[0217] Furthermore, the division of functional blocks in the block diagram is just one example; multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple parts, or some functions can be moved to other functional blocks. In addition, the functions of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.
[0218] Furthermore, the order in which each step in the flowchart is performed is illustrative for the purpose of specifically illustrating this disclosure, and may be in a different order. Also, some of the above steps may be performed simultaneously (in parallel) with other steps.
[0219] Although a three-dimensional data encoding device and a three-dimensional data decoding device, etc., relating to one or more embodiments have been described above based on embodiments, this disclosure is not limited to these embodiments. Without departing from the spirit of this disclosure, various modifications that a person skilled in the art can conceive of may be applied to these embodiments, and forms constructed by combining components from different embodiments may also be included within the scope of one or more embodiments. [Industrial applicability]
[0220] This disclosure is applicable to three-dimensional data encoding devices and three-dimensional data decoding devices. [Explanation of Symbols]
[0221] 100 3D data encoding device 101 Encoding section 102, 202 Motion compensation unit 103, 203 First buffer 104, 204 Second buffer 105, 205 switching section 106, 206 Interpretation Unit 200 Three-dimensional data decoding device 201 Decoding section
Claims
1. Motion compensation is performed to generate a first set of reference points by correcting the positional information of one or more first three-dimensional points to match the coordinate system of the target three-dimensional point to be encoded. A predicted point of the target three-dimensional point is selected from either the first group of reference points or the second group of reference points which includes the one or more first three-dimensional points including the position information before correction. The position information of the target three-dimensional point is encoded by referring to at least a portion of the position information of the predicted point. Three-dimensional data encoding method.
2. In the correction described above, the positional information of the one or more first three-dimensional points is aligned with the coordinate system of the target three-dimensional point based on first information indicating the displacement between the coordinate system of the one or more first three-dimensional points and the coordinate system of the target three-dimensional point. The three-dimensional data encoding method according to claim 1.
3. In the correction described above, the positional information of one or more second three-dimensional points included in the first reference point group is derived by projecting the one or more first three-dimensional points onto the coordinate origin of the target three-dimensional point according to the displacement. The three-dimensional data encoding method according to claim 2.
4. The first information includes at least one of second information relating to movement parallel to the horizontal plane and third information relating to rotation about a vertical axis. The three-dimensional data encoding method according to claim 2.
5. The position information includes a distance component, a horizontal angle component, and an elevation angle component. The correction involves correcting at least one of the distance component and the horizontal angle component. The three-dimensional data encoding method according to claim 1.
6. The aforementioned three-dimensional data encoding method further, Decide whether or not to make the aforementioned correction, A bitstream is generated that includes the encoded position information of the target three-dimensional point and a fourth piece of information indicating whether or not to perform the correction. The three-dimensional data encoding method according to claim 1.
7. If the above correction is not performed, the predicted point is selected from the second group of reference points. The three-dimensional data encoding method according to claim 6.
8. The one or more first three-dimensional points are included in the first processing unit. If the correction is not performed, the predicted point is selected from the second reference point group and a third reference point group which is one or more third three-dimensional points included in a second processing unit different from the first processing unit, and which include position information before correction. The three-dimensional data encoding method according to claim 6.
9. The first group of reference points is generated by correcting one or more fourth three-dimensional points, which are part of the one or more first three-dimensional points. The three-dimensional data encoding method according to claim 1.
10. The position information includes a distance component, a horizontal angle component, and an elevation angle component. The fourth three-dimensional point is a first three-dimensional point among the one or more first three-dimensional points whose elevation angle component is larger than a predetermined value. The three-dimensional data encoding method according to claim 9.
11. The fourth three-dimensional point is a first three-dimensional point among the one or more first three-dimensional points whose vertical position is higher than a predetermined position. The three-dimensional data encoding method according to claim 9.
12. Motion compensation is performed to generate a first set of reference points by correcting the positional information of one or more first three-dimensional points to match the coordinate system of the target three-dimensional point to be decoded. A predicted point of the target three-dimensional point is selected from either the first group of reference points or the second group of reference points which includes the one or more first three-dimensional points including the position information before correction. The position information of the target three-dimensional point is decoded by referring to at least a portion of the position information of the predicted point. Three-dimensional data decoding method.
13. In the correction described above, the positional information of the one or more first three-dimensional points is aligned with the coordinate system of the target three-dimensional point based on first information indicating the displacement between the coordinate system of the one or more first three-dimensional points and the coordinate system of the target three-dimensional point. The method for decoding three-dimensional data according to claim 12.
14. In the correction described above, the positional information of one or more second three-dimensional points included in the first reference point group is derived by projecting the one or more first three-dimensional points onto the coordinate origin of the target three-dimensional point according to the displacement. The method for decoding three-dimensional data according to claim 13.
15. The first information includes at least one of second information relating to movement parallel to the horizontal plane and third information relating to rotation about a vertical axis. The method for decoding three-dimensional data according to claim 13.
16. The position information includes a distance component, a horizontal angle component, and an elevation angle component. The correction involves correcting at least one of the distance component and the horizontal angle component. The method for decoding three-dimensional data according to claim 12.
17. The aforementioned three-dimensional data decoding method further includes, From the bitstream, a fourth piece of information is obtained indicating whether or not the correction should be performed. Based on the fourth piece of information, a decision is made as to whether or not to perform the correction. The method for decoding three-dimensional data according to claim 12.
18. If the above correction is not performed, the predicted point is selected from the second group of reference points. The method for decoding three-dimensional data according to claim 17.
19. The one or more first three-dimensional points are included in the first processing unit. If the correction is not performed, the predicted point is selected from either the second group of reference points or the third group of reference points, which consists of one or more third three-dimensional points included in a second processing unit different from the first processing unit, and which include the position information before correction. The method for decoding three-dimensional data according to claim 17.
20. The first group of reference points is generated by correcting one or more fourth three-dimensional points, which are part of the one or more first three-dimensional points. The method for decoding three-dimensional data according to claim 12.
21. The position information includes a distance component, a horizontal angle component, and an elevation angle component. The fourth three-dimensional point is a first three-dimensional point among the one or more first three-dimensional points whose elevation angle component is larger than a predetermined value. The method for decoding three-dimensional data according to claim 20.
22. The fourth three-dimensional point is a first three-dimensional point among the one or more first three-dimensional points whose vertical position is higher than a predetermined position. The method for decoding three-dimensional data according to claim 20.
23. Processor and Equipped with memory, The processor uses the memory to: Motion compensation is performed to generate a first set of reference points by correcting the positional information of one or more first three-dimensional points to match the coordinate system of the target three-dimensional point to be encoded. A predicted point of the target three-dimensional point is selected from either the first group of reference points or the second group of reference points which includes the one or more first three-dimensional points including the position information before correction. The position information of the target three-dimensional point is encoded by referring to at least a portion of the position information of the predicted point. Three-dimensional data encoding device.
24. Processor and Equipped with memory, The processor uses the memory to: Motion compensation is performed to generate a first set of reference points by correcting the positional information of one or more first three-dimensional points to match the coordinate system of the target three-dimensional point to be decoded. A predicted point of the target three-dimensional point is selected from either the first group of reference points or the second group of reference points which includes the one or more first three-dimensional points including the position information before correction. The position information of the target three-dimensional point is decoded by referring to at least a portion of the position information of the predicted point. Three-dimensional data decoding device.
Citation Information
Patent Citations
Motion-compensated compression of dynamic voxelized point clouds
US20170347120A1
US2021/99711A1
Map display device
WO2014020663A1