Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

By encoding three-dimensional data in smaller units with common additional information, the method addresses delay issues in existing technologies, enabling faster and more flexible data output and processing.

JP7897233B2Active Publication Date: 2026-07-29PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2022-06-10
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing three-dimensional data encoding and decoding processes suffer from significant delay times due to the large amount of data and the need for complete encoding units to be finished before output, limiting flexibility and efficiency.

Method used

The method generates multiple encoded data units smaller than the initial unit, each with common additional information, allowing for independent processing and output without waiting for the entire initial unit to be completed, enhancing flexibility and reducing processing load.

Benefits of technology

This approach reduces delay times in data output and improves flexibility by allowing for adjustable data sizes suitable for transmission, while maintaining efficient processing without individual additional information for each unit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007897233000002
    Figure 0007897233000002
  • Figure 0007897233000003
    Figure 0007897233000003
  • Figure 0007897233000004
    Figure 0007897233000004
Patent Text Reader

Abstract

The three-dimensional data encoding method includes encoding information regarding a plurality of positions of a plurality of three-dimensional points included in one unit which is an encoding unit, with the encoding made in units of a second unit smaller than the first unit, so as to generate a plurality of encoded data (S111) and outputting the plurality of encoded data (S112), each of the plurality of encoded data not having individual additional information. The plurality of encoded data may have, for example, common additional information. The common additional information may include, for example, first information that indicates the size of first encoded data included in the plurality of encoded data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, and a three-dimensional data decoding apparatus.

Background Art

[0002] In the future, the spread of devices or services that utilize three-dimensional data is expected in a wide range of fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots. Three-dimensional data is acquired by various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras.

[0003] As one method of expressing three-dimensional data, there is a method called point cloud that represents the shape of a three-dimensional structure by a point group in a three-dimensional space. In a point cloud, the positions and colors of the point group are stored. Although the point cloud is expected to become mainstream as a method of expressing three-dimensional data, the amount of data of the point group is extremely large. Therefore, in the accumulation or transmission of three-dimensional data, compression of the amount of data by encoding is essential, similar to two-dimensional moving images (for example, MPEG-4 AVC or HEVC standardized by MPEG).

[0004] In addition, compression of the point cloud is partially supported by a publicly available library (Point Cloud Library) that performs point cloud-related processing.

[0005] In addition, a technique for searching and displaying facilities located around a vehicle using three-dimensional map data is known (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

[0007] In the encoding and decoding processes of three-dimensional data, it is desirable to reduce the delay time from the generation of encoded data to its output.

[0008] The purpose of this disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can shorten the delay time from the generation of encoded data to its output. [Means for solving the problem]

[0009] A three-dimensional data encoding method according to one aspect of this disclosure generates multiple encoded data by encoding information about multiple positions of multiple three-dimensional points included in a first unit which is an encoding unit, for each second unit which is smaller than the first unit, outputs the multiple encoded data, and each of the multiple encoded data has individual additional information. Furthermore, the plurality of encoded data have common additional information, and the common additional information includes first information indicating the size of the first encoded data included in the plurality of encoded data. .

[0010] A three-dimensional data decoding method according to one aspect of this disclosure involves obtaining a plurality of encoded data sets generated by encoding information about a plurality of positions of a plurality of three-dimensional points contained in a first unit, which is an encoding unit, for each second unit smaller than the first unit, and decoding the plurality of encoded data sets to generate information about a plurality of positions of a plurality of three-dimensional points contained in the first unit, and each of the plurality of encoded data sets has individual additional information. Furthermore, the plurality of encoded data have common additional information, and the common additional information includes first information indicating the size of the first encoded data included in the plurality of encoded data. . [Effects of the Invention]

[0011] This disclosure provides a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can shorten the delay time from the generation of encoded data to its output. [Brief explanation of the drawing]

[0012] [Figure 1] FIG. 1 is a block diagram of a three-dimensional data encoding device according to an embodiment. [Figure 2] FIG. 2 is a block diagram of a three-dimensional data decoding device according to an embodiment. [Figure 3] FIG. 3 is a diagram showing the encoding order of a plurality of three-dimensional points according to an embodiment. [Figure 4] FIG. 4 is a diagram showing an example of the syntax of position information according to an embodiment. [Figure 5] FIG. 5 is a diagram showing an example of the syntax of position information according to an embodiment. [Figure 6] FIG. 6 is a diagram showing an example of a plurality of divided bitstreams according to an embodiment. [Figure 7] FIG. 7 is a diagram showing the syntax for byte alignment processing according to an embodiment. [Figure 8] FIG. 8 is a diagram showing an example of the syntax of a header according to an embodiment. [Figure 9] FIG. 9 is a diagram showing an example of the syntax of position information corresponding to subset division according to an embodiment. [Figure 10] FIG. 10 is a flowchart of processing related to subset boundaries according to an embodiment. [Figure 11] FIG. 11 is a flowchart of three-dimensional data encoding processing according to an embodiment. [Figure 12] FIG. 12 is a flowchart of three-dimensional data decoding processing according to an embodiment. <000—089>

BEST MODE FOR CARRYING OUT THE INVENTION

[0013] A three-dimensional data encoding method according to one aspect of the present disclosure generates a plurality of encoded data by encoding information about the positions of a plurality of three-dimensional points included in a first unit, which is an encoding unit, for each second unit smaller than the first unit, outputs the plurality of encoded data, and each of the plurality of encoded data does not have individual additional information.

[0014] According to this, the three-dimensional data encoding method can output the encoded data of the second unit without waiting for the encoded data of the first unit to be complete. Thereby, the delay time from when the encoded data is generated until it is output can be shortened. Further, since the encoded data of the second unit does not have individual additional information, an increase in the processing amount for generating the plurality of encoded data of the second unit can be suppressed compared to the case where the encoded data of the second unit has individual additional information. Also, since the second unit is not restricted to being a unit having additional information, the degree of freedom in generating the encoded data of the second unit can be improved. Thereby, for example, the plurality of encoded data can be adjusted to a size suitable for transmission.

[0015] Note that the size of the second unit may be variable. That is, the size of the encoded data only needs to be smaller than the size of the first unit and may be different from the sizes of other encoded data.

[0016] For example, the plurality of encoded data may have common additional information. For example, the common additional information may include first information indicating the size of first encoded data included in the plurality of encoded data.

[0017] According to this, the three-dimensional data decoding device can specify the end of the first encoded data using the first information.

[0018] For example, the information regarding the multiple positions of the multiple three-dimensional points may be expressed by representing each of the multiple positions with a distance component, a first directional component, and a second directional component, and the information regarding the multiple positions of the multiple three-dimensional points may be encoded using a predetermined set of reference positions, each of the set of reference positions including the first directional component and the second directional component, and the first information may indicate the magnitude of the first directional component of the first encoded data.

[0019] According to this method, the size of the encoded data can be represented by the size of the first direction component, which reduces the amount of data for the first information compared to when the size of the encoded data is represented by the distance component, the first direction component, and the second direction component.

[0020] For example, each of the plurality of encoded data may include second information indicating whether or not termination processing is performed on the encoded data.

[0021] According to this, it is possible to choose whether or not to perform terminal processing at each second unit identified by the first information. Therefore, the degree of flexibility in data partitioning can be improved.

[0022] For example, the common additional information may include a third piece of information indicating whether or not other encoded data included in the plurality of encoded data is used to encode the first encoded data included in the plurality of encoded data.

[0023] This allows switching whether encoded data depends on other encoded data or not.

[0024] For example, the common additional information may include a fourth piece of information indicating whether the context used for arithmetic coding of the first coded data included in the plurality of coded data depends on other coded data included in the plurality of coded data.

[0025] This allows for switching whether encoded data depends on other encoded data or not. Furthermore, switching whether the context depends on other encoded data depending on the data may improve encoding efficiency.

[0026] A three-dimensional data decoding method according to one aspect of this disclosure acquires a plurality of encoded data sets generated by encoding information about a plurality of positions of a plurality of three-dimensional points contained in a first unit, which is an encoding unit, for each second unit smaller than the first unit, and decodes the plurality of encoded data sets to generate information about a plurality of positions of a plurality of three-dimensional points contained in the first unit, wherein each of the plurality of encoded data sets does not have individual additional information.

[0027] According to this, the three-dimensional data decoding method can begin decoding the second unit of encoded data without waiting for all of the first unit's encoded data to be completed. This reduces the delay time from receiving the encoded data to starting decoding. Furthermore, because the second unit's encoded data does not contain individual additional information, the increase in processing load required to analyze multiple second unit encoded data can be suppressed compared to cases where the second unit's encoded data contains individual additional information. In addition, since the second unit is not restricted by the fact that it contains additional information, the degree of freedom in generating the second unit's encoded data can be improved. This allows, for example, multiple encoded data to be adjusted to a size suitable for transmission.

[0028] For example, the plurality of encoded data may have common additional information. For example, the common additional information may include first information indicating the size of the first encoded data included in the plurality of encoded data.

[0029] According to this, the three-dimensional data decoding method can identify the end of the first encoded data using the first information.

[0030] For example, the information regarding the multiple positions of the multiple three-dimensional points may be expressed by representing each of the multiple positions with a distance component, a first directional component, and a second directional component, and the information regarding the multiple positions of the multiple three-dimensional points may be encoded using a predetermined set of reference positions, each of the set of reference positions including the first directional component and the second directional component, and the first information may indicate the magnitude of the first directional component of the first encoded data.

[0031] According to this method, the size of the encoded data can be represented by the size of the first direction component, which reduces the amount of data for the first information compared to when the size of the encoded data is represented by the distance component, the first direction component, and the second direction component.

[0032] For example, each of the plurality of encoded data may include second information indicating whether or not termination processing is performed on the encoded data.

[0033] According to this, it is possible to choose whether or not to perform terminal processing at each second unit identified by the first information. Therefore, the degree of flexibility in data partitioning can be improved.

[0034] For example, the common additional information may include a third piece of information indicating whether or not other encoded data included in the plurality of encoded data is used to decode the first encoded data included in the plurality of encoded data.

[0035] This allows switching whether encoded data depends on other encoded data or not.

[0036] For example, the common additional information may include a fourth piece of information indicating whether the context used for arithmetic decoding of the first encoded data included in the plurality of encoded data depends on other encoded data included in the plurality of encoded data.

[0037] This allows for switching whether encoded data depends on other encoded data or not. Furthermore, switching whether the context depends on other encoded data depending on the data may improve encoding efficiency.

[0038] Furthermore, a three-dimensional data encoding device according to one aspect of the present disclosure comprises a processor and a memory, wherein the processor generates a plurality of encoded data by encoding information about a plurality of positions of a plurality of three-dimensional points contained in a first unit which is an encoding unit, for each second unit which is smaller than the first unit, using the memory, and outputs the plurality of encoded data, each of the plurality of encoded data having no individual additional information.

[0039] According to this, the three-dimensional data encoding device can output the encoded data of the second unit without waiting for all the encoded data of the first unit to be completed. This reduces the delay time from the generation of encoded data to its output. Furthermore, because the encoded data of the second unit does not contain individual additional information, the increase in processing load required to generate multiple encoded data of the second unit can be suppressed compared to cases where the encoded data of the second unit contains individual additional information. In addition, since the second unit is not restricted by the fact that it is a unit containing additional information, the degree of freedom in generating the encoded data of the second unit can be improved. This allows, for example, multiple encoded data to be adjusted to a size suitable for transmission.

[0040] Furthermore, a three-dimensional data decoding device according to one aspect of the present disclosure comprises a processor and a memory, wherein the processor uses the memory to acquire a plurality of encoded data sets generated by encoding information about a plurality of positions of a plurality of three-dimensional points contained in a first unit, which is an encoding unit, for each second unit smaller than the first unit, and generates information about a plurality of positions of a plurality of three-dimensional points contained in the first unit by decoding the plurality of encoded data sets, and each of the plurality of encoded data sets does not have individual additional information.

[0041] According to this, the three-dimensional data decoding device can begin decoding the second unit of encoded data without waiting for all of the first unit's encoded data to be completed. This reduces the delay time from receiving the encoded data to starting decoding. Furthermore, because the second unit's encoded data does not contain individual additional information, the increase in processing load required to analyze multiple second unit encoded data can be suppressed compared to cases where the second unit's encoded data contains individual additional information. In addition, since the second unit is not restricted by the presence of additional information, the degree of freedom in generating the second unit's encoded data can be improved. This allows, for example, multiple encoded data to be adjusted to a size suitable for transmission.

[0042] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.

[0043] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, components in the following embodiments that are not described in an independent claim will be described as optional components.

[0044] (Embodiment) First, the configuration of the three-dimensional data encoding device 100 according to this embodiment will be described. Figure 1 is a block diagram of the three-dimensional data encoding device 100 according to this embodiment. The three-dimensional data encoding device 100 generates a bitstream (encoded stream) by encoding point cloud data, which is three-dimensional data.

[0045] Point cloud data contains positional information for multiple three-dimensional points. This positional information indicates the three-dimensional position of each point. This positional information is sometimes also referred to as geometric information.

[0046] For example, positional information is expressed in polar coordinates and includes one distance component and two directional components (angle components). Specifically, positional information includes distance d, elevation angle θ, and horizontal angle φ. Point cloud data is, for example, data obtained from a laser sensor such as LiDAR.

[0047] Furthermore, the point cloud data may include attribute information (color, reflectance, etc.) for each three-dimensional point in addition to positional information. Also, although Figure 1 shows a processing unit for encoding the positional information of the point cloud data, the three-dimensional data encoding device 100 may include other processing units, such as a processing unit for encoding attribute information.

[0048] The three-dimensional data encoding device 100 includes a conversion unit 101, a subtraction unit 102, a quantization unit 103, an entropy encoding unit 104, an inverse quantization unit 105, an addition unit 106, a buffer 108, an intra prediction unit 109, a buffer 110, a motion detection compensation unit 111, an inter prediction unit 112, a switching unit 113, and a control unit 114.

[0049] The conversion unit 101 generates conversion information by converting the position information contained in the input point cloud data to be encoded. Specifically, the conversion unit 101 generates information for associating multiple reference positions with three-dimensional points. The conversion unit 101 also converts the position information of the three-dimensional points using the reference positions. For example, the conversion information is the difference between the position information of the reference positions and the position information of the three-dimensional points. Details of this will be described later. The conversion unit 101 may also have a buffer to hold the converted position information. The conversion unit 101 can also be described as a calculation unit that calculates the value to be encoded.

[0050] The subtraction unit 102 generates a residual signal (also called a predicted residual) by subtracting a predicted value from the converted position information. The quantization unit 103 quantizes the residual signal. The entropy coding unit 104 generates a bitstream by entropy coding the quantized residual signal. The entropy coding unit 104 also entropy codes control information, such as the information generated by the conversion unit 101, and adds the coded information to the bitstream.

[0051] The inverse quantization unit 105 generates a residual signal by inverse quantization of the quantized residual signal obtained by the quantization unit 103. The summing unit 106 restores the transformation information by adding the predicted value to the residual signal generated by the inverse quantization unit 105. The buffer 108 holds the restored transformation information as a reference point set for intra-prediction. The buffer 110 holds the restored transformation information as a reference point set for inter-prediction.

[0052] Note that the reconstructed transformation information may not perfectly match the original transformation information because it contains quantization errors. The three-dimensional points reconstructed through this encoding and decoding process are referred to as the encoded three-dimensional points, the decoded three-dimensional points, or the processed three-dimensional points.

[0053] The intra prediction unit 109 calculates a predicted value using the transformation information of one or more reference points, which are other processed three-dimensional points belonging to the same frame as the three-dimensional point to be processed (hereinafter referred to as the target point).

[0054] The motion detection and compensation unit 111 detects the displacement (motion detection) between the target frame, which is the frame containing the target points, and the reference frame, which is a different frame from the target frame, and corrects (motion compensation) the transformation information of the point cloud contained in the reference frame based on the detected displacement. The information indicating the detected displacement (motion information) is stored, for example, in a bitstream.

[0055] The interpretation unit 112 calculates predicted values ​​using the transformation information of one or more reference points included in the motion-compensated point cloud. Motion detection and motion compensation are not required.

[0056] The switching unit 113 selects either the predicted value calculated by the intra-prediction unit 109 or the predicted value obtained by the inter-prediction unit 112, and outputs the selected predicted value to the subtraction unit 102 and the addition unit 106. In other words, the switching unit 113 switches between using intra-prediction or inter-prediction. For example, the switching unit 113 calculates the cost value when using intra-prediction and the cost value when using inter-prediction, and selects the prediction method that results in a smaller cost value. The cost value is, for example, a value based on the code amount after encoding, and the smaller the code amount, the smaller the cost value. Similarly, if there are multiple methods (multiple prediction modes) for both intra-prediction and inter-prediction, the prediction mode to be used is determined based on the cost value. The method for determining the prediction method (intra-prediction or inter-prediction) and prediction mode is not limited to this, and may be determined based on externally specified settings, or based on the characteristics of the point cloud data, or the selectable candidates may be narrowed down.

[0057] The control unit 114 determines the division boundary of the encoded data based on the conversion information output from the conversion unit 101. The control unit 114 also controls the entropy encoding unit 104 to perform the arithmetic encoding termination process described later at the division boundary.

[0058] The three-dimensional data encoding device 100 may acquire position information represented in a Cartesian coordinate system, convert the acquired position information in a Cartesian coordinate system into position information in a polar coordinate system, and perform the above encoding process on the obtained position information in a polar coordinate system. For example, the three-dimensional data encoding device 100 may be equipped with a coordinate transformation unit that performs this coordinate transformation process before the transformation unit 101. In this case, the three-dimensional data encoding device 100 may generate position information in a polar coordinate system by performing the inverse transformation of the transformation process performed by the transformation unit 101 on the transformed information restored by the addition unit 106, convert the generated position information in a polar coordinate system into position information in a Cartesian coordinate system, calculate the difference between the obtained position information in a Cartesian coordinate system and the input original position information in a Cartesian coordinate system, and store information indicating the calculated difference in a bitstream.

[0059] Next, the configuration of the three-dimensional data decoding device 200, which decodes the bitstream generated by the three-dimensional data encoding device 100, will be described. Figure 2 is a block diagram of the three-dimensional data decoding device 200 according to this embodiment. Although Figure 2 shows the processing unit for decoding the position information of the point cloud, the three-dimensional data decoding device 200 may also include other processing units, such as a processing unit for decoding the attribute information of the point cloud. For example, the three-dimensional data decoding device 200 generates decoded point cloud data by decoding the bitstream generated by the three-dimensional data encoding device 100 shown in Figure 1.

[0060] The three-dimensional data decoding device 200 comprises an entropy decoding unit 201, an inverse quantization unit 202, an addition unit 203, an inverse transformation unit 204, a buffer 205, an intra prediction unit 206, a buffer 207, a motion compensation unit 208, an inter-prediction unit 209, a switching unit 210, and a control unit 211.

[0061] The three-dimensional data decoding device 200 acquires multiple bitstreams generated in the three-dimensional data encoding device 100.

[0062] The control unit 211 monitors multiple received bitstreams (encoded data) and detects bitstream division boundaries based on the information contained in the bitstreams. The control unit 211 controls the entropy decoding unit 201 to perform arithmetic decoding termination processing, described later, at the detected division boundaries. This bitstream corresponds to the bitstream output from the entropy coding unit 104 in the three-dimensional data encoding device 100.

[0063] The entropy decoding unit 201 generates quantized residual signals and control information, etc., by entropy decoding the received bitstream.

[0064] The inverse quantization unit 202 generates a residual signal by inverse quantization of the quantized residual signal obtained by the entropy decoding unit 201. The summing unit 203 restores the conversion information by adding a predicted value to the residual signal generated by the inverse quantization unit 202.

[0065] The inverse transformation unit 204 restores the position information by performing an inverse transformation of the transformation process performed by the transformation unit 101 on the transformed information. Specifically, the inverse transformation unit 204 obtains information from the bitstream to associate multiple reference positions with three-dimensional points, and associates the multiple reference positions with three-dimensional points based on the obtained information. The inverse transformation unit 204 also converts the transformed information of the three-dimensional points into position information using the reference positions. For example, the inverse transformation unit 204 calculates the position information by adding the transformed information and the reference positions. The inverse transformation unit 204 can also be described as a calculation unit that calculates position information from the decoded values. This position information is output as decoded point cloud data.

[0066] Buffer 205 holds the transformation information restored by the summing unit 203 as a reference point group for intra-prediction. Buffer 207 holds the transformation information restored by the summing unit 203 as a reference point group for intra-prediction. The intra-prediction unit 206 calculates the predicted value using the transformation information of one or more reference points, which are other three-dimensional points belonging to the same frame as the target point.

[0067] The motion compensation unit 208 acquires motion information from the bitstream indicating the displacement between the target frame and the reference frame, and corrects (motion compensates for) the transformation information of the point cloud included in the reference frame based on the displacement indicated by the motion information. The interpretation unit 209 calculates a predicted value using the transformation information of one or more reference points included in the motion-compensated point cloud. Motion compensation is not required.

[0068] The switching unit 210 selects either the predicted value calculated by the intra-prediction unit 206 or the predicted value obtained by the inter-prediction unit 209, and outputs the selected predicted value to the addition unit 203. For example, the switching unit 210 obtains information indicating the prediction method (intra-prediction or inter-prediction) from the bitstream, and determines the prediction method to be used based on the obtained information. Similarly, if there are multiple methods (multiple prediction modes) for both intra-prediction and inter-prediction, information indicating the prediction mode is obtained from the bitstream, and the prediction mode to be used is determined based on the obtained information.

[0069] The three-dimensional data decoding device 200 may convert the decoded position information expressed in polar coordinates to position information expressed in Cartesian coordinates and output the position information expressed in Cartesian coordinates. For example, the three-dimensional data decoding device 200 may include a coordinate transformation unit that performs this coordinate transformation after the inverse transformation unit 204. In this case, the three-dimensional data decoding device 200 obtains from the bitstream information indicating the difference between the original position information in Cartesian coordinates before encoding and decoding and the position information in Cartesian coordinates after decoding. The three-dimensional data decoding device 200 may convert the position information in polar coordinates restored by the inverse transformation unit 204 to position information in Cartesian coordinates, add the difference indicated by the above information to the obtained position information in Cartesian coordinates, and output the obtained position information in Cartesian coordinates.

[0070] Next, the operation of the three-dimensional data encoding device 100 will be explained. Figure 3 is a diagram showing the operation of the conversion unit 101, and is a diagram showing the encoding order (processing order) of multiple three-dimensional points (multiple reference positions) in the encoding process.

[0071] In Figure 3, the horizontal direction represents the horizontal angle φ in polar coordinates, and the vertical direction represents the elevation angle θ in polar coordinates. The conversion unit 101 sets up multiple reference positions rm (where m=0,1,2,…) (also called reference points). Here, a reference position rm is represented by the horizontal angle φ and the elevation angle θ. In other words, a reference position rm is represented by two components (θ, φ) of the three components (d, θ, φ) that represent the position information of a three-dimensional point. In the example shown in Figure 3, the reference position rm, indicated by the squares in the figure, is set according to the horizontal sampling interval Δφ of the LiDAR and the scanline interval Δθk of the LiDAR (where k=1,2,3). In other words, multiple reference positions are set by a predetermined combination of multiple horizontal angles and multiple elevation angles, and are arranged in a matrix on the plane represented by the horizontal angle φ and the elevation angle θ. In the example shown in Figure 3, the interval Δφ of the multiple horizontal angles φj (where j=0,1,2,…) of the multiple reference positions is constant. Furthermore, the intervals between multiple elevation angles θk (where k=0,1,2,3) of multiple reference positions can be set individually.

[0072] Furthermore, the conversion unit 101 performs encoding (conversion) of points pn (where n=0,1,2,…) located near each reference position, indicated by the dashed arrows in the figure. Note that shaded squares indicate a first reference position where a point referencing that reference position exists, and unshaded squares indicate a second reference position where no point referencing that reference position exists.

[0073] A point that references the reference position is a point that is based on the reference position and is associated with the reference position (encoded (transformed) using the reference position) as described later. A point that references the reference position is a point whose horizontal angle φ and elevation angle θ values ​​fall within a predetermined range that includes the reference position. For example, a point that references the reference position is a point pn on the same scanline (with the same elevation angle) that has a horizontal angle of φj or greater and less than φj + Δφ. Note that the range of the horizontal angle is not limited to this, and may be φj - Δφ / 2 or greater and less than φj + Δφ / 2, etc.

[0074] Furthermore, the processing order (encoding order) shown in Figure 3 uses multiple reference positions (e.g., r0 to r3) with the same horizontal angle value as processing units (corresponding to each column in Figure 3), and within each processing unit, multiple reference positions are processed (encoded) in an order based on elevation angle (ascending order in Figure 3). Also, multiple processing units (corresponding to each column in Figure 3) are processed in an order based on horizontal angle (ascending order in Figure 3). In other words, for each set of multiple reference positions with the same horizontal angle value, multiple reference positions are processed in ascending order of elevation angle. Alternatively, multiple reference positions with the same elevation angle value may be processed in ascending order of horizontal angle.

[0075] The conversion unit 101 generates information to identify the position (φj, θk) of the reference position rm that the target point pn refers to during the encoding (conversion) of the target point. The conversion unit 101 also generates information to identify the offset (φ) from the reference position to the target point. o n, θ o Information is generated to identify n) and the distance information dn of the target point. Here, φ o n is the difference between the horizontal angle φj of the reference position and the horizontal angle of the target point, and θ o n is the difference between the elevation angle θk of the reference position and the elevation angle of the target point.

[0076] Furthermore, information to identify the position of the reference point that the target point refers to, and the offset (φ) from the reference point to the point in question. o n, θ o n) The information used to identify the distance information dn of the target point may be information that identifies the difference value from the predicted value generated based on the processed information, or it may be information that identifies the value itself.

[0077] Furthermore, the three-dimensional data encoding device 100 may store the horizontal sampling interval Δφ and the scanline interval Δθk of the LiDAR in a bitstream. For example, the three-dimensional data encoding device may store Δφ and Δθk in a header such as an SPS or GPS. This allows the three-dimensional data decoding device 200 to set multiple reference positions using Δφ and Δθk.

[0078] Here, SPS (Sequence Parameter Set) is a parameter set (control information) for a sequence containing multiple frames. Furthermore, SPS is a parameter set common to both position information and attribute information. GPS (Geometry Parameter Set) is a frame-specific parameter set, specifically for position information.

[0079] Furthermore, the conversion unit 101 may convert the horizontal sampling interval Δφ and the scanline interval Δθk of the LiDAR into integer values ​​with a predetermined bit width, and store the converted values ​​in a bitstream. Also, although the example shown in Figure 3 shows an example where the number of scanlines (number of elevation angles) is 4, the same can be carried out when other scanline counts such as 16, 64, or 128 are used.

[0080] Next, the syntax for position information will be explained. Figure 4 shows an example of the syntax for the position information of each point. In the syntax examples shown in Figures 4 and 5, the parameters (signals) stored in the bitstream are shown in bold. The three-dimensional data encoding device 100 repeatedly applies this syntax for each reference position rm to generate column_pos, which indicates the index of the horizontal angle φj of the reference position rm that will be used as the reference for the next point pn to be processed, and row_pos, which indicates the index of the elevation angle θk, and further generates parameters related to point pn.

[0081] In this example, the three-dimensional data encoding device 100 initializes variables before processing the first point. Specifically, it sets first_point_in_column to 1, column_pos to 0, and row_pos to 0, indicating that it is the first syntax corresponding to the horizontal angle φj. Alternatively, the three-dimensional data encoding device 100 may notify the three-dimensional data decoding device 200 of the column_pos and row_pos values ​​of the first point prior to the syntax corresponding to the first point. In this case, the three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 may set first_point_in_column to 0 and then apply this syntax using these values.

[0082] Next, the three-dimensional data encoding device 100 generates a next_column_flag at the reference position rm corresponding to the elevation angle θ0 (i.e., when first_point_in_column is 1). The next_column_flag indicates whether there is one or more points that reference the horizontal angle φj corresponding to the position of the reference position rm. In other words, the next_column_flag indicates whether there is a point that references one of several reference positions with the same horizontal angle φj as the reference position rm. For example, if there is one or more points that reference the horizontal angle φj corresponding to the position of the reference position rm (for example, horizontal angles φ0, φ1, φ2, and φ4 shown in Figure 3), the next_column_flag is set to 0, and if there are no points that reference the horizontal angle φj corresponding to the position of the reference position rm (for example, horizontal angle φ3 shown in Figure 3), the next_column_flag is set to 1. Furthermore, a next_column_flag is provided for each horizontal angle φj (for each column in Figure 3).

[0083] The three-dimensional data encoding device 100 can generate information that identifies the horizontal angle φj (φ0 + column_pos × Δφ) corresponding to the next point pn to be processed by repeatedly generating next_column_flag until next_column_flag becomes 0. This may reduce the amount of coding required to notify next_row_flag. As shown in Figure 5 below, it is also possible to determine whether or not to notify next_column_flag based on whether row_pos is 0 or not. However, by making the determination based on first_point_in_column, it is possible to avoid unnecessary notification of next_column_flag even when there are multiple points at the position where row_pos is 0, thus potentially reducing the amount of coding.

[0084] The three-dimensional data encoding device 100 generates a next_row_flag at each candidate position of the reference position rm, which serves as the reference for the next point pn to be processed. The next_row_flag indicates whether or not the point pn to be processed exists at the elevation angle θk. In other words, the next_row_flag indicates whether or not there is a point that references the reference position rm. For example, if the point pn to be processed exists at the elevation angle θk, the next_row_flag is set to 0 (for example, r0 and r1 in Figure 3), and if the point pn to be processed does not exist at the elevation angle θk (for example, r2 and r3 in Figure 3), the next_row_flag is set to 1. Furthermore, a next_row_flag is provided for each reference position.

[0085] The three-dimensional data encoding device 100, when next_row_flag is 1, repeatedly applies the syntax shown in Figure 4 to generate a next_row_flag corresponding to each candidate position. The three-dimensional data encoding device 100 repeats this process until next_row_flag becomes 0, thereby generating information that can identify the elevation angle θk corresponding to the next point pn to be processed. For example, the elevation angle θk corresponding to the next point pn to be processed is expressed by the following equation (Equation 1).

[0086]

number

[0087] When row_pos reaches the number of scan lines (num_rows shown in Figure 4), processing moves to the next horizontal angle φj. At this time, the three-dimensional data encoding device 100 sets row_pos to 0, increases column_pos by 1, and sets first_point_in_column to 1.

[0088] As a result, the three-dimensional data encoding device 100 can generate information (next_column_flag, next_row_flag) that can identify the horizontal angle φj and elevation angle θk of the reference position rm which serves as the reference for the point pn to be processed.

[0089] Next, the three-dimensional data encoding device 100 generates information regarding the distance to the target point pn, information regarding the horizontal angle offset from the reference position rm to the target point pn, and pred_mode, which is information regarding the prediction method for these parameters. Here, the distance information is, for example, the residual residual_radius, which shows the difference between the distance to the target point and the predicted value generated by a predetermined method. The information regarding the horizontal angle offset is, for example, the horizontal angle offset φ o n is the residual_phi, which represents the difference between n and the predicted value generated by a predetermined method.

[0090] The predicted value is calculated based on, for example, information on processed three-dimensional points. For example, the predicted value is at least a part of the parameters of one or more processed three-dimensional points located in the vicinity of the target point. In this example, the three-dimensional data encoding device 100 omits the generation of information regarding the elevation angle offset, assuming that the elevation angle offset is always 0. However, it may generate information regarding the elevation angle offset from the reference position rm to the point pn to be processed and store it in the bitstream. For example, the information regarding the elevation angle offset is the elevation angle offset θ. on is the residual_theta, which represents the difference between n and the predicted value generated by a predetermined method.

[0091] The three-dimensional data encoding device 100 may convert the input position information in a Cartesian coordinate system into position information expressed in a polar coordinate system, and then perform the above encoding process on the obtained position information expressed in a polar coordinate system. In this case, the three-dimensional data encoding device 100 may convert the encoded and decoded position information in a polar coordinate system (for example, position information generated by performing an inverse transformation on the output signal of the adder 106 shown in Figure 1) back into position information in a Cartesian coordinate system, calculate the difference between the obtained position information in a Cartesian coordinate system and the original input position information in a Cartesian coordinate system, and store information indicating this difference in the bitstream. This information indicating the difference is, for example, the correction values ​​for the X, Y, and Z axes residual_x, residual_y, and residual_z. In other words, if no coordinate system transformation is performed, residual_x, residual_y, and residual_z do not need to be included in the bitstream.

[0092] Furthermore, the next_column_flag, next_row_flag, pred_mode, residual_radius, residual_phi, residual_theta, residual_x, residual_y, and residual_z generated above are stored in a bitstream and sent to the three-dimensional data decoding device 200. Note that all or part of these signals may be entropy coded (arithmetic coded) by the entropy coding unit 104 before being stored in the bitstream.

[0093] As described above, the three-dimensional data encoding device 100 can determine the values ​​of syntax elements for each candidate position of the reference position rm by using flags (next_column_flag, next_row_flag) associated with each candidate position as information for identifying the horizontal angle φj and elevation angle θk of the reference position rm that will serve as the reference for the next point pn to be processed. Furthermore, it may be possible to reduce the latency of encoding, decoding, or data transmission processes.

[0094] Note that the syntax of `next_column_flag` and `next_row_flag`, as well as the assignment of values ​​to variables such as `first_point_in_column`, as described above are just examples; you may change the assignments, such as reversing the assignment of 0 to 1. In this case, it is possible to do so by aligning the related conditional checks.

[0095] Next, another example of the syntax will be explained. Figure 5 shows an example of the syntax for the position information of each point. The three-dimensional data encoding device 100 repeatedly applies this syntax for each reference position rm to generate column_pos, which indicates the index of the horizontal angle φj of the reference position rm that will be used as the reference for the next point pn to be processed, and row_pos, which indicates the index of the elevation angle θk, and further generates parameters related to point pn. Note that the example shown in Figure 5 differs from the example shown in Figure 4 in the method of generating next_row_flag and next_column_flag, which are used to determine the values ​​of column_pos and row_pos.

[0096] In this example, the three-dimensional data encoding device 100 first initializes variables before applying this syntax to the first point. Specifically, the three-dimensional data encoding device 100 notifies the three-dimensional data decoding device 200 of the column_pos and row_pos values ​​of the first point prior to the syntax corresponding to the first point. That is, the three-dimensional data encoding device 100 stores, for example, the column_pos and row_pos values ​​of the first point in the bitstream. The three-dimensional data encoding device 100 and the three-dimensional data decoding device 200 then apply this syntax using these values.

[0097] Next, the three-dimensional data encoding device 100 generates a next_row_flag for the reference position rm indicated by next_row_flag and next_column_flag, and notifies the three-dimensional data decoding device 200 whether or not a point pn exists with respect to said rm.

[0098] If next_row_flag is 1, the three-dimensional data encoding device 100 first increments row_pos by 1. Next, the three-dimensional data encoding device 100 determines whether row_pos has reached the number of scan lines (num_rows shown in Figure 5). If row_pos has reached the number of scan lines, the three-dimensional data encoding device 100 determines that the candidate position will move to the next horizontal angle φj, sets row_pos to 0, and increments column_pos by 1. Next, the three-dimensional data encoding device 100 determines whether row_pos is 0. If row_pos is 0, the three-dimensional data encoding device 100 generates one or more next_column_flags and repeatedly increments column_pos by 1 until next_column_flag becomes 0. After that, the three-dimensional data encoding device 100 repeatedly applies the syntax shown in Figure 5 until next_row_flag becomes 0.

[0099] Furthermore, if next_row_flag is 0, the three-dimensional data encoding device 100 determines that the values ​​indicated by next_row_flag and next_column_flag at that time are the index of the horizontal angle φj and the index of the elevation angle θk of the reference position rm which will serve as the reference for the next point pn to be processed, and stores the parameters related to the point pn (for example, pred_mode, residual_radius, residual_phi, residual_x, residual_y, and residual_z shown in Figure 5) in the bitstream, similar to the example shown in Figure 4. The horizontal angle φj can be calculated using φ0 + column_pos × Δφ, using each index value and the horizontal sampling interval Δφ of the LiDAR. The elevation angle θk can be calculated using the above (Equation 1), using each index value and the scanline interval Δθk of the LiDAR.

[0100] Note that if no coordinate system transformation is performed, residual_x, residual_y, and residual_z do not need to be included in the bitstream. Also, residual_theta may be included in the bitstream.

[0101] As a result, the notification of next_column_flag can be limited to only the cases where row_pos=0 and next_row_flag=1, which may reduce the amount of code.

[0102] As described above, the three-dimensional data encoding device 100 can determine the values ​​of syntax elements for each candidate position of the reference position rm by using flags (next_column_flag, next_row_flag) associated with each candidate position as information for identifying the horizontal angle φj and elevation angle θk of the reference position rm that will serve as the reference for the next point pn to be processed. Furthermore, it may be possible to reduce the latency of encoding, decoding, or data transmission processes.

[0103] Note that the assignment of values ​​to syntax elements such as next_column_flag and next_row_flag in the above explanation is just an example; you may change the assignment, for example, by reversing the assignment of 0 to 1. In this case, it is possible to do so by aligning the related conditional checks.

[0104] The processing of the control unit 114 will be described below. Figure 6 is a diagram illustrating the processing of the control unit 114, and shows an example of multiple divided bitstreams. Figure (a) shows a bitstream generated by encoding point cloud data for one frame, and for example, shows a bitstream output from the entropy encoding unit 104. The bitstream shown in Figure (a) includes encoded data (arithmetic encoded data for one frame) obtained by encoding one frame without interruption. The bitstream also includes the frame header and the frame footer.

[0105] Figure (b) shows an example of multiple divided bitstreams. The control unit 114 controls the division of one frame of encoded data into multiple subsets of encoded data. Here, the division of encoded data does not occur after all of the encoded data for one frame has been generated, but rather the subsets of encoded data are generated sequentially while the encoded data for one frame is being generated. Furthermore, the three-dimensional data encoding device 100 may transmit the generated encoded data sequentially without waiting for all of the subsets of encoded data to be generated.

[0106] Here, when encoding information including the positions of three-dimensional point clouds using arithmetic coding, one codeword is constructed for each processing unit, such as a frame, as shown in Figure 6(a). In addition, codeword termination is generally performed at the end of each processing unit so that all information can be extracted during decoding. Furthermore, byte alignment processing is performed as needed, adding a specific bit pattern so that the end of the data is at an integer multiple of the byte length from the beginning of the bitstream. Here, we will refer to this process, including byte alignment processing, as termination processing.

[0107] On the other hand, in this embodiment, as shown in Figure 6(b), the control unit 114 divides the data of one frame into multiple subsets of encoded data by performing codeword termination processing on the encoded data of the frame at predetermined positions. At this time, the encoded data of each subset does not have a header and a footer. In other words, the three-dimensional data encoding device 100 does not add a subset-level header and a footer to the bitstream. Note that a subset-level header and a footer are not necessarily required, but the three-dimensional data encoding device 100 may add at least one of a subset-level header and a footer to the bitstream. In this case, the subset-level header and footer include, for example, some of the information contained in the frame header or the frame footer.

[0108] In this way, the three-dimensional data encoding device 100 divides the data of one frame into multiple subsets of encoded data by performing codeword termination processing at predetermined positions on the encoded data of the frame. As a result, the three-dimensional data encoding device 100 can, for example, packetize the encoded data for each subset and immediately transmit the obtained data. This reduces the delay time from when the encoded data is generated until it is output. The three-dimensional data decoding device 200 can also reduce the delay time from when it receives the encoded data until it starts decoding. Therefore, it is possible to reduce the time required from encoding to decoding.

[0109] Note that in Figures 6(a) and 6(b), the bitstream has both a frame header and a frame footer, but it may have only one of them.

[0110] Figure 7 shows the syntax for byte alignment processing. This syntax is the byte alignment syntax specified in Recommendation ITU-T H.266 | International Standard ISO / IEC 23090-3 Versatile Video Coding.

[0111] For example, the three-dimensional data encoding device 100 may perform byte alignment processing by inserting consecutive "0"s with "1"s using the syntax shown in Figure 7. However, the three-dimensional data encoding device 100 is not necessarily required to use this syntax and may use other syntax. For example, the three-dimensional data encoding device 100 may use a syntax that omits alignment_bit_equal_to_one and aligns by inserting only consecutive "0"s.

[0112] While the example of dividing a frame into subsets has been described, other processing units that encode information including the positions of three-dimensional points, such as slices, may also be divided into subsets. Furthermore, the processing unit before division is, for example, an encoding unit. Here, an encoding unit is a unit of encoding and decoding processing, and consists of one or more random access units. In other words, the data in a processing unit is individually decodeable data. Also, the data in a subset does not contain additional information; that is, the data in a subset is not individually decodeable.

[0113] Next, we will explain the structure of the header information. Figure 8 is a diagram showing an example of the syntax of the header included in the bitstream shown in Figure 6(b), and is a diagram showing an example of the syntax of a sequence parameter set (SPS).

[0114] Furthermore, an example of the semantics of each signal shown in Figure 8 is given below. A value of 1 for sps_subset_enabled_flag indicates that subset splitting is available for multiple frames in the bitstream that reference the SPS in question. A value of 0 for sps_subset_enabled_flag indicates that subset splitting is disabled for multiple frames in the bitstream that reference the SPS in question.

[0115] The value obtained by adding 1 to sps_subset_size_minus1 indicates the number of reference position columns corresponding to the data contained in one subset (subset partition) of the frame referencing the SPS in the bitstream.

[0116] A value of 1 for sps_subset_dependency_exist_flag indicates that the first subset of frames in the bitstream that reference the SPS may depend on the second subset of the frame, which precedes the first subset in encoding order. A value of 0 for sps_subset_dependency_exist_flag indicates that the subset of frames in the bitstream that reference the SPS is independent of other subsets of the frame. Here, dependency means that information from other subsets is used (referenced) during the coding or decoding of the subset being processed. Independence means that information from other subsets is not used (referenced) during the coding or decoding of the subset being processed.

[0117] A value of 1 for sps_subset_dependent_cabac_flag indicates that the initial entropy context state of a subset of frames referencing the SPS in the bitstream may depend on the final entropy context state of a preceding subset of that frame, and that the determination of the context of that subset of that frame may also depend on the decoded parameters of a preceding subset of that frame. A value of 0 for sps_subset_dependent_cabac_flag indicates that the determination of the context of the first subset of frames referencing the SPS in the bitstream is independent of the second subset of that frame and the second subset preceding the first subset. If subset partitioning is not available, sps_subset_dependent_cabac_flag is set to a value of 0. Here, dependency means that information (context or parameters) of other subsets is used (referenced) during arithmetic coding or arithmetic decoding of the subset being processed. Furthermore, independence means that during arithmetic coding or arithmetic decoding of the subset being processed, information (context or parameters) from other subsets is not used (referenced).

[0118] As shown in these examples, the three-dimensional data encoding device 100 may store information in the sequence parameter set (SPS) indicating whether or not to divide the data of one frame into multiple subsets, for example, sps_subset_enabled_flag. Furthermore, if the three-dimensional data encoding device 100 divides the data of one frame into multiple subsets, it may store in the SPS information regarding the size of the subsets, for example, sps_subset_size_minus1, and information indicating whether or not to allow the processing subset to reference information of processed subsets within the same frame, for example, sps_subset_dependency_exist_flag. Additionally, if the three-dimensional data encoding device 100 allows the processing subset to reference information of processed subsets within the same frame, it may store in the SPS information indicating whether or not to allow the processing subset to reference the context of processed subsets within the same frame, for example, sps_subset_dependent_cabac_flag.

[0119] For example, information regarding the size of a subset indicates the range or number of reference positions rm included in one subset. For example, information regarding the size of a subset indicates the number of columns of reference positions included in one subset from among the multiple reference positions shown in Figure 3. Specifically, information regarding a subset indicates the number of column_pos, which represent the index of the horizontal angle φj of the reference position rm. Note that information regarding the size of a subset can be any information that allows a predetermined size to be identified in both the encoding and decoding processes, such as information regarding the size of the encoded data. In other words, information regarding the size of a subset may be information indicating the reference positions or three-dimensional points included in the subset, or it may be information indicating the size of the encoded data.

[0120] Alternatively, sps_subset_size can be used instead of sps_subset_size_minus1. sps_subset_size is set to 0 if subset partitioning is not applied, and to a value of 1 or greater that indicates the size of the subset if subset partitioning is applied.

[0121] Furthermore, if it is permitted to reference information from processed subsets within the same frame from the subset to be processed, the three-dimensional data encoding device 100 may, for example, refer to information on reference positions or three-dimensional points included in other processed subsets within the same frame to select the context of the subset to be processed, or it may inherit and use the context obtained as a result of encoding the previous subset in the encoding order within the same frame. In addition, the three-dimensional data encoding device 100 may, for example, perform intra-prediction by referencing information on reference positions or three-dimensional points included in other processed subsets within the same frame. That is, the three-dimensional data encoding device 100 may calculate a predicted value using information on reference positions or three-dimensional points included in other processed subsets within the same frame, and calculate the difference between the position information of the target point and the predicted value.

[0122] Furthermore, if it is not permitted to reference information from processed subsets within the same frame from the subset to be processed, the three-dimensional data encoding device 100 selects the context of the subset to be processed by referencing information of reference positions or three-dimensional points included in the same subset and without referencing information of reference positions or three-dimensional points included in other subsets. Also, the three-dimensional data encoding device 100 does not inherit the context obtained as a result of encoding the immediately preceding subset in the encoding order within the same frame. In addition, the three-dimensional data encoding device 100 may perform intra-prediction by referencing information of reference positions or three-dimensional points included in the same subset and without referencing information of reference positions or three-dimensional points included in other subsets.

[0123] Furthermore, the three-dimensional data encoding device 100 can restrict the access of information that affects context selection (such as the last context of the previous subset, or information referenced during context selection, such as the reference position and three-dimensional point information) across subset boundaries by storing other information, such as sps_subset_dependent_cabac_flag, in the bitstream.

[0124] As described above, the three-dimensional data encoding device 100 allows access to information of processed subsets even when performing subset partitioning, potentially minimizing the decrease in encoding efficiency due to partitioning and shortening the time required from encoding to decoding. Furthermore, by prohibiting access to information that affects context selection across subset boundaries, arithmetic encoding of individual subsets can be performed independently, potentially improving error tolerance while suppressing the decrease in encoding efficiency. Moreover, by prohibiting access to information of processed subsets, parallel processing on a subset-by-subset basis becomes possible while suppressing partitioning overhead compared to partitioning with headers such as slices, potentially improving the throughput of encoding and decoding processes.

[0125] Although an example of dividing a frame into subsets has been described, other processing units that encode information including the positions of three-dimensional points, such as slices, may also be divided into subsets. Furthermore, although an example of syntax in SPS was shown above, the three-dimensional data encoding device 100 may store all or some of the parameters shown above in the geometry parameter set (GPS), the frame header, or the slice header.

[0126] Furthermore, the same processing as in the three-dimensional data encoding device 100 is performed in the three-dimensional data decoding device 200 for context selection and intra prediction, etc.

[0127] Next, we will explain an example of location information syntax that supports subset partitioning. Figure 9 shows an example of location information syntax that supports subset partitioning. The syntax shown in Figure 9 is the same as the syntax shown in Figure 5, but with added information regarding subset boundaries. Also, the syntax elements from pred_mode onwards are the same as in Figure 5, so they are omitted in Figure 9.

[0128] In the example shown in Figure 9, the three-dimensional data encoding device 100 determines the subset boundary based on column_pos. If a subset boundary is determined, it stores subset termination information (for example, end_of_subset shown in Figure 9) in the bitstream and performs arithmetic code termination processing, including byte alignment. The parameter subset_size shown in Figure 9 is set to the value obtained by adding 1 to sps_subset_size_minus1 shown in Figure 8. The variable subset_bounddary, used for determining the subset boundary, is a variable that represents the next subset boundary position. subset_bounddary is initialized to subset_size at the beginning of the frame. Thereafter, subset_size is added to subset_bounddary each time a subset boundary is determined. In other words, for the multiple reference positions shown in Figure 3, the encoded data is divided according to the number of columns indicated by subset_size. For example, if subset_size is 2, multiple subsets are set up so that each subset contains two columns of encoded data for the multiple reference positions shown in Figure 3.

[0129] Furthermore, the subset end information (end_of_subset) is fixed to a value of 1, and the three-dimensional data encoding device 100 may always perform arithmetic code termination processing when storing subset end information. In this case, the partitioning process (termination processing) is always performed for each number of columns indicated by subset_size.

[0130] Alternatively, the three-dimensional data encoding device 100 may set the subset end information (end_of_subset) to either value 0 or value 1. If the value is 1, it performs arithmetic code termination processing; if the value is 0, it does not need to perform termination processing. Note that if arithmetic code termination processing is not performed (when the subset end information is value 0), the three-dimensional data encoding device 100 also omits the byte alignment (byte_alignment()) immediately following end_of_subset as shown in Figure 9. In this case, by setting the subset end information to value 0, the two subsets can be treated as one subset without performing a splitting process (termination processing) between them. Therefore, the degree of freedom for the size of each subset (e.g., the number of columns) can be improved.

[0131] Alternatively, the three-dimensional data encoding device 100 may not store subset termination information and may always perform arithmetic code termination processing when it determines that a subset boundary has been reached.

[0132] Furthermore, the three-dimensional data encoding device 100 may switch the subset boundary determination on and off according to sps_subset_enabled_flag. If sps_subset_enabled_flag indicates that subset partitioning should not be performed (for example, a value of 0), the three-dimensional data encoding device 100 may turn off subset boundary determination and not perform partitioning.

[0133] Furthermore, the syntax shown in Figure 9 is an example of continuing to process this syntax even after arithmetic code termination processing has been performed, but processing may also be resumed from the beginning of this syntax after arithmetic code termination processing has been performed. In this case, the three-dimensional data encoding device 100 may notify the three-dimensional data decoding device 200 of the column_pos and row_pos values ​​of the starting point after the resumption, prior to the syntax corresponding to the starting point after the resumption. In this case, the three-dimensional data decoding device 200 may apply this syntax using these values.

[0134] Furthermore, while we have explained an example here where content corresponding to subset partitioning has been added to the syntax shown in Figure 5, subset partitioning can also be supported by inserting the same content as in this example immediately after the increment operation of column_pos into the syntax shown in Figure 4.

[0135] The following describes an example of processing at a subset boundary. Figure 10 is a flowchart showing an example of a processing procedure corresponding to the content related to subset boundaries, using the syntax example shown in Figure 9.

[0136] First, the three-dimensional data encoding device 100 determines, for example, whether the reference position to be processed is a subset boundary using the method described above (S101). If the reference position to be processed is a subset boundary (Yes in S101), the three-dimensional data encoding device 100 adds end_of_subset to the bitstream (S102). Next, the three-dimensional data encoding device 100 performs arithmetic code termination processing (S103).

[0137] Next, the three-dimensional data encoding device 100 determines whether the target subset can be processed by referring to the information of the processed subset based on sps_subset_dependency_exist_flag (S104). If the information of the processed subset cannot be referred to (No in S104), the three-dimensional data encoding device 100 resets the buffer that stores the processed parameters referred to in prediction (S105), and resets the buffer and context that store the processed parameters referred to in context selection (S107).

[0138] On the other hand, if information about the processed subset is available (Yes in S104), the three-dimensional data encoding device 100 determines whether it is possible to perform context selection of the subset to be processed by referring to the information about the processed subset based on sps_subset_dependent_cabac_flag (S106). If information about the processed subset is not available (No in S106), the three-dimensional data encoding device 100 resets the buffer and context that store the processed parameters referenced in context selection (S107).

[0139] As mentioned above, whether or not termination processing is performed may be switched depending on the value of end_of_subset. If termination processing is not performed, the processing from step S103 onwards is omitted. Also, if end_of_subset is not used and termination processing is always performed at the subset boundary, step S102 may be omitted.

[0140] Furthermore, the same processing as described above is performed in the three-dimensional data decoding device 200. In step S102, the three-dimensional data decoding device 200 obtains end_of_subset from the bitstream and determines whether or not to perform the processing from step S103 onwards depending on the value of end_of_subset. Also, if the processing in step S102 is omitted in the three-dimensional data encoding device 100, the processing in step S102 is similarly omitted in the three-dimensional data decoding device 200.

[0141] Furthermore, the termination process in the three-dimensional data decoding device 200 includes the process of removing bit patterns added by the byte alignment process performed by the three-dimensional data encoding device 100.

[0142] As described above, by allowing access to information of processed subsets when performing subset partitioning, it may be possible to minimize the decrease in encoding efficiency caused by partitioning while shortening the time required from encoding to decoding. Furthermore, by prohibiting access to information that affects context selection across subset boundaries, it becomes possible to perform arithmetic decoding of individual subsets independently, potentially improving error tolerance while suppressing the decrease in encoding efficiency. Moreover, by prohibiting access to information of processed subsets, it becomes possible to suppress partitioning overhead compared to partitioning with headers such as slices, while enabling parallel processing on a subset-by-subset basis, potentially improving the throughput of encoding and decoding processes.

[0143] The three-dimensional data encoding device 100 may also packetize the encoded data for each subset and transmit it in subset units. This allows for a faster start time for packet transmission, enabling delayed transmission. In this case, the three-dimensional data decoding device 200 can generate the original bitstream by combining the encoded data for each subset contained in the multiple received packets.

[0144] Furthermore, the three-dimensional data encoding device 100 may store subset identification information in the encoded data for each subset to identify the order of the subsets. For example, the subset identification information may be a sequential number or the like.

[0145] This allows the three-dimensional data decoding device 200 to identify the order of the column_pos of a subset based on the subset identification number and information about the size of the subset (for example, the number of column_pos indicating the index of the horizontal angle φj). Therefore, even if some packets that have been packetized in subset units are lost during transmission, the three-dimensional data decoding device 200 can use the subset identification information to determine the lost subset and skip processing for each column of the lost subset.

[0146] Furthermore, because the three-dimensional data decoding device 200 can identify the columns of a subset, it can decode each subset independently even if the subsets are not received in order. In other words, the three-dimensional data decoding device 200 can process multiple encoded data from multiple subsets in parallel. In addition, the three-dimensional data decoding device 200 can use the positions of the columns of each decoded subset to identify the position of each subset within the entire frame, and can reconstruct the point cloud data of a single frame by combining the point cloud data (position information) of multiple subsets.

[0147] In the above explanation, we have used the example of using multiple reference positions as shown in Figure 3, etc., but the above subset partitioning may also be applied when using other encoding methods. For example, the above subset partitioning may be applied when using a prediction tree that shows the reference relationships in prediction. In this case, for example, the stream (encoded data) is divided into multiple subsets by performing termination processing on the stream at predetermined horizontal angles.

[0148] As described above, the three-dimensional data encoding device according to this embodiment performs the processing shown in Figure 11. The three-dimensional data encoding device generates multiple encoded data by encoding information about multiple positions of multiple three-dimensional points included in a first unit (e.g., slice, tile, or frame), which is an encoding unit, for each second unit (e.g., the subset described above) that is smaller than the first unit (S111), and outputs the multiple encoded data (S112). Each of the multiple encoded data does not have individual additional information (e.g., a header and a footer). For example, encoded data for the second unit may be generated sequentially while the encoded data for the first unit is being generated. In this case, the three-dimensional data encoding device may sequentially transmit the generated encoded data for the second unit without waiting for the generation of all the encoded data for the first unit.

[0149] According to this, the three-dimensional data encoding device can output the encoded data of the second unit without waiting for all the encoded data of the first unit to be completed. This reduces the delay time from the generation of encoded data to its output. Furthermore, because the encoded data of the second unit does not contain individual additional information, the increase in processing load required to generate multiple encoded data of the second unit can be suppressed compared to cases where the encoded data of the second unit contains individual additional information. In addition, since the second unit is not restricted by the fact that it is a unit containing additional information, the degree of freedom in generating the encoded data of the second unit can be improved. This allows, for example, multiple encoded data to be adjusted to a size suitable for transmission.

[0150] For example, multiple encoded data sets may share common additional information (e.g., slice header, tile header, frame header, SPS, or GPS).

[0151] For example, common additional information includes first information (e.g., sps_subset_size_minus1) indicating the size of the first encoded data contained in multiple encoded data. In other words, the three-dimensional data encoding device stores the first information in the common additional information. With this, the three-dimensional data decoding device can use the first information to identify the end of the first encoded data. Note that the sizes of the multiple encoded data may be the same or different. If the sizes of the multiple encoded data are different, the first information may indicate the size of each of the multiple encoded data.

[0152] For example, information about multiple positions of multiple three-dimensional points is represented for each of the multiple positions by a distance component, a first directional component, and a second directional component (e.g., distance, horizontal angle, and elevation angle). For example, each of the multiple positions is represented in polar coordinates. The three-dimensional data encoding device encodes information about multiple positions of multiple three-dimensional points using a predetermined set of reference positions. Each of the reference positions includes a first directional component and a second directional component (e.g., horizontal angle and elevation angle). The first information represents the size of the second unit, specifically the size of the first directional component of the first encoded data (e.g., the number of columns included in one second unit in Figure 3). This allows the size of the encoded data to be represented by the size of the first directional component, thus reducing the amount of data in the first information compared to when the size of the encoded data is represented by the distance component, first directional component, and second directional component.

[0153] The first information may indicate a reference position or three-dimensional point included in the first encoded data, or it may indicate the data size (e.g., number of bytes) of the first encoded data. For example, the first information may indicate the number of reference positions or three-dimensional points included in the first encoded data from among a plurality of reference positions. The number of reference positions may also be indicated by the number of columns or rows of reference positions.

[0154] For example, each of the multiple encoded data contains second information (end_of_subset) indicating whether or not to perform termination on that encoded data. In other words, the three-dimensional data encoding device adds this second information to each of the multiple encoded data. For example, the three-dimensional data encoding device switches whether or not to perform termination on the second encoded data to be processed based on the second information contained in that second encoded data.

[0155] According to this, it is possible to choose whether or not to perform terminal processing at each second unit identified by the first information. Therefore, the degree of flexibility in data partitioning can be improved.

[0156] For example, common additional information includes third information (e.g., sps_subset_dependency_exist_flag) indicating whether to use other encoded data included in the multiple encoded data to encode the first encoded data included in the multiple encoded data. In other words, the 3D data encoding device stores the third information in the common additional information. For example, the third information indicates whether the first encoded data refers to other encoded data included in the multiple encoded data. For example, if the third information indicates that other encoded data should be used to encode the first encoded data, the 3D data encoding device will refer to the information of the 3D point corresponding to the other encoded data when encoding the 3D point corresponding to the first encoded data (for example, used for prediction). Alternatively, the 3D data encoding device will determine the context used for arithmetic encoding of the information of the 3D point corresponding to the first encoded data using the information of the 3D point corresponding to the other encoded data. Alternatively, the 3D data encoding device will continue to use the context used for arithmetic encoding of the information of the 3D point corresponding to the other encoded data. On the other hand, if the third information indicates that other encoded data will not be used to encode the first encoded data, the three-dimensional data encoding device will not refer to the information of the three-dimensional points corresponding to the other encoded data when encoding the three-dimensional points corresponding to the first encoded data. Alternatively, the three-dimensional data encoding device will not use the information of the three-dimensional points corresponding to the other encoded data to determine the context used for arithmetic encoding of the information of the three-dimensional points corresponding to the first encoded data. Alternatively, the three-dimensional data encoding device will not continue to use the context used for arithmetic encoding of the information of the three-dimensional points corresponding to the other encoded data.

[0157] This allows switching whether encoded data depends on other encoded data or not. For example, encoding efficiency can be improved by making encoded data dependent on other encoded data. Alternatively, by making encoded data independent, each of multiple encoded data can be processed independently, enabling parallel processing, etc.

[0158] For example, common additional information includes a fourth piece of information (sps_subset_dependent_cabac_flag) indicating whether the context used for arithmetic coding of the first coded data included in multiple coded data depends on other coded data included in the multiple coded data. In other words, the 3D data encoding device stores the fourth piece of information in the common additional information. For example, if the fourth piece of information indicates that the context used for arithmetic coding of the first coded data depends on other coded data, the 3D data encoding device determines the context used for arithmetic coding of the 3D point information corresponding to the first coded data using the information of the 3D point corresponding to the other coded data. Alternatively, the 3D data encoding device continues to use the context used for arithmetic coding of the 3D point information corresponding to the other coded data. On the other hand, if the fourth piece of information indicates that the context used for arithmetic coding of the first coded data does not depend on other coded data, the 3D data encoding device does not determine the context used for arithmetic coding of the 3D point information corresponding to the first coded data using the information of the 3D point corresponding to the other coded data. Alternatively, the three-dimensional data encoding device does not continue to use the context used for arithmetic encoding of the three-dimensional point information corresponding to other encoded data.

[0159] This allows switching whether encoded data depends on other encoded data or not. For example, encoding efficiency can be improved by making encoded data dependent on other encoded data. Alternatively, by making encoded data independent, each of multiple encoded data can be processed independently, enabling parallel processing, etc.

[0160] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0161] Furthermore, the three-dimensional data decoding device according to this embodiment performs the processing shown in Figure 12. The three-dimensional data decoding device acquires multiple encoded data (S121) generated by encoding information about multiple positions of multiple three-dimensional points contained in a first unit (e.g., slice, tile, or frame), which is an encoding unit, for each second unit (e.g., the subset described above) that is smaller than the first unit. By decoding the multiple encoded data, it generates information about multiple positions of multiple three-dimensional points contained in the first unit (S122). Each of the multiple encoded data does not have individual additional information. Note that the acquisition of the multiple encoded data is performed sequentially, and the three-dimensional data decoding device may perform the decoding process sequentially on the received encoded data without waiting for all the encoded data to be received. The three-dimensional data decoding device may also process the multiple encoded data in parallel.

[0162] According to this, the three-dimensional data decoding device can begin decoding the second unit of encoded data without waiting for all of the first unit's encoded data to be completed. This reduces the delay time from receiving the encoded data to starting decoding. Furthermore, because the second unit's encoded data does not contain individual additional information, the increase in processing load required to analyze multiple second unit encoded data can be suppressed compared to cases where the second unit's encoded data contains individual additional information. In addition, since the second unit is not restricted by the presence of additional information, the degree of freedom in generating the second unit's encoded data can be improved. This allows, for example, multiple encoded data to be adjusted to a size suitable for transmission.

[0163] For example, multiple encoded data sets share common additional information (e.g., slice header, tile header, frame header, SPS, or GPS).

[0164] For example, common additional information includes first information (e.g., sps_subset_size_minus1) indicating the size of the first encoded data contained in multiple encoded data. For example, a three-dimensional data decoder obtains the first information from the common additional information and uses the first information to identify the end of the first encoded data. Thus, the three-dimensional data decoder can identify the end of the first encoded data using the first information. Note that the sizes of the multiple encoded data may be the same or different. If the sizes of the multiple encoded data are different, the first information may indicate the size of each of the multiple encoded data.

[0165] For example, each of the multiple location information includes a distance component, a first directional component, and a second directional component (e.g., distance, horizontal angle, and elevation angle). For example, each of the multiple location information is represented in polar coordinates. The three-dimensional data encoding device encodes multiple location information of multiple three-dimensional points using a predetermined set of reference positions. Each of the reference positions includes a first directional component and a second directional component (e.g., horizontal angle and elevation angle). The first information represents the size of the second unit, specifically the size of the first directional component of the second unit (e.g., the number of columns contained in one second unit in Figure 3). This allows the size of the second unit to be represented by the size of the first directional component, thus reducing the amount of data in the first information.

[0166] For example, each of multiple encoded data contains a second piece of information (end_of_subset) indicating whether or not to perform termination on that encoded data. For example, a three-dimensional data decoding device switches whether or not to perform termination on the encoded data based on the second piece of information contained in the encoded data to be processed.

[0167] According to this, in a three-dimensional data encoding device, it is possible to choose whether or not to perform termination processing at the second unit identified by the first information. Therefore, the degree of freedom in data partitioning can be improved.

[0168] For example, common additional information includes third information (e.g., sps_subset_dependency_exist_flag) indicating whether or not other encoded data included in the multiple encoded data is used to encode the first encoded data included in the multiple encoded data. For example, a three-dimensional data decoder obtains third information from common additional information. For example, third information indicates whether or not the first encoded data refers to other decoded encoded data included in the multiple encoded data. For example, if the third information indicates that other encoded data is used to encode the first encoded data, the three-dimensional data decoder refers to the information of the three-dimensional point corresponding to the other decoded encoded data when decoding the three-dimensional point corresponding to the first encoded data (for example, used for prediction). Alternatively, the three-dimensional data decoder determines the context used for arithmetic decoding of the information of the three-dimensional point corresponding to the first encoded data using the information of the three-dimensional point corresponding to the other decoded encoded data. Alternatively, the three-dimensional data decoder continues to use the context used for arithmetic decoding of the information of the three-dimensional point corresponding to the other decoded encoded data. On the other hand, if the third information indicates that other encoded data is not used to encode the first encoded data, the three-dimensional data decoding device will not refer to the information of the three-dimensional points corresponding to the other encoded data that have already been decoded when decoding the three-dimensional points corresponding to the first encoded data. Alternatively, the three-dimensional data decoding device will not use the information of the three-dimensional points corresponding to the other encoded data that have already been decoded to determine the context used for arithmetic decoding of the information of the three-dimensional points corresponding to the first encoded data. Alternatively, the three-dimensional data decoding device will not continue to use the context that was used for arithmetic decoding of the information of the three-dimensional points corresponding to the other encoded data that have already been decoded.

[0169] According to this, a three-dimensional data encoding device can switch whether encoded data depends on other encoded data or not. For example, encoding efficiency can be improved by making encoded data dependent on other encoded data. Alternatively, by making encoded data independent of other encoded data, each of the multiple encoded data can be processed independently, thus enabling parallel processing, etc.

[0170] For example, common additional information includes a fourth piece of information (sps_subset_dependent_cabac_flag) indicating whether the context used for arithmetic decoding of the first encoded data contained in multiple encoded data depends on other encoded data contained in the multiple encoded data. For example, a three-dimensional data decoder obtains the fourth piece of information from the common additional information. For example, if the fourth piece of information indicates that the context used for arithmetic decoding of the first encoded data depends on other encoded data, the three-dimensional data decoder determines the context used for arithmetic decoding of the three-dimensional point information corresponding to the first encoded data using the information of the three-dimensional point corresponding to the other decoded encoded data. Alternatively, the three-dimensional data decoder continues to use the context used for arithmetic decoding of the information of the three-dimensional point corresponding to the other decoded encoded data. On the other hand, if the fourth piece of information indicates that the context used for arithmetic decoding of the first encoded data does not depend on other encoded data, the three-dimensional data decoder does not determine the context used for arithmetic decoding of the three-dimensional point information corresponding to the first encoded data using the information of the three-dimensional point corresponding to the other decoded encoded data. Alternatively, the three-dimensional data decoding device does not continue to use the context used for arithmetic decoding of the three-dimensional point information corresponding to other encoded data that has already been decoded.

[0171] According to this, a three-dimensional data encoding device can switch whether encoded data depends on other encoded data or not. For example, encoding efficiency can be improved by making encoded data dependent on other encoded data. Alternatively, by making encoded data independent of other encoded data, each of the multiple encoded data can be processed independently, thus enabling parallel processing, etc.

[0172] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0173] Although embodiments and modified examples of the three-dimensional data encoding device and three-dimensional data decoding device, etc., of the present disclosure have been described above, the present disclosure is not limited to these embodiments.

[0174] Furthermore, each processing unit included in the three-dimensional data encoding device and the three-dimensional data decoding device according to the above embodiment is typically implemented as an integrated circuit (LSI). These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip.

[0175] Furthermore, integrated circuit implementation is not limited to LSIs; it may also be achieved using dedicated circuits or general-purpose processors. Field-Programmable Gate Arrays (FPGAs), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow for the reconfiguration of the connections and settings of circuit cells within the LSI, may also be used.

[0176] Furthermore, in each of the above embodiments, each component may be implemented by being composed of dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0177] Furthermore, this disclosure may be implemented as a three-dimensional data encoding method or a three-dimensional data decoding method, etc., performed by a three-dimensional data encoding device and a three-dimensional data decoding device, etc.

[0178] Furthermore, the division of functional blocks in the block diagram is just one example; multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple parts, or some functions can be moved to other functional blocks. In addition, the functions of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.

[0179] Furthermore, the order in which each step in the flowchart is performed is illustrative for the purpose of specifically illustrating this disclosure, and may be in a different order. Also, some of the above steps may be performed simultaneously (in parallel) with other steps.

[0180] Although a three-dimensional data encoding device and a three-dimensional data decoding device, etc., relating to one or more embodiments have been described above based on embodiments, this disclosure is not limited to these embodiments. Without departing from the spirit of this disclosure, various modifications that a person skilled in the art can conceive of may be applied to these embodiments, and forms constructed by combining components from different embodiments may also be included within the scope of one or more embodiments. [Industrial applicability]

[0181] This disclosure is applicable to three-dimensional data encoding devices and three-dimensional data decoding devices. [Explanation of Symbols]

[0182] 100 3D data encoding device 101 Conversion Unit 102 Subtraction Unit 103 Quantization section 104 Entropy coding unit 105, 202 Inverse quantization section 106, 203 Addition section 108, 110, 205, 207 buffers 109, 206 Intra Prediction Unit 111 Motion detection compensation unit 112, 209 Interpretation Unit 113, 210 Switching section 114, 211 Control Unit 200 Three-dimensional data decoding device 201 Entropy Decoder 204 Inverse Transform Section 208 Motion compensation unit

Claims

1. Multiple encoded data are generated by encoding information about multiple positions of multiple three-dimensional points contained in a first unit, which is an encoding unit, into second units that are smaller than the first unit. Output the multiple encoded data mentioned above, Each of the aforementioned multiple encoded data does not have individual additional information. The aforementioned multiple encoded data have common additional information, The aforementioned common additional information includes first information indicating the size of the first encoded data included in the plurality of encoded data. Three-dimensional data encoding method.

2. The information regarding the plurality of positions of the plurality of three-dimensional points is expressed by representing each of the plurality of positions with a distance component, a first directional component, and a second directional component. The information regarding the multiple positions of the multiple three-dimensional points is encoded using a predetermined set of reference positions. Each of the plurality of reference positions includes the first directional component and the second directional component, The first information indicates the magnitude of the first directional component of the first encoded data. The three-dimensional data encoding method according to claim 1.

3. Each of the plurality of encoded data includes second information indicating whether or not termination processing is performed on the encoded data. The three-dimensional data encoding method according to claim 1.

4. The aforementioned common additional information includes third information indicating whether or not other encoded data included in the plurality of encoded data is used to encode the first encoded data included in the plurality of encoded data. The three-dimensional data encoding method according to claim 1.

5. The aforementioned common additional information includes a fourth piece of information indicating whether the context used for arithmetic coding of the first coded data included in the plurality of coded data depends on other coded data included in the plurality of coded data. The three-dimensional data encoding method according to claim 1.

6. The size indicated by the first information is a horizontal angle. The three-dimensional data encoding method according to claim 1.

7. Multiple encoded data are obtained by encoding information about multiple positions of multiple three-dimensional points contained in a first unit, which is an encoding unit, in units smaller than the first unit, for each second unit. By decoding the plurality of encoded data, information about the multiple positions of the multiple three-dimensional points included in the first unit is generated. Each of the aforementioned multiple encoded data does not have individual additional information. The aforementioned multiple encoded data have common additional information, The aforementioned common additional information includes first information indicating the size of the first encoded data included in the plurality of encoded data. Three-dimensional data decoding method.

8. The information regarding the plurality of positions of the plurality of three-dimensional points is expressed by representing each of the plurality of positions with a distance component, a first directional component, and a second directional component. The information regarding the multiple positions of the multiple three-dimensional points is encoded using a predetermined set of reference positions. Each of the plurality of reference positions includes the first directional component and the second directional component, The first information indicates the magnitude of the first directional component of the first encoded data. The method for decoding three-dimensional data according to claim 7.

9. Each of the plurality of encoded data includes second information indicating whether or not termination processing is performed on the encoded data. The method for decoding three-dimensional data according to claim 7.

10. The aforementioned common additional information includes third information indicating whether or not other encoded data included in the plurality of encoded data is used to decode the first encoded data included in the plurality of encoded data. The method for decoding three-dimensional data according to claim 7.

11. The aforementioned common additional information includes a fourth piece of information indicating whether the context used for arithmetic decoding of the first encoded data included in the plurality of encoded data depends on other encoded data included in the plurality of encoded data. The method for decoding three-dimensional data according to claim 7.

12. The size indicated by the first information is a horizontal angle. The method for decoding three-dimensional data according to claim 7.

13. Processor and Equipped with memory, The processor uses the memory to: Multiple encoded data are generated by encoding information about multiple positions of multiple three-dimensional points contained in a first unit, which is an encoding unit, into second units that are smaller than the first unit. Output the multiple encoded data mentioned above, Each of the aforementioned multiple encoded data does not have individual additional information. The aforementioned multiple encoded data have common additional information, The aforementioned common additional information includes first information indicating the size of the first encoded data included in the plurality of encoded data. Three-dimensional data encoding device.

14. Processor and Equipped with memory, The processor uses the memory to: Multiple encoded data are obtained by encoding information about multiple positions of multiple three-dimensional points contained in a first unit, which is an encoding unit, in units smaller than the first unit, for each second unit. By decoding the plurality of encoded data, information about the multiple positions of the multiple three-dimensional points included in the first unit is generated. Each of the aforementioned multiple encoded data does not have individual additional information. The aforementioned multiple encoded data have common additional information, The aforementioned common additional information includes first information indicating the size of the first encoded data included in the plurality of encoded data. Three-dimensional data decoding device.