Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
By using the possession style of the N forktree structure in three-dimensional data encoding and decoding, the encoding method is determined, and the problem of low efficiency of three-dimensional data encoding in the prior art is solved, and more efficient data compression and recovery is achieved.
Patent Information
- Application Number
- CN202411961444.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-27
- Filing Date
- 2019-06-26
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is relatively inefficient in encoding and decoding of three-dimensional data, and it is difficult to effectively compress large-scale three-dimensional point cloud data.
A new three-dimensional data encoding method is adopted to generate an occupation style adjacent to the nodes in the N-forktree structure of multiple three-dimensional points, and determine whether to use the first encoding or the second encoding to encode and decode the three-dimensional data.
It improves the efficiency of 3D data encoding and decoding, and can compress and recover 3D point cloud data more effectively.
Smart Images

Figure CN120014079A_ABST
Abstract
Description
[0001] This application is a division of an invention patent application with an application date of June 26, 2019, application number 201980040980.9, and invention name “Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device”. Technical Field
[0002] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. Background Art
[0003] In the future, devices and services that make use of 3D data will become more common in large fields such as computer vision, map information, monitoring, infrastructure inspection, or image distribution, which are used for autonomous operation of cars or robots. 3D data is obtained by various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple single-lens reflex cameras.
[0004] As a method of expressing three-dimensional data, there is a method called point cloud, which expresses the shape of a three-dimensional structure through a group of points in a three-dimensional space. The position and color of the point group are stored in the point cloud. Although point cloud is expected to become the mainstream method of expressing three-dimensional data, the amount of point group data is very large. Therefore, in the accumulation or transmission of three-dimensional data, it is necessary to compress the data volume through encoding, just like two-dimensional dynamic images (as an example, there are MPEG-4AVC or HEVC standardized by MPEG).
[0005] Furthermore, compression of point clouds is partially supported by a public library (PointCloud Library) that performs point cloud association processing.
[0006] Furthermore, there is a known technique for searching for facilities around a vehicle using three-dimensional map data and displaying the facilities (for example, refer to Patent Document 1).
[0007] Prior art literature
[0008] Patent Literature
[0009] Patent Document 1 International Publication No. 2014 / 020663 Summary of the invention
[0010] Problem that the invention aims to solve
[0011] It is desirable to improve coding efficiency in encoding and decoding of three-dimensional data.
[0012] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency.
[0013] Means used to solve problems
[0014] A three-dimensional data encoding method in one form disclosed herein generates an occupancy pattern of adjacent nodes adjacent to a node in an N (N is an integer greater than 2) fork tree structure of multiple three-dimensional points, wherein the multiple three-dimensional points are included in the three-dimensional data, and determines whether to set a candidate node that can use the first encoding based on the occupancy pattern.
[0015] A three-dimensional data decoding method in one form disclosed herein obtains parameters from a bit stream, obtains an occupancy pattern of adjacent nodes adjacent to nodes in an N (N is an integer greater than 2) fork tree structure of multiple three-dimensional points, wherein the multiple three-dimensional points are included in the three-dimensional data, and determines whether to set a candidate node that can be used for the first decoding based on the occupancy pattern.
[0016] Effects of the Invention
[0017] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 The structure of the encoded three-dimensional data according to the first embodiment is shown.
[0019] Figure 2 An example of a prediction structure between SPCs belonging to the lowest layer of the GOS according to the first embodiment is shown.
[0020] Figure 3 An example of an inter-layer prediction structure in embodiment 1 is shown.
[0021] Figure 4 An example of the encoding order of the GOS according to the first embodiment is shown.
[0022] Figure 5 An example of the encoding order of the GOS according to the first embodiment is shown.
[0023] Figure 6 This is a block diagram of a three-dimensional data encoding device according to Embodiment 1.
[0024] Figure 7 This is a flowchart of the encoding process of implementation mode 1.
[0025] Figure 8 This is a block diagram of a three-dimensional data decoding device according to Embodiment 1.
[0026] Fig. 9 This is a flowchart of the decoding process in implementation mode 1.
[0027] Fig.10 An example of meta-information according to the first embodiment is shown.
[0028] Fig.11 A configuration example of a SWLD according to the second embodiment is shown.
[0029] Fig.12 An operation example of the server and the client according to the second embodiment is shown.
[0030] Fig.13 An operation example of the server and the client according to the second embodiment is shown.
[0031] Fig.14 An operation example of the server and the client according to the second embodiment is shown.
[0032] Fig.15 An operation example of the server and the client according to the second embodiment is shown.
[0033] Fig.16 This is a block diagram of a three-dimensional data encoding device according to Embodiment 2.
[0034] Fig.17 This is a flowchart of the encoding process of implementation mode 2.
[0035] Fig.18 This is a block diagram of a three-dimensional data decoding device according to Embodiment 2.
[0036] Fig.19 This is a flowchart of the decoding process of implementation mode 2.
[0037] Fig. 20 A configuration example of a WLD according to the second embodiment is shown.
[0038] Fig.21 An example of the octree structure of the WLD according to the second embodiment is shown.
[0039] Fig. 22 A configuration example of a SWLD according to the second embodiment is shown.
[0040] Fig.23 An example of the octree structure of the SWLD according to the second embodiment is shown.
[0041] Fig.24 This is a block diagram of a three-dimensional data creation device according to the third embodiment.
[0042] Fig.25 This is a block diagram of a three-dimensional data transmitting device according to Embodiment 3.
[0043] Fig.26This is a block diagram of a three-dimensional information processing device according to a fourth embodiment.
[0044] Fig. 27 This is a block diagram of a three-dimensional data creation device according to a fifth embodiment.
[0045] Fig.28 The configuration of the system according to the sixth embodiment is shown.
[0046] Fig.29 This is a block diagram of a client device according to a sixth embodiment.
[0047] Fig.30 This is a block diagram of a server in implementation mode 6.
[0048] Fig.31 This is a flowchart of the three-dimensional data creation process performed by the client device of the sixth embodiment.
[0049] Fig.32 This is a flowchart of the sensor information transmission process performed by the client device according to the sixth embodiment.
[0050] Fig.33 This is a flowchart of the three-dimensional data production processing performed by the server of the sixth embodiment.
[0051] Fig.34 This is a flowchart of the three-dimensional map transmission processing performed by the server in the sixth embodiment.
[0052] Fig.35 The configuration of a modified example of the system of the sixth embodiment is shown.
[0053] Fig.36 The configuration of the server and client device according to the sixth embodiment is shown.
[0054] Fig.37 This is a block diagram of a three-dimensional data encoding device according to embodiment 7.
[0055] Fig.38 An example of the prediction residual in Embodiment 7 is shown.
[0056] Fig.39 An example of volume in Embodiment 7 is shown.
[0057] Fig.40 An example of octree representation of volume in Embodiment 7 is shown.
[0058] Fig.41 An example of a bit string of volume in Implementation Example 7 is shown.
[0059] Fig.42 An example of octree representation of volume in Embodiment 7 is shown.
[0060] Fig.43 An example of volume in Embodiment 7 is shown.
[0061] Fig.44 This is a diagram for explaining the intra-frame prediction processing of embodiment 7.
[0062] Fig.45 This is a diagram used to illustrate the rotation and translation processing of embodiment 7.
[0063] Fig.46 An example of the syntax of the RT application flag and RT information according to the seventh embodiment is shown.
[0064] Fig.47 This is a diagram used to illustrate the inter-frame prediction processing of embodiment 7.
[0065] Fig.48 This is a block diagram of a three-dimensional data decoding device according to the seventh embodiment.
[0066] Fig.49 This is a flowchart of a three-dimensional data encoding process performed by the three-dimensional data encoding device according to the seventh embodiment.
[0067] Fig.50 This is a flowchart of a three-dimensional data decoding process performed by the three-dimensional data decoding device according to the seventh embodiment.
[0068] Fig.51 This is a diagram showing the reference relationship in the octree structure of implementation mode 8.
[0069] Fig.52 This is a diagram showing the reference relationship in the spatial area of Implementation Example 8.
[0070] Fig.53 This is a diagram showing an example of adjacent reference nodes according to the eighth embodiment.
[0071] Fig.54 This is a diagram showing the relationship between a parent node and a node in implementation mode 8.
[0072] Fig.55 This is a diagram showing an example of occupancy coding of a parent node in implementation mode 8.
[0073] Fig.56 This is a block diagram showing a three-dimensional data encoding device according to an eighth embodiment.
[0074] Fig.57 This is a block diagram showing a three-dimensional data decoding device according to an eighth embodiment.
[0075] Fig.58 It is a flowchart showing the three-dimensional data encoding processing of implementation mode 8.
[0076] Fig.59It is a flowchart showing the three-dimensional data encoding processing of implementation mode 8.
[0077] Fig.60 This is a diagram showing an example of switching the coding table in implementation mode 8.
[0078] Fig.61 This is a diagram showing the reference relationship in the spatial region of Modification 1 of Implementation Example 8.
[0079] Fig.62 This is a diagram showing a syntax example of header information according to variant example 1 of implementation example 8.
[0080] Fig.63 This is a diagram showing a syntax example of header information according to variant example 1 of implementation example 8.
[0081] Fig.64 This is a diagram showing an example of adjacent reference nodes according to variant example 2 of implementation example 8.
[0082] Fig.65 This is a diagram showing an example of a target node and adjacent nodes according to variation 2 of implementation example 8.
[0083] Fig.66 This is a diagram of the reference relationship in the octree structure of variant example 3 of implementation example 8.
[0084] Fig.67 This is a diagram showing the reference relationship in the spatial region of variant example 3 of implementation example 8.
[0085] Fig.68 This is a diagram showing examples and processing of adjacent nodes in implementation mode 9.
[0086] Fig.69 This is a flowchart of the three-dimensional data encoding process of the ninth embodiment.
[0087] Fig.70 This is a flowchart of the three-dimensional data encoding process of the ninth embodiment.
[0088] Fig.71 This is a flowchart of a modified example of the three-dimensional data encoding process of the ninth embodiment.
[0089] Fig.72 This is a flowchart of the three-dimensional data decoding process according to the ninth embodiment.
[0090] Fig.73 This is a flowchart of a modified example of the three-dimensional data decoding process of the ninth embodiment.
[0091] Fig.74 This is a diagram showing a syntax example of the header of Implementation Example 9.
[0092] Fig.75This is a diagram showing a syntax example of node information in implementation mode 9.
[0093] Fig.76 This is a block diagram of a three-dimensional data encoding device according to embodiment 9.
[0094] Fig.77 This is a block diagram of a three-dimensional data decoding device according to the ninth embodiment.
[0095] Fig.78 This is a flowchart of a modified example of the three-dimensional data encoding process of the ninth embodiment.
[0096] Fig.79 This is a flowchart of a modified example of the three-dimensional data encoding process of the ninth embodiment.
[0097] Fig.80 This is a flowchart of a modified example of the three-dimensional data decoding process of the ninth embodiment.
[0098] Fig.81 This is a flowchart of a modified example of the three-dimensional data decoding process of the ninth embodiment.
[0099] Fig.82 This is a flowchart of the three-dimensional data encoding process of the ninth embodiment.
[0100] Fig.83 This is a flowchart of the three-dimensional data decoding process according to the ninth embodiment. DETAILED DESCRIPTION
[0101] A three-dimensional data encoding method in one form of the present disclosure generates a first occupancy pattern representing the occupancy status of multiple second neighboring nodes when a first flag represents a first value, the multiple second neighboring nodes include a first neighboring node whose parent node is different from the parent node of an object node, the object node is included in an N-ary tree structure of multiple three-dimensional points included in the three-dimensional data, N is an integer greater than 2, based on the first occupancy pattern, determines whether a first code can be used to encode multiple three-dimensional position information included in the object node without dividing the object node into multiple child nodes, generates a second occupancy pattern representing the occupancy status of multiple third neighboring nodes when the first flag represents a second value different from the first value, the multiple third neighboring nodes do not include the first neighboring node whose parent node is different from the parent node of the object node, determines whether the first code can be used based on the second occupancy pattern, and generates a bit stream including the first flag.
[0102] Thus, the three-dimensional data encoding method can switch the adjacent node occupation pattern for determining whether the first encoding can be used according to the first flag. Thus, it is possible to appropriately determine whether the first encoding can be used, thereby improving encoding efficiency.
[0103] For example, it may also be that when it is determined that the first encoding can be used, whether to use the first encoding is determined based on a specified condition, and when it is determined that the first encoding is used, the object node is encoded using the first encoding; when it is determined that the first encoding is not used, the object node is encoded using a second encoding that divides the object node into a plurality of child nodes, and the bit stream further includes a second flag indicating whether the first encoding is used.
[0104] For example, in the determination of whether the first code can be used based on the first occupancy pattern or the second occupancy pattern, it may be determined whether the first code can be used based on the first occupancy pattern or the second occupancy pattern and the number of nodes in the occupancy state included in the parent node.
[0105] For example, in the determination of whether the first coding can be used based on the first occupancy pattern or the second occupancy pattern, it may be determined whether the first coding can be used based on the number of nodes in the occupancy state included in the first occupancy pattern or the second occupancy pattern and the grandparent node of the object node.
[0106] For example, in the determination of whether the first code can be used based on the first occupancy pattern or the second occupancy pattern, whether the first code can be used may be determined based on the first occupancy pattern or the second occupancy pattern and the layer to which the target node belongs.
[0107] A three-dimensional data decoding method according to one aspect of the present disclosure obtains a first flag from a bit stream, generates a first occupancy pattern indicating occupancy states of a plurality of second neighboring nodes when the first flag indicates a first value, the plurality of second neighboring nodes include a first neighboring node whose parent node is different from the parent node of an object node, the object node is included in an N-ary tree structure of a plurality of three-dimensional points included in the three-dimensional data, N being an integer greater than 2, determines based on the first occupancy pattern whether a first decoding can be used for decoding a plurality of three-dimensional position information included in the object node without dividing the object node into a plurality of child nodes, generates a second occupancy pattern indicating occupancy states of a plurality of third neighboring nodes when the first flag indicates a second value different from the first value, the plurality of third neighboring nodes do not include the first neighboring node whose parent node is different from the parent node of the object node, and determines based on the second occupancy pattern whether the first decoding can be used.
[0108] Thus, the three-dimensional data decoding method can switch the adjacent node occupation pattern for determining whether the first code can be used according to the first flag. Thus, it is possible to appropriately determine whether the first code can be used, thereby improving coding efficiency.
[0109] For example, when it is determined that the first decoding can be used, a second flag indicating whether the first decoding is used is obtained from the bit stream, and when the second flag indicates that the first decoding is used, the first decoding is used to decode the object node; when the second flag indicates that the first decoding is not used, the second decoding that divides the object node into a plurality of child nodes is used to decode the object node.
[0110] For example, in the determination of whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern, it may be determined whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern and the number of nodes in an occupied state included in the parent node.
[0111] For example, in the determination of whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern, it may be determined whether the first decoding can be used based on the number of nodes in an occupied state included in the first occupancy pattern or the second occupancy pattern and a grandparent node of the object node.
[0112] For example, in the determination of whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern, whether the first decoding can be used may be determined based on the first occupancy pattern or the second occupancy pattern and the layer to which the target node belongs.
[0113] In addition, a three-dimensional data encoding device according to one aspect of the present disclosure is a three-dimensional data encoding device that encodes multiple three-dimensional points having attribute information, and includes a processor and a memory. The processor uses the memory to generate a first occupancy pattern indicating an occupancy state of multiple second neighboring nodes when a first flag indicates a first value, the multiple second neighboring nodes including a first neighboring node whose parent node is different from a parent node of an object node, the object node being included in an N-ary tree structure of multiple three-dimensional points included in the three-dimensional data, N being an integer greater than or equal to 2, based on the first occupancy pattern, determine whether a first code that encodes multiple three-dimensional position information included in the object node without dividing the object node into multiple child nodes can be used, and when the first flag indicates a second value different from the first value, generate a second occupancy pattern indicating an occupancy state of multiple third neighboring nodes, the multiple third neighboring nodes do not include the first neighboring node whose parent node is different from the parent node of the object node, and determine whether the first code can be used based on the second occupancy pattern to generate a bit stream including the first flag.
[0114] Thus, the three-dimensional data encoding device can switch the adjacent node occupation pattern for determining whether the first encoding can be used according to the first flag. Thus, it is possible to appropriately determine whether the first encoding can be used, thereby improving encoding efficiency.
[0115] In addition, a three-dimensional data decoding device according to one aspect of the present disclosure is a three-dimensional data decoding device that decodes multiple three-dimensional points having attribute information, and includes a processor and a memory. The processor uses the memory to generate a first occupancy pattern indicating occupancy states of multiple second neighboring nodes when the first flag indicates a first value, the multiple second neighboring nodes including a first neighboring node whose parent node is different from the parent node of an object node, the object node being included in an N-ary tree structure of multiple three-dimensional points included in the three-dimensional data, N being an integer greater than or equal to 2, based on the first occupancy pattern, determine whether a first decoding that decodes multiple three-dimensional position information included in the object node without dividing the object node into multiple child nodes can be used, and when the first flag indicates a second value different from the first value, generate a second occupancy pattern indicating occupancy states of multiple third neighboring nodes, the multiple third neighboring nodes do not include the first neighboring node whose parent node is different from the parent node of the object node, and determine whether the first decoding can be used based on the second occupancy pattern.
[0116] Thus, the three-dimensional data decoding device can switch the adjacent node occupation pattern for determining whether the first code can be used according to the first flag. Thus, it is possible to appropriately determine whether the first code can be used, thereby improving coding efficiency.
[0117] In addition, these general or specific forms can be implemented by systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, and can be implemented by any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0118] The following detailed description of the implementation mode is given with reference to the accompanying drawings. In addition, the implementation modes to be described below are all specific examples of the present disclosure. The numerical values, shapes, materials, constituent elements, configuration positions of constituent elements, connection forms, steps, order of steps, etc. shown in the following implementation modes are all examples, and the main purpose is not to limit the present disclosure. Furthermore, the constituent elements of the following implementation modes that are not recorded in the technical solution showing the highest concept are described as arbitrary constituent elements.
[0119] (Implementation Method 1)
[0120] First, the data structure of encoded three-dimensional data (hereinafter also referred to as encoded data) according to the present embodiment will be described. Figure 1 The structure of the encoded three-dimensional data involved in this embodiment is shown.
[0121] In this embodiment, the three-dimensional space is divided into a space (SPC) equivalent to a picture in the encoding of a dynamic image, and the three-dimensional data is encoded in units of space. The space is further divided into volumes (VLM) equivalent to macroblocks in dynamic image encoding, and prediction and conversion are performed in units of VLM. The volume includes a minimum unit corresponding to a position coordinate, namely a plurality of voxels (VXL). In addition, prediction means that, similar to the prediction performed in a two-dimensional image, predicted three-dimensional data similar to the processing unit of the processing object is generated with reference to other processing units, and the difference between the predicted three-dimensional data and the processing unit of the processing object is encoded. Furthermore, the prediction includes not only spatial prediction with reference to other prediction units at the same time, but also temporal prediction with reference to prediction units at different times.
[0122] For example, when encoding a three-dimensional space represented by point group data such as a point cloud, a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes each point of the point group or a plurality of points contained in a voxel according to the size of the voxel. If the voxel is subdivided, the three-dimensional shape of the point group can be expressed with high accuracy, and if the size of the voxel is increased, the three-dimensional shape of the point group can be roughly expressed.
[0123] In addition, although the following description takes the case where the three-dimensional data is a point cloud as an example, the three-dimensional data is not limited to the point cloud, and can also be three-dimensional data in any form.
[0124] Furthermore, voxels of a hierarchical structure may be used. In this case, in the n-order hierarchy, whether or not a sampling point exists in the n-1-order or lower hierarchy (the lower hierarchy of the n-order hierarchy) may be sequentially indicated. For example, when decoding only the n-order hierarchy, if a sampling point exists in the n-1-order or lower hierarchy, decoding may be performed by assuming that a sampling point exists at the center of a voxel in the n-order hierarchy.
[0125] Furthermore, the encoding device obtains point group data through a distance sensor, a stereo camera, a monocular camera, a gyroscope, or an inertial sensor.
[0126] As with the encoding of moving images, the space is classified into at least one of the following three prediction structures: an intra-frame space (I-SPC) that can be decoded independently, a prediction space (P-SPC) that can only be referenced unidirectionally, and a bidirectional space (B-SPC) that can be referenced bidirectionally. In addition, the space has two types of time information: decoding time and display time.
[0127] And, if Figure 1 As shown, as a processing unit including a plurality of spaces, there is a GOS (Group Of Space) which is a random access unit. Also, as a processing unit including a plurality of GOS, there is a world space (WLD).
[0128] The spatial area occupied by the world space is associated with an absolute position on the earth through GPS or latitude and longitude information. This position information is stored as meta information. In addition, the meta information can be included in the encoded data or transmitted separately from the encoded data.
[0129] Furthermore, within the GOS, all SPCs may be three-dimensionally adjacent, or there may be SPCs that are not three-dimensionally adjacent to other SPCs.
[0130] In addition, the encoding, decoding or referencing of the three-dimensional data included in the processing unit such as GOS, SPC or VLM is also simply referred to as encoding, decoding or referencing the processing unit. The three-dimensional data included in the processing unit includes at least one set of spatial positions such as three-dimensional coordinates and characteristic values such as color information.
[0131] Next, the prediction structure of the SPC in the GOS will be described. Although multiple SPCs in the same GOS or multiple VLMs in the same SPC occupy different spaces, they have the same time information (decoding time and display time).
[0132] Furthermore, in the GOS, the first SPC in the decoding order is the I-SPC. Furthermore, there are two types of GOS, the closed GOS and the open GOS. The closed GOS is a GOS that can decode all the SPCs in the GOS when decoding starts from the first I-SPC. In the open GOS, in the GOS, some SPCs earlier than the display time of the first I-SPC refer to different GOSs and can only be decoded in the GOS.
[0133] In addition, in the case of coded data such as map information, the WLD may be decoded in the reverse direction of the coding order. If there is a dependency between GOS, it is difficult to reproduce the data in the reverse direction. Therefore, in this case, a closed GOS is basically used.
[0134] Furthermore, the GOS has a layer structure in the height direction, and encoding or decoding is performed sequentially starting from the SPC of the bottom layer.
[0135] Figure 2 An example of a prediction structure between SPCs belonging to the lowest layer of the GOS is shown. Figure 3 An example of an inter-layer prediction structure is shown.
[0136] There are more than one I-SPC in the GOS. Although there are objects such as people, animals, cars, bicycles, traffic lights, or buildings that serve as land landmarks in the three-dimensional space, it is particularly effective to encode small-sized objects as I-SPCs. For example, when a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes the GOS at a low processing amount or high speed, it only decodes the I-SPC in the GOS.
[0137] Furthermore, the encoding device may switch the encoding interval or the occurrence frequency of the I-SPC according to the density of the objects in the WLD.
[0138] And, in Figure 3 In the configuration shown, the encoding device or decoding device encodes or decodes the plurality of layers sequentially from the lower layer (layer 1). This allows, for example, autonomous vehicles to prioritize data near the ground with a large amount of information.
[0139] In addition, in the coded data used by a drone or the like, coding or decoding may be performed sequentially starting from the SPC of the upper layer in the height direction within the GOS.
[0140] Furthermore, the encoding device or decoding device may encode or decode multiple layers in such a way that the decoding device roughly grasps the GOS and can gradually increase the resolution. For example, the encoding device or decoding device may encode or decode in the order of layers 3, 8, 1, 9, ...
[0141] Next, the corresponding method of the static object and the dynamic object is described.
[0142] In three-dimensional space, there are static objects or scenes such as buildings and roads (hereinafter collectively referred to as static objects), and dynamic objects such as vehicles and people (hereinafter referred to as dynamic objects). Object detection can be performed by extracting feature points from point cloud data or images captured by stereo cameras. Here, an example of a method for encoding dynamic objects is described.
[0143] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects using identification information.
[0144] For example, GOS is used as the identification unit. In this case, GOS including SPCs constituting static objects and GOS including SPCs constituting dynamic objects are distinguished within the coded data or by identification information stored separately from the coded data.
[0145] Alternatively, SPC is used as the identification unit. In this case, the SPC including only the VLM constituting the static object and the SPC including the VLM constituting the dynamic object are distinguished by the above-mentioned identification information.
[0146] Alternatively, VLM or VXL may be used as the identification unit. In this case, the VLM or VXL including the static object and the VLM or VXL including the dynamic object are distinguished by the above-mentioned identification information.
[0147] Furthermore, the encoding device may encode the dynamic object as one or more VLMs or SPCs, and encode the VLM or SPC including the static object and the SPC including the dynamic object as different GOSs. Furthermore, when the size of the GOS becomes variable according to the size of the dynamic object, the encoding device may store the size of the GOS separately as meta-information.
[0148] Furthermore, the encoding device encodes the static object and the dynamic object independently of each other, and the dynamic object can be overlapped with respect to the world space composed of the static object. In this case, the dynamic object is composed of one or more SPCs, and each SPC corresponds to one or more SPCs constituting the static object overlapped with the SPC. In addition, the dynamic object may not be represented by an SPC, but may be represented by one or more VLMs or VXLs.
[0149] Furthermore, the encoding device may encode static objects and dynamic objects as different streams.
[0150] Furthermore, the encoding device may generate a GOS including one or more SPCs constituting a dynamic object. Furthermore, the encoding device may set the GOS (GOS_M) including the dynamic object and the GOS of the static object corresponding to the spatial region of the GOS_M to be of the same size (occupying the same spatial region). In this way, overlapping processing can be performed in units of GOS.
[0151] The P-SPC or B-SPC constituting the dynamic object may also refer to the SPC included in the encoded different GOS. When the position of the dynamic object changes over time and the same dynamic object is encoded as the GOS at different times, cross-GOS reference is effective from the perspective of compression rate.
[0152] Furthermore, the first method and the second method may be switched according to the purpose of the encoded data. For example, when the encoded three-dimensional data is used as a map, it is desirable to separate the three-dimensional data from the dynamic objects, so the encoding device uses the second method. In addition, when the encoding device encodes the three-dimensional data of an event such as a concert or sports, if it is not necessary to separate the dynamic objects, the first method is used.
[0153] Furthermore, the decoding time and display time of GOS or SPC can be stored in the encoded data or as meta-information. Furthermore, the time information of static objects can all be the same. At this time, the actual decoding time and display time can be determined by the decoding device. Alternatively, different values can be assigned to each GOS or SPC as the decoding time, and the same value can be assigned to all the display times. Moreover, as shown in the decoder mode in dynamic image encoding such as HEVC's HRD (Hypothetical Reference Decoder), the decoder has a buffer of a specified size. As long as the bit stream is read at a specified bit rate according to the decoding time, a model that will not be destroyed and is guaranteed to be decodable can be imported.
[0154] Next, the configuration of the GOS in the world space is described. The coordinates of the three-dimensional space in the world space are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, and z-axis). By setting a prescribed rule in the encoding order of the GOS, spatially adjacent GOS can be encoded continuously in the encoded data. For example, Figure 4 In the example shown, the GOS in the xz plane is continuously encoded. After the encoding of all GOS in an xz plane is completed, the value of the y axis is updated. That is, as the encoding continues, the world space extends in the y axis direction. And the index number of the GOS is set as the encoding order.
[0155] Here, the three-dimensional space of the world space corresponds one-to-one to absolute geographical coordinates such as GPS, latitude and longitude. Alternatively, the three-dimensional space can be represented by a relative position relative to a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, etc., and the direction vectors are stored together with the encoded data as meta information.
[0156] Furthermore, the size of the GOS is set to be fixed, and the encoding device stores the size as meta-information. Furthermore, the size of the GOS can be switched, for example, depending on whether it is in the city, indoors, or outdoors. That is, the size of the GOS can be switched according to the amount or nature of objects that have value as information. Alternatively, the encoding device can appropriately switch the size of the GOS or the interval of the I-SPC in the GOS according to the density of the object, etc. in the same world space. For example, the encoding device sets the size of the GOS to be smaller and the interval of the I-SPC in the GOS to be shorter when the density of the object is higher.
[0157] exist Figure 5 In the example, in the area from the 3rd to the 10th GOS, since the density of objects is high, the GOS is subdivided to achieve random access of fine granularity. And, the 7th to the 10th GOS exist on the back of the 3rd to the 6th GOS, respectively.
[0158] Next, the configuration and operation flow of the three-dimensional data encoding device according to this embodiment will be described. Figure 6 It is a block diagram of the three-dimensional data encoding device 100 involved in this embodiment. Figure 7 : is a flowchart showing an operation example of the three-dimensional data encoding device 100 .
[0159] Figure 6 The three-dimensional data encoding device 100 shown generates encoded three-dimensional data 112 by encoding three-dimensional data 111. The three-dimensional data encoding device 100 includes an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.
[0160] like Figure 7 As shown, first, the acquisition unit 101 acquires three-dimensional data 111 as point group data ( S101 ).
[0161] Next, the coding region determination unit 102 determines a coding target region from the spatial region corresponding to the obtained point cloud data (S102). For example, the coding region determination unit 102 determines a spatial region around the position as the coding target region according to the position of the user or vehicle.
[0162] Next, the segmentation unit 103 segments the point group data included in the area of the encoding object into respective processing units. Here, the processing units are the above-mentioned GOS and SPC, etc. And, the area of the encoding object corresponds to the above-mentioned world space, for example. Specifically, the segmentation unit 103 segments the point group data into processing units according to the size of the pre-set GOS, the presence or size of the dynamic object (S103). And, the segmentation unit 103 determines the starting position of the SPC that becomes the beginning in the encoding order in each GOS.
[0163] Next, the encoding unit 104 generates the encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs in each GOS ( S104 ).
[0164] In addition, here, after the area to be coded is divided into GOS and SPC, although an example of coding each GOS is shown, the order of processing is not limited to the above. For example, after the composition of a GOS is determined, the GOS may be coded, and then the composition of the GOS may be determined.
[0165] In this way, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into random access units, that is, divides it into first processing units (GOS) corresponding to three-dimensional coordinates, divides the first processing unit (GOS) into a plurality of second processing units (SPC), and divides the second processing unit (SPC) into a plurality of third processing units (VLM). In addition, the third processing unit (VLM) includes more than one voxel (VXL), and the voxel (VXL) is the smallest unit corresponding to the position information.
[0166] Next, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Furthermore, the three-dimensional data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).
[0167] For example, when the first processing unit (GOS) of the processing object is a closed GOS, the three-dimensional data encoding device 100 performs encoding with reference to other second processing units (SPCs) included in the first processing unit (GOS) of the processing object for the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object. That is, the three-dimensional data encoding device 100 does not refer to the second processing unit (SPC) included in the first processing unit (GOS) different from the first processing unit (GOS) of the processing object.
[0168] Furthermore, when the first processing unit (GOS) of the processing object is an open GOS, the second processing unit (SPC) of the processing object contained in the first processing unit (GOS) of the processing object is encoded with reference to other second processing units (SPCs) contained in the first processing unit (GOS) of the processing object, or a second processing unit (SPC) contained in a first processing unit (GOS) different from the first processing unit (GOS) of the processing object.
[0169] Furthermore, the three-dimensional data encoding device 100 selects one as the type of the second processing unit (SPC) of the processing object from among the first type (I-SPC) that does not refer to other second processing units (SPC), the second type (P-SPC) that refers to one other second processing unit (SPC), and the third type that refers to two other second processing units (SPC), and encodes the second processing unit (SPC) of the processing object according to the selected type.
[0170] Next, the configuration and operation flow of the three-dimensional data decoding device according to the present embodiment will be described. Figure 8 It is a block diagram of the three-dimensional data decoding device 200 according to this embodiment. Fig. 9 : is a flowchart showing an example of the operation of the three-dimensional data decoding device 200 .
[0171] Figure 8 The three-dimensional data decoding device 200 shown generates decoded three-dimensional data 212 by decoding the encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. The three-dimensional data decoding device 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.
[0172] First, the acquisition unit 201 acquires the encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to the meta information stored in the encoded three-dimensional data 211 or separately from the encoded three-dimensional data, and determines the GOS including the spatial position, object, or SPC corresponding to the time at which decoding is started as the GOS to be decoded.
[0173] Next, the decoded SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded in the GOS (S203). For example, the decoded SPC determination unit 203 determines (1) whether to decode only the I-SPC, (2) whether to decode the I-SPC and the P-SPC, or (3) whether to decode all types. In addition, when the type of the SPC to be decoded is predetermined, such as when all SPCs are decoded, this step may not be performed.
[0174] Next, the decoding unit 204 obtains the SPC that is the first in the decoding order (the same as the encoding order) in the GOS, and the address position that starts in the encoded three-dimensional data 211, obtains the encoded data of the first SPC from the address position, and decodes each SPC in sequence from the first SPC (S204). And, the above-mentioned address position is stored in meta information, etc.
[0175] In this way, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates the decoded three-dimensional data 212 of the first processing unit (GOS) as a random access unit by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS) corresponding to the three-dimensional coordinates. More specifically, the three-dimensional data decoding device 200 decodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). And, the three-dimensional data decoding device 200 decodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).
[0176] The meta information for random access is described below. This meta information is generated by the three-dimensional data encoding device 100 and included in the encoded three-dimensional data 112 (211).
[0177] In conventional random access of two-dimensional moving images, decoding is started from the head frame of a random access unit near a designated time. However, in the world space, random access to (coordinates or objects, etc.) is also envisioned in addition to time.
[0178] Therefore, in order to realize random access to at least three elements, namely coordinates, objects, and time, a table is prepared in which each element is associated with the index number of the GOS. Furthermore, the index number of the GOS is associated with the address of the I-SPC at the beginning of the GOS. Fig.10 An example of a table included in the meta information is shown. Fig.10 Of all the tables shown, at least one table may be used.
[0179] As an example, random access starting from a coordinate is described below. When accessing the coordinates (x2, y2, z2), the coordinate-GOS table is first referenced, and it can be known that the location with the coordinates (x2, y2, z2) is included in the second GOS. Next, the GOS address table is referenced, and since it can be known that the address of the first I-SPC in the second GOS is addr(2), the decoding unit 204 obtains data from this address and starts decoding.
[0180] In addition, the address may be an address in a logical format or a physical address of a HDD or a memory. Furthermore, information for identifying a file segment may be used instead of an address. For example, a file segment is a unit after segmenting one or more GOSs.
[0181] Furthermore, when the object spans across multiple GOSs, the GOSs to which the multiple objects belong may also be shown in the object GOS table. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. In addition, if the multiple GOSs are open GOSs, the compression efficiency can be further improved by having the multiple GOSs refer to each other.
[0182] Examples of objects include people, animals, cars, bicycles, traffic lights, or buildings that serve as landmarks on land, etc. For example, when encoding in world space, the three-dimensional data encoding device 100 extracts feature points unique to the object from a three-dimensional point cloud, etc., detects the object based on the feature points, and can set the detected object as a random access point.
[0183] In this way, the three-dimensional data encoding device 100 generates the first information, which indicates the plurality of first processing units (GOS) and the three-dimensional coordinates corresponding to each of the plurality of first processing units (GOS). And the encoded three-dimensional data 112 (211) includes the first information. And the first information further indicates at least one of the object, time, and data storage destination corresponding to each of the plurality of first processing units (GOS).
[0184] The three-dimensional data decoding device 200 obtains the first information from the encoded three-dimensional data 211 , uses the first information to determine the first processing unit of the encoded three-dimensional data 211 corresponding to the specified three-dimensional coordinates, object or time, and decodes the encoded three-dimensional data 211 .
[0185] Other examples of meta information are described below. In addition to the meta information for random access, the three-dimensional data encoding device 100 can also generate and store the following meta information. Furthermore, the three-dimensional data decoding device 200 can also use this meta information during decoding.
[0186] In the case where three-dimensional data is used as map information, a profile is specified according to the purpose, and information indicating the profile may be included in the meta-information. For example, a profile for urban areas or suburbs is specified, or a profile for flying objects is specified, and the maximum or minimum size of the world space, SPC or VLM is defined respectively. For example, in the profile for urban areas, more detailed information is required than in the suburbs, so the minimum size of VLM is set smaller.
[0187] The meta information may also include a tag value indicating the type of the object. The tag value corresponds to the VLM, SPC, or GOS that constitutes the object. The tag value may be set according to the type of the object, for example, a tag value of "0" indicates a "person", a tag value of "1" indicates a "car", and a tag value of "2" indicates a "traffic light". Alternatively, when the type of the object is difficult to determine or does not need to be determined, a tag value indicating the size, or whether it is a dynamic object or a static object, etc. may be used.
[0188] Furthermore, the meta-information may include information indicating the range of the spatial region occupied by the world space.
[0189] Furthermore, the meta-information may be used as header information common to the entire stream of coded data or to a plurality of SPCs such as an SPC in a GOS, and may store the size of the SPC or VXL.
[0190] Furthermore, the meta-information may include identification information of a distance sensor or a camera used in generating the point cloud, or information indicating the positional accuracy of a point group in the point cloud.
[0191] Also, the meta information may include information showing whether the world space is composed of only static objects or contains dynamic objects.
[0192] Modifications of this embodiment will be described below.
[0193] The encoding device or decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on meta information indicating the spatial position of the GOSs.
[0194] In the case where three-dimensional data is used as a spatial map for the movement of a vehicle or a flying object, or such a spatial map is generated, the encoding device or decoding device can encode or decode the GOS or SPC contained in the space determined based on GPS, path information, or a zoom factor, etc.
[0195] Furthermore, the decoding device may also start decoding from the space close to its own position or walking path. The encoding device or decoding device may also make the priority of the space far from its own position or walking path lower than the priority of the space close to it to perform encoding or decoding. Here, lowering the priority means lowering the processing order, lowering the resolution (screening post-processing), or lowering the image quality (improving the encoding efficiency. For example, increasing the quantization step size), etc.
[0196] Furthermore, when decoding the coded data which is hierarchically coded in space, the decoding device may decode only the lower layer.
[0197] Furthermore, the decoding device may start decoding from a lower layer according to the zoom ratio or purpose of the map.
[0198] Furthermore, in applications such as estimating the self-position of a car or robot during automatic driving or identifying an object, the encoding device or decoding device may also reduce the resolution of an area outside the area within a specified height from the road surface (the area to be identified) for encoding or decoding.
[0199] Furthermore, the encoding device may also encode the point clouds representing the indoor and outdoor spatial shapes separately, for example, by separating the GOS representing the indoor space (indoor GOS) from the GOS representing the outdoor space (outdoor GOS), so that when the decoding device uses the encoded data, it can select the GOS to be decoded according to the viewpoint position.
[0200] Furthermore, the encoding device may encode the indoor GOS and the outdoor GOS whose coordinates are close to each other by placing them adjacent to each other in the coded stream. For example, the encoding device may correspond the identifiers of the two and store information indicating that the corresponding identifiers are established in the coded stream or in meta-information stored separately. Accordingly, the decoding device can refer to the information in the meta-information to identify the indoor GOS and the outdoor GOS whose coordinates are close to each other.
[0201] Furthermore, the encoding device may switch the size of the GOS or SPC between the indoor GOS and the outdoor GOS. For example, the encoding device sets the size of the GOS to be smaller indoors than outdoors. Furthermore, the encoding device may change the accuracy of extracting feature points from the point cloud or the accuracy of object detection between the indoor GOS and the outdoor GOS.
[0202] Furthermore, the encoding device may add information for the decoding device to distinguish and display dynamic objects from static objects to the encoded data. Accordingly, the decoding device can represent the dynamic objects by combining them with red frames or explanatory texts. In addition, the decoding device may replace the dynamic objects and represent them with only red frames or explanatory texts. Furthermore, the decoding device may represent more detailed object categories. For example, a car may be represented by a red frame, and a person may be represented by a yellow frame.
[0203] Furthermore, the encoding device or decoding device may determine whether to perform encoding or decoding by treating the dynamic object and the static object as different SPCs or GOSs according to the frequency of occurrence of the dynamic object or the ratio of the static object to the dynamic object. For example, when the frequency of occurrence or the ratio of the dynamic object exceeds a threshold, the SPC or GOS in which the dynamic object and the static object are mixed is allowed, and when the frequency of occurrence or the ratio of the dynamic object does not exceed the threshold, the SPC or GOS in which the dynamic object and the static object are mixed is not allowed.
[0204] When the dynamic object is detected from the two-dimensional image information of the camera instead of the point cloud, the encoding device can obtain the information (frame or text, etc.) for identifying the detection result and the object position separately, and encode the information as part of the three-dimensional encoded data. In this case, the decoding device overlaps the auxiliary information (frame or text) representing the dynamic object with the decoding result of the static object.
[0205] Furthermore, the coding device may change the density of VXL or VLM according to the complexity of the shape of the static object, etc. For example, the coding device sets VXL or VLM to be denser when the shape of the static object is more complex. Furthermore, the coding device may determine the quantization step size when quantizing the spatial position or color information according to the density of VXL or VLM. For example, the coding device sets the quantization step size to be smaller when the VXL or VLM is denser.
[0206] As described above, the encoding device or decoding device according to the present embodiment performs spatial encoding or decoding in a spatial unit having coordinate information.
[0207] Furthermore, the encoding device and the decoding device perform encoding or decoding in volume units in space. The volume includes voxels which are the minimum units corresponding to the position information.
[0208] Furthermore, the encoding device and the decoding device establish a table of correspondence between each element of spatial information including coordinates, objects, and time and GOP, or a table of correspondence between each element, so as to establish correspondence between arbitrary elements to perform encoding or decoding. Furthermore, the decoding device determines the coordinates using the value of the selected element, determines the volume, voxel, or space based on the coordinates, and decodes the space including the volume or voxel, or the determined space.
[0209] Then, the encoding device determines a volume, voxel, or space that can be selected by an element by extracting feature points or recognizing an object, and encodes it as a volume, voxel, or space that can be randomly accessed.
[0210] Spaces are divided into three types: I-SPC which can be encoded or decoded as a single space, P-SPC which can be encoded or decoded with reference to any one of the processed spaces, and B-SPC which can be encoded or decoded with reference to any two of the processed spaces.
[0211] One or more volumes correspond to static objects or dynamic objects. Spaces containing static objects and spaces containing dynamic objects are encoded or decoded as different GOSs. That is, SPCs containing static objects and SPCs containing dynamic objects are assigned to different GOSs.
[0212] Dynamic objects are encoded or decoded for each object and correspond to one or more spaces containing only static objects. That is, multiple dynamic objects are encoded separately, and the encoded data of multiple dynamic objects obtained correspond to SPCs containing only static objects.
[0213] The encoding device and the decoding device increase the priority of the I-SPC in the GOS to perform encoding or decoding. For example, the encoding device performs encoding in a manner that reduces the degradation of the I-SPC (after decoding, the original three-dimensional data can be reproduced more faithfully). And, the decoding device, for example, only decodes the I-SPC.
[0214] The encoding device may change the frequency of using the I-SPC according to the density or value (number) of objects in the world space to perform encoding. That is, the encoding device changes the frequency of selecting the I-SPC according to the number or density of objects included in the three-dimensional data. For example, the encoding device increases the frequency of using the I space when the density of objects in the world space is greater.
[0215] Furthermore, the encoding device sets the random access point in units of GOS, and stores information indicating the spatial area corresponding to the GOS in the header information.
[0216] The coding device, for example, adopts a default value as the spatial size of the GOS. In addition, the coding device may also change the size of the GOS according to the value (number) or density of the objects or dynamic objects. For example, the coding device sets the spatial size of the GOS to be smaller when the objects or dynamic objects are denser or more in number.
[0217] Furthermore, the space or volume includes a group of feature points derived from information obtained by sensors such as a depth sensor, a gyroscope, or a camera. The coordinates of the feature points are set to the center position of the voxel. Furthermore, by subdividing the voxels, the position information can be highly accurate.
[0218] The feature point group is derived using a plurality of pictures. The plurality of pictures have at least the following two types of time information: actual time information and the same time information in the plurality of pictures corresponding to the space (for example, encoding time for rate control, etc.).
[0219] Furthermore, encoding or decoding is performed in units of GOS including one or more spaces.
[0220] The encoding device and the decoding device refer to the space in the processed GOS to predict the P space or the B space in the GOS to be processed.
[0221] Alternatively, the encoding device and the decoding device predict the P space or B space in the GOS to be processed by using the processed space in the GOS to be processed without referring to different GOSs.
[0222] Furthermore, the encoding device and the decoding device transmit or receive the encoded stream in units of a world space including one or more GOSs.
[0223] Furthermore, the GOS has a layer structure in at least one direction in the world space, and the encoding device and the decoding device perform encoding or decoding from the lower layer. For example, the GOS that can be randomly accessed belongs to the lowest layer. The GOS belonging to the upper layer only refers to the GOS belonging to the layer below the same layer. That is, the GOS is spatially divided in a predetermined direction and includes a plurality of layers each having one or more SPCs. The encoding device and the decoding device perform encoding or decoding for each SPC by referring to the SPC included in the layer that is the same layer as the SPC or the layer below the SPC.
[0224] Furthermore, the encoding device and the decoding device continuously encode or decode the GOS in the world space unit including the plurality of GOS. The encoding device and the decoding device write or read information indicating the order (direction) of encoding or decoding as metadata. That is, the encoded data includes information indicating the encoding order of the plurality of GOS.
[0225] Furthermore, the encoding device and the decoding device encode or decode two or more different spaces or GOS in parallel.
[0226] Furthermore, the encoding device and the decoding device encode or decode the space information (coordinates, size, etc.) of the space or the GOS.
[0227] Furthermore, the encoding device and the decoding device encode or decode the space or GOS included in the specific space determined based on external information related to the own position and / or area size, such as GPS, path information, or magnification.
[0228] The encoding device or decoding device performs encoding or decoding by giving a lower priority to a space far from the own position than to a space close to the own position.
[0229] The encoding device sets a direction in the world space according to the magnification or purpose, and encodes the GOS having a layer structure in the direction. And the decoding device preferentially decodes the GOS having a layer structure in the direction of the world space set according to the magnification or purpose from the lower layer.
[0230] The encoding device changes the feature point extraction, object recognition accuracy, or space area size contained in the indoor and outdoor spaces. However, the encoding device and the decoding device encode or decode the indoor GOS and the outdoor GOS with close coordinates adjacent to each other in the world space, and also encode or decode these identifiers in correspondence.
[0231] (Implementation Method 2)
[0232] When point cloud coded data is used in actual devices or services, it is desirable to transmit and receive required information according to the application in order to reduce network bandwidth. However, such a function does not exist in the existing 3D data coding structure, and therefore there is no corresponding coding method.
[0233] What will be described in this embodiment is a three-dimensional data encoding method and a three-dimensional data encoding device for providing the function of sending and receiving required information according to the purpose in the encoded data of a three-dimensional point cloud, as well as a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data.
[0234] A voxel (VXL) having a certain feature value or more is defined as a feature voxel (FVXL), and a world space (WLD) formed by the FVXL is defined as a sparse world space (SWLD). Fig.11The sparse world space and the configuration example of the world space are shown. SWLD includes: FGOS, which is a GOS composed of FVXL; FSPC, which is an SPC composed of FVXL; and FVLM, which is a VLM composed of FVXL. The data structure and prediction structure of FGOS, FSPC and FVLM can be the same as those of GOS, SPC and VLM.
[0235] The feature quantity refers to a feature quantity that expresses the three-dimensional position information of VXL or the visible light information of the position of VXL, and in particular, a feature quantity that can detect more corners and edges of a three-dimensional object. Specifically, the feature quantity is a three-dimensional feature quantity or a visible light feature quantity described below, but other than this, any feature quantity can be used as long as it expresses the position, brightness, or color information of VXL.
[0236] As the three-dimensional feature, a SHOT feature (Signature of Histograms of Orientations), a PFH feature (Point Feature Histograms), or a PPF feature (Point Pair Feature) is used.
[0237] The SHOT feature is obtained by segmenting the area around the VXL, calculating the inner product between the reference point and the normal vector of the segmented area, and converting it into a histogram. The SHOT feature has a high dimension and high feature expression.
[0238] The PFH feature is obtained by selecting a plurality of two-point groups near VXL, calculating the normal vector etc. from the two points, and histogramming them. Since the PFH feature is a histogram feature, it is robust to small amounts of interference and has a high feature expression.
[0239] The PPF feature is a feature calculated according to the VXL of two points using a normal vector, etc. Since all VXLs are used in this PPF feature, it is robust to occlusion.
[0240] Furthermore, as feature quantities of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients) that utilize information such as brightness gradient information of an image can be used.
[0241] SWLD is generated by calculating the above-mentioned feature amount from each VXL of WLD and extracting FVXL. Here, SWLD may be updated every time WLD is updated, or may be updated regularly after a certain period of time has passed regardless of the update timing of WLD.
[0242] SWLD can be generated for each feature quantity. For example, as shown in SWLD1 based on SHOT feature quantity and SWLD2 based on SIFT feature quantity, SWLD can be generated for each feature quantity and used according to the purpose. In addition, the feature quantity of each FVXL calculated can be stored as feature quantity information in each FVXL.
[0243] Next, a method of using the sparse world space (SWLD) will be described. Since the SWLD only includes feature voxels (FVXL), the data size is generally smaller than that of the WLD including all VXLs.
[0244] In applications that use feature quantities to achieve a certain purpose, by using SWLD information instead of WLD, it is possible to reduce the time required to read from the hard disk, and to reduce the bandwidth and transmission time during network transmission. For example, as map information, WLD and SWLD are stored in the server in advance, and the map information to be sent is switched to WLD or SWLD according to the needs of the client, thereby reducing the network bandwidth and transmission time. A specific example is shown below.
[0245] Fig.12 as well as Fig.13 The following shows the use cases of SWLD and WLD. Fig.12 As shown, when the client 1 as a vehicle-mounted device needs map information for the purpose of determining its own position, the client 1 sends a request for obtaining map data for estimating its own position to the server (S301). The server sends SWLD to the client 1 according to the acquisition request (S302). The client 1 uses the received SWLD to determine its own position (S303). At this time, the client 1 obtains the VXL information of the surrounding area of the client 1 through various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras, and estimates its own position information based on the obtained VXL information and SWLD. Here, the own position information includes the three-dimensional position information and orientation of the client 1.
[0246] like Fig.13As shown, when the client 2 as a vehicle-mounted device needs map information for the purpose of drawing a three-dimensional map or the like, the client 2 sends a request for obtaining map data for drawing a map to the server (S311). The server sends a WLD to the client 2 in accordance with the acquisition request (S312). The client 2 uses the received WLD to draw a map (S313). At this time, the client 2 uses, for example, an image taken by itself with a visible light camera and the WLD obtained from the server to create a conceptual image, and draws the created image on a screen such as a car navigation system.
[0247] As described above, the server sends SWLD to the client when the feature amount of each VXL is mainly needed, such as estimating the own position, and sends WLD to the client when detailed VXL information is needed, such as drawing a map. This makes it possible to efficiently send and receive map data.
[0248] In addition, the client can determine which one of SWLD and WLD it needs and request the server to send SWLD or WLD. And the server can determine which one of SWLD and WLD should be sent according to the status of the client or the network.
[0249] Next, a method of switching the transmission and reception between the sparse world space (SWLD) and the world space (WLD) will be described.
[0250] The reception of WLD or SWLD can be switched according to the network bandwidth. Fig.14 An example of operation in this case is shown. For example, when a low-speed network that can use the network bandwidth in an LTE (Long Term Evolution) environment is used, when the client accesses the server via the low-speed network (S321), the SWLD as map information is obtained from the server (S322). In addition, when a high-speed network with sufficient network bandwidth is used in a WiFi environment, the client accesses the server via the high-speed network (S323) and obtains the WLD from the server (S324). Accordingly, the client can obtain appropriate map information according to the network bandwidth of the client.
[0251] Specifically, the client receives SWLD via LTE outdoors, and acquires WLD via WiFi when entering a facility or the like indoors. This allows the client to acquire more detailed map information indoors.
[0252] In this way, the client can request WLD or SWLD from the server according to the frequency band of the network it uses. Alternatively, the client can send information showing the frequency band of the network it uses to the server, and the server can send appropriate data (WLD or SWLD) to the client according to the information. Alternatively, the server can determine the network bandwidth of the client and send appropriate data (WLD or SWLD) to the client.
[0253] Furthermore, the reception of WLD or SWLD can be switched according to the moving speed. Fig.15 An example of operation in this case is shown. For example, when the client moves at high speed (S331), the client receives SWLD from the server (S332). In addition, when the client moves at low speed (S333), the client receives WLD from the server (S334). Accordingly, the client can suppress the network bandwidth and obtain map information according to the speed. Specifically, when the client is driving on a highway, by receiving SWLD with a small amount of data, the map information can be updated at a roughly appropriate speed. In addition, when the client is driving on a general road, by receiving WLD, more detailed map information can be obtained.
[0254] In this way, the client can request WLD or SWLD from the server according to its own moving speed. Alternatively, the client can send information showing its own moving speed to the server, and the server can send appropriate data (WLD or SWLD) to the client according to the information. Alternatively, the server can determine the moving speed of the client and send appropriate data (WLD or SWLD) to the client.
[0255] Furthermore, the client may first obtain the SWLD from the server, and then obtain the WLD of the important area. For example, when the client obtains map data, it first obtains the general map information with the SWLD, selects the area where the features such as buildings, signs, or people appear more, and then obtains the WLD of the selected area. In this way, the client can suppress the amount of data received from the server and obtain the detailed information of the required area.
[0256] Furthermore, the server may create SWLDs for each object based on the WLD, and the client may receive them separately according to the purpose. In this way, the network bandwidth can be suppressed. For example, the server pre-identifies a person or a car from the WLD, and creates a SWLD for the person and a SWLD for the car. When the client wants to obtain information about the people around it, it receives the SWLD for the person, and when it wants to obtain information about the car, it receives the SWLD for the car. Furthermore, the type of such SWLD can be distinguished based on information (flag or type, etc.) attached to the header, etc.
[0257] Next, the configuration and operation flow of the three-dimensional data encoding device (for example, a server) according to the present embodiment will be described. Fig.16 It is a block diagram of the three-dimensional data encoding device 400 involved in this embodiment. Fig.17 This is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device 400.
[0258] Fig.16 The three-dimensional data encoding device 400 shown in the figure generates encoded three-dimensional data 413 and 414 as encoded streams by encoding input three-dimensional data 411. Here, the encoded three-dimensional data 413 is encoded three-dimensional data corresponding to WLD, and the encoded three-dimensional data 414 is encoded three-dimensional data corresponding to SWLD. The three-dimensional data encoding device 400 includes: an acquisition unit 401, a coding area determination unit 402, a SWLD extraction unit 403, a WLD encoding unit 404, and a SWLD encoding unit 405.
[0259] like Fig.17 As shown, first, the acquisition unit 401 acquires input three-dimensional data 411 which is point group data in a three-dimensional space (S401).
[0260] Next, the encoding region determination unit 402 determines the encoding target spatial region based on the spatial region where the point cloud data exists ( S402 ).
[0261] Next, the SWLD extraction unit 403 defines the spatial region of the encoding object as WLD, and calculates the feature value from each VXL included in the WLD. In addition, the SWLD extraction unit 403 extracts the VXL whose feature value is greater than a predetermined threshold value, defines the extracted VXL as FVXL, and adds the FVXL to the SWLD to generate the extracted three-dimensional data 412 (S403). That is, the extracted three-dimensional data 412 whose feature value is greater than the threshold value is extracted from the input three-dimensional data 411.
[0262] Next, the WLD encoding unit 404 generates encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 adds information for distinguishing that the encoded three-dimensional data 413 is a stream containing the WLD to the header of the encoded three-dimensional data 413.
[0263] Then, the SWLD encoding unit 405 generates encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information for distinguishing that the encoded three-dimensional data 414 is a stream containing the SWLD to the header of the encoded three-dimensional data 414.
[0264] Furthermore, the processing order of the process of generating the encoded three-dimensional data 413 and the process of generating the encoded three-dimensional data 414 may be reversed from the above. Furthermore, part or all of the above-mentioned processes may be executed in parallel.
[0265] Information assigned to the header of the encoded three-dimensional data 413 and 414 is defined as a parameter such as "world_type". When world_type = 0, it means that the stream contains WLD, and when world_type = 1, it means that the stream contains SWLD. In the case of defining more other categories, the assigned value can be increased, such as world_type = 2. In addition, a specific flag can be included in one of the encoded three-dimensional data 413 and 414. For example, the encoded three-dimensional data 414 can be assigned a flag indicating that the stream contains SWLD. In this case, the decoding device can determine whether it is a stream containing WLD or a stream containing SWLD based on the presence or absence of the flag.
[0266] Furthermore, the encoding method used by the WLD encoding unit 404 when encoding the WLD may be different from the encoding method used by the SWLD encoding unit 405 when encoding the SWLD.
[0267] For example, since data is decimated in SWLD, the correlation with surrounding data may be lower than that in WLD. Therefore, in the encoding method for SWLD, inter-frame prediction is prioritized over intra-frame prediction and inter-frame prediction in comparison with the encoding method for WLD.
[0268] Furthermore, the encoding method used for SWLD and the encoding method used for WLD may have different methods of expressing the three-dimensional position. For example, the three-dimensional position of FVXL may be expressed by three-dimensional coordinates in FWLD, and the three-dimensional position may be expressed by an octree described later in WLD, or vice versa.
[0269] Furthermore, the SWLD coding unit 405 performs coding in such a manner that the data size of the coded three-dimensional data 414 of SWLD is smaller than the data size of the coded three-dimensional data 413 of WLD. For example, as described above, the correlation between data in SWLD may be reduced compared to that in WLD. Accordingly, the coding efficiency is reduced, and the data size of the coded three-dimensional data 414 may be larger than the data size of the coded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained coded three-dimensional data 414 is larger than the data size of the coded three-dimensional data 413 of WLD, the SWLD coding unit 405 performs re-coding to generate the coded three-dimensional data 414 with a reduced data size again.
[0270] For example, the SWLD extraction unit 403 generates the extracted three-dimensional data 412 again with the number of extracted feature points reduced, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be made coarser. For example, in the octree structure described later, the degree of quantization may be made coarser by rounding the data of the lowest layer.
[0271] Furthermore, when the data size of the SWLD encoded three-dimensional data 414 cannot be made smaller than the data size of the WLD encoded three-dimensional data 413, the SWLD encoding unit 405 may not generate the SWLD encoded three-dimensional data 414. Alternatively, the WLD encoded three-dimensional data 413 may be copied to the SWLD encoded three-dimensional data 414. That is, the WLD encoded three-dimensional data 413 may be used directly as the SWLD encoded three-dimensional data 414.
[0272] Next, the configuration and operation flow of the three-dimensional data decoding device (eg, client) according to the present embodiment will be described. Fig.18 It is a block diagram of a three-dimensional data decoding device 500 according to this embodiment. Fig.19 3D data decoding processing performed by the 3D data decoding apparatus 500 is shown in FIG.
[0273] Fig.18 The three-dimensional data decoding device 500 shown generates decoded three-dimensional data 512 or 513 by decoding the encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0274] The three-dimensional data decoding device 500 includes an acquisition unit 501 , a header analysis unit 502 , a WLD decoding unit 503 , and a SWLD decoding unit 504 .
[0275] like Fig.19 As shown, first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Then, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 to determine whether the encoded three-dimensional data 511 is a stream containing WLD or a stream containing SWLD (S502). For example, the determination is made by referring to the above-mentioned world_type parameter.
[0276] When the coded three-dimensional data 511 is a stream including WLD ("Yes" in S503), the WLD decoding unit 503 decodes the coded three-dimensional data 511 to generate decoded three-dimensional data 512 of WLD (S504). In addition, when the coded three-dimensional data 511 is a stream including SWLD ("No" in S503), the SWLD decoding unit 504 decodes the coded three-dimensional data 511 to generate decoded three-dimensional data 513 of SWLD (S505).
[0277] Also, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding the WLD may be different from the decoding method used by the SWLD decoding unit 504 when decoding the SWLD. For example, in the decoding method for SWLD, priority may be given to the inter-frame prediction of intra-frame prediction and inter-frame prediction over the decoding method for WLD.
[0278] Furthermore, the decoding method for SWLD and the decoding method for WLD may have different representation methods for the three-dimensional position. For example, the three-dimensional position of FVXL may be represented by three-dimensional coordinates in SWLD and by an octree described later in WLD, or vice versa.
[0279] Next, octree expression as a method of expressing three-dimensional positions will be described. VXL data included in three-dimensional data is converted into an octree structure and encoded. Fig. 20 An example of a VXL of a WLD is shown. Fig.21 Shows Fig. 20 The octree structure of WLD is shown in Figure 1. Fig. 20 In the example shown, there are three VXLs 1 to 3 that are VXLs (hereinafter, effective VXLs) that include point groups. Fig.21 As shown, the octree structure consists of nodes and leaf nodes. Each node has a maximum of 8 nodes or leaf nodes. Each leaf node has VXL information. Here, Fig.21 Among the leaf nodes shown, leaf nodes 1, 2, and 3 represent Fig. 20 VXL1, VXL2, VXL3 shown.
[0280] Specifically, each node and leaf node corresponds to a three-dimensional position. Fig. 20 The block corresponding to node 1 is divided into 8 blocks, and among the 8 blocks, the block including the valid VXL is set as a node, and the other blocks are set as leaf nodes. The block corresponding to the node is further divided into 8 nodes or leaf nodes, and this process is repeated the same number of times as the number of levels in the tree structure. And all the blocks at the bottom are set as leaf nodes.
[0281] and, Fig. 22 Shows from Fig. 20 An example of SWLD generated by WLD is shown. Fig. 20 The feature extraction results of VXL1 and VXL2 shown are determined to be FVXL1 and FVXL2 and are included in SWLD. In addition, VXL3 is not determined to be FVXL and is therefore not included in SWLD. Fig.23 Shows Fig. 22 The octree structure of SWLD is shown in Figure 1. Fig.23 In the octree structure shown, Fig.21 The leaf node 3 shown, which is equivalent to VXL3, is deleted. Fig.21 The node 3 shown has no valid VXL and is changed to a leaf node. In this way, generally speaking, the number of leaf nodes of SWLD is smaller than that of WLD, and the encoded three-dimensional data of SWLD is also smaller than that of WLD.
[0282] Modifications of this embodiment will be described below.
[0283] For example, when a client such as a vehicle-mounted device estimates its own position, it receives SWLD from a server, uses SWLD to estimate its own position, and performs obstacle detection. It then uses various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple monocular cameras to perform obstacle detection based on the three-dimensional information of the surrounding area obtained by itself.
[0284] In general, it is difficult to include VXL data of a flat area in SWLD. For this reason, the server maintains a subsampled world space (SubWLD) that is a downsampled version of WLD for detecting stationary obstacles, and can send SWLD and SubWLD to the client. This can suppress network bandwidth while enabling the client to estimate its own position and detect obstacles.
[0285] Furthermore, when the client quickly depicts three-dimensional map data, it is convenient if the map information is in a grid structure. Therefore, the server can generate a grid based on the WLD and store it in advance as a grid world space (MWLD). For example, when the client needs to perform a rough three-dimensional depiction, the MWLD is received, and when a detailed three-dimensional depiction is required, the WLD is received. In this way, the network bandwidth can be suppressed.
[0286] Furthermore, although the server sets the VXL with a feature value above the threshold value as FVXL from each VXL, FVXL can also be calculated by different methods. For example, if the server determines that VXL, VLM, SPC, or GOS constituting a signal or intersection is required for self-position estimation, driving assistance, or automatic driving, it can be included in SWLD as FVXL, FVLM, FSPC, FGOS. Furthermore, the above judgment can be performed manually. In addition, FVXL obtained by the above method can be added to FVXL set based on the feature value. That is, the SWLD extraction unit 403 can further extract data corresponding to an object with predetermined attributes from the input three-dimensional data 411 as extracted three-dimensional data 412.
[0287] Furthermore, labels different from feature quantities may be assigned to situations that require these purposes. The server may separately maintain FVXL required for self-position estimation such as signals or intersections, driving assistance, or autonomous driving as an upper layer of SWLD (e.g., lane world space).
[0288] Furthermore, the server may also attach attributes to the VXL in the WLD in random access units or specified units. Attributes include, for example, information indicating whether the location is required or not required for estimating the location, or information indicating whether traffic information such as signals or intersections is important. Furthermore, attributes may also include the correspondence between lane information (GDF: Geographic Data Files, etc.) and features (intersections or roads, etc.).
[0289] Furthermore, as a method of updating WLD or SWLD, the following method can be adopted.
[0290] Update information showing changes in people, construction, or street trees (trajectory orientation) is uploaded to the server as a point group or metadata. The server updates the WLD based on the upload, and then updates the SWLD using the updated WLD.
[0291] Furthermore, when the client detects a mismatch between the three-dimensional information generated by itself and the three-dimensional information received from the server when estimating its own position, the client can send the three-dimensional information generated by itself to the server together with the update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is old.
[0292] Furthermore, although information for distinguishing between WLD and SWLD is added as the header information of the coded stream, when there are multiple world spaces such as a grid world space or a lane world space, information for distinguishing them may be added to the header information. Furthermore, when there are multiple SWLDs with different feature quantities, information for distinguishing them may also be added to the header information.
[0293] Furthermore, although SWLD is composed of FVXL, it may also include VXL that is not determined to be FVXL. For example, SWLD may include adjacent VXL used when calculating the feature quantity of FVXL. Accordingly, even if each FVXL of SWLD does not have feature quantity information attached, the client can calculate the feature quantity of FVXL when receiving SWLD. In addition, at this time, SWLD may include information for distinguishing whether each VXL is FVXL or VXL.
[0294] As described above, the three-dimensional data encoding device 400 extracts extracted three-dimensional data 412 (second three-dimensional data) whose feature value is above a threshold value from the input three-dimensional data 411 (first three-dimensional data), and generates encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.
[0295] According to this, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding the data whose feature quantity is greater than the threshold value. In this way, the amount of data can be reduced compared to the case where the input three-dimensional data 411 is directly encoded. Therefore, the three-dimensional data encoding device 400 can reduce the amount of data during transmission.
[0296] Furthermore, the three-dimensional data encoding device 400 further encodes the input three-dimensional data 411 to generate encoded three-dimensional data 413 (second encoded three-dimensional data).
[0297] According to this, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414 according to the purpose of use, for example.
[0298] Furthermore, the extracted three-dimensional data 412 is encoded by a first encoding method, and the input three-dimensional data 411 is encoded by a second encoding method different from the first encoding method.
[0299] Accordingly, the three-dimensional data encoding device 400 can adopt appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.
[0300] Furthermore, in the first encoding method, compared with the second encoding method, priority is given to the inter-frame prediction between the intra-frame prediction and the inter-frame prediction.
[0301] According to this, the three-dimensional data encoding device 400 can increase the priority of inter-frame prediction for the extracted three-dimensional data 412 where the correlation between adjacent data is likely to be low.
[0302] Furthermore, the first encoding method and the second encoding method have different methods of expressing the three-dimensional position. For example, in the second encoding method, the three-dimensional position is expressed by an octree, while in the first encoding method, the three-dimensional position is expressed by three-dimensional coordinates.
[0303] According to this, the three-dimensional data encoding device 400 can adopt a more appropriate three-dimensional position expression method for three-dimensional data having different data numbers (number of VXLs or FVXLs).
[0304] Furthermore, at least one of the coded three-dimensional data 413 and 414 includes an identifier indicating whether the coded three-dimensional data is coded three-dimensional data obtained by encoding the input three-dimensional data 411 or coded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, the identifier indicates whether the coded three-dimensional data is the coded three-dimensional data 413 of the WLD or the coded three-dimensional data 414 of the SWLD.
[0305] Based on this, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0306] Furthermore, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 in such a manner that the data amount of the encoded three-dimensional data 414 is smaller than the data amount of the encoded three-dimensional data 413 .
[0307] According to this, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 smaller than the data amount of the encoded three-dimensional data 413 .
[0308] Furthermore, the three-dimensional data encoding device 400 further extracts data corresponding to an object having a predetermined attribute from the input three-dimensional data 411 as extracted three-dimensional data 412. For example, the object having a predetermined attribute is an object required for self-position estimation, driving assistance, or automatic driving, such as a signal or an intersection.
[0309] Accordingly, the three-dimensional data encoding device 400 can generate the encoded three-dimensional data 414 including the data required by the decoding device.
[0310] Furthermore, the three-dimensional data encoding device 400 (server) further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the status of the client.
[0311] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the status of the client.
[0312] Furthermore, the status of the client includes the communication status of the client (eg, network bandwidth) or the moving speed of the client.
[0313] Furthermore, the three-dimensional data encoding device 400 further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the request of the client.
[0314] Thereby, the three-dimensional data encoding device 400 can send appropriate data according to the request of the client.
[0315] Furthermore, the three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400 .
[0316] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature amount extracted from the input three-dimensional data 411 is greater than the threshold value by the first decoding method. Furthermore, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 by using a second decoding method different from the first decoding method.
[0317] According to this, the three-dimensional data decoding device 500 can selectively receive the encoded three-dimensional data 414 and the encoded three-dimensional data 413 obtained by encoding data having a feature value greater than a threshold value, for example, according to the purpose of use. According to this, the three-dimensional data decoding device 500 can reduce the amount of data during transmission. Furthermore, the three-dimensional data decoding device 500 can adopt appropriate decoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.
[0318] Furthermore, in the first decoding method, compared with the second decoding method, priority is given to the inter prediction between the intra prediction and the inter prediction.
[0319] According to this, the three-dimensional data decoding apparatus 500 can increase the priority of inter-frame prediction for extracting three-dimensional data in which the correlation between adjacent data is likely to be low.
[0320] Furthermore, the first decoding method and the second decoding method use different methods for expressing the three-dimensional position. For example, the second decoding method expresses the three-dimensional position by an octree, while the first decoding method expresses the three-dimensional position by three-dimensional coordinates.
[0321] According to this, the three-dimensional data decoding apparatus 500 can adopt a more appropriate three-dimensional position expression method for three-dimensional data having different data numbers (number of VXLs or FVXLs).
[0322] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 by referring to the identifier.
[0323] Based on this, the three-dimensional data decoding device 500 can easily determine whether the obtained encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414 .
[0324] Furthermore, the 3D data decoding device 500 notifies the server of the state of the client (the 3D data decoding device 500). The 3D data decoding device 500 receives one of the encoded 3D data 413 and 414 transmitted from the server according to the state of the client.
[0325] Accordingly, the three-dimensional data decoding apparatus 500 can receive appropriate data according to the status of the client.
[0326] Furthermore, the status of the client includes the communication status of the client (eg, network bandwidth) or the moving speed of the client.
[0327] Furthermore, the three-dimensional data decoding apparatus 500 further requests one of the encoded three-dimensional data 413 and 414 from the server, and receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in accordance with the request.
[0328] Thereby, the three-dimensional data decoding device 500 can receive appropriate data according to the application.
[0329] (Implementation method 3)
[0330] In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles will be described. For example, the three-dimensional data is transmitted and received between the own vehicle and surrounding vehicles.
[0331] Fig.24 This is a block diagram of a three-dimensional data production device 620 according to this embodiment. The three-dimensional data production device 620 is included in the vehicle, for example, and produces denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 with the first three-dimensional data 632 produced by the three-dimensional data production device 620.
[0332] The three-dimensional data creation device 620 includes a three-dimensional data creation unit 621 , a request range determination unit 622 , a search unit 623 , a reception unit 624 , a decoding unit 625 , and a synthesis unit 626 .
[0333] First, the three-dimensional data creation unit 621 uses sensor information 631 detected by a sensor of the own vehicle to create first three-dimensional data 632. Next, the request range determination unit 622 determines a request range, which is a three-dimensional space range where the created first three-dimensional data 632 does not have enough data.
[0334] Next, the search unit 623 searches for surrounding vehicles that have three-dimensional data of the requested range, and sends request range information 633 showing the requested range to the surrounding vehicles determined by the search. Next, the receiving unit 624 receives the encoded three-dimensional data 634 (S624) as the encoded stream of the requested range from the surrounding vehicles. In addition, the search unit 623 can indiscriminately issue requests to all vehicles existing in the determined range, and receive the encoded three-dimensional data 634 from the responding party. In addition, the search unit 623 is not limited to vehicles, and can also issue requests to objects such as traffic lights or signs, and receive the encoded three-dimensional data 634 from the object.
[0335] Next, the decoder 625 decodes the received encoded three-dimensional data 634 to obtain second three-dimensional data 635. Next, the synthesizer 626 synthesizes the first three-dimensional data 632 and the second three-dimensional data 635 to create denser third three-dimensional data 636.
[0336] Next, the configuration and operation of the three-dimensional data transmitting device 640 according to this embodiment will be described. Fig.25 It is a block diagram of the three-dimensional data transmitting device 640.
[0337] The three-dimensional data sending device 640 is included in the above-mentioned surrounding vehicles, for example, and processes the fifth three-dimensional data 652 produced by the surrounding vehicles into the sixth three-dimensional data 654 requested by the own vehicle, and generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and sends the encoded three-dimensional data 634 to the own vehicle.
[0338] The three-dimensional data transmitting device 640 includes a three-dimensional data creating unit 641 , a receiving unit 642 , an extracting unit 643 , a coding unit 644 , and a transmitting unit 645 .
[0339] First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 using sensor information 651 detected by sensors provided in surrounding vehicles. Next, the receiving unit 642 receives the requested range information 633 transmitted from the own vehicle.
[0340] Next, the extraction unit 643 extracts the three-dimensional data of the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652, and processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654. Next, the encoding unit 644 encodes the sixth three-dimensional data 654, thereby generating the encoded three-dimensional data 634 as an encoded stream. Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to the own vehicle.
[0341] In addition, although an example is described here in which the own vehicle has the three-dimensional data creation device 620 and the surrounding vehicles have the three-dimensional data transmission device 640, each vehicle may also have the functions of the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.
[0342] (Implementation 4)
[0343] In this embodiment, the operation related to abnormal situations in the estimation of the own position based on the three-dimensional map will be described.
[0344] The use of autonomous movement of moving objects such as automatic driving of automobiles, robots, or flying objects such as drones will expand in the future. An example of a method for achieving such autonomous movement is a method in which a moving object estimates its position in a three-dimensional map (self-position estimation) and drives according to the map.
[0345] The self-position estimation is achieved by matching the three-dimensional map with the three-dimensional information around the own vehicle obtained by sensors such as the rangefinder (LIDAR, etc.) or stereo camera mounted on the own vehicle (hereinafter referred to as the own vehicle detection three-dimensional data), and estimating the own vehicle position within the three-dimensional map.
[0346] As shown in the HD map proposed by HERE, a 3D map is not only a 3D point cloud, but also includes 2D map data such as road and intersection shape information, and information such as congestion and accidents that change in real time. The 3D map is composed of multiple layers such as 3D data, 2D data, and metadata that changes in real time, and the device can obtain only the required data or refer to the required data.
[0347] The point cloud data may be the above-mentioned SWLD, or may include point group data that is not a feature point. Furthermore, the transmission and reception of the point cloud data is basically performed in one or more random access units.
[0348] As a matching method for the three-dimensional map and the three-dimensional data detected by the own vehicle, the following method can be adopted. For example, the device compares the shapes of the point groups in each other's point clouds and determines the parts with high similarity between the feature points as the same position. In addition, when the three-dimensional map is composed of SWLD, the device compares and matches the feature points constituting the SWLD with the three-dimensional feature points extracted from the three-dimensional data detected by the own vehicle.
[0349] Here, in order to estimate the own position with high accuracy, the following (A) and (B) need to be met: (A) a three-dimensional map and three-dimensional data for detecting the own vehicle are available, and (B) their accuracy meets a predetermined benchmark. However, in the following abnormal situation, (A) or (B) cannot be met.
[0350] (1) A three-dimensional map cannot be obtained through the communication path.
[0351] (2) There is no three-dimensional map or the obtained three-dimensional map is damaged.
[0352] (3) The sensor of the own vehicle fails, or due to bad weather, the accuracy of the three-dimensional data generated by the own vehicle is insufficient.
[0353] The following describes the operation for dealing with these abnormal situations. Although the operation is described below using a vehicle as an example, the following method can also be applied to all moving objects that move autonomously, such as robots and drones.
[0354] The following describes the configuration and operation of the three-dimensional information processing device according to the present embodiment for detecting abnormalities in three-dimensional data corresponding to a three-dimensional map or a vehicle. Fig.26 It is a block diagram showing a configuration example of a three-dimensional information processing device 700 according to this embodiment.
[0355] The three-dimensional information processing device 700 is mounted on a mobile object such as a motor vehicle. Fig.26 As shown, the three-dimensional information processing device 700 includes a three-dimensional map acquisition unit 701 , a vehicle detection data acquisition unit 702 , an abnormality determination unit 703 , a countermeasure operation determination unit 704 , and an operation control unit 705 .
[0356] In addition, the three-dimensional information processing device 700 may also include a camera for obtaining a two-dimensional image, or may include a sensor for one-dimensional data using ultrasonic waves or lasers, which is used to detect structural objects or moving objects around the vehicle, which is not shown in the figure. In addition, the three-dimensional information processing device 700 may also include a communication unit (not shown) for obtaining a three-dimensional map through a mobile communication network such as 4G or 5G, or communication between vehicles, or communication between roads and vehicles.
[0357] The three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 near the driving route. For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 through a mobile communication network, communication between vehicles, or communication between roads and vehicles.
[0358] Next, the own vehicle detection data acquisition unit 702 acquires the own vehicle detection three-dimensional data 712 based on the sensor information. For example, the own vehicle detection data acquisition unit 702 generates the own vehicle detection three-dimensional data 712 based on the sensor information acquired by the sensor included in the own vehicle.
[0359] Next, the abnormality determination unit 703 detects an abnormality by performing a predetermined check on at least one of the obtained three-dimensional map 711 and the vehicle detection three-dimensional data 712. That is, the abnormality determination unit 703 determines whether at least one of the obtained three-dimensional map 711 and the vehicle detection three-dimensional data 712 is abnormal.
[0360] When an abnormal situation is detected, the countermeasure action determination unit 704 determines a countermeasure action for the abnormal situation. Next, the action control unit 705 controls the operation of each processing unit required for executing the countermeasure action, such as the three-dimensional map acquisition unit 701.
[0361] In addition, when no abnormality is detected, the three-dimensional information processing device 700 ends the processing.
[0362] The three-dimensional information processing device 700 estimates the own position of the vehicle having the three-dimensional information processing device 700 using the three-dimensional map 711 and the own vehicle detection three-dimensional data 712. Then, the three-dimensional information processing device 700 uses the result of the own position estimation to make the vehicle automatically drive.
[0363] Accordingly, the three-dimensional information processing device 700 obtains map data (three-dimensional map 711) including the first three-dimensional position information via the channel. For example, the first three-dimensional position information is encoded in units of partial spaces having three-dimensional coordinate information, and the first three-dimensional position information includes a plurality of random access units, each of which is a collection of more than one partial space and can be independently decoded. For example, the first three-dimensional position information is data (SWLD) encoded with a feature point whose three-dimensional feature quantity is above a predetermined threshold.
[0364] Furthermore, the three-dimensional information processing device 700 generates the second three-dimensional position information (the vehicle detection three-dimensional data 712) based on the information detected by the sensor. Next, the three-dimensional information processing device 700 performs an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information to determine whether the first three-dimensional position information or the second three-dimensional position information is abnormal.
[0365] When determining that the first three-dimensional position information or the second three-dimensional position information is abnormal, the three-dimensional information processing device 700 determines a countermeasure for the abnormality. Next, the three-dimensional information processing device 700 executes control required for executing the countermeasure.
[0366] According to this, the three-dimensional information processing device 700 can detect abnormality in the first three-dimensional position information or the second three-dimensional position information and perform a corresponding operation.
[0367] (Implementation method 5)
[0368] In this embodiment, a method of transmitting three-dimensional data to a rear vehicle and the like will be described.
[0369] Fig. 27 1 is a block diagram showing a configuration example of a three-dimensional data production device 810 according to the present embodiment. The three-dimensional data production device 810 is mounted on a vehicle, for example. The three-dimensional data production device 810 transmits and receives three-dimensional data with external traffic cloud monitoring, a vehicle in front, or a vehicle behind, and produces and accumulates the three-dimensional data.
[0370] The three-dimensional data production device 810 includes: a data receiving unit 811, a communication unit 812, a receiving control unit 813, a format conversion unit 814, multiple sensors 815, a three-dimensional data production unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a sending control unit 820, a format conversion unit 821, and a data sending unit 822.
[0371] The data receiving unit 811 receives three-dimensional data 831 from traffic cloud monitoring or the vehicle ahead. The three-dimensional data 831 includes, for example, a point cloud including information on areas that cannot be detected by the vehicle's sensor 815, visible light images, depth information, sensor position information, or speed information.
[0372] The communication unit 812 communicates with the traffic cloud monitoring or the vehicle ahead, and sends a data transmission request or the like to the traffic cloud monitoring or the vehicle ahead.
[0373] The reception control unit 813 exchanges information such as a corresponding format with the communication partner via the communication unit 812, and establishes communication with the communication partner.
[0374] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data receiving unit 811. Furthermore, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.
[0375] The plurality of sensors 815 are a group of sensors such as LiDAR, a visible light camera, or an infrared camera that obtains information outside the vehicle and generates sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point group data). In addition, the number of sensors 815 may not be multiple.
[0376] The three-dimensional data generating unit 816 generates three-dimensional data 834 based on the sensor information 833. The three-dimensional data 834 includes, for example, point cloud, visible light image, depth information, sensor position information, or speed information.
[0377] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 produced by traffic cloud monitoring or the front vehicle, etc., into the three-dimensional data 834 produced based on the sensor information 833 of the own vehicle, thereby constructing three-dimensional data 835 that also includes the space in front of the front vehicle that cannot be detected by the sensor 815 of the own vehicle.
[0378] The three-dimensional data accumulation unit 818 accumulates the generated three-dimensional data 835 and the like.
[0379] The communication unit 819 communicates with the traffic cloud monitoring or the vehicle behind, and sends a data transmission request and the like to the traffic cloud monitoring or the vehicle behind.
[0380] The transmission control unit 820 exchanges information such as the corresponding format with the communication partner via the communication unit 819 to establish communication with the communication partner. In addition, the transmission control unit 820 determines the transmission area of the space of the three-dimensional data to be transmitted based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication partner.
[0381] Specifically, the transmission control unit 820 determines the transmission area including the space in front of the vehicle that cannot be detected by the sensor of the rear vehicle according to the data transmission request from the traffic cloud monitoring or the rear vehicle. In addition, the transmission control unit 820 determines the transmission area by judging whether the space that can be transmitted or the transmitted space has been updated based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the area that is both specified by the data transmission request and the area where the corresponding three-dimensional data 835 exists as the transmission area. In addition, the transmission control unit 820 notifies the format corresponding to the communication partner and the transmission area to the format conversion unit 821.
[0382] The format conversion unit 821 generates three-dimensional data 837 by converting the three-dimensional data 836 of the transmission area in the three-dimensional data 835 accumulated in the three-dimensional data accumulation unit 818 into a format corresponding to the receiving side. In addition, the three-dimensional data 837 may be compressed or encoded by the format conversion unit 821 to reduce the data amount.
[0383] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic cloud monitoring or the rear vehicle. The three-dimensional data 837 includes, for example, a point cloud in front of the vehicle including information on the area that becomes a blind spot for the rear vehicle, a visible light image, depth information, or sensor position information.
[0384] In addition, although the example in which the format conversion units 814 and 821 perform format conversion is described here, format conversion may not be performed.
[0385] With this configuration, the three-dimensional data production device 810 obtains three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the own vehicle from the outside, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 with three-dimensional data 834 based on sensor information 833 detected by the sensor 815 of the own vehicle. In this way, the three-dimensional data production device 810 can generate three-dimensional data of a range that cannot be detected by the sensor 815 of the own vehicle.
[0386] In addition, the three-dimensional data production device 810 can send three-dimensional data of the space in front of its own vehicle that cannot be detected by the sensors of the rear vehicle to the traffic cloud monitoring or the rear vehicle, etc. according to the data sending request from the traffic cloud monitoring or the rear vehicle.
[0387] (Implementation 6)
[0388] In the fifth embodiment, a client device such as a vehicle sends three-dimensional data to another vehicle or a server such as a traffic cloud monitoring device. In this embodiment, the client device sends sensor information obtained by the sensor to the server or other client devices.
[0389] First, the configuration of a system according to this embodiment will be described. Fig.28 1 shows the structure of the system for transmitting and receiving three-dimensional maps and sensor information according to the present embodiment. The system includes a server 901 and client devices 902A and 902B. In addition, when the client devices 902A and 902B are not specifically distinguished, they are also referred to as client devices 902.
[0390] The client device 902 is, for example, an in-vehicle device mounted on a mobile body such as a vehicle. The server 901 is, for example, a traffic cloud monitoring system, and can communicate with a plurality of client devices 902 .
[0391] The server 901 transmits the three-dimensional map composed of the point cloud to the client device 902. In addition, the composition of the three-dimensional map is not limited to the point cloud, and can also be expressed by other three-dimensional data such as a mesh structure.
[0392] The client device 902 sends the sensor information obtained by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR information, visible light images, infrared images, depth images, sensor position information, and speed information.
[0393] The data sent and received between the server 901 and the client device 902 may be compressed when it is desired to reduce the data, and may not be compressed when it is desired to maintain the accuracy of the data. When compressing the data, a three-dimensional compression method based on an octree may be used in the point cloud, for example. In addition, a two-dimensional image compression method may be used in the visible light image, the infrared image, and the depth image. The two-dimensional image compression method is, for example, MPEG-4AVC or HEVC standardized by MPEG.
[0394] Furthermore, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in accordance with the three-dimensional map transmission request from the client device 902. In addition, the server 901 may transmit the three-dimensional map without waiting for the three-dimensional map transmission request from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a predetermined space. Furthermore, the server 901 may transmit the three-dimensional map suitable for the location of the client device 902 at regular intervals to the client device 902 that has received a transmission request once. Furthermore, the server 901 may transmit the three-dimensional map to the client device 902 every time the three-dimensional map managed by the server 901 is updated.
[0395] The client device 902 sends a three-dimensional map transmission request to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 sends a three-dimensional map transmission request to the server 901.
[0396] In addition, in the following cases, the client device 902 may also issue a three-dimensional map transmission request to the server 901. In the case where the three-dimensional map held by the client device 902 is relatively old, the client device 902 may also issue a three-dimensional map transmission request to the server 901. For example, in the case where the client device 902 obtains the three-dimensional map and a certain period of time has passed, the client device 902 may also issue a three-dimensional map transmission request to the server 901.
[0397] Alternatively, the client device 902 may send a three-dimensional map transmission request to the server 901 before a certain time when the client device 902 is about to leave the space shown on the three-dimensional map held by the client device 902. For example, the client device 902 may send a three-dimensional map transmission request to the server 901 when the client device 902 is within a predetermined distance from the boundary of the space shown on the three-dimensional map held by the client device 902. Furthermore, when the moving path and moving speed of the client device 902 are known, the time when the client device 902 leaves the space shown on the three-dimensional map held by the client device 902 may be predicted based on the known moving path and moving speed.
[0398] When the error between the three-dimensional data generated by the client device 902 based on the sensor information and the position of the three-dimensional map is greater than a certain range, the client device 902 may send a request to the server 901 to send the three-dimensional map.
[0399] The client device 902 transmits the sensor information to the server 901 in accordance with the sensor information transmission request transmitted from the server 901. In addition, the client device 902 may transmit the sensor information to the server 901 without waiting for the sensor information transmission request from the server 901. For example, when the client device 902 has received a sensor information transmission request from the server 901 once, the client device 902 may periodically transmit the sensor information to the server 901 within a certain period. Furthermore, when the error between the three-dimensional data generated by the client device 902 based on the sensor information and the position of the three-dimensional map obtained from the server 901 is greater than a certain range, the client device 902 may determine that there is a possibility that the three-dimensional map around the client device 902 has changed, and transmit the determination result together with the sensor information to the server 901.
[0400] The server 901 issues a request to send sensor information to the client device 902. For example, the server 901 receives location information of the client device 902 such as GPS from the client device 902. When the server 901 determines that the client device 902 is close to a space with little information in the three-dimensional map managed by the server 901 based on the location information of the client device 902, the server 901 issues a request to send sensor information to the client device 902 in order to regenerate the three-dimensional map. In addition, the server 901 may issue a request to send sensor information when it is desired to update the three-dimensional map, when it is desired to check the road conditions such as when snow is accumulated or when a disaster occurs, when it is desired to check the congestion conditions or the accident conditions.
[0401] Furthermore, the client device 902 may set the data volume of the sensor information to be sent to the server 901 according to the communication state or the frequency band when receiving the sensor information sending request received from the server 901. Setting the data volume of the sensor information to be sent to the server 901 means, for example, increasing or decreasing the data itself or selecting an appropriate compression method.
[0402] Fig.29 is a block diagram showing an example of the configuration of the client device 902. The client device 902 receives a three-dimensional map composed of a point cloud or the like from the server 901, and estimates the position of the client device 902 itself based on the three-dimensional data produced based on the sensor information of the client device 902. The client device 902 then transmits the obtained sensor information to the server 901.
[0403] The client device 902 includes: a data receiving unit 1011, a communication unit 1012, a receiving control unit 1013, a format conversion unit 1014, multiple sensors 1015, a three-dimensional data production unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a sending control unit 1021, and a data sending unit 1022.
[0404] The data receiving unit 1011 receives a three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including a point cloud such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0405] The communication unit 1012 communicates with the server 901 and transmits a data transmission request (for example, a transmission request of a three-dimensional map) and the like to the server 901 .
[0406] The reception control unit 1013 exchanges information such as the corresponding format with the communication partner via the communication unit 1012, and establishes communication with the communication partner.
[0407] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion on the three-dimensional map 1031 received by the data receiving unit 1011. Furthermore, the format conversion unit 1014 performs decompression or decoding processing when the three-dimensional map 1031 is compressed or encoded. In addition, the format conversion unit 1014 does not perform decompression or decoding processing when the three-dimensional map 1031 is non-compressed data.
[0408] The plurality of sensors 1015 are a group of sensors mounted on the client device 902 such as LiDAR, a visible light camera, an infrared camera, or a depth sensor for obtaining information outside the vehicle, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point group data). In addition, the number of sensors 1015 may not be multiple.
[0409] The three-dimensional data generating unit 1016 generates three-dimensional data 1034 of the surroundings of the own vehicle based on the sensor information 1033. For example, the three-dimensional data generating unit 1016 generates point cloud data having color information of the surroundings of the own vehicle using information obtained by LiDAR and visible light images obtained by a visible light camera.
[0410] The three-dimensional image processing unit 1017 uses the received three-dimensional map 1032 such as the point cloud and the three-dimensional data 1034 of the surroundings of the own vehicle generated based on the sensor information 1033 to perform the own vehicle's own position estimation processing, etc. Alternatively, the three-dimensional image processing unit 1017 may synthesize the three-dimensional map 1032 and the three-dimensional data 1034 to produce three-dimensional data 1035 of the surroundings of the own vehicle, and use the produced three-dimensional data 1035 to perform the own position estimation processing.
[0411] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032 , the three-dimensional data 1034 , the three-dimensional data 1035 , and the like.
[0412] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format corresponding to the receiving side. In addition, the format conversion unit 1019 can reduce the amount of data by compressing or encoding the sensor information 1037. In addition, when format conversion is not required, the format conversion unit 1019 can omit the processing. In addition, the format conversion unit 1019 can control the amount of data to be transmitted according to the designation of the transmission range.
[0413] The communication unit 1020 communicates with the server 901 , and receives a data transmission request (a sensor information transmission request) and the like from the server 901 .
[0414] The transmission control unit 1021 exchanges information such as a corresponding format with the communication partner via the communication unit 1020, thereby establishing communication.
[0415] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information obtained by multiple sensors 1015, such as information obtained by LiDAR, a brightness image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information.
[0416] Next, the configuration of the server 901 will be described. Fig.30 1 is a block diagram showing an example of the configuration of the server 901. The server 901 receives sensor information sent from the client device 902, and creates three-dimensional data based on the received sensor information. The server 901 updates the three-dimensional map managed by the server 901 using the created three-dimensional data. Furthermore, the server 901 sends the updated three-dimensional map to the client device 902 in accordance with the three-dimensional map sending request from the client device 902.
[0417] The server 901 includes: a data receiving unit 1111, a communication unit 1112, a receiving control unit 1113, a format conversion unit 1114, a three-dimensional data production unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a sending control unit 1121, and a data sending unit 1122.
[0418] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information obtained by LiDAR, a brightness image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information.
[0419] The communication unit 1112 communicates with the client device 902 and transmits a data transmission request (for example, a sensor information transmission request) and the like to the client device 902 .
[0420] The reception control unit 1113 exchanges information such as the corresponding format with the communication partner via the communication unit 1112, thereby establishing communication.
[0421] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 performs decompression or decoding processing to generate the sensor information 1132. When the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.
[0422] The three-dimensional data generating unit 1116 generates three-dimensional data 1134 of the surroundings of the client device 902 based on the sensor information 1132. For example, the three-dimensional data generating unit 1116 generates point cloud data having color information of the surroundings of the client device 902 using information obtained by LiDAR and visible light images obtained by a visible light camera.
[0423] The three-dimensional data synthesis unit 1117 synthesizes the three-dimensional data 1134 generated based on the sensor information 1132 with the three-dimensional map 1135 managed by the server 901 , thereby updating the three-dimensional map 1135 .
[0424] The three-dimensional data storage unit 1118 stores the three-dimensional map 1135 and the like.
[0425] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format corresponding to the receiving side. In addition, the format conversion unit 1119 can also reduce the amount of data by compressing or encoding the three-dimensional map 1135. In addition, when format conversion is not required, the format conversion unit 1119 can also omit the processing. In addition, the format conversion unit 1119 can control the amount of data to be sent according to the designation of the sending range.
[0426] The communication unit 1120 communicates with the client device 902 , and receives a data transmission request (a transmission request of a three-dimensional map) and the like from the client device 902 .
[0427] The transmission control unit 1121 exchanges information such as a corresponding format with the communication partner via the communication unit 1120 to establish communication.
[0428] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including a point cloud such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0429] Next, the workflow of the client device 902 will be described. Fig.31 1 is a flowchart showing the operation performed by the client device 902 when obtaining a three-dimensional map.
[0430] First, the client device 902 requests the server 901 to send a three-dimensional map (point cloud, etc.) (S1001). At this time, the client device 902 also sends the location information of the client device 902 obtained by GPS, etc., and accordingly, the server 901 can be requested to send a three-dimensional map related to the location information.
[0431] Next, the client device 902 receives the three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate a non-compressed three-dimensional map (S1003).
[0432] Next, the client device 902 creates three-dimensional data 1034 of the surroundings of the client device 902 based on the sensor information 1033 obtained from the plurality of sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created based on the sensor information 1033 (S1005).
[0433] Fig.32 101 is a flowchart showing the operation of the client device 902 when transmitting sensor information. First, the client device 902 receives a request to transmit sensor information from the server 901 (S1011). The client device 902 that has received the transmission request transmits the sensor information 1037 to the server 901 (S1012). In addition, when the sensor information 1033 includes a plurality of information obtained by a plurality of sensors 1015, the client device 902 compresses each information in a compression method suitable for each information, thereby generating the sensor information 1037.
[0434] Next, the workflow of the server 901 is described. Fig.33 1 is a flowchart showing the operation of the server 901 when acquiring sensor information. First, the server 901 requests the client device 902 to send sensor information (S1021). Next, the server 901 receives the sensor information 1037 sent from the client device 902 in accordance with the request (S1022). Next, the server 901 uses the received sensor information 1037 to create three-dimensional data 1134 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 on the three-dimensional map 1135 (S1024).
[0435] Fig.341031 is a flowchart showing the work of the server 901 when sending a three-dimensional map. First, the server 901 receives a request to send a three-dimensional map from the client device 902 (S1031). The server 901 that has received the request to send a three-dimensional map sends the three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 can extract a three-dimensional map in the vicinity corresponding to the location information of the client device 902, and send the extracted three-dimensional map. Furthermore, the server 901 can compress the three-dimensional map composed of the point cloud, for example, using an octree compression method, and send the compressed three-dimensional map.
[0436] Hereinafter, modified examples of the present embodiment will be described.
[0437] The server 901 uses the sensor information 1037 received from the client device 902 to create three-dimensional data 1134 near the location of the client device 902. Next, the server 901 matches the created three-dimensional data 1134 with a three-dimensional map 1135 of the same area managed by the server 901, and calculates the difference between the three-dimensional data 1134 and the three-dimensional map 1135. When the difference is greater than a predetermined threshold, the server 901 determines that some abnormality has occurred around the client device 902. For example, when the ground subsidence occurs due to a natural disaster such as an earthquake, it can be considered that there will be a large difference between the three-dimensional map 1135 managed by the server 901 and the three-dimensional data 1134 created based on the sensor information 1037.
[0438] The sensor information 1037 may also include at least one of the type of sensor, the performance of the sensor, and the model of the sensor. In addition, a category ID corresponding to the performance of the sensor may be added to the sensor information 1037. For example, when the sensor information 1037 is information obtained by LiDAR, an identifier may be assigned in consideration of the performance of the sensor, for example, category 1 may be assigned to a sensor that can obtain information with an accuracy of several mm, category 2 may be assigned to a sensor that can obtain information with an accuracy of several cm, and category 3 may be assigned to a sensor that can obtain information with an accuracy of several m. In addition, the server 901 may estimate the performance information of the sensor from the model of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 may determine the specification information of the sensor according to the model of the vehicle. In this case, the server 901 may obtain the information of the model of the vehicle in advance, or include the information in the sensor information. In addition, the server 901 may switch the degree of correction of the three-dimensional data 1134 produced using the sensor information 1037 using the obtained sensor information 1037. For example, when the sensor performance is high accuracy (category 1), the server 901 does not perform correction on the three-dimensional data 1134. When the sensor performance is low accuracy (category 3), the server 901 applies correction suitable for the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (intensity) of correction as the accuracy of the sensor is lower.
[0439] The server 901 may also simultaneously send a request to send sensor information to multiple client devices 902 existing in a certain space. When the server 901 receives multiple sensor information from multiple client devices 902, it is not necessary to use all the sensor information to create the three-dimensional data 1134. For example, the server 901 may select the sensor information to be used according to the performance of the sensor. For example, when the server 901 updates the three-dimensional map 1135, it may select high-precision sensor information (category 1) from the received multiple sensor information and use the selected sensor information to create the three-dimensional data 1134.
[0440] The server 901 is not limited to servers such as traffic cloud monitoring, but can also be other client devices (car-mounted). Fig.35 The system configuration in this case is shown.
[0441] For example, the client device 902C sends a request to send sensor information to the client device 902A that is nearby, and obtains the sensor information from the client device 902A. Then, the client device 902C uses the obtained sensor information of the client device 902A to create three-dimensional data, and updates the three-dimensional map of the client device 902C. In this way, the client device 902C can utilize the performance of the client device 902C to generate a three-dimensional map of the space that can be obtained from the client device 902A. For example, when the performance of the client device 902C is high, this situation can be considered to occur.
[0442] In this case, the client device 902A that has provided the sensor information is given the right to obtain the high-precision three-dimensional map generated by the client device 902C. The client device 902A receives the high-precision three-dimensional map from the client device 902C in accordance with the right.
[0443] Furthermore, the client device 902C may send a request to send sensor information to multiple client devices 902 (client device 902A and client device 902B) in the vicinity. If the sensor of the client device 902A or the client device 902B is high-performance, the client device 902C can use the sensor information obtained by the high-performance sensor to create three-dimensional data.
[0444] Fig.36 1 is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example, a 3D map compression / decoding processing unit 1201 that compresses and decodes a 3D map and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.
[0445] The client device 902 includes: a three-dimensional map decoding processing unit 1211, and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives the coded data of the compressed three-dimensional map, decodes the coded data and obtains the three-dimensional map. The sensor information compression processing unit 1212 does not compress the three-dimensional data produced by the obtained sensor information, but compresses the sensor information itself, and sends the compressed coded data of the sensor information to the server 901. According to this structure, the client device 902 can keep the processing unit (device or LSI) for decoding the three-dimensional map (point cloud, etc.) inside, without having to keep the processing unit for compressing the three-dimensional data of the three-dimensional map (point cloud, etc.) inside. In this way, the cost and power consumption of the client device 902 can be suppressed.
[0446] As described above, the client device 902 involved in this embodiment is mounted on the mobile body, and generates three-dimensional data 1034 of the surroundings of the mobile body based on the sensor information 1033 indicating the surrounding conditions of the mobile body obtained by the sensor 1015 mounted on the mobile body. The client device 902 estimates the own position of the mobile body using the generated three-dimensional data 1034. The client device 902 transmits the obtained sensor information 1033 to the server 901 or another mobile body 902.
[0447] Based on this, the client device 902 transmits the sensor information 1033 to the server 901 or the like. In this way, the amount of data to be transmitted may be reduced compared to the case of transmitting three-dimensional data. Furthermore, since it is not necessary to perform processing such as compression or encoding of three-dimensional data on the client device 902, the amount of processing on the client device 902 can be reduced. Therefore, the client device 902 can reduce the amount of data to be transmitted or simplify the configuration of the device.
[0448] Furthermore, the client device 902 further sends a request to send a three-dimensional map to the server 901, and receives the three-dimensional map 1031 from the server 901. The client device 902 estimates its own position using the three-dimensional data 1034 and the three-dimensional map 1032.
[0449] Furthermore, the sensor information 1033 includes at least one of information obtained by a laser sensor, a brightness image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.
[0450] Also, the sensor information 1033 includes information showing the performance of the sensor.
[0451] Furthermore, the client device 902 encodes or compresses the sensor information 1033, and transmits the encoded or compressed sensor information 1037 to the server 901 or another mobile body 902. Thus, the client device 902 can reduce the amount of data to be transmitted.
[0452] For example, the client device 902 includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0453] Furthermore, the server 901 according to the present embodiment can communicate with the client device 902 mounted on the mobile body, and receive sensor information 1037 indicating the surrounding conditions of the mobile body obtained by the sensor 1015 mounted on the mobile body from the client device 902. The server 901 generates three-dimensional data 1134 of the surroundings of the mobile body based on the received sensor information 1037.
[0454] Accordingly, the server 901 uses the sensor information 1037 sent from the client device 902 to create the three-dimensional data 1134. In this way, compared with the case where the client device 902 sends the three-dimensional data, it is possible to reduce the amount of data to be sent. In addition, since it is not necessary to perform processing such as compression or encoding of the three-dimensional data on the client device 902, the processing amount of the client device 902 can be reduced. In this way, the server 901 can reduce the amount of data to be transmitted or simplify the configuration of the device.
[0455] Furthermore, the server 901 further transmits a request for transmitting sensor information to the client device 902 .
[0456] Furthermore, the server 901 further updates the three-dimensional map 1135 using the created three-dimensional data 1134 , and transmits the three-dimensional map 1135 to the client device 902 in response to a transmission request for the three-dimensional map 1135 from the client device 902 .
[0457] Furthermore, the sensor information 1037 includes at least one of information obtained by a laser sensor, a brightness image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.
[0458] Also, the sensor information 1037 includes information showing the performance of the sensor.
[0459] Furthermore, the server 901 further calibrates the three-dimensional data according to the performance of the sensor. Accordingly, the three-dimensional data production method can improve the quality of the three-dimensional data.
[0460] Furthermore, when receiving sensor information, the server 901 receives a plurality of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 used in the production of the three-dimensional data 1134 based on a plurality of information indicating the performance of the sensors included in the plurality of sensor information 1037. In this way, the server 901 can improve the quality of the three-dimensional data 1134.
[0461] Furthermore, the server 901 decodes or decompresses the received sensor information 1137, and creates three-dimensional data 1134 based on the decoded or decompressed sensor information 1132. In this way, the server 901 can reduce the amount of data to be transmitted.
[0462] For example, the server 901 includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0463] (Implementation 7)
[0464] In this embodiment, a method for encoding and a method for decoding three-dimensional data using an inter-frame prediction process will be described.
[0465] Fig.37 3D data encoding device 1300 according to the present embodiment is a block diagram. The 3D data encoding device 1300 generates a coded bit stream (hereinafter also simply referred to as a bit stream) as a coded signal by encoding 3D data. Fig.37 As shown, the three-dimensional data encoding device 1300 includes: a segmentation unit 1301, a subtraction unit 1302, a transformation unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transformation unit 1306, an addition unit 1307, a reference volume memory 1308, an intra-frame prediction unit 1309, a reference space memory 1310, an inter-frame prediction unit 1311, a prediction control unit 1312, and an entropy coding unit 1313.
[0466] The segmentation unit 1301 segments each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM) as coding units. Furthermore, the segmentation unit 1301 performs octree representation (Octree conversion) on the voxels in each volume. In addition, the segmentation unit 1301 may make the space and the volume the same size and perform octree representation on the space. Furthermore, the segmentation unit 1301 may also attach information required for octree conversion (depth information, etc.) to the header of the bitstream, etc.
[0467] The subtraction unit 1302 calculates the difference between the volume (coding target volume) output from the division unit 1301 and the prediction volume generated by intra prediction or inter prediction described later, and outputs the calculated difference as a prediction residual to the transformation unit 1303 . Fig.38 An example of calculating the prediction residual is shown. The bit strings of the encoding target volume and the prediction volume shown here are, for example, position information indicating the positions of three-dimensional points (for example, point clouds) included in the volume.
[0468] The following describes the octree representation and the scanning order of voxels. The volume is transformed into an octree structure (octreeization) and then encoded. The octree structure consists of nodes and leaf nodes. Each node has 8 nodes or leaf nodes, and each leaf node has voxel (VXL) information. Fig.39 An example of the configuration of a volume including a plurality of voxels is shown. Fig.40 Shows the Fig.39 The volume shown is transformed into an example of an octree structure. Here, Fig.40 The leaf nodes 1, 2, and 3 in the leaf nodes shown represent Fig.39 The voxels VXL1, VXL2, and VXL3 shown express the VXL (hereinafter referred to as effective VXL) including the point group.
[0469] The octree is represented by a binary sequence of 0 and 1. For example, when a node or valid VXL is set to a value of 1 and the others are set to a value of 0, each node and leaf node is assigned Fig.40 The binary sequence shown in FIG. Then, the binary sequence is scanned in a width-first or depth-first scanning order. For example, in the case of a width-first scan, Fig.41 The binary sequence shown in A. When the depth-first scan is performed, we get Fig.41 The binary sequence shown in B. The binary sequence obtained by this scanning is encoded by entropy coding, so that the amount of information is reduced.
[0470] Next, the depth information in the octree representation is explained. The depth in the octree representation is used to control the granularity of the point cloud information contained in the volume. If the depth is set to a large value, the point cloud information can be reproduced at a finer level, but the amount of data used to represent the nodes and leaf nodes will increase. On the contrary, if the depth is set to a small value, although the amount of data can be reduced, multiple point cloud information at different positions and colors will be regarded as the same position and the same color, so the original information of the point cloud information will be lost.
[0471] For example, Fig.42 Shows the Fig.40 The octree with depth = 2 is shown as an example of expressing an octree with depth = 1. Fig.42 The octree shown is Fig.40 The octree shown has a small amount of data. Fig.42 The octree shown is similar to Fig.42 Compared with the octree shown, the number of bits after binary serialization is less. Fig.40 The leaf nodes 1 and 2 shown in the figure become Fig.41 The leaf node 1 shown is shown. That is, Fig.40 The leaf node 1 and the leaf node 2 shown are information of different locations.
[0472] Fig.43 Shown with Fig.42 The volume corresponding to the octree shown. Fig.39 The VXL1 and VXL2 shown are Fig.43 In this case, the three-dimensional data encoding device 1300 corresponds to VXL12 shown in FIG. Fig.39 The color information of VXL1 and VXL2 shown in the figure generates Fig.43For example, the three-dimensional data encoding device 1300 calculates the color information of VXL1 and VXL2 as the color information of VXL12 using the average value, median value, or weighted average value. In this way, the three-dimensional data encoding device 1300 can control the reduction of the data amount by changing the depth of the octree.
[0473] The three-dimensional data encoding device 1300 may also use any one of the world space units, space units, and volume units to set the depth information of the octree. In addition, at this time, the three-dimensional data encoding device 1300 may also attach the depth information to the header information of the world space, the header information of the space, or the header information of the volume. In addition, the same value may be used as the depth information in all world spaces, spaces, and volumes at different times. In this case, the three-dimensional data encoding device 1300 may also attach the depth information to the header information that manages the world space at all times.
[0474] When the voxel contains color information, the transformation unit 1303 applies a frequency transformation such as an orthogonal transformation to the prediction residual of the color information of the voxels in the volume. For example, the transformation unit 1303 scans the prediction residual in a certain scanning order to produce a one-dimensional arrangement. Thereafter, the transformation unit 1303 transforms the one-dimensional arrangement into the frequency domain by applying a one-dimensional orthogonal transformation to the produced one-dimensional arrangement. Accordingly, when the value of the prediction residual in the volume is close, the value of the frequency component of the low frequency band becomes larger, and the value of the frequency component of the high frequency band becomes smaller. Therefore, the quantization unit 1304 can more effectively reduce the amount of coding.
[0475] Furthermore, the transformation unit 1303 may use orthogonal transformation of two or more dimensions instead of one-dimensional orthogonal transformation. For example, the transformation unit 1303 maps the prediction residual into a two-dimensional arrangement in a certain scanning order, and applies a two-dimensional orthogonal transformation to the obtained two-dimensional arrangement. Furthermore, the transformation unit 1303 may select the orthogonal transformation method to be used from a plurality of orthogonal transformation methods. In this case, the three-dimensional data encoding device 1300 attaches information indicating which orthogonal transformation method is used to the bit stream. Furthermore, the transformation unit 1303 may select the orthogonal transformation method to be used from a plurality of orthogonal transformation methods with different dimensions. In this case, the three-dimensional data encoding device 1300 attaches information indicating which dimensional orthogonal transformation method is used to the bit stream.
[0476] For example, the transformation unit 1303 matches the scanning order of the prediction residual with the scanning order (width-first or depth-first, etc.) in the octree within the volume. Accordingly, since it is not necessary to attach information showing the scanning order of the prediction residual to the bitstream, the additional overhead can be reduced. In addition, the transformation unit 1303 may also apply a scanning order different from the scanning order of the octree. In this case, the three-dimensional data encoding device 1300 attaches information showing the scanning order of the prediction residual to the bitstream. Accordingly, the three-dimensional data encoding device 1300 can efficiently encode the prediction residual. In addition, the three-dimensional data encoding device 1300 may attach information (flag, etc.) indicating whether the scanning order of the octree is applicable to the bitstream, and when the scanning order of the octree is not applicable, the information showing the scanning order of the prediction residual is attached to the bitstream.
[0477] The conversion unit 1303 may convert not only the prediction residual of the color information but also other attribute information of the voxel. For example, the conversion unit 1303 may convert and encode information such as reflectivity obtained when a point cloud is obtained by LiDAR or the like.
[0478] The transform unit 1303 may skip the process when the space does not have attribute information such as color information. Furthermore, the three-dimensional data encoding device 1300 may add information (flag) indicating whether to skip the process of the transform unit 1303 to the bit stream.
[0479] The quantization unit 1304 quantizes the frequency components of the prediction residual generated by the transformation unit 1303 using the quantization control parameters to generate quantization coefficients. The amount of information is reduced accordingly. The generated quantization coefficients are output to the entropy coding unit 1313. The quantization unit 1304 can control the quantization control parameters according to world space units, space units, or volume units. At this time, the three-dimensional data encoding device 1300 attaches the quantization control parameters to the respective header information, etc. In addition, the quantization unit 1304 can also change the weight according to the frequency components of each prediction residual to perform quantization control. For example, the quantization unit 1304 can perform detailed quantization on the low-frequency components and coarse quantization on the high-frequency components. In this case, the three-dimensional data encoding device 1300 can attach parameters representing the weights of each frequency component to the header.
[0480] The quantization unit 1304 may skip the process when the space does not have attribute information such as color information. In addition, the three-dimensional data encoding device 1300 may add information (flag) indicating whether the process of the quantization unit 1304 is skipped to the bit stream.
[0481] The inverse quantization unit 1305 inversely quantizes the quantization coefficients generated by the quantization unit 1304 using the quantization control parameters, thereby generating inverse quantization coefficients of the prediction residual, and outputs the generated inverse quantization coefficients to the inverse transformation unit 1306 .
[0482] The inverse transform unit 1306 generates an inverse transform applied prediction residual by applying inverse transform to the inverse quantized coefficients generated by the inverse quantization unit 1305. Since the inverse transform applied prediction residual is a prediction residual generated after quantization, it may not be completely consistent with the prediction residual output by the transform unit 1303.
[0483] The adding unit 1307 adds the prediction residual after inverse transformation generated by the inverse transform unit 1306 and the prediction volume generated by intra-frame prediction or inter-frame prediction described later, which is used to generate the prediction residual before quantization, to generate a reconstructed volume. The reconstructed volume is stored in the reference volume memory 1308 or the reference space memory 1310.
[0484] The intra prediction unit 1309 generates a predicted volume of the encoding target volume using the attribute information of the adjacent volume stored in the reference volume memory 1308. The attribute information includes the color information or reflectance of the voxel. The intra prediction unit 1309 generates a predicted value of the color information or reflectance of the encoding target volume.
[0485] Fig.44 1309 is a diagram for explaining the operation of the intra prediction unit 1309. For example, Fig.44 As shown, the intra-frame prediction unit 1309 generates a predicted volume of the encoding target volume (volume idx=3) based on the adjacent volume (volume idx=0). Here, volume idx is identifier information added to the volume in the space, and different values are assigned to each volume. The order of assigning volume idx can be the same as the encoding order or different from the encoding order. For example, as Fig.44 The intra prediction unit 1309 uses the average value of the color information of the voxels contained in the volume idx=0 which is the adjacent volume as the predicted value of the color information of the encoding object volume shown. In this case, the prediction residual is generated by subtracting the predicted value of the color information from the color information of each voxel contained in the encoding object volume. The processing after the transformation unit 1303 is performed on the prediction residual. And, in this case, the three-dimensional data encoding device 1300 adds the adjacent volume information and the prediction mode information to the bit stream. Here, the adjacent volume information is information showing the adjacent volume used in the prediction, for example, the volume idx showing the adjacent volume used in the prediction. And the prediction mode information shows the mode used in the generation of the prediction volume. The mode refers to, for example, an average value mode for generating a prediction value based on the average value of the voxels in the adjacent volume, or an intermediate value mode for generating a prediction value based on the intermediate value of the voxels in the adjacent volume.
[0486] The intra-frame prediction unit 1309 may also generate a prediction volume based on a plurality of adjacent volumes. Fig.44 In the illustrated configuration, the intra prediction unit 1309 generates prediction volume 0 based on volume idx=0, and generates prediction volume 1 based on volume idx=1. Then, the intra prediction unit 1309 generates the final prediction volume by averaging the prediction volume 0 and the prediction volume 1. In this case, the three-dimensional data encoding device 1300 may also attach multiple volume idxs of the multiple volumes used in generating the prediction volume to the bitstream.
[0487] Fig.45 The inter-frame prediction process involved in this embodiment is shown in the mode. The inter-frame prediction unit 1311 uses the coded space at a different time T_LX to encode (inter-frame prediction) for the space (SPC) at a certain time T_Cur. In this case, the inter-frame prediction unit 1311 applies rotation and translation processing to the coded space at different time T_LX to perform encoding processing.
[0488] Furthermore, the three-dimensional data encoding device 1300 adds RT information related to the rotation and translation processing of the space at the different time T_LX to the bitstream. The different time T_LX is, for example, the time T_L0 before the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 may also add RT information RT_L0 related to the rotation and translation processing of the space at the time T_L0 to the bitstream.
[0489] Alternatively, the different time T_LX is, for example, time T_L1 after the certain time T_Cur. In this case, the three-dimensional data encoding apparatus 1300 may add RT information RT_L1 about the spatial rotation and translation processing applied at time T_L1 to the bit stream.
[0490] Alternatively, the inter prediction unit 1311 performs encoding by referring to spaces at different times T_L0 and T_L1 (bi-prediction). In this case, the three-dimensional data encoding device 1300 may add both RT information RT_L0 and RT_L1 about rotation and translation applied to the space to the bitstream.
[0491] In addition, although T_L0 is set as the time before T_Cur and T_L1 is set as the time after T_Cur, it is not limited to this. For example, T_L0 and T_L1 can both be the time before T_Cur. Or, T_L0 and T_L1 can both be the time after T_Cur.
[0492] Furthermore, the three-dimensional data encoding device 1300 may add RT information related to the rotation and translation applied to each space to the bitstream when encoding with reference to multiple spaces at different times. For example, the three-dimensional data encoding device 1300 manages the multiple encoded spaces referred to by two reference lists (L0 list and L1 list). When the first reference space in the L0 list is set to L0R0, the second reference space in the L0 list is set to L0R1, the first reference space in the L1 list is set to L1R0, and the second reference space in the L1 list is set to L1R1, the three-dimensional data encoding device 1300 adds RT information RT_L0R0 of L0R0, RT information RT_L0R1 of L0R1, RT information RT_L1R0 of L1R0, and RT information RT_L1R1 of L1R1 to the bitstream. For example, the three-dimensional data encoding device 1300 adds these RT information to the header of the bitstream, etc.
[0493] Furthermore, the three-dimensional data encoding device 1300 may determine whether rotation and translation are applicable to each reference space when encoding with reference to a plurality of reference spaces at different times. In this case, the three-dimensional data encoding device 1300 may attach information (RT applicable flag, etc.) indicating whether rotation and translation are applicable to each reference space to header information of the bitstream, etc. For example, the three-dimensional data encoding device 1300 calculates RT information and ICP error value using an ICP (Interactive Closest Point) algorithm according to the encoding object space and each reference space to be referenced. When the ICP error value is below a predetermined certain value, the three-dimensional data encoding device 1300 determines that rotation and translation are not required and sets the RT applicable flag to OFF (invalid). In addition, when the ICP error value is larger than the above-mentioned certain value, the three-dimensional data encoding device 1300 sets the RT applicable flag to ON (valid) and attaches the RT information to the bitstream.
[0494] Fig.46 An example of a syntax in which RT information and an RT applicable flag are attached to a header is shown. In addition, the number of bits allocated to each syntax can be determined based on the range that the syntax can take. For example, when the number of reference spaces included in the reference list L0 is 8, 3 bits can be allocated in MaxRefSpc_l0. The number of allocated bits can be changed according to the values that each syntax can take, or the number of allocated bits can be fixed without being affected by the possible values. In the case of fixing the number of allocated bits, the three-dimensional data encoding device 1300 can attach the fixed number of bits to other header information.
[0495] Here, Fig.46MaxRefSpc_10 shown shows the number of reference spaces included in reference list L0. RT_flag_10[i] is the RT application flag for reference space i in reference list L0. When RT_flag_10[i] is 1, rotation and translation are applied to reference space i. When RT_flag_10[i] is 0, rotation and translation are not applied to reference space i.
[0496] R_l0[i] and T_l0[i] are RT information of reference space i in reference list L0. R_l0[i] is rotation information of reference space i in reference list L0. The rotation information indicates the content of the applied rotation processing, such as a rotation matrix or quaternion. T_l0[i] is translation information of reference space i in reference list L0. The translation information indicates the content of the applied translation processing, such as a translation vector.
[0497] MaxRefSpc_l1 indicates the number of reference spaces included in reference list L1. RT_flag_l1[i] is an RT application flag for reference space i in reference list L1. When RT_flag_l1[i] is 1, rotation and translation are applied to reference space i. When RT_flag_l1[i] is 0, rotation and translation are not applied to reference space i.
[0498] R_l1[i] and T_l1[i] are RT information of reference space i in reference list L1. R_l1[i] is rotation information of reference space i in reference list L1. The rotation information indicates the content of the applied rotation processing, such as a rotation matrix or quaternion. T_l1[i] is translation information of reference space i in reference list L1. The translation information indicates the content of the applied translation processing, such as a translation vector.
[0499] The inter-frame prediction unit 1311 generates a predicted volume of the encoding target volume using information of the encoded reference space stored in the reference space memory 1310. As described above, before generating the predicted volume of the encoding target volume, the inter-frame prediction unit 1311 uses the ICP (Interactive Closest Point) algorithm to obtain RT information in the encoding target space and the reference space in order to make the positional relationship between the encoding target space and the entire reference space close. Then, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space using the obtained RT information, thereby obtaining the reference space B. Thereafter, the inter-frame prediction unit 1311 generates a predicted volume of the encoding target volume in the encoding target space using information in the reference space B. Here, the three-dimensional data encoding device 1300 adds the RT information used to obtain the reference space B to the header information of the encoding target space, etc.
[0500] In this way, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space, so as to make the overall positional relationship between the encoding object space and the reference space close, and then uses the information of the reference space to generate a prediction volume, thereby improving the accuracy of the prediction volume. In addition, since the prediction residual can be suppressed, the amount of encoding can be reduced. In addition, although an example of performing ICP using the encoding object space and the reference space is shown here, it is not limited to this. For example, in order to reduce the amount of processing, the inter-frame prediction unit 1311 can also perform ICP using at least one of the encoding object space from which the number of voxels or point clouds is extracted, and the reference space from which the number of voxels or point clouds is extracted, so as to obtain RT information.
[0501] Furthermore, when the ICP error value obtained from the ICP result is smaller than a predetermined first threshold value, that is, when the positional relationship between the encoding object space and the reference space is close, the inter-frame prediction unit 1311 may determine that rotation and translation processing are not required, and does not perform rotation and translation. In this case, the three-dimensional data encoding device 1300 may not add RT information to the bitstream, thereby suppressing additional overhead.
[0502] Furthermore, when the ICP error value is greater than a predetermined second threshold, the inter-frame prediction unit 1311 determines that the shape change in space is large, and intra-frame prediction can be applied to all volumes of the encoding object space. Hereinafter, the space to which intra-frame prediction is applied is referred to as intra-frame space. Furthermore, the second threshold is a value greater than the above-mentioned first threshold. Furthermore, it is not limited to ICP, and any method can be applied as long as the method of obtaining RT information from two voxel sets or two point cloud sets.
[0503] Furthermore, when the three-dimensional data contains attribute information such as shape or color, the inter-frame prediction unit 1311 searches, for example, a volume in the reference space that is closest to the shape or color attribute information of the encoding target volume as a prediction volume of the encoding target volume in the encoding target space. Furthermore, the reference space is, for example, a reference space after the above-mentioned rotation and translation processing. The inter-frame prediction unit 1311 generates a prediction volume based on the volume (reference volume) obtained by the search. Fig.47 is a diagram for explaining the generation of the prediction volume. Fig.47When the encoding target volume (volume idx=0) shown in the figure is encoded by using inter-frame prediction, the reference volumes in the reference space are scanned in sequence while searching for a volume in which the difference between the encoding target volume and the reference volume, i.e., the prediction residual, is the smallest. The inter-frame prediction unit 1311 selects the volume with the smallest prediction residual as the prediction volume. The prediction residual between the encoding target volume and the prediction volume is encoded by the processing after the transformation unit 1303. Here, the prediction residual refers to the difference between the attribute information of the encoding target volume and the attribute information of the prediction volume. In addition, the three-dimensional data encoding device 1300 adds the volume idx of the reference volume in the reference space referred to as the prediction volume to the header of the bit stream, etc.
[0504] exist Fig.47 In the example shown, the reference volume idx=4 of the reference space L0R0 is selected as the prediction volume of the encoding target volume. Then, the prediction residual between the encoding target volume and the reference volume and the reference volume idx=4 are encoded and added to the bit stream.
[0505] In addition, although the description here is made by taking the generation of the predicted volume of the attribute information as an example, the same processing can be performed on the predicted volume of the position information.
[0506] The prediction control unit 1312 controls whether to use intra-frame prediction or inter-frame prediction to encode the encoding target volume. Here, a mode including intra-frame prediction and inter-frame prediction is referred to as a prediction mode. For example, the prediction control unit 1312 calculates the prediction residual when the encoding target volume is predicted by intra-frame prediction and the prediction residual when it is predicted by inter-frame prediction as an evaluation value, and selects the prediction mode with the smaller evaluation value. Alternatively, the prediction control unit 1312 may apply orthogonal transformation, quantization, and entropy coding to the prediction residual of intra-frame prediction and the prediction residual of inter-frame prediction, respectively, to calculate the actual amount of code, and select the prediction mode using the calculated amount of code as the evaluation value. Furthermore, overhead information other than the prediction residual (reference volume idx information, etc.) may be added to the evaluation value. Furthermore, the prediction control unit 1312 may also generally select intra-frame prediction when the encoding target space is predetermined to be encoded in the intra-frame space.
[0507] The entropy coding unit 1313 generates a coded signal (coded bit stream) by performing variable length coding on the quantized coefficients input from the quantization unit 1304. Specifically, the entropy coding unit 1313 binarizes the quantized coefficients and performs arithmetic coding on the obtained binary signal, for example.
[0508] Next, a three-dimensional data decoding device that decodes the encoded signal generated by the three-dimensional data encoding device 1300 will be described. Fig.4814 is a block diagram of a three-dimensional data decoding device 1400 according to the present embodiment. The three-dimensional data decoding device 1400 includes an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transformation unit 1403, an addition unit 1404, a reference volume memory 1405, an intra-frame prediction unit 1406, a reference space memory 1407, an inter-frame prediction unit 1408, and a prediction control unit 1409.
[0509] The entropy decoding unit 1401 performs variable length decoding on the coded signal (coded bit stream). For example, the entropy decoding unit 1401 performs arithmetic decoding on the coded signal to generate a binary signal, and generates a quantized coefficient based on the generated binary signal.
[0510] The inverse quantization unit 1402 performs inverse quantization on the quantized coefficients input from the entropy decoding unit 1401 using a quantization parameter added to a bit stream or the like, thereby generating inverse quantized coefficients.
[0511] The inverse transform unit 1403 generates a prediction residual by performing an inverse transform on the inverse quantized coefficient input from the inverse quantization unit 1402. For example, the inverse transform unit 1403 generates a prediction residual by performing an inverse orthogonal transform on the inverse quantized coefficient based on information added to the bit stream.
[0512] The adder 1404 adds the prediction residual generated by the inverse transform unit 1403 to the prediction volume generated by intra prediction or inter prediction to generate a reconstructed volume. The reconstructed volume is output as decoded three-dimensional data and stored in the reference volume memory 1405 or the reference space memory 1407.
[0513] The intra prediction unit 1406 generates a prediction volume by intra prediction using the reference volume in the reference volume memory 1405 and the information added to the bitstream. Specifically, the intra prediction unit 1406 obtains the prediction mode information and the adjacent volume information (e.g., volume idx) added to the bitstream, and generates a prediction volume using the adjacent volume indicated by the adjacent volume information in the mode indicated by the prediction mode information. In addition, the details of these processes are the same as the processes of the intra prediction unit 1309 described above, except that the information added to the bitstream is used.
[0514] The inter-frame prediction unit 1408 generates a prediction volume through inter-frame prediction using the reference space in the reference space memory 1407 and the information attached to the bitstream. Specifically, the inter-frame prediction unit 1408 uses the RT information of each reference space attached to the bitstream, applies rotation and translation processing to the reference space, and generates a prediction volume using the reference space after application. In addition, when the RT application flag of each reference space exists in the bitstream, the inter-frame prediction unit 1408 applies rotation and translation processing to the reference space according to the RT application flag. In addition, the details of the above-mentioned processing are the same as the processing of the above-mentioned inter-frame prediction unit 1311, except that the information attached to the bitstream is used.
[0515] Whether to decode the decoding target volume by intra prediction or inter prediction is controlled by the prediction control unit 1409. For example, the prediction control unit 1409 selects intra prediction or inter prediction according to information indicating the prediction mode to be used, which is added to the bit stream. In addition, when it is predetermined that the decoding target space is decoded as the intra space, the prediction control unit 1409 may normally select the intra prediction.
[0516] The following is a description of a variation of the present embodiment. In the present embodiment, although the application of rotation and translation in spatial units is described as an example, rotation and translation in finer units may also be applied. For example, the three-dimensional data encoding device 1300 may divide the space into subspaces and apply rotation and translation in subspace units. In this case, the three-dimensional data encoding device 1300 generates RT information according to each subspace, and attaches the generated RT information to the header of the bitstream, etc. In addition, the three-dimensional data encoding device 1300 may use volume units as encoding units to apply rotation and translation. In this case, the three-dimensional data encoding device 1300 generates RT information in encoding volume units, and attaches the generated RT information to the header of the bitstream, etc. In addition, the above may be combined. That is, the three-dimensional data encoding device 1300 may apply rotation and translation in finer units after applying rotation and translation in large units. For example, the three-dimensional data encoding device 1300 may apply rotation and translation in spatial units, and apply different rotations and translations to each of the multiple volumes contained in the obtained space.
[0517] Furthermore, although the present embodiment is described by taking the application of rotation and translation to the reference space as an example, the present invention is not limited thereto. For example, the three-dimensional data encoding device 1300 may apply scaling processing to change the size of the three-dimensional data. Furthermore, the three-dimensional data encoding device 1300 may also apply any one or two of rotation, translation, and scaling. Furthermore, as described above, when the processing is applied in different units in multiple stages, the type of processing applied in each unit may be different. For example, rotation and translation may be applied in the spatial unit, and translation may be applied in the volume unit.
[0518] Note that these modifications are also applicable to the three-dimensional data decoding device 1400 .
[0519] As described above, the three-dimensional data encoding device 1300 according to this embodiment performs the following processing. Fig.48 3D data encoding apparatus 1300 performs an inter-frame prediction process.
[0520] First, the three-dimensional data encoding device 1300 generates predicted position information (e.g., predicted volume) using the position information of three-dimensional points included in the target three-dimensional data (e.g., encoding target space) and the reference three-dimensional data (e.g., reference space) at different times (S1301). Specifically, the three-dimensional data encoding device 1300 generates the predicted position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.
[0521] In addition, the three-dimensional data encoding device 1300 performs rotation and translation processing with a first unit (e.g., space), and generates predicted position information with a second unit (e.g., volume) that is smaller than the first unit. For example, the three-dimensional data encoding device 1300 searches for a volume in which the difference between the encoding object volume and the position information contained in the encoding object space is the smallest from among a plurality of volumes contained in the reference space after the rotation and translation processing, and uses the obtained volume as the predicted volume. In addition, the three-dimensional data encoding device 1300 may perform the rotation and translation processing and the generation of the predicted position information with the same unit.
[0522] Furthermore, the three-dimensional data encoding device 1300 may apply a first rotation and translation process to the position information of the three-dimensional points contained in the reference three-dimensional data using a first unit (e.g., space), and may apply a second rotation and translation process to the position information of the three-dimensional points obtained by the first rotation and translation process using a second unit (e.g., volume) that is smaller than the first unit, thereby generating predicted position information.
[0523] Here, the position information of the three-dimensional point and the predicted position information are as follows: Fig.41As shown, the octree structure is used for representation. For example, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the width among the depth and width in the octree structure. Alternatively, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the depth and width among the depth and width in the octree structure.
[0524] And, if Fig.46 As shown, the three-dimensional data encoding device 1300 encodes the RT applicable flag indicating whether the rotation and translation processing is applied to the position information of the three-dimensional point contained in the reference three-dimensional data. That is, the three-dimensional data encoding device 1300 generates a coded signal (coded bit stream) including the RT applicable flag. In addition, the three-dimensional data encoding device 1300 encodes the RT information indicating the content of the rotation and translation processing. That is, the three-dimensional data encoding device 1300 generates a coded signal (coded bit stream) including the RT information. Alternatively, the three-dimensional data encoding device 1300 may encode the RT information when the RT applicable flag indicates that the rotation and translation processing is applicable, and may not encode the RT information when the RT applicable flag indicates that the rotation and translation processing is not applicable.
[0525] The three-dimensional data includes, for example, position information of three-dimensional points and attribute information (color information, etc.) of each three-dimensional point. The three-dimensional data encoding device 1300 generates predicted attribute information by using the attribute information of the three-dimensional points included in the reference three-dimensional data (S1302).
[0526] Next, the three-dimensional data encoding device 1300 uses the predicted position information to encode the position information of the three-dimensional points included in the object three-dimensional data. Fig.38 As shown, differential position information which is a difference between the position information of the three-dimensional point included in the target three-dimensional data and the predicted position information is calculated (S1303).
[0527] Furthermore, the three-dimensional data encoding device 1300 uses the predicted attribute information to encode the attribute information of the three-dimensional points included in the target three-dimensional data. For example, the three-dimensional data encoding device 1300 calculates the difference between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information, that is, differential attribute information (S1304). Next, the three-dimensional data encoding device 1300 transforms and quantizes the calculated differential attribute information (S1305).
[0528] Finally, the 3D data encoding device 1300 encodes (for example, entropy encoding) the differential position information and the quantized differential attribute information (S1306). That is, the 3D data encoding device 1300 generates an encoded signal (encoded bit stream) including the differential position information and the differential attribute information.
[0529] In addition, when the three-dimensional data does not include attribute information, the three-dimensional data encoding device 1300 may not perform steps S1302, S1304, and S1305. In addition, the three-dimensional data encoding device 1300 may only encode the position information of the three-dimensional point or encode the attribute information of the three-dimensional point.
[0530] and, Fig.49 The processing sequence shown is only an example and is not limited thereto. For example, since the processing of location information (S1301, S1303) and the processing of attribute information (S1302, S1304, S1305) are independent of each other, they can be executed in any order, or some of them can be processed in parallel.
[0531] As described above, in this embodiment, the three-dimensional data encoding device 1300 generates predicted position information using the position information of the three-dimensional points included in the target three-dimensional data and the reference three-dimensional data at different times, and encodes the difference between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information, that is, the differential position information. Accordingly, since the amount of data of the encoded signal can be reduced, the encoding efficiency can be improved.
[0532] Furthermore, in this embodiment, the three-dimensional data encoding device 1300 generates predicted attribute information by using attribute information of three-dimensional points included in the reference three-dimensional data, and encodes the difference between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information, that is, the differential attribute information. Accordingly, since the data amount of the encoded signal can be reduced, the encoding efficiency can be improved.
[0533] For example, the three-dimensional data encoding device 1300 includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0534] Fig.48 3D data decoding apparatus 1400 performs an inter-frame prediction process.
[0535] First, the three-dimensional data decoding apparatus 1400 decodes (eg, performs entropy decoding) the difference position information and the difference attribute information according to the coded signal (coded bit stream) ( S1401 ).
[0536] Furthermore, the three-dimensional data decoding device 1400 decodes the RT application flag indicating whether the rotation and translation processing is applied to the position information of the three-dimensional point included in the reference three-dimensional data based on the coded signal. Furthermore, the three-dimensional data decoding device 1400 decodes the RT information indicating the content of the rotation and translation processing. In addition, the three-dimensional data decoding device 1400 decodes the RT information when the RT application flag indicates that the rotation and translation processing is applied, and does not decode the RT information when the RT application flag indicates that the rotation and translation processing is not applied.
[0537] Next, the three-dimensional data decoding apparatus 1400 performs inverse quantization and inverse transformation on the decoded differential attribute information ( S1402 ).
[0538] Next, the three-dimensional data decoding device 1400 generates predicted position information (e.g., predicted volume) using the position information of the three-dimensional points included in the target three-dimensional data (e.g., decoding target space) and the reference three-dimensional data (e.g., reference space) at different times (S1403). Specifically, the three-dimensional data decoding device 1400 generates the predicted position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.
[0539] More specifically, when the RT application flag indicates that the rotation and translation processing is applied, the three-dimensional data decoding device 1400 applies the rotation and translation processing to the position information of the three-dimensional point included in the reference three-dimensional data indicated by the RT information. Also, when the RT application flag indicates that the rotation and translation processing is not applied, the three-dimensional data decoding device 1400 does not apply the rotation and translation processing to the position information of the three-dimensional point included in the reference three-dimensional data.
[0540] In addition, the three-dimensional data decoding device 1400 may perform the rotation and translation processing in a first unit (e.g., space), and may generate the predicted position information in a second unit (e.g., volume) that is smaller than the first unit. In addition, the three-dimensional data decoding device 1400 may also perform the rotation and translation processing, and the generation of the predicted position information in the same unit.
[0541] Furthermore, the three-dimensional data decoding device 1400 may apply a first rotation and translation process with a first unit (e.g., space) to the position information of the three-dimensional points included in the reference three-dimensional data, and may apply a second rotation and translation process with a second unit (e.g., volume) that is smaller than the first unit to the position information of the three-dimensional points obtained by the first rotation and translation process, thereby generating predicted position information.
[0542] Here, the position information of the three-dimensional point and the predicted position information are as follows: Fig.41As shown, the octree structure is used for representation. For example, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the width among the depth and width in the octree structure. Alternatively, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the depth among the depth and width in the octree structure.
[0543] The three-dimensional data decoding apparatus 1400 generates predicted attribute information by using the attribute information of the three-dimensional points included in the reference three-dimensional data ( S1404 ).
[0544] Next, the three-dimensional data decoding device 1400 decodes the coded position information contained in the coded signal by using the predicted position information, thereby restoring the position information of the three-dimensional point contained in the target three-dimensional data. Here, the coded position information is, for example, differential position information, and the three-dimensional data decoding device 1400 restores the position information of the three-dimensional point contained in the target three-dimensional data by adding the differential position information to the predicted position information (S1405).
[0545] Furthermore, the three-dimensional data decoding device 1400 decodes the coded attribute information included in the coded signal by using the predicted attribute information, thereby restoring the attribute information of the three-dimensional point included in the target three-dimensional data. Here, the coded attribute information is, for example, differential attribute information, and the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional point included in the target three-dimensional data by adding the differential attribute information to the predicted attribute information (S1406).
[0546] Alternatively, if the three-dimensional data does not include attribute information, the three-dimensional data decoding device 1400 may not perform steps S1402, S1404, and S1406. Furthermore, the three-dimensional data decoding device 1400 may only decode the position information of the three-dimensional point or decode the attribute information of the three-dimensional point.
[0547] and, Fig.50 The order of processing shown is an example and is not limited to this. For example, since the processing of location information (S1403, S1405) and the processing of attribute information (S1402, S1404, S1406) are independent of each other, they can be performed in any order, and some of them can be processed in parallel.
[0548] (Implementation 8)
[0549] A method of controlling reference during encoding of the occupancy rate encoding in this embodiment will be described. In addition, the following mainly describes the operation of the three-dimensional data encoding device, but the same processing can be performed in the three-dimensional data decoding device.
[0550] Fig.51 as well as Fig.52 is a diagram showing a reference relationship involved in this embodiment, Fig.51 It is a graph that represents the reference relationship on the octree structure. Fig.52 It is a graph that represents reference relationships in a spatial area.
[0551] In this embodiment, when the three-dimensional data encoding device encodes the encoding information of the node of the encoding object (hereinafter referred to as the object node), the encoding information of each node in the parent node to which the object node belongs is referenced. However, the encoding information of each node in other nodes (hereinafter referred to as parent adjacent nodes) in the same layer as the parent node is not referenced. In other words, the three-dimensional data encoding device sets the parent adjacent node to be unable to be referenced, or prohibits reference.
[0552] In addition, the three-dimensional data encoding device may also allow reference to the encoding information in the parent node (hereinafter referred to as the grandparent node) to which the parent node belongs. That is, the three-dimensional data encoding device may also refer to the encoding information of the parent node and the grandparent node to which the object node belongs, and encode the encoding information of the object node.
[0553] Here, the coding information is, for example, an occupancy code. When encoding the occupancy code of the object node, the three-dimensional data coding device refers to information indicating whether each node in the parent node to which the object node belongs contains a point group (hereinafter referred to as occupancy information). In other words, when encoding the occupancy code of the object node, the three-dimensional data coding device refers to the occupancy code of the parent node. On the other hand, the three-dimensional data coding device does not refer to the occupancy information of each node in the parent adjacent node. That is, the three-dimensional data coding device does not refer to the occupancy code of the parent adjacent node. In addition, the three-dimensional data coding device may also refer to the occupancy information of each node in the grandparent node. That is, the three-dimensional data coding device may also refer to the occupancy information of the parent node and the parent adjacent node.
[0554] For example, when encoding the occupancy code of the object node, the three-dimensional data encoding device uses the occupancy code of the parent node or the grandparent node to which the object node belongs, and switches the coding table used when entropy coding the occupancy code of the object node. In addition, the details are described later. At this time, the three-dimensional data encoding device may not refer to the occupancy code of the parent adjacent node. Thus, when encoding the occupancy code of the object node, the three-dimensional data encoding device can appropriately switch the coding table according to the information of the occupancy code of the parent node or the grandparent node, thereby improving the coding efficiency. In addition, the three-dimensional data encoding device does not refer to the parent adjacent node, thereby suppressing the confirmation processing of the information of the parent adjacent node and the memory capacity used to store the processing. In addition, it becomes easy to scan the occupancy code of each node of the octree in depth-first order and encode it.
[0555] Hereinafter, an example of switching the coding table using the occupancy rate coding of the parent node will be described. Fig.53 is a diagram showing an example of a target node and adjacent reference nodes. Fig.54 It is a graph representing the relationship between parent nodes and nodes. Fig.55 is a diagram showing an example of encoding the occupancy rate of a parent node. Here, the adjacent reference node refers to a node that is spatially adjacent to the target node and is referenced when encoding the target node. Fig.53 In the example shown, the adjacent nodes are nodes belonging to the same layer as the target node. In addition, as reference adjacent nodes, the node X adjacent to the target block in the x direction, the node Y adjacent to the y direction, and the node Z adjacent to the z direction are used. That is, one adjacent block is set as the reference adjacent block in each of the x, y, and z directions.
[0556] also, Fig.54 The node numbers shown are examples, and the relationship between the node numbers and the positions of the nodes is not limited to this. Fig.55 In , the node 0 is allocated to the lower bit and the node 7 is allocated to the upper bit, but the allocation may be in the reverse order. In addition, each node may be allocated to any bit.
[0557] The three-dimensional data encoding device determines a coding table when entropy coding is performed on the occupancy coding of the target node, for example, by the following equation.
[0558] CodingTable=(FlagX<<2)+(FlagY<<1)+(FlagZ)
[0559] Here, CodingTable represents a coding table for encoding the occupancy of the object node, and represents any value from 0 to 7. FlagX is the occupancy information of the adjacent node X, and represents 1 if the adjacent node X contains (occupies) the point group, and represents 0 if not. FlagY is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Y contains (occupies) the point group, and represents 0 if not. FlagZ is the occupancy information of the adjacent node Z, and represents 1 if the adjacent node Z contains (occupies) the point group, and represents 0 if not.
[0560] Furthermore, since information indicating whether an adjacent node is occupied is included in the occupancy code of the parent node, the three-dimensional data encoding device may select the encoding table using the value indicated in the occupancy code of the parent node.
[0561] As can be seen from the above, the three-dimensional data encoding device can improve the encoding efficiency by switching the encoding table using information indicating whether the neighboring nodes of the target node include a point group.
[0562] In addition, if Fig.53 As shown, the three-dimensional data encoding device can also switch the adjacent reference node according to the spatial position of the object node in the parent node. That is, the three-dimensional data encoding device can also switch the adjacent node for reference among multiple adjacent nodes according to the spatial position of the parent node of the object node.
[0563] Next, configuration examples of a three-dimensional data encoding device and a three-dimensional data decoding device will be described. Fig.56 This is a block diagram of a three-dimensional data encoding device 2100 according to this embodiment. Fig.56 The three-dimensional data encoding device 2100 shown includes an octree generation unit 2101 , a geometric information calculation unit 2102 , a coding table selection unit 2103 , and an entropy encoding unit 2104 .
[0564] The octree generation unit 2101 generates, for example, an octree based on the input three-dimensional points (point cloud), and generates an occupancy code for each node included in the octree. The geometry information calculation unit 2102 obtains occupancy information indicating whether the adjacent reference node of the object node is occupied. For example, the geometry information calculation unit 2102 obtains the occupancy information of the adjacent reference node from the occupancy code of the parent node to which the object node belongs. In addition, Fig.53 As shown, the geometric information calculation unit 2102 may switch the adjacent reference node according to the position of the target node in the parent node. In addition, the geometric information calculation unit 2102 does not refer to the occupancy information of each node in the parent adjacent node.
[0565] The coding table selection unit 2103 selects a coding table used in entropy coding of the occupancy coding of the target node using the occupancy information of the adjacent reference nodes calculated by the geometric information calculation unit 2102. The entropy coding unit 2104 generates a bit stream by entropy coding the occupancy coding using the selected coding table. In addition, the entropy coding unit 2104 may also attach information indicating the selected coding table to the bit stream.
[0566] Fig.57 This is a block diagram of a three-dimensional data decoding device 2110 according to this embodiment. Fig.57 The three-dimensional data decoding device 2110 shown includes an octree generation unit 2111 , a geometric information calculation unit 2112 , a coding table selection unit 2113 , and an entropy decoding unit 2114 .
[0567] The octree generation unit 2111 generates an octree of a certain space (node) using the header information of the bitstream, etc. The octree generation unit 2111 generates a large space (root node) using the size of the x-axis, y-axis, and z-axis directions of the certain space added to the header information, and generates an octree by dividing the space into two in the x-axis, y-axis, and z-axis directions, respectively, to generate eight small spaces A (nodes A0 to A7). In addition, nodes A0 to A7 are sequentially set as object nodes.
[0568] The geometry information calculation unit 2112 obtains the occupancy information indicating whether the adjacent reference node of the object node is occupied. For example, the geometry information calculation unit 2112 obtains the occupancy information of the adjacent reference node from the occupancy rate code of the parent node to which the object node belongs. Fig.53 As shown, the geometric information calculation unit 2112 may switch the adjacent reference node according to the position of the target node in the parent node. In addition, the geometric information calculation unit 2112 does not refer to the occupancy information of each node in the parent adjacent node.
[0569] The coding table selection unit 2113 selects a coding table (decoding table) used for entropy decoding of the occupancy code of the target node using the occupancy information of the adjacent reference nodes calculated by the geometric information calculation unit 2112. The entropy decoding unit 2114 generates a three-dimensional point by entropy decoding the occupancy code using the selected coding table. In addition, the coding table selection unit 2113 decodes and obtains the information of the selected coding table attached to the bit stream, and the entropy decoding unit 2114 may also use the coding table indicated by the obtained information.
[0570] Each bit of the occupancy code (8 bits) included in the bit stream indicates whether a point group is included in each of the eight small spaces A (nodes A0 to A7). Furthermore, the three-dimensional data decoding device divides the small space node A0 into eight small spaces B (nodes B0 to B7) and generates an octree, decodes the occupancy code, and obtains information indicating whether each node of the small space B contains a point group. In this way, the three-dimensional data decoding device decodes the occupancy code of each node while generating an octree from a large space to a small space.
[0571] The following describes the flow of processing by the three-dimensional data encoding device and the three-dimensional data decoding device. Fig.58 3D data encoding processing in a 3D data encoding device. First, the 3D data encoding device determines (defines) a space (object node) containing part or all of the input 3D point group (S2101). Next, the 3D data encoding device divides the object node 8 into 8 small spaces (nodes) (S2102). Next, the 3D data encoding device generates an occupancy code of the object node according to whether each node contains a point group (S2103).
[0572] Next, the three-dimensional data encoding device calculates (obtains) the occupancy information of the adjacent reference nodes of the object node from the occupancy code of the parent node of the object node (S2104). Next, the three-dimensional data encoding device selects a coding table used in entropy coding based on the occupancy information of the adjacent reference nodes of the determined object node (S2105). Next, the three-dimensional data encoding device performs entropy coding on the occupancy code of the object node using the selected coding table (S2106).
[0573] Furthermore, the three-dimensional data encoding device repeatedly divides each node 8 and encodes the occupancy rate code of each node until the node cannot be divided (S2107). That is, the processing of steps S2102 to S2106 is recursively repeated.
[0574] Fig.59 This is a flowchart of a three-dimensional data decoding method in a three-dimensional data decoding device. First, the three-dimensional data decoding device uses the header information of the bit stream to determine (define) the space (object node) to be decoded (S2111). Next, the three-dimensional data decoding device divides the object node 8 to generate 8 small spaces (nodes) (S2112). Next, the three-dimensional data decoding device calculates (obtains) the occupancy information of the adjacent reference nodes of the object node from the occupancy code of the parent node of the object node (S2113).
[0575] Next, the three-dimensional data decoding apparatus selects a coding table used for entropy decoding based on the occupancy information of the adjacent reference nodes (S2114). Next, the three-dimensional data decoding apparatus entropy decodes the occupancy code of the target node using the selected coding table (S2115).
[0576] Furthermore, the three-dimensional data decoding device repeatedly divides each node 8 and decodes the occupancy code of each node until the node cannot be divided (S2116). That is, the processing of steps S2112 to S2115 is recursively repeated.
[0577] Next, an example of switching the coding table will be described. Fig.60 is a diagram showing an example of switching of the coding table. Fig.60 As shown in the coding table 0, the same context model can be applied to multiple occupancy codes. In addition, each occupancy code can also be assigned a different context model. Thus, the context model can be assigned according to the probability of occurrence of the occupancy code, thereby improving the coding efficiency. In addition, a context model that updates the probability table according to the frequency of occurrence of the occupancy code can also be used. In addition, a context model that fixes the probability table can also be used.
[0578] Hereinafter, Modification 1 of the present embodiment will be described. Fig.61In the above embodiment, the three-dimensional data encoding device does not refer to the occupancy rate encoding of the parent adjacent node, but it is also possible to switch whether to refer to the occupancy rate encoding of the parent adjacent node according to a specific condition.
[0579] For example, when the three-dimensional data encoding device encodes the occupancy rate of the object node by referring to the occupancy information of the node in the parent adjacent node while encoding the octree with a width-first scan. On the other hand, when the three-dimensional data encoding device encodes the occupancy rate of the object node by referring to the occupancy information of the node in the parent adjacent node while encoding the octree with a depth-first scan. In this way, according to the scanning order (encoding order) of the nodes of the octree, the nodes that can be referenced are appropriately switched, so that the encoding efficiency can be improved and the processing load can be suppressed.
[0580] Furthermore, the three-dimensional data encoding device may add information such as whether the octree is encoded with width first or depth first to the header of the bit stream. Fig.62 This is a diagram showing an example of the syntax of header information in this case. Fig.62 The octree_scan_order shown is encoding order information (encoding order flag) indicating the encoding order of the octree. For example, when octree_scan_order is 0, it indicates width priority, and when it is 1, it indicates depth priority. Thus, by referring to octree_scan_order, the three-dimensional data decoding device can know whether the bit stream is encoded in width priority or depth priority, and can appropriately decode the bit stream.
[0581] Furthermore, the three-dimensional data encoding device may add information indicating whether or not to prohibit reference to a parent adjacent node to header information of the bit stream. Fig.63 The figure shows a syntax example of the header information in this case. limit_refer_flag is the prohibition switching information (prohibition switching flag) indicating whether to prohibit the reference to the parent adjacent node. For example, when limit_refer_flag is 1, it prohibits the reference to the parent adjacent node, and when it is 0, it indicates that there is no reference restriction (the reference to the parent adjacent node is permitted).
[0582] That is, the three-dimensional data encoding device determines whether to prohibit reference to the parent adjacent node, and switches whether to prohibit or allow reference to the parent adjacent node based on the result of the above determination. In addition, the three-dimensional data encoding device generates a bit stream including prohibition switching information, the prohibition switching information is the result of the above determination, and indicates whether to prohibit reference to the parent adjacent node.
[0583] Furthermore, the three-dimensional data decoding device obtains prohibition switching information indicating whether to prohibit referring to the parent adjacent node from the bit stream, and switches whether to prohibit or permit referring to the parent adjacent node based on the prohibition switching information.
[0584] Thus, the three-dimensional data encoding device can control the reference of the parent adjacent node and generate a bit stream. In addition, the three-dimensional data decoding device can obtain information indicating whether to prohibit the reference of the parent adjacent node from the header of the bit stream.
[0585] In addition, in this embodiment, as an example of prohibiting the coding process of referring to the parent adjacent node, the coding process of occupancy coding is recorded as an example, but it is not necessarily limited to this. For example, the same method can also be applied when encoding other information of the node of the octree. For example, when encoding other attribute information such as color, normal vector, or reflectivity attached to the node, the method of this embodiment can also be applied. In addition, the same method can also be applied when encoding the coding table or the predicted value.
[0586] Next, a second variation of the present embodiment will be described. Fig.53 Although an example using three reference adjacent nodes is shown, four or more reference adjacent nodes may be used. Fig.64 This is a diagram showing an example of a target node and reference adjacent nodes.
[0587] For example, the three-dimensional data encoding device calculates the Fig.64 The occupancy coding of the object node shown is a coding table when entropy coding is performed.
[0588] CodingTable=(FlagX0<<3)+(FlagX1<<2)+(FlagY<<1)+(FlagZ)
[0589] Here, CodingTable represents a coding table for encoding the occupancy of the object node, and represents any value from 0 to 15. FlagXN is the occupancy information of the adjacent node XN (N=0…1), and represents 1 if the adjacent node XN contains (occupies) a point group, and represents 0 if not. FlagY is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Y contains (occupies) a point group, and represents 0 if not. FlagZ is the occupancy information of the adjacent node Z, and represents 1 if the adjacent node Z contains (occupies) a point group, and represents 0 if not.
[0590] At this time, if the adjacent node is, for example Fig.64 In the case where the adjacent node X0 is not referenceable (prohibited from reference), the three-dimensional data encoding device may use a fixed value such as 1 (occupied) or 0 (non-occupied) as a substitute value.
[0591] Fig.65 is a graph showing examples of object nodes and adjacent nodes. Fig.65As shown, when it is impossible to refer to (prohibited to refer to) the adjacent nodes, the occupancy rate code of the grandparent node of the object node can also be referred to to calculate the occupancy information of the adjacent nodes. Fig.65 The three-dimensional data encoding device can also use the occupancy information of the adjacent node G0 to calculate the FlagX0 of the above formula, and use the calculated FlagX0 to determine the value of the encoding table. Fig.65 The neighboring node G0 shown is a neighboring node that can be determined by the occupancy rate coding of the grandparent node. The neighboring node X1 is a neighboring node that can be determined by the occupancy rate coding of the parent node.
[0592] Hereinafter, Modification 3 of the present embodiment will be described. Fig.66 as well as Fig.67 is a diagram showing a reference relationship involved in this modification example, Fig.66 It is a graph that represents the reference relationship on the octree structure. Fig.67 It is a graph that represents reference relationships in a spatial area.
[0593] In this variant, when encoding the encoding information of the node of the encoding object (hereinafter referred to as the object node 2), the three-dimensional data encoding device refers to the encoding information of each node in the parent node to which the object node 2 belongs. That is, the three-dimensional data encoding device allows reference to the information (for example, occupancy information) of the child nodes of the first node whose parent node is the same as the parent node of the object node among the plurality of adjacent nodes. Fig.66 When encoding the occupancy rate code of the object node 2 shown in FIG. 1 , reference is made to the nodes existing in the parent node to which the object node 2 belongs, for example, Fig.66 The occupancy rate of the object node is shown in . Fig.67 As shown, Fig.66 The occupancy code of the object node shown indicates whether each node in the object node adjacent to the object node 2 is occupied. Therefore, the three-dimensional data encoding device can switch the encoding table of the occupancy code of the object node 2 according to the finer shape of the object node, thereby improving the encoding efficiency.
[0594] The three-dimensional data encoding device may calculate a coding table for entropy coding the occupancy coding of the target node 2 by, for example, the following equation.
[0595] CodingTable=(FlagX1<<5)+(FlagX2<<4)+(FlagX3<<3)+(FlagX4<<2)+(FlagY<<1)+(FlagZ)
[0596] Here, CodingTable represents a coding table for encoding the occupancy of the object node 2, and represents any value from 0 to 63. FlagXN is the occupancy information of the adjacent node XN (N=1…4), and represents 1 if the adjacent node XN contains (occupies) a point group, and represents 0 if not. FlagY is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Y contains (occupies) a point group, and represents 0 if not. FlagZ is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Z contains (occupies) a point group, and represents 0 if not.
[0597] Furthermore, the three-dimensional data encoding device may change the method of calculating the encoding table according to the node position of the object node 2 within the parent node.
[0598] In addition, if it is not prohibited to refer to the parent adjacent node, the three-dimensional data encoding device can refer to the encoding information of each node in the parent adjacent node. For example, if it is not prohibited to refer to the parent adjacent node, it is allowed to refer to the information (such as occupancy information) of the child node of the third node whose parent node is different from the parent node of the object node. Fig.65 In the example shown, the three-dimensional data encoding device refers to the occupancy code of the adjacent node X0 whose parent node is different from the parent node of the object node to obtain the occupancy information of the child node of the adjacent node X0. The three-dimensional data encoding device switches the encoding table used in the entropy encoding of the occupancy code of the object node based on the obtained occupancy information of the child node of the adjacent node X0.
[0599] As described above, the three-dimensional data encoding device according to the present embodiment encodes information (such as occupancy rate encoding) of object nodes included in an N-way tree structure (N is an integer greater than or equal to 2) of a plurality of three-dimensional points included in three-dimensional data. Fig.51 as well as Fig.52 As shown, in the above encoding, the three-dimensional data encoding device allows reference to information (e.g., occupancy information) of a first node whose parent node is the same as the parent node of the object node among a plurality of adjacent nodes spatially adjacent to the object node, and prohibits reference to information (e.g., occupancy information) of a second node whose parent node is different from the parent node of the object node. In other words, in the above encoding, the three-dimensional data encoding device allows reference to information (e.g., occupancy rate encoding) of the parent node, and prohibits reference to information (e.g., occupancy rate encoding) of other nodes (parent adjacent nodes) at the same layer as the parent node.
[0600] Thus, the three-dimensional data encoding device can improve the encoding efficiency by referring to the information of the first node whose parent node is the same as the parent node of the object node among the plurality of adjacent nodes that are spatially adjacent to the object node. In addition, the three-dimensional data encoding device can reduce the processing amount because it does not refer to the information of the second node whose parent node is different from the parent node of the object node among the plurality of adjacent nodes. In this way, the three-dimensional data encoding device can improve the encoding efficiency and reduce the processing amount.
[0601] For example, the three-dimensional data encoding device further determines whether to prohibit reference to the second node information, and in the above encoding, based on the result of the above determination, switches whether to prohibit or permit reference to the second node information. The three-dimensional data encoding device further generates a switching prohibition information (for example, Fig.63 The limit_refer_flag shown in the bit stream is configured such that the prohibition switching information is the result of the above determination and indicates whether reference to the second node is prohibited.
[0602] This allows the three-dimensional data encoding device to switch whether to prohibit reference to the information of the second node. In addition, the three-dimensional data decoding device can appropriately perform decoding processing using the prohibition switching information.
[0603] For example, the information of the object node is information indicating whether there is a three-dimensional point in each of the child nodes belonging to the object node (for example, occupancy coding), the information of the first node is information indicating whether there is a three-dimensional point in the first node (occupancy information of the first node), and the information of the second node is information indicating whether there is a three-dimensional point in the second node (occupancy information of the second node).
[0604] For example, in the above encoding, the three-dimensional data encoding device selects a coding table based on whether a three-dimensional point exists at the first node, and uses the selected coding table to entropy encode the information of the target node (eg, occupancy code).
[0605] For example, Fig.66 as well as Fig.67 As shown, the three-dimensional data encoding device allows reference to information (for example, occupancy information) of child nodes of a first node among a plurality of adjacent nodes during the encoding.
[0606] Therefore, the three-dimensional data encoding device can improve encoding efficiency because it can refer to more detailed information of adjacent nodes.
[0607] For example, Fig.53 As shown, in the above encoding, the three-dimensional data encoding device switches the adjacent node to be referenced among the plurality of adjacent nodes according to the spatial position of the parent node of the object node.
[0608] Thus, the three-dimensional data encoding device can refer to appropriate adjacent nodes according to the spatial position of the target node in the parent node.
[0609] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0610] In addition, the three-dimensional data decoding device according to the present embodiment decodes information (such as occupancy code) of the object node included in the N (N is an integer greater than or equal to 2) tree structure of a plurality of three-dimensional points included in the three-dimensional data. Fig.51 as well as Fig.52 As shown, in the above decoding, the three-dimensional data decoding device allows reference to information (e.g., occupancy information) of a first node whose parent node is the same as the parent node of the object node among a plurality of adjacent nodes spatially adjacent to the object node, and prohibits reference to information (e.g., occupancy information) of a second node whose parent node is different from the parent node of the object node. In other words, in the above decoding, the three-dimensional data decoding device allows reference to information (e.g., occupancy code) of the parent node, and prohibits reference to information (e.g., occupancy code) of other nodes (parent adjacent nodes) in the same layer as the parent node.
[0611] Thus, the three-dimensional data decoding device can improve the coding efficiency by referring to the information of the first node whose parent node is the same as the parent node of the object node among the plurality of adjacent nodes that are spatially adjacent to the object node. In addition, the three-dimensional data decoding device can reduce the processing amount because it does not refer to the information of the second node whose parent node is different from the parent node of the object node among the plurality of adjacent nodes. In this way, the three-dimensional data decoding device can improve the coding efficiency and reduce the processing amount.
[0612] For example, the three-dimensional data decoding device further obtains, from the bit stream, prohibition switching information indicating whether to prohibit reference to the information of the second node (for example, Fig.63 limit_refer_flag) is shown in the above decoding. Based on the prohibition switching information, whether to prohibit or permit reference to the information of the second node is switched.
[0613] Thus, the three-dimensional data decoding device can appropriately perform decoding processing using the switching inhibition information.
[0614] For example, the information of the object node is information indicating whether there is a three-dimensional point in each of the child nodes belonging to the object node (for example, occupancy coding), the information of the first node is information indicating whether there is a three-dimensional point in the first node (occupancy information of the first node), and the information of the second node is information indicating whether there is a three-dimensional point in the second node (occupancy information of the second node).
[0615] For example, in the above decoding, the three-dimensional data decoding device selects a coding table based on whether a three-dimensional point exists at the first node, and uses the selected coding table to entropy decode the information of the target node (eg, occupancy code).
[0616] For example, Fig.66 as well as Fig.67 As shown, the three-dimensional data decoding device allows reference to information (for example, occupancy information) of child nodes of a first node among a plurality of adjacent nodes during the above decoding.
[0617] As a result, the three-dimensional data decoding device can refer to more detailed information of adjacent nodes, thereby improving encoding efficiency.
[0618] For example, Fig.53 As shown, in the above decoding, the three-dimensional data decoding device switches the adjacent node to be referenced among the plurality of adjacent nodes according to the spatial position of the target node in the parent node.
[0619] Thus, the three-dimensional data decoding device can refer to appropriate adjacent nodes according to the spatial position of the target node in the parent node.
[0620] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0621] (Implementation method 9)
[0622] When encoding the encoding information of the node of the encoding object (hereinafter referred to as the object node), the three-dimensional data encoding device can improve the encoding efficiency by using the adjacent node information of the object node. For example, the three-dimensional data encoding device uses the adjacent node information to switch the encoding table (probability table, etc.) used to entropy encode the occupancy code (occupancy code) of the object node. Here, the adjacent node information is, for example, information indicating whether a plurality of nodes (adjacent nodes) spatially adjacent to the object node are occupied nodes (occupied nodes) (whether the adjacent nodes contain a point group), etc.
[0623] For example, the three-dimensional data encoding device may also switch the encoding table using information indicating the number of occupied nodes (adjacent occupied nodes) among a plurality of adjacent nodes. For example, the number of occupied nodes among the six adjacent nodes (left, right, top, bottom, near front, and back) adjacent to the object node may be calculated, and the encoding table used for entropy encoding of the occupancy rate encoding of the object node may be switched according to the number of occupied nodes.
[0624] Furthermore, the three-dimensional data encoding device may use a pattern (hereinafter referred to as a neighboring pattern) of neighboring occupied nodes instead of the number of neighboring occupied nodes. Fig.68is a diagram showing examples of adjacent nodes and the processing of this embodiment. Fig.68 In the example shown, the adjacent nodes X0, Y0, and Z0 are occupied nodes, and the adjacent nodes X1, Y1, and Z1 are non-occupied nodes that are not occupied nodes. In this case, the three-dimensional data encoding device converts the bit string (Z1 Z0 Y1Y0 X1 X0) = (010101) into a decimal number, calculates the value 21 of the adjacent occupation pattern, and uses the 21st encoding table to encode the occupancy code of the object node. In addition, the three-dimensional data encoding device can also use other values calculated using the value 21 as the value of the encoding table.
[0625] Here, the neighboring node information of the object node is used to calculate the neighboring occupancy pattern, the coding table is switched according to the value, and the NeighbourPatternCodingFlag (neighboring pattern coding flag) is set. The NeighbourPatternCodingFlag is a flag for switching whether to perform arithmetic coding on the occupancy coding of the object node. The three-dimensional data coding device adds the NeighbourPatternCodingFlag to the header of the bit stream, etc.
[0626] In the case of NeighbourPatternCodingFlag = 1, the three-dimensional data encoding device calculates the neighboring occupancy pattern by using the neighboring node information of the object node, switches the coding table according to the value, and performs arithmetic coding on the occupancy coding of the object node. In the case of NeighbourPatternCodingFlag = 0, the three-dimensional data encoding device does not use the neighboring node information of the object node, and performs arithmetic coding on the occupancy coding of the object node.
[0627] For example, when NeighbourPatternCodingFlag = 1, Fig.68 As shown, the three-dimensional data encoding device uses the six nodes adjacent to the target node to calculate the adjacent occupancy pattern. In this case, the value of the adjacent occupancy pattern can take a value from 0 to 63. Therefore, the three-dimensional data encoding device switches a total of 64 coding tables and performs arithmetic coding on the occupancy code of the target node.
[0628] For example, in Fig.68 In the example shown, the value of the neighboring occupancy pattern is 21, and the three-dimensional data encoding device uses the 21st coding table to entropy encode the occupancy code of the target node. Alternatively, the three-dimensional data encoding device may use a coding table of an index number calculated based on the value 21.
[0629] In addition, for example, when NeighbourPatternCodingFlag=0, the three-dimensional data encoding device determines the coding table without using the value of the neighboring node information. For example, the three-dimensional data encoding device sets the value of the neighboring occupancy pattern to 0, determines the coding table without using the value of the neighboring node information, and performs arithmetic coding on the occupancy coding of the object node. That is, the three-dimensional data encoding device sets the neighboring occupancy pattern to 0 and uses the 0th coding table. In other words, the three-dimensional data encoding device uses a predetermined coding table.
[0630] In this way, the three-dimensional data coding device calculates the neighboring occupancy pattern according to the value of NeighbourPatternCodingFlag, and switches whether to switch the coding table according to the calculated neighboring occupancy pattern to perform coding. This makes it possible to strike a balance between coding efficiency and low-process quantization.
[0631] Here, as a mode for encoding the object node, there may be, for example, a normal node (or normal mode) and an early terminated node (or direct coding mode), wherein the normal node further divides the node into 8 child nodes and encodes them in an octree structure, and the early terminated node stops the encoding in the 8-divided octree structure and directly encodes the position information of the three-dimensional points in the node.
[0632] For example, when the number of point groups in the object node is below a certain threshold value A, the three-dimensional data encoding device sets the object node as an early terminal node to stop the octree segmentation. Alternatively, when the number of point groups in the parent node is below a certain threshold value B, the three-dimensional data encoding device sets the object node as an early terminal node to stop the octree segmentation. Alternatively, when the number of point groups included in the adjacent node is below a certain threshold value C, the three-dimensional data encoding device sets the object node as an early terminal node to stop the octree segmentation.
[0633] In this way, the three-dimensional data encoding device can also use the number of point groups contained in the object node, parent node, or adjacent node to determine whether the object node is an early terminal node. If it is true, the octree segmentation is stopped, and if it is false, the octree segmentation is continued for encoding. Thus, the three-dimensional data encoding device stops the octree segmentation when the number of point groups contained in the object node, parent node, or adjacent node decreases, thereby reducing the processing time. In addition, the three-dimensional data encoding device can also use entropy coding, etc. to encode the three-dimensional position information of each point group contained in the node for the early terminal node.
[0634] Fig.694 is a flowchart of the three-dimensional data encoding process of this embodiment. First, the three-dimensional data encoding device determines whether the object node satisfies the condition I of becoming an early terminal node (S4401). In other words, the determination is to determine whether the object node is likely to be encoded as an early terminal node, that is, whether the early terminal node can be used.
[0635] When condition I is true ("Yes" in S4401), next, the three-dimensional data encoding device determines whether the object node satisfies the determination condition J of whether it is an early terminal node (S4402). In other words, this determination is whether the object node is actually encoded as an early terminal node, that is, whether the early terminal node is used.
[0636] When condition J is true ("Yes" in S4402), the three-dimensional data encoding device sets early_terminated_node_flag to 1 and encodes the early_terminated_node_flag (S4403). Next, the three-dimensional data encoding device directly encodes the position information of the three-dimensional point included in the object node (S4404). That is, the three-dimensional data encoding device applies the early terminal node to the object node.
[0637] On the other hand, when condition J is false (No in S4402), the three-dimensional data encoding device sets early_terminated_node_flag to 0 and encodes the early_terminated_node_flag (S4405). Next, the three-dimensional data encoding device sets the target node to a normal node and continues encoding based on octree partitioning (S4406).
[0638] In addition, when condition I is false ("No" in S4401), the three-dimensional data encoding device does not encode early_terminated_node_flag, sets the target node to a normal node, and continues encoding based on octree partitioning (S4406).
[0639] For example, condition J includes a condition such as whether the number of three-dimensional points in the object node is below a threshold value (e.g., value 2). For example, if the number of three-dimensional points in the object node is below the threshold value, the three-dimensional data encoding device determines the object node as an early terminal node, otherwise it determines that the object node is not an early terminal node.
[0640] In addition, condition I includes, for example, whether the layer to which the object node belongs is a layer or higher than a predetermined layer of the octree. For example, condition I may also include a condition such as whether the object node is a layer larger than a node having a leaf node (the lowest layer) (whether it contains a space of a certain size or higher, etc.).
[0641] In addition, condition I may also include conditions on the nodes (sibling nodes) contained in the parent node of the object node, or the occupancy information of the sibling nodes of the parent node. That is, the three-dimensional data encoding device may also determine whether the object node is likely to become an early terminal node based on the occupancy information of the sibling nodes or the sibling nodes of the parent node. For example, the three-dimensional data encoding device counts the number of nodes in an occupied state among the sibling nodes in the parent node that is the same as the object node. Condition I includes conditions such as whether the counted value is less than a predetermined value. Alternatively, the three-dimensional data encoding device counts the number of nodes in an occupied state among the sibling nodes of the parent node of the object node. Condition I includes conditions such as whether the counted value is less than a predetermined value.
[0642] In this way, the three-dimensional data encoding device first determines whether the object node is likely to become an early terminal node using the hierarchy in the octree structure of the object node, the occupancy information of the sibling nodes of the object node, or the occupancy information of the sibling nodes of the parent node. If possible, the three-dimensional data encoding device encodes early_terminated_node_flag, otherwise it does not encode early_terminated_node_flag. Thus, the three-dimensional data encoding device can select the early terminal node while encoding while suppressing the additional overhead.
[0643] In addition, the three-dimensional data encoding device may set an EarlyTerminatdCodingFlag (early terminal encoding flag), which is a flag indicating whether to use an early terminal node (direct encoding mode) for encoding, and attach the flag to a header or the like.
[0644] In addition, condition I may include a determination of whether EarlyTerminatdCodingFlag = 1 is satisfied. In this way, by setting a structure for switching whether to use an early termination node (direct coding mode) according to the value of EarlyTerminatdCodingFlag, a balance between coding efficiency and low-process quantization can be achieved.
[0645] Alternatively, the three-dimensional data encoding device may calculate the neighboring occupancy pattern of the object node, and the condition I may include whether the neighboring occupancy pattern = 0 is satisfied. Thus, when the neighboring node is in a non-occupied state, that is, when the three-dimensional points are sparse, the possibility of selecting an early terminal node becomes high. Therefore, the three-dimensional data encoding device can improve the encoding efficiency by efficiently selecting the early terminal node.
[0646] Fig.70 1 is a flowchart of a three-dimensional data encoding process (early terminal node determination process) of the three-dimensional data encoding device of this embodiment. The three-dimensional data encoding device uses NeighbourPatternCodingFlag and EarlyTerminatdCodingFlag.
[0647] First, the three-dimensional data encoding device determines whether NeighbourPatternCodingFlag is 1 (S4411). The NeighbourPatternCodingFlag is generated in the three-dimensional data encoding device, for example. For example, the three-dimensional data encoding device determines the value of NeighbourPatternCodingFlag based on a coding mode specified externally or an input three-dimensional point.
[0648] When NeighbourPatternCodingFlag=1 (Yes in S4411), the three-dimensional data encoding device calculates the neighbor occupancy pattern of the object node (S4412). For example, the three-dimensional data encoding device uses the calculated neighbor occupancy pattern for selecting a coding table for performing arithmetic coding on the occupancy code.
[0649] On the other hand, when NeighbourPatternCodingFlag=0 (No in S4411), the three-dimensional data encoding device does not calculate the neighboring occupation pattern, but sets the value of the neighboring occupation pattern to 0 (S4413).
[0650] Alternatively, the three-dimensional data encoding device may initialize the neighboring occupation pattern to 0, and if NeighbourPatternCodingFlag=1, update the value of the neighboring occupation pattern.
[0651] Next, the three-dimensional data encoding device determines whether condition I is satisfied (S4414). For example, when EarlyTerminatdCodingFlag=1, the three-dimensional data encoding device may also use the value of the set adjacent occupancy pattern to determine whether the object node is likely to become an early terminal node. That is, condition I may also include a condition of whether the set adjacent occupancy pattern is 0.
[0652] That is, condition I may include whether the condition of EarlyTerminatedCodingFlag = 1 is satisfied. In addition, condition I may include whether the condition of adjacent occupation pattern = 0 is satisfied. For example, condition I may be true when EarlyTerminatedCodingFlag = 1 and adjacent occupation pattern = 0, and false otherwise.
[0653] Thus, when NeighbourPatternCodingFlag=1, the three-dimensional data encoding device can also use the neighboring occupancy pattern calculated and set for the coding table switching in the judgment of the early terminal node (condition I). As a result, the amount of processing for recalculating the neighboring occupancy pattern can be reduced. In addition, when NeighbourPatternCodingFlag=0, the three-dimensional data encoding device can determine that at least one condition of condition I is satisfied by setting the neighboring occupancy pattern to 0. Therefore, the three-dimensional data encoding device does not need to calculate the neighboring occupancy pattern separately, so the amount of processing can be reduced.
[0654] In addition, condition I may also include, for example, whether the layer to which the object node belongs is a layer or higher than a predetermined layer of the octree. For example, condition I may also include a condition such as whether the object node is a layer larger than a node having a leaf node (the lowest layer) (whether it contains a space of a certain size or higher, etc.).
[0655] In addition, condition I may also include conditions on the nodes (brother nodes) contained in the parent node of the object node, or the occupancy information of the sibling nodes of the parent node. That is, the three-dimensional data encoding device may also determine whether the object node is likely to become an early terminal node based on the occupancy information of the sibling nodes or the sibling nodes of the parent node. For example, the three-dimensional data encoding device counts the number of nodes in an occupied state among the sibling nodes in the parent node that has the same parent node as the object node. Condition I may also include conditions such as whether the counted value is less than a predetermined value. Alternatively, the three-dimensional data encoding device counts the number of nodes in an occupied state among the sibling nodes of the parent node of the object node. Condition I may also include conditions such as whether the counted value is less than a predetermined value.
[0656] In addition, condition I may include any one of the above-mentioned multiple conditions, or may include multiple conditions. Alternatively, when multiple conditions are included as condition I, for example, when all conditions are met, it is determined that condition I is met (true), and in other cases, it is determined that condition I is not met (false). Alternatively, it is also possible to determine that condition I is met (true) when at least one of the multiple conditions is met.
[0657] In addition, the processing of steps S4415 to S4419 is the same as Fig.69 The processing of steps S4402 to S4406 shown is the same, and repeated description is omitted.
[0658] Fig.71 This is a flowchart of a modified example of the three-dimensional data encoding process (early terminal node determination process) of the three-dimensional data encoding device according to the present embodiment. Fig.71 The processing shown is similar to Fig.70 The processing shown is different in that step S4411 is changed to step S4411A.
[0659] The three-dimensional data encoding device determines whether at least one of NeighbourPatternCodingFlag=1 and EarlyTerminatdCodingFlag=1 is satisfied (S4411A). The NeighbourPatternCodingFlag and EarlyTerminatdCodingFlag are generated in the three-dimensional data encoding device, for example. For example, the three-dimensional data encoding device determines the values of NeighbourPatternCodingFlag and EarlyTerminatdCodingFlag based on a coding mode specified externally or an input three-dimensional point.
[0660] When at least one of NeighbourPatternCodingFlag = 1 and EarlyTerminatdCodingFlag = 1 is satisfied ("Yes" in S4411A), the three-dimensional data encoding device calculates the neighboring occupancy pattern of the object node and sets the value (S4412). When neither NeighbourPatternCodingFlag = 1 nor EarlyTerminatdCodingFlag = 1 is satisfied ("No" in S4411A), the three-dimensional data encoding device does not calculate the neighboring occupancy pattern and sets the value of the neighboring occupancy pattern to 0 (S4413).
[0661] In addition, the three-dimensional data encoding device may also initialize the neighboring occupied pattern with a value of 0, and if NeighbourPatternCodingFlag = 1 or EarlyTerminatdCodingFlag = 1, update the value of the neighboring occupied pattern. Fig.70 same.
[0662] That is, when EarlyTerminatdCodingFlag=1, the three-dimensional data coding device uses the value of the set neighboring occupation pattern to determine whether the target node is likely to become an early terminal node. That is, condition I may also include a condition of whether the set neighboring occupation pattern is zero.
[0663] Thus, when NeighbourPatternCodingFlag=1 or EarlyTerminatdCodingFlag=1, the three-dimensional data encoding device calculates the neighboring occupancy pattern of the target node, and uses the calculated neighboring occupancy pattern value to determine the possibility that the target node is an early terminal node. Therefore, the three-dimensional data encoding device can appropriately select an early terminal node, thereby improving encoding efficiency.
[0664] Next, the processing of the three-dimensional data decoding device according to this embodiment will be described. Fig.72 This is a flowchart of the three-dimensional data decoding process (early terminal node determination process) of the three-dimensional data decoding device according to the present embodiment.
[0665] First, the three-dimensional data decoding apparatus decodes NeighbourPatternCodingFlag from the header of the bit stream (S4421). Next, the three-dimensional data decoding apparatus decodes EarlyTerminatdCodingFlag from the header of the bit stream (S4422).
[0666] Next, the three-dimensional data decoding device determines whether the decoded NeighbourPatternCodingFlag is 1 (S4423).
[0667] When NeighbourPatternCodingFlag is 1 (Yes in S4423), the three-dimensional data decoding device calculates the neighbor occupancy pattern of the object node (S4424). In addition, the three-dimensional data decoding device may also use the calculated neighbor occupancy pattern for selecting a coding table for performing arithmetic decoding on the occupancy code.
[0668] When NeighbourPatternCodingFlag is 0 (No in S4423), the 3D data decoding device sets the neighbor occupancy pattern to 0 (S4425). Alternatively, the 3D data decoding device may initialize the neighbor occupancy pattern to 0, and if NeighbourPatternCodingFlag=1, update the value of the neighbor occupancy pattern.
[0669] Next, the three-dimensional data decoding device determines whether condition I is true (S4426). The details of this process are the same as the process of step S4414 in the three-dimensional data encoding device.
[0670] When condition I is true (Yes in S4426), the three-dimensional data decoding device decodes early_terminated_node_flag from the bit stream (S4427). Next, the three-dimensional data decoding device determines whether early_terminated_node_flag is 1 (S4428).
[0671] When early_terminated_node_flag is 1 ("Yes" in S4428), the three-dimensional data decoding device decodes the position information of the three-dimensional point in the object node (S4429). That is, the three-dimensional data decoding device applies the early terminal node to the object node. When early_terminated_node_flag is 0 ("No" in S4428), the three-dimensional data decoding device sets the object node as a normal node and continues decoding based on octree partitioning (S4430).
[0672] In addition, when condition I is false ("No" in S4426), the three-dimensional data decoding device does not decode early_terminated_node_flag from the bit stream, but sets the target node to a normal node and continues decoding based on octree partitioning (S4430).
[0673] Fig.73 This is a flowchart of a modified example of the three-dimensional data decoding process (early terminal node determination process) of the three-dimensional data decoding device according to the present embodiment. Fig.73 The processing shown is similar to Fig.72 The processing shown is different in that step S4423 is changed to step S4423A.
[0674] In step S4423A, the three-dimensional data decoding apparatus determines whether at least one of NeighbourPatternCodingFlag=1 and EarlyTerminatdCodingFlag=1 is satisfied (S4423A).
[0675] When at least one of NeighbourPatternCodingFlag=1 and EarlyTerminatdCodingFlag=1 is satisfied ("Yes" in S4423A), the three-dimensional data decoding device calculates the neighboring occupancy pattern of the object node (S4424). When neither NeighbourPatternCodingFlag=1 nor EarlyTerminatdCodingFlag=1 is satisfied ("No" in S4423A), the three-dimensional data decoding device does not calculate the neighboring occupancy pattern and sets the value of the neighboring occupancy pattern to 0 (S4425).
[0676] In addition, the three-dimensional data decoding device may also initialize the neighboring occupancy pattern with a value of 0, and if NeighbourPatternCodingFlag = 1 or EarlyTerminatdCodingFlag = 1, update the value of the neighboring occupancy pattern. Fig.72 same.
[0677] Next, a syntax example of a bit stream generated by the three-dimensional data encoding device according to this embodiment will be described. Fig.74 This is a diagram showing a syntax example of pc_header included in a bitstream. This pc_header() is, for example, header information of a plurality of input three-dimensional points. That is, the information included in pc_header() is commonly used for a plurality of three-dimensional points (nodes).
[0678] pc_header includes NeighbourPatternCodingFlag (neighboring pattern coding flag) and EarlyTerminatdCodingFlag (early terminal coding flag).
[0679] NeighbourPatternCodingFlag is information indicating whether to use neighboring node information (neighboring occupancy pattern) to switch the coding table used for arithmetic coding of occupancy coding. For example, NeighbourPatternCodingFlag = 1 indicates that the coding table is switched using neighboring node information, and NeighbourPatternCodingFlag = 0 indicates that the coding table is not switched using neighboring node information.
[0680] EarlyTerminatdCodingFlag is information indicating whether (can be used) the early termination node (direct coding mode) is used. For example, EarlyTerminatdCodingFlag=1 indicates that the early termination node is used, and EarlyTerminatdCodingFlag=0 indicates that the early termination node is not used.
[0681] Fig.75 This is a diagram showing a syntax example of node information (node (depth, index)). This node information is information of a node included in the octree and is set for each node. The node information includes occupancy_code (occupancy code), early_terminated_node_flag (early terminal node flag), and coordinate_of_3Dpoint (three-dimensional coordinates).
[0682] The occupancy_code is information indicating whether a child node of a node is in an occupied state. The three-dimensional data encoding device may switch the encoding table according to the value of NeighbourPatternCodingFlag and perform arithmetic encoding on the occupancy_code.
[0683] early_terminated_node_flag is information indicating whether the node is an early terminal node. For example, early_terminated_node_flag = 1 indicates that the node is an early terminal node, and early_terminated_node_flag = 0 indicates that the node is not an early terminal node. In addition, when the early_terminated_node_flag of the target node is not encoded in the bitstream, the three-dimensional data decoding device may also estimate the value of the early_terminated_node_flag of the target node to be 0.
[0684] coordinate_of_3Dpoint is the position information of the point group included in the node when the node is an early terminal node. In addition, when a plurality of point groups are included in the node, coordinate_of_3Dpoint may include the position information of each point group.
[0685] In addition, the 3D data encoding device may not add NeighbourPatternCodingFlag or EarlyTerminatdCodingFlag to the header, but may define the value of NeighbourPatternCodingFlag or EarlyTerminatdCodingFlag according to a standard or a standard profile or level, etc. Thus, the 3D data decoding device can correctly restore the bitstream by determining the value of NeighbourPatternCodingFlag or EarlyTerminatdCodingFlag by referring to the standard information included in the bitstream.
[0686] In addition, the three-dimensional data encoding device may also perform entropy encoding on at least one of the above-mentioned NeighbourPatternCodingFlag, EarlyTerminatdCodingFlag, early_terminated_node_flag, and coordinate_of_3Dpoint. For example, the three-dimensional data encoding device performs arithmetic encoding after binarizing each value.
[0687] In addition, in this embodiment, an octree structure is used as an example, but the present invention is not limited to this, and the above-mentioned method can also be applied to an N-ary tree structure (N is an integer greater than or equal to 2) such as a quadtree or a hexadecimal tree, or other tree structures.
[0688] Next, a configuration example of a three-dimensional data encoding device according to this embodiment will be described. Fig.76 4 is a block diagram of a three-dimensional data encoding device 4400 according to the present embodiment. The three-dimensional data encoding device 4400 includes an octree generation unit 4401 , a geometric information calculation unit 4402 , a coding table selection unit 4403 , and an entropy coding unit 4404 .
[0689] The octree generation unit 4401 generates, for example, an octree from the input three-dimensional points (point cloud), and generates occupancy coding of each node of the octree. In addition, when EarlyTerminatdCodingFlag=1, the octree generation unit 4401 can also use the judgment of conditions I and J to determine whether the object node is an early terminal node. If it is true, the octree segmentation is stopped, and if it is false, the octree segmentation is continued for encoding. In addition, the octree generation unit 4401 can also attach a flag (early_terminated_node_flag) indicating whether each node is an early terminal node to the bit stream. As a result, the three-dimensional data decoding device can correctly determine whether the node is an early terminal node.
[0690] The geometric information calculation unit 4402 obtains information indicating whether the neighboring nodes of the target node are occupied, and calculates the neighboring occupied pattern based on the obtained information. Fig.68etc. to calculate the neighboring occupancy pattern. In addition, the geometric information calculation unit 4402 may also calculate the neighboring occupancy pattern based on the occupancy coding of the parent node to which the object node belongs. In addition, the geometric information calculation unit 4402 may save the coded nodes in a list and search for neighboring nodes from the list. In addition, the geometric information calculation unit 4402 may also switch the neighboring nodes based on the position within the parent node of the object node. In addition, the geometric information calculation unit 4402 may switch whether to calculate the neighboring occupancy pattern based on the values of NeighbourPatternCodingFlag and EarlyTerminatdCodingFlag.
[0691] The coding table selection unit 4403 selects a coding table used for entropy coding of the target node using the adjacent node occupancy information (adjacent occupancy pattern) calculated by the geometric information calculation unit 4402. For example, the coding table selection unit 4403 selects a coding table of an index number calculated based on the value of the adjacent occupancy pattern.
[0692] The entropy coding unit 4404 performs entropy coding on the occupancy code of the target node by using the coding table of the selected index number, thereby generating a bit stream. The entropy coding unit 4404 may add information of the selected coding table to the bit stream.
[0693] Next, a configuration example of a three-dimensional data decoding device according to this embodiment will be described. Fig.77 4 is a block diagram of a three-dimensional data decoding device 4410 according to the present embodiment. The three-dimensional data decoding device 4410 includes an octree generation unit 4411 , a geometric information calculation unit 4412 , a coding table selection unit 4413 , and an entropy decoding unit 4414 .
[0694] The octree generation unit 4411 generates an octree of a certain space (node) using the header information of the bitstream, etc. For example, the octree generation unit 4411 generates a large space (root node) using the size of the x-axis, y-axis, and z-axis directions of a certain space attached to the header information, and divides the space into two in the x-axis, y-axis, and z-axis directions, thereby generating eight small spaces A (nodes A0 to A7) to generate an octree. In addition, nodes A0 to A7 are set in sequence as object nodes.
[0695] In addition, when the value of EarlyTerminatdCodingFlag obtained by decoding the header is 1, the octree generation unit 4411 may use the judgment of condition I and condition J to determine whether the target node is an early terminal node, and if true, stop the octree division, and if false, continue the octree division and decoding. In addition, the octree generation unit 4411 may also decode a flag indicating whether each node is an early terminal node.
[0696] The geometric information calculation unit 4412 obtains information indicating whether the neighboring nodes of the target node are occupied, and calculates the neighboring occupied pattern based on the obtained information. For example, the geometric information calculation unit 4412 can calculate the neighboring occupied pattern by using Fig.68 etc. to calculate the adjacent occupancy pattern. In addition, the geometric information calculation unit 4412 can calculate the occupancy information of the adjacent node based on the occupancy code of the parent node to which the object node belongs. In addition, the geometric information calculation unit 4412 can also save the decoded nodes in a list and search for adjacent nodes from the list. In addition, the geometric information calculation unit 4412 can also switch the adjacent nodes according to the position in the parent node of the object node. In addition, the geometric information calculation unit 4412 can also switch whether to calculate the adjacent occupancy pattern based on the values of NeighbourPatternCodingFlag and EarlyTerminatdCodingFlag obtained by decoding the header.
[0697] The coding table selection unit 4413 selects a coding table used for entropy decoding of the target node using the adjacent node occupancy information (adjacent occupancy pattern) calculated by the geometric information calculation unit 4412. For example, the three-dimensional data decoding apparatus selects a coding table of an index number calculated based on the value of the adjacent occupancy pattern.
[0698] The entropy decoding unit 4414 generates a three-dimensional point (point cloud) by entropy decoding the occupancy code of the object node using the selected coding table. The entropy decoding unit 4414 may also decode from the bit stream and obtain information indicating the selected coding table, and entropy decode the occupancy code of the object node using the coding table indicated by the information.
[0699] In addition, each bit of the occupancy code (8 bits) contained in the bit stream indicates whether a point group is included in each of the 8 small spaces A (node A0 to node A7). In addition, the three-dimensional data decoding device further divides the small space node A0 into 8 small spaces B (node B0 to node B7) to generate an octree, and decodes the occupancy code to calculate information indicating whether a point group is included in each node of the small space B. In this way, the three-dimensional data decoding device decodes the occupancy code of each node while generating an octree from a large space to a small space. In addition, when the object node is an early terminal node, the three-dimensional data decoding device can also directly decode the three-dimensional information encoded in the bit stream and stop the octree division at this node.
[0700] Hereinafter, a modification example of the three-dimensional data encoding process and the three-dimensional data decoding process (early terminal node determination process) will be described.
[0701] The method for calculating the neighboring occupancy pattern of the object node is not limited to using Fig.68The method of calculating the occupancy information of the six adjacent nodes shown in the figure may also be other methods. For example, the three-dimensional data encoding device may also refer to the adjacent nodes of the object node that exist in the parent node (the brother nodes of the object node) to calculate the adjacent occupancy pattern. For example, the three-dimensional data encoding device may also calculate the adjacent occupancy pattern composed of the object node and position in the parent node and the occupancy information of the three adjacent nodes in the parent node. When six adjacent nodes are used, the information of the adjacent nodes whose parent nodes are different from the parent nodes of the object node is used. When three adjacent nodes are used, the information of the adjacent nodes whose parent nodes are different from the parent nodes of the object node is not used. Thus, the three-dimensional data encoding device can calculate the adjacent occupancy pattern with reference to the occupancy coding of the parent node, thereby reducing the amount of processing.
[0702] In addition, the three-dimensional data encoding device can also switch between the above-mentioned method of using 6 adjacent nodes (a method of referring to adjacent nodes whose parent nodes are different from the parent nodes of the object node) and the method of using 3 adjacent nodes (a method of not referring to adjacent nodes whose parent nodes are different from the parent nodes of the object node). For example, the three-dimensional data encoding device uses the method of using 6 adjacent nodes to calculate the adjacent occupancy pattern A in the switching judgment of the coding table used for arithmetic coding of the occupancy code. In addition, the three-dimensional data encoding device uses the method of using 3 adjacent nodes to calculate the adjacent occupancy pattern B in the possibility judgment of the early terminal node (condition I). In addition, the three-dimensional data encoding device can also use the method of using 3 adjacent nodes in the switching judgment of the coding table, and use the method of using 6 adjacent nodes in the possibility judgment of the early terminal node (condition I). In this way, the three-dimensional data encoding device can control the balance between coding efficiency and processing volume by switching the calculation method of the adjacent occupancy pattern.
[0703] Fig.78 4441 is a flowchart of the three-dimensional data encoding process. First, the three-dimensional data encoding device determines whether NeighbourPatternCodingFlag is 1 (S4441).
[0704] When NeighbourPatternCodingFlag=1 (Yes in S4441), the three-dimensional data encoding device calculates the neighbor occupancy pattern A of the target node (S4442). The three-dimensional data encoding device may use the calculated neighbor occupancy pattern A for selecting a coding table for performing arithmetic coding on the occupancy code.
[0705] On the other hand, when NeighbourPatternCodingFlag=0 (No in S4441), the three-dimensional data encoding device does not calculate the neighboring occupation pattern A, but sets the value of the neighboring occupation pattern A to 0 (S4443).
[0706] Next, the three-dimensional data encoding device determines whether EarlyTerminatdCodingFlag is 1 (S4444).
[0707] When EarlyTerminatdCodingFlag=1 (Yes in S4444), the three-dimensional data coding device calculates the neighboring occupancy pattern B of the target node (S4445). The neighboring occupancy pattern B is used for determination of condition I, for example.
[0708] On the other hand, when EarlyTerminatdCodingFlag=0 (No in S4444), the three-dimensional data encoding device does not calculate the adjacent occupation pattern B, but sets the value of the adjacent occupation pattern B to 0 (S4446).
[0709] For example, the three-dimensional data encoding device uses a method different from the calculation of the adjacent occupancy pattern A and the calculation of the adjacent occupancy pattern B. For example, the three-dimensional data encoding device uses a method using six adjacent nodes to calculate the adjacent occupancy pattern A, and uses a method using three adjacent nodes to calculate the adjacent occupancy pattern B. In addition, in the three-dimensional data decoding device, similarly, a method different from the calculation of the adjacent occupancy pattern A and the adjacent occupancy pattern B may be used.
[0710] In addition, the 3D data encoding device may also initialize the neighboring occupancy pattern A to a value of 0, and if NeighbourPatternCodingFlag=1, update the value of the neighboring occupancy pattern A. In addition, the 3D data encoding device may also initialize the neighboring occupancy pattern B to a value of 0, and if EarlyTerminatedCodingFlag=1, update the value of the neighboring occupancy pattern B.
[0711] Next, the three-dimensional data encoding device determines whether condition I is satisfied (S4447). The details of this process are as follows. Fig.70 The condition I is the same as the S4414 shown. However, the difference is that the adjacent occupation pattern B is used as the adjacent occupation pattern in step S4414. That is, condition I may also include whether the condition of EarlyTerminatedCodingFlag=1 is satisfied. In addition, condition I may also include whether the condition of adjacent occupation pattern B=0 is satisfied. For example, it may be that when EarlyTerminatedCodingFlag=1 and adjacent occupation pattern B=0, condition I is true, and in other cases, condition I is false.
[0712] In addition, the processing of steps S4448 to S4452 is the same as Fig.70The processing of steps S4415 to S4419 shown is the same, and repeated description is omitted.
[0713] Fig.79 This is a flowchart of a modified example of the three-dimensional data encoding process (early terminal node determination process) of the three-dimensional data encoding device according to the present embodiment. Fig.79 The processing shown is similar to Fig.78 The processing shown is different in that steps S4453 and S4454 are added.
[0714] When EarlyTerminatdCodingFlag=1 (Yes in S4444), the three-dimensional data encoding device determines whether NeighbourPatternCodingFlag is 1 (S4453).
[0715] When NeighbourPatternCodingFlag=1 (Yes in S4453), the three-dimensional data encoding device sets the value of the neighboring occupation pattern A to the value of the neighboring occupation pattern B (S4454).
[0716] When NeighbourPatternCodingFlag=0 (No in S4453), the three-dimensional data encoding device calculates the neighboring occupation pattern B (S4445).
[0717] In this way, when the adjacent occupancy pattern A is calculated, the three-dimensional data encoding device uses the adjacent occupancy pattern A as the adjacent occupancy pattern B. That is, the adjacent occupancy pattern A is used in the determination of the condition I. Thus, when the adjacent occupancy pattern A is calculated, the three-dimensional data encoding device does not calculate the adjacent occupancy pattern B, and thus the amount of processing can be reduced.
[0718] Fig.80 This is a flowchart of a modified example of the three-dimensional data decoding process (early terminal node determination process) of the three-dimensional data decoding device according to the present embodiment.
[0719] First, the three-dimensional data decoding apparatus decodes NeighbourPatternCodingFlag from the header of the bit stream (S4461). Next, the three-dimensional data decoding apparatus decodes EarlyTerminatdCodingFlag from the header of the bit stream (S4462).
[0720] Next, the three-dimensional data decoding device determines whether the decoded NeighbourPatternCodingFlag is 1 (S4463).
[0721] When NeighbourPatternCodingFlag is 1 (Yes in S4463), the three-dimensional data decoding device calculates the neighboring occupancy pattern A of the object node (S4464). In addition, the three-dimensional data decoding device may use the calculated neighboring occupancy pattern for selecting a coding table for performing arithmetic decoding on the occupancy code.
[0722] When NeighbourPatternCodingFlag is 0 (No in S4463), the three-dimensional data decoding device sets the neighboring occupation pattern A to 0 (S4465).
[0723] Next, the three-dimensional data decoding apparatus determines whether EarlyTerminatdCodingFlag is 1 (S4466).
[0724] When EarlyTerminatdCodingFlag=1 (Yes in S4466), the three-dimensional data decoding apparatus calculates the neighboring occupancy pattern B of the target node (S4467). The neighboring occupancy pattern B is used for determination of condition I, for example.
[0725] On the other hand, when EarlyTerminatdCodingFlag=0 (No in S4466), the three-dimensional data decoding apparatus does not calculate the adjacent occupation pattern B, but sets the value of the adjacent occupation pattern B to 0 (S4468).
[0726] For example, the three-dimensional data decoding apparatus uses different methods for calculating the neighboring occupancy pattern A and calculating the neighboring occupancy pattern B. For example, the three-dimensional data decoding apparatus calculates the neighboring occupancy pattern A using a method using six neighboring nodes, and calculates the neighboring occupancy pattern B using a method using three neighboring nodes.
[0727] In addition, the 3D data decoding device may initialize the neighboring occupancy pattern A to a value of 0, and if NeighbourPatternCodingFlag=1, update the value of the neighboring occupancy pattern A. In addition, the 3D data decoding device may initialize the neighboring occupancy pattern B to a value of 0, and if EarlyTerminatedCodingFlag=1, update the value of the neighboring occupancy pattern B.
[0728] Next, the three-dimensional data encoding device determines whether condition I is satisfied (S4469). The details of this process are as follows. Fig.72The condition I is the same as the S4426 shown. However, the difference is that the adjacent occupation pattern B is used as the adjacent occupation pattern in step S4426. That is, condition I may also include whether the condition of EarlyTerminatedCodingFlag=1 is satisfied. In addition, condition I may also include whether the condition of adjacent occupation pattern B=0 is satisfied. For example, it may be that when EarlyTerminatedCodingFlag=1 and adjacent occupation pattern B=0, condition I is true, and in other cases, condition I is false.
[0729] In addition, the processing of steps S4470 to S4473 is the same as Fig.72 The processing of steps S4427 to S4430 shown is the same, and repeated description is omitted.
[0730] Fig.81 This is a flowchart of a modified example of the three-dimensional data decoding process (early terminal node determination process) of the three-dimensional data decoding device according to the present embodiment. Fig.81 The processing shown is similar to Fig.80 The processing shown is different in that steps S4474 and S4475 are added.
[0731] When EarlyTerminatdCodingFlag=1 (Yes in S4466), the three-dimensional data decoding device next determines whether NeighbourPatternCodingFlag is 1 (S4474).
[0732] When NeighbourPatternCodingFlag=1 (Yes in S4474), the three-dimensional data decoding apparatus sets the value of the neighboring occupation pattern A to the value of the neighboring occupation pattern B (S4475).
[0733] When NeighbourPatternCodingFlag=0 (No in S4474), the three-dimensional data encoding device calculates the neighboring occupation pattern B (S4467).
[0734] In this way, when the adjacent occupancy pattern A is calculated, the three-dimensional data decoding device uses the adjacent occupancy pattern A as the adjacent occupancy pattern B. That is, the adjacent occupancy pattern A is used in the determination of the condition I. Thus, when the adjacent occupancy pattern A is calculated, the three-dimensional data decoding device does not calculate the adjacent occupancy pattern B, and thus the amount of processing can be reduced.
[0735] In addition, the correspondence between the values of the various flags (0 or 1) and their meanings described above is an example, and the correspondence between the values of the various flags and their meanings may be the opposite of the above.
[0736] As described above, the three-dimensional data encoding device of this embodiment enters the hole Fig.82 Processing shown.
[0737] The three-dimensional data encoding device determines whether the first flag (e.g., NeighbourPatternCodingFlag) indicates a first value (e.g., 1) (S4481). When the first flag indicates the first value ("Yes" in S4481), the three-dimensional data encoding device generates a first occupancy pattern (e.g., neighbor occupancy pattern A) indicating an occupancy state including a plurality of second neighboring nodes, wherein the plurality of second neighboring nodes include a first neighboring node whose parent node is different from the parent node of the object node, and the object node is included in an N (N is an integer greater than or equal to 2) fork tree structure of a plurality of three-dimensional points included in the three-dimensional data (S4482) (e.g., Fig.79 S4442 and S4454).
[0738] Next, the three-dimensional data encoding device determines, based on the first occupancy pattern, whether it is possible to use a first encoding (e.g., an early terminal node) (S4483) that encodes multiple three-dimensional position information included in the object node without dividing the object node into multiple child nodes (e.g., Fig.79 S4447).
[0739] The three-dimensional data encoding device generates a second occupancy pattern (for example, adjacent occupancy pattern B) representing the occupancy state of multiple third adjacent nodes when the first flag indicates a second value (for example, 0) different from the first value ("No" in S4481), and the multiple third adjacent nodes do not include the first adjacent node whose parent node is different from the parent node of the object node (S4484) (for example, Fig.79 Next, the three-dimensional data encoding device determines whether the first encoding can be used (S4485) based on the second occupancy pattern (for example, Fig.79 S4447).
[0740] In addition, the three-dimensional data encoding device generates a bit stream including the first flag (S4486).
[0741] Thus, the three-dimensional data encoding device can switch the adjacent node occupation pattern for determining whether the first encoding can be used according to the first flag. Thus, it is possible to appropriately determine whether the first encoding can be used, thereby improving encoding efficiency.
[0742] For example, when it is determined that the first code can be used, the three-dimensional data encoding device determines whether to use the first code (for example, Fig.79In S4448), if it is determined that the first code is used, the object node is encoded using the first code (for example, Fig.79 In S4450), when it is determined that the first encoding is not to be used, the object node is encoded using a second encoding that divides the object node into a plurality of child nodes (for example, Fig.79 S4452). The bit stream also includes a second flag (e.g., early_terminated_node_flag) indicating whether the first encoding is used.
[0743] For example, in determining whether the first code can be used based on the first occupancy pattern or the second occupancy pattern, the three-dimensional data encoding device determines whether the first code can be used based on the first occupancy pattern or the second occupancy pattern and the number of nodes in the occupancy state included in the parent node. For example, when the number of nodes in the occupancy state included in the parent node is less than a predetermined number, the three-dimensional data encoding device determines that the first code can be used, and when the number of nodes in the occupancy state included in the parent node is more than a predetermined number, the three-dimensional data encoding device determines that the first code cannot be used.
[0744] For example, in determining whether the first code can be used based on the first occupancy pattern or the second occupancy pattern, the three-dimensional data encoding device determines whether the first code can be used based on the first occupancy pattern or the second occupancy pattern and the number of nodes in the occupancy state included in the grandparent node of the object node. For example, when the number of nodes in the occupancy state included in the grandparent node is less than a predetermined number, the three-dimensional data encoding device determines that the first code can be used, and when the number of nodes in the occupancy state included in the grandparent node is more than a predetermined number, the three-dimensional data encoding device determines that the first code cannot be used.
[0745] For example, in determining whether the first code can be used based on the first occupancy pattern or the second occupancy pattern, the three-dimensional data encoding device determines whether the first code can be used based on the first occupancy pattern or the second occupancy pattern and the layer to which the object node belongs. For example, when the layer to which the object node belongs is lower than a predetermined layer, the three-dimensional data encoding device determines that the first code can be used, and when the layer to which the object node belongs is higher than a predetermined layer, the three-dimensional data encoding device determines that the first code cannot be used.
[0746] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0747] In addition, the three-dimensional data decoding device of this embodiment performs Fig.83First, the three-dimensional data decoding device obtains the first flag (for example, NeighbourPatternCodingFlag) from the bit stream (S4491). The three-dimensional data decoding device determines whether the first flag indicates a first value (for example, 1) (S4492).
[0748] The three-dimensional data decoding device generates a first occupancy pattern (e.g., adjacent occupancy pattern A) indicating occupancy states of a plurality of second adjacent nodes when the first flag indicates a first value ("No" in S4492), wherein the plurality of second adjacent nodes include a first adjacent node whose parent node is different from the parent node of the object node, and the object node is included in an N (N is an integer greater than or equal to 2) fork tree structure of a plurality of three-dimensional points included in the three-dimensional data (S4493) (e.g., Fig.81 Next, the three-dimensional data decoding device determines, based on the first occupancy pattern, whether it is possible to use the first decoding (e.g., early terminal node) (S4494) that decodes the multiple three-dimensional position information included in the object node without dividing the object node into multiple child nodes (e.g., Fig.81 S4469).
[0749] The three-dimensional data decoding device generates a second occupancy pattern (for example, adjacent occupancy pattern B) indicating the occupancy state of a plurality of third adjacent nodes when the first flag indicates a second value (for example, 0) different from the first value ("No" in S4492), wherein the plurality of third adjacent nodes does not include a first adjacent node whose parent node is different from the parent node of the object node (S4495) (for example, Fig.81 Next, the three-dimensional data decoding device determines whether the first decoding can be used (S4496) based on the second occupancy pattern (for example, Fig.81 S4469).
[0750] Thus, the three-dimensional data decoding device can switch the neighbor node occupation pattern for determining whether the first code can be used according to the first flag. Thus, it is possible to appropriately determine whether the first code can be used, thereby improving coding efficiency.
[0751] For example, when it is determined that the first decoding can be used, the three-dimensional data decoding apparatus obtains a second flag (for example, Fig.81 In S4470), when the second flag indicates that the first decoding is used, the object node is decoded using the first decoding (for example, Fig.81 In S4472), when the second flag indicates that the first decoding is not to be used, the object node is decoded using the second decoding that divides the object node into a plurality of child nodes (for example, Fig.81 S4473).
[0752] For example, in determining whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern, the three-dimensional data decoding device determines whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern and the number of nodes in the occupancy state included in the parent node. For example, the three-dimensional data decoding device determines that the first encoding can be used when the number of nodes in the occupancy state included in the parent node is less than a predetermined number, and determines that the first encoding cannot be used when the number of nodes in the occupancy state included in the parent node is more than a predetermined number.
[0753] For example, in determining whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern, the three-dimensional data decoding device determines whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern and the number of nodes in the occupancy state included in the grandparent node of the target node. For example, the three-dimensional data decoding device determines that the first encoding can be used when the number of nodes in the occupancy state included in the grandparent node is less than a predetermined number, and determines that the first encoding cannot be used when the number of nodes in the occupancy state included in the grandparent node is more than a predetermined number.
[0754] For example, in determining whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern, the three-dimensional data decoding device determines whether the first decoding can be used based on the first occupancy pattern or the second occupancy pattern and the layer to which the target node belongs. For example, the three-dimensional data decoding device determines that the first encoding can be used when the layer to which the target node belongs is lower than a predetermined layer, and determines that the first encoding cannot be used when the layer to which the target node belongs is higher than a predetermined layer.
[0755] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0756] As mentioned above, the three-dimensional data encoding device and the three-dimensional data decoding device and the like according to the embodiments of the present disclosure have been described, but the present disclosure is not limited to these embodiments.
[0757] Furthermore, each processing unit included in the three-dimensional data encoding device and three-dimensional data decoding device of the above-mentioned embodiment can be typically implemented as an LSI of an integrated circuit. These can be made into one chip separately, or part or all of them can be made into one chip.
[0758] Furthermore, integrated circuits are not limited to LSIs, and can be implemented by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays) that are programmable after LSI manufacturing, or reconfigurable processors that can reconfigure the connections or settings of circuits within LSIs can also be used.
[0759] Furthermore, in each of the above-mentioned embodiments, each component may be formed by dedicated hardware, or may be implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.
[0760] Furthermore, the present disclosure can be implemented as a three-dimensional data encoding method or a three-dimensional data decoding method, etc., which are executed by a three-dimensional data encoding device or a three-dimensional data decoding device, etc.
[0761] Furthermore, the division of the functional blocks in the block diagram is an example, and multiple functional blocks can be implemented as one functional block, and one functional block can also be divided into multiple blocks, and a part of the functions can also be moved to other functional blocks. Furthermore, the functions of multiple functional blocks with similar functi...
Claims
1. A three-dimensional data encoding method, wherein: generating an occupation pattern of adjacent nodes adjacent to a node included in an N-ary tree structure of a plurality of three-dimensional points included in three-dimensional data, wherein N is an integer greater than or equal to 2, Whether to set a candidate node that can use the first encoding is determined based on the occupancy pattern.
2. The three-dimensional data encoding method according to claim 1, wherein: The candidate node is not a root node in the N-ary tree structure.
3. The three-dimensional data encoding method according to claim 1, wherein: The number of points contained in the parent node of the candidate node is used in determining whether to set the candidate node.
4. The three-dimensional data encoding method according to claim 1, wherein: encoding parameters indicating neighboring nodes that are the subject of the occupancy pattern, A bitstream including the encoded parameters is generated.
5. The three-dimensional data encoding method according to claim 1, wherein: When it is determined not to set the candidate node, the node is encoded using a second encoding for dividing the node into a plurality of child nodes.
6. The three-dimensional data encoding method according to claim 1, wherein: The first encoding is a direct mode.
7. A three-dimensional data decoding method, wherein: Obtaining an occupation pattern of adjacent nodes adjacent to a node included in an N-ary tree structure of a plurality of three-dimensional points, wherein the plurality of three-dimensional points are included in the three-dimensional data, and N is an integer greater than or equal to 2, Whether to set a candidate node that can use the first decoding is determined based on the occupancy pattern.
8. The three-dimensional data decoding method according to claim 7, wherein: The candidate node is not a root node in the N-ary tree structure.
9. The three-dimensional data decoding method according to claim 7, wherein: The number of points contained in the parent node of the candidate node is used in determining whether to set the candidate node.
10. The three-dimensional data decoding method according to claim 7, wherein: When it is determined not to set the candidate node, the node is decoded using a second decoding method of dividing the node into a plurality of child nodes.
11. The three-dimensional data decoding method according to claim 7, wherein: The first decoding is a direct mode.
12. The three-dimensional data decoding method according to claim 7, wherein: Get parameters from the bitstream, The adjacent nodes that become the object of the occupation pattern change according to the value indicated by the parameter.
13. A three-dimensional data encoding device, wherein: have: processor; as well as Memory, The processor uses the memory, generating an occupation pattern of adjacent nodes adjacent to a node included in an N-ary tree structure of a plurality of three-dimensional points included in three-dimensional data, wherein N is an integer greater than or equal to 2, Whether to set a candidate node that can use the first encoding is determined based on the occupancy pattern.
14. A three-dimensional data decoding device, wherein: have: Processor; and Memory, The processor uses the memory, Obtaining an occupation pattern of adjacent nodes adjacent to a node included in an N-ary tree structure of a plurality of three-dimensional points, wherein the plurality of three-dimensional points are included in the three-dimensional data, and N is an integer greater than or equal to 2, Whether to set a candidate node that can use the first decoding is determined based on the occupancy pattern.
Citation Information
Patent Citations
Map display device
WO2014020663A1