Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device and three-dimensional data decoding device
By encoding multiple segmented data separately and generating a bit stream containing identifiers, the problem of large processing volume of the three-dimensional data decoding device is solved, and a more efficient decoding process is achieved.
Patent Information
- Application Number
- CN201980040789.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-08-10
- Filing Date
- 2019-08-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2039-08-09
AI Technical Summary
In the encoding and decoding process of three-dimensional data, it is difficult for the prior art to effectively reduce the processing amount of the three-dimensional data decoding device.
By encoding multiple segmented data separately, a bit stream of corresponding coded data and control information is generated, and an identifier is saved in the control information so that the object space can be easily restored during decoding.
This method can significantly reduce the processing amount of the three-dimensional data decoding device and improve the decoding efficiency.
Smart Images

Figure CN112313709B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device and a three-dimensional data decoding device. Background Art
[0002] In the future, devices and services that make use of 3D data will become more common in large fields such as computer vision, map information, monitoring, infrastructure inspection, or image distribution, which are used for autonomous operation of cars or robots. 3D data is obtained by various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple single-lens reflex cameras.
[0003] As one of the methods of expressing three-dimensional data, there is a method called point cloud, which expresses the shape of a three-dimensional structure through a group of points in a three-dimensional space. The position and color of the point group are saved in the point cloud. Although point cloud is expected to become the mainstream method of expressing three-dimensional data, the amount of point group data is very large. Therefore, in the accumulation or transmission of three-dimensional data, it is necessary to compress the data volume through encoding, just like two-dimensional moving images (for example, MPEG-4AVC or HEVC standardized by MPEG).
[0004] In addition, compression of point clouds is partially supported by a public library (PointCloud Library) that performs point cloud association processing.
[0005] Furthermore, there is known a technique for searching for facilities located around a vehicle using three-dimensional map data and displaying the facilities (for example, see Patent Document 1).
[0006] Prior art literature
[0007] Patent Literature
[0008] Patent Document 1: International Publication No. 2014 / 020663 Summary of the invention
[0009] Problems to be solved by the invention
[0010] In the encoding and decoding of three-dimensional data, it is desirable to reduce the amount of processing in a three-dimensional data decoding device.
[0011] An object of the present invention is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device or a three-dimensional data decoding device capable of reducing the amount of processing in a three-dimensional data decoding device.
[0012] Means used to solve problems
[0013] A three-dimensional data encoding method according to a technical solution of the present invention comprises the following steps: encoding a plurality of segmentation data respectively to generate a plurality of encoded data respectively corresponding to the plurality of segmentation data, wherein the plurality of segmentation data are contained in a plurality of subspaces obtained by dividing an object space including a plurality of three-dimensional points, and each of the subspaces includes one or more three-dimensional points; generating a bit stream including the plurality of encoded data and a plurality of control information respectively corresponding to the plurality of encoded data; and storing a first identifier and a second identifier in each of the plurality of control information, wherein the first identifier indicates a subspace corresponding to the encoded data corresponding to the control information, and the second identifier indicates the segmentation data corresponding to the encoded data corresponding to the control information.
[0014] A three-dimensional data decoding method according to a technical solution of the present invention obtains a first identifier and a second identifier stored in the plurality of control information from a bit stream including a plurality of coded data and a plurality of control information respectively corresponding to the plurality of coded data, wherein the plurality of coded data are generated by respectively encoding a plurality of segmentation data, the plurality of segmentation data are included in a plurality of subspaces obtained by segmenting an object space including a plurality of three-dimensional points, and each includes more than one three-dimensional point, the first identifier indicates a subspace corresponding to the coded data corresponding to the control information, and the second identifier indicates segmentation data corresponding to the coded data corresponding to the control information; the plurality of segmentation data are restored by decoding the plurality of coded data; and the object space is restored by combining the plurality of segmentation data using the first identifier and the second identifier.
[0015] Effects of the Invention
[0016] The present invention can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device capable of reducing the amount of processing in a three-dimensional data decoding device. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a diagram showing the structure of a three-dimensional data encoding and decoding system according to the first embodiment.
[0018] Figure 2 This is a diagram showing an example of the structure of point cloud data related to Implementation Example 1.
[0019] Figure 3 This is a diagram showing an example of the structure of a data file recording point group data information related to Implementation Example 1.
[0020] Figure 4 This is a diagram showing types of point cloud data related to Implementation Example 1.
[0021] Figure 5 This is a diagram showing the structure of the first encoding unit according to the first embodiment.
[0022] Figure 6 This is a block diagram of the first encoding unit related to implementation mode 1.
[0023] Figure 7 This is a diagram showing the structure of the first decoding unit according to the first embodiment.
[0024] Figure 8 This is a block diagram of the first decoding unit related to implementation mode 1.
[0025] Fig. 9 This is a diagram showing the structure of the second encoding unit related to Implementation Example 1.
[0026] Fig.10 This is a block diagram of the second encoding unit related to implementation mode 1.
[0027] Fig.11 This is a diagram showing the structure of the second decoding unit related to Implementation Example 1.
[0028] Fig.12 This is a block diagram of the second decoding unit related to implementation mode 1.
[0029] Fig.13 This is a diagram showing a protocol stack regarding PCC coded data according to the first embodiment.
[0030] Fig.14 This is a block diagram of the encoding unit related to implementation mode 1.
[0031] Fig.15 This is a block diagram of a decoding unit related to implementation mode 1.
[0032] Fig.16 This is a flowchart of the encoding process related to implementation mode 1.
[0033] Fig.17 This is a flowchart of the decoding process related to implementation mode 1.
[0034] Fig.18 This is a diagram showing the basic structure of ISOBMFF related to implementation example 2.
[0035] Fig.19 This is a diagram showing a protocol stack related to Implementation Example 2.
[0036] Fig. 20 This is a diagram showing an example of storing NAL units in a file for codec 1 according to the second embodiment.
[0037] Fig.21 This is a diagram showing an example of storing the NAL unit related to Implementation Example 2 in a file for codec 2.
[0038] Fig. 22 This is a diagram showing the structure of the first multiplexing unit related to Implementation Example 2.
[0039] Fig.23 This is a diagram showing the structure of the first inverse multiplexing unit related to Implementation Example 2.
[0040] Fig.24 This is a diagram showing the structure of the second multiplexing unit related to Implementation Example 2.
[0041] Fig.25 This is a diagram showing the structure of the second inverse multiplexing unit related to Implementation Example 2.
[0042] Fig.26 This is a flowchart of the processing performed by the first multiplexing unit in the second implementation mode.
[0043] Fig. 27 This is a flowchart of the processing performed by the second multiplexing unit in implementation mode 2.
[0044] Fig.28 This is a flowchart of the processing performed by the first demultiplexing unit and the first decoding unit in accordance with the second embodiment.
[0045] Fig.29 This is a flowchart showing the processing performed by the second demultiplexing unit and the second decoding unit related to implementation mode 2.
[0046] Fig.30 This is a diagram showing the structure of the encoding unit and the third multiplexing unit related to the third implementation mode.
[0047] Fig.31 This is a diagram showing the structure of the third inverse multiplexing unit and decoding unit related to implementation mode 3.
[0048] Fig.32 This is a flowchart of the processing performed by the third multiplexing unit in implementation mode 3.
[0049] Fig.33 This is a flowchart of the processing performed by the third demultiplexing unit and decoding unit in accordance with the third implementation mode.
[0050] Fig.34 This is a flowchart of the processing performed by the three-dimensional data storage device according to the third embodiment.
[0051] Fig.35 This is a flowchart of the processing performed by the three-dimensional data acquisition device according to the third embodiment.
[0052] Fig.36 This is a diagram showing the structure of the encoding unit and the multiplexing unit related to Implementation Example 4.
[0053] Fig.37This is a diagram showing an example of the structure of encoded data related to Implementation Example 4.
[0054] Fig.38 This is a diagram showing an example of the structure of coded data and NAL units related to Implementation Example 4.
[0055] Fig.39 This is a diagram showing a semantic example of pcc_nal_unit_type related to Implementation Example 4.
[0056] Fig.40 This is a diagram showing an example of the order in which NAL units are sent according to the fourth embodiment.
[0057] Fig.41 This is a flowchart of the processing performed by the three-dimensional data encoding device according to the fourth embodiment.
[0058] Fig.42 This is a flowchart of the processing performed by the three-dimensional data decoding device according to the fourth embodiment.
[0059] Fig.43 This is a flowchart of the multiplexing process related to implementation mode 4.
[0060] Fig.44 This is a flowchart of the inverse multiplexing processing related to implementation mode 4.
[0061] Fig.45 This is a flowchart of the processing performed by the three-dimensional data encoding device according to the fourth embodiment.
[0062] Fig.46 This is a flowchart of the processing performed by the three-dimensional data decoding device according to the fourth embodiment.
[0063] Fig.47 This is a block diagram of the first encoding unit related to implementation mode 5.
[0064] Fig.48 This is a block diagram of the first decoding unit related to implementation mode 5.
[0065] Fig.49 This is a block diagram of a division unit according to the fifth embodiment.
[0066] Fig.50 This is a diagram showing an example of dividing slices and tiles according to the fifth embodiment.
[0067] Fig.51 This is a diagram showing an example of a division pattern of slices and tiles according to the fifth embodiment.
[0068] Fig.52 This is a diagram showing an example of dependency relationship related to implementation mode 5.
[0069] Fig.53 This is a diagram showing an example of the decoding order of data related to Implementation Example 5.
[0070] Fig.54 This is a flowchart of the encoding process related to implementation mode 5.
[0071] Fig.55 This is a block diagram of the combining part related to implementation mode 5.
[0072] Fig.56 This is a diagram showing an example of the structure of coded data and NAL units related to implementation mode 5.
[0073] Fig.57 This is a flowchart of the encoding process related to implementation mode 5.
[0074] Fig.58 This is a flowchart of the decoding process related to implementation mode 5.
[0075] Fig.59 This is a flowchart of the encoding process related to implementation mode 5.
[0076] Fig.60 This is a flowchart of the decoding process related to implementation mode 5. DETAILED DESCRIPTION
[0077] A three-dimensional data encoding method according to a technical solution of the present invention comprises the following steps: encoding a plurality of segmentation data respectively to generate a plurality of encoded data respectively corresponding to the plurality of segmentation data, wherein the plurality of segmentation data are contained in a plurality of subspaces obtained by dividing an object space containing a plurality of three-dimensional points, and each of the subspaces contains one or more three-dimensional points; generating a bit stream containing the plurality of encoded data and a plurality of control information respectively corresponding to the plurality of encoded data; and storing a first identifier and a second identifier in each of the plurality of control information, wherein the first identifier indicates a subspace corresponding to the encoded data corresponding to the control information, and the second identifier indicates the segmentation data corresponding to the encoded data corresponding to the control information.
[0078] Thus, a 3D data decoding device that decodes a bit stream generated by the 3D data encoding method can easily restore the object space by combining data of a plurality of divided data using the first identifier and the second identifier. Therefore, the amount of processing in the 3D data decoding device can be reduced.
[0079] For example, in the above-mentioned encoding, the position information and attribute information of the three-dimensional point contained in each of the above-mentioned multiple segmentation data are encoded; the above-mentioned multiple encoded data respectively include the encoded data of the above-mentioned position information and the encoded data of the above-mentioned attribute information; the above-mentioned multiple control information respectively include the control information of the encoded data of the above-mentioned position information and the control information of the encoded data of the above-mentioned attribute information; the above-mentioned first identifier and the above-mentioned second identifier are stored in the control information of the encoded data of the above-mentioned position information.
[0080] For example, in the bit stream, the plurality of control information may be respectively arranged before the coded data corresponding to the control information.
[0081] A three-dimensional data decoding method according to a technical solution of the present invention obtains a first identifier and a second identifier stored in the plurality of control information from a bit stream including a plurality of coded data and a plurality of control information respectively corresponding to the plurality of coded data, wherein the plurality of coded data are generated by respectively encoding a plurality of segmentation data, the plurality of segmentation data are included in a plurality of subspaces obtained by segmenting an object space including a plurality of three-dimensional points, and each includes more than one three-dimensional point, the first identifier indicates a subspace corresponding to the coded data corresponding to the control information, and the second identifier indicates segmentation data corresponding to the coded data corresponding to the control information; the plurality of segmentation data are restored by decoding the plurality of coded data; and the object space is restored by combining the plurality of segmentation data using the first identifier and the second identifier.
[0082] Thus, the three-dimensional data decoding method can easily restore the object space by combining the data of a plurality of divided data using the first identifier and the second identifier. Therefore, the amount of processing in the three-dimensional data decoding device can be reduced.
[0083] For example, it may also be that the above-mentioned multiple encoded data are generated by encoding the position information and attribute information of the three-dimensional points contained in the corresponding segmentation data, respectively, including the encoded data of the above-mentioned position information and the encoded data of the above-mentioned attribute information; the above-mentioned multiple control information respectively include the control information of the encoded data of the above-mentioned position information and the control information of the encoded data of the above-mentioned attribute information; the above-mentioned first identifier and the above-mentioned second identifier are stored in the control information of the encoded data of the above-mentioned position information.
[0084] For example, in the bit stream, the control information may be arranged before the corresponding coded data.
[0085] In addition, a three-dimensional data encoding device according to a technical solution of the present invention comprises a processor and a memory; the processor uses the memory to perform the following processing: by respectively encoding a plurality of segmented data, a plurality of encoded data respectively corresponding to the plurality of segmented data are generated, the plurality of segmented data are contained in a plurality of subspaces obtained by dividing an object space including a plurality of three-dimensional points, and each includes one or more three-dimensional points; a bit stream including the plurality of encoded data and a plurality of control information respectively corresponding to the plurality of encoded data is generated; in each of the plurality of control information, a first identifier and a second identifier are stored, the first identifier represents the subspace corresponding to the encoded data corresponding to the control information, and the second identifier represents the segmented data corresponding to the encoded data corresponding to the control information.
[0086] Thus, a 3D data decoding device that decodes the bit stream generated by the 3D data encoding device can easily restore the object space by combining the data of the plurality of divided data using the first identifier and the second identifier. Therefore, the amount of processing in the 3D data decoding device can be reduced.
[0087] In addition, a three-dimensional data decoding device according to a technical solution of the present invention includes a processor and a memory; the processor uses the memory to perform the following processing: obtaining a first identifier and a second identifier stored in the multiple control information from a bit stream including multiple encoded data and multiple control information corresponding to the multiple encoded data, the multiple encoded data being generated by encoding multiple segmented data respectively, the multiple segmented data being included in multiple subspaces obtained by segmenting an object space including multiple three-dimensional points, and each including more than one three-dimensional point, the first identifier representing the subspace corresponding to the encoded data corresponding to the control information, and the second identifier representing the segmented data corresponding to the encoded data corresponding to the control information; restoring the multiple segmented data by decoding the multiple encoded data; and restoring the object space by combining the multiple segmented data using the first identifier and the second identifier.
[0088] Thus, the three-dimensional data decoding device can easily restore the target space by combining the data of a plurality of divided data using the first identifier and the second identifier. Therefore, the amount of processing in the three-dimensional data decoding device can be reduced.
[0089] In addition, these inclusive or specific technical solutions may also be implemented by systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or by any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0090] Hereinafter, the embodiments are described in detail with reference to the accompanying drawings. In addition, the embodiments described below all represent a specific example of the present invention. The numerical values, shapes, materials, constituent elements, configuration positions and connection forms of constituent elements, steps, and the order of steps shown in the following embodiments are examples and are not intended to limit the present invention. In addition, among the constituent elements of the following embodiments, constituent elements that are not described in the independent claims representing the highest concept are described as arbitrary constituent elements.
[0091] (Implementation Method 1)
[0092] When point cloud coded data is used in actual devices or services, it is desirable to transmit and receive required information according to the application in order to reduce network bandwidth. However, such a function does not exist in the existing 3D data coding structure, and therefore there is no corresponding coding method.
[0093] In this embodiment, a three-dimensional data encoding method and a three-dimensional data encoding device are described for providing the function of sending and receiving required information according to the purpose in the encoded data of a three-dimensional point cloud, a three-dimensional data decoding method and a three-dimensional data decoding device are described for decoding the encoded data, a three-dimensional data multiplexing method for multiplexing the encoded data, and a three-dimensional data transmission method for transmitting the encoded data.
[0094] In particular, currently, the first encoding method and the second encoding method are studied as encoding methods (encoding methods) for point group data, but the composition of the encoded data and the method of saving the encoded data in the system format are not defined. In this case, there is a problem that MUX processing (multiplexing) in the encoding unit, or transmission or accumulation cannot be directly performed.
[0095] Furthermore, there is no method that supports a format in which two codecs, namely the first encoding method and the second encoding method, exist in a mixed form, such as PCC (Point Cloud Compression).
[0096] In this embodiment, a method of configuring PCC coded data in which two codecs, namely, a first coding method and a second coding method, are mixed and storing the coded data in a system format is described.
[0097] First, the configuration of the three-dimensional data (point cloud data) encoding and decoding system according to the present embodiment will be described. Figure 1 2 is a diagram showing an example of the configuration of a three-dimensional data encoding and decoding system according to the present embodiment. Figure 1As shown, the three-dimensional data encoding and decoding system includes a three-dimensional data encoding system 4601, a three-dimensional data decoding system 4602, a sensor terminal 4603 and an external connection unit 4604.
[0098] The three-dimensional data encoding system 4601 generates encoded data or multiplexed data by encoding point group data as three-dimensional data. In addition, the three-dimensional data encoding system 4601 can be a three-dimensional data encoding device implemented by a single device, or a system implemented by multiple devices. In addition, the three-dimensional data encoding device can also be included in a part of the multiple processing units included in the three-dimensional data encoding system 4601.
[0099] The three-dimensional data encoding system 4601 includes a point cloud data generating system 4611, a presentation unit 4612, an encoding unit 4613, a multiplexing unit 4614, an input / output unit 4615, and a control unit 4616. The point cloud data generating system 4611 includes a sensor information acquiring unit 4617 and a point cloud data generating unit 4618.
[0100] The sensor information acquisition unit 4617 acquires sensor information from the sensor terminal 4603 and outputs the sensor information to the point cloud data generation unit 4618. The point cloud data generation unit 4618 generates point cloud data based on the sensor information and outputs the point cloud data to the encoding unit 4613.
[0101] The presentation unit 4612 presents the sensor information or point group data to the user. For example, the presentation unit 4612 displays information or an image based on the sensor information or point group data.
[0102] The encoding unit 4613 encodes (compresses) the point cloud data, and outputs the obtained encoded data, control information obtained in the encoding process, and other additional information to the multiplexing unit 4614. The additional information includes, for example, sensor information.
[0103] The multiplexing unit 4614 generates multiplexed data by multiplexing the coded data, control information, and additional information input from the coding unit 4613. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.
[0104] The input / output unit 4615 (e.g., a communication unit or an interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as an internal memory. The control unit 4616 (or an application execution unit) controls each processing unit. That is, the control unit 4616 performs control such as encoding and multiplexing.
[0105] Furthermore, the sensor information may be input to the encoding unit 4613 or the multiplexing unit 4614. Furthermore, the input / output unit 4615 may directly output the point cloud data or the encoded data to the outside.
[0106] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 4601 is input to the three-dimensional data decoding system 4602 via the external connection unit 4604 .
[0107] The three-dimensional data decoding system 4602 generates point group data as three-dimensional data by decoding the coded data or the multiplexed data. In addition, the three-dimensional data decoding system 4602 may be a three-dimensional data decoding device implemented by a single device, or may be a system implemented by multiple devices. In addition, the three-dimensional data decoding device may also include a part of the multiple processing units included in the three-dimensional data decoding system 4602.
[0108] The three-dimensional data decoding system 4602 includes a sensor information acquisition unit 4621 , an input / output unit 4622 , an inverse multiplexing unit 4623 , a decoding unit 4624 , a prompting unit 4625 , a user interface 4626 , and a control unit 4627 .
[0109] The sensor information acquisition unit 4621 acquires sensor information from the sensor terminal 4603 .
[0110] The input / output unit 4622 obtains the transmission signal, decodes the multiplexed data (file format or packet) according to the transmission signal, and outputs the multiplexed data to the demultiplexing unit 4623 .
[0111] The inverse multiplexing unit 4623 obtains the encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 4624 .
[0112] The decoding unit 4624 reconstructs the point cloud data by decoding the encoded data.
[0113] The prompt unit 4625 prompts the user with the point group data. For example, the prompt unit 4625 displays information or images based on the point group data. The user interface 4626 obtains instructions based on the user's operation. The control unit 4627 (or the application execution unit) controls each processing unit. That is, the control unit 4627 controls demultiplexing, decoding, and prompting.
[0114] In addition, the input / output unit 4622 may directly obtain point group data or coded data from the outside. In addition, the prompt unit 4625 may obtain additional information such as sensor information and prompt information based on the additional information. In addition, the prompt unit 4625 may also perform prompts based on the user's instructions obtained by the user interface 4626.
[0115] The sensor terminal 4603 generates sensor information, which is information acquired by a sensor. The sensor terminal 4603 is a terminal equipped with a sensor or a camera, and includes, for example, a moving object such as a car, a flying object such as an airplane, a mobile terminal or a camera.
[0116] The sensor information that can be obtained by the sensor terminal 4603 is, for example, (1) the distance between the sensor terminal 4603 and the object, or the reflectivity of the object, obtained by LIDAR, millimeter wave radar, or infrared sensor, (2) the distance between the camera and the object, or the reflectivity of the object, obtained from multiple monocular camera images or stereo camera images. In addition, the sensor information may also include the posture, direction, rotation (angular velocity), position (GPS information or altitude), speed or acceleration of the sensor. In addition, the sensor information may also include temperature, air pressure, humidity, or magnetism.
[0117] The external connection unit 4604 is implemented by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the Internet, or broadcasting.
[0118] Next, point group data will be described. Figure 2 It is a diagram showing the structure of point group data. Figure 3 It is a diagram showing a configuration example of a data file that describes information on point group data.
[0119] Point group data includes data of multiple points. The data of each point includes position information (three-dimensional coordinates) and attribute information corresponding to the position information. A group of multiple such points is called a point group. For example, a point group represents the three-dimensional shape of an object.
[0120] Sometimes, position information (Position) such as three-dimensional coordinates is also called geometry (Geometry). In addition, the data of each point can also include attribute information (attribute) of multiple attribute categories. Attribute categories are, for example, color or reflectivity.
[0121] One piece of attribute information may be associated with one piece of location information, or a plurality of pieces of attribute information having different attribute categories may be associated with one piece of location information. In addition, a plurality of pieces of attribute information having the same attribute category may be associated with one piece of location information.
[0122] Figure 3 The illustrated data file configuration example is an example of a case where position information and attribute information correspond to each other on a one-to-one basis, and indicates position information and attribute information of N points constituting point cloud data.
[0123] The position information is, for example, information on three axes, x, y, and z. The attribute information is, for example, color information of RGB. Representative data files include ply files and the like.
[0124] Next, types of point cloud data will be described. Figure 4 is a graph showing the types of point group data. Figure 4As shown, the point group data includes static objects and dynamic objects.
[0125] A static object is a 3D point group data at any time (a certain moment). A dynamic object is a 3D point group data that changes over time. Hereinafter, the 3D point group data at a certain moment is referred to as a PCC frame or frame.
[0126] The object may be a point group whose area is limited to a certain extent, such as normal image data, or a large-scale point group whose area is not limited, such as map information.
[0127] In addition, there are point group data of various densities, and there may be sparse point group data and dense point group data.
[0128] The details of each processing unit are described below. The sensor information is obtained by various methods such as a distance sensor such as LIDAR or a rangefinder, a stereo camera, or a combination of multiple monocular cameras. The point group data generation unit 4618 generates point group data based on the sensor information obtained by the sensor information acquisition unit 4617. The point group data generation unit 4618 generates position information as point group data and adds attribute information for the position information to the position information.
[0129] The point group data generation unit 4618 may also process the point group data when generating the position information or the additional attribute information. For example, the point group data generation unit 4618 may also reduce the amount of data by deleting point groups with repeated positions. In addition, the point group data generation unit 4618 may also transform the position information (position conversion, rotation or standardization, etc.) and may also render the attribute information.
[0130] In addition, Figure 1 In the figure, the point group data generating system 4611 is included in the three-dimensional data encoding system 4601, but can also be independently set outside the three-dimensional data encoding system 4601.
[0131] The coding unit 4613 encodes the point cloud data based on a predetermined coding method, thereby generating coded data. There are generally two types of coding methods. The first type is a coding method using position information, which is hereinafter referred to as the first coding method. The second type is a coding method using a video codec, which is hereinafter referred to as the second coding method.
[0132] The decoding unit 4624 decodes the encoded data based on a predetermined encoding method, thereby decoding the point cloud data.
[0133] The multiplexing unit 4614 generates multiplexed data by multiplexing the coded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to the PCC coded data, the multiplexing unit 4614 also multiplexes other media such as images, sounds, subtitles, applications, files, or reference time information. In addition, the multiplexing unit 4614 can also multiplex attribute information associated with sensor information or point group data.
[0134] As multiplexing methods or file formats, there are ISOBMFF, MPEG-DASH which is a transmission method based on ISOBMFF, MMT, MPEG-2TS Systems, RMP, etc.
[0135] The demultiplexing unit 4623 extracts PCC coded data, other media, time information, etc. from the multiplexed data.
[0136] The input / output unit 4615 transmits the multiplexed data using a method consistent with a transmission medium or storage medium such as broadcasting or communication. The input / output unit 4615 can communicate with other devices via the Internet, or can communicate with a storage unit such as a cloud server.
[0137] As the communication protocol, http, ftp, TCP, UDP, etc. can be used. Either a PULL type communication method or a PUSH type communication method can be used.
[0138] Any of wired transmission and wireless transmission can be used. As wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), or coaxial cable, etc. are used. As wireless transmission, wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), or millimeter wave, etc. are used.
[0139] In addition, as a broadcasting method, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 is used.
[0140] Figure 5 This diagram shows the configuration of a first coding unit 4630 which is an example of the coding unit 4613 that performs coding using the first coding method. Figure 6 4 is a block diagram of the first coding unit 4630. The first coding unit 4630 generates coded data (coded stream) by coding the point cloud data using the first coding method. The first coding unit 4630 includes a position information coding unit 4631, an attribute information coding unit 4632, an additional information coding unit 4633, and a multiplexing unit 4634.
[0141] The first coding unit 4630 is characterized in that the coding is performed with awareness of the three-dimensional structure. In addition, the first coding unit 4630 is characterized in that the attribute information coding unit 4632 performs coding using information obtained from the position information coding unit 4631. The first coding method is also called GPCC (Geometry based PCC).
[0142] The point group data is PCC point group data such as a PLY file, or PCC point group data generated based on sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData). The position information is input to the position information encoding unit 4631, the attribute information is input to the attribute information encoding unit 4632, and the additional information is input to the additional information encoding unit 4633.
[0143] The position information encoding unit 4631 generates encoded position information (Compressed Geometry) as encoded data by encoding the position information. For example, the position information encoding unit 4631 uses an N-ary tree structure such as an octree to encode the position information. Specifically, in the octree, the object space is divided into 8 nodes (subspaces), and 8 bits of information (occupancy code) are generated to indicate whether each node contains a point group. In addition, the node containing the point group is further divided into 8 nodes, and 8 bits of information are generated to indicate whether each of the 8 nodes contains a point group. This process is repeated until it becomes below the threshold of the number of point groups contained in a predetermined layer or node.
[0144] The attribute information encoding unit 4632 generates the encoded attribute information (Compressed Attribute) as the encoded data by encoding using the composition information generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 determines the reference point (reference node) to be referred to in the encoding of the object point (object node) of the processing object based on the octree structure generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 refers to a node whose parent node in the octree is the same as the parent node of the object node among the surrounding nodes or adjacent nodes. In addition, the method of determining the reference relationship is not limited to this.
[0145] In addition, the encoding process of the attribute information may include at least one of a quantization process, a prediction process, and an arithmetic coding process. In this case, the reference means using the reference node in the calculation of the predicted value of the attribute information, or using the state of the reference node in the determination of the encoding parameter (for example, indicating whether the occupancy information of the point group is included in the reference node). For example, the encoding parameter is a quantization parameter in the quantization process, or a context in the arithmetic coding, etc.
[0146] The additional information encoding unit 4633 generates encoded additional information (Compressed MetaData) as encoded data by encoding compressible data in the additional information.
[0147] The multiplexing unit 4634 generates a coded stream (Compressed Stream) as coded data by multiplexing the coding position information, the coding attribute information, the coding additional information, and other additional information. The generated coded stream is output to a processing unit of the system layer (not shown).
[0148] Next, the first decoding unit 4640 which is an example of the decoding unit 4624 that performs decoding according to the first encoding method is described. Figure 7 This is a diagram showing the structure of the first decoding unit 4640. Figure 8 46 is a block diagram of a first decoding unit 4640. The first decoding unit 4640 generates point cloud data by decoding coded data (coded stream) coded by the first coding method by the first coding method. The first decoding unit 4640 includes a demultiplexing unit 4641, a position information decoding unit 4642, an attribute information decoding unit 4643, and an additional information decoding unit 4644.
[0149] A coded stream (Compressed Stream) as coded data is input to the first decoding unit 4640 from a processing unit of a system layer (not shown).
[0150] The inverse multiplexing unit 4641 separates the encoded position information (Compressed Geometry), the encoded attribute information (Compressed Attribute), the encoded additional information (Compressed MetaData), and other additional information from the encoded data.
[0151] The position information decoding unit 4642 generates position information by decoding the encoded position information. For example, the position information decoding unit 4642 restores the position information of the point group represented by three-dimensional coordinates based on the encoded position information represented by an N-ary tree structure such as an octree.
[0152] The attribute information decoding unit 4643 decodes the encoded attribute information based on the composition information generated by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 determines a reference point (reference node) to be referred to in decoding the object point (object node) of the processing object based on the octree structure obtained by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 refers to a node whose parent node in the octree is the same as the parent node of the object node among the surrounding nodes or adjacent nodes. In addition, the method of determining the reference relationship is not limited to this.
[0153] In addition, the decoding process of the attribute information may also include at least one of an inverse quantization process, a prediction process, and an arithmetic decoding process. In this case, the reference means that the reference node is used in the calculation of the predicted value of the attribute information, or the state of the reference node is used in the determination of the decoded parameter (for example, indicating whether the occupancy information of the point group is included in the reference node). For example, the decoded parameter is a quantization parameter in the inverse quantization process, or a context in the arithmetic decoding, etc.
[0154] The additional information decoding unit 4644 generates additional information by decoding the encoded additional information. In addition, the first decoding unit 4640 uses the additional information required for decoding processing of the position information and the attribute information during decoding, and outputs the additional information required for the application to the outside.
[0155] Next, the second encoding unit 4650 which is an example of the encoding unit 4613 that performs encoding using the second encoding method is described. Fig. 9 This is a diagram showing the structure of the second encoding unit 4650. Fig.10 This is a block diagram of the second encoding unit 4650.
[0156] The second coding unit 4650 generates coded data (coded stream) by coding the point cloud data using the second coding method. The second coding unit 4650 includes an additional information generating unit 4651 , a position image generating unit 4652 , an attribute image generating unit 4653 , a video coding unit 4654 , an additional information coding unit 4655 , and a multiplexing unit 4656 .
[0157] The second coding unit 4650 generates a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and codes the generated position image and attribute image using an existing video coding method. The second coding method is also called VPCC (video based PCC).
[0158] The point group data is PCC point group data such as a PLY file or PCC point group data generated based on sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData).
[0159] The additional information generating unit 4651 generates mapping information of a plurality of two-dimensional images by projecting a three-dimensional structure onto a two-dimensional image.
[0160] The position image generation unit 4652 generates a position image (Geometry Image) based on the position information and the mapping information generated by the additional information generation unit 4651. The position image is, for example, a distance image representing the distance (Depth) as a pixel value. In addition, the distance image can be an image of multiple point groups observed from one viewpoint (an image of multiple point groups projected on one two-dimensional plane), or multiple images of multiple point groups observed from multiple viewpoints, or a single image formed by merging these multiple images.
[0161] The attribute image generation unit 4653 generates an attribute image based on the attribute information and the mapping information generated by the additional information generation unit 4651. The attribute image is, for example, an image that represents the attribute information (e.g., color (RGB)) as a pixel value. In addition, the image may be an image of a plurality of point groups observed from a single viewpoint (an image in which a plurality of point groups are projected on a single two-dimensional plane), or may be a plurality of images of a plurality of point groups observed from a plurality of viewpoints, or may be a single image formed by merging these plurality of images.
[0162] The image encoding unit 4654 encodes the position image and the attribute image using an image encoding method, thereby generating a compressed position image (Compressed Geometry Image) and a compressed attribute image (Compressed Attribute Image) as encoded data. In addition, as the image encoding method, any known encoding method can be used. For example, the image encoding method is AVC or HEVC.
[0163] The additional information encoding unit 4655 generates coded additional information (Compressed MetaData) by encoding the additional information and mapping information included in the point cloud data.
[0164] The multiplexing unit 4656 generates a coded stream (Compressed Stream) as coded data by multiplexing the coding position image, the coding attribute image, the coding additional information, and other additional information. The generated coded stream is output to a processing unit of the system layer (not shown).
[0165] Next, the second decoding unit 4660 which is an example of the decoding unit 4624 that performs decoding according to the second encoding method is described. Fig.11 This is a diagram showing the structure of the second decoding unit 4660. Fig.1246 is a block diagram of a second decoding unit 4660. The second decoding unit 4660 generates point cloud data by decoding coded data (coded stream) coded by the second coding method by the second coding method. The second decoding unit 4660 includes an inverse multiplexing unit 4661, a video decoding unit 4662, an additional information decoding unit 4663, a position information generating unit 4664, and an attribute information generating unit 4665.
[0166] A coded stream (Compressed Stream) as coded data is input to the second decoding unit 4660 from a processing unit of a system layer (not shown).
[0167] The inverse multiplexing unit 4661 separates the coded position image (Compressed Geometry Image), the coded attribute image (Compressed Attribute Image), the coded additional information (Compressed MetaData), and other additional information from the coded data.
[0168] The video decoding unit 4662 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding method. In addition, as the video encoding method, any known encoding method can be used. For example, the video encoding method is AVC or HEVC.
[0169] The additional information decoding unit 4663 generates additional information including mapping information and the like by decoding the encoded additional information.
[0170] The position information generating unit 4664 generates position information using the position image and the mapping information. The attribute information generating unit 4665 generates attribute information using the attribute image and the mapping information.
[0171] The second decoding unit 4660 uses the additional information required for decoding during decoding, and outputs the additional information required for the application to the outside.
[0172] The following describes the issues in the PCC coding method. Fig.13 It is a diagram showing a protocol stack related to PCC coded data. Fig.13 This shows an example of multiplexing, transmitting, or accumulating data of other media such as images (for example, HEVC) or audio into PCC coded data.
[0173] The multiplexing method and file format have functions for multiplexing, transmitting or accumulating various coded data. In order to transmit or accumulate coded data, the coded data must be converted into a format of the multiplexing method. For example, HEVC stipulates a technology for storing coded data in a data structure called a NAL unit and storing the NAL unit in ISOBMFF.
[0174] On the other hand, currently, as a coding method for point group data, the first coding method (Codec1) and the second coding method (Codec2) are being studied, but the composition of the coded data and the method of saving the coded data in the system format are not defined, and there is a problem that MUX processing (multiplexing), transmission and accumulation in the coding unit cannot be directly performed.
[0175] In the following, unless a specific encoding method is mentioned, either the first encoding method or the second encoding method is indicated.
[0176] The following is an explanation of the method for defining the NAL unit of the present embodiment. For example, in previous codecs such as HEVC, a NAL unit of one format is defined for one codec. However, there is no method to support a mixed format of two codecs (hereinafter referred to as PCC codecs) such as the first coding method and the second coding method, as in PCC.
[0177] First, the encoding unit 4670 having the functions of both the first encoding unit 4630 and the second encoding unit 4650 described above, and the decoding unit 4680 having the functions of both the first decoding unit 4640 and the second decoding unit 4660 are described.
[0178] Fig.14 4 is a block diagram of a coding unit 4670 according to the present embodiment. The coding unit 4670 includes the first coding unit 4630 and the second coding unit 4650 described above, and a multiplexing unit 4671. The multiplexing unit 4671 multiplexes the coded data generated by the first coding unit 4630 and the coded data generated by the second coding unit 4650, and outputs the obtained coded data.
[0179] Fig.15 This is a block diagram of a decoding unit 4680 related to the present embodiment. The decoding unit 4680 includes the first decoding unit 4640 and the second decoding unit 4660 described above and an inverse multiplexing unit 4681. The inverse multiplexing unit 4681 extracts the coded data using the first coding method and the coded data using the second coding method from the input coded data. The inverse multiplexing unit 4681 outputs the coded data using the first coding method to the first decoding unit 4640, and outputs the coded data using the second coding method to the second decoding unit 4660.
[0180] With the above configuration, the encoding unit 4670 can selectively encode the point cloud data using the first encoding method and the second encoding method. In addition, the decoding unit 4680 can decode the encoded data encoded using the first encoding method, the encoded data encoded using the second encoding method, and the encoded data encoded using both the first encoding method and the second encoding method.
[0181] For example, the coding unit 4670 may switch the coding method (the first coding method and the second coding method) in units of point cloud data or frames. In addition, the coding unit 4670 may switch the coding method in units of codable data.
[0182] The encoding unit 4670 generates, for example, encoded data (encoded stream) including identification information of the PCC codec.
[0183] The demultiplexing unit 4681 included in the decoding unit 4680 identifies the data using identification information of the PCC codec, for example. The demultiplexing unit 4681 outputs the data to the first decoding unit 4640 if the data is data encoded using the first encoding method, and outputs the data to the second decoding unit 4660 if the data is data encoded using the second encoding method.
[0184] Furthermore, the coding unit 4670 may send information indicating whether both coding methods or one of the coding methods is used as control information, in addition to the identification information of the PCC codec.
[0185] Next, the encoding process according to this embodiment will be described. Fig.16 2 is a flowchart of the encoding process according to the present embodiment. By using the identification information of the PCC codec, encoding processes corresponding to a plurality of codecs can be performed.
[0186] First, the encoding unit 4670 encodes the PCC data using a codec of one or both of the first encoding method and the second encoding method (S4681).
[0187] When the codec used is the second encoding method (the second encoding method in S4682), the encoder 4670 sets the pcc_codec_type included in the NAL unit header to a value indicating that the data included in the payload of the NAL unit is data encoded using the second encoding method (S4683). Next, the encoder 4670 sets the identifier of the NAL unit used for the second encoding method for the pcc_nal_unit_type of the NAL unit header (S4684). The encoder 4670 generates a NAL unit having the set NAL unit header and including the encoded data in the payload. The encoder 4670 transmits the generated NAL unit (S4685).
[0188] On the other hand, when the codec used is the first encoding method (the first encoding method in S4682), the encoder 4670 sets the pcc_codec_type included in the NAL unit header to a value indicating that the data included in the payload of the NAL unit is data encoded using the first encoding method (S4686). Next, the encoder 4670 sets the identifier of the NAL unit used for the first encoding method for the pcc_nal_unit_type of the NAL unit header (S4687). Next, the encoder 4670 generates a NAL unit having the set NAL unit header and including the encoded data in the payload. And, the encoder 4670 sends the generated NAL unit (S4685).
[0189] Next, the decoding process according to this embodiment will be described. Fig.17 This is a flowchart of the decoding process according to the present embodiment. By using the identification information of the PCC codec, it is possible to perform decoding processes corresponding to a plurality of codecs.
[0190] First, the decoding unit 4680 receives a NAL unit (S4691). For example, the NAL unit is generated by the process in the encoding unit 4670 described above.
[0191] Next, the decoding unit 4680 determines whether pcc_codec_type included in the NAL unit header indicates the first encoding method or the second encoding method ( S4692 ).
[0192] When pcc_codec_type indicates the second coding method (the second coding method in S4692), the decoding unit 4680 determines that the data included in the payload of the NAL unit is data encoded using the second coding method (S4693). Furthermore, the second decoding unit 4660 identifies the data by assuming that pcc_nal_unit_type included in the NAL unit header is an identifier of the NAL unit for the second coding method (S4694). Furthermore, the decoding unit 4680 decodes the PCC data using the decoding process of the second coding method (S4695).
[0193] On the other hand, when pcc_codec_type indicates the first coding method (the first coding method in S4692), the decoding unit 4680 determines that the data included in the payload of the NAL unit is data encoded using the first coding method (S4696). Furthermore, the decoding unit 4680 identifies the data by assuming that pcc_nal_unit_type included in the NAL unit header is an identifier of the NAL unit for the first coding method (S4697). Furthermore, the decoding unit 4680 decodes the PCC data using the decoding process of the first coding method (S4698).
[0194] As described above, a three-dimensional data encoding device according to a technical solution of the present invention generates a coding stream by encoding three-dimensional data (e.g., point group data), and stores information indicating a coding method used in the above coding, among the first coding method and the second coding method, in control information (e.g., a parameter set) of the above coding stream (e.g., identification information of a codec).
[0195] Thus, when decoding the coded stream generated by the 3D data coding device, the 3D data decoding device can use the information stored in the control information to determine the coding method used in the coding. Therefore, the 3D data decoding device can correctly decode the coded stream even when multiple coding methods are used.
[0196] For example, the three-dimensional data includes position information. The three-dimensional data encoding device encodes the position information in the encoding. The three-dimensional data encoding device stores information indicating the encoding method used in encoding the position information in the control information of the position information in the storage.
[0197] For example, the three-dimensional data includes position information and attribute information. The three-dimensional data encoding device encodes the position information and the attribute information in the encoding. The three-dimensional data encoding device stores information indicating the encoding method used in encoding the position information in the control information of the position information and the encoding method used in encoding the attribute information in the control information of the attribute information.
[0198] This allows different encoding methods to be used for position information and attribute information, thereby improving encoding efficiency.
[0199] For example, the three-dimensional data encoding method further stores the encoding stream in one or more units (eg, NAL units).
[0200] For example, the above-mentioned unit includes the following information (such as pcc_nal_unit_type), which has a common format in the above-mentioned first encoding method and the above-mentioned second encoding method, indicates the type of data included in the above-mentioned unit, and has independent definitions in the above-mentioned first encoding method and the above-mentioned second encoding method.
[0201] For example, the above-mentioned unit includes the following information (such as codec1_nal_unit_type or codec2_nal_unit_type), which has a format independent of the above-mentioned first encoding method and the above-mentioned second encoding method, indicates the type of data included in the above-mentioned unit, and has a definition independent of the above-mentioned first encoding method and the above-mentioned second encoding method.
[0202] For example, the above-mentioned unit includes the following information (such as pcc_nal_unit_type), which has a common format in the above-mentioned first encoding method and the above-mentioned second encoding method, indicates the type of data contained in the above-mentioned unit, and has a common definition in the above-mentioned first encoding method and the above-mentioned second encoding method.
[0203] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0204] In addition, the three-dimensional data decoding device related to the present embodiment determines the encoding method used in encoding the encoded stream based on information indicating the encoding method used in encoding the three-dimensional data among the first encoding method and the second encoding method (e.g., identification information of the codec) contained in control information (e.g., a parameter set) of the encoded stream generated by encoding the three-dimensional data, and decodes the encoded stream using the determined encoding method.
[0205] Thus, when decoding the coded stream, the 3D data decoding device can use the information stored in the control information to determine the coding method used in the coding. Therefore, the 3D data decoding device can correctly decode the coded stream even when multiple coding methods are used.
[0206] For example, the three-dimensional data includes position information, and the coded stream includes coded data of the position information. In the determination, the three-dimensional data decoding device determines the coding method used in coding the position information based on information indicating the coding method used in coding the position information, of the first coding method and the second coding method, included in the control information of the position information included in the coded stream. In the decoding, the three-dimensional data decoding device decodes the coded data of the position information using the coding method used in coding the determined position information.
[0207] For example, the three-dimensional data includes position information and attribute information, and the coded stream includes coded data of the position information and coded data of the attribute information. In the determination, the three-dimensional data decoding device determines the coding method used in coding the position information based on information indicating the coding method used in coding the position information between the first coding method and the second coding method, which is included in the control information of the position information included in the coded stream, and determines the coding method used in coding the attribute information based on information indicating the coding method used in coding the attribute information between the first coding method and the second coding method, which is included in the control information of the attribute information included in the coded stream. In the decoding, the three-dimensional data decoding device decodes the coded data of the position information using the coding method determined to be used in coding the position information, and decodes the coded data of the attribute information using the coding method determined to be used in coding the attribute information.
[0208] This allows different encoding methods to be used for position information and attribute information, thereby improving encoding efficiency.
[0209] For example, the coded stream is stored in one or more units (for example, NAL units), and the three-dimensional data decoding device further obtains the coded stream from the one or more units.
[0210] For example, the above-mentioned unit includes the following information (such as pcc_nal_unit_type), which has a common format in the above-mentioned first encoding method and the above-mentioned second encoding method, indicates the type of data included in the above-mentioned unit, and has independent definitions in the above-mentioned first encoding method and the above-mentioned second encoding method.
[0211] For example, the above-mentioned unit includes the following information (such as codec1_nal_unit_type or codec2_nal_unit_type), which has an independent format in the above-mentioned first encoding method and the above-mentioned second encoding method, indicates the type of data included in the above-mentioned unit, and has an independent definition in the above-mentioned first encoding method and the above-mentioned second encoding method.
[0212] For example, the above-mentioned unit includes the following information (such as pcc_nal_unit_type), which has a common format in the above-mentioned first encoding method and the above-mentioned second encoding method, indicates the type of data contained in the above-mentioned unit, and has a common definition in the above-mentioned first encoding method and the above-mentioned second encoding method.
[0213] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0214] (Implementation Method 2)
[0215] In this embodiment, a method of storing a NAL unit in an ISOBMFF file is described.
[0216] ISOBMFF (ISO based media file format) is a file format standard defined by ISO / IEC 14496-12. ISOBMFF defines a format that can multiplex and store various media such as video, audio, and text, and is a standard that is independent of the media.
[0217] The basic structure (file) of ISOBMFF is explained. The basic unit in ISOBMFF is a box. A box consists of type, length, and data. The combination of boxes of various types is called a file.
[0218] Fig.18 This is a diagram showing the basic structure (file) of ISOBMFF. The ISOBMFF file mainly includes boxes such as ftyp which indicates the file version (brand) using 4CC (4 character code), moov which stores metadata such as control information, and mdat which stores data.
[0219] The method of saving each media in the ISOBMFF file is separately specified. For example, the method of saving AVC video and HEVC video is specified by ISO / IEC14496-15. Here, in order to store or transmit PCC coded data, it is conceivable to use the ISOBMFF function extension, but there is no provision for saving PCC coded data in the ISOBMFF file. Therefore, in this embodiment, the method of saving PCC coded data in the ISOBMFF file is described.
[0220] Fig.19 This is a diagram showing a protocol stack when a NAL unit common to a PCC codec is saved in an ISOBMFF file. Here, a NAL unit common to a PCC codec is saved in an ISOBMFF file. NAL units are common to PCC codecs, but since multiple PCC codecs are saved in a NAL unit, it is desirable to define a saving method corresponding to each codec (Carriage of Codec1, Carriage of Codec2).
[0221] Next, a method of storing a common PCC NAL unit supporting a plurality of PCC codecs in an ISOBMFF file is described. Fig. 20This diagram shows an example of storing a common PCC NAL unit in a file of ISOBMFF of the storage method of codec 1 (Carriage of Codec1). Fig.21 This diagram shows an example of storing a common PCC NAL unit in an ISOBMFF file of the storage method (Carriage of Codec2) of codec 2.
[0222] Here, ftyp is important information for identifying the file format, and a different identifier for each codec is defined as ftyp. When PCC coded data encoded by the first coding method (coding mode) is saved in a file, ftyp=pcc1 is set. When PCC coded data encoded by the second coding method is saved in a file, ftyp=pcc2 is set.
[0223] Here, pcc1 indicates codec 1 using PCC (first encoding method). pcc2 indicates codec 2 using PCC (second encoding method). That is, pcc1 and pcc2 indicate that the data is PCC (coded data of three-dimensional data (point cloud data)), and indicate PCC codec (first encoding method and second encoding method).
[0224] The following describes a method of storing the NAL unit in an ISOBMFF file. The multiplexing unit parses the NAL unit header, and when pcc_codec_type=Codec1, pcc1 is written in the ISOBMFF ftyp.
[0225] Furthermore, the multiplexing unit parses the NAL unit header, and when pcc_codec_type=Codec2, pcc2 is recorded in ftyp of ISOBMFF.
[0226] Furthermore, when pcc_nal_unit_type is metadata, the multiplexing unit stores the NAL unit in, for example, moov or mdat using a predetermined method. When pcc_nal_unit_type is data, the multiplexing unit stores the NAL unit in, for example, moov or mdat using a predetermined method.
[0227] For example, the multiplexing unit may store the NAL unit size in the NAL unit similarly to HEVC.
[0228] By parsing the ftyp contained in the file in the inverse multiplexing unit (system layer) using this storage method, it is possible to determine whether the PCC coded data is encoded using the first coding method or the second coding method. Furthermore, as described above, by determining whether the PCC coded data is encoded using the first coding method or the second coding method, it is possible to extract coded data encoded using one of the coding methods from data in which coded data encoded using both coding methods are mixed. Thus, when transmitting coded data, the amount of data transmitted can be suppressed. In addition, by using this storage method, it is possible to use a common data format instead of setting different data (file) formats in the first coding method and the second coding method.
[0229] In addition, when the identification information of the codec is indicated in metadata of the system layer such as ftyp in ISOBMFF, the multiplexing unit may store the NAL unit after deleting pcc_nal_unit_type in the ISOBMFF file.
[0230] Next, the structure and operation of the multiplexing unit of the three-dimensional data encoding system (three-dimensional data encoding device) according to the present embodiment and the demultiplexing unit of the three-dimensional data decoding system (three-dimensional data decoding device) according to the present embodiment are described.
[0231] Fig. 22 47 is a diagram showing the structure of the first multiplexing unit 4710. The first multiplexing unit 4710 includes a file conversion unit 4711 that generates multiplexed data (file) by storing the coded data and control information (NAL unit) generated by the first coding unit 4630 in an ISOBMFF file. The first multiplexing unit 4710 includes, for example, Figure 1 In the multiplexing unit 4614 shown.
[0232] Fig.23 47 is a diagram showing the structure of the first demultiplexing unit 4720. The first demultiplexing unit 4720 includes a file inverse conversion unit 4721 that obtains coded data and control information (NAL unit) from the multiplexed data (file) and outputs the obtained coded data and control information to the first decoding unit 4640. The first demultiplexing unit 4720 includes, for example, Figure 1 In the inverse multiplexing unit 4623 shown.
[0233] Fig.24 4730 is a diagram showing the structure of the second multiplexing unit 4730. The second multiplexing unit 4730 includes a file conversion unit 4731 that generates multiplexed data (file) by storing the coded data and control information (NAL unit) generated by the second coding unit 4650 in an ISOBMFF file. The second multiplexing unit 4730 is included in, for example, Figure 1 In the multiplexing unit 4614 shown.
[0234] Fig.25 The second demultiplexing unit 4740 is a diagram showing the structure of the second demultiplexing unit 4740. The second demultiplexing unit 4740 includes a file inverse conversion unit 4741 that obtains coded data and control information (NAL unit) from the multiplexed data (file) and outputs the obtained coded data and control information to the second decoding unit 4660. The second demultiplexing unit 4740 includes, for example, Figure 1 In the inverse multiplexing unit 4623 shown.
[0235] Fig.26 This is a flowchart of the multiplexing process performed by the first multiplexing unit 4710. First, the first multiplexing unit 4710 analyzes pcc_codec_type included in the NAL unit header to determine whether the codec used is the first encoding method or the second encoding method (S4701).
[0236] When pcc_codec_type indicates the second encoding method (the second encoding method in S4702), the first multiplexing unit 4710 does not process the NAL unit (S4703).
[0237] On the other hand, when pcc_codec_type indicates the second coding method (the first coding method in S4702), the first multiplexer 4710 records pcc1 in ftyp (S4704). That is, the first multiplexer 4710 records information indicating that data encoded by the first coding method is stored in the file in ftyp.
[0238] Next, the first multiplexing unit 4710 parses the pcc_nal_unit_type included in the NAL unit header, and saves the data into a box (moov or mdat, etc.) using a predetermined method corresponding to the data type indicated by pcc_nal_unit_type (S4705). In addition, the first multiplexing unit 4710 creates an ISOBMFF file including the ftyp and the box (S4706).
[0239] Fig. 27 This is a flowchart of the multiplexing process performed by the second multiplexing unit 4730. First, the second multiplexing unit 4730 analyzes pcc_codec_type included in the NAL unit header to determine whether the codec used is the first encoding method or the second encoding method (S4711).
[0240] When pcc_unit_type indicates the second coding method (the second coding method in S4712), the second multiplexing unit 4730 records pcc2 in ftyp (S4713). That is, the second multiplexing unit 4730 records information indicating that data encoded by the second coding method is stored in the file in ftyp.
[0241] Next, the second multiplexing unit 4730 parses the pcc_nal_unit_type included in the NAL unit header, and saves the data into a box (moov or mdat, etc.) using a predetermined method corresponding to the data type indicated by pcc_nal_unit_type (S4714). Furthermore, the second multiplexing unit 4730 creates an ISOBMFF file including the ftyp and the box (S4715).
[0242] On the other hand, when pcc_unit_type indicates the first encoding method (the first encoding method in S4712), the second multiplexing unit 4730 does not process the NAL unit (S4716).
[0243] In addition, the above-mentioned process shows an example of encoding PCC data using either the first encoding method or the second encoding method. The first multiplexing unit 4710 and the second multiplexing unit 4730 save the desired NAL unit to the file by identifying the codec type of the NAL unit. In addition, in the case where the identification information of the PCC codec is included in addition to the NAL unit header, the first multiplexing unit 4710 and the second multiplexing unit 4730 may also use the identification information of the PCC codec included in addition to the NAL unit header to identify the codec type (the first encoding method or the second encoding method) in steps S4701 and S4711.
[0244] Furthermore, when the first multiplexing unit 4710 and the second multiplexing unit 4730 save the data in the file in steps S4706 and S4714, they may delete pcc_nal_unit_type from the NAL unit header and save the data in the file.
[0245] Fig.28This is a flowchart showing the processing performed by the first demultiplexing unit 4720 and the first decoding unit 4640. First, the first demultiplexing unit 4720 parses the ftyp contained in the ISOBMFF file (S4721). In the case where the codec represented by ftyp is the second coding method (pcc2) (it is the second coding method in S4722), the first demultiplexing unit 4720 determines that the data contained in the payload of the NAL unit is data encoded using the second coding method (S4723). In addition, the first demultiplexing unit 4720 passes the determination result to the first decoding unit 4640. The first decoding unit 4640 does not process the NAL unit (S4724).
[0246] On the other hand, when the codec indicated by ftyp is the first coding method (pcc1) (the first coding method in S4722), the first demultiplexing unit 4720 determines that the data included in the payload of the NAL unit is data encoded using the first coding method (S4725). In addition, the first demultiplexing unit 4720 transmits the determination result to the first decoding unit 4640.
[0247] The first decoding unit 4640 identifies the data by assuming that pcc_nal_unit_type included in the NAL unit header is an identifier of the NAL unit for the first encoding method (S4726). Then, the first decoding unit 4640 decodes the PCC data using the decoding process of the first encoding method (S4727).
[0248] Fig.29 4731. This is a flowchart showing the processing performed by the second demultiplexing unit 4740 and the second decoding unit 4660. First, the second demultiplexing unit 4740 parses the ftyp contained in the ISOBMFF file (S4731). When the codec represented by ftyp is the second coding method (pcc2) (the second coding method in S4732), the second demultiplexing unit 4740 determines that the data contained in the payload of the NAL unit is data encoded using the second coding method (S4733). In addition, the second demultiplexing unit 4740 passes the determination result to the second decoding unit 4660.
[0249] The second decoding unit 4660 identifies the data by assuming that pcc_nal_unit_type included in the NAL unit header is an identifier of the NAL unit for the second encoding method (S4734). Then, the second decoding unit 4660 decodes the PCC data using the decoding process of the second encoding method (S4735).
[0250] On the other hand, when the codec indicated by ftyp is the first coding method (pcc1) (the first coding method in S4732), the second demultiplexing unit 4740 determines that the data included in the payload of the NAL unit is data encoded using the first coding method (S4736). In addition, the second demultiplexing unit 4740 passes the determination result to the second decoding unit 4660. The second decoding unit 4660 does not process the NAL unit (S4737).
[0251] In this way, for example, by identifying the codec type of the NAL unit in the first demultiplexing unit 4720 or the second demultiplexing unit 4740, the codec type can be identified at an earlier stage. Furthermore, the desired NAL unit can be input to the first decoding unit 4640 or the second decoding unit 4660, and unnecessary NAL units can be removed. In this case, in the first decoding unit 4640 or the second decoding unit 4660, it may be unnecessary to parse the identification information of the codec. In addition, it is also possible to parse the identification information of the codec again by referring to the NAL unit type in the first decoding unit 4640 or the second decoding unit 4660.
[0252] Furthermore, when pcc_nal_unit_type is deleted from the NAL unit header in the first multiplexing unit 4710 or the second multiplexing unit 4730 , the first inverse multiplexing unit 4720 or the second inverse multiplexing unit 4740 may output the NAL unit to the first decoding unit 4640 or the second decoding unit 4660 after assigning pcc_nal_unit_type to the NAL unit.
[0253] (Implementation method 3)
[0254] In this embodiment, a multiplexing unit and a demultiplexing unit corresponding to the encoding unit 4670 and the decoding unit 4680 corresponding to a plurality of codecs described in Embodiment 1 are described. Fig.30 This is a diagram showing the configuration of the encoding unit 4670 and the third multiplexing unit 4750 according to this embodiment.
[0255] The coding unit 4670 codes the point cloud data using one or both of the first coding method and the second coding method. The coding unit 4670 may switch the coding method (the first coding method and the second coding method) in units of point cloud data or frames. In addition, the coding unit 4670 may switch the coding method in units that can be coded.
[0256] The encoding unit 4670 generates encoded data (encoded stream) including identification information of the PCC codec.
[0257] The third multiplexing unit 4750 includes a file conversion unit 4751. The file conversion unit 4751 converts the NAL unit output from the encoding unit 4670 into a file of PCC data. The file conversion unit 4751 parses the codec identification information included in the NAL unit header to determine whether the PCC coded data is data encoded by the first coding method, data encoded by the second coding method, or data encoded by both methods. The file conversion unit 4751 records the version name that can identify the codec in ftyp. For example, in the case of indicating encoding by both methods, pcc3 is recorded in ftyp.
[0258] Furthermore, when the encoding unit 4670 describes the identification information of the PCC codec other than the NAL unit, the file conversion unit 4751 may determine the PCC codec (encoding method) using the identification information.
[0259] Fig.31 This is a diagram showing the configuration of the third inverse multiplexing unit 4760 and the decoding unit 4680 according to the present embodiment.
[0260] The third demultiplexing unit 4760 includes a file inverse conversion unit 4761. The file inverse conversion unit 4761 analyzes ftyp included in the file and determines whether the PCC coded data is data coded using the first coding method, data coded using the second coding method, or data coded using both methods.
[0261] When the PCC coded data is coded using one of the coding methods, the data is input to the corresponding decoding unit of the first decoding unit 4640 and the second decoding unit 4660, and no data is input to the other decoding unit. When the PCC coded data is coded using both coding methods, the data is input to the decoding unit 4680 corresponding to both methods.
[0262] The decoding unit 4680 decodes the PCC coded data using one or both of the first coding method and the second coding method.
[0263] Fig.32 It is a flowchart showing the processing performed by the third multiplexing unit 4750 related to this embodiment.
[0264] First, the third multiplexing unit 4750 analyzes pcc_codec_type included in the NAL unit header to determine whether the codec used is the first encoding method, the second encoding method, or both the first encoding method and the second encoding method (S4741).
[0265] When the second encoding method is used (Yes in S4742, and the second encoding method is used in S4743), the third multiplexing unit 4750 records pcc2 in ftyp (S4744). That is, the third multiplexing unit 4750 records information indicating that data encoded by the second encoding method is stored in the file in ftyp.
[0266] Next, the third multiplexing unit 4750 parses the pcc_nal_unit_type included in the NAL unit header, and saves the data into a box (moov or mdat, etc.) using a predetermined method corresponding to the data type indicated by pcc_unit_type (S4745). Furthermore, the third multiplexing unit 4750 creates an ISOBMFF file including the ftyp and the box (S4746).
[0267] On the other hand, when the first encoding method is used ("Yes" in S4742, and the first encoding method is used in S4743), the third multiplexing unit 4750 records pcc1 in ftyp (S4747). That is, the third multiplexing unit 4750 records information indicating that data encoded by the first encoding method is stored in the file in ftyp.
[0268] Next, the third multiplexing unit 4750 parses the pcc_nal_unit_type included in the NAL unit header, and saves the data into a box (moov or mdat, etc.) using a predetermined method corresponding to the data type indicated by pcc_unit_type (S4748). Furthermore, the third multiplexing unit 4750 creates an ISOBMFF file including the ftyp and the box (S4746).
[0269] On the other hand, when both the first encoding method and the second encoding method are used (No in S4742), the third multiplexing unit 4750 records pcc3 in ftyp (S4749). That is, the third multiplexing unit 4750 records information indicating that data encoded by both encoding methods is stored in the file in ftyp.
[0270] Next, the third multiplexing unit 4750 parses the pcc_nal_unit_type included in the NAL unit header, and saves the data into a box (moov or mdat, etc.) using a predetermined method corresponding to the data type indicated by pcc_unit_type (S4750). Furthermore, the third multiplexing unit 4750 creates an ISOBMFF file including the ftyp and the box (S4746).
[0271] Fig.334760 and the decoding unit 4680. First, the third demultiplexing unit 4760 parses the ftyp contained in the ISOBMFF file (S4761). In the case where the codec represented by ftyp is the second coding method (pcc2) ("yes" in S4762, and it is the second coding method in S4763), the third demultiplexing unit 4760 determines that the data contained in the payload of the NAL unit is data encoded using the second coding method (S4764). In addition, the third demultiplexing unit 4760 transmits the determination result to the decoding unit 4680.
[0272] The decoding unit 4680 identifies the data by assuming that pcc_nal_unit_type included in the NAL unit header is an identifier of the NAL unit for the second encoding method (S4765). Then, the decoding unit 4680 decodes the PCC data using the decoding process of the second encoding method (S4766).
[0273] On the other hand, when the codec indicated by ftyp is the first coding method (pcc1) ("Yes" in S4762, and the first coding method in S4763), the third demultiplexing unit 4760 determines that the data included in the payload of the NAL unit is data encoded using the first coding method (S4767). In addition, the third demultiplexing unit 4760 transmits the determination result to the decoding unit 4680.
[0274] The decoding unit 4680 identifies the data by assuming that pcc_nal_unit_type included in the NAL unit header is an identifier of the NAL unit for the first encoding method (S4768). Then, the decoding unit 4680 decodes the PCC data using the decoding process of the first encoding method (S4769).
[0275] On the other hand, when ftyp indicates that both coding methods (pcc3) are used (No in S4762), the third demultiplexing unit 4760 determines that the data included in the payload of the NAL unit is data encoded using both the first coding method and the second coding method (S4770). In addition, the third demultiplexing unit 4760 transmits the determination result to the decoding unit 4680.
[0276] The decoding unit 4680 identifies the data by setting the pcc_nal_unit_type included in the NAL unit header to be the identifier of the NAL unit for the codec described in pcc_codec_type (S4771). Furthermore, the decoding unit 4680 decodes the PCC data using the decoding processing of both encoding methods (S4772). That is, the decoding unit 4680 decodes the data encoded by the first encoding method using the decoding processing of the first encoding method, and decodes the data encoded by the second encoding method using the decoding processing of the second encoding method.
[0277] The following describes a modified example of the present embodiment. As the type of version represented by ftyp, the following types may be represented by identification information. In addition, a combination of the following types may be represented by identification information.
[0278] The identification information indicates whether the object of the original data before PCC encoding is a point group with a limited area or a large-scale point group with an unlimited area like map information.
[0279] The identification information may also indicate whether the original data before PCC encoding is a static object or a dynamic object.
[0280] As described above, the identification information may indicate whether the PCC coded data is data coded using the first coding method or data coded using the second coding method.
[0281] The identification information may indicate an algorithm used in PCC encoding. Here, the algorithm is, for example, an encoding method that can be used in the first encoding method or the second encoding method.
[0282] The identification information may also indicate the difference in the method of storing the PCC coded data in the ISOBMFF file. For example, the identification information may indicate whether the storage method used is a storage method for accumulation or a storage method for real-time transmission such as dynamic streaming.
[0283] In addition, in Embodiments 2 and 3, an example of using ISOBMFF as a file format is described, but other methods may also be used. For example, the same method as in this embodiment may be used when saving PCC encoded data to MPEG-2TS Systems, MPEG-DASH, MMT, or RMP.
[0284] In the above, an example is shown in which metadata such as identification information is stored in ftyp, but these metadata may be stored outside of ftyp. For example, these metadata may be stored in moov.
[0285] As described above, the three-dimensional data storage device (or three-dimensional data multiplexing device, or three-dimensional data encoding device) performs Fig.34 Processing shown.
[0286] First, the three-dimensional data storage device (for example, including the first multiplexing unit 4710, the second multiplexing unit 4730, or the third multiplexing unit 4750) obtains one or more units (for example, NAL units) storing a coded stream obtained by encoding point group data (S4781). Next, the three-dimensional data storage device saves the one or more units to a file (for example, an ISOBMFF file) (S4782). In addition, during the saving (S4782), the three-dimensional data storage device saves information indicating that the data stored in the file is data obtained by encoding point group data (for example, pcc1, pcc2, or pcc3) in the control information (for example, ftyp) of the above file.
[0287] Thus, in a device that processes a file generated by the three-dimensional data storage device, it is possible to refer to the control information of the file and determine at an early stage whether the data stored in the file is coded data of point group data, thereby reducing the processing amount of the device or speeding up the processing.
[0288] For example, the information also indicates the encoding method used in encoding the point group data in the first encoding method and the second encoding method. In addition, the data stored in the file is the data obtained by encoding the point group data, and the encoding method used in encoding the point group data in the first encoding method and the second encoding method, which can be represented by a single information or different information.
[0289] Thus, in a device that processes a file generated by the three-dimensional data storage device, the codec used for the data stored in the file can be determined earlier by referring to the file control information, thereby reducing the amount of processing of the device or speeding up the processing.
[0290] For example, the first encoding method is a method (GPCC) of encoding position information of the point group data represented by an N (N is an integer greater than 2) fork tree and using the position information to encode attribute information, and the second encoding method is a method (VPCC) of generating a two-dimensional image based on the point group data and encoding the two-dimensional image using an image encoding method.
[0291] For example, the above file is based on ISOBMFF (ISO based media file format: ISO base media file format).
[0292] For example, the three-dimensional data storage device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0293] In addition, as described above, the three-dimensional data acquisition device (or three-dimensional data demultiplexing device, or three-dimensional data decoding device) performs Fig.35 Processing shown.
[0294] The three-dimensional data acquisition device (for example, including the first inverse multiplexing unit 4720, the second inverse multiplexing unit 4740, or the third inverse multiplexing unit 4760) acquires a file (for example, an ISOBMFF file) storing one or more units (for example, NAL units), wherein the one or more units store a coded stream obtained by encoding point group data (S4791). Next, the three-dimensional data acquisition device acquires one or more units from the file (S4792). In addition, the control information of the file (for example, ftyp) includes information indicating that the data stored in the file is data obtained by encoding the point group data (for example, pcc1, pcc2, or pcc3).
[0295] For example, the three-dimensional data acquisition device refers to the above information to determine whether the data stored in the file is data obtained by encoding the point group data. In addition, when the three-dimensional data acquisition device determines that the data stored in the file is data obtained by encoding the point group data, it generates the point group data by decoding the data obtained by encoding the point group data contained in one or more units. Alternatively, when the three-dimensional data acquisition device determines that the data stored in the file is data obtained by encoding the point group data, it outputs (notifies) information indicating that the data contained in one or more units is data obtained by encoding the point group data to a subsequent processing unit (for example, the first decoding unit 4640, the second decoding unit 4660, or the decoding unit 4680).
[0296] Thus, the 3D data acquisition device can refer to the control information of the file and determine whether the data stored in the file is the coded data of the point group data at an early stage, thereby reducing the processing amount of the 3D data acquisition device or the subsequent device or speeding up the processing.
[0297] For example, the information also indicates the encoding method used in the encoding of the first encoding method and the second encoding method. In addition, the data stored in the file is the data obtained by encoding the point group data, and the encoding method used in the encoding of the point group data of the first encoding method and the second encoding method, which can be represented by a single information or different information.
[0298] Thus, the 3D data acquisition device can refer to the control information of the file and determine the codec used for the data stored in the file at an early stage, thereby reducing the amount of processing of the 3D data acquisition device or a subsequent device or increasing the processing speed.
[0299] For example, the three-dimensional data acquisition device acquires data encoded by one of the encoding methods from the encoded point cloud data including data encoded by the first encoding method and data encoded by the second encoding method based on the above information.
[0300] For example, the first encoding method is a method (GPCC) of encoding position information of the point group data represented by an N (N is an integer greater than 2) fork tree and using the position information to encode attribute information, and the second encoding method is a method (VPCC) of generating a two-dimensional image based on the point group data and encoding the two-dimensional image using an image coding method.
[0301] For example, the above file is based on ISOBMFF (ISO based media file format).
[0302] For example, the three-dimensional data acquisition device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0303] (Implementation 4)
[0304] In this embodiment, the types of coded data (position information (Geometry), attribute information (Attribute), additional information (Metadata)) generated by the first coding unit 4630 or the second coding unit 4650, the generation method of the additional information (metadata), and the multiplexing process in the multiplexing unit are described. In addition, the additional information (metadata) may be expressed as a parameter set or control information.
[0305] In this embodiment, Figure 4 The dynamic object (three-dimensional point group data that changes with time) described in the above is used as an example, but the same method can also be used in the case of a static object (three-dimensional point group data at an arbitrary time).
[0306] Fig.36 The diagram shows the configuration of a coding unit 4801 and a multiplexing unit 4802 included in the three-dimensional data coding device according to the present embodiment. The coding unit 4801 corresponds to, for example, the first coding unit 4630 or the second coding unit 4650 described above. The multiplexing unit 4802 corresponds to the multiplexing unit 4634 or 46456 described above.
[0307] The encoding unit 4801 encodes the point cloud data of a plurality of PCC (Point Cloud Compression) frames to generate encoded data (Multiple Compressed Data) of a plurality of position information, attribute information, and additional information.
[0308] The multiplexing unit 4802 converts the data of a plurality of data types (position information, attribute information, and additional information) into a data structure that takes data access in the decoding device into consideration by converting the data into NAL units.
[0309] Fig.37 4801. The arrows in the figure represent dependency relationships related to the decoding of the coded data, and the root of the arrow depends on the data at the tip of the arrow. That is, the decoding device decodes the data at the tip of the arrow, and uses the decoded data to decode the data at the root of the arrow. In other words, dependency means that the data of the dependency target is referenced (used) in the processing (encoding or decoding, etc.) of the data of the dependency source.
[0310] First, the generation process of the encoded data of the position information is described. The encoding unit 4801 generates the encoded position data (Compressed Geometry Data) of each frame by encoding the position information of each frame. In addition, the encoded position data is represented by G(i). i represents the frame number or the time of the frame, etc.
[0311] In addition, the encoding unit 4801 generates a position parameter set (GPS(i)) corresponding to each frame. The position parameter set includes parameters that can be used in decoding the encoded position data. In addition, the encoded position data of each frame depends on the corresponding position parameter set.
[0312] In addition, the coded position data composed of a plurality of frames is defined as a position sequence (Geometry Sequence). The coding unit 4801 generates a position sequence parameter set (Geometry Sequence PS: also referred to as position SPS) that stores parameters commonly used in decoding processing of a plurality of frames in the position sequence. The position sequence depends on the position SPS.
[0313] Next, the generation process of the encoded data of the attribute information is described. The encoding unit 4801 generates the compressed attribute data (Compressed Attribute Data) of each frame by encoding the attribute information of each frame. In addition, the compressed attribute data is represented by A(i). In addition, Fig.37 , an example in which attribute X and attribute Y exist is shown, and the encoded attribute data of attribute X is represented by AX(i), and the encoded attribute data of attribute Y is represented by AY(i).
[0314] In addition, the encoding unit 4801 generates an attribute parameter set (APS(i)) corresponding to each frame. In addition, the attribute parameter set of attribute X is represented by AXPS(i), and the attribute parameter set of attribute Y is represented by AYPS(i). The attribute parameter set includes parameters that can be used in decoding the encoded attribute information. The encoded attribute data depends on the corresponding attribute parameter set.
[0315] In addition, the coded attribute data consisting of a plurality of frames is defined as an attribute sequence (Attribute Sequence). The coding unit 4801 generates an attribute sequence parameter set (Attribute Sequence PS: also referred to as attribute SPS) that stores parameters commonly used in decoding processing of a plurality of frames in the attribute sequence. The attribute sequence depends on the attribute SPS.
[0316] Furthermore, in the first encoding method, the encoded attribute data depends on the encoded position data.
[0317] In addition, Fig.37 , an example is shown in which there are two types of attribute information (attribute X and attribute Y). In the case of two types of attribute information, for example, two encoding units generate respective data and metadata. In addition, for example, an attribute sequence is defined for each type of attribute information, and an attribute SPS is generated for each type of attribute information.
[0318] In addition, Fig.37 In FIG. 4 , an example is shown in which there is one type of position information and two types of attribute information, but the present invention is not limited thereto. The attribute information may be one type or three or more types. In this case, the encoded data can be generated in the same way. In addition, in the case of point cloud data without attribute information, there may be no attribute information. In this case, the encoding unit 4801 may not generate a parameter set associated with the attribute information.
[0319] Next, the generation process of the additional information (metadata) is described. The coding unit 4801 generates a parameter set of the entire PCC stream, namely, a PCC stream PS (PCC Stream PS: also referred to as stream PS). The coding unit 4801 stores parameters that can be used in common in the decoding process of one or more position sequences and one or more attribute sequences in the stream PS. For example, the stream PS contains identification information of the codec representing the point group data and information representing the algorithm used in the encoding. The position sequence and the attribute sequence depend on the stream PS.
[0320] Next, the access unit and GOF are described. In this embodiment, the concept of access unit (AccessUnit: AU) and GOF (Group of Frame) is newly introduced.
[0321] An access unit is a basic unit used to access data during decoding, and is composed of one or more data and one or more metadata. For example, an access unit is composed of location information and one or more attribute information at the same time. A GOF is a random access unit, and is composed of one or more access units.
[0322] The coding unit 4801 generates an access unit header (AU Header) as identification information indicating the beginning of the access unit. The coding unit 4801 stores parameters related to the access unit in the access unit header. For example, the access unit header includes: the structure or information of the encoded data included in the access unit. In addition, the access unit header includes parameters commonly used for the data included in the access unit, such as parameters related to the decoding of the encoded data.
[0323] In addition, the encoder 4801 may generate an access unit delimiter that does not include parameters related to the access unit instead of the access unit header. The access unit delimiter is used as identification information indicating the beginning of the access unit. The decoding device recognizes the beginning of the access unit by detecting the access unit header or the access unit delimiter.
[0324] Next, the generation of identification information at the beginning of GOF is described. The coding unit 4801 generates a GOF header (GOFHeader) as identification information indicating the beginning of GOF. The coding unit 4801 stores parameters related to GOF in the GOF header. For example, the GOF header includes: the structure or information of the encoded data included in the GOF. In addition, the GOF header includes parameters commonly used for the data included in the GOF, such as parameters related to the decoding of the encoded data.
[0325] In addition, the encoder 4801 may generate a GOF delimiter that does not include parameters related to GOF instead of the GOF header. The GOF delimiter is used as identification information indicating the beginning of GOF. The decoding device recognizes the beginning of GOF by detecting the GOF header or the GOF delimiter.
[0326] In PCC coded data, for example, an access unit is defined as a PCC frame unit. The decoding device accesses the PCC frame based on identification information at the head of the access unit.
[0327] In addition, for example, GOF is defined as one random access unit. The decoding device accesses the random access unit based on the identification information at the beginning of GOF. For example, if PCC frames have no dependency on each other and can be decoded independently, the PCC frame can also be defined as a random access unit.
[0328] Furthermore, two or more PCC frames may be allocated to one access unit, and a plurality of random access units may be allocated to one GOF.
[0329] Furthermore, the encoder 4801 may define and generate parameter sets or metadata other than those described above. For example, the encoder 4801 may generate SEI (Supplemental Enhancement Information) storing parameters (optional parameters) that may not necessarily be used during decoding.
[0330] Next, the structure of the coded data and the method of storing the coded data in the NAL unit are described.
[0331] For example, a data format is specified for each type of coded data. Fig.38 This is a diagram showing examples of coded data and NAL units.
[0332] For example, Fig.38 As shown, the coded data includes a header and a payload. In addition, the coded data may also include coded data, a header, or length information indicating the length (data amount) of the payload. In addition, the coded data may not include a header.
[0333] The header includes, for example, identification information for identifying data. The identification information indicates, for example, the data type or the frame number.
[0334] The header includes, for example, identification information indicating a reference relationship. The identification information is, for example, information that is saved in the header when there is a dependency relationship between data and is used to refer to the reference target from the reference source. For example, the header of the reference target includes identification information for identifying the data. The header of the reference source includes identification information indicating the reference target.
[0335] Furthermore, when the reference destination or the reference source can be identified or derived from other information, identification information for specifying data or identification information indicating a reference relationship may be omitted.
[0336] The multiplexing unit 4802 stores the coded data in the payload of the NAL unit. The NAL unit header includes pcc_nal_unit_type, which is identification information of the coded data. Fig.39 This is a diagram showing a semantic example of pcc_nal_unit_type.
[0337] like Fig.39As shown, when pcc_codec_type is codec 1 (Codec 1: first coding method), pcc_nal_unit_type values 0 to 10 are allocated to the coded position data (Geometry), coded attribute X data (AttributeX), coded attribute Y data (AttributeY), position PS (Geom.PS), attribute XPS (AttrX.PS), attribute YPS (AttrX.PS), position SPS (Geometry Sequence PS), attribute XSPS (AttributeX Sequence PS), attribute YSPS (AttributeY Sequence PS), AU Header, and GOF Header in codec 1. In addition, values 11 and later are allocated as a reserve for codec 1.
[0338] When pcc_codec_type is codec2 (Codec2: second encoding method), pcc_nal_unit_type values 0 to 2 are assigned to codec data A (DataA), metadata A (MetaDataA), and metadata B (MetaDataB). In addition, values 3 and later are assigned as a reserve for codec2.
[0339] Next, the order in which data is transmitted is described. Next, the constraints on the order in which NAL units are transmitted are described.
[0340] The multiplexing unit 4802 sends out the NAL units in GOF or AU units. The multiplexing unit 4802 arranges the GOF header at the beginning of the GOF and arranges the AU header at the beginning of the AU.
[0341] The multiplexing unit 4802 may configure a sequence parameter set (SPS) for each AU so that even if data is lost due to packet loss or the like, the decoding device can perform decoding from the next AU.
[0342] When there is a dependency relationship related to decoding in the coded data, the decoding device decodes the reference source data after decoding the reference target data. In the decoding device, in order to enable decoding in the order received without rearranging the data, the multiplexing unit 4802 sends the reference target data first.
[0343] Fig.40 This is a diagram showing an example of the order in which NAL units are sent. Fig.40 It shows three examples: position information priority, parameter priority and data merging.
[0344] The sending order with priority on position information is an example in which information related to position information and information related to attribute information are sent together. In the case of this sending order, the sending of information related to position information is completed earlier than the sending of information related to attribute information.
[0345] For example, by using this transmission order, a decoding device that does not decode attribute information may be able to set a time when no processing is performed by ignoring the decoding of attribute information. In addition, for example, in the case of a decoding device that wants to decode position information early, it may be possible to decode the position information earlier by obtaining the encoded data of the position information earlier.
[0346] In addition, Fig.40 In the example, the attribute XSPS and the attribute YSPS are combined and recorded as the attribute SPS, but the attribute XSPS and the attribute YSPS may be configured separately.
[0347] In the parameter set priority sending order, the parameter set is sent first, and then the data is sent.
[0348] As described above, as long as the constraints of the NAL unit sending order are followed, the multiplexing unit 4802 can send the NAL units in any order. For example, the order identification information can also be defined, and the multiplexing unit 4802 has the function of sending NAL units in a plurality of styles. For example, the order identification information of the NAL unit is stored in the stream PS.
[0349] The 3D data decoding device may also perform decoding based on the sequence identification information. The 3D data decoding device may also indicate a desired transmission sequence to the 3D data encoding device, and the 3D data encoding device (multiplexing unit 4802) controls the transmission sequence according to the indicated transmission sequence.
[0350] In addition, as long as the order of sending data is within the range of the constraints of the sending order, such as the sending order of data merging, the multiplexing unit 4802 may generate coded data that combines multiple functions. Fig.40 As shown, the GOF header and the AU header may be combined, or the AXPS and the AYPS may be combined. In this case, an identifier indicating that the data has multiple functions is defined in pcc_nal_unit_type.
[0351] The following describes a modified example of the present embodiment. PS has levels, such as frame-level PS, sequence-level PS, and PCC sequence-level PS. If the PCC sequence level is set as the upper level and the frame level is set as the lower level, the parameter storage method can also use the following method.
[0352] The default PS value is represented by the higher-order PS. In addition, when the value of the lower-order PS is different from the value of the upper-order PS, the value of the PS is represented by the lower-order PS. Alternatively, the value of the PS is not recorded in the upper-order PS, and the value of the PS is recorded in the lower-order PS. Alternatively, one or both of the lower-order PS and the upper-order PS are used to represent the information of whether the value of the PS is represented by the lower-order PS, the upper-order PS, or both. Alternatively, the lower-order PS may be merged into the upper-order PS. Alternatively, when the lower-order PS and the upper-order PS are repeated, the multiplexing unit 4802 may omit the sending of one of them.
[0353] In addition, the encoding unit 4801 or the multiplexing unit 4802 may divide the data into slices or tiles, etc., and send the divided data. The divided data includes information for identifying the divided data, and the parameter set includes parameters used in decoding the divided data. In this case, in pcc_nal_unit_type, an identifier indicating that it is data related to a tile or a slice or data storing parameters is defined.
[0354] Next, the processing of the sequence identification information will be described. Fig.41 This is a flowchart of processing performed by the three-dimensional data encoding device (encoding unit 4801 and multiplexing unit 4802) regarding the order in which NAL units are sent.
[0355] First, the 3D data encoding device determines the order of NAL unit transmission (position information priority or parameter set priority) (S4801). For example, the 3D data encoding device determines the transmission order based on a designation from a user or an external device (eg, a 3D data decoding device).
[0356] When the determined sending order is position information priority (position information priority in S4802), the three-dimensional data encoding device sets the order identification information included in the stream PS to position information priority (S4803). That is, in this case, the order identification information indicates that the NAL units are sent in the order of position information priority. And the three-dimensional data encoding device sends the NAL units in the order of position information priority (S4804).
[0357] On the other hand, when the determined transmission order is parameter set priority (parameter set priority in S4802), the three-dimensional data encoding device sets the sequence identification information included in the stream PS to parameter set priority (S4805). That is, in this case, the sequence identification information indicates that the NAL units are transmitted in the order of parameter set priority. And the three-dimensional data encoding device transmits the NAL units in the order of parameter set priority (S4806).
[0358] Fig.4248 is a flowchart of a process performed by the three-dimensional data decoding device regarding the order in which NAL units are sent. First, the three-dimensional data decoding device analyzes the order identification information included in the stream PS (S4811).
[0359] When the sending order indicated by the sequence identification information is position information-first (position information-first in S4812), the three-dimensional data decoding apparatus decodes the NAL units in the position information-first sending order (S4813).
[0360] On the other hand, when the sending order indicated by the sequence identification information is parameter set priority (parameter set priority in S4812), the three-dimensional data decoding device sets the sending order of the NAL unit to parameter set priority and decodes the NAL unit (S4814).
[0361] For example, when the three-dimensional data decoding apparatus does not decode the attribute information, in step S4813 , it may obtain the NAL unit related to the position information instead of all the NAL units, and decode the position information from the obtained NAL unit.
[0362] Next, the processing related to the generation of AU and GOF is described. Fig.43 This is a flowchart showing the processing performed by the three-dimensional data encoding device (multiplexing unit 4802) regarding the generation of AUs and GOFs in the multiplexing of NAL units.
[0363] First, the three-dimensional data encoding device determines the type of encoded data (S4821). Specifically, the three-dimensional data encoding device determines whether the encoded data to be processed is data starting with AU, data starting with GOF, or other data.
[0364] When the coded data is data starting with GOF (starting with GOF in S4822), the three-dimensional data coding apparatus arranges the GOF header and the AU header at the beginning of the coded data belonging to GOF to generate a NAL unit (S4823).
[0365] When the coded data is data starting with an AU (starting with an AU in S4822), the three-dimensional data coding apparatus arranges the AU header at the beginning of the coded data belonging to the AU and generates a NAL unit (S4824).
[0366] When the coded data does not start with either GOF or AU (other than GOF or AU in S4822), the three-dimensional data coding apparatus arranges the coded data after the AU header of the AU to which the coded data belongs and generates a NAL unit (S4825).
[0367] Next, the processing related to access to AU and GOF is described. Fig.44 This is a flowchart of the processing of the three-dimensional data decoding device related to access to AU and GOF in the inverse multiplexing of NAL units.
[0368] First, the 3D data decoding device determines the type of coded data contained in the NAL unit by parsing nal_unit_type contained in the NAL unit (S4831). Specifically, the 3D data decoding device determines whether the coded data contained in the NAL unit is data starting with AU, data starting with GOF, or other data.
[0369] When the coded data included in the NAL unit is data starting with GOF (starting with GOF in S4832), the three-dimensional data decoding device determines that the NAL unit is the start position of random access, accesses the NAL unit, and starts decoding processing (S4833).
[0370] On the other hand, when the encoded data included in the NAL unit is data starting with an AU (starting with an AU in S4832), the three-dimensional data decoding device determines that the NAL unit starts with an AU, accesses the data included in the NAL unit, and decodes the AU (S4834).
[0371] On the other hand, when the coded data included in the NAL unit is not at the beginning of either GOF or AU (other than the beginning of GOF or AU in S4832), the three-dimensional data decoding apparatus does not process the NAL unit.
[0372] As described above, the three-dimensional data encoding device performs Fig.45 The three-dimensional data encoding device encodes time-series three-dimensional data (for example, point group data of a dynamic object). The three-dimensional data includes position information and attribute information at each time.
[0373] First, the three-dimensional data encoding device encodes the position information (S4841). Next, the three-dimensional data encoding device encodes the attribute information of the processing object with reference to the position information at the same time as the attribute information of the processing object (S4842). Fig.37 As shown, the position information and attribute information at the same time constitute an access unit (AU). That is, the three-dimensional data encoding device encodes the attribute information of the processing object with reference to the position information included in the same access unit as the attribute information of the processing object.
[0374] Thus, the three-dimensional data encoding device can use the access unit to facilitate the control of references in encoding. Therefore, the three-dimensional data encoding device can reduce the amount of processing in the encoding process.
[0375] For example, the three-dimensional data encoding device generates a bit stream including information on encoded position information (encoded position data), encoded attribute information (encoded attribute data), and position information indicating a reference destination of the attribute information of a processing target.
[0376] For example, the bitstream includes a position parameter set (position PS) including control information of position information at each time point, and an attribute parameter set (attribute PS) including control information of attribute information at each time point.
[0377] For example, the bitstream includes a position sequence parameter set (position SPS) including control information common to position information at multiple times, and an attribute sequence parameter set (attribute SPS) including control information common to attribute information at multiple times.
[0378] For example, the bitstream includes a stream parameter set (stream PS) including position information at multiple times and control information common to attribute information at multiple times.
[0379] For example, the bitstream includes an access unit header (AU header) including common control information in the access unit.
[0380] For example, the three-dimensional data encoding device encodes GOF (Group of Frames) composed of one or more access units so that they can be decoded independently. That is, GOF is a random access unit.
[0381] For example, the bitstream contains a GOF header containing common control information within the GOF.
[0382] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0383] In addition, as described above, the three-dimensional data decoding device performs Fig.46 The processing shown. The three-dimensional data decoding device decodes the time series three-dimensional data (for example, point group data of a dynamic object). The three-dimensional data includes position information and attribute information at each time. The position information and attribute information at the same time constitute an access unit (AU).
[0384] First, the three-dimensional data decoding device decodes the position information from the bit stream (S4851). That is, the three-dimensional data decoding device generates the position information by decoding the encoded position information (encoded position data) included in the bit stream.
[0385] Next, the three-dimensional data decoding device decodes the attribute information of the processing object from the bit stream by referring to the position information at the same time as the attribute information of the processing object (S4852). That is, the three-dimensional data decoding device generates the attribute information by decoding the encoded attribute information (encoded attribute data) contained in the bit stream. At this time, the three-dimensional data decoding device refers to the decoded position information contained in the same access unit as the attribute information.
[0386] Thus, the three-dimensional data decoding device can use the access unit to facilitate the control of references in decoding. Therefore, the three-dimensional data decoding method can reduce the processing amount of the decoding process.
[0387] For example, the three-dimensional data decoding apparatus obtains information indicating the position information of the reference target of the attribute information of the processing target from the bit stream, and decodes the attribute information of the processing target by referring to the position information of the reference target indicated by the obtained information.
[0388] For example, the bitstream includes a position parameter set (position PS) including control information of position information at each time point, and an attribute parameter set (attribute PS) including control information of attribute information at each time point. That is, the 3D data decoding device decodes the position information at the processing target time point using the control information included in the position parameter set at the processing target time point, and decodes the attribute information at the processing target time point using the control information included in the attribute parameter set at the processing target time point.
[0389] For example, the bitstream includes: a position sequence parameter set (position SPS) including common control information in position information at multiple times, and an attribute sequence parameter set (attribute SPS) including common control information in attribute information at multiple times. That is, the three-dimensional data decoding device decodes the position information at multiple times using the control information included in the position sequence parameter set, and decodes the attribute information at multiple times using the control information included in the attribute sequence parameter set.
[0390] For example, the bitstream includes a stream parameter set (stream PS) including control information common to position information at multiple times and attribute information at multiple times. That is, the 3D data decoding device decodes the position information at multiple times and attribute information at multiple times using the control information included in the stream parameter set.
[0391] For example, the bitstream includes an access unit header (AU header) including common control information in the access unit. That is, the 3D data decoding device decodes the position information and attribute information included in the access unit using the control information included in the access unit header.
[0392] For example, the three-dimensional data decoding device independently decodes GOF (Group of Frames) composed of one or more access units. That is, GOF is a random access unit.
[0393] For example, the bitstream includes a GOF header including common control information in the GOF. That is, the three-dimensional data decoding device decodes the position information and attribute information included in the GOF using the control information included in the GOF header.
[0394] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0395] (Implementation method 5)
[0396] In HEVC encoding, there is a tool for data segmentation such as slices or tiles in order to enable parallel processing by a decoding device, but there is no such tool in PCC (Point Cloud Compression) encoding.
[0397] In PCC, various data segmentation methods can be considered based on parallel processing, compression efficiency, and compression algorithm. Here, the definition, data structure, and sending and receiving methods of slices and tiles are explained.
[0398] Fig.47 This is a block diagram showing the structure of the first coding unit 4910 included in the three-dimensional data coding device related to this embodiment. The first coding unit 4910 generates coded data (coded stream) by coding the point cloud data using the first coding method (GPCC (Geometry based PCC)). The first coding unit 4910 includes a segmentation unit 4911, a plurality of position information coding units 4912, a plurality of attribute information coding units 4913, an additional information coding unit 4914, and a multiplexing unit 4915.
[0399] The segmentation unit 4911 generates a plurality of segmentation data by segmenting the point group data. Specifically, the segmentation unit 4911 generates a plurality of segmentation data by segmenting the space of the point group data into a plurality of subspaces. Here, the subspace refers to one of a tile and a slice, or a combination of a tile and a slice. More specifically, the point group data includes position information, attribute information, and additional information. The segmentation unit 4911 segments the position information into a plurality of segmentation position information, and segments the attribute information into a plurality of segmentation attribute information. In addition, the segmentation unit 4911 generates additional information about the segmentation.
[0400] The plurality of position information encoding units 4912 encode the plurality of divided position information to generate a plurality of encoded position information. For example, the plurality of position information encoding units 4912 processes the plurality of divided position information in parallel.
[0401] The multiple attribute information encoding unit 4913 generates multiple coded attribute information by encoding multiple pieces of split attribute information. For example, the multiple attribute information encoding unit 4913 processes multiple pieces of split attribute information in parallel.
[0402] The additional information encoding unit 4914 generates encoded additional information by encoding the additional information included in the point group data and the additional information on data division generated by the division unit 4911 during division.
[0403] The multiplexing unit 4915 generates coded data (coded stream) by multiplexing a plurality of coding position information, a plurality of coding attribute information, and coded additional information, and transmits the generated coded data. The coded additional information is used at the time of decoding.
[0404] In addition, Fig.47 , an example is shown in which the number of position information coding units 4912 and attribute information coding units 4913 is two, but the number of position information coding units 4912 and attribute information coding units 4913 may be one or more than three. In addition, multiple split data may be processed in parallel in the same chip like multiple cores in a CPU, may be processed in parallel by cores of multiple chips, or may be processed in parallel by multiple cores of multiple chips.
[0405] Fig.48 49 is a block diagram showing the configuration of the first decoding unit 4920. The first decoding unit 4920 restores the point group data by decoding the coded data (coded stream) generated by encoding the point group data using the first coding method (GPCC). The first decoding unit 4920 includes an inverse multiplexing unit 4921, a plurality of position information decoding units 4922, a plurality of attribute information decoding units 4923, an additional information decoding unit 4924, and a combining unit 4925.
[0406] The demultiplexing unit 4921 generates a plurality of encoding position information, a plurality of encoding attribute information, and encoding additional information by demultiplexing the encoded data (encoded stream).
[0407] The plurality of position information decoding units 4922 generate a plurality of divided position information by decoding a plurality of coded position information. For example, the plurality of position information decoding units 4922 process a plurality of coded position information in parallel.
[0408] The multiple attribute information decoding unit 4923 generates multiple pieces of segmented attribute information by decoding multiple pieces of coded attribute information. For example, the multiple attribute information decoding unit 4923 processes multiple pieces of coded attribute information in parallel.
[0409] The plurality of additional information decoding units 4924 generate additional information by decoding the encoded additional information.
[0410] The combining unit 4925 combines a plurality of pieces of segment position information using the additional information to generate position information. The combining unit 4925 combines a plurality of pieces of segment attribute information using the additional information to generate attribute information.
[0411] In addition, Fig.48 , an example is shown in which the number of position information decoding units 4922 and attribute information decoding units 4923 is two, but the number of position information decoding units 4922 and attribute information decoding units 4923 may be one, or may be three or more. In addition, a plurality of split data may be processed in parallel in the same chip like multiple cores in a CPU, may be processed in parallel by cores of multiple chips, or may be processed in parallel by multiple cores of multiple chips.
[0412] Next, the structure of the dividing unit 4911 will be described. Fig.49 49 is a block diagram of the dividing unit 4911. The dividing unit 4911 includes a slice dividing unit 4931 (Slice Divider), a position information tile dividing unit 4932 (Geometry Tile Divider), and an attribute information tile dividing unit 4933 (Attribute Tile Divider).
[0413] The slice segmentation unit 4931 generates a plurality of slice position information by segmenting the position information (Position (Geometry)) into slices. In addition, the slice segmentation unit 4931 generates a plurality of slice attribute information by segmenting the attribute information (Attribute) into slices. In addition, the slice segmentation unit 4931 outputs information on the slice segmentation and slice additional information (SliceMetaData) including information generated in the slice segmentation.
[0414] The position information tile partitioning unit 4932 generates a plurality of partition position information (a plurality of tile position information) by partitioning a plurality of slice position information into tiles. In addition, the position information tile partitioning unit 4932 outputs position tile additional information (GeometryTile MetaData) including information related to tile partitioning of position information and information generated in tile partitioning of position information.
[0415] The attribute information tile partitioning unit 4933 generates multiple partition attribute information (multiple tile attribute information) by partitioning multiple slice attribute information into tiles. In addition, the attribute information tile partitioning unit 4933 outputs attribute tile additional information (Attribute Tile MetaData) including information related to tile partitioning of attribute information and information generated in tile partitioning of attribute information.
[0416] In addition, the number of divided slices or tiles is 1 or more. That is, the division of slices or tiles may not be performed.
[0417] In addition, although an example of performing tile division after slice division is shown here, slice division may be performed after tile division. In addition, in addition to slices and tiles, new division types may be defined to perform division with three or more division types.
[0418] The following describes a method for segmenting point cloud data. Fig.50 The diagram shows an example of slice and tile division.
[0419] First, the method of slice segmentation is described. The segmentation unit 4911 segments the three-dimensional point group data into arbitrary point groups in slice units. The segmentation unit 4911 does not segment the position information and attribute information of the constituent points during the slice segmentation, but segments the position information and the attribute information together. That is, the segmentation unit 4911 performs slice segmentation in a manner that the position information and the attribute information at any point belong to the same slice. In addition, as long as these are followed, any method of segmentation number and segmentation method can be used. In addition, the minimum unit of segmentation is a point. For example, the number of segmentations of the position information and the attribute information is the same. For example, the three-dimensional points corresponding to the position information after the slice segmentation and the three-dimensional points corresponding to the attribute information are included in the same slice.
[0420] In addition, the segmentation unit 4911 generates slice additional information as additional information related to the number of segments and the segmentation method when the slice is segmented. The slice additional information is the same in position information and attribute information. For example, the slice additional information includes information indicating the reference coordinate position, size, or side length of the bounding box after segmentation. In addition, the slice additional information includes information indicating the number of segments and the segmentation type, etc.
[0421] Next, a method of tile division is described: The division unit 4911 divides the data after the slice division into slice position information (G slice) and slice attribute information (A slice), and divides the slice position information and the slice attribute information into tile units.
[0422] In addition, Fig.50 An example of division using an octree structure is shown in the figure, but any number of divisions and any division method may be used.
[0423] Furthermore, the segmentation unit 4911 may segment the position information and the attribute information using different segmentation methods or using the same segmentation method. Furthermore, the segmentation unit 4911 may segment the plurality of slices into tiles using different segmentation methods or using the same segmentation method.
[0424] In addition, the segmentation unit 4911 generates tile additional information related to the number of segments and the segmentation method when the tile is segmented. The tile additional information (position tile additional information and attribute tile additional information) is independent of the position information and the attribute information. For example, the tile additional information includes information indicating the reference coordinate position, size or side length of the bounding box after segmentation. In addition, the tile additional information includes information indicating the number of segments and the segmentation type, etc.
[0425] Next, an example of a method of dividing the point cloud data into slices or tiles will be described. As a method of dividing into slices or tiles, the dividing unit 4911 may use a preset method or may adaptively switch the method to be used according to the point cloud data.
[0426] When performing slice segmentation, the segmentation unit 4911 segments the three-dimensional space with respect to both the position information and the attribute information. For example, the segmentation unit 4911 determines the shape of the object and segments the three-dimensional space into slices according to the shape of the object. For example, the segmentation unit 4911 extracts objects such as trees or buildings and performs segmentation in object units. For example, the segmentation unit 4911 performs slice segmentation so that the entirety of one or more objects is included in one slice. Alternatively, the segmentation unit 4911 segments one object into a plurality of slices.
[0427] In this case, the encoding device may also change the encoding method for each slice, for example. For example, the encoding device may also use a high-quality compression method for a specific object or a specific part of an object. In this case, the encoding device may also save information indicating the encoding method for each slice in additional information (metadata).
[0428] Furthermore, the segmentation unit 4911 may perform slice segmentation based on map information or location information so that each slice corresponds to a preset coordinate space.
[0429] When segmenting the tiles, the segmentation unit 4911 segments the position information and the attribute information independently. For example, the segmentation unit 4911 segments the slice into tiles according to the data volume or the processing volume. For example, the segmentation unit 4911 determines whether the data volume of the slice (e.g., the number of three-dimensional points included in the slice) is greater than a preset threshold. The segmentation unit 4911 segments the slice into tiles when the data volume of the slice is greater than the threshold. The segmentation unit 4911 does not segment the slice into tiles when the data volume of the slice is less than the threshold.
[0430] For example, the division unit 4911 divides the slice into tiles so that the processing amount or processing time in the decoding device is within a certain range (less than a preset value). This makes the processing amount of each tile of the decoding device constant, and facilitates distributed processing in the decoding device.
[0431] Furthermore, when the processing amounts of the position information and the attribute information are different, for example, when the processing amount of the position information is greater than that of the attribute information, the division unit 4911 divides the position information into a larger number of parts than the attribute information.
[0432] Furthermore, for example, depending on the content, if the decoding device can decode and display the position information earlier and decode and display the attribute information later, the division unit 4911 can also divide the position information into more parts than the attribute information. Thus, the decoding device can increase the number of parallel operations of the position information, so that the processing of the position information can be accelerated compared to the processing of the attribute information.
[0433] Furthermore, the decoding device does not necessarily need to process the sliced or tiled data in parallel, and may determine whether to process them in parallel based on the number or capabilities of the decoding processing units.
[0434] By dividing the data using the above method, adaptive encoding corresponding to the content or object can be realized. In addition, parallel processing in the decoding process can be realized. As a result, the flexibility of the point group encoding system or the point group decoding system is improved.
[0435] Fig.51 This is a diagram showing an example of the style of segmentation of slices and tiles. The DU in the figure is a data unit (DataUnit), which represents the data of a tile or slice. In addition, each DU includes a slice index (Slice Index) and a tile index (TileIndex). The value on the upper right of the DU in the figure represents the slice index, and the value on the lower left of the DU represents the tile index.
[0436] In pattern 1, in slice partitioning, the number of partitions and the partitioning method are the same in G slices and A slices. In tile partitioning, the number of partitions and the partitioning method for G slices are different from the number of partitions and the partitioning method for A slices. In addition, the same number of partitions and the partitioning method are used among multiple G slices. The same number of partitions and the partitioning method are used among multiple A slices.
[0437] In pattern 2, in slice partitioning, the number of partitions and the partitioning method are the same in G slices and A slices. In tile partitioning, the number of partitions and the partitioning method for G slices are different from those for A slices. In addition, the number of partitions and the partitioning method are different between multiple G slices. The number of partitions and the partitioning method are different between multiple A slices.
[0438] Next, the encoding method of the segmented data is described. The three-dimensional data encoding device (first encoding unit 4910) encodes the segmented data separately. When encoding the attribute information, the three-dimensional data encoding device generates dependency information indicating which composition information (position information, additional information or other attribute information) is encoded as additional information. That is, the dependency information, for example, indicates composition information of a reference target (dependent target). In this case, the three-dimensional data encoding device generates the dependency information based on the composition information corresponding to the segmented shape of the attribute information. In addition, the three-dimensional data encoding device can also generate the dependency information based on the composition information corresponding to multiple segmented shapes.
[0439] The dependency information may also be generated by the three-dimensional data encoding device, and the generated dependency information may be sent to the three-dimensional data decoding device. Alternatively, the dependency information may be generated by the three-dimensional data decoding device, and the three-dimensional data encoding device may not send the dependency information. In addition, the dependency used by the three-dimensional data encoding device may be preset, and the three-dimensional data encoding device may not send the dependency information.
[0440] Fig.52 : is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the figure represents the dependency target, and the root of the arrow represents the dependency source. The three-dimensional data decoding device decodes the data in the order from the dependency target to the dependency source. In addition, the data represented by the solid line in the figure is the data actually sent, and the data represented by the dotted line is the data not sent.
[0441] In this figure, G represents position information and A represents attribute information. s1 Indicates the location information of slice number 1, G s2 Indicates the location information of slice number 2. G s1t1 Indicates the location information of slice number 1 and tile number 1, G s1t2 Indicates the location information of slice number 1 and tile number 2, G s2t1 Indicates the location information of slice number 2 and tile number 1, G s2t2 Indicates the location information of slice number 2 and tile number 2. Similarly, A s1 Indicates the attribute information of slice number 1, A s2 Indicates the attribute information of slice number 2. s1t1 Indicates the attribute information of slice number 1 and tile number 1, A s1t2 Indicates the attribute information of slice number 1 and tile number 2, A s2t1 Indicates the attribute information of slice number 2 and tile number 1, A s2t2 Indicates the attribute information of slice number 2 and tile number 2.
[0442] Mslice indicates additional information about slices, MGtile indicates additional information about position tiles, and MAtile indicates additional information about attribute tiles. s1t1 Indicates attribute information A s1t1 Dependency information of D s2t1 Indicates attribute information A s2t1 Dependency information.
[0443] Furthermore, the 3D data encoding device may rearrange the data in a decoding order so that the 3D data decoding device does not need to rearrange the data. Furthermore, the data may be rearranged in the 3D data decoding device, or the data may be rearranged in both the 3D data encoding device and the 3D data decoding device.
[0444] Fig.53 is a diagram showing an example of the order in which data is decoded. Fig.53 In the example, the decoding is performed sequentially from the data on the left. The three-dimensional data decoding device starts decoding from the data of the dependent target among the data in a dependent relationship. For example, the three-dimensional data encoding device rearranges the data in advance into this order and sends it out. In addition, any order is acceptable as long as the data of the dependent target comes first. In addition, the three-dimensional data encoding device may also send the additional information and the dependency relationship information before the data.
[0445] Fig.54 4 is a flowchart showing the process flow of the three-dimensional data symbolization device. First, the three-dimensional data encoding device encodes the data of a plurality of slices or tiles as described above (S4901). Next, the three-dimensional data encoding device encodes the data of a plurality of slices or tiles as described above (S4902). Fig.53 As shown, the data is rearranged in such a manner that the data depending on the target is prioritized (S4902). Next, the three-dimensional data encoding device multiplexes (converts into NAL units) the rearranged data (S4903).
[0446] Next, the configuration of the combining unit 4925 included in the first decoding unit 4920 will be described. Fig.55 4 is a block diagram showing the configuration of the combining unit 4925. The combining unit 4925 includes a position information tile combining unit 4941 (Geometry Tile Combiner), an attribute information tile combining unit 4942 (Attribute Tile Combiner), and a slice combining unit (Slice Combiner).
[0447] The position information tile combining unit 4941 combines a plurality of split position information using the position tile additional information to generate a plurality of slice position information. The attribute information tile combining unit 4942 combines a plurality of split attribute information using the attribute tile additional information to generate a plurality of slice attribute information.
[0448] The slice combining unit 4943 combines a plurality of slice position information using the slice additional information to generate position information. In addition, the slice combining unit 4943 combines a plurality of slice attribute information using the slice additional information to generate attribute information.
[0449] In addition, the number of divided slices or tiles is 1 or more. That is, the division of slices or tiles may not be performed.
[0450] In addition, although the example of performing tile division after slice division is shown here, slice division may be performed after tile division. In addition, new division types may be defined in addition to slices and tiles, and division may be performed with three or more division types.
[0451] Next, the structure of the coded data divided into slices or tiles and the method of storing the coded data in the NAL unit (multiplexing method) are described. Fig.56 This is a diagram for explaining the structure of coded data and a method of storing coded data in a NAL unit.
[0452] The encoded data (segment position information and segment attribute information) is stored in the payload of the NAL unit.
[0453] The encoded data includes a header and a payload. The header is identification information used to determine the data contained in the payload. The identification information includes, for example, the type of slice segmentation or tile segmentation (slice_type, tile_type), index information used to determine the slice or tile (slice_idx, tile_idx), location information of the data (slice or tile), or the address of the data (address), etc. The index information used to determine the slice is also recorded as a slice index (Slice Index). The index information used to determine the tile is also recorded as a tile index (Tile Index). In addition, the type of segmentation is, for example, a method based on object shape as described above, a method based on map information or location information, or a method based on data volume or processing volume, etc.
[0454] In addition, all or part of the above information may be stored in one of the headers of the segmented position information and the headers of the segmented attribute information, but not in the other. For example, when the same segmentation method is used in the position information and the attribute information, the type of segmentation (slice_type, tile_type) and index information (slice_idx, tile_idx) are the same in the position information and the attribute information. Therefore, this information may also be included in the header of one of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, this information may also be included in the header of the position information, and not included in the header of the attribute information. In this case, the three-dimensional data decoding device, for example, determines that the attribute information of the dependent source belongs to the same slice or tile as the slice or tile of the position information of the dependent target.
[0455] In addition, additional information about slice segmentation or tile segmentation (slice additional information, position tile additional information or attribute tile additional information) and dependency information indicating the dependency relationship may also be stored in an existing parameter set (GPS, APS, position SPS or attribute SPS, etc.) and sent out. In the case where the segmentation method changes for each frame, the information indicating the segmentation method may also be stored in the parameter set (GPS or APS, etc.) of each frame. In the case where the segmentation method does not change within a sequence, the information indicating the segmentation method may also be stored in the parameter set (position SPS or attribute SPS) of each sequence. Furthermore, in the case where the same segmentation method is used in the position information and the attribute information, the information indicating the segmentation method may also be stored in the parameter set (stream PS) of the PCC stream.
[0456] In addition, the above information can be stored in one of the above parameter sets or in multiple parameter sets. In addition, a parameter set for tile segmentation or slice segmentation can also be defined, and the above information can be stored in the parameter set. In addition, this information can also be stored in the header of the encoded data.
[0457] In addition, the header of the encoded data contains identification information indicating the dependency relationship. That is, when there is a dependency relationship between the data, the header contains identification information for referencing the dependent target from the dependent source. For example, the header of the data of the dependent target contains identification information for determining the data. The header of the data of the dependent source contains identification information indicating the dependent target. In addition, in the case where the identification information for determining the data, the additional information related to the slice segmentation or tile segmentation, and the identification information indicating the dependency relationship can be identified or derived based on other information, this information can also be omitted.
[0458] Next, the flow of encoding and decoding processing of point cloud data according to the present embodiment will be described. Fig.57This is a flowchart of the encoding process of the point cloud data according to the present embodiment.
[0459] First, the three-dimensional data encoding device determines the segmentation method to be used (S4911). The segmentation method includes whether to perform slice segmentation and whether to perform tile segmentation. In addition, the segmentation method may also include the number of segmentations and the type of segmentation when performing slice segmentation or tile segmentation. The type of segmentation refers to the above-mentioned method based on object shape, the method based on map information or location information, or the method based on data volume or processing volume. In addition, the segmentation method may also be preset.
[0460] When slice segmentation is performed ("Yes" in S4912), the three-dimensional data encoding device generates a plurality of slice position information and a plurality of slice attribute information by segmenting the position information and the attribute information together (S4913). In addition, the three-dimensional data encoding device generates slice additional information related to the slice segmentation. In addition, the three-dimensional data encoding device may also segment the position information and the attribute information independently.
[0461] When tile segmentation is performed ("Yes" in S4914), the three-dimensional data encoding device generates a plurality of segmentation position information and a plurality of segmentation attribute information (S4915) by independently segmenting a plurality of slice position information and a plurality of slice attribute information (or position information and attribute information). In addition, the three-dimensional data encoding device generates position tile additional information and attribute tile additional information related to tile segmentation. In addition, the three-dimensional data encoding device may also segment the slice position information and the slice attribute information together.
[0462] Next, the three-dimensional data encoding device generates a plurality of coded position information and a plurality of coded attribute information by encoding the plurality of segmentation position information and the plurality of segmentation attribute information respectively (S4916). In addition, the three-dimensional data encoding device generates dependency relationship information.
[0463] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) the plurality of encoding position information, the plurality of encoding attribute information and the additional information (S4917). In addition, the three-dimensional data encoding device sends the generated encoded data.
[0464] Fig.58 4921. The flowchart of the decoding process of the point cloud data of the present embodiment. First, the three-dimensional data decoding device determines the segmentation method by parsing the additional information (slice additional information, position tile additional information and attribute tile additional information) about the segmentation method contained in the coded data (coded stream) (S4921). The segmentation method includes whether to perform slice segmentation and whether to perform tile segmentation. In addition, the segmentation method may also include the number of segmentations and the type of segmentation when performing slice segmentation or tile segmentation.
[0465] Next, the three-dimensional data decoding apparatus generates segmentation position information and segmentation attribute information by decoding the plurality of pieces of encoding position information and the plurality of pieces of encoding attribute information included in the encoded data using the dependency information included in the encoded data (S4922).
[0466] When the additional information indicates that tile segmentation has been performed ("Yes" in S4923), the three-dimensional data decoding device combines the plurality of segmentation position information and the plurality of segmentation attribute information by respective methods based on the position tile additional information and the attribute tile additional information, thereby generating a plurality of slice position information and a plurality of slice attribute information (S4924). Alternatively, the three-dimensional data decoding device may combine the plurality of segmentation position information and the plurality of segmentation attribute information by the same method.
[0467] When the additional information indicates that the slice segmentation has been performed ("Yes" in S4925), the three-dimensional data decoding device combines the plurality of slice position information and the plurality of slice attribute information (the plurality of segmentation position information and the plurality of segmentation attribute information) in the same method based on the slice additional information, thereby generating position information and attribute information (S4926). Alternatively, the three-dimensional data decoding device may combine the plurality of slice position information and the plurality of slice attribute information in different methods.
[0468] As described above, the three-dimensional data encoding device according to this embodiment performs Fig.60 The processing shown. First, the three-dimensional data encoding device is divided into a plurality of segmented data (e.g., tiles), and the plurality of segmented data are contained in a plurality of subspaces (e.g., slices) obtained by dividing the object space containing a plurality of three-dimensional points, and each of the plurality of segmented data contains one or more three-dimensional points (S4932). Here, the segmented data is one or more data sets contained in the subspace and containing one or more three-dimensional points. In addition, the segmented data is also a space, and may also include a space that does not contain three-dimensional points. In addition, a plurality of segmented data may be contained in one subspace, or one segmented data may be contained in one subspace. In addition, a plurality of subspaces may be set in the object space, or one subspace may be set in the object space.
[0469] Next, the three-dimensional data encoding device generates a plurality of coded data corresponding to the plurality of segmented data by encoding the plurality of segmented data respectively (S4931). The three-dimensional data encoding device generates a plurality of coded data and a plurality of control information corresponding to the plurality of coded data respectively (e.g. Fig.56In each of the plurality of control information, a first identifier (e.g., slice_idx) indicating a subspace corresponding to the coded data corresponding to the control information and a second identifier (e.g., tile_idx) indicating segmented data corresponding to the coded data corresponding to the control information are stored.
[0470] Thus, a 3D data decoding device that decodes a bit stream generated by a 3D data encoding device can easily restore the target space by combining data of a plurality of divided data using the first identifier and the second identifier. Therefore, the amount of processing in the 3D data decoding device can be reduced.
[0471] For example, in the above encoding, the three-dimensional data encoding device encodes the position information and attribute information of the three-dimensional point contained in each of the plurality of segmented data. The plurality of encoded data respectively include encoded data of the position information and encoded data of the attribute information. The plurality of control information respectively include control information of the encoded data of the position information and control information of the encoded data of the attribute information. The first identifier and the second identifier are stored in the control information of the encoded data of the position information.
[0472] For example, in a bit stream, a plurality of control information are respectively arranged before the coded data corresponding to the control information.
[0473] In addition, the three-dimensional data encoding device may be that an object space including a plurality of three-dimensional points is set as one or more subspaces, and the above-mentioned subspace includes one or more segmentation data including one or more three-dimensional points; by respectively encoding the above-mentioned segmentation data, a plurality of encoded data respectively corresponding to the above-mentioned plurality of segmentation data are generated, and a bit stream including the above-mentioned plurality of encoded data and a plurality of control information respectively corresponding to the above-mentioned plurality of encoded data is generated; in each of the above-mentioned plurality of control information, a first identifier representing the subspace corresponding to the encoded data corresponding to the control information and a second identifier representing the segmentation data corresponding to the encoded data corresponding to the control information are stored.
[0474] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0475] In addition, the three-dimensional data decoding device of this embodiment performs Fig.60 First, the three-dimensional data decoding device receives a plurality of coded data and a plurality of control information (eg Fig.56The 3D data decoding device obtains a bit stream of a header (shown in FIG. 4 ), obtains a first identifier (e.g., slice_idx) and a second identifier (e.g., tile_idx) stored in the plurality of control information, the plurality of coded data being generated by respectively encoding a plurality of segmented data (e.g., tiles), the plurality of segmented data being included in a plurality of subspaces (e.g., slices) obtained by segmenting an object space including a plurality of three-dimensional points, and each including one or more three-dimensional points, the first identifier indicating the subspace corresponding to the coded data corresponding to the control information, and the second identifier indicating the segmented data corresponding to the coded data corresponding to the control information (S4941). Next, the 3D data decoding device restores the plurality of segmented data by decoding the plurality of coded data (S4942). Next, the 3D data decoding device restores the object space by combining the plurality of segmented data using the first identifier and the second identifier (S4943). For example, the 3D data coding device restores the plurality of subspaces by combining the plurality of segmented data using the second identifier, and restores the object space (a plurality of three-dimensional points) by combining the plurality of subspaces using the first identifier. Furthermore, the three-dimensional data decoding device may obtain coded data of a desired subspace or divided data from the bit stream using at least one of the first identifier and the second identifier, and selectively decode or preferentially decode the obtained coded data.
[0476] Thus, the three-dimensional data decoding device can easily restore the target space by combining the data of the plurality of divided data using the first identifier and the second identifier. Therefore, the amount of processing in the three-dimensional data decoding device can be reduced.
[0477] For example, the plurality of coded data are generated by encoding the position information and the attribute information of the three-dimensional point included in the corresponding segmented data, and include the coded data of the position information and the coded data of the attribute information. The plurality of control information include the control information of the coded data of the position information and the control information of the coded data of the attribute information. The first identifier and the second identifier are stored in the control information of the coded data of the position information.
[0478] For example, in a bitstream, the control information is arranged before the corresponding coded data.
[0479] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0480] As mentioned above, although the three-dimensional data encoding device and the three-dimensional data decoding device etc. which concern on embodiment of this invention are demonstrated, this invention is not limited to this embodiment.
[0481] In addition, each processing unit included in the three-dimensional data encoding device and the three-dimensional data decoding device according to the above-mentioned embodiment is typically implemented as an integrated circuit, that is, LSI. They may be formed individually as one chip, or part or all of them may be included as one chip.
[0482] In addition, integrated circuits are not limited to LSIs, and can also be implemented by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays) that can be programmed after LSI manufacturing, or reconfigurable processors that can reconfigure the connections or settings of circuit cells within LSIs can also be used.
[0483] In addition, in each of the above-mentioned embodiments, each component may be formed by dedicated hardware, or implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.
[0484] Furthermore, the present invention may be implemented as a three-dimensional data encoding method or a three-dimensional data decoding method executed by a three-dimensional data encoding device or a three-dimensional data decoding device.
[0485] In addition, the division of the functional blocks in the block diagram is an example, and multiple functional blocks can also be implemented as one functional block, or one functional block can be divided into multiple blocks, or a part of the function can be transferred to other functional blocks. In addition, it is also possible that a single hardware or software processes the functions of multiple functional blocks with similar functions in parallel or in a time-sharing manner.
[0486] In addition, the execution order of each step in the flowchart is exemplified for the purpose of specifically describing the present invention, and may be an order other than the above. In addition, a part of the above steps may be executed simultaneously (in parallel) with other steps.
[0487] In the above, the three-dimensional data encoding device and the three-dimensional data decoding device related to one or more technical solutions are described based on the implementation mode, but the present invention is not limited to the implementation mode. As long as it does not deviate from the main purpose of the present invention, various deformation forms that can be thought of by those skilled in the art to the present embodiment, or forms constructed by combining the constituent elements of different implementation modes can also be included in the scope of one or more technical solutions.
[0488] Industrial Applicability
[0489] The present invention can be applied to a three-dimensional data encoding device and a three-dimensional data decoding device.
[0490] Description of symbols
[0491] 4601 Three-dimensional data encoding system
[0492] 4602 3D data decoding system
[0493] 4603 Sensor Terminal
[0494] 4604 External connection
[0495] 4611 Point Group Data Generation System
[0496] 4612 Prompt Department
[0497] 4613 Coding Department
[0498] 4614 Multiplexing Department
[0499] 4615 Input and Output
[0500] 4616 Control Department
[0501] 4617 Sensor information acquisition unit
[0502] 4618 Point Group Data Generation Department
[0503] 4621 Sensor information acquisition unit
[0504] 4622 Input and Output
[0505] 4623 Inverse Multiplexing Department
[0506] 4624 Decoding Department
[0507] 4625 Prompt Department
[0508] 4626 User Interface
[0509] 4627 Control Department
[0510] 4630 1st Coding Department
[0511] 4631 Position Information Coding Unit
[0512] 4632 Attribute Information Coding Unit
[0513] 4633 Additional Information Coding Department
[0514] 4634 Multiplexing Department
[0515] 4640 Decoding Unit 1
[0516] 4641 Inverse Multiplexing Department
[0517] 4642 Position information decoding unit
[0518] 4643 Attribute information decoding unit
[0519] 4644 Additional information decoding unit
[0520] 4650 No. 2 Coding Department
[0521] 4651 Additional Information Generation Department
[0522] 4652 Position Image Generation Unit
[0523] 4653 Attribute Image Generation Unit
[0524] 4654 Video Coding Department
[0525] 4655 Additional Information Coding Department
[0526] 4656 Multiplexing Department
[0527] 4660 Decoding Unit 2
[0528] 4661 Inverse Multiplexing Unit
[0529] 4662 Image Decoding Department
[0530] 4663 Additional information decoding unit
[0531] 4664 Position Information Generation Unit
[0532] 4665 Attribute Information Generation Unit
[0533] 4670 Coding Department
[0534] 4671 Multiplexing Department
[0535] 4680 Decoding Department
[0536] 4681 Inverse Multiplexing Department
[0537] 4710 No.1 Multiplexing Department
[0538] 4711 File Conversion Department
[0539] 4720 1st Inverse Multiplexing Unit
[0540] 4721 File Inverse Transformation Unit
[0541] 4730 Second Multiplexing Unit
[0542] 4731 File Conversion Department
[0543] 4740 Second inverse multiplexing unit
[0544] 4741 File Inverse Conversion Unit
[0545] 4750 No. 3 Multiplexing Unit
[0546] 4751 File Conversion Department
[0547] 4760 The 3rd Inverse Multiplexing Unit
[0548] 4761 File Inverse Transformation Unit
[0549] 4801 Coding Department
[0550] 4802 Multiplexing Department
[0551] 4910 1st Coding Department
[0552] 4911 Division
[0553] 4912 Position Information Coding Unit
[0554] 4913 Attribute Information Coding Unit
[0555] 4914 Additional Information Coding Department
[0556] 4915 Multiplexing Department
[0557] 4920 Decoding Unit 1
[0558] 4921 Inverse Multiplexing Department
[0559] 4922 Position information decoding unit
[0560] 4923 Attribute Information Decoding Unit
[0561] 4924 Additional information decoding unit
[0562] 4925 Joint
[0563] 4931 Slice Division
[0564] 4932 Position Information Tile Segmentation Unit
[0565] 4933 Attribute Information Tile Division
[0566] 4941 Position information tile joint
[0567] 4942 Attribute information tile joint
[0568] 4943 Slice joint
Claims
1. A three-dimensional data encoding method, wherein: encoding position information and attribute information of three-dimensional points included in each of a plurality of segmentation data, wherein the plurality of segmentation data is included in a plurality of subspaces obtained by segmenting an object space including the plurality of three-dimensional points; generating a plurality of coded data corresponding to the plurality of divided data, respectively, and respectively including the coded data of the position information and the coded data of the attribute information; Generate a bit stream including the plurality of coded data and a plurality of control information respectively corresponding to the plurality of coded data; In each of the plurality of control information, a first identifier indicating a subspace corresponding to the coded data corresponding to the control information and a second identifier indicating divided data corresponding to the coded data corresponding to the control information are stored.
2. The three-dimensional data encoding method according to claim 1, wherein: The plurality of control information respectively include control information of the coded data of the position information and control information of the coded data of the attribute information; The first identifier and the second identifier are stored in the control information of the encoded data of the position information.
3. The three-dimensional data encoding method according to claim 1 or 2, wherein: In the bit stream, the plurality of control information are respectively arranged before the coded data corresponding to the control information.
4. A three-dimensional data decoding method, wherein: Obtaining, from a bit stream including a plurality of coded data and a plurality of control information respectively corresponding to the plurality of coded data, a first identifier and a second identifier stored in the plurality of control information, wherein the plurality of coded data are generated by respectively encoding a plurality of segmented data, the plurality of segmented data being included in a plurality of subspaces obtained by segmenting an object space including a plurality of three-dimensional points, the first identifier indicating the subspace corresponding to the coded data corresponding to the control information, and the second identifier indicating the segmented data corresponding to the coded data corresponding to the control information; By decoding the plurality of coded data, the plurality of segmented data are restored; restoring the object space by combining the plurality of divided data using the first identifier and the second identifier; The plurality of encoded data respectively include encoded data of position information of three-dimensional points and encoded data of attribute information.
5. The three-dimensional data decoding method according to claim 4, wherein: The plurality of control information respectively include control information of the coded data of the position information and control information of the coded data of the attribute information; The first identifier and the second identifier are stored in the control information of the encoded data of the position information.
6. The three-dimensional data decoding method according to claim 4 or 5, wherein: In the above bit stream, the above control information is arranged before the corresponding coded data.
7. A three-dimensional data encoding device, wherein: have: Processor; and Memory; The processor uses the memory to perform the following processing: encoding position information and attribute information of three-dimensional points included in each of a plurality of segmentation data, wherein the plurality of segmentation data is included in a plurality of subspaces obtained by segmenting an object space including the plurality of three-dimensional points; generating a plurality of coded data corresponding to the plurality of divided data, respectively, and respectively including the coded data of the position information and the coded data of the attribute information; Generate a bit stream including the plurality of coded data and a plurality of control information respectively corresponding to the plurality of coded data; In each of the plurality of control information, a first identifier indicating a subspace corresponding to the coded data corresponding to the control information and a second identifier indicating divided data corresponding to the coded data corresponding to the control information are stored.
8. A three-dimensional data decoding device, wherein: have: Processor; and Memory; The processor uses the memory to perform the following processing: Obtaining, from a bit stream including a plurality of coded data and a plurality of control information respectively corresponding to the plurality of coded data, a first identifier and a second identifier stored in the plurality of control information, wherein the plurality of coded data are generated by respectively encoding a plurality of segmented data, the plurality of segmented data being included in a plurality of subspaces obtained by segmenting an object space including a plurality of three-dimensional points, the first identifier indicating the subspace corresponding to the coded data corresponding to the control information, and the second identifier indicating the segmented data corresponding to the coded data corresponding to the control information; By decoding the plurality of coded data, the plurality of segmented data are restored; restoring the object space by combining the plurality of divided data using the first identifier and the second identifier; The plurality of encoded data respectively include encoded data of position information of three-dimensional points and encoded data of attribute information.
Citation Information
Patent Citations
Map display device
WO2014020663A1
Scalable point cloud compression with transform, and corresponding decompression
US20170347122A1
Display method and display device
WO2018083999A1