3D data encoding method, decoding method, encoding device, decoding device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2017-05-23
- Publication Date
- 2026-08-14
AI Technical Summary
虽然预想点云数据作为三维数据的表现方法将成为主流,但是,点群的数据量非常大
[0017]本申请能够提供一种能够减少传输时的数据量的三维数据编码方法、三维数据解码方法、三维数据编码装置、或三维数据解码装置。
Smart Images

Figure CN116630452B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on May 23, 2017, with application number 201780036423.0 (international application number PCT / JP2017 / 019114) and entitled "Three-dimensional data encoding method, decoding method, encoding device, decoding device". Technical Field
[0002] This application relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. Background Technology
[0003] In major fields such as computer vision, mapping, surveillance, infrastructure inspection, and image distribution for autonomous operation of automobiles or robots, devices and services that flexibly utilize 3D data will become increasingly common in the future. 3D data is obtained through various methods, including distance sensors such as rangefinders, stereo cameras, or combinations of multiple monocular cameras.
[0004] One method of representing 3D data is point cloud data, which uses groups of points in 3D space to represent the shape of a 3D structure (see, for example, Non-Patent Document 1). The point cloud data stores the position and color of the point groups. While point cloud data is expected to become the mainstream method for representing 3D data, the data volume of point groups is extremely large. Therefore, in the accumulation or transmission of 3D data, similar to 2D dynamic images (for example, MPEG-4 AVC or HEVC standardized by MPEG), data compression through encoding is necessary.
[0005] Furthermore, the compression of point cloud data is partly supported by publicly available libraries such as the Point Cloud Library, which perform point cloud data association processing.
[0006] Existing technical documents
[0007] Non-patent literature
[0008] Non-patent document 1: "Octree-Based Progressive Geometry Coding of PointClouds", Eurographics Symposium on Point-Based Graphics (2006) Summary of the Invention
[0009] The problem that the invention aims to solve
[0010] Compared to two-dimensional data, this type of three-dimensional data is much larger in volume, and the amount of three-dimensional encoded data to be transmitted is also enormous.
[0011] The purpose of this application is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can reduce the amount of data transmitted.
[0012] The methods used to solve the problem
[0013] One embodiment of this application relates to a three-dimensional data encoding method comprising: an extraction step of extracting a second three-dimensional data from a first three-dimensional data having a feature quantity of a threshold or higher; and a first encoding step of generating first encoded three-dimensional data by encoding the second three-dimensional data.
[0014] One embodiment of this application relates to a three-dimensional data decoding method comprising: a first decoding step, wherein the first encoded three-dimensional data is obtained by encoding the first three-dimensional data, wherein the feature quantity extracted from the first three-dimensional data is above a threshold; and a second decoding step, wherein the second encoded three-dimensional data obtained by encoding the first three-dimensional data is decoded by a second decoding method different from the first decoding method.
[0015] Furthermore, all or specific forms of these can be realized as systems, methods, integrated circuits, computer programs, or computer-readable CD-ROMs and other recording media, and can be realized by combining systems, methods, integrated circuits, computer programs, and recording media.
[0016] The effects of the invention
[0017] This application can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can reduce the amount of data transmitted. Attached Figure Description
[0018] Figure 1 The structure of the encoded three-dimensional data involved in Implementation 1 is shown.
[0019] Figure 2 An example of the prediction structure between SPCs at the lowest level of the GOS involved in Implementation 1 is shown.
[0020] Figure 3 An example of the interlayer prediction structure involved in Implementation 1 is shown.
[0021] Figure 4 An example of the encoding order of the GOS involved in Implementation 1 is shown.
[0022] Figure 5 An example of the encoding order of the GOS involved in Implementation 1 is shown.
[0023] Figure 6 This is a block diagram of the three-dimensional data encoding device according to Embodiment 1.
[0024] Figure 7 This is a flowchart of the encoding process involved in Implementation Method 1.
[0025] Figure 8 This is a block diagram of the three-dimensional data decoding device according to Embodiment 1.
[0026] Figure 9 This is a flowchart of the decoding process involved in Implementation Method 1.
[0027] Figure 10 An example of the metadata involved in Implementation 1 is shown.
[0028] Figure 11 An example of the configuration of the SWLD involved in Embodiment 2 is shown.
[0029] Figure 12 An example of the operation of the server and client involved in Implementation Method 2 is shown.
[0030] Figure 13 An example of the operation of the server and client involved in Implementation Method 2 is shown.
[0031] Figure 14 An example of the operation of the server and client involved in Implementation Method 2 is shown.
[0032] Figure 15 An example of the operation of the server and client involved in Implementation Method 2 is shown.
[0033] Figure 16 This is a block diagram of the three-dimensional data encoding device involved in Embodiment 2.
[0034] Figure 17 This is a flowchart of the encoding process involved in Implementation Method 2.
[0035] Figure 18 This is a block diagram of the three-dimensional data decoding device involved in Embodiment 2.
[0036] Figure 19 This is a flowchart of the decoding process involved in Implementation Method 2.
[0037] Figure 20 An example of the configuration of the WLD involved in Embodiment 2 is shown.
[0038] Figure 21 An example of an octree structure of the WLD involved in Embodiment 2 is shown.
[0039] Figure 22 An example of the configuration of the SWLD involved in Embodiment 2 is shown.
[0040] Figure 23 An example of the octree structure of the SWLD involved in Implementation 2 is shown. Detailed Implementation
[0041] When using coded data such as point cloud data in actual devices or services, random access is required for the desired spatial location or target object. However, random access as a function does not exist in 3D coded data to date, and therefore, there is no encoding method for this purpose.
[0042] This application provides a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can provide random access functionality in encoded three-dimensional data.
[0043] One aspect of this application relates to a three-dimensional data encoding method for encoding three-dimensional data. The three-dimensional data encoding method includes: a partitioning step, dividing the three-dimensional data into first processing units corresponding to three-dimensional coordinates, wherein the first processing unit is a random access unit; and an encoding step, generating encoded data by encoding each of the plurality of first processing units.
[0044] Accordingly, random access for each first processing unit becomes possible. Thus, this three-dimensional data encoding method can provide random access functionality in encoded three-dimensional data.
[0045] For example, the three-dimensional data encoding method may include a generation step in which first information is generated, the first information showing a plurality of first processing units and three-dimensional coordinates corresponding to each of the plurality of first processing units, the encoded data including the first information.
[0046] For example, the first information may further indicate at least one of the objects, times, and data storage destinations corresponding to each of the plurality of first processing units.
[0047] For example, in the partitioning step, the first processing unit may be further divided into a plurality of second processing units, and in the encoding step, each of the plurality of second processing units may be encoded.
[0048] For example, in the encoding step, the encoding can be performed with reference to other second processing units contained in the first processing unit of the processing object, for the second processing unit of the processing object contained in the first processing unit of the processing object.
[0049] Therefore, by referring to other second processing units, coding efficiency can be improved.
[0050] For example, in the encoding step, the type of the second processing unit of the processing object can be selected from a first type that does not refer to other second processing units, a second type that refers to one other second processing unit, and a third type that refers to two other second processing units, and the second processing unit of the processing object can be encoded according to the selected type.
[0051] For example, in the encoding step, the frequency of selecting the first type can be varied according to the number or density of objects contained in the three-dimensional data.
[0052] Therefore, it is possible to appropriately set the random access and coding efficiency in the compromise relationship.
[0053] For example, in the encoding step, the size of the first processing unit can be determined according to the number or density of objects contained in the three-dimensional data, or the number or density of dynamic objects.
[0054] Therefore, it is possible to appropriately set the random access and coding efficiency in the compromise relationship.
[0055] For example, the first processing unit may include multiple layers spatially divided in a predetermined direction, each of which includes one or more second processing units. In the encoding step, the second processing unit is encoded by referring to the second processing units included in the same layer as or below the second processing unit.
[0056] Accordingly, for example, it is possible to improve the random accessibility of important layers in the system and suppress the reduction in coding efficiency.
[0057] For example, in the partitioning step, the second processing unit that includes only static objects and the second processing unit that includes only dynamic objects may be assigned to different first processing units.
[0058] Therefore, dynamic and static objects can be easily controlled.
[0059] For example, in the encoding step, multiple dynamic objects may be encoded separately, and the encoded data of the multiple dynamic objects may correspond to a second processing unit that includes only static objects.
[0060] Therefore, dynamic and static objects can be easily controlled.
[0061] For example, in the partitioning step, the second processing unit may be further divided into a plurality of third processing units, and in the encoding step, each of the plurality of third processing units may be encoded.
[0062] For example, the third processing unit may include one or more voxels, which are the smallest units corresponding to location information.
[0063] For example, the second processing unit may include a group of feature points derived from information obtained from the sensor.
[0064] For example, the encoded data may include information indicating the encoding order of the plurality of the first processing units.
[0065] For example, the encoded data may include information indicating the size of the plurality of the first processing units.
[0066] For example, in the encoding step, multiple first processing units may be encoded in parallel.
[0067] Furthermore, one embodiment of this application relates to a three-dimensional data decoding method including a decoding step, in which three-dimensional data of the first processing unit is generated by decoding each of the encoded data of the first processing unit corresponding to the three-dimensional coordinates, wherein the first processing unit is a random access unit.
[0068] Accordingly, random access to each first processing unit becomes possible. Thus, this 3D data decoding method can provide random access functionality within encoded 3D data.
[0069] Alternatively, one embodiment of this application may include a three-dimensional data encoding device comprising: a division unit that divides the three-dimensional data into first processing units corresponding to three-dimensional coordinates, the first processing units being random access units; and an encoding unit that generates encoded data by encoding each of the plurality of the first processing units.
[0070] Accordingly, random access per first processing unit becomes possible. Thus, the three-dimensional data encoding device is able to provide random access functionality in encoding three-dimensional data.
[0071] Alternatively, one embodiment of this application may involve a three-dimensional data decoding device that decodes three-dimensional data. The three-dimensional data decoding device includes a decoding unit that generates three-dimensional data of the first processing unit by decoding each of the encoded data of the first processing unit corresponding to the three-dimensional coordinates. The first processing unit is a random access unit.
[0072] Accordingly, random access per first processing unit becomes possible. Thus, the three-dimensional data decoding device is able to provide random access functionality for encoded three-dimensional data.
[0073] Furthermore, by dividing and encoding the space, this application enables the quantification and prediction of the space, which is effective even without random access.
[0074] Furthermore, one embodiment of this application relates to a three-dimensional data encoding method comprising: an extraction step of extracting a second three-dimensional data from the first three-dimensional data, wherein the feature quantity is above a threshold; and a first encoding step of generating first encoded three-dimensional data by encoding the second three-dimensional data.
[0075] Accordingly, this 3D data encoding method generates first-encoded 3D data by encoding data with feature values above a certain threshold. This reduces the amount of data required to encode the 3D data compared to directly encoding the first-encoded data. Therefore, this 3D data encoding method can reduce the amount of data transmitted.
[0076] For example, the three-dimensional data encoding method may further include a second encoding step, in which the first three-dimensional data is encoded to generate the second encoded three-dimensional data.
[0077] Accordingly, this three-dimensional data encoding method can selectively transmit the first encoded three-dimensional data and the second encoded three-dimensional data according to their intended use, for example.
[0078] For example, the second three-dimensional data may be encoded by the first encoding method, or the first three-dimensional data may be encoded by a second encoding method different from the first encoding method.
[0079] Accordingly, this three-dimensional data encoding method can employ appropriate encoding methods for the first three-dimensional data and the second three-dimensional data respectively.
[0080] For example, in the first coding method, compared with the second coding method, intra-frame prediction and inter-frame prediction are given priority.
[0081] Accordingly, this three-dimensional data encoding method can improve the priority of inter-frame prediction for second-dimensional data where the correlation between adjacent data is prone to decrease.
[0082] For example, the representation of three-dimensional position may differ between the first encoding method and the second encoding method.
[0083] Accordingly, this three-dimensional data encoding method can employ more appropriate representations of three-dimensional positions for three-dimensional data with varying amounts of data.
[0084] For example, at least one of the first encoded three-dimensional data and the second encoded three-dimensional data may include an identifier indicating whether the encoded three-dimensional data is obtained by encoding the first three-dimensional data or by encoding a portion of the first three-dimensional data.
[0085] Therefore, the decoding device can easily determine whether the obtained encoded three-dimensional data is the first encoded three-dimensional data or the second encoded three-dimensional data.
[0086] For example, in the first encoding step, the second three-dimensional data can be encoded in such a way that the amount of data in the first encoded three-dimensional data is smaller than the amount of data in the second encoded three-dimensional data.
[0087] Accordingly, this three-dimensional data encoding method enables the amount of data in the first encoded three-dimensional data to be smaller than the amount of data in the second encoded three-dimensional data.
[0088] For example, in the extraction step, data corresponding to objects with predefined properties can be further extracted from the first three-dimensional data as the second three-dimensional data.
[0089] Accordingly, the three-dimensional data encoding method can generate first encoded three-dimensional data including the data required by the decoding device.
[0090] For example, the three-dimensional data encoding method may further include a sending step, in which one of the first encoded three-dimensional data and the second encoded three-dimensional data is sent to the client according to the client's state.
[0091] Accordingly, this three-dimensional data encoding method can send appropriate data according to the client's state.
[0092] For example, the client's status could include the client's communication status or the client's movement speed.
[0093] For example, the three-dimensional data encoding method may further include a sending step, in which, according to the client's request, one of the first encoded three-dimensional data and the second encoded three-dimensional data is sent to the client.
[0094] Therefore, this three-dimensional data encoding method can send appropriate data according to the client's request.
[0095] Furthermore, one embodiment of this application relates to a three-dimensional data decoding method comprising: a first decoding step, wherein a first decoder method is used to decode first encoded three-dimensional data, the first encoded three-dimensional data being obtained by encoding second three-dimensional data from the first three-dimensional data, wherein the feature quantity extracted from the first three-dimensional data is above a threshold; and a second decoding step, wherein a second decoder method, different from the first decoder method, is used to decode the second encoded three-dimensional data obtained by encoding the first three-dimensional data.
[0096] Accordingly, this 3D data decoding method can selectively receive, for example, data obtained by encoding first-encoded 3D data and second-encoded 3D data whose feature values are above a threshold, according to their intended use. Consequently, this 3D data decoding method can reduce the amount of data transmitted. Furthermore, this 3D data decoding method can employ appropriate decoding methods for the first-encoded 3D data and the second-encoded 3D data respectively.
[0097] For example, in the first decoding method, inter-frame prediction is prioritized over intra-frame prediction in the second decoding method.
[0098] Accordingly, this 3D data decoding method can improve the priority of inter-frame prediction for second-dimensional data where the correlation between adjacent data is prone to decrease.
[0099] For example, the representation of three-dimensional position may differ between the first decoding method and the second decoding method.
[0100] Accordingly, this 3D data decoding method can employ more appropriate 3D position representation techniques for 3D data with different data volumes.
[0101] For example, at least one of the first encoded three-dimensional data and the second encoded three-dimensional data may include an identifier indicating whether the encoded three-dimensional data is obtained by encoding the first three-dimensional data or by encoding a portion of the first three-dimensional data, and the first encoded three-dimensional data and the second encoded three-dimensional data are identified with reference to the identifier.
[0102] Therefore, this three-dimensional data decoding method can easily determine whether the obtained encoded three-dimensional data is first-encoded three-dimensional data or second-encoded three-dimensional data.
[0103] For example, the three-dimensional data decoding method may further include: a notification step, which notifies the server of the client's status; and a receiving step, which receives, according to the client's status, one of the first encoded three-dimensional data and the second encoded three-dimensional data sent from the server.
[0104] Therefore, this 3D data decoding method can receive appropriate data according to the client's state.
[0105] For example, the client's status could include the client's communication status or the client's movement speed.
[0106] For example, the three-dimensional data decoding method may further include: a request step, requesting one of the first encoded three-dimensional data and the second encoded three-dimensional data from a server; and a receiving step, receiving the first encoded three-dimensional data and the second encoded three-dimensional data sent from the server according to the request.
[0107] Accordingly, this three-dimensional data decoding method can receive appropriate data corresponding to the intended use.
[0108] Furthermore, one embodiment of the present application relates to a three-dimensional data encoding apparatus comprising: an extraction unit that extracts second three-dimensional data from first three-dimensional data having a feature quantity of a threshold or higher; and a first encoding unit that generates first encoded three-dimensional data by encoding the second three-dimensional data.
[0109] Accordingly, the three-dimensional data encoding device generates first encoded three-dimensional data by encoding data whose feature values are above a threshold. This reduces the amount of data compared to directly encoding the first three-dimensional data. Therefore, the three-dimensional data encoding device can reduce the amount of data transmitted.
[0110] Furthermore, one embodiment of the present application relates to a three-dimensional data decoding apparatus comprising: a first decoding unit that decodes first encoded three-dimensional data using a first decoding method, the first encoded three-dimensional data being obtained by encoding second three-dimensional data from the first three-dimensional data, wherein the feature quantity extracted from the first three-dimensional data is above a threshold; and a second decoding unit that decodes the second encoded three-dimensional data obtained by encoding the first three-dimensional data using a second decoding method different from the first decoding method.
[0111] Accordingly, this 3D data decoding device can selectively receive, for example, first encoded 3D data and second encoded 3D data obtained by encoding data with feature values above a threshold, according to their intended use. This reduces the amount of data transmitted. Furthermore, the device can employ appropriate decoding methods for both the first and second 3D data.
[0112] Furthermore, these general or specific forms can be realized by systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, and can be realized by any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0113] The embodiments will now be described in detail with reference to the accompanying drawings. Furthermore, the embodiments described below are all specific examples illustrating this application. The numerical values, shapes, materials, constituent elements, arrangement positions of constituent elements, connection methods, steps, and order of steps shown in the following embodiments are all examples and are not intended to limit this application. Moreover, constituent elements not described in the technical solution illustrating the highest-level concept among the constituent elements of the following embodiments are described as arbitrary constituent elements.
[0114] (Implementation Method 1)
[0115] First, the data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) involved in this embodiment will be explained. Figure 1 The structure of the encoded three-dimensional data involved in this embodiment is shown.
[0116] In this embodiment, the three-dimensional space is divided into spatial units (SPCs) equivalent to image space in motion image coding, and the three-dimensional data is encoded on a spatial basis. The space is further divided into volumes (VLMs) equivalent to macroblocks in motion image coding, and prediction and transformation are performed on a VLM basis. A volume includes multiple voxels (VXLs), the smallest unit corresponding to position coordinates. Furthermore, prediction refers to generating predicted three-dimensional data similar to the processing unit of the object being processed, referencing other processing units, similar to the prediction performed in two-dimensional images, and encoding the difference between the predicted three-dimensional data and the processing unit of the object being processed. Moreover, this prediction includes not only spatial prediction referencing other prediction units at the same time, but also temporal prediction referencing prediction units at different times.
[0117] For example, when encoding a three-dimensional space represented by point cloud data or other point group data, a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes each point of the point group or multiple points contained within a voxel according to the size of the voxel. If the voxels are subdivided, the three-dimensional shape of the point group can be represented with high precision; if the size of the voxels is increased, the three-dimensional shape of the point group can be represented with coarse precision.
[0118] Furthermore, although the following explanation uses point cloud data as an example, 3D data is not limited to point cloud data and can be any form of 3D data.
[0119] Furthermore, voxels with hierarchical structures can be utilized. In this case, within an nth-order hierarchy, the presence or absence of sampling points in the (n-1)th-order and lower-order hierarchy (the lower layers of the nth-order hierarchy) can be sequentially shown. For example, when decoding only the nth-order hierarchy, if sampling points exist in the (n-1)th-order and lower-order hierarchy, the decoding can be performed by considering the center of the voxel in the nth-order hierarchy as the location of the sampling point.
[0120] Furthermore, the encoding device obtains point group data through distance sensors, stereo cameras, SLR cameras, gyroscopes, or inertial sensors.
[0121] Regarding spatial representation, similar to the coding of moving images, it is classified into at least one of the following three prediction structures: intra-frame spatial representation (I-SPC) which can be decoded independently, prediction spatial representation (P-SPC) which can only be referenced in one direction, and bidirectional spatial representation (B-SPC) which can be referenced in both directions. Furthermore, the spatial representation has both decoding time and display time information.
[0122] And, as Figure 1 As shown, as a processing unit comprising multiple spaces, there is a Group of Spaces (GOS) which serves as a random access unit. Furthermore, as a processing unit comprising multiple GOS, there exists a World Space (WLD).
[0123] The spatial regions occupied by the world are mapped to absolute locations on Earth using GPS, latitude, and longitude information. This location information is stored as metadata. Furthermore, metadata can be included within coded data or transmitted separately from it.
[0124] Furthermore, within GOS, all SPCs can be three-dimensionally adjacent, or they can exist in a way that is not three-dimensionally adjacent to other SPCs.
[0125] Furthermore, the encoding, decoding, or referencing processes corresponding to the 3D data contained in processing units such as GOS, SPC, or VLM will also be simply referred to as encoding, decoding, or referencing the processing unit. The 3D data contained in the processing unit includes, for example, at least one set of characteristic values such as spatial position (3D coordinates) and color information.
[0126] Next, the prediction structure of SPC in GOS will be explained. Multiple SPCs within the same GOS, or multiple VLMs within the same SPC, although occupying different spaces, hold the same timing information (decoding timing and display timing).
[0127] Furthermore, within the GOS, the SPC that begins the decoding order is the I-SPC. There are also two types of GOS: closed GOS and open GOS. A closed GOS is a GOS that can decode all SPCs within the GOS starting from the first I-SPC. In an open GOS, a subset of SPCs whose display time is earlier than the first I-SPC references a different GOS and can only be decoded within that GOS.
[0128] Furthermore, in encoded data such as map information, there are cases where WLD is decoded in the reverse direction of the encoding order. If there is a dependency between GOS, it is difficult to perform reverse regeneration. Therefore, in such cases, a closed GOS is generally used.
[0129] Furthermore, GOS has a layered structure in the height direction, and encoding or decoding is performed sequentially starting from the SPC of the bottom layer.
[0130] Figure 2 An example of the prediction structure between SPCs in the lowest layer of GOS is shown. Figure 3 An example of the predicted structure between layers is shown.
[0131] There are more than one I-SPC within a GOS. Although objects such as people, animals, cars, bicycles, traffic lights, or buildings that serve as land landmarks exist in three-dimensional space, encoding small objects as I-SPCs is particularly effective. For example, a 3D data decoding device (hereinafter also referred to as the decoding device) decodes only the I-SPCs within the GOS when decoding the GOS with low processing power or high speed.
[0132] Furthermore, the encoding device can switch the encoding interval or occurrence frequency of I-SPC according to the density of objects within the WLD.
[0133] Furthermore, in Figure 3In the configuration shown, the encoding or decoding device encodes or decodes multiple layers sequentially, starting from the lower layer (layer 1). Accordingly, for example, for autonomous vehicles, the priority of data near the ground, which contains a large amount of information, can be increased.
[0134] In addition, in the encoded data used by drones, etc., within GOS, encoding or decoding can be performed sequentially starting from the SPC of the upper layer in the height direction.
[0135] Furthermore, the encoding or decoding device can encode or decode multiple layers in a manner that allows the decoding device to roughly grasp the GOS and gradually increase the resolution. For example, the encoding or decoding device can encode or decode in the order of layers 3, 8, 1, 9, etc.
[0136] Next, the corresponding methods for static and dynamic objects will be explained.
[0137] In three-dimensional space, there exist static objects or scenes such as buildings or roads (hereinafter referred to as static objects) and dynamic objects such as vehicles or people (hereinafter referred to as dynamic objects). Object detection can be performed separately by extracting feature points from point cloud data or images captured by stereo cameras. Here, an example of an encoding method for dynamic objects is illustrated.
[0138] The first method is to encode objects without distinguishing between static and dynamic objects. The second method is to distinguish between static and dynamic objects by identifying information.
[0139] For example, GOS is used as the identification unit. In this case, the GOS that constitutes the SPC of a static object and the GOS that constitutes the SPC of a dynamic object are distinguished either within the encoded data or by identification information stored separately from the encoded data.
[0140] Alternatively, the SPC can be used as the identification unit. In this case, the SPC that includes only the VLM constituting a static object and the SPC that includes the VLM constituting a dynamic object are distinguished by the identification information described above.
[0141] Alternatively, VLM or VXL can be used as the identification unit. In this case, VLM or VXL including static objects and VLM or VXL including dynamic objects are distinguished by the identification information described above.
[0142] Furthermore, the encoding device can encode a dynamic object as one or more VLMs or SPCs, and encode VLMs or SPCs including static objects and SPCs including dynamic objects as different GOSs. Moreover, when the size of the GOS becomes variable according to the size of the dynamic object, the encoding device stores the size of the GOS separately as metadata.
[0143] Furthermore, the encoding device encodes static and dynamic objects independently, allowing dynamic objects to overlap within a world space composed of static objects. In this case, a dynamic object is composed of one or more SPCs, and each SPC corresponds to one or more SPCs of the static object that overlaps with it. Alternatively, a dynamic object may not be represented by an SPC, but rather by one or more VLMs or VXLs.
[0144] Furthermore, the encoding device can encode static objects and dynamic objects as distinct streams.
[0145] Furthermore, the encoding device can also generate GOS that includes one or more SPCs constituting a dynamic object. Moreover, the encoding device can set the GOS (GOS_M) including the dynamic object and the GOS of the static object corresponding to the spatial region of GOS_M to be of the same size (occupying the same spatial region). In this way, overlapping processing can be performed on a GOS-by-GOS basis.
[0146] The P-SPC or B-SPC that constitutes a dynamic object can also refer to the SPCs contained in different encoded GOS. Since the position of a dynamic object changes over time, and the same dynamic object is encoded as a GOS at different times, referencing across GOSs is effective from a compression ratio perspective.
[0147] Furthermore, the first and second methods described above can be switched depending on the intended use of the encoded data. For example, when encoding 3D data for use as a map, the encoding device uses the second method because it is desirable to separate it from dynamic objects. Conversely, when encoding 3D data for events such as concerts or sporting events, the encoding device uses the first method if separation of dynamic objects is not required.
[0148] Furthermore, the decoding and display times of GOS or SPC can be stored within the encoded data or as metadata. Also, the timing information for static objects can all be identical. In this case, the actual decoding and display times can be determined by the decoding device. Alternatively, different values can be assigned to each GOS or SPC as the decoding time, while the same value can be assigned to all display times. Moreover, as shown in the decoder modes of motion graphics coding such as HEVC's HRD (Hypothetical Reference Decoder), the decoder has a buffer of a specified size. As long as the bitstream is read at a specified bit rate according to the decoding time, a model that will not be corrupted and can be decoded can be imported.
[0149] Next, the configuration of GOS within world space will be explained. The three-dimensional coordinates in world space are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, z-axis). By setting prescribed rules in the encoding order of GOS, spatially adjacent GOS can be encoded consecutively within the encoded data. For example, in... Figure 4 In the example shown, the World Space (GOS) within the xz plane is encoded sequentially. After encoding all GOS within an xz plane, the y-axis value is updated. That is, as encoding continues, the world space extends towards the y-axis. Furthermore, the index number of the GOS is set as the encoding order.
[0150] Here, the three-dimensional space of the world corresponds one-to-one with absolute geographical coordinates such as GPS, latitude, and longitude. Alternatively, the three-dimensional space can be represented by relative positions relative to a pre-defined reference position. The directions of the x, y, and z axes of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, and these direction vectors are stored as metadata along with the encoded data.
[0151] Furthermore, the size of the GOS is set to a fixed value, and the encoding device stores this size as metadata. The size of the GOS can be switched, for example, depending on whether it is indoors or outdoors, or whether it is in a city. That is, the size of the GOS can be switched according to the quantity or nature of objects with informational value. Alternatively, the encoding device can appropriately switch the size of the GOS or the interval of the I-SPCs within the GOS according to the density of objects within the same world space. For example, the higher the object density, the smaller the size of the GOS and the shorter the interval of the I-SPCs within the GOS.
[0152] exist Figure 5In the example, in the region from the 3rd to the 10th GOS, due to the high density of objects, the GOS is subdivided to achieve fine-grained random access. Furthermore, the 7th to 10th GOS are located on the back sides of the 3rd to 6th GOS, respectively.
[0153] Next, the structure and operation flow of the three-dimensional data encoding device involved in this embodiment will be explained. Figure 6 This is a block diagram of the three-dimensional data encoding device 100 according to this embodiment. Figure 7 This is a flowchart illustrating an example of the operation of the three-dimensional data encoding device 100.
[0154] Figure 6 The 3D data encoding apparatus 100 shown generates encoded 3D data 112 by encoding 3D data 111. The 3D data encoding apparatus 100 includes: an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.
[0155] like Figure 7 As shown, firstly, the acquisition unit 101 acquires three-dimensional data 111 as point group data (S101).
[0156] Next, the encoding region determination unit 102 determines the region of the encoding object from the spatial region corresponding to the obtained point group data (S102). For example, the encoding region determination unit 102 determines the spatial region surrounding the user or vehicle's location as the region of the encoding object.
[0157] Next, the partitioning unit 103 divides the point group data contained in the region of the encoding object into various processing units. Here, the processing units are the aforementioned GOS and SPC, etc. Furthermore, the region of the encoding object corresponds, for example, to the aforementioned world space. Specifically, the partitioning unit 103 divides the point group data into processing units based on a preset GOS size, the presence or size of dynamic objects (S103). Furthermore, the partitioning unit 103 determines the starting position of the SPC that will be the first in the encoding sequence within each GOS.
[0158] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding multiple SPCs within each GOS (S104).
[0159] Furthermore, although an example of encoding each GOS is shown here after dividing the region of the encoded object into GOS and SPC, the processing order is not limited to the above. For example, a GOS can be encoded after its structure is determined, and then the order of GOS structure can be determined afterward.
[0160] In this way, the 3D data encoding device 100 generates encoded 3D data 112 by encoding the 3D data 111. Specifically, the 3D data encoding device 100 divides the 3D data into random access units, that is, into first processing units (GOS) corresponding to 3D coordinates, then divides the first processing units (GOS) into multiple second processing units (SPC), and then divides the second processing units (SPC) into multiple third processing units (VLM). Furthermore, each third processing unit (VLM) includes one or more voxels (VXL), where a voxel (VXL) is the smallest unit corresponding to position information.
[0161] Next, the 3D data encoding device 100 generates encoded 3D data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the 3D data encoding device 100 encodes each of the plurality of second processing units (SPC) in each of the first processing units (GOS). Furthermore, the 3D data encoding device 100 encodes each of the plurality of third processing units (VLM) in each of the second processing units (SPC).
[0162] For example, when the first processing unit (GOS) of the object being processed is a closed GOS, the 3D data encoding apparatus 100 encodes the second processing unit (SPC) of the object being processed, which is contained within the first processing unit (GOS), by referring to other second processing units (SPCs) contained within the first processing unit (GOS). That is, the 3D data encoding apparatus 100 does not refer to the second processing units (SPCs) contained in a first processing unit (GOS) that is different from the first processing unit (GOS) of the object being processed.
[0163] Furthermore, when the first processing unit (GOS) of the processing object is an open GOS, the second processing unit (SPC) of the processing object contained in the first processing unit (GOS) of the processing object is encoded with reference to other second processing units (SPCs) contained in the first processing unit (GOS) of the processing object, or the second processing units (SPCs) contained in a first processing unit (GOS) different from the first processing unit (GOS) of the processing object.
[0164] Furthermore, the three-dimensional data encoding device 100 selects one of the following as the type of the second processing unit (SPC) of the processing object: a first type (I-SPC) that does not refer to other second processing units (SPCs), a second type (P-SPC) that refers to one other second processing unit (SPC), and a third type that refers to two other second processing units (SPCs), and encodes the second processing unit (SPC) of the processing object according to the selected type.
[0165] Next, the configuration and operation flow of the three-dimensional data decoding device involved in this embodiment will be described. Figure 8 This is a block diagram of the three-dimensional data decoding device 200 involved in this embodiment. Figure 9 This is a flowchart illustrating an example of the operation of the three-dimensional data decoding device 200.
[0166] Figure 8 The 3D data decoding apparatus 200 shown generates decoded 3D data 212 by decoding encoded 3D data 211. Here, encoded 3D data 211 is, for example, encoded 3D data 112 generated by the 3D data encoding apparatus 100. The 3D data decoding apparatus 200 includes: an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.
[0167] First, the acquisition unit 201 acquires the encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS of the decoding object (S202). Specifically, the decoding start GOS determination unit 202 refers to the metadata stored in or separately from the encoded three-dimensional data 211, and determines the GOS of the decoding object, including the spatial position of the start of decoding, the object, or the SPC corresponding to the time.
[0168] Next, the SPC decoding decision unit 203 determines the type (I, P, B) of the SPC to be decoded within the GOS (S203). For example, the SPC decoding decision unit 203 determines (1) whether to decode only I-SPC, (2) whether to decode both I-SPC and P-SPC, and (3) whether to decode all types. Alternatively, if the type of SPC to be decoded is predetermined, such as when decoding all SPCs, this step may be omitted.
[0169] Next, the decoding unit 204 obtains the address position of the SPC that starts in the decoding order (same as the encoding order) within the GOS, starting in the encoded three-dimensional data 211, obtains the encoded data of the starting SPC from that address position, and decodes each SPC sequentially from that starting SPC (S204). Furthermore, the aforementioned address position is stored in metadata, etc.
[0170] Thus, the 3D data decoding device 200 decodes the decoded 3D data 212. Specifically, the 3D data decoding device 200 generates decoded 3D data 212, which serves as a random access unit, by decoding each of the encoded 3D data 211 of a first processing unit (GOS) corresponding to 3D coordinates. More specifically, the 3D data decoding device 200 decodes each of a plurality of second processing units (SPCs) in each first processing unit (GOS). Furthermore, the 3D data decoding device 200 decodes each of a plurality of third processing units (VLMs) in each second processing unit (SPC).
[0171] The metadata used for random access is described below. This metadata is generated by the three-dimensional data encoding device 100 and is contained in the encoded three-dimensional data 112 (211).
[0172] In conventional two-dimensional random access to moving images, decoding begins with the first frame of a random access unit near a specified time. However, in world space, random access based on coordinates or objects is envisioned in addition to time.
[0173] Therefore, in order to achieve at least random access to the three elements of coordinates, object, and time, a table was prepared that corresponds each element to the index number of the GOS. Furthermore, a correspondence was established between the index number of the GOS and the address of the I-SPC that becomes the beginning of the GOS. Figure 10 An example of a table included in the metadata is shown. Additionally, there is no need to use... Figure 10 Of all the tables shown, at least one table is required.
[0174] The following example illustrates random access starting from coordinates. When accessing coordinates (x2, y2, z2), the coordinate-GOS table is consulted first. It is known that the location with coordinates (x2, y2, z2) is included in the second GOS. Next, referring to the GOS address table, it is known that the address of the I-SPC at the beginning of the second GOS is addr(2). Therefore, the decoding unit 204 obtains the data from this address and begins decoding.
[0175] Furthermore, the address can be a logical address or a physical address of the HDD or memory. Also, information defining file segments can be used instead of addresses. For example, a file segment is a unit of data after dividing one or more GOS (Gateway Operating System) units.
[0176] Furthermore, when an object spans multiple GOS, the GOS to which multiple objects belong can be displayed in the object GOS table. If these multiple GOS are closed GOS, the encoding and decoding devices can perform encoding or decoding in parallel. Additionally, if these multiple GOS are open GOS, the compression efficiency can be further improved by referencing each other.
[0177] Examples of objects include people, animals, cars, bicycles, traffic lights, or buildings that serve as landmarks on land. For example, when encoding in world space, the 3D data encoding device 100 extracts feature points unique to the object from 3D point cloud data, detects the object based on these feature points, and can set the detected object as a random access point.
[0178] Thus, the three-dimensional data encoding device 100 generates first information showing a plurality of first processing units (GOS) and three-dimensional coordinates corresponding to each of the plurality of first processing units (GOS). Furthermore, the encoded three-dimensional data 112 (211) includes this first information. The first information further shows at least one of the object, time, and data storage destination corresponding to each of the plurality of first processing units (GOS).
[0179] The three-dimensional data decoding device 200 obtains first information from the encoded three-dimensional data 211, uses the first information to determine the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object or time, and decodes the encoded three-dimensional data 211.
[0180] Examples of other metadata are described below. In addition to metadata for random access, the 3D data encoding device 100 can also generate and store the following metadata. Furthermore, the 3D data decoding device 200 can utilize this metadata during decoding.
[0181] When using 3D data as map information, profiles are defined according to their purpose, and the information for that profile can be included in the metadata. For example, profiles may be defined for urban or suburban areas, or for flying objects, and the maximum or minimum size of world space, SPC, or VLM may be defined respectively. For example, in an urban-oriented profile, more detailed information is needed than in a suburban area, so the minimum size of the VLM is set to be smaller.
[0182] Meta-information can also include label values indicating the type of object. These label values correspond to the VLM, SPC, or GOS that constitute the object. Label values can be set according to the type of object, for example, label value "0" represents "person," label value "1" represents "car," and label value "2" represents "traffic light." Alternatively, when the type of object is difficult to determine or does not need to be determined, label values representing properties such as size, or whether the object is dynamic or static, can be used.
[0183] Furthermore, metadata can also include information showing the extent of the spatial region occupied by the world space.
[0184] Furthermore, metadata can also be used as header information shared by the entire stream of encoded data or multiple SPCs such as SPCs within GOS to store the size of an SPC or VXL.
[0185] Furthermore, metadata may also include identification information such as distance sensors or cameras used in the generation of point cloud data, or information showing the positional accuracy of point groups within the point cloud data.
[0186] Furthermore, meta-information can include information indicating whether the world space consists only of static objects or contains dynamic objects.
[0187] The following describes variations of this embodiment.
[0188] The encoding or decoding device can encode or decode two or more SPCs or GOSs that are different from each other in parallel. The GOSs that are encoded or decoded in parallel can be determined based on metadata indicating the spatial location of the GOSs.
[0189] In cases where three-dimensional data is used as a spatial map of moving vehicles or flying objects, or in the generation of such spatial maps, the encoding or decoding device can encode or decode the GOS or SPC contained in the space determined based on GPS, path information, or zoom level.
[0190] Furthermore, the decoding device can also start decoding sequentially from the spaces closest to its own position or path. The encoding or decoding device can also prioritize spaces farther from its own position or path over closer spaces when encoding or decoding. Here, lowering priority means reducing the processing order, reducing resolution (post-processing), or reducing image quality (to improve encoding efficiency, such as by increasing the quantization step size).
[0191] Furthermore, when decoding encoded data that is hierarchically encoded in space, the decoding device can also decode only the lower levels.
[0192] Furthermore, the decoding device can also start decoding from the lower levels, depending on the map's zoom level or purpose.
[0193] Furthermore, in applications such as self-position estimation or object recognition during the autonomous movement of cars or robots, the encoding or decoding device can reduce the resolution of the area outside the area within a specified height of the road surface (the area to be identified) for encoding or decoding.
[0194] Furthermore, the encoding device can also encode point cloud data representing indoor and outdoor spatial shapes independently. For example, by separating the GOS representing the interior (indoor GOS) from the GOS representing the exterior (outdoor GOS), the decoding device can select the GOS to be decoded according to the viewpoint position when using the encoded data.
[0195] Furthermore, the encoding device can encode indoor and outdoor GOS with close coordinates adjacently in the encoding stream. For example, the encoding device maps the identifiers of the two together and stores information showing that the corresponding identifiers have been established in the encoding stream or in separately stored metadata. Accordingly, the decoding device can identify indoor and outdoor GOS with close coordinates by referring to the information in the metadata.
[0196] Furthermore, the encoding device can switch the size of GOS or SPC between indoor and outdoor GOS. For example, the encoding device can set a smaller GOS size indoors compared to outdoors. Additionally, the encoding device can also change the accuracy of feature point extraction from point cloud data or the accuracy of object detection between indoor and outdoor GOS.
[0197] Furthermore, the encoding device can append information used by the decoding device to distinguish between dynamic and static objects to the encoded data. Accordingly, the decoding device can combine dynamic objects with red boxes or explanatory text to represent them. Alternatively, the decoding device can replace dynamic objects with only red boxes or explanatory text. Moreover, the decoding device can represent more detailed object categories. For example, a car can be represented with a red box, and a person with a yellow box.
[0198] Furthermore, the encoding or decoding device can determine whether to encode or decode by classifying dynamic objects and static objects as different SPCs or GOSs based on factors such as the frequency of occurrence of dynamic objects or the ratio of static objects to dynamic objects. For example, if the frequency or ratio of dynamic objects exceeds a threshold, an SPC or GOS in which dynamic and static objects are mixed is allowed; if the frequency or ratio of dynamic objects does not exceed the threshold, an SPC or GOS in which dynamic and static objects are mixed is not allowed.
[0199] When a dynamic object is detected not from point cloud data but from two-dimensional image information from a camera, the encoding device can obtain the information (boxes or text, etc.) used to identify the detection result and the object's position separately, and encode this information as part of the three-dimensional encoded data. In this case, the decoding device overlays the auxiliary information (boxes or text) representing the dynamic object onto the decoding result of the static object.
[0200] Furthermore, the encoding device can adjust the density of VXL or VLM according to the complexity of the static object's shape. For example, the more complex the shape of the static object, the denser the VXL or VLM will be. Moreover, the encoding device can determine the quantization step size when quantizing spatial location or color information based on the density of VXL or VLM. For example, the denser the VXL or VLM, the smaller the quantization step size will be.
[0201] As shown above, the encoding or decoding device involved in this embodiment uses spatial units with coordinate information to encode or decode space.
[0202] Furthermore, the encoding and decoding devices perform encoding or decoding in space using volume units. Volume includes the smallest unit corresponding to positional information, namely a voxel.
[0203] Furthermore, the encoding and decoding devices establish correspondences between arbitrary elements by creating tables that correspond to each element, including spatial information such as coordinates, objects, and time, with the Group of Pictures (GOP), or tables that correspond between elements. The decoding device uses the value of the selected element to determine the coordinates, and determines the volume, voxel, or space based on the coordinates, then decodes the space including that volume or voxel, or the determined space.
[0204] Furthermore, the encoding device determines the volume, voxel, or space that can be selected by the elements through feature point extraction or object recognition, and encodes it as a volume, voxel, or space that can be randomly accessed.
[0205] The space is divided into three types: I-SPC, which can be encoded or decoded by a single space unit; P-SPC, which is encoded or decoded by referring to any one processed space; and B-SPC, which is encoded or decoded by referring to any two processed spaces.
[0206] More than one volume corresponds to either a static object or a dynamic object. The space containing static objects and the space containing dynamic objects are encoded or decoded as different GOS. That is, the SPC containing static objects and the SPC containing dynamic objects are assigned to different GOS.
[0207] Dynamic objects are encoded or decoded individually, corresponding to one or more spaces containing only static objects. That is, multiple dynamic objects are encoded separately, and the resulting encoded data of multiple dynamic objects corresponds to an SPC containing only static objects.
[0208] The encoding and decoding devices prioritize I-SPCs within the GOS for encoding or decoding. For example, the encoding device encodes in a way that reduces I-SPC degradation (so that the original 3D data can be reproduced more faithfully after decoding). The decoding device, for example, decodes only the I-SPCs.
[0209] The encoding device can adjust the frequency of I-SPC utilization based on the density or quantity of objects in world space. In other words, the encoding device changes the frequency of I-SPC selection according to the number or density of objects contained in the 3D data. For example, the higher the density of objects in world space, the more frequently the encoding device uses I-space.
[0210] Furthermore, the encoding device sets random access points in units of GOS and stores information showing the spatial region corresponding to the GOS in the header information.
[0211] The encoding device may use a default value as the size of the GOS. Alternatively, the encoding device may change the size of the GOS according to the number or density of objects or dynamic objects. For example, the encoding device will set the size of the GOS smaller when the objects or dynamic objects are denser or more numerous.
[0212] Furthermore, the space or volume includes a group of feature points derived from information obtained using sensors such as depth sensors, gyroscopes, or cameras. The coordinates of the feature points are set as the center position of the voxels. Moreover, through voxel subdivision, high precision of positional information can be achieved.
[0213] Feature point groups are derived using multiple images. The multiple images have at least two types of temporal information: the actual temporal information and the spatially corresponding temporal information of the same time in the multiple images (e.g., the encoded temporal information used for rate control, etc.).
[0214] Furthermore, encoding or decoding is performed in units of GOS that include more than one space.
[0215] The encoding and decoding devices, with reference to the space within the processed GOS, predict the P space or B space within the GOS of the object being processed.
[0216] Alternatively, the encoding and decoding devices do not refer to different GOS, but use the processed space within the GOS of the object being processed to predict the P space or B space within the GOS of the object being processed.
[0217] Furthermore, the encoding and decoding devices transmit or receive encoded streams in units of world space comprising one or more GOS.
[0218] Furthermore, the GOS has a layered structure in at least one direction within world space, and the encoding and decoding devices encode or decode starting from the lower layer. For example, a GOS capable of random access belongs to the lowest layer. A GOS belonging to a higher layer only refers to GOS belonging to layers below the same layer. That is, the GOS is spatially divided in a predefined direction, including multiple layers, each with more than one SPC. The encoding and decoding devices encode or decode for each SPC by referring to SPCs contained in layers that are in the same layer as or lower than that SPC.
[0219] Furthermore, the encoding and decoding devices continuously encode or decode GOS within a world space unit comprising multiple GOS. The encoding and decoding devices write or read information indicating the order (direction) of encoding or decoding as metadata. That is, the encoded data includes information indicating the encoding order of multiple GOS.
[0220] Furthermore, the encoding and decoding devices encode or decode two or more different spaces or GOS in parallel.
[0221] Furthermore, the encoding and decoding devices encode or decode the spatial information (coordinates, size, etc.) of the space or GOS.
[0222] Furthermore, the encoding and decoding devices encode or decode the space or GOS contained in a specific space determined based on external information related to their own location and / or area size, such as GPS, path information, or magnification.
[0223] Encoding or decoding devices prioritize spaces farther away from themselves over spaces closer to themselves when performing encoding or decoding.
[0224] The encoding device sets a direction in world space according to a magnification or purpose, and encodes GOS with a layered structure in that direction. The decoding device, for a GOS with a layered structure in one direction of world space set according to a magnification or purpose, preferentially decodes it starting from the lower layer.
[0225] The encoding device alters the accuracy of feature point extraction, object recognition, and spatial area size contained in indoor and outdoor spaces. However, the encoding and decoding devices encode or decode indoor and outdoor GOS that are close in coordinates, placing them adjacent in world space, and also map these identifiers together for encoding or decoding.
[0226] (Implementation Method 2)
[0227] When using encoded point cloud data for practical devices or services, it is desirable to transmit and receive the required information according to its intended purpose in order to conserve network bandwidth. However, current encoding structures for 3D data do not possess this functionality, and therefore, there is no corresponding encoding method.
[0228] This embodiment will describe a three-dimensional data encoding method and a three-dimensional data encoding apparatus for providing the function of sending and receiving required information according to the purpose in the encoded data of three-dimensional point cloud data, as well as a three-dimensional data decoding method and a three-dimensional data decoding apparatus for decoding the encoded data.
[0229] A voxel with a certain number of characteristics (VXL) is defined as a characteristic voxel (FVXL), and the world space (WLD) composed of FVXL is defined as a sparse world space (SWLD). Figure 11 This illustrates a sparse world space and examples of its composition. SWLD includes: FGOS, a GOS constructed from FVXL; FSPC, an SPC constructed from FVXL; and FVLM, a VLM constructed from FVXL. The data structures and prediction structures of FGOS, FSPC, and FVLM can be the same as those of GOS, SPC, and VLM.
[0230] Feature quantities refer to the characteristic quantities that represent the three-dimensional position information of VXL, or the visible light information of VXL position, especially the corners and edges of three-dimensional objects where more feature quantities can be detected. Specifically, although the feature quantity is the three-dimensional feature quantity or visible light feature quantity described below, it can be any feature quantity that represents the position, brightness, or color information of VXL.
[0231] As three-dimensional feature quantities, SHOT (Signature of Histograms of Orientations), PFH (Point Feature Histograms), or PPF (Point Pair Feature) features are used.
[0232] The SHOT feature is obtained by dividing the area around the VXL region, calculating the inner product of the reference point and the normal vector of the divided region, and then performing histogram generation. This SHOT feature is characterized by high dimensionality and high feature expressiveness.
[0233] The PFH feature is obtained by selecting multiple pairs of points near VXL, calculating the normal vector from these points, and then performing histogram transformation. Because it is a histogram feature, the PFH feature is robust against a small amount of interference and exhibits high feature expressiveness.
[0234] The PPF feature is a feature calculated using the VXL of a 2-point vector and the normal vector, etc. Because all VXLs are used, the PPF feature is robust for occlusion.
[0235] Furthermore, as a feature quantity of visible light, it is possible to use SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients), which incorporate information such as the brightness gradient of the image.
[0236] SWLD is generated by calculating the aforementioned feature quantities from each VXL of WLD and extracting FVXL. Here, SWLD can be updated every time WLD is updated, or it can be updated periodically after a certain period of time, regardless of the update timing of WLD.
[0237] SWLDs can be generated for each feature. For example, SWLD1 based on SHOT features and SWLD2 based on SIFT features can be generated separately for each feature, and the SWLDs can be used according to their intended purpose. Furthermore, the calculated features of each FVXL can also be stored as feature information in each FVXL.
[0238] Next, the utilization method of Sparse World Space (SWLD) will be explained. Since SWLD only contains feature voxels (FVXL), its data size is generally smaller compared to WLD, which includes all VXLs.
[0239] In applications that utilize feature quantities to achieve a certain purpose, using SWLD information instead of WLD can suppress hard drive read time and reduce network transmission bandwidth and transmission time. For example, as map information, by pre-storing WLD and SWLD on the server and switching the sent map information to WLD or SWLD according to the client's request, network bandwidth and transmission time can be reduced. Specific examples are shown below.
[0240] Figure 12 as well as Figure 13 Examples of the use of SWLD and WLD are shown. For example... Figure 12 As shown, when client 1, acting as a vehicle-mounted device, needs map information for determining its own location, client 1 sends a request to the server for map data for location estimation (S301). The server sends the SWLD (Site Layout Data) to client 1 according to the request (S302). Client 1 uses the received SWLD to determine its own location (S303). At this time, client 1 uses various methods such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras to acquire VXL (Very Large Scale) information of its surroundings, and estimates its own location information based on the obtained VXL information and SWLD. Here, the location information includes client 1's three-dimensional location information and orientation, etc.
[0241] like Figure 13 As shown, when client 2, acting as a vehicle-mounted device, needs map information for map drawing purposes such as 3D maps, client 2 sends a request to the server for map data acquisition (S311). The server, in accordance with this request, sends a WLD (Wide Image Data) to client 2 (S312). Client 2 uses the received WLD to perform map drawing (S313). At this time, client 2, for example, uses images captured by itself with a visible light camera, and the WLD acquired from the server, to create a conceptual image, which is then displayed on a screen such as a car navigation system.
[0242] As shown above, the server sends the SWLD to the client when it primarily needs VXL feature quantities such as its own location estimation, similar to map drawing. When more detailed VXL information is required, the server sends the WLD to the client. This enables efficient sending and receiving of map data.
[0243] Additionally, the client can determine whether it needs a SWLD or a WLD and request the server to send either one. Furthermore, the server can determine which SWLD or WLD to send based on the client's or network conditions.
[0244] Next, we will explain the method for switching between sending and receiving Sparse World Space (SWLD) and World Space (WLD).
[0245] The reception of WLD or SWLD can be switched according to the network bandwidth. Figure 14 An example of this operation is shown. For instance, when a low-speed network with sufficient bandwidth, such as an LTE (Long Term Evolution) network, is used, the client accesses the server via the low-speed network (S321) and obtains the SWLD (Wide Layout Map) as map information from the server (S322). Conversely, when a high-speed network with ample bandwidth, such as a WiFi network, is used, the client accesses the server via the high-speed network (S323) and obtains the SWLD from the server (S324). Accordingly, the client can obtain appropriate map information based on its network bandwidth.
[0246] Specifically, the client receives SWLD via LTE outdoors, and obtains WLD via WiFi when entering indoor facilities. Based on this, the client can obtain more detailed indoor map information.
[0247] In this way, the client can request WLD or SWLD from the server according to the frequency band of its network. Alternatively, the client can send information indicating the frequency band of its network to the server, and the server can send the appropriate data (WLD or SWLD) to the client based on that information. Or, the server can determine the client's network bandwidth and send the appropriate data (WLD or SWLD) to the client.
[0248] Furthermore, the reception of WLD or SWLD can be switched according to the movement speed. Figure 15 An example of operation in this scenario is shown. For instance, when the client is moving at high speed (S331), the client receives a SWLD from the server (S332). Conversely, when the client is moving at low speed (S333), the client receives a WLD from the server (S334). Accordingly, the client can both conserve network bandwidth and obtain map information according to speed. Specifically, when the client is traveling on a highway, by receiving a small amount of SWLD, it can update the map information at a roughly appropriate speed. Furthermore, when the client is traveling on a regular road, by receiving a WLD, it can obtain more detailed map information.
[0249] In this way, the client can request WLD or SWLD from the server based on its own movement speed. Alternatively, the client can send information indicating its own movement speed to the server, and the server can send the appropriate data (WLD or SWLD) to the client based on that information. Or, the server can determine the client's movement speed and send the appropriate data (WLD or SWLD) to the client.
[0250] Alternatively, the client can first obtain the SWLD from the server, and then obtain the WLD for the more important areas within it. For example, when acquiring map data, the client can first obtain a general map information from the SWLD, filter out areas with a high frequency of features such as buildings, signs, or people, and then obtain the WLD for the filtered areas. In this way, the client can both control the amount of data received from the server and obtain the detailed information for the required areas.
[0251] Furthermore, the server can create separate SWLDs for each object based on the WLD, and the client can receive them according to their intended use. This can reduce network bandwidth usage. For example, the server can pre-identify people or vehicles from the WLD and create separate SWLDs for people and vehicles. The client receives the person SWLD when it wants to obtain information about people around it, and the vehicle SWLD when it wants to obtain information about vehicles. Moreover, the types of SWLDs can be distinguished based on information attached to the header (such as logos or types).
[0252] Next, the configuration and operation flow of the three-dimensional data encoding device (e.g., a server) involved in this embodiment will be described. Figure 16 This is a block diagram of the three-dimensional data encoding device 400 involved in this embodiment. Figure 17 This is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device 400.
[0253] Figure 16 The illustrated 3D data encoding apparatus 400 generates encoded 3D data 413 and 414 as encoded streams by encoding input 3D data 411. Here, encoded 3D data 413 is encoded 3D data corresponding to WLD (Wide Area Decoding), and encoded 3D data 414 is encoded 3D data corresponding to SWLD (Simplified Swing Data Decoding). The 3D data encoding apparatus 400 includes: an acquisition unit 401, an encoding region determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.
[0254] like Figure 17 As shown, firstly, the acquisition unit 401 acquires input three-dimensional data 411 as point group data in three-dimensional space (S401).
[0255] Next, the encoding region determination unit 402 determines the spatial region of the encoding object based on the spatial region where the point group data exists (S402).
[0256] Next, the SWLD extraction unit 403 defines the spatial region of the encoded object as WLD and calculates the feature quantity based on each VXL contained in the WLD. Furthermore, the SWLD extraction unit 403 extracts VXLs whose feature quantity is above a predetermined threshold, defines the extracted VXLs as FVXLs, and appends these FVXLs to the SWLD to generate extracted three-dimensional data 412 (S403). That is, extracted three-dimensional data 412 with feature quantities above the threshold is extracted from the input three-dimensional data 411.
[0257] Next, the WLD encoding unit 404 generates encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 appends information used to distinguish whether the encoded three-dimensional data 413 is a stream containing the WLD to the header of the encoded three-dimensional data 413.
[0258] Furthermore, the SWLD encoding unit 405 generates encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 appends information used to distinguish whether the encoded three-dimensional data 414 is a stream containing the SWLD to the header of the encoded three-dimensional data 414.
[0259] Furthermore, the processing order of generating coded 3D data 413 and generating coded 3D data 414 can also be reversed as described above. Additionally, some or all of the above processes can be executed in parallel.
[0260] Information assigned to the headers of encoded 3D data 413 and 414 is, for example, defined as a parameter such as "world_type". When world_type = 0, it indicates that the stream contains WLD; when world_type = 1, it indicates that the stream contains SWLD. When defining other categories, the assigned value can be increased, such as world_type = 2. Furthermore, specific flags can be included in either encoded 3D data 413 or 414. For example, encoded 3D data 414 can be assigned a flag indicating that the stream contains SWLD. In this case, the decoding device can determine whether the stream contains WLD or SWLD based on the presence or absence of the flag.
[0261] Furthermore, the encoding method used by the WLD encoding unit 404 when encoding WLD can be different from the encoding method used by the SWLD encoding unit 405 when encoding SWLD.
[0262] For example, because SWLD data is sampled, its correlation with surrounding data may be lower compared to WLD. Therefore, in coding methods used for SWLD, inter-frame prediction is prioritized over intra-frame prediction and inter-frame prediction compared to coding methods used for WLD.
[0263] Furthermore, the representation of 3D position can differ between the encoding methods used for SWLD and WLD. For example, in FWLD, the 3D position of FVXL can be represented by 3D coordinates, while in WLD, the 3D position can be represented by an octree (described later), or vice versa.
[0264] Furthermore, the SWLD encoding unit 405 encodes the SWLD encoded three-dimensional data 414 in a manner where the data size of the SWLD encoded three-dimensional data 413 is smaller than the data size of the WLD encoded three-dimensional data 413. For example, as described above, the correlation between data in SWLD may be reduced compared to WLD. Consequently, the encoding efficiency decreases, and the data size of the encoded three-dimensional data 414 may be larger than the data size of the WLD encoded three-dimensional data 413. Therefore, when the obtained encoded three-dimensional data 414 has a larger data size than the WLD encoded three-dimensional data 413, the SWLD encoding unit 405 re-encodes it to generate encoded three-dimensional data 414 with a reduced data size.
[0265] For example, the SWLD extraction unit 403 regenerates extracted 3D data 412 with a reduced number of extracted feature points, and the SWLD encoding unit 405 encodes the extracted 3D data 412. Alternatively, the quantization level in the SWLD encoding unit 405 can be made coarser. For example, in the octree structure described later, the quantization level can be made coarser by rounding the data at the lowest level.
[0266] Furthermore, if the SWLD encoding unit 405 cannot make the data size of the SWLD encoded three-dimensional data 414 smaller than the data size of the WLD encoded three-dimensional data 413, it may choose not to generate the SWLD encoded three-dimensional data 414. Alternatively, the WLD encoded three-dimensional data 413 may be copied to the SWLD encoded three-dimensional data 414. That is, the WLD encoded three-dimensional data 413 can be directly used as the SWLD encoded three-dimensional data 414.
[0267] Next, the configuration and operation flow of the three-dimensional data decoding device (e.g., client) involved in this embodiment will be described. Figure 18 This is a block diagram of the three-dimensional data decoding device 500 involved in this embodiment. Figure 19 This is a flowchart of the three-dimensional data decoding process performed by the three-dimensional data decoding device 500.
[0268] Figure 18 The illustrated 3D data decoding apparatus 500 generates decoded 3D data 512 or 513 by decoding encoded 3D data 511. Here, encoded 3D data 511 is, for example, encoded 3D data 413 or 414 generated by the 3D data encoding apparatus 400.
[0269] The 3D data decoding device 500 includes: an acquisition unit 501, a head parsing unit 502, a WLD decoding unit 503, and a SWLD decoding unit 504.
[0270] like Figure 19 As shown, firstly, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Next, the header parsing unit 502 parses the header of the encoded three-dimensional data 511 and determines whether the encoded three-dimensional data 511 is a stream containing WLD or a stream containing SWLD (S502). For example, the determination is made by referring to the world_type parameter mentioned above.
[0271] If the encoded 3D data 511 is a stream containing WLD (S503 "Yes"), the WLD decoding unit 503 decodes the encoded 3D data 511 to generate decoded 3D data 512 of WLD (S504). Alternatively, if the encoded 3D data 511 is a stream containing SWLD (S503 "No"), the SWLD decoding unit 504 decodes the encoded 3D data 511 to generate decoded 3D data 513 of SWLD (S505).
[0272] Furthermore, similar to the encoding apparatus, the decoding method used by the WLD decoding unit 503 when decoding the WLD can be different from the decoding method used by the SWLD decoding unit 504 when decoding the SWLD. For example, in the decoding method for the SWLD, compared with the decoding method for the WLD, inter-frame prediction in intra-frame prediction and inter-frame prediction can be prioritized.
[0273] Furthermore, the methods used to represent the three-dimensional position can differ between the decoding methods used for SWLD and WLD. For example, SWLD can represent the three-dimensional position of FVXL using three-dimensional coordinates, while WLD can represent the three-dimensional position using an octree (described later), and vice versa.
[0274] Next, the octree representation as a method of representing three-dimensional location will be explained. The VXL data contained in the three-dimensional data is converted into an octree structure and then encoded. Figure 20 An example of VXL for WLD is shown. Figure 21 It shows Figure 20 The octree structure of WLD is shown. Figure 20 In the example shown, there are three VXL1 to VXL3 that constitute the VXL (hereinafter, valid VXL) of the point group. Figure 21 As shown, the octree structure consists of nodes and leaves. Each node has a maximum of 8 nodes or leaves. Each leaf contains VXL information. Here, Figure 21 Among the leaves shown, leaves 1, 2, and 3 represent... Figure 20 VXL1, VXL2, and VXL3 are shown.
[0275] Specifically, each node and leaf corresponds to a 3D position. Node 1 and... Figure 20 All the blocks shown correspond to each other. The block corresponding to node 1 is divided into 8 blocks. Among these 8 blocks, the block with a valid VXL is set as a node, and the other blocks are set as leaves. The block corresponding to a node is further divided into 8 nodes or leaves, and this process is repeated the same number of times as the number of levels in the tree structure. Furthermore, all the blocks at the bottom level are set as leaves.
[0276] and, Figure 22 It shows from Figure 20 The example shown is a SWLD generated from a WLD. Figure 20 The feature extraction results of VXL1 and VXL2 shown are identified as FVXL1 and FVXL2 and added to SWLD. VXL3, however, is not identified as FVXL and therefore is not included in SWLD. Figure 23 It shows Figure 22 The octree structure of SWLD is shown. Figure 23 In the octree structure shown, Figure 21 The leaf 3 shown, equivalent to VXL3, has been deleted. Accordingly, Figure 21 Node 3 shown does not have a valid VXL and has been changed to a leaf. Thus, generally speaking, SWLD has fewer leaves than WLD, and the encoded 3D data of SWLD is also smaller than that of WLD.
[0277] The following describes variations of this embodiment.
[0278] For example, when a vehicle-mounted device or other client is estimating its own position, it receives a SWLD from the server, uses the SWLD to estimate its own position, and performs obstacle detection. Then, it uses various methods such as distance sensors such as rangefinders, stereo cameras, or combinations of multiple monocular cameras to perform obstacle detection based on the three-dimensional information of the surrounding environment it has obtained.
[0279] Furthermore, it is generally difficult to include VXL data for flat areas in a SWLD. Therefore, the server maintains a downsampled world space (subWLD) that is downsampled from the WLD for detecting stationary obstacles, and can send both the SWLD and subWLD to the client. This allows for both network bandwidth management and client-side position estimation and obstacle detection.
[0280] Furthermore, a grid-structured map is advantageous when the client rapidly draws 3D map data. Therefore, the server can generate a grid based on the World Layout (WLD) and maintain it beforehand as a Grid World Space (MWLD). For example, the client receives the MWLD when a rough 3D drawing is needed, and the WLD when a detailed 3D drawing is required. This helps to control network bandwidth usage.
[0281] Furthermore, although the server sets VXLs with feature values above a threshold as FVXLs from each VXL, FVXLs can also be calculated using different methods. For example, if the server determines that VXLs, VLMs, SPCs, or GOSs constituting signals or intersections are needed in self-position estimation, driving assistance, or autonomous driving, they can be included in the SWLD as FVXLs, FVLMs, FSPCs, or FGOSs. Moreover, the above determination can be performed manually. In addition, FVXLs obtained by the above method can be added to FVXLs, etc., set based on feature values. That is, the SWLD extraction unit 403 can further extract data corresponding to objects with predefined attributes from the input 3D data 411 as extracted 3D data 412.
[0282] Furthermore, different labels can be assigned to the feature values depending on the situation requiring these uses. The server can maintain the FVXL required for self-position estimation of signals or intersections, driver assistance, or autonomous driving as a higher layer of SWLD (e.g., lane world space).
[0283] Furthermore, the server can also append attributes to the VXL within the WLD according to random access units or specified units. Attributes may include, for example, information indicating whether they are needed or not in the self-location estimation, or information indicating whether they are important as traffic information such as signals or intersections. Additionally, attributes may include the correspondence between lane information (GDF: Geographic Data Files, etc.) and features (intersections or roads, etc.).
[0284] Furthermore, the following methods can be used as an update method for WLD or SWLD.
[0285] Updated information such as changes in people, construction, or street trees (track-oriented) is loaded onto the server as point clusters or metadata. The server updates the WLD based on this loading, and then uses the updated WLD to update the SWLD.
[0286] Furthermore, if the client detects a mismatch between the 3D information it generates when estimating its own position and the 3D information received from the server, it can send its generated 3D information along with an update notification to the server. In this case, the server uses the WLD to update the SWLD. If the SWLD has not been updated, the server determines that the WLD itself is outdated.
[0287] Furthermore, as header information of the encoded stream, information for distinguishing between WLD and SWLD is appended. For example, in cases where multiple world spaces exist, such as grid world space or lane world space, information for differentiating them can be appended to the header information. Also, when multiple SWLDs with different feature values exist, information for distinguishing them individually can be appended to the header information.
[0288] Furthermore, although the SWLD is composed of FVXLs, it can also include VXLs that are not identified as FVXLs. For example, the SWLD can include adjacent VXLs used when calculating the feature values of the FVXLs. Accordingly, even if each FVXL in the SWLD does not have additional feature value information, the client can calculate the feature values of the FVXLs when receiving the SWLD. In addition, in this case, the SWLD can include information for distinguishing whether each VXL is an FVXL or a VXL.
[0289] As described above, the three-dimensional data encoding device 400 extracts three-dimensional data 412 (second three-dimensional data) from the input three-dimensional data 411 (first three-dimensional data) with a feature quantity of more than a threshold, and generates encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.
[0290] Accordingly, the 3D data encoding device 400 generates encoded 3D data 414 by encoding data whose feature values are above a threshold. This reduces the amount of data compared to directly encoding the input 3D data 411. Therefore, the 3D data encoding device 400 can reduce the amount of data transmitted.
[0291] Furthermore, the three-dimensional data encoding device 400 generates encoded three-dimensional data 413 (second encoded three-dimensional data) by encoding the input three-dimensional data 411.
[0292] Accordingly, the three-dimensional data encoding device 400 can selectively transmit encoded three-dimensional data 413 and encoded three-dimensional data 414, for example, according to its intended use.
[0293] Furthermore, the extracted 3D data 412 is encoded by the first encoding method, and the input 3D data 411 is encoded by the second encoding method, which is different from the first encoding method.
[0294] Accordingly, the three-dimensional data encoding device 400 can employ appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.
[0295] Furthermore, in the first coding method, compared to the second coding method, inter-frame prediction is prioritized in both intra-frame prediction and inter-frame prediction.
[0296] Accordingly, the 3D data encoding device 400 can extract 3D data 412 for adjacent data where the correlation between them is likely to decrease, thereby increasing the priority of inter-frame prediction.
[0297] Furthermore, the representation of 3D position differs between the first and second encoding methods. For example, the second encoding method uses an octree to represent 3D position, while the first encoding method uses 3D coordinates.
[0298] Accordingly, the three-dimensional data encoding device 400 can adopt a more appropriate three-dimensional position representation method for three-dimensional data with different data numbers (number of VXL or FVXL).
[0299] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the input three-dimensional data 411 or by encoding a portion of the input three-dimensional data 411. That is, the identifier indicates whether the encoded three-dimensional data is WLD encoded three-dimensional data 413 or SWLD encoded three-dimensional data 414.
[0300] Accordingly, the decoding device can easily determine whether the acquired encoded three-dimensional data is encoded three-dimensional data 413 or encoded three-dimensional data 414.
[0301] Furthermore, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 in such a way that the amount of data encoded in the three-dimensional data 414 is less than the amount of data encoded in the three-dimensional data 413.
[0302] Accordingly, the three-dimensional data encoding device 400 can encode three-dimensional data 414 in a smaller amount than the amount of data encodes three-dimensional data 413.
[0303] Furthermore, the 3D data encoding device 400 further extracts data corresponding to objects with predefined attributes from the input 3D data 411 as extracted 3D data 412. For example, objects with predefined attributes refer to objects needed in self-position estimation, driving assistance, or autonomous driving, such as signals or intersections.
[0304] Accordingly, the three-dimensional data encoding device 400 is able to generate encoded three-dimensional data 414, including the data required by the decoding device.
[0305] Furthermore, the three-dimensional data encoding device 400 (server) further sends one of the encoded three-dimensional data 413 and 414 to the client according to the client's state.
[0306] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the client's status.
[0307] Furthermore, the client's status includes the client's communication status (such as network bandwidth) or the client's movement speed.
[0308] Furthermore, the three-dimensional data encoding device 400 further sends one of the encoded three-dimensional data 413 and 414 to the client according to the client's request.
[0309] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the client's request.
[0310] Furthermore, the three-dimensional data decoding device 500 according to this embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0311] That is, the 3D data decoding device 500 decodes the encoded 3D data 414 obtained by encoding the extracted 3D data 412, in which the feature quantity extracted from the input 3D data 411 is above a threshold, using the first decoding method. Furthermore, the 3D data decoding device 500 decodes the encoded 3D data 413 obtained by encoding the input 3D data 411 using a second decoding method different from the first decoding method.
[0312] Accordingly, the 3D data decoding device 500 can selectively receive encoded 3D data 414 and encoded 3D data 413 obtained by encoding data with feature values above a threshold, for example, according to their intended use. This reduces the amount of data transmitted. Furthermore, the 3D data decoding device 500 can employ appropriate decoding methods for the input 3D data 411 and the extracted 3D data 412 respectively.
[0313] Furthermore, in the first decoding method, compared to the second decoding method, intra-frame prediction and inter-frame prediction are given priority.
[0314] Accordingly, the 3D data decoding device 500 can improve the priority of inter-frame prediction by extracting 3D data where the correlation between adjacent data is easily reduced.
[0315] Furthermore, the methods used to represent 3D position differ between the first and second decoding methods. For example, the second decoding method uses an octree to represent 3D position, while the first decoding method uses 3D coordinates.
[0316] Accordingly, the 3D data decoding device 500 can employ more appropriate 3D position representation methods for 3D data with different data numbers (number of VXL or FVXL).
[0317] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the input three-dimensional data 411 or by encoding a portion of the input three-dimensional data 411. The three-dimensional data decoding device 500 refers to this identifier to identify the encoded three-dimensional data 413 and 414.
[0318] Accordingly, the three-dimensional data decoding device 500 can easily determine whether the obtained encoded three-dimensional data is encoded three-dimensional data 413 or encoded three-dimensional data 414.
[0319] Furthermore, the 3D data decoding device 500 notifies the server of the client's (3D data decoding device 500) status. The 3D data decoding device 500 receives encoded 3D data 413 and 414 sent from the server according to the client's status.
[0320] Accordingly, the 3D data decoding device 500 can receive appropriate data according to the client's status.
[0321] Furthermore, the client's status includes the client's communication status (such as network bandwidth) or the client's movement speed.
[0322] Furthermore, the 3D data decoding device 500 further requests the server to encode one of the 3D data 413 and 414, and receives the encoded 3D data 413 and 414 sent from the server in accordance with the request.
[0323] Accordingly, the 3D data decoding device 500 is able to receive appropriate data corresponding to its purpose.
[0324] Although the three-dimensional data encoding apparatus and the three-dimensional data decoding apparatus involved in the embodiments of this application have been described above, this application is not limited to these embodiments.
[0325] Furthermore, the processing units included in the three-dimensional data encoding or decoding apparatus according to the above embodiments are typically implemented as LSIs of integrated circuits. These can be fabricated as a single chip, or some or all of them can be fabricated as a single chip.
[0326] Furthermore, integrated circuitry is not limited to LSIs; it can also be achieved using dedicated circuits or general-purpose processors. After LSI manufacturing, programmable FPGAs (Field Programmable Gate Arrays) or reconfigurable processors capable of reconfiguring the connections or settings of the internal circuitry of the LSI can be utilized.
[0327] Furthermore, in the various embodiments described above, each component can be constructed using dedicated hardware or implemented by executing software programs suitable for each component. Each component is implemented by a program execution unit such as a CPU or processor reading and executing software programs recorded on a recording medium such as a hard disk or semiconductor memory.
[0328] Furthermore, this application can be implemented as a three-dimensional data encoding method or a three-dimensional data decoding method executed by a three-dimensional data encoding device or a three-dimensional data decoding device.
[0329] Furthermore, taking the division of functional blocks in the block diagram as an example, multiple functional blocks can be implemented as a single functional block, or a single functional block can be divided into multiple functional blocks, or a portion of the functionality can be transferred to other functional blocks. Moreover, the functions of multiple functional blocks with similar capabilities can be processed in parallel by a single hardware or software component or through time-division processing.
[0330] Furthermore, the order in which the steps in the flowchart are executed is an example shown for the purpose of illustrating this application, and the order may be different from the one described above. Also, some of the steps described above may be executed simultaneously (in parallel) with other steps.
[0331] The above description addresses one or more embodiments of the three-dimensional data encoding and decoding apparatus, based on specific implementations. This application is not limited to these embodiments. Without departing from the spirit of this application, various modifications conceivable to those skilled in the art, implemented in this embodiment, and combinations of constituent elements from different embodiments are all included within the scope of the one or more embodiments.
[0332] Industrial availability
[0333] This application is applicable to both three-dimensional data encoding devices and three-dimensional data decoding devices.
[0334] Symbol Explanation
[0335] 100, 400 three-dimensional data encoding device
[0336] 101, 201, 401, and 501 received departmental approval.
[0337] 102, 402 coding area determination department
[0338] Division 103
[0339] 104 Coding Department
[0340] 111 Three-dimensional data
[0341] 112, 211, 413, 414, 511 Encode 3D Data
[0342] 200 and 500 3D data decoding devices
[0343] 202 Decoding Begins: GOS Decision Department
[0344] 203 Decoding SPC Decision Department
[0345] 204 Decoding Department
[0346] Decoding 3D Data (212, 512, 513)
[0347] 403SWLD Extraction Section
[0348] 404WLD Encoding Department
[0349] 405SWLD Encoding Section
[0350] 411 Input 3D data
[0351] 412 Extracting 3D Data
[0352] 502 Head Analysis Section
[0353] 503WLD Decoding Section
[0354] 504SWLD Decoding Section
Claims
1. A three-dimensional data encoding method, The process includes a sending step, in which one of the first encoded 3D data and the second encoded 3D data is sent to the client based on the client's movement speed. The first encoded three-dimensional data is generated by extracting second three-dimensional data with feature values above a threshold from the first three-dimensional data and then encoding the second three-dimensional data. The second encoded three-dimensional data is generated by encoding the first three-dimensional data. The feature quantity is a feature quantity based on three-dimensional position information or visible light information.
2. The three-dimensional data encoding method as described in claim 1, In the sending step, If the client's movement speed is above a predetermined threshold, the first encoded 3D data is sent to the client. If the client's movement speed is less than the threshold, the second encoded 3D data is sent to the client.
3. The three-dimensional data encoding method as described in claim 1 or 2, The second three-dimensional data is encoded by the first encoding method. The first three-dimensional data is encoded by a second encoding method, which is different from the first encoding method.
4. The three-dimensional data encoding method as described in claim 3, In the first coding method, compared with the second coding method, intra-frame prediction and inter-frame prediction are given priority.
5. The three-dimensional data encoding method as described in claim 3, The representation of three-dimensional position differs between the first encoding method and the second encoding method.
6. The three-dimensional data encoding method as described in claim 1 or 2, At least one of the first encoded three-dimensional data and the second encoded three-dimensional data includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the first three-dimensional data or by encoding a portion of the first three-dimensional data.
7. The three-dimensional data encoding method as described in claim 1 or 2, The second three-dimensional data is encoded in such a way that the amount of data in the first encoded three-dimensional data is smaller than the amount of data in the second encoded three-dimensional data.
8. A three-dimensional data decoding method, comprising: The receiving step involves receiving, based on the client's movement speed, one of the following sent from the server: (i) first encoded three-dimensional data obtained by encoding second three-dimensional data whose feature quantity extracted from the first three-dimensional data is above a threshold, and (ii) second encoded three-dimensional data obtained by encoding the first three-dimensional data. as well as The decoding step involves decoding one of the received first-encoded three-dimensional data and second-encoded three-dimensional data. The feature quantity is a feature quantity based on three-dimensional position information or visible light information.
9. The three-dimensional data decoding method as described in claim 8, In the receiving step, If the client's movement speed is above a predetermined threshold, the first encoded three-dimensional data is received; If the client's movement speed is less than the threshold, the second encoded three-dimensional data is received.
10. The three-dimensional data decoding method as described in claim 8 or 9, The first encoded three-dimensional data is decoded by the first decoding method. The second encoded three-dimensional data is decoded by a second decoding method that is different from the first decoding method.
11. The three-dimensional data decoding method as described in claim 10, In the first decoding method, compared with the second decoding method, intra-frame prediction and inter-frame prediction are given priority.
12. The three-dimensional data decoding method as described in claim 10, The representation of three-dimensional position differs between the first decoding method and the second decoding method.
13. The three-dimensional data decoding method as described in claim 8 or 9, At least one of the first encoded three-dimensional data and the second encoded three-dimensional data includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the first three-dimensional data or by encoding a portion of the first three-dimensional data. The first coded three-dimensional data and the second coded three-dimensional data are identified by referring to the identifier.
14. The three-dimensional data decoding method as described in claim 8 or 9, The three-dimensional data decoding method further includes: The notification step involves informing the server of the client's movement speed.
15. A three-dimensional data encoding device, comprising: The transmitting unit, based on the client's movement speed, transmits one of the first encoded 3D data and the second encoded 3D data to the client; and One of the first encoding unit and the second encoding unit, (i) the first encoding unit extracts second three-dimensional data with a feature quantity of more than a threshold from the first three-dimensional data, encodes the second three-dimensional data, and generates the first encoded three-dimensional data; (ii) the second encoding unit encodes the first three-dimensional data to generate the second encoded three-dimensional data. The feature quantity is a feature quantity based on three-dimensional position information or visible light information.
16. A three-dimensional data decoding device, comprising: The receiving unit, based on the client's movement speed, receives one of the following from the server: (i) first encoded three-dimensional data obtained by encoding second three-dimensional data whose feature quantity extracted from the first three-dimensional data is above a threshold, and (ii) second encoded three-dimensional data obtained by encoding the first three-dimensional data; and The decoding unit decodes one of the received first-encoded three-dimensional data and second-encoded three-dimensional data. The feature quantity is a feature quantity based on three-dimensional position information or visible light information.
Citation Information
Patent Citations
Simplified attributive analysis method for automatic coding of three-dimensional ship modeling part
CN101727521A
Wavelet transformation based unequal fault-tolerant protection method for three-dimensional data transmission
CN103001726A