Three-dimensional information processing method and three-dimensional information processing device

By dividing the three-dimensional data into random access units and encoding them, the problem of abnormal three-dimensional position information is solved, and the effects of random access and data volume reduction are achieved.

CN114627664BActive Publication Date: 2025-09-30PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202210453296.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-08-26
Filing Date
2017-08-23
Publication Date
2025-09-30
Estimated Expiration
2037-08-23

AI Technical Summary

Technical Problem

The existing technology cannot perform appropriate response when the three-dimensional position information is abnormal, and the three-dimensional coded data lacks random access function.

Method used

By dividing three-dimensional data into random access units corresponding to three-dimensional coordinates and adopting encoding and decoding methods, encoded data is generated to achieve random access function, and at the same time, the data source or mode is switched to respond in abnormal situations.

Benefits of technology

It achieves appropriate correspondence when the three-dimensional position information is abnormal, and provides a random access function in the encoded three-dimensional data, reducing the amount of data and improving the encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627664B_ABST
    Figure CN114627664B_ABST
Patent Text Reader

Abstract

A three-dimensional information processing method and a three-dimensional information processing device are provided. The three-dimensional information processing method includes the following processing: determining whether first three-dimensional position information within a first range can be obtained via a communication path; and if the first three-dimensional position information cannot be obtained via the communication path, obtaining second three-dimensional position information within a second range narrower than the first range via the communication path.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application was filed on August 23, 2017, with Chinese patent application number 201780051543.8 (international application number PCT / JP2017 / 030034), and is a divisional application of the patent application entitled “Three-dimensional information processing method and three-dimensional information processing device”. Technical Field

[0002] The present application relates to a three-dimensional information processing method and a three-dimensional information processing device. Background Art

[0003] Devices and services that utilize three-dimensional data will become increasingly common in a wide range of fields, including computer vision for autonomous vehicles and robots, mapping, surveillance, infrastructure inspection, and image distribution. Three-dimensional data is acquired using various methods, including distance sensors such as rangefinders, stereo cameras, and combinations of multiple single-lens reflex cameras.

[0004] One method of representing three-dimensional data is point cloud data, which represents the shape of a three-dimensional structure through a group of points in a three-dimensional space (for example, see Non-Patent Document 1). The position and color of the point group are stored in the point cloud data. Although point cloud data is expected to become the mainstream method of representing three-dimensional data, the amount of point group data is very large. Therefore, in the storage or transmission of three-dimensional data, as with two-dimensional dynamic images (for example, MPEG-4AVC or HEVC standardized by MPEG), it is necessary to compress the data volume through encoding.

[0005] Furthermore, compression of point cloud data is partially supported by a publicly available library (Point Cloud Library) that performs point cloud data association processing.

[0006] (Prior art literature)

[0007] (Non-patent literature)

[0008] Non-patent document 1 "Octree-Based Progressive Geometry Coding of PointClouds", Eurographics Symposium on Point-Based Graphics (2006)

[0009] In a three-dimensional information processing method or a three-dimensional information processing apparatus that processes such three-dimensional information, it is desired to be able to take appropriate measures even when an abnormality occurs in the three-dimensional position information. Summary of the Invention

[0010] An object of the present application is to provide a three-dimensional information processing method or a three-dimensional information processing apparatus that can appropriately respond even when an abnormality occurs in three-dimensional position information.

[0011] A three-dimensional information processing method involved in one form of the present application includes the following processing: determining whether first three-dimensional position information of a first range can be obtained via a communication path; and in the case that the first three-dimensional position information cannot be obtained via the communication path, obtaining second three-dimensional position information of a second range narrower than the first range via the communication path.

[0012] In addition, a three-dimensional information processing method involved in one form of the present application includes the following processing: determining whether the generation accuracy of the data of the three-dimensional position information generated based on the information detected by the sensor is above a reference value; and when the generation accuracy of the data of the three-dimensional position information is not above the reference value, generating a second three-dimensional position information based on information detected by an alternative sensor different from the sensor.

[0013] In addition, a three-dimensional information processing method involved in one form of the present application includes the following processing: using three-dimensional position information generated based on information detected by a sensor to estimate the self-position of a moving body having the sensor; using the result of the self-position estimation to enable the moving body to perform automatic driving; judging whether the generation accuracy of the data of the three-dimensional position information is above a benchmark value; and switching the mode of the automatic driving when the generation accuracy of the data of the three-dimensional position information is not above the benchmark value.

[0014] In addition, a three-dimensional information processing method involved in one form of the present application includes the following processing: obtaining map data including first three-dimensional position information through a communication path; obtaining second three-dimensional position information generated based on information detected by a sensor; judging whether the first three-dimensional position information can be obtained via the communication path; in a case where the first three-dimensional position information can be obtained via the communication path, using the first three-dimensional position information and the second three-dimensional position information to estimate the self-position of the mobile body having the sensor; and in a case where the first three-dimensional position information cannot be obtained via the communication path, obtaining map data including two-dimensional position information via the communication path, and using the two-dimensional position information and the second three-dimensional position information to estimate the self-position of the mobile body.

[0015] A three-dimensional information processing device involved in one form of the present application comprises: a judgment unit, which judges whether first three-dimensional position information of a first range can be obtained via a communication path; and an acquisition unit, which obtains second three-dimensional position information of a second range narrower than the first range via the communication path when the first three-dimensional position information cannot be obtained via the communication path.

[0016] In addition, a three-dimensional information processing device involved in one form of the present application comprises: an acquisition unit, which obtains map data including first three-dimensional position information through a communication path; and obtains second three-dimensional position information generated based on information detected by a sensor; a judgment unit, which judges whether the first three-dimensional position information can be obtained via the communication path; and a self-position estimation unit, which estimates the self-position of a moving body having the sensor by using the first three-dimensional position information and the second three-dimensional position information when the first three-dimensional position information can be obtained via the communication path, and obtains map data including two-dimensional position information via the communication path when the first three-dimensional position information cannot be obtained via the communication path, and estimates the self-position of the moving body by using the two-dimensional position information and the second three-dimensional position information.

[0017] A three-dimensional information processing method involved in one form of the present application includes: an acquisition step of obtaining map data including first three-dimensional position information through a communication path; a generation step of generating second three-dimensional position information based on information detected by a sensor; a judgment step of performing abnormality judgment processing on the first three-dimensional position information or the second three-dimensional position information, thereby judging whether the first three-dimensional position information or the second three-dimensional position information is abnormal; a decision step of deciding a response to the abnormality when it is judged that the first three-dimensional position information or the second three-dimensional position information is abnormal; and a work control step of executing the control required for implementing the response work.

[0018] In addition, all or specific forms of these can be implemented as systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, and can be implemented by combining systems, methods, integrated circuits, computer programs, and recording media.

[0019] The present application can provide a three-dimensional information processing method or a three-dimensional information processing device that can appropriately respond even when an abnormality occurs in three-dimensional position information. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 The structure of the encoded three-dimensional data according to the first embodiment is shown.

[0021] Figure 2 An example of a prediction structure between SPCs belonging to the lowest layer of the GOS according to the first embodiment is shown.

[0022] Figure 3 An example of the inter-layer prediction structure according to the first embodiment is shown.

[0023] Figure 4 An example of the coding order of the GOS according to the first embodiment is shown.

[0024] Figure 5 An example of the coding order of the GOS according to the first embodiment is shown.

[0025] Figure 6 This is a block diagram of the three-dimensional data encoding device involved in Embodiment 1.

[0026] Figure 7 This is a flowchart of the encoding process involved in Implementation 1.

[0027] Figure 8 This is a block diagram of the three-dimensional data decoding device according to the first embodiment.

[0028] Figure 9 This is a flowchart of the decoding process involved in Implementation 1.

[0029] Figure 10 An example of meta-information according to the first embodiment is shown.

[0030] Figure 11 A configuration example of a SWLD according to the second embodiment is shown.

[0031] Figure 12 An operation example of the server and the client according to the second embodiment is shown.

[0032] Figure 13 An operation example of the server and the client according to the second embodiment is shown.

[0033] Figure 14 An operation example of the server and the client according to the second embodiment is shown.

[0034] Figure 15 An operation example of the server and the client according to the second embodiment is shown.

[0035] Figure 16 This is a block diagram of a three-dimensional data encoding device according to the second embodiment.

[0036] Figure 17 This is a flowchart of the encoding process involved in Implementation 2.

[0037] Figure 18 This is a block diagram of a three-dimensional data decoding device according to the second embodiment.

[0038] Figure 19 This is a flowchart of the decoding process involved in Implementation 2.

[0039] Figure 20 A configuration example of a WLD according to the second embodiment is shown.

[0040] Figure 21 An example of the octree structure of the WLD according to the second embodiment is shown.

[0041] Figure 22 A configuration example of a SWLD according to the second embodiment is shown.

[0042] Figure 23 An example of the octree structure of the SWLD according to the second embodiment is shown.

[0043] Figure 24 This is a schematic diagram showing a state of transmission and reception of three-dimensional data between vehicles according to the third embodiment.

[0044] Figure 25 An example of three-dimensional data transmitted between vehicles according to the third embodiment is shown.

[0045] Figure 26 This is a block diagram of a three-dimensional data creation device according to the third embodiment.

[0046] Figure 27 This is a flowchart of the three-dimensional data creation process involved in the third embodiment.

[0047] Figure 28 This is a block diagram of a three-dimensional data transmitting device according to Embodiment 3.

[0048] Figure 29 This is a flowchart of the three-dimensional data transmission process involved in the third embodiment.

[0049] Figure 30 This is a block diagram of a three-dimensional data creation device according to the third embodiment.

[0050] Figure 31 This is a flowchart of the three-dimensional data creation process involved in the third embodiment.

[0051] Figure 32 This is a block diagram of a three-dimensional data transmitting device according to Embodiment 3.

[0052] Figure 33 This is a flowchart of the three-dimensional data transmission process involved in the third embodiment.

[0053] Figure 34 This is a block diagram of a three-dimensional information processing device according to the fourth embodiment.

[0054] Figure 35 This is a flowchart of the three-dimensional information processing method involved in the fourth embodiment.

[0055] Figure 36 This is a flowchart of the three-dimensional information processing method involved in the fourth embodiment.

[0056] Figure 37 The structure of the image information processing system is shown.

[0057] Figure 38 An example of a notification screen displayed when the camera is activated is shown.

[0058] Figure 39 This is a diagram of the overall structure of the content provision system that implements content distribution services.

[0059] Figure 40 This is a diagram showing the overall configuration of a digital broadcasting system.

[0060] Figure 41 An example of a smartphone is shown.

[0061] Figure 42 This is a block diagram showing an example configuration of a smartphone.

[0062] Explanation of symbols

[0063] 100, 400 three-dimensional data encoding device

[0064] 101, 201, 401, 501 obtain department

[0065] 102, 402 Coding region determination unit

[0066] 103 Division

[0067] 104, 644 Coding Department

[0068] 111, 607 3D data

[0069] 112, 211, 413, 414, 511, 634 coded 3D data

[0070] 200, 500 3D data decoding device

[0071] 202 Decoding start GOS determination unit

[0072] 203 Decoding SPC determination unit

[0073] 204, 625 Decoding Department

[0074] 212, 512, 513 decoding 3D data

[0075] 403 SWLD Extraction Department

[0076] 404 WLD Coding Department

[0077] 405 SWLD Coding Department

[0078] 411 Input 3D Data

[0079] 412 Extracting 3D Data

[0080] 502 Header Parsing

[0081] 503 WLD Decoding Department

[0082] 504 SWLD decoding unit

[0083] 600 own vehicle

[0084] 601 Surrounding vehicles

[0085] 602, 605 sensor detection range

[0086] Areas 603 and 606

[0087] 604 Occlusion Area

[0088] 620, 620A 3D data production device

[0089] 621, 641 3D Data Production Department

[0090] 622 Request Scope Determination Department

[0091] 623 Search Department

[0092] 624, 642 Receiving Department

[0093] 626 Synthesis Department

[0094] 627 Detection Area Determination Unit

[0095] 628 Surrounding Condition Detection Department

[0096] 629 Autonomous Movement Control Unit

[0097] 631, 651 sensor information

[0098] 632 1st 3D data

[0099] 633 Request range information

[0100] 635 Second 3D Data

[0101] 636 3rd dimensional data

[0102] 637 Order Signal

[0103] 638 Send Data

[0104] 639 Surrounding Condition Detection Results

[0105] 640, 640A 3D data transmission device

[0106] 643 Extraction Department

[0107] 645 Sending Department

[0108] 646 Transmission Permission Judgment Unit

[0109] 652 5th 3D data

[0110] 654 6th 3D Data

[0111] 700 Three-dimensional information processing device

[0112] 701 Three-dimensional Map Acquisition Department

[0113] 702 Vehicle detection data acquisition unit

[0114] 703 Abnormal Situation Judgment Department

[0115] 704 Response Work Decision Department

[0116] 705 Work Control Department

[0117] 711 3D Map

[0118] 712 3D data of vehicle detection DETAILED DESCRIPTION

[0119] When using coded data such as point cloud data in actual devices or services, random access to a desired spatial position or target object is required. However, to date, random access as a function in three-dimensional coded data does not exist, and therefore, a coding method for this purpose does not exist.

[0120] In the present application, a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device or a three-dimensional data decoding device is provided, which can provide a random access function in encoding three-dimensional data.

[0121] A three-dimensional data encoding method involved in one form of the present application encodes three-dimensional data, and the three-dimensional data encoding method includes: a division step, dividing the three-dimensional data into first processing units corresponding to three-dimensional coordinates respectively, and the first processing units are random access units; and an encoding step, generating encoded data by encoding each of a plurality of the first processing units.

[0122] According to this, random access is possible for each first processing unit. Thus, the three-dimensional data encoding method can provide a random access function in encoding three-dimensional data.

[0123] For example, the three-dimensional data encoding method may include a generation step, in which first information is generated, the first information showing multiple first processing units and three-dimensional coordinates corresponding to each of the multiple first processing units, and the encoded data includes the first information.

[0124] For example, the first information may further indicate at least one of an object, a time, and a data storage destination corresponding to each of the plurality of first processing units.

[0125] For example, in the dividing step, the first processing unit may be further divided into a plurality of second processing units, and in the encoding step, each of the plurality of second processing units may be encoded.

[0126] For example, in the encoding step, encoding may be performed on a second processing unit of a processing target included in a first processing unit of a processing target with reference to another second processing unit included in the first processing unit of the processing target.

[0127] According to this, by referring to other second processing units, the encoding efficiency can be improved.

[0128] For example, in the encoding step, as the type of the second processing unit of the processing object, one may be selected from a first type that does not refer to other second processing units, a second type that refers to one other second processing unit, and a third type that refers to two other second processing units, and the second processing unit of the processing object may be encoded according to the selected type.

[0129] For example, in the encoding step, the frequency of selecting the first type may be changed according to the number or density of objects included in the three-dimensional data.

[0130] This makes it possible to appropriately set the trade-off between random access and coding efficiency.

[0131] For example, in the encoding step, the size of the first processing unit may be determined according to the number or density of objects included in the three-dimensional data, or the number or density of dynamic objects.

[0132] This makes it possible to appropriately set the trade-off between random access and coding efficiency.

[0133] For example, it may also be that the first processing unit includes multiple layers that are spatially divided in a predetermined direction, each of the multiple layers includes one or more second processing units, and in the encoding step, the second processing unit is encoded with reference to the second processing unit included in the layer that is the same layer as the second processing unit or lower than the second processing unit.

[0134] This makes it possible to improve random access to important layers in the system, for example, and to suppress a decrease in coding efficiency.

[0135] For example, in the dividing step, the second processing unit including only static objects and the second processing unit including only dynamic objects may be allocated to different first processing units.

[0136] This makes it possible to easily control dynamic objects and static objects.

[0137] For example, in the encoding step, a plurality of dynamic objects may be encoded separately, and the encoded data of the plurality of dynamic objects may correspond to the second processing unit including only static objects.

[0138] This makes it possible to easily control dynamic objects and static objects.

[0139] For example, in the dividing step, the second processing unit may be further divided into a plurality of third processing units, and in the encoding step, each of the plurality of third processing units may be encoded.

[0140] For example, the third processing unit may include one or more voxels, where the voxel is the minimum unit corresponding to the position information.

[0141] For example, the second processing unit may include a feature point group derived from information obtained from a sensor.

[0142] For example, the encoded data may include information indicating an encoding order of a plurality of the first processing units.

[0143] For example, the encoded data may include information indicating the sizes of a plurality of the first processing units.

[0144] For example, in the encoding step, a plurality of the first processing units may be encoded in parallel.

[0145] Furthermore, a three-dimensional data decoding method involved in one form of the present application includes a decoding step, in which three-dimensional data of a first processing unit corresponding to a three-dimensional coordinate is decoded respectively, thereby generating the first processing unit, which is a random access unit.

[0146] Thus, random access to each first processing unit becomes possible. Thus, the three-dimensional data decoding method can provide a random access function in the encoded three-dimensional data.

[0147] Furthermore, it may also be that a three-dimensional data encoding device involved in one form of the present application includes: a division unit, which divides the three-dimensional data into first processing units corresponding to three-dimensional coordinates, respectively, and the first processing units are random access units; and an encoding unit, which generates encoded data by encoding each of a plurality of the first processing units.

[0148] This enables random access per first processing unit, and thus the three-dimensional data encoding device can provide a random access function in encoding three-dimensional data.

[0149] Furthermore, it may also be that a three-dimensional data decoding device involved in one form of the present application decodes three-dimensional data, and the three-dimensional data decoding device includes a decoding unit that generates three-dimensional data of the first processing unit by decoding each of the encoded data of the first processing unit corresponding to the three-dimensional coordinates, wherein the first processing unit is a random access unit.

[0150] This enables random access per first processing unit, and thus the three-dimensional data decoding apparatus can provide a random access function for encoded three-dimensional data.

[0151] In addition, the present application divides and encodes the space, thereby enabling quantization and prediction of the space, and is effective even when random access is not performed.

[0152] Furthermore, a three-dimensional data encoding method involved in one form of the present application includes: an extraction step of extracting second three-dimensional data having a feature value greater than a threshold value from the first three-dimensional data; and a first encoding step of generating first encoded three-dimensional data by encoding the second three-dimensional data.

[0153] Based on this, the 3D data encoding method generates first encoded 3D data by encoding data with a feature value greater than a threshold value. This reduces the amount of encoded 3D data compared to directly encoding the first 3D data. Therefore, the 3D data encoding method can reduce the amount of data required for transmission.

[0154] For example, the three-dimensional data encoding method may further include a second encoding step of encoding the first three-dimensional data to generate second encoded three-dimensional data.

[0155] According to this, the three-dimensional data encoding method can selectively transmit the first encoded three-dimensional data and the second encoded three-dimensional data according to the intended use, for example.

[0156] For example, the second three-dimensional data may be encoded using a first encoding method, and the first three-dimensional data may be encoded using a second encoding method different from the first encoding method.

[0157] Therefore, the three-dimensional data encoding method can adopt appropriate encoding methods for the first three-dimensional data and the second three-dimensional data, respectively.

[0158] For example, in the first encoding method, priority may be given to inter-frame prediction among intra-frame prediction and inter-frame prediction, compared with the second encoding method.

[0159] According to this, the three-dimensional data encoding method can increase the priority of inter-frame prediction for the second three-dimensional data where the correlation between adjacent data tends to be low.

[0160] For example, the first encoding method and the second encoding method may have different methods of expressing three-dimensional positions.

[0161] According to this, the three-dimensional data encoding method can adopt a more appropriate three-dimensional position expression method for three-dimensional data with different data amounts.

[0162] For example, at least one of the first encoded three-dimensional data and the second encoded three-dimensional data may include an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the first three-dimensional data or encoded three-dimensional data obtained by encoding a part of the first three-dimensional data.

[0163] With this, the decoding device can easily determine whether the obtained encoded three-dimensional data is the first encoded three-dimensional data or the second encoded three-dimensional data.

[0164] For example, in the first encoding step, the second three-dimensional data may be encoded so that the data size of the first encoded three-dimensional data is smaller than the data size of the second encoded three-dimensional data.

[0165] According to this, the three-dimensional data encoding method can make the data size of the first encoded three-dimensional data smaller than the data size of the second encoded three-dimensional data.

[0166] For example, in the extraction step, data corresponding to an object having a predetermined attribute may be further extracted from the first three-dimensional data as the second three-dimensional data.

[0167] According to this, the three-dimensional data encoding method can generate the first encoded three-dimensional data including the data required by the decoding device.

[0168] For example, the three-dimensional data encoding method may further include a sending step, in which one of the first encoded three-dimensional data and the second encoded three-dimensional data is sent to the client according to the state of the client.

[0169] Accordingly, the three-dimensional data encoding method can send appropriate data according to the status of the client.

[0170] For example, the status of the client may include a communication status of the client or a moving speed of the client.

[0171] For example, the three-dimensional data encoding method may further include a sending step, in which one of the first encoded three-dimensional data and the second encoded three-dimensional data is sent to the client in response to a request from the client.

[0172] Accordingly, the three-dimensional data encoding method can send appropriate data according to the client's request.

[0173] Furthermore, a three-dimensional data decoding method involved in one form of the present application includes: a first decoding step of decoding first encoded three-dimensional data using the first decoding method, wherein the first encoded three-dimensional data is obtained by encoding second three-dimensional data whose feature quantity extracted from the first three-dimensional data is greater than a threshold value; and a second decoding step of decoding second encoded three-dimensional data obtained by encoding the first three-dimensional data using a second decoding method different from the first decoding method.

[0174] Thus, the 3D data decoding method can selectively receive first encoded 3D data obtained by encoding data having a feature value exceeding a threshold value, and second encoded 3D data, for example, based on the intended use. This reduces the amount of data transmitted. Furthermore, the 3D data decoding method can employ appropriate decoding methods for each of the first and second 3D data.

[0175] For example, in the first decoding method, priority may be given to inter-frame prediction among intra-frame prediction and inter-frame prediction, compared with the second decoding method.

[0176] According to this, the three-dimensional data decoding method can increase the priority of inter-frame prediction for the second three-dimensional data where the correlation between adjacent data tends to be low.

[0177] For example, the first decoding method and the second decoding method may use different methods for expressing three-dimensional positions.

[0178] According to this, the three-dimensional data decoding method can adopt a more appropriate three-dimensional position expression method for three-dimensional data with different data amounts.

[0179] For example, at least one of the first encoded three-dimensional data and the second encoded three-dimensional data may include an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the first three-dimensional data or encoded three-dimensional data obtained by encoding a part of the first three-dimensional data, and the first encoded three-dimensional data and the second encoded three-dimensional data are identified with reference to the identifier.

[0180] According to this, the three-dimensional data decoding method can easily determine whether the obtained encoded three-dimensional data is the first encoded three-dimensional data or the second encoded three-dimensional data.

[0181] For example, the three-dimensional data decoding method may further include: a notification step of notifying the server of the client status; and a receiving step of receiving one of the first encoded three-dimensional data and the second encoded three-dimensional data sent from the server according to the client status.

[0182] Accordingly, the three-dimensional data decoding method can receive appropriate data according to the status of the client.

[0183] For example, the status of the client may include a communication status of the client or a moving speed of the client.

[0184] For example, the three-dimensional data decoding method may further include: a requesting step of requesting one of the first encoded three-dimensional data and the second encoded three-dimensional data from the server; and a receiving step of receiving one of the first encoded three-dimensional data and the second encoded three-dimensional data sent from the server in accordance with the request.

[0185] According to this, the three-dimensional data decoding method can receive appropriate data corresponding to the application.

[0186] Furthermore, a three-dimensional data encoding device according to one embodiment of the present application includes: an extraction unit that extracts second three-dimensional data having a feature value greater than a threshold value from first three-dimensional data; and a first encoding unit that generates first encoded three-dimensional data by encoding the second three-dimensional data.

[0187] In this manner, the 3D data encoding device generates first encoded 3D data by encoding data with a feature value greater than or equal to a threshold value. This reduces the amount of data compared to directly encoding the first 3D data. Therefore, the 3D data encoding device can reduce the amount of data required for transmission.

[0188] Furthermore, a three-dimensional data decoding device involved in one form of the present application comprises: a first decoding unit, which decodes first encoded three-dimensional data using a first decoding method, wherein the first encoded three-dimensional data is obtained by encoding second three-dimensional data whose feature quantity extracted from the first three-dimensional data is greater than a threshold value; and a second decoding unit, which decodes second encoded three-dimensional data obtained by encoding the first three-dimensional data using a second decoding method different from the first decoding method.

[0189] Thus, the 3D data decoding device can selectively receive first coded 3D data and second coded 3D data, obtained by encoding data with a feature value exceeding a threshold value, based on, for example, the intended use. This reduces the amount of data required for transmission. Furthermore, the 3D data decoding device can employ appropriate decoding methods for each of the first and second 3D data.

[0190] Furthermore, a three-dimensional data production method involved in one form of the present application includes: a production step of producing first three-dimensional data based on information detected by a sensor; a receiving step of receiving encoded three-dimensional data obtained by encoding second three-dimensional data; a decoding step of obtaining the second three-dimensional data by decoding the received encoded three-dimensional data; and a synthesis step of producing third three-dimensional data by synthesizing the first three-dimensional data and the second three-dimensional data.

[0191] According to this, the three-dimensional data creation method can create detailed third three-dimensional data using the created first three-dimensional data and the received second three-dimensional data.

[0192] For example, the combining step may combine the first three-dimensional data and the second three-dimensional data to create third three-dimensional data having a higher density than the first three-dimensional data and the second three-dimensional data.

[0193] For example, the second three-dimensional data may be three-dimensional data generated by extracting data having a feature value equal to or greater than a threshold value from fourth three-dimensional data.

[0194] According to this, the three-dimensional data production method can reduce the amount of three-dimensional data to be transmitted.

[0195] For example, the three-dimensional data production method may further include a searching step in which a transmitting device serving as a transmitting source of the encoded three-dimensional data is searched for, and in the receiving step, the encoded three-dimensional data is received from the searched transmitting device.

[0196] According to this, the three-dimensional data creation method can identify a transmitting device holding necessary three-dimensional data by searching, for example.

[0197] For example, the three-dimensional data production method may further include: a determination step of determining a request range, wherein the request range is the range of the three-dimensional space for requesting three-dimensional data; and a sending step of sending information showing the request range to the sending device, wherein the second three-dimensional data includes three-dimensional data of the request range.

[0198] Accordingly, the three-dimensional data production method can not only receive the required three-dimensional data but also reduce the amount of the transmitted three-dimensional data.

[0199] For example, in the determining step, a spatial range including an obstruction area that cannot be detected by the sensor may be determined as the requested range.

[0200] A three-dimensional data sending method involved in one form of the present application includes: a production step of producing fifth three-dimensional data based on information detected by a sensor; an extraction step of producing sixth three-dimensional data by extracting a portion of the fifth three-dimensional data; an encoding step of generating encoded three-dimensional data by encoding the sixth three-dimensional data; and a sending step of sending the encoded three-dimensional data.

[0201] According to this, the three-dimensional data transmission method can not only transmit the three-dimensional data produced by the device itself to other devices, but also reduce the amount of the three-dimensional data to be transmitted.

[0202] For example, in the creating step, seventh three-dimensional data may be created based on information detected by the sensor, and the fifth three-dimensional data may be created by extracting data having a feature value greater than a threshold value from the seventh three-dimensional data.

[0203] According to this, the three-dimensional data transmission method can reduce the amount of three-dimensional data to be transmitted.

[0204] For example, the three-dimensional data sending method may further include a receiving step, in which information indicating a requested range is received from a receiving device, and the requested range is the range of the three-dimensional space for requesting three-dimensional data; in the extraction step, the sixth three-dimensional data is produced by extracting the three-dimensional data of the requested range from the fifth three-dimensional data; and in the sending step, the encoded three-dimensional data is sent to the receiving device.

[0205] According to this, the three-dimensional data transmission method can reduce the amount of three-dimensional data to be transmitted.

[0206] Furthermore, a three-dimensional data production device involved in one form of the present application comprises: a production unit, which produces first three-dimensional data based on information detected by a sensor; a receiving unit, which receives encoded three-dimensional data obtained by encoding second three-dimensional data; a decoding unit, which obtains the second three-dimensional data by decoding the received encoded three-dimensional data; and a synthesis unit, which produces third three-dimensional data by synthesizing the first three-dimensional data and the second three-dimensional data.

[0207] With this, the three-dimensional data creation device can create detailed third three-dimensional data using the created first three-dimensional data and the received second three-dimensional data.

[0208] Furthermore, a three-dimensional data sending device involved in one form of the present application comprises: a production unit, which produces fifth three-dimensional data based on information detected by a sensor; an extraction unit, which produces sixth three-dimensional data by extracting a portion of the fifth three-dimensional data; an encoding unit, which generates encoded three-dimensional data by encoding the sixth three-dimensional data; and a sending unit, which sends the encoded three-dimensional data.

[0209] According to this, the three-dimensional data transmitting device can not only transmit the three-dimensional data produced by itself to other devices, but also reduce the amount of the transmitted three-dimensional data.

[0210] Furthermore, a three-dimensional information processing method involved in one form of the present application includes: an acquisition step of obtaining map data including first three-dimensional position information via a communication path; a generation step of generating second three-dimensional position information based on information detected by a sensor; a judgment step of judging whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing abnormality judgment processing on the first three-dimensional position information or the second three-dimensional position information; a decision step of deciding a response to the abnormality when the first three-dimensional position information or the second three-dimensional position information is judged to be abnormal; and a work control step of executing the control required for implementing the response work.

[0211] According to this, the three-dimensional information processing method can detect abnormality in the first three-dimensional position information or the second three-dimensional position information and can perform a corresponding operation.

[0212] For example, the first three-dimensional position information may be encoded in units of partial spaces having three-dimensional coordinate information, and the first three-dimensional position information includes multiple random access units, each of the multiple random access units being a collection of more than one partial spaces and capable of being independently decoded.

[0213] According to this, the three-dimensional information processing method can reduce the data volume of the obtained first three-dimensional position information.

[0214] For example, the first three-dimensional position information may be data in which a feature point having a three-dimensional feature value equal to or greater than a predetermined threshold value is encoded.

[0215] According to this, the three-dimensional information processing method can reduce the data volume of the obtained first three-dimensional position information.

[0216] For example, in the judging step, it may be judged whether the first three-dimensional position information can be obtained via the communication path, and if the first three-dimensional position information cannot be obtained via the communication path, the first three-dimensional position information is judged to be abnormal.

[0217] According to this, the three-dimensional information processing method performs appropriate handling when the first three-dimensional position information cannot be generated according to the communication status and the like.

[0218] For example, the three-dimensional information processing method may further include a self-position estimating step. In the self-position estimating step, the self-position of the mobile body having the sensor is estimated using the first three-dimensional position information and the second three-dimensional position information. In the judging step, it is predicted whether the mobile body will enter an area with poor communication status. In the operation control step, when it is predicted that the mobile body will enter an area with poor communication status, the mobile body obtains the first three-dimensional information before entering the area.

[0219] According to this, the three-dimensional information processing method can obtain the first three-dimensional position information in advance when there is a possibility that the first three-dimensional position information cannot be obtained.

[0220] For example, in the operation control step, when the first three-dimensional position information cannot be obtained via the communication path, third three-dimensional position information having a narrower range than the first three-dimensional position information may be obtained via the communication path.

[0221] According to this, the three-dimensional information processing method can reduce the amount of data obtained via a communication path, and can obtain three-dimensional position information even in a poor communication condition.

[0222] For example, the three-dimensional information processing method may further include a self-position estimating step, in which the self-position of the moving body having the sensor is estimated using the first three-dimensional position information and the second three-dimensional position information. In the operation control step, when the first three-dimensional position information cannot be obtained via the communication path, map data including two-dimensional position information is obtained via the communication path, and the self-position of the moving body having the sensor is estimated using the two-dimensional position information and the second three-dimensional position information.

[0223] According to this, the three-dimensional information processing method can reduce the amount of data obtained via the communication path, and thus can obtain three-dimensional position information even in a poor communication condition.

[0224] For example, the three-dimensional information processing method may further include an automatic driving step, in which the result of the self-position estimation is used to enable the mobile body to perform automatic driving, and in the judgment step, further, based on the moving environment of the mobile body, it is judged whether to perform automatic driving of the mobile body using the result of the self-position estimation of the mobile body based on the two-dimensional position information and the second three-dimensional position information.

[0225] Based on this, the three-dimensional information processing method can determine whether to continue autonomous driving according to the moving environment of the moving object.

[0226] For example, the three-dimensional information processing method may further include an automatic driving step, in which the result of the self-position estimation is used to enable the moving body to perform automatic driving, and in the working control step, the automatic driving mode is switched according to the moving environment of the moving body.

[0227] Based on this, the three-dimensional information processing method can set an appropriate autonomous driving mode according to the moving environment of the moving object.

[0228] For example, the determining step may determine whether the data of the first three-dimensional position information is complete, and if the data of the first three-dimensional position information is incomplete, determine that the first three-dimensional position information is abnormal.

[0229] According to this, the three-dimensional information processing method can perform appropriate countermeasures when, for example, the first three-dimensional position information is destroyed.

[0230] For example, in the judgment step, it is judged whether the generation accuracy of the data of the second three-dimensional position information is above a reference value. If the generation accuracy of the data of the second three-dimensional position information is not above the reference value, the second three-dimensional position information is judged to be abnormal.

[0231] According to this, the three-dimensional information processing method can perform appropriate countermeasures when the accuracy of the second three-dimensional position information is low.

[0232] For example, in the operation control step, when the generation accuracy of the second three-dimensional position information data is not equal to or greater than the reference value, the fourth three-dimensional position information may be generated based on information detected by an alternative sensor different from the sensor.

[0233] According to this, the three-dimensional information processing method can obtain three-dimensional position information using a substitute sensor, for example, when a sensor fails.

[0234] For example, the three-dimensional information processing method may further include: a self-position estimating step, using the first three-dimensional position information and the second three-dimensional position information to estimate the self-position of the moving body having the sensor; and an automatic driving step, using the result of the self-position estimation to enable the moving body to perform automatic driving, and in the work control step, when the generation accuracy of the data of the second three-dimensional position information is not above the benchmark value, the automatic driving mode is switched.

[0235] According to this, the three-dimensional information processing method performs an appropriate countermeasure when the accuracy of the second three-dimensional position information is low.

[0236] For example, in the operation control step, when the generation accuracy of the data of the second three-dimensional position information is not equal to or greater than the reference value, the operation calibration of the sensor may be performed.

[0237] According to this, the three-dimensional information processing method can improve the accuracy of the second three-dimensional position information when the accuracy of the second three-dimensional position information is low.

[0238] Furthermore, a three-dimensional information processing device involved in one form of the present application comprises: an acquisition unit, which acquires map data including first three-dimensional position information via a communication path; a generation unit, which generates second three-dimensional position information based on information detected by a sensor; a judgment unit, which judges whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing abnormality judgment processing on the first three-dimensional position information or the second three-dimensional position information; a decision unit, which decides a response work for the abnormality when it is judged that the first three-dimensional position information or the second three-dimensional position information is abnormal; and a work control unit, which performs the control required for implementing the response work.

[0239] With this, the three-dimensional information processing device can detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and perform a corresponding operation.

[0240] In addition, these general or specific forms can be implemented by systems, methods, integrated circuits, computer programs or computer-readable recording media such as CD-ROMs, and can be implemented by any combination of systems, methods, integrated circuits, computer programs and recording media.

[0241] The following detailed description of the embodiments is given with reference to the accompanying drawings. In addition, the embodiments to be described below are all specific examples of the present application. The numerical values, shapes, materials, constituent elements, configuration positions of constituent elements, connection forms, steps, order of steps, etc. shown in the following embodiments are all examples, and their purpose is not to limit the present application. Moreover, among the constituent elements of the following embodiments, the constituent elements that are not recorded in the technical solution showing the highest concept are described as arbitrary constituent elements.

[0242] (Implementation 1)

[0243] First, the data structure of encoded three-dimensional data (hereinafter also referred to as encoded data) according to this embodiment will be described. Figure 1 The structure of the encoded three-dimensional data according to this embodiment is shown.

[0244] In this embodiment, the three-dimensional space is divided into spaces (SPC) equivalent to the spaces of pictures in the encoding of dynamic images, and the three-dimensional data is encoded in units of space. The space is further divided into volumes (VLM) equivalent to macroblocks in dynamic image encoding, and prediction and conversion are performed in units of VLM. The volume includes a plurality of voxels (VXL), which are the smallest units corresponding to position coordinates. In addition, prediction means generating predicted three-dimensional data similar to the processing unit of the processing object with reference to other processing units, and encoding the difference between the predicted three-dimensional data and the processing unit of the processing object, similar to the prediction performed in the two-dimensional image. Furthermore, the prediction includes not only spatial prediction with reference to other prediction units at the same time, but also temporal prediction with reference to prediction units at different times.

[0245] For example, when encoding a three-dimensional space represented by point cloud data, such as point cloud data, a three-dimensional data encoding device (hereinafter referred to as an encoding device) encodes each point in the point cloud or multiple points contained within a voxel, all at once, according to the size of the voxel. Subdividing the voxels allows for a high-precision representation of the point cloud's three-dimensional shape, while increasing the voxel size allows for a coarse representation of the point cloud's three-dimensional shape.

[0246] In addition, although the following description uses the case where the three-dimensional data is point cloud data as an example, the three-dimensional data is not limited to point cloud data and can also be three-dimensional data in any form.

[0247] Furthermore, voxels with a hierarchical structure can be utilized. In this case, within the nth level, it is possible to sequentially indicate whether a sampling point exists in the n-1th level or lower levels (the levels below the nth level). For example, when decoding only the nth level, if a sampling point exists in the n-1th level or lower levels, decoding can be performed as if the sampling point exists at the center of the voxel in the nth level.

[0248] Furthermore, the encoding device obtains point group data through a distance sensor, a stereo camera, a monocular camera, a gyroscope, or an inertial sensor.

[0249] Similar to video encoding, spaces are classified into at least one of the following three prediction structures: independently decodable intra-frame space (I-SPC), unidirectionally referenced prediction space (P-SPC), and bidirectionally referenced prediction space (B-SPC). Spaces also contain two types of time information: decoding time and display time.

[0250] And, as Figure 1As shown, as a processing unit including a plurality of spaces, there is a GOS (Group of Space) which is a random access unit. Furthermore, as a processing unit including a plurality of GOS, there is a world space (WLD).

[0251] The spatial area occupied by world space is associated with an absolute position on Earth using GPS or latitude and longitude information. This location information is stored as metadata. Metadata can be included in the encoded data or transmitted separately.

[0252] Furthermore, within the GOS, all SPCs may be three-dimensionally adjacent, or there may be SPCs that are not three-dimensionally adjacent to other SPCs.

[0253] In addition, below, the encoding, decoding, or referencing of three-dimensional data contained in a processing unit such as a GOS, SPC, or VLM will also be referred to simply as encoding, decoding, or referencing the processing unit. Furthermore, the three-dimensional data contained in the processing unit includes, for example, at least one pair of spatial positions such as three-dimensional coordinates and characteristic values ​​such as color information.

[0254] Next, the prediction structure of the SPC in the GOS will be described. Although multiple SPCs in the same GOS or multiple VLMs in the same SPC occupy different spaces, they have the same time information (decoding time and display time).

[0255] Furthermore, within a GOS, the first SPC in decoding order is the I-SPC. Furthermore, there are two types of GOS: closed GOS and open GOS. A closed GOS is one that can decode all SPCs within the GOS when decoding starts from the first I-SPC. In an open GOS, some SPCs within the GOS that are earlier than the display time of the first I-SPC refer to a different GOS and can only be decoded in that GOS.

[0256] In addition, in coded data such as map information, the WLD may be decoded in the reverse order of the coding order. If there is a dependency between GOS, it will be difficult to reproduce the data in the reverse order. Therefore, in this case, a closed GOS is basically used.

[0257] Furthermore, the GOS has a layer structure in the height direction, and encoding or decoding is performed sequentially starting from the SPC of the bottom layer.

[0258] Figure 2 An example of the prediction structure between SPCs belonging to the lowest layer of the GOS is shown. Figure 3 An example of an inter-layer prediction structure is shown.

[0259] There are one or more I-SPCs within a GOS. While objects such as people, animals, cars, bicycles, traffic lights, and landmarks exist within a three-dimensional space, encoding small objects as I-SPCs is particularly effective. For example, a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes only the I-SPCs within the GOS when decoding the GOS with low throughput or high speed.

[0260] Furthermore, the encoding device may switch the encoding interval or the appearance frequency of the I-SPC according to the density of objects within the WLD.

[0261] And, in Figure 3 In the configuration shown, the encoding device or decoding device encodes or decodes multiple layers sequentially starting from the lower layer (layer 1). This allows, for example, autonomous vehicles to prioritize data near the ground, which contains a large amount of information.

[0262] In addition, in the coded data used by drones, etc., coding or decoding can be performed sequentially starting from the SPC of the upper layer in the height direction within the GOS.

[0263] Furthermore, the encoding device or decoding device may encode or decode multiple layers in such a manner that the decoding device can roughly grasp the GOS and gradually increase the resolution. For example, the encoding device or decoding device may encode or decode layers in the order of 3, 8, 1, 9, etc.

[0264] Next, the corresponding method of static objects and dynamic objects is explained.

[0265] In three-dimensional space, there are static objects or scenes such as buildings and roads (hereinafter referred to as static objects), as well as dynamic objects such as vehicles and people (hereinafter referred to as dynamic objects). Object detection can also be performed by extracting feature points from point cloud data or images captured by stereo cameras. Here, an example of a method for encoding dynamic objects is described.

[0266] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects using identification information.

[0267] For example, GOS is used as the identification unit. In this case, GOS including SPCs constituting static objects and GOS including SPCs constituting dynamic objects are distinguished within the coded data or by identification information stored separately from the coded data.

[0268] Alternatively, SPC is used as the identification unit. In this case, the SPC including only the VLM constituting the static object and the SPC including the VLM constituting the dynamic object are distinguished by the above-mentioned identification information.

[0269] Alternatively, VLM or VXL may be used as the identification unit. In this case, a VLM or VXL including a static object and a VLM or VXL including a dynamic object are distinguished by the above-mentioned identification information.

[0270] Furthermore, the encoding device may encode dynamic objects as one or more VLMs or SPCs, and may encode the VLMs or SPCs containing static objects and the SPCs containing dynamic objects as separate GOSs. Furthermore, if the size of the GOS is variable depending on the size of the dynamic object, the encoding device may store the size of the GOS separately as meta-information.

[0271] Furthermore, the encoding device encodes static objects and dynamic objects independently, allowing dynamic objects to be superimposed on the world space composed of static objects. In this case, a dynamic object is composed of one or more SPCs, each of which corresponds to one or more SPCs that constitute the static object with which it is superimposed. Furthermore, dynamic objects can be represented not by SPCs but by one or more VLMs or VXLs.

[0272] Furthermore, the encoding device may encode static objects and dynamic objects as different streams.

[0273] Furthermore, the encoding device may generate a GOS that includes one or more SPCs constituting a dynamic object. Furthermore, the encoding device may set the GOS (GOS_M) including the dynamic object and the GOS of the static object corresponding to the spatial region of the GOS_M to be of the same size (occupying the same spatial region). This enables overlapping processing on a GOS basis.

[0274] The P-SPC or B-SPC that constitutes a dynamic object can also refer to the SPC contained in a different coded GOS. When the position of a dynamic object changes over time and the same dynamic object is coded as a GOS at different times, cross-GOS reference is effective from the perspective of compression rate.

[0275] Furthermore, the first and second methods can be switched depending on the intended use of the encoded data. For example, when the encoded 3D data is used as a map, it is desirable to separate the data from dynamic objects, so the encoding device uses the second method. Alternatively, when encoding 3D data of events such as concerts or sports, where separation of dynamic objects is not necessary, the encoding device uses the first method.

[0276] Furthermore, the decoding time and display time of the GOS or SPC can be stored in the encoded data or as meta-information. Furthermore, the time information of static objects can all be the same. In this case, the actual decoding time and display time can be determined by the decoding device. Alternatively, different values ​​can be assigned to each GOS or SPC as the decoding time, and the same value can be assigned to all as the display time. Moreover, as shown in the decoder mode in dynamic image encoding such as HEVC's HRD (Hypothetical Reference Decoder), the decoder has a buffer of a specified size. As long as the bit stream is read at a specified bit rate according to the decoding time, a model that will not be destroyed and is guaranteed to be decodable can be imported.

[0277] Next, the configuration of GOS in the world space is explained. The coordinates of the three-dimensional space in the world space are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, z-axis). By setting a prescribed rule in the encoding order of GOS, spatially adjacent GOS can be encoded continuously in the encoded data. For example, Figure 4 In the example shown, the GOS within the xz plane are encoded continuously. After encoding all GOS within a single xz plane, the y-axis value is updated. In other words, as encoding continues, world space extends in the y-axis direction. Furthermore, the GOS index numbers are set to the encoding order.

[0278] Here, the three-dimensional world space corresponds one-to-one to absolute geographic coordinates such as GPS or latitude and longitude. Alternatively, the three-dimensional space can be represented by relative positions relative to a predetermined reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, and these direction vectors are stored as metadata along with the encoded data.

[0279] Furthermore, the size of the GOS is set to be fixed, and the encoding device stores this size as meta-information. Furthermore, the size of the GOS can be switched, for example, depending on whether the location is urban, indoors, or outdoors. In other words, the size of the GOS can be switched based on the quantity or nature of objects with informational value. Alternatively, within the same world space, the encoding device can appropriately switch the size of the GOS or the spacing of I-SPCs within the GOS based on, for example, the density of objects. For example, the encoding device can set the GOS size to be smaller and the spacing of I-SPCs within the GOS to be shorter when the density of objects is higher.

[0280] exist Figure 5In the example, in the area from the 3rd to the 10th GOS, the density of objects is high, so the GOS is subdivided to achieve fine-grained random access. In addition, the 7th to the 10th GOS exist on the back of the 3rd to the 6th GOS, respectively.

[0281] Next, the configuration and operation flow of the three-dimensional data encoding device according to this embodiment will be described. Figure 6 This is a block diagram of the three-dimensional data encoding device 100 according to this embodiment. Figure 7 3D data encoding apparatus 100 is a flowchart showing an example of its operation.

[0282] Figure 6 The three-dimensional data encoding device 100 shown generates encoded three-dimensional data 112 by encoding three-dimensional data 111. The three-dimensional data encoding device 100 includes an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.

[0283] like Figure 7 As shown, first, the acquisition unit 101 acquires three-dimensional data 111 as point cloud data ( S101 ).

[0284] Next, the coding region determination unit 102 determines a coding target region from the spatial region corresponding to the obtained point cloud data (S102). For example, the coding region determination unit 102 determines a spatial region around the position of the user or vehicle as the coding target region.

[0285] Next, the division unit 103 divides the point cloud data contained in the encoding target region into processing units. Here, the processing units are the aforementioned GOS and SPCs, for example. Furthermore, the encoding target region corresponds to the aforementioned world space, for example. Specifically, the division unit 103 divides the point cloud data into processing units based on the size of the pre-set GOS and the presence or size of dynamic objects (S103). Furthermore, the division unit 103 determines the starting position of the SPC in each GOS, which is the first SPC in the encoding order.

[0286] Next, the encoding unit 104 generates the encoded three-dimensional data 112 by sequentially encoding the plurality of SPCs in each GOS ( S104 ).

[0287] In addition, although an example is shown here in which the encoding target area is divided into GOS and SPC and each GOS is encoded, the processing order is not limited to the above. For example, the GOS may be encoded after the composition of the GOS is determined, and then the GOS composition may be determined.

[0288] In this manner, the 3D data encoding device 100 generates encoded 3D data 112 by encoding the 3D data 111. Specifically, the 3D data encoding device 100 divides the 3D data into random access units, namely, first processing units (GOS) corresponding to respective 3D coordinates. The first processing units (GOS) are divided into a plurality of second processing units (SPCs), and the second processing units (SPCs) are divided into a plurality of third processing units (VLMs). Furthermore, each third processing unit (VLM) includes one or more voxels (VXLs), which are the smallest units corresponding to positional information.

[0289] Next, the 3D data encoding device 100 generates encoded 3D data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the 3D data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Furthermore, the 3D data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).

[0290] For example, when the first processing unit (GOS) to be processed is a closed GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) to be processed included in the first processing unit (GOS) to be processed by referring to other second processing units (SPC) included in the first processing unit (GOS). In other words, the three-dimensional data encoding device 100 does not refer to the second processing unit (SPC) included in the first processing unit (GOS) different from the first processing unit (GOS) to be processed.

[0291] Furthermore, when the first processing unit (GOS) of the processing object is an open GOS, the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object is encoded with reference to other second processing units (SPC) included in the first processing unit (GOS) of the processing object, or to a second processing unit (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) of the processing object.

[0292] Furthermore, the three-dimensional data encoding device 100 selects one type of the second processing unit (SPC) of the processing object from among a first type (I-SPC) that does not refer to other second processing units (SPC), a second type (P-SPC) that refers to one other second processing unit (SPC), and a third type that refers to two other second processing units (SPC), and encodes the second processing unit (SPC) of the processing object according to the selected type.

[0293] Next, the configuration and operation flow of the three-dimensional data decoding device according to this embodiment will be described. Figure 8 This is a block diagram of a three-dimensional data decoding device 200 according to this embodiment. Figure 9 3D data decoding apparatus 200 is a flowchart showing an example of its operation.

[0294] Figure 8 The illustrated three-dimensional data decoding device 200 generates decoded three-dimensional data 212 by decoding encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. The three-dimensional data decoding device 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.

[0295] First, the acquisition unit 201 acquires the encoded 3D data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to metadata stored within or separately from the encoded 3D data 211 and determines the GOS that includes the SPC corresponding to the spatial position, object, or time at which decoding should start as the decoding start GOS.

[0296] Next, the decoded SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded within the GOS (S203). For example, the decoded SPC determination unit 203 determines (1) whether to decode only the I-SPC, (2) whether to decode both the I-SPC and the P-SPC, or (3) whether to decode all types. Alternatively, this step may be omitted if the type of SPC to be decoded is predetermined, such as when all SPCs are to be decoded.

[0297] Next, the decoding unit 204 obtains the first SPC in the decoding order (the same as the encoding order) within the GOS, and the address position at the beginning of the encoded 3D data 211. From this address position, it obtains the encoded data of the first SPC and decodes each SPC sequentially from the first SPC (S204). The address position is stored in meta-information, etc.

[0298] In this manner, the 3D data decoding apparatus 200 decodes the decoded 3D data 212. Specifically, the 3D data decoding apparatus 200 decodes each of the encoded 3D data 211 of the first processing unit (GOS) corresponding to the three-dimensional coordinates, thereby generating the decoded 3D data 212 of the first processing unit (GOS), which is a random access unit. More specifically, the 3D data decoding apparatus 200 decodes each of the plurality of second processing units (SPC) for each first processing unit (GOS). Furthermore, the 3D data decoding apparatus 200 decodes each of the plurality of third processing units (VLM) for each second processing unit (SPC).

[0299] The following describes the meta-information for random access. This meta-information is generated by the three-dimensional data encoding device 100 and included in the encoded three-dimensional data 112 (211).

[0300] In conventional random access of two-dimensional moving images, decoding starts from the first frame of the random access unit near a specified time. However, in world space, random access to (coordinates, objects, etc.) is also possible in addition to time.

[0301] Therefore, in order to realize random access to at least the three elements of coordinates, objects, and time, a table is prepared that associates each element with the index number of the GOS. Furthermore, the index number of the GOS is associated with the address of the I-SPC at the beginning of the GOS. Figure 10 An example of a table included in the meta information is shown. Figure 10 Of all the tables shown, at least one table may be used.

[0302] The following example illustrates random access using coordinates as the starting point. When accessing coordinates (x2, y2, z2), the system first references the coordinate-GOS table and determines that the location (x2, y2, z2) is included in the second GOS. Next, the system references the GOS address table and determines that the address of the first I-SPC in the second GOS is addr(2). Therefore, the decoding unit 204 obtains data from this address and begins decoding.

[0303] In addition, the address can be an address in a logical format or a physical address of an HDD or memory. Furthermore, information identifying a file segment can be used instead of an address. For example, a file segment is a unit obtained by segmenting one or more GOSs.

[0304] Furthermore, if an object spans multiple GOSs, the object GOS table may also indicate the GOSs to which the multiple objects belong. If the multiple GOSs are closed, the encoding and decoding devices can perform encoding and decoding in parallel. Furthermore, if the multiple GOSs are open, the GOSs can reference each other, further improving compression efficiency.

[0305] Examples of objects include people, animals, cars, bicycles, traffic lights, and buildings that serve as landmarks on land. For example, when encoding in world space, the three-dimensional data encoding device 100 extracts feature points unique to an object from three-dimensional point cloud data, detects the object based on the feature points, and can set the detected object as a random access point.

[0306] In this manner, the three-dimensional data encoding device 100 generates first information indicating a plurality of first processing units (GOS) and the three-dimensional coordinates corresponding to each of the plurality of first processing units (GOS). The encoded three-dimensional data 112 (211) includes the first information. Furthermore, the first information further indicates at least one of an object, a time, and a data storage destination corresponding to each of the plurality of first processing units (GOS).

[0307] The 3D data decoding device 200 obtains first information from the encoded 3D data 211 , uses the first information to determine the first processing unit of encoded 3D data 211 corresponding to the specified 3D coordinates, object or time, and decodes the encoded 3D data 211 .

[0308] Other examples of meta-information are described below. In addition to the meta-information for random access, the 3D data encoding device 100 can also generate and store the following meta-information. The 3D data decoding device 200 can also use this meta-information during decoding.

[0309] When using 3D data as map information, profiles are defined based on the intended use, and information indicating the profiles can be included in the meta-information. For example, profiles for urban areas or suburban areas, or for flying objects, can be defined, with the maximum and minimum sizes of the world space, SPC, or VLM defined for each. For example, in an urban area profile, more detailed information is required than in a suburban area, so the minimum size of the VLM is set smaller.

[0310] Meta-information may also include a tag value indicating the type of object. This tag value corresponds to the VLM, SPC, or GOS that constitutes the object. Tag values ​​can be set based on the object type, for example, a tag value of "0" representing a "person," a tag value of "1" representing a "car," and a tag value of "2" representing a "traffic light." Alternatively, when the object type is difficult or unnecessary to determine, a tag value indicating properties such as size or whether the object is dynamic or static can be used.

[0311] Furthermore, the meta-information may include information indicating the range of the spatial region occupied by the world space.

[0312] Furthermore, the meta-information may be used as header information shared by the entire stream of coded data or a plurality of SPCs such as an SPC within a GOS, and may store the size of the SPC or VXL.

[0313] Furthermore, the meta-information may include identification information of a distance sensor, a camera, or the like used in generating the point cloud data, or information indicating the positional accuracy of a point group within the point cloud data.

[0314] Also, the meta information may include information showing whether the world space is composed of only static objects or contains dynamic objects.

[0315] Modifications of this embodiment will be described below.

[0316] The encoding device or decoding device can encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on meta-information indicating the spatial position of the GOSs.

[0317] When three-dimensional data is used as a spatial map for a moving vehicle or flying object, or when such a spatial map is generated, the encoding device or decoding device can encode or decode the GOS or SPC contained in a space determined based on GPS, path information, or zoom factor, etc.

[0318] Furthermore, the decoding device may also decode sequentially starting from spaces closest to its own position or path. The encoding device or decoding device may also prioritize spaces farther from its own position or path, lowering the priority of spaces closer to it, and perform encoding or decoding accordingly. Here, lowering the priority means reducing the processing order, reducing the resolution (post-filtering processing), or reducing image quality (to improve coding efficiency, such as increasing the quantization step size).

[0319] Furthermore, when decoding coded data that has been spatially hierarchically coded, the decoding device may decode only the lower layer.

[0320] Furthermore, the decoding device may start decoding from a lower layer according to the zoom ratio or purpose of the map.

[0321] Furthermore, in applications such as estimating the self-position of a car or robot during autonomous driving or identifying objects, the encoding device or decoding device may also reduce the resolution of areas outside the area within a specified height from the road surface (the area for identification) for encoding or decoding.

[0322] Furthermore, the encoding device may also independently encode point cloud data representing indoor and outdoor spatial shapes. For example, by separating the GOS representing indoor spaces (indoor GOS) from the GOS representing outdoor spaces (outdoor GOS), the decoding device can select the GOS to decode based on the viewpoint position when using the encoded data.

[0323] Furthermore, the encoding device can encode indoor and outdoor GOSs that are close to each other in the coded stream. For example, the encoding device can associate identifiers between the two and store information indicating the corresponding identifiers within the coded stream or in separately stored meta-information. This allows the decoding device to identify the indoor and outdoor GOSs that are close to each other by referring to the information in the meta-information.

[0324] Furthermore, the encoding device may switch the size of the GOS or SPC between indoor and outdoor GOS. For example, the encoding device may set the size of the GOS to be smaller indoors than outdoors. Furthermore, the encoding device may change the accuracy of feature point extraction from point cloud data or the accuracy of object detection between indoor and outdoor GOS.

[0325] Furthermore, the encoding device can add information to the encoded data that allows the decoding device to distinguish between dynamic and static objects. This allows the decoding device to display dynamic objects in combination with a red frame or explanatory text. Alternatively, the decoding device can replace dynamic objects with a red frame or explanatory text alone. Furthermore, the decoding device can display more detailed object categories. For example, a car can be displayed with a red frame, while a person can be displayed with a yellow frame.

[0326] Furthermore, the encoding or decoding device may determine whether to encode or decode dynamic objects and static objects as different SPCs or GOSs based on the frequency of occurrence of dynamic objects or the ratio of static objects to dynamic objects. For example, if the frequency or ratio of dynamic objects exceeds a threshold, an SPC or GOS containing a mixture of dynamic and static objects is permitted. If the frequency or ratio of dynamic objects does not exceed the threshold, an SPC or GOS containing a mixture of dynamic and static objects is not permitted.

[0327] When dynamic objects are detected from 2D camera image information rather than point cloud data, the encoding device can obtain information identifying the detection result (such as a box or text) and the object position separately, encoding this information as part of the 3D encoded data. In this case, the decoding device displays auxiliary information representing the dynamic object (such as a box or text) overlaid on the decoded result of the static object.

[0328] Furthermore, the encoding device may change the density of VXL or VLM based on factors such as the complexity of the shape of the static object. For example, the encoding device may set VXL or VLM to be denser when the shape of the static object is more complex. Furthermore, the encoding device may determine the quantization step size when quantizing spatial position or color information based on the density of VXL or VLM. For example, the encoding device may set the quantization step size to be smaller when the VXL or VLM is denser.

[0329] As described above, the encoding device or decoding device according to the present embodiment performs spatial encoding or decoding in spatial units having coordinate information.

[0330] Furthermore, the encoding device and the decoding device perform encoding or decoding in units of volume in space. The volume includes voxels, which are the minimum units corresponding to the position information.

[0331] Furthermore, the encoding device and decoding device perform encoding or decoding by creating a table that associates various elements of spatial information, including coordinates, objects, and time, with GOPs, or by creating a table that associates various elements, thereby establishing a correspondence between arbitrary elements. Furthermore, the decoding device uses the values ​​of the selected elements to determine coordinates, identifies a volume, voxel, or space based on the coordinates, and decodes the space including the volume or voxel, or the identified space.

[0332] Furthermore, the encoding device determines a volume, voxel, or space that can be selected by an element through feature point extraction or object recognition, and encodes it as a volume, voxel, or space that can be randomly accessed.

[0333] Spaces are divided into three types: I-SPC, which can be encoded or decoded by the space itself; P-SPC, which can be encoded or decoded with reference to any one processed space; and B-SPC, which can be encoded or decoded with reference to any two processed spaces.

[0334] One or more volumes correspond to static objects or dynamic objects. The space containing static objects and the space containing dynamic objects are encoded or decoded as different GOSs. That is, the SPC containing static objects and the SPC containing dynamic objects are assigned to different GOSs.

[0335] Dynamic objects are encoded or decoded for each object and correspond to one or more spaces containing only static objects. In other words, multiple dynamic objects are encoded separately, and the resulting encoded data of multiple dynamic objects corresponds to an SPC containing only static objects.

[0336] The encoding and decoding devices increase the priority of the I-SPC within the GOS to perform encoding or decoding. For example, the encoding device performs encoding to minimize degradation of the I-SPC (enabling more faithful reproduction of the original 3D data after decoding). Furthermore, the decoding device, for example, only decodes the I-SPC.

[0337] The encoding device can change the frequency of using I-SPCs to perform encoding according to the density or number of objects in world space. Specifically, the encoding device changes the frequency of selecting I-SPCs according to the number or density of objects contained in the three-dimensional data. For example, the encoding device may increase the frequency of using I-space as the density of objects in world space increases.

[0338] Furthermore, the encoding device sets a random access point in units of GOS, and stores information indicating a spatial area corresponding to the GOS in the header information.

[0339] The encoding device, for example, uses a default value as the spatial size of the GOS. Alternatively, the encoding device may change the size of the GOS based on the number (quantity) or density of objects or dynamic objects. For example, the encoding device may set the spatial size of the GOS to be smaller when the objects or dynamic objects are denser or more numerous.

[0340] Furthermore, the space or volume includes a cluster of feature points derived from information obtained by sensors such as depth sensors, gyroscopes, or cameras. The coordinates of the feature points are set to the center of the voxel. Furthermore, by subdividing the voxels, it is possible to achieve higher precision in positional information.

[0341] The feature point cluster is derived using multiple pictures. The multiple pictures have at least two types of time information: actual time information and spatially corresponding time information common to multiple pictures (for example, encoding time used for rate control, etc.).

[0342] Furthermore, encoding or decoding is performed in units of GOS including one or more spaces.

[0343] The encoding device and the decoding device refer to the space in the processed GOS to predict the P space or B space in the GOS to be processed.

[0344] Alternatively, the encoding device and the decoding device predict the P space or B space in the GOS to be processed by using the processed space in the GOS to be processed without referring to different GOSs.

[0345] Furthermore, the encoding device and the decoding device transmit or receive the encoded stream in units of a world space including one or more GOSs.

[0346] Furthermore, the GOS has a layer structure in at least one direction within world space, and the encoding and decoding devices perform encoding or decoding starting from the lowest layer. For example, the randomly accessible GOS belongs to the lowest layer. The GOS belonging to a higher layer only references the GOS belonging to the layers below it. In other words, the GOS is spatially divided in a predetermined direction and includes multiple layers, each having one or more SPCs. The encoding and decoding devices perform encoding or decoding for each SPC by referencing SPCs contained in layers in the same layer or layers below it.

[0347] Furthermore, the encoding device and decoding device continuously encode or decode the GOS within a world space unit that includes multiple GOS. The encoding device and decoding device write or read information indicating the order (direction) of encoding or decoding as metadata. In other words, the encoded data includes information indicating the order in which the multiple GOS were encoded.

[0348] Furthermore, the encoding device and the decoding device encode or decode two or more different spaces or GOSs in parallel.

[0349] Furthermore, the encoding device and the decoding device encode or decode the space information (coordinates, size, etc.) of the space or the GOS.

[0350] Furthermore, the encoding device and the decoding device encode or decode the space or GOS included in the specific space specified based on external information related to the own position and / or area size, such as GPS, route information, or magnification.

[0351] The encoding device or decoding device performs encoding or decoding by giving a lower priority to a space far from the own position than to a space close to the own position.

[0352] The encoding device sets a direction in the world space according to a magnification or application, and encodes a GOS having a layer structure in that direction. Furthermore, the decoding device decodes the GOS having a layer structure in the world space direction set according to the magnification or application, preferentially starting from the lower layer.

[0353] The encoding device varies the accuracy of feature point extraction and object recognition, the size of the spatial area, etc. in indoor and outdoor spaces. However, the encoding device and decoding device encode or decode indoor GOS and outdoor GOS with close coordinates adjacent to each other in world space, and also encode or decode these identifiers in correspondence.

[0354] (Implementation Method 2)

[0355] When using encoded point cloud data in actual devices or services, it is desirable to transmit and receive only the required information according to the intended use in order to reduce network bandwidth. However, existing 3D data encoding structures do not have this capability, and therefore no corresponding encoding methods exist.

[0356] What will be described in this embodiment is a three-dimensional data encoding method and a three-dimensional data encoding device for providing the function of sending and receiving required information according to the purpose in the encoded data of three-dimensional point cloud data, as well as a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data.

[0357] A voxel (VXL) having a certain feature value or more is defined as a feature voxel (FVXL), and a world space (WLD) formed by the FVXL is defined as a sparse world space (SWLD). Figure 11 This section shows an example of a sparse world space and its configuration. SWLD includes: FGOS, a GOS constructed from FVXL; FSPC, an SPC constructed from FVXL; and FVLM, a VLM constructed from FVXL. The data structures and prediction structures of FGOS, FSPC, and FVLM can be the same as those of GOS, SPC, and VLM.

[0358] Feature quantities represent the three-dimensional positional information of the VXL or visible light information at the VXL position. In particular, feature quantities are more likely to be detected at corners and edges of three-dimensional objects. Specifically, these feature quantities may be three-dimensional feature quantities or visible light feature quantities as described below, but any feature quantity may be used as long as it represents the position, brightness, or color information of the VXL.

[0359] As the three-dimensional feature, a SHOT feature (Signature of Histograms of Orientations), a PFH feature (Point Feature Histograms), or a PPF feature (Point Pair Feature) is used.

[0360] SHOT features are obtained by dividing the area around the VXL, calculating the inner product between the reference point and the normal vector of the divided area, and then histogramming the results. This SHOT feature has high dimensionality and high expressiveness.

[0361] PFH features are obtained by selecting multiple pairs of points near VXL, calculating normal vectors and other parameters from these two points, and then histogramming them. Since these PFH features are histogram features, they are robust to small amounts of noise and have high expressiveness.

[0362] The PPF feature is calculated based on the VXL of two points using the normal vector, etc. Since all VXLs are used in this PPF feature, it is insensitive to occlusion.

[0363] Furthermore, as feature quantities of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients) that utilize information such as brightness gradient information of an image can be used.

[0364] SWLD is generated by calculating the above-mentioned feature quantity from each VXL of WLD and extracting FVXL. Here, SWLD can be updated every time WLD is updated, or it can be updated regularly after a certain period of time regardless of the update timing of WLD.

[0365] SWLDs can be generated for each feature. For example, as shown in SWLD1 based on SHOT features and SWLD2 based on SIFT features, SWLDs can be generated for each feature and used according to the application. Furthermore, the calculated feature values ​​for each FVXL can be stored as feature value information in each FVXL.

[0366] Next, the method of using the sparse world space (SWLD) will be described. Since the SWLD only contains feature voxels (FVXL), the data size is generally smaller than that of the WLD which includes all VXLs.

[0367] In applications that utilize feature quantities to achieve a specific purpose, using SWLD information instead of WLD can reduce the time required to read data from the hard disk and reduce the bandwidth and transmission time during network transmission. For example, by pre-storing WLD and SWLD as map information on a server, switching the transmitted map information to WLD or SWLD based on client requests can reduce network bandwidth and transmission time. A specific example is shown below.

[0368] Figure 12 as well as Figure 13 The following shows the use of SWLD and WLD. Figure 12 As shown, when client 1, which is an in-vehicle device, needs map information for determining its own position, client 1 sends a request to the server to obtain map data for estimating its own position (S301). The server sends SWLD to client 1 in accordance with the acquisition request (S302). Client 1 uses the received SWLD to determine its own position (S303). At this time, client 1 obtains VXL information of the surrounding area of ​​client 1 through various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple monocular cameras, and estimates its own position information based on the obtained VXL information and SWLD. Here, the own position information includes the three-dimensional position information and orientation of client 1.

[0369] like Figure 13 As shown, when client 2, an in-vehicle device, requires map information for mapping purposes such as three-dimensional maps, client 2 sends a request to the server to obtain map data for mapping (S311). In response to this request, the server sends a WLD to client 2 (S312). Client 2 uses the received WLD to perform map rendering (S313). At this point, client 2 creates a conceptual image using, for example, images captured by a visible light camera and the WLD received from the server, and then renders the created image on a screen such as a car navigation system.

[0370] As described above, the server sends SWLD to the client for applications that primarily require the characteristic values ​​of each VXL, such as estimating its own position, and sends WLD to the client when detailed VXL information is required, such as for map drawing. This enables efficient transmission and reception of map data.

[0371] In addition, the client can determine whether it needs SWLD or WLD and request the server to send SWLD or WLD. Moreover, the server can determine whether to send SWLD or WLD based on the status of the client or the network.

[0372] Next, a method for switching between the transmission and reception of the sparse world space (SWLD) and the world space (WLD) will be described.

[0373] The reception of WLD or SWLD can be switched according to the network bandwidth. Figure 14 An example of operation in this case is shown. For example, when a low-speed network with sufficient network bandwidth, such as in an LTE (Long Term Evolution) environment, is used, the client accesses the server via the low-speed network (S321) and obtains the SWLD (map information) from the server (S322). Alternatively, when a high-speed network with sufficient network bandwidth, such as in a WiFi environment, is used, the client accesses the server via the high-speed network (S323) and obtains the WLD from the server (S324). This allows the client to obtain appropriate map information based on the client's network bandwidth.

[0374] Specifically, the client receives SWLD via LTE outdoors, and acquires WLD via WiFi when entering a facility or other indoor location. This allows the client to acquire more detailed map information indoors.

[0375] In this way, the client can request WLD or SWLD from the server according to the frequency band of the network it is using. Alternatively, the client can send information indicating the frequency band of the network it is using to the server, and the server can send appropriate data (WLD or SWLD) to the client based on this information. Alternatively, the server can determine the network bandwidth of the client and send appropriate data (WLD or SWLD) to the client.

[0376] Furthermore, the reception of WLD or SWLD can be switched according to the moving speed. Figure 15 An example of operation in this case is shown. For example, when the client is moving at high speed (S331), the client receives SWLD from the server (S332). In addition, when the client is moving at low speed (S333), the client receives WLD from the server (S334). Accordingly, the client can suppress the network bandwidth and obtain map information according to the speed. Specifically, when the client is driving on a highway, by receiving SWLD with a small amount of data, the map information can be updated at a roughly appropriate speed. In addition, when the client is driving on a general road, by receiving WLD, more detailed map information can be obtained.

[0377] In this way, the client can request a WLD or SWLD from the server based on its own moving speed. Alternatively, the client can send information indicating its own moving speed to the server, and the server can send appropriate data (WLD or SWLD) to the client based on this information. Alternatively, the server can determine the client's moving speed and send appropriate data (WLD or SWLD) to the client.

[0378] Alternatively, the client can first obtain the SWLD from the server and then obtain the WLD of important areas within it. For example, when acquiring map data, the client can first use the SWLD to obtain general map information, filter out areas with a high incidence of features such as buildings, signs, or people, and then obtain the WLD of these filtered areas. This allows the client to reduce the amount of data received from the server while still obtaining detailed information about the desired area.

[0379] Alternatively, the server can create a separate SWLD for each object based on the WLD, and the client can receive each SWLD based on its intended purpose. This can reduce network bandwidth. For example, the server can pre-identify a person or a car from the WLD and create a SWLD for the person and a SWLD for the car. If the client wants to obtain information about people around it, it receives the SWLD for the person; if it wants to obtain information about the car, it receives the SWLD for the car. Furthermore, the type of SWLD can be distinguished based on information (flag or type, etc.) attached to the header.

[0380] Next, the configuration and operation flow of the three-dimensional data encoding device (for example, a server) according to this embodiment will be described. Figure 16 This is a block diagram of a three-dimensional data encoding device 400 according to this embodiment. Figure 17 This is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device 400.

[0381] Figure 16 The illustrated three-dimensional data encoding device 400 encodes input three-dimensional data 411 to generate encoded three-dimensional data 413 and 414 as encoded streams. Encoded three-dimensional data 413 corresponds to the WLD, and encoded three-dimensional data 414 corresponds to the SWLD. The three-dimensional data encoding device 400 includes an acquisition unit 401, a coding region determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and a SWLD encoding unit 405.

[0382] like Figure 17 As shown, first, the acquisition unit 401 acquires input three-dimensional data 411 which is point cloud data in a three-dimensional space ( S401 ).

[0383] Next, the coding region determination unit 402 determines a spatial region to be coded based on the spatial region where the point cloud data exists ( S402 ).

[0384] Next, the SWLD extraction unit 403 defines the spatial region to be encoded as a WLD and calculates features from each VXL contained in the WLD. Furthermore, the SWLD extraction unit 403 extracts VXLs whose features exceed a predetermined threshold, defines these VXLs as FVXLs, and appends these FVXLs to the SWLD to generate extracted three-dimensional data 412 (S403). Specifically, extracted three-dimensional data 412 with features exceeding the threshold is extracted from the input three-dimensional data 411.

[0385] Next, the WLD encoder 404 encodes the input 3D data 411 corresponding to the WLD to generate encoded 3D data 413 corresponding to the WLD (S404). At this time, the WLD encoder 404 appends information to the header of the encoded 3D data 413 that identifies the encoded 3D data 413 as a stream containing the WLD.

[0386] Then, the SWLD encoding unit 405 encodes the extracted 3D data 412 corresponding to the SWLD to generate encoded 3D data 414 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information to the header of the encoded 3D data 414 to distinguish that the encoded 3D data 414 is a stream containing the SWLD.

[0387] Furthermore, the order of the process of generating the encoded three-dimensional data 413 and the process of generating the encoded three-dimensional data 414 may be reversed. Furthermore, part or all of the above processes may be executed in parallel.

[0388] The information assigned to the header of the encoded three-dimensional data 413 and 414 is defined as a parameter such as "world_type". In the case of world_type = 0, it indicates that the stream contains WLD, and in the case of world_type = 1, it indicates that the stream contains SWLD. When more categories are defined, the assigned value can be increased, such as world_type = 2. In addition, a specific flag can be included in one of the encoded three-dimensional data 413 and 414. For example, the encoded three-dimensional data 414 can be assigned a flag indicating that the stream contains SWLD. In this case, the decoding device can determine whether it is a stream containing WLD or a stream containing SWLD based on the presence or absence of the flag.

[0389] Furthermore, the encoding method used by the WLD encoding unit 404 when encoding the WLD may be different from the encoding method used by the SWLD encoding unit 405 when encoding the SWLD.

[0390] For example, since data is decimated in SWLD, the correlation with surrounding data may be lower than that in WLD. Therefore, in the encoding method for SWLD, inter-frame prediction is prioritized over intra-frame prediction and inter-frame prediction in comparison with the encoding method for WLD.

[0391] Furthermore, the encoding method used for SWLD and the encoding method used for WLD may differ in the representation of the three-dimensional position. For example, the three-dimensional position of FVXL may be represented by three-dimensional coordinates in FWLD, while the three-dimensional position may be represented by an octree (described later) in WLD, or vice versa.

[0392] Furthermore, the SWLD encoding unit 405 encodes the SWLD encoded three-dimensional data 414 so that the data size is smaller than the WLD encoded three-dimensional data 413. For example, as described above, the correlation between data in SWLD and WLD may be lower. This may reduce the encoding efficiency, and the data size of the encoded three-dimensional data 414 may be larger than the data size of the WLD encoded three-dimensional data 413. Therefore, if the data size of the obtained encoded three-dimensional data 414 is larger than the data size of the WLD encoded three-dimensional data 413, the SWLD encoding unit 405 re-encodes the data to generate the encoded three-dimensional data 414 with a reduced data size.

[0393] For example, the SWLD extraction unit 403 generates extracted three-dimensional data 412 again with a reduced number of extracted feature points, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be coarsened. For example, in the octree structure described later, the degree of quantization can be coarsened by rounding the data in the lowest layer.

[0394] Furthermore, if the data size of the SWLD encoded three-dimensional data 414 cannot be made smaller than the data size of the WLD encoded three-dimensional data 413, the SWLD encoding unit 405 may not generate the SWLD encoded three-dimensional data 414. Alternatively, the WLD encoded three-dimensional data 413 may be copied to the SWLD encoded three-dimensional data 414. In other words, the WLD encoded three-dimensional data 413 may be used as the SWLD encoded three-dimensional data 414.

[0395] Next, the configuration and operation flow of the three-dimensional data decoding device (eg, client) according to this embodiment will be described. Figure 18 This is a block diagram of a three-dimensional data decoding device 500 according to this embodiment. Figure 19 3D data decoding processing performed by the 3D data decoding apparatus 500 is shown in FIG.

[0396] Figure 18 The illustrated three-dimensional data decoding apparatus 500 generates decoded three-dimensional data 512 or 513 by decoding encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding apparatus 400.

[0397] The three-dimensional data decoding device 500 includes an acquisition unit 501 , a header analysis unit 502 , a WLD decoding unit 503 , and a SWLD decoding unit 504 .

[0398] like Figure 19 As shown, first, the acquisition unit 501 obtains the encoded 3D data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded 3D data 511 to determine whether the encoded 3D data 511 is a stream containing a WLD or a stream containing a SWLD (S502). For example, this determination is made by referring to the world_type parameter described above.

[0399] If the encoded 3D data 511 is a stream containing a WLD ("Yes" in S503), the WLD decoding unit 503 decodes the encoded 3D data 511 to generate decoded 3D WLD data 512 (S504). If the encoded 3D data 511 is a stream containing a SWLD ("No" in S503), the SWLD decoding unit 504 decodes the encoded 3D data 511 to generate decoded 3D SWLD data 513 (S505).

[0400] Furthermore, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding the WLD may be different from the decoding method used by the SWLD decoding unit 504 when decoding the SWLD. For example, in the decoding method for the SWLD, priority may be given to the inter-frame prediction of intra-frame prediction and inter-frame prediction over the decoding method for the WLD.

[0401] Furthermore, the three-dimensional position representation method can be different between the decoding method used for SWLD and the decoding method used for WLD. For example, the three-dimensional position of FVXL can be represented by three-dimensional coordinates in SWLD, while the three-dimensional position can be represented by an octree (described later) in WLD, or vice versa.

[0402] Next, octree representation as a method of representing three-dimensional positions will be described. VXL data included in three-dimensional data is converted into an octree structure and then encoded. Figure 20 An example of VXL of WLD is shown. Figure 21 Shown Figure 20 The octree structure of WLD is shown in Figure 2. Figure 20 In the example shown, there are three VXLs 1 to 3 that are VXLs containing point groups (hereinafter referred to as valid VXLs). Figure 21 As shown, the octree structure consists of nodes and leaves. Each node has a maximum of 8 nodes or leaves. Each leaf has VXL information. Figure 21 Among the leaves shown, leaves 1, 2, and 3 represent Figure 20 VXL1, VXL2, and VXL3 shown.

[0403] Specifically, each node and leaf corresponds to a three-dimensional position. Figure 20 The block corresponding to node 1 is divided into eight blocks. Among the eight blocks, the blocks containing valid VXLs are set as nodes, and the remaining blocks are set as leaves. The blocks corresponding to the nodes are further divided into eight nodes or leaves. This process is repeated as many times as the number of levels in the tree structure. Finally, all the blocks at the bottom level are set as leaves.

[0404] and, Figure 22 Shown from Figure 20 An example of SWLD generated by WLD is shown. Figure 20 The feature extraction results of VXL1 and VXL2 shown are determined to be FVXL1 and FVXL2 and are included in SWLD. In addition, VXL3 is not determined to be FVXL and is therefore not included in SWLD. Figure 23 Shown Figure 22 The octree structure of SWLD is shown in Figure 1. Figure 23 In the octree structure shown, Figure 21 The leaf 3 shown, which corresponds to VXL3, is deleted. Figure 21 The node 3 shown does not have a valid VXL and is changed to a leaf. In this way, generally speaking, the number of leaves of SWLD is smaller than that of WLD, and the encoded three-dimensional data of SWLD is also smaller than that of WLD.

[0405] Modifications of this embodiment will be described below.

[0406] For example, when a client such as a vehicle-mounted device estimates its own position, it receives SWLD from the server, uses SWLD to estimate its own position, and performs obstacle detection. It then uses various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple monocular cameras to perform obstacle detection based on the three-dimensional information of the surrounding area obtained by itself.

[0407] Furthermore, it's generally difficult to include VXL data for flat areas in the SWLD. Therefore, the server maintains a subsampled world space (subWLD) that downsamples the WLD for stationary obstacle detection and can send both the SWLD and the subWLD to the client. This reduces network bandwidth while enabling client-side position estimation and obstacle detection.

[0408] Furthermore, when clients need to quickly render 3D map data, a grid-like structure is often more convenient. Therefore, the server can generate a grid based on the WLD and store it in advance as a grid world space (MWLD). For example, if the client requires a coarse 3D rendering, it receives the MWLD; if it requires a detailed 3D rendering, it receives the WLD. This helps reduce network bandwidth.

[0409] Furthermore, while the server sets the VXL with a feature value above a threshold value from each VXL as the FVXL, the FVXL can also be calculated using different methods. For example, if the server determines that the VXL, VLM, SPC, or GOS that constitute a signal or intersection are necessary for self-position estimation, driving assistance, or autonomous driving, they can be included in the SWLD as FVXL, FVLM, FSPC, or FGOS. Furthermore, this determination can be made manually. In addition, the FVXL, etc. obtained by the above method can be added to the FVXL, etc. set based on the feature value. That is, the SWLD extraction unit 403 can further extract data corresponding to objects with predetermined attributes from the input three-dimensional data 411 as extracted three-dimensional data 412.

[0410] Furthermore, labels different from feature values ​​can be assigned to situations required for these purposes. The server can separately maintain FVXL, which is required for estimating the position of itself, such as signals or intersections, driving assistance, or autonomous driving, as a higher layer of SWLD (e.g., lane world space).

[0411] Furthermore, the server can also add attributes to the VXL within the WLD in random access units or specified units. Attributes include, for example, information indicating whether a location is required or not for estimating the location, or information indicating whether traffic information such as signals or intersections is important. Attributes can also include the correspondence between lane information (GDF: Geographic Data Files, etc.) and features (such as intersections or roads).

[0412] Furthermore, as a method of updating WLD or SWLD, the following method can be adopted.

[0413] Update information showing changes in people, construction, or street trees (trajectory orientation) is uploaded to the server as a point group or metadata. The server updates the WLD based on this upload, and then updates the SWLD using the updated WLD.

[0414] Furthermore, if the client detects a mismatch between the 3D information generated by itself when estimating its own position and the 3D information received from the server, it can send the generated 3D information to the server along with an update notification. In this case, the server uses the WLD to update the SWLD. If the SWLD is not updated, the server determines that the WLD itself is outdated.

[0415] Furthermore, while information distinguishing between WLD and SWLD is added to the coded stream header, if there are multiple world spaces, such as a grid world space or a lane world space, information distinguishing between them can be added to the header. Furthermore, if there are multiple SWLDs with different feature values, information distinguishing between them can also be added to the header.

[0416] Furthermore, while the SWLD consists of FVXLs, it can also include VXLs that are not identified as FVXLs. For example, the SWLD can include adjacent VXLs used to calculate the characteristics of the FVXLs. This allows the client to calculate the characteristics of the FVXLs upon receiving the SWLD, even if the FVXLs in the SWLD do not have feature information attached. Furthermore, the SWLD can include information to distinguish whether each VXL is an FVXL or a VXL.

[0417] As described above, the three-dimensional data encoding device 400 extracts extracted three-dimensional data 412 (second three-dimensional data) whose feature value is above a threshold value from the input three-dimensional data 411 (first three-dimensional data), and generates encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.

[0418] Based on this, the three-dimensional data encoding device 400 generates encoded three-dimensional data 414 by encoding data with a feature value greater than or equal to a threshold value. This reduces the amount of data compared to directly encoding the input three-dimensional data 411. Therefore, the three-dimensional data encoding device 400 can reduce the amount of data required for transmission.

[0419] Furthermore, the three-dimensional data encoding device 400 further encodes the input three-dimensional data 411 to generate encoded three-dimensional data 413 (second encoded three-dimensional data).

[0420] With this, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414 according to the intended use, for example.

[0421] Furthermore, the extracted three-dimensional data 412 is encoded using a first encoding method, and the input three-dimensional data 411 is encoded using a second encoding method that is different from the first encoding method.

[0422] Accordingly, the three-dimensional data encoding device 400 can adopt appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.

[0423] Furthermore, in the first encoding method, compared with the second encoding method, priority is given to inter-frame prediction among intra-frame prediction and inter-frame prediction.

[0424] According to this, the three-dimensional data encoding apparatus 400 can increase the priority of inter-frame prediction for the extracted three-dimensional data 412 where the correlation between adjacent data tends to be low.

[0425] Furthermore, the first encoding method and the second encoding method use different methods to represent three-dimensional positions. For example, the second encoding method represents three-dimensional positions using an octree, while the first encoding method represents three-dimensional positions using three-dimensional coordinates.

[0426] With this, the three-dimensional data encoding device 400 can adopt a more appropriate three-dimensional position expression method for three-dimensional data having different data amounts (number of VXLs or FVXLs).

[0427] Furthermore, at least one of the coded 3D data 413 and 414 includes an identifier indicating whether the coded 3D data is obtained by encoding the input 3D data 411 or by encoding a portion of the input 3D data 411. In other words, the identifier indicates whether the coded 3D data is the coded 3D data 413 of the WLD or the coded 3D data 414 of the SWLD.

[0428] Based on this, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.

[0429] Furthermore, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 so that the data amount of the encoded three-dimensional data 414 is smaller than the data amount of the encoded three-dimensional data 413 .

[0430] As a result, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 smaller than the data amount of the encoded three-dimensional data 413 .

[0431] Furthermore, the three-dimensional data encoding device 400 extracts data corresponding to objects having predetermined attributes from the input three-dimensional data 411 as extracted three-dimensional data 412. For example, the objects having predetermined attributes may be objects required for self-position estimation, driving assistance, or autonomous driving, such as signals or intersections.

[0432] Thus, the three-dimensional data encoding device 400 can generate the encoded three-dimensional data 414 including the data required by the decoding device.

[0433] Furthermore, the three-dimensional data encoding device 400 (server) further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the status of the client.

[0434] Thus, the three-dimensional data encoding device 400 can send appropriate data according to the status of the client.

[0435] Furthermore, the client status includes the client's communication status (eg, network bandwidth) or the client's moving speed.

[0436] Furthermore, the three-dimensional data encoding device 400 further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the client's request.

[0437] Thus, the three-dimensional data encoding device 400 can send appropriate data according to the client's request.

[0438] Furthermore, the three-dimensional data decoding device 500 according to this embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400 .

[0439] Specifically, the three-dimensional data decoding apparatus 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature quantity extracted from the input three-dimensional data 411 is greater than or equal to a threshold value using a first decoding method. Furthermore, the three-dimensional data decoding apparatus 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 using a second decoding method that is different from the first decoding method.

[0440] Thus, the 3D data decoding device 500 can selectively receive the encoded 3D data 414 and 413, which are obtained by encoding data with feature quantities exceeding a threshold value, according to, for example, the intended use. This allows the 3D data decoding device 500 to reduce the amount of data transmitted. Furthermore, the 3D data decoding device 500 can employ appropriate decoding methods for each of the input 3D data 411 and the extracted 3D data 412.

[0441] Furthermore, in the first decoding method, compared with the second decoding method, priority is given to inter prediction between intra prediction and inter prediction.

[0442] With this, the three-dimensional data decoding apparatus 500 can increase the priority of inter-frame prediction for extracting three-dimensional data in which the correlation between adjacent data tends to be low.

[0443] Furthermore, the first decoding method and the second decoding method use different methods to represent the three-dimensional position. For example, the second decoding method represents the three-dimensional position using an octree, while the first decoding method represents the three-dimensional position using three-dimensional coordinates.

[0444] Thus, the three-dimensional data decoding apparatus 500 can adopt a more appropriate three-dimensional position expression method for three-dimensional data having different data amounts (number of VXLs or FVXLs).

[0445] Furthermore, at least one of the encoded 3D data 413 and 414 includes an identifier indicating whether the encoded 3D data is obtained by encoding the input 3D data 411 or by encoding a portion of the input 3D data 411. The 3D data decoding apparatus 500 identifies the encoded 3D data 413 and 414 by referring to the identifier.

[0446] Based on this, the three-dimensional data decoding apparatus 500 can easily determine whether the obtained encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414 .

[0447] Furthermore, the 3D data decoding device 500 notifies the server of the status of the client (the 3D data decoding device 500 ) and receives one of the encoded 3D data 413 and 414 transmitted from the server according to the status of the client.

[0448] Thus, the three-dimensional data decoding apparatus 500 can receive appropriate data according to the status of the client.

[0449] Furthermore, the client status includes the client's communication status (eg, network bandwidth) or the client's moving speed.

[0450] Furthermore, the three-dimensional data decoding apparatus 500 further requests one of the encoded three-dimensional data 413 and 414 from the server, and receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in accordance with the request.

[0451] This allows the three-dimensional data decoding apparatus 500 to receive appropriate data corresponding to the intended use.

[0452] (Implementation 3)

[0453] In this embodiment, a method of transmitting and receiving three-dimensional data between vehicles will be described.

[0454] Figure 24 It is a schematic diagram showing the state of transmission and reception of three-dimensional data 607 between the own vehicle 600 and the surrounding vehicles 601 .

[0455] When acquiring three-dimensional data using sensors mounted on the vehicle 600 (e.g., distance sensors such as rangefinders, stereo cameras, or a combination of multiple monocular cameras), obstacles such as surrounding vehicles 601 may result in areas within the sensor detection range 602 of the vehicle 600 where three-dimensional data cannot be generated (hereinafter referred to as occlusion areas 604). Furthermore, while the larger the area in which three-dimensional data is acquired, the higher the accuracy of autonomous operation, the sensor detection range of the vehicle 600 itself is limited.

[0456] The sensor detection range 602 of the host vehicle 600 includes an area 603 where three-dimensional data can be obtained and an obstruction area 604. The area where the host vehicle 600 intends to obtain three-dimensional data includes the sensor detection range 602 of the host vehicle 600 and other areas. Furthermore, the sensor detection range 605 of the surrounding vehicle 601 includes the obstruction area 604 and an area 606 not included in the sensor detection range 602 of the host vehicle 600.

[0457] Surrounding vehicles 601 transmit information detected by surrounding vehicles 601 to own vehicle 600. By acquiring information detected by surrounding vehicles 601, such as vehicles traveling ahead, own vehicle 600 can obtain three-dimensional data 607 for an obstructed area 604 and an area 606 outside the sensor detection range 602 of own vehicle 600. Using the information acquired from surrounding vehicles 601, own vehicle 600 supplements the three-dimensional data for obstructed area 604 and area 606 outside the sensor detection range.

[0458] The use of three-dimensional data in autonomous vehicle or robot operations is to estimate the vehicle's position, detect surrounding conditions, or both. For example, to estimate the vehicle's position, three-dimensional data generated by the vehicle 600 based on sensor information from the vehicle 600 is used. To detect surrounding conditions, three-dimensional data obtained from surrounding vehicles 601 is used in addition to the three-dimensional data generated by the vehicle 600.

[0459] The surrounding vehicles 601 for transmitting three-dimensional data 607 to the vehicle 600 can be determined based on the state of the vehicle 600. For example, the surrounding vehicles 601 may be the vehicle in front of the vehicle 600 when the vehicle 600 is traveling straight, the oncoming vehicle when the vehicle 600 is turning right, or the vehicle behind the vehicle when the vehicle 600 is backing up. Alternatively, the driver of the vehicle 600 may directly designate the surrounding vehicles 601 for transmitting three-dimensional data 607 to the vehicle 600.

[0460] Furthermore, the vehicle 600 may search for surrounding vehicles 601 that hold 3D data for areas that the vehicle 600 cannot obtain, and that are within the space for which the 3D data 607 is to be obtained. Areas that the vehicle 600 cannot obtain include, for example, an occlusion area 604 and an area 606 outside the sensor detection range 602.

[0461] Furthermore, the vehicle 600 may determine the blocked area 604 based on sensor information of the vehicle 600. For example, the vehicle 600 may determine the blocked area 604 as an area within the sensor detection range 602 of the vehicle 600 where three-dimensional data cannot be generated.

[0462] The following describes an operation example in which the vehicle traveling ahead transmits the three-dimensional data 607 . Figure 25 An example of three-dimensional data transmitted in this case is shown.

[0463] like Figure 25 As shown, the three-dimensional data 607 transmitted from the preceding vehicle is, for example, a sparse world space (SWLD) of point cloud data. Specifically, the preceding vehicle generates three-dimensional WLD data (point cloud data) based on information detected by its sensors. The preceding vehicle generates the SWLD data (point cloud data) by extracting data with feature quantities exceeding a threshold from the WLD data. The preceding vehicle then transmits the generated SWLD data to the vehicle 600.

[0464] The own vehicle 600 receives the SWLD and merges the received SWLD into the point cloud data created by the own vehicle 600 .

[0465] The transmitted SWLD includes information on absolute coordinates (the position of the SWLD in the coordinate system of the three-dimensional map). The vehicle 600 can perform a merging process by superimposing the point cloud data generated by the vehicle 600 based on the absolute coordinates.

[0466] The SWLD transmitted from the surrounding vehicle 601 may be the SWLD of an area 606 outside the sensor detection range 602 of the own vehicle 600 and within the sensor detection range 605 of the surrounding vehicle 601, or the SWLD of an obstruction area 604 relative to the own vehicle 600, or both. Furthermore, the transmitted SWLD may be the SWLD of an area used by the surrounding vehicle 601 for detecting surrounding conditions, among the aforementioned SWLDs.

[0467] Furthermore, the surrounding vehicle 601 can vary the density of the transmitted point cloud data according to the communication time based on the speed difference between the own vehicle 600 and the surrounding vehicle 601. For example, when the speed difference is large and the communication time is short, the surrounding vehicle 601 can reduce the density (data volume) of the point cloud data by extracting three-dimensional points with large feature values ​​from the SWLD.

[0468] Furthermore, the detection of surrounding conditions refers to determining the presence of people, vehicles, road construction equipment, etc., identifying their types, and detecting their positions, moving directions, and moving speeds.

[0469] Furthermore, the own vehicle 600 may obtain the braking information of the surrounding vehicles 601 instead of the three-dimensional data 607 generated by the surrounding vehicles 601, or in addition to the three-dimensional data 607. Here, the braking information of the surrounding vehicles 601 is, for example, information indicating whether the accelerator or brake of the surrounding vehicles 601 is depressed, or the degree to which the accelerator or brake is depressed.

[0470] Furthermore, the point cloud data generated by each vehicle is segmented into random access units, taking into account low-latency communication between vehicles. Furthermore, map data downloaded from the server, such as 3D maps, is segmented into larger random access units than in the case of inter-vehicle communication.

[0471] Data in areas that are prone to being blocked, such as an area in front of a preceding vehicle or an area behind a following vehicle, is divided into small random access units as data for low latency.

[0472] When traveling at high speed, the importance of the front becomes higher, so each vehicle generates SWLD in a narrow viewing angle range in small random access units when traveling at high speed.

[0473] When the SWLD created for transmission by the preceding vehicle includes an area where point cloud data can be obtained by the own vehicle 600, the preceding vehicle can reduce the transmission amount by removing the point cloud data of the area.

[0474] Next, the configuration and operation of the three-dimensional data creation device 620 as the three-dimensional data receiving device according to this embodiment will be described.

[0475] Figure 26 This is a block diagram of a three-dimensional data creation device 620 according to this embodiment. This three-dimensional data creation device 620 is included in, for example, the aforementioned vehicle 600 and creates denser third three-dimensional data 636 by combining received second three-dimensional data 635 with first three-dimensional data 632 created by the three-dimensional data creation device 620.

[0476] The three-dimensional data creation device 620 includes a three-dimensional data creation unit 621 , a request range determination unit 622 , a search unit 623 , a reception unit 624 , a decoding unit 625 , and a synthesis unit 626 . Figure 27 This is a flowchart showing the operation of the three-dimensional data creation device 620.

[0477] First, the three-dimensional data creation unit 621 creates first three-dimensional data 632 (S621) using sensor information 631 detected by sensors included in the vehicle 600. Next, the request range determination unit 622 determines a request range, which is the three-dimensional spatial range where the data in the created first three-dimensional data 632 is insufficient (S622).

[0478] Next, the search unit 623 searches for surrounding vehicles 601 that hold three-dimensional data for the requested range and transmits request range information 633 indicating the requested range to the surrounding vehicles 601 identified through the search (S623). Next, the receiving unit 624 receives the encoded three-dimensional data 634 as a coded stream for the requested range from the surrounding vehicles 601 (S624). Alternatively, the search unit 623 can issue a request indiscriminately to all vehicles within the identified range and receive the encoded three-dimensional data 634 from any responding vehicles. Furthermore, the search unit 623 is not limited to vehicles; it can also issue a request to objects such as traffic lights and signs and receive the encoded three-dimensional data 634 from those objects.

[0479] Next, the decoding unit 625 decodes the received encoded 3D data 634 to obtain second 3D data 635 (S625). Next, the synthesizing unit 626 synthesizes the first 3D data 632 and the second 3D data 635 to create denser third 3D data 636 (S626).

[0480] Next, the configuration and operation of the three-dimensional data transmitting device 640 according to this embodiment will be described. Figure 28 It is a block diagram of the three-dimensional data transmitting device 640.

[0481] The three-dimensional data sending device 640 is included in the above-mentioned surrounding vehicle 601, for example, and processes the fifth three-dimensional data 652 produced by the surrounding vehicle 601 into the sixth three-dimensional data 654 requested by the own vehicle 600, and generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and sends the encoded three-dimensional data 634 to the own vehicle 600.

[0482] The three-dimensional data transmitting device 640 includes a three-dimensional data creating unit 641 , a receiving unit 642 , an extracting unit 643 , an encoding unit 644 , and a transmitting unit 645 . Figure 29 3D data transmitting apparatus 640 is a flowchart showing the operation of the apparatus 640 .

[0483] First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 ( S641 ) using sensor information 651 detected by sensors included in the surrounding vehicles 601. Next, the receiving unit 642 receives the requested range information 633 transmitted from the own vehicle 600 ( S642 ).

[0484] Next, the extraction unit 643 extracts the three-dimensional data within the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652, processing the fifth three-dimensional data 652 into sixth three-dimensional data 654 (S643). The encoding unit 644 then encodes the sixth three-dimensional data 654 to generate encoded three-dimensional data 634 as an encoded stream (S644). The transmission unit 645 then transmits the encoded three-dimensional data 634 to the vehicle 600 (S645).

[0485] In addition, although the example in which the own vehicle 600 has the three-dimensional data creation device 620 and the surrounding vehicle 601 has the three-dimensional data transmission device 640 is described here, each vehicle may also have the functions of the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.

[0486] Hereinafter, the configuration and operation of the three-dimensional data creation device 620 in the case of being a surrounding situation detection device that realizes a process of detecting the surrounding situation of the own vehicle 600 will be described. Figure 30 3D data creation device 620A is a block diagram showing the configuration of the 3D data creation device 620A in this case. Figure 30 The three-dimensional data production device 620A shown in FIG. Figure 26 In addition to the configuration of the three-dimensional data creation device 620 shown, the three-dimensional data creation device 620A further includes a detection area determination unit 627 , a surrounding situation detection unit 628 , and an autonomous operation control unit 629 .

[0487] Figure 31This is a flowchart of the surrounding situation detection process of the vehicle 600 performed by the three-dimensional data creation device 620A.

[0488] First, the three-dimensional data creation unit 621 creates first three-dimensional data 632 (S661) as point cloud data using sensor information 631 of the detection range of the vehicle 600, detected by sensors included in the vehicle 600. Furthermore, the three-dimensional data creation device 620A can also use the sensor information 631 to estimate its own position.

[0489] Next, the detection area determination unit 627 determines a detection target range as a spatial area for detecting surrounding conditions (S662). For example, the detection area determination unit 627 calculates the area required for detecting surrounding conditions for safe autonomous operation based on the autonomous operation (autonomous driving) conditions such as the driving direction and speed of the vehicle 600, and determines this area as the detection target range.

[0490] Next, the request range determination unit 622 determines the blocked area 604 and a spatial area that is outside the detection range of the sensor of the own vehicle 600 but is necessary for the detection of the surrounding situation as the request range ( S663 ).

[0491] If the request range determined in step S663 exists ("Yes" in S664), the search unit 623 searches for surrounding vehicles that have information related to the request range. For example, the search unit 623 can inquire about the surrounding vehicles whether they have information related to the request range, and determine whether the surrounding vehicles have information related to the request range based on the request range and the position of the surrounding vehicles. Next, the search unit 623 sends a request signal 637 for requesting the transmission of three-dimensional data to the surrounding vehicles 601 determined by the search. In addition, after receiving the permission signal sent from the surrounding vehicle 601 indicating acceptance of the request of the request signal 637, the search unit 623 sends the request range information 633 indicating the request range to the surrounding vehicles 601 (S665).

[0492] Next, the receiving unit 624 detects a transmission notification of the transmission data 638 as information on the request range, and receives the transmission data 638 ( S666 ).

[0493] Alternatively, the three-dimensional data creation device 620A may issue a request to all vehicles within a specified range without discriminating, rather than searching for a destination to send a request to, and receive data 638 from any vehicle that responds with information related to the requested range. Furthermore, the search unit 623 is not limited to vehicles; it may also issue a request to an object such as a traffic light or sign and receive data 638 from that object.

[0494] Furthermore, the transmitted data 638 includes at least one of the encoded three-dimensional data 634 generated by the surrounding vehicle 601 and obtained by encoding the three-dimensional data of the requested range, and the surrounding situation detection results 639 of the requested range. The surrounding situation detection results 639 indicate the position, movement direction, and movement speed of people and vehicles detected by the surrounding vehicle 601. Furthermore, the transmitted data 638 may also include information indicating the position and movement of the surrounding vehicle 601. For example, the transmitted data 638 may also include braking information of the surrounding vehicle 601.

[0495] If the received transmission data 638 includes the encoded three-dimensional data 634 ("Yes" in S667), the decoder 625 decodes the encoded three-dimensional data 634 to obtain the second three-dimensional SWLD data 635 (S668). In other words, the second three-dimensional data 635 is three-dimensional data (SWLD) generated by extracting data having a feature value greater than a threshold value from the fourth three-dimensional data (WLD).

[0496] Next, the first three-dimensional data 632 and the second three-dimensional data 635 are synthesized by the synthesizing unit 626 to generate third three-dimensional data 636 ( S669 ).

[0497] Next, the surrounding situation detection unit 628 detects the surrounding situation of the own vehicle 600 using the third three-dimensional data 636, which is point cloud data of the spatial region required for surrounding situation detection (S670). Furthermore, if the received transmission data 638 includes surrounding situation detection results 639, the surrounding situation detection unit 628 uses these results in addition to the third three-dimensional data 636 to detect the surrounding situation of the own vehicle 600. Furthermore, if the received transmission data 638 includes braking information of the surrounding vehicle 601, the surrounding situation detection unit 628 uses this braking information in addition to the third three-dimensional data 636 to detect the surrounding situation of the own vehicle 600.

[0498] Next, the autonomous operation control unit 629 controls the autonomous operation (automatic driving) of the vehicle 600 (S671) based on the surrounding situation detection result performed by the surrounding situation detection unit 628. In addition, the surrounding situation detection result can be presented to the driver through a UI (user interface) or the like.

[0499] If the requested range does not exist in step S663 ("No" in S664), that is, if all the spatial area information required for surrounding situation detection has been generated based on the sensor information 631, the surrounding situation detection unit 628 detects the surrounding situation of the host vehicle 600 using the point cloud data of the spatial area required for surrounding situation detection, that is, the first three-dimensional data 632 (S672). Then, the autonomous operation control unit 629 controls the autonomous operation (automatic driving) of the host vehicle 600 based on the surrounding situation detection results of the surrounding situation detection unit 628 (S671).

[0500] If the received transmission data 638 does not include the encoded three-dimensional data 634 (No in S667), that is, if the transmission data 638 only includes the surrounding condition detection result 639 or braking information of the surrounding vehicle 601, the surrounding condition detection unit 628 detects the surrounding conditions of the own vehicle 600 using the first three-dimensional data 632 and the surrounding condition detection result 639 or braking information (S673). The autonomous operation control unit 629 then controls the autonomous operation (autonomous driving) of the own vehicle 600 based on the surrounding condition detection results of the surrounding condition detection unit 628 (S671).

[0501] Next, the three-dimensional data transmitting device 640A that transmits the transmission data 638 to the three-dimensional data creating device 620A described above will be described. Figure 32 FIG. 6 is a block diagram of the three-dimensional data transmitting device 640A.

[0502] Figure 32 The three-dimensional data transmitting device 640A shown in FIG. Figure 28 In addition to the configuration of the three-dimensional data transmitting device 640 shown, the device further includes a transmission availability determination unit 646. The three-dimensional data transmitting device 640A is included in the surrounding vehicle 601.

[0503] Figure 33 3D data transmission device 640A operates as follows: First, the 3D data creation unit 641 creates fifth 3D data 652 using sensor information 651 detected by sensors included in the surrounding vehicle 601 (S681).

[0504] Next, the receiving unit 642 receives a delegation signal 637 for requesting the transmission of three-dimensional data from the own vehicle 600 (S682). Next, the transmission feasibility determination unit 646 decides whether to respond to the delegation indicated by the delegation signal 637 (S683). For example, the transmission feasibility determination unit 646 decides whether to respond to the delegation based on the content pre-set by the user. Alternatively, the receiving unit 642 may first accept a request from the other party such as a request range, and the transmission feasibility determination unit 646 may decide whether to respond to the delegation based on the content. For example, the transmission feasibility determination unit 646 may decide to respond to the delegation when it holds three-dimensional data within the requested range, and may decide not to respond to the delegation when it does not have three-dimensional data within the requested range.

[0505] If the request is accepted ("Yes" in S683), the three-dimensional data transmitting device 640A transmits a permission signal to the vehicle 600, and the receiving unit 642 receives the requested range information 633 indicating the requested range (S684). Next, the extraction unit 643 extracts the point cloud data of the requested range from the fifth three-dimensional data 652, which is point cloud data, and creates the transmission data 638, which includes the SWLD of the extracted point cloud data, namely the sixth three-dimensional data 654 (S685).

[0506] Specifically, the three-dimensional data transmitting device 640A may generate seventh three-dimensional data (WLD) based on the sensor information 651 and extract data having a feature value greater than a threshold value from the seventh three-dimensional data (WLD) to generate fifth three-dimensional data 652 (SWLD). Furthermore, the three-dimensional data generating unit 641 may pre-generate three-dimensional data for the SWLD, and the extracting unit 643 may extract three-dimensional data for the SWLD within the requested range from the three-dimensional data for the SWLD. Alternatively, the extracting unit 643 may generate three-dimensional data for the SWLD within the requested range based on the three-dimensional data for the WLD within the requested range.

[0507] Furthermore, the transmitted data 638 may include the surrounding condition detection result 639 of the requested range performed by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601. Furthermore, the transmitted data 638 may not include the sixth three-dimensional data 654, but may include only at least one of the surrounding condition detection result 639 of the requested range performed by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601.

[0508] When the transmission data 638 includes the sixth three-dimensional data 654 (Yes in S686 ), the encoding unit 644 encodes the sixth three-dimensional data 654 to generate the encoded three-dimensional data 634 ( S687 ).

[0509] Then, the transmitting unit 645 transmits the transmission data 638 including the encoded three-dimensional data 634 to the own vehicle 600 ( S688 ).

[0510] Furthermore, in the case where the sending data 638 does not include the sixth three-dimensional data 654 ("No" in S686), the sending unit 645 sends the sending data 638 including at least one of the surrounding condition detection result 639 of the requested range performed by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601 to the own vehicle 600 (S688).

[0511] Modifications of this embodiment will be described below.

[0512] For example, the information transmitted from surrounding vehicle 601 may not be the three-dimensional data or surrounding condition detection results generated by the surrounding vehicle, but rather accurate feature point information of surrounding vehicle 601 itself. Own vehicle 600 uses this feature point information of surrounding vehicle 601 to correct the feature point information of the preceding vehicle within the point cloud data obtained by own vehicle 600. This improves the matching accuracy of own vehicle 600's position estimation.

[0513] Furthermore, the feature point information of the preceding vehicle is three-dimensional point information composed of color information and coordinate information, for example. Therefore, even if the sensor of the own vehicle 600 is a laser sensor or a stereo camera, the feature point information of the preceding vehicle can be used regardless of its type.

[0514] Furthermore, vehicle 600 is not limited by transmission time and can use SWLD point cloud data to calculate the accuracy of its own position estimate. For example, if vehicle 600's sensor is an imaging device such as a stereo camera, it detects two-dimensional points in the image captured by the camera and uses these two-dimensional points to estimate its own position. Furthermore, while estimating its own position, vehicle 600 also generates point cloud data of surrounding objects. Vehicle 600 re-projects the SWLD three-dimensional points onto the two-dimensional image and evaluates the accuracy of its own position estimate based on the error between the detected points on the two-dimensional image and the re-projected points.

[0515] In addition, when the sensor of the own vehicle 600 is a laser sensor such as LIDAR, the own vehicle 600 evaluates the accuracy of its own position estimation based on the error calculated by the iterative closest point algorithm (Iterative Closest Point) using the SWLD of the produced point cloud data and the SWLD of the three-dimensional map.

[0516] Furthermore, when the communication status via a base station or server such as 5G is poor, the own vehicle 600 can obtain a three-dimensional map from the surrounding vehicles 601.

[0517] Furthermore, vehicle-to-vehicle communication can be used to obtain distant information that cannot be obtained from vehicles surrounding the vehicle 600. For example, information about a recent traffic accident several hundred meters or kilometers ahead can be obtained by communicating with oncoming vehicles or by receiving information from surrounding vehicles. The data transmitted in this manner is meta-information above the dynamic three-dimensional map.

[0518] Furthermore, the detection results of the surrounding conditions and the information detected by the vehicle 600 can be presented to the user through a user interface, for example, by overlaying the navigation screen or the windshield.

[0519] Furthermore, it is also possible that, without supporting autonomous driving, a vehicle with cruise control can track surrounding vehicles when it detects surrounding vehicles traveling in autonomous driving mode.

[0520] Furthermore, when a three-dimensional map cannot be obtained or the own position cannot be estimated due to too many blocked areas, the own vehicle 600 can switch the working mode from the automatic driving mode to the tracking mode of the surrounding vehicles.

[0521] Furthermore, the tracked vehicle can warn the user of the tracked situation, and a user interface can be provided in which the user can specify whether to allow tracking. At this point, advertisements can be displayed on the tracked vehicle, and payment rewards can be provided on the tracked side.

[0522] Furthermore, while the information transmitted is basically SWLD data in the form of three-dimensional data, it may also be information corresponding to a request setting set for the vehicle 600 or a public setting set for a preceding vehicle. For example, the information transmitted may be SWLD data in the form of dense point cloud data, or it may be the detection results of surrounding conditions by a preceding vehicle, or braking information of the preceding vehicle.

[0523] The vehicle 600 then receives the WLD, visualizes the WLD's 3D data, and displays the visualized 3D data to the driver using a GUI. In this case, the vehicle 600 can use different colors to display information, allowing the user to distinguish between the point cloud data generated by the vehicle 600 and the received point cloud data.

[0524] In addition, when the own vehicle 600 uses the GUI to prompt the driver with the information detected by the own vehicle 600 and the detection results of the surrounding vehicles 601, the information is prompted by color differentiation, etc. in a way that the user can distinguish the information detected by the own vehicle 600 and the received detection results.

[0525] As described above, in the three-dimensional data creation device 620 according to this embodiment, the three-dimensional data creation unit 621 creates first three-dimensional data 632 based on sensor information 631 detected by the sensor. The receiving unit 624 receives encoded three-dimensional data 634 obtained by encoding the second three-dimensional data 635. The decoding unit 625 decodes the received encoded three-dimensional data 634 to obtain the second three-dimensional data 635. The synthesis unit 626 synthesizes the first three-dimensional data 632 and the second three-dimensional data 635 to create third three-dimensional data 636.

[0526] Thus, the three-dimensional data creation device 620 can create detailed third three-dimensional data 636 using the created first three-dimensional data 632 and the received second three-dimensional data 635 .

[0527] Furthermore, the synthesis unit 626 can synthesize the first three-dimensional data 632 and the second three-dimensional data 635 to create third three-dimensional data 636 having a higher density than the first three-dimensional data 632 and the second three-dimensional data 635 .

[0528] Furthermore, the second three-dimensional data 635 (for example, SWLD) is three-dimensional data generated by extracting data having a feature value equal to or greater than a threshold value from the fourth three-dimensional data (for example, WLD).

[0529] As a result, the three-dimensional data creation device 620 can reduce the amount of three-dimensional data to be transmitted.

[0530] Furthermore, the three-dimensional data creation device 620 further includes a search unit 623 that searches for a transmission device that is a transmission source of the encoded three-dimensional data 634. The reception unit 624 receives the encoded three-dimensional data 634 from the searched transmission device.

[0531] With this, the three-dimensional data creation device 620 can identify the transmission device holding the required three-dimensional data by searching, for example.

[0532] The three-dimensional data creation device further includes a request range determination unit 622 that determines a request range, which is the range of the three-dimensional space for which three-dimensional data is requested. The search unit 623 transmits request range information 633 indicating the request range to the transmission device. The second three-dimensional data 635 includes the three-dimensional data within the request range.

[0533] As a result, the three-dimensional data creation device 620 can not only receive the required three-dimensional data but also reduce the amount of three-dimensional data to be transmitted.

[0534] Furthermore, the request range determination unit 622 determines a spatial range including the blocked area 604 that cannot be detected by the sensor as the request range.

[0535] Furthermore, in the three-dimensional data transmitting device 640 according to this embodiment, a three-dimensional data generating unit 641 generates fifth three-dimensional data 652 based on sensor information 651 detected by a sensor. An extracting unit 643 extracts a portion of the fifth three-dimensional data 652 to generate sixth three-dimensional data 654. An encoding unit 644 encodes the sixth three-dimensional data 654 to generate encoded three-dimensional data 634. A transmitting unit 645 transmits the encoded three-dimensional data 634.

[0536] As a result, the three-dimensional data transmitting device 640 can not only transmit the three-dimensional data produced by itself to other devices, but also reduce the amount of the transmitted three-dimensional data.

[0537] Furthermore, the three-dimensional data generating unit 641 generates seventh three-dimensional data (e.g., WLD) based on sensor information 651 detected by the sensor, and generates fifth three-dimensional data 652 (e.g., SWLD) by extracting data having a feature value greater than a threshold value from the seventh three-dimensional data.

[0538] Accordingly, the three-dimensional data transmitting apparatus 640 can reduce the amount of three-dimensional data to be transmitted.

[0539] The three-dimensional data transmitting device 640 further includes a receiving unit 642 that receives, from the receiving device, request range information 633 indicating a requested range, the range of the three-dimensional space for which three-dimensional data is requested. An extraction unit 643 extracts the three-dimensional data within the requested range from the fifth three-dimensional data 652 to create sixth three-dimensional data 654. A transmission unit 645 transmits the encoded three-dimensional data 634 to the receiving device.

[0540] Accordingly, the three-dimensional data transmitting apparatus 640 can reduce the amount of three-dimensional data to be transmitted.

[0541] (Implementation 4)

[0542] In this embodiment, an operation related to abnormal situations in self-position estimation based on a three-dimensional map will be described.

[0543] Applications such as autonomous driving of vehicles and autonomous movement of mobile objects such as robots and drones are expected to expand in the future. One example of a method for achieving such autonomous movement is a method in which a mobile object estimates its own position within a three-dimensional map (self-position estimation) and drives according to the map.

[0544] Self-position estimation is achieved by matching the three-dimensional map with the three-dimensional information around the own vehicle obtained by sensors such as the rangefinder (LiDAR, etc.) or stereo camera installed on the own vehicle (hereinafter referred to as the own vehicle detection three-dimensional data), and estimating the own vehicle position within the three-dimensional map.

[0545] 3D maps, such as those offered by HERE's HD maps, are not just 3D point clouds; they also include 2D map data such as road and intersection shape information, as well as real-time information such as traffic jams and accidents. 3D maps are constructed from multiple layers of data, including 3D and 2D data and metadata that changes over time. Devices can retrieve only the data they need, or they can reference only the data they need.

[0546] The point cloud data may be the aforementioned SWLD, or may include point group data that is not a feature point. Furthermore, the transmission and reception of the point cloud data is basically performed in one or more random access units.

[0547] The following methods can be used to match a 3D map with the 3D data detected by the own vehicle. For example, the device compares the shapes of the point clusters in the respective point clouds and determines areas with high similarity between feature points as being at the same location. Furthermore, if the 3D map is constructed using SWLD, the device compares and matches the feature points that make up the SWLD with the 3D feature points extracted from the 3D data detected by the own vehicle.

[0548] To accurately estimate the vehicle's position, the following conditions (A) and (B) must be met: (A) a 3D map and 3D vehicle detection data must be available, and (B) their accuracy must meet a predetermined standard. However, in the following exceptional circumstances, either (A) or (B) may not be met.

[0549] (1) A three-dimensional map cannot be obtained through the communication path.

[0550] (2) There is no three-dimensional map, or the obtained three-dimensional map is damaged.

[0551] (3) The sensor of the own vehicle fails, or due to bad weather, the accuracy of the generated three-dimensional data of the own vehicle detection is insufficient.

[0552] The following describes the operations for dealing with these abnormal situations. Although the following operation is described using a vehicle as an example, the following method can also be applied to all moving objects that move autonomously, such as robots and drones.

[0553] The following describes the configuration and operation of the three-dimensional information processing device according to this embodiment for detecting abnormalities in three-dimensional data corresponding to a three-dimensional map or a vehicle. Figure 34 This is a block diagram showing a configuration example of a three-dimensional information processing device 700 according to this embodiment. Figure 35 3D information processing method performed by the 3D information processing apparatus 700 is a flowchart.

[0554] The three-dimensional information processing device 700 is mounted on a mobile object such as a motor vehicle. Figure 34 As shown, the three-dimensional information processing device 700 includes a three-dimensional map acquisition unit 701 , a vehicle detection data acquisition unit 702 , an abnormality determination unit 703 , a countermeasure operation determination unit 704 , and an operation control unit 705 .

[0555] The 3D information processing device 700 may also include a camera that captures 2D images, or a sensor (not shown) that uses ultrasonic or laser-based 1D data to detect structures or moving objects around the vehicle. Furthermore, the 3D information processing device 700 may include a communication unit (not shown) that generates a 3D map via a mobile communication network such as 4G or 5G, or between vehicles or between roads and vehicles.

[0556] like Figure 35 As shown, the three-dimensional map obtaining unit 701 obtains a three-dimensional map 711 near the driving route (S701). For example, the three-dimensional map obtaining unit 701 obtains the three-dimensional map 711 through a mobile communication network, vehicle-to-vehicle communication, or road-to-vehicle communication.

[0557] Next, the vehicle detection data acquisition unit 702 acquires vehicle detection three-dimensional data 712 based on the sensor information ( S702 ). For example, the vehicle detection data acquisition unit 702 generates vehicle detection three-dimensional data 712 based on sensor information acquired by sensors included in the vehicle.

[0558] Next, the abnormality determination unit 703 detects an abnormality by performing a predetermined check on at least one of the obtained three-dimensional map 711 and the vehicle's own three-dimensional detection data 712 (S703). In other words, the abnormality determination unit 703 determines whether at least one of the obtained three-dimensional map 711 and the vehicle's own three-dimensional detection data 712 is abnormal.

[0559] In step S703, if an abnormality is detected ("Yes" in S704), the response action determination unit 704 determines the response action for the abnormality (S705). Next, the action control unit 705 controls the operation of each processing unit required to implement the response action, such as the three-dimensional map acquisition unit 701 (S706).

[0560] If no abnormality is detected in step S703 (No in step S704 ), the three-dimensional information processing apparatus 700 ends the processing.

[0561] The 3D information processing device 700 estimates the vehicle's own position using the 3D map 711 and the own vehicle detection 3D data 712. The 3D information processing device 700 then uses the result of the own position estimation to enable the vehicle to autonomously drive.

[0562] Based on this, the three-dimensional information processing device 700 obtains map data (three-dimensional map 711) including the first three-dimensional position information via the communication path. For example, the first three-dimensional position information is encoded in units of partial spaces having three-dimensional coordinate information. The first three-dimensional position information includes multiple random access units, each of which is a collection of one or more partial spaces and can be independently decoded. For example, the first three-dimensional position information is data (SWLD) encoded at feature points whose three-dimensional feature quantities exceed a predetermined threshold.

[0563] The three-dimensional information processing device 700 generates second three-dimensional position information (vehicle detection three-dimensional data 712) based on the information detected by the sensor. The three-dimensional information processing device 700 then performs abnormality determination processing on the first three-dimensional position information or the second three-dimensional position information to determine whether the first three-dimensional position information or the second three-dimensional position information is abnormal.

[0564] When the three-dimensional information processing device 700 determines that the first three-dimensional position information or the second three-dimensional position information is abnormal, it determines a countermeasure for the abnormality and then executes control required for executing the countermeasure.

[0565] With this, the three-dimensional information processing device 700 can detect abnormality in the first three-dimensional position information or the second three-dimensional position information and perform a corresponding operation.

[0566] The following describes the operation to be performed in abnormal situation 1, that is, in the case where the three-dimensional map 711 cannot be obtained through communication.

[0567] Three-dimensional map 711 is required for estimating the vehicle's own position. However, if the vehicle does not already have three-dimensional map 711 corresponding to the route to the destination, it must obtain three-dimensional map 711 through communication. However, due to congestion on the communication path or poor radio reception, the vehicle may not be able to obtain three-dimensional map 711 along the route.

[0568] Abnormality determination unit 703 checks whether three-dimensional map 711 has been obtained for all sections on the route to the destination, or for sections within a predetermined range from the current location. If it cannot be obtained, it determines that abnormality 1 has occurred. Specifically, abnormality determination unit 703 determines whether three-dimensional map 711 (first three-dimensional location information) can be obtained via the communication path. If it cannot be obtained via the communication path, it determines that three-dimensional map 711 is abnormal.

[0569] When it is determined that the abnormal situation is 1, the countermeasure decision unit 704 selects one of the following two countermeasures: (1) continuing the self-position estimation, and (2) stopping the self-position estimation.

[0570] First, (1) a specific example of the operation to be performed when the self-position estimation is continued will be described. When the self-position estimation is continued, a three-dimensional map 711 on the route to the destination is required.

[0571] For example, the vehicle determines a location within the range of the already acquired three-dimensional map 711 where a communication path is available, moves to that location, and acquires the three-dimensional map 711. In this case, the vehicle can acquire all three-dimensional maps 711 up to the destination, or it can acquire three-dimensional maps 711 per random access unit within the maximum size that can be stored in the vehicle's memory, HDD, or other storage device.

[0572] Alternatively, the vehicle may obtain the communication status of the route. If the communication status is predicted to deteriorate along the route, the vehicle may obtain a three-dimensional map 711 for that section before reaching the section with poor communication status, or obtain a three-dimensional map 711 for the maximum range that can be obtained. In other words, the three-dimensional information processing device 700 predicts whether the vehicle will enter an area with poor communication status. If the three-dimensional information processing device 700 predicts that the vehicle will enter an area with poor communication status, the three-dimensional map 711 is obtained before the vehicle enters the area.

[0573] Alternatively, the vehicle may identify a random access unit of the three-dimensional map 711 that is narrower than normal and constitutes the minimum required for estimating its own position on the route, and receive the identified random access unit. In other words, if the three-dimensional information processing device 700 cannot obtain the three-dimensional map 711 (first three-dimensional position information) via the communication path, it can obtain third three-dimensional position information with a narrower range than the first three-dimensional position information via the communication path.

[0574] Furthermore, when a vehicle cannot access the distribution server of the three-dimensional map 711, it can obtain the three-dimensional map 711 from other vehicles or other mobile bodies traveling around its own vehicle. At this time, the other vehicles or other mobile bodies are mobile bodies that have already obtained the three-dimensional map 711 on the path to the destination and can communicate with its own vehicle.

[0575] Next, a specific example of the operation in the case of (2) stopping the own position estimation will be described. In this case, the three-dimensional map 711 on the route to the destination is not necessary.

[0576] For example, the vehicle will notify the driver that functions such as autonomous driving based on its own position estimation cannot continue to be executed, and the operating mode will be transferred to a manual operation mode performed by the driver.

[0577] While the level of self-position estimation may differ depending on the presence of a human, autonomous driving is typically performed. Furthermore, the results of self-position estimation may also be used for navigation purposes, such as when a human is driving. Therefore, the results of self-position estimation are not necessarily used for autonomous driving.

[0578] Furthermore, when the vehicle cannot utilize commonly used communication paths such as mobile communication networks such as 4G or 5G, it is confirmed whether the three-dimensional map 711 can be obtained via other communication paths such as Wi-Fi (registered trademark) or millimeter wave communication between the road and the vehicle, or communication between vehicles, and the communication path used can be switched to a communication path that can obtain the three-dimensional map 711.

[0579] Furthermore, even if the vehicle cannot obtain the three-dimensional map 711, it can obtain a two-dimensional map and continue autonomous driving using the two-dimensional map and the vehicle's own three-dimensional detection data 712. Specifically, even if the three-dimensional information processing device 700 cannot obtain the three-dimensional map 711 via the communication path, it can obtain map data (a two-dimensional map) including two-dimensional position information via the communication path and estimate the vehicle's own position using the two-dimensional position information and the vehicle's own three-dimensional detection data 712.

[0580] Specifically, the vehicle uses the two-dimensional map and the vehicle's own three-dimensional detection data 712 to estimate its own position, and uses the vehicle's own three-dimensional detection data 712 to detect surrounding vehicles, pedestrians, and obstacles.

[0581] Here, map data such as HD maps can include not only 3D maps 711 composed of 3D point clouds, but also 2D map data (2D maps), simplified map data that extracts characteristic information such as road shapes and intersections from 2D map data, and metadata indicating real-time information such as traffic jams, accidents, and construction. For example, map data has a layered structure with 3D data (3D maps 711), 2D data (2D maps), and metadata arranged in order from the lower layer.

[0582] Here, two-dimensional data is smaller than three-dimensional data. Therefore, even in poor communication conditions, the vehicle can obtain a two-dimensional map. Furthermore, the vehicle can uniformly obtain a large-scale two-dimensional map in areas with good communication conditions. Therefore, if the communication path is poor and it is difficult to obtain three-dimensional map 711, the vehicle can receive a layer containing a two-dimensional map instead of the three-dimensional map 711. Furthermore, because metadata is small in size, the vehicle can, for example, receive metadata constantly regardless of communication conditions.

[0583] Among the methods of estimating the own position using a two-dimensional map and the own vehicle detection three-dimensional data 712 , there are, for example, the following two methods.

[0584] The first method is a method of matching two-dimensional feature quantities. Specifically, the vehicle extracts two-dimensional feature quantities from the vehicle detection three-dimensional data 712 and matches the extracted two-dimensional feature quantities with the two-dimensional map.

[0585] For example, the vehicle projects its own vehicle detection three-dimensional data 712 onto the same plane as the two-dimensional map and matches the obtained two-dimensional data with the two-dimensional map. The matching is performed using two-dimensional image features extracted from both.

[0586] If 3D map 711 includes SWLD, it can simultaneously store 3D feature quantities for feature points within 3D space, as well as 2D feature quantities within the same plane as the 2D map. For example, identification information can be assigned to the 2D feature quantities. Alternatively, the 2D feature quantities can be stored in a separate layer from the 3D data and the 2D map, allowing the vehicle to obtain the 2D feature quantity data simultaneously with the 2D map.

[0587] When a two-dimensional map represents information on positions of white lines, guardrails, and buildings on the road at different heights from the ground (not in the same plane) on the same map, the vehicle extracts feature values ​​from multiple height data of the three-dimensional data 712 detected by its own vehicle.

[0588] Furthermore, information indicating the correspondence between the feature points in the two-dimensional map and the feature points in the three-dimensional map 711 may be stored as meta-information of the map data.

[0589] The second method is to match three-dimensional feature quantities. Specifically, the vehicle obtains three-dimensional feature quantities corresponding to feature points in the two-dimensional map and matches the obtained three-dimensional feature quantities with the three-dimensional feature quantities of the vehicle's detected three-dimensional data 712.

[0590] Specifically, the three-dimensional feature values ​​corresponding to the feature points of the two-dimensional map are stored in the map data. When the vehicle obtains the two-dimensional map, it also obtains the three-dimensional feature values. In addition, when the three-dimensional map 711 includes SWLD, by assigning information for identifying the feature points of the SWLD that correspond to the feature points of the two-dimensional map, the vehicle can determine the location corresponding to the two-dimensional map based on the identification information. Figure 1 In addition, in this case, since it is sufficient to express the two-dimensional position, the amount of data can be reduced compared to the case of expressing the three-dimensional position.

[0591] Furthermore, when estimating its own position using a two-dimensional map, the accuracy of its own position estimation is lower than that of the three-dimensional map 711. Therefore, the vehicle determines whether to continue autonomous driving even when the estimation accuracy is reduced, and only continues autonomous driving if it is determined that it can continue.

[0592] Whether autonomous driving can continue is influenced by factors such as whether the vehicle is traveling on an urban area or a highway or other road with few vehicles or pedestrians, as well as the driving environment, such as road width and road clutter (vehicle or pedestrian density). Furthermore, markers for sensors such as cameras to identify can be placed on business premises, on streets, or inside buildings. In these specific areas, the markers can be accurately identified by two-dimensional sensors, for example, by including the marker's location information in a two-dimensional map, enabling highly accurate self-position estimation.

[0593] Furthermore, by including identification information indicating whether each area is a specific area within the map, the vehicle can determine whether it is within a specific area. If the vehicle is within the specific area, the vehicle determines to continue autonomous driving. In this way, the vehicle can determine whether to continue autonomous driving based on the accuracy of its own position estimation when using a two-dimensional map or the vehicle's driving environment.

[0594] In this way, the three-dimensional information processing device 700 can determine whether to perform automatic driving of the vehicle based on the vehicle's driving environment (the moving environment of the moving body), which is the result of estimating the vehicle's own position using a two-dimensional map and the vehicle's own detection three-dimensional data 712.

[0595] Furthermore, the vehicle may not determine whether to continue autonomous driving, but may instead switch the autonomous driving level (mode) based on the accuracy of its own position estimate or the vehicle's driving environment. Switching the autonomous driving level (mode) here may include, for example, limiting speed, increasing the amount of driver input (lowering the autonomous driving level), obtaining driving information from a preceding vehicle and switching the mode based on that information, or obtaining driving information from a vehicle set to the same destination and using that information to switch the autonomous driving mode.

[0596] Furthermore, the map can include information, associated with location information, indicating the recommended level of autonomous driving when estimating the vehicle's position using a two-dimensional map. The recommended level can be metadata that changes dynamically based on factors such as traffic volume. This eliminates the need for vehicles to determine the level based on the surrounding environment and can instead determine the level solely based on information within the map. Furthermore, by having multiple vehicles referencing the same map, the level of autonomous driving for each vehicle can be maintained consistently. Furthermore, the recommended level can be a mandatory level rather than a recommendation.

[0597] Furthermore, the vehicle can switch the level of autonomous driving depending on whether there is a driver (manned or unmanned). For example, the vehicle can reduce the level of autonomous driving when there is a driver and stop when there is no one. The vehicle determines where it can safely stop by identifying pedestrians, vehicles, and traffic signs in the surrounding area. Alternatively, the map may include location information showing where the vehicle can safely stop, and the vehicle can refer to this location information to determine where it can safely stop.

[0598] Next, the operation to be performed in abnormal situation 2, that is, when the three-dimensional map 711 does not exist or the obtained three-dimensional map 711 is damaged, will be described.

[0599] Abnormality determination unit 703 determines which of the following situations (1) and (2) it falls under, and determines that abnormality 2 is the case if (1) three-dimensional map 711 for part or all of the sections on the route to the destination is not available on the distribution server that is the access destination and cannot be obtained, or (2) part or all of the obtained three-dimensional map 711 is damaged. In other words, abnormality determination unit 703 determines whether the data of three-dimensional map 711 is complete. If the data of three-dimensional map 711 is incomplete, it determines that three-dimensional map 711 is abnormal.

[0600] When it is determined to be abnormal situation 2, the following countermeasures are performed. First, an example of the countermeasures when (1) the three-dimensional map 711 cannot be obtained will be described.

[0601] For example, the vehicle sets a route that does not pass through a section without the three-dimensional map 711 .

[0602] If the vehicle cannot set an alternative route due to reasons such as the absence of an alternative route or the significant increase in distance required for an alternative route, the vehicle sets a route that includes a section where the three-dimensional map 711 is not available. Furthermore, the vehicle notifies the driver of a change in driving mode in this section, thereby switching the driving mode to manual operation mode.

[0603] (2) When part or all of the acquired three-dimensional map 711 is destroyed, the following countermeasures are performed.

[0604] The vehicle identifies the damaged area in three-dimensional map 711, requests data about the damaged area through communication, obtains the data, and updates three-dimensional map 711 using the obtained data. The vehicle can specify the damaged area using positional information such as absolute or relative coordinates in three-dimensional map 711, or by index number of the random access unit that constitutes the damaged area. In this case, the vehicle replaces the random access unit containing the damaged area with the obtained random access unit.

[0605] Next, an explanation will be given of the operation to be performed in abnormal situation 3, that is, when the sensors of the own vehicle cannot generate the own vehicle detection three-dimensional data 712 due to a malfunction or bad weather.

[0606] The abnormality determination unit 703 checks whether the generation error of the own vehicle detection three-dimensional data 712 is within the allowable range. If not, it determines that abnormality 3 has occurred. Specifically, the abnormality determination unit 703 determines whether the generation accuracy of the own vehicle detection three-dimensional data 712 is above a reference value. If the generation accuracy of the own vehicle detection three-dimensional data 712 is not above the reference value, the own vehicle detection three-dimensional data 712 is determined to be abnormal.

[0607] As a method for confirming whether the generation error of the own vehicle detection three-dimensional data 712 is within the allowable range, the following method can be adopted.

[0608] The spatial resolution of the vehicle's detected three-dimensional data 712 during normal operation is predetermined based on the resolution in the depth and scanning directions of the vehicle's three-dimensional sensors, such as range finders and stereo cameras, or the density of the point cloud that can be generated. Furthermore, the vehicle obtains the spatial resolution of the three-dimensional map 711 based on metadata included in the three-dimensional map 711.

[0609] The vehicle uses the spatial resolution of both to estimate a baseline value for matching error when matching the vehicle's detected 3D data 712 with the 3D map 711 based on 3D feature quantities. Matching error can be achieved using statistics such as the error in the 3D feature quantity for each feature point, the average of the errors between multiple feature points, or the error in the spatial distance between multiple feature points. The allowable range for deviation from the baseline value is pre-set.

[0610] If the matching error between the vehicle detection three-dimensional data 712 generated before or during the vehicle starts traveling and the three-dimensional map 711 is not within the allowable range, it is determined to be abnormal situation 3.

[0611] Alternatively, the vehicle may use a test pattern with a known three-dimensional shape for accuracy inspection to obtain its own vehicle detection three-dimensional data 712 for a test pattern such as before driving begins, and determine whether it is an abnormal situation 3 based on whether the shape error is within an allowable range.

[0612] For example, the vehicle performs the above determination each time before starting a journey. Alternatively, the vehicle performs the above determination at regular intervals while driving, thereby obtaining a time-series variation in the matching error. If the matching error shows a tendency to increase, the vehicle may determine that an abnormality is occurring, even if the error is within the allowable range. Furthermore, if the vehicle can predict an abnormality based on the time-series variation, the user may be notified of the predicted abnormality by displaying a message urging an inspection or repair. Furthermore, by distinguishing between abnormalities caused by temporary factors such as bad weather and abnormalities due to sensor failures based on the time-series variation, the user may be notified of only abnormalities due to sensor failures.

[0613] Furthermore, when the vehicle is judged to be in abnormal situation 3, any one of the following three response actions is selected or selectively executed: (1) operating an emergency substitute sensor (rescue mode), (2) switching the operating mode, and (3) performing a three-dimensional sensor operation calibration.

[0614] First, (1) the case of operating an emergency substitute sensor will be described. The vehicle operates an emergency substitute sensor that is different from the three-dimensional sensor used during normal operation. Specifically, when the generation accuracy of the own vehicle detection three-dimensional data 712 is not above a reference value, the three-dimensional information processing device 700 generates own vehicle detection three-dimensional data 712 (fourth three-dimensional position information) based on information detected by the substitute sensor that is different from the normal sensor.

[0615] Specifically, when a vehicle uses multiple cameras or LiDARs to obtain 3D vehicle detection data 712, the vehicle identifies a malfunctioning sensor based on, for example, the direction in which the matching error in 3D vehicle detection data 712 exceeds the allowable range. The vehicle then activates a replacement sensor corresponding to the malfunctioning sensor.

[0616] The replacement sensor can be a three-dimensional sensor, a camera that obtains two-dimensional images, or a one-dimensional sensor such as ultrasound. If the replacement sensor is a sensor other than a three-dimensional sensor, the accuracy of the vehicle's position estimation may be reduced or even impossible. Therefore, the vehicle can switch the autonomous driving mode based on the type of replacement sensor.

[0617] For example, if the replacement sensor is a three-dimensional sensor, the vehicle continues in autonomous driving mode. Furthermore, if the replacement sensor is a two-dimensional sensor, the vehicle switches from fully autonomous driving to semi-autonomous driving mode, which requires human control. Furthermore, if the replacement sensor is a one-dimensional sensor, the vehicle switches to manual braking mode, which disables automatic braking control.

[0618] Furthermore, the vehicle can switch between autonomous driving modes depending on the driving environment. For example, if the vehicle uses a two-dimensional sensor instead of a 2D sensor, it can continue in fully autonomous driving mode on highways and switch to semi-autonomous driving mode in urban areas.

[0619] Furthermore, if the vehicle can continue to estimate its position without a replacement sensor and can obtain a sufficient number of feature points using only the sensors currently in operation, it can still do so. However, since it cannot detect specific directions, the vehicle will switch to semi-autonomous driving or manual operation.

[0620] Next, (2) the response to switching the operating mode is described. The vehicle switches the operating mode from the automatic driving mode to the manual operation mode. Alternatively, the vehicle may continue to drive automatically until it reaches a place such as a roadside where it can stop safely, and then stop. Furthermore, the vehicle may switch the operating mode to the manual operation mode after stopping. In this way, the three-dimensional information processing device 700 switches the automatic driving mode when the generation accuracy of the vehicle's own detection three-dimensional data 712 is not above the reference value.

[0621] Next, (3) will describe the response work for correcting the operation of the three-dimensional sensor. The vehicle determines the three-dimensional sensor that is malfunctioning based on the direction in which the matching error occurs, and calibrates the determined sensor. Specifically, when multiple LiDARs or cameras are used as sensors, a portion of the three-dimensional space reconstructed by each sensor overlaps. That is, the data of the overlapping portion is obtained by multiple sensors. The three-dimensional point group data obtained for the overlapping portion is different between the normal sensor and the malfunctioning sensor. Therefore, the vehicle performs LiDAR origin correction, or camera exposure, focus, etc., in a manner that enables the malfunctioning sensor to obtain the same three-dimensional point group data as the normal sensor, and adjusts the operation of the predetermined part.

[0622] After adjustment, if the matching error is within the allowable range, the vehicle continues the previous operating mode. Otherwise, if the matching accuracy is not within the allowable range after adjustment, the vehicle performs the above-mentioned response of (1) operating the emergency replacement sensor or (2) switching the operating mode.

[0623] In this manner, the three-dimensional information processing device 700 performs sensor operation calibration when the data generation accuracy of the own vehicle detection three-dimensional data 712 is not equal to or higher than the reference value.

[0624] The following describes how to select a countermeasure action. The countermeasure action can be selected by a user such as the driver, or it can be selected automatically by the vehicle without the user's intervention.

[0625] Furthermore, the vehicle can also switch control depending on whether a driver is aboard. For example, if a driver is aboard, the vehicle prioritizes switching to manual operation mode. Otherwise, if a driver is not aboard, the vehicle prioritizes stopping at a safe location.

[0626] The information indicating the stopping place may be included as meta-information in the three-dimensional map 711. Alternatively, the vehicle may send a response request for the stopping place to a service that manages the operation information of the autonomous driving, thereby obtaining the information indicating the stopping place.

[0627] Furthermore, when the vehicle is operating along a prescribed route, for example, the vehicle's operating mode can be shifted to a mode where an operator manages the vehicle's operations via a communication path. In particular, in vehicles operating in fully autonomous mode, the risk of a malfunction in the vehicle's position estimation function is high. Therefore, when a vehicle detects an abnormality, or if the detected abnormality cannot be corrected, it notifies the service managing operational information via a communication path of the abnormality. This service can notify nearby vehicles of the presence of an abnormal vehicle or issue instructions to vacate nearby parking spaces.

[0628] Furthermore, when an abnormal situation is detected, the vehicle can travel slower than usual.

[0629] If a self-driving vehicle used in a taxi-like ride-sharing service experiences an abnormality, it will notify the operation management center and stop at a safe location. Alternatively, the ride-sharing service can dispatch a replacement vehicle. Alternatively, the user of the ride-sharing service can drive the vehicle. In these situations, discounts on fares or special points can be offered.

[0630] Furthermore, in the method for dealing with abnormal situation 1, a method of estimating the own position based on a two-dimensional map has been described. However, the own position can also be estimated using a two-dimensional map in normal situations. Figure 36 This is a flowchart of the self-position estimation process in this case.

[0631] First, the vehicle obtains a three-dimensional map 711 of the vicinity of the driving route (S711). Next, the vehicle obtains its own vehicle detection three-dimensional data 712 based on sensor information (S712).

[0632] Next, the vehicle determines whether a three-dimensional map 711 is necessary for estimating its own position (S713). Specifically, the vehicle determines whether a three-dimensional map 711 is necessary based on the accuracy of its own position estimation when using a two-dimensional map, as well as the driving environment. For example, the same method as described above for abnormal situation 1 is used.

[0633] If the vehicle determines that 3D map 711 is not necessary ("No" in S714), it obtains a 2D map (S715). At this point, the vehicle can also obtain the additional information described in the handling method for abnormal situation 1. Furthermore, the vehicle can generate a 2D map based on 3D map 711. For example, the vehicle can extract an arbitrary plane from 3D map 711 to generate a 2D map.

[0634] Next, the vehicle estimates its own position using the vehicle detection three-dimensional data 712 and the two-dimensional map (S716). The method for estimating the own position using the two-dimensional map is the same as that described in the method for dealing with abnormal situation 1 above.

[0635] If the vehicle determines that the three-dimensional map 711 is necessary (Yes in S714 ), the vehicle obtains the three-dimensional map 711 ( S717 ) and estimates its own position using the vehicle detection three-dimensional data 712 and the three-dimensional map 711 ( S718 ).

[0636] Furthermore, the vehicle can switch between primarily using a two-dimensional map and a three-dimensional map 711 based on the speed of its own communication device or the conditions of the communication path. For example, while receiving three-dimensional map 711, the vehicle can pre-set the communication speed required for driving. If the communication speed is below the set value, the vehicle primarily uses the two-dimensional map. If the communication speed is greater than the set value, the vehicle primarily uses the three-dimensional map 711. Alternatively, the vehicle may not determine whether to switch between using a two-dimensional map and a three-dimensional map and primarily use the two-dimensional map.

[0637] The three-dimensional information processing device according to the embodiments of the present application has been described above, but the present application is not limited to these embodiments.

[0638] Furthermore, the processing units included in the three-dimensional information processing device according to the above-mentioned embodiment are typically implemented as LSIs, which are integrated circuits. These may be individually formed into a single chip, or a portion or all of them may be formed into a single chip.

[0639] Furthermore, integrated circuits are not limited to LSIs and can also be implemented using dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays) that can be programmed after LSI fabrication, or reconfigurable processors that can reconfigure the connections and settings of circuit cells within the LSI, can also be used.

[0640] Furthermore, in each of the above-described embodiments, each component may be formed by dedicated hardware, or may be implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0641] Furthermore, the present application can be implemented as a three-dimensional information processing method executed by a three-dimensional information processing device.

[0642] The functional block divisions shown in the block diagram are merely examples. Multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple blocks, and some functions can be moved to other functional blocks. Furthermore, the functions of multiple functional blocks with similar functions can be processed in parallel or in a time-sliced ​​manner by a single hardware or software.

[0643] Furthermore, the execution order of each step in the flowchart is an example given for the purpose of specifically explaining the present application, and may be an order other than the above. Furthermore, part of the above steps may be executed simultaneously (in parallel) with other steps.

[0644] The above description of one or more aspects of a three-dimensional information processing device is based on the embodiments. However, this application is not limited to these embodiments. Within the scope of this application, various variations that can be imagined by those skilled in the art are applied to this embodiment, as well as aspects that are obtained by combining components from different embodiments. All aspects are included within the scope of one or more aspects.

[0645] (Implementation 5)

[0646] Other application examples of the image processing methods and apparatuses shown in the above-mentioned embodiments and systems using these methods and apparatuses are described. The system can be applied to imaging systems that are progressing in intelligence and widening of target space, for example, the following (1) to (4): (1) surveillance systems using surveillance cameras installed in stores or factories, or police vehicle-mounted cameras, (2) traffic information systems using private cameras, vehicle-mounted cameras, or cameras installed on roads, (3) environmental survey or transmission systems using remotely operated or automatically controlled devices such as drones, and (4) content transmission and reception systems using cameras installed in entertainment facilities or stadiums, mobile cameras such as drones, or private cameras.

[0647] Figure 37 The configuration of the image information processing system ex100 in this embodiment is shown. In this embodiment, examples of preventing the occurrence of blind spots and prohibiting the imaging of specific areas are described.

[0648] Figure 37 The video information processing system ex100 shown includes a video information processing device ex101, a plurality of cameras ex102, and a video receiving device ex103. However, the video receiving device ex103 does not necessarily need to be included in the video information processing system ex100.

[0649] The image information processing device ex101 includes a storage unit ex111 and an analysis unit ex112. Each of the N cameras ex102 has the function of capturing an image and transmitting the captured image data to the image information processing device ex101. Furthermore, the camera ex102 may also have the function of displaying the captured image. Furthermore, the camera ex102 may encode the captured image signal using a coding method such as HEVC or H.264 and transmit it to the image information processing device ex101, or may transmit unencoded image data to the image information processing device ex101.

[0650] Here, each camera ex102 is a fixed camera such as a surveillance camera, a mobile camera mounted on an unmanned aerial vehicle (UAV) radio remote control or a vehicle, or a user camera owned by a user.

[0651] The mobile camera receives the instruction signal transmitted from the image information processing device ex101 and changes its own position or imaging direction according to the received instruction signal.

[0652] Before starting to shoot, the timing of the multiple cameras ex102 is calibrated using time information from a server or a reference camera. Furthermore, the spatial positions of the multiple cameras ex102 are calibrated based on how the object appears in the space being shot or their relative positions to the reference camera.

[0653] The storage unit ex111 included in the information processing device ex101 stores image data transmitted from the N cameras ex102.

[0654] The analyzing unit ex112 detects blind spots based on the image data stored in the storing unit ex111 and sends an instruction signal to the mobile camera to prevent the occurrence of blind spots. The mobile camera moves according to the instruction signal and continues shooting.

[0655] The analysis unit ex112 uses, for example, SfM (Structure from Motion) to detect blind spots. SfM is a method of reconstructing the three-dimensional shape of an object from multiple images captured at different locations. It is a widely known shape recovery technique that simultaneously estimates the object shape and camera position. For example, the analysis unit ex112 uses SfM to reconstruct the three-dimensional shape of a facility or stadium from the image data stored in the storage unit ex111, detecting areas that cannot be reconstructed as blind spots.

[0656] Furthermore, if the position and shooting direction of the camera ex102 are fixed and known, the analysis unit ex112 can use this known information to perform SfM. Furthermore, if the position and shooting direction of a mobile camera can be obtained using GPS and an angle sensor included in the mobile camera, the mobile camera can transmit this information to the analysis unit ex112, which can then use this information to perform SfM.

[0657] Furthermore, blind spot detection methods are not limited to the aforementioned SfM method. For example, the analysis unit ex112 can use information from a depth sensor, such as a photoelectric rangefinder, to determine the spatial distance of the captured object. Furthermore, the analysis unit ex112 can determine whether a pre-defined marker or specific object in the space is included in the image. If so, the analysis unit ex112 can detect information such as the camera position, shooting direction, and zoom factor based on the size of the marker or specific object. In this way, the analysis unit ex112 can detect blind spots using any method capable of detecting the shooting area of ​​each camera. Furthermore, the analysis unit ex112 can obtain information such as the relative position of multiple captured objects from image data or proximity sensors, and identify areas with a high probability of blind spots based on this obtained positional relationship.

[0658] Here, blind spots include not only parts of the desired area where no image is captured, but also areas where the image quality is inferior to other areas, or areas where the image quality does not meet the predetermined quality standards. The detection target can be appropriately set based on the system's configuration and purpose. For example, the required image quality can be set higher for specific objects within the captured space. Furthermore, the required image quality can be set lower for specific areas within the captured space, meaning that even if no image is captured, it is not considered a blind spot.

[0659] In addition, the above-mentioned image quality refers to various information related to the image, including the area occupied by the subject in the image (such as the number of pixels), or whether the focus is on the subject. This information or their combination can be used as a basis to determine whether it is a blind spot.

[0660] While the above description focuses on detecting areas that are actually blind spots, preventing blind spots requires more than just detecting those areas. For example, if there are multiple subjects, and at least some of them are moving, the presence of other subjects between a subject and the camera could create new blind spots. To address this, the analysis unit ex112 can detect the motion of the multiple subjects from captured image data, and estimate areas that could potentially become new blind spots based on the detected motion of the multiple subjects and the position information of the camera ex102. In this case, the image information processing device ex101 sends an instruction signal to the mobile camera, instructing it to capture the areas that could potentially become blind spots, thereby preventing blind spots.

[0661] Furthermore, if there are multiple mobile cameras, the image information processing device ex101 needs to select which mobile camera to send a command signal to in order to capture a blind spot or an area that could become a blind spot. Furthermore, if there are multiple mobile cameras and multiple blind spots or areas that could become blind spots, the image information processing device ex101 needs to determine which blind spot or area that could become a blind spot to capture for each of the multiple mobile cameras. For example, the image information processing device ex101 selects the mobile camera closest to the blind spot or area that could become a blind spot based on the blind spot or area that could become a blind spot and the position of each mobile camera at the time of capture. Furthermore, the image information processing device ex101 can also determine whether a new blind spot has occurred for each mobile camera even if it has not yet obtained image data currently being captured. It can then select a mobile camera that has been determined to have no blind spots, even if it has not yet obtained image data currently being captured.

[0662] With the above configuration, the image information processing device ex101 detects blind spots and sends an instruction signal to the mobile camera to prevent blind spots, thereby preventing the occurrence of blind spots.

[0663] (Variation 1)

[0664] While the above description describes an example of sending a command signal to a mobile camera to instruct movement, the command signal can also be a signal instructing the user to move the user's camera. For example, in response to the command signal, the user camera displays an instruction image to the user, indicating the direction to change the camera. Alternatively, the user camera can display an instruction image showing the movement path on a map to indicate the user's movement. Furthermore, to improve the quality of captured images, the user camera may not display detailed shooting instructions, such as shooting direction, angle, picture angle, image quality, and movement of the shooting area. Furthermore, if control is possible on the image information processing device ex101, the image information processing device ex101 can automatically control the feature quantities of the camera ex102 related to such shooting.

[0665] Here, the user camera is, for example, a smartphone, a tablet terminal, a wearable terminal, or an HMD (Head Mounted Display) held by a spectator in a stadium or a security guard in a facility.

[0666] Furthermore, the display terminal that displays the instruction image does not necessarily need to be the same as the user camera that captured the image data. For example, the user camera may transmit an instruction signal or instruction image to a display terminal that has previously established a connection with the user camera, causing the display terminal to display the instruction image. Alternatively, the information of the display terminal corresponding to the user camera may be pre-registered with the image information processing device ex101. In this case, the image information processing device ex101 directly transmits an instruction signal to the display terminal corresponding to the user camera, causing the display terminal to display the instruction image.

[0667] (Variation 2)

[0668] The analyzing unit ex112 can, for example, use SfM to restore the three-dimensional shape of a facility or stadium from the image data stored in the storage unit ex111, thereby generating a free-viewpoint image (3D reconstruction data). This free-viewpoint image is stored in the storage unit ex111. The image information processing device ex101 reads image data corresponding to the field of view information (and / or viewpoint information) transmitted from the image receiving device ex103 from the storage unit ex111 and transmits it to the image receiving device ex103. The image receiving device ex103 can be one of multiple cameras.

[0669] (Variation 3)

[0670] The image information processing device ex101 can also detect a prohibited area. In this case, the analyzing unit ex112 analyzes the captured image and sends a prohibition signal to the mobile camera if it captures the prohibited area. The mobile camera stops capturing images while receiving the prohibition signal.

[0671] The analysis unit ex112 determines whether a mobile camera is capturing a predetermined prohibited shooting area within the space by, for example, mapping the three-dimensional virtual space restored using SfM with the captured image. Alternatively, the analysis unit ex112 uses a marker or a characteristic object placed within the space as a trigger to determine whether the mobile camera is capturing a prohibited shooting area. Examples of prohibited shooting areas include restrooms within a facility or a stadium.

[0672] Furthermore, when the user's camera captures a photographing prohibited area, the user's camera may display a message on a display connected wirelessly or wired, or output sound or voice from a speaker or earphone to notify the user that the current location is a photographing prohibited area.

[0673] For example, as the above-mentioned message, it is displayed that the direction in which the camera is currently facing is prohibited from shooting. Alternatively, the shooting prohibited area and the current shooting area are shown on the displayed map. In addition, regarding the resumption of shooting, it is automatically executed when, for example, no shooting prohibition signal is output. Alternatively, if the shooting prohibition signal is not output and the user performs an operation to restart shooting, shooting can be restarted. In addition, if shooting is stopped and restarted multiple times in a short period of time, calibration can be performed again. Alternatively, a notification can be issued to allow the user to confirm the current location or urge the user to move.

[0674] Furthermore, in the case of special missions such as police, a verification password or fingerprint authentication that disables these functions can be used for recording. Moreover, even in such cases, when images of prohibited areas are displayed externally or saved, image processing such as mosaicking can be automatically applied.

[0675] With the above configuration, the image information processing device ex101 determines whether imaging is prohibited and notifies the user to stop imaging, thereby prohibiting imaging of a certain area.

[0676] (Variation 4)

[0677] To construct a three-dimensional virtual space from images, it is necessary to collect images from multiple viewpoints. Therefore, the image information processing system ex100 offers incentives to users who share captured images. For example, the image information processing device ex101 can distribute the images for free or at a discount, award monetary points that can be used in online or physical stores or within games, or award points with non-monetary value, such as social status within virtual spaces like games. Furthermore, the image information processing device ex101 awards special high points to users who share images with valuable fields of view (and / or viewpoints), such as those with high requests.

[0678] (Variant 5)

[0679] The image information processing device ex101 can send additional information to the user camera based on the analysis results of the analysis unit ex112. In this case, the user camera overlays the additional information onto the captured image and displays it on the screen. For example, in the case of a game in a stadium, player information such as player names and heights can be associated with each player in the image, displaying the player's name or facial image. Furthermore, the image information processing device ex101 can extract the additional information based on a partial or complete area of ​​the image data by searching the internet. The camera ex102 can receive this additional information via short-range wireless communication such as Bluetooth (registered trademark) or visible light communication using lighting in a stadium, etc., and map the received additional information onto the image data. The camera ex102 can perform this mapping based on certain rules, such as a table stored in a storage unit connected to the camera ex102 via wired or wireless communication. The table can also be a table showing the correspondence between information obtained through visible light communication technology and the additional information. Alternatively, the mapping can be performed using an internet search and the most accurate combination.

[0680] Furthermore, in the surveillance system, information on, for example, persons requiring attention can be superimposed on user cameras held by security guards within the facility, thereby enabling an improvement in the accuracy of the surveillance system to be expected.

[0681] (Variation 6)

[0682] The analysis unit ex112 can determine which area within the facility or stadium the user's camera is capturing by matching the free viewpoint image with the image captured by the user's camera. The method for determining the capturing area is not limited to this method; the various capturing area determination methods described in the aforementioned embodiments or other capturing area determination methods may also be used.

[0683] The image information processing device ex101 transmits the past image to the user camera based on the analysis result of the analyzing unit ex112. The user camera superimposes the past image on the captured image or replaces the captured image with the past image and displays the result on the screen.

[0684] For example, during halftime, the highlights of the first half can be displayed as past videos. This allows users to enjoy the highlights of the first half as images in the direction they are looking. Furthermore, past videos are not limited to the highlights of the first half; they can also be highlights from past games played at the same stadium. Furthermore, the timing for distributing past videos by the image information processing device ex101 is not limited to halftime; for example, it can be after the game or during the game itself. Especially during games, based on the analysis results of the analysis unit ex112, the image information processing device ex101 can distribute important scenes that the user has missed. Furthermore, the image information processing device ex101 can distribute past videos only upon user request, or it can issue a message allowing distribution of past videos before distributing them.

[0685] (Variant 7)

[0686] The image information processing device ex101 may also transmit advertising information to the user's camera based on the analysis result of the analyzing unit ex112. The user's camera may superimpose the advertising information on the captured image and display it on the screen.

[0687] For example, advertising information can be distributed before distributing past video footage during halftime or after a game, as described in Modification 6. This allows the distributor to receive advertising fees from the advertiser, thereby providing users with inexpensive or free video distribution services. Furthermore, the video information processing device ex101 can send a message allowing the distribution of advertising information before distributing the advertising information, provide free service only if the user views the advertisement, or offer a cheaper service than if the user does not view the advertisement.

[0688] Furthermore, when a user clicks "Buy Now" in response to an advertisement, the purchased drinks are delivered to their seat by a service representative or the venue's automated delivery system, based on the user's location, using this system or arbitrary location information. Payment can be made directly to the service representative or using credit card information pre-set in the mobile terminal's app. Furthermore, advertisements can include links to e-commerce websites, enabling standard online shopping options such as home delivery.

[0689] (Variation 8)

[0690] The image receiving device ex103 may also be a camera ex102 (user camera). In this case, the analyzing unit ex112 determines which area within the facility or stadium the user camera is capturing by matching the free viewpoint image with the image captured by the user camera. The method for determining the captured area is not limited to this.

[0691] For example, when a user swipes in the direction of an arrow displayed on the screen, the user's camera generates viewpoint information indicating that the viewpoint has moved in that direction. The image information processing device ex101 reads from the storage unit ex111 image data capturing the area where the user's camera's shooting area has moved by the viewpoint information, as determined by the analysis unit ex112, and begins transmitting this image data to the user's camera. The user's camera then displays not only the captured image but also the image delivered by the image information processing device ex101.

[0692] As described above, users within a facility or stadium can view images from their preferred viewpoint simply by sliding the screen. For example, a spectator watching a game from third base at a baseball stadium can view images from first base. Furthermore, in a surveillance system, security guards within the facility can simply swipe the screen to view the viewpoint they want to confirm or to view images requested from a control center. This ability to change viewpoints as needed promises enhanced surveillance system accuracy.

[0693] Furthermore, the distribution of images to users within a facility or stadium is effective even when, for example, an obstruction exists between the user's camera and the subject, obstructing visibility. In this case, the user's camera can switch from the captured image to the distributed image from the image information processing device ex101, displaying a portion of the user's camera's shooting area that includes the obstruction, or it can switch the entire screen from the captured image to the distributed image. Furthermore, the user's camera can synthesize the captured image and the distributed image to display an image that shows the subject through the obstruction. This configuration allows users to view images distributed from the image information processing device ex101 even when the subject is obstructed from their position, thereby mitigating the effects of the obstruction.

[0694] Furthermore, in the case where the distributed image is displayed as an image of an area that cannot be seen due to an obstacle, the input processing performed by the user, such as the sliding screen described above, may perform display switching control that is different from the display switching control according to the input processing. For example, based on information about the movement and shooting direction of the user's camera, and previously obtained position information of the obstacle, when it is determined that the shooting area contains an obstacle, the display switching from the captured image to the distributed image may be automatically performed. Furthermore, when it is determined through analysis of the captured image data that an obstacle is reflected instead of the shooting object, the display switching from the captured image to the distributed image may be automatically performed. Furthermore, when the area of ​​the obstacle contained in the captured image (e.g., the number of pixels) exceeds a specified threshold, or when the ratio of the area of ​​the obstacle to the area of ​​the shooting object exceeds a specified ratio, the display switching from the captured image to the distributed image may be automatically performed.

[0695] Furthermore, the display of the captured image can be switched to the display of the distributed image, and vice versa, in accordance with the user's input processing.

[0696] (Variant 9)

[0697] The speed of transmitting the image data to the image information processing device ex101 may be instructed based on the importance of the image data captured by each camera ex102.

[0698] In this case, the analyzing unit ex112 determines the importance of the image data stored in the storage unit ex111 or the camera ex102 that captured the image data. The importance determination is performed based on, for example, the number of people or moving objects included in the image, the image quality of the image data, or a combination thereof.

[0699] Furthermore, the importance of image data can be determined based on the location of the camera ex102 that captured the image data or the area in which the image data was captured. For example, if multiple other cameras ex102 are currently capturing images near the target camera ex102, the importance of the image data captured by the target camera ex102 will be lowered. Furthermore, even if the target camera ex102 is located farther away from other cameras ex102, if multiple other cameras ex102 are capturing images of the same area, the importance of the image data captured by the target camera ex102 will be lowered. Furthermore, the importance of image data can be determined based on the number of requests received by the image distribution service. The method for determining importance is not limited to the above methods or combinations; any method that is consistent with the configuration and purpose of the surveillance system or image distribution system may be used.

[0700] Furthermore, the importance determination does not need to be based on the captured image data. For example, the importance of camera ex102 that transmits image data to terminals other than image information processing device ex101 can be set higher. Conversely, the importance of camera ex102 that transmits image data to terminals other than image information processing device ex101 can be set lower. This increases the freedom to control the communication band in accordance with the purpose and characteristics of each service, for example, when multiple services requiring image data transmission share a common communication band. This prevents degradation of the quality of each service due to the inability to obtain the required image data.

[0701] Furthermore, the analyzing unit ex112 may determine the importance of the image data using the free viewpoint images and the images captured by the camera ex102.

[0702] Based on the importance determination results from the analysis unit ex112, the image information processing device ex101 sends a communication speed instruction signal to the camera ex102. For example, the image information processing device ex101 instructs the camera ex102 that captured a highly important image to use a higher communication speed. Furthermore, the image information processing device ex101 not only controls the speed but can also send a signal instructing multiple transmissions of important information to minimize the disadvantages of missing information. This enables efficient communication throughout a facility or stadium. Furthermore, communication between the camera ex102 and the image information processing device ex101 can be either wired or wireless. Furthermore, the image information processing device ex101 can control only one of the two methods.

[0703] The camera ex102 transmits the captured image data to the image information processing device ex101 at the communication speed specified by the communication speed instruction signal. Furthermore, if the camera ex102 fails to retransmit the captured image data after a predetermined number of retransmissions, it can stop retransmitting the captured image data and start transmitting the next captured image data. This allows for efficient communication throughout the facility or stadium, and accelerates processing in the analysis unit ex112.

[0704] Furthermore, when the bandwidth allocated to each communication speed is insufficient to transmit the captured image data, the camera ex102 may convert the captured image data into image data at a bit rate that can be transmitted at the allocated communication speed, transmit the converted image data, or stop transmitting the image data.

[0705] Furthermore, to prevent blind spots, when using image data, only a portion of the captured area contained in the captured image data may need to be corrected. In this case, the camera ex102 needs to extract at least the area needed to prevent blind spots from the image data, generate extracted image data, and then transmit the extracted image data to the image information processing device ex101. This configuration allows for the reduction of blind spots while utilizing a smaller communication bandwidth.

[0706] Furthermore, for example, when displaying superimposed information or distributing images, camera ex102 needs to transmit its position information and shooting direction information to the image information processing device ex101. In this case, even if camera ex102 is allocated only a bandwidth insufficient for transmitting image data, it can simply transmit the position information and shooting direction information detected by camera ex102. Furthermore, after the image information processing device ex101 estimates the position information and shooting direction information of camera ex102, it can convert the captured image data to the resolution required for the estimated position information and shooting direction information and transmit the converted image data to the image information processing device ex101. This configuration enables even cameras ex102 allocated only a small communication bandwidth to provide services such as superimposed display of supplementary information or image distribution. Furthermore, the image information processing device ex101 can also effectively utilize the shooting area information to obtain shooting area information from more cameras ex102, for example, to detect areas of interest.

[0707] Furthermore, the switching of image data transmission processing in accordance with the allocated communication band can be performed by the camera ex102 based on the notified communication band, or the image information processing device ex101 can determine the operation of each camera ex102 and notify each camera ex102 of a control signal indicating the determined operation. In this way, the processing load required for the switching of operations can be appropriately distributed based on the required computational effort, the processing capabilities of the camera ex102, and the required communication band.

[0708] (Variation 10)

[0709] The analyzing unit ex112 can determine the importance of the image data based on the field of view information (and / or viewpoint information) transmitted from the image receiving device ex103. For example, the analyzing unit ex112 assigns a higher importance to image data that includes a larger area indicated by field of view information (and / or viewpoint information). Alternatively, the analyzing unit ex112 can determine the importance of the image data by considering the number of people or moving objects in the image. The method for determining importance is not limited to this.

[0710] Furthermore, the communication control method described in this embodiment is not necessarily applicable to a system that reconstructs a three-dimensional shape from a plurality of image data. For example, in an environment where multiple cameras ex102 are present, the communication control method described in this embodiment is effective as long as image data is transmitted selectively or at different transmission speeds via wired and / or wireless communication.

[0711] (Variation 11)

[0712] In the video distribution system, the video information processing device ex101 may transmit an overview video showing the entire captured scene to the video receiving device ex103 .

[0713] Specifically, upon receiving a distribution request from the video receiving device ex103, the video information processing device ex101 reads an overview image of the entire facility or stadium from the storage unit ex111 and transmits this overview image to the video receiving device ex103. This overview image can be updated at a long interval (possibly at a low frame rate) and of low quality. When a viewer touches a desired portion of the overview image displayed on the screen of the video receiving device ex103, the video receiving device ex103 transmits field of view information (and / or viewpoint information) corresponding to the touched portion to the video information processing device ex101.

[0714] The video information processing device ex101 reads video data corresponding to the field of view information (and / or viewpoint information) from the storage unit ex111 and transmits the video data to the video receiving device ex103.

[0715] Furthermore, the analysis unit ex112 prioritizes 3D shape restoration (3D reconstruction) for the area indicated by the field of view information (and / or viewpoint information), thereby generating a free-viewpoint image. The analysis unit ex112 restores the 3D shape of the entire facility or stadium with sufficient accuracy to provide an overview. This allows the image information processing device ex101 to efficiently restore the 3D shape. This allows for high frame rates and high image quality in the free-viewpoint image of the area the viewer desires to view.

[0716] (Variation 12)

[0717] Alternatively, the image information processing device ex101 may store, as pre-images, three-dimensional shape reconstruction data of a facility or stadium, generated in advance from design drawings, for example. Pre-images are not limited to these data types and may also be virtual space data, which maps the spatial concavity and convexity obtained from a depth sensor for each object, or images derived from past or calibration images or video data.

[0718] For example, in a soccer match at a stadium, the analysis unit ex112 can restrict 3D shape reconstruction to only the players and the ball, then synthesize the resulting reconstruction data with pre-processed footage to generate a free-viewpoint image. Alternatively, the analysis unit ex112 can prioritize 3D shape reconstruction for the players and the ball. This allows the image information processing device ex101 to efficiently perform 3D shape reconstruction. This allows for high frame rates and high image quality for the free-viewpoint images of the players and the ball that the viewer is interested in. Furthermore, in a surveillance system, the analysis unit ex112 can restrict 3D shape reconstruction to only people and moving objects, or prioritize them for 3D shape reconstruction.

[0719] (Variant 13)

[0720] The time of each device can also be calibrated at the start of shooting, based on a server reference time or the like. The analysis unit ex112 uses the multiple image data captured by the multiple cameras ex102, captured at times within a predetermined time range with a set accuracy, to perform three-dimensional shape reconstruction. This time is detected, for example, by using the time when the captured image data was stored in the storage unit ex111. The method for detecting the time is not limited to this. Consequently, the image information processing device ex101 can efficiently perform three-dimensional shape reconstruction, thereby achieving a high frame rate and high image quality for free-viewpoint images.

[0721] Alternatively, the analyzing unit ex112 may restore the three-dimensional shape by using only the high-quality data among the plurality of image data stored in the storage unit ex111 or by giving priority to the high-quality data.

[0722] (Variant 14)

[0723] The analysis unit ex112 can use the camera attribute information to restore the three-dimensional shape. For example, the analysis unit ex112 can use the camera attribute information to generate a three-dimensional image using methods such as visual morphology or multi-view stereo. In this case, the camera ex102 transmits captured image data and camera attribute information to the image information processing device ex101. The camera attribute information includes, for example, the shooting position, shooting angle, shooting time, and zoom factor.

[0724] As a result, the video information processing apparatus ex101 can efficiently restore the three-dimensional shape, thereby achieving a higher frame rate and higher image quality for free viewpoint videos.

[0725] Specifically, camera ex102 defines three-dimensional coordinates within a facility or stadium. It then transmits information indicating the coordinates, the angle, zoom level, and time of capture, along with the image, to the image information processing device ex101 as camera attribute information. Furthermore, when camera ex102 is activated, the clock on the communication network within the facility or stadium is synchronized with the clock within the camera, thereby generating time information.

[0726] Furthermore, when the camera ex102 is activated or at an arbitrary timing, the camera ex102 is directed toward a specific point in the facility or the stadium, thereby obtaining position and angle information of the camera ex102. Figure 38 The figure shows an example of a notification displayed on the screen of camera ex102 when camera ex102 is activated. Following this notification, the user overlaps the "+" displayed in the center of the screen with the "+" on the center of the soccer ball in the advertisement on the north side of the stadium. When the user touches the display of camera ex102, camera ex102 obtains vector information from camera ex102 to the advertisement, determining a reference for the camera's position and angle. Subsequently, the camera's coordinates and angle at each moment are determined based on the camera's motion information. Of course, this display is not limited to this; arrows or other indicators can also be used to indicate coordinates, angles, or the speed of movement of the captured area, even during shooting.

[0727] The coordinates of the camera ex102 can be determined using radio waves from GPS, WiFi (registered trademark), 3G, LTE (Long Term Evolution), or 5G (Wireless LAN), or using short-range wireless technologies such as beacons (Bluetooth (registered trademark), ultrasonic waves). Furthermore, information indicating the base station within a facility or stadium to which captured image data was transmitted can be used.

[0728] (Variant 15)

[0729] This system can be provided as an application program that operates on a mobile terminal such as a smartphone.

[0730] To log in to the system, various SNS accounts can be used. Alternatively, application-specific accounts or customer accounts with limited functionality can be used. Using accounts in this way allows users to evaluate favorite images or accounts. Furthermore, by prioritizing bandwidth allocation for image data similar to the image data currently being filmed or viewed, or for image data with viewpoints similar to those currently being filmed or viewed, the resolution of this image data can be improved. This allows for highly accurate reconstruction of three-dimensional shapes from these viewpoints.

[0731] Furthermore, users can use the app to select their favorite images and follow others to see their selected images before other users. With the other party's approval, they can connect with them through text chat, thus creating new groups.

[0732] In this way, users are connected to each other in the group, and the shooting itself can be activated. The sharing of the captured images is also activated, and thus the restoration of a three-dimensional shape with higher accuracy can be promoted.

[0733] Furthermore, based on the connections within the group, users can edit images or videos taken by others, creating new images or videos by combining them with their own. This allows only those within the group to share the new images or videos, making it possible to create new video works. Furthermore, by inserting CG animations into these edits, video works can also be used in augmented reality games.

[0734] Furthermore, since this system allows for sequential output of 3D model data, 3D printers at facilities can output three-dimensional objects based on the 3D model data of characteristic scenes, such as goal-scoring scenes. This allows objects based on the game scenes to be sold as gifts, such as keychains, and distributed to participating players. Of course, images from optimal viewpoints can also be printed, as well as regular photos.

[0735] (Variant 16)

[0736] Using the above system, for example, it is possible to manage the general status of the entire area through a control center connected to the system based on images from police car cameras and police wearable cameras.

[0737] During typical patrols, still images are transmitted and received every few minutes, for example. Furthermore, the control center identifies areas with a high probability of crime based on crime maps generated through analysis of past crime data, or maintains regional data associated with the identified crime probability. The frequency of image transmission and reception can be increased in areas with a high probability of crime, or the images can be converted to animations. Furthermore, animations or three-dimensional reconstructions created using techniques such as SfM can be used when an incident occurs. Furthermore, the control center or each terminal can simultaneously utilize information from other sensors, such as depth sensors or temperature sensors, to calibrate images or virtual space, enabling police to more accurately grasp the situation.

[0738] The control center then uses the three-dimensional reconstruction data to feed back information about the object to multiple terminals, allowing individuals with each terminal to track the object.

[0739] Furthermore, aerial photography using flying devices such as quadcopters and drones has recently become popular for purposes such as surveying buildings and the environment, or capturing immersive sports footage. While image shake can be a problem with these autonomous mobile devices, SfM can correct for this shake and create three-dimensional images by adjusting position and tilt. This improves image quality and spatial reconstruction accuracy.

[0740] Furthermore, the installation of onboard cameras that capture images outside the vehicle is mandatory due to national regulations. Even with these onboard cameras, the use of three-dimensional data modeled from multiple images allows for more accurate understanding of the weather in the direction of the destination, road conditions ahead, and traffic congestion.

[0741] (Variant 17)

[0742] The above system can also be applied to a system that uses multiple cameras to measure the distance or model a building or equipment, for example.

[0743] For example, when using a drone to photograph buildings from above for distance measurement or modeling, if a moving object is reflected in the camera during distance measurement, the accuracy of the distance measurement will be reduced. Furthermore, distance measurement and modeling of moving objects may become impossible.

[0744] Furthermore, as described above, by utilizing multiple cameras (fixed cameras, smartphones, wearable cameras, drones, etc.), it is possible to achieve stable and accurate distance measurement and modeling of buildings, regardless of whether there are moving objects. Furthermore, distance measurement and modeling of moving objects are also possible.

[0745] Specifically, at a construction site, for example, cameras can be attached to workers' helmets. This allows for building distance measurement in parallel with the workers' work. This can improve operational efficiency and prevent errors. Furthermore, images captured by the workers' cameras can be used to model the building. Furthermore, remote managers can view the modeled building to confirm progress.

[0746] Furthermore, the system can be used to inspect equipment that cannot be stopped, such as equipment in factories or power plants, and to check for abnormalities in the opening and closing of bridges and reservoirs, or the operation of amusement park rides.

[0747] Furthermore, by monitoring the congestion conditions and traffic volumes on roads through this system, a map showing the congestion conditions and traffic volumes on roads at each time period can be generated.

[0748] (Implementation 6)

[0749] By recording a program for implementing the configuration of the image processing method described in each of the above embodiments onto a storage medium, the processing described in each of the above embodiments can be easily implemented on a standalone computer system. The storage medium may be a magnetic disk, an optical disk, a magneto-optical disk, an IC card, a semiconductor memory, or any other medium capable of recording a program.

[0750] Furthermore, here we describe application examples of the image processing methods shown in the above embodiments and a system that uses them. The system is characterized by having an apparatus that uses the image processing method. The other configurations of the system can be modified as appropriate depending on the situation.

[0751] Figure 39 This figure shows the overall configuration of a content providing system ex200 that implements content distribution services. The area where communication services are provided is divided into cells of desired sizes, and base stations ex206, ex207, ex208, ex209, and ex210, which are fixed wireless stations, are installed in each cell.

[0752] The content providing system ex200 is connected to various devices such as a computer ex211, a PDA (Personal Digital Assistant) ex212, a camera ex213, a smartphone ex214, and a game console ex215 on the Internet ex201 via an Internet service provider ex202, a communication network ex204, and base stations ex206 to ex210.

[0753] However, the content providing system ex200 is not subject to Figure 39The illustrated configuration is limited, and any combination of elements may be connected. Furthermore, each device may be directly connected to a communication network ex204, such as a telephone line, cable television, or optical communication, without going through the fixed wireless base stations ex206 to ex210. Furthermore, each device may be directly connected to another via short-range wireless communication or the like.

[0754] The camera ex213 is a device such as a digital video camera capable of capturing moving images, and the camera ex216 is a device such as a digital camera capable of capturing both still and moving images. Furthermore, the smartphone ex214 may be any of the following: a smartphone using the GSM (registered trademark) standard, the CDMA (Code Division Multiple Access) standard, the W-CDMA (Wideband-Code Division Multiple Access) standard, the LTE (Long Term Evolution) standard, the HSPA (High Speed ​​Packet Access) standard, or a communication standard utilizing high frequency bands, or a PHS (Personal Handyphone System).

[0755] In the content providing system ex200, live performance distribution is enabled by connecting a camera ex213 and other devices to a streaming server ex203 via a base station ex209 and a communication network ex204. For live performance distribution, content captured by a user using the camera ex213 (e.g., images of a live music performance) is encoded and transmitted to the streaming server ex203. The streaming server ex203, in turn, streams the received content data to requesting clients. Clients include computers ex211, PDAs ex212, cameras ex213, smartphones ex214, and game consoles ex215, capable of decoding the encoded data. Each device that receives the distributed data decodes and reproduces the data.

[0756] Furthermore, the encoding of captured data can be performed by either the camera ex213 or the streaming server ex203, which transmits the data, or they can be shared. Similarly, the decoding of distributed data can be performed by either the client or the streaming server ex203, or they can be shared. Furthermore, not only the camera ex213 but also still and / or moving image data captured by the camera ex216 can be transmitted to the streaming server ex203 via the computer ex211. In this case, the encoding process can be performed by any of the camera ex216, the computer ex211, or the streaming server ex203, or they can be shared. Furthermore, the display of the decoded image can be coordinated across multiple devices connected to the system to display the same image, the entire image can be displayed on a device with a large display, or a zoomed-in display of a portion of the image can be performed on a smartphone ex214 or the like.

[0757] These encoding and decoding processes are typically handled by the computer ex211 or the LSI ex500 included in various devices. The LSI ex500 can be a single chip or a structure composed of multiple chips. Alternatively, video encoding and decoding software can be loaded onto a recording medium (CD-ROM, floppy disk, hard disk, etc.) readable by the computer ex211 or other devices, and the encoding and decoding processes can be performed using this software. Furthermore, if the smartphone ex214 has a camera, video data captured by the camera can also be transmitted. In this case, the video data is encoded and processed by the LSI ex500 included in the smartphone ex214.

[0758] Furthermore, the streaming media server ex203 may be a plurality of servers or a plurality of computers that process, record, and distribute data in a distributed manner.

[0759] As described above, in the content provision system ex200, the encoded data is received and played back by the client. Thus, in the content provision system ex200, the client can receive, decode, and play back user-sent information in real time, enabling personal broadcasting even for users without special rights or equipment.

[0760] In addition, the example is not limited to the content providing system ex200. Figure 40As shown, the above embodiments can also be applied to the digital broadcasting system ex300. Specifically, a broadcasting station ex301 transmits multiplexed data, obtained by multiplexing music data onto video data, via radio waves to a communication or satellite ex302. This video data is encoded using the moving image encoding methods described in the above embodiments. The broadcasting satellite ex302, receiving this data, emits broadcast radio waves, which are then received by an antenna ex304 in a home capable of satellite broadcast reception. The received multiplexed data is decoded and reproduced by a device such as a television (receiver) ex400 or a set-top box (STB) ex317.

[0761] Furthermore, the video decoding device or video encoding device described in the above embodiments can also be incorporated into a reader / recorder ex318 to read and decode multiplexed data recorded on a recording medium ex315 such as a DVD or BD, or a memory ex316 such as an SD card, or to encode video signals from the recording medium ex315 or memory ex316, and optionally multiplex the audio signals and write them to the reader / recorder ex318. In this case, the reproduced video signal can be displayed on a monitor ex319, and the video signal can be reproduced in another device or system using the recording medium ex315 or memory ex316 on which the multiplexed data is recorded. Alternatively, the video decoding device can be incorporated into a set-top box ex317 connected to a cable TV cable ex303 or a satellite / terrestrial broadcast antenna ex304, and displayed on a television monitor ex319. In this case, the video decoding device can be incorporated into a television set instead of a set-top box.

[0762] Figure 41 A smartphone ex214 is shown. Figure 42 The following shows an example of the configuration of a smartphone ex214. The smartphone ex214 includes an antenna ex450 for transmitting and receiving radio waves with a base station ex210, a camera unit ex465 capable of shooting videos and still images, and a display unit ex458 such as a liquid crystal display that displays decoded data such as videos shot by the camera unit ex465 and videos received by the antenna ex450. The smartphone ex214 also includes an operation unit ex466 such as a touch screen, a sound output unit ex457 such as a speaker for outputting sound, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing coded or decoded data such as captured videos, still images, recorded sounds, or received videos, still images, emails, etc., and a slot unit ex464 as an interface unit that is connected to the Figure 40The memory ex316 shown, or the interface of SIMex468 for identifying users and authenticating access to various data represented by the network.

[0763] The smartphone ex214 includes a main control unit ex460 that integrates control of the display unit ex458 and the operation unit ex466, a power supply circuit unit ex461, an operation input control unit ex462, an image signal processing unit ex455, a camera interface unit ex463, an LCD (Liquid Crystal Display) control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467, which are interconnected via a bus ex470.

[0764] When the user ends a call and turns on the power button, the power supply circuit unit ex461 supplies power from the battery pack to various components, thereby activating the smartphone ex214 to be operational.

[0765] Under the control of the main control unit ex460, which includes a CPU, ROM, and RAM, the smartphone ex214, in voice call mode, receives audio signals collected by the audio input unit ex456 and converts them into digital audio signals through the audio signal processing unit ex454. The modulator / demodulator ex452 performs spread spectrum processing on the digital audio signals, which are then subjected to digital-to-analog conversion and frequency conversion by the transmitter / receiver ex451 before being transmitted from the antenna ex450. Furthermore, in voice call mode, the smartphone ex214 amplifies the data received by the antenna ex450, performs frequency conversion and analog-to-digital conversion on the data, performs inverse spread spectrum processing by the modulator / demodulator ex452, converts the data into analog audio signals through the audio signal processing unit ex454, and outputs the signals via the audio output unit ex457.

[0766] Furthermore, when sending an email in data communication mode, the text data of the email, input via the main unit's operation unit ex466 or other input methods, is sent to the main control unit ex460 via the operation input control unit ex462. The main control unit ex460 then performs spectrum spreading processing on the text data using the modulation / demodulation unit ex452. The transmission / reception unit ex451 then performs digital-to-analog conversion and frequency conversion, and then transmits the data via the antenna ex450 to the base station ex210. When receiving an email, the received data undergoes processing similar to the inverse of the aforementioned processing and is then output to the display unit ex458.

[0767] In data communication mode, when transmitting video, still images, or both video and audio, the video signal processing unit ex455 compresses and encodes the video signal provided by the camera unit ex465 using the video encoding method described in the above embodiments, and then sends the encoded video data to the multiplexing / demultiplexing unit ex453. Furthermore, the audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 during the process of capturing video or still images with the camera unit ex465, and then sends the encoded audio data to the multiplexing / demultiplexing unit ex453.

[0768] The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data provided by the video signal processing unit ex455 and the encoded audio data provided by the audio signal processing unit ex454 in a prescribed manner. The resulting multiplexed data is subjected to spectrum spreading processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452. After digital-to-analog conversion and frequency conversion are performed by the transmitting / receiving unit ex451, the data is transmitted via the antenna ex450.

[0769] In data communication mode, when receiving data linking to a video file such as a homepage, or an email with attached video and / or audio, the multiplexing / demultiplexing unit ex453 demultiplexes the multiplexed data received via antenna ex450 into a bit stream of video data and a bit stream of audio data to decode the multiplexed data. The multiplexing / demultiplexing unit ex453 then supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronous bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method described in the above embodiments. The video signal, such as a video or still image included in a video file linking to a homepage, is then displayed on the display unit ex458 via the LCD control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and the audio output unit ex457 outputs the audio signal.

[0770] Furthermore, similar to the television ex400, terminals such as the smartphone ex214 can be used as both a transmitting and receiving terminal equipped with both an encoder and a decoder. Alternatively, three implementation options are possible: a transmitting terminal with only an encoder and a receiving terminal with only a decoder. Furthermore, while the digital broadcasting system ex300 illustrates the transmission and reception of multiplexed data, where music data and other data are multiplexed with video data, it is also possible to multiplex data such as text data associated with the video in addition to audio data, or to use video data itself rather than multiplexed data.

[0771] Furthermore, the present application is not limited to the above-described embodiments, and various modifications and corrections can be made without departing from the scope of the present application.

[0772] The present application can be applied to a three-dimensional information processing device.

Claims

1. A three-dimensional information processing method, comprising the following processing: determining whether first map data including first three-dimensional position information of a first range can be obtained via the communication path; and When the first map data cannot be obtained via the communication path, second map data required for estimating the self-position of the mobile body is obtained via the communication path, the second map data including third three-dimensional position information of a second range narrower than the first range. The own position of the mobile object is estimated based on the first map data or the second map data.

2. The three-dimensional information processing method according to claim 1, further comprising: The self-position of the mobile body having the sensor is estimated using the first three-dimensional position information and the second three-dimensional position information generated based on the information detected by the sensor. In the judgment, it is predicted whether the mobile object will enter an area with poor communication conditions, When it is predicted that the mobile object will enter an area with a poor communication status, the mobile object obtains the first map data including the first three-dimensional position information before the mobile object enters the area.

3. The three-dimensional information processing method according to claim 1, further comprising: determining whether at least a portion of the data of the first three-dimensional position information is damaged, When at least a portion of the data of the first three-dimensional position information is damaged, data including the portion is requested via the communication path.

4. The three-dimensional information processing method according to claim 1, further comprising: Determine the random access unit of the three-dimensional position information required for estimating the self-position of the mobile object, If the first map data cannot be obtained via the communication path, the second map data is obtained via the communication path based on the determined random access unit.

5. A three-dimensional information processing device comprising: a determination unit that determines whether first map data including first three-dimensional position information of a first range can be obtained via the communication path; and an obtaining unit for obtaining, via the communication path, second map data required for estimating the self-position of the mobile body when the first map data cannot be obtained via the communication path, the second map data including third three-dimensional position information of a second range narrower than the first range; The own position of the mobile object is estimated based on the first map data or the second map data.

Citation Information

Cited By

  • Three-dimensional information processing method and three-dimensional information processing apparatus

    CN121236934A