Three-dimensional space information processing system

The method generates a three-dimensional point cloud by adding attribute values from two-dimensional images, addressing inefficiencies in existing point cloud data, enabling efficient and accurate self-position estimation.

JP2025147222APending Publication Date: 2025-10-06PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025132384
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-11-20
Filing Date
2025-08-07
Publication Date
2025-10-06

AI Technical Summary

Technical Problem

Existing three-dimensional point cloud data is inefficient in terms of data storage and processing, necessitating a need to generate compressed 3D point cloud data and enable self-localization using reduced-data-size 3D point cloud data.

Method used

A method and device that generate a three-dimensional point cloud by adding attribute values from two-dimensional images to three-dimensional points, allowing efficient position estimation without sensing new point clouds, and adjust data amount based on criteria and importance levels.

Benefits of technology

Enables accurate and efficient self-position estimation using reduced three-dimensional point cloud data, reducing data transmission and processing requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025147222000001_ABST
    Figure 2025147222000001_ABST
Patent Text Reader

Abstract

To provide a three-dimensional space information processing system which can achieve further improvement.SOLUTION: A three-dimensional space information processing system includes a memory and a processor, acquires a first three-dimensional point group sensed by a distance sensor, and a two-dimensional image imaged by a camera, detects an attribute value corresponding to each point of the first three-dimensional point group, from the two-dimensional image, adds the detected attribute value to each of the points of the first three-dimensional point group and thereby generates a three-dimensional point group associated with the attribute value, processes the generated three-dimensional point group so as to adjust an information amount on the basis of a predetermined reference, and executes space recognition processing, using the processed three-dimensional point group.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a three-dimensional spatial information processing system. [Background technology]

[0002] Patent Document 1 discloses a method for transferring three-dimensional shape data. In Patent Document 1, three-dimensional shape data is sent over a network for each element, such as a polygon or voxel. The receiving side then imports the three-dimensional shape data and displays an image of each received element. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 9-237354 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the technology disclosed in the above patent document requires further improvement. [Means for solving the problem]

[0005] In order to achieve the above-mentioned object, a three-dimensional spatial information processing system according to one embodiment of the present disclosure is an information processing system having a memory and a processor, and is configured to acquire a first three-dimensional point cloud sensed by a distance sensor and a two-dimensional image captured by a camera, detect attribute values ​​corresponding to each point of the first three-dimensional point cloud from the two-dimensional image, add the detected attribute values ​​to each point of the first three-dimensional point cloud to generate a three-dimensional point cloud associated with the attribute values, process the generated three-dimensional point cloud in a manner that allows the amount of information to be adjusted based on predetermined criteria, and perform spatial recognition processing using the processed three-dimensional point cloud.

[0006] In order to achieve the above object, a three-dimensional point cloud data generation method according to one embodiment of the present disclosure acquires a first three-dimensional point cloud obtained by sensing a three-dimensional object using a distance sensor, detects attribute values ​​corresponding to each of a plurality of first three-dimensional points included in the first three-dimensional point cloud using a processor, and generates a second three-dimensional point cloud including a plurality of second three-dimensional points by adding the attribute values ​​to each of the plurality of first three-dimensional points, each of which is composed of the three-dimensional coordinates and the attribute value of the corresponding first three-dimensional point.

[0007] In order to achieve the above-mentioned object, a three-dimensional point cloud data generation method according to one embodiment of the present disclosure is a three-dimensional point cloud generation method that uses a processor to generate a three-dimensional point cloud including a plurality of three-dimensional points, and includes the steps of: acquiring a two-dimensional image obtained by capturing an image of a three-dimensional object using a camera; and acquiring a first three-dimensional point cloud obtained by sensing the three-dimensional object using a distance sensor; detecting one or more attribute values ​​of the two-dimensional image corresponding to a position on the two-dimensional image from the acquired two-dimensional image; and, for each of the detected one or more attribute values, (i) identifying one or more first three-dimensional points among the plurality of three-dimensional points that constitute the first three-dimensional point cloud to which the position on the two-dimensional image of the attribute value corresponds; and (ii) adding the attribute value to the identified one or more first three-dimensional points to generate a second three-dimensional point cloud including one or more second three-dimensional points, each having the attribute value.

[0008] Furthermore, a position estimation method according to one embodiment of the present disclosure is a position estimation method for estimating the current position of a moving body using a processor, which includes the steps of: acquiring a three-dimensional point cloud including a plurality of three-dimensional points, each having a first attribute value pre-assigned thereto, which is an attribute value of a first two-dimensional image obtained by capturing an image of a three-dimensional object; acquiring a second two-dimensional image of the surroundings of the moving body captured by a camera equipped on the moving body; detecting one or more second attribute values, which are attribute values ​​of the second two-dimensional image corresponding to a position on the second two-dimensional image, from the acquired second two-dimensional image; generating one or more combinations consisting of the second attribute value and the one or more fifth three-dimensional points for each of the detected one or more second attribute values; acquiring the position and orientation of the camera relative to the moving body from a storage device; and calculating the position and orientation of the moving body using the generated one or more combinations and the acquired position and orientation of the camera.

[0009] These general or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. [Effects of the Invention]

[0010] The present disclosure can be further improved. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram illustrating an overview of a location estimation system. [Figure 2] FIG. 2 is a block diagram illustrating an example of a functional configuration of the position estimation system. [Figure 3] FIG. 3 is a block diagram showing an example of the functional configuration of a vehicle serving as a client device. [Figure 4] FIG. 4 is a sequence diagram illustrating an example of the operation of the position estimation system. [Figure 5] FIG. 5 is a diagram for explaining the operation of the position estimation system. [Figure 6] FIG. 6 is a block diagram illustrating an example of the functional configuration of the mapping unit. [Figure 7] FIG. 7 is a flowchart showing an example of detailed mapping processing. [Figure 8] FIG. 8 is a flowchart showing an example of the operation of the three-dimensional point cloud data generating device in the first modification. [Figure 9] FIG. 9 is a flowchart showing an example of detailed importance calculation processing. [Figure 10] FIG. 10 is a diagram illustrating an example of the data configuration of the third three-dimensional point. [Figure 11] FIG. 11 is a block diagram illustrating an example of a functional configuration of the encoding unit. [Figure 12] FIG. 12 is a flowchart showing an example of detailed encoding processing. DETAILED DESCRIPTION OF THE INVENTION

[0012] (Findings that formed the basis of this disclosure) Three-dimensional point cloud data is widespread in manufacturing and construction and has recently become important in information technology applications such as autonomous driving. Three-dimensional point cloud data is typically large and inefficient in terms of data storage and processing. Therefore, there is a need to generate compressed 3D point cloud data in a compact size in order to utilize 3D point cloud data in practical applications. Thus, there is a need to be able to reduce the amount of 3D point cloud data. There is also a need to be able to estimate self-localization using reduced-data-size 3D point cloud data.

[0013] The present disclosure aims to provide a three-dimensional point cloud data generation method and a three-dimensional point cloud data generation device that can effectively reduce the amount of three-dimensional point cloud data, as well as a position estimation method and a position estimation device that can estimate a self-position using three-dimensional point cloud data with reduced data amount.

[0014] A three-dimensional point cloud data generation method according to one embodiment of the present disclosure is a three-dimensional point cloud generation method that uses a processor to generate a three-dimensional point cloud including a plurality of three-dimensional points, and includes the steps of: acquiring a two-dimensional image obtained by capturing an image of a three-dimensional object using a camera; and acquiring a first three-dimensional point cloud obtained by sensing the three-dimensional object using a distance sensor; detecting one or more attribute values ​​of the two-dimensional image corresponding to a position on the two-dimensional image from the acquired two-dimensional image; and, for each of the detected one or more attribute values, (i) identifying one or more first three-dimensional points among the plurality of three-dimensional points that constitute the first three-dimensional point cloud to which the position on the two-dimensional image of the attribute value corresponds; and (ii) adding the attribute value to the identified one or more first three-dimensional points to generate a second three-dimensional point cloud including one or more second three-dimensional points, each having the attribute value.

[0015] According to this, a second three-dimensional point cloud is generated including second three-dimensional points obtained by adding attribute values ​​of a two-dimensional image corresponding to positions on a two-dimensional image obtained using a camera to three-dimensional points corresponding to positions on the two-dimensional image, among a plurality of three-dimensional points constituting the first three-dimensional point cloud obtained using a distance sensor. Therefore, a position estimation device that estimates its own position can efficiently estimate its own position without sensing a new three-dimensional point cloud by simply capturing a new two-dimensional image of its surroundings and comparing the attribute values ​​corresponding to positions in the captured two-dimensional image with the attribute values ​​added to each second three-dimensional point of the second three-dimensional point cloud.

[0016] Furthermore, the acquiring step may include acquiring a plurality of the two-dimensional images, each of which is obtained by capturing an image with the camera at a different position and / or orientation from each other; the detecting step may include detecting the one or more attribute values ​​for each of the acquired two-dimensional images; the three-dimensional point cloud data generation method may further include matching corresponding attribute values ​​in two of the plurality of two-dimensional images using the one or more attribute values ​​detected for each of the plurality of two-dimensional images, thereby outputting one or more pairs of matched attribute values; the identifying step may include identifying, for each of the one or more pairs, one or more first three-dimensional points whose positions on the two-dimensional image correspond to the two attribute values ​​that constitute the pair, using the positions on the two-dimensional image of each of the two attribute values ​​that constitute the pair and the position and orientation of the camera when each of the two two-dimensional images was captured; and the generating step may include adding, for each of the one or more pairs, attribute values ​​based on the two attribute values ​​that constitute the pair to the identified one or more first three-dimensional points, thereby generating the second three-dimensional point cloud.

[0017] According to this method, a pair of attribute values ​​is identified by matching the attribute values ​​of two two-dimensional images, a three-dimensional point corresponding to the position is identified using the positions corresponding to the two-dimensional images of the attribute values ​​of the pair, and the two attribute values ​​that make up the pair are added to the identified three-dimensional point. Therefore, it is possible to accurately identify the three-dimensional point corresponding to the attribute values.

[0018] Furthermore, in the generation, the second 3D point may be generated by adding a plurality of the attribute values ​​to the identified one of the first 3D points.

[0019] According to this, since the number of attribute values ​​added to one first three-dimensional point is multiple, in position estimation, the attribute values ​​of the two-dimensional image corresponding to the position on the two-dimensional image obtained for position estimation can be accurately associated with the first three-dimensional point.

[0020] Furthermore, the plurality of attribute values ​​added to the one first three-dimensional point in the generation may be attribute values ​​detected from a plurality of the two-dimensional images, respectively.

[0021] A single first 3D point can be captured from multiple viewpoints. Furthermore, in position estimation, two-dimensional images are captured at different positions, and thus the resulting two-dimensional images are captured from multiple different viewpoints. Therefore, by adding multiple attribute values ​​detected from multiple two-dimensional images, even when two-dimensional images captured from multiple different viewpoints are used in position estimation, it is possible to easily associate the attribute values ​​obtained from each two-dimensional image with the first 3D point. Therefore, position estimation can be easily performed.

[0022] Furthermore, the plurality of attribute values ​​added to the one first 3D point in the generation may be attribute values ​​of different attribute types.

[0023] According to this, since the attribute values ​​added to one first three-dimensional point are different types of attribute values, in position estimation, the attribute values ​​of the two-dimensional image corresponding to the position on the two-dimensional image obtained for position estimation can be accurately associated with the first three-dimensional point.

[0024] In addition, during the detection, the feature amounts calculated for each of the multiple regions that make up the acquired two-dimensional image may be detected as the one or more attribute values ​​of the two-dimensional image that correspond to positions on the two-dimensional image.

[0025] Therefore, the feature amount at a position in the two-dimensional image can be easily calculated using a predetermined method.

[0026] Furthermore, the generation may further include (i) calculating the importance of each of the one or more second three-dimensional points based on the attribute value added to the second three-dimensional point, and (ii) adding the calculated importance to the second three-dimensional point to generate a third three-dimensional point cloud including one or more third three-dimensional points, each having the attribute value and the importance.

[0027] Therefore, for example, a client device can prioritize the use of third three-dimensional points with greater importance, and can adjust the amount of three-dimensional point cloud data used for processing so as not to adversely affect the accuracy of processing.

[0028] Furthermore, the method may further receive a threshold value from a client device, extract one or more fourth three-dimensional points from the one or more third three-dimensional points that have been assigned an importance level that exceeds the received threshold value, and transmit a fourth three-dimensional point cloud including the extracted one or more fourth three-dimensional points to the client device.

[0029] Therefore, the amount of three-dimensional point cloud data to be transmitted can be adjusted according to the request of the client device.

[0030] Furthermore, a position estimation method according to one embodiment of the present disclosure is a position estimation method for estimating the current position of a moving body using a processor, which includes the steps of: acquiring a three-dimensional point cloud including a plurality of three-dimensional points, each having a first attribute value pre-assigned thereto, which is an attribute value of a first two-dimensional image obtained by capturing an image of a three-dimensional object; acquiring a second two-dimensional image of the surroundings of the moving body captured by a camera equipped on the moving body; detecting one or more second attribute values, which are attribute values ​​of the second two-dimensional image corresponding to a position on the second two-dimensional image, from the acquired second two-dimensional image; generating one or more combinations consisting of the second attribute value and the one or more fifth three-dimensional points for each of the detected one or more second attribute values; acquiring the position and orientation of the camera relative to the moving body from a storage device; and calculating the position and orientation of the moving body using the generated one or more combinations and the acquired position and orientation of the camera.

[0031] According to this, the attribute values ​​of the two-dimensional image captured of the three-dimensional object are added to the three-dimensional point cloud in advance, so that the device can efficiently estimate its own position by capturing a new second two-dimensional image of the surroundings of the device without having to sense a new three-dimensional point cloud, and comparing the second attribute value corresponding to the position of the captured second two-dimensional image and the combination of one or more fifth three-dimensional points corresponding to the second attribute value with the first attribute value added to each three-dimensional point of the three-dimensional point cloud.

[0032] These general or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or as any combination of a system, a method, an integrated circuit, a computer program, or a recording medium.

[0033] Hereinafter, a three-dimensional point cloud data generation method, a position estimation method, a three-dimensional point cloud data generation device, and a position estimation device according to one embodiment of the present disclosure will be described in detail with reference to the drawings.

[0034] Note that the embodiments described below each illustrate a specific example of the present disclosure. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components.

[0035] (Embodiment 1) First, an outline of the present embodiment will be described. In this embodiment, a position estimation system that estimates the self-position of a moving body such as a vehicle using three-dimensional point cloud data will be described.

[0036] FIG. 1 is a diagram illustrating an overview of a location estimation system.

[0037] 1 shows a three-dimensional point cloud data generation device 100, vehicles 200 and 300, a communication network 400, and a base station 410 of a mobile communication system. For example, a position estimation system 1 includes, among these components, the three-dimensional point cloud data generation device 100 and the vehicles 200 and 300. The vehicles 200 and 300 are different types of vehicles. The position estimation system 1 is not limited to including one vehicle 200, but may include two or more vehicles 200, and similarly, is not limited to including one vehicle 300, but may include three or more vehicles 300.

[0038] The 3D point cloud data generating device 100 uses a plurality of 2D images captured by a camera 210 equipped in the vehicle 200 and a LiDAR (Light Detection and Ranging Apparatus) equipped in the vehicle 200. , and a 3D point cloud obtained by Laser Imaging Detection and Ranging (LAR)230. The 3D point cloud data generating device 100 is a device for generating 3D point cloud data used for estimating the self-position of the vehicle 300 by the vehicle 300. Here, the 3D point cloud data is data indicating a 3D point cloud including a plurality of 3D points to which feature amounts of feature points are added. The plurality of 3D points to which feature amounts of feature points are added will be described later. The 3D point cloud data generating device 100 is, for example, a server.

[0039] Vehicle 200 is equipped with camera 210 and LiDAR 230, and is a vehicle that captures two-dimensional images of three-dimensional objects around vehicle 200 and detects a three-dimensional point cloud of the three-dimensional objects. The three-dimensional point cloud is composed of, for example, a plurality of three-dimensional coordinates (X, Y, Z) that indicate a plurality of positions on the surface of the three-dimensional objects around vehicle 200. For example, vehicle 200 captures images using camera 210 at different times while traveling on a road, thereby generating two-dimensional images captured by camera 210 of vehicle 200 from the position of vehicle 200 at each time. In other words, vehicle 200 captures images using camera 210 at different times while traveling along a travel route. Therefore, the plurality of two-dimensional images are obtained by capturing images using camera 210 at different positions and / or orientations. Furthermore, the vehicle 200 performs sensing using the LiDAR 230 at a plurality of different times while traveling on a road, for example, and generates a three-dimensional point cloud sensed by the LiDAR 230 of the vehicle 200 from the position of the vehicle 200 at each of the times.

[0040] The sensor data from the LiDAR 230 may be used to identify a two-dimensional image captured by the camera 210 of the vehicle 200 and the position and attitude of the camera 210 when the two-dimensional image was captured. In this case, an accurate 3D map used by a general vehicle to estimate its own position is obtained by a device separate from the vehicle 200. The accurate 3D map may be generated using a 3D point cloud obtained by the above process. The accurate 3D map is an example of a first 3D point cloud and is composed of a plurality of 3D points. The plurality of 3D points may be represented by coordinates in three different directions, for example, the X-axis direction, the Y-axis direction, and the Z-axis direction.

[0041] The sensor data from LiDAR 230 may be used to generate an accurate three-dimensional map used by a general vehicle to estimate its own position. In this case, LiDAR 230 is a distance sensor that can measure, for example, more three-dimensional points with higher accuracy than LiDARs used by general vehicles for autonomous driving or driving assistance. In other words, vehicle 200 is a sensing device that acquires a three-dimensional point cloud and two-dimensional images to generate an accurate three-dimensional map used for self-position estimation.

[0042] The timing of imaging by camera 210 and the timing of sensing by LiDAR 230 may or may not be synchronized. The timing of imaging by camera 210 and the timing of sensing by LiDAR 230 may be the same timing or different timings.

[0043] The vehicle 300 is equipped with a camera 310 and captures two-dimensional images of three-dimensional objects around the vehicle 300. The vehicle 300 estimates its own position using the two-dimensional image captured by the camera 310 and the three-dimensional point cloud acquired from the three-dimensional point cloud data generation device 100. The vehicle 300 performs automatic driving or driving assistance using the result of estimating its own position. In other words, the vehicle 300 functions as a client device.

[0044] The communication network 400 may be a general-purpose network such as the Internet, or may be a dedicated network. The base station 410 is a base station used in a mobile communication system such as a third-generation mobile communication system (3G), a fourth-generation mobile communication system (4G), LTE (registered trademark), or a fifth-generation mobile communication system (5G).

[0045] Next, a specific example of the functional configuration of the position estimation system will be described with reference to FIG.

[0046] FIG. 2 is a block diagram illustrating an example of a functional configuration of the position estimation system.

[0047] First, the functional configuration of the three-dimensional point cloud data generating device 100 will be described.

[0048] The 3D point cloud data generating device 100 includes a communication unit 110, a mapping unit 120, an encoding unit 130, and an external memory 140.

[0049] The communication unit 110 communicates with the vehicles 200 and 300 via the communication network 400. The communication unit 110 acquires a two-dimensional image and a three-dimensional point cloud by receiving them from the vehicle 200 via communication via the communication network 400. The communication unit 110 may also acquire the relative position and attitude of the camera 210 with respect to the LiDAR 230 by receiving them from the vehicle 200 via communication via the communication network 400. The communication unit 110 transmits the three-dimensional point cloud to which the feature amounts of the feature points have been added to the vehicle 300 via communication via the communication network 400. The communication unit 110 may also transmit the three-dimensional point cloud to which the feature amounts of the feature points have been added to a higher-level server (not shown) via communication via the communication network 400.

[0050] The three-dimensional point cloud data generation device 100 may store the relative position and orientation of the camera 210 with respect to the LiDAR 230 in advance in a non-volatile storage device such as the external memory 140 as a table for each vehicle type or for each vehicle 200. In this case, the three-dimensional point cloud data generation device 100 may acquire the relative position and orientation of the camera 210 with respect to the LiDAR 230 in the vehicle 200 by acquiring information for identifying the vehicle type or vehicle from the vehicle 200, and identifying the relative position and orientation of the camera 210 with respect to the LiDAR 230 in the vehicle 200 from the table stored in the non-volatile storage device.

[0051] The communication unit 110 is realized by a communication interface that can be communicatively connected to the communication network 400. Specifically, the communication unit 110 is communicatively connected to the communication network 400 through a communication connection with a base station 410 of a mobile communication system. The communication unit 110 may be realized by a wireless communication interface that complies with communication standards used in mobile communication systems such as a third generation mobile communication system (3G), a fourth generation mobile communication system (4G), LTE (registered trademark), or a fifth generation mobile communication system (5G). The communication unit 110 may also be realized by a wireless LAN (Local Area Network) interface that complies with the IEEE 802.11a, b, g, n, or ac standard, or by a communication interface that communicatively connects to the communication network 400 through a communication connection with a router (not shown) (e.g., a mobile wireless LAN router).

[0052] The mapping unit 120 performs a mapping process of adding attribute values ​​detected from the two-dimensional image to a plurality of three-dimensional points that make up the three-dimensional point cloud, using the two-dimensional image and the three-dimensional point cloud acquired by the communication unit 110. Details of the mapping unit will be described later with reference to FIGS. 6 and 7.

[0053] The encoding unit 130 generates an encoded stream by encoding the three-dimensional point cloud obtained by the mapping process. The encoding unit 130 may store the generated encoded stream in the external memory 140. The encoding unit 130 may also cause the communication unit 110 to transmit the generated encoded stream to the vehicle 300 via the communication network 400. Details of the encoding process by the encoding unit 130 will be described later using Figures 8 and 9.

[0054] The mapping unit 120 and the encoding unit 130 may each be realized by a processor and memory, or by a dedicated circuit, i.e., the mapping unit 120 and the encoding unit 130 may each be realized by software or hardware.

[0055] The external memory 140 may store information required by the processor, such as a program, etc. The external memory 140 may also store data generated by the processor's processing, such as a three-dimensional point cloud with attribute values ​​added, an encoded stream, etc. The external memory 140 may be realized by a non-volatile storage device, such as a flash memory or an HDD (Hard Disk Drive).

[0056] The vehicle 200 includes a camera 210, an acquisition unit 220, a LiDAR 230, an acquisition unit 240, a memory 250, and a communication unit 260.

[0057] Camera 210 captures a three-dimensional object in a three-dimensional space around vehicle 200 to obtain a two-dimensional image. Camera 210 is disposed at a predetermined position and in a predetermined orientation relative to vehicle 200. For example, camera 210 may be disposed inside or outside the cabin of vehicle 200. Camera 210 may also capture an image in front of vehicle 200, an image on the right or left side of vehicle 200, an image behind vehicle 200, or an image in all 360-degree directions of vehicle 200. Camera 210 may be configured with one camera or two or more cameras. Camera 210 may capture images at different times to obtain multiple two-dimensional images. Camera 210 may also capture a moving image including multiple frames as multiple two-dimensional images at a predetermined frame rate.

[0058] Acquisition unit 220 is provided corresponding to camera 210, stores two-dimensional images obtained by capturing images at different times by camera 210, and outputs the stored two-dimensional images to communication unit 260. The two-dimensional images are associated with the times at which the two-dimensional images were captured. Acquisition unit 220 may be a processing unit built into camera 210. In other words, camera 210 may have the functions of acquisition unit 220.

[0059] The LiDAR 230 obtains a three-dimensional point cloud by sensing three-dimensional objects in a three-dimensional space around the vehicle 200. The LiDAR 230 is a laser sensor that is held by the vehicle 200 and detects the distance to a three-dimensional object within a detection range of 360 degrees in the horizontal direction of the vehicle 200 and a predetermined angle (e.g., 30 degrees) in the vertical direction. The LiDAR 230 is an example of a distance sensor. The LiDAR 230 measures the distance from the LiDAR 230 to the three-dimensional object by emitting a laser beam into the surroundings and detecting the laser beam reflected by the surrounding objects. The LiDAR 230 measures the distance, for example, on the order of centimeters. In this way, the LiDAR 230 detects the three-dimensional coordinates of each of multiple points on the terrain surface around the vehicle 200. In other words, the LiDAR 230 detects the three-dimensional shape of the terrain including the objects around the vehicle 200 by detecting the multiple three-dimensional coordinates of the surrounding terrain surface. In this way, LiDAR 230 obtains a three-dimensional point cloud that is composed of three-dimensional coordinates of a plurality of points and indicates the three-dimensional shape of the terrain including objects around vehicle 200. Note that the distance sensor is not limited to LiDAR 230, and may be other distance sensors such as a millimeter-wave radar, an ultrasonic sensor, a ToF (Time of Flight) camera, a stereo camera, or the like.

[0060] The acquisition unit 240 is provided corresponding to the LiDAR 230, stores a three-dimensional point cloud obtained by sensing by the LiDAR 230 at a plurality of different timings, and outputs the stored three-dimensional point cloud to the communication unit 260. The three-dimensional point cloud is associated with the timing at which the three-dimensional point cloud was sensed. The acquisition unit 240 may be a processing unit built into the LiDAR 230. In other words, the LiDAR 230 may have the function of the acquisition unit 240.

[0061] Memory 250 stores the position and attitude of camera 210 with respect to vehicle 200, and the position and attitude of LiDAR 230 with respect to vehicle 200. Memory 250 may also store the relative position and attitude of camera 210 with respect to LiDAR 230. The relative position and attitude of camera 210 with respect to LiDAR 230 may be detected by matching the intensity of laser light output by LiDAR 230 with a two-dimensional image captured by camera 210, the two-dimensional image including reflected light of laser light from LiDAR 230 reflected from a three-dimensional object, using a method such as mutual information, or may be detected using equipment other than camera 210 and LiDAR 230. Memory 250 is realized by, for example, a non-volatile storage device.

[0062] The communication unit 260 communicates with the three-dimensional point cloud data generation device 100 via the communication network 400. The communication unit 260 transmits the two-dimensional image and the three-dimensional point cloud to the three-dimensional point cloud data generation device 100 by communication via the communication network 400. The communication unit 260 may also transmit the relative position and orientation of the camera 210 with respect to the LiDAR 230 to the three-dimensional point cloud data generation device 100 by communication via the communication network 400.

[0063] The communication unit 260 is realized by a communication interface that can be communicatively connected to the communication network 400. Specifically, the communication unit 260 is communicatively connected to the communication network 400 through a communication connection with a base station 410 of a mobile communication system. The communication unit 260 may be realized by a wireless communication interface that complies with communication standards used in mobile communication systems such as the third generation mobile communication system (3G), the fourth generation mobile communication system (4G), LTE (registered trademark), and the fifth generation mobile communication system (5G).

[0064] Next, the functional configuration of the vehicle 300 will be described.

[0065] FIG. 3 is a block diagram showing an example of the functional configuration of a vehicle serving as a client device.

[0066] The vehicle 300 includes a camera 310 , an acquisition unit 320 , a communication unit 330 , a decoding unit 340 , a position estimation unit 350 , a control unit 360 , and an external memory 370 .

[0067] Camera 310 is disposed at a predetermined position and in a predetermined attitude relative to vehicle 300, and captures images of three-dimensional objects in the three-dimensional space around vehicle 300. Camera 310 differs from camera 210 in that it is disposed on vehicle 300, but otherwise has the same configuration as camera 210, and therefore can be described by substituting camera 210 for camera 310 and vehicle 200 for vehicle 300 in the description of camera 210. For this reason, a detailed description of camera 310 will be omitted.

[0068] Acquisition unit 320 is provided corresponding to camera 310, stores two-dimensional images obtained by capturing images with camera 310, and outputs the stored two-dimensional images to position estimation unit 350. Acquisition unit 320 may be a processing unit built into camera 310. In other words, camera 310 may have the functions of acquisition unit 320.

[0069] The communication unit 330 communicates with the 3D point cloud data generation device 100 via the communication network 400. The communication unit 330 acquires the encoded stream from the 3D point cloud data generation device 100 by receiving it through communication via the communication network 400. In addition, the communication unit 330 may transmit a threshold value related to importance (described later) to the 3D point cloud data generation device 100 via the communication network 400 in order to reduce the communication load.

[0070] The communication unit 330 may acquire all of the multiple 3D points included in the 3D point cloud stored in the 3D point cloud data generation device 100. When acquiring the 3D point cloud from the 3D point cloud data generation device 100, the communication unit 330 may use a GPS (Global Positioning System) (not shown). Alternatively, all of the above 3D points may be acquired as a 3D point cloud of a predetermined area based on the position of the vehicle 300 detected by a position detection device with low accuracy, such as a 3D point cloud of a predetermined area. This eliminates the need to acquire a 3D point cloud of the periphery of the moving object's path, such as all roads in the world, and reduces the amount of 3D point cloud data to be acquired. Furthermore, the communication unit 330 does not need to acquire all 3D points. Instead, the communication unit 330 may transmit the threshold to the 3D point cloud data generation device 100 to acquire, from among the above 3D points, multiple 3D points with importance greater than the threshold, and not acquire multiple 3D points with importance equal to or less than the threshold.

[0071] The communication unit 330 is realized by a communication interface that can be communicatively connected to the communication network 400. Specifically, the communication unit 330 is communicatively connected to the communication network 400 through a communication connection with a base station 410 of a mobile communication system. The communication unit 330 may be realized by a wireless communication interface that complies with communication standards used in mobile communication systems such as the third generation mobile communication system (3G), the fourth generation mobile communication system (4G), LTE (registered trademark), and the fifth generation mobile communication system (5G).

[0072] The decoding unit 340 decodes the coded stream acquired by the communication unit 330 to generate a three-dimensional point group to which attribute values ​​are added.

[0073] The position estimation unit 350 estimates the position and attitude of the vehicle 300 by using the two-dimensional image obtained by the camera 310 and the three-dimensional point cloud with attribute values ​​added, obtained by decoding in the decoding unit 340, to estimate the position and attitude of the camera 310 in the three-dimensional point cloud.

[0074] The control unit 360 controls the operation of the vehicle 300. Specifically, the control unit 360 performs automatic driving or driving assistance of the vehicle 300 by controlling the steering that steers the wheels, the engine that drives the wheels to rotate, a power source such as a motor, and the brakes that brake the wheels. For example, the control unit 360 determines a route indicating which roads the vehicle 300 will travel on, using the current position of the vehicle 300, the destination of the vehicle 300, and information about surrounding roads. The control unit 360 also controls the steering, power source, and brakes so that the vehicle 300 travels on the determined route.

[0075] The decoding unit 340, the position estimation unit 350, and the control unit 360 may be realized by a processor or a dedicated circuit, that is, the decoding unit 340, the position estimation unit 350, and the control unit 360 may be realized by software or hardware.

[0076] The external memory 370 may store information required by the processor, such as a program. The external memory 370 may also store data generated by the processor's processing, such as a three-dimensional point cloud with attribute values ​​added, an encoded stream, etc. The external memory 370 may be a non-volatile storage device, such as a flash memory or an HDD (Hard Disk Drive). This is achieved by:

[0077] Next, the operation of the position estimation system 1 will be described.

[0078] Fig. 4 is a sequence diagram showing an example of the operation of the position estimation system, and Fig. 5 is a diagram for explaining the operation of the position estimation system.

[0079] First, the vehicle 200 transmits, via the communication network 400, to the three-dimensional point cloud data generation device 100 (S1), a two-dimensional image obtained by the camera 210 capturing an image of a three-dimensional object in the three-dimensional space around the vehicle 200, and a three-dimensional point cloud obtained by the LiDAR 230 sensing the three-dimensional object in the three-dimensional space around the vehicle 200. Note that detailed processing in the vehicle 200 has already been explained using Fig. 2 and will not be repeated here.

[0080] Next, in the 3D point cloud data generation device 100, the communication unit 110 acquires the 2D image and the 3D point cloud from the vehicle 200 via the communication network 400 (S3).

[0081] Then, the mapping unit 120 performs a mapping process using the two-dimensional image and the three-dimensional point cloud acquired by the communication unit 110 to add attribute values ​​detected from the two-dimensional image to multiple three-dimensional points that make up the three-dimensional point cloud (S4).

[0082] Next, the details of the mapping process will be described later together with the detailed configuration of mapping section 120 using FIGS.

[0083] Fig. 6 is a block diagram showing an example of the functional configuration of a mapping unit, and Fig. 7 is a flowchart showing an example of detailed mapping processing.

[0084] As shown in FIG. 6, the mapping unit 120 includes a feature detection module 121, a feature matching module 122, a point cloud registration module 123, a triangulation module 124, a position and orientation calculation module 125, and a memory 126.

[0085] The following description is based on the assumption that the mapping unit 120 receives a two-dimensional image and a three-dimensional point cloud sensed at a timing corresponding to the timing at which the two-dimensional image was captured. That is, the mapping unit 120 performs mapping processing by associating the sensed two-dimensional image and the sensed three-dimensional point cloud at corresponding timings. The three-dimensional point cloud obtained at the timing at which the two-dimensional image was captured may be the three-dimensional point cloud sensed at the timing closest to the timing at which the two-dimensional image was captured, or the most recent three-dimensional point cloud sensed before the timing at which the two-dimensional image was captured. Furthermore, the two-dimensional image and the three-dimensional point cloud obtained at corresponding timings may be the two-dimensional image and the three-dimensional image obtained at the same timing if the timing of the image capture by the camera 210 and the timing of the sensing by the LiDAR 230 are synchronized.

[0086] The feature detection module 121 detects feature points in each of the multiple two-dimensional images acquired by the communication unit 110 (S11). The feature detection module 121 detects feature quantities at the feature points. The feature quantities may be, for example, ORB (Oriented FAST and Rotated BRIEF), SIF (Simple Interpolated Feature Extraction), or the like. T, DAISY, etc. The feature amount may be represented by a 256-bit data string. The feature amount may be the luminance value of each pixel in the two-dimensional image, or a color expressed by RGB values, etc. The feature amount of a feature point is an example of an attribute value of a two-dimensional image corresponding to a position on the two-dimensional image. The feature amount is not limited to the feature amount of a feature point as long as it corresponds to a position on the two-dimensional image, but may be a feature amount calculated for each of a plurality of regions. The plurality of regions are regions that make up the two-dimensional image, and each region may be one pixel of the two-dimensional image, or may be a block composed of a set of a plurality of pixels.

[0087] The position and orientation calculation module 125 acquires an accurate three-dimensional map from the memory 126 or the external memory 140, and determines the position and orientation of the LiDAR 230 in the accurate three-dimensional map using the acquired accurate three-dimensional map and the three-dimensional point cloud acquired by the communication unit 110, that is, the three-dimensional point cloud that is the detection result of the LiDAR 230 (S12). The position and orientation calculation module 125 uses a parameter such as ICP (Iterative Closest Point), for example. Using a turn matching algorithm, the 3D point cloud detected by the LiDAR 230 is matched with the 3D point cloud constituting the accurate 3D map. As a result, the position and orientation calculation module 125 identifies the position and orientation of the LiDAR 230 in the accurate 3D map. Note that the position and orientation calculation module 125 may obtain the accurate 3D map from a higher-level server external to the 3D point cloud data generation device 100.

[0088] Next, the position and orientation calculation module 125 acquires the position and orientation of the LiDAR 230 in the accurate three-dimensional map and the relative position and orientation of the camera 210 with respect to the LiDAR 230 stored in the memory 250. Then, the position and orientation calculation module 125 calculates the position and orientation of the camera 210 at the time when each of the multiple two-dimensional images was captured, using the position and orientation of the LiDAR 230 in the acquired accurate three-dimensional map and the relative position and orientation of the camera 210 with respect to the LiDAR 230 stored in the memory 250 (S13).

[0089] Next, the feature matching module 122 matches a pair of the multiple two-dimensional images, i.e., the feature points in each of the two two-dimensional images, using the feature points of each of the multiple two-dimensional images detected by the feature detection module 121 and the position and orientation of the camera 210 at the timing when each of the multiple two-dimensional images was captured, calculated by the position and orientation calculation module 125 (S14). For example, as shown in "Matching / Triangulation" in FIG. 5, the feature matching module 122 matches a feature point P1 among the multiple feature points detected by the feature detection module 121 in a two-dimensional image I1 captured by the camera 210 at timing T with a feature point P2 among the multiple feature points detected in a two-dimensional image I2 captured by the camera 210 at timing T+1 after timing T. For example, timing T+1 may be the timing of capturing the image next to timing T. In this way, the feature matching module 122 associates the multiple feature points detected in each of the multiple two-dimensional images between different two-dimensional images. In this way, the feature matching module 122 uses the multiple feature points detected for each of the multiple two-dimensional images to match corresponding feature points in two of the multiple two-dimensional images, thereby outputting multiple pairs of matched feature points.

[0090] Next, for each of the multiple feature point pairs matched by the feature matching module 122, the triangulation module 124 calculates the three-dimensional position of the matched feature point pair on the two-dimensional image corresponding to the position of the matched feature point pair on the two-dimensional image, for example, by triangulation using the positions of the two feature points constituting the pair on the two-dimensional image and the position and orientation of the camera 210 when each of the two two-dimensional images from which the feature points of the pair were obtained was captured (S15). For example, as shown in "Matching / Triangulation" in FIG. 5, the triangulation module 124 calculates the three-dimensional position P10 by triangulating the feature point P1 detected in the two-dimensional image I1 and the feature point P2 detected in the two-dimensional image I2. The position and orientation of the camera 210 used here are the position and orientation of the camera 210 at the time when each of the two two-dimensional images from which the feature point pairs were obtained was captured.

[0091] Next, the point cloud registration module 123 identifies a plurality of first three-dimensional points, which are three-dimensional points corresponding to the three-dimensional positions calculated by the triangulation module 124, from among the plurality of three-dimensional points constituting the first three-dimensional point cloud as an accurate three-dimensional map (S16). The point cloud registration module 123 may, for example, identify a three-dimensional point that is closest to one of the plurality of three-dimensional points constituting the first three-dimensional point cloud as the first three-dimensional point. For example, the point cloud registration module 123 may identify, as the first three-dimensional points, one or more three-dimensional points that are included within a predetermined range based on the one three-dimensional position from among the plurality of three-dimensional points constituting the first three-dimensional point cloud. In other words, the point cloud registration module 123 may identify, as the first three-dimensional points corresponding to one three-dimensional position, a plurality of three-dimensional points that satisfy a predetermined condition based on the one three-dimensional position.

[0092] Then, the point cloud registration module 123 generates a second three-dimensional point cloud composed of a plurality of second three-dimensional point clouds by adding, for each of the plurality of pairs, the feature amounts of the two feature points constituting the pair to the plurality of identified first three-dimensional points (S17). Note that, for each of the plurality of pairs, the point cloud registration module 123 may generate the second three-dimensional point cloud by (i) adding, for each of the plurality of pairs, to the plurality of identified first three-dimensional points, the feature amount of one of the two feature points constituting the pair, or (ii) adding, for each of the plurality of pairs, a feature amount calculated from the two feature amounts of the two feature points constituting the pair. In other words, the point cloud registration module 123 may generate the second three-dimensional point cloud by adding, for each of the plurality of identified first three-dimensional points, a feature amount based on the two feature amounts of the plurality of pairs. Note that, for example, the one feature amount calculated from the two feature amounts is an average value. It should be noted that the "feature amount" referred to here is an example of an attribute value, and therefore the processing in the point cloud registration module 123 can also be applied to other attribute values.

[0093] In steps S16 and S17, the point cloud registration module 123 identifies a 3D point corresponding to the 3D position P10 obtained by the triangulation module 124, for example, as shown in "Correspondence" in FIG. 5, and performs correspondence by adding the feature amounts of each of the feature points P1 and P2 corresponding to the 3D position P10 to the 3D point in the accurate 3D map corresponding to the 3D position P10. As a result, a 3D point cloud C1 is generated, including a plurality of 3D points, each having the feature amounts of the feature points P1 and P2 added thereto, as shown in FIG. 5. That is, the 3D points added with attribute values ​​included in the 3D point cloud C1 are composed of 3D coordinates and the illuminance (brightness) and feature amounts of the feature points P1 and P2 in the 2D images as attribute values. That is, in addition to the feature amounts of the feature points being added to the 3D points, attribute values ​​such as the illuminance (brightness) or color (e.g., RGB values) of the feature points in the 2D images may also be added.

[0094] The point cloud registration module 123 may generate a second 3D point by adding multiple attribute values ​​to one first 3D point. For example, as described above, the point cloud registration module 123 may add, as multiple attribute values, feature quantities of multiple feature points detected from multiple 2D images to one first 3D point. Furthermore, the point cloud registration module 123 may add, as multiple attribute values, attribute values ​​of different types to one first 3D point.

[0095] An accurate three-dimensional map may be stored in memory 126. An accurate three-dimensional map may be stored in advance in memory 126, or an accurate three-dimensional map received from a higher-level server by communication unit 110 may be stored in memory 126. Memory 126 is realized by a non-volatile storage device such as a flash memory or an HDD (Hard Disk Drive), for example.

[0096] In this way, by performing the mapping process in step S4, the point cloud registration module 123 generates a second three-dimensional point cloud including a plurality of second three-dimensional points, which are a plurality of three-dimensional points to which feature amounts of feature points have been added.

[0097] Returning to FIG. 4, after step S4, the encoding unit 130 generates an encoded stream by encoding the second 3D point group generated by the mapping unit 120 (S5).

[0098] Then, the communication unit 110 transmits the encoded stream generated by the encoding unit 130 to the vehicle 300 (S6).

[0099] In the vehicle 300, the communication unit 330 acquires the encoded stream from the 3D point cloud data generation device 100 (S7).

[0100] Next, the decoding unit 340 obtains a second three-dimensional point cloud by decoding the encoded stream obtained by the communication unit 330 (S8). That is, the vehicle 300 obtains a three-dimensional point cloud including a plurality of three-dimensional points to which attribute values ​​of two-dimensional images obtained by capturing a three-dimensional object are respectively added in advance.

[0101] Next, the position estimation unit 350 estimates the position and attitude of the vehicle 300 by estimating the position and attitude of the camera 310 in the three-dimensional point cloud using the two-dimensional image obtained by the camera 310 and the three-dimensional point cloud with attribute values ​​added, obtained by decoding in the decoding unit 340 (S9).

[0102] The position estimation process performed by the position estimation unit 350 will now be described in detail.

[0103] As shown in FIG. 3, the position estimation unit 350 includes a feature detection module 351, a feature matching module 352, and a memory 353.

[0104] The feature detection module 351 acquires a second two-dimensional image of the surroundings of the vehicle 300, captured by the camera 310 equipped on the vehicle 300. Then, the feature detection module 351 detects, from the acquired second two-dimensional image, a plurality of second attribute values, which are attribute values ​​of the second two-dimensional image corresponding to positions on the second two-dimensional image. The processing in the feature detection module 351 is similar to the processing in the feature detection module 121 of the mapping unit 120 in the vehicle 200. Note that the processing in the feature detection module 351 does not need to be the same for all of the processing in the feature detection module 121 of the mapping unit 120. For example, when the feature detection module 121 detects a plurality of types of first attribute values ​​as attribute values, the feature detection module 351 may detect one or more types of attribute values ​​included in the plurality of types as second attribute values.

[0105] Next, for each of the plurality of second attribute values ​​detected by the feature detection module 351, the feature matching module 352 identifies one or more fifth 3D points among the plurality of second 3D points corresponding to the second attribute value, thereby generating one or more combinations of the second attribute value and one or more fifth 3D points. Specifically, for each of the plurality of two-dimensional positions on the two-dimensional image corresponding to the plurality of attribute values ​​detected by the feature detection module 351, the feature matching module 352 identifies one or more second 3D points to which an attribute value closest to the attribute value associated with the two-dimensional position is assigned as the fifth 3D point. In this way, the feature matching module 352 generates one or more combinations of the two-dimensional positions on the two-dimensional image and one or more fifth 3D points. The feature matching module 352 acquires the position and orientation of the camera 310 relative to the vehicle 300 from the memory 353. The feature matching module 352 calculates the position and orientation of the vehicle 300 using the one or more identified combinations and the acquired position and orientation of the camera 310 relative to the vehicle 300. Note that the more combinations the feature matching module 352 generates, the more accurately it can calculate the position and orientation of the vehicle 300.

[0106] The memory 353 may store the position and orientation of the camera 310 relative to the vehicle 300. The memory 353 may store a second three-dimensional point cloud obtained by decoding by the decoding unit 340. The memory 353 is realized by, for example, a non-volatile storage device such as a flash memory or an HDD (Hard Disk Drive).

[0107] According to the three-dimensional point cloud data generation method of the present embodiment, the three-dimensional point cloud data generation device 100 generates a second three-dimensional point cloud including second three-dimensional points obtained by adding, to the three-dimensional points corresponding to positions on the two-dimensional image, feature amounts of feature points as attribute values ​​of the two-dimensional image corresponding to positions on the two-dimensional image, among the plurality of three-dimensional points constituting the first three-dimensional point cloud obtained using a distance sensor such as the LiDAR 230. Therefore, the vehicle 300 as a position estimation device that estimates its own position can efficiently estimate its own position without sensing the three-dimensional objects around the vehicle 300, by capturing two-dimensional images of the three-dimensional objects around the vehicle 300 and comparing the feature amounts of the feature points corresponding to the positions of the captured two-dimensional images with the feature amounts of the feature points added to each second three-dimensional point of the second three-dimensional point cloud obtained from the three-dimensional point cloud data generation device 100.

[0108] Furthermore, in the 3D point cloud data generation method according to this embodiment, a pair of feature points is identified by matching feature points of two 2D images, a 3D point corresponding to the position is identified using the positions of the feature points of the pair corresponding to the 2D images, and feature amounts of the two feature points constituting the pair are added to the identified 3D point. As a result, the 3D point corresponding to the feature points can be identified with high accuracy.

[0109] Furthermore, in the 3D point cloud data generation method according to this embodiment, in generating the second 3D point cloud, the feature amounts of each of the plurality of feature points are added to one identified first 3D point to generate the second 3D point. In this way, since a plurality of feature points are added to one first 3D point, in position estimation in vehicle 300, it is possible to accurately associate the feature points of the two-dimensional image corresponding to positions on the two-dimensional image acquired for position estimation with the first 3D point.

[0110] Furthermore, in the 3D point cloud data generation method according to this embodiment, the multiple attribute values ​​added to one first 3D point in generating the second 3D point cloud are each attribute values ​​detected from multiple 2D images. Here, one first 3D point is a point that can be captured from multiple viewpoints. Furthermore, in position estimation in vehicle 300, camera 310 captures two-dimensional images at different positions. Therefore, the two-dimensional images obtained by camera 310 are captured from multiple different viewpoints. As a result, by adding multiple attribute values ​​detected from multiple two-dimensional images, even when two-dimensional images captured from multiple different viewpoints are used in position estimation, it is possible to easily associate the attribute values ​​obtained from each two-dimensional image with the first 3D point. Therefore, position estimation can be easily performed.

[0111] Furthermore, in the 3D point cloud data generation method according to this embodiment, the multiple attribute values ​​added to one first 3D point in generating the second 3D point cloud are attribute values ​​with different attribute types. In this way, since the attribute values ​​added to one first 3D point are multiple different types of attribute values, in position estimation in vehicle 300, it is possible to accurately associate the attribute values ​​of the two-dimensional image corresponding to the position on the two-dimensional image acquired for position estimation with the first 3D point.

[0112] Furthermore, according to the position estimation method of this embodiment, the attribute values ​​of the two-dimensional image of the three-dimensional object are added to the three-dimensional point cloud in advance. Therefore, without sensing a new three-dimensional point cloud, it is possible to efficiently estimate the self-position of the vehicle 300 by simply capturing a new second two-dimensional image of the surroundings of the vehicle 300 and comparing the second attribute value corresponding to the position of the captured second two-dimensional image and the combination of one or more fifth three-dimensional points corresponding to the second attribute value with the first attribute value added to each three-dimensional point of the three-dimensional point cloud.

[0113] (Variation) (Variation 1) The 3D point cloud data generation device 100 according to the above embodiment may further calculate an importance for each of the plurality of second 3D points included in the second 3D point cloud, add the calculated importance to the corresponding second 3D point, and generate a third 3D point cloud including the plurality of third 3D points obtained by adding the importance. Furthermore, the 3D point cloud data generation device 100 may narrow down the number of third 3D points to be transmitted to the vehicle 300 from the plurality of third 3D points included in the third 3D point cloud according to the added importance, and transmit the narrowed number of third 3D points to the vehicle 300.

[0114] The operation will be described with reference to Fig. 8. Fig. 8 is a flowchart showing an example of the operation of the three-dimensional point cloud data generating device in Modification 1. This operation is performed in place of step S5 in the sequence diagram of Fig. 4.

[0115] In the 3D point cloud data generation device 100, the point cloud registration module 123 of the mapping unit 120 further calculates the importance of each of the one or more second 3D points based on the attribute value added to the second 3D point (S21). Then, the point cloud registration module 123 adds the calculated importance to each of the one or more second 3D points to generate a third 3D point cloud including one or more 3D points each having an attribute value and an importance. Details of the importance calculation process will be described later using FIG. 9.

[0116] Next, the encoding unit 130 of the 3D point cloud data generation device 100 generates an encoded stream using the third 3D point cloud (S22). Details of the encoded stream generation process will be described later with reference to FIGS. 10 and 11.

[0117] FIG. 9 is a flowchart showing an example of detailed importance calculation processing.

[0118] The point cloud registration module 123 performs a calculation process of importance for each of the second 3D points included in the generated second 3D point cloud. The following describes the process performed for one second 3D point. In the calculation process of importance, the same process is performed for each of all second 3D points.

[0119] First, the point cloud registration module 123 calculates the number of two-dimensional images in which the second three-dimensional point to be processed can be seen (S31). The point cloud registration module 123 may calculate the number by counting the number of two-dimensional images in which the second three-dimensional point to be processed is reflected among multiple two-dimensional images.

[0120] Furthermore, the point cloud registration module 123 calculates a matching error in matching of a plurality of feature points associated with the second 3D point to be processed (S32). The point cloud registration module 123 may obtain the matching error from the feature matching module 122.

[0121] In addition, the point cloud registration module 123 calculates the matching error in matching between the second 3D point to be processed (i.e., the 3D point in the accurate 3D map) and the feature point associated with the second 3D point (S33).

[0122] The point cloud registration module 123 calculates the importance of the second 3D point to be processed using the number of 2D images in which the second 3D point to be processed can be seen, the matching error between the multiple feature points of the 2D images, and the matching error between the 3D map and the feature point, calculated in steps S31 to S33 (S34). For example, the point cloud registration module 123 calculates the importance to a higher value as the number of 2D images in which the second 3D point to be processed can be seen increases. Also, for example, the point cloud registration module 123 calculates the importance to a lower value as the matching error between the multiple feature points of the 2D images increases. Also, for example, the point cloud registration module 123 calculates the importance to a lower value as the matching error between the 3D map and the feature point increases. In this way, the importance is an index that indicates that the higher the value, the more important it is.

[0123] The point cloud registration module 123 acquires the three-dimensional coordinates of the second three-dimensional point to be processed (S35), and acquires the feature points of the two-dimensional image added to the second three-dimensional point to be processed (S36).

[0124] Then, the point cloud registration module 123 generates a third 3D point by adding the calculated importance and feature points to the acquired 3D coordinates (S37). Note that, although the 3D coordinates and feature points are respectively acquired from the second 3D point to be processed in steps S35 and S36, this is not limitative, and the third 3D point may be generated by adding the calculated importance to the second 3D point to be processed in step S37.

[0125] The point cloud registration module 123 performs the processes of steps S31 to S37 for each of the second 3D points to generate a third 3D point cloud, which is a 3D point cloud to which importance is further added. As shown in Fig. 10, the third 3D point includes 3D coordinates, feature points of each of N 2D images in which the second 3D point can be seen, and importance.

[0126] Here, among the N feature quantities, if any have similar values, they may be integrated. For example, if the difference between the values ​​F0 and F1 of two feature quantities is equal to or less than a threshold, they may be integrated into a single feature quantity whose value is the average value of F0 and F1. This reduces the amount of data in the second 3D point cloud. Also, an upper limit may be set on the value of the number N of feature quantities that can be assigned to the second 3D point. For example, if the number of feature quantities is greater than N, the top N may be selected using the feature quantity values, and the selected N feature quantities may be assigned to the 3D point.

[0127] Next, the details of the encoding process will be described later together with a detailed configuration of the encoding unit 130 using FIGS.

[0128] Fig. 11 is a block diagram showing an example of the functional configuration of the encoding unit, and Fig. 12 is a flowchart showing an example of detailed encoding processing.

[0129] As shown in FIG. 11, the encoding unit 130 includes a feature sorting module 131, a feature combination module 132, a memory 133, and an entropy encoding module 134.

[0130] In the encoding process by the encoding unit 130, the process is performed on a plurality of third 3D points included in the third 3D point group.

[0131] The feature sorting module 131 sorts the third 3D points included in the third 3D point group in descending order of the importance assigned to each of the third 3D points (S41).

[0132] Next, the feature combination module 132 starts a loop 1 in which the following steps S43 and S44 are executed for each of the plurality of third 3D points (S42). The feature combination module 132 executes the loop 1 in descending order of importance.

[0133] The feature combination module 132 determines whether the importance assigned to the third 3D point to be processed exceeds a threshold (S43). The threshold is a value received from the vehicle 300. The threshold may be set to a different value depending on the specifications of the vehicle 300. In other words, the threshold may be set to a larger value as the information processing capability and / or detection capability of the vehicle 300 increases.

[0134] When the feature combination module 132 determines that the importance attached to the third 3D point to be processed exceeds the threshold (Yes in S43), it executes the process of step S44.

[0135] On the other hand, if the feature combination module 132 determines that the importance assigned to the third 3D point to be processed is less than the threshold (No in S43), it executes the process of step S45, thereby ending loop 1.

[0136] In step S44, the feature combination module 132 adds the third 3D point to be processed as data to be encoded (S44). After step S44, the feature combination module 132 executes loop 1 with the next third 3D point to be processed.

[0137] In this way, by executing steps S43 and S44, feature combination module 132 extracts a plurality of fourth 3D points to which importance exceeding the threshold is assigned from among a plurality of third 3D points.

[0138] Note that the process of step S41 does not necessarily have to be executed. In this case, loop 1 is executed for all the third 3D points, and even if step S43 returns No, loop 1 is repeated for the next third 3D point.

[0139] In step S45, the entropy encoding module 134 performs entropy encoding on the third 3D points that are the encoding target data, and generates an encoded stream (S45). For example, the entropy encoding module 134 may represent the encoding target data in an octree structure, binarize it, and perform arithmetic encoding to generate the encoded stream.

[0140] The generated encoded stream is transmitted to the vehicle 300 that transmitted the threshold value in step S6 of FIG.

[0141] According to the 3D point cloud data generation method of Variation 1, the generation of the second 3D points further includes (i) calculating the importance of each of the second 3D points based on the attribute values ​​added to the second 3D points, and (ii) adding the calculated importance to the second 3D points to generate a third 3D point cloud including one or more third 3D points, each of which has an attribute value and an importance. Therefore, for example, third 3D points with greater importance can be used preferentially, and the amount of 3D point cloud data used in processing can be adjusted so as not to adversely affect the processing accuracy.

[0142] Furthermore, in the 3D point cloud data generation method according to the first modification, a threshold value is received from the vehicle 300, which is a client device, and a plurality of fourth 3D points, which have an importance level exceeding the received threshold value, are extracted from the one or more third 3D points, and a fourth 3D point cloud including the extracted plurality of fourth 3D points is transmitted to the client device. Therefore, the amount of 3D point cloud data to be transmitted can be adjusted in response to a request from the vehicle 300.

[0143] Although the encoding unit 130 of the 3D point cloud data generation device 100 in the first modification acquires the threshold value from the vehicle 300, it may acquire the number of 3D points from the vehicle 300. In this case, the encoding unit 130 may encode the third 3D point of the acquired number of 3D points as encoding target data in descending order of importance.

[0144] (Variation 2) In the above-described first modification, the 3D point cloud data generation device 100 receives a threshold value from the vehicle 300, which is a client device, and thereby excludes third 3D points whose importance is equal to or less than the threshold value from the data to be encoded. However, this is not limiting, and the encoded stream may be generated by encoding all of the plurality of third 3D points as the data to be encoded. In other words, the 3D point cloud data generation device 100 may transmit data obtained by encoding all of the plurality of third 3D points to the vehicle 300.

[0145] In this case, the vehicle 300 obtains a plurality of third 3D points by decoding the received encoded stream. The vehicle 300 may then use, among the plurality of third 3D points, a third 3D point having an importance exceeding a threshold value in a process of estimating the self-position of the vehicle 300.

[0146] Note that, when the communication unit 330 acquires all 3D points, the decoding unit 340 of the vehicle 300 may prioritize decoding of all 3D points in descending order of importance. This allows the decoding unit 340 to decode only the required number of 3D points, thereby reducing the processing load required for the decoding process. Furthermore, in this case, if the encoded stream is encoded in descending order of importance, the decoding unit 340 may decode the encoded stream and stop the decoding process when it decodes a 3D point with an importance equal to or less than a threshold, or may stop the decoding process when it acquires the required number of 3D points. This allows the 3D point cloud data generation device 100 to reduce the processing load required for the decoding process without having to generate an encoded stream for each vehicle 300. This reduces the processing load required for the 3D point cloud data generation device 100 to generate an encoded stream.

[0147] Note that vehicle 300 does not need to use the multiple third 3D points in the process of estimating its own position, and may use them in the process of reconstructing a 3D image of a 3D map. That is, vehicle 300 as a client device may switch the threshold depending on the application of vehicle 300. For example, when vehicle 300 estimates its own position, it may determine that only 3D points with high importance are necessary and set the threshold to a first threshold, whereas when the client draws a map, it may determine that 3D points with low importance are also necessary and set the threshold to a second threshold smaller than the first threshold.

[0148] (others) The mapping unit 120 may add the three-dimensional position calculated by matching the feature points in the two two-dimensional images to an accurate three-dimensional map as a three-dimensional point having the feature amount of the feature point.

[0149] The mapping unit 120 may match feature points in one two-dimensional image with three-dimensional points in an accurate three-dimensional map and add feature quantities of the feature points to the three-dimensional points matched with the feature points. In this case, the feature points may be matched with three-dimensional points in an accurate three-dimensional map by back-projecting the feature points into three-dimensional space using the orientation of the camera 210 or by projecting the three-dimensional points onto the two-dimensional image. In other words, the mapping unit 120 matches feature points with three-dimensional points by identifying three-dimensional points close to three-dimensional positions identified by a pair of matching feature points in two two-dimensional images among multiple two-dimensional images. However, the present invention is not limited to this, and the mapping unit 120 may also match feature points in one two-dimensional image with three-dimensional points. In other words, for each of the detected one or more attribute values, the mapping unit 120 may (i) identify one or more first three-dimensional points among the multiple three-dimensional points that make up the first three-dimensional point cloud, whose positions on the two-dimensional image correspond to the attribute value, and (ii) add the attribute value to the identified one or more first three-dimensional points, thereby generating a second three-dimensional point cloud that includes one or more second three-dimensional points, each having the attribute value.

[0150] Furthermore, the encoding of the 3D points is not limited to entropy encoding, and any encoding method may be applied. For example, the 3D point group that is the data to be encoded may be encoded using an octree structure (octree coding).

[0151] Although the three-dimensional point cloud data generation device 100 is described as being a server separate from the vehicle 200, it may be provided in the vehicle 200. In other words, the processing by the three-dimensional point cloud data generation device 100 may be executed by the vehicle 200. In this case, the second three-dimensional point cloud or the third three-dimensional point cloud obtained by the vehicle 200 may be transmitted to an upper server, and the upper server may collect the second three-dimensional point cloud or the third three-dimensional point cloud from multiple vehicles 200 at various locations.

[0152] In each of the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Here, the software that realizes the 3D point cloud data generation method or position estimation method of each of the above embodiments is a program such as the following.

[0153] In other words, this program causes a computer to execute a three-dimensional point cloud generation method that uses a processor to generate a three-dimensional point cloud containing a plurality of three-dimensional points, and includes acquiring a two-dimensional image obtained by capturing an image of a three-dimensional object using a camera and a first three-dimensional point cloud obtained by sensing the three-dimensional object using a distance sensor, detecting one or more attribute values ​​of the two-dimensional image corresponding to positions on the two-dimensional image from the acquired two-dimensional image, and for each of the detected one or more attribute values, (i) identifying one or more first three-dimensional points among the multiple three-dimensional points that make up the first three-dimensional point cloud to which the position on the two-dimensional image of the attribute value corresponds, and (ii) adding the attribute value to the identified one or more first three-dimensional points to generate a second three-dimensional point cloud containing one or more second three-dimensional points each having the attribute value.

[0154] The program also causes a computer to execute a position estimation method using a processor to estimate the current position of a moving body, the method comprising: acquiring a three-dimensional point cloud including a plurality of three-dimensional points, each having a first attribute value pre-assigned, which is an attribute value of a first two-dimensional image obtained by capturing an image of a three-dimensional object; acquiring a second two-dimensional image of the surroundings of the moving body captured by a camera equipped on the moving body; detecting from the acquired second two-dimensional image one or more second attribute values, which are attribute values ​​of the second two-dimensional image, that correspond to positions on the second two-dimensional image; generating one or more combinations consisting of the second attribute value and the one or more fifth three-dimensional points for each of the detected one or more second attribute values; acquiring the position and orientation of the camera relative to the moving body from a storage device; and calculating the position and orientation of the moving body using the generated one or more combinations and the acquired position and orientation of the camera.

[0155] While the 3D point cloud data generation method, the position estimation method, the 3D point cloud data generation device, and the position estimation device according to one or more aspects of the present disclosure have been described based on the embodiments, the present disclosure is not limited to these embodiments. As long as they do not deviate from the spirit of the present disclosure, various modifications conceivable by those skilled in the art to the present embodiments and configurations constructed by combining components of different embodiments may also be included within the scope of the present disclosure. [Industrial Applicability]

[0156] The present disclosure is useful as a three-dimensional point cloud data generation method, a position estimation method, a three-dimensional point cloud data generation device, and a position estimation device that can be further improved. [Explanation of symbols]

[0157] 1. Location estimation system 100 3D point cloud data generator 110 Communications Department 120 Mapping Section 121, 351 Feature Detection Module 122,352 Feature Matching Module 123 Point Cloud Registration Module 124 Triangulation Module 125 Position and Orientation Calculation Module 126, 133, 250, 353 memory 130 Encoding section 131 Feature Sorting Module 132 Feature Combination Module 134 Entropy Encoding Module 140, 370 external memory 200, 300 vehicles 210, 310 camera 220, 240, 320 acquisition part 230 LiDAR 260, 330 Communications Department 340 Decoding Unit 350 Position estimation part 360 Control Unit 400 Communication Network 410 base station

Claims

1. An information processing system including a memory and a processor, Acquiring a first three-dimensional point cloud sensed by a distance sensor and a two-dimensional image captured by a camera; Detecting attribute values ​​corresponding to each point of the first three-dimensional point cloud from the two-dimensional image; adding the detected attribute values ​​to each point of the first 3D point cloud to generate a 3D point cloud associated with the attribute values; processing the generated three-dimensional point cloud in a manner that allows an amount of information to be adjusted based on a predetermined criterion; configured to perform spatial recognition processing using the processed three-dimensional point cloud. Three-dimensional spatial information processing system.

2. The information processing system, a data generating device that generates a third 3D point cloud associated with the importance by adding an importance calculated based on the attribute value to each point of the generated 3D point cloud; a client device that receives a threshold value for adjusting the amount of data of the third 3D point cloud from the data generating device, and extracts a fourth 3D point cloud including 3D points whose importance exceeds the threshold value based on the received threshold value; the data generator is configured to transmit the extracted fourth 3D point cloud to the client device; The client device performs the spatial recognition process using the received fourth three-dimensional point cloud. The three-dimensional space information processing system according to claim 1 .

3. the spatial recognition processing is performed by a device that executes the spatial recognition processing, The device, a means for externally acquiring a three-dimensional point cloud including a plurality of three-dimensional points to which attribute values ​​of the three-dimensional points included in the generated three-dimensional point cloud have been added in advance, and in which a plurality of different types of attribute values ​​and importance calculated based on the plurality of attribute values ​​are associated with each of the three-dimensional points; means for acquiring a second two-dimensional image of the surroundings of the device, the second image being captured by a camera included in the device; means for detecting, from the acquired second two-dimensional image, an attribute value of the second two-dimensional image corresponding to a position on the second two-dimensional image; means for calculating a position and orientation of the device by comparing attribute values ​​of the three-dimensional points included in the three-dimensional point cloud with attribute values ​​of the detected second two-dimensional image; 3. The three-dimensional space information processing system according to claim 1.

4. the space recognition process is a process of estimating a self-position of a moving object, the processed 3D point cloud has an importance calculated based on the attribute value; In the process of estimating the self-position, the position and orientation of the moving body are calculated using only the three-dimensional points whose importance exceeds a predetermined threshold. The three-dimensional space information processing system according to claim 1 .

5. The attribute value includes brightness information, color information, or a predetermined image feature amount of an object in the two-dimensional image. The three-dimensional space information processing system according to any one of claims 2 to 4.

6. The importance is calculated to be larger as the number of two-dimensional images in which the three-dimensional point is reflected increases, and as the matching error between a plurality of attribute values ​​of the three-dimensional point decreases. The three-dimensional space information processing system according to any one of claims 2 to 5.

Citation Information

Patent Citations

  • Method for transferring and displaying three-dimensional shape data

    JP1997237354A