Information processing device, information processing method, and information processing program
The image data is captured by multiple cameras and integrated and generated comprehensive point cloud information, which solves the problem of insufficient position information of surrounding objects in VSLAM processing and improves detection stability.
Patent Information
- Application Number
- JP2023529380
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-24
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-06-24
AI Technical Summary
In VSLAM processing, the position information of surrounding objects is insufficient, resulting in unstable object position and self-position detection.
Multiple cameras are used to capture image data from different locations, generate first and second point cloud information, and generate comprehensive point cloud information through alignment and integration processing.
Through multi-camera integration technology, the accuracy and stability of the location information of surrounding objects are improved, and the problem of insufficient information in VSLAM processing is solved.
Smart Images

Figure 0007673802000001 
Figure 0007673802000002 
Figure 0007673802000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]
[0002] There is a technology called SLAM (Simultaneous Localization and Mapping) that acquires point cloud information about three-dimensional objects around a moving object such as a vehicle and estimates the position information of the vehicle and surrounding three-dimensional objects. There is also a technology called Visual SLAM (Simultaneous Localization and Mapping: VSLAM) that performs SLAM using images captured by a camera. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-205949 [Patent Document 2] Japanese Patent Application Laid-Open No. 2016-045874 [Patent Document 3] Japanese Patent Application Laid-Open No. 2016-123021 [Patent Document 4] International Publication No. 2019 / 073795 [Patent Document 5] International Publication No. 2020 / 246261 [Non-patent literature]
[0004] [Non-Patent Document 1] “Vision SLAM Using Omni·Directional Visual Scan Matching” 2008 IEEE / RSJ International Conference on Intelligent Robots and Systems Sept. 22·26, 2008 Summary of the Invention [Problem to be solved by the invention]
[0005] However, for example, in VSLAM processing, the position information of surrounding objects obtained by VSLAM processing may be insufficient, which may result in unstable detection of the positions of surrounding objects and the vehicle's own position by VSLAM.
[0006] In one aspect, the present invention aims to provide an information processing device, an information processing method, and an information processing program that solve the problem of insufficient position information of surrounding objects obtained by VSLAM processing. [Means for solving the problem]
[0007] In one aspect, the information processing device disclosed in the present application includes an acquisition unit that acquires first point cloud information based on first image data obtained from a first imaging unit provided at a first position of a moving body, and acquires second point cloud information based on second image data obtained from a second imaging unit provided at a second position different from the first position of the moving body, an alignment processing unit that performs alignment processing between the first point cloud information and the second point cloud information, and an integration processing unit that generates integrated point cloud information using the first point cloud information and the second point cloud information on which the alignment processing has been performed. [Effects of the Invention]
[0008] According to one aspect of the information processing device disclosed in the present application, it is possible to resolve the lack of position information of surrounding objects obtained by VSLAM processing. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the overall configuration of an information processing system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of the information processing device according to the embodiment. [Figure 3]FIG. 3 is a diagram illustrating an example of a functional configuration of the information processing device according to the embodiment. [Figure 4] FIG. 4 is a schematic diagram illustrating an example of environment map information according to the embodiment. [Figure 5] FIG. 5 is a plan view showing an example of a situation in which a mobile object is parked backward in a parking space. [Figure 6] FIG. 6 is a plan view showing an example of the imaging range of an imaging unit provided in front of the moving body when the moving body moves forward. [Figure 7] FIG. 7 is a plan view showing an example of the imaging range of an imaging unit provided behind the moving body when the moving body moves backward. [Figure 8] FIG. 8 is a diagram showing an example of point cloud information relating to the front of a moving object, generated by VSLAM processing, when the moving object moves forward once along a trajectory. [Figure 9] FIG. 9 is a diagram showing an example of point cloud information relating to the rear of a moving object, generated by VSLAM processing, when the moving object moves backward once along the trajectory. [Figure 10] FIG. 10 is a diagram showing an example of integrated point cloud information generated by the integration process in the backward parking of the moving object shown in FIG. [Figure 11] FIG. 11 is a flowchart showing an example of the flow of the integration process shown in FIGS. [Figure 12] FIG. 12 is an explanatory diagram of an asymptotic curve generated by the determination unit. [Figure 13] FIG. 13 is a schematic diagram showing an example of the reference projection plane. [Figure 14] FIG. 14 is a schematic diagram showing an example of a projection shape determined by the determination unit. [Figure 15] FIG. 15 is a schematic diagram illustrating an example of the functional configuration of the integration processing unit and the determination unit. [Figure 16] FIG. 16 is a flowchart showing an example of the flow of information processing executed by the information processing device. [Figure 17]FIG. 17 is a flowchart showing an example of the flow of the point cloud integration process in step S27 of FIG. [Figure 18] FIG. 18 is a diagram showing point cloud information related to the backward VSLAM processing of an information processing device according to a comparative example. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, with reference to the accompanying drawings, embodiments of the information processing device, information processing method, and information processing program disclosed herein will be described in detail. Note that the following embodiments do not limit the disclosed technology. Furthermore, each embodiment can be appropriately combined within a range that does not cause contradiction in the processing content.
[0011] 1 is a diagram showing an example of the overall configuration of an information processing system 1 according to this embodiment. The information processing system 1 includes an information processing device 10, an imaging unit 12, a detection unit 14, and a display unit 16. The information processing device 10, the imaging unit 12, the detection unit 14, and the display unit 16 are connected to each other so as to be able to exchange data or signals.
[0012] In this embodiment, the information processing device 10, the image capturing unit 12, the detection unit 14, and the display unit 16 are mounted on a moving object 2 as an example.
[0013] The moving body 2 is an object that can move. The moving body 2 is, for example, a vehicle, a flyable object (a manned airplane, an unmanned airplane (e.g., a UAV (Unmanned Aerial Vehicle), a drone)), a robot, etc. The moving body 2 is, for example, a moving body that moves through human driving operation, or a moving body that can move automatically (autonomously) without human driving operation. In this embodiment, a case where the moving body 2 is a vehicle will be described as an example. Examples of vehicles include a two-wheeled vehicle, a three-wheeled vehicle, and a four-wheeled vehicle. In this embodiment, a case where the vehicle is an autonomously moving four-wheeled vehicle will be described as an example.
[0014] Note that the information processing device 10, the photographing unit 12, the detection unit 14, and the display unit 16 are not limited to being all mounted on the moving object 2. The information processing device 10 may be mounted on a stationary object. A stationary object is an object fixed to the ground. A stationary object is an object that cannot be moved or an object that is stationary relative to the ground. Examples of stationary objects include traffic lights, parked vehicles, and road signs. Furthermore, the information processing device 10 may be mounted on a cloud server that executes processing on the cloud.
[0015] The photographing unit 12 photographs the periphery of the moving object 2 and acquires photographed image data. In the following description, the photographed image data will be simply referred to as a photographed image. The photographing unit 12 is, for example, a digital camera capable of shooting video. Photographing refers to converting an image of a subject formed by an optical system such as a lens into an electrical signal. The photographing unit 12 outputs the photographed image to the information processing device 10. In addition, in this embodiment, the photographing unit 12 will be described assuming that it is a monocular fisheye camera (for example, with a viewing angle of 195 degrees).
[0016] In this embodiment, an example will be described in which four photographing units 12 (photographing units 12A to 12D) are mounted on a moving object 2. The multiple photographing units 12 (photographing units 12A to 12D) photograph subjects in their respective photographing areas E (photographing areas E1 to E4) and acquire photographed images. Note that the photographing directions of these multiple photographing units 12 are different from each other. Furthermore, the photographing directions of these multiple photographing units 12 are adjusted in advance so that at least a portion of the photographing area E between adjacent photographing units 12 overlaps.
[0017] Furthermore, the four photographing units 12A to 12D are merely an example, and there is no limitation on the number of photographing units 12. For example, if the moving object 2 has a vertically long shape like a bus or truck, it is possible to arrange one photographing unit 12 at the front, rear, front of the right side, rear of the right side, front of the left side, and rear of the left side of the moving object 2, thereby using a total of six photographing units 12. In other words, the number and arrangement positions of the photographing units 12 can be set arbitrarily depending on the size and shape of the moving object 2. Note that the present invention can be realized by providing at least two photographing units 12.
[0018] The detection unit 14 detects position information of each of a plurality of detection points around the moving object 2. In other words, the detection unit 14 detects position information of each of the detection points in the detection area F. The detection points refer to each of the points in real space that are individually observed by the detection unit 14. The detection points correspond to, for example, three-dimensional objects around the moving object 2.
[0019] The position information of the detection point is information indicating the position of the detection point in real space (three-dimensional space). For example, the position information of the detection point is information indicating the distance from the detection unit 14 (i.e., the position of the moving object 2) to the detection point and the direction of the detection point relative to the detection unit 14. These distances and directions can be expressed, for example, by position coordinates indicating the relative position of the detection point relative to the detection unit 14, position coordinates indicating the absolute position of the detection point, or a vector.
[0020] The detection unit 14 may be, for example, a 3D (Three-Dimensional) scanner, a 2D (Two-Dimensional) scanner, a distance sensor (millimeter-wave radar, laser sensor), a sonar sensor that detects objects using sound waves, an ultrasonic sensor, or the like. The laser sensor may be, for example, a three-dimensional LiDAR (Laser Imaging Detection and Ranging) sensor. The detection unit 14 may also be a device that uses a technology that measures distance from images captured by a stereo camera or a monocular camera, such as SfM (Structure from Motion) technology. A plurality of image capture units 12 may also be used as the detection unit 14. One of the plurality of image capture units 12 may also be used as the detection unit 14.
[0021] The display unit 16 displays various types of information and is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.
[0022] In this embodiment, the information processing device 10 is communicably connected to an electronic control unit (ECU) 3 mounted on the moving object 2. The ECU 3 is a unit that performs electronic control of the moving object 2. In this embodiment, the information processing device 10 is capable of receiving CAN (Controller Area Network) data such as the speed and moving direction of the moving object 2 from the ECU 3.
[0023] Next, the hardware configuration of the information processing device 10 will be described.
[0024] FIG. 2 is a diagram illustrating an example of a hardware configuration of the information processing device 10. As shown in FIG.
[0025] The information processing device 10 is, for example, a computer, and includes a CPU (Central Processing Unit) 10A, a ROM (Read Only Memory) 10B, a RAM (Random Access Memory) 10C, and an I / F (Interface) 10D. The CPU 10A, ROM 10B, RAM 10C, and I / F 10D are interconnected by a bus 10E, and have a hardware configuration similar to that of a typical computer.
[0026] The CPU 10A is a calculation device that controls the information processing device 10. The CPU 10A corresponds to an example of a hardware processor. The ROM 10B stores programs and the like that realize various processes by the CPU 10A. The RAM 10C stores data necessary for various processes by the CPU 10A. The I / F 10D is an interface that is connected to the imaging unit 12, the detection unit 14, the display unit 16, the ECU 3, etc., and is used to send and receive data.
[0027] A program for executing information processing executed by the information processing device 10 of this embodiment is provided by being pre-installed in a ROM 10B or the like. The program executed by the information processing device 10 of this embodiment may be provided by being recorded on a recording medium in a format that can be installed on the information processing device 10 or in a format that can be executed. The recording medium is a computer-readable medium. Examples of the recording medium include a CD (Compact Disc)-ROM, a flexible disk (FD), a CD-R (Recordable), a DVD (Digital Versatile Disk), a USB (Universal Serial Bus) memory, and an SD (Secure Digital) card.
[0028] Next, a functional configuration of the information processing device 10 according to this embodiment will be described. The information processing device 10 uses Visual SLAM processing to simultaneously estimate position information of the detection point and self-position information of the moving object 2 from images captured by the imaging unit 12. The information processing device 10 stitches together multiple spatially adjacent captured images to generate and display a composite image overlooking the periphery of the moving object 2. In this embodiment, the imaging unit 12 is used as the detection unit 14.
[0029] Fig. 3 is a diagram showing an example of the functional configuration of the information processing device 10. In Fig. 3, in order to clarify the data input / output relationship, an image capturing unit 12 and a display unit 16 are also shown in addition to the information processing device 10.
[0030] The information processing device 10 includes an acquisition unit 20, a selection unit 23, a VSLAM processing unit 24, an integration processing unit 29, a determination unit 30, a deformation unit 32, a virtual viewpoint line of sight determination unit 34, a projection transformation unit 36, and an image synthesis unit 38.
[0031] Some or all of the above-mentioned multiple units may be realized by, for example, causing a processing device such as CPU 10A to execute a program, that is, by software. Also, some or all of the above-mentioned multiple units may be realized by hardware such as an IC (Integrated Circuit), or may be realized by a combination of software and hardware.
[0032] The acquiring section 20 acquires the photographed images from the photographing section 12. The acquiring section 20 acquires the photographed images from each of the photographing sections 12 (photographing sections 12A to 12D).
[0033] The acquisition unit 20 outputs the acquired captured image to the projection transformation unit 36 and the selection unit 23 every time the acquisition unit 20 acquires the captured image.
[0034] The selection unit 23 selects a detection area of the detection point. In this embodiment, the selection unit 23 selects at least one of the multiple imaging units 12 (imaging units 12A to 12D), thereby selecting the detection area.
[0035] In this embodiment, the selection unit 23 selects at least one of the photographing units 12 using vehicle state information, detection direction information contained in the CAN data received from the ECU 3, or instruction information input by an operation instruction from the user.
[0036] The vehicle state information is information indicating, for example, the traveling direction of the moving object 2, the state of the direction indication of the moving object 2, the gear state of the moving object 2, etc. The vehicle state information can be derived from CAN data. The detected direction information is information indicating the direction in which noteworthy information is detected, and can be derived using POI (Point of Interest) technology. The instruction information is information input by a user's operational instruction, for example, in an automatic parking mode, assuming a case in which the type of parking to be performed, such as perpendicular parking or parallel parking, is selected.
[0037] For example, the selection unit 23 selects the detection area E (E1 to D4) using the vehicle state information. Specifically, the selection unit 23 identifies the traveling direction of the moving object 2 using the vehicle state information. The selection unit 23 associates the traveling direction with the identification information of any one of the image capturing units 12 and stores them in advance. For example, the selection unit 23 associates the identification information of the image capturing unit 12D (see FIG. 1) that captures the rear of the moving object 2 with the back movement information and stores it in advance. The selection unit 23 also associates the identification information of the image capturing unit 12A (see FIG. 1) that captures the front of the moving object 2 with the forward movement information and stores it in advance.
[0038] Then, the selection unit 23 selects the detection area E by selecting the photographing unit 12 corresponding to the parking information derived from the received vehicle state information.
[0039] The selection unit 23 may also select the photographing unit 12 whose photographing area E is the direction indicated by the detected direction information. The selection unit 23 may also select the photographing unit 12 whose photographing area E is the direction indicated by the detected direction information derived using POI technology.
[0040] The selection unit 23 outputs the photographed image captured by the selected photographing unit 12 from among the photographed images acquired by the acquisition unit 20 to the VSLAM processing unit 24.
[0041] The VSLAM processing unit 24 acquires first point cloud information based on first image data obtained from a first imaging unit, which is one of the imaging units 12A to 12D. The VSLAM processing unit 24 acquires second point cloud information based on second image data obtained from a second imaging unit, which is one of the imaging units 12A to 12D and different from the first imaging unit. That is, the VSLAM processing unit 24 receives from the selection unit 23 an image captured by one of the imaging units 12A to 12D, performs VSLAM processing using the image to generate environment map information, and outputs the generated environment map information to the determination unit 30. The VSLAM processing unit 24 is an example of an acquisition unit.
[0042] More specifically, the VSLAM processing unit 24 includes a matching unit 25, a storage unit 26, a self-position estimation unit 27A, a three-dimensional reconstruction unit 27B, and a correction unit .
[0043] The matching unit 25 performs a feature extraction process and a matching process between multiple captured images (multiple captured images in different frames) captured at different times. In detail, the matching unit 25 performs a feature extraction process from these multiple captured images. The matching unit 25 performs a matching process for identifying corresponding points between the multiple captured images captured at different times, using the feature amounts between the multiple captured images. The matching unit 25 outputs the matching process result to the storage unit 26.
[0044] The self-position estimation unit 27A estimates the self-position relative to the captured image by projective transformation or the like using the multiple matching points acquired by the matching unit 25. Here, the self-position includes information on the position (three-dimensional coordinates) and tilt (rotation) of the image capturing unit 12. The self-position estimation unit 27 stores the self-position information as point cloud information in the environmental map information 26A.
[0045] The three-dimensional reconstruction unit 27B performs perspective projection transformation processing using the movement amount (translation amount and rotation amount) of the self-position estimated by the self-position estimation unit 27A, and determines the three-dimensional coordinates of the matching points (coordinates relative to the self-position). The three-dimensional reconstruction unit 27B stores the peripheral position information, which is the determined three-dimensional coordinates, in the environmental map information 26A as point cloud information.
[0046] As a result, new peripheral position information and self-position information are sequentially added to the environment map information 26A as the moving object 2 on which the image capturing unit 12 is mounted moves.
[0047] The storage unit 26 stores various types of data. The storage unit 26 is, for example, a semiconductor memory element such as RAM or flash memory, a hard disk, an optical disk, or the like. The storage unit 26 may be a storage device provided outside the information processing device 10. The storage unit 26 may also be a storage medium. Specifically, the storage medium may store or temporarily store programs and various types of information downloaded via a LAN (Local Area Network), the Internet, or the like.
[0048] The environmental map information 26A is information in which point cloud information, which is peripheral position information calculated by the three-dimensional restoration unit 27B, and point cloud information, which is self-position information calculated by the self-position estimation unit 27A, are registered in a three-dimensional coordinate space with a predetermined position in the real space as the origin (reference position). The predetermined position in the real space may be determined based on, for example, a predetermined condition.
[0049] For example, the predetermined position is the position of the moving object 2 when the information processing device 10 executes the information processing of this embodiment. For example, assume that the information processing is executed at a predetermined timing, such as a parking scene of the moving object 2. In this case, the information processing device 10 may set the position of the moving object 2 when it determines that the predetermined timing has been reached as the predetermined position. For example, the information processing device 10 may determine that the predetermined timing has been reached when it determines that the behavior of the moving object 2 has become behavior indicative of a parking scene. Examples of behavior indicative of a parking scene due to backing up include when the speed of the moving object 2 drops below a predetermined speed, when the gear of the moving object 2 is shifted into reverse gear, or when a signal indicating the start of parking is received by a user's operation instruction, etc. Note that the predetermined timing is not limited to a parking scene.
[0050] Fig. 4 is a schematic diagram of an example of the environment map information 26A. As shown in Fig. 4, the environment map information 26A is information in which point cloud information, which is the position information (peripheral position information) of each of the detection points P, and point cloud information, which is the self-position information of the self-position S of the moving object 2, are registered at corresponding coordinate positions in the three-dimensional coordinate space. Note that Fig. 4 shows the self-position S of self-positions S1 to S3 as an example. The larger the value of the number following S, the closer the self-position S is to the current timing.
[0051] The correction unit 28 corrects the peripheral position information and self-position information registered in the environmental map information 26A using, for example, the least squares method, so that the sum of the differences in distance in three-dimensional space between previously calculated three-dimensional coordinates and newly calculated three-dimensional coordinates for points that have been matched multiple times between multiple frames is minimized. Note that the correction unit 28 may also correct the amount of movement (translation amount and rotation amount) of the self-position used in the process of calculating the self-position information and the peripheral position information.
[0052] The timing of the correction process by the correction unit 28 is not limited. For example, the correction unit 28 may execute the correction process at a predetermined timing. The predetermined timing may be determined based on a predetermined condition, for example. In the present embodiment, the information processing device 10 is described as being configured to include the correction unit 28 as an example. However, the information processing device 10 may not be configured to include the correction unit 28.
[0053] The integration processing unit 29 executes alignment processing between the first point cloud information and the second point cloud information received from the VSLAM processing unit 24. The integration processing unit 29 executes the integration processing using the first point cloud information and the second point cloud information on which alignment processing has been executed. Here, the integration processing is processing of aligning and integrating the first point cloud information, which is point cloud information of peripheral position information and self-position information acquired using images captured using a first imaging unit, with the second point cloud information, which is point cloud information of peripheral position information and self-position information acquired using images captured using a second imaging unit different from the first imaging unit, to generate integrated point cloud information including at least both of the point cloud information.
[0054] The integration process executed by the integration processing unit 29 will be described below with reference to FIGS.
[0055] FIG. 5 is a plan view showing an example of a situation in which a moving object 2 is rear-parked in a parking space PA. FIG. 6 is a plan view showing an example of the imaging range E1 of the imaging unit 12A (hereinafter also referred to as the "front imaging unit 12A") provided in front of the moving object 2 when the moving object 2 moves forward. FIG. 7 is a plan view showing an example of the imaging range E4 of the imaging unit 12D (hereinafter also referred to as the "rear imaging unit 12D") provided behind the moving object 2 when the moving object 2 moves backward. Note that FIG. 5 illustrates a case in which the moving object 2 moves forward once along the track OB1, then switches the gear of the moving object 2 from drive "D" to reverse "R" and moves backward along the track OB2, thereby being rear-parked in the parking space PA. Note that car1, car2, and car3 each represent other moving objects parked in parking spaces other than the parking space PA.
[0056] For the sake of specificity, the following description will be made taking as an example the integration process when the moving object 2 performs rear parking as shown in FIG.
[0057] When the moving object 2 moves forward along the trajectory OB1, as shown in FIG. 6, images of the shooting range E1 are sequentially acquired by the forward imaging unit 12A as the moving object 2 moves. The VSLAM processing unit 24 executes VSLAM processing using the images of the shooting range E1 sequentially output from the selection unit 23, and generates point cloud information about the area ahead of the moving object 2. Note that the VSLAM processing using the images of the shooting range E1 by the forward imaging unit 12A is also referred to as "forward VSLAM processing" below. The forward VSLAM processing is an example of the first processing or the second processing. The VSLAM processing unit 24 also executes forward VSLAM processing using newly input images of the shooting range E1, and updates the point cloud information about the area around the moving object 2.
[0058] FIG. 8 is a diagram showing an example of point cloud information M1 related to the periphery of the moving object 2, generated by the VSLAM processing unit 24, when the moving object 2 moves forward once along the trajectory OB1. In FIG. 8, the point cloud present in region R1 of the point cloud information M1 related to the forward VSLAM processing is the point cloud corresponding to car1 in FIG. 5. Because the moving object 2 is moving forward along the trajectory OB1, point cloud information corresponding to car1 in FIG. 5 can be acquired during the period when car1 enters the imaging range E1 and is captured in the image. Therefore, as shown in FIG. 8, it can be seen that many point clouds exist in region R1 corresponding to car1.
[0059] 7, after the moving object 2 moves forward along the trajectory OB1, images of the shooting range E4 are sequentially acquired by the rearward photographing unit 12D as the moving object 2 moves backward after shifting gears. The VSLAM processing unit 24 executes VSLAM processing using the images of the shooting range E4 sequentially output from the selection unit 23, and generates point cloud information relating to the rear of the moving object 2. Note that the VSLAM processing using the images of the shooting range E4 by the rearward photographing unit 12D is also referred to as "rearward VSLAM processing" hereinafter. The rearward VSLAM processing is an example of the first processing or the second processing. The VSLAM processing unit 24 also executes rearward VSLAM processing using newly input images of the shooting range E4, and updates the point cloud information relating to the rear of the moving object 2.
[0060] FIG. 9 is a diagram showing an example of point cloud information M2 relating to the rear of the moving object 2, generated by the VSLAM processing unit 24, when the moving object 2 moves backward along the trajectory OB2 after shifting gears. In FIG. 9, the point cloud existing in region R2 of the point cloud information M2 related to the rear VSLAM processing is the point cloud corresponding to car2 in FIG. 5. Because the moving object 2 is moving backward along the trajectory OB2, point cloud information corresponding to car2 in FIG. 5 can be acquired during the period when car2 enters the shooting range E4 and is captured in the image. Therefore, as shown in FIG. 9, it can be seen that many point clouds exist in region R2.
[0061] The integration processing unit 29 executes, for example, a point cloud alignment process between point cloud information M1 related to the forward VSLAM process and point cloud information M2 related to the backward VSLAM process. Here, the point cloud alignment process is a process for aligning multiple point clouds with each other by executing a calculation process including at least one of parallel movement (translational movement) and rotational movement for at least one of the multiple point clouds to be aligned. For example, when executing the point cloud alignment process for two pieces of point cloud information, first, the self-position coordinates of both are aligned, peripheral point clouds within a certain range from the self-position are set as alignment targets, the difference in position between corresponding points is calculated as a distance, and the amount of parallel movement of one reference position relative to the other reference position when the sum of the distances is equal to or less than a predetermined threshold is calculated.
[0062] The point cloud alignment process may be any process that aligns target point cloud information. Examples of the point cloud alignment process include scan matching processes that use algorithms such as ICP (Iterative Closest Point) and NDT (Normal Distribution Transform).
[0063] The integration processing unit 29 uses the point cloud information M1 and the point cloud information M2 that have been subjected to the point cloud alignment processing to generate integrated point cloud information in which the point cloud information M1 and the point cloud information M2 are integrated.
[0064] Fig. 10 is a diagram showing an example of integrated point cloud information M3 generated by the integration process for rear parking of the moving object 2 shown in Fig. 5. As shown in Fig. 10, the integrated point cloud information M3 includes both point cloud information M1 related to the forward VSLAM process and point cloud information M2 related to the rear VSLAM process. Therefore, it includes a large amount of point cloud information for area R3 corresponding to car1, area R5 corresponding to car2, etc.
[0065] The above-described integration process is executed in conjunction with the input of vehicle state information indicating a change in the operating state (in this embodiment, a change in gear).
[0066] Fig. 11 is a flowchart showing an example of the flow of the integration process shown in Fig. 5 to Fig. 10. As shown in Fig. 11, the integration processing unit 29 determines whether the gear is in a forward state (for example, drive "D") or a reverse state (for example, reverse "R") (step S1). Hereinafter, an example will be described in which the forward gear is drive "D" and the reverse gear is reverse "R".
[0067] When the integrated processing unit 29 determines that the gear is in drive "D" (D in step S1), it executes the above-described forward VSLAM process (step S2a). The integrated processing unit 29 repeatedly executes the forward VSLAM process until the gear is changed (No in step S3a).
[0068] On the other hand, if the gear is changed from drive "D" to reverse "R" (Yes in step S3a), the integrated processing unit 29 executes the rear VSLAM process described above (step S4a).
[0069] The integration processing unit 29 performs point cloud registration using the rear point cloud information obtained by the rear VSLAM processing and the front point cloud information obtained by the front VSLAM processing (step S5a).
[0070] The integration processing unit 29 generates integrated point cloud information using the rear point cloud information and the front point cloud information after the point cloud alignment processing (step S6a).
[0071] The integration processing unit 29 executes backward VSLAM processing as the moving object 2 moves backward, and successively updates the integrated point cloud information (step S7a).
[0072] On the other hand, if the integrated processing unit 29 determines in step S1 that the gear is in reverse "R" (R in step S1), it executes the rear VSLAM process described above (step S2b). The integrated processing unit 29 repeatedly executes the rear VSLAM process until the gear is changed (No in step S3b).
[0073] On the other hand, if the gear is changed from reverse "R" to drive "D" (Yes in step S3b), the integrated processing unit 29 executes the above-described front VSLAM processing (step S4b).
[0074] The integration processing unit 29 performs alignment processing using the forward point cloud information obtained by the forward VSLAM processing and the backward point cloud information obtained by the backward VSLAM processing to align both point clouds (step S5b).
[0075] The integration processing unit 29 generates integrated point cloud information using the front point cloud information and the rear point cloud information after the point cloud alignment processing (step S6b).
[0076] The integration processing unit 29 executes forward VSLAM processing as the moving object 2 moves forward, and successively updates the integrated point cloud information (step S7b).
[0077] Returning to Figure 3, the determination unit 30 receives environmental map information including integrated point cloud information from the integration processing unit 29, and calculates the distance between the moving body 2 and surrounding three-dimensional objects using the surrounding position information and self-position information stored in the environmental map information 26A.
[0078] The determining unit 30 determines the projection shape of the projection surface using the distance between the moving object 2 and the surrounding three-dimensional objects, and generates projection shape information. The determining unit 30 outputs the generated projection shape information to the transforming unit 32.
[0079] Here, the projection surface is a three-dimensional surface onto which an image of the periphery of the moving object 2 is projected. The image of the periphery of the moving object 2 is a captured image of the periphery of the moving object 2, which is captured by each of the image capturing units 12A to 12D. The projection shape of the projection surface is a three-dimensional (3D) shape that is virtually formed in a virtual space corresponding to the real space. In this embodiment, the determination of the projection shape of the projection surface executed by the determination unit 30 is referred to as a projection shape determination process.
[0080] Furthermore, the determination unit 30 calculates an asymptotic curve of the peripheral position information relative to the self-position, using the peripheral position information of the moving object 2 and the self-position information stored in the environment map information 26A.
[0081] FIG. 12 is an explanatory diagram of an asymptotic curve Q generated by the determination unit 30. Here, the asymptotic curve is an asymptotic curve of a plurality of detection points P in the environment map information 26A. FIG. 12 shows an example of an asymptotic curve Q shown in a projected image obtained by projecting a captured image onto a projection surface when the moving object 2 is viewed from above. For example, it is assumed that the determination unit 30 has identified three detection points P in descending order of proximity to the self-position S of the moving object 2. In this case, the determination unit 30 generates the asymptotic curves Q of these three detection points P.
[0082] The determination unit 30 outputs the self-position and the asymptotic curve information to the virtual viewpoint line-of-sight determination unit 34.
[0083] The deformation unit 32 deforms the projection surface based on the projection shape information determined using the environment map information including the integrated point cloud information received from the determination unit 30. The deformation unit 32 is an example of a deformation unit.
[0084] FIG. 13 is a schematic diagram showing an example of a reference projection plane 40. FIG. 14 is a schematic diagram showing an example of a projection shape 41 determined by the determination unit 30. That is, the deformation unit 32 deforms the reference projection plane shown in FIG. 13, which is stored in advance, based on the projection shape information, and determines a deformed projection plane 42 as the projection shape 41 shown in FIG. 14. The deformation unit 32 generates deformed projection plane information based on the projection shape 41. This deformation of the reference projection plane is performed, for example, using the detection point P closest to the moving object 2 as a reference. The deformation unit 32 outputs the deformed projection plane information to the projection conversion unit 36.
[0085] Furthermore, for example, the deformation unit 32 deforms the reference projection plane into a shape that follows the asymptotic curve of a predetermined number of detection points P in order of proximity to the moving object 2, based on the projection shape information.
[0086] The virtual viewpoint line of sight determining unit 34 determines virtual viewpoint line of sight information based on the self-position and the asymptotic curve information.
[0087] Determination of virtual viewpoint line-of-sight information will be described with reference to FIGS. 12 and 14. The virtual viewpoint line-of-sight determination unit 34 determines, for example, a direction passing through a detection point P closest to the self-position S of the moving object 2 and perpendicular to the deformed projection plane as the line-of-sight direction. Furthermore, the virtual viewpoint line-of-sight determination unit 34 fixes, for example, the direction of the line-of-sight direction L, and determines the coordinates of the virtual viewpoint O as an arbitrary Z coordinate and arbitrary XY coordinates in a direction away from the asymptotic curve Q toward the self-position S. In this case, the XY coordinates may be coordinates of a position farther away from the asymptotic curve Q than the self-position S. The virtual viewpoint line-of-sight determination unit 34 then outputs virtual viewpoint line-of-sight information indicating the virtual viewpoint O and the line-of-sight direction L to the projection transformation unit 36. Note that, as shown in FIG. 14, the line-of-sight direction L may be a direction from the virtual viewpoint O toward the position of the vertex W of the asymptotic curve Q.
[0088] The projection conversion unit 36 generates a projection image by projecting the captured image acquired from the image capturing unit 12 onto the deformed projection surface based on the deformed projection surface information and the virtual viewpoint line of sight information. The projection conversion unit 36 converts the generated projection image into a virtual viewpoint image and outputs it to the image synthesis unit 38. Here, the virtual viewpoint image is an image obtained by viewing the projection image in an arbitrary direction from a virtual viewpoint.
[0089] With reference to FIG. 14, the projection image generation process by the projection transformation unit 36 will be described in detail. The projection transformation unit 36 projects the captured image onto the modified projection surface 42. Then, the projection transformation unit 36 generates a virtual viewpoint image (not shown), which is an image obtained by viewing the captured image projected onto the modified projection surface 42 from an arbitrary virtual viewpoint O in a line of sight direction L. The position of the virtual viewpoint O may be set to, for example, the latest self-position S of the moving object 2. In this case, the X and Y coordinate values of the virtual viewpoint O may be set to the X and Y coordinate values of the latest self-position S of the moving object 2. Furthermore, the Z coordinate value (vertical position) of the virtual viewpoint O may be set to the Z coordinate value of the detection point P closest to the self-position S of the moving object 2. The line of sight direction L may be determined based on, for example, a predetermined criterion.
[0090] The line of sight direction L may be, for example, a direction from the virtual viewpoint O toward the detection point P that is closest to the self-position S of the moving object 2. The line of sight direction L may also be a direction that passes through the detection point P and is perpendicular to the deformed projection plane 42. Virtual viewpoint line of sight information indicating the virtual viewpoint O and the line of sight direction L is created by the virtual viewpoint line of sight determination unit 34.
[0091] For example, the virtual viewpoint line-of-sight determination unit 34 may determine, as the line-of-sight direction L, a direction that passes through the detection point P closest to the self-position S of the moving object 2 and is perpendicular to the deformed projection plane 42. Alternatively, the virtual viewpoint line-of-sight determination unit 34 may fix the direction of the line-of-sight direction L and determine the coordinates of the virtual viewpoint O as an arbitrary Z coordinate and arbitrary X and Y coordinates in a direction away from the asymptotic curve Q toward the self-position S. In this case, the X and Y coordinates may be coordinates of a position farther away from the asymptotic curve Q than the self-position S. The virtual viewpoint line-of-sight determination unit 34 then outputs virtual viewpoint line-of-sight information indicating the virtual viewpoint O and the line-of-sight direction L to the projection transformation unit 36. Note that, as shown in FIG. 14 , the line-of-sight direction L may be a direction from the virtual viewpoint O toward the position of the vertex W of the asymptotic curve Q.
[0092] The projection transformation unit 36 receives virtual viewpoint line-of-sight information from the virtual viewpoint line-of-sight determination unit 34. By receiving the virtual viewpoint line-of-sight information, the projection transformation unit 36 identifies a virtual viewpoint O and a line-of-sight direction L. Then, the projection transformation unit 36 generates a virtual viewpoint image, which is an image viewed from the virtual viewpoint O in the line-of-sight direction L, from the captured image projected onto the modified projection surface 42. The projection transformation unit 36 outputs the virtual viewpoint image to the image synthesis unit 38.
[0093] The image synthesis unit 38 generates a synthesized image by extracting a part or all of the virtual viewpoint images. For example, the image synthesis unit 38 performs a process of joining together a plurality of virtual viewpoint images (here, four virtual viewpoint images corresponding to the imaging units 12A to 12D) in the boundary region between the imaging units.
[0094] The image synthesis unit 38 outputs the generated synthetic image to the display unit 16. The synthetic image may be a bird's-eye view image with a virtual viewpoint O above the moving object 2, or may be an image in which the moving object 2 is displayed semi-transparently with a virtual viewpoint O inside the moving object 2.
[0095] The projection transformation unit 36 and the image synthesis unit 38 constitute an image generation unit 37. The image generation unit 37 is an example of an image generation unit.
[0096] [Configuration example of the integration processing unit 29 and the determination unit 30] Next, an example of the detailed configuration of the integration processing unit 29 and the determination unit 30 will be described.
[0097] Fig. 15 is a schematic diagram showing an example of the functional configuration of the integration processing unit 29 and the determination unit 30. As shown in Fig. 15, the integration processing unit 29 includes a past map storage unit 291, a difference calculation unit 292, an offset processing unit 293, and an integration unit 294. The determination unit 30 also includes an absolute distance conversion unit 30A, an extraction unit 30B, a nearest neighbor identification unit 30C, a reference projection surface shape selection unit 30D, a scale determination unit 30E, an asymptotic curve calculation unit 30F, a shape determination unit 30G, and a boundary region determination unit 30H.
[0098] The past map storage unit 291 acquires and stores (holds) the environmental map information output from the VSLAM processing unit 24 in accordance with changes in the vehicle state information of the moving body 2. For example, the past map storage unit 291 holds point cloud information included in the latest environmental map information output from the VSLAM processing unit 24, triggered by input of gear information (vehicle state information) indicating a gear shift (at the timing of the gear shift).
[0099] The difference calculation unit 292 performs a point cloud alignment process between the point cloud information included in the environmental map information output from the VSLAM processing unit 24 and the point cloud information held by the past map holding unit 291. For example, the difference calculation unit 292 calculates, as an offset amount Δ, the amount of parallel movement of one origin relative to the other origin when the sum of the inter-point distances between the point cloud information included in the environmental map information output from the VSLAM processing unit 24 and the point cloud information held by the past map holding unit 291 is smallest.
[0100] The offset processing unit 293 offsets the point cloud information (coordinates) held by the past map holding unit 291 using the offset amount calculated by the difference calculation unit 292. For example, the offset processing unit 293 adds an offset amount Δ to the point cloud information held by the past map holding unit 291 to translate it.
[0101] The integrating unit 294 generates integrated point cloud information using the point cloud information included in the environment map information output from the VSLAM processing unit 24 and the point cloud information output from the offset processing unit 293. For example, the integrating unit 294 generates integrated point cloud information by superimposing the point cloud information output from the offset processing unit 293 on the point cloud information included in the environment map information output from the VSLAM processing unit 24. The integrating unit 294 is an example of a combining unit.
[0102] The absolute distance conversion unit 30A converts the relative positional relationship between the self-position and the surrounding three-dimensional object, which can be known from the environmental map information 26A, into the absolute value of the distance from the self-position to the surrounding three-dimensional object.
[0103] Specifically, for example, the speed data of the moving object 2 included in the CAN data received from the ECU 3 of the moving object 2 is used. For example, in the case of the environment map information 26A shown in FIG. 4, the relative positional relationship between the self-position S and multiple detection points P can be known, but the absolute value of the distance is not calculated. Here, the distance between the self-position S3 and the self-position S2 can be calculated based on the inter-frame period for calculating the self-position and the speed data during that period based on the CAN data. Since the relative positional relationship in the environment map information 26A is similar to that in real space, by knowing the distance between the self-position S3 and the self-position S2, the absolute values of the distances from the self-position S to all other detection points P can also be calculated. Note that when the detection unit 14 acquires the distance information of the detection points P, the absolute distance conversion unit 30A may be omitted.
[0104] Then, the absolute distance conversion unit 30A outputs the calculated measured distances of each of the multiple detection points P to the extraction unit 30B. In addition, the absolute distance conversion unit 30A outputs the calculated current position of the moving object 2 to the virtual viewpoint line of sight determination unit 34 as self-position information of the moving object 2.
[0105] The extraction unit 30B extracts detection points P that are present within a specific range from among the multiple detection points P for which the measurement distances have been received from the absolute distance conversion unit 30A. The specific range is, for example, a range from the road surface on which the moving object 2 is placed to a height equivalent to the vehicle height of the moving object 2. Note that the range is not limited to this range.
[0106] By extracting detection points P within the range by the extraction unit 30B, it is possible to extract detection points P of objects that may be obstacles to the progress of the moving body 2, objects located adjacent to the moving body 2, and the like.
[0107] Then, the extraction unit 30B outputs the measured distance of each of the extracted detection points P to the nearest neighbor identification unit 30C.
[0108] The nearest neighbor identification unit 30C divides the surroundings of the self-position S of the moving body 2 into specific ranges (for example, angular ranges), and for each range, identifies the detection point P closest to the moving body 2, or multiple detection points P in order of proximity to the moving body 2. The nearest neighbor identification unit 30C identifies the detection point P using the measured distance received from the extraction unit 30B. In this embodiment, a form in which the nearest neighbor identification unit 30C identifies multiple detection points P in order of proximity to the moving body 2 for each range will be described as an example.
[0109] The nearest neighbor specifying unit 30C outputs the measurement distance of the detection point P specified for each range to the reference projection surface shape selecting unit 30D, the scale determining unit 30E, the asymptotic curve calculating unit 30F, and the boundary region determining unit 30H.
[0110] The reference projection surface shape selection unit 30D selects the shape of the reference projection surface.
[0111] Here, the reference projection surface will be described with reference to Fig. 13. The reference projection surface 40 is, for example, a projection surface having a shape that serves as a reference when changing the shape of the projection surface. The shape of the reference projection surface 40 is, for example, bowl-shaped, cylindrical, etc. Note that Fig. 13 shows an example of a bowl-shaped reference projection surface 40.
[0112] The bowl-shaped container has a bottom surface 40A and a side wall surface 40B, one end of which is continuous with the bottom surface 40A and the other end of which is open. The width of the horizontal cross section of the side wall surface 40B increases from the bottom surface 40A toward the open end of the other end. The bottom surface 40A is, for example, circular. Here, a circular shape includes a perfect circle and other circular shapes such as an ellipse. The horizontal cross section is an orthogonal plane perpendicular to the vertical direction (arrow Z direction). The orthogonal plane is a two-dimensional plane along the arrow X direction, which is perpendicular to the arrow Z direction, and the arrow Y direction, which is perpendicular to the arrow Z direction and the arrow X direction. Hereinafter, the horizontal cross section and the orthogonal plane may be referred to as the XY plane. The bottom surface 40A may have a shape other than a circle, such as an egg shape.
[0113] The cylindrical shape is a shape consisting of a circular bottom surface 40A and a side wall surface 40B that is continuous with the bottom surface 40A. The side wall surface 40B that constitutes the cylindrical reference projection surface 40 has a cylindrical shape with an opening at one end that is continuous with the bottom surface 40A and an open other end. However, the side wall surface 40B that constitutes the cylindrical reference projection surface 40 has a shape in which the diameter in the XY plane is approximately constant from the bottom surface 40A side toward the opening at the other end. The bottom surface 40A may have a shape other than a circle, such as an egg shape.
[0114] In this embodiment, the case where the shape of the reference projection plane 40 is bowl-shaped as shown in Fig. 13 will be described as an example. The reference projection plane 40 is a three-dimensional model virtually formed in a virtual space with a bottom surface 40A that is a surface that substantially coincides with the road surface below the moving object 2 and the center of the bottom surface 40A being the self-position S of the moving object 2.
[0115] The reference projection surface shape selection unit 30D selects the shape of the reference projection surface 40 by reading one specific shape from multiple types of reference projection surfaces 40. For example, the reference projection surface shape selection unit 30D selects the shape of the reference projection surface 40 based on the positional relationship between the self-position and surrounding three-dimensional objects, the stabilization distance, etc. The shape of the reference projection surface 40 may also be selected based on an operational instruction from the user. The reference projection surface shape selection unit 30D outputs shape information of the determined reference projection surface 40 to the shape determination unit 30G. In this embodiment, as described above, an embodiment in which the reference projection surface shape selection unit 30D selects a bowl-shaped reference projection surface 40 will be described as an example.
[0116] The scale determination unit 30E determines the scale of the reference projection plane 40 of the shape selected by the reference projection plane shape selection unit 30D. For example, the scale determination unit 30E makes a decision to reduce the scale when there are multiple detection points P within a predetermined distance range from the self-position S. The scale determination unit 30E outputs scale information of the determined scale to the shape determination unit 30G.
[0117] The asymptotic curve calculation unit 30F uses each of the stabilization distances of the detection points P closest to the self-position S for each range from the self-position S received from the nearest neighbor identification unit 30C, and outputs asymptotic curve information of the calculated asymptotic curve Q to the shape determination unit 30G and the virtual viewpoint line of sight determination unit 34. Note that the asymptotic curve calculation unit 30F may calculate the asymptotic curve Q of the detection points P accumulated for each of multiple portions of the reference projection plane 40. Then, the asymptotic curve calculation unit 30F may output the asymptotic curve information of the calculated asymptotic curve Q to the shape determination unit 30G and the virtual viewpoint line of sight determination unit 34.
[0118] The shape determination unit 30G enlarges or reduces the reference projection plane 40, which has a shape indicated by the shape information received from the reference projection plane shape selection unit 30D, to the scale of the scale information received from the scale determination unit 30E. Then, the shape determination unit 30G determines, as the projection shape, a shape obtained by deforming the reference projection plane 40 after enlarging or reducing it so that it becomes a shape that follows the asymptotic curve information of the asymptotic curve Q received from the asymptotic curve calculation unit 30F.
[0119] Here, the determination of the projection shape will be described in detail with reference to Fig. 14. As shown in Fig. 14, the shape determination unit 30G determines, as the projection shape 41, a shape obtained by deforming the reference projection plane 40 into a shape that passes through a detection point P that is closest to the self-position S of the moving object 2, which is the center of the bottom surface 40A of the reference projection plane 40. A shape that passes through the detection point P means that the deformed side wall surface 40B is a shape that passes through the detection point P. The self-position S is the latest self-position S calculated by the self-position estimation unit 27.
[0120] That is, the shape determination unit 30G identifies the detection point P that is closest to the self-position S among the multiple detection points P registered in the environmental map information 26A. In detail, the XY coordinates of the center position (self-position S) of the moving object 2 are set to (X, Y) = (0, 0). Then, the shape determination unit 30G determines the X 2 +Y 2 The detected point P where the value of is the smallest is identified as the detected point P closest to the self-position S. Then, the shape determination unit 30G determines, as the projected shape 41, a shape obtained by deforming the side wall surface 40B of the reference projection plane 40 so that it passes through the detected point P.
[0121] More specifically, the shape determination unit 30G determines the deformed shape of the bottom surface 40A and a portion of the side wall surface 40B as the projected shape 41 so that, when the reference projection plane 40 is deformed, a portion of the side wall surface 40B becomes a wall surface passing through the detection point P closest to the moving object 2. The deformed projected shape 41 is, for example, a shape that is raised from a rising line 44 on the bottom surface 40A in a direction toward the center of the bottom surface 40A when viewed from the XY plane (planar view). "Raising" means, for example, bending or folding a portion of the side wall surface 40B and the bottom surface 40A in a direction toward the center of the bottom surface 40A so that the angle formed between the side wall surface 40B of the reference projection plane 40 and the bottom surface 40A becomes smaller. Note that in the raised shape, the rising line 44 may be located between the bottom surface 40A and the side wall surface 40B, and the bottom surface 40A may remain undeformed.
[0122] The shape determination unit 30G determines to deform the specific region on the reference projection plane 40 so that it protrudes to a position that passes through the detection point P when viewed from the viewpoint (planar view) of the XY plane. The shape and range of the specific region may be determined based on predetermined criteria. Then, the shape determination unit 30G determines to deform the reference projection plane 40 so that the distance from the self-position S continuously increases from the protruding specific region toward regions other than the specific region on the side wall surface 40B.
[0123] For example, it is preferable to determine the projected shape 41 so that the outer periphery of the cross section along the XY plane has a curved shape, as shown in Fig. 14. Note that the outer periphery of the cross section of the projected shape 41 is, for example, a circle, but may have a shape other than a circle.
[0124] The shape determination unit 30G may determine, as the projected shape 41, a shape obtained by deforming the reference projection plane 40 so that the shape follows an asymptotic curve. The shape determination unit 30G generates an asymptotic curve of a predetermined number of detection points P in a direction away from the detection point P closest to the self-position S of the moving object 2. The number of detection points P may be more than one. For example, the number of detection points P is preferably three or more. In this case, the shape determination unit 30G preferably generates an asymptotic curve of a plurality of detection points P located at positions that are at least a predetermined angle away from the self-position S. For example, the shape determination unit 30G may determine, as the projected shape 41, a shape obtained by deforming the reference projection plane 40 so that the shape follows the generated asymptotic curve Q for the asymptotic curve Q shown in FIG. 12.
[0125] The shape determination unit 30G may divide the surroundings of the self-position S of the moving body 2 into specific ranges, and for each range, identify the detection point P closest to the moving body 2, or multiple detection points P in order of proximity to the moving body 2. Then, the shape determination unit 30G may determine, as the projection shape 41, a shape obtained by deforming the reference projection plane 40 so as to become a shape that passes through the detection points P identified for each range, or a shape that follows the asymptotic curve Q of the identified multiple detection points P.
[0126] Then, the shape determination unit 30G outputs the projection shape information of the determined projection shape 41 to the deformation unit 32.
[0127] Next, an example of the flow of information processing including point cloud integration processing executed by the information processing device 10 according to this embodiment will be described.
[0128] FIG. 16 is a flowchart showing an example of the flow of information processing executed by the information processing device 10.
[0129] The acquisition unit 20 acquires the captured image from the image capturing unit 12 (step S10). The acquisition unit 20 also directly acquires the specified content (for example, the gear of the moving object 2 has shifted to reverse gear) and the vehicle state.
[0130] The selection unit 23 selects at least one of the imaging units 12A to 12D (step S12).
[0131] The matching unit 25 extracts features and performs matching processing using a plurality of captured images, which were selected in step S12 and captured by the imaging unit 12 at different timings among the captured images acquired in step S10 (step S14). The matching unit 25 also registers information on corresponding points between the plurality of captured images at different timings, which have been identified by the matching processing, in the storage unit 26.
[0132] The self-position estimation unit 27 reads the matching points and the environmental map information 26A (peripheral position information and self-position information) from the storage unit 26 (step S16). The self-position estimation unit 27 estimates the self-position relative to the captured image by projective transformation or the like using the multiple matching points acquired from the matching unit 25 (step S18), and registers the calculated self-position information in the environmental map information 26A (step S20).
[0133] The three-dimensional restoration unit 26B reads the environment map information 26A (peripheral position information and self-position information) (step S22). The three-dimensional restoration unit 26B performs perspective projection transformation processing using the movement amount (translation amount and rotation amount) of the self-position estimated by the self-position estimation unit 27, determines the three-dimensional coordinates of the matching point (coordinates relative to the self-position), and registers them as peripheral position information in the environment map information 26A (step S24).
[0134] The correction unit 28 reads the environment map information 26A (peripheral position information and self-position information). The correction unit 28 corrects (step S26) the peripheral position information and self-position information already registered in the environment map information 26A using, for example, the least squares method so that the total difference in distance in three-dimensional space between previously calculated three-dimensional coordinates and newly calculated three-dimensional coordinates for points that have been matched multiple times across multiple frames is minimized, and updates the environment map information 26A.
[0135] The integration processing unit 29 receives the environment map information 26A output from the VSLAM processing unit 24, and executes point cloud integration processing (step S27).
[0136] Fig. 17 is a flowchart showing an example of the flow of the point cloud integration process in step S27 of Fig. 16. That is, in response to a gear shift, the past map storage unit 291 stores the point cloud information included in the latest environmental map information output from the VSLAM processing unit 24 (step S113a).
[0137] The difference calculation unit 292 performs a scan matching process using the point cloud information included in the environmental map information output from the VSLAM processing unit 24 and the point cloud information held by the past map holding unit 291, and calculates the offset amount (step S113b).
[0138] The offset processing unit 293 applies an offset amount to the point cloud information held by the past map holding unit 291, moves it in parallel, and aligns the point cloud information (step S113c).
[0139] The integration unit 294 generates integrated point cloud information using the point cloud information included in the environment map information output from the VSLAM processing unit 24 and the point cloud information output from the offset processing unit 293 (step S113d).
[0140] Returning to FIG. 16 , the absolute distance conversion unit 30A takes in the speed data (host vehicle speed) of the moving object 2 included in the CAN data received from the ECU 3 of the moving object 2. Using the speed data of the moving object 2, the absolute distance conversion unit 30A converts the peripheral position information included in the environmental map information 26A into distance information from the current position, which is the latest host position S of the moving object 2, to each of the plurality of detection points P (step S28). The absolute distance conversion unit 30A outputs the calculated distance information of each of the plurality of detection points P to the extraction unit 30B. In addition, the absolute distance conversion unit 30A outputs the calculated current position of the moving object 2 to the virtual viewpoint line of sight determination unit 34 as host position information of the moving object 2.
[0141] The extraction unit 30B extracts detection points P that exist within a specific range from among the plurality of detection points P for which distance information has been received (step S30).
[0142] The nearest neighbor identification unit 30C divides the surroundings of the self-position S of the moving body 2 into specific ranges, and for each range, identifies the detection point P closest to the moving body 2, or multiple detection points P in order of closest to the moving body 2, and extracts the distance to the nearest object (step S32). The nearest neighbor identification unit 30C outputs the measured distance d of the detection point P identified for each range (the measured distance between the moving body 2 and the nearest object) to the reference projection surface shape selection unit 30D, the scale determination unit 30E, the asymptotic curve calculation unit 30F, and the boundary region determination unit 30H.
[0143] The asymptotic curve calculation unit 30F calculates an asymptotic curve (step S34), and outputs it to the shape determination unit 30G and the virtual viewpoint line of sight determination unit 34 as asymptotic curve information.
[0144] The reference projection plane shape selection unit 30D selects the shape of the reference projection plane 40 (step S36), and outputs shape information of the selected reference projection plane 40 to the shape determination unit 30G.
[0145] The scale determination unit 30E determines the scale of the reference projection plane 40 of the shape selected by the reference projection plane shape selection unit 30D (step S38), and outputs scale information of the determined scale to the shape determination unit 30G.
[0146] The shape determination unit 30G determines a projection shape, which is how to deform the shape of the reference projection plane, based on the scale information and the asymptotic curve information (step S40). The shape determination unit 30G outputs projection shape information of the determined projection shape 41 to the deformation unit 32.
[0147] The transformation unit 32 transforms the shape of the reference projection plane based on the projection shape information (step S42), and outputs the transformed transformed projection plane information to the projection conversion unit .
[0148] The virtual viewpoint line-of-sight determination unit 34 determines virtual viewpoint line-of-sight information based on the self-position and the asymptotic curve information (step S44). The virtual viewpoint line-of-sight determination unit 34 outputs the virtual viewpoint line-of-sight information indicating the virtual viewpoint O and the line-of-sight direction L to the projection transformation unit 36.
[0149] The projection conversion unit 36 generates a projection image by projecting the captured image acquired from the image capturing unit 12 onto the deformed projection surface based on the deformed projection surface information and the virtual viewpoint line of sight information. The projection conversion unit 36 converts the generated projection image into a virtual viewpoint image (step S46) and outputs it to the image synthesis unit 38.
[0150] The boundary area determination unit 30H determines a boundary area based on the distance to the nearest object identified for each range. That is, the boundary area determination unit 30H determines a boundary area as an overlap area of spatially adjacent peripheral images based on the position of the object nearest to the moving object 2 (step S48). The boundary area determination unit 30H outputs the determined boundary area to the image synthesis unit 38.
[0151] The image synthesis unit 38 generates a synthesized image by stitching spatially adjacent perspective projection images together using the boundary area (step S50). That is, the image synthesis unit 38 generates a synthesized image by stitching together perspective projection images from four directions according to the boundary area set at the angle of the nearest object direction. In the boundary area, the spatially adjacent perspective projection images are blended at a predetermined ratio.
[0152] The display unit 16 displays the composite image (step S52).
[0153] The information processing device 10 determines whether to end the information processing (step S54). For example, the information processing device 10 makes the determination in step S54 by determining whether or not a signal indicating that the position of the moving object 2 has stopped moving has been received from the ECU 3. Alternatively, for example, the information processing device 10 may make the determination in step S54 by determining whether or not an instruction to end the information processing has been received through an operation instruction by a user or the like.
[0154] If a negative determination is made in step S54 (step S54: No), the processes from step S10 to step S54 are repeatedly executed.
[0155] On the other hand, if the determination in step S54 is affirmative (step S54: Yes), this routine ends.
[0156] When returning from step S54 to step S10 after performing the correction process of step S26, the correction process of the subsequent step S26 may be omitted. Also, when returning from step S54 to step S10 without performing the correction process of step S26, the correction process of the subsequent step S26 may be performed.
[0157] As described above, the information processing device 10 according to the embodiment includes the VSLAM processing unit 24 as an acquisition unit, the difference calculation unit 292 and offset processing unit 293 as alignment processing units, and the integrating unit 294 as an integration unit. The VSLAM processing unit 24 acquires point cloud information related to forward VSLAM processing based on image data obtained from the imaging unit 12A provided in front of the moving object 2, and acquires point cloud information related to backward VSLAM processing based on image data obtained from the imaging unit 12D provided behind the moving object 2. The difference calculation unit 292 and the offset processing unit 293 perform alignment processing between the point cloud information related to the forward VSLAM processing and the point cloud information related to the backward VSLAM processing. The integrating unit 294 generates integrated point cloud information using the point cloud information related to the forward VSLAM processing and the point cloud information related to the backward VSLAM processing on which alignment processing has been performed.
[0158] Therefore, it is possible to generate integrated point cloud information by integrating point cloud information acquired using images captured in the past with different image capture units and point cloud information acquired using images captured with the current image capture unit, and to generate an image of the periphery of the moving object using this integrated point cloud information. Therefore, even in cases such as when a car is parking by turning around, it is possible to resolve the lack of position information of surrounding objects obtained by VSLAM processing.
[0159] Fig. 18 is a diagram showing point cloud information M5 related to rear VSLAM processing of an information processing device according to a comparative example. That is, Fig. 18 shows point cloud information M5 acquired only by the rear VSLAM when a moving object 2 moves forward once along trajectory OB1, then shifts gears and moves backward along trajectory OB2 to be rear-parked in a parking space PA. If the point cloud integration processing according to this embodiment is not used, an image of the periphery of the moving object will be generated and displayed using point cloud information M5 acquired only by the rear VSLAM.
[0160] Therefore, as shown in Figure 18, the point cloud information M5 acquired only by the rear VSLAM becomes sparse in the area R6 corresponding to car1 because, as the moving body 2 moves backward, car1 shown in Figure 5 immediately leaves the shooting range D4, and the detection of objects around the moving body such as car1 may become unstable.
[0161] In contrast, the information processing device 10 according to this embodiment generates the integrated point cloud information shown in Fig. 10 through integration processing. As can be seen by comparing the region R3 corresponding to car1 in the integrated point cloud information M3 in Fig. 10 with the region R6 corresponding to car1 in the point cloud information M5 in Fig. 18, for example, there are more point clouds corresponding to car1 in the integrated point cloud information M3 than in the integrated point cloud information M5. Therefore, the information processing device 10 according to this embodiment can more stably detect objects surrounding a moving body such as car1 than the information processing device according to the comparative example.
[0162] The VSLAM processing unit 24 transitions from forward VSLAM processing, which acquires forward point cloud information, to rearward VSLAM processing, which acquires rearward point cloud information, based on vehicle state information indicating the state of the moving body 2. After transitioning from forward VSLAM processing to rearward VSLAM processing, the difference calculation unit 292 and the offset processing unit 293 perform alignment processing using the forward point cloud information and the rearward point cloud information.
[0163] Therefore, there is no need to execute both forward VSLAM processing and backward VSLAM processing, and the calculation load can be reduced.
[0164] (Variation 1) In the above embodiment, an example was given of a case in which the point cloud integration process is triggered by the input of gear information (vehicle state information) indicating that the gear of the moving object 2 has been switched from drive "D" to reverse "R." The vehicle state information that can be used as a trigger is not limited to gear information. For example, the point cloud integration process can also be triggered by information indicating that the steering wheel (handle) has been turned more than a certain amount to change the direction of travel of the moving object 2, or information indicating that the speed of the moving object 2 has dropped below a certain amount in preparation for parking.
[0165] (Variation 2) In the above embodiment, the point cloud integration process has been described as an example in which the moving body 2 is parking backward. However, the point cloud integration process can also be performed in cases in which the moving body 2 is parallel parking or forward parking. For example, in cases in which the moving body 2 is parallel parking, the point cloud integration process can be performed using point cloud information related to side VSLAM processing, which uses images captured by the image capturing units 12B and 12C arranged on the sides of the moving body 2, and point cloud information related to rear VSLAM processing.
[0166] In the above embodiment, the point cloud integration process is described as an example using point cloud information related to forward VSLAM processing and point cloud information related to backward VSLAM processing. However, the point cloud integration process may be performed using point cloud information related to VSLAM processing in three or more directions (or three or more different locations), such as point cloud information related to forward VSLAM processing, point cloud information related to backward VSLAM processing, and point cloud information related to side VSLAM processing.
[0167] Furthermore, if the moving body 2 is a drone, point cloud integration processing can be performed using point cloud information related to upward VSLAM processing using images acquired by a camera unit provided on the upper surface of the moving body 2, point cloud information related to downward VSLAM processing using images acquired by a camera unit provided on the lower surface of the moving body 2, and point cloud information related to lateral VSLAM processing.
[0168] (Variation 3) In the above embodiment, the point cloud integration process is described as an example in which the input of vehicle state information is used as a trigger to switch from the forward VSLAM process to the rearward VSLAM process (or vice versa), and the point cloud information related to the forward VSLAM process is integrated with the point cloud information related to the rearward VSLAM process. However, the point cloud integration process can also be performed using each piece of point cloud information obtained by executing multiple VSLAM processes in parallel.
[0169] For example, a forward VSLAM processing unit 24 and a rearward VSLAM processing unit 24 are provided. The forward imaging unit 12A and the rearward imaging unit 12D capture images in different directions relative to the moving object 2, and each VSLAM processing unit 24 acquires forward point cloud information and rearward point cloud information in parallel. The difference calculation unit 292 and offset processing unit 293, and the alignment processing unit, perform the alignment processing using the forward point cloud information and rearward point cloud information acquired in parallel.
[0170] With this configuration, the multi-directional VSLAM processes complement each other, which further resolves the lack of detection information and enables the generation of a highly reliable surroundings map.
[0171] Although the embodiments and modifications have been described above, the information processing device, information processing method, and information processing program disclosed herein are not limited to the above-described embodiments, and the components can be modified and embodied in each implementation stage without departing from the spirit of the invention. Furthermore, various inventions can be created by appropriately combining multiple components disclosed in the above-described embodiments and modifications. For example, some components may be deleted from all of the components shown in the embodiments.
[0172] The information processing device 10 of the above embodiment and each modified example can be applied to various devices. For example, the information processing device 10 of the above embodiment and each modified example can be applied to a surveillance camera system that processes images obtained from a surveillance camera, or an in-vehicle system that processes images of the surrounding environment outside the vehicle. [Explanation of symbols]
[0173] 10. Information processing equipment 12, 12A~12D Photography Department 14 Detector 20 Acquisition Department 23 Selection section 24 VSLAM processing unit 25 Matching Section 26 Memory section 26A Environmental Map Information 27A Self-position estimation part 27B 3D reconstruction section 28 Correction section 29 Integrated Processing Unit 30 Decision Section 30A Absolute distance conversion unit 30B Extraction part 30C Nearest neighbor identification part 30D Reference projection surface shape selection section 30E Scale determination section 30F Asymptotic curve calculation section 30G shape determining section 30H Boundary area determination part 32 Deformed part 34 Virtual viewpoint line of sight determination unit 36 Projection transformation unit 37 Image generation unit 38 Image synthesis unit 291 Past Map Storage Department 292 Difference calculation part 293 Offset Processing Unit 294 Integration Department
Claims
1. an acquisition unit that acquires first point cloud information based on first image data acquired from a first imaging unit provided at a first position of the moving body, and acquires second point cloud information based on second image data acquired from a second imaging unit provided at a second position different from the first position of the moving body; a registration processing unit that executes registration processing between the first point cloud information and the second point cloud information; an integration processing unit that generates integrated point cloud information using the first point cloud information and the second point cloud information on which the alignment processing has been performed; Equipped with the acquisition unit transitions from a first process of acquiring the first point cloud information to a second process of acquiring the second point cloud information based on status information indicating a status of the moving object; the registration processing unit executes the registration processing using the first point cloud information acquired in the first processing and the second point cloud information acquired in the second processing. Information processing device.
2. the registration processing unit calculates difference information including at least one of a translation amount and a rotation amount of at least one of the first point cloud information and the second point cloud information, and performs the registration processing based on the difference information. The information processing device according to claim 1 .
3. The state information includes at least one of information for changing a traveling direction of the moving object and speed information of the moving object.
3. The information processing device according to claim 1 or 2.
4. The first imaging unit and the second imaging unit capture images of different directions with respect to the moving object. The information processing device according to claim 1 .
5. the acquisition unit executes a first process of acquiring the first point cloud information and a second process of acquiring the second point cloud information in parallel; the registration processing unit performs the registration process using the first point cloud information and the second point cloud information acquired in parallel; The information processing device according to claim 1 .
6. an image generating unit that projects an image of the periphery of the moving object, the image including the first image data and the second image data, onto a projection surface; a deformation unit that deforms the projection surface based on the integrated point cloud information; The information processing device according to claim 1 , further comprising:
7. 1. A computer-implemented information processing method, comprising: acquiring first point cloud information based on first image data obtained from a first photographing unit provided at a first position of the moving body, and acquiring second point cloud information based on second image data obtained from a second photographing unit provided at a second position different from the first position of the moving body; performing a registration process between the first point cloud information and the second point cloud information; generating integrated point cloud information using the first point cloud information and the second point cloud information on which the registration process has been performed; In the acquiring step, a transition is made from a first process of acquiring the first point cloud information to a second process of acquiring the second point cloud information based on status information indicating a status of the moving body; In the step of performing the registration process, the registration process is performed using the first point cloud information acquired in the first process and the second point cloud information acquired in the second process. Information processing methods.
8. On the computer, acquiring first point cloud information based on first image data obtained from a first photographing unit provided at a first position of the moving body, and acquiring second point cloud information based on second image data obtained from a second photographing unit provided at a second position different from the first position of the moving body; performing a registration process between the first point cloud information and the second point cloud information; generating integrated point cloud information using the first point cloud information and the second point cloud information on which the registration process has been performed; Run the command, In the acquiring step, a transition is made from a first process of acquiring the first point cloud information to a second process of acquiring the second point cloud information based on status information indicating a status of the moving body; In the step of executing the registration process, the registration process is executed using the first point cloud information acquired in the first process and the second point cloud information acquired in the second process. Information processing program.
Citation Information
Patent Citations
Travel support system, travel support program, and travel support method
JP2012118871A
Information processor, method for information processing, and program
JP2016045874A
Image processing system and image processor
JP2016123021A
Clearance limit determination device
JP2017083245A
Environment map generation method, environment map generation apparatus, and environment map generation program
JP2018205949A