Information processing device, information processing method, and information processing program

The information processing device addresses unstable VSLAM position detection by buffering and selectively processing image data from multiple sources, enhancing the accuracy of surrounding object and vehicle position estimation.

JP7761057B2Active Publication Date: 2025-10-28SOCIONEXT INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2023559277
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-10-28
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

Existing VSLAM processing often results in unstable detection of surrounding object positions and the vehicle's own position due to insufficient position information.

Method used

An information processing device with a buffering unit that buffers image data from multiple imaging units and a VSLAM processing unit that executes VSLAM processing using extracted image data, incorporating features like image thinning and determination information to enhance position estimation.

Benefits of technology

Resolves the issue of insufficient position information in VSLAM processing by stabilizing the detection of surrounding objects and the vehicle's position through improved image data handling and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007761057000001
    Figure 0007761057000001
  • Figure 0007761057000002
    Figure 0007761057000002
  • Figure 0007761057000003
    Figure 0007761057000003
Patent Text Reader

Abstract

In one embodiment, an information processing device (10) comprises a buffering unit (23) and a VSLAM processing unit (24). The buffering unit (23) buffers image data of moving body surroundings, said image data being obtained by an imaging unit (12) of a moving body (2), and transmits extraction image data extracted from among the buffered image data on the basis of extraction image determination information. The VSLAM processing unit (24) performs VSLAM processing using the extraction image data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] There is a technology called SLAM (Simultaneous Localization and Mapping) that acquires point cloud information about three-dimensional objects around a moving object such as a vehicle and estimates the position information of the vehicle and surrounding three-dimensional objects. There is also a technology called Visual SLAM (Simultaneous Localization and Mapping: VSLAM) that performs SLAM using images captured by a camera. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-062684 [Patent Document 2] Patent Publication No. 2021-082181 [Patent Document 3] International Publication No. 2020 / 246261 [Patent Document 4] Japanese Patent Application Publication No. 2018-205949 [Patent Document 5] Japanese Patent Application Laid-Open No. 2016-045874 [Patent Document 6] Japanese Patent Application Laid-Open No. 2016-123021 [Patent Document 7] International Publication No. 2019 / 073795 [Patent Document 8] International Publication No. 2020 / 246261 [Non-patent literature]

[0004] [Non-Patent Document 1] “Vision SLAM Using Omni·Directional Visual Scan Matching” 2008 IEEE / RSJ International Conference on Intelligent Robots and Systems Sept. 22·26, 2008 Summary of the Invention [Problem to be solved by the invention]

[0005] However, for example, in VSLAM processing, the position information of surrounding objects obtained by VSLAM processing may be insufficient, which may result in unstable detection of the positions of surrounding objects and the vehicle's own position by VSLAM.

[0006] In one aspect, the present invention aims to provide an information processing device, an information processing method, and an information processing program that solve the problem of insufficient position information of surrounding objects obtained by VSLAM processing. [Means for solving the problem]

[0007] In one aspect, the information processing device disclosed herein includes a buffering unit and a VSLAM processing unit. The buffering unit buffers image data of the area around a moving object obtained by an imaging unit of the moving object, and outputs extracted image data extracted from the buffered image data based on extracted image determination information. The VSLAM processing unit executes VSLAM processing using the extracted image data. [Effects of the Invention]

[0008] According to one aspect of the information processing device disclosed in the present application, it is possible to resolve the lack of position information of surrounding objects obtained by VSLAM processing. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the overall configuration of an information processing system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of the information processing device according to the embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of a functional configuration of the information processing device according to the embodiment. [Figure 4] FIG. 4 is a diagram showing an example of the configuration of the image buffering unit. [Figure 5] FIG. 5 is a schematic diagram illustrating an example of environment map information according to the embodiment. [Figure 6] FIG. 6 is a plan view showing the trajectory of a moving object when the moving object moves forward, stops temporarily, and then moves backward to park in a parking space. [Figure 7] FIG. 7 is a diagram for explaining the timing at which the buffering VSLAM process is started when the moving object moves along the trajectory shown in FIG. [Figure 8] FIG. 8 is a diagram for explaining the spatial range for acquiring the left photographed image used in the buffering VSLAM process when the moving object moves along the trajectory shown in FIG. [Figure 9] FIG. 9 is a diagram for explaining the buffering VSLAM process that starts at the trigger occurrence time. [Figure 10] FIG. 10 is a diagram for explaining the buffering VSLAM process when one second has elapsed since the timing of generation of the trigger information shown in FIG. [Figure 11] FIG. 11 is a diagram for explaining the buffering VSLAM process when one second has elapsed since the trigger information generation timing shown in FIG. [Figure 12] FIG. 12 is a diagram for explaining the buffering VSLAM process when one second has elapsed from the point in time when two seconds have elapsed since the timing of generation of the trigger information shown in FIG. [Figure 13] FIG. 13 is a diagram for explaining the buffering VSLAM process when 2.5 seconds have passed since 3 seconds had passed since the timing of occurrence of the trigger information shown in FIG. [Figure 14] FIG. 14 is an explanatory diagram of an asymptotic curve generated by the determination unit. [Figure 15] FIG. 15 is a schematic diagram showing an example of the reference projection plane. [Figure 16] FIG. 16 is a schematic diagram showing an example of a projection shape determined by the determination unit. [Figure 17] FIG. 17 is a schematic diagram illustrating an example of a functional configuration of the determination unit. [Figure 18] FIG. 18 is a flowchart showing an example of the flow of information processing executed by the information processing device. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, with reference to the accompanying drawings, embodiments of the information processing device, information processing method, and information processing program disclosed herein will be described in detail. Note that the following embodiments do not limit the disclosed technology. Furthermore, each embodiment can be appropriately combined within a range that does not cause contradiction in the processing content.

[0011] 1 is a diagram showing an example of the overall configuration of an information processing system 1 according to this embodiment. The information processing system 1 includes an information processing device 10, an imaging unit 12, a detection unit 14, and a display unit 16. The information processing device 10, the imaging unit 12, the detection unit 14, and the display unit 16 are connected to each other so as to be able to exchange data or signals.

[0012] In this embodiment, the information processing device 10, the image capturing unit 12, the detection unit 14, and the display unit 16 are mounted on a moving object 2 as an example.

[0013] The moving body 2 is an object that can move. The moving body 2 is, for example, a vehicle, a flyable object (a manned airplane, an unmanned airplane (e.g., a UAV (Unmanned Aerial Vehicle), a drone)), a robot, etc. The moving body 2 is, for example, a moving body that moves through human driving operation, or a moving body that can move automatically (autonomously) without human driving operation. In this embodiment, a case where the moving body 2 is a vehicle will be described as an example. Examples of vehicles include a two-wheeled vehicle, a three-wheeled vehicle, and a four-wheeled vehicle. In this embodiment, a case where the vehicle is an autonomously moving four-wheeled vehicle will be described as an example.

[0014] Note that the information processing device 10, the photographing unit 12, the detection unit 14, and the display unit 16 are not limited to being all mounted on the moving object 2. The information processing device 10 may be mounted on a stationary object. A stationary object is an object fixed to the ground. A stationary object is an object that cannot be moved or an object that is stationary relative to the ground. Examples of stationary objects include traffic lights, parked vehicles, and road signs. Furthermore, the information processing device 10 may be mounted on a cloud server that executes processing on the cloud.

[0015] The photographing unit 12 photographs the periphery of the moving object 2 and acquires photographed image data. In the following description, the photographed image data will be simply referred to as a photographed image. The photographing unit 12 is, for example, a digital camera capable of shooting video. Photographing refers to converting an image of a subject formed by an optical system such as a lens into an electrical signal. The photographing unit 12 outputs the photographed image to the information processing device 10. In addition, in this embodiment, the photographing unit 12 will be described assuming that it is a monocular fisheye camera (for example, with a viewing angle of 195 degrees).

[0016] In this embodiment, an example will be described in which four imaging units 12, namely, a front imaging unit 12A, a left imaging unit 12B, a right imaging unit 12C, and a rear imaging unit 12D, are mounted on a moving object 2. The multiple imaging units 12 (front imaging unit 12A, left imaging unit 12B, right imaging unit 12C, and rear imaging unit 12D) each capture an image of a subject in a different imaging area E (front imaging area E1, left imaging area E2, right imaging area E3, and rear imaging area E4) and acquire a captured image. That is, the imaging directions of the multiple imaging units 12 are different from each other. Furthermore, the imaging directions of the multiple imaging units 12 are adjusted in advance so that at least a portion of the imaging area E of adjacent imaging units 12 overlaps. Furthermore, for convenience of explanation, the imaging area E in FIG. 1 is shown as being the same size as in FIG. 1, but in reality, it may include an area further away from the moving object 2.

[0017] Furthermore, the four photographing units 12A, 12B, 12C, and 12D are merely examples, and there is no limit to the number of photographing units 12. For example, if the moving body 2 has a vertically long shape like a bus or truck, it is possible to use a total of six photographing units 12, by arranging one photographing unit 12 at the front, one at the rear, one at the front of the right side, one at the rear of the right side, one at the front of the left side, and one at the rear of the left side of the moving body 2. In other words, the number and arrangement positions of the photographing units 12 can be set arbitrarily depending on the size and shape of the moving body 2.

[0018] The detection unit 14 detects position information of each of a plurality of detection points around the moving object 2. In other words, the detection unit 14 detects position information of each of the detection points in the detection area F. The detection points refer to each of the points in real space that are individually observed by the detection unit 14. The detection points correspond to, for example, three-dimensional objects around the moving object 2.

[0019] The position information of the detection point is information indicating the position of the detection point in real space (three-dimensional space). For example, the position information of the detection point is information indicating the distance from the detection unit 14 (i.e., the position of the moving object 2) to the detection point and the direction of the detection point relative to the detection unit 14. These distances and directions can be expressed, for example, by position coordinates indicating the relative position of the detection point relative to the detection unit 14, position coordinates indicating the absolute position of the detection point, or a vector.

[0020] The detection unit 14 may be, for example, a 3D (Three-Dimensional) scanner, a 2D (Two-Dimensional) scanner, a distance sensor (millimeter-wave radar, laser sensor), a sonar sensor that detects objects using sound waves, an ultrasonic sensor, or the like. The laser sensor may be, for example, a three-dimensional LiDAR (Laser Imaging Detection and Ranging) sensor. The detection unit 14 may also be a device that uses a technology that measures distance from images captured by a stereo camera or a monocular camera, such as SfM (Structure from Motion) technology. A plurality of image capture units 12 may also be used as the detection unit 14. One of the plurality of image capture units 12 may also be used as the detection unit 14.

[0021] The display unit 16 displays various types of information and is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.

[0022] In this embodiment, the information processing device 10 is communicably connected to an electronic control unit (ECU) 3 mounted on the moving object 2. The ECU 3 is a unit that performs electronic control of the moving object 2. In this embodiment, the information processing device 10 is capable of receiving CAN (Controller Area Network) data such as the speed and moving direction of the moving object 2 from the ECU 3.

[0023] Next, the hardware configuration of the information processing device 10 will be described.

[0024] FIG. 2 is a diagram illustrating an example of a hardware configuration of the information processing device 10. As shown in FIG.

[0025] The information processing device 10 is, for example, a computer, and includes a CPU (Central Processing Unit) 10A, a ROM (Read Only Memory) 10B, a RAM (Random Access Memory) 10C, and an I / F (Interface) 10D. The CPU 10A, ROM 10B, RAM 10C, and I / F 10D are interconnected by a bus 10E, and have a hardware configuration similar to that of a typical computer.

[0026] The CPU 10A is a calculation device that controls the information processing device 10. The CPU 10A corresponds to an example of a hardware processor. The ROM 10B stores programs and the like that realize various processes by the CPU 10A. The RAM 10C stores data necessary for various processes by the CPU 10A. The I / F 10D is an interface that is connected to the imaging unit 12, the detection unit 14, the display unit 16, the ECU 3, etc., and is used to send and receive data.

[0027] A program for executing information processing executed by the information processing device 10 of this embodiment is provided by being pre-installed in a ROM 10B or the like. The program executed by the information processing device 10 of this embodiment may be provided by being recorded on a recording medium in a format that can be installed on the information processing device 10 or in a format that can be executed. The recording medium is a computer-readable medium. Examples of the recording medium include a CD (Compact Disc)-ROM, a flexible disk (FD), a CD-R (Recordable), a DVD (Digital Versatile Disk), a USB (Universal Serial Bus) memory, and an SD (Secure Digital) card.

[0028] Next, a functional configuration of the information processing device 10 according to this embodiment will be described. The information processing device 10 uses Visual SLAM processing to simultaneously estimate position information of the detection point and self-position information of the moving object 2 from images captured by the imaging unit 12. The information processing device 10 stitches together multiple spatially adjacent captured images to generate and display a composite image overlooking the periphery of the moving object 2. In this embodiment, the imaging unit 12 is used as the detection unit 14.

[0029] Fig. 3 is a diagram showing an example of the functional configuration of the information processing device 10. In Fig. 3, in order to clarify the data input / output relationship, an image capturing unit 12 and a display unit 16 are also shown in addition to the information processing device 10.

[0030] The information processing device 10 includes an acquisition unit 20, an image buffering unit 23, a VSLAM processing unit 24, a determination unit 30, a deformation unit 32, a virtual viewpoint line of sight determination unit 34, a projection transformation unit 36, and an image synthesis unit 38.

[0031] Some or all of the above-mentioned multiple units may be realized by, for example, causing a processing device such as CPU 10A to execute a program, that is, by software. Also, some or all of the above-mentioned multiple units may be realized by hardware such as an IC (Integrated Circuit), or may be realized by a combination of software and hardware.

[0032] The acquisition unit 20 acquires captured images from the imaging unit 12. That is, the acquisition unit 20 acquires captured images from each of the front imaging unit 12A, the left imaging unit 12B, the right imaging unit 12C, and the rear imaging unit 12D.

[0033] The acquisition unit 20 outputs the acquired captured image to the projection transformation unit 36 ​​and the image buffering unit 23 every time it acquires the captured image.

[0034] The image buffering unit 23 buffers the captured images sent from the image capture unit 12, thins them out, and then sends them to the VSLAM processing unit 24. Furthermore, in the buffering VSLAM processing described below, the image buffering unit 23 buffers image data of the surroundings of the moving object 2 captured by the image capture unit 12 of the moving object 2, and sends extracted image data from the buffered image data. Here, the extracted image data is image data extracted from the buffered image based on the extracted image determination information for a predetermined capture period, thinning interval, capture direction, and area within the image. Furthermore, the extracted image determination information includes, for example, information determined based on at least one of information on the movement state of the moving object 2, instruction information from the occupant of the moving object 2, information on objects surrounding the moving object 2 identified by a surrounding object detection sensor mounted on the moving object 2, and information on the surroundings of the moving object 2 recognized based on the image data obtained by the image capture unit 12.

[0035] Fig. 4 is a diagram showing an example of the configuration of image buffering unit 23. As shown in Fig. 4, image buffering unit 23 has a first storage unit 230a, a first thinning unit 231a, a second storage unit 230b, a second thinning unit 231b, a sending unit 232, and a sending data determination unit 233. In this embodiment, for the sake of concrete explanation, an example will be given in which images are sent from each of left imaging unit 12B and right imaging unit 12C to image buffering unit 23 at a frame rate of 30 fps via acquisition unit 20.

[0036] The transmission data determination unit 233 receives inputs of vehicle state information included in the CAN data received from the ECU 3, instruction information from the occupant of the mobile body, information identified by a surrounding object detection sensor mounted on the mobile body, information on the recognition of a specific image, etc. Here, the vehicle state information is information including, for example, the traveling direction of the mobile body 2, the state of direction instructions of the mobile body 2, the gear state of the mobile body 2, etc. The instruction information from the occupant of the mobile body is, for example, information input by a user's operational instruction, assuming a case in which the type of parking to be performed, such as perpendicular parking or parallel parking, is selected in automatic parking mode.

[0037] The transmission data determination unit 233 generates extraction image determination information based on vehicle state information, instruction information from a passenger in the moving object 2, information identified by a surrounding object detection sensor mounted on the moving object 2, information on the recognition of a specific image, and the like. The extraction image determination information includes, for example, information on the shooting period, thinning interval, shooting direction, and specific area within the image of the image data to be subjected to VSLAM processing. The specific area within the image can be derived using POI (Point of Interest) technology, for example. The extraction image determination information is output to the first thinning unit 231a, the second thinning unit 231b, and the transmission unit 232.

[0038] First storage unit 230a receives images captured by left imaging unit 12B and sent at a frame rate of, for example, 30 fps, and stores the images for, for example, one second (i.e., 30 frames at 30 fps). First storage unit 230a also updates the images it stores at a predetermined cycle.

[0039] The first thinning unit 231a thins out and reads out a plurality of frames of images stored in the first storage unit 230a. The first thinning unit 231a controls the rate at which images are thinned out (thinning intervals) based on extraction image determination information. The first thinning unit 231a also temporarily stores the images read out from the first storage unit 230a.

[0040] The second storage unit 230b receives images captured by the right imaging unit 12C and sent at a frame rate of, for example, 30 fps, and stores the images for, for example, one second (i.e., 30 frames at 30 fps). The first storage unit 230a updates the images it stores at a predetermined cycle.

[0041] The second thinning unit 231b thins out and reads out a plurality of frames of images stored in the second storage unit 230b. The second thinning unit 231b controls the rate at which images are thinned out (thinning intervals) based on the extraction image determination information. The second thinning unit 231b also temporarily stores the images read out from the second storage unit 230b.

[0042] The sending unit 232 determines the shooting direction (also called "target direction") to be used in the buffering VSLAM processing described later based on the extracted image determination information. The sending unit 232 reads out images corresponding to the determined target direction from the first thinning unit 231a and the second thinning unit 231b one by one in chronological order, and sends them sequentially to the VSLAM processing unit 24.

[0043] 3, the VSLAM processing unit 24 executes VSLAM processing using the captured image sent from the image buffering unit 23. That is, the VSLAM processing unit 24 receives the captured image from the image buffering unit 23, executes VSLAM processing using the captured image to generate environment map information, and outputs the generated environment map information to the determination unit 30.

[0044] More specifically, the VSLAM processing unit 24 includes a matching unit 25, a storage unit 26, a self-position estimation unit 27A, a three-dimensional reconstruction unit 27B, and a correction unit .

[0045] The matching unit 25 performs a feature extraction process and a matching process between multiple captured images (multiple captured images in different frames) captured at different times. In detail, the matching unit 25 performs a feature extraction process from these multiple captured images. The matching unit 25 performs a matching process for identifying corresponding points between the multiple captured images captured at different times, using the feature amounts between the multiple captured images. The matching unit 25 outputs the matching process result to the storage unit 26.

[0046] The self-position estimation unit 27A estimates the self-position relative to the captured image by projective transformation or the like using the multiple matching points acquired by the matching unit 25. Here, the self-position includes information on the position (three-dimensional coordinates) and tilt (rotation) of the image capturing unit 12. The self-position estimation unit 27 stores the self-position information as point cloud information in the environmental map information 26A.

[0047] The three-dimensional reconstruction unit 27B performs perspective projection transformation processing using the movement amount (translation amount and rotation amount) of the self-position estimated by the self-position estimation unit 27A, and determines the three-dimensional coordinates of the matching points (coordinates relative to the self-position). The three-dimensional reconstruction unit 27B stores the peripheral position information, which is the determined three-dimensional coordinates, in the environmental map information 26A as point cloud information.

[0048] As a result, new peripheral position information and self-position information are sequentially added to the environment map information 26A as the moving object 2 on which the image capturing unit 12 is mounted moves.

[0049] The storage unit 26 stores various types of data. The storage unit 26 is, for example, a semiconductor memory element such as RAM or flash memory, a hard disk, an optical disk, or the like. The storage unit 26 may be a storage device provided outside the information processing device 10. The storage unit 26 may also be a storage medium. Specifically, the storage medium may store or temporarily store programs and various types of information downloaded via a LAN (Local Area Network), the Internet, or the like.

[0050] The environmental map information 26A is information in which point cloud information, which is peripheral position information calculated by the three-dimensional restoration unit 27B, and point cloud information, which is self-position information calculated by the self-position estimation unit 27A, are registered in a three-dimensional coordinate space with a predetermined position in the real space as the origin (reference position). The predetermined position in the real space may be determined based on, for example, a predetermined condition.

[0051] For example, the predetermined position is the position of the moving object 2 when the information processing device 10 executes the information processing of this embodiment. For example, assume that the information processing is executed at a predetermined timing, such as a parking scene of the moving object 2. In this case, the information processing device 10 may set the position of the moving object 2 when it determines that the predetermined timing has been reached as the predetermined position. For example, the information processing device 10 may determine that the predetermined timing has been reached when it determines that the behavior of the moving object 2 has become behavior indicative of a parking scene. Examples of behavior indicative of a parking scene due to backing up include when the speed of the moving object 2 drops below a predetermined speed, when the gear of the moving object 2 is shifted into reverse gear, or when a signal indicating the start of parking is received by a user's operation instruction, etc. Note that the predetermined timing is not limited to a parking scene.

[0052] Fig. 5 is a schematic diagram of an example of the environment map information 26A. As shown in Fig. 5, the environment map information 26A is information in which point cloud information, which is the position information (peripheral position information) of each of the detection points P, and point cloud information, which is the self-position information of the self-position S of the moving object 2, are registered at corresponding coordinate positions in the three-dimensional coordinate space. Note that Fig. 5 shows the self-positions S of the self-positions S1 to S3 as an example. The larger the value of the number following S, the closer the self-position S is to the current timing.

[0053] The correction unit 28 corrects the peripheral position information and the self-position information registered in the environmental map information 26A using, for example, the least squares method, so that the sum of the differences in distance in three-dimensional space between the three-dimensional coordinates calculated previously and the three-dimensional coordinates calculated newly for points that have been matched multiple times between multiple frames is minimized. Note that the correction unit 28 may also correct the amount of movement (translation amount and rotation amount) of the self-position used in the process of calculating the self-position information and the peripheral position information.

[0054] The timing of the correction process by the correction unit 28 is not limited. For example, the correction unit 28 may execute the correction process at a predetermined timing. The predetermined timing may be determined based on a predetermined condition, for example. In the present embodiment, the information processing device 10 is described as being configured to include the correction unit 28 as an example. However, the information processing device 10 may not be configured to include the correction unit 28.

[0055] (Buffering VSLAM processing) 6 to 13, the buffering VSLAM processing realized by the image buffering unit 23 and the VSLAM processing unit 24 will be described. The buffering VSLAM processing is a process in which image data of the surroundings of the moving object 2 obtained by image capture by the image capture unit 12 of the moving object 2 is buffered, and VSLAM processing is performed using extracted image data from the buffered image.

[0056] In the following, for the sake of specificity, an example will be described in which the buffering VSLAM process is used when the moving object 2 is parked backward.

[0057] FIG. 6 is a plan view showing paths OB1 to OB4 when a moving object 2 moves forward, stops temporarily, and then moves backward to park backward into a parking space PA. That is, in the example shown in FIG. 6, the moving object 2 travels through the parking lot from the left side of the drawing toward the parking space PA while decelerating (paths OB1 and OB2). The moving object 2 turns to the right of the traveling direction by more than a predetermined angle in order to park backward into the parking space PA (path OB3), then stops temporarily, and the gear of the moving object 2 is switched from drive "D" to reverse "R". Then, the moving object 2 moves backward to be parked backward into the parking space PA (path OB4). Note that car1, car2, and car3 each represent other moving objects parked in parking spaces other than the parking space PA.

[0058] Fig. 7 is a diagram for explaining the timing of starting the buffering VSLAM process when the moving object 2 moves along the trajectories OB1 to OB4 shown in Fig. 6. That is, when the speed of the moving object 2 traveling through the parking lot at position P1 becomes equal to or less than a first predetermined speed, the first storage unit 230a starts buffering left-side captured images of the left-side capturing area E2 by the left-side capturing unit 12B. Also, the second storage unit 230b starts buffering right-side captured images of the right-side capturing area E by the right-side capturing unit 12C. Thereafter, the first storage unit 230a and the second storage unit 230b continue buffering the captured images at a frame rate of 30 fps.

[0059] The determination of whether the speed of the moving object 2 has become equal to or lower than the first predetermined speed can be made based on the vehicle state information received by the image buffering unit 23.

[0060] The first thinning section 231a and the second thinning section 231b each perform thinning processing and output the captured image to the sending section 232.

[0061] 7, when the moving object 2 continues to travel and turns to the right by a predetermined angle or more (position P2), the sending unit 232 determines that the target direction is to the left based on the extracted image determination information including the vehicle state information, and starts sending the left captured image of the left shooting area E2 to the VSLAM processing unit 24. In this embodiment, the extracted image determination information that triggers the sending unit 232 to start sending the left captured image to the VSLAM processing unit 24 is also called "trigger information." The timing at which the trigger information is generated is an example of a predetermined timing.

[0062] The left photographed image to be transmitted by the transmission unit 232 corresponds to an image acquired and buffered during a predetermined period prior to the timing of occurrence of the trigger information.

[0063] The VSLAM processing unit 24 executes VSLAM processing using the left-side photographed image sent from the sending unit 232. The VSLAM processing using the left-side photographed image is executed until the moving object 2 travels on the path OB3 while slowing down, then stops temporarily and switches the gear from drive "D" to reverse "R" (position P3).

[0064] After the gear is switched from drive "D" to reverse "R", when the moving body 2 moves backward along the trajectory OB4 and is parked backward in the parking space PA, VSLAM processing is executed at a frame rate of 3 fps using rearward captured images of the rearward captured area E4 taken by the rearward capturing unit 12D.

[0065] In this way, when the captured image on which the VSLAM process is performed changes, the point cloud information obtained by each VSLAM process may be matched to generate integrated point cloud information. For example, in this embodiment, the point cloud information generated by the VSLAM process using the left-side captured image may be integrated with the point cloud information generated by the VSLAM process using the rearward captured image. In other words, map information obtained by the VSLAM process based on image data of the area around the moving object 2 before the change in the moving direction may be integrated with map information obtained by the VSLAM process based on image data of the area around the moving object 2 after the change in the moving direction.

[0066] FIG. 8 is a diagram for explaining the spatial range for acquiring the left photographed image used in the buffering VSLAM process when the moving object 2 moves along the trajectories OB1 to OB4 shown in FIG.

[0067] In Figure 8, position Tr indicates the position of the left photographing unit 12B at the time when the trigger information was generated, position PTr-1 indicates the position of the left photographing unit 12B at the time 1 second before the trigger information was generated, position PTr+1 indicates the position of the left photographing unit 12B at the time 1 second after the trigger information was generated, and position PTr+3.5 indicates the position of the left photographing unit 12B at the stopped position (3.5 seconds after the trigger information was generated).

[0068] As shown in FIG. 8, when trigger information is generated while the left-side image capture unit 12B is at position Tr, multiple frames of left-side captured images acquired during a period of one second prior to the generation of the trigger information are stored in the first storage unit 230a. These multiple frames of left-side captured images correspond to images captured over a range from position PTr-1 to position Tr in FIG. 8. In response to the generation of the trigger information (i.e., when the left-side image capture unit 12B reaches position Tr), the transmission unit 232 extracts the left-side captured images stored in the first storage unit 230a and stores them in the first thinning unit 231a. The first thinning unit 231a then begins transmitting the images sequentially in chronological order to the VSLAM processing unit 24 via the transmission unit 232. Therefore, the VSLAM processing unit 24 can perform VSLAM processing for reverse parking using multiple left-side captured images of car1 and car2 before the gear is shifted from drive (D) to reverse (R).

[0069] 8, it is assumed that the traveling speed of the moving object 2 falls below a predetermined threshold value one second after the trigger information is generated. In this case, the first thinning unit 231a changes the frame rate of the thinning process from 5 fps to 2 fps, for example. Therefore, VSLAM processing is performed using left-side captured images transmitted at a frame rate of 5 fps for the section L1 from position PTr-1 to position PTr+1, and using left-side captured images transmitted at a frame rate of 2 fps for the section L2 from position PTr+1 to position PTr+3.5.

[0070] 9A to 9C are diagrams for explaining the buffering VSLAM process that starts at the trigger occurrence time. In Fig. 9 to Fig. 13, the trigger occurrence time is represented as Tr, and the reference time is 0 s.

[0071] In response to the occurrence of a trigger, the first thinning unit 231a reads out a plurality of frames of left-side photographed images, which have been thinned out to correspond to 5 fps, from the first storage unit at a predetermined cycle. The first thinning unit 231a outputs the plurality of frames of left-side photographed images read out from the first storage unit 230a to the sending unit 232.

[0072] Similarly, in response to the occurrence of a trigger, the second thinning unit 231b reads out, at a predetermined interval, from the second storage unit 230b, a plurality of frames of right-side photographed images that have been thinned out to correspond to 5 fps. The second thinning unit 231b outputs the plurality of frames of right-side photographed images read out from the second storage unit 230b to the sending unit 232.

[0073] In this way, the left photographed image read by the first thinning unit 231a and the right photographed image read by the second thinning unit 231b are photographed images acquired during a period (an example of a first period) going back one second from the trigger occurrence time and stored in the first storage unit 230a and the second storage unit 230b. Note that Fig. 9 illustrates a case where the thinning rate is 1 / 6 and six frames of images #0, #6, #12, #18, #24, and #30 are read out.

[0074] The sending unit 232 starts sending the left-side photographed image (frame #0) corresponding to the determined photographing direction from among the left-side photographed image acquired from the first thinning unit 231a and the right-side photographed image acquired from the second thinning unit 231b to the VSLAM processing unit 24.

[0075] Fig. 10 is a diagram for explaining the buffering VSLAM processing when one second has elapsed since the timing of occurrence of the trigger information shown in Fig. 9. That is, Fig. 10 shows that, during one second from the timing of occurrence of the trigger information, the first thinning unit 231a thins out to a level equivalent to 5 fps, while the VSLAM processing is performed at a pace of 3 fps.

[0076] The first thinning unit 231a reads out a plurality of frames of left-side captured images (#36, #42, #48, #54, #60) thinned to correspond to 5 fps from the first storage unit 230a at a predetermined cycle and outputs the readout to the transmission unit 232.

[0077] Similarly, the second thinning unit 231b reads out multiple frames of right-side captured images (#36, #42, #48, #54, #60) that have been thinned out to correspond to 5 fps from the second storage unit 230b at a predetermined cycle and outputs them to the transmission unit 232.

[0078] The sending unit 232 sends the left photographed images (#0, #6, #12) corresponding to the determined target direction, out of the left photographed images acquired from the first thinning unit 231a and the right photographed images acquired from the second thinning unit 231b, to the VSLAM processing unit 24. The VSLAM processing unit 24 performs VSLAM processing using the multiple frames of the left photographed images (#0, #6, #12) received from the sending unit 232.

[0079] Fig. 11 is a diagram for explaining the buffering VSLAM process when one second has elapsed since the trigger information was generated as shown in Fig. 10, and another second has elapsed as the moving object 2 moves forward. That is, Fig. 11 shows the buffering VSLAM process that is executed until one second has elapsed (i.e., time 2s) after the thinning frame rate is changed from 5 fps to 2 fps as the moving object 2 decelerates, one second after the trigger information was generated.

[0080] The first thinning section 231a reads out a plurality of frames of left-side captured images (#75, #90) that have been thinned out to correspond to 2 fps from the first storage section 230a at a predetermined cycle, and outputs the readout to the sending section 232.

[0081] Similarly, the second thinning section 231b reads out a plurality of frames of right photographed images (#75, #90) that have been thinned out to correspond to 2 fps from the second storage section 230b at a predetermined cycle, and outputs the readout to the sending section 232.

[0082] The sending unit 232 sends the left photographed images (#18, #24, #30) corresponding to the determined target direction, out of the left photographed images acquired from the first thinning unit 231a and the right photographed images acquired from the second thinning unit 231b, to the VSLAM processing unit 24. The VSLAM processing unit 24 sequentially executes VSLAM processing using the multiple frames of the left photographed images (#18, #24, #30) received from the sending unit 232.

[0083] Fig. 12 is a diagram for explaining the buffering VSLAM process when one second has elapsed since two seconds had elapsed since the trigger information was generated as shown in Fig. 11. That is, Fig. 12 shows the buffering VSLAM process that is executed up to the point in time when two seconds have elapsed (i.e., time 3s) after one second has elapsed since the trigger information was generated and the thinning frame rate has been changed from equivalent to 5 fps to equivalent to 2 fps.

[0084] The first thinning unit 231a reads out multiple frames of left-side captured images (#105, #120) that have been thinned out to correspond to 2 fps during the one second period from time 2s to time 3s from the first storage unit 230a at a predetermined cycle and outputs them to the transmission unit 232.

[0085] Similarly, the second thinning unit 231b reads out multiple frames of right-side captured images (#105, #120) that have been thinned out to correspond to 2 fps during the one second period from time 2s to time 3s from the second storage unit 230b at a predetermined cycle and outputs them to the transmission unit 232.

[0086] During the one second period from two seconds after the trigger information is generated to three seconds after the trigger information is generated, the sending unit 232 sends the left side captured images (#36, #42, #48) corresponding to the determined target direction, out of the left side captured images acquired from the first thinning unit 231a and the right side captured images acquired from the second thinning unit 231b, to the VSLAM processing unit 24. The VSLAM processing unit 24 performs VSLAM processing using the multiple frames of the left side captured images (#36, #42, #48) received from the sending unit 232.

[0087] Fig. 13 is a diagram for explaining the buffering VSLAM process when two seconds have elapsed since three seconds had elapsed since the trigger information was generated as shown in Fig. 12. That is, Fig. 13 shows the buffering VSLAM process that is executed from the time when the moving object 2 stops 3.5 seconds after the trigger information was generated until the time when the moving object 2 starts moving backward 1.5 seconds after the trigger information was generated. Note that the period from the time when the trigger information was generated until the time when the moving object 2 stops 3.5 seconds after the trigger information was generated is an example of the second period.

[0088] During the two seconds from the time three seconds after the trigger information is generated to the time five seconds before reverse motion begins, the sending unit 232 sends the left side captured images (#54, #60, #75, #90, #105, #120) corresponding to the determined target direction out of the left side captured images acquired from the first thinning unit 231a and the right side captured images acquired from the second thinning unit 231b to the VSLAM processing unit 24. The VSLAM processing unit 24 performs VSLAM processing using the multiple frames of the left side captured images (#54, #60, #75, #90, #105, #120) received from the sending unit 232.

[0089] Therefore, VSLAM processing unit 24 sets the target period for acquiring left-side captured images as the period from the point in time preceding the occurrence of trigger information by the first period to the point in time when moving object 2 stops. In the five seconds from the occurrence of trigger information to the start of reverse movement, VSLAM processing unit 24 can complete VSLAM processing using 15 frames of left-side captured images (i.e., VSLAM processing at an average of 3 fps) before the start of reverse movement in order to reverse park into parking space PA.

[0090] Returning to FIG. 3, the determination unit 30 receives the environmental map information from the VSLAM processing unit 24, and calculates the distance between the moving body 2 and surrounding three-dimensional objects using the surrounding position information and self-position information stored in the environmental map information 26A.

[0091] The determining unit 30 determines the projection shape of the projection surface using the distance between the moving object 2 and the surrounding three-dimensional objects, and generates projection shape information. The determining unit 30 outputs the generated projection shape information to the transforming unit 32.

[0092] Here, the projection surface is a three-dimensional surface onto which an image of the periphery of the moving object 2 is projected. The image of the periphery of the moving object 2 is a captured image of the periphery of the moving object 2, which is captured by each of the image capturing units 12A to 12D. The projection shape of the projection surface is a three-dimensional (3D) shape that is virtually formed in a virtual space corresponding to the real space. In this embodiment, the determination of the projection shape of the projection surface executed by the determination unit 30 is referred to as a projection shape determination process.

[0093] Furthermore, the determination unit 30 calculates an asymptotic curve of the peripheral position information relative to the self-position, using the peripheral position information of the moving object 2 and the self-position information stored in the environment map information 26A.

[0094] FIG. 14 is an explanatory diagram of an asymptotic curve Q generated by the determination unit 30. Here, the asymptotic curve is an asymptotic curve of a plurality of detection points P in the environment map information 26A. FIG. 14 shows an example in which the asymptotic curve Q is displayed in a projected image obtained by projecting a captured image onto a projection surface when the moving object 2 is viewed from above. For example, it is assumed that the determination unit 30 has identified three detection points P in descending order of proximity to the self-position S of the moving object 2. In this case, the determination unit 30 generates the asymptotic curve Q of these three detection points P.

[0095] The determination unit 30 outputs the self-position and the asymptotic curve information to the virtual viewpoint line-of-sight determination unit 34.

[0096] The deformation unit 32 deforms the projection surface based on the projection shape information determined using the environment map information including the integrated point cloud information received from the determination unit 30. The deformation unit 32 is an example of a deformation unit.

[0097] FIG. 15 is a schematic diagram showing an example of a reference projection plane 40. FIG. 16 is a schematic diagram showing an example of a projection shape 41 determined by the determination unit 30. That is, the deformation unit 32 deforms the reference projection plane shown in FIG. 15, which is stored in advance, based on the projection shape information, and determines a deformed projection plane 42 as the projection shape 41 shown in FIG. 16. The deformation unit 32 generates deformed projection plane information based on the projection shape 41. The deformation of this reference projection plane is performed, for example, using the detection point P closest to the moving object 2 as a reference. The deformation unit 32 outputs the deformed projection plane information to the projection conversion unit 36.

[0098] Furthermore, for example, the deformation unit 32 deforms the reference projection plane into a shape that follows the asymptotic curve of a predetermined number of detection points P in order of proximity to the moving object 2, based on the projection shape information.

[0099] The virtual viewpoint line of sight determining unit 34 determines virtual viewpoint line of sight information based on the self-position and the asymptotic curve information.

[0100] Determination of virtual viewpoint line-of-sight information will be described with reference to FIGS. 14 and 16. The virtual viewpoint line-of-sight determination unit 34 determines, for example, a direction passing through a detection point P closest to the self-position S of the moving object 2 and perpendicular to the deformed projection plane as the line-of-sight direction. Furthermore, the virtual viewpoint line-of-sight determination unit 34 fixes, for example, the direction of the line-of-sight direction L, and determines the coordinates of the virtual viewpoint O as an arbitrary Z coordinate and arbitrary XY coordinates in a direction away from the asymptotic curve Q toward the self-position S. In this case, the XY coordinates may be coordinates of a position farther away from the asymptotic curve Q than the self-position S. The virtual viewpoint line-of-sight determination unit 34 then outputs virtual viewpoint line-of-sight information indicating the virtual viewpoint O and the line-of-sight direction L to the projection transformation unit 36. Note that, as shown in FIG. 16, the line-of-sight direction L may be a direction from the virtual viewpoint O toward the position of the vertex W of the asymptotic curve Q.

[0101] The projection conversion unit 36 ​​generates a projection image by projecting the captured image acquired from the image capturing unit 12 onto the deformed projection surface based on the deformed projection surface information and the virtual viewpoint line of sight information. The projection conversion unit 36 ​​converts the generated projection image into a virtual viewpoint image and outputs it to the image synthesis unit 38. Here, the virtual viewpoint image is an image obtained by viewing the projection image in an arbitrary direction from a virtual viewpoint.

[0102] With reference to FIG. 16, the projection image generation process by the projection transformation unit 36 ​​will be described in detail. The projection transformation unit 36 ​​projects the captured image onto the modified projection surface 42. Then, the projection transformation unit 36 ​​generates a virtual viewpoint image (not shown), which is an image obtained by viewing the captured image projected onto the modified projection surface 42 from an arbitrary virtual viewpoint O in a line of sight direction L. The position of the virtual viewpoint O may be set to, for example, the latest self-position S of the moving object 2. In this case, the X and Y coordinate values ​​of the virtual viewpoint O may be set to the X and Y coordinate values ​​of the latest self-position S of the moving object 2. Furthermore, the Z coordinate value (vertical position) of the virtual viewpoint O may be set to the Z coordinate value of the detection point P closest to the self-position S of the moving object 2. The line of sight direction L may be determined based on, for example, a predetermined criterion.

[0103] The line of sight direction L may be, for example, a direction from the virtual viewpoint O toward the detection point P that is closest to the self-position S of the moving object 2. The line of sight direction L may also be a direction that passes through the detection point P and is perpendicular to the deformed projection plane 42. Virtual viewpoint line of sight information indicating the virtual viewpoint O and the line of sight direction L is created by the virtual viewpoint line of sight determination unit 34.

[0104] For example, the virtual viewpoint line-of-sight determination unit 34 may determine, as the line-of-sight direction L, a direction that passes through the detection point P closest to the self-position S of the moving object 2 and is perpendicular to the deformed projection plane 42. Alternatively, the virtual viewpoint line-of-sight determination unit 34 may fix the direction of the line-of-sight direction L and determine the coordinates of the virtual viewpoint O as an arbitrary Z coordinate and arbitrary X and Y coordinates in a direction away from the asymptotic curve Q toward the self-position S. In this case, the X and Y coordinates may be coordinates of a position farther away from the asymptotic curve Q than the self-position S. The virtual viewpoint line-of-sight determination unit 34 then outputs virtual viewpoint line-of-sight information indicating the virtual viewpoint O and the line-of-sight direction L to the projection transformation unit 36. Note that, as shown in FIG. 16 , the line-of-sight direction L may be a direction from the virtual viewpoint O toward the position of the vertex W of the asymptotic curve Q.

[0105] The projection transformation unit 36 ​​receives virtual viewpoint line-of-sight information from the virtual viewpoint line-of-sight determination unit 34. By receiving the virtual viewpoint line-of-sight information, the projection transformation unit 36 ​​identifies a virtual viewpoint O and a line-of-sight direction L. Then, the projection transformation unit 36 ​​generates a virtual viewpoint image, which is an image viewed from the virtual viewpoint O in the line-of-sight direction L, from the captured image projected onto the modified projection surface 42. The projection transformation unit 36 ​​outputs the virtual viewpoint image to the image synthesis unit 38.

[0106] The image synthesis unit 38 generates a synthesized image by extracting a part or all of the virtual viewpoint images. For example, the image synthesis unit 38 performs a process of stitching together a plurality of virtual viewpoint images (here, four virtual viewpoint images corresponding to the imaging units 12A to 12D) in the boundary region between the imaging units.

[0107] The image synthesis unit 38 outputs the generated synthetic image to the display unit 16. The synthetic image may be a bird's-eye view image with a virtual viewpoint O above the moving object 2, or may be an image in which the moving object 2 is displayed semi-transparently with a virtual viewpoint O inside the moving object 2.

[0108] The projection transformation unit 36 ​​and the image synthesis unit 38 constitute an image generation unit 37. The image generation unit 37 is an example of an image generation unit.

[0109] [Configuration example of the determination unit 30] Next, an example of the detailed configuration of the determination unit 30 shown in FIG. 3 will be described.

[0110] Fig. 17 is a schematic diagram showing an example of the functional configuration of the determination unit 30. As shown in Fig. 17, the determination unit 30 includes a CAN buffering unit 29, an absolute distance conversion unit 30A, an extraction unit 30B, a nearest neighbor identification unit 30C, a reference projection surface shape selection unit 30D, a scale determination unit 30E, an asymptotic curve calculation unit 30F, a shape determination unit 30G, and a boundary region determination unit 30H.

[0111] The CAN buffering unit 29 buffers the vehicle status information included in the CAN data sent from the ECU 3, and sends the information to the absolute distance conversion unit 30A after thinning out the data. The vehicle status information buffered by the CAN buffering unit 29 and the image data buffered by the image buffering unit 23 can be associated with each other by time information.

[0112] The absolute distance conversion unit 30A converts the relative positional relationship between the self-position and the surrounding three-dimensional object, which can be known from the environmental map information, into the absolute value of the distance from the self-position to the surrounding three-dimensional object.

[0113] Specifically, for example, the speed data of the moving object 2 included in the vehicle state information sent from the CAN buffering unit 29 is used. For example, in the case of the environment map information 26A shown in FIG. 5, the relative positional relationship between the self-position S and multiple detection points P can be known, but the absolute value of the distance is not calculated. Here, the distance between the self-position S3 and the self-position S2 can be calculated based on the inter-frame period for calculating the self-position and the speed data during that period based on the vehicle state information. Since the relative positional relationship in the environment map information 26A is similar to that in real space, knowing the distance between the self-position S3 and the self-position S2 also makes it possible to calculate the absolute values ​​of the distances from the self-position S to all other detection points P. Note that when the detection unit 14 acquires the distance information of the detection points P, the absolute distance conversion unit 30A may be omitted.

[0114] Then, the absolute distance conversion unit 30A outputs the calculated measured distances of each of the multiple detection points P to the extraction unit 30B. In addition, the absolute distance conversion unit 30A outputs the calculated current position of the moving object 2 to the virtual viewpoint line of sight determination unit 34 as self-position information of the moving object 2.

[0115] The extraction unit 30B extracts detection points P that are present within a specific range from among the multiple detection points P for which the measurement distances have been received from the absolute distance conversion unit 30A. The specific range is, for example, a range from the road surface on which the moving object 2 is placed to a height equivalent to the vehicle height of the moving object 2. Note that the range is not limited to this range.

[0116] By extracting detection points P within the range by the extraction unit 30B, it is possible to extract detection points P of objects that may be obstacles to the progress of the moving body 2, objects located adjacent to the moving body 2, and the like.

[0117] Then, the extraction unit 30B outputs the measured distance of each of the extracted detection points P to the nearest neighbor identification unit 30C.

[0118] The nearest neighbor identification unit 30C divides the surroundings of the self-position S of the moving body 2 into specific ranges (for example, angular ranges), and for each range, identifies the detection point P closest to the moving body 2, or multiple detection points P in order of proximity to the moving body 2. The nearest neighbor identification unit 30C identifies the detection point P using the measured distance received from the extraction unit 30B. In this embodiment, a form in which the nearest neighbor identification unit 30C identifies multiple detection points P in order of proximity to the moving body 2 for each range will be described as an example.

[0119] The nearest neighbor specifying unit 30C outputs the measurement distance of the detection point P specified for each range to the reference projection surface shape selecting unit 30D, the scale determining unit 30E, the asymptotic curve calculating unit 30F, and the boundary region determining unit 30H.

[0120] The reference projection surface shape selection unit 30D selects the shape of the reference projection surface.

[0121] Here, the reference projection surface will be described with reference to Fig. 15. The reference projection surface 40 is, for example, a projection surface having a shape that serves as a reference when changing the shape of the projection surface. The shape of the reference projection surface 40 is, for example, bowl-shaped, cylindrical, etc. Note that Fig. 15 shows an example of a bowl-shaped reference projection surface 40.

[0122] The bowl-shaped container has a bottom surface 40A and a side wall surface 40B, one end of which is continuous with the bottom surface 40A and the other end of which is open. The width of the horizontal cross section of the side wall surface 40B increases from the bottom surface 40A toward the open end of the other end. The bottom surface 40A is, for example, circular. Here, a circular shape includes a perfect circle and other circular shapes such as an ellipse. The horizontal cross section is an orthogonal plane perpendicular to the vertical direction (arrow Z direction). The orthogonal plane is a two-dimensional plane along the arrow X direction, which is perpendicular to the arrow Z direction, and the arrow Y direction, which is perpendicular to the arrow Z direction and the arrow X direction. Hereinafter, the horizontal cross section and the orthogonal plane may be referred to as the XY plane. The bottom surface 40A may have a shape other than a circle, such as an egg shape.

[0123] The cylindrical shape is a shape consisting of a circular bottom surface 40A and a side wall surface 40B that is continuous with the bottom surface 40A. The side wall surface 40B that constitutes the cylindrical reference projection surface 40 has a cylindrical shape with an opening at one end that is continuous with the bottom surface 40A and an open other end. However, the side wall surface 40B that constitutes the cylindrical reference projection surface 40 has a shape in which the diameter in the XY plane is approximately constant from the bottom surface 40A side toward the opening at the other end. The bottom surface 40A may have a shape other than a circle, such as an egg shape.

[0124] In this embodiment, the case where the shape of the reference projection plane 40 is bowl-shaped as shown in Fig. 15 will be described as an example. The reference projection plane 40 is a three-dimensional model virtually formed in a virtual space with a bottom surface 40A that is a surface that substantially coincides with the road surface below the moving object 2 and the center of the bottom surface 40A being the self-position S of the moving object 2.

[0125] The reference projection surface shape selection unit 30D selects the shape of the reference projection surface 40 by reading one specific shape from multiple types of reference projection surfaces 40. For example, the reference projection surface shape selection unit 30D selects the shape of the reference projection surface 40 based on the positional relationship between the self-position and surrounding three-dimensional objects, the stabilization distance, etc. The shape of the reference projection surface 40 may also be selected based on an operational instruction from the user. The reference projection surface shape selection unit 30D outputs shape information of the determined reference projection surface 40 to the shape determination unit 30G. In this embodiment, as described above, an embodiment in which the reference projection surface shape selection unit 30D selects a bowl-shaped reference projection surface 40 will be described as an example.

[0126] The scale determination unit 30E determines the scale of the reference projection plane 40 of the shape selected by the reference projection plane shape selection unit 30D. For example, the scale determination unit 30E makes a decision to reduce the scale when there are multiple detection points P within a predetermined distance range from the self-position S. The scale determination unit 30E outputs scale information of the determined scale to the shape determination unit 30G.

[0127] The asymptotic curve calculation unit 30F uses each of the stabilization distances of the detection points P closest to the self-position S for each range from the self-position S received from the nearest neighbor identification unit 30C, and outputs asymptotic curve information of the calculated asymptotic curve Q to the shape determination unit 30G and the virtual viewpoint line of sight determination unit 34. Note that the asymptotic curve calculation unit 30F may calculate the asymptotic curve Q of the detection points P accumulated for each of multiple portions of the reference projection plane 40. Then, the asymptotic curve calculation unit 30F may output the asymptotic curve information of the calculated asymptotic curve Q to the shape determination unit 30G and the virtual viewpoint line of sight determination unit 34.

[0128] The shape determination unit 30G enlarges or reduces the reference projection plane 40, which has a shape indicated by the shape information received from the reference projection plane shape selection unit 30D, to the scale of the scale information received from the scale determination unit 30E. Then, the shape determination unit 30G determines, as the projection shape, a shape obtained by deforming the reference projection plane 40 after enlarging or reducing it so that it becomes a shape that follows the asymptotic curve information of the asymptotic curve Q received from the asymptotic curve calculation unit 30F.

[0129] Here, the determination of the projection shape will be described in detail with reference to Fig. 16. As shown in Fig. 16, the shape determination unit 30G determines, as the projection shape 41, a shape obtained by deforming the reference projection plane 40 into a shape that passes through a detection point P that is closest to the self-position S of the moving object 2, which is the center of the bottom surface 40A of the reference projection plane 40. A shape that passes through the detection point P means that the deformed side wall surface 40B is a shape that passes through the detection point P. The self-position S is the latest self-position S calculated by the self-position estimation unit 27.

[0130] That is, the shape determination unit 30G identifies the detection point P that is closest to the self-position S among the multiple detection points P registered in the environmental map information 26A. In detail, the XY coordinates of the center position (self-position S) of the moving object 2 are set to (X, Y) = (0, 0). Then, the shape determination unit 30G determines the X 2 +Y 2 The detected point P where the value of is the smallest is identified as the detected point P closest to the self-position S. Then, the shape determination unit 30G determines, as the projected shape 41, a shape obtained by deforming the side wall surface 40B of the reference projection plane 40 so that it passes through the detected point P.

[0131] More specifically, the shape determination unit 30G determines the deformed shape of the bottom surface 40A and a portion of the side wall surface 40B as the projected shape 41 so that, when the reference projection plane 40 is deformed, a portion of the side wall surface 40B becomes a wall surface passing through the detection point P closest to the moving object 2. The deformed projected shape 41 is, for example, a shape that is raised from a rising line 44 on the bottom surface 40A in a direction toward the center of the bottom surface 40A when viewed from the XY plane (planar view). "Raising" means, for example, bending or folding a portion of the side wall surface 40B and the bottom surface 40A in a direction toward the center of the bottom surface 40A so that the angle formed between the side wall surface 40B of the reference projection plane 40 and the bottom surface 40A becomes smaller. Note that in the raised shape, the rising line 44 may be located between the bottom surface 40A and the side wall surface 40B, and the bottom surface 40A may remain undeformed.

[0132] The shape determination unit 30G determines to deform the specific region on the reference projection plane 40 so that it protrudes to a position that passes through the detection point P when viewed from the viewpoint (planar view) of the XY plane. The shape and range of the specific region may be determined based on predetermined criteria. Then, the shape determination unit 30G determines to deform the reference projection plane 40 so that the distance from the self-position S continuously increases from the protruding specific region toward regions other than the specific region on the side wall surface 40B.

[0133] For example, it is preferable to determine the projected shape 41 so that the outer periphery of the cross section along the XY plane has a curved shape, as shown in Fig. 16. Note that the outer periphery of the cross section of the projected shape 41 is, for example, a circle, but may have a shape other than a circle.

[0134] The shape determination unit 30G may determine, as the projected shape 41, a shape obtained by deforming the reference projection plane 40 so that the shape follows an asymptotic curve. The shape determination unit 30G generates an asymptotic curve of a predetermined number of detection points P in a direction away from the detection point P closest to the self-position S of the moving object 2. The number of detection points P may be more than one. For example, the number of detection points P is preferably three or more. In this case, the shape determination unit 30G preferably generates an asymptotic curve of a plurality of detection points P located at positions that are at least a predetermined angle away from the self-position S. For example, the shape determination unit 30G may determine, as the projected shape 41, a shape obtained by deforming the reference projection plane 40 so that the shape follows the generated asymptotic curve Q for the asymptotic curve Q shown in FIG. 14.

[0135] The shape determination unit 30G may divide the surroundings of the self-position S of the moving body 2 into specific ranges, and for each range, identify the detection point P closest to the moving body 2, or multiple detection points P in order of proximity to the moving body 2. Then, the shape determination unit 30G may determine, as the projection shape 41, a shape obtained by deforming the reference projection plane 40 so as to become a shape that passes through the detection points P identified for each range, or a shape that follows the asymptotic curve Q of the identified multiple detection points P.

[0136] Then, the shape determination unit 30G outputs the projection shape information of the determined projection shape 41 to the deformation unit 32.

[0137] Next, an example of the flow of information processing including buffering VSLAM processing executed by the information processing device 10 according to this embodiment will be described.

[0138] FIG. 18 is a flowchart showing an example of the flow of information processing executed by the information processing device 10.

[0139] The first storage unit 230a and the second storage unit 230b of the image buffering unit 23 acquire and store the left captured image from the left capturing unit 12B and the right captured image from the right capturing unit 12C via the acquisition unit 20 (step S2). The image buffering unit 23 also generates extracted image determination information based on vehicle state information included in the CAN data received from the ECU 3, instruction information from a passenger in the moving object 2, information identified by a surrounding object detection sensor mounted on the moving object 2, information that a specific image has been recognized, and the like (step S4).

[0140] The first thinning section 231a and the second thinning section 231b of the image buffering section 23 perform the thinning process at a frame rate based on the extracted image determination information (step S6).

[0141] The sending unit 232 determines the target direction based on the extracted image determination information (step S8).

[0142] The sending unit 232 sends the captured image corresponding to the determined target direction (for example, the left captured image) to the matching unit 25 (step S9).

[0143] The matching unit 25 extracts features and performs matching processing (step S10) using a plurality of captured images taken at different times that were selected in step S12 and captured by the imaging unit 12 from among the captured images acquired in step S10. The matching unit 25 also registers information on corresponding points between the plurality of captured images taken at different times, which have been identified by the matching processing, in the storage unit 26.

[0144] The self-position estimation unit 27 reads the matching points and the environmental map information 26A (peripheral position information and self-position information) from the storage unit 26 (step S12). The self-position estimation unit 27 estimates the self-position relative to the captured image by projective transformation or the like using the multiple matching points acquired from the matching unit 25 (step S14), and registers the calculated self-position information in the environmental map information 26A (step S16).

[0145] The three-dimensional restoration unit 26B reads the environment map information 26A (peripheral position information and self-position information) (step S18). The three-dimensional restoration unit 26B performs perspective projection transformation processing using the movement amount (translation amount and rotation amount) of the self-position estimated by the self-position estimation unit 27, determines the three-dimensional coordinates of the matching point (coordinates relative to the self-position), and registers them as peripheral position information in the environment map information 26A (step S20).

[0146] The correction unit 28 reads the environment map information 26A (peripheral position information and self-position information). The correction unit 28 corrects (step S22) the peripheral position information and self-position information already registered in the environment map information 26A using, for example, the least squares method, so that the sum of the differences in distance in three-dimensional space between previously calculated three-dimensional coordinates and newly calculated three-dimensional coordinates for points that have been matched multiple times across multiple frames is minimized, and updates the environment map information 26A.

[0147] The absolute distance conversion unit 30A acquires the vehicle state information from the CAN buffering unit 29 (step S24), and performs a thinning process on the vehicle state information to correspond to the thinning process of the first thinning unit 231a or the second thinning unit 231b (step S26).

[0148] The absolute distance conversion unit 30A takes in the speed data (host vehicle speed) of the moving object 2 included in the CAN data received from the ECU 3 of the moving object 2. Using the speed data of the moving object 2, the absolute distance conversion unit 30A converts the peripheral position information included in the environmental map information 26A into distance information from the current position, which is the latest host position S of the moving object 2, to each of the plurality of detection points P (step S28). The absolute distance conversion unit 30A outputs the calculated distance information of each of the plurality of detection points P to the extraction unit 30B. In addition, the absolute distance conversion unit 30A outputs the calculated current position of the moving object 2 to the virtual viewpoint line of sight determination unit 34 as host position information of the moving object 2.

[0149] The extraction unit 30B extracts detection points P that exist within a specific range from among the plurality of detection points P for which distance information has been received (step S30).

[0150] The nearest neighbor identification unit 30C divides the surroundings of the self-position S of the moving body 2 into specific ranges, and for each range, identifies the detection point P closest to the moving body 2, or multiple detection points P in order of closest to the moving body 2, and extracts the distance to the nearest object (step S32). The nearest neighbor identification unit 30C outputs the measured distance d of the detection point P identified for each range (the measured distance between the moving body 2 and the nearest object) to the reference projection surface shape selection unit 30D, the scale determination unit 30E, the asymptotic curve calculation unit 30F, and the boundary region determination unit 30H.

[0151] The asymptotic curve calculation unit 30F calculates an asymptotic curve (step S34), and outputs it to the shape determination unit 30G and the virtual viewpoint line of sight determination unit 34 as asymptotic curve information.

[0152] The reference projection plane shape selection unit 30D selects the shape of the reference projection plane 40 (step S36), and outputs shape information of the selected reference projection plane 40 to the shape determination unit 30G.

[0153] The scale determination unit 30E determines the scale of the reference projection plane 40 of the shape selected by the reference projection plane shape selection unit 30D (step S38), and outputs scale information of the determined scale to the shape determination unit 30G.

[0154] The shape determination unit 30G determines a projection shape, which is how to deform the shape of the reference projection plane, based on the scale information and the asymptotic curve information (step S40). The shape determination unit 30G outputs projection shape information of the determined projection shape 41 to the deformation unit 32.

[0155] The transformation unit 32 transforms the shape of the reference projection plane based on the projection shape information (step S42), and outputs the transformed transformed projection plane information to the projection conversion unit .

[0156] The virtual viewpoint line-of-sight determination unit 34 determines virtual viewpoint line-of-sight information based on the self-position and the asymptotic curve information (step S44). The virtual viewpoint line-of-sight determination unit 34 outputs the virtual viewpoint line-of-sight information indicating the virtual viewpoint O and the line-of-sight direction L to the projection transformation unit 36.

[0157] The projection conversion unit 36 ​​generates a projection image by projecting the captured image acquired from the image capturing unit 12 onto the deformed projection surface based on the deformed projection surface information and the virtual viewpoint line of sight information. The projection conversion unit 36 ​​converts the generated projection image into a virtual viewpoint image (step S46) and outputs it to the image synthesis unit 38.

[0158] The boundary area determination unit 30H determines a boundary area based on the distance to the nearest object identified for each range. That is, the boundary area determination unit 30H determines a boundary area as an overlap area of ​​spatially adjacent peripheral images based on the position of the object nearest to the moving object 2 (step S48). The boundary area determination unit 30H outputs the determined boundary area to the image synthesis unit 38.

[0159] The image synthesis unit 38 generates a synthesized image by stitching spatially adjacent perspective projection images together using the boundary area (step S50). That is, the image synthesis unit 38 generates a synthesized image by stitching together perspective projection images from four directions according to the boundary area set at the angle of the nearest object direction. In the boundary area, the spatially adjacent perspective projection images are blended at a predetermined ratio.

[0160] The display unit 16 displays the composite image (step S52).

[0161] The information processing device 10 determines whether to end the information processing (step S54). For example, the information processing device 10 makes the determination in step S54 by determining whether or not a signal indicating that the position of the moving object 2 has stopped moving has been received from the ECU 3. Alternatively, for example, the information processing device 10 may make the determination in step S54 by determining whether or not an instruction to end the information processing has been received through an operation instruction by a user or the like.

[0162] If a negative determination is made in step S54 (step S54: No), the processes from step S2 to step S54 are repeatedly executed.

[0163] On the other hand, if the determination in step S54 is affirmative (step S54: Yes), this routine ends.

[0164] When returning from step S54 to step S2 after executing the correction process of step S22, the correction process of the subsequent step S22 may be omitted. Also, when returning from step S54 to step S2 without executing the correction process of step S22, the correction process of the subsequent step S22 may be executed.

[0165] As described above, the information processing device 10 according to the embodiment includes the image buffering unit 23 as a buffering unit, and the VSLAM processing unit 24 as a VSLAM processing unit. The image buffering unit 23 buffers image data of the surroundings of a moving object obtained by the imaging unit 12. The image buffering unit 23 sends extracted image data from the buffered images. The VSLAM processing unit 24 performs VSLAM processing using the sent image data.

[0166] Therefore, for example, during the period from the generation of trigger information to the start of reverse motion, the VSLAM processing unit 24 can complete VSLAM processing using captured images containing a large amount of three-dimensional information about the surroundings of the parking space before the vehicle starts to reverse. As a result, for example, when a vehicle is parking while turning around, information about objects near the parking position that are the last to enter the frame and the first to exit the frame (e.g., car1, car2, etc., shown in FIGS. 6 to 8) can be increased compared to conventional methods. Even when there are few objects near the parking position, the captured images obtained by capturing the area in front of the parking space are used, thereby increasing information about three-dimensional objects compared to conventional methods. Furthermore, even when the vehicle 2 is moving at a certain speed or faster, buffered images can be used at the desired frame rate, thereby substantially increasing information about three-dimensional objects compared to conventional methods. As a result, the lack of position information about surrounding objects obtained by VSLAM processing can be resolved, and VSLAM can stabilize detection of the positions of surrounding objects and the vehicle's own position.

[0167] The image buffering unit 23 sends image data obtained during a target period that includes a first period around the time the trigger information is generated and a second period from the first period until the moving object 2 stops. Therefore, even during the period until the moving object 2 decelerates and stops, during which the gear of the moving object 2 is switched from drive "D" to reverse "R," the VSLAM process can be continued using the buffered images.

[0168] The image buffering unit 23 sends out the extracted image data.

[0169] Therefore, for sections where the moving object 2 moves at a certain speed or faster, VSLAM processing can be performed using temporally adjacent captured images at a relatively high frame rate, for example. As a result, it is possible to substantially increase the amount of information on three-dimensional objects compared to conventional methods.

[0170] The image buffering unit 23 buffers at least a leftward photographed image obtained by photographing a first direction (left direction) and a rightward photographed image obtained by photographing a second direction (right direction) different from the first direction. The image buffering unit 23 sends out the leftward photographed image or the rightward photographed image based on extracted image determination information including an operation status such as the speed and turning angle of the mobile body, instruction information from the passenger of the mobile body, information identified by a surrounding object detection sensor mounted on the mobile body, and information that a specific image has been recognized.

[0171] Therefore, the VSLAM processing unit 24 can perform VSLAM processing using only the left-side captured image, which contains a large amount of information about three-dimensional objects near the parking space PA, out of the left-side captured image and the right-side captured image. As a result, a significant reduction in processing load can be achieved compared to VSLAM processing that uses both the left-side captured image and the right-side captured image.

[0172] (Variation 1) The captured images to be subjected to the buffering VSLAM process, i.e., the captured images from a position in front of the parking space PA to be subjected to the VSLAM process, can be adjusted arbitrarily by adjusting the length of the first period going back from the time when the trigger information was generated. For example, if the first period going back from the time when the trigger information was generated is set long, the VSLAM process can be performed using information on many three-dimensional objects that are in front of the parking space PA as seen from the mobile object 2. The first period can also be set to 0 as necessary. Furthermore, the buffering VSLAM process can also be performed after a delay of a third period from the time when the trigger information was generated.

[0173] (Variation 2) In the above embodiment, the buffering VSLAM process is started when the moving speed of the moving object 2 falls below a threshold value and the steering wheel (handle) is turned by more than a certain amount to turn, which are used as trigger information. In contrast to this, for example, the buffering VSLAM process may be started when an input instruction from a user is used as a trigger. Such an example of using an input instruction from a user as a trigger can be used, for example, when performing automatic parking. Furthermore, the trigger information may be an operating status such as the speed of the moving object 2, information identified by a surrounding object detection sensor mounted on the moving object 2, information that a specific image is recognized, or the like.

[0174] (Variation 3) In the above embodiment, the target direction to be used in the buffering VSLAM processing is determined to be left based on vehicle state information including the fact that the moving speed of the moving object 2 has fallen below a threshold and that the steering wheel (handle) has been turned by more than a certain amount to turn. However, for example, the target direction to be used in the buffering VSLAM processing may be determined using an input instruction from a user as a trigger. Such an example of determining the target direction based on an input instruction from a user can be used, for example, in automatic parking. Furthermore, the trigger information may include the operating status of the moving object 2, such as the speed of the moving object 2, information identified by a surrounding object detection sensor mounted on the moving object 2, or information that a specific image has been recognized.

[0175] (Variation 4) In the above embodiment, in order to reduce the load of VSLAM processing, the target direction is determined to be left based on vehicle state information, and buffering VSLAM processing is performed using only left-side captured images captured in the left-side shooting area E2. However, if the load of VSLAM processing is not a problem, it is also possible to perform buffering VSLAM processing using both left-side captured images captured in the left-side shooting area E2 and right-side captured images captured in the right-side shooting area E3.

[0176] (Variation 5) In the above embodiment, the left imaging area E2 and the right imaging area E3 are stored in the first storage unit 230a and the second storage unit 230b, respectively. Alternatively, the first storage unit 230a, the second storage unit 230b, or a new storage unit may be provided to input and store images captured by the front imaging unit 12A and the rear imaging unit 12D. Furthermore, if the moving body 2 is a drone, it is also possible to store upward captured images acquired by a capture unit provided on the top surface of the moving body 2, or downward captured images acquired by a capture unit provided on the bottom surface of the moving body 2, and use these to perform buffering VSLAM processing.

[0177] (Variation 6) In the above embodiment, an example has been described in which the buffering VSLAM process is used when the moving object 2 is parked backward. However, the buffering VSLAM process may also be used when the moving object 2 is parked forward.

[0178] With this configuration, the buffering VSLAM process and the normal VSLAM process complement each other, which further resolves the lack of detection information and enables the generation of a highly reliable surroundings map.

[0179] (Variation 7) In the above embodiment, an example was given in which all 30 fps images from the imaging unit are buffered and then thinned out in the buffering unit 23. In contrast, the buffering unit 23 may take images into the first storage unit 230a and the second storage unit 230b while thinning out the images at the maximum frame rate (e.g., 5 fps) used in VSLAM processing, and the first thinning unit 231a and the second thinning unit 231b may thin out the images from the first storage unit 230a and the second storage unit 230b for a VSLAM processing section with an even lower rate.

[0180] Although the embodiments and modifications have been described above, the information processing device, information processing method, and information processing program disclosed herein are not limited to the above-described embodiments, and the components can be modified and embodied in each implementation stage without departing from the spirit of the invention. Furthermore, various inventions can be created by appropriately combining multiple components disclosed in the above-described embodiments and modifications. For example, some components may be deleted from all of the components shown in the embodiments.

[0181] The information processing device 10 of the above embodiment and each modified example can be applied to various devices. For example, the information processing device 10 of the above embodiment and each modified example can be applied to a surveillance camera system that processes images obtained from a surveillance camera, or an in-vehicle system that processes images of the surrounding environment outside the vehicle. [Explanation of symbols]

[0182] 10. Information processing equipment 12, 12A~12D Photography Department 14 Detector 20 Acquisition Department 23 Image buffering unit 24 VSLAM processing unit 25 Matching Section 26 Memory section 26A Environmental Map Information 27A Self-position estimation part 27B 3D reconstruction section 28 Correction section 29 CAN buffering section 30 Decision Section 30A Absolute distance conversion unit 30B Extraction part 30C Nearest neighbor identification part 30D Reference projection surface shape selection section 30E Scale determination section 30F Asymptotic curve calculation section 30G shape determining section 30H Boundary area determination part 32 Deformed part 34 Virtual viewpoint line of sight determination unit 36 Projection transformation unit 37 Image generation unit 38 Image synthesis unit 230a 1st storage section 230b 2nd storage section 231a 1st thinning section 231b Second thinning section 232 Transmission Unit 233 Transmission data determination unit

Claims

1. a buffering unit that buffers image data of the area around the moving object obtained by the image capturing unit of the moving object, and that outputs extracted image data extracted at different thinning intervals from the buffered image data in chronological order in accordance with extracted image determination information; a VSLAM processing unit that executes VSLAM processing using the extracted image data; An information processing device comprising:

2. The extracted image determination information includes information determined based on at least one of information on the state of movement of the moving body, instruction information from a passenger of the moving body, information on objects in the vicinity of the moving body identified by a surrounding object detection sensor mounted on the moving body, and information on the vicinity of the moving body recognized based on image data obtained by a photographing unit. The information processing device according to claim 1 .

3. the buffering unit extracts, as the extracted image data, image data for a first period from the buffered image data based on the extracted image determination information; 3. The information processing device according to claim 1 or 2.

4. the buffering unit performs a thinning process on the buffered image data based on the extracted image determination information to generate the extracted image data; In the thinning process, the time interval for thinning out the buffered image data is different between the first period and a second period different from the first period. The information processing device according to claim 3 .

5. The buffering unit buffering first image data obtained by capturing an image in a first direction and second image data obtained by capturing an image in a second direction different from the first direction, among the image data around the moving object; transmitting, as the extracted image data, first extracted image data extracted from the first image data and second extracted image data extracted from the second image data based on the extracted image determination information; The information processing device according to claim 1 .

6. the buffered image data includes a plurality of frames; the extracted image data includes image data obtained by cutting out a partial area of ​​at least one frame of the buffered image data; The information processing device according to claim 1 .

7. the image data of the surroundings of the moving object is image data obtained by a plurality of the photographing units, When the moving object changes its direction of travel, the buffering unit buffers the image data of the area around the moving object acquired by the imaging unit that differs before and after the change. The information processing device according to claim 1 .

8. The VSLAM processing unit map information obtained by VSLAM processing based on image data of the surroundings of the moving object before the moving object changes its direction of travel; and integrating the image data of the surroundings of the moving body after the change in the moving direction of the moving body with map information obtained by VSLAM processing based on the image data. The information processing device according to claim 7 .

9. 1. A computer-implemented information processing method, comprising: Buffering image data of the surroundings of the moving object obtained by the imaging unit of the moving object; a step of transmitting extracted image data extracted at different thinning intervals according to extracted image determination information from the buffered image data in a time series; performing a VSLAM process using the extracted image data; An information processing method including:

10. On the computer, Buffering image data of the surroundings of the moving object obtained by the imaging unit of the moving object; a step of transmitting extracted image data from the buffered image data in time series at different thinning intervals according to extracted image determination information; performing a VSLAM process using the extracted image data; An information processing program for executing the above.

Citation Information

Patent Citations

  • Information processor, method for information processing, and program

    JP2016045874A

  • Image processing system and image processor

    JP2016123021A

  • Environment map generation method, environment map generation apparatus, and environment map generation program

    JP2018205949A

  • Mobile location estimation system and mobile location method

    JP2020153956A

  • Parking support apparatus

    JP2021062684A