Information processing device, information processing method, and program

The information processing apparatus improves self-position estimation by extracting key points from pedestrian skeletons to enhance localization accuracy in congested areas.

JP7876706B2Active Publication Date: 2026-06-19HONDA MOTOR CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HONDA MOTOR CO LTD
Filing Date
2023-03-22
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Conventional techniques for self-position estimation of mobile bodies fail to generate accurate environmental maps in congested areas with moving objects, leading to low robustness in self-position estimation.

Method used

An information processing apparatus that extracts feature points from stationary objects and key points from pedestrian skeletons, using skeleton recognition to improve self-position estimation by correcting the estimated position with key points, even when stationary or moving pedestrians are present.

Benefits of technology

Enhances the robustness of self-localization by utilizing both feature points from stationary objects and key points from pedestrian skeletons, especially in crowded environments where feature points are scarce.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007876706000001
    Figure 0007876706000001
  • Figure 0007876706000002
    Figure 0007876706000002
  • Figure 0007876706000003
    Figure 0007876706000003
Patent Text Reader

Abstract

The present invention provides an information processing device comprising: an image acquisition unit that acquires one or more images, captured in a time series, of the surrounding conditions of a mobile object; a feature point extraction unit that extracts feature points of a stationary object captured in the images; a key point extraction unit that extracts key points, which are features of a different type than the feature points, from skeleton information about a pedestrian captured in the images; and a localization unit that estimates the location of the mobile object in the surrounding conditions on the basis of the extracted feature points and the extracted key points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, a technique for estimating the self-position of a mobile body that autonomously travels has been known. For example, in Patent Document 1, a technique is described in which a laser is irradiated around a mobile body, the reflected light of the laser is received, an object existing around the mobile body is detected, an environmental map of the mobile body is generated, and when a moving object is detected, the moving object is deleted from the environmental map.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The technique described in Patent Document 1 generates an environmental map based only on stationary objects without considering moving objects. However, for example, when a mobile body travels on a road congested by people, it is not always possible to generate a highly accurate environmental map based only on stationary objects, and the conventional technique may have low robustness in self-position estimation.

[0005] The present invention has been made in consideration of such circumstances, and one of the objectives is to provide an information processing apparatus, an information processing method, and a program that can improve the robustness of self-position estimation.

Means for Solving the Problems

[0006] The vehicle control device according to this invention employs the following configuration. (1) An information processing device according to one aspect of the present invention comprises: an image acquisition unit that acquires one or more images of the surrounding situation of a moving object taken in a time series; a feature point extraction unit that extracts feature points of stationary objects captured in the images; a key point extraction unit that extracts key points, which are of a different type of feature from the feature points, from the skeleton information of a pedestrian captured in the images; and a self-position estimation unit that estimates the position of the moving object in the surrounding situation based on the extracted feature points and key points.

[0007] (2): In the embodiment of (1) above, the key point extraction unit extracts the location representing the ankle of the pedestrian as the key point.

[0008] (3) In the embodiment of (1) above, the self-position estimation unit estimates the position of the moving body based on the extracted feature points and corrects the estimated position of the moving body using the key points.

[0009] (4) In the embodiment of (3) above, the self-position estimation unit corrects the position of the moving body using the amount of displacement of the key points between the one or more images of the pedestrian captured in time series, when the pedestrian is stationary and the moving body equipped with a camera that takes one or more images in time series is moving.

[0010] (5) In the embodiment of (3) above, when the pedestrian and the mobile body equipped with a camera that takes one or more images in a time series are moving, the self-position estimation unit predicts the future position of the pedestrian from the amount of displacement of the key point between the one or more images of the pedestrian taken in a time series, and corrects the position of the mobile body using the predicted future position.

[0011] (6) An information processing method according to one aspect of the present invention involves a computer acquiring one or more images taken in chronological order of the surrounding conditions of a moving object, extracting feature points of stationary objects captured in the images, extracting key points which are of a different type of feature from the feature points from the skeleton information of pedestrians captured in the images, and estimating the position of the moving object in the surrounding conditions based on the extracted feature points and key points.

[0012] (7): A program according to one aspect of the present invention causes a computer to acquire one or more images of the surrounding situation of a moving object in a time series, to extract feature points of stationary objects captured in the images, to extract key points which are of a different type of feature from the feature points from the skeleton information of pedestrians captured in the images, and to estimate the position of the moving object in the surrounding situation based on the extracted feature points and key points. [Effects of the Invention]

[0013] According to the embodiments described in (1) to (7) above, the robustness of self-localization can be improved. [Brief explanation of the drawing]

[0014] [Figure 1] This figure shows an example of the configuration of the mobile body 1 and control device 100 according to the embodiment. [Figure 2] This is a perspective view of the moving object 1, seen from above. [Figure 3] This figure shows an example of an image IM acquired by the image acquisition unit 110. [Figure 4] This diagram illustrates the self-position correction performed by the self-position estimation unit 140 when pedestrian P is stationary. [Figure 5] This figure shows another example of an image IM acquired by the image acquisition unit 110. [Figure 6] This diagram illustrates the correction of the self-position performed by the self-position estimation unit 140 when a pedestrian P is moving. [Figure 7]This is a diagram showing an example of map information 72. [Figure 8] This is a flowchart showing an example of the flow of processing executed by the control device 100.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, embodiments of the information processing apparatus, information processing method, and program of the present invention will be described with reference to the drawings. The information processing apparatus of the present invention is mounted on, for example, a moving body. The moving body moves on both a roadway and a predetermined area different from the roadway. The moving body may be referred to as micromobility. An electric kick scooter is a type of micromobility. Also, the moving body may be a vehicle on which a passenger can board, or may be an autonomous moving body capable of autonomous driving without a driver. The latter autonomous moving body is used, for example, for transporting luggage and the like. The predetermined area is, for example, a sidewalk. Also, the predetermined area may be part or all of a roadside strip, a bicycle lane, an open space, etc., or may include all of a sidewalk, a roadside strip, a bicycle lane, an open space, etc.

[0016] FIG. 1 is a diagram showing an example of the configuration of the moving body 1 and the control device 100 according to the embodiment. The moving body 1 is equipped with, for example, an external detection device 10, a moving body sensor 12, an operator 14, an internal camera 16, a positioning device 18, an acceleration sensor 20, a mode changeover switch 22, a dial switch 24, a moving mechanism 30, a drive device 40, an external notification device 50, a storage device 70, and a control device 100. Note that some of these configurations that are not essential for realizing the functions of the present invention may be omitted.

[0017] The external detection device 10 is various devices that set the traveling direction of the moving body 1 as the detection range. The external detection device 10 includes an external camera, a radar device, LIDAR (Light Detection and Ranging), a sensor fusion device, and the like. The external detection device 10 outputs information (images, positions of objects, etc.) indicating the detection results to the control device 100.

[0018] The mobile body sensor 12 includes, for example, a speed sensor, a yaw rate (angular velocity) sensor, a direction sensor, and an operation amount detection sensor attached to the operator 14. The operator 14 includes, for example, an operator for instructing acceleration and deceleration (such as an accelerator pedal or a brake pedal) and an operator for instructing steering (such as a steering wheel). In this case, the mobile body sensor 12 may include an accelerator opening sensor, a brake depression amount sensor, a steering torque sensor, etc. The mobile body 1 may be provided with an operator in a form other than the above as the operator 14 (for example, a non-annular rotary operator, a joystick, a button, etc.).

[0019] The internal camera 16 images at least the head of the passenger of the mobile body 1 from the front. The internal camera 16 is a digital camera using an imaging device such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor). The internal camera 16 outputs the captured image to the control device 100.

[0020] The positioning device 18 is a device that positions the position of the mobile body 1. The positioning device 18 is, for example, a GNSS (Global Navigation Satellite System) receiver, and based on the signal received from the GNSS satellite, it identifies the position of the mobile body 1 and outputs it as position information. Note that the position information of the mobile body 1 may be estimated from the position of the Wi-Fi base station to which the communication device described later is connected.

[0021] The acceleration sensor 20 detects the acceleration of the mobile body 1 and outputs a signal corresponding to the detected acceleration to the control device 100. The acceleration sensor 20 detects the acceleration acting in the vertical direction (height direction) in addition to the horizontal direction of the mobile body 1.

[0022] The mode selector switch 22 is a switch operated by the occupant. The mode selector switch 22 may be a mechanical switch or a GUI (Graphical User Interface) switch set on a touch panel. The mode selector switch 22 accepts an operation to switch the driving mode to one of the following: Mode A: an assist mode in which either steering operation or acceleration / deceleration control is performed by the occupant, and the other is performed automatically, and there may be Mode A-1 in which steering operation is performed by the occupant and acceleration / deceleration control is performed automatically, and Mode A-2 in which acceleration / deceleration operation is performed by the occupant and steering control is performed automatically; Mode B: a manual driving mode in which steering operation and acceleration / deceleration operation are performed by the occupant; or Mode C: an automatic driving mode in which operation control and acceleration / deceleration control are performed automatically.

[0023] The mobility mechanism 30 is a mechanism for moving the mobile body 1 on a road. The mobility mechanism 30 is, for example, a group of wheels including steering wheels and drive wheels. Alternatively, the mobility mechanism 30 may be legs for multi-legged walking.

[0024] The drive unit 40 outputs force to the moving mechanism 30 to move the moving body 1. For example, the drive unit 40 includes a motor that drives the drive wheels, a battery that stores the power supplied to the motor, and a steering device that adjusts the steering angle of the steering wheels. The drive unit 40 may also be equipped with an internal combustion engine or a fuel cell as a means of outputting driving force or generating power. Furthermore, the drive unit 40 may also be equipped with a braking device that uses frictional force or air resistance.

[0025] The external notification device 50 is, for example, a lamp, display device, or speaker installed on the outer panel of the mobile body 1 to notify information to the outside of the mobile body 1. The external notification device 50 operates differently depending on whether the mobile body 1 is moving on a sidewalk or on a roadway. For example, the external notification device 50 is controlled to illuminate a lamp when the mobile body 1 is moving on a sidewalk, and not to illuminate a lamp when the mobile body 1 is moving on a roadway. The color of the light emitted by this lamp is preferably a color specified by law. The external notification device 50 may also be controlled to illuminate a green lamp when the mobile body 1 is moving on a sidewalk, and a blue lamp when the mobile body 1 is moving on a roadway. If the external notification device 50 is a display device, it displays a message such as "The mobile body 1 is traveling on a sidewalk" in text or graphics when the mobile body 1 is traveling on a sidewalk.

[0026] Figure 2 is a perspective view of the mobile body 1 from above. In the figure, FW is the steering wheel, RW is the drive wheel, SD is the steering mechanism, MT is the motor, and BT is the battery. The steering mechanism SD, motor MT, and battery BT are included in the drive system 40. AP is the accelerator pedal, BP is the brake pedal, WH is the steering wheel, SP is the speaker, and MC is the microphone. The mobile body 1 shown is a single-seater, and the occupant P is seated in the driver's seat DS wearing a seat belt SB. Arrow D1 is the direction of travel (velocity vector) of the mobile body 1. The external environment detection device 10 is located near the front end of the mobile body 1, the internal camera 16 is located in a position that can capture the occupant P's head from in front of the occupant P, and the mode switching switch 22 is located on the boss of the steering wheel WH. An external notification device 50, which serves as a display device, is also located near the front end of the mobile body 1.

[0027] Returning to Figure 1, the storage device 70 is a non-transient storage device such as an HDD (Hard Disk Drive), flash memory, or RAM (Random Access Memory). The storage device 70 stores map information 72 and other programs 74 executed by the control device 100. In the figure, the storage device 70 is shown outside the frame of the control device 100, but the storage device 70 may be included within the control device 100.

[0028] [Control device] The control device 100 includes, for example, an image acquisition unit 110, a feature point extraction unit 120, a key point extraction unit 130, a self-localization unit 140, and a control unit 150. For example, it is realized by a hardware processor such as a CPU (Central Processing Unit) executing a program 74 (software). Some or all of these components may be realized by hardware (including circuitry) such as an LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or GPU (Graphics Processing Unit), or by the cooperation of software and hardware. The program may be stored in the storage device 70 in advance, or it may be stored on a removable storage medium (non-transient storage medium) such as a DVD or CD-ROM and installed in the storage device 70 when the storage medium is mounted on a drive device. The combination of the image acquisition unit 110, the feature point extraction unit 120, the key point extraction unit 130, and the self-position estimation unit 140 is an example of an "information processing device" as defined in the claims.

[0029] The image acquisition unit 110 acquires image IMs in chronological order from the external camera, which is the external environment detection device 10, capturing the surrounding environment of the moving object 1. In particular, the image acquisition unit 110 acquires image IMs in chronological order from the external camera, which is the external environment detection device 10, capturing the area in front of the moving object 1 in the direction of travel. Figure 3 shows an example of an image IM acquired by the image acquisition unit 110. The left side of Figure 3 represents the image IM acquired by the image acquisition unit 110 at time t-1, and the right side of Figure 3 represents the image IM acquired by the image acquisition unit 110 at time t.

[0030] The feature point extraction unit 120 identifies stationary objects captured by the image acquisition unit 110 based on the time-series image IM acquired by the image acquisition unit 110 and the velocity of the moving object 1, and extracts one or more feature points FP from the stationary objects in the image IM using a predetermined method. Here, the predetermined method is, for example, an extraction method using a trained model that has been trained to output the edges of objects (e.g., buildings or road structures) captured in the image IM as a point cloud when an image IM is input. Alternatively, the predetermined method may be any feature point extraction method used in Visual SLAM (Simultaneous Localization and Mapping), a technology that grasps its own position in three dimensions from image data captured by a camera. For example, in Figure 3, the feature point extraction unit 120 identifies building B as a stationary object and extracts one or more feature points FP using the predetermined method.

[0031] Conventional Visual SLAM methods, however, treat dynamic objects such as cars and pedestrians as noise and estimate their own position using only feature points (FPs) extracted from stationary objects. However, for example, when a moving object 1 travels on a crowded path (in other words, a path where it is difficult to extract feature points FPs from stationary objects), it is not always possible to generate a highly accurate environment map based solely on stationary objects. Therefore, conventional techniques sometimes exhibited low robustness in self-localization.

[0032] Against this backdrop, the keypoint extraction unit 130 extracts keypoints from the skeleton information of the pedestrian captured in the image IM, which are of a different type from the feature points of the stationary object extracted by the feature point extraction unit 120. The self-position estimation unit 140 then estimates the self-position of the moving object 1 based on the extracted feature points and keypoints. More specifically, when the keypoint extraction unit 130 detects a pedestrian P from the image IM, it applies arbitrary skeleton recognition processing to the pedestrian P and recognizes, for example, the ankle of the pedestrian P (more specifically, the center position of the left and right ankles) as a keypoint KP. The keypoint KP is not limited to the ankle, but may be, for example, the elbow, knee, or pelvis. Furthermore, the keypoint extraction unit 130 may change the type of keypoint KP to be extracted for each scene in which the moving object 1 is traveling. For example, the key point extraction unit 130 may change the key point KP to be extracted from the ankle to the pelvis if the inclination or unevenness of the road surface on which the moving body 1 travels exceeds a threshold (i.e., if it is expected that the position change of the ankle will become larger).

[0033] If the pedestrian P from which the keypoint KP has been extracted remains stationary between frames (i.e., the position of the keypoint KP remains unchanged between frames, taking into account the velocity of the moving object 1), the self-position estimation unit 140 first identifies the position of the moving object 1 based on the extracted feature point FP, and then estimates the position of the moving object 1 by correcting the identified position of the moving object 1 using the keypoint KP.

[0034] Figure 4 illustrates the self-position correction performed by the self-position estimation unit 140 when the pedestrian P is stationary. In Figure 4, code EP(t-1) indicates the self-position estimated one frame earlier, code EP'(t) indicates the self-position identified based on the feature point FP at time t, and code EP(t) indicates the self-position finally estimated by correcting the identified EP'(t).

[0035] The self-position estimation unit 140 first measures the displacement d1 between the previous feature point FP(t-1) and the current feature point FP(t) between frames, and determines the self-position EP'(t) by shifting the previous self-position EP(t-1) by the measured displacement d1. Next, the self-position estimation unit 140 measures the displacement d2 between the previous key point KP(t-1) and the current key point KP(t) between frames, and estimates the final self-position EP(t) by shifting the self-position EP'(t) by, for example, a correction amount d2-d1. The correction amount at this time may be, for example, an intermediate value (d2-d1) / 2 between the displacement of the feature point and the displacement of the key point, and at least it should take into account the displacement of the key point.

[0036] Thus, according to this embodiment, in addition to feature points FP, a small number of key points KP are extracted from the skeleton information of dynamic objects such as pedestrians P and used for self-position estimation. Generally, the number of key points extracted from skeleton information is smaller than the number of feature points FP detected by conventional Visual SLAM, and there is a tendency for less error to occur in terms of time-series tracking. Therefore, by estimating self-position using both feature points FP and key points KP, the robustness of self-position estimation can be improved. In particular, for example, when a moving object 1 is traveling on a road crowded with people (in other words, a road where it is difficult to extract feature points FP from stationary objects), even if sufficient feature points FP cannot be obtained when estimating self-position, the self-position can be robustly estimated by utilizing key points KP.

[0037] The estimation method shown in Figure 4 in the situation shown in Figure 3 assumes that pedestrian P is stationary. However, the present invention can also be applied when pedestrian P is moving. Below, we will describe a method for estimating the self-position using feature points FP and key points KP when pedestrian P is moving.

[0038] Figure 5 shows another example of an image IM acquired by the image acquisition unit 110. The left part of Figure 5 represents the image IM acquired by the image acquisition unit 110 at time t-1, the center part of Figure 5 represents the image IM acquired by the image acquisition unit 110 at time t, and the right part of Figure 5 represents the image IM acquired by the image acquisition unit 110 at time t+1. Furthermore, in the right part of Figure 5, EKP(t+1) represents the position of keypoint KP(t+1) at time t+1, estimated based on the position of keypoint KP(t-1) at time t-1 and the position of keypoint KP(t) at time t.

[0039] More specifically, when the keypoint extraction unit 130 detects a pedestrian P from the image IM, it applies arbitrary skeleton recognition processing to the pedestrian P and recognizes, for example, the pelvis of the pedestrian P as a keypoint KP. If the pedestrian P from which the keypoint KP was extracted has moved between frames (in Figure 5, between frames at time t-1 and time t) (i.e., if the position of the keypoint KP has moved between frames, taking into account the velocity of the moving object 1), the self-position estimation unit 140 first predicts the position of the keypoint KP at time t+1 based on the displacement amount and time of the keypoint KP between frames.

[0040] Figure 6 is a diagram illustrating the self-position correction performed by the self-position estimation unit 140 when a pedestrian P is moving. The left side of Figure 6 shows the method for predicting the position of keypoint KP, and the right side of Figure 6 shows the method for correcting the self-position according to the predicted position of keypoint KP.

[0041] As shown in the left part of Figure 6, the self-position estimation unit 140 assumes, for example, that pedestrian P is moving in a straight line at a constant speed, and predicts the position of EKP(t+1) on a two-dimensional plane, where EKP(t) lies on the straight line connecting the position of keypoint KP(t-1) and the position of keypoint KP(t) at time t, and where the distance between keypoint KP(t-1) and keypoint KP(t) is the same as the distance between keypoint KP(t) and keypoint EKP(t+1). Next, the self-position estimation unit 140 calculates the discrepancy d between the predicted position EKP(t+1) and the actual position KP(t+1) at time t+1, and determines whether the calculated discrepancy d is within a threshold.

[0042] If the deviation d is determined to be within the threshold, the self-position estimation unit 140 determines that the pedestrian P is moving in a straight line at a constant speed (i.e., moving linearly), and therefore the moving object 1, which measured the linear movement by the pedestrian P, is also moving linearly. For this reason, if the self-position estimation unit 140 does not have a linear relationship with the self-position EP(t-1) from two steps prior, the self-position EP(t) from the previous step, and the self-position EP'(t+1) identified based on the feature point FP(t+1) from the current step, the self-position estimation unit 140 corrects the self-position EP'(t+1) to have a linear relationship with the self-position EP(t-1) from two steps prior and the self-position EP(t) from the previous step, and estimates the self-position EP(t+1) from the current step. As with the case in Figure 4, the correction amount in this case only needs to take into account the linear relationship described above. For example, the self-position EP'(t+1) identified based on the feature point FP(t+1) in this case may be corrected to the midpoint of the self-position with which the linear relationship holds.

[0043] Thus, even when pedestrian P is moving, the self-position estimation unit 140 determines whether pedestrian P is moving linearly based on the displacement of pedestrian P's key point KP. If it is determined that pedestrian P is moving linearly, the self-position of the moving object 1 is corrected so that it also changes linearly over time. In other words, even if sufficient feature points FP cannot be obtained when estimating the self-position, the self-position can be robustly estimated by utilizing the key point KP.

[0044] The self-localization unit 140 stores information including the extracted feature points FP and key points KP, and the estimated self-localization, as map information 72 in the storage device 70. Figure 7 shows an example of map information 72. Map information 72 includes, for example, feature points FP(t) and key points KP(t) extracted from the image IM at each time point t, as well as the estimated self-localization EP(t), as point cloud data. In this case, the point cloud data may be the extracted raw 3D data, or it may be 2D data obtained by projecting the 3D data onto a bird's-eye view coordinate system.

[0045] The control unit 150 controls the drive unit 40 or provides driving assistance to the occupant, while referring to the map information 72, according to the set driving mode. For example, when the driving mode is set to automatic driving mode, the control unit 150 detects free space from the map information 72 and controls the drive unit 40 so that the mobile body 1 travels within the free space starting from its estimated self-position. For example, in the case of the three-dimensional map MP1 in Figure 8, the control unit 150 may detect the space between the left and right planes as free space. Also, for example, in the case of the two-dimensional map MP2 in Figure 8, the control unit 150 may detect the space between the left and right straight lines as free space. When the driving mode is set to manual driving mode, the control unit 150 displays the map information 72 on the external notification device 50, which is a display device, and the occupant can drive while referring to the map information 72 displayed on the display device.

[0046] Next, with reference to Figure 8, the processing flow performed by the control device 100 will be described. Figure 8 is a flowchart showing an example of the processing flow performed by the control device 100.

[0047] First, the image acquisition unit 110 acquires image IMs in chronological order from the external camera, which is the external environment detection device 10, capturing the surrounding environment of the moving object 1 (step S100). Next, the feature point extraction unit 120 extracts one or more feature points from the acquired image IMs using a predetermined method (step S102). Next, the key point extraction unit 130 extracts key points of pedestrians from the acquired image IMs (step S104). Next, the self-position estimation unit 140 determines the self-position based on the extracted feature points (step S106).

[0048] Next, the self-position estimation unit 140 determines whether the pedestrian is stationary or moving linearly based on the time-series image IM (step S108). If it is determined that the pedestrian is stationary or moving linearly, the self-position estimation unit 140 corrects its own position based on the extracted keypoints (step S110). Next, the control unit 150 either drives the mobile body 1 or provides driving assistance for the mobile body 1 based on the corrected self-position (step S112). On the other hand, if it is determined that the pedestrian is neither stationary nor moving linearly, the self-position estimation unit 140 drives the mobile body 1 or provides driving assistance for the mobile body 1 based on the self-position identified in step S106. This completes the processing of this flowchart.

[0049] As described above, according to this embodiment, one or more images of the surrounding environment of a moving object are acquired in a time series, feature points of stationary objects captured in the images are extracted, key points, which are of a different type of feature than the feature points, are extracted from the skeleton information of pedestrians captured in the images, and the position of the moving object in the surrounding environment is estimated based on the extracted feature points and key points. This improves the robustness of self-localization.

[0050] The embodiments described above can be expressed as follows. A memory device that stores the program, Equipped with a hardware processor, The hardware processor executes the program, Obtain one or more images in chronological order that capture the surrounding environment of a moving object. Extract the feature points of the stationary objects captured in the aforementioned image, From the skeleton information of the pedestrian captured in the aforementioned image, key points, which are of a different type of feature than the aforementioned feature points, are extracted. Based on the extracted feature points and key points, the position of the moving object in the surrounding environment is estimated. A vehicle control device configured in such a way.

[0051] Although embodiments for carrying out the present invention have been described above using examples, the present invention is not limited in any way to these embodiments, and various modifications and substitutions can be made without departing from the spirit of the present invention. [Explanation of Symbols]

[0052] 10. External detection devices 12 Mobile Sensors 14 Operators 16 Internal Cameras 18 Positioning device 20 Accelerometer 22 Mode selector switch 30 Moving mechanism 40 Drive unit 50 External notification device 70 Storage device 100 Control device 110 Image acquisition unit 120 Feature point extraction unit 130 Keypoint Extraction Section 140 Self-position estimation part 150 Control Unit

Claims

1. An image acquisition unit that acquires one or more images in chronological order of the surrounding conditions of a moving object, A feature point extraction unit that extracts feature points of stationary objects captured in the aforementioned image, A keypoint extraction unit extracts keypoints, which are of a different type of feature from the aforementioned feature points, from the skeleton information of the pedestrian captured in the aforementioned image. The system includes a self-position estimation unit that estimates the position of the moving object in the surrounding environment based on the extracted feature points and key points, Information processing device.

2. The key point extraction unit extracts the location representing the pedestrian's ankle as the key point. The information processing apparatus according to claim 1.

3. The self-position estimation unit estimates the position of the moving object based on the extracted feature points and corrects the estimated position of the moving object using the key points. The information processing apparatus according to claim 1.

4. The self-position estimation unit, when the pedestrian is stationary and the mobile body equipped with a camera that takes one or more images in a time series is moving, corrects the position of the mobile body using the displacement amount of the keypoint between the one or more images of the pedestrian taken in a time series. The information processing apparatus according to claim 3.

5. The self-position estimation unit, when the pedestrian and the mobile body equipped with a camera that takes one or more images in a time series are moving, predicts the future position of the pedestrian from the amount of displacement of the key points between the one or more images of the pedestrian taken in a time series, and corrects the position of the mobile body using the predicted future position. The information processing apparatus according to claim 3.

6. Computers Obtain one or more images in chronological order that capture the surrounding environment of a moving object. Extract the feature points of the stationary objects captured in the aforementioned image, From the skeleton information of the pedestrian captured in the aforementioned image, key points, which are of a different type of feature than the aforementioned feature points, are extracted. Based on the extracted feature points and key points, the position of the moving object in the surrounding environment is estimated. Information processing methods.

7. On the computer, The system captures one or more images in chronological order of the surrounding environment of a moving object. Extract the feature points of the stationary object captured in the aforementioned image. From the skeleton information of the pedestrian captured in the aforementioned image, key points, which are of a different type of feature than the aforementioned feature points, are extracted. Based on the extracted feature points and key points, the position of the moving object in the surrounding environment is estimated. program.