Self-position estimation model learning method, self-position estimation model learning device, self-position estimation method, self-position estimation device, recording medium, and robot
By taking local and overlooking images in a dynamic environment, combining trajectory information and feature quantity calculations, learning one's own position estimation model is solved, and the problem of the existing technology being difficult to estimate one's own position in complex dynamic environments is achieved, and accurate position estimation in environments such as dense populations is achieved.
Patent Information
- Application Number
- CN202080076842.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-13
- Filing Date
- 2020-10-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-10-21
AI Technical Summary
In complex dynamic environments such as populations, existing feature point-based self-position estimation algorithms are difficult to stabilize the position, and existing processing methods cannot be effectively applied to dense dynamic environments.
By taking local images from the viewpoint of the object's own position estimation in a dynamic environment and shooting a top-view image from a top-view angle, combining trajectory information calculation, feature quantity calculation and distance calculation, the self-position estimation model is learned to output position information.
Even in dynamic environments where it is difficult to estimate one's own position, it can accurately estimate one's own position, improving the autonomous navigation ability in complex environments.
Smart Images

Figure CN114698388B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for learning a self-position estimation model, a device for learning a self-position estimation model, a program for learning a self-position estimation model, a self-position estimation method, a self-position estimation device, a self-position estimation program, and a robot. Background Art
[0002] In existing feature point-based self-position estimation (Simultaneously Localization and Mapping: SLAM) algorithms (for example, refer to Non-Patent Document 1 "ORB-SLAM2: an Open-Source SLAM System for Monocular, Stereo and RGB-D Cameras https: / / 128.84.21.199 / pdf / 1610.06475.pdf"), movement information such as rotation or translation is calculated by observing static feature points in a three-dimensional space from multiple viewpoints.
[0003] However, in an environment such as a crowd scene that includes many moving objects and occlusions, geometric constraints fail, stable position restoration cannot be performed, and the self-position on the map is frequently lost (for example, refer to Non-Patent Document 2 "Getting Robots Unfrozen and Unlost in Dense Pedestrian Crowds https: / / arxiv.org / pdf / 1810.00352.pdf").
[0004] As other methods for dealing with moving objects, there are methods for directly modeling the movement of moving objects and robust estimation methods that use an error function to reduce the influence of parts equivalent to moving objects, but neither can be applied to a complex and dense dynamic environment such as a crowd.
[0005] In addition, in feature point-based SLAM represented by the technology described in Non-Patent Document 1, by creating a visual vocabulary based on the feature points of a scene and storing it in a database, the same scene can be recognized.
[0006] In addition, Non-Patent Document 3 ([N.N+, ECCV’16] Localizing and Orienting Street Views Using Overhead Imagery https: / / lugiavn.github.io / gatech / crossview_ecev2016 / nam_ecev 2016.pdf) and Non-Patent Document 4 ([S. Workman+, ICCV’15] Wide-Area Image Geolocalization with Aerial Reference Imagery https: / / www.cv-foundation.org / openaccess / content_iccv_2015 / papers / Workman_Wide-Area_Image_Geolocalization_ICCV_2015_paper.pdf) disclose techniques that can extract features from overhead images and local images respectively, and retrieve which block of the overhead image the local image corresponds to respectively. Summary of the Invention
[0007] Problems to be Solved by the Invention
[0008] However, in the techniques described in Non-Patent Documents 3 and 4 above, since only the image similarity between static scenes is used as a clue for matching, the matching accuracy is low and a large number of candidate regions appear.
[0009] The technology of the present disclosure is completed in view of the above points, and its purpose is to provide a method for learning a self-position estimation model, a self-position estimation model learning device, a self-position estimation model learning program, a self-position estimation method, a self-position estimation device, a self-position estimation program, and a robot that can estimate the self-position of a self-position estimation object even in a dynamic environment where it has been difficult to estimate the self-position of the self-position estimation object in the past.
[0010] Means for Solving the Problems
[0011] A first aspect of the present disclosure is a method for learning an own-position estimation model, which is executed by a computer and includes the following steps: an acquisition step of acquiring, in time series, a local image and an aerial image synchronized with the local image, where the local image is a local image captured from the viewpoint of an own-position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the own-position estimation object; and a learning step of learning an own-position estimation model, where the own-position estimation model takes the local image and the aerial image acquired in time series as inputs and outputs the position of the own-position estimation object.
[0012] In the above first aspect, it may also be that the learning step includes: a trajectory information calculation step of calculating first trajectory information based on the local image and second trajectory information based on the aerial image; a feature amount calculation step of calculating a first feature amount based on the first trajectory information and a second feature amount based on the second trajectory information; a distance calculation step of calculating the distance between the first feature amount and the second feature amount; an estimation step of estimating the position of the own-position estimation object based on the distance; and an update step of updating the parameters of the own-position estimation model in such a manner that the higher the similarity between the first feature amount and the second feature amount, the smaller the distance.
[0013] In the above first aspect, it may also be that, in the feature amount calculation step, the second feature amount is calculated based on the second trajectory information in a plurality of partial regions selected from a region near the position of the own-position estimation object estimated last time, in the distance calculation step, the distance is calculated for each of the plurality of partial regions, and in the estimation step, the predetermined position of the partial region with the smallest distance among the distances calculated for each of the plurality of partial regions is estimated as the position of the own-position estimation object.
[0014] A second aspect of the present disclosure is an own-position estimation model learning device, which includes: an acquisition unit that acquires, in time series, a local image and an aerial image synchronized with the local image, where the local image is a local image captured from the viewpoint of an own-position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the own-position estimation object; and a learning unit that learns an own-position estimation model, where the own-position estimation model takes the local image and the aerial image acquired in time series as inputs and outputs the position of the own-position estimation object.
[0015] A third aspect of the present disclosure is a self-position estimation model learning program for causing a computer to execute a process including the following steps: an acquisition step of acquiring local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of a self-position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the self-position estimation object; and a learning step of learning a self-position estimation model that takes as input the local images and the aerial image acquired in time series and outputs the position of the self-position estimation object.
[0016] A fourth aspect of the present disclosure is a self-position estimation method for causing a computer to execute a process including the following steps: an acquisition step of acquiring local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of a self-position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the self-position estimation object; and an estimation step of estimating the self-position of the self-position estimation object based on the local images and the aerial image acquired in time series and a self-position estimation model learned by the self-position estimation model learning method according to the first aspect described above.
[0017] A fifth aspect of the present disclosure is a self-position estimation device including: an acquisition unit that acquires local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of a self-position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the self-position estimation object; and an estimation unit that estimates the self-position of the self-position estimation object based on the local images and the aerial image acquired in time series and a self-position estimation model learned by the self-position estimation model learning device according to the second aspect described above.
[0018] A sixth aspect of the present disclosure is a self-position estimation program for causing a computer to execute a process including the following steps: an acquisition step of acquiring local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of a self-position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the self-position estimation object; and an estimation step of estimating the self-position of the self-position estimation object based on the local images and the aerial image acquired in time series and a self-position estimation model learned by the self-position estimation model learning method according to the first aspect described above.
[0019] A seventh aspect of the present disclosure is a robot, comprising: an acquisition unit that acquires local images and aerial images synchronized with the local images in time series, where the local images are local images captured from the viewpoint of the robot in a dynamic environment, and the aerial images are aerial images captured from an aerial view of the position of the robot; an estimation unit that estimates the self-position of the robot based on the local images and the aerial images acquired in time series and a self-position estimation model learned by the self-position estimation model learning device described in the above second aspect; an autonomous driving unit that causes the robot to drive autonomously; and a control unit that controls the autonomous driving unit to move the robot towards a destination based on the position estimated by the estimation unit.
[0020] Advantages of the Invention
[0021] According to the technology of the present disclosure, it is possible to estimate the self-position of an object to be estimated of the self-position even in a dynamic environment where it has been difficult to estimate the self-position in the past. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a diagram showing a schematic configuration of a self-position estimation model learning system.
[0023] Figure 2 It is a block diagram showing the hardware configuration of a self-position estimation model learning device.
[0024] Figure 3 It is a block diagram showing the functional configuration of a self-position estimation model learning device.
[0025] Figure 4 It is a diagram showing a situation where a robot moves towards a destination in a crowd.
[0026] Figure 5 It is a block diagram showing the functional configuration of a learning unit of a self-position estimation model learning device.
[0027] Figure 6 It is a diagram for explaining a partial area.
[0028] Figure 7 It is a flowchart showing the process of self-position estimation model learning processing of a self-position estimation model learning device.
[0029] Figure 8 It is a block diagram showing the functional configuration of a self-position estimation device.
[0030] Figure 9 It is a block diagram showing the hardware configuration of a self-position estimation device.
[0031] Figure 10It is a flowchart showing the process of robot control processing of the own position estimation device. Detailed implementation mode
[0032] Hereinafter, an example of an implementation mode of the technology of the present disclosure will be described with reference to the drawings. In addition, in each drawing, the same or equivalent components and parts are labeled with the same reference numerals. In addition, the dimensional ratios of the drawings are sometimes exaggerated for ease of explanation and are sometimes different from the actual ratios.
[0033] Figure 1 It is a diagram showing the schematic structure of the own position estimation model learning system 1.
[0034] As Figure 1 shown, the own position estimation model learning system 1 includes an own position estimation model learning device 10 and a simulator 20. The simulator 20 will be described later.
[0035] Next, the own position estimation model learning device 10 will be described.
[0036] Figure 2 It is a block diagram showing the hardware structure of the own position estimation model learning device 10.
[0037] As Figure 2 shown, the own position estimation model learning device 10 has a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), a storage device, an input unit, a monitor, a CD drive device, and a communication interface. Each structure is connected via a bus 19 so as to be able to communicate with each other.
[0038] In the present embodiment, an own position estimation model learning program is stored in the storage device 14. The CPU 11 is a central arithmetic processing unit that executes various programs or controls each structure. That is, the CPU 11 reads a program from the storage device 14 and executes the program using the RAM 13 as a work area. The CPU 11 controls each of the above structures and performs various arithmetic processes according to the program recorded in the storage device 14.
[0039] The ROM 12 stores various programs and various data. The RAM 13 temporarily stores programs or data as a work area. The storage device 14 is composed of an HDD (Hard Disk Drive) or an SSD (Solid State Drive) and stores various programs including an operating system and various data.
[0040] The input unit 15 includes pointing devices such as a keyboard 151 and a mouse 152, and is used for various inputs. The monitor 16 is, for example, a liquid crystal display, which displays various information. The monitor 16 can also adopt a touch panel method and function as the input unit 15. The optical disc drive device 17 reads data stored in various recording media (such as CD-ROM or Blu-ray Disc), writes data to the recording media, and so on.
[0041] The communication interface 18 is an interface for communicating with other devices such as the simulator 20, and uses standards such as Ethernet (registered trademark), FDDI, or Wi-Fi (registered trademark), for example.
[0042] Next, the functional structure of the own position estimation model learning device 10 will be described.
[0043] Figure 3 It is a block diagram showing an example of the functional structure of the own position estimation model learning device 10.
[0044] As Figure 3 shown, the own position estimation model learning device 10 has a acquisition unit 30 and a learning unit 32 as its functional structure. Each functional structure is realized by the CPU 11 reading the own position estimation program stored in the storage device 14, expanding it in the RAM 13, and executing it.
[0045] The acquisition unit 30 acquires destination information, a local image, and an aerial image from the simulator 20. The simulator 20 outputs, for example, as Figure 4 shown, the local image and the aerial image synchronized with the local image when the autonomous driving type robot RB moves toward the destination p indicated by the destination information in time series. g
[0046] In addition, in the present embodiment, as Figure 4 shown, the robot RB moves toward the destination p in a dynamic environment including moving objects such as people HB existing around. g In the present embodiment, the case where the moving object is a person HB, that is, the case where the dynamic environment is a crowd, will be described, but it is not limited thereto. For example, as other examples of the dynamic environment, environments where cars, autonomous driving type robots, drones, airplanes, ships, etc. exist can be cited.
[0047] Here, the local image is Figure 4 In a dynamic environment such as shown in FIG. 1 , an image is captured from the viewpoint of the robot RB, which is the object of self-position estimation. In addition, the following description is made of a case where the local image is captured by an optical camera, but the invention is not limited thereto. That is, as long as motion information indicating how an object existing within the field of view of the robot RB moves can be obtained, for example, motion information obtained by an event-based camera can be used, or motion information obtained by performing image processing on the local image using a known method such as optical flow can be used.
[0048] In addition, the bird's-eye view image is an image captured from a position overlooking the robot RB. Specifically, the bird's-eye view image is, for example, an image obtained by capturing a range including the robot RB from above the robot RB, and is an image obtained by capturing a range larger than the range represented by the partial image. In addition, the bird's-eye view image may be a RAW (Raw image format) image, or a dynamic image such as an image processed by an image.
[0049] The learning unit 32 receives as input the partial image and the bird's-eye view image acquired in time series by the acquisition unit 30 , and learns a self-position estimation model that outputs the position of the robot RB.
[0050] Next, the learning unit 32 will be described in detail.
[0051] like Figure 5 As shown, the learning unit 32 includes a first trajectory information calculation unit 33 - 1 , a second trajectory information calculation unit 33 - 2 , a first feature vector calculation unit 34 - 1 , a second feature vector calculation unit 34 - 2 , a distance calculation unit 35 , and a self-position estimation unit 36 .
[0052] The first trajectory information calculation unit 33 - 1 calculates the trajectory information of N (N is a plurality of) temporally continuous partial images I1 (= {I1 1 、I1 2 ,…,I1 N}), calculate the first trajectory information t of person HB 1 In the first trajectory information t 1 In the calculation of , for example, known methods such as the above-mentioned optical flow and MOT (Multi Object Tracking) can be used, but it is not limited to these.
[0053] The second trajectory information calculation unit 33 - 2 calculates the trajectory information based on the N bird's-eye view images I2 (= {I2 1 、I2 2 ,…,I2 N}), calculate the second trajectory information t of human HB 2 In the calculation of the second trajectory information t 2 , similar to the calculation of the first trajectory information, known methods such as optical flow can be used, but it is not limited to this.
[0054] The first feature vector calculation unit 34-1 calculates the K 1 -dimensional first feature vector φ 1 (t 1 ). Specifically, the first feature vector calculation unit 34-1, for example, calculates the K 1 -dimensional first feature vector φ 1 by inputting the first trajectory information t 1 into the first convolutional neural network (CNN: Convolutional neural network). 1 (t 1 ). In addition, the first feature vector φ 1 (t 1 ) is an example of the first feature quantity. It is not limited to the feature vector, and other feature quantities can also be calculated.
[0055] The second feature vector calculation unit 34-2 calculates the K 2 -dimensional second feature vector φ 2 (t 2 ). Specifically, the second feature vector calculation unit 34-2, similar to the first feature vector calculation unit 34-1, for example, calculates the K 2 -dimensional second feature vector φ 2 by inputting the second trajectory information t 2 into a second convolutional neural network different from the first convolutional neural network used in the first feature vector calculation unit 34-1. 2 (t 2 ). In addition, the second feature vector φ 2 (t 2 ) is an example of the second feature quantity. It is not limited to the feature vector, and other feature quantities can also be calculated.
[0056] Here, as Figure 6 shown, the second trajectory information t 2 input into the second convolutional neural network is not the trajectory information of the entire overhead image I2, but M (M is a plurality) partial regions W t randomly selected from the local region L near the position p 1 ~W M of the previous detected robot RB. Among them, the second trajectory information t 21 ~t 2M . Thus, for each partial region W 1 ~W M Calculate the second feature vector φ 2 (t 21 )~φ 2 (t 2M ). Hereinafter, without distinguishing the second trajectory information t 21 ~t 2M , it is sometimes simply referred to as the second trajectory information t 2 . Similarly, without distinguishing the second feature vector φ 2 (t 21 )~φ 2 (t 2M ), it is sometimes simply referred to as the second feature vector φ 2 (t 2 ).
[0057] In addition, the local area L is set to include the range within which the robot RB can move from the position p of the robot RB detected last time t-1 . In addition, a partial area W 1 ~W M is randomly selected from the local area L 1 ~W M . In addition, the number of the partial areas W 1 ~W M and the size of the partial areas W 1 ~W M affect the processing speed and the estimation accuracy of the self-position. Therefore, the number of the partial areas W 1 ~W M and the size of the partial areas W 1 ~W M are set to arbitrary values according to the desired processing speed and the estimation accuracy of the self-position. Hereinafter, without particularly distinguishing the partial areas W 1 ~W M , it is sometimes simply referred to as the partial area W. In addition, in the present embodiment, the case of randomly selecting the partial areas W 1 ~W M from the local area L is described, but it is not limited thereto. For example, the local area L can be equally divided to set the partial areas W
[0058] The distance calculation unit 35 calculates, for example, using a neural network, the distance g(φ 1 (t 1 ) representing the similarity between each of the first feature vector φ 1 ~W M and the second feature vectors φ 2 (t 21 )~φ 2 (t 2M ) of the partial areas W 1(t 1 )、 φ 2 (t 21 )) ~ g(φ 1 (t 1 )、 φ 2 (t 2M ))。 Moreover, the neural network is learned such that: the higher the similarity between the first feature vector φ 1 (t 1 ) and the second feature vector φ 2 (t 2 ), the smaller the distance g(φ 1 (t 1 ), φ 2 (t 2 ))。
[0059] In addition, the first feature vector calculation unit 34-1, the second feature vector calculation unit 34-2, and the distance calculation unit 35 can use, for example, a well-known learning model such as a Siamese Network using contrastive loss or triplet loss. In this case, the parameters of the neural network used in the first feature vector calculation unit 34-1, the second feature vector calculation unit 34-2, and the distance calculation unit 35 are learned in such a way that the higher the similarity between the first feature vector φ 1 (t 1 ) and the second feature vector φ 2 (t 2 ), the smaller the distance g(φ 1 (t 1 ), φ 2 (t 2 ))。 Additionally, as a method for calculating the distance, not limited to the case of using a neural network, Mahalanobis distance learning, which is an example of a metric learning method, can also be used.
[0060] The self-position estimation unit 36 estimates the position of a predetermined part region W of the second feature vector φ 1 (t 1 ) corresponding to the smallest distance among the distances g(φ 2 (t 21 )) ~ g(φ 1 (t 1 ), φ 2 (t 2M )) as the self-position p 2 (t 2 ), for example, the central position. t .
[0061] Thus, the own - position estimation model learning device 10 can be functionally said to be a device that learns an own - position estimation model for estimating and outputting the own position based on local images and aerial images.
[0062] Next, the operation of the own - position estimation model learning device 10 will be described.
[0063] Figure 7 FIG. is a flowchart showing the process of the own - position estimation model learning process of the own - position estimation model learning device 10. The CPU 11 reads the own - position estimation model learning program from the storage device 14, expands and executes it in the RAM 13, thereby performing the own - position estimation model learning process.
[0064] In step S100, the CPU 11, as the acquisition unit 30, acquires the position information of the destination p g from the simulator 20.
[0065] In step S102, the CPU 11, as the acquisition unit 30, acquires N time - series local images I1(={I1 1 , I1 2 , …, I1 N}) from the simulator 20.
[0066] In step S104, the CPU 11, as the acquisition unit 30, acquires N time - series aerial images I2(={I2 1 , I2 2 , …, I2 N}) synchronized with the local image I1 from the simulator 20.
[0067] In step S106, the CPU 11, as the first trajectory information calculation unit 33 - 1, calculates the first trajectory information t 1 based on the local image I1.
[0068] In step S108, the CPU 11, as the second trajectory information calculation unit 33 - 2, calculates the second trajectory information t 2 based on the aerial image I2.
[0069] In step S110, the CPU 11, as the first feature vector calculation unit 34 - 1, calculates the first feature vector φ 1 based on the first trajectory information t 1 (t 1 ).
[0070] In step S112, the CPU 11, as the second feature vector calculation unit 34 - 2, based on a partial region W 2 in the second trajectory information t 1 ~W MThe second trajectory information t 21 ~t 2M , calculate the second feature vector φ 2 (t 21 )~φ 2 (t 2M ).
[0071] In step S114, the CPU 11 acts as the distance calculation unit 35 and calculates the distance g(φ 1 (t 1 ) and each second feature vector φ 2 (t 21 )~φ 2 (t 2M )) that represents the similarity. That is, the distance is calculated for each partial region W. 1 (t 1 ), φ 2 (t 21 ))~g(φ 1 (t 1 ), φ 2 (t 2M ))).
[0072] In step S116, the CPU 11 acts as the self-position estimation unit 36 and estimates the representative position, such as the center position, of the partial region W of the second feature vector φ 1 (t 1 ) corresponding to the minimum distance among the distances g(φ 2 (t 21 ))~g(φ 1 (t 1 ), φ 2 (t 2M )) as its own position p 2 (t 2 ), and outputs it to the simulator 20. t
[0073] In step S118, the CPU 11 acts as the learning unit 32 and updates the parameters of the self-position estimation model. That is, if a siamese network is used as the learning model included in the self-position estimation model, the parameters of the siamese network are updated.
[0074] In step S120, the CPU 11 acts as the self-position estimation unit 36 and determines whether the robot RB has reached the destination p g . That is, it determines whether the position p t of the robot RB estimated in step S116 is the same as the destination p g It is consistent. Then, when it is determined that the robot RB has reached the destination p g the process proceeds to step S122. On the other hand, when it is determined that the robot RB has not reached the destination p g the process proceeds to step S102, and the processing of steps S102 to S120 is repeated until it is determined that the robot RB has reached the destination p g That is, the learning model is learned. In addition, the processing of steps S102 and S104 is an example of an acquisition step. In addition, the processing of steps S108 to S118 is an example of a learning step.
[0075] In step S122, the CPU 11, as the own-position estimation unit 36, determines whether the end condition for ending learning is satisfied. In the present embodiment, the end condition is, for example, the case where a predetermined number (for example, 100) of rounds end when the robot RB reaches the destination p g from the starting point. The CPU 11 ends this routine when it determines that the end condition is satisfied. On the other hand, when the end condition is not satisfied, the process proceeds to step S100, the destination p g is changed, and the processing of steps S100 to S122 is repeated until the end condition is satisfied.
[0076] In this way, in the present embodiment, local images captured from the viewpoint of the robot RB and aerial images synchronized with the local images captured from a position overlooking the robot RB are acquired in time series, and an own-position estimation model that takes the local images and aerial images acquired in time series as inputs and outputs the position of the robot RB is learned. Thus, even in a dynamic environment where it has been difficult to estimate the own position of the robot RB in the past, the position of the robot RB can be estimated.
[0077] In addition, there may be a case where the minimum distance calculated in step S116 above is too large, that is, a case where the own position cannot be estimated. Therefore, it may be that when the minimum distance calculated in step S116 is equal to or greater than a predetermined threshold, it is determined that the own position cannot be estimated, and a partial area W t-1 is reselected from the local area L near the position p 1 of the robot RB detected last time M to W
[0078] In addition, as another example of the case where the estimation of its own position cannot be performed, there is a case where trajectory information cannot be calculated. For example, it is a case where there are no people HB around the robot RB at all and it becomes a completely static environment. In such a case, the estimation of its own position can also be re-performed by executing the processes of steps S112 to S116 again.
[0079] Next, a robot RB that estimates its own position using the self-position estimation model learned by the self-position estimation model learning device 10 will be described.
[0080] In Figure 8 a schematic structure of the robot RB is shown. As Figure 8 shown, the robot RB includes a self-position estimation device 40, a camera 42, a robot information acquisition unit 44, a notification unit 46, and an autonomous driving unit 48. The self-position estimation device 40 includes an acquisition unit 50 and a control unit 52.
[0081] The camera 42 takes pictures of the surroundings of the robot RB at a predetermined interval during the period from the starting point to the destination p g and outputs the captured partial images to the acquisition unit 50 of the self-position estimation device 40.
[0082] The acquisition unit 50 requests and acquires an aerial image taken from a position overlooking the robot RB from an external device (not shown) via wireless communication.
[0083] The control unit 52 has the function of the self-position estimation model learned by the self-position estimation model learning device 10. That is, the control unit 52 estimates the position of the robot RB based on the time-series synchronized partial images and aerial images acquired by the acquisition unit 50.
[0084] The robot information acquisition unit 44 acquires the speed of the robot RB as robot information. The speed of the robot RB is acquired using a speed sensor, for example. The robot information acquisition unit 44 outputs the acquired speed of the robot RB to the acquisition unit 50.
[0085] The acquisition unit 50 acquires the state of the person HB based on the partial images captured by the camera 42. Specifically, the captured images are analyzed using a known method, and the position and speed of the person HB existing around the robot RB are calculated.
[0086] The control unit 52 has the function of a learned robot control model for controlling the robot RB to autonomously travel to the destination p g
[0087] The robot control model is, for example, a model that takes as input robot information related to the state of the robot RB, environmental information related to the surrounding environment of the robot RB, and destination information related to the destination that the robot RB should reach, and selects and outputs an action corresponding to the state of the robot RB. For example, a model learned through reinforcement learning is used. Here, the robot information includes information on the position and speed of the robot RB. In addition, the environmental information includes information related to the dynamic environment. Specifically, for example, it includes information on the position and speed of the person HB existing around the robot RB.
[0088] The control unit 52 takes as input the destination information, the position and speed of the robot RB, and the state information of the person HB, selects an action corresponding to the state of the robot RB, and controls at least one of the notification unit 46 and the autonomous driving unit 48 based on the selected action.
[0089] The notification unit 46 has a function of notifying the presence of the robot RB to the surrounding person HB by outputting a sound or a warning sound.
[0090] The autonomous driving unit 48 has a function of autonomously driving the robot RB, such as tires and a motor for driving the tires.
[0091] In the case where the selected action is an action of moving the robot RB in a specified direction and speed, the control unit 52 controls the autonomous driving unit 48 to move the robot RB in the specified direction and speed.
[0092] In addition, in the case where the selected action is an intervention action, the control unit 52 controls the notification unit 46 to output a message such as "Please make way" or emit a warning sound.
[0093] Next, the hardware structure of the own position estimation device 40 will be described.
[0094] As Figure 9 shown, the own position estimation device 40 includes a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), a storage device 64, and a communication interface 65. Each structure is connected via a bus 66 so as to be able to communicate with each other.
[0095] In this embodiment, a self-position estimation program is stored in the storage device 64. The CPU 61 is a central processing unit that executes various programs and controls each component. That is, the CPU 61 reads a program from the storage device 64 and executes the program using the RAM 63 as a work area. The CPU 61 controls each of the above components and performs various arithmetic processes according to the program recorded in the storage device 64.
[0096] The ROM 62 stores various programs and various data. The RAM 63 temporarily stores programs or data as a work area. The storage device 64 is composed of an HDD (Hard Disk Drive) or an SSD (Solid State Drive) and stores various programs including an operating system and various data.
[0097] The communication interface 65 is an interface for communicating with other devices and uses standards such as Ethernet (registered trademark), FDDI, or Wi-Fi (registered trademark).
[0098] Next, the operation of the self-position estimation device 40 will be described.
[0099] Figure 10 is a flowchart showing the process of the self-position estimation process of the self-position estimation device 40. The CPU 61 reads the self-position estimation program from the storage device 64, expands and executes it in the RAM 63, thereby performing the self-position estimation process.
[0100] In step S200, the CPU 61, as the acquisition unit 50, acquires the position information of the destination p from an external device (not shown) through wireless communication. g of.
[0101] In step S202, the CPU 61, as the acquisition unit 50, acquires N time-series partial images I1 (= {I1 1 、I1 2 、…、I1 N}) from the camera 42.
[0102] In step S204, the CPU 61, as the acquisition unit 50, requests and acquires N time-series aerial images I2 (= {I2 1 、I2 2 、…、I2 N}) synchronized with the partial image I1 from an external device (not shown). At this time, the position p of the robot RB estimated by the previous execution of this routine t-1 is sent to the external device, and an aerial image of the periphery including the position p of the robot RB estimated last time is acquired from the external device. t-1 of.
[0103] In step S206, the CPU 61, acting as the control unit 52, calculates the first trajectory information t based on the partial image I1 1 .
[0104] In step S208, the CPU 61, acting as the control unit 52, calculates the second trajectory information t based on the bird's-eye view image I2 2 .
[0105] In step S210, the CPU 61, acting as the control unit 52, calculates the first feature vector φ based on the first trajectory information t 1 , calculating the first feature vector φ 1 (t 1 ).
[0106] In step S212, the CPU 61, acting as the control unit 52, based on the second trajectory information t in the partial area W 2 ~W 1 ~W M of the second trajectory information t 21 ~t 2M , calculates the second feature vector φ 2 (t 21 )~φ 2 (t 2M ).
[0107] In step S214, the CPU 61, acting as the control unit 52, calculates the distance g(φ representing the similarity between the first feature vector φ 1 (t 1 ) and each second feature vector φ 2 (t 21 )~φ 2 (t 2M ). That is, the distance is calculated for each partial area W 1 (t 1 ), φ 2 (t 21 ))~g(φ 1 (t 1 ), φ 2 (t 2M ))
[0108] In step S216, the CPU 61, acting as the control unit 52, sets the second feature vector φ corresponding to the minimum distance among the distances g(φ 1 (t 1 ), φ 2 (t 21 ))~g(φ 1 (t 1 ), φ 2 (t 2M )) calculated in step S2142 (t 2 ) is estimated as the representative position of the partial area W, for example, the center position is t .
[0109] In step S218, the CPU 61 as the acquisition unit 50 acquires the speed of the robot as the state of the robot RB from the robot information acquisition unit 44. In addition, the partial image acquired in step S202 is analyzed using a known method to calculate state information related to the state of the person HB existing around the robot RB, that is, the position and speed of the person HB.
[0110] In step S220, the CPU 61, as the control unit 52, selects an action corresponding to the state of the robot RB based on the destination information obtained in step S200, the position of the robot RB estimated in step S216, the speed of the robot RB obtained in step S218, and the state information of the person HB obtained in step S218, and controls at least one of the notification unit 46 and the autonomous driving unit 48 based on the selected action.
[0111] In step S222, the CPU 61 as the control unit 52 determines whether the robot RB has reached the destination p. g That is, determine the position p of the robot RB t Is it related to the destination? g Then, when it is determined that the robot RB has reached the destination p g On the other hand, if it is determined that the robot RB has not reached the destination p, this routine ends. g In the case of a failure, the process proceeds to step S202, and the processing of steps S202 to S222 is repeated until it is determined that the robot RB has reached the destination p. g In addition, the processing of steps S202 and S204 is an example of an acquisition step. In addition, the processing of steps S206 to S216 is an example of an estimation step.
[0112] In this way, the robot RB autonomously drives to the destination while estimating its own position based on the own position estimation model learned by the own position estimation model learning device 10 .
[0113] In addition, in the present embodiment, the case where the robot RB is equipped with the own position estimation device 40 has been described. However, the function of the own position estimation device 40 may be provided in an external server. In this case, the robot RB transmits the partial image captured by the camera 42 to the external server. The external server estimates the position of the robot RB based on the partial image transmitted from the robot RB and the bird's-eye view image obtained from the device providing the bird's-eye view image, and transmits it to the robot RB. Then, the robot RB selects an action based on the own position received from the external server and autonomously travels to the destination.
[0114] In addition, in the present embodiment, the case where the own position estimation target is the autonomous driving type robot RB has been described. However, it is not limited thereto, and the own position estimation target may also be a portable terminal device carried by a person. In this case, the function of the own position estimation device 40 is provided in the portable terminal device.
[0115] In addition, the robot control process executed by the CPU reading software (program) in the above-described embodiments may be executed by various processors other than the CPU. Examples of the processor in this case include a PLD (Programmable Logic Device) such as an FPGA (Field-Programmable Gate Array) that can change the circuit structure after manufacturing, and a dedicated circuit such as an ASIC (Application Specific Integrated Circuit) that has a circuit structure specifically designed to execute a specific process. In addition, the own position estimation model learning process and the own position estimation process may be executed by one of these various processors, or may be executed by a combination of two or more processors of the same type or different types (for example, a combination of multiple FPGAs, and a combination of a CPU and an FPGA). More specifically, the hardware structure of these various processors is a circuit combining circuit elements such as semiconductor elements.
[0116] In addition, in each of the above-described embodiments, a method of pre-storing the own-position estimation model learning program in the storage device 14 and pre-storing the own-position estimation program in the storage device 64 has been described, but the present invention is not limited thereto. The program may also be provided in a form recorded on a recording medium such as a CD-ROM (Compact Disc Read Only Memory), a DVD-ROM (Digital Versatile Disc Read Only Memory), or a USB (Universal Serial Bus) memory. In addition, the program may be configured to be downloaded from an external device via a network.
[0117] Regarding all the documents, patent applications, and technical standards described in this specification, they are incorporated herein by reference to the extent that the incorporation by reference of each document, patent application, and technical standard becomes the same as the case where it is specifically and separately described.
[0118] Reference Numeral Explanation
[0119] 1: Own-Position Estimation Model Learning System; 10: Own-Position Estimation Model Learning Device; 20: Simulator; 30: Acquisition Unit; 32: Learning Unit; 33: Trajectory Information Calculation Unit; 34: Feature Vector Calculation Unit; 35: Distance Calculation Unit; 36: Own-Position Estimation Unit; 40: Own-Position Estimation Device; 42: Camera; 44: Robot Information Acquisition Unit; 46: Notification Unit; 48: Autonomous Driving Unit; 50: Acquisition Unit; 52: Control Unit; HB: Human; RB: Robot.
Claims
1. A method for learning an own - position estimation model, which is executed by a computer and includes the following steps: An acquisition step of acquiring local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of an own - position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the own - position estimation object ; and A learning step of learning an own - position estimation model, where the own - position estimation model takes the local images and the aerial image acquired in time series as inputs and outputs the position of the own - position estimation object, The learning step includes: A distance calculation step of calculating the distance between a first feature quantity calculated based on the local image and a second feature quantity calculated based on the aerial image; An estimation step of estimating the position of the own - position estimation object based on the distance; and An update step of updating the parameters of the own - position estimation model in such a way that the higher the similarity between the first feature quantity and the second feature quantity, the smaller the distance.
2. The method for learning an own - position estimation model according to claim 1, wherein, The learning step further includes: A trajectory information calculation step of calculating first trajectory information based on the local image and second trajectory information based on the aerial image; A feature quantity calculation step of calculating the first feature quantity based on the first trajectory information and calculating the second feature quantity based on the second trajectory information.
3. The method for learning an own - position estimation model according to claim 2, wherein, In the feature quantity calculation step, the second feature quantity is calculated based on the second trajectory information in a plurality of partial regions selected from a region near the position of the own - position estimation object estimated last time, In the distance calculation step, the distance is calculated for each of the plurality of partial regions, In the estimation step, the predetermined position of the partial region with the smallest distance among the distances calculated for each of the plurality of partial regions is estimated as the position of the own - position estimation object.
4. An own - position estimation model learning device, which includes: An acquisition unit that acquires local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of an own - position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the own - position estimation object; and A learning unit that learns an own - position estimation model, where the own - position estimation model takes the local images and the aerial image acquired in time series as inputs and outputs the position of the own - position estimation object, The learning unit includes: A distance calculation unit that calculates the distance between a first feature quantity calculated based on the local image and a second feature quantity calculated based on the aerial image; and An estimation unit that estimates the position of the own - position estimation object based on the distance, The learning unit updates the parameters of the own position estimation model such that the higher the similarity between the first feature quantity and the second feature quantity, the smaller the distance.
5. A recording medium that records an own position estimation model learning program for causing a computer to execute a process including the following steps: An acquisition step of acquiring local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of an own position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the own position estimation object ; and A learning step of learning an own position estimation model that takes as input the local images and the aerial image acquired in time series and outputs the position of the own position estimation object, The learning step includes: A distance calculation step of calculating the distance between a first feature quantity calculated based on the local image and a second feature quantity calculated based on the aerial image; An estimation step of estimating the position of the own position estimation object based on the distance; and An update step of updating the parameters of the own position estimation model such that the higher the similarity between the first feature quantity and the second feature quantity, the smaller the distance.
6. An own position estimation method for causing a computer to execute a process including the following steps: An acquisition step of acquiring local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of an own position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the own position estimation object; and An estimation step of estimating the own position of the own position estimation object based on the local images and the aerial image acquired in time series and an own position estimation model learned by the own position estimation model learning method according to any one of claims 1 to 3.
7. An own position estimation device that includes: An acquisition unit that acquires local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of an own position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the own position estimation object; and An estimation unit that estimates the own position of the own position estimation object based on the local images and the aerial image acquired in time series and an own position estimation model learned by the own position estimation model learning device according to claim 4.
8. A recording medium that records an own position estimation program for causing a computer to execute a process including the following steps: An acquisition step of acquiring local images and an aerial image synchronized with the local images in time series, where the local images are local images captured from the viewpoint of an own position estimation object in a dynamic environment, and the aerial image is an aerial image captured from an aerial view of the position of the own position estimation object; and An estimation step of estimating the self-position of the self-position estimation object based on the local images and the bird's-eye view images obtained in time series, and the self-position estimation model learned by the self-position estimation model learning method according to any one of claims 1 to 3.
9. A robot, which comprises: an acquisition unit that acquires local images and bird's-eye view images synchronized with the local images in time series, where the local images are local images captured from the viewpoint of the robot in a dynamic environment, and the bird's-eye view images are bird's-eye view images captured from a position overlooking the robot; an estimation unit that estimates the self-position of the robot based on the local images and the bird's-eye view images obtained in time series, and the self-position estimation model learned by the self-position estimation model learning device according to claim 4; an autonomous driving unit that causes the robot to drive autonomously; and a control unit that controls the autonomous driving unit to move the robot to a destination based on the position estimated by the estimation unit.