Mobile object control system and mobile object control method
The movement control system corrects for camera distortions and predicts feature reliability to create accurate three-dimensional maps, enhancing robot navigation and task performance in challenging environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- HITACHI GE NUCLEAR ENERGY LTD
- Filing Date
- 2024-10-24
- Publication Date
- 2026-05-12
AI Technical Summary
Wide-angle cameras like fisheye or monocular stereo cameras capture images with distortion, leading to inaccurate three-dimensional maps when feature points from distorted regions are used.
A movement control system that includes image acquisition, two-dimensional and three-dimensional feature extraction, map update, trajectory generation, and control units to create an accurate three-dimensional map by correcting for camera distortions and predicting feature reliability.
Enables accurate three-dimensional mapping, allowing robots to navigate and perform tasks in unfamiliar environments with improved feature point matching accuracy.
Smart Images

Figure 2026076837000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a movement control system and a movement control method.
Background Art
[0002] Regarding the control of a moving body, for example, the technique described in Patent Document 1 is known. That is, Patent Document 1 describes "generating the map including the three-dimensional coordinates of the feature points using the three-dimensional coordinates of the moving body, the distance information, and the camera parameters of the imaging device".
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] For example, a wide-angle camera such as a fisheye camera or a monocular stereo camera can capture a wide range in one shot because of its wide field of view, but the captured image tends to be distorted in a predetermined manner due to lens characteristics and the like. If feature points extracted in a region with large distortion in the captured image are used for creating a three-dimensional map, there is a possibility of causing a decrease in the accuracy of the three-dimensional map. Such a situation is not considered in Patent Document 1, and there is room for improvement.
[0005] Therefore, an object of the present disclosure is to provide a movement control system or the like that creates an accurate three-dimensional map based on a captured image.
Means for Solving the Problems
[0006] To solve the aforementioned problems, the mobile body control system according to this disclosure comprises: an image acquisition unit that acquires images captured by a camera of a mobile body; a two-dimensional feature extraction unit that extracts two-dimensional features from the captured images; a three-dimensional feature extraction unit that extracts three-dimensional features indicating the position of a point in real space corresponding to the two-dimensional features based on the two-dimensional features and predetermined camera parameters indicating the characteristics of the camera; a map update unit that generates or updates a three-dimensional map by registering the three-dimensional features; a trajectory generation unit that generates a plurality of trajectory candidates when the mobile body moves; and a predetermined trajectory candidate among the plurality of trajectory candidates. The system comprises: a 2D feature prediction unit that calculates a predicted 2D feature indicating the position where the 3D feature will appear on the captured image when the moving body moves, based on the camera parameters; a prediction reliability calculation unit that calculates a predicted 2D feature reliability, which is the reliability when the predicted 2D feature is used for feature point matching, based on the camera parameters; a trajectory evaluation unit that calculates an evaluation value for each of the trajectory candidates based on the predicted 2D feature reliability; a trajectory determination unit that determines the trajectory of the moving body based on the evaluation value; and a control unit that controls the moving body to move along the trajectory. [Effects of the Invention]
[0007] According to this disclosure, it is possible to provide a mobile object control system, etc., that creates an accurate three-dimensional map based on captured images. [Brief explanation of the drawing]
[0008] [Figure 1] This is an explanatory diagram including a robot, which is the controlled object of the mobile control system according to the first embodiment. [Figure 2] This is an explanatory diagram relating to candidate trajectories for a robot in a mobile control system according to the first embodiment. [Figure 3] This figure shows the hardware configuration of the control device of the mobile body control system according to the first embodiment. [Figure 4] This is a functional block diagram of the control device for the mobile body control system according to the first embodiment. [Figure 5] This is an explanatory diagram relating to the predicted two-dimensional features in the mobile body control system according to the first embodiment. [Figure 6] This is an explanatory diagram relating to the predicted two-dimensional features when it is assumed that the robot moves along a predetermined trajectory candidate in the mobile body control system according to the first embodiment. [Figure 7] This flowchart shows the processes executed by the control device of the mobile body control system according to the first embodiment. [Figure 8] This is a functional block diagram of the control device for the mobile body control system according to the second embodiment. [Figure 9] This is a functional block diagram of the control device for the mobile body control system according to the third embodiment. [Figure 10] This is a functional block diagram of the control device for the mobile body control system according to the fourth embodiment. [Figure 11] This is a functional block diagram of the control device for a modified mobile vehicle control system. [Modes for carrying out the invention]
[0009] ≪First Embodiment≫ Below, we will first briefly describe the robot 20 (mobile body: see Figure 1), which is the object controlled by the control device 10 (see Figure 1), and then describe the control device 10 in detail.
[0010] Figure 1 is an explanatory diagram including a robot 20, which is the controlled object of the mobile control system 100 according to the first embodiment. The mobile control system 100 shown in Figure 1 is a system for controlling a robot 20 (mobile body) and includes a control device 10. The robot 20 is a mobile body that moves autonomously based on commands from the control device 10. It is not particularly necessary for the robot 20 to be completely autonomous; it may be controlled remotely as needed. In the example in Figure 1, the robot 20 includes a housing 21, wheels 22, and a camera 23.
[0011] The housing 21 is a box-shaped structure that forms the outer casing of the robot 20. The wheels 22 are the means of movement for the robot 20 and are driven by a motor (not shown) to rotate at a predetermined rate. Note that the configuration of the robot 20 shown in Figure 1 is just one example and is not limited to this. For example, a robot equipped with other means of movement such as walking legs or crawlers (tracks) may be used. Alternatively, a multi-jointed snake-like robot or an unmanned aerial vehicle such as a drone may be used as the robot.
[0012] The camera 23 shown in Figure 1 is used to image the area around the robot 20, and although not shown, it includes a lens and an image sensor. The lens is an optical element that refracts light and focuses it onto the image sensor. The image sensor converts the light incident through the lens into an electric image to generate an image. The camera 23 repeats imaging at a predetermined control cycle (for example, several times, tens of times, or tens of times per second). The images captured by the camera 23 are transmitted sequentially to the control device 10.
[0013] In the example shown in Figure 1, a single camera 23 is installed on the front of the housing 21 of the robot 20. Such a camera 23 could be a wide-angle camera, such as a fisheye camera or a monocular stereo camera, or a perspective projection camera. Regarding the number of lenses, a monocular camera could be used as camera 23, or a stereo camera with two lenses could be used as appropriate. Furthermore, multiple cameras could be mounted on a single robot 20. It is not particularly necessary for the orientation of the robot 20 (direction of movement) and the orientation of the camera 23 (optical axis) to coincide; the camera 23 simply needs to be installed at a predetermined position on the robot 20.
[0014] The control device 10 shown in FIG. 1 is a device that controls the robot 20 and is configured to communicate with the robot 20 either wired (such as a cable) or wirelessly. Note that the control device 10 may be composed of a single computer, or may be configured such that a plurality of computers are connected in a predetermined manner via signal lines and a network. The control device 10 determines the trajectory when the robot 20 moves based on the captured image of the camera 23, and generates a predetermined control command value to move the robot 20 along the trajectory. The control command value generated by the control device 10 is sequentially transmitted to the robot 20.
[0015] In the first embodiment and the like, the control device 10 creates a predetermined three-dimensional map based on the captured image of the camera 23 of the robot 20. Such a movement control system 100 is appropriately used, for example, when investigating the damage situation or removing fuel debris in the decommissioning work of a nuclear power plant using the robot 20.
[0016] Note that the use of the robot 20 is not limited to the above examples. For example, when performing maintenance inspection of a plant using the robot 20, or when the robot 20 moves autonomously in an environment where it is difficult to use GPS (Global Positioning System) (such as indoors where it is difficult to receive radio waves or planetary exploration), the movement control system 100 can also be applied. In short, the movement control system 100 is used when creating a three-dimensional map of an environment unknown to the robot 20, or when moving the robot 20 autonomously in an environment where it is difficult to use GPS.
[0017] FIG. 2 is an explanatory diagram regarding the trajectory candidates of the robot 20. In Figure 2, the object B1 surrounding the robot 20 is shown in a simplified form. Time t in Figure 2 is the current time when the robot 20 acquired the latest image. Time t+1 and time t+2 indicate future times based on a predetermined control cycle (for example, 0.1 seconds and 0.01 seconds after time t). The position and direction of the arrows at time t, time t+1, and time t+2 represent the position and orientation (direction) of the robot 20 at that time.
[0018] The control device 10 (see Figure 1) generates candidate trajectories for the robot 20 from time t+1 onward, based on the robot 20's position and orientation at the current time t. Here, "candidate trajectories" refers to candidate trajectories for when the robot 20 moves. In the example in Figure 2, seven candidate trajectories R1 to R7 are generated based on the robot 20's position and orientation at the current time t. Each of these candidate trajectories R1 to R7 includes waypoints at time t+1 and time t+2.
[0019] The control device 10 (see Figure 1) calculates a predetermined evaluation value for each of the multiple trajectory candidates R1 to R7, and determines the actual trajectory when the robot 20 moves based on this evaluation value. This process is repeated sequentially while the robot 20 is moving.
[0020] Figure 3 shows the hardware configuration of the control device 10. As shown in Figure 3, the control device 10 has a hardware configuration that includes a processor 10a, RAM 10b (Random Access Memory), ROM 10c (Read Only Memory), HDD 10d (Hard Disk Drive), and a communication interface 10e, which are predeterminedly connected via an internal bus 10f.
[0021] The processor 10a is hardware that executes a predetermined program. For example, a CPU (Central Processing Unit) is used as such a processor 10a. RAM 10b is volatile memory for temporarily storing predetermined data. ROM 10c and HDD 10d are non-volatile memories where predetermined data is stored. The processor 10a reads the predetermined program stored in ROM 10c and HDD 10d and loads it into RAM 10b to execute the predetermined process. The communication interface 10e is an interface for the control device 10 to exchange predetermined data with the robot 20. Note that the hardware configuration shown in Figure 3 is an example and is not limited to this.
[0022] Figure 4 is a functional block diagram of the control device 10. The control device 10 shown in Figure 4 has the following functional configuration: an image acquisition unit 111, a two-dimensional feature extraction unit 112, a two-dimensional feature storage unit 113, a three-dimensional feature extraction unit 114, and a map update unit 115. In addition to the above configurations, the control device 10 also includes a trajectory generation unit 116, a two-dimensional feature prediction unit 117, a prediction reliability calculation unit 118, a trajectory evaluation unit 119, a trajectory determination unit 120, and a control unit 121.
[0023] The image acquisition unit 111 sequentially acquires images captured by the camera 23 (see Figure 1) from the robot 20. The two-dimensional feature extraction unit 112 extracts two-dimensional features from the images captured by the camera 23. Here, "two-dimensional features" refers to unique points (feature points) in the captured image that can be distinguished from others. Such two-dimensional features are represented by a combination of the position of the feature point on the captured image (a position represented by two-dimensional pixel coordinates) and a predetermined feature descriptor at that position.
[0024] The aforementioned "feature descriptor" describes a predetermined feature based on the brightness values of multiple pixels surrounding a feature point. For extracting 2D features including such feature descriptors, methods such as SIFT (Scale-Invariant Feature Transform) and BRIEF (Binary Robust Independent Elementary Features) can be used. In addition, methods such as DISK (Discrete Keypoints) and SuperPoint, which utilize deep learning, may be used as appropriate for extracting 2D features.
[0025] The two-dimensional feature storage unit 113 shown in Figure 4 stores the aforementioned two-dimensional features and the captured images used to extract the two-dimensional features in an associated manner. The two-dimensional features stored in the two-dimensional feature storage unit 113 are then read out as appropriate by the subsequent three-dimensional feature extraction unit 114.
[0026] The 3D feature extraction unit 114 extracts 3D features corresponding to the 2D features based on the 2D features and predetermined camera parameters. Here, "3D features" are coordinate values indicating the position of a point in real space corresponding to the 2D features, and are expressed in a 3D coordinate system based on the position of the camera 23 (see Figure 1), for example.
[0027] More specifically, the 3D feature extraction unit 114 first associates physically identical locations between two captured images taken from different viewpoints (the position of the camera 23). In this association process, the 2D features of each captured image are used. Specifically, if the feature descriptors of the 2D features in the two captured images are identical (or correspond to a predetermined value), the 3D feature extraction unit 114 associates these pairs of 2D features with each other. This process is called "feature point matching." As a feature point matching method, for example, a brute-force method may be used, or a deep learning method such as SuperGlue may be used as appropriate.
[0028] The two images used in feature point matching may be taken at different times, or they may be taken simultaneously by two cameras. For example, the two images obtained from taking images at different times may be the image obtained in the current scan and the image obtained in the previous scan. Alternatively, the image obtained in the current scan may be used along with an image from a previous scan that has a relatively large number of 2D features.
[0029] The 3D feature extraction unit 114 then calculates the values of the 3D coordinates in real space corresponding to these points (i.e., 3D features) by triangulation, based on the positional relationship of points that correspond to each other in the two captured images.
[0030] The aforementioned triangulation uses predetermined camera parameters and a camera model, which is a predetermined mathematical formula that includes these camera parameters. Here, "camera parameters" are predetermined parameters that represent the characteristics of camera 23 (see Figure 1). For example, when a monocular camera is used as camera 23 (see Figure 1), internal parameters such as the lens distortion coefficient, focal length, and field of view of camera 23 are used as "camera parameters" for extracting three-dimensional features.
[0031] Furthermore, when a stereo camera with two lenses is used as camera 23 (see Figure 1), the 3D feature extraction unit 114 extracts 3D features based on the aforementioned internal parameters as well as values indicating the parallax (coordinate system shift) between captured images. The aforementioned values indicating parallax are determined based on external parameters such as the distance between the lenses. Such external parameters are also included in the "camera parameters."
[0032] Then, based on a pair of mutually corresponding 2D features between the two captured images, one 3D feature is identified. This 3D feature indicates the position of a point on the object surface around the robot 20. A predetermined 3D map representing the environment around the robot 20 is a point cloud data containing numerous 3D features.
[0033] The map update unit 115 shown in Figure 4 generates or updates a 3D map by registering the aforementioned 3D features. As described above, the "3D map" is point cloud data representing the surrounding environment (surface shape of objects) of the robot 20. Furthermore, "updating" the 3D map means rewriting the generated 3D map data to a predetermined format. When updating such a 3D map, a map update rule based on SLAM (Simultaneous Localization and Mapping) may be used as appropriate.
[0034] The trajectory generation unit 116 shown in Figure 4 generates one or more trajectory candidates for when the robot 20 (mobile body: see Figure 1) moves. As mentioned above, "trajectory candidate" refers to a candidate trajectory for when the robot 20 moves (see Figure 2). For generating such trajectory candidates, a sampling-based trajectory generation method such as MPPI (Model Predictive Path Integral Control) is used as appropriate. The trajectory candidate data generated by the trajectory generation unit 116 includes the position of the robot 20 and the attitude (orientation) of the robot 20 at a future time (for example, time t+1 or time t+2 in Figure 2).
[0035] The 2D feature prediction unit 117 calculates predicted 2D features based on camera parameters, indicating the positions where 3D features will appear on the captured image when the robot 20 (mobile body) moves along a predetermined trajectory candidate from among multiple trajectory candidates. The aforementioned "predicted 2D features" indicate the positions where 3D features will appear on the captured image, assuming that the image is captured at a waypoint included in the predetermined trajectory candidate. To give a specific example, assuming that the robot 20 moves along trajectory candidate R1 in Figure 2, and that the image is captured at a waypoint at time t+1, the positions where a predetermined 3D feature (a 3D feature based on the image captured at the current time t) will appear on that captured image are calculated as predicted 2D features.
[0036] Such predicted two-dimensional features are calculated using a geometric method based on three-dimensional features and predetermined camera parameters (e.g., distortion coefficient, focal length, and field of view). The predicted two-dimensional features are represented by the two-dimensional pixel coordinate values in the captured image.
[0037] Figure 5 is an explanatory diagram regarding the predicted two-dimensional features. Note that arrow w0 in Figure 5 represents the position and orientation (direction) of robot 20 at the current time t (when the latest image was acquired). Arrows w1 and w2 represent the position and orientation (direction) of robot 20 at future times t+1 and t+2, assuming that robot 20 moves along a predetermined trajectory candidate.
[0038] The 3D features k1 to k6 shown in Figure 5 are predetermined 3D features that have already been extracted at the current time t. The captured image G represents the image obtained assuming that the image was captured at time t+1 at the position and orientation indicated by arrow w1. In the example in Figure 5, the 3D features k1 to k4 that are within the field of view F of camera 23 (see Figure 1) are captured in the captured image G. The positions of these 3D features k1 to k4 on the captured image G1 correspond to the "predicted 2D features" mentioned above.
[0039] For example, if a fisheye camera is used as camera 23 (see Figure 1), the image distortion is small in the central region of the captured image, but the image distortion is large in the peripheral region of the captured image. Also, if a monocular stereo camera is used as camera 23 (see Figure 1), the magnitude of image distortion changes to a predetermined value in the vertical direction of the captured image.
[0040] Furthermore, in a given region of an captured image, the greater the image distortion in that region, and the lower the spatial resolution, the more inaccurate the feature descriptor values of the 2D features extracted from that region tend to be. If the feature descriptor values are inaccurate, the accuracy of the feature point matching (corresponding 2D features between two captured images) will decrease, resulting in a lower accuracy for the 3D map.
[0041] Therefore, in the first embodiment, the prediction confidence calculation unit 118 shown in Figure 4 calculates the prediction two-dimensional feature confidence based on the camera parameters. Here, "prediction two-dimensional feature confidence" refers to the confidence level when a predetermined prediction two-dimensional feature is used for feature point matching. The prediction confidence calculation unit 118 (see Figure 4) calculates the change amount α of the position of the prediction two-dimensional feature on the captured image when lens distortion correction (coordinate transformation) is performed on the captured image from the camera 23 (see Figure 1). The change amount α is an index value that indicates the degree of image distortion at the position (pixel) of the prediction two-dimensional feature, and is calculated based on the following equation (1).
[0042]
number
[0043] Note that the x and y in equation (1) represent the x and y values of the predicted 2D features in the pixel coordinates on the captured image, in that order. Also, the x in equation (1) C ,y C This shows the x and y values of the predicted 2D features at the pixel coordinates when the image distortion is almost eliminated by correcting for lens distortion (coordinate transformation). Note that lens distortion correction is performed based on predetermined camera parameters. The greater the image distortion at point (x,y), the greater the change in the predicted 2D features before and after lens distortion correction, and therefore the larger the value of the change amount α.
[0044] Furthermore, the prediction confidence calculation unit 118 (see Figure 4) calculates the distance in real space between the target 3D feature and the camera 23 (see Figure 1) as an index value d indicating the level of spatial resolution. Generally, the longer the distance between the camera 23 and the target object, the lower the spatial resolution tends to be.
[0045] The prediction confidence calculation unit 118 (see Figure 4) then calculates the predicted two-dimensional feature confidence R based on the following equation (2). The value k included in equation (2) is a predetermined coefficient that is set in advance by the user. The denominator on the right side of equation (2) is the product of an index value d indicating the height of spatial resolution and a value indicating the degree of image distortion (a change amount α based on equation (1)).
[0046]
number
[0047] The smaller the image distortion in the captured image, and the higher the spatial resolution, the higher the value of the predicted 2D feature confidence R in equation (2), and the higher the accuracy of feature point matching tends to be. Therefore, among multiple trajectory candidates (for example, trajectory candidates R1 to R7 in Figure 2), it is desirable to select trajectory candidates that have a large number of predicted 2D features with a high value of the predicted 2D feature confidence R.
[0048] The trajectory evaluation unit 119 shown in Figure 4 calculates an evaluation value for each trajectory candidate based on the predicted two-dimensional feature confidence R. One evaluation value is calculated for each trajectory candidate. For example, the trajectory evaluation unit 119 calculates the number of predicted two-dimensional feature confidence values at one or more waypoints included in a given trajectory candidate (for example, the waypoints at time t+1 and time t+2 in Figure 2) that are equal to or greater than a predetermined value as an evaluation value.
[0049] The method for calculating the evaluation value described above is merely an example and is not limited thereto. For example, the trajectory evaluation unit 119 may calculate the evaluation value based on the sum of the predicted two-dimensional feature confidence scores for each of the one or more waypoints included in the predetermined trajectory candidate.
[0050] Furthermore, the trajectory evaluation unit 119 may calculate a weighted average of the predicted two-dimensional feature confidence scores for one or more waypoints included in a predetermined trajectory candidate, such that a greater weight is given to the waypoint closer to the acquisition time of the latest image (for example, time t in Figure 2), and calculate an evaluation value based on this weighted average. This makes it possible to calculate an evaluation value for a trajectory candidate from the perspective of giving more emphasis to a more reliable prediction of the near future.
[0051] The higher the evaluation value of a given trajectory candidate, the more 2D features can be acquired that improve the accuracy of feature point matching when the robot 20 moves along that trajectory candidate. In other words, a large number of 2D features can be acquired from regions in the captured image where distortion is small and spatial resolution is high.
[0052] Figure 6 is an explanatory diagram of the predicted two-dimensional features assuming that the robot 20 moves along a predetermined trajectory candidate. The captured image G0 shown in Figure 6 represents the 3D features k1 to k5 that appear in the latest captured image at the current time t. Here, "appearing" the 3D features k1 to k5 in the captured image means that as the position and orientation (orientation) of the robot 20 changes, the positions corresponding to the 3D features k1 to k5 are identified on the captured image (calculated as predicted 2D features).
[0053] Image G1 shows the positions where 3D features k1 to k5 appear in the image at time t+1 when the robot 20 moves along a predetermined trajectory candidate R1 (i.e., predicted 2D features). Image G7 shows the positions where 3D features k1 to k5 appear in the image at time t+1 when the robot 20 moves along another trajectory candidate R7 (i.e., predicted 2D features). It is assumed that trajectory candidate R1 has a higher evaluation value than trajectory candidate R7.
[0054] In the example in Figure 6, among the 3D features k1 to k5 in the captured images G0, G1, and G7, those with a confidence level (a value based on equation (2); in the case of captured images G1 and G7 at time t+1, the predicted 2D feature confidence level) of the corresponding 2D features are shown in white, while those with a confidence level below the predetermined value are shown in black.
[0055] As mentioned above, the predicted 2D feature confidence is the confidence of predicted 2D features in future captured images, but it is also possible to calculate the confidence of 2D features extracted from actual captured images based on equation (2). In the following explanation, the confidence of 2D features corresponding to 3D features (including the case of predicted 2D feature confidence) will sometimes simply be referred to as the confidence of 3D features.
[0056] For example, in the image captured at the current time t, there are 3D features k1 to k3 whose confidence level is below a predetermined value. Subsequently, if the robot 20 moves along the candidate trajectory R1 (the candidate trajectory with a high evaluation value), it is predicted that the confidence level of 3D features k1 to k3 will increase at the position at time t+1. For example, if the 3D features k1 to k3 are captured in a region with small image distortion in the image G1 at time t+1, their confidence level tends to increase. In the example in Figure 6, the number of 3D features with high confidence level increases from 3 (time t) to 5 (time t+1 at candidate trajectory R1) due to the movement of the robot 20, making it possible to perform feature point matching with high accuracy.
[0057] On the other hand, if the robot 20 moves along another trajectory candidate R7 (a trajectory candidate with a low evaluation value), it is predicted that the confidence level of the 3D features k1 to k3 at time t+1 will remain low. For example, if the 3D features k1 to k3 are located in a region with high image distortion in the captured image G7 at time t+1, their confidence level tends to be low. Therefore, it is desirable to select a trajectory candidate R1 (or a predetermined trajectory close to trajectory candidate R1) with a relatively high evaluation value.
[0058] The trajectory determination unit 120 shown in Figure 4 determines the trajectory of the robot (mobile body) based on the evaluation values described above. For example, the trajectory determination unit 120 identifies the trajectory with the highest evaluation value among several trajectory candidates and sets this trajectory candidate as the actual trajectory. Alternatively, the trajectory determination unit 120 may set the trajectory by performing a weighting calculation based on the evaluation values of each of the multiple trajectory candidates and taking a weighted average of the trajectory candidates (taking a weighted average of the speed and acceleration of the robot 20 when moving along the trajectory candidate).
[0059] The control unit 121 shown in Figure 4 controls the robot 20 (mobile body) to move along a predetermined trajectory determined by the trajectory determination unit 120. In controlling the robot 20 in this way, for example, a PID (Proportional Integral Derivative) controller may be used as appropriate. Also, in the trajectory generation unit 116 (see Figure 4), if the aforementioned MPPI is used to generate trajectory candidates, a predetermined control amount calculated at the time of trajectory candidate generation may be used for the actual control of the robot 20. The control command values generated by the control unit 121 are transmitted sequentially to the robot 20.
[0060] Figure 7 is a flowchart showing the processes performed by the control device (see also Figure 4 as appropriate). In Figure 7, during the "START" phase, it is assumed that the robot 20 is moving while the camera 23 repeatedly takes images at a predetermined control cycle. In step S101, the control device 10 acquires the image captured by the camera 23 (see Figure 1) of the robot 20 using the image acquisition unit 111 (image acquisition step). In step S102, the control device 10 extracts two-dimensional features from the captured image using the two-dimensional feature extraction unit 112 (two-dimensional feature extraction step). The two-dimensional features extracted in this way are stored in the two-dimensional feature storage unit 113 in association with the captured image.
[0061] In step S103, the control device 10 extracts three-dimensional features using the three-dimensional feature extraction unit 114 (three-dimensional feature extraction step). Specifically, the control device 10 first associates two-dimensional features between two captured images by feature point matching. Then, the control device 10 extracts three-dimensional features based on the positional relationship of the corresponding two-dimensional features between the two captured images and predetermined camera parameters.
[0062] In step S104, the control device 10 generates or updates a 3D map using the map update unit 115 (map update step). That is, the control device 10 generates or updates a 3D map as appropriate by registering the 3D features extracted in step S103. In step S105, the control device 10 generates one or more trajectory candidates using the trajectory generation unit 116 (trajectory generation step). For example, the control device 10 generates trajectory candidates R1 to R7 as shown in Figure 2.
[0063] In step S106, the control device 10 calculates predicted two-dimensional features using the two-dimensional feature prediction unit 117 (two-dimensional feature prediction step). That is, the control device 10 calculates predicted two-dimensional features based on camera parameters that indicate where three-dimensional features will appear on the captured image when the robot 20 moves along a predetermined trajectory candidate.
[0064] In step S107, the control device 10 calculates the predicted two-dimensional feature reliability using the prediction reliability calculation unit 118 (prediction reliability calculation step). That is, the control device 10 calculates the predicted two-dimensional feature reliability, which indicates how reliable the predicted two-dimensional features are when performing feature point matching, based on the camera parameters. As mentioned above, if the predicted two-dimensional features are located in areas of the captured image with large image distortion or low spatial resolution, the predicted two-dimensional feature reliability tends to be low.
[0065] In step S108, the control device 10 calculates an evaluation value for the trajectory candidate using the trajectory evaluation unit 119 (trajectory evaluation step). That is, the control device 10 calculates an evaluation value for the trajectory candidate based on the predicted two-dimensional feature confidence level at each waypoint included in the trajectory candidate. In step S109, the control device 10 determines the trajectory of the robot 20 when it moves using the trajectory determination unit 120 (trajectory determination step). For example, the control device 10 determines the trajectory candidate with the highest evaluation value as the actual trajectory of the robot 20.
[0066] In step S110, the control device 10 controls the robot 20 by the control unit 121 to move the robot 20 along the aforementioned trajectory (control step). After the processing in step S110, the control device returns to "START" (RETURN). The series of processes shown in Figure 7 are to be repeated at predetermined intervals.
[0067] <Effects> According to the first embodiment, assuming that the robot 20 moves along a predetermined trajectory candidate, the positions of 2D features on the captured image are calculated as predicted 2D features. Then, the confidence level of the predicted 2D features, which is the predicted 2D feature confidence level, is calculated based on predetermined camera parameters, and an evaluation value of the trajectory candidate is calculated based on this predicted 2D feature confidence level. This makes it possible to quantitatively evaluate which trajectory candidate is desirable even when using special cameras where the degree of image distortion differs depending on the position on the captured image (for example, a fisheye camera or a monocular stereo camera). Therefore, since feature point matching based on the captured image can be performed with high accuracy, the accuracy of the 3D map can be improved.
[0068] Furthermore, by creating an accurate 3D map, the robot 20 can move appropriately and perform predetermined tasks such as decommissioning work, even in unfamiliar environments such as the inside of a reactor building where rubble, damaged items, or fuel debris are present.
[0069] ≪Second Embodiment≫ The second embodiment differs from the first embodiment in that, among the multiple three-dimensional features extracted by the three-dimensional feature extraction unit 114 (see Figure 8), those with a confidence level equal to or greater than a predetermined value are registered in the three-dimensional map. Other aspects are the same as the first embodiment. Therefore, the differences from the first embodiment will be explained, and the overlapping parts will be omitted.
[0070] Figure 8 is a functional block diagram of the control device 10A of the mobile body control system 100A according to the second embodiment. The control device 10A shown in Figure 8 includes a reliability determination unit 122 in addition to the configurations described in the first embodiment (see Figure 4). The reliability determination unit 122 calculates the reliability of using the 2D features extracted by the 2D feature extraction unit 114 for feature point matching, and determines whether this reliability is above a predetermined value. The predetermined value is a confidence threshold that serves as a criterion for deciding whether to register 3D features in a 3D map or to use them when calculating predicted 2D features, and is set in advance.
[0071] The calculation of the confidence level of the 2D features is the same as the processing content of the prediction confidence level calculation unit 118. Specifically, the confidence level determination unit 122 calculates the confidence level of the target 2D feature based on the equation (2) described above. This confidence level indicates how reliable the 2D features in the captured image obtained from actual imaging can be when used for feature point matching (how much they contribute to the creation of a highly accurate 3D map).
[0072] Incidentally, the confidence value calculated by the confidence determination unit 122 is approximately the same as the value calculated when the target 2D feature was previously treated as a predicted 2D feature. Therefore, for example, the confidence determination unit 122 may be omitted as appropriate, and the past predicted 2D feature confidence calculated by the prediction confidence calculation unit 118 may be used as the confidence value of the 2D feature in the current captured image.
[0073] In this way, the confidence level determination unit 122 calculates the confidence level for each of the multiple three-dimensional features extracted by the three-dimensional feature extraction unit 114. The confidence level determination unit 122 then outputs to the map update unit 115 any of the multiple three-dimensional features whose confidence level is equal to or greater than a predetermined value.
[0074] The map update unit 115 registers the 3D features output from the confidence determination unit 122 into the 3D map. Specifically, the map update unit 115 registers into the 3D map only those 3D features extracted by the 3D feature extraction unit 114 whose corresponding 2D features have a confidence level equal to or greater than a predetermined value. This allows for the registration of 3D features corresponding to 2D features (2D features that enable high-precision feature point matching) in regions of the captured image where image distortion is small and spatial resolution is high into the 3D map.
[0075] On the other hand, the reliability determination unit 122 outputs to the 2D feature prediction unit 117, but does not output to the map update unit 115, any 3D features whose reliability is below a predetermined value. This prevents 3D features based on 2D features with low feature point matching accuracy from being registered in the 3D map. In addition, since the frequency of updating the 3D map is reduced compared to the first embodiment, the computational load on the control device 10A can be reduced.
[0076] The 2D feature prediction unit 117 calculates predicted 2D features for 3D features extracted by the 3D feature extraction unit 114, focusing on those 3D features whose corresponding 2D features have a confidence level below a predetermined value. This makes it possible to set trajectories (trajectories that increase confidence) that will allow 3D features that currently have low confidence and have not been registered in the 3D map to be registered in the future. In other words, it is possible to set trajectories so that 3D features are captured in areas of the captured image where image distortion is small and spatial resolution is high.
[0077] Furthermore, if a predetermined object (not shown) is newly placed in the field of view of camera 23 (see Figure 1) such that it appears in an area where the distortion of the captured image is relatively large, the control unit 121 may change the trajectory of the robot 20 (mobile body) so that camera 23 faces the object directly. This causes the object (the object newly placed by a worker, etc.) to appear in an area of the captured image where the image distortion is small and the spatial resolution is high, thus enabling the generation of highly accurate 3D features of this object. As a result, an accurate 3D map including the shape of the outer surface of this object can be generated. Incidentally, "an area where the distortion of the captured image is relatively large" refers, for example, to the area near the edge of the captured image if camera 23 is a fisheye camera, or to the area where the image distortion is large in the vertical direction of the captured image if it is a monocular stereo camera.
[0078] <Effects> According to the second embodiment, among the multiple 3D features, those with a confidence level above a predetermined value are registered in the 3D map. Therefore, a 3D map with even higher accuracy can be created than in the first embodiment. In addition, among the multiple 3D features, those with a confidence level below a predetermined value are used as targets when calculating predicted 2D features. This makes it possible to set a predetermined trajectory so that 3D features that currently have low confidence levels and are not registered in the 3D map will be registered in the 3D map in the future.
[0079] ≪Third Embodiment≫ The third embodiment differs from the first embodiment in that it registers novel 3D features from among the multiple 3D features extracted by the 3D feature extraction unit 114 (see Figure 9) in a 3D map, and calculates predicted 2D features based on these novel features. Other aspects are the same as the first embodiment. Therefore, we will explain the parts that differ from the first embodiment, and omit explanations of overlapping parts.
[0080] Figure 9 is a functional block diagram of the control device 10B of the mobile body control system 100B according to the third embodiment. The control device 10B shown in Figure 9 includes a new feature determination unit 123 in addition to the configurations described in the first embodiment (see Figure 4). The new feature determination unit 123 determines whether the 3D features extracted by the 3D feature extraction unit 114 are new 3D features that have not yet been registered in the 3D map.
[0081] If a 3D feature has not yet been registered in the 3D map, the new feature determination unit 123 outputs the 3D feature as a new 3D feature to the map update unit 115 and also outputs it to the 2D feature prediction unit 117. The map update unit 115 registers the 3D features that correspond to the new 3D features from among the multiple 3D features extracted by the 3D feature extraction unit 114 in the 3D map.
[0082] The 2D feature prediction unit 117 calculates predicted 2D features by targeting the 3D features that correspond to new 3D features from among the multiple 3D features extracted by the 3D feature extraction unit 114. Note that 3D features already registered in the 3D map do not need to be used in the processing of the 2D feature prediction unit.
[0083] <Effects> According to the third embodiment, by generating trajectories that prioritize the reliability of novel two-dimensional features, a three-dimensional map of an unknown environment can be efficiently created.
[0084] ≪Fourth Embodiment≫ The fourth embodiment is a combination of the second embodiment (see Figure 8) and the third embodiment (see Figure 9). In other words, the fourth embodiment differs from the first embodiment in that it reflects novel 3D features with a confidence level of or higher than a predetermined value, among the multiple 3D features extracted by the 3D feature extraction unit 114 (see Figure 10), onto the 3D map. Furthermore, the fourth embodiment differs from the first embodiment in that it calculates predicted 2D features targeting novel 3D features with a confidence level below a predetermined value. Other aspects are the same as the first embodiment. Therefore, we will explain the parts that differ from the first embodiment, and omit explanations of overlapping parts.
[0085] Figure 10 is a functional block diagram of the control device 10C of the mobile body control system 100C according to the fourth embodiment. The control device 10C shown in Figure 10 includes, in addition to the configurations described in the first embodiment (see Figure 4), a new feature determination unit 123 and a reliability determination unit 122. The new feature determination unit 123 determines whether the 3D feature extracted by the 3D feature extraction unit 114 is a new 3D feature that has not yet been registered in the 3D map. If the 3D feature has not yet been registered in the 3D map, the new feature determination unit 123 outputs the 3D feature as a new 3D feature to the map update unit 115, and also outputs it to the reliability determination unit 122 along with the 2D feature (new 2D feature) corresponding to the new 3D feature.
[0086] The confidence level determination unit 122 calculates the confidence level of the new 2D features used to extract the new 3D features. The confidence level determination unit 122 outputs to the map update unit 115 any new 3D features whose confidence level is above a predetermined value. The confidence level determination unit 122 outputs to the 2D feature prediction unit 117 any new 3D features whose confidence level is below the predetermined value, instead of outputting to the map update unit 115.
[0087] <Effects> According to the fourth embodiment, an accurate 3D map can be efficiently created by registering new 3D features with a confidence level of or higher than a predetermined value in the 3D map. Furthermore, by calculating predicted 2D features for new 3D features with a confidence level below a predetermined value, it is possible to set a trajectory that will allow new 3D features that currently have low confidence levels and have not been registered in the 3D map to be registered in the 3D map in the future.
[0088] ≪Variations≫ Although the mobile body control system 100 and mobile body control method, etc. related to this disclosure have been described in detail in each embodiment, this disclosure is not limited to these descriptions, and various modifications can be made.
[0089] Figure 11 is a functional block diagram of the control device 10D of a modified mobile body control system. In the modified example shown in Figure 11, the robot 20D is configured to include a camera 23 and a control device 10. The configuration of the control device 10 may be any of the first to fourth embodiments. The robot 20 moves along a trajectory determined by the control device 10. This configuration also achieves the same effects as the other embodiments.
[0090] Furthermore, in each embodiment, the case in which the prediction confidence calculation unit 118 (see Figure 4) calculates the prediction two-dimensional feature confidence based on both the amount of change in the position of the prediction two-dimensional feature due to lens distortion correction and the spatial resolution in the captured image (i.e., equation (2) described above), has been described. However, one of these may be used and the other may be omitted as appropriate. That is, the prediction confidence calculation unit 118 may calculate the prediction two-dimensional feature confidence based on the amount of change in the position of the prediction two-dimensional feature on the captured image when lens distortion correction is performed on the captured image of the camera 23. Alternatively, the prediction confidence calculation unit 118 may calculate the prediction two-dimensional feature confidence based on the spatial resolution at the position of the prediction two-dimensional feature in the captured image of the camera 23.
[0091] Furthermore, while each embodiment describes a case where the "mobile body" whose trajectory is set by the control device 10 is a robot 20, it is not limited to this. For example, each embodiment can be applied to various other types of "mobile bodies," such as automobiles. Furthermore, while each embodiment describes a case where the control device 10 creates a 3D map, it is not limited to this. For example, it is also possible to create a predetermined 2D map based on the 3D map. In addition, the data included in the 3D map is not limited to 3D features. For example, predetermined color data represented by RGB (data indicating the color of pixels on the captured image) may be associated with each 3D feature.
[0092] Furthermore, the processing (such as the mobile object control method) of the mobile object control system 100 and the control device 10 may be executed as a predetermined program on a computer. The aforementioned program can be provided via a communication line, or it can be written to a recording medium such as a CD-ROM and distributed.
[0093] Furthermore, this disclosure is not limited to the embodiments and includes various modifications. For example, the embodiments are described in detail for illustrative purposes and are not necessarily limited to having all the configurations described. Also, some of the configurations of the embodiments can be added, deleted, or replaced with other configurations.
[0094] Furthermore, each of the aforementioned configurations, functions, processing units, processing means, etc., may be implemented in hardware, either partially or entirely, by designing them as integrated circuits, for example. Alternatively, each of the aforementioned configurations, functions, etc., may be implemented in software by having the processor interpret and execute programs that realize each function. Information such as programs, tables, and files that realize each function can be stored in memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD. Furthermore, the control lines and information lines shown are those deemed necessary for explanatory purposes, and not all control lines and information lines are necessarily shown in the actual product. In reality, it can be assumed that almost all components are interconnected. [Explanation of Symbols]
[0095] 10, 10A, 10B, 10C control devices 20,20D Robot (Mobile) 21 cabinets 22 wheels 23 Cameras 100, 100A, 100B, 100C Mobile Control System 111 Image acquisition unit 112 Two-dimensional feature extraction unit 113 Two-dimensional feature memory unit 114 3D Feature Extraction Unit 115 Map Update Department 116 Trajectory generation part 117 2D Feature Prediction Unit 118 Prediction Confidence Calculation Unit 119 Orbital Evaluation Department 120 Orbit determination part 121 Control Unit 122 Confidence Determination Unit 123 Novel Feature Determination Unit B1 Object R1,R2,R3,R4,R5,R6,R7 Orbit candidates S101 Step (Image acquisition step) S102 Step (2D Feature Extraction Step) S103 Step (3D Feature Extraction Step) S104 Step (Map Update Step) S105 Step (Orbital Generation Step) S106 Step (2D Feature Prediction Step) S107 Step (Prediction Confidence Calculation Step) S108 Step (Trajectory Evaluation Step) S109 Step (Trajectory Determination Step) S110 Step (Control Step)
Claims
1. An image acquisition unit that acquires images captured by a camera on a mobile device, It includes a two-dimensional feature extraction unit that extracts two-dimensional features from the captured image, A three-dimensional feature extraction unit extracts three-dimensional features indicating the position of a point in real space corresponding to the two-dimensional features, based on the two-dimensional features and predetermined camera parameters indicating the characteristics of the camera. A map update unit that generates or updates a 3D map by registering the aforementioned 3D features, A trajectory generation unit that generates a plurality of trajectory candidates when the moving object moves, A two-dimensional feature prediction unit calculates a predicted two-dimensional feature based on the camera parameters that indicates the position where the three-dimensional feature will appear on the captured image when the moving object moves along a predetermined trajectory candidate from among a plurality of trajectory candidates, A prediction confidence calculation unit calculates a prediction two-dimensional feature confidence score, which is the confidence score when the aforementioned prediction two-dimensional features are used for feature point matching, based on the camera parameters. A trajectory evaluation unit calculates an evaluation value for each of the trajectory candidates based on the predicted two-dimensional feature reliability, A trajectory determination unit that determines the trajectory of the moving body based on the evaluation value, A mobile body control system comprising: a control unit that controls the mobile body to move the mobile body along the aforementioned trajectory.
2. The prediction confidence calculation unit calculates the prediction two-dimensional feature confidence based on the amount of change in the position of the prediction two-dimensional feature on the captured image when lens distortion correction is performed on the captured image of the camera. A mobile body control system according to claim 1, characterized by the following:
3. The prediction confidence calculation unit calculates the prediction two-dimensional feature confidence based on the spatial resolution at the location of the prediction two-dimensional feature in the image captured by the camera. A mobile body control system according to claim 1, characterized by the following:
4. The trajectory evaluation unit calculates the number of predicted two-dimensional feature confidence values at each of the one or more waypoints included in the trajectory candidate that are equal to or greater than a predetermined value, and uses this number as the evaluation value. A mobile body control system according to claim 1, characterized by the following:
5. The trajectory evaluation unit calculates the evaluation value based on the sum of the predicted two-dimensional feature confidence scores for each of the one or more waypoints included in the trajectory candidate. A mobile body control system according to claim 1, characterized by the following:
6. The trajectory evaluation unit calculates a weighted average of the predicted two-dimensional feature reliability for one or more waypoints included in the predetermined trajectory candidate, such that a greater weight is given to points closer to the acquisition time of the latest image, and calculates the evaluation value based on the weighted average. A mobile body control system according to claim 1, characterized by the following:
7. The system includes a reliability determination unit that calculates the reliability of using the two-dimensional features extracted by the two-dimensional feature extraction unit for feature point matching, and determines whether the reliability is equal to or greater than a predetermined value. The map update unit registers, among the plurality of three-dimensional features extracted by the three-dimensional feature extraction unit, those in which the reliability of the two-dimensional feature corresponding to the three-dimensional feature is equal to or greater than the predetermined value, into the three-dimensional map. A mobile body control system according to claim 1, characterized by the following:
8. The two-dimensional feature prediction unit calculates the predicted two-dimensional feature by targeting, among the plurality of three-dimensional features extracted by the three-dimensional feature extraction unit, those for which the confidence level of the two-dimensional feature corresponding to the three-dimensional feature is less than the predetermined value. A mobile body control system according to claim 7, characterized by the following:
9. If a predetermined object is newly placed in the camera's field of view such that it appears in an area where the distortion of the captured image is relatively large, the control unit changes the trajectory of the moving object so that the camera faces the object directly. A mobile body control system according to claim 1, characterized by the following:
10. The system includes a new feature determination unit that determines whether the three-dimensional feature extracted by the three-dimensional feature extraction unit is a new three-dimensional feature that has not yet been registered in the three-dimensional map. The map update unit registers the three-dimensional features that correspond to the new three-dimensional features from among the multiple three-dimensional features into the three-dimensional map. A mobile body control system according to claim 1, characterized by the following:
11. The two-dimensional feature prediction unit calculates the predicted two-dimensional features by targeting the ones among the plurality of three-dimensional features that correspond to the new three-dimensional features. A mobile body control system according to claim 10, characterized by the above.
12. Image acquisition step to acquire images captured by a camera on a moving object, The process includes a two-dimensional feature extraction step for extracting two-dimensional features from the captured image, A three-dimensional feature extraction step, which extracts a three-dimensional feature indicating the position of a point in real space corresponding to the two-dimensional feature, based on the two-dimensional feature and predetermined camera parameters indicating the characteristics of the camera. A map update step that generates or updates a 3D map by registering the aforementioned 3D features, A trajectory generation step that generates a plurality of trajectory candidates when the moving object moves, A two-dimensional feature prediction step is to calculate a predicted two-dimensional feature based on the camera parameters that indicates the position in the captured image where the three-dimensional feature is visible when the moving object moves along a predetermined trajectory candidate from among a plurality of trajectory candidates, A prediction confidence calculation step, which calculates the prediction two-dimensional feature confidence, which is the confidence level when the prediction two-dimensional features are used for feature point matching, based on the camera parameters, A trajectory evaluation step in which an evaluation value is calculated for each of the trajectory candidates based on the predicted two-dimensional feature confidence, A trajectory determination step in which the trajectory of the moving body is determined based on the evaluation value, A method for controlling a moving body, comprising a control step of controlling the moving body to move the moving body along the aforementioned trajectory.