Device for performing distance estimation and method for performing distance estimation

The device improves distance estimation accuracy in unknown environments by selecting and retraining images using preset conditions, reducing user burden and complexity, and preventing overfitting.

WO2026094254A1PCT designated stage Publication Date: 2026-05-07YAMAHA MOTOR CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
YAMAHA MOTOR CO LTD
Filing Date
2024-11-01
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing distance estimation devices using images face accuracy issues when deployed in environments different from the training environment, necessitating user-prepared retraining images, which increases user burden.

Method used

A device and method that selects images for retraining a learning model based on preset conditions, allowing retraining without additional user effort, using an imaging unit, distance estimation unit, and learning unit to improve accuracy in unknown environments.

Benefits of technology

Enhances distance estimation accuracy in unknown environments by retraining the learning model with selected images, reducing user burden and complexity, and preventing overfitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024039083_07052026_PF_FP_ABST
    Figure JP2024039083_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A device 100 for performing distance estimation comprises: an imaging unit 1; a distance estimation unit 20 for inputting a captured image 41 to a learning model 40 and estimating the distance Dt to an object 90; an image selection unit 21 for selecting a retraining image 42, used for retraining the learning model 40, from the captured image 41 on the basis of a preset condition; and a training unit 22 for retraining the learning model 40 by performing machine learning on the basis of the selected retraining image 42.
Need to check novelty before this filing date? Find Prior Art

Description

Device for distance estimation and method for distance estimation

[0001] This invention relates to a device for estimating distance and a method for estimating distance, and more particularly to a device for estimating the distance to an object using an image and a method for estimating distance.

[0002] Conventionally, devices that use images to estimate the distance to an object are known. Such a device is disclosed, for example, in Japanese Patent Publication No. 6922399.

[0003] Japanese Patent Publication No. 6922399 discloses an image processing device that measures distance using two images. The image processing device disclosed in Japanese Patent Publication No. 6922399 is configured to acquire the parallax between two images and to acquire the distance to an object based on the acquired parallax.

[0004] Patent No. 6922399

[0005] Although not disclosed in the aforementioned Japanese Patent Publication No. 6922399, it is known that the distance to an object can be accurately obtained (estimated) using an image by using a learning model trained with a deep learning algorithm. However, when using an image taken in an unknown environment (location) other than the environment (location) in which the images used for training to generate the learning model were taken, the accuracy of estimating the distance to the object decreases. In this case, the accuracy of estimating the distance to the object can be improved by retraining the learning model using images taken in the unknown environment. However, in order to retrain the learning model, the user needs to prepare training images (retraining images) in the unknown environment, which increases the burden on the user. Therefore, there is a need for a device and method for performing distance estimation that can improve the accuracy of distance estimation while suppressing an increase in the burden on the user when estimating distance from images taken in an unknown environment other than the environment in which the images used for training to generate the learning model were taken.

[0006] This invention was made to solve the above-mentioned problems, and one objective of this invention is to provide a distance estimation device and a distance estimation method that can improve the accuracy of distance estimation while suppressing an increase in the burden on the user when estimating distance from images taken in an unknown environment other than the environment in which the images used for training to generate a learning model were taken.

[0007] The apparatus for performing distance estimation according to the first aspect of this invention comprises: an imaging unit that images an object; a distance estimation unit that inputs the images captured by the imaging unit into a learning model and estimates the distance to the object; an image selection unit that selects images for retraining the learning model from the images captured by the imaging unit based on preset conditions; and a learning unit that retrains the learning model by performing machine learning based on the selected images for retraining the learning model.

[0008] The distance estimation apparatus according to the first aspect of this invention comprises, as described above, an image selection unit that selects images for retraining from captured images captured by the imaging unit based on preset conditions, and a learning unit that retrains the learning model by performing machine learning based on the selected images for retraining. As a result, since images for retraining to be used for retraining the learning model are selected from captured images captured by the imaging unit, the learning model can be retrained without the user having to prepare images for retraining. Therefore, an increase in the user's burden can be suppressed. Furthermore, the learning unit retrains the learning model using the images for retraining selected by the image selection unit from captured images captured by the imaging unit based on preset conditions. Therefore, even if the apparatus is placed in an unknown environment different from the environment in which the images used to generate the learning model were captured, the learning model can be retrained using images for retraining selected based on preset conditions that allow for the selection of images capable of improving the distance estimation accuracy from captured images taken in that unknown environment. Therefore, even if the apparatus is placed in an unknown environment, the distance estimation accuracy can be improved by retraining the learning model using the images for retraining. As a result, when estimating distance from images taken in an unknown environment other than the environment in which the images used for training the learning model were taken, it is possible to improve the accuracy of distance estimation while suppressing an increase in the user's burden.

[0009] In the device for estimating distance according to the first phase described above, preferably, a mobile body equipped with an imaging unit is further provided, the imaging unit is configured to capture images while the mobile body is moving, the image selection unit selects retraining images from the captured images taken while the mobile body is moving, and the distance estimation unit estimates the distance between the mobile body and the object while the mobile body is moving. With this configuration, the learning model can be retrained using retraining images selected from captured images taken while the mobile body is moving in various environments (locations). As a result, the accuracy of the distance estimation to the object by the learning model can be improved when the mobile body is moving in various environments (locations).

[0010] In this case, preferably, the system further includes a control unit that controls the movement of the moving object based on the distance between the moving object and the target object estimated by the distance estimation unit. With this configuration, it is possible to estimate the distance accurately, and therefore the movement of the moving object can be controlled accurately.

[0011] In the device that performs distance estimation according to the first phase described above, preferably, it further includes a storage unit that stores pre-training images used for training before the learning model is retrained. The image selection unit obtains a first similarity score, which is the degree of similarity between the captured image and the pre-training images, and determines whether the first similarity score satisfies a preset condition. If it satisfies the condition, the captured image is selected as the retraining image to be used for retraining the learning model. With this configuration, by using the first similarity score between the captured image and the pre-training images as the selection criterion for the retraining image, excessive optimization (overfitting) for the environment in which the captured image was taken can be suppressed. As a result, the accuracy of distance estimation can be easily improved when retraining the learning model using the retraining image. Furthermore, since the device that performs distance estimation itself has a storage unit that stores the retraining images, unlike a configuration in which, for example, the retraining images are stored on the cloud, the device that performs distance estimation does not need to have a communication unit or the like. Therefore, the complexity of the device configuration of the device that performs distance estimation can be suppressed. In addition, long-distance communication is not required, and the speed of information exchange can be increased.

[0012] In this case, preferably, the image selection unit obtains a second similarity score, which is the degree of similarity between the captured image and the retraining image, and further determines whether the second similarity score satisfies a preset condition. If both the first and second similarity scores satisfy the preset conditions, the captured image is selected as the retraining image to be used for retraining the learning model. With this configuration, by using the second similarity score between the captured image and the retraining image as the selection criterion, captured images with extremely low similarity (outlier images) can be excluded from the retraining images. As a result, the accuracy of distance estimation can be further improved when retraining the learning model using the retraining images.

[0013] In a configuration in which the image selection unit acquires a second similarity and selects an image to be used as a retraining image for retraining the learning model if both the first and second similarity values ​​satisfy preset conditions, preferably, the image selection unit acquires a second similarity if the acquired first similarity value is within a predetermined first range, and selects the image to be used as a retraining image for retraining the learning model if the acquired second similarity value is within a predetermined second range. With this configuration, it is possible to easily select an image to be used as a retraining image that can suppress overfitting and prevent the learning model from being trained using outlier images. As a result, it is possible to easily select a retraining image that can improve the accuracy of distance estimation.

[0014] In this case, preferably, the learning unit retrains the learning model using the retraining images each time a predetermined number of retraining images have been stored in the memory unit since the last retraining. With this configuration, the increase in processing load associated with retraining can be easily suppressed compared to a configuration in which the learning model is retrained each time an image is captured. Furthermore, when the learning model is retrained each time an image is captured, the objects depicted in the retraining images do not change much (the degree of change is small). Therefore, in a configuration in which the learning model is retrained each time an image is captured, the degree of improvement in distance estimation accuracy due to retraining the learning model is small. Thus, by configuring as described above, the learning model is retrained each time a predetermined number of retraining images have been stored in the memory unit since the last retraining, so the decrease in the degree of improvement in distance estimation accuracy when retraining the learning model can be suppressed.

[0015] In a configuration in which an captured image is selected as a retraining image to be used for retraining the learning model when the first and second similarity values ​​described above satisfy pre-set conditions, preferably, the learning unit retrains the learning model using the captured image, the retraining image, and the pre-training image. With this configuration, unlike a configuration in which the learning model is retrained using only the captured image and the retraining image, it becomes possible to use the pre-training image for retraining the learning model, thereby suppressing overfitting. As a result, it is possible to easily suppress the decrease in the accuracy of distance estimation by the learning model due to overfitting.

[0016] In a configuration comprising a control unit that controls the movement of the moving body described above, preferably, the object is a plurality of obstacles present in the movement area of ​​the moving body, and the moving body has a moving body main body that is controlled by the control unit to move while avoiding the plurality of obstacles, and a work unit provided in the moving body main body that performs a predetermined operation on a work object present in the movement area of ​​the moving body. With this configuration, even when a device that moves the moving body to the position of a work object while avoiding a plurality of obstacles and performs the operation by the work unit is placed in an unknown environment, the accuracy of estimating the distance between the plurality of obstacles and the moving body can be improved. As a result, even when a device that performs a predetermined operation by a work unit provided in the moving body main body is placed in an unknown environment, the accuracy of estimating the distance between the plurality of obstacles and the moving body can be improved while suppressing an increase in the burden on the user. Furthermore, since it is possible to improve the accuracy of estimating the distance between the plurality of obstacles and the moving body, the work efficiency of the operation performed by the work unit after moving the moving body to the position of a work object while avoiding a plurality of obstacles can be improved.

[0017] In this case, preferably, the imaging unit includes a first imaging unit provided on the mobile body and a second imaging unit provided on the work unit. With this configuration, it is possible to retrain the learning model using the images captured by the first imaging unit, and also to retrain the learning model using the images captured by the second imaging unit. Therefore, by retraining the learning model using the images captured by the first imaging unit, the accuracy of estimating the distance from the mobile body to an obstacle can be improved. Furthermore, by retraining the learning model using the images captured by the second imaging unit, the accuracy of estimating the distance from the work unit to the work target can be improved. As a result, even when a device that performs a predetermined task using a work unit provided on the mobile body is placed in an unknown environment, it is easy to improve both the accuracy of estimating the distance between multiple obstacles and the mobile body, and the accuracy of estimating the distance from the work unit to the work target.

[0018] In the device that performs distance estimation using the first phase described above, preferably, the imaging unit is a stereo camera, and the distance estimation unit inputs the captured image, which includes three-dimensional information captured by the stereo camera, into a learning model to estimate the distance to the object. With this configuration, for example, compared to a configuration in which the imaging unit includes only one monocular camera, it becomes possible to use the difference information between the two cameras in training the learning model, thereby improving the accuracy of distance estimation to the object.

[0019] The method for estimating distance in the second aspect of this invention comprises the steps of: capturing an image of an object and acquiring an image; inputting the acquired image into a learning model and estimating the distance to the object; selecting images for retraining the learning model from the captured images based on pre-set conditions; and retraining the learning model by performing machine learning based on the selected images for retraining.

[0020] The method for performing distance estimation according to the second aspect of this invention comprises, as described above, the steps of selecting images for retraining the learning model from captured images based on pre-set conditions, and retraining the learning model by performing machine learning based on the selected images for retraining. This provides a method for performing distance estimation that, similar to the apparatus for performing distance estimation according to the first aspect, can improve the accuracy of distance estimation while suppressing an increase in the user's burden when estimating distance from images captured in an unknown environment other than the environment in which the images used for training to generate the learning model were captured.

[0021] In the method for estimating distance using the second phase described above, preferably, in the step of acquiring an image, an image is acquired while the moving body is moving by an imaging unit provided on the moving body; in the step of selecting a retraining image, a retraining image is selected from the image captured while the moving body is moving; and in the step of estimating the distance to the object, the distance between the moving body and the object is estimated based on the image while the moving body is moving. With this configuration, similar to the device for estimating distance using the first phase described above, it is possible to provide a method for estimating distance using a learning model that can improve the accuracy of distance estimation to the object by a learning model for a moving body moving in various environments (locations).

[0022] In this case, preferably, the method further includes a step of controlling the movement of the moving object based on the estimated distance between the moving object and the target object. With this configuration, it is possible to provide a method for distance estimation that can accurately control the movement of a moving object, similar to the device that performs distance estimation using the first phase described above.

[0023] In the method for estimating distance using the second phase described above, preferably, the method further includes the steps of obtaining a first similarity, which is the degree of similarity between the captured image and a pre-training image used for training before the retraining of the learning model, and obtaining a second similarity, which is the degree of similarity between the captured image and the retraining image. In the step of selecting the retraining image, it is determined whether the first and second similarities satisfy a predetermined condition, and if they do, the captured image is selected as the retraining image to be used for retraining the learning model. With this configuration, similar to the apparatus for estimating distance using the first phase described above, it is possible to provide a method for estimating distance that can easily improve the accuracy of distance estimation when retraining the learning model using the retraining image.

[0024] In this case, preferably, if the acquired first similarity is within a predetermined first range, the second similarity is acquired in the step of acquiring the second similarity, and if the acquired second similarity is within a predetermined second range, the captured image is selected as the retraining image to be used for training the learning model in the step of selecting the retraining image. With this configuration, it is possible to provide a distance estimation method that allows for easy selection of a retraining image capable of improving the accuracy of distance estimation, similar to the apparatus that performs distance estimation using the first surface described above.

[0025] In a configuration in which a second similarity is obtained when the first similarity is within a predetermined first range, and an acquired image is selected as a retraining image when the acquired second similarity is within a predetermined second range, preferably, the step of retraining the learning model is performed each time a predetermined number of retraining images have been selected since the last training. With this configuration, it is possible to provide a distance estimation method that can easily suppress the increase in processing load associated with retraining compared to a configuration in which the learning model is retrained each time an image is captured, similar to the device that performs distance estimation using the first surface. Furthermore, similar to the device that performs distance estimation using the first surface, it is possible to provide a distance estimation method that can suppress the decrease in the degree of improvement in distance estimation accuracy when retraining the learning model.

[0026] In a configuration in which an captured image is selected as a retraining image to be used for retraining the learning model when the first and second similarity values ​​satisfy the conditions set in advance, preferably, in the step of retraining the learning model, the learning model is retrained using the captured image, the retraining image, and the pre-training image. With this configuration, it is possible to provide a distance estimation method that can easily suppress the decrease in the accuracy of distance estimation by the learning model due to overfitting, similar to the apparatus that performs distance estimation using the first surface described above.

[0027] According to the present invention, when estimating distance from images taken in an unknown environment other than the environment in which the images used for training to generate a learning model were taken, it is possible to provide a device and a method for performing distance estimation that can improve the accuracy of distance estimation while suppressing an increase in the burden on the user.

[0028] This is a block diagram showing the configuration of a device that performs distance estimation according to the first embodiment. This is a schematic diagram showing an example of a moving object in a device that performs distance estimation according to the first embodiment. This is a schematic diagram for explaining the difference between an object in an image captured by an imaging unit and two images according to the first embodiment. This is a schematic diagram for explaining the parallax acquired by the distance estimation unit according to the first embodiment. This is a schematic diagram for explaining the configuration in which a distance measurement unit estimates the distance to an object according to the first embodiment. This is a schematic diagram for explaining the configuration in which a learning unit according to the first embodiment acquires the error between a first image and a second image. This is a schematic diagram for explaining the configuration in which a learning unit according to the first embodiment retrains a learning model using a parallax map and an error. This is a schematic diagram for explaining the configuration in which an image selection unit according to the first embodiment selects images for retraining. This is a schematic diagram for explaining the configuration in which a distance estimation unit estimates distance and an image selection unit selects reconstructed images according to the first embodiment. This is a flowchart showing the distance estimation process in a device that performs distance estimation according to the first embodiment. This is a flowchart showing the process in which an image selection unit according to the first embodiment selects images for retraining. This is a block diagram showing the configuration of a distance estimation device according to the second embodiment. This is a schematic diagram showing an example of a moving object in a distance estimation device according to the second embodiment.

[0029] The following describes embodiments of the present invention based on the drawings.

[0030] [First Embodiment] Referring to Figures 1 to 11, the configuration of the distance estimation device 100 according to the first embodiment of the present invention, and the configuration by which the distance estimation device 100 estimates the distance Dt to the object 90 will be described.

[0031] (Configuration of the distance estimation device) As shown in Figure 1, the distance estimation device 100 comprises an imaging unit 1, a control unit 2, a moving body 3, and a storage unit 4.

[0032] The imaging unit 1 is configured to capture images of the object 90 (see Figure 2). In the first embodiment, the imaging unit 1 is a stereo camera.

[0033] The control unit 2 is configured to control each part of the device 100 that estimates distance. The control unit 2 includes a distance estimation unit 20, an image selection unit 21, a learning unit 22, and a movement control unit 23 as functional blocks. The movement control unit 23 is an example of the "control unit" in the claims. The control unit 2 includes a CPU (Central Processing Unit), a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) configured for image processing, and a memory having a ROM (Read Only Memory) and a RAM (Random Access Memory), etc.

[0034] The control unit 2 functions as the distance estimation unit 20, the image selection unit 21, the learning unit 22, and the movement control unit 23 by executing various programs stored in the storage unit 4.

[0035] The distance estimation unit 20 estimates the distance Dt to the object 90.

[0036] The image selection unit 21 selects a relearning image 42 for use in relearning the learning model 40 from the captured image 41 captured by the imaging unit 1. The learning model 40 is, for example, a neural network. The learning model 40 is generated by learning to output a disparity map 45 (see FIG. 5) showing the disparity d (see FIG. 4) between two images (pre-learning images 43).

[0037] The learning unit 22 relearns the learning model 40 using the relearning image 42.

[0038] The movement control unit 23 controls the movement operation of the moving body 3.

[0039] Details of each configuration of the distance estimation unit 20, the image selection unit 21, the learning unit 22, and the movement control unit 23 will be described later.

[0040] The moving body 3 has a moving body main body 30. Details of the moving body 3 will be described later.

[0041] The memory unit 4 stores the learning model 40, the captured image 41, the relearning image 42, and the pre-training image 43 used for learning prior to the relearning of the learning model 40. The memory unit 4 also stores the first range 50 and the second range 51 used for selecting the relearning image 42, which will be described later. The memory unit 4 also stores various programs (not shown) executed by the control unit 2. The memory unit 4 is a non-volatile storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive).

[0042] (Arrangement of the mobile body and imaging unit) As shown in Figure 2, the mobile body 3 is, for example, a motorcycle (saddle-type vehicle). The mobile body main unit 30 includes a body 30a and wheels 30b.

[0043] As shown in Figure 2, the moving body 3 is provided with an imaging unit 1. Specifically, the imaging unit 1 is provided on the body 30a. The imaging unit 1 (see Figure 1) is configured to capture an image 41 (see Figure 1) while the moving body 3 is moving. The imaging unit 1 is configured to capture the image 41 as a moving image at a predetermined frame rate. The predetermined frame rate is, for example, 30 fps (frames per second).

[0044] The distance estimation unit 20 (see Figure 1) estimates the distance Dt between the moving body 3 and the object 90 based on the captured image 41 captured by the imaging unit 1. In the first embodiment, the distance estimation unit 20 estimates the distance between the moving body 3 and the object 90 while the moving body 3 is in motion. The object 90 includes, for example, obstacles, other motorcycles, other automobiles, pedestrians, etc.

[0045] In the first embodiment, the mobile body control unit 23 (see Figure 1) controls the movement of the mobile body 3 based on the distance Dt between the mobile body 3 and the object 90 estimated by the distance estimation unit 20. Based on the distance Dt, the mobile body control unit 23 performs control to decelerate the mobile body 3, for example. In other words, in the first embodiment, the distance Dt is used for collision avoidance of the mobile body 3.

[0046] (Estimation of distance to object) Next, with reference to Figures 3 to 5, the configuration in which the distance estimation unit 20 estimates the distance Dt to the object 90 will be described.

[0047] As shown in Figure 3, the imaging unit 1 includes a first optical system 1a and a second optical system 1b. The first optical system 1a images the object 90 projected at the position where a straight line 50a connecting the focal point 1c and the object 90 intersects with the projection plane 1d. The second optical system 1b images the object 90 projected at the position where a straight line 50b connecting the focal point 1e and the object 90 intersects with the projection plane 1f.

[0048] The first captured image 41a is an image captured by the first optical system a located on one side (left side) of the stereo camera. The second captured image 41b is an image captured by the second optical system 1b located on the other side (right side) of the stereo camera.

[0049] From the first image 41a and the second image 41b captured by the stereo camera, the distance Dt from the imaging unit 1 to the object 90 can be estimated as shown in equation (1) below. Note that the distance from the imaging unit 1 to the object 90 is the distance from the focal point 1c to the object 90 (from the focal point 1e to the object 90). Here, Df is the distance between focal point 1c and focal point 1e. Also, f is the distance between focal point 1c and projection plane 1d (between focal point 1e and projection plane 1f). Also, d is the parallax between the first captured image 41a and the second captured image 41b.

[0050] As shown in Figure 4, the parallax d is the difference in the position where the same part of the object 90 is captured in the superimposed image 44, which is obtained by superimposing the first captured image 41a and the second captured image 41b. Specifically, as shown by the dashed line in Figure 4, when the focal point 1c of the first optical system 1a is moved to the position of the focal point 1e of the second optical system 1b, the parallax d is the difference between the position where the line 50c connecting the object 90 and the focal point 1c intersects with the projection plane 1d and the position where the line 50b intersects with the projection plane 1d. The distance f between the focal point 1c and the projection plane 1d, and the distance between the focal point 1c and the focal point 1e are known values. Therefore, if the parallax d between the first captured image 41a and the second captured image 41b can be obtained, the distance Dt from the imaging unit 1 to the object 90 can be estimated.

[0051] Therefore, in the first embodiment, as shown in Figure 5, the distance estimation unit 20 inputs the captured image 41 captured by the imaging unit 1 into the learning model 40 (see Figure 1) and estimates the distance Dt to the object 90. Specifically, the distance estimation unit 20 inputs the captured image 41 (first captured image 41a and second captured image 41b) containing three-dimensional information captured by the stereo camera into the learning model 40 and estimates the distance Dt to the object 90. More specifically, the distance estimation unit 20 inputs the first captured image 41a and the second captured image 41b into the learning model 40 and obtains a disparity map 45. Then, the distance estimation unit 20 obtains the disparity d from the disparity map 45. After that, the distance estimation unit 20 estimates the distance Dt from the imaging unit 1 to the object 90 based on the obtained disparity d and the above equation (1). The parallax map 45 is an image obtained by converting the size of the parallax d between the first captured image 41a and the second captured image 41b into pixel values. In the parallax map 45, areas with a larger parallax d appear brighter.

[0052] The distance estimation unit 20 outputs the estimated distance Dt to the mobile body control unit 23.

[0053] Here, the learning model 40 is generated by being trained using pre-training images 43 (see Figure 1). The environment (location) in which the pre-training images 43 were taken and the environment (location) in which the captured images 41 were taken may be different from each other. In this case, the difference between the pre-training images 43 used to generate the learning model 40 and the captured images 41 becomes large, so the accuracy of the distance Dt estimation by the learning model 40 decreases.

[0054] (Retraining the Learning Model) In the first embodiment, the learning unit 22 (see Figure 1) retrains the learning model 40 by performing machine learning based on the selected retraining images 42. The learning unit 22 retrains the learning model 40 using the captured image 41, the retraining image 42, and the pre-training image 43. In the first embodiment, the learning unit 22 retrains the learning model 40 by self-supervised learning using the captured image 41, the retraining image 42, and the pre-training image 43.

[0055] As shown in Figure 6, the learning unit 22 obtains the error 46 between the first captured image 41a and the second captured image 41b. Specifically, the learning unit 22 uses the second captured image 41b and the disparity map 45 to obtain an image (warp image 47) that has been transformed so that the second captured image 41b corresponds to the first captured image 41a. Then, the learning unit 22 obtains the error 46 by subtracting the warp image 47 from the first captured image 41a.

[0056] Subsequently, as shown in Figure 7, the learning unit 22 adjusts the parameters of the learning model 40 using the disparity map 45 and the error 46 by backpropagation, thereby obtaining the retrained learning model 40a. The obtained retrained learning model 40a is stored in the memory unit 4. Backpropagation is a method of retraining a learning model by adjusting the parameters of each layer of the neural network—the output layer, the hidden layer, and the input layer—based on the output of the learning model and its error.

[0057] (Selection of Relearning Images) Here, when the learning model 40 is relearned using only the captured image 41, overfitting may occur, and the estimation accuracy of the distance Dt may decrease. Therefore, in the first embodiment, as shown in FIG. 8, the image selection unit 21 is configured to acquire the similarity S from the captured image 41 and select the relearning image 42 to be used for the relearning of the learning model 40. The image selection unit 21 selects the relearning image 42 from the captured images 41 captured during the movement of the moving body 3. Note that overfitting means that a learning model 40 that can acquire a learning model 40 suitable for the captured image 41 is constructed, but the estimation accuracy of the distance Dt decreases for the captured image 41 captured in an unknown environment.

[0058] In the first embodiment, the image selection unit 21 selects the relearning image 42 to be used for the relearning of the learning model 40 from the captured images 41 captured by the imaging unit 1 based on a preset condition. Specifically, the image selection unit 21 acquires a first similarity S1, which is the degree of similarity between the captured image 41 and the prelearning image 43, based on the following formula (2). Here, t c is the feature amount of the captured image 41. The feature amount is, for example, the activation amount or gradient of a specific layer of the neural network with respect to the captured image 41. Also, μ1 c is the feature amount of the prelearning image 43 and is represented by the following formula (3). Also, σ1 c 2 is the variance of the feature amount of the prelearning image 43 and is represented by the following formula (4). Here, T1 k,c is the activation amount or gradient of a specific layer of the neural network with respect to k prelearning images 43. Specifically, T1 k,c is obtained by reducing the dimension of the tensor of the vertical × horizontal × channel dimension in the prelearning image 43 by taking the expected value (average) of the vertical × horizontal corresponding to the space. Also, μ1 c is the average of the activation amounts or gradients of a specific layer of the neural network with respect to k prelearning images 43. Also, σ1 c 2This is the variance of the activation amount or gradient of a specific layer of the neural network for k pre-training images 43. In other words, the first similarity S1 is a value that indicates the similarity between the captured image 41 and the pre-training images 43, calculated by the normalized Euclidean distance. The smaller the first similarity S1, the more similar the captured image 41 and the pre-training images 43 are. In the first embodiment, the image selection unit 21 determines whether the first similarity S1 satisfies a preset condition, and if it does, selects the captured image 41 as a retraining image 42 to be used for retraining the learning model 40.

[0059] Furthermore, the image selection unit 21 obtains a second similarity score S2, which is the degree of similarity between the captured image 41 and the retraining image 42, based on the following formula (5). Here, μ2 c σ² is a feature quantity of the relearning image 42 already stored in the memory unit 4, and is expressed by the following equation (6). c 2 This is the variance of the features of the retraining image 42, and is expressed by equation (7) shown below. Here, T2 k,c This is the activation amount or gradient of a specific layer of the neural network for k retraining images 42. Specifically, T2 k,c This is obtained by reducing the dimensionality of the tensor with dimensions of height × width × channels in the retraining image 42 by taking the expected value (average) of height × width, which corresponds to the space. Also, μ2 c σ² is the average of the activation amount or gradient of a specific layer of the neural network for k retraining images 42. c 2 This is the variance of the activation amount or gradient of a particular layer of the neural network for k retraining images 42. In other words, the second similarity S2 is a value that indicates the similarity between the captured image 41 and the retraining images 42, calculated by the normalized Euclidean distance. A smaller second similarity S2 means that the captured image 41 and the retraining images 42 are more similar.

[0060] Here, even if the learning model 40 is retrained using captured images 41 that have a high similarity to the pre-training images 43, the degree of improvement in the accuracy of distance Dt estimation by the learning model 40 is small. Therefore, it is preferable to retrain the learning model 40 using captured images 41 that have a relatively low first similarity S1 with the pre-training images 43. However, if the first similarity S1 with the pre-training images 43 is extremely low, it is not suitable for retraining the learning model 40. The same can be said for the second similarity S2 between the retraining images 42 and the captured images 41.

[0061] Therefore, in the first embodiment, the image selection unit 21 further determines whether the second similarity S2 satisfies the preset conditions, and selects the captured image 41 as a retraining image 42 to be used for retraining the learning model 40 if both the first similarity S1 and the second similarity S2 satisfy the preset conditions. Specifically, the image selection unit 21 acquires the second similarity S2 if the acquired first similarity S1 is within a predetermined first range 50. The first range 50 is the threshold range of the first similarity S1.

[0062] The image selection unit 21 then selects the captured image 41 as a retraining image 42 to be used for retraining the learning model 40 if the acquired second similarity S2 is within a predetermined second range 51. That is, the image selection unit 21 selects the captured image 41 as a retraining image 42 if the first similarity S1 is within a predetermined first range 50 and the second similarity S2 is within a predetermined second range 51. The image selection unit 21 does not select the captured image 41 as a retraining image 42 if the first similarity S1 is outside the first range 50 or the second similarity S2 is outside the second range 51. The selected retraining image 42 is stored in the storage unit 4 in association with the disparity map 45. The second range 51 is the threshold range of the second similarity S2.

[0063] Figure 9 shows the flow of distance estimation and image selection processing when the imaging unit 1 captures images 41 at a predetermined frame rate. Figure 9 shows the case where 2n images 41 are captured, from frame No. 1 to frame No. n.

[0064] As shown in Figure 9, the distance estimation unit 20 estimates the distance Dt each time a frame of the captured image 41 is captured. In the example shown in Figure 9, the distance estimation unit 20 estimates the distance Dt from each of the captured images 41 each time the first captured image 41a and the second captured image 41b, captured image 41c and captured image 41d, captured image 41e and captured image 41f, captured image 41g and captured image 41h, and captured image 41i and captured image 41j are captured.

[0065] The image selection unit 21 then selects a retraining image 42 each time a frame of the captured image 41 is captured. In the example shown in Figure 9, each time the first captured image 41a and the second captured image 41b, captured image 41c and captured image 41d, captured image 41e and captured image 41f, captured image 41g and captured image 41h, and captured image 41i and captured image 41j are captured, the image selection unit 21 obtains a similarity S (first similarity S1 and second similarity S2) from each captured image 41, and selects a retraining image 42 based on the obtained similarity S. In Figure 9, the captured images 41 other than captured images 41c and captured image 41d are selected as retraining images 42a to 42g.

[0066] Furthermore, when a predetermined number of relearning images 42 are selected and stored in the storage unit 4, the learning unit 22 (see Figure 1) retrains the learning model 40 using the relearning images 42. Specifically, each time a predetermined number of relearning images 42 are stored in the storage unit 4 since the last relearning, the learning unit 22 retrains the learning model 40 using the relearning images 42.

[0067] (Distance Estimation Process) Next, referring to Figure 10, the process by which the distance estimation device 100 estimates the distance Dt to the object 90 will be described.

[0068] In step S10, the imaging unit 1 images the object 90 and acquires an image 41. In the first embodiment, in step S10, the imaging unit 1 provided on the moving body 3 acquires the image 41 while the moving body 3 is moving.

[0069] Next, in step S11, the distance estimation unit 20 inputs the acquired image 41 into the learning model 40 and estimates the distance Dt to the object 90. In the first embodiment, in step S11, the distance estimation unit 20 estimates the distance between the moving body 3 and the object 90 based on the image 41 while the moving body 3 is moving.

[0070] Next, in step S12, the mobile body control unit 23 controls the movement of the mobile body 3 based on the estimated distance Dt between the mobile body 3 and the object 90.

[0071] Furthermore, the processing of step S13 is performed in parallel with the processing of steps S11 and S12. That is, in step S13, the image selection unit 21 selects a retraining image 42 to be used for retraining the learning model 40 from the captured images 41 based on preset conditions. In the first embodiment, in step S13, the image selection unit 21 selects a retraining image 42 from the captured images 41 taken while the moving body 3 is moving.

[0072] Next, the learning unit 22 determines whether the number of relearning images 42 stored in the storage unit 4 has reached a predetermined number. If the number of relearning images 42 stored in the storage unit 4 has reached a predetermined number, the process proceeds to step S15. If the number of relearning images 42 stored in the storage unit 4 has not reached a predetermined number, the process proceeds to step S10.

[0073] If the process proceeds from step S14 to step S15, in step S15, the learning unit 22 retrains the learning model 40 by performing machine learning based on the selected retraining images 42. That is, step S15 is executed each time a predetermined number of retraining images 42 have been selected since the last training. In step S15, the learning unit 22 retrains the learning model 40 using the captured image 41, the retraining images 42, and the pre-training images 43.

[0074] Next, in step S16, the mobile unit control 23 determines whether or not to stop the mobile unit 3. For example, if there is an input to stop the movement of the mobile unit 3, the mobile unit control 23 determines to stop the mobile unit 3. If the mobile unit 3 is to be stopped, the process ends. If the mobile unit 3 is not to be stopped, the process proceeds to step S10.

[0075] (Process for selecting images for retraining) Next, with reference to Figure 11, the process by which the image selection unit 21 selects the images 42 for retraining will be explained.

[0076] In step S13a, the image selection unit 21 obtains a first similarity S1 between the captured image 41 and the pre-training image 43 used for training before the retraining of the learning model 40.

[0077] Next, in step S13b, the image selection unit 21 determines whether the first similarity S1 is within a predetermined first range 50. If the first similarity S1 is within the first range 50, the process proceeds to step S13c. If the first similarity S1 is not within the first range 50, the process proceeds to step S13f.

[0078] If the process proceeds from step S13b to step S13c, in step S13c, the image selection unit 21 obtains the second similarity S2 between the captured image 41 and the retraining image 42. That is, the image selection unit 21 obtains the second similarity S2 in step S13c if the obtained first similarity S1 is within a predetermined first range 50.

[0079] Next, in step S13d, the image selection unit 21 determines whether the second similarity S2 is within a predetermined second range 51. If the second similarity S2 is within the second range 51, the process proceeds to step S13e. If the second similarity S2 is not within the second range 51, the process proceeds to step S13f.

[0080] If the process proceeds from step S13d to step S13e, in step S13e, the image selection unit 21 selects the captured image 41 as the retraining image 42 to be used for retraining the learning model 40. Specifically, the image selection unit 21 selects the captured image 41 as the retraining image 42 if the second similarity S2 obtained in step S13c is within a predetermined second range 51. That is, in step 13e, the image selection unit 21 determines whether the first similarity S1 and the second similarity S2 satisfy the pre-set conditions, and if they do, selects the captured image 41 as the retraining image 42. After that, the process ends.

[0081] Furthermore, if the process proceeds from step S13b or step S13d to step S13f, in step S13f, the image selection unit 21 performs a process of not selecting the captured image 41 as the retraining image 42. Specifically, if the first similarity S1 is not within the first range 50, or if the first similarity S1 is within the first range 50 but the second similarity S2 is not within the second range 51, the image selection unit 21 does not select the captured image 41 as the retraining image 42 and does not store it in the storage unit 4.

[0082] (Effects of the First Embodiment) In the first embodiment, the following effects can be obtained.

[0083] In the first embodiment, as described above, the system includes an image selection unit 21 that selects a retraining image 42 to be used for retraining the learning model 40 from captured images 41 captured by the imaging unit 1 based on preset conditions, and a learning unit 22 that retrains the learning model 40 by performing machine learning based on the selected retraining image 42. As a result, the retraining image 42 to be used for retraining the learning model 40 is selected from captured images 41 captured by the imaging unit 1, so the learning model 40 can be retrained without the user having to prepare the retraining image 42. Therefore, an increase in the user's burden can be suppressed. The learning unit 22 also retrains the learning model 40 using the retraining image 42 selected by the image selection unit 21 from captured images 41 captured by the imaging unit 1 based on preset conditions. Therefore, even if the device is placed in an unknown environment different from the environment in which the images used to generate the learning model 40 were taken, the learning model 40 can be retrained using retraining images 42 selected based on pre-set conditions, which allow for the selection of images from captured images 41 taken in that unknown environment that can improve the accuracy of distance Dt estimation. As a result, even if the device is placed in an unknown environment, the accuracy of distance Dt estimation can be improved by retraining the learning model 40 using retraining images 42.

[0084] Furthermore, in the first embodiment, as described above, the imaging unit 1 is configured to capture an image 41 while the moving body 3 is moving, the image selection unit 21 selects a retraining image 42 from the image 41 captured while the moving body 3 is moving, and the distance estimation unit 20 estimates the distance between the moving body 3 and the object 90 while the moving body 3 is moving. As a result, the learning model 40 can be retrained using the retraining image 42 selected from the image 41 captured while the moving body 3 is moving in various environments (locations). Consequently, the accuracy of the distance Dt to the object 90 estimated by the learning model 40 can be improved when the moving body 3 is moving in various environments (locations).

[0085] Furthermore, in the first embodiment, as described above, a mobile body control unit 23 is further provided that controls the movement of the mobile body 3 based on the distance Dt between the mobile body 3 and the object 90 estimated by the distance estimation unit 20. This makes it possible to estimate the distance Dt with high accuracy, and therefore to control the movement of the mobile body 3 with high accuracy.

[0086] Furthermore, in the first embodiment, as described above, the image selection unit 21 acquires a first similarity S1, which is the degree of similarity between the captured image 41 and the pre-training image 43, determines whether the first similarity S1 satisfies a preset condition, and if it does, selects the captured image 41 as the retraining image 42 to be used for retraining the learning model 40. By using the first similarity S1 between the captured image 41 and the pre-training image 43 as the selection criterion for the retraining image 42, excessive optimization (overfitting) to the environment in which the captured image 41 was taken can be suppressed. As a result, when retraining the learning model 40 using the retraining image 42, the estimation accuracy of the distance Dt can be easily improved. In addition, since the distance estimation device 100 itself is equipped with a storage unit 4 for storing the retraining image 42, unlike a configuration in which the retraining image 42 is stored on the cloud, for example, the distance estimation device 100 does not need to be equipped with a communication unit or the like. Therefore, the complexity of the device configuration of the distance estimation device 100 can be suppressed. Furthermore, it eliminates the need for long-distance communication, allowing for faster information exchange.

[0087] Furthermore, in the first embodiment, as described above, the image selection unit 21 obtains a second similarity S2, which is the degree of similarity between the captured image 41 and the retraining image 42. If both the first similarity S1 and the second similarity S2 satisfy the conditions set in advance, the captured image 41 is selected as the retraining image 42 to be used for retraining the learning model 40. By using the second similarity S2 between the captured image 41 and the retraining image 42 as a selection criterion, captured images 41 with extremely low similarity (outlier images) can be excluded from the retraining images 42. As a result, when retraining the learning model 40 using the retraining images 42, the estimation accuracy of the distance Dt can be further improved.

[0088] Furthermore, in the first embodiment, as described above, the image selection unit 21 acquires a second similarity S2 when the acquired first similarity S1 is within a predetermined first range 50, and selects the captured image 41 as a retraining image 42 to be used for retraining the learning model 40 when the acquired second similarity S2 is within a predetermined second range 51. This makes it possible to easily select a captured image 41 as a retraining image 42 that can suppress overfitting and prevent the learning model 40 from being trained using outlier images. As a result, it is possible to easily select a retraining image 42 that can improve the estimation accuracy of distance Dt.

[0089] Furthermore, in the first embodiment, as described above, the learning unit 22 retrains the learning model 40 using the retraining images 42 each time a predetermined number of retraining images 42 are stored in the storage unit 4 since the last retraining. This makes it easy to suppress the increase in processing load associated with retraining compared to a configuration in which the learning model 40 is retrained each time an image 41 is captured. Also, when the learning model 40 is retrained each time an image 41 is captured, the object 90 in the retraining images 42 does not change much (the degree of change is small). Therefore, in a configuration in which the learning model 40 is retrained each time an image 41 is captured, the degree of improvement in the distance Dt estimation accuracy due to retraining the learning model 40 becomes small. Therefore, by configuring it as described above, the learning model 40 is retrained each time a predetermined number of retraining images 42 are stored in the storage unit 4 since the last retraining, so it is possible to suppress the decrease in the degree of improvement in the distance Dt estimation accuracy when retraining the learning model 40.

[0090] Furthermore, in the first embodiment, as described above, the learning unit 22 retrains the learning model 40 using the captured image 41, the retraining image 42, and the pre-training image 43. This makes it possible to use the pre-training image 43 for retraining the learning model 40, unlike the configuration in which the learning model 40 is retrained using only the captured image 41 and the retraining image 42, thus suppressing overfitting. As a result, it is possible to easily suppress the decrease in the accuracy of distance Dt estimation by the learning model 40 due to overfitting.

[0091] Furthermore, in the first embodiment, as described above, the distance estimation unit 20 inputs the captured image 41, which includes three-dimensional information captured by the stereo camera, into the learning model 40 to estimate the distance Dt to the object 90. This makes it possible to use the difference information between the two cameras in training the learning model 40, compared to a configuration in which the imaging unit 1 includes only one monocular camera, for example, thus improving the accuracy of estimating the distance Dt to the object 90.

[0092] Furthermore, in the first embodiment, as described above, the method includes the steps of selecting a retraining image 42 to be used for retraining the learning model 40 from the captured image 41 based on pre-set conditions, and retraining the learning model 40 by performing machine learning based on the selected retraining image 42. This makes it possible to provide a method for performing distance estimation that, similar to the distance estimation device 100 in the first embodiment, can improve the accuracy of distance estimation while suppressing an increase in the user's burden when estimating distance Dt from an image captured in an unknown environment other than the environment in which the images used for training to generate the learning model 40 were captured.

[0093] Furthermore, in the first embodiment, as described above, in the step of selecting the retraining image 42, the retraining image 42 is selected from the captured image 41 taken while the mobile body 3 is moving, and in the step of estimating the distance Dt to the object 90, the distance between the mobile body 3 and the object 90 is estimated based on the captured image 41 while the mobile body 3 is moving. This makes it possible to provide a distance estimation method that can improve the accuracy of the estimation of the distance Dt to the object 90 by the learning model 40 for a mobile body 3 moving in various environments (locations), similar to the distance estimation device 100 according to the first embodiment.

[0094] Furthermore, the first embodiment further includes a step of controlling the movement of the moving body 3 based on the estimated distance Dt between the moving body 3 and the object 90, as described above. This provides a method for performing distance estimation that enables accurate control of the movement of the moving body 3, similar to the distance estimation device 100 according to the first embodiment.

[0095] Furthermore, in the first embodiment, as described above, the method further includes the steps of obtaining a first similarity S1, which is the degree of similarity between the captured image 41 and a pre-training image 43 used for training before the retraining of the learning model 40, and obtaining a second similarity S2, which is the degree of similarity between the captured image 41 and the retraining image 42. In the step of selecting the retraining image 42, it is determined whether the first similarity S1 and the second similarity S2 satisfy a preset condition, and if they do, the captured image 41 is selected as the retraining image 42 to be used for retraining the learning model 40. This makes it possible to provide a method for performing distance estimation that can easily improve the accuracy of distance Dt estimation when retraining the learning model 40 using the retraining image 42, similar to the distance estimation device 100 according to the first embodiment.

[0096] Furthermore, in the first embodiment, as described above, if the acquired first similarity S1 is within a predetermined first range 50, the second similarity S2 is acquired in the step of acquiring the second similarity S2, and if the acquired second similarity S2 is within a predetermined second range 51, the captured image 41 is selected as the retraining image 42 to be used for training the learning model 40 in the step of selecting the retraining image 42. This makes it possible to provide a distance estimation method that makes it possible to easily select a retraining image 42 capable of improving the estimation accuracy of distance Dt, similar to the distance estimation device 100 according to the first embodiment.

[0097] Furthermore, in the first embodiment, as described above, the step of retraining the learning model 40 is performed each time a predetermined number of retraining images 42 have been selected since the last training. This makes it possible to provide a distance estimation method that can easily suppress the increase in processing load associated with retraining, compared to a configuration in which the learning model 40 is retrained each time an image 41 is captured, similar to the distance estimation device 100 in the first embodiment. Also, similar to the distance estimation device 100 in the first embodiment, it is possible to provide a distance estimation method that can suppress the degree of improvement in the accuracy of distance Dt estimation when retraining the learning model 40.

[0098] Furthermore, in the first embodiment, as described above, in the step of retraining the learning model 40, the learning model 40 is retrained using the captured image 41, the retraining image 42, and the pre-training image 43. This makes it possible to provide a distance estimation method that can easily suppress the decrease in the accuracy of distance Dt estimation by the learning model 40 due to overfitting, similar to the distance estimation device 100 according to the first embodiment.

[0099] [Second Embodiment] Next, a distance estimation device 200 according to the second embodiment will be described with reference to Figures 12 and 13. Components similar to those in the distance estimation device 100 according to the first embodiment are denoted by the same reference numerals, and detailed descriptions are omitted.

[0100] The distance estimation device 200 according to the second embodiment comprises an imaging unit 201, a control unit 202, a moving body 203, and a storage unit 4.

[0101] The imaging unit 201 includes a first imaging unit 10 and a second imaging unit 11. The first imaging unit 10 captures a first image 240a. The second imaging unit 11 captures a second image 240b. Each of the first imaging unit 10 and the second imaging unit 11 is a stereo camera.

[0102] The control unit 202 includes a distance estimation unit 202a, an image selection unit 202b, a learning unit 202c, and a mobile object control unit 202d.

[0103] The distance estimation unit 202a estimates the distance Dt2 (see Figure 13) from the first imaging unit 10 to the object 90 (see Figure 13) based on the first image 240a captured by the first imaging unit 10 and the first learning model 40b. The distance estimation unit 202a also estimates the distance Dt3 (see Figure 13) from the second imaging unit 201a to the object 90 based on the second image 240b captured by the second imaging unit 11 and the second learning model 40c. The configuration in which the distance estimation unit 202a estimates distances Dt2 and Dt3 is the same as the configuration in which the distance estimation unit 20 estimates distance Dt according to the first embodiment.

[0104] Furthermore, the first learning model 40b is generated by training it to output a disparity map 45 showing the disparity d of two images, similar to the learning model 40, using the first pre-training image 242a. Similarly, the second learning model 40c is generated by training it to output a disparity map 45 showing the disparity d of two images, similar to the learning model 40, using the second pre-training image 242b.

[0105] The image selection unit 202b selects a first retraining image 241a from the first captured image 240a to be used for retraining the first learning model 40b. The image selection unit 202b also selects a second retraining image 241b from the second captured image 240b to be used for retraining the second learning model 40c. Similar to the image selection unit 21 in the first embodiment described above, the image selection unit 202b acquires a similarity score S and uses the acquired similarity score S, the first range 50, and the second range 51 to select the first retraining image 241a. The image selection unit 202b also acquires a similarity score S and uses the acquired similarity score S, the third range 52, and the fourth range 53 to select a second retraining image 241b in the same configuration as the image selection unit 21. The third range 52 is the threshold range used to determine the similarity between the second captured image 240b and the second pre-training image 242b. The fourth range 53 is the threshold range used to determine the similarity between the second captured image 240b and the second retraining image 241b.

[0106] The learning unit 202c retrains the first learning model 40b using the first retraining image 241a. The learning unit 202c also retrains the second learning model 40c using the second retraining image 241b. The configuration in which the learning unit 202c retrains the first learning model 40b and the second learning model 40c is the same as in the first embodiment in which the learning unit 22 retrains the learning model 40.

[0107] The mobile unit control unit 202d controls the movement of the mobile unit body 31 and the work unit 32 provided by the mobile unit 203.

[0108] The mobile body 203 has a mobile body main section 31 and a work section 32. In this embodiment, the mobile body 203 is a fruit tree harvesting device configured to move autonomously by an AGV (Automatic Guided Vehicle) or AMR (Autonomous Mobile Robot).

[0109] As shown in Figure 13, the object 90 is a plurality of obstacles 90a present in the movement area of ​​the mobile body 3. The object 90 also includes the work target 90b of the work unit 32. The work target 90b is a fruit or the like. The mobile body 203 moves within the movement area while avoiding the plurality of obstacles 90a, and the work unit 32 performs tasks such as harvesting on the work target 90b present in this movement area.

[0110] Furthermore, the mobile body 31 includes a body 31a and wheels 31b. A storage compartment 33 for housing the work object 90b is mounted on the body 31a.

[0111] Furthermore, the mobile body control unit 202d controls the movement of the mobile body 31, such as acceleration, deceleration, and stopping, based on the distance Dt2 estimated by the distance estimation unit 202a. Specifically, the mobile body control unit 202d decelerates the mobile body 31 when the distance Dt2 between the mobile body 31 and the obstacle 90a falls below a preset threshold. The mobile body control unit 202d also accelerates the mobile body 31 when the distance Dt2 exceeds the threshold. In addition, the mobile body control unit 202d stops the mobile body 31 when the distance Dt2 falls below the threshold and reaches the limit value. As a result, the mobile body 31 is controlled by the mobile body control unit 202d to move while avoiding multiple obstacles 90a.

[0112] As shown in Figure 13, the first imaging unit 10 is provided on the mobile body 31. The first imaging unit 10 captures a first image 240a while the mobile body 31 is moving.

[0113] Furthermore, the work unit 32 is provided on the mobile body 31. The work unit 20b performs predetermined operations on the work object 90b that is within the movement range of the mobile body 3. The work unit 32 includes an arm 32a and a gripping 32b. The arm 32a is configured to be extendable and retractable by being controlled by the mobile body control unit 202d. The arm 32a is configured to be movable by being controlled by the mobile body control unit 202d. Furthermore, the gripping 32b is configured to grip the work object 90b by being controlled by the mobile body control unit 202d. Furthermore, the gripping 32b is provided on the end of the arm 32a opposite to the mobile body 31. The work unit 32 is controlled by the mobile body control unit 202d to harvest the work object 90b, such as fruit, and store it in the storage unit 33.

[0114] As shown in Figure 13, the second imaging unit 11 is provided on the work unit 32. Specifically, the second imaging unit 11 is provided near the gripping unit 32b at the end of the arm unit 32a opposite to the movable body unit 31. The second imaging unit 11 captures a second image 240b while the arm unit 32a is moving.

[0115] Furthermore, the other configurations of the distance estimation device 200 according to the second embodiment are the same as those of the distance estimation device 100 according to the first embodiment.

[0116] (Effects of the second embodiment) In the second embodiment, the following effects can be obtained.

[0117] In the second embodiment, as described above, the object 90 is a plurality of obstacles 90a present in the movement area of ​​the mobile body 203, and the mobile body 203 has a mobile body main body 31 controlled by a mobile body control unit 202d to move while avoiding the plurality of obstacles 90a, and a work unit 32 provided on the mobile body main body 31 that performs predetermined work on a work target 90b present in the movement area of ​​the mobile body 203. As a result, even when a device that moves the mobile body 203 to the position of the work target 90b while avoiding the plurality of obstacles 90a and performs work by the work unit 32 is placed in an unknown environment, the accuracy of estimating the distance Dt2 between the plurality of obstacles 90a and the mobile body 203 can be improved. Furthermore, since it is possible to improve the accuracy of estimating the distance Dt2 between multiple obstacles 90a and the mobile body 203, the work efficiency of the work performed by the work unit 32 can be improved by moving the mobile body 203 to the position of the work target 90b while avoiding multiple obstacles 90a.

[0118] Furthermore, in the second embodiment, as described above, the imaging unit 201 includes a first imaging unit 10 provided on the mobile body 31 and a second imaging unit 11 provided on the work unit 32. This makes it possible to retrain the first learning model 40b using the first image 240a captured by the first imaging unit 10, and to retrain the second learning model 40c using the second image 240b captured by the second imaging unit 11. Therefore, by retraining the first learning model 40b using the first image 240a captured by the first imaging unit 10, the estimation accuracy of the distance Dt2 from the mobile body 31 to the obstacle 90a can be improved. Also, by retraining the second learning model 40c using the second image 240b captured by the second imaging unit 11, the estimation accuracy of the distance Dt3 from the work unit 32 to the work target 90b can be improved. As a result, even when a device that performs a predetermined task using a work unit 32 provided on the mobile body 31 is placed in an unknown environment, it is easy to simultaneously improve the estimation accuracy of the distance Dt2 between multiple obstacles 90a and the mobile body 203, and improve the estimation accuracy of the distance Dt2 from the work unit 32 to the work target 90b.

[0119] Other effects of the second embodiment are the same as those of the first embodiment described above.

[0120] [Modifications] It should be noted that the embodiments disclosed herein are illustrative and not restrictive in all respects. The scope of the present invention is indicated by the claims rather than the description of the embodiments above, and further includes all modifications (modifications) within the meaning and scope equivalent to the claims.

[0121] For example, in the first and second embodiments described above, an example was shown in which the distance estimation device 100 (distance estimation device 200) includes a moving body 3 (moving body 203), but the present invention is not limited thereto. The distance estimation device does not have to include a moving body. Furthermore, if the distance estimation device does not include a moving body, the distance estimation device does not have to include a moving body control unit. However, if the distance estimation device does not include a moving body, the changes in the captured images (first captured image, second captured image) become less significant, and the degree of improvement in distance estimation accuracy due to retraining of the learning model (first learning model, second learning model) becomes smaller. Therefore, the present invention is preferably applied to a distance estimation device that includes a moving body.

[0122] Furthermore, in the first and second embodiments described above, an example was shown in which the image selection unit 21 (image selection unit 202b) selects a retraining image 42 (first retraining image 241a, second retraining image 241b) when the first similarity S1 between the captured image 41 (first captured image 240a, second captured image 240b) and the pretraining image 43 (first pretraining image 242a, second pretraining image 242b) and the second similarity S2 between the captured image 41 (first captured image 240a, second captured image 240b) and the retraining image 42 (first retraining image 241a, second retraining image 241b) satisfies predetermined conditions. However, the present invention is not limited thereto. The image selection unit may be configured to select a retraining image when either the first similarity or the second similarity satisfies predetermined conditions. Furthermore, the image selection unit may be configured to select images for retraining without using the first and second similarity metrics.

[0123] Furthermore, in the first and second embodiments described above, an example was shown in which the image selection unit 21 (image selection unit 202b) acquires a second similarity S2 when the first similarity S1 is within a predetermined first range 50, and selects the captured image 41 as a retraining image 42 when the second similarity S2 is within a predetermined second range 51. However, the present invention is not limited thereto. For example, the image selection unit may be configured to acquire a first similarity when the second similarity is within a second range, and select the captured image as a retraining image when the first similarity is within a first range.

[0124] Furthermore, while the first and second embodiments described above show an example of a configuration in which the image selection unit 21 (image selection unit 202b) uses normalized Euclidean distance when acquiring the first similarity S1 and the second similarity S2, the present invention is not limited thereto. For example, the image selection unit may be configured to acquire the first similarity and the second similarity based on Euclidean distance, Mahalanobis distance, or cosine distance, etc.

[0125] Furthermore, in the first and second embodiments described above, an example was shown in which the learning unit 22 (learning unit 202c) retrains the learning model 40 (first learning model 40b, second learning model 40c) each time a predetermined number of retraining images 42 (first retraining image 241a, second retraining image 241b) have been stored in the storage unit 4 since the last retraining. However, the present invention is not limited to this. For example, the learning unit may be configured to retrain the learning model 40 each time a retraining image is selected.

[0126] Furthermore, although the second embodiment described above shows an example in which the imaging unit 201 comprises a first imaging unit 10 and a second imaging unit 11, the present invention is not limited thereto. For example, the imaging unit does not need to be provided in the work unit, as long as it is provided in at least the main body of the mobile unit. In this case, the distance between the work unit and the work object may be estimated based on the image captured by the imaging unit provided in the main body of the mobile unit, or a distance sensor or the like may be provided in the work unit, and the distance between the work unit and the work object may be obtained based on the output value of the distance sensor.

[0127] Furthermore, while the first embodiment described above shows the mobile body 3 as a motorcycle (saddle-type vehicle) and the second embodiment described above shows the mobile body 203 as a fruit tree harvesting device, the present invention is not limited thereto. For example, the mobile body may be an automobile moving along a road, a ship moving on water, an industrial robot moving within a factory (such as an articulated robot and a feeder insertion / removal device), and a pesticide spraying device that sprays pesticides on objects such as tree leaves. If the mobile body is an articulated robot, the mobile body may move within the factory and perform tasks such as supplying workpieces for production continuation (consumable parts such as bolts and assembly chips) necessary for the production line equipment to continue production, or supplying workpieces for assembly to workers performing assembly work within the factory. Also, if the mobile body is a feeder insertion / removal device, the mobile body may move within the factory and perform tasks such as handing over a feeder that holds multiple parts to a component mounting machine that mounts parts onto a circuit board, and receiving the feeder from the component mounting machine (insertion / removal work). Furthermore, if the mobile unit is a pesticide spraying device, it may have a work unit that moves around the farm and performs the task of spraying pesticides on target objects such as tree leaves.

[0128] Furthermore, although the first and second embodiments described above show examples where the imaging unit 1 (first imaging unit 10, second imaging unit 11) is a stereo camera, the present invention is not limited to this. For example, the imaging unit may be a three-dimensional camera, multiple monocular cameras, etc. The imaging unit may be configured in any way as long as it is possible to obtain the distance from the captured image to the object.

[0129] Furthermore, while the first and second embodiments described above show an example of a configuration in which the distance estimation unit 20 (distance estimation unit 202a) estimates the distance Dt (distance Dt2, distance Dt3) and the image selection unit 21 selects the retraining images 42 (first retraining image 241a, second retraining image 241b) in parallel, the present invention is not limited thereto. The distance estimation by the distance estimation unit and the selection of retraining images by the image selection unit may be configured to be performed sequentially in a predetermined order. In addition, the distance estimation by the distance estimation unit and the selection of retraining images by the image selection unit may be performed by different control units (control devices).

[0130] Furthermore, in the first and second embodiments described above, an example was shown in which the learning unit 22 (learning unit 202c) acquires an image (warp image 47) by converting the second captured image 41b to correspond to the first captured image 41a from the second captured image 41b and the disparity map 45 when acquiring the error 46, but the present invention is not limited thereto. The learning unit may be configured to acquire a warp image by converting the first captured image to correspond to the second captured image from the first captured image and the disparity map.

[0131] Furthermore, in the first and second embodiments described above, for the sake of explanation, the control processing of the control unit 2 (control unit 202) was explained using a flow-driven flowchart that processes sequentially according to the processing flow, but the present invention is not limited thereto. In the present invention, the control processing of the control unit may be performed by event-driven processing, which executes processing on an event-by-event basis. In this case, it may be performed as a completely event-driven system, or a combination of event-driven and flow-driven systems may be used.

[0132] 1, 201 Imaging unit 3, 203 Mobile unit 4 Memory unit 10 First imaging unit 11 Second imaging unit 20, 202a Distance estimation unit 21, 202b Image selection unit 22, 202c Learning unit 23, 202d Mobile unit control unit (control unit) 30, 31 Mobile unit main body 32 Working unit 40 Learning model 41, 41c-41j Captured images 41a, 240a First captured image (captured image) 41b, 240b Second captured image (captured image) 42, 42a-42g Images for retraining 43 Images for pretraining 50 First range 51 Second range 90 Target object 90a Multiple obstacles 90b Work target 100, 200 Device for distance estimation 241a First retraining image (retraining image) 241b Second retraining image (retraining image) 242a First pretraining image (pretraining image) 242b Second pretraining image (pretraining image) Dt, Dt2, Dt3 Distance (distance to object) S1 First similarity S2 Second similarity

Claims

1. A distance estimation device comprising: an imaging unit for imaging an object; a distance estimation unit for inputting the image captured by the imaging unit into a learning model and estimating the distance to the object; an image selection unit for selecting images for retraining the learning model from the images captured by the imaging unit based on pre-set conditions; and a learning unit for retraining the learning model by performing machine learning based on the selected images for retraining.

2. The distance estimation apparatus according to claim 1, further comprising a moving body provided with the imaging unit, wherein the imaging unit is configured to capture the image while the moving body is moving, the image selection unit selects the retraining image from the image captured while the moving body is moving, and the distance estimation unit estimates the distance between the moving body and the object while the moving body is moving.

3. The distance estimation apparatus according to claim 2, further comprising a control unit that controls the movement of the moving body based on the distance between the moving body and the object estimated by the distance estimation unit.

4. The distance estimation apparatus according to claim 1, further comprising a storage unit for storing pre-training images used for training prior to the retraining of the learning model, wherein the image selection unit obtains a first similarity, which is the degree of similarity between the captured image and the pre-training images, determines whether the first similarity satisfies the preset conditions, and if it satisfies them, selects the captured image as the retraining image to be used for retraining the learning model.

5. The distance estimation apparatus according to claim 4, wherein the image selection unit obtains a second similarity, which is the degree of similarity between the captured image and the retraining image, and further determines whether the second similarity satisfies the preset conditions, and if both the first similarity and the second similarity satisfy the preset conditions, the captured image is selected as the retraining image to be used for retraining the learning model.

6. The distance estimation apparatus according to claim 5, wherein the image selection unit acquires a second similarity when the acquired first similarity is within a predetermined first range, and selects the captured image as the retraining image to be used for retraining the learning model when the acquired second similarity is within a predetermined second range.

7. The distance estimation apparatus according to claim 6, wherein the learning unit retrains the learning model using the retraining images each time a predetermined number of the retraining images have been stored in the storage unit since the last retraining.

8. The distance estimation apparatus according to claim 5, wherein the learning unit retrains the learning model using the captured image, the retraining image, and the pre-training image.

9. The distance estimation device according to claim 3, wherein the object is a plurality of obstacles present in the movement area of ​​the moving body, the moving body comprises a moving body main body portion controlled by the control unit to move while avoiding the plurality of obstacles, and a work portion provided in the moving body main body portion for performing a predetermined work on a work object present in the movement area of ​​the moving body.

10. The distance estimation apparatus according to claim 9, wherein the imaging unit includes a first imaging unit provided on the main body of the mobile unit and a second imaging unit provided on the work unit.

11. The distance estimation apparatus according to claim 1, wherein the imaging unit is a stereo camera, and the distance estimation unit inputs the image captured by the stereo camera, which includes three-dimensional information, into the learning model to estimate the distance to the object.

12. A method for estimating distance, comprising the steps of: capturing an image of an object and acquiring an image; inputting the acquired image into a learning model and estimating the distance to the object; selecting a retraining image to be used for retraining the learning model from the image based on pre-set conditions; and retraining the learning model by performing machine learning based on the selected retraining image.

13. A method for estimating distance according to claim 12, comprising the steps of acquiring the captured image, acquiring the captured image while the moving body is moving using an imaging unit provided on the moving body, selecting the retraining image, selecting the retraining image from the captured image acquired while the moving body is moving, and estimating the distance to the object, estimating the distance between the moving body and the object based on the captured image while the moving body is moving.

14. A method for distance estimation according to claim 13, further comprising the step of controlling the movement of the moving body based on the estimated distance between the moving body and the object.

15. A method for performing distance estimation according to claim 12, further comprising the steps of: obtaining a first similarity, which is the degree of similarity between the captured image and a pre-training image used for training before the retraining of the learning model; and obtaining a second similarity, which is the degree of similarity between the captured image and the retraining image, wherein in the step of selecting the retraining image, it is determined whether the first similarity and the second similarity satisfy the pre-set conditions, and if they do, the captured image is selected as the retraining image to be used for retraining the learning model.

16. A method for performing distance estimation according to claim 15, wherein, if the acquired first similarity is within a predetermined first range, the second similarity is acquired in the step of acquiring the second similarity, and if the acquired second similarity is within a predetermined second range, the captured image is selected as the retraining image to be used for training the learning model in the step of selecting the retraining image.

17. The method for performing distance estimation according to claim 15, wherein the step of retraining the learning model is performed each time a predetermined number of images for retraining have been selected since the last training.

18. A method for performing distance estimation according to claim 15, wherein in the step of retraining the learning model, the learning model is retrained using the captured image, the retraining image, and the pre-training image.

Citation Information

Patent Citations

  • Object measurement method, measuring device, program, and computer-readable recording medium

    JP2020197983A

  • Learning method, program, and image processing apparatus

    JP2021165944A