System, method, and program
The system uses a two-dimensional distance sensor and self-supervised learning to accurately estimate a person's three-dimensional position on a mobile robot, overcoming sparse LiDAR and labeling challenges.
Patent Information
- Application Number
- JP2023214818
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-07-02
- Estimated Expiration
- 2043-12-20
AI Technical Summary
Existing systems using 2D LiDAR for detecting the position of a person on a mobile robot face issues with sparse LiDAR signals, limited calculation amount, and the need for manual large-scale labeling.
A system employing a two-dimensional distance sensor, a camera, and a conversion neural network trained by self-supervised learning using knowledge distillation to estimate the three-dimensional position of a person, utilizing a bounding box detection and binary classification of detection points to determine correspondence with the person.
Enables accurate and efficient estimation of a person's position with reduced computational load and without large-scale manual labeling, allowing for precise tasks like conveyance.
Smart Images

Figure 2025098584000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a system, method, and program for estimating a three-dimensional position of a person.
Background Art
[0002] Patent Document 1 discloses a system having a camera and LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging). An image segmentation mapper performs segmentation of an image. Each of the segmentations is associated with spatial coordinates. A depth mapper generates a depth map of a scene based on depth values and spatial coordinates.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In this system, when using a 2D LiDAR for detecting the position of a person on a mobile robot, there are generally three major problems. Specifically, the LiDAR signals are sparse, the amount of calculation is limited, and manual large-scale labeling is required.
[0005] The present disclosure has been made in view of the above background, and an object thereof is to provide a system, method, and program capable of accurately estimating the position of a person around a robot.
Means for Solving the Problems
[0006] The system according to the present disclosure includes a mobile robot equipped with a two-dimensional distance sensor that detects the distance to a peripheral point, a camera that images an image of the periphery of the mobile robot, a detection unit that detects a bounding box surrounding a person included in the image, a determination unit that determines whether each detection point of the two-dimensional distance sensor included in the bounding box corresponds to the person, and an estimation unit that estimates the three-dimensional position of the person based on the distance to the detection point determined to correspond to the person.
[0007] In the above system, the determination unit may be a conversion neural network that takes, as input, position data including the detection direction and distance from the two-dimensional distance sensor and outputs binary data indicating whether each of the detection points corresponds to a person.
[0008] In the above system, the conversion neural network may be a machine learning model trained by self-supervised learning using knowledge distillation.
[0009] In the above system, the training data of the conversion neural network may be data obtained by a feature estimator combined with a segmentation network that segments a person from an image of a camera and a clustering algorithm that extracts detection points on the ankles of the person.
[0010] The method according to the present disclosure is a method for estimating the three-dimensional position of a person using a computer, and includes steps of detecting the distance to a peripheral point using a two-dimensional distance sensor mounted on a mobile robot, imaging an image of the periphery of the mobile robot with a camera, detecting a bounding box surrounding a person included in the image, determining whether each detection point of the two-dimensional distance sensor included in the bounding box corresponds to the person, and estimating the three-dimensional position of the person based on the distance to the detection point determined to correspond to the person.
[0011] In the above method, the conversion neural network may determine whether the detection point corresponds to the person, and the conversion neural network may output binary data indicating whether each of the detection points corresponds to a person, using the position data including the detection direction and distance from the two-dimensional distance sensor as an input.
[0012] In the above method, the conversion neural network may be a machine learning model trained by self-supervised learning using knowledge distillation.
[0013] In the above method, the training data of the conversion neural network may be data obtained by a feature estimator combined with a segmentation network that segments a person from an image of a camera and a clustering algorithm that extracts detection points on the ankle of the person.
[0014] The program according to the present disclosure causes a computer to execute the above method.
Advantages of the Invention
[0015] According to the present disclosure, it is possible to provide a system, a method, and a program capable of accurately estimating the position of a person around a robot.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0017] Hereinafter, with reference to FIG. 1, the configuration and method of the system according to this embodiment will be described. FIG. 1 is a schematic diagram showing the overall configuration of a system 1 having a mobile robot 100 (also simply referred to as robot 100). Here, the robot 100 is an autonomous mobile robot having wheels 11. Therefore, the robot 100 can autonomously move along the route to the destination.
[0018] The robot 100 includes a main body 10, wheels 11, a distance sensor 13, a support column 20, and a camera 21. The main body 10 is a chassis that rotatably holds the wheels 11. The main body 10 is a housing that houses a battery, a wheel motor, a control unit, etc. (not shown). The main body 10 may be a carrier for carrying luggage or the like. A support column 20 for supporting the camera 21 is attached to the main body 10. That is, the camera 21 is installed on the support column 20.
[0019] The camera 21 is a CMOS (Complementary Metal Oxide Semiconductor) image sensor, a CCD (Charge Coupled Device) image sensor, or the like. The camera 21 may be mounted on a smartphone, a tablet computer, or the like. The camera 21 may be a color camera such as an RGB camera. The camera 21 captures an image of the periphery of the robot 100. For example, since the camera 21 faces the front of the robot 100, it captures an image in front of the moving direction of the robot 100. Therefore, when there is a person P in front of the robot 100, the camera 21 captures an image including the person P. The camera 21 outputs the captured data to a control unit described later.
[0020] A distance sensor 13 is mounted on the side surface of the main body 10. The distance sensor 13 is, for example, an optical sensor that measures the distance D to a person P in the vicinity. The distance sensor 13 is preferably a two-dimensional distance sensor such as a two-dimensional LiDAR (Light Detection And Ranging). The distance sensor 13 may be provided on the four side surfaces, i.e., the front, rear, left, and right of the main body 10, or may be provided on only some of the side surfaces. The distance sensor 13 has a light source and a photosensor. The distance sensor 13 emits measurement light, for example, toward the front in the moving direction. Then, the distance sensor 13 detects the reflected light reflected by the person P or an object in the vicinity (hereinafter also collectively referred to as a surrounding point).
[0021] For example, the distance sensor 13 measures the distance to the surrounding point as point cloud data. The distance sensor 13 scans the measurement light to detect the distance to the surrounding point in each direction. The surrounding points include a wall, an obstacle, another robot, a person, etc. For example, the distance sensor 13 scans the laser light at a constant angular interval in an arbitrary plane such as a horizontal plane. The distance sensor 13 changes the scanning angle, i.e., the detection direction, and detects the distance to the surrounding point.
[0022] The distance sensor 13 can acquire position data indicating the distance to the detection point. The position data is data in which the distance is associated with the detection direction. That is, the position data of each detection point includes the detection direction (scanning angle) and the distance value. The distance sensor 13 outputs the position data to a control unit described later. Here, the correspondence between the detection direction in the distance sensor 13 and the angle of view in the camera 21 is known. That is, in the robot 100, since the installation positions of the distance sensor 13 and the camera 21 are fixed, the position data of the distance sensor 13 can be converted into xy coordinates in the image. That is, the detection point of the distance sensor 13 indicated by the three-dimensional coordinates can be projected onto the two-dimensional image of the camera 21.
[0023] Next, the configuration and processing of the control unit will be described. FIG. 2 is a block diagram showing the configuration of the processing device 30. The processing device 30 includes a detection unit 31, a determination unit 32, and an estimation unit 33.
[0024] The detection unit 31 detects a bounding box (boundary box) that surrounds a person included in the image captured by the camera 21. As shown in FIG. 3, the bounding box B is shown as a rectangular frame that surrounds the person P in the image I.
[0025] The detection unit 31 executes a bounding box detection network to detect the bounding box B that surrounds the person P. Known methods can be used for bounding box detection. For example, the detection unit 31 detects the person P included in the image I by performing object detection by image processing. Then, the detection unit 31 specifies a rectangular frame that surrounds the person P in the image I as the bounding box B. The bounding box B is indicated by xy coordinates or the like in the image I. The detection unit 31 may detect the bounding box B using a machine learning model such as deep learning.
[0026] By trimming the image I using the bounding box B, downsampling can be performed at high speed. Since the bounding box detection is only used to adjust the network, it is not necessary to make the bounding box detection very tight. Therefore, the detection process of the bounding box can be executed at high speed.
[0027] The determination unit 32 determines whether each detection point of the distance sensor 13 included in the bounding box B corresponds to the person P. For example, the determination unit 32 uses a transformer that takes as input position data including the detection direction and distance from the distance sensor 13 and outputs binary data indicating whether each detection point corresponds to a person. The transformer may be a conversion neural network generated by machine learning. The conversion neural network may be a machine learning model trained by self-supervised learning (SSL) using knowledge distillation.
[0028] For example, the determination unit 32 generates binary data indicating the determination result. For example, when the detection point DP of the distance sensor 13 corresponds to the person P, the data value becomes "1", and when the detection point does not correspond to the person P, the data value becomes "0". Specifically, the determination unit 32 binarizes each detection point included in the bounding box B to generate binary data.
[0029] FIG. 4 is a schematic diagram for explaining the determination process in the determination unit 32. As shown in FIG. 4, the position data of the detection point DP in the bounding box B is input to the transformer 321. As described above, the position data of the detection point DP includes the detection direction and distance in the distance sensor 13.
[0030] The transformer 321 receives the position data at N (N is an integer of 2 or more) detection points. Then, the transformer 321 generates binary data for each detection point DP. That is, the transformer 321 outputs an N-bit data sequence. In the output of the transformer 321 shown in FIG. 4, the detection point DP corresponding to the person P is shown as the detection point DP1, and the detection point DP not corresponding to the person is shown as the detection point DP0. The output of the transformer 321 is a sequence of the same length that passes through the sigmoid layer because it is treated as a binary classification of whether each detection point belongs to a person.
[0031] The estimation unit 33 estimates the three-dimensional position of the person P based on the distance to the detection point DP1 determined to correspond to a person. For example, the median value of the distances of a plurality of detection points DP1 is calculated as the distance from the distance sensor 13 to the person P. The estimation unit 33 calculates the three-dimensional coordinates of the person based on the distance of the person P. The estimation unit 33 identifies the direction of the person P based on an image or position data. The estimation unit 33 estimates the three-dimensional position based on the distance and direction. For example, the estimation unit 33 can estimate the three-dimensional coordinates of the person P based on the xy coordinates of the person P in the image I of the camera 21 and the direction in the position data.
[0032] In this way, the determination unit 32 determines whether each detection point DP is a detection point DP1 corresponding to the person P or a detection point DP0 not corresponding to the person P. Therefore, the determination can be made with high accuracy with a limited amount of calculation. Further, the estimation unit 33 can estimate the three-dimensional position with high accuracy based on the sparse signal from the two-dimensional distance sensor 13. The transformer 321 can be constructed without performing large-scale manual labeling.
[0033] The transformer 321 is a conversion neural network that takes as input position data including the detection direction and distance from the distance sensor 13 and outputs binary data indicating whether each detection point DP corresponds to a person. Specifically, the transformer 321 performs a binary classification as to whether each detection point belongs to a person. The determination can be made with high accuracy with a limited amount of calculation. The transformer 321 can be constructed without performing large-scale manual labeling.
[0034] The transformer 321 may be a machine learning model trained by self-supervised learning using knowledge distillation. By using such a transformer 321, it is possible to make a highly accurate determination while suppressing the computational load. This enables the estimation of the three-dimensional position of a human in a robot 100 equipped with a two-dimensional distance sensor. In the processing device 30, even when there are computational constraints, highly accurate position estimation can be performed. Since the robot 100 can estimate the position of a person with high accuracy, tasks such as conveyance can be appropriately carried out. Also, highly accurate position estimation is possible from the sparse signals detected by the two-dimensional distance sensor 13.
[0035] The machine learning of the transformer 321 will be described below. First, the training data used for machine learning will be described. Data obtained by a segmentation network that segments a person from a camera image and a keypoint estimator (feature estimator) combined with a clustering algorithm that extracts detection points on a person's ankle is used as the training data. That is, the data obtained by the processing in the segmentation network, the keypoint estimator, and the clustering algorithm becomes the training data. Then, by performing machine learning using the training data, various parameters etc. of the transformer 321 are optimized.
[0036] As shown below, a processing device such as a server generates learning data. FIG. 4 is a block diagram showing the configuration of a processing device 40 that generates learning data (also referred to as training data). The processing device 40 includes a segmentation network 41, a keypoint estimator 42, and a clustering unit 43. The processing device 40 stores algorithms for performing processes such as segmentation, keypoint (feature point) estimation, and clustering. Each algorithm can be realized by a machine learning model of a deep neural network or the like. CNN (Convolutional neural network), FCN (Fully Convolutional Network), etc. can be used. The processing device 40 may be the same device as the processing device 30 or a different device. The segmentation network 41, the keypoint estimator 42, and the clustering unit 43 may be existing models obtained by deep learning or the like.
[0037] The processing device 40 extracts a data set of the image of the camera 21 and the position data of the detection points detected by the distance sensor 13 from the log of the robot 100. Note that the robot equipped with various sensors for detecting data may be the same robot as the robot 100 shown in FIG. 1 or a different robot. That is, the robot that collects the learning data may be the same robot as the robot 100 that performs the three-dimensional position estimation of the person P or a different robot. Also, the camera and the distance sensor used for data collection may be of the same type as the camera 21 and the distance sensor 13 used for position estimation or of a different type. For example, the distance sensor 13 may be a 2D LiDAR or a 3D LiDAR. Also, various types of RGB cameras can be used for the camera 21.
[0038] The segmentation network 41 is a machine learning model that segments the image I. For example, the segmentation network 41 classifies each pixel of the image I into a class (object). The segmentation network 41 predicts a class label for all the objects included in the image I by instance segmentation. Therefore, the segmentation network 41 can identify the pixels of the class corresponding to the person P in the image I.
[0039] The keypoint estimator 42 is a machine learning model that estimates the keypoints (feature points) of the person P from the image I. The keypoint estimator 42 estimates the keypoints of the person P by keypoint estimation (feature point estimation) processing. The keypoint estimator 42 estimates a specific part of the person P from among the pixels corresponding to the person P. For example, the keypoint estimator 42 estimates the joint positions such as the elbows, shoulders, pelvis, wrists, ankles, and knees as keypoints by image processing. The keypoint estimator 42 estimates the xy coordinates of the keypoints in the image I.
[0040] The processing device 40 can estimate the posture and skeleton of the person P from the estimation result of the keypoint estimator 42. FIG. 6 is a schematic diagram showing the skeleton F obtained from the feature points of the person P obtained by keypoint estimation. For example, the processing device 40 can estimate the skeleton F of the person P by connecting the positions of the joints.
[0041] The clustering unit 43 is an algorithm for clustering the detection points of the distance sensor 13. For example, the clustering unit 43 uses the DBSCAN (Density-Based Spatial Clustering of Applications with noise) clustering algorithm to cluster the detection points of the distance sensor 13. In the DBSCAN clustering algorithm, clustering is performed according to the density of the detection points. That is, the clustering unit 43 expands the area of the cluster until the density of the detection points becomes less than a certain value. A plurality of detection points included in an area with a density of a certain value or more belong to one cluster. The processing device 40 clusters a plurality of detection points and divides them into a plurality of clusters.
[0042] The processing device 40 extracts the cluster corresponding to the ankle from among the plurality of clusters obtained by clustering. The processing device 40 compares the keypoint estimation result and the clustering result to identify the ankle cluster. Specifically, the cluster closest to the keypoint estimated to be the ankle is identified as the cluster corresponding to the ankle from among the plurality of clusters. Note that the distance sensor 13 is preferably set to detect the height and direction of the ankle.
[0043] The processing device 40 extracts the position data of the detection points included in the cluster corresponding to the ankle and the image data as teacher data. The position data of the detection points at the position of the ankle becomes accurate information (correct label, teacher label). The processing device 40 can automatically label the learning data. That is, labeled detection points can be automatically generated. Incorrect marking due to inconsistency between the distance sensor 13 and the camera 21 can be suppressed.
[0044] The processing device 40 performs supervised learning using the teacher data in which the correct information is associated with the image data. The processing device 40 updates each parameter of the transformer 321. That is, the parameters of the transformer 321 are tuned so that the network of the transformer 321 is optimized.
[0045] In this way, the processing device 40 generates the transformer 321 by performing machine learning using the learning data. The processing device 40 can construct the transformer 321 by supervised learning. It is possible to obtain high-precision segmentation and the effect of knowledge distillation from a human keypoint estimator. The transformer 321 can be constructed by self-supervised learning via knowledge distillation. The transformer 321 may be mounted on the robot 100 or may be executed online.
[0046] In this way, efficient annotation and machine learning can be performed. Further, since the determination unit 32 makes a determination based on the position data of the detection points, the determination unit 32 can make an appropriate determination even between different sensor models and environments.
[0047] The processes of segmentation, keypoint estimation, and clustering in the processing device 40 have a large computational load and are thus difficult to execute online. In order to generate detection points with ground truth as learning data, the processing device 40 executes the above processes offline. The processing device 40 can generate learning data for a large amount of data without manually labeling it. At the detection points on the ankle, it is possible to determine whether it belongs to a person with high determination accuracy. Also, for detection points other than a person, by labeling them, they can be used as learning data.
[0048] Note that the present invention is not limited to the above-described embodiments and can be appropriately modified without departing from the gist. Furthermore, the present disclosure can be realized by causing a processor such as a CPU (Central Processing Unit) to execute a computer program for part or all of the processing in the processing device 30. For example, the processing devices 30, 40, etc. can be implemented as devices capable of executing a program such as a central processing unit of a computer. And various functions can also be realized by a program.
Explanation of Symbols
[0049] DP detection point I image P person F skeleton DP detection point 1 system 10 main body part 11 wheel 13 distance sensor 20 support column 21 camera 30 processing device 31 detection unit 32 determination unit 321 transformer 33 estimation unit 40 processing device 41 segmentation network 42 keypoint estimator 43 clustering unit 100 robot
Claims
1. A mobile robot equipped with a two-dimensional distance sensor for detecting the distance to surrounding points, a camera for imaging an image of the surroundings of the mobile robot, a detection unit for detecting a bounding box surrounding a person included in the image, a determination unit for determining whether each of the detection points of the two-dimensional distance sensor included in the bounding box corresponds to the person, and an estimation unit for estimating the three-dimensional position of the person based on the distance to the detection point determined to correspond to the person. A system comprising:
2. The system according to claim 1, wherein the determination unit is a conversion neural network that takes as input position data including the detection direction and distance from the two-dimensional distance sensor and outputs binary data indicating whether each of the detection points corresponds to a person.
3. The system according to claim 2, wherein the conversion neural network is a machine learning model trained by self-supervised learning using knowledge distillation.
4. The system according to claim 3, wherein the learning data of the conversion neural network is data obtained by a keypoint estimator combined with a segmentation network for segmenting a person from an image of a camera and a clustering algorithm for extracting detection points on the ankles of the person.
5. A method for estimating the three-dimensional position of a person using a computer, comprising: detecting the distance to surrounding points using a two-dimensional distance sensor mounted on a mobile robot; imaging an image of the surroundings of the mobile robot with a camera; detecting a bounding box surrounding a person included in the image; determining whether each of the detection points of the two-dimensional distance sensor included in the bounding box corresponds to the person; and estimating the three-dimensional position of the person based on the distance to the detection point determined to correspond to the person. A method comprising:
6. A conversion neural network determines whether the detection point corresponds to the person, The method according to claim 5, wherein the conversion neural network is a conversion neural network that takes as input position data including the detection direction and distance from the two-dimensional distance sensor and outputs binary data indicating whether each of the detection points corresponds to a person.
7. The method according to claim 6, wherein the conversion neural network is a machine learning model learned by self-supervised learning using knowledge distillation. **Claim 8** The method according to claim 6, wherein the training data of the conversion neural network is data obtained by a feature estimator combined with a segmentation network that segments a person from an image of a camera and a clustering algorithm that extracts detection points on the ankle of the person. **Claim 9** A program that causes a computer to execute the method according to any one of claims 5 to 8.
Citation Information
Patent Citations
Action recognition system
JP2006221379A
Driving assist system
JP2020166485A
Systems and methods for laser depth map sampling
JP2018526641A