System, method, and program product
By carrying a two-dimensional distance sensor and camera on a mobile robot, combined with a transformed neural network, the problem of sparse signal and limited computing volume of two-dimensional LiDAR in human position detection is solved, high-precision three-dimensional position inference is achieved, and manual tag attachment is avoided.
Patent Information
- Application Number
- CN202411858788.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-20
AI Technical Summary
When using two-dimensional LiDAR for human position detection, the prior art has problems such as sparse signals, limited computing volume and the need to manually attach large-scale tags.
Using a mobile robot equipped with a two-dimensional distance sensor, combined with a camera and a transform neural network, the detection unit detects the bounding box, the determination unit determines the binary data corresponding to the detection point and the human being, and the inference unit infers the three-dimensional position of the human being based on the distance.
High-precision inference of the three-dimensional position of the person located around the robot is achieved, reducing the computational load and avoiding the steps of manual tag attachment.
Smart Images

Figure CN120178263A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a system, a method, and a program for inferring the three-dimensional position of a person. Background Art
[0002] A system having a camera and LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) is disclosed in Patent Document 1. An image segmentation mapper performs segmentation of an image. Each of the segments is associated with spatial coordinates. A depth mapper generates a depth map of a scene based on depth values and spatial coordinates.
[0003] Patent Document 1: Japanese Patent Application Laid-Open No. 2018-526641
[0004] In this system, when a two-dimensional LiDAR is used for detecting the position of a human by a mobile robot, generally there are three major problems. Specifically, it can be cited that the signals of the LiDAR are sparse, the amount of calculation is limited, and it is necessary to manually attach large-scale labels. Summary of the Invention
[0005] The present disclosure has been made in view of the above background, and an object thereof is to provide a system, a method, and a program capable of accurately inferring the position of a person located around a robot.
[0006] The system according to the present disclosure includes: a mobile robot equipped with a two-dimensional distance sensor that detects distances to surrounding points; a camera that captures an image of the surroundings of the mobile robot; a detection unit that detects a bounding box surrounding a person included in the image; a determination unit that determines whether each of the detection points of the two-dimensional distance sensor included in the bounding box corresponds to the person; and an inference unit that infers the three-dimensional position of the person based on the distances up to the detection points determined to correspond to the person.
[0007] In the above system, the determination unit may be a transformation neural network that takes as input position data including a detection direction and a distance from the two-dimensional distance sensor and outputs binary data indicating whether each of the detection points corresponds to a person.
[0008] In the above system, the transformation neural network may be a machine learning model that is trained by using knowledge distillation and self-supervised learning.
[0009] In the above system, the learning data of the above transformation neural network can be data obtained by a segmentation network that segments a person from an image of a camera and a feature inference device that combines a clustering algorithm for extracting detection points located at the ankles of the person.
[0010] The method according to the present disclosure is a method for a computer to infer the three-dimensional position of a person, and includes: a step of using a two-dimensional distance sensor mounted on a mobile robot to detect the distance to surrounding points; a step of capturing an image of the surroundings of the mobile robot by a camera; a step of detecting a bounding box surrounding a person included in the image; a step of determining whether each of the detection points of the two-dimensional distance sensor included in the bounding box corresponds to the person; and a step of inferring the three-dimensional position of the person based on the distance to the detection points determined to correspond to the person.
[0011] In the above method, a transformation neural network determines whether the detection points correspond to the person, and the transformation neural network may be a transformation neural network that takes position data including a detection direction and a distance from the two-dimensional distance sensor as input and outputs binary data indicating whether each of the detection points corresponds to a person.
[0012] In the above method, the transformation neural network may be a machine learning model that uses knowledge distillation and is trained by self-teaching learning.
[0013] In the above method, the learning data of the transformation neural network can be data obtained by a segmentation network that segments a person from an image of a camera and a feature inference device that combines a clustering algorithm for extracting detection points located at the ankles of the person.
[0014] The program according to the present disclosure causes a computer to execute the above method.
[0015] According to the present disclosure, it is possible to provide a system, a method, and a program capable of accurately inferring the position of a person located around a robot.
[0016] The above and other objects, features, and advantages of the present disclosure will be more fully understood from the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram showing the overall configuration of the system.
[0018] Figure 2 It is a block diagram showing the control system of the system.
[0019] Figure 3 It is a diagram schematically showing the image I of the camera 21.
[0020] Figure 4 It is a schematic diagram for explaining the determination process in the determination unit.
[0021] Figure 5 It is a block diagram showing the configuration of a processing device for generating learning data.
[0022] Figure 6 It is a diagram for explaining the results of segmentation and keypoint inference. Detailed implementation mode
[0023] Hereinafter, with reference to Figure 1 The configuration and method of the system according to this embodiment will be described. Figure 1 It is a schematic diagram showing the overall configuration of a system 1 having a mobile robot 100 (also simply referred to as robot 100). Here, the robot 100 is an autonomous mobile robot having wheels 11. Therefore, the robot 100 can autonomously move along the path to the destination.
[0024] The robot 100 includes a main body 10, wheels 11, a distance sensor 13, a pillar 20, and a camera 21. The main body 10 serves as a chassis that holds the wheels 11 rotatably. The main body 10 is a housing that houses a battery, a wheel motor, a control unit, etc. not shown. The main body 10 can be a transport cart for transporting goods or the like. A pillar 20 for supporting the camera 21 is installed on the main body 10. That is, the camera 21 is provided on the pillar 20.
[0025] The camera 21 is a CMOS (Complementary Metal Oxide Semiconductor) image sensor, a CCD (Charge Coupled Device) image sensor, etc. The camera 21 can be mounted on a smartphone, a tablet computer, etc. The camera 21 can be a color camera such as an RGB camera. The camera 21 captures an image of the periphery of the robot 100. For example, since the camera 21 faces the front of the robot 100, it captures an image in front of the moving direction of the robot 100. Therefore, when a person P is in front of the robot 100, the camera 21 captures an image including the person P. The camera 21 outputs the captured data to the control unit described later.
[0026] A distance sensor 13 is mounted on the side surface of the main body 10. The distance sensor 13 is, for example, an optical sensor that measures the distance D to a person P in the vicinity. Preferably, the distance sensor 13 is a two-dimensional distance sensor such as a two-dimensional LiDAR (Light Detection And Ranging). The distance sensor 13 can be provided on the four side surfaces of the front, back, left, and right of the main body 10, or can be provided only on a part of the side surfaces. The distance sensor 13 has a light source and a photo sensor. The distance sensor 13, for example, emits measurement light toward the front in the moving direction. Moreover, the distance sensor 13 detects the reflected light reflected by the person P, an object in the vicinity, etc. (hereinafter, also collectively referred to as surrounding points).
[0027] For example, the distance sensor 13 measures the distance to the surrounding points as point cloud data. The distance sensor 13 detects the distance to the surrounding points in each direction by scanning the measurement light. The surrounding points include walls, obstacles, other robots, people, etc. For example, the distance sensor 13 scans the laser at a certain angular interval in an arbitrary plane such as a horizontal plane. The distance sensor 13 gradually changes the scanning angle, that is, the detection direction, and detects the distance to the surrounding points.
[0028] The distance sensor 13 can obtain position data indicating the distance to the detection points. In the position data, data in which the distance and the detection direction are corresponded is formed. That is, the position data of each detection point includes the detection direction (scanning angle) and the distance value. The distance sensor 13 outputs the position data to a control unit described later. Here, the correspondence relationship between the detection direction of the distance sensor 13 and the viewing angle of the camera 21 is known. That is, in the robot 100, since the installation positions of the distance sensor 13 and the camera 21 are fixed, the position data of the distance sensor 13 can be transformed into xy coordinates in the image. That is, the detection points of the distance sensor 13 shown in three-dimensional coordinates can be projected onto the two-dimensional image of the camera 21.
[0029] Next, the configuration and processing of the control unit will be described. Figure 2 It is a block diagram showing the configuration of the processing device 30. The processing device 30 includes a detection unit 31, a determination unit 32, and an inference unit 33.
[0030] The detection unit 31 detects a bounding box that encloses a person included in the image captured by the camera 21. As Figure 3 shown, the bounding box B is shown as a rectangular box that encloses the person P in the image I.
[0031] The detection unit 31 executes a bounding box detection network to detect the bounding box B that encloses the person P. Bounding box detection can use well-known techniques. For example, the detection unit 31 detects the person P included in the image I by performing object detection based on image processing. Moreover, the detection unit 31 determines the rectangular box that encloses the person P in the image I as the bounding box B. The bounding box B is shown by xy coordinates in the image I, etc. The detection unit 31 can use a machine learning model based on deep learning, etc. to detect the bounding box B.
[0032] By using the bounding box B to crop the image I, downsampling can be performed at high speed. Since the bounding box detection is only used for adjusting the network, the bounding box detection does not need to be very precise. Therefore, the detection process of the bounding box can be executed at high speed.
[0033] The determination unit 32 determines whether each detection point of the distance sensor 13 included in the bounding box B corresponds to the person P. For example, the determination unit 32 uses a converter that takes the position data including the detection direction and distance from the distance sensor 13 as input and outputs binary data indicating whether each detection point corresponds to the person. The converter can be a transformation neural network generated by machine learning. The transformation neural network can be a machine learning model trained by self-supervised learning (SSL) using knowledge distillation.
[0034] For example, the determination unit 32 generates binary data representing the determination result. For example, when the detection point DP of the distance sensor 13 corresponds to the person P, the data value is "1", and when the detection point does not correspond to the person P, the data value is "0". Specifically, the determination unit 32 binarizes each detection point included in the bounding box B to generate binary data.
[0035] Figure 4 is a schematic diagram for explaining the determination process in the determination unit 32. As Figure 4 shown, the position data of the detection point DP in the bounding box B is input to the transformer 321. As described above, the position data of the detection point DP includes the detection direction and distance in the distance sensor 13.
[0036] The transformer 321 is input with the position data of N (N is an integer of 2 or more) detection points. Moreover, the transformer 321 generates binary data for each detection point DP. That is, the transformer 321 outputs a data string of N bits. In Figure 4In the output of the converter 321 shown, the detection point DP corresponding to the person P is shown as the detection point DP1, and the detection point DP not corresponding to the person is shown as the detection point DP0. Since the output of the converter 321 is processed as a binary classification of whether each detection point belongs to a person, it becomes a sequence of the same length passing through the activation layer (sigmoid layer).
[0037] The inference unit 33 infers the three-dimensional position of the person P based on the distance up to the detection point DP1 determined to correspond to the person. For example, the median value of the distances of the plurality of detection points DP1 is calculated as the distance from the distance sensor 13 to the person P. The inference unit 33 calculates the three-dimensional coordinates of the person based on the distance of the person P. The inference unit 33 determines the direction of the person P based on the image or the position data. The inference unit 33 infers the three-dimensional position based on the distance and the direction. For example, the inference unit 33 can infer the three-dimensional coordinates of the person P based on the xy coordinates of the person P in the image I of the camera 21 and the direction in the position data.
[0038] In this way, the determination unit 32 determines whether each detection point DP is the detection point DP1 corresponding to the person P or the detection point DP0 not corresponding to the person P. Therefore, it is possible to perform the determination with high accuracy with a limited amount of calculation. In addition, the inference unit 33 can accurately infer the three-dimensional position based on the sparse signal from the two-dimensional distance sensor 13. The converter 321 can be constructed even without manually performing large-scale label attachment.
[0039] The converter 321 is a transformation neural network that takes as input the position data including the detection direction and the distance from the distance sensor 13 and outputs binary data indicating whether each detection point DP corresponds to a person. Specifically, the converter 321 performs a binomial classification of whether each detection point belongs to a person. It is possible to perform the determination with high accuracy with a limited amount of calculation. The converter 321 can be constructed even without manually performing large-scale label attachment.
[0040] The converter 321 can be a machine learning model trained by self-teaching learning using knowledge distillation. By using such a converter 321, it is possible to achieve highly accurate determination while suppressing the computational load. It is possible to infer the three-dimensional position of a human using the robot 100 equipped with a two-dimensional distance sensor. In the processing device 30, even when there are computational constraints, it is possible to perform highly accurate position inference. Since the robot 100 can accurately infer the position of a person, it is possible to appropriately perform tasks such as transportation. In addition, it is possible to perform highly accurate position inference based on the sparse signal detected by the two-dimensional distance sensor 13.
[0041] Hereinafter, the machine learning of the converter 321 will be described. First, the learning data used in the machine learning will be described. The data obtained by using a segmentation network that segments a person from an image of a camera and a key point inferencer (feature inferencer) that combines a clustering algorithm for extracting detection points located at the ankles of a person is used as the learning data. That is, the data obtained by the processes of the segmentation network, the key point inferencer, and the clustering algorithm becomes the learning data. Moreover, by performing machine learning using the learning data, various parameters, etc. of the converter 321 are optimized.
[0042] As shown below, a processing device such as a server generates learning data. Figure 5 It is a block diagram showing the configuration of a processing device 40 that generates learning data (also referred to as training data). The processing device 40 includes a segmentation network 41, a key point inferencer 42, and a clustering unit 43. The processing device 40 stores algorithms for performing processes such as segmentation, key point (feature point) inference, and clustering. Each algorithm can be implemented by a machine learning model of a deep neural network, etc. CNN (Convolutional neural network), FCN (Fully Convolutional Network), etc. can be used. The processing device 40 can be the same device as the processing device 30 or a different device. The segmentation network 41, the key point inferencer 42, and the clustering unit 43 can be existing models obtained by deep learning, etc.
[0043] The processing device 40 extracts a data set of an image of the camera 21 and the position data of the detection points detected by the distance sensor 13 from the log of the robot 100. Among them, the robot equipped with various sensors for detecting data can be the same robot as the Figure 1 robot 100 shown, or a different robot. That is, the robot that collects the learning data and the robot 100 that performs the three-dimensional position inference of the person P can be the same robot or different robots. In addition, the camera and the distance sensor used in the data collection can be of the same type or different types as the camera 21 and the distance sensor 13 used in the position inference. For example, the distance sensor 13 can be a two-dimensional LiDAR or a three-dimensional LiDAR. In addition, regarding the camera 21, various types of RGB cameras can also be used.
[0044] The segmentation network 41 is a machine learning model that segments the image I. For example, the segmentation network 41 classifies each pixel of the image I into a class (object). The segmentation network 41 predicts class labels for all the objects included in the image I through instance segmentation. Therefore, the segmentation network 41 can identify the pixels of the class corresponding to the person P in the image I.
[0045] The keypoint estimator 42 is a machine learning model that infers the keypoints (feature points) of the person P from the image I. The keypoint estimator 42 infers the keypoints of the person P through keypoint inference (feature point inference) processing. The keypoint estimator 42 infers specific parts of the person P from among the pixels corresponding to the person P. For example, the keypoint estimator 42 infers the joint positions such as the elbows, shoulders, pelvis, arms, ankles, and knees as keypoints through image processing. The keypoint estimator 42 infers the xy coordinates of the keypoints in the image I.
[0046] The processing device 40 can infer the pose and skeleton of the person P based on the inference results of the keypoint estimator 42. Figure 6 is a schematic diagram showing the skeleton F obtained based on the feature points of the person P obtained through keypoint inference. For example, the processing device 40 can infer the skeleton F of the person P by connecting the positions of the joints.
[0047] The clustering unit 43 is an algorithm that clusters the detection points of the distance sensor 13. For example, the clustering unit 43 uses the DBSCAN (Density-Based Spatial Clustering of Applications with noise) clustering algorithm to cluster the detection points of the distance sensor 13. In the DBSCAN clustering algorithm, clustering is performed based on the density of the detection points. That is, the clustering unit 43 gradually expands the area of the cluster until the density of the detection points is less than a certain value. Multiple detection points included in the area with a density of a certain value or more belong to one cluster. The processing device 40 clusters multiple detection points into multiple clusters.
[0048] The processing device 40 extracts the cluster corresponding to the ankle from among the multiple clusters obtained through clustering. The processing device 40 compares the keypoint inference results and the clustering results to determine the cluster of the ankle. Specifically, the cluster closest to the keypoint inferred to be the ankle is determined as the cluster corresponding to the ankle from among the multiple clusters. Among them, it is preferable that the distance sensor 13 is set to detect the height and direction of the ankle.
[0049] The processing device 40 extracts the position data and image data of the detection points included in the cluster corresponding to the ankle as teaching data. The position data of the detection points located at the position of the ankle becomes correct information (correct label, teaching label). The processing device 40 can automatically label the learning data. That is, it can automatically generate labeled detection points. It is possible to suppress incorrect labeling caused by mismatches for the distance sensor 13 and the camera 21.
[0050] The processing device 40 performs supervised learning with the teaching data corresponding to the image data as the correct information. The processing device 40 updates the respective parameters of the converter 321. That is, the processing device 40 adjusts the parameters to optimize the network of the converter 321.
[0051] In this way, the processing device 40 generates the converter 321 by performing machine learning using the learning data. The processing device 40 can construct the converter 321 through supervised learning. It is possible to obtain the effects of high-precision segmentation and knowledge distillation from the human key point inferencer. It is possible to construct the converter 321 through self-supervised learning with knowledge distillation. The converter 321 can be carried by the robot 100 or executed online.
[0052] In this way, efficient annotation and machine learning can be performed. In addition, since the determination unit 32 makes a determination based on the position data of the detection points, the determination unit 32 can appropriately make a determination even in different sensor models and environments.
[0053] The processes of segmentation, key point inference, and clustering in the processing device 40 are difficult to execute online due to the large computational load. In order to generate detection points with correct labels (ground truth) as learning data, the processing device 40 executes the above processes offline. Without manually attaching labels to a large amount of data, the processing device 40 can generate learning data. It is possible to determine whether it belongs to a person with high determination accuracy through the detection points at the ankle. In addition, for detection points other than a person, it is also possible to use them as learning data by attaching labels.
[0054] Furthermore, the present invention is not limited to the above-described embodiments and can be appropriately modified without departing from the gist. And, the present disclosure can be implemented by causing a processor such as a CPU (Central Processing Unit) to execute a computer program to perform part or all of the processing in the processing device 30. For example, the processing devices 30, 40, etc. can be installed as a device that can execute a program such as a central arithmetic processing unit of a computer. Moreover, various functions can also be implemented by a program.
[0055] Any type of non-transitory computer-readable medium can be used to store and provide a program to a computer. Non-transitory computer-readable media include any type of tangible storage medium. Examples of non-transitory computer-readable media include magnetic storage media (such as floppy disks, magnetic tapes, hard disk drives, etc.), magneto-optical storage media (such as magneto-optical discs), CD-ROM (Compact Disc Read-Only Memory), CD-R (Recordable Compact Disc), CD-R / W (Rewritable Compact Disc), and semiconductor memories (such as mask ROM, PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, RAM (Random Access Memory), etc.). Any type of transitory computer-readable medium can be used to provide a program to a computer. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable media can provide a program to a computer via a wired communication line (such as wires and optical fibers) or a wireless communication line.
[0056] In the above-described disclosure, it is obvious that the embodiments of the disclosure can vary in many ways. Such variations should not be regarded as a departure from the spirit and scope of the present disclosure, and it is obvious to those skilled in the art that all such modifications are intended to be included within the scope claimed in this application.
Claims
1. A system, wherein: have: The mobile robot is equipped with a two-dimensional distance sensor that detects the distance to the surrounding points; a camera for capturing images of the surroundings of the mobile robot; A detection unit, detecting a bounding box surrounding a person included in the image; a determination unit that determines whether each of the detection points of the two-dimensional distance sensor included in the bounding box corresponds to the person; and The estimating unit estimates the three-dimensional position of the person based on the distance to the detection point determined to correspond to the person.
2. The system according to claim 1, wherein: The determination unit is a transformation neural network that receives position data including a detection direction and a distance from the two-dimensional distance sensor as input and outputs binary data indicating whether each of the detection points corresponds to a person.
3. The system according to claim 2, wherein: The Transformer Neural Network is a machine learning model that uses knowledge distillation and is trained through self-taught learning.
4. The system according to claim 3, wherein: The learning data of the transformation neural network is data obtained by a segmentation network that segments a person from an image of a camera and a key point inferencer that is combined with a clustering algorithm that extracts detection points located at ankles of a person.
5. A method for using a computer to infer the three-dimensional position of a person, wherein: have: a step of detecting a distance to a peripheral point using a two-dimensional distance sensor mounted on the mobile robot; The step of capturing an image of the surroundings of the mobile robot by a camera; a step of detecting a bounding box surrounding a person included in said image; a step of determining whether each of the detection points of the two-dimensional distance sensor included in the bounding box corresponds to the person; and A step of estimating the three-dimensional position of the person based on the distance to the detection point determined to correspond to the person.
6. The method according to claim 5, wherein: The transformation neural network determines whether the detection point corresponds to the person, The transformation neural network is a transformation neural network that receives position data including detection direction and distance from the two-dimensional distance sensor as input and outputs binary data indicating whether each of the detection points corresponds to a person.
7. The method according to claim 6, wherein: The Transformer Neural Network is a machine learning model that uses knowledge distillation and is trained through self-taught learning.
8. The method according to claim 6, wherein: The learning data of the transformation neural network is data obtained by a segmentation network that segments a person from an image of a camera and a feature inference device that is combined with a clustering algorithm that extracts detection points located at ankles of a person.
9. A program product, wherein A computer is caused to execute the method according to any one of claims 5 to 8.
Citation Information
Patent Citations
Systems and methods for laser depth map sampling
JP2018526641A