An Adaptive Following Method for Mobile Robots Based on Deep Fusion Ranging

By introducing Mask R-CNN and monocular depth estimation algorithm into mobile robots, combined with deep fusion ranging technology, the artifact problem of monocular depth sensors during pedestrian depth measurement is solved, and the stable follow-up of the target pedestrian is achieved.

CN114937070BActive Publication Date: 2025-05-30CHANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210695752.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2025-05-30
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

In the prior art, monocular depth sensors have artifact problems when measuring the depth of pedestrians, resulting in low depth ranging accuracy and it is difficult to achieve stable follow-up of mobile robots to pedestrians.

Method used

Adaptive follow-up method of mobile robots based on deep fusion ranging is adopted, combined with Mask R-CNN instance segmentation algorithm and monocular depth estimation algorithm, the distance and angle of the target pedestrian and robot are calculated by the fusion of camera depth information, and stable follow-up is achieved using a proportional differential integral controller.

Benefits of technology

By introducing Mask R-CNN and monocular depth estimation algorithm, pedestrian distance is accurately measured, invalid depth pixel points are suppressed, ranging accuracy is improved, and stable follow-up to the target pedestrian is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114937070B_ABST
    Figure CN114937070B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of robot applications, and in particular to an adaptive following method for a mobile robot based on depth fusion ranging, including obtaining a depth image and a color image by using a monocular camera; introducing the MaskR-CNN algorithm to obtain the total number of depth pixels of the pedestrian mask and the pedestrian mask area; introducing a monocular depth estimation algorithm to output an inference depth image and replace the invalid pixel points in the camera depth image; using the mask to extract the depth pixel points of the pedestrian area from the camera depth image to accurately measure the human-robot distance; using a proportional integral differential controller to adjust the distance and angle deviation between the target pedestrian and the robot. The present invention introduces the MaskR-CNN instance segmentation algorithm and the monocular depth estimation algorithm, and then fuses the camera depth information to calculate the distance and angle between the target pedestrian and the robot, and sends the position information to the proportional differential integral control module of the robot to achieve stable following of the target pedestrian.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot applications, and particularly to a mobile robot adaptive following method based on depth fusion ranging. Background Art

[0002] With the rapid development of robot perception technology, the tracking accuracy of following robots has been continuously improved. The target following method for moving robots according to instructions can achieve target following through the position and angle deviations between the target and the robot, which greatly facilitates people's production and life. However, in actual scenarios, complex scenarios such as people passing through each other, obstacles, and changes in light intensity will occur, posing new challenges to the precise following of mobile robots.

[0003] In the prior art, millimeter-wave radar and lidar can directly obtain the target position, but both have low imaging resolution and it is difficult to distinguish the object boundary. High-resolution lidar is difficult to popularize in mobile robots due to its high price. By learning from the biological mechanism of human vision, the effect of human-like visual perception can be achieved only by using a low-cost monocular depth sensor. The depth information of pedestrians is obtained by using the color image texture information output by the monocular depth sensor. Due to the artifact problem in the depth image of the monocular depth sensor, it is difficult to improve the accuracy of depth ranging, resulting in the inability of the mobile robot to stably follow pedestrians. The present invention proposes a mobile robot adaptive following method based on depth fusion ranging, aiming at the two core contents of pedestrian detection and precise ranging in the following system. Summary of the Invention

[0004] Aiming at the deficiencies of the existing algorithms, the present invention proposes a mobile robot adaptive following method based on depth fusion ranging, introducing the Mask R-CNN instance segmentation algorithm and the monocular depth estimation algorithm, and then fusing the camera depth information to calculate the distance and angle between the target pedestrian and the robot, and sending the position information to the proportional-integral-derivative control module of the robot to achieve stable following of the target pedestrian.

[0005] The technical solution adopted by the present invention is as follows: A mobile robot adaptive following method based on depth fusion ranging includes the following steps:

[0006] S1. Obtain a depth image T l and a color image x using a monocular camera;

[0007] S2. Introduce the Mask R-CNN algorithm to perform pixel-level instance segmentation on the color image x to obtain the pedestrian mask and the total number n of depth pixels in the pedestrian mask area;

[0008] Further, the specific steps include:

[0009] The color image x is input into the trained ResNeXt neural network to obtain the corresponding feature map;

[0010] For each point in the feature map, a candidate region of interest (ROI) is set to obtain multiple candidate regions of interest;

[0011] The candidate regions of interest are used as the input image and sent to the RPN network for foreground and background classification and bounding box regression, and the softmax probability distribution predicted by the classifier is adopted;

[0012] The true class label of the target is u, and the regression parameter of u predicted by the bounding box regressor is t u , and the true target bounding box regression parameter (v x , v y , v w , v h ) is v;

[0013] The classification loss function L c , the bounding box regression loss function L b and the loss function L of the mask m are calculated respectively. By defining a multi-task loss function L for each RoI, a part of the candidate regions of interest are filtered out, and the remaining ROIs are subjected to ROIAlign operation to correspond the pixels of the collected original camera color image and the obtained feature map, and then the pixels of the feature map and the extracted features are corresponded; finally, p(x) is output.

[0014] S3. Introduce a monocular depth estimation algorithm to output the inference depth image T t , and replace the invalid pixel points in the camera depth image T l ;

[0015] The specific steps include:

[0016] Define E 1 , G 1 to represent the encoder and decoder of the structure extraction module. The encoder maps the color image X to the latent space c, and the decoder generates the structure diagram M s ;

[0017] Furthermore, the formula of the latent space c is as follows:

[0018]

[0019] Among them, x represents the input color image; W c , W s represent the weight matrices; since there is a lot of useless low-dimensional information in the structure diagram M s ;

[0020] Encoder E in the deep attention module 2 Using an extended residual network, after the decoder G 2 Add a sigmoid function to output the deep attention M 2 , E 2 = DRNet(x, W D ), M A = σ(G 2 (DRNet(x, W D ), W A )); where σ represents the sigmoid activation function; w D , w A represent weight matrices; the inference depth is: T t = f(M s ×M A )); where f() represents the depth prediction module;

[0021] S4. Use the mask to extract the depth pixel points of the pedestrian area from the camera depth image T l to accurately measure the human-robot distance E dis ;

[0022] Furthermore, the specific steps include:

[0023] Define P li to represent the depth value of the pedestrian area of the camera depth image T l , P ti to represent the depth value of the pedestrian area of the monocular depth estimation algorithm network inference depth image T t , calculate the measured human-robot distance E dis ;

[0024] Furthermore, the calculation formula of E dis is as follows: where n is the total number of depth pixels in the pedestrian mask area.

[0025] S5. Use a proportional-integral-derivative controller (PID) to adjust the distance e d (k) and angle deviation e w (k) between the target pedestrian and the robot, adjust the error between the set value and the actual value, and use the measured distance and angle as inputs, with the outputs being the linear velocity and angular velocity.

[0026] Advantages of the present invention:

[0027] 1. By introducing the Mask R-CNN network, pixel-level instance segmentation of the target pedestrian is performed to obtain a pedestrian mask, and thus depth pixel points are extracted from the depth image according to the mask area to accurately measure the pedestrian distance;

[0028] 2. Use the monocular depth estimation algorithm to perform depth inference on the invalid pixel points in the depth camera. At the same time, fuse the inferred depth image and the camera depth information to suppress the invalid depth pixel points to improve the pedestrian distance measurement accuracy;

[0029] 3. Use a proportional-integral-derivative controller to adjust the error between the actual human-robot distance and the set distance of the mobile robot, improve the motion state of the mobile robot, and achieve stable following of the target pedestrian. Description of the Drawings

[0030] Figure 1 is the flowchart of the mobile robot adaptive following method based on depth fusion ranging of the present invention;

[0031] Figure 2 is the schematic diagram of generating the pedestrian foreground mask by the Mask R-CNN network structure of the present invention;

[0032] Figure 3 is the structure diagram of inferring depth using the monocular depth estimation algorithm of the present invention;

[0033] Figure 4 is the following effect diagram of the proportional-integral-derivative controller adjusting the deviation between the set target value and the actual value of the present invention;

[0034] Figure 5 is the comparison result of the measured distance and the actual distance between the method of the present invention and the YOLOv5 algorithm;

[0035] Figure 6 is the pedestrian detection of YOLOv5 and Mask R-CNN of the present invention;

[0036] Figure 7 is the visualization of indoor and outdoor color images and depth images of the present invention. Detailed Embodiments

[0037] The present invention will be further described below with reference to the drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.

[0038] The present invention includes two parts: a depth fusion ranging algorithm and a pedestrian following control. Combining the Mask R-CNN instance segmentation algorithm and the monocular depth estimation algorithm model, and then fusing the camera depth information, sending the information to the proportional-derivative-integral control module of the robot to achieve stable following of the pedestrian.

[0039] First, run the pedestrian detection module and determine whether a pedestrian can be detected. When the system determines that a pedestrian is detected, the depth fusion ranging module obtains the distance information and detects whether the mobile robot has reached the set distance. If the distance between the target pedestrian and the mobile robot is not equal to the set distance, the proportional-integral-derivative controller outputs the corresponding linear velocity and angular velocity, so that the trolley stably follows the target pedestrian to move to the target distance. When the target pedestrian continues to move, the trolley will continue to follow; otherwise, the trolley will stop moving.

[0040] As Figure 1 shown, a mobile robot adaptive following method based on depth fusion ranging includes the following steps:

[0041] S1. Use a monocular camera to obtain a depth image T l and a color image x;

[0042] Through depth ranging, the robot can follow the target at a set distance. The mobile robot uses an RGB-D camera to collect color images and depth images, obtains the pedestrian position information based on the depth image, adjusts the relative position between the human and the machine accordingly, and screens pixel points according to the obtained depth image information.

[0043] S2. Introduce the Mask R-CNN algorithm to perform pixel-level instance segmentation on the color image x, and obtain the total number n of depth pixels in the pedestrian mask and the pedestrian mask area. Figure 2 Schematic diagram of generating a pedestrian foreground mask for the Mask R-CNN network structure;

[0044] The color image x is preprocessed and input into the trained ResNeXt neural network to obtain the corresponding feature map. Then, a candidate region of interest (ROI) is set for each point in the feature map, so as to obtain multiple candidate regions of interest. These candidate regions of interest are used as the input map and sent to the RPN network for foreground and background classification and boundary regression. The softmax probability distribution predicted by the classifier is used, the label of the foreground is 1, the label of the background is 0, the true class label of the target is u, and the regression parameter of u predicted by the bounding box regressor is t u , and the true target bounding box regression parameter (v x , v y , v w , v h ) is v. During the training process, the classification loss function L c , the bounding box regression loss function L b and the mask loss function L m, the Mask RCNN network model outputs a binary mask for each RoI. By defining a multi-task loss function L for each RoI, a part of the candidate regions of interest are filtered out, and the remaining ROIs are subjected to ROIAlign operation to correspond the pixels of the collected original camera color image and the obtained feature map, and then the pixels of the feature map and the extracted features are corresponded; finally, the model output p(x) is performed, where the true label is y; among them, the classification loss function is L c =-logs u , the bounding box regression loss function The loss function of the mask is L m =-ylog(p(x))-(1-y)log(1-p(x)). A multi-task loss function is defined for each RoI as follows: L = L c +L b +L m . At the same time, in order to reduce the false positive rate, the pedestrian confidence of the algorithm in the present invention is set to 0.92;

[0045] Depth image ranging generally detects and selects the region where the object is located, calculates the depth value for the selected region, and then completes the ranging and tracking task. The accuracy of depth ranging depends on the accuracy of object detection. However, in actual scenarios, factors such as lighting conditions and target size will affect the accuracy of object detection, resulting in inaccurate object bounding boxes, missed detections and false detections of targets; such as Figure 6 By comparing the YOLOv5 algorithm and the Mask R-CNN model in the present invention, when affected by light, the YOLOv5 algorithm will have obvious false detections and missed detections and other phenomena, and it cannot accurately select pedestrians, affecting the accuracy of depth ranging; when detecting pedestrians at a long distance, the YOLOv5 algorithm is prone to problems such as false detections of pedestrians, repeated detections of single pedestrians, and missed detections of pedestrians at a long distance, while the Mask R-CNN algorithm has significant advantages in long-distance instance segmentation; in a crowded scene, the YOLOv5 algorithm has problems of false detections and missed detections at the image edge and occlusion, while the Mask R-CNN algorithm can achieve pixel-level segmentation through the strategy of candidate regions of interest (ROIs) identified by the classifier, thereby improving the accuracy of pedestrian detection. Therefore, the effect of introducing the Mask R-CNN algorithm to segment the human body target is better, and it can still output a pedestrian mask well and accurately obtain the depth points of the pedestrian region in some complex scenarios; such as Figure 5The Mask RCNN network model and the YOLOV5 network model are selected to compare the measured distance with the actual distance. The distances detected by the YOLOv5 algorithm are: 1.02 m, 2.47 m, 3.74 m, and the distances detected by the Mask RCNN algorithm are: 0.6 m, 2.38 m, 4.35 m; the corresponding actual distances are: 0.6 m, 2.40 m, 4.80 m. From this set of experimental data, it can be seen that the Mask RCNN algorithm selected by the present invention is relatively accurate for pedestrian target detection.

[0046] S3. Introduce a monocular depth estimation algorithm to output the inference depth image T t , replacing the invalid pixel points in the camera depth image T l ; Figure 3 The structural diagram of inferring depth using the monocular depth estimation algorithm (S2R-DepthNet) is given:

[0047] The present invention introduces a monocular depth estimation algorithm for depth fusion ranging. During the process of the robot following a pedestrian, due to the uncertainty of the pedestrian's motion state and the imaging mechanism of the monocular depth camera, the mobile robot generates a distance error. The depth ranging error stems from the depth pixel points with a distance of 0 in the depth image. The monocular depth estimation algorithm introduced by the present invention infers the depth and is used to replace the invalid depth points in the pedestrian mask area of the camera depth image, thereby improving the ranging accuracy and completing the stable following task of the pedestrian; define E 1 , G 1 represent the encoder and decoder of the structure extraction module. The former maps the color image X to the latent space c, and the latter generates the structure diagram M s , and the formula is as follows: Among them, x represents the input color image; W c , W s represent the weight matrices; since there is a lot of useless low-dimensional information in the structure diagram M s , the monocular depth estimation algorithm combines the attention mechanism with the structure diagram M s to eliminate the influence of low-dimensional information on depth estimation and further establish the mapping relationship between color pixels and depth pixels; the encoder E 2 in the depth attention module uses an extended residual network to increase the receptive field of the encoder network; a sigmoid function is added after the decoder G 2 to output the depth attention M 2 , and the specific process is as follows: E 2 =DRNet(x, W D ), M A =σ(G 2 (DRNet(x, W D ), W A)); where σ represents the sigmoid activation function; w D and w A represent the weight matrix; the inferred depth of the monocular depth estimation algorithm is: T t = f(M s × M A ); where f() represents the depth prediction module, M s is the structure diagram, and M A is the depth attention map;

[0048] S4. Extract the depth pixel points of the pedestrian area from the camera depth image using a mask to accurately measure the human-robot distance;

[0049] According to the above three steps, output a pedestrian mask through Mask R-CNN, and then use the monocular depth estimation algorithm to infer the depth, which is used to replace the invalid depth points in the pedestrian mask area of the camera depth image, reduce the invalid depth points, and thus improve the ranging accuracy; combined with the depth image information collected by the monocular camera, define P li to represent the depth value of the pedestrian area in the camera depth image T l , and P ti to represent the depth value of the pedestrian area in the depth image T t inferred by the monocular depth estimation algorithm network, measure the human-robot distance E dis , and the formula is as follows: where n is the total number of depth pixels in the pedestrian mask area;

[0050] To verify the role of the monocular estimation algorithm in depth fusion, Figure 7 shows the visualization results of the color image and the depth image. The redder the color of the depth image, the farther the human-robot distance, and the bluer the color, the closer the human-robot distance. Figure 7 In (a), it is an indoor scene. There are invalid points in the camera depth image, while the S2R-DepthNet algorithm has depth continuity in its inferred image due to the addition of the structure consistency module, avoiding the appearance of invalid depth points. Figure 7 In (b), it is an outdoor scene. It is not difficult to see that due to the strong outdoor light and the excessive human-robot distance, the depth image of the depth camera shows blue (that is, the distance represented by each pixel point is 0), while the S2R-DepthNet inferred depth image can obtain the human-robot distance, but compared with the camera depth image, the inferred depth distance is not accurate. In the outdoor scene, the depth fusion algorithm can obtain approximate human-robot distance information, thereby improving the robustness of the mobile robot following system; therefore, the present invention uses the ranging method of Mask R-CNN at close range and the depth fusion ranging method at long range, replaces the invalid depth points with the depth image after monocular depth estimation for depth ranging, and finally improves the accuracy of depth ranging.

[0051] S5. Use a proportional-integral-derivative controller (PID) to adjust the distance and angle deviation between the target pedestrian and the robot, regulate the error between the set value and the actual value, take the measured distance and angle as inputs, and the outputs are the linear velocity and angular velocity;

[0052] Adopt a discrete proportional-integral-derivative controller (PID), and its calculation formula is:

[0053]

[0054] where, e d (k) and e w (k) are the distance deviation and angle deviation at the k-th moment respectively; e d (k - 1) and e w (k - 1) are the distance deviation and angle deviation at the (k - 1)-th moment respectively; K P , K I and K D are the proportional coefficient, integral coefficient and derivative coefficient of the velocity proportional-integral-derivative controller respectively; K p , K i and K d are the proportional coefficient, integral coefficient and derivative coefficient of the angle proportional-integral-derivative controller respectively; and respectively represent the cumulative sum of the deviations of e d (k) and e w (k); to ensure the safety of the mobile robot when following the target person, the following distance set in the present invention is 0.6 m, the following angle is set to 0°, the maximum values of the angular velocity and linear velocity are 0.6 m / s, and by adjusting appropriate PID parameters, it can stably follow the pedestrian and meet the requirements of the control system;

[0055] Adopt the displacement error D err to evaluate the accuracy of the distance between the robot and the target. The displacement error is calculated from the actual distance R dis and the measured distance E dis as follows: D err = R dis - E dis .

[0056] Figure 4 The following figure shows the following effect diagram of the proportional-integral-derivative controller adjusting the deviation between the set target value and the actual value.

[0057] Experimental result analysis

[0058] A simulation experiment is designed for a robot to follow a target pedestrian. The mobile robot consists of a LeTV LeTMC-520 RGB-D camera, a Jetson Nano, an STM32 controller, and a wheeled mobile platform, integrating functions such as visual perception, motion planning, and control execution. Among them, the LeTV monocular depth camera is used as a visual sensor, enabling the robot to follow the target at a set distance through depth ranging. The Jetson Nano has the advantages of low power consumption and high computing power. It can not only cooperate with the monocular depth camera to measure the distance but also connect to the STM32 to control the movement of the trolley;

[0059] The displacement error D err (Displacement Error) is used to evaluate the accuracy of the distance between the robot and the target. The displacement error is calculated from the actual distance R dis (Real Distance) and the measured distance E dis and can be obtained as: D err =R dis -E dis ; By comparing the tracking effects of the mobile robot indoors and outdoors, the indoor mobile robot uses a monocular depth camera to collect color images and depth images, and obtains the pedestrian position information based on the depth image to adjust the relative position between the target pedestrian and the mobile robot; In the outdoor scenario, due to the relatively complex outdoor light conditions, the mobile robot can still stably track the target pedestrian.

[0060] Inspired by the above ideal embodiments of the present invention, through the above description, relevant staff can make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. An adaptive following method for a mobile robot based on depth fusion ranging, characterized in that, it includes the following steps: S1. Use a monocular camera to obtain a depth image and a color image; S2. Introduce the Mask R-CNN algorithm to perform pixel-level instance segmentation on the color image, and obtain the total number of depth pixels in the pedestrian mask and the pedestrian mask area; The step S2 includes: The color image x is input into the trained ResNeXt neural network to obtain the corresponding feature map; Set candidate regions of interest (ROIs) for each point in the feature map to obtain multiple candidate regions of interest; Use the candidate regions of interest as inputs and send them into the RPN network for foreground and background classification and boundary regression, and adopt the softmax probability distribution predicted by the classifier; Calculate the classification loss function $L$ separately c 、the bounding box regression loss function $L$ b and the loss function $L$ of the mask m , define a multi-task loss function $L$ for each RoI, filter out a part of the candidate regions of interest, perform ROIAlign operation on the remaining ROIs, correspond the pixels of the collected original camera color image and the obtained feature map, and then correspond the pixels of the feature map and the extracted features; finally, output $p(x)$; S3. Introduce a monocular depth estimation algorithm to output an inferred depth image and replace the invalid pixel points in the camera depth image; The step S3 includes: Define E 1 , G 1 represent the encoder and decoder of the structure extraction module. The encoder maps the color image X to the latent space c, and the decoder generates the structure diagram M s ; Encoder E in the depth attention module 2 Use an extended residual network. After the decoder G 2 Add a sigmoid function to output the depth attention M 2 , E 2 = DRNet(x, W D ), M A = σ(G 2 (DRNet(x, W D ), W A )); where σ represents the sigmoid activation function; w D , w A represent weight matrices; the inference depth is: T t = f(M s ×M A ); where f() represents the depth prediction module; S4. Use the mask to extract the depth pixel points of the pedestrian area from the camera depth image to accurately measure the distance between the person and the robot; S5. Use a proportional-integral-derivative controller to adjust the distance and angle deviation between the target pedestrian and the robot, and adjust the error between the set value and the actual value. Use the measured distance and angle as inputs, and the outputs are the linear velocity and the angular velocity.

2. The adaptive following method for a mobile robot based on depth fusion ranging according to claim 1, characterized in that, the calculation formula for the latent space c is: Among them, x represents the input color image; W c , W s represent weight matrices.

3. The adaptive following method for a mobile robot based on depth fusion ranging according to claim 2, characterized in that, the step S4 includes: Define P li Represent the depth value of the pedestrian area in the camera depth image T l P ti Represent the depth value of the pedestrian area in the monocular depth estimation algorithm network inference depth image T t Calculate the measured human-machine distance E dis .

4. The adaptive following method for a mobile robot based on depth fusion ranging according to claim 3, characterized in that, The measured human-machine distance E dis is calculated by the formula: where n is the total number of depth pixels in the pedestrian mask area.

Citation Information

Patent Citations

  • A mobile robot target tracking method based on a depth map region of interest

    CN109949375A

  • 6D pose estimation method based on monocular RGB camera regression depth information

    CN113393522A