A guide robot and its guide control method

Through reinforcement learning and timing prediction models, the path planning and control of guide robots is optimized, and the existing guide robots are solved, and more efficient and safe guide tasks are achieved.

CN116035874BActive Publication Date: 2025-07-25WEST LAKE ROBOTICS TECH (HANGZHOU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310032362.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2025-07-25
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

The existing mobile guide robots have shortcomings in terrain adaptability and safety, and cannot effectively meet the travel needs of visually impaired users.

Method used

The four-legged guide robot control method based on reinforcement learning and timing prediction models is adopted to obtain terrain information through radar, RGB-D camera and encoder, and combine path planning and robot dynamics model to optimize the rope tension, rope length and yaw angle to achieve accurate guide tasks.

Benefits of technology

The terrain adaptability and safety of the guide robot are improved, and the path planning can be adjusted according to the speed of the blind, ensuring the safety and efficiency of the guide tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116035874B_ABST
    Figure CN116035874B_ABST
Patent Text Reader

Abstract

A guide dog robot, comprising a robot platform and a control device. The robot platform is a guide dog robot platform, and the control device is arranged above the robot platform. The control device consists of a brushless DC motor, a motor controller, a gimbal with an encoder, and a traction rope. The motor controller communicates with a computing unit through a CAN communication network. The motor controller controls the brushless DC motor, and a force sensor is integrally arranged on the motor controller. The traction rope is arranged on the force sensor. The beneficial effect of the present invention is that this solution uses a four-legged guide dog robot control based on reinforcement learning and a time series prediction model to accurately and efficiently complete the guide dog task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of blind guiding, and particularly to a blind guiding robot and a blind guiding control method thereof. Background Art

[0002] Visually impaired people refer to those whose vision has declined to the extent that they cannot correct it through traditional methods to ensure their daily living standards. Patients with a certain degree of vision decline are called blind people; researchers around the world have been constantly trying to develop guiding robots to help blind people navigate and improve their quality of life. Guiding robots in various forms and with various functions have continuously come into view; among them, the mainstream development means are mainly divided into intelligent blind canes, wearable blind guiding devices, handheld blind guiding instruments, intelligent terminal-based blind guiding systems, and mobile blind guiding robots.

[0003] Intelligent electronic blind canes often add various sensors and microcontrollers on the basis of traditional blind canes, and can provide road surface information to visually impaired people; the disadvantage is that they cannot give specific navigation information, have a single function, and a low safety factor.

[0004] Wearable blind guiding devices assemble corresponding blind guiding devices on the coats, glasses, backpacks, shoes, earphones and other equipment of visually impaired people, and use the voices in the left and right ears to give people a sense of direction; but they cannot give a physical pulling feeling and too many devices bring too much weight, which is easy to cause fatigue to people.

[0005] The handheld blind guiding instrument is also one of the five mainstream development means. The biggest advantage is that it is light, convenient and easy to carry. This system is simple, low-cost and not affected by electromagnetic interference, but it is greatly affected by obstacles, reflectors and other light sources, and is only suitable for simple and dark indoor environments.

[0006] The intelligent terminal-based blind guiding system makes full use of the advantages that intelligent terminals integrate a large number of sensors and have built-in navigation functions; however, there is a problem that most of the existing relatively perfect blind guiding systems require long-term maintenance and are too costly.

[0007] Therefore, the mobile blind guiding robot has become the most potential and applicable blind guiding method. The most common mobile blind guiding robot at present is a combination of a wheeled cart and a towing rod. This kind of blind guiding method can easily and safely guide the blind without interactive information such as sound and vibration. Such mobile blind guiding robots have significantly improved the travel efficiency and safety of visually impaired people; however, such robots have high requirements for terrain, and the current wheeled blind guiding robots do not meet the terrain adaptability requirements; at the same time, the existing mobile blind guiding robots also need to further enhance safety, and these factors increase the hidden dangers of travel for visually impaired users. Summary of the Invention

[0008] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide a blind-guiding robot and a blind-guiding control method thereof.

[0009] The present invention is achieved through the following technical solutions: a guide robot for the blind, comprising a robot platform and a control device, wherein the robot platform is a guide robot platform, the control device is arranged above the robot platform, and the control device comprises a brushless DC motor, a motor controller, a universal joint with an encoder, and a traction rope; the motor controller communicates with a computing unit through a CAN communication network, the motor controller controls the brushless DC motor, a force sensor is integrated on the motor controller, and a traction rope is arranged on the force sensor.

[0010] Preferably, a radar, an RGB-D camera and an encoder are arranged on the robot platform, and the radar, RGB-D camera and encoder obtain point cloud, the distance between the camera and the blind person and the deflection angle of the blind person relative to the guide robot, and the point cloud, the distance between the camera and the blind person and the deflection angle of the blind person relative to the guide robot are sent to the mapping and positioning system by the encoder; the mapping and positioning system generates a grid map and sends it to the path planning system; the blind person path planner plans a collision-free path according to the position of the person and the target position and transmits it to the path correction module; the path correction module modifies the path according to the speed of the blind person and the guide robot according to the trained strategy output k The value is used to modify the step size of the path planning to obtain a planned path that meets the speed of the blind person; k The value is the sampling rate of the original path planning, and its value range is [0, 1, 2, 3,4]; the reinforcement learning network obtains the sampling rate of path planning by inputting the speed of the blind and the guide robot; the blind motion prediction model in the blind motion planner obtains the position of the blind according to the blind position information and feedback force input by the blind positioning module, and transmits it to the blind motion planner; the blind motion planner optimizes the output rope tension, rope length and yaw angle of the guide robot through model predictive control according to the revised planned path and the predicted blind position; the expected position of the guide robot can be obtained by the position of the person, the direction of the rope tension, the rope length and the yaw angle of the guide robot; the guide robot motion prediction model predicts the position of the guide robot according to the positioning information, expected speed, feedback speed, feedback force and encoder angle of the guide robot positioning module; the guide robot motion planner optimizes the output command speed of the guide robot through the MPC method according to the predicted guide robot position and the expected guide robot position through model predictive control; the guide robot speed controller controls the guide robot movement according to the command speed, and the motor PID controller controls the motor according to the command tension to pull the blind through the guide rope to complete the guide task.

[0011] A guiding control method for a guiding robot, including pedestrian motion planning, robot motion control, and a motion model prediction model;

[0012] The pedestrian motion planning is implemented through a pedestrian motion planner, and the pedestrian motion planner plans an expected position for the robot and an expected force for the rope motor ; First, the planner will first call the pedestrian state predictor HSE to predict the future speed of the pedestrian , and the SAC algorithm will select suitable points from the global path planning points as the expected positions for the next steps of the person; Then construct a matrix , where is a unit vector representing the force direction, represents the rope length, is the yaw angle, and predict the position of the person in the next steps according to the pedestrian state, and its formula is:

[0013]

[0014]

[0015] where represents the time step; T To optimize and obtain the best

[0016] , minimize the loss function according to the formula:

[0017] where, , , is a weight parameter, , and are set threshold parameters;

[0018] For the loss function, minimizing can make the pedestrian approach the expected end point; minimizing can make the predicted trajectory of the pedestrian approach the planned trajectory ; minimizing and can make the magnitude and direction of the rope tension not change suddenly; minimizing and can make the length and direction of the rope not change suddenly;

[0019] Finally, is used to control the motor; meanwhile, according to the optimization matrix , the Robot Dynamics Model (RDM) is used to predict the path of the robot motion controller:

[0020]

[0021]

[0022] The described robot motion controller plans the desired speed for the guide robot ; Since the underlying controller of the robot can resist interference, the dynamics model of the robot is related to the disturbances received in the past several steps and the real speed; therefore, the force F and the speed change received by the dog cannot be simplified to a linear relationship;

[0023] To better express F and relationship, a dataset is collected by the West Lake guide robot, and a robot kinematics model RDM based on Transformer is trained;

[0024] First, an optimization matrix to be constructed , according to the formula:

[0025]

[0026]

[0027]

[0028] predicts the position of the guide robot in the next steps;

[0029] To optimize and obtain the best , the following loss function is minimized:

[0030]

[0031] where , is the weight parameter;

[0032] For the above loss function, minimizing can make the predicted trajectory of the robot close to the planned trajectory ; Minimizing can ensure the efficiency of the robot's walking;

[0033] Finally, it is obtained Control the robot with the optimal linear velocity and angular velocity.

[0034] The described motion prediction model is a Transformer based on the existing sequence-to-sequence model, using an encoder-decoder architecture. In the encoder-decoder architecture, the encoder converts the input sequence (x1, …, xn) into a continuous representation z = (z1, …, zn), and then the decoder generates the output sequence (y1, …, ym) based on this representation. The model is stacked by an encoder and a decoder, and each layer has the same structure. The encoder consists of 2 layers, and each layer includes two sub-layers: the first layer is a multi-head self-attention layer, and the second layer is a simple fully connected feed-forward network. After each sub-layer, a residual connection and normalization are connected. , that is, the output of each sub-layer is. For the convenience of the residual connection, the output vector dimensions of all sub-layers in the model, including the embedding layer, are .

[0035] The attention mechanism used in Transformer is called "Scaled Dot-Product Attention". The input of this module includes three vectors: the query vector , the key vector and the value vector . The three vectors are all calculated based on the input vector. The dimensions of the query vector and the key vector are , and the dimension of the value vector is . First, calculate the dot product of a single query vector and all key vectors, then divide it by , and finally obtain the corresponding weights through a softmax function, and then weight them with the value vector. This process can be expressed by the following formula:

[0036]

[0037] where the query vector Q, the key vector K, and the value vector V; the dimensions of the query vector and the key vector are . Given an input matrix, multiple groups of V, K, and Q matrices are calculated based on different parameter matrices, and then multiple weighted V matrices are calculated through multiple attention functions. Finally, these matrices are concatenated and the final output is obtained through a weight matrix W. This is the multi-head attention mechanism of Transformer. The above process can be expressed by the formula as follows:

[0038] .

[0039] Preferably, the prediction method of the pedestrian state predictor HSE is as follows:

[0040] The Transformer model is used; the Transformer proposes an encoder-decoder structure for sequential data; the speed of a person (x, y) and the force exerted on the person ( , ) are used as inputs; the input window is 5 and the output window is 1; before entering the encoder, the input first enters a linear layer, the input dimension of the linear layer is 4, and the output dimension is 40; then the output of the linear layer is encoded with position information, and then enters the encoder layer; the number of encoder layers is 2, parameters , , dropout = 0.01, and the activation function uses the relu function; finally, it enters the decoder, and the decoder uses a linear layer structure, the input dimension of the decoder is 40, and the output dimension is 126; the output of the decoder is concatenated with the force at the next moment optimized by the MPC as the input of the last linear layer, and its output is the speed of the person (x, y) and the force coordinate of the person at the next moment ( , ); the learning rate of the model is 0.05, and the optimizer is SGD.

[0041] Preferably, the method for establishing the robot dynamics model RDM is as follows:

[0042] The desired speed, feedback speed of the guide robot, the angle of the encoder, and the magnitude of the feedback force are used as inputs; the input window is 10 and the output window is 1; before entering the encoder, the input first enters a linear layer, the input dimension of the linear layer is 4, and the output dimension is 40; then the output of the linear layer is encoded with position information, and then enters the encoder layer; the number of encoder layers is 1, parameters , , dropout = 0.01, and the activation function uses the relu function; finally, it enters the decoder, and the decoder uses a linear layer structure, the input dimension of the decoder is 40, and the output dimension is 4; the speed discount coefficient D is solved according to the desired speed and feedback speed output by the decoder; the learning rate of the model is 0.05, and the optimizer is SGD.

[0043] Preferably, the calculation method of the supplementary k value in the path correction module is as follows:

[0044] Step 1: Build a virtual simulation environment for the guide robot with the ability to train neural networks and construct a path correction network;

[0045] Step 2: Initialize the virtual simulation environment;

[0046] Step 3: Continuously update the simulation environment. In each simulation environment, the path correction network combines each simulation environment and outputs path correction parameters k , and based on the output of the path planning, perform path correction operations to obtain the final corrected planned path, and calculate the reward function according to the decisions of the robot;

[0047] Step 4: Judge the termination condition of the environment training, and collect the training data set in the current environment;

[0048] Step 5: Use the training data set to train the control network, obtain an optimized path correction network, and deploy it to the real guide robot for path planning correction;

[0049] Further, the first part of the path correction network is a fully connected network, which includes two hidden layers, each layer containing 256 nodes, and the activation function selects the relu function.

[0050] Further, the initialization of the virtual simulation environment includes initializing the simulation environment where the guide robot is located, as well as initializing the initial position, attitude, and environmental terrain information of the robot, and setting the initial roll angle, pitch angle, and yaw angle of the guide robot to 0.

[0051] Further, the specific process of calculating the reward function according to the decision response of the robot is as follows: In the simulation environment, the guide robot makes corresponding single decisions according to the path correction network, calculates the reward function of each decision in real time, and designs a threshold to judge whether the robot and the blind person fall; Repeat the decision instructions of the guide robot until reaching the set destination or reaching the upper limit of the training times in the current environment, and exit the current environment simulation;

[0052] The calculation formula of the reward function r is as follows:

[0053]

[0054]

[0055]

[0056] Among them, represents the speed of the person, represents the speed of the robot; is the reward for the moving speed of the blind person, encouraging the blind person to move at a reasonable speed; is the reward for the moving speed of the robot, and its purpose is to encourage the guide robot to move at a reasonable speed.

[0057] Even further, the specific process of Step 5 is as follows:

[0058] Collect the current states s, actions a, desired states of the guide robot and the blind person in the simulation environment , reward results r, and termination judgment conditions d in the simulation environment, and record them as the decision instruction dataset in the current environment , where N is the size of the dataset; and use the action instruction dataset in the current environment to train the path correction network. The optimizer uses Adam, and the learning rate is 0.001; repeat the above operations to train the path correction network until the total training times limit is reached.

[0059] Preferably, the specific method of the path planning is as follows:

[0060] Use the A* planner on the grid map to generate a collision-free set of human waypoints; the coordinates of the person are obtained from the camera depth and the encoder angle, and extended to ; then calculate the movement cost and the heuristic cost and the total cost on the grid map; find a path with the lowest total cost, and the collision-free set of human waypoints will be passed to the pedestrian motion planner; then, the pedestrian motion planner will pass in a target point of the robot. Similarly, path planning is performed from the coordinates of the robot to the target point, and the set of waypoints is passed to the robot motion planner.

[0061] The beneficial effects of the present invention are as follows: This solution uses a quadruped guide robot control based on reinforcement learning and a time series prediction model to accurately and efficiently complete the guide task;

[0062] Reinforcement learning generates a corresponding k value according to the states of the blind person and the guide robot through the learned policy to modify the step length of the path planning; the motion planner MPC of the blind person optimizes the pulling force F of the rope, the rope length l, and the yaw angle theta of the guide robot according to the modified planned path, so as to obtain the planned path of the guide robot, and obtains the speed of the robot through the motion planner to complete the task; the present invention better describes the states of the blind person and the guide robot through the time series prediction model, and the guide robot continuously explores based on reinforcement learning; learns the strategy of how to modify the step length of the planned path according to the speed of the blind person to cooperate with the subjective speed of the person, and is stably applied on the physical guide robot to cooperate with different blind people and speeds. 5. It can allow people to have a greater degree of subjective movement, and the robot can adjust its own speed and other states to cooperate well with the subjective movement of people, and can complete the guide task while ensuring safety;

[0063] 6. The modified path inherits the correctness of the original path direction and incorporates the speed information of the blind person, so that it can ensure the safe and correct completion of the guide task while appropriately adjusting the speed of the robot to cooperate with the subjective speed of the blind person;

[0064] 7. The linear equation of the original solution cannot analyze and predict the states of the blind and the guide robot in a fine-grained manner. The time series prediction model as a multi-layer non-linear function can accurately analyze the state of the blind at the next moment based on historical data with a certain step size. Description of the Drawings

[0065] Figure 1 It is a schematic diagram of the control principle flow of a guide control method for a guide robot.

[0066] Figure 2 It is a schematic diagram of the structure of a guide robot and its guide control method.

[0067] Figure 3 It is a schematic diagram of the rear structure of a guide robot and its guide control method.

[0068] Figure 4 It is a schematic diagram of the side structure of a guide robot and its guide control method.

[0069] Wherein: 1. Robot platform; 2. Control device; 3. Towing rope. Detailed Embodiments

[0070] In the description of the present invention, it should also be noted that unless otherwise clearly defined and limited, the terms "set", "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0072] Such as Figures 1-4As shown, a guide robot for the blind comprises a robot platform and a control device, wherein the robot platform is a guide robot platform, the control device is arranged above the robot platform, and the control device comprises a brushless DC motor, a motor controller, a universal joint with an encoder, and a traction rope; the motor controller communicates with a computing unit via a CAN communication network, the motor controller controls the brushless DC motor, a force sensor is integrated on the motor controller, and a traction rope is arranged on the force sensor.

[0073] Preferably, a radar, an RGB-D camera and an encoder are arranged on the robot platform, and the radar, RGB-D camera and encoder obtain point cloud, the distance between the camera and the blind person and the deflection angle of the blind person relative to the guide robot, and the point cloud, the distance between the camera and the blind person and the deflection angle of the blind person relative to the guide robot are sent to the mapping and positioning system by the encoder; the mapping and positioning system generates a grid map and sends it to the path planning system; the blind person path planner plans a collision-free path according to the position of the person and the target position and transmits it to the path correction module; the path correction module modifies the k value of the path according to the trained strategy output based on the speed of the blind person and the guide robot, thereby modifying the step size of the path planning to obtain a planned path that meets the speed of the blind person; the K value is the sampling rate of the original path planning, and its value range is [0, 1, 2, 3,4]; the reinforcement learning network obtains the sampling rate of path planning by inputting the speed of the blind and the guide robot; the blind motion prediction model in the blind motion planner obtains the position of the blind according to the blind position information and feedback force input by the blind positioning module, and transmits it to the blind motion planner; the blind motion planner optimizes the output rope tension, rope length and yaw angle of the guide robot through model predictive control according to the revised planned path and the predicted blind position; the expected position of the guide robot can be obtained by the position of the person, the direction of the rope tension, the rope length and the yaw angle of the guide robot; the guide robot motion prediction model predicts the position of the guide robot according to the positioning information, expected speed, feedback speed, feedback force and encoder angle of the guide robot positioning module; the guide robot motion planner optimizes the output command speed of the guide robot according to the predicted guide robot position and the expected guide robot position through the model predictive control (MPC) method; the guide robot speed controller controls the guide robot movement according to the command speed, and the motor PID controller controls the motor according to the command tension to pull the blind through the guide rope to complete the guide task.

[0074] A blind guide control method of a blind guide robot, including pedestrian motion planning, robot motion control, and motion model prediction model;

[0075] The described pedestrian motion planning is achieved through a pedestrian motion planner, which plans the desired position for the robot and the desired force for the rope motor ; The pedestrian motion planner includes a Human State Estimator (HSE) and a motion planning function. First, the planner calls the HSE to predict the future speed of the pedestrian ; Then, the soft actor-critic (SAC) algorithm selects appropriate points from the global path planning points as the desired positions for the next

[0076] steps of the person; Then, a matrix is constructed, where is a unit vector representing the force direction, , represents the rope length, is the yaw angle, and its formula is predicted according to the pedestrian state:

[0077]

[0078]

[0079] to predict the position of the person in the next steps, where T represents the time step;

[0080] To optimize and obtain the best , the motion planning function is used to design the following minimized loss function:

[0081]

[0082] where , , are weight parameters, , and are set threshold parameters;

[0083] For the loss function, minimizing can make the pedestrian approach the desired end point; minimizing can make the predicted trajectory of the pedestrian close to the planned trajectory ; minimizing and can prevent sudden changes in the magnitude and direction of the rope tension; minimizing and The length and direction of the rope cannot change suddenly;

[0084] By solving the above loss function, the optimal result is finally obtained. It is used to control the motor; at the same time, according to the optimization matrix , the Robot Dynamics Model (RDM) is used to predict the path of the robot motion controller as follows:

[0085]

[0086]

[0087] The described robot motion controller plans the desired speed for the guide robot ; Since the underlying controller of the robot can resist interference, the dynamics model of the robot is related to the disturbances received in the past several steps and the true speed; therefore, the force F and the speed change received by the dog cannot be simplified to a linear relationship;

[0088] To better express the F and relationship, a dataset was collected using the West Lake guide robot, and a robot kinematic model RDM based on Transformer was trained;

[0089] First, the matrix to be optimized was constructed. According to the formula:

[0090]

[0091]

[0092]

[0093] The position of the guide robot in the next steps was predicted;

[0094] To optimize and obtain the best , the following loss function was minimized:

[0095]

[0096] where , is the weight parameter;

[0097] For the above loss function, minimizing can make the predicted trajectory of the robot close to the planned trajectory ; Minimize It can ensure the efficiency of the robot's movement;

[0098] Finally, it is obtained As the optimal linear velocity and angular velocity to control the robot.

[0099] The described motion prediction model is that Transformer is based on the existing sequence-to-sequence model and uses the encoder-decoder architecture; in the encoder-decoder architecture, the encoder converts the input sequence (x1,…,xn) into a continuous representation z=(z1,…,zn), and then the decoder generates the output sequence (y1,…,ym) based on this representation; the model is stacked by the encoder and the decoder, and each layer has the same structure; the encoder consists of 2 layers, and each layer includes two sub-layers: the first layer is the multi-head self-attention layer, and the second layer is a simple fully connected feed-forward network; after each sub-layer, a residual connection and normalization are connected , that is, the output of each sub-layer is. For the convenience of the residual connection, the output vector dimensions of all sub-layers in the model, including the embedding layer (initial word embedding), are .

[0100] The attention mechanism used in Transformer is called "Scaled Dot-Product Attention"; the input of this module includes three vectors: the query vector , the key vector and the value vector ; the three vectors are all calculated based on the input vector (the initial input vector is the word embedding), the dimensions of the query vector and the key vector are , and the dimension of the value vector is ; first calculate the dot product of a single query vector and all key vectors, then divide it by , and finally obtain the corresponding weight through a softmax function, and then weight it with the value vector; this process can be expressed by the following formula:

[0101]

[0102] where the query vector Q, the key vector K and the value vector V; the dimensions of the query vector and the key vector are ; Given an input matrix, multiple sets of V, K, and Q matrices are calculated based on different parameter matrices, and then multiple weighted V matrices are obtained through multiple attention functions. Finally, these matrices are concatenated and the final output is obtained through a weight matrix W; this is the multi-head attention mechanism of the Transformer. The above process can be expressed by the following formula:

[0103] Preferably, the method of the pedestrian state predictor (Human State Estimator, HSE) is as follows:

[0104] The Transformer model is used; the Transformer proposes an encoder-decoder structure for sequential data; the human speed (x, y) and the force exerted on the human ( , ) are used as inputs; the input window is 5 and the output window is 1; before entering the encoder, the input first enters a linear layer, the input dimension of the linear layer is 4, and the output dimension is 40; then the output of the linear layer is encoded with position information, and then enters the encoder layer; the number of encoder layers is 2, parameters , , dropout = 0.01, and the activation function uses the relu function; finally, it enters the decoder, the decoder uses a linear layer structure, the input dimension of the decoder is 40, and the output dimension is 126; the output of the decoder is concatenated with the force at the next moment optimized by MPC as the input of the last linear layer, and its output is the human speed (x, y) and the force coordinate of the human at the next moment ( , ); the learning rate of the model is 0.05, and the optimizer is SGD.

[0105] Preferably, the method of the robot dynamics model (Robot Dynamic Model, RDM) is as follows:

[0106] The desired speed, feedback speed of the guide robot, the angle of the encoder, and the magnitude of the feedback force are used as inputs; the input window is 10 and the output window is 1; before entering the encoder, the input first enters a linear layer, the input dimension of the linear layer is 4, and the output dimension is 40; then the output of the linear layer is encoded with position information, and then enters the encoder layer; the number of encoder layers is 1, parameters , , dropout = 0.01, and the activation function is the relu function; finally, it enters the decoder. The decoder uses a linear layer structure. The input dimension of the decoder is 40, and the output dimension is 4. The speed discount coefficient D is obtained by solving the expected speed and feedback speed output by the decoder; the learning rate of the model is 0.05, and the optimizer is SGD.

[0107] Preferably, in the path correction module, the calculation of the supplementary k value is as follows:

[0108] Step 1: Build a virtual simulation environment for a guide robot with the ability of neural network training, and construct a path correction network;

[0109] Step 2: Initialize the virtual simulation environment;

[0110] Step 3: Continuously update the simulation environment. In each simulation environment, the path correction network combines each simulation environment and outputs path correction parameters k , and according to the output of path planning, perform path correction operations to obtain the final corrected planned path, and calculate the reward function according to the decision of the robot;

[0111] Step 4: Judge the termination condition of environment training, and collect the training data set in the current environment;

[0112] Step 5: Use the training data set to train the control network, obtain an optimized path correction network, and deploy it to a real guide robot for path planning correction;

[0113] Further, the first part of the path correction network is a fully connected network, which includes two hidden layers, each layer containing 256 nodes, and the activation function selects the relu function.

[0114] Further, the initialization of the virtual simulation environment includes initializing the simulation environment where the guide robot is located, as well as initializing the initial position, attitude and environmental terrain information of the robot, and setting the initial roll angle, pitch angle and yaw angle of the guide robot to 0.

[0115] Further, the specific process of calculating the reward function according to the decision response of the robot is as follows: In the simulation environment, the guide robot makes corresponding single decisions according to the path correction network, calculates the reward function of each decision in real time, and designs a threshold to judge whether the robot and the blind person fall; repeatedly execute the decision instruction of the guide robot until reaching the set destination or reaching the upper limit of the training times in the current environment, and exit the current environment simulation;

[0116] The calculation formula of the reward function r is as follows:

[0117]

[0118]

[0119]

[0120] Among them, represents the speed of the representative person, represents the speed of the robot; is the reward for the moving speed of the blind person, which encourages the blind person to move at a reasonable speed; is the reward for the moving speed of the robot, and its purpose is to encourage the guide robot to move at a reasonable speed.

[0121] Furthermore, the specific process of the fifth step is as follows:

[0122] Collect the current states s, actions a, desired states , reward results r, and termination judgment conditions d of the guide robot and the blind person in the simulation environment, and record them as the decision instruction data set under the current environment, where N is the size of the data set; and use the action instruction data set under the current environment to train the path correction network, use Adam as the optimizer, and the learning rate is 0.001; repeat the above operations to train the path correction network until the total training times limit is reached.

[0123] Preferably, the specific method of the path planning is:

[0124] Use the A* planner on the grid map to generate a collision-free set of human waypoints; the coordinates of the person are obtained from the camera depth and the encoder angle, and extended to ; then calculate the movement cost and the heuristic cost and the total cost on the grid map; find a path with the lowest total cost, and the collision-free set of human waypoints will be passed to the pedestrian motion planner; then, the pedestrian motion planner will pass in a target point of the robot. Similarly, path planning is performed from the coordinates of the robot to the target point, and the set of waypoints is passed to the robot motion planner.

[0125] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A guiding control method for a guide robot, characterized in that: It includes a robot platform and a control device, wherein the robot platform is a guide robot platform, the control device is arranged above the robot platform, and the control device includes a brushless DC motor, a motor controller, a universal joint with an encoder, and a traction rope; the motor controller communicates with a computing unit via a CAN communication network, the motor controller controls the brushless DC motor, a force sensor is integrated on the motor controller, and a traction rope is arranged on the force sensor; A radar, an RGB-D camera and an encoder are arranged on the robot platform, and the radar, RGB-D camera and encoder obtain point clouds, the distance between the camera and the blind person, and the deflection angle of the blind person relative to the guide robot, and the point clouds, the distance between the camera and the blind person, and the deflection angle of the blind person relative to the guide robot are sent to the mapping and positioning system by the encoder; the mapping and positioning system generates a grid map and sends it to the path planning system; the blind person path planner plans a collision-free path according to the position of the person and the target position and transmits it to the path correction module; The path correction module modifies the step size of path planning according to the speed of the blind person and the guide robot by outputting the value for modifying the path according to the trained strategy, so as to obtain a planned path that conforms to the speed of the blind person; k The value is the sampling rate of the original path planning, and its value range is [0, 1, 2, 3, 4]; k The value is the sampling rate of the original path planning, and its value range is [0, 1, 2, 3, 4]; The reinforcement learning network obtains the sampling rate of path planning by inputting the speed of the blind person and the guide robot; the blind person movement prediction model in the blind person movement planner obtains the position of the blind person according to the blind person position information input by the blind person positioning module and the feedback force processing and analysis, and transmits it to the blind person movement planner; The blind motion planner optimizes the output rope tension, rope length and yaw angle of the guide robot through model predictive control according to the corrected planning path and the predicted position of the blind person; the expected position of the guide robot can be obtained from the position of the person, the direction of the rope tension, the rope length and the yaw angle of the guide robot; the guide robot motion prediction model predicts the position of the guide robot according to the positioning information, expected speed, feedback speed, feedback force and encoder angle of the guide robot positioning module; the guide robot motion planner optimizes the output command speed of the guide robot through the MPC method according to the predicted guide robot position and the expected guide robot position; the guide robot speed controller controls the guide robot movement according to the command speed, and the motor PID controller controls the motor according to the command tension to pull the blind person through the guide rope to complete the guide task; The blind guiding control method of the blind guiding robot includes pedestrian motion planning, robot motion control, and motion model prediction model; The pedestrian motion planning described above is implemented by a pedestrian motion planner, which plans the desired position for the robot , and plans the desired force for the rope motor ; First, the planner will first call the pedestrian state predictor HSE to predict the future speed of the pedestrian , and the SAC algorithm will select appropriate from the global path planning points points as the desired position for the person's next steps; Then construct a matrix , where is a unit vector representing the force direction, represents the rope length, is the yaw angle, and its formula is predicted according to the pedestrian state: to predict the position of the next person, where T represents the time step; To optimize and obtain the best , the loss function was minimized according to the formula: Among them, , , are weight parameters, , and are set threshold parameters; For the loss function, minimizing can bring the pedestrian closer to the desired end point; minimizing can make the predicted trajectory of the pedestrian close to the planned trajectory ; minimizing and can prevent sudden changes in the magnitude and direction of the rope tension; minimizing and can prevent sudden changes in the length and direction of the rope; Finally, is used to control the motor; meanwhile, according to the optimization matrix , the robot dynamics model RDM is used to predict the path of the robot motion controller: The described robot motion controller plans an expected speed for the guide robot ; Since the underlying controller of the robot can resist interference, the dynamic model of the robot is related to the disturbances received in the past several steps and the true speed; Therefore, the force F and the speed change cannot be simplified to a linear relationship; To better express F and relationship, a dataset was collected by the West Lake guide robot, and a robot kinematic model RDM based on Transformer was trained; First, construct the matrix to be optimized , according to the formula: Predicted the position of the guide robot in the next steps; To optimize for the best , the following loss function is minimized: Among them, , is a weight parameter; For the above loss function, minimizing can make the predicted trajectory of the robot close to the planned trajectory ; minimizing can ensure the efficiency of the robot's movement; Finally, control the robot with the optimal linear velocity and angular velocity; The described motion prediction model is a Transformer based on the existing sequence-to-sequence model, using the encoder-decoder architecture; in the encoder-decoder architecture, the encoder converts the input sequence (x1, …, xn) into a continuous representation z = (z1, …, zn), and then the decoder generates the output sequence (y1, …, ym) based on this representation; the model is stacked by encoders and decoders, and each layer has the same structure; the encoder consists of 2 layers, and each layer includes two sub-layers: the first layer is the multi-head self-attention layer, and the second layer is a simple fully-connected feed-forward network; after each sub-layer, a residual connection and normalization are connected , that is, for the convenience of residual connection of the output of each sub-layer, the output vector dimensions of all sub-layers in the model, including the embedding layer, are ; The attention mechanism used in the Transformer is called "Scaled Dot-Product Attention"; the input to this module consists of three vectors: the query vector , the key vector , and the value vector ; all three vectors are calculated based on the input vector. The query vector and the key vector have a dimension of , and the value vector has a dimension of ; first, the dot product of a single query vector and all key vectors is calculated, then it is divided by , and finally, the corresponding weights are obtained through a softmax function and then weighted with the value vector; this process can be expressed by the following formula: Among them are the query vector Q, the key vector K, and the value vector V; the dimensions of the query vector and the key vector are ; given an input matrix, multiple groups of V, K, and Q matrices are calculated based on different parameter matrices, and then multiple weighted V matrices are calculated through multiple attention functions. Finally, these matrices are concatenated and passed through a weight matrix W to obtain the final output; this is the multi-head attention mechanism of the Transformer; the above process can be expressed by the following formula: 。 2. The blind guiding control method of a blind guiding robot according to claim 1, characterized in that: The prediction method of the pedestrian state predictor HSE is: The Transformer model is used; the Transformer proposes an encoder-decoder structure for sequential data; the human speed (x, y) and the force applied to the human ( , ) are used as inputs; the input window is 5 and the output window is 1; the input first enters a linear layer before entering the encoder, the input dimension of the linear layer is 4, and the output dimension is 40; then the output of the linear layer is encoded with position information and then enters the encoder layer; the number of encoder layers is 2, and the parameters , , dropout = 0.01, and the activation function uses the relu function; finally, it enters the decoder, and the decoder uses a linear layer structure, the input dimension of the decoder is 40, and the output dimension is 126; the output of the decoder is concatenated with the force at the next moment optimized by MPC as the input of the last linear layer, and its output is the human speed (x, y) and the force coordinates applied to the human at the next moment ( , ); the learning rate of the model is 0.05, and the optimizer is SGD.

3. The guiding control method of a guiding robot according to claim 1, characterized in that: The method for establishing the robot dynamics model RDM is: Take the desired speed, feedback speed of the blind robot, the angle of the encoder, and the magnitude of the feedback force as inputs; the input window is 10 and the output window is 1; the input first enters a linear layer before entering the encoder, the input dimension of the linear layer is 4, and the output dimension is 40; then encode the position information of the output of the linear layer, and then enter the encoder layer; the number of encoder layers is 1, parameter , , dropout = 0.01, and the activation function uses the relu function; finally enter the decoder, the decoder uses a linear layer structure, the input dimension of the decoder is 40, and the output dimension is 4; solve the speed discount coefficient D according to the desired speed and feedback speed output by the decoder; the learning rate of the model is 0.05, and the optimizer is SGD.

4. The blind guiding control method of a blind guiding robot according to claim 1, wherein: The calculation method for supplementing the k value in the path correction module is as follows: Step 1: Build a virtual simulation environment for the guide robot with neural network training capabilities and construct a path correction network; Step 2: Initialize the virtual simulation environment; Step 3: Continuously update the simulation environment. In each simulation environment, the path correction network combines each simulation environment and outputs path correction parameters k , and according to the output of path planning, perform path correction operations to obtain the final corrected planned path, and calculate the reward function according to the decision of the robot. Step 4: Determine the termination conditions of environmental training and collect the training data set in the current environment; Step 5: Use the training data set to train the control network, obtain an optimized path correction network, and deploy it on a real guide robot for path planning correction; The specific method of the path planning is: Use the A* planner on the grid map to generate a collision-free set of human waypoints; the coordinates of the person are obtained from the camera depth and encoder angle, and extended to ; Subsequently, calculate the movement cost on the grid map and the heuristic cost as well as the total cost ; Find the path with the lowest total cost, and the collision-free set of human waypoints will be passed to the pedestrian motion planner; Subsequently, the pedestrian motion planner will pass in a target point for the robot. Similarly, path planning is performed from the coordinates of the robot to the target point, and the set of waypoints is passed to the robot motion planner.

5. A guiding control method for a guiding robot according to claim 4, characterized in that: The first part of the path correction network is a fully connected network, which includes two hidden layers, each layer contains 256 nodes, and the activation function selects the relu function.

6. A guiding control method for a guiding robot according to claim 4, characterized in that: The described initialization of the virtual simulation environment includes initializing the simulation environment where the guide robot is located, as well as initializing the initial position, attitude, and environmental terrain information of the robot, and setting the initial roll angle, pitch angle, and yaw angle of the guide robot to 0.

7. A guiding control method for a guiding robot according to claim 4, characterized in that: The specific process of calculating the reward function according to the decision response of the robot is as follows: In the simulation environment, the guide robot makes corresponding single decisions according to the path correction network, calculates the reward function of each decision in real time, and designs a threshold to judge whether the robot and the blind person fall; Repeat the decision instructions of the guide robot until the set destination is reached or the upper limit of the training times in the current environment is reached, and exit the current environment simulation; The calculation formula of the described reward function r is as follows: Among them, represents the speed of the person, represents the speed of the robot; is the reward for the moving speed of the blind, which encourages the blind to move at a reasonable speed; is the reward for the moving speed of the robot, whose purpose is to encourage the guide robot to move at a reasonable speed.

8. A guiding control method for a guiding robot according to claim 7, characterized in that: The specific process of the described step five is as follows: Collect the current states s, actions a, desired states of the guide robot and the blind person in the simulation environment , reward results r, and termination judgment conditions d, and record them as the decision instruction dataset in the current environment , where N is the size of the dataset; and use the action instruction dataset in the current environment to train the path correction network. The optimizer uses Adam, and the learning rate is 0.001; repeat the above operations to train the path correction network until the total training times limit is reached.

Citation Information

Patent Citations

  • Blind guiding robot system

    CN115079690A