Human motion intention estimation-based collaborative robot-human interaction method and device
By using Actor-Critic reinforcement learning and recursive least squares online estimation algorithms, the robot can actively estimate and follow human intentions, solving the problem that robots have difficulty accurately estimating motion intentions and achieving high-precision, low-cost, and robust human-computer interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2025-12-23
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, robots have difficulty accurately estimating human movement intentions, resulting in unsmooth human-computer interaction and increased operator burden. Furthermore, reliance on external force sensors and model uncertainties affect control accuracy and robustness.
A robot system dynamics model is constructed using an Actor-Critic reinforcement learning controller and a recursive least squares online estimation algorithm. The human-robot interaction force is estimated through internal state information, and an adaptive update law for the desired trajectory is designed to enable the robot to actively follow human intentions and reduce the impact of model uncertainty.
It achieves high-precision human-computer interaction without the need for external force sensors, reducing system cost and complexity. The robot can actively predict and follow human intentions, improving the naturalness and fluency of the interaction and possessing strong robustness.
Smart Images

Figure CN121468675B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot intelligent control technology, and in particular to a collaborative robot human-machine interaction method and device based on human motion intention estimation. Background Technology
[0002] With the rapid development of robotics technology, robot applications have expanded from industrial production to military, medical, and service sectors. In the future society, communication between humans and robots will become increasingly frequent, and human-computer interaction technology, as an information bridge between humans and robots, has become a crucial part of human life. In collaborative tasks between humans, people typically estimate the other's movement intentions and coordinate their own actions accordingly, thereby improving the smoothness and efficiency of collaboration. In human-robot collaboration scenarios, if robots can similarly estimate human movement intentions, they can proactively respond to human movements. This will lead to more effective collaboration; however, how robots can accurately detect human movement intentions remains a pressing challenge. During human-robot collaboration, if robots cannot accurately understand human movement intentions, it becomes an additional burden for human operators and may even hinder the smooth progress of collaboration. Conversely, if robots can capture and understand human movement intentions in real time, humans will be able to guide the robot's movements with less effort and fewer demands. Therefore, how to efficiently and accurately estimate human movement intentions has attracted widespread research interest.
[0003] In existing technologies, research has effectively improved robots' ability to adapt to human uncertainty and achieve compliant and precise physical interaction by combining state observers to estimate interaction forces to replace sensors, reinforcement learning, and modern optimization techniques. These advancements utilize sophisticated algorithms to compensate for, adapt to, and respond to human movements, but they do not fundamentally solve the problem of anticipation. Therefore, developing a control method that enables robots to proactively anticipate and follow the movement intentions of human operators is a problem that awaits resolution in this field.
[0004] In dynamic environments, safe human-robot interaction relies on good mutual understanding between participants and the timely and real-time generation of interactive behaviors such as timing of entry and behavioral responses. The implementation methods can be summarized as human motion intention prediction and behavioral learning. For example, an adaptive impedance controller can be used to implement human-robot collaborative handling in a task space. By using vision-force perception for hand localization and human-robot interaction-force measurement, the movement of a human partner can be tracked under conditions of limited actuator input, uncertain initial states, and unmodeled robot dynamics. Robot systems and real-world environments contain various uncertainties, making it difficult to obtain accurate robot dynamics models. This leads to problems such as completely unknown models, model mismatch, and model-based nonlinear control strategies being unsuitable for real-world robot systems. Reinforcement learning (RL) algorithms, through the collaborative work of Actor and Critic networks, adapt to complex environments with model uncertainties, effectively solving the uncertainty problem in robots. However, these methods lack an active following mechanism; even if the interaction force can be estimated, the robot cannot actively adjust its target trajectory based on the estimated intention, making it difficult to fundamentally reduce the human effort required to guide the robot. Summary of the Invention
[0005] To address the technical problems of reliance on external force sensors in human-computer interaction and poor tracking performance and unsmooth interaction caused by model uncertainty in existing technologies, this invention provides a collaborative robot human-computer interaction method and apparatus based on human motion intention estimation. The technical solution is as follows:
[0006] On the one hand, a collaborative robot human-computer interaction method based on human motion intention estimation is provided. This method is implemented by a collaborative robot human-computer interaction device based on human motion intention estimation, and includes: S1. Construct a dynamic model of a robot system with human-machine interaction force, and design an Actor-Critic reinforcement learning-based controller. S2. Based on the Actor-Critic reinforcement learning controller, output the optimal control strategy; and approximate the uncertain dynamics of the robot system's dynamic model online based on the optimal control strategy. S3. An online estimation algorithm based on recursive least squares is used to estimate the force applied by the human body and output the interaction force between the human and the computer. S4. Design the desired trajectory adaptive update law; based on the interaction force between human and machine, adopt the desired trajectory adaptive update law to dynamically adjust the desired joint trajectory of the robot system, so that the desired joint trajectory of the robot actively follows the human's movement intention.
[0007] On the other hand, a collaborative robot human-computer interaction device based on human motion intention estimation is provided. This device is applied to a collaborative robot human-computer interaction method based on human motion intention estimation. The device includes: Building units are used to construct dynamic models of robot systems with human-robot interaction forces and to design Actor-Critic reinforcement learning-based controllers. The output unit is used to output the optimal control strategy based on the Actor-Critic reinforcement learning controller; and to approximate the uncertain dynamics of the dynamic model of the robot system online based on the optimal control strategy. The estimation unit is used to estimate the force applied by the human body using an online estimation algorithm based on recursive least squares, and output the interaction force between the human and the computer. An adjustment unit is used to design the adaptive update law of the desired trajectory. Based on the interaction force between human and machine, the desired joint trajectory of the robot system is dynamically adjusted according to the adaptive update law of the desired trajectory, so that the desired joint trajectory of the robot actively follows the human's movement intention.
[0008] On the other hand, a collaborative robot human-computer interaction device based on human motion intention estimation is provided. The collaborative robot human-computer interaction device based on human motion intention estimation includes: a processor; a memory, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, any one of the methods of the collaborative robot human-computer interaction method based on human motion intention estimation described above is implemented.
[0009] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement any of the above-described collaborative robot human-computer interaction methods based on human motion intention estimation.
[0010] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: This invention successfully estimates human-robot interaction forces online from the robot's internal state information and torque input using an intention estimator based on a recursive least squares online estimation algorithm and a human control model. This achieves accurate human-robot interaction without the need for external force sensors, significantly reducing system cost and complexity.
[0011] After adopting the embodiments of the present invention, the robot can proactively predict the human's movement intentions and start moving in advance. This makes the interaction process more natural and smooth, just like collaboration between people.
[0012] Even when the robot model is inaccurate or there are unmodeled dynamics, the embodiments of this invention can achieve high-precision trajectory tracking. It exhibits strong robustness to uncertainties in the robot system's dynamics, ensuring high-precision trajectory tracking performance. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart of a collaborative robot human-computer interaction method based on human motion intention estimation provided by an embodiment of the present invention; Figure 2 This is a schematic diagram of a human-computer interaction process structure provided in an embodiment of the present invention; Figure 3 This is a diagram showing the experimental results of comparing an embodiment of the present invention with a traditional neural network method. Figure 4 This is a diagram of a complex experimental result provided in an embodiment of the present invention; Figure 5 This is a force estimation result diagram on the X, Y, and Z axes provided by an embodiment of the present invention; Figure 6 This is a tracking performance diagram in joint space provided by an embodiment of the present invention; Figure 7 This is a tracking error map in joint space provided by an embodiment of the present invention; Figure 8 This is a block diagram of a collaborative robot human-computer interaction device based on human motion intention estimation provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of a collaborative robot human-computer interaction device based on human motion intention estimation provided in an embodiment of the present invention. Detailed Implementation
[0015] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0016] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0017] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0018] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0019] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0020] This invention provides a collaborative robot human-computer interaction method based on human motion intention estimation. This method can be implemented by a collaborative robot human-computer interaction device based on human motion intention estimation, which can be a terminal or a server. Figure 1 The flowchart shown is for a collaborative robot human-robot interaction method based on human motion intention estimation. The processing flow of this method may include the following steps:
[0021] S1. Construct a dynamic model of the robot system with human-machine interaction force, and design an Actor-Critic reinforcement learning-based controller.
[0022] The process of constructing the dynamic model of the robot system with human-machine interaction is represented by the following formula (1): (1) in, Represents a position vector. Represents the inertia matrix. Representing the Coriolis matrix and the eccentric matrix, Represents the gravitational matrix. Indicates a robot system; This indicates the input torque.
[0023] In order to describe human control behavior, in physical human-computer interaction tasks, the proportional-derivative (PD) controller is represented by the following formula (2): (2) in, This represents the first control gain matrix; This represents the second control gain matrix; This represents the desired human trajectory in joint space, which is unknown.
[0024] In one feasible implementation, to address model uncertainties and dynamic changes in robot systems, this invention designs a reinforcement learning controller based on the Actor-Critic framework. The Critic neural network primarily provides evaluations based on the current system state and input variables. The Actor neural network adjusts weights from the feedback of the Critic neural network, thereby adjusting the control strategy to better approximate the uncertainties in human-machine interaction systems. This controller can learn online and adapt to system changes, achieving high-performance trajectory tracking control.
[0025] Optionally, the Actor-Critic reinforcement learning controller includes: a Critic neural network and an Actor neural network; Among them, the Critic neural network evaluates the long-term performance of the current control strategy based on the current state variables and input variables of the robot system; Among them, the Actor neural network adjusts the current control strategy through the feedback information of the Critic neural network, and by adjusting the current control strategy, it approximates the uncertain dynamics of the dynamic model in the robot system. The control torque based on the Actor-Critic reinforcement learning controller consists of a feedforward term and a feedback term. The feedforward term is used to compensate for the interaction force and the uncertainty of the robot system, further improving the trajectory tracking accuracy and interaction compliance. The feedback term is used to ensure the stability of the closed-loop robot system and suppress residual disturbances that the feedforward term could not fully compensate for.
[0026] The control torque is expressed by the following formula (3): (3) in, Indicates control torque; Indicates the feedforward term; Indicates a feedforward term.
[0027] In one feasible implementation, the feedforward term is represented by the following formula (4): (4) in, Indicates human-computer interaction force; This represents the approximation value of the Actor neural network for the uncertain dynamics of the system; This represents the position matrix of the system.
[0028] In this embodiment of the invention, the controller is constructed using an Actor-Critic dual neural network structure to achieve online learning and dynamic compensation for system uncertainties.
[0029] In the Actor-Critic framework-based reinforcement learning controller, the Critic neural network acts as the evaluator. Its core task is to assess the long-term performance of the current control policy, i.e., to predict the cumulative value (or cost) over all future time steps if the current control policy is continued in the current state. The Critic network does not directly output control signals; instead, it addresses the problem through a long-term cost function, aiming to minimize the evaluation error. The Critic network's learning does not rely on a pre-known standard answer but rather utilizes online and bootstrapping learning through temporal difference methods.
[0030] The essence of Critic neural network design is to build an intelligent evaluation system capable of learning online and accurately predicting the long-term performance of control strategies. It continuously self-corrects by calculating temporal difference errors, and its output evaluation signal directly drives the strategy optimization of the Actor network. This is the key to enabling the entire reinforcement learning controller to adapt to uncertain environments and achieve high-performance control.
[0031] Optionally, S1 is designed based on an Actor-Critic reinforcement learning controller, including: S11. Define a long-term cost function; based on the long-term cost function, obtain the cost function error; update the weights of the Critic neural network by constructing a minimized temporal difference error; In one feasible implementation, the long-run cost function is expressed by the following formula (5): (5) in, Indicates the discount factor; The instantaneous cost function is represented by h; the integral variable represents a future moment; and t represents the current moment. This represents the discount factor.
[0032] The cost function error is expressed by the following formula (6): (6) in, Indicates the error of the cost function; This represents an estimate of the long-run cost function; Indicates the discount factor; The minimization of the timing difference error is expressed by the following formula (7): (7) in, This represents minimizing the timing difference error; This represents the transpose of the cost function error.
[0033] S12. Based on the updated weights, design the weight update law of the Critic neural network, and output the evaluation signal through the weight update rate. S13. Based on the evaluation signal output by the Critic neural network, design the weight update law of the Actor neural network; construct an Actor-Critic reinforcement learning controller based on the weight update laws of the Critic neural network and the Actor neural network.
[0034] Alternatively, the weight update law of the Critic neural network is expressed by the following formula (8): (8) in, This represents the weight update rate of the Critic neural network; This represents the learning rate of the Critic neural network; Represents the instantaneous cost function; Indicates the output vector; This represents the transpose of the weight vector of the Critic neural network.
[0035] in, ,in, This represents the activation function vector of the Critic neural network; , .
[0036] In the Actor-Critic framework, the goal of the Actor network is to find an optimal policy that minimizes the long-term cost predicted by the Critic network through the control actions generated by that policy. The Actor network updates following the idea of policy gradients. The core idea is: if an action (policy) results in a low long-term cost, the probability of selecting that action is increased; if it results in a high long-term cost, the probability of selecting that action is decreased. In this embodiment of the invention, the signal of high or low long-term cost is the temporal difference error calculated by the Critic network. .
[0037] The update of the Actor neural network is a policy optimization process based on evaluation signals. It utilizes the signals provided by the Critic. By continuously fine-tuning its internal parameters, the control strategy it outputs gradually approaches the optimal strategy that minimizes long-term costs and achieves high-precision tracking and compliant interaction. The continuous "execution-evaluation-learning" cycle between the Actor and Critic gives the entire control system a powerful ability to learn online and adapt to uncertainty.
[0038] Alternatively, the weight update law of the Actor neural network is expressed by the following formula (9): (9) in, Represents the weights of the Actor neural network; This represents the learning rate of the Actor neural network; This represents the activation function vector of the Actor neural network; Represents a positive constant; This represents the evaluation from the Critic neural network.
[0039] S2. Based on the Actor-Critic reinforcement learning controller, output the optimal control strategy; based on the optimal control strategy, approximate the uncertain dynamics of the dynamic model in the robot system online.
[0040] S3. An online estimation algorithm based on recursive least squares is used to estimate the force applied to the human body and output the interaction force between the human and the computer.
[0041] Among them, in the estimation At that time, the main challenge faced by system dynamics methods is the acceleration signal. Because this signal cannot be directly measured in practical applications and is prone to introducing system noise, it poses difficulties for real-time robot control. Although offline filtering methods can solve the above problems, they are difficult to meet the stringent requirements of real-time control systems. Therefore, this invention proposes an observer based on the least squares method for estimating the force applied by a human. The algorithm consists of two stages: first, filtering the system dynamic signal, especially the acceleration signal, to reduce noise interference; second, using RLS to approximate the force applied by the human body.
[0042] Optionally, the specific implementation process of S3 includes S31-S32: S31. Design a filter; Based on the designed filter, filter the dynamic equations of the robot system to obtain the filtered dynamic equations. In one feasible implementation, to avoid affecting the acceleration signal The direct measurement and differentiation of the filter are used to design a filter, which can be expressed by the following formula (10): (10) in, t represents the cutoff frequency of the filter; t represents time.
[0043] In one feasible implementation, the dynamic process is filtered, wherein the filter can be defined by the following formula (11): (11) in, Indicates signal The output after filtering; express Filtering.
[0044] Among them, after filtering both sides of the original dynamic equation, the linear parameterization characteristics of system dynamics are utilized, and mathematical transformations such as integration by parts are performed to finally rewrite the filtered equation as the following formula (12): (12) in, It is a new regression matrix composed of the filtered state signals; This represents the filtered human-computer interaction torque; Represents a vector of physical parameters.
[0045] In one feasible implementation, the recursive least squares-based online estimation algorithm estimates the interaction force between the human operator and the robot in real time and accurately without relying on external force sensors. It calculates this force by processing the robot's internal data, thus achieving sensorless force perception. In this embodiment, it transforms the complex dynamic parameter estimation and force estimation problem into an online optimization problem through recursion and iteration, enabling the robot to perceive human pushing or pulling forces in real time and accurately, as if it had an intrinsic sense, laying a crucial foundation for subsequent active following control. Its greatest advantages are high computational efficiency, suitability for real-time control, and the elimination of the need for expensive force sensors.
[0046] S32. Based on the filtered dynamic equation, a cost function is constructed; an online estimation algorithm based on recursive least squares is adopted to solve the interaction force between human and computer through the cost function, and the interaction force between human and computer is updated in real time through recursion to obtain the final interaction force between human and computer.
[0047] In order to estimate the interaction forces The cost function is constructed as follows (13): (13) in, , It is the forgetting factor; the process of solving the cost function is represented by the following formula (14): (14) in, This represents an estimate of the interaction force.
[0048] In one feasible implementation, the process of updating the interaction force between humans and computers in real time through recursion is represented by the following formula (15): (15) in, The sampling period is It is the gain matrix. It is the covariance matrix. It's about adjusting parameters. Ultimately, through Obtain the estimated value of the joint space interaction torque. .
[0049] In most cases, the advantage of this estimation algorithm lies in its ability to effectively avoid matrix manipulation. sum vector The computational requirement for inversion is eliminated, and it does not require obtaining a vector. The second derivative information. This design not only enhances the practicality of the algorithm but also improves its adaptability in various application scenarios.
[0050] S4. Design the desired trajectory adaptive update law; based on the interaction force between human and machine, adopt the desired trajectory adaptive update law to dynamically adjust the desired joint trajectory of the robot system, so that the desired joint trajectory of the robot actively follows the human's movement intention.
[0051] The adaptive update law of the desired trajectory enables robots to proactively perceive and adapt to human intentions, rather than passively responding as in traditional control. This mechanism adjusts the robot's desired trajectory in real time, actively approaching the human's actual intended trajectory, thereby fundamentally reducing interaction conflicts and achieving smooth and effortless human-robot collaboration. Specifically, based on the estimated direction and magnitude of the interaction force, the robot's desired motion target is dynamically adjusted, thereby reducing motion conflicts between humans and robots at the source and achieving proactive following.
[0052] Alternatively, the adaptive update law of the desired trajectory of S4 is expressed by the following formula (16): (16) in, The derivative of the desired trajectory; Represents a positive definite gain matrix; This represents the estimated human-computer interaction torque.
[0053] Specifically, based on the adaptive update law of the desired trajectory, the desired joint trajectory of the robot is... Actively follow human movement intentions .
[0054] In one feasible implementation, the stability of the robot system is proven by constructing a Lyapunov function.
[0055] The constructed Lyapunov function is expressed by the following formula (17): (17) in, Represents the control system function; The system function represents the force estimation function; This represents the trajectory generation system function; This represents the Lyapunov function.
[0056] in, , as well as This can be expressed by the following formula (18): (18) in, This represents the transpose of the joint position tracking error. This represents the transpose of the filtering tracking error; This represents the transpose of the estimation error of the weights in the Critic neural network. This represents the transpose of the estimation error of the weights in the Actor neural network. This represents the estimation error of the weights in the Critic neural network. This represents the estimation error of the weights in the Actor neural network; Indicates the actual joint positions of the robot; Indicates the joint position expected by humans; This represents a positive definite weight matrix.
[0057] In one feasible implementation, the function described above can be transformed through a series of derivatives to ultimately obtain the result when... When, all errors asymptotically converge to zero; when At that time, by using inequality scaling, it is proved that all errors are uniformly bounded eventually. Using Lyapunov functions, it is rigorously proven that all signals of the closed-loop system are stable under the combined action of the entire controller, including reinforcement learning, force estimation, and trajectory updates, and that the tracking error is confined to a small range.
[0058] The human-computer interaction system platform was built based on the embodiments of the present invention; the experimental platform was selected as the Baxter robot and its corresponding ROS system, such as... Figure 2 As shown, the Baxter robot, manufactured by the American company Rethink Robotics, is a research and experimental platform for dual-arm control, applicable to both academic research and industrial applications. The Baxter robot uses elastic actuators instead of direct motor shaft connections. The motors are connected to the joints via springs, so the torque generated by the motors does not directly act on the links. In the event of a collision, the springs enhance shock absorption to reduce risk. The Baxter robot's head has a 360-degree sonar sensor and a forward-facing camera. Each of the two robotic arms has seven flexible joints, providing full degrees of freedom. The Baxter robot features an electric parallel gripper and a mobile base, serving as a primary experimental platform for human-robot interaction control. The Baxter robot uses the ROS-based open-source robot operating system, running on a Linux platform. Users can connect to the robot's internal computer via a network to read information or send commands. The host computer control system uses Ubuntu 20.04. The computer control platform and the physical robot are connected to the same IP address via USB and a router. The entire robot interaction system framework is based on the ROS noetic version of the robot operating system. The Baxter robot can sense collisions, making it an excellent choice for human-robot interaction control.
[0059] In one feasible implementation, trajectory adaptive testing is conducted based on the established platform. First, the parameters of the control system are set as follows: and The robot's initial position is set to... The desired location for humans is...
[0060] .
[0061] in, Figure 3 The results are shown in the comparison with traditional neural network methods. The control group demonstrates that the rise time and settling time of the Active Reinforcement Learning (RLP) method are shorter than those of traditional neural network methods. Therefore, the RLP method proposed in this embodiment enables the robot's actual position to converge quickly to the desired position.
[0062] The desired trajectory of a person is represented as an ellipse as follows: .
[0063] Among them, such as Figure 4The figure shown is a complex experimental result diagram provided by an embodiment of the present invention; it can be seen from the result diagram that the actual position of the robot can converge to the human's expected trajectory, with an average error of less than 0.02 meters.
[0064] Among them, such as Figure 5 The figure shows a force estimation result on the X, Y and Z axes provided by an embodiment of the present invention. Figure 5 The comparison between the estimated interaction forces and the actual interaction forces in the X, Y, and Z directions using the recursive least squares method is shown. The online estimation algorithm based on recursive least squares ensures that the estimated force values accurately track the actual force values, with an average error of only 0.002 N.
[0065] In one feasible implementation, joint space tracking tests are conducted based on the constructed platform. The control parameters are set as follows: and The initial joint position is... The execution network and evaluation network have 256 and 64 nodes, respectively. The radial basis function centers are located at... ,width Learning rate and The values are 1000 and 0.01 respectively. The cost function parameters are... and .like Figure 6 and Figure 7 The tracking performance and error in joint space are shown, with the x-axis representing time and the y-axis representing radians. The results demonstrate that the RLP controller effectively drives the joint motors, enabling the robot to accurately execute the actions expected by the human. The minimal error indicates a high degree of consistency between the robot's actual movement and the human's expectations, while the system exhibits no sustained oscillations or divergences.
[0066] Traditional physical human-computer interaction systems heavily rely on external force sensors to perceive human intentions, and most existing control methods (such as impedance control) are passively reactive, meaning the robot only moves after being subjected to force, leading to choppy interaction and a heavy workload for the operator. Furthermore, the robot's dynamics model contains uncertainties, affecting the accuracy and robustness of control. This invention proposes a sensorless, reinforcement learning-based active control framework that enables the robot to estimate human intentions online and actively predict and follow those intentions, while compensating for the uncertainties in the system model.
[0067] This invention presents a sensorless intent estimator based on a recursive least squares online estimation algorithm and a human control model. As an interactive force estimation device for physical human-computer interaction systems, it achieves high-precision interactive force estimation without the need for a force sensor, laying the foundation for subsequent active control.
[0068] This invention presents a dynamic uncertainty compensator based on Actor-Critic reinforcement learning. A dual neural network structure achieves high-precision tracking and robustness. By handling model uncertainty, the system achieves ideal tracking accuracy on real robots.
[0069] The embodiments of the present invention integrate an active control framework and a controller structure, enabling a synergistic effect through a specific combination of core modules that achieve active, sensorless, high-precision, and robust physical human-computer interaction.
[0070] This invention successfully estimates human-robot interaction forces online from the robot's internal state information and torque input using an intention estimator based on a recursive least squares online estimation algorithm and a human control model. This achieves accurate human-robot interaction without the need for external force sensors, significantly reducing system cost and complexity.
[0071] After adopting the embodiments of the present invention, the robot can proactively predict the human's movement intentions and start moving in advance. This makes the interaction process more natural and smooth, just like collaboration between people.
[0072] Even when the robot model is inaccurate or there are unmodeled dynamics, the embodiments of this invention can achieve high-precision trajectory tracking. It exhibits strong robustness to uncertainties in the robot system's dynamics, ensuring high-precision trajectory tracking performance.
[0073] Figure 8 This is a block diagram of a collaborative robot human-computer interaction device based on human motion intention estimation, provided by an embodiment of the present invention. This device is used in a collaborative robot human-computer interaction method based on human motion intention estimation. (Refer to...) Figure 8 The device includes a construction unit 810, an output unit 820, an estimation unit 830, and an adjustment unit 840. Wherein:
[0074] Building unit 810 is used to build a dynamic model of a robot system with human-robot interaction force and to design an Actor-Critic reinforcement learning-based controller. Output unit 820 is used to output the optimal control strategy based on the Actor-Critic reinforcement learning controller; and to approximate the uncertain dynamics of the dynamic model of the robot system online based on the optimal control strategy. The estimation unit 830 is used to estimate the force applied by the human body using an online estimation algorithm based on recursive least squares, and output the interaction force between the human and the computer. The adjustment unit 840 is used to design the adaptive update law of the desired trajectory. Based on the interaction force between human and machine, the desired joint trajectory of the robot system is dynamically adjusted according to the adaptive update law of the desired trajectory, so that the desired joint trajectory of the robot actively follows the human's movement intention.
[0075] Optionally, the Actor-Critic-based reinforcement learning controller includes: a Critic neural network and an Actor neural network; Among them, the Critic neural network evaluates the long-term performance of the current control strategy based on the current state variables and input variables of the robot system; Among them, the Actor neural network adjusts the current control strategy through the feedback information of the Critic neural network, and by adjusting the current control strategy, it approximates the uncertain dynamics of the dynamic model in the robot system. The control torque based on the Actor-Critic reinforcement learning controller consists of a feedforward term and a feedback term. The feedforward term is used to compensate for the interaction force and the uncertainty of the robot system, further improving the trajectory tracking accuracy and interaction compliance. The feedback term is used to ensure the stability of the closed-loop robot system and suppress residual disturbances that the feedforward term could not fully compensate for.
[0076] Optionally, the building unit is used for: Define a long-term cost function; obtain the cost function error based on the long-term cost function; update the weights of the Critic neural network by constructing a minimized temporal difference error; Based on the updated weights, a weight update law for the Critic neural network is designed, and an evaluation signal is output through the weight update rate. Based on the evaluation signal output by the Critic neural network, the weight update law of the Actor neural network is calculated; based on the weight update laws of the Critic neural network and the Actor neural network, an Actor-Critic reinforcement learning controller is constructed.
[0077] Optionally, the weight update law of the Critic neural network is expressed by the following formula (1): (1) in, This represents the weight update rate of the Critic neural network; Represents the learning law; Represents the instantaneous cost function; Indicates the output vector; This represents the transpose of the weight vector of the Critic neural network.
[0078] Optionally, the weight update law of the Actor neural network is expressed by the following formula (2): (2) in, Represents the weights of the Actor neural network; This represents the learning rate of the Actor neural network; This represents the activation function vector of the Actor neural network; Represents a positive constant; This represents the evaluation from the Critic neural network.
[0079] Optionally, the estimation unit 830 is used for: Design a filter; use the designed filter to filter the dynamic equations of the robot system to obtain the filtered dynamic equations. Based on the filtered dynamic equations, a cost function is constructed. An online estimation algorithm based on recursive least squares is used to solve for the interaction force between humans and computers through the cost function. The interaction force between humans and computers is updated in real time through recursion to obtain the final interaction force between humans and computers.
[0080] Optionally, the adaptive update law of the desired trajectory is expressed by the following formula (3): (3) in, The derivative of the desired trajectory; Represents a positive definite gain matrix; This represents the estimated human-computer interaction torque.
[0081] This invention successfully estimates human-robot interaction forces online from the robot's internal state information and torque input using an intention estimator based on a recursive least squares online estimation algorithm and a human control model. This achieves accurate human-robot interaction without the need for external force sensors, significantly reducing system cost and complexity.
[0082] After adopting the embodiments of the present invention, the robot can proactively predict the human's movement intentions and start moving in advance. This makes the interaction process more natural and smooth, just like collaboration between people.
[0083] Even when the robot model is inaccurate or there are unmodeled dynamics, the embodiments of this invention can achieve high-precision trajectory tracking. It exhibits strong robustness to uncertainties in the robot system's dynamics, ensuring high-precision trajectory tracking performance.
[0084] Figure 9 This is a schematic diagram of the structure of a collaborative robot human-computer interaction device based on human motion intention estimation provided in an embodiment of the present invention, as shown below. Figure 9 As shown, collaborative robot human-computer interaction devices based on human motion intention estimation can include the above-mentioned... Figure 8 The illustrated collaborative robot human-machine interaction device is based on human motion intention estimation. Optionally, the collaborative robot human-machine interaction device 410 based on human motion intention estimation may include a first processor 2001.
[0085] Optionally, the collaborative robot human-computer interaction device 410 based on human motion intention estimation may also include a memory 2002 and a transceiver 2003.
[0086] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.
[0087] The following is combined with Figure 9 The various components of the collaborative robot human-computer interaction device 410 based on human motion intention estimation are described in detail below: The first processor 2001 is the control center of the collaborative robot human-machine interaction device 410 based on human motion intention estimation. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0088] Optionally, the first processor 2001 can execute various functions of the collaborative robot human-machine interaction device 410 based on human motion intention estimation by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0089] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 9 CPU0 and CPU1 are shown in the diagram.
[0090] In a specific implementation, as one example, the collaborative robot human-computer interaction device 410 based on human motion intention estimation may also include multiple processors, for example... Figure 9The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0091] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0092] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected via the interface circuit of the collaborative robot human-machine interaction device 410 based on human motion intention estimation. Figure 9 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0093] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0094] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 9 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0095] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and can be connected to the interface circuit of the collaborative robot human-machine interaction device 410 based on human motion intention estimation. Figure 9(Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0096] It should be noted that, Figure 9 The structure of the collaborative robot human-machine interaction device 410 based on human motion intention estimation shown in the figure does not constitute a limitation on the router. Actual collaborative robot human-machine interaction devices based on human motion intention estimation may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0097] Furthermore, the technical effects of the collaborative robot human-computer interaction device 410 based on human motion intention estimation can be referred to the technical effects of the collaborative robot human-computer interaction method based on human motion intention estimation described in the above method embodiments, and will not be repeated here.
[0098] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or it may be any conventional processor, etc.
[0099] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0100] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0101] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0102] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0103] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0104] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0105] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0106] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0107] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0108] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0109] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0110] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for human-robot interaction based on human motion intention estimation, characterized in that, The method includes: S1. Construct a dynamic model of a robot system with human-machine interaction force, and design an Actor-Critic reinforcement learning-based controller. The Actor-Critic reinforcement learning controller includes: a Critic neural network and an Actor neural network; Among them, the Critic neural network evaluates the long-term performance of the current control strategy based on the current state variables and input variables of the robot system; Among them, the Actor neural network adjusts the current control strategy through the feedback information of the Critic neural network, and by adjusting the current control strategy, it approximates the uncertain dynamics of the dynamic model in the robot system. The control torque based on the Actor-Critic reinforcement learning controller consists of a feedforward term and a feedback term. The feedforward term is used to compensate for the interaction force and the uncertainty of the robot system, further improving the trajectory tracking accuracy and interaction compliance. The feedback term is used to ensure the stability of the closed-loop robot system and suppress residual disturbances that the feedforward term could not fully compensate for. The design of S1 is based on an Actor-Critic reinforcement learning controller, including: S11. Define a long-term cost function; based on the long-term cost function, obtain the cost function error; update the weights of the Critic neural network by constructing a minimized temporal difference error; The long-run cost function is expressed by the following formula (1): (1) in, Indicates the discount factor; The instantaneous cost function is represented by h; the integral variable represents a future moment; and t represents the current moment. Indicates the discount factor; The cost function error is expressed by the following formula (2): (2) in, Indicates the error of the cost function; This represents an estimate of the long-run cost function; Indicates the discount factor; The minimization of the timing difference error is expressed by the following formula (3): (3) in, This represents minimizing the timing difference error; This represents the transpose of the cost function error; S12. Based on the updated weights, design the weight update law of the Critic neural network and output the evaluation signal through the weight update rate. S13. Based on the evaluation signal output by the Critic neural network, calculate the weight update law of the Actor neural network; construct an Actor-Critic reinforcement learning controller based on the weight update law of the Critic neural network and the weight update law of the Actor neural network. S2. Based on the Actor-Critic reinforcement learning controller, output the optimal control strategy; and approximate the uncertain dynamics of the robot system's dynamic model online based on the optimal control strategy. S3. An online estimation algorithm based on recursive least squares is used to estimate the force applied by the human body and output the interaction force between the human and the computer. S4. Design the desired trajectory adaptive update law; based on the interaction force between human and machine, adopt the desired trajectory adaptive update law to dynamically adjust the desired joint trajectory of the robot system, so that the desired joint trajectory of the robot actively follows the human's movement intention. Among them, according to the expected trajectory adaptive update law, the robot's expected joint trajectory actively follows the human's movement intention; The adaptive update law of the expected trajectory of S4 is expressed by the following formula (4): (4) in, The derivative of the desired trajectory; Represents a positive definite gain matrix; This represents the estimated human-computer interaction torque.
2. The collaborative robot human-computer interaction method based on human motion intention estimation according to claim 1, characterized in that, The weight update law of the Critic neural network is expressed by the following formula (5): (5) in, This represents the weight update rate of the Critic neural network; Represents the learning law; Represents the instantaneous cost function; Indicates the output vector; This represents the transpose of the weight vector of the Critic neural network.
3. The collaborative robot human-computer interaction method based on human motion intention estimation according to claim 1, characterized in that, The weight update law of the Actor neural network is expressed by the following formula (6): (6) in, Represents the weights of the Actor neural network; This represents the learning rate of the Actor neural network; This represents the activation function vector of the Actor neural network function; Represents a positive constant; This represents the evaluation from the Critic neural network.
4. The collaborative robot human-computer interaction method based on human motion intention estimation according to claim 1, characterized in that, The S3 employs a recursive least squares online estimation algorithm to estimate the force applied by the human body and outputs the interaction force between the human and the computer, including: S31. Design a filter; Based on the designed filter, filter the dynamic equations of the robot system to obtain the filtered dynamic equations. In order to avoid affecting the acceleration signal The direct measurement and differentiation of the filter are used to design a filter, which can be expressed by the following formula (7): (7) in, The cutoff frequency of the filter is represented by t; time is represented by t. The dynamics of the robot system are filtered, as expressed by the following formula (8): (8) in, Indicates signal The output after filtering; express Filtering; Among them, after filtering both sides of the original dynamic equation, the linear parameterization characteristics of the system dynamics are utilized, and through the mathematical transformation of integral by parts, the filtered dynamic equation is finally rewritten as the following formula (9): (9) in, It is a new regression matrix composed of the filtered state signals; This represents the filtered human-computer interaction torque; Represents a vector of physical parameters; S32. Based on the filtered dynamic equation, a cost function is constructed; an online estimation algorithm based on recursive least squares is adopted to solve the interaction force between human and computer through the cost function, and the interaction force between human and computer is updated in real time through recursion to obtain the final interaction force between human and computer.
5. A collaborative robot human-computer interaction device based on human motion intention estimation, wherein the collaborative robot human-computer interaction device based on human motion intention estimation is used to implement the collaborative robot human-computer interaction method based on human motion intention estimation as described in any one of claims 1-4, characterized in that, The device includes: Building units are used to construct dynamic models of robot systems with human-robot interaction forces and to design Actor-Critic reinforcement learning-based controllers. The output unit is used to output the optimal control strategy based on the Actor-Critic reinforcement learning controller; and to approximate the uncertain dynamics of the dynamic model of the robot system online based on the optimal control strategy. The estimation unit is used to estimate the force applied by the human body using an online estimation algorithm based on recursive least squares, and output the interaction force between the human and the computer. An adjustment unit is used to design the adaptive update law of the desired trajectory. Based on the interaction force between human and machine, the desired joint trajectory of the robot system is dynamically adjusted according to the adaptive update law of the desired trajectory, so that the desired joint trajectory of the robot actively follows the human's movement intention.
6. A collaborative robot human-computer interaction device based on human motion intention estimation, characterized in that, The collaborative robot human-computer interaction device based on human motion intention estimation includes: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 4.