A human-machine shared tracking control method based on differential game

By adopting a human-machine shared tracking control method based on differential game theory, the problems of mutual adaptability and safety in human-machine systems are solved, enabling robots to autonomously avoid obstacles and cooperate in a safe environment. This method is applicable to a variety of robot systems.

CN116774628BActive Publication Date: 2025-11-18UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310678022.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-07
Publication Date
2025-11-18
Estimated Expiration
2043-06-07

AI Technical Summary

Technical Problem

Existing human-machine system models fail to fully consider the mutual adaptability between human operators and robots, and the control obstacle function is insufficient in terms of safety, failing to effectively adjust autonomously when unsafe factors occur to ensure the safety and collaborative operation of the human-machine system.

Method used

A human-machine shared tracking control method based on differential game theory is adopted. By constructing a human-machine differential game model, a robot shared control strategy is designed. Combining minimum control energy and control obstacle function, the robot can track the desired trajectory in a safe environment and autonomously avoid obstacles when they appear.

Benefits of technology

It effectively depicts the dynamic negotiation process between humans and machines. The robot's shared control strategy can adaptively adjust the control energy to ensure collaborative operation of the human-machine system in a safe environment. It can also autonomously avoid obstacles when they appear, ensuring system safety. It is applicable to any robot system that can be modeled as an affine nonlinear system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116774628B_ABST
    Figure CN116774628B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of robot tracking control, and discloses a kind of man-machine shared tracking control method based on differential game, comprising: establishing man-machine differential game model;Estimate human control strategy and objective function parameter;Design robot shared control strategy;To the shared control strategy determined by robot, realize man-machine shared tracking control;The present application uses differential dynamic game to model the dynamic negotiation relationship between human participants and robot, so that the robot adjusts its own willingness under the game framework by estimating the willingness of human to eliminate tracking error, and designs robot tracking control strategy on this basis;Further, the present application considers the safety of man-machine cooperation, designs robot safety control strategy based on control barrier function, so that it can realize safe shared control together with robot tracking control strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot tracking and control technology, specifically to a human-machine shared tracking and control method based on differential game theory. Background Technology

[0002] With the development of intelligent technologies such as vision, 5G communication, and AI, robots are increasingly being applied in smart factories, logistics warehousing, and other scenarios to assist human operators in completing complex or demanding tasks. The key to achieving natural and efficient human-robot interaction lies in fully leveraging humans' natural advantages in perception and decision-making, as well as robots' ability to repeatedly execute high-precision control tasks. To this end, robots need to recognize the human operator's movement intentions and provide timely and appropriate assistance accordingly. One of the key scientific issues in this process is establishing the interactive relationship between humans and robots. Most existing human-machine system models and methods do not fully consider the mutual adaptability between human operators and robots, and cannot accurately depict the dynamic negotiation relationship between them.

[0003] Furthermore, in human-robot shared workspaces, the safety of human-robot collaboration is a cutting-edge and hot topic in human-robot collaborative control. Due to potential human negligence and the possibility of emergencies in the work environment, human operators' commands are not always safe. In such cases, it is desirable for robots to autonomously identify environmental unsafety and temporarily disobey human commands to ensure human-robot safety. Control Barrier Functions (CBFs) are widely used to solve obstacle avoidance problems in robotic systems, but related research has not fully considered responses to human operators. How to design a highly scalable continuous safety controller that functions only when unsafe factors occur, without affecting human-robot collaborative operations in safe environments, is a key issue worthy of further research. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a human-machine shared tracking control method based on differential game theory.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0006] A human-machine shared tracking control method based on differential game theory includes the following steps:

[0007] Step A: Establish a human-machine differential game model:

[0008] A two-degree-of-freedom robotic arm model with physical contact and interaction with humans is constructed to obtain a human-machine system model. The human-machine system model is an affine nonlinear system with two control inputs, namely the human control input u1(t) and the robot control input u2(t) in the human-machine system model.

[0009] The augmented system state ξ(t) is defined by the desired bounded reference trajectory and the robot's trajectory tracking error at time t. Based on the human-machine system model and the evolution of the reference trajectory, a human-machine augmented system model is constructed.

[0010] The objective function J1(ξ(0),u1) for humans is constructed with the goal of achieving trajectory tracking with minimum control energy, and the objective function J2(ξ(0),u2) for robots is constructed with the goal of assisting humans in achieving trajectory tracking with minimum control energy.

[0011] Step B, Estimate the human control policy and objective function parameters: Based on the properties of the Nash equilibrium policy, the estimated human control policy is... Modeling and through Estimate the parameter θ1 of the human objective function;

[0012] Step C, Design the robot shared control strategy: Design the robot tracking control strategy based on the optimal control principle. 2,t (t), design a robot safety control strategy u based on a suitable control barrier function. 2,s Combining the robot tracking control strategy and the safety control strategy, we obtain the robot shared control strategy u2(t)=u 2,t (t)+u 2,s (t), the robot shared control strategy is the robot's control input in the human-machine system model;

[0013] Step D: Apply the shared control strategy u2(t) determined in step C to the robot to achieve human-robot shared tracking control.

[0014] Furthermore, the two-degree-of-freedom robotic arm includes two links connected in series, with the link farther from the end effector being link one and the link closer to the end effector being link two; in step A, when constructing a two-degree-of-freedom robotic arm model that has physical contact and interaction with humans to obtain the human-machine system model:

[0015] The human-machine system model is as follows:

[0016]

[0017] In the formula, q(t) = [q1(t)q2(t)] T Let q1(t) represent the angular position vector of the robotic arm at time t, and let q2(t) represent the angular position of link one at time t. This represents the angular velocity vector of the robotic arm at time t. This represents the time derivative of q1(t). This represents the time derivative of q2(t). Let represent the angular acceleration vector of the robotic arm at time t, and M(q(t)) represent the inertia matrix of the robotic arm at time t. Let G(q(t)) represent the centrifugal force and Coriolis force matrix of the robotic arm at time t, G(q(t)) represent the gravitational effect matrix of the robotic arm at time t, τ(t) represent the torque of the robotic arm at time t, and f(t) represent the interaction force between the human and the robotic arm at time t.

[0018] make The augmented vectors representing the angular position and angular velocity of the robotic arm are used to construct the human-machine system model as follows:

[0019]

[0020] In the formula, This represents the time derivative of z(t). This represents the dynamics of the human-machine system model shift. This represents the input dynamics of a human-machine system. In the human-machine system model, the human's control input is u1(t) = f(t), and the robot's control input is u2(t) = τ(t). M -1 (q(t)) denotes the inverse matrix of M(q(t)).

[0021] Furthermore, in step A, a human-machine augmentation system model is constructed. hour:

[0022] Use z d (t) represents the desired bounded reference trajectory, z d The derivative of (t) satisfy:

[0023]

[0024] In the formula Indicate z d The time derivative of f(t), d (z d (t) is about z d A Lipschitz continuous function of (t) and satisfying f d (0) = 0;

[0025] Define the state of the augmented system as ξ(t) = [e(t)] T z d (t) T ] T Where e(t) = z(t) - z d (t) represents the robot's trajectory tracking error at time t. Based on the human-machine system model and the evolution of the reference trajectory, the human-machine augmented system model is... Build as:

[0026]

[0027] in The first derivative of ξ(t) is represented by... This represents the offset dynamics of the human-machine augmentation system model. This represents the input dynamics of the human-machine augmentation system.

[0028] Furthermore, in step A, when constructing the human objective function J1(ξ(0),u1) with the goal of achieving trajectory tracking with minimum control energy:

[0029]

[0030] In the formula, γ∈(0,1) represents the discount factor, and θ1 represents the human willingness to eliminate tracking errors. Let θ1 denote the transpose of θ1, and ψ(t) denote the pairwise matrix ξ(t). T The vector obtained by vectorizing the upper half matrix.

[0031] Furthermore, in step A, when constructing the robot's objective function J2(ξ(0),u2) with the goal of minimizing control energy to assist humans in achieving trajectory tracking:

[0032]

[0033] In the formula, γ∈(0,1) represents the discount factor, θ2 represents the robot's willingness to eliminate tracking errors, θ2=θ-θ1, θ1 represents the human's willingness to eliminate tracking errors, and θ is a constant vector set in advance according to the complexity of the task.

[0034] Furthermore, step B specifically includes:

[0035] Step B1, Estimate human control strategy:

[0036] Based on the properties of the Nash equilibrium policy, the estimated human control policy is modeled as follows: Where G(ξ(t)) T Denotes the transpose of G(ξ(t)). Let Φ(ξ(t)) denote the gradient of Φ(ξ(t)) with respect to ξ(t), and let Φ(ξ(t)) denote the activation function of the neural network. This represents the weights of a human actor neural network, and its update rate. for:

[0037]

[0038] Where β>0 represents the learning rate of the neural network weights for the human control strategy. This represents the estimated state of the human-machine augmentation system model. The evolutionary process satisfies: in This represents the offset dynamics of the human-machine augmented system model under the estimated human control strategy. Let Λ represent the augmented vector of the estimated angular position and angular velocity of the robotic arm under the human control strategy, where Λ is a positive definite constant matrix.

[0039] Step B2, estimate the parameters of the human objective function:

[0040] use This represents an estimate of the parameter θ1 of the human objective function. update rate for:

[0041]

[0042] In the formula, α1>0 represents the learning rate of the human objective function parameters, and ψ(t) represents the pairwise matrix ξ(t). T The vector obtained by vectorizing the upper half matrix.

[0043] Furthermore, step C specifically includes:

[0044] Step C1: Design the robot tracking control strategy;

[0045] Shared control strategy for robots 2,t (t) is designed as follows:

[0046]

[0047] In the formula This represents the weights of the robot actor neural network, with an update rate of:

[0048]

[0049] In the formula α3>0 indicates the first learning rate for the robot actor neural network weights, and α4>0 indicates the second learning rate for the robot actor neural network weights. This represents the weights of the robot's critic neural network. update rate satisfy:

[0050]

[0051] Where α2>0 represents the learning rate of the robot's critic neural network weights, θ1 represents the human's willingness to eliminate tracking errors, and ψ(t) represents the pairing matrix ξ(t). T The vector obtained by vectorizing the upper half matrix;

[0052] Step C2, Design robot safety control strategy:

[0053] Set a safety set C = {ξ(t), |h(ξ(t))>0}, where h(ξ(t)) represents the collision function, and I represents a four-dimensional identity matrix, and [II] represents a 4×8 matrix formed by combining two four-dimensional identity matrices; z ob (t) represents the obstacle position x at time t. ob The augmented vector r corresponding to (t) h For the safety radius, ||·|| 2 The square of the vector's second norm is used to define the robot safety control strategy u. 2,s (t) is designed as follows:

[0054]

[0055] Where K(ξ(t)) is the gain matrix that satisfies G(ξ(t))K(ξ(t)) with the smallest eigenvalue greater than 0. Let Π(ξ(t)) represent the gradient of ξ(t). in

[0056] Step C3: Combining the robot tracking control strategy and the safety control strategy, we obtain the robot shared control strategy u2(t): u2(t) = u 2,t (t)+u 2,s (t).

[0057] Compared with the prior art, the beneficial technical effects of the present invention are:

[0058] 1. This invention uses differential game theory to effectively characterize the human-machine dynamic negotiation process. The designed robot shared control strategy adjusts its own objective function parameters by estimating the human objective function parameters, and can adaptively adjust its own control energy according to the human control energy.

[0059] 2. The robot shared control strategy designed in this invention can track the desired trajectory in a safe environment and autonomously avoid obstacles when they appear, thus ensuring the safety of the human-machine system.

[0060] 3. The shared tracking control strategy designed in this invention is applicable to any robot system that can be modeled as an affine nonlinear system, and can be applied to a variety of scenarios, with wide applicability. Attached Figure Description

[0061] Figure 1 This is a flowchart of the control method in this invention. Detailed Implementation

[0062] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.

[0063] This embodiment presents a human-robot shared tracking control method based on differential game theory. It utilizes differential dynamic game theory to model the dynamic negotiation relationship between human participants and the robot. Within the game framework, the robot adjusts its own intentions by estimating the human's willingness to eliminate tracking errors, and designs a robot tracking control strategy based on this. Furthermore, this invention considers the safety of human-robot collaboration, designing a robot safety control strategy based on a control barrier function, enabling it to work together with the robot tracking control strategy to achieve safe shared control. The specific implementation process is as follows: Figure 1 As shown, the specific steps include the following.

[0064] Step A, establish a human-machine differential game model, specifically including:

[0065] Step A1, Establish a human-machine system model:

[0066] The two-degree-of-freedom robotic arm model that has physical contact and interaction with humans, i.e., the human-machine system model, is constructed as follows:

[0067]

[0068] In the formula, q(t) = [q1(t)q2(t)] T Let q1(t) represent the angular position vector of the robotic arm at time t, and let q2(t) represent the angular position of link one at time t. This represents the angular velocity vector of the robotic arm at time t. This represents the time derivative of q1(t). This represents the time derivative of q2(t). Let represent the angular acceleration vector of the robotic arm at time t, and M(q(t)) represent the inertia matrix of the robotic arm at time t. Let C represent the centrifugal force and Coriolis force matrix of the robotic arm at time t, G(q(t)) represent the gravitational effect matrix of the robotic arm at time t, τ(t) represent the torque of the robotic arm at time t, and f(t) represent the interaction force between the human and the robotic arm at time t. C actually comprises two parts: centrifugal force and Coriolis force, and is usually represented by a single C.

[0069] make The augmented vectors representing the angular position and angular velocity of the robotic arm, representing the two-degree-of-freedom robotic arm model that has physical contact and interaction with a human, are used to construct the human-machine system model as follows:

[0070]

[0071] In the formula, This represents the time derivative of z(t). This represents the dynamics of the human-machine system model shift. This represents the input dynamics of a human-machine system. In the human-machine system model, the human's control input is u1(t) = f(t), and the robot's control input is u2(t) = τ(t). M -1 (q(t)) denotes the inverse matrix of M(q(t)).

[0072] In this step, the human-machine system is modeled as an affine nonlinear system with two control inputs. In subsequent steps, the human-machine system is analyzed and designed. Therefore, the method proposed in this invention is applicable to a variety of robot systems, as long as the corresponding human-machine system can be modeled as an affine nonlinear system of the above form.

[0073] Step A2, establish a human-machine augmentation system model:

[0074] Use z d (t) represents the desired bounded reference trajectory, z d The derivative of (t) satisfy:

[0075]

[0076] In the formula Indicate z d The time derivative of f(t), d (z d (t) is about z d (t) is a Lipschitz continuous function that satisfies f d (0) = 0.

[0077] Define the state of the augmented system as ξ(t) = [e(t)] T z d (t) T ] T Where e(t) = z(t) - z d (t) represents the robot's trajectory tracking error at time t. Based on the human-machine system model and the evolution of the reference trajectory, the human-machine augmented system model is... Build as:

[0078]

[0079] in The first derivative of ξ(t) is represented by... This represents the offset dynamics of the human-machine augmentation system model. This represents the input dynamics of the human-machine augmentation system.

[0080] Step A3, establish the human-machine objective function:

[0081] In this invention, it is assumed that both human and robot behaviors are rational. The human goal is to achieve trajectory tracking with minimal control energy, while the robot's goal is to assist the human in achieving better trajectory tracking with minimal control energy. Based on this, the human objective function is considered to be:

[0082]

[0083] In the formula, γ∈(0,1) represents the discount factor, and θ1 represents the human willingness to eliminate tracking errors. Let θ1 denote the transpose of θ1, and ψ(t) denote the pairwise matrix ξ(t). T The vector obtained by vectorizing the upper half matrix. Note that the robot cannot directly observe the human control input u1(t) and its objective function parameters θ1.

[0084] Consider the robot's objective function as:

[0085]

[0086] In the formula, θ2 represents the robot's intention to eliminate tracking errors, and its value is θ2 = θ - θ1, where θ is a constant vector set in advance according to the complexity of the task.

[0087] Since humans and robots are physically coupled in the human-machine system model and are reflected in the human-machine objective function through the augmented system state ξ(t), the human-machine augmented system model established in step A2 and the human-machine objective function established in step A3 constitute a differential game.

[0088] Step B involves estimating the human control strategy and objective function parameters, specifically including:

[0089] Step B1, Estimate human control strategy:

[0090] Based on the properties of the Nash equilibrium policy, the estimated human control policy is modeled as follows: Where G(ξ(t)) T Denotes the transpose of G(ξ(t)). Let Φ(ξ(t)) denote the gradient of Φ(ξ(t)) with respect to ξ(t), and let Φ(ξ(t)) denote the activation function of the neural network. This represents the weights of the human actor neural network, with an update rate of:

[0091]

[0092] Where β>0 represents the learning rate of the neural network weights for the human control strategy. The estimated state of the human-machine augmentation system is expressed, and its evolution process satisfies: in This represents the estimated offset dynamics of the human-machine augmented system under the human control strategy. Let Λ represent the augmented vector of the robotic arm's angular position and angular velocity under the estimated human control strategy, where Λ is a positive definite constant matrix.

[0093] Step B2, estimate the parameters of the human objective function:

[0094] use This represents an estimate of the parameter θ1 of the human objective function. update rate for:

[0095]

[0096] In the formula, α1>0 represents the learning rate of the human objective function parameters.

[0097] Step C, design the robot's shared control strategy, which includes the following steps:

[0098] The robot control strategy needs to achieve two goals: 1) to respond optimally to human behavior when there are no obstacles (safety); 2) to reduce the priority of human goals and take obstacle avoidance measures autonomously when unsafe factors that are not perceived by humans appear in the working environment, so as to ensure human-machine safety, and to continue to respond optimally to human behavior after the unsafe factors disappear.

[0099] Step C1, Design the robot tracking control strategy:

[0100] The robot shared control strategy is designed as follows:

[0101]

[0102] In the formula This represents the weights of the robot actor neural network, with an update rate of:

[0103]

[0104] In the formula α3>0 indicates the first learning rate for the robot actor neural network weights, and α4>0 indicates the second learning rate for the robot actor neural network weights. The weights of the robot's critic neural network are represented by the update rate that satisfies:

[0105]

[0106] Where α2>0 represents the learning rate of the robot's critic neural network weights.

[0107] Step C2, Design robot safety control strategy:

[0108] Set a safety set C = {ξ(t), |h(ξ(t))>0}, where h(ξ(t)) represents the collision function, and z ob (t) represents the obstacle position x at time t. ob The augmented vector r corresponding to (t) h To determine the safety radius, the robot safety control strategy is designed as follows:

[0109]

[0110] Where K(ξ(t)) is the gain matrix that satisfies G(ξ(t))K(ξ(t)) with the smallest eigenvalue greater than 0. Let Π)ξ(t) represent the gradient of Π)ξ(t) with respect to ξ(t). in

[0111] Step C3: Combining the robot tracking control strategy and the safety control strategy, we obtain the robot shared control strategy, which has the form: u2(t) = u 2,t (t)+u 2,s (t).

[0112] Step D: Apply the shared control strategy u2(t) determined in step C to the robot to achieve human-robot shared tracking control.

[0113] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.

[0114] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A human-machine shared tracking control method based on differential game theory, comprising the following steps: Step A: Establish a human-machine differential game model: A two-degree-of-freedom robotic arm model with physical contact and interaction with humans is constructed to obtain a human-machine system model. The human-machine system model is an affine nonlinear system with two control inputs, namely the human's control input in the human-machine system model. Control input of the robot in the human-machine system model A two-degree-of-freedom robotic arm consists of two links connected in series. The link furthest from the end effector is link one, and the link closest to the end effector is link two. The human-machine system model is as follows: In the formula Indicates robotic arm The angular position vector at time t. Indicates that link one is in Angular position at time, Indicates that link two is in Angular position at time, Indicates that the robotic arm is in Angular velocity vector at time t, express Time derivative, express Time derivative, Indicates that the robotic arm is in The angular acceleration vector at time t. Indicates that the robotic arm is in The inertia matrix at time t, Indicates that the robotic arm is in The centrifugal force and Coriolis force matrix at time t. Indicates that the robotic arm is in The gravitational effect matrix at time t, Indicates in The constant torque of the robotic arm Indicates the interaction between the human and the robotic arm The power of interaction in time; The augmented vectors representing the angular position and angular velocity of the robotic arm are used to construct the human-machine system model as follows: In the formula, express Time derivative, This represents the dynamics of the human-machine system model shift. This represents the input dynamics of a human-machine system, specifically the human control input in the human-machine system model. Control input of the robot in the human-machine system model , express The inverse matrix; The augmented system state is defined by the desired bounded reference trajectory and the robot's trajectory tracking error at time t. Based on the human-machine system model and the evolution of the reference trajectory, a human-machine augmentation system model is constructed. : use Represents the expected bounded reference trajectory. derivative satisfy: In the formula express Time derivative, For about A Lipschitz continuous function that satisfies Define the augmented system state as ,in Represents robots The trajectory tracking error at any given time is determined based on the human-machine system model and the evolution of the reference trajectory, thus affecting the human-machine augmented system model. Build as: ;in express The first derivative, This represents the offset dynamics of the human-machine augmentation system model. This represents the input dynamics of a human-machine augmentation system; Construct a human objective function with the goal of achieving trajectory tracking with minimum control energy. : In the formula, Indicates the discount factor. This represents humanity's desire to eliminate tracking errors. express transpose, Representing the pairing matrix The vector obtained by vectorizing the upper half matrix; Construct a robot objective function with the goal of assisting humans in trajectory tracking with minimal control energy. : In the formula, Indicates the discount factor. This indicates the robot's willingness to eliminate tracking errors. , This represents humanity's desire to eliminate tracking errors. A constant vector pre-defined based on the complexity of the task; Step B, Estimate the human control policy and objective function parameters: Based on the properties of the Nash equilibrium policy, the estimated human control policy is... Modeling and through Parameters of the human objective function The estimation process specifically includes: Step B1, estimating the human control policy: Based on the properties of the Nash equilibrium policy, the estimated human control policy is modeled as follows: ,in express transpose, express about gradient, Represents the activation function of a neural network. This represents the weights of a human actor neural network, and its update rate. for: ; in This represents the learning rate of the neural network weights representing the human control strategy. This represents the estimated state of the human-machine augmentation system model. The evolutionary process satisfies: ,in This represents the offset dynamics of the human-machine augmented system model under the estimated human control strategy. Let represent the augmented vectors of the estimated angular position and angular velocity of the robotic arm under the human control strategy. It is a positive definite constant matrix; Step B2, estimate the parameters of the human objective function: using Represents the parameters of the human objective function. The estimate, update rate for: ; In the formula This represents the learning rate of the human objective function parameters. Representing the pairing matrix The vector obtained by vectorizing the upper half matrix; Step C, Design the robot shared control strategy: Design the robot tracking control strategy based on the optimal control principle. Design robot safety control strategies based on control obstacle functions. By combining robot tracking control strategy and safety control strategy, a robot shared control strategy is obtained. The robot shared control strategy refers to the robot's control input in the human-machine system model. Step D: Apply the shared control strategy determined in step C to the robot. This enables human-machine shared tracking and control.

2. The human-machine shared tracking control method based on differential game theory according to claim 1, characterized in that, Step C specifically includes: Step C1: Design the robot tracking control strategy; Shared control strategy for robots Designed as follows: ; In the formula This represents the weights of the robot actor neural network, with an update rate of: ; In the formula , This indicates the first learning rate of the robot actor neural network weights. This represents the second learning rate of the robot actor neural network weights. This represents the weights of the robot's critic neural network. update rate satisfy: ; in This represents the learning rate of the robot's critic neural network weights. This represents humanity's desire to eliminate tracking errors. Representing the pairing matrix The vector obtained by vectorizing the upper half matrix; Step C2, Design robot safety control strategy: Set security set ,in, Represents the collision function, and , Represents a four-dimensional identity matrix. This represents the combination of two four-dimensional identity matrices. 3D matrix; express Obstacle position at any given time The corresponding augmented vector, For the safety radius, The square of the vector's 2-norm is used in robot safety control strategies. Designed as follows: ; in To meet The gain matrix whose smallest eigenvalue is greater than 0. express about gradient, ,in ; Step C3: Combine the robot tracking control strategy and the safety control strategy to obtain the robot shared control strategy. : .

Citation Information

Patent Citations

  • Local trajectory adjustment and man-machine sharing control method and system suitable for robot

    CN114454157A

  • Human-unmanned aerial vehicle group safety interactive motion planning method based on dynamic game

    CN115933748A