A motion-behavior-criticism neural network learning control method for human-computer interaction process

Through multi-layered technical means, the technical problems in the human-computer interaction process in the existing technology have been solved, and adaptive dynamic programming technology has been realized. This technology has solved the technical problems in the existing technology and addressed the technical challenges in the human-computer interaction process.

CN119065504BActive Publication Date: 2025-11-21NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411255416.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-11-21
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

Existing observer methods have limitations in improving estimation accuracy, and traditional control methods cannot effectively address the nonlinearity and dynamic coupling issues of series force haptic feedback devices, leading to confusion of control forces during human-computer interaction.

Method used

A multi-layer behavioral network is used to approximate the human interaction behavior force in the human-computer interaction process. By constructing a human-computer interaction dynamics model and an augmented system, combined with radial basis functions and multi-layer RBF neural networks, the weight estimation vector update law of the behavioral network is designed, and the output torque and performance index function are approximated by action networks and critique networks to achieve adaptive dynamic programming.

Benefits of technology

The control precision of the human-computer interaction system has been improved by approximating the optimal solution through an adaptive dynamic programming process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119065504B_ABST
    Figure CN119065504B_ABST
Patent Text Reader

Abstract

The application relates to a motion-behavior-critic neural network learning control method for a human-machine interaction process, comprising the following steps: constructing a human-machine interaction dynamics model and obtaining a state vector of an augmented system; constructing a behavior network, obtaining a weight estimation vector update law of the behavior network, updating the behavior network, and obtaining a behavior moment by using the updated behavior network; obtaining a performance index function of the system in the human-machine interaction process; constructing a motion network and a critic network, obtaining a weight estimation vector update law of the motion network and the critic network, updating the motion network, and obtaining an output moment according to the updated motion network; and updating a state observation error term and a joint angular velocity observation value of a state observer, and a joint angle position and a joint angular velocity of the human-machine interaction dynamics model. The application realizes learning of a human-machine interaction system and improves control precision of the learned system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human-computer interaction collaborative teleoperation system, and particularly relates to a motion-behavior-criticism neural network learning control method for human-computer interaction process. BACKGROUND

[0002] In the process of researching the behavior stability of human interacting with a series force tactile feedback device in a teleoperation system, the behavior of human in the human-computer interaction process is taken as the research object, wherein the behavior of human is described by the interaction force of human acting on the end of the force tactile feedback device. Since most force tactile feedback devices do not have force sensors, the interaction force of human cannot be directly obtained from the force tactile feedback device, resulting in the confusion between the interaction force of human and the control force output by the device.

[0003] The existing processing method estimates the behavior force by using an observer, which is effective. However, since the observer needs to set an observer gain parameter, and this parameter does not have self-adaptability, the observer method has limitations in further improving the estimation accuracy. In addition, since the adopted series force tactile feedback device has the characteristics of nonlinearity, dynamic coupling and model parameter uncertainty, the traditional control method cannot achieve the ideal effect.

[0004] Therefore, it is necessary to provide a motion-behavior-criticism neural network learning control method for human-computer interaction process to solve the above problems. SUMMARY

[0005] The present application provides a motion-behavior-criticism neural network learning control method for human-computer interaction process to solve the problem that the existing observer method has limitations in further improving the estimation accuracy. In addition, since the adopted series force tactile feedback device has the characteristics of nonlinearity, dynamic coupling and model parameter uncertainty, the traditional control method cannot achieve the ideal effect.

[0006] The motion-behavior-criticism neural network learning control method for human-computer interaction process provided by the present application adopts the following technical scheme, comprising:

[0007] A human-computer interaction dynamics model is constructed according to the joint angle position, joint angular velocity, joint angular acceleration, inertia matrix, centripetal force and Coriolis force matrix, gravity matrix of the force tactile feedback device in the joint space under the human-computer interaction process, and the behavior torque of the interaction behavior force of the operator at the end of the force tactile feedback device on each joint; the state vector of the augmented system is obtained according to the joint angle position, joint angular velocity and expected trajectory of the force tactile feedback device in the joint space under the human-computer interaction process by using the augmented system to represent the human-computer interaction dynamics model;

[0008] The Gaussian function is taken as a radial basis function, and a behavior network is constructed according to a multi-layer RBF neural network; a weight estimation vector update law of the behavior network is obtained according to a system state observer in a human-machine interaction process, a radial basis function and a human-machine interaction dynamics model; the behavior network is updated based on the weight estimation vector update law; and behavior torque of an interactive behavior force of an operator at a terminal of the force tactile feedback device on each joint in the human-machine interaction process is obtained by using the updated behavior network.

[0009] A performance index function of the system in the human-machine interaction process is obtained according to output torque of each joint of the force tactile feedback device in joint space in the human-machine interaction process, behavior torque of the interactive behavior force of the operator at the terminal of the force tactile feedback device on each joint and a state vector of the augmented system.

[0010] An action network and a critic network are constructed; a weight estimation vector update law of the action network and the critic network is obtained according to the performance index function, an activation vector corresponding to the action network and the critic network, a weight estimation vector and behavior torque output by the behavior network; the action network is updated according to the weight estimation vector update law of the action network and the critic network; and output torque of each joint of the force tactile feedback device is obtained according to the updated action network.

[0011] Behavior torque of the interactive behavior force of the operator at the terminal of the force tactile feedback device on each joint in the human-machine interaction process is obtained according to the updated behavior network, and output torque of each joint of the force tactile feedback device is obtained according to the updated action network; a state observation error term and a joint angular velocity observation value of the state observer are updated; and joint angular position and joint angular velocity of the human-machine interaction dynamics model are updated.

[0012] Preferably, the expression of the human-machine interaction dynamics model is:

[0013]

[0014] In the formula, q represents joint angular position of the force tactile feedback device in joint space in the human-machine interaction process; represents joint angular velocity of the force tactile feedback device in joint space in the human-machine interaction process; represents joint angular acceleration of the force tactile feedback device in joint space in the human-machine interaction process; M represents an inertia matrix of the force tactile feedback device in joint space in the human-machine interaction process; C represents a centripetal force and Coriolis force matrix of the force tactile feedback device in joint space in the human-machine interaction process; G represents a gravity matrix of the force tactile feedback device in joint space in the human-machine interaction process; τ s represents output torque of each joint of the force tactile feedback device in joint space in the human-machine interaction process; τ hrepresents the behavior torque of the interactive behavior force exerted by the operator at the end of the force tactile feedback device on each joint.

[0015] Preferably, the expression of the augmented system is:

[0016]

[0017]

[0018]

[0019] wherein, represents the state vector of the augmented system at time t, e represents the joint angle position tracking error of the system; represents the derivative of the state vector r(t) of the augmented system with respect to time at time t, represents a Lipschitz continuous function transpose; represents the derivative of the joint angle position tracking error of the system; q d represents the desired joint angle position of each joint; represents the derivative of the desired joint angle position of each joint with respect to time, wherein the subscript d represents desired; F(r(t)) represents a nonlinear dynamics function; H(r(t)) represents a gain function; τ s represents the output torque of each joint of the force tactile feedback device in the joint space during human-machine interaction; h represents the behavior torque of the interactive behavior force exerted by the operator at the end of the force tactile feedback device on each joint; represents the transpose of the system state function; h T (e+q d ) represents the transpose of the system output gain function; represents a row vector of length n composed of 0.

[0020] Preferably, the expression of the system state observer during human-machine interaction is:

[0021]

[0022] wherein, represents the derivative of the joint angle velocity observation value; M -1 represents the inverse matrix of the inertia matrix of the force tactile feedback device in the joint space during human-machine interaction; C represents the centripetal and Coriolis force matrix of the force tactile feedback device in the joint space during human-machine interaction; G represents the gravity matrix of the force tactile feedback device in the joint space during human-machine interaction; τ s represents the output torque of each joint of the force tactile feedback device in the joint space during human-machine interaction; represents the estimated behavior torque; σ represents the joint angular velocity observation; represents the state observation error; D represents a positive diagonal matrix; ζ represents the state observation error term; L represents the error correction matrix.

[0023] Preferably, the expression of the weight estimation vector update law of the behavior network is:

[0024]

[0025] wherein the behavior network comprises an initial behavior network and j layers of extended behavior networks; represents the weight estimation vector update law of the initial behavior network in the behavior network; represents the weight estimation vector update law of the i-th layer of extended behavior networks in the behavior network; represents the weight estimation vector update law of the j-th layer of extended behavior networks in the behavior network; represents the weight estimation vector of the i+1-th layer of extended behavior networks in the behavior network; represents the weight estimation vector of the initial behavior network in the behavior network; represents the weight estimation vector of the i-th layer of extended behavior networks in the behavior network; represents the weight estimation vector of the j-th layer of extended behavior networks in the behavior network; represents the weight estimation vector of the 1-th layer of extended behavior networks in the behavior network; represents the activation vector of the behavior network; ζ represents the state observation error term; M -1 represents the inverse matrix of the inertia matrix of the force tactile feedback device in the joint space in the human-machine interaction process; δ h represents the learning rate of the initial behavior network in the behavior network, δ hi represents the learning rate of the i-th layer of extended behavior networks in the behavior network, δ hj represents the learning rate of the j-th layer of extended behavior networks in the behavior network; k h is an adjustable parameter of the initial behavior network, k hi represents the adjustable parameter of the i-th layer of extended behavior networks in the behavior network, k hj represents the adjustable parameter of the j-th layer of extended behavior networks in the behavior network.

[0026] Preferably, the expression of the performance index function of the system in the human-machine interaction process is:

[0027]

[0028] wherein V represents the performance index function of the system in the human-machine interaction process; α>0 represents a discount factor; Q=(Q1,0 2n×2n ;0 2n×2n ,0 2n×2n) is a positive semi-definite matrix, is a positive definite matrix; denotes a set of real matrices with dimension 2n x 2n; R, K are positive definite matrices; τ s denotes the output torque of the haptic feedback device in joint space at each joint during human-machine interaction; h denotes the behavior torque of the interaction force exerted by the operator at the end of the haptic feedback device on each joint; r(i) denotes the state vector of the augmented system at time i, where i is an integral variable in the time interval [t,∞].

[0029] Preferably, the expressions for constructing the action network and the critic network are:

[0030] The expression for the action network is:

[0031]

[0032] The expression for the critic network is:

[0033]

[0034] wherein, denotes the output torque estimate of the haptic feedback device in joint space at each joint during human-machine interaction; denotes the performance index function estimate of the system during human-machine interaction; β denotes the output peak value limit; denotes the weight estimate vector of the action network; denotes the activation vector of the action network; denotes the weight estimate vector of the critic network; denotes the activation vector of the critic network.

[0035] Preferably, the expression for the weight estimate vector update law of the action network and the critic network is:

[0036] The weight estimate vector update law of the action network and the critic network is defined as then,

[0037]

[0038] wherein, denotes the weight estimate vector update law of the action network and the critic network; denotes the weight estimate vector of the action network; denotes the weight estimate vector of the critic network; δ denotes the learning rate of the action network and the critic network; ψ denotes the residual error caused by the action network and the critic network in the performance index function part; φ denotes the residual error caused by the action network and the critic network in the remaining part.

[0039] Preferably, the expression of ψ is:

[0040]

[0041] wherein ψ represents the residual error caused by the action network and the critic network in the performance index function part; Q = (Q1, Q2n ×2 n; Q2n ×2 n, Q2n ×2 n) is a semi-positive definite matrix, Q1 is a positive definite matrix; the superscript T represents the transpose of the matrix or the function output; δ represents the learning rate; K is a positive definite matrix; α represents an adjustable parameter; β represents an output peak limit; represents the weight estimation vector of the initial behavior network; represents the transpose of the activation vector of the behavior network; r(ι) represents the state vector of the augmented system at time ι, wherein ι is the integral variable on the time interval [t-T p ]; κ is the integral variable on the output interval .

[0042] Preferably, the expression of φ is:

[0043]

[0044]

[0045]

[0046] wherein φ c represents the residual error caused by the critic network; φ a represents the residual error caused by the action network; e represents the joint angle position tracking error of the system; R is a positive definite matrix; r(ι) represents the state vector of the augmented system at time ι, wherein ι is the integral variable on the time interval [t-T p , t]; T p represents the operation interval time; represents the activation vector of the action network; represents the activation vector of the critic network; α represents an adjustable parameter; β represents an output peak limit; represents the matrix Kronecker product operator; r(t-T p ) represents the state vector of the augmented system at time t-T p ; r(t) represents the state vector of the augmented system at time t.

[0047] The beneficial effects of the present application are:

[0048] This invention approximates the behavioral torques of human interaction forces at various joints during human-computer interaction (HCI) through a multi-layered behavioral network. Specifically, based on the system state observer, radial basis functions, and HCI dynamics model, the weight estimation vector update law of the behavioral network is obtained. This law is then used to update the behavioral network. Simultaneously, an action network and a critique network are used to approximate the output torques and performance index functions. The performance index functions are then used to evaluate the current system state, action network strategy, and the quality of the behavioral network strategy, guiding the action network to learn. Specifically, the performance index of the adaptive dynamic programming process is designed based on the behavioral network, enabling the action network to approach the optimal solution of the system while considering behavioral forces. This yields the weight estimation vector update laws for the action network and critique network. These laws are then used to update the corresponding action network. Finally, the updated output values ​​of the action and behavioral networks are used to update the state observation error term and joint angular velocity observations of the state observer, as well as the joint angle positions and angular velocities of the HCI dynamics model. This achieves learning for the HCI system and improves the control accuracy of the learned system. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart of a human-computer interaction process action-behavior-critique neural network learning control method according to the present invention;

[0051] Figure 2 This is a schematic diagram of the action-behavior-critique neural network structure in an embodiment of the present invention;

[0052] Figure 3 This is a schematic diagram of the error trajectory of the behavioral network and the system state observer estimating the behavioral torque during the simulation process of this embodiment;

[0053] Figure 4 This is a schematic diagram of the behavioral torque of joint 1 and the estimated behavioral torque trajectory during the simulation process of this embodiment;

[0054] Figure 5 This is a schematic diagram of the behavioral torque of joint 2 and the estimated behavioral torque trajectory during the simulation process of this embodiment;

[0055] Figure 6Trajectory of joint angle position tracking error in the simulation process of the embodiment;

[0056] Figure 7 Trajectory of behavior network node weight in the simulation process of the embodiment;

[0057] Figure 8 Trajectory of joint angle position tracking error in the simulation process of the embodiment;

[0058] Figure 9 Trajectory of output torque estimated value estimated by the action network in the simulation process of the embodiment;

[0059] Figure 10 Trajectory of action network weight in the simulation process of the embodiment;

[0060] Figure 11 Trajectory of critic network weight in the simulation process of the embodiment. DETAILED DESCRIPTION

[0061] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0062] An embodiment of a motion-behavior-critic neural network learning control method for a human-computer interaction process of the present application, as shown in Figure 1 , includes:

[0063] S1, constructing a human-computer interaction dynamics model and obtaining a state vector of an augmented system;

[0064] According to the joint angle position, joint angular velocity, joint angular acceleration, inertia matrix, centripetal force and Coriolis force matrix, gravity matrix of the force tactile feedback device in the joint space in the human-computer interaction process, and the behavior torque of the interaction behavior force of the operator exerted on the end of the force tactile feedback device on each joint, a human-computer interaction dynamics model is constructed; the human-computer interaction dynamics model is represented by an augmented system, and the state vector of the augmented system is obtained according to the joint angle position, joint angular velocity and expected trajectory of the force tactile feedback device in the joint space in the human-computer interaction process.

[0065] Wherein, the expression of the human-computer interaction dynamics model is:

[0066]

[0067] where q represents joint angle position of the force tactile feedback device in joint space during human-machine interaction process, represents joint angular velocity of the force tactile feedback device in joint space during human-machine interaction process, represents joint angular acceleration of the force tactile feedback device in joint space during human-machine interaction process M represents inertia matrix of the force tactile feedback device in joint space during human-machine interaction process, C represents centripetal and Coriolis force matrix of the force tactile feedback device in joint space during human-machine interaction process, G represents gravity matrix of the force tactile feedback device in joint space during human-machine interaction process, s represents output torque of each joint of the force tactile feedback device in joint space during human-machine interaction process, h represents behavior torque of the interaction behavior force exerted by the operator at the end of the force tactile feedback device on each joint, n is the number of joints; represents a set of vectors with dimension n; represents a set of matrices with dimension n x n.

[0068] The human-machine interaction system can be described by an uncertain model in the form of formula (2) as follows:

[0069]

[0070] where, represents unknown nonlinear dynamics function of the system, h(q) is unknown output gain function,

[0071] Joint angle position tracking error of the control system can be represented as e = q - q d , and its derivative is represented as where, represents expected joint angle position of each joint, is expected joint angle position q d derivative with respect to time, is expected joint angle position q d second-order derivative with respect to time. It is assumed that there is a Lipschitz continuous function such that and the function is equal to 0 at zero point, represented as m(0) = 0. The human-machine interaction dynamics model under trajectory tracking problem can be described by an augmented system, specifically, the expression of the augmented system is: ​​

[0072]

[0073]

[0074]

[0075] where r(t) represents the state vector of the augmented system at time t, represents a set of vectors with dimension 4n; e represents the joint angle position tracking error of the system; represents the time derivative of the state vector r(t) of the augmented system at time t, represents a Lipschitz continuous function transpose; represents the derivative of the joint angle position tracking error of the system; q d represents the desired joint angle position of each joint; represents the derivative of the desired joint angle position of each joint with respect to time, where subscript d represents desired; F(r(t)) represents a nonlinear dynamics function; H(r(t)) represents a gain function; τ s represents the output torque of the force tactile feedback device in joint space for each joint during human-robot interaction; τ h represents the behavior torque of the interactive behavior force exerted by the operator at the end of the force tactile feedback device on each joint; represents the transpose of the system state function; h T (e+q d ) represents the transpose of the system output gain function; represents an n-length row vector composed of 0; it is noted that (·) T represents the transpose of a matrix vector, for example, the transpose of the inertia matrix M is M T , and the transpose of a matrix vector is uniformly expressed in this way hereinafter.

[0076] S2, a behavior network is constructed, a weight estimation vector update law of the behavior network is obtained, the behavior network is updated, and a behavior torque is obtained by using the updated behavior network;

[0077] Specifically, a Gaussian function is taken as a radial basis function, and a behavior network is constructed according to a multi-layer RBF neural network; a weight estimation vector update law of the behavior network is obtained according to a system state observer during human-robot interaction, a radial basis function, and a human-robot interaction dynamics model; the behavior network is updated based on the weight estimation vector update law; and a behavior torque of the interactive behavior force exerted by the operator at the end of the force tactile feedback device on each joint during human-robot interaction is obtained by using the updated behavior network.

[0078] In the embodiment, the multi-layer RBF neural network is used to approximate the behavior torque. The radial basis function of the RBF neural network is Gaussian function. The Gaussian function of the behavior network is expressed as:

[0079]

[0080] In the formula, a, b and c are constants; and is the activation function of the behavior network; represents a set of vectors with the dimension of c; c represents the number of nodes in the hidden layer of the behavior neural network; is the center point of the Gaussian function, represents a set of vectors with the dimension of 2n; λ is the width of the Gaussian function; and exp is a natural constant; is the input vector of the Gaussian function.

[0081] Therefore, in the subsequent design, the behavior torque τ h is described by using the RBF neural network as follows:

[0082]

[0083] In the formula, a, b and c are constants; and is the ideal weight vector of the behavior network; represents a set of real number matrices with the dimension of n x r; is the approximation error of the behavior network; and n is the dimension of the basis function. Since the ideal weight vector of the behavior network cannot be directly obtained, the ideal weight vector is estimated as follows: is the weight estimation vector of the behavior network, and the estimated behavior torque is In the application, the ideal weight vector of the behavior network is assumed to be is the activation vector, and the neural network estimation error ε h is bounded, and the derivative of the ideal weight vector of the behavior network is also unavailable. In the conventional method, it is generally considered that is equal to 0, but this method will bring additional estimation error. In order to reduce the estimation error, the behavior network is expanded here. On the basis of the original behavior network, the layer neural network is introduced to represent the derivative of the ideal weight matrix of the behavior network represents a set of real numbers; the original behavior network is collectively referred to as the initial behavior network; and the introduced neural network is collectively referred to as the expanded behavior network; and the ideal weight matrix of the expanded behavior network is expressed as υ = 1, 2,..., j, and is defined as follows:

[0084]

[0085] wherein, is the ideal weight estimation vector of the first layer extended behavior network in the behavior network; is the ideal weight estimation vector of the i+1th layer extended behavior network in the behavior network; is the derivative of the ideal weight estimation vector of the jth layer extended behavior network in the behavior network; is the derivative of the ideal weight estimation vector of the ith layer extended behavior network in the behavior network; is the derivative of the ideal weight estimation vector of the initial behavior network; is the weight estimation vector introduced by the extended behavior network to estimate the ideal weight vector At this time, the behavior network has a total of j+1 layers, including the initial behavior network, whose weight estimation vector is denoted as and the jth layer extended behavior network, whose weight estimation vector is denoted as Then the problem is converted into designing the weight estimation matrix of the initial behavior network and the weight estimation matrix of the extended behavior network and The update law of the weight estimation vector of the behavior network is obtained by introducing the system state observer. The system state observer in the human-machine interaction process is as follows:

[0086]

[0087] wherein, denotes the derivative of the joint angular velocity observation value, and has denotes the state observation error, ζ denotes the state observation error term, is the state observation error feedback correction term, is a positive diagonal matrix, and by adjusting the parameters of the matrix D, the accumulation of the state observation error in the observation process of the state observer can be suppressed; M -1 denotes the inverse matrix of the inertia matrix of the force tactile feedback device in the joint space in the human-machine interaction process; C denotes the centripetal force and Coriolis force matrix of the force tactile feedback device in the joint space in the human-machine interaction process; G denotes the gravity matrix of the force tactile feedback device in the joint space in the human-machine interaction process; τ s denotes the output torque of each joint of the force tactile feedback device in the joint space in the human-machine interaction process; denotes the estimated behavior torque; σ denotes the joint angular velocity observation value; D denotes a positive diagonal matrix; L denotes an error correction matrix, I n denotes an identity matrix with a dimension of n x n; denotes a set of real number vectors with dimension n, denotes a set of real number matrices with dimension n x n. In combination with the system state observer, in order to avoid the behavior network estimating redundant information, X in equation (6) is selected as

[0088] So far, according to equation (1), equation (6) and equation (9), the weight estimation vector update law of the multi-layer behavior network is designed as:

[0089]

[0090] where the behavior network includes the initial behavior network and the j-layer extended behavior network; denotes the weight estimation vector update law of the initial behavior network in the behavior network; denotes the weight estimation vector update law of the i-layer extended behavior network in the behavior network; denotes the weight estimation vector update law of the j-layer extended behavior network in the behavior network; denotes the weight estimation vector of the i+1-layer extended behavior network in the behavior network; denotes the weight estimation vector of the initial behavior network in the behavior network; denotes the weight estimation vector of the i-layer extended behavior network in the behavior network; denotes the weight estimation vector of the j-layer extended behavior network in the behavior network; denotes the weight estimation vector of the 1-layer extended behavior network in the behavior network; denotes the activation vector of the behavior network; ζ denotes the state observation error term; M -1 denotes the inverse matrix of the inertia matrix of the force tactile feedback device in the joint space in the human-computer interaction process; δ h denotes the learning rate of the initial behavior network in the behavior network, δ h > 0; δ hi denotes the learning rate of the i-layer extended behavior network in the behavior network, δ hi > 0, i = 1, 2, L, j-1; δ hj denotes the learning rate of the j-layer extended behavior network in the behavior network, δ hj > 0; by adjusting the size of the learning rate, the convergence speed of the behavior network is changed; k h is the adjustable parameter of the initial behavior network, k h > 0; k hi denotes the adjustable parameter of the i-layer extended behavior network in the behavior network, k hi > 0; k hj denotes the adjustable parameter of the j-layer extended behavior network in the behavior network, k hj > 0.

[0091] So far, the behavior network is updated by using the weight estimation vector update law, and the behavior torque of the interactive behavior force of the operator at the end of the force tactile feedback device on each joint in the human-machine interaction process can be obtained by using the updated behavior network.

[0092] S3, obtaining a performance index function of the system in the human-machine interaction process;

[0093] According to the output torque of the force tactile feedback device in the joint space of each joint, the behavior torque of the interactive behavior force of the operator at the end of the force tactile feedback device on each joint, and the augmented state vector of the system, the performance index function of the system in the human-machine interaction process is obtained.

[0094] In this embodiment, after the behavior network is designed, the value of the output torque τ s is obtained, that is, the control law is designed. The adaptive dynamic programming control method is selected to design the control law, and the performance index function in the method is designed first. The traditional performance index function only contains the quadratic function of the system state and the quadratic function of the control output. However, in the human-machine interaction process, the interactive behavior force of the human also affects the system performance, and the use of the traditional performance index function may affect the performance of the control algorithm. Therefore, the performance index function of the system in the human-machine interaction process is designed as follows:

[0095]

[0096] In the formula, V represents the performance index function of the system in the human-machine interaction process; α>0 represents a discount factor; Q=(Q1,0 2n×2n ;0 2n×2n ,0 2n×2n ) is a semi-positive definite matrix, is a positive definite matrix; represents a set of real matrices with a dimension of 2n*2n; R and K are both positive definite matrices; τ s represents the output torque of the force tactile feedback device in the joint space of each joint in the human-machine interaction process; τ h represents the behavior torque of the interactive behavior force of the operator at the end of the force tactile feedback device on each joint; r(i) represents the state vector of the augmented system at time i, where i is an integral variable in the time interval [t,∞].

[0097] S4, constructing an action network and a criticism network, obtaining a weight estimation vector update law of the action network and the criticism network and updating the action network, and obtaining the output torque according to the updated action network;

[0098] Specifically, the action network and the critic network are constructed, the weight estimation vector update law of the action network and the critic network is obtained according to the performance index function, the corresponding activation vector, the weight estimation vector of the action network and the critic network and the behavior moment output by the behavior network, the action network is updated according to the weight estimation vector update law of the action network and the critic network, and the output moment of each joint of the force tactile feedback device is obtained according to the updated action network.

[0099] Step 41, the step of constructing the action network and the critic network is:

[0100] Since the performance index function of formula (11) is complex in form, the performance index function value cannot be obtained by directly integrating. The neural network is introduced to solve this problem, that is, the action network is introduced to approximate τ s The critic network is introduced to approximate V respectively. The hyperbolic tangent function tanh(·) is introduced to describe the output restriction problem, which is described in the following form:

[0101]

[0102] Wherein, is the ideal weight vector of the action network; is the ideal weight vector of the critic network, l is the number of hidden layer nodes of the action network, and m is the number of hidden layer nodes of the critic network; is the activation vector of the action network; is the activation vector of the critic network; is the approximation error of the action network, is the approximation error of the critic network; denotes a set of real number matrices with the dimension of l x n; denotes a set of real number matrices with the dimension of m x n; denotes a set of real number vectors with the dimension of l; denotes a set of real number vectors with the dimension of m; denotes a set of real number vectors with the dimension of n; β is the output peak limit, and it is assumed that ε a and ε c are bounded, since the ideal weight vector cannot be obtained, therefore, the weight estimation vector of the action network is introduced to estimate the ideal weight vector of the action network, the weight estimation vector of the critic network is introduced to estimate the ideal weight vector of the critic network At this time, formula (12) is rewritten as:

[0103] The expression of the action network is:

[0104]

[0105] Expression of critic network:

[0106]

[0107] In the formula, represents the estimated value of the output torque of the force tactile feedback device in the joint space of the human-computer interaction process at each joint; represents the estimated value of the performance index function of the system in the human-computer interaction process; β represents the output peak value limit; represents the weight estimation vector of the action network; represents the activation vector of the action network; represents the weight estimation vector of the critic network; represents the activation vector of the critic network.

[0108] Step 42, the step of obtaining the weight estimation vector update law of the action network and the critic network:

[0109] Based on formulas (13) and (14), the problem is converted into designing the weight estimation vector update law of the action network and the weight estimation vector update law of the critic network , where t represents the operation time, and T represents the operation interval time p , in the time interval [t-T p , t], the model-free integral reinforcement learning algorithm is applied to the performance index function of formula (11), and the equation is as follows:

[0110]

[0111] In the formula, r(t-T p ) represents the state vector of the augmented system at t-T p ; r(t) represents the state vector of the augmented system at t; and r(ι) represents the state vector of the augmented system at ι, where ι is the integral variable in the time interval [t-T p , t].

[0112] In the present application, since the ideal value of the output torque τ s and the ideal value of the behavior torque τ h cannot be directly obtained, the estimated output torque estimation value the behavior torque estimation value and the performance index function estimation value are used to replace τ s , τ h and V. There is an error between the estimated value and the ideal value, and formula (15) can be rewritten as:

[0113]

[0114] Where ε represents the error term generated after replacing the ideal value with the estimated value, and when ε = 0, the estimated value can be considered as the ideal value; α > 0 is an adjustable parameter; Indicates tT p The weight estimation vector of the action network at time t; r(tT) p ) represents tT p The augmented system state vector at time T; p Indicates the interval between operations; The matrix Kronecker product operation is represented; vec(·) represents the function that converts a matrix into a vector by rows; r(ι) represents the state vector of the augmented system at time ι, where ι is the time interval [tT]. p The integral variable on ]; κ is the output interval. The integral variable on.

[0115] With ε approaching 0 as the design objective, the gradient descent method is used to design the update law of the weight estimation vectors of the action network and the critique network according to formula (16), and the weight estimation vectors of the action network and the critique network are defined. This includes the weight estimation vectors of the action network and the critique network. The update law for the weight estimation vectors of the action network and the critique network is then expressed as:

[0116]

[0117] In the formula, The update law for the weight estimation vectors of the action network and the critique network is represented. This represents the weight estimation vector of the action network; The weight estimation vector of the critique network; δ represents the learning rate of the action network and the critique network; ψ represents the residual error brought by the action network and the critique network in the performance index function part; φ represents the residual error brought by the action network and the critique network in the rest part.

[0118] Where δ>0 represents the learning rate, and the expression for ψ is:

[0119]

[0120] In the formula, ψ represents the residual error introduced by the action network and the critique network in the performance index function part; Q=(Q1,02n) ×2 n;02n ×2 n,02n ×2 n) is a positive semi-definite matrix, Q1 is a positive definite matrix; the superscript T indicates the transpose of the matrix or function output; δ represents the learning rate; K is a positive definite matrix; α represents an adjustable parameter; β represents the peak output limit; a weight estimation vector representing an initial behavior network; a transpose of an activation vector representing a behavior network; r(i) represents a state vector of an augmented system at i-th time.

[0121] wherein the expression of φ is:

[0122]

[0123]

[0124]

[0125] wherein φ c represents a residual error caused by a critic network; φ a represents a residual error caused by an actor network; e represents a joint angle position tracking error of the system; R is a positive definite matrix; r(i) represents a state vector of an augmented system at i-th time, wherein i is an integral variable on a time interval [t-T p , t]; T p represents an operation interval time; represents an activation vector of an actor network; represents an activation vector of a critic network; α represents an adjustable parameter; β represents an output peak value limit; represents a matrix Kronecker product operator; r(t-T p ) represents a state vector of an augmented system at t-T p ; r(t) represents a state vector of an augmented system at t-th time.

[0126] S5, updating a state observation error term and a joint angle velocity observation value of a state observer, and a joint angle position and a joint angle velocity of a human-machine interaction dynamics model;

[0127] Specifically, the behavior torque of the interaction behavior force of the operator at the end of the force touch feedback device on each joint in the human-machine interaction process is obtained according to the updated behavior network, and the output torque of each joint of the force touch feedback device is obtained according to the updated actor network, and the state observation error term and the joint angle velocity observation value of the state observer, and the joint angle position and the joint angle velocity of the human-machine interaction dynamics model are updated.

[0128] The embodiment will be described below in combination with the drawings:

[0129] As Figure 2 shown, step one, the initial time is recorded as t0, the system is initialized, and the system initial parameters are recorded as Taking the joint angle position as an example, the initial joint angle position is

[0130] wherein the steps of system initialization are:

[0131] 1.1, parameter initialization:

[0132] Initialize the operation interval time T p , the desired trajectory of the system, the joint angle position Joint angular velocity Observer output And the neural network parameters. Wherein the neural network parameters include activation function parameters, learning rate parameters and initial weight matrix parameters. At this time the system state observer output Respectively equal to the initial joint angle position

[0133] And joint angular velocity And the parameters M, C, G of the human-robot interaction dynamics model are known.

[0134] 1.2, behavior network initialization: according to the initial joint angle position Joint angular velocity And Get the state observation error term And get According to formula (10)

[0135] 1.3, action network and critic network initialization: according to the initial joint angle position Joint angular velocity And the desired trajectory to obtain the augmented system state vector r(t0), and combine the corresponding activation function of the action network and the critic network, the initial weight vector and the initial output of the behavior network According to formula (17) Thus the output of the action network is obtained

[0136] 1.4, state observer and human-robot interaction dynamics model update: according to the output of the behavior network And the output of the action network Update the state observation error term of the state observer And the joint angular velocity observation value And the joint angle position of the human-robot interaction dynamics model Joint angular velocity

[0137] Step two, the current time is recorded as t. Assume that there exists a matrix A, A t Indicates the matrix A at the current time, The matrix A represents the next time, and the subsequent unified expression method is adopted. The state observation error term ζ of the system state observer is obtained through the state observation value and the current joint angle position and angular velocity t Secondly, the behavior network weight update law is determined based on the state observation error and system dynamics The behavior network weight estimation update law is obtained by recursion And the estimated behavior torque is calculated

[0138] Step three, based on the current trajectory tracking error information, the current behavior network output The last action network estimated torque output Get action, critical neural network weight estimation update law, finally, calculate the action network output, and add the system output limit to get the output torque estimate value

[0139] Step four, the output torque estimate value of the action network And the behavior torque estimate value of the behavior network Commonly input to the system to realize the update learning of the system.

[0140] The simulation results are shown in Figures 3 to 7 , wherein Figure 3 The error trajectory diagram of the behavior network and the system state observer estimated behavior torque is shown in Figure 3 It is found that the behavior network estimated behavior torque error is small and can be kept within a small range of fluctuation and stable for a long time, and Figure 3 The error trajectory of the behavior torque estimated by the system state observer is described, and it can be seen that the behavior network approximation behavior torque method is better, which has higher approximation accuracy, Figure 4 The behavior torque and estimated behavior torque trajectory diagram of joint 1 is shown in Figure 5 The behavior torque and estimated behavior torque trajectory diagram of joint 2 is shown in Figure 6 The behavior torque and estimated behavior torque trajectory diagram of joint 3 is shown in Figures 4 to 6 It can be seen that Figure 7 The behavior network node weight trajectory diagram is shown in Figure 7 It can be seen that the behavior network node weight value can converge within a limited time.

[0141] Figure 8 The joint angle position tracking error trajectory diagram is shown in Figure 9 The trajectory diagram of the output torque estimate value estimated by the action network is shown in Figure 10 The action network weight trajectory diagram is shown in Figure 11 The critical network weight trajectory diagram is shown in Figures 8 to 11The four figures describe the joint angle trajectory tracking error, the action network estimated torque trajectory under the output limited condition, and the weight trajectory of the action network and the critic network under the adaptive dynamic programming control method based on action, behavior, and critic neural network. It can be found that the state error can quickly converge and remain stable. The weights of the action network and the critic network can also converge. At the same time Figure 8 A PID method comparison group is added, from Figure 8 It can be seen that the proposed control method has smaller trajectory tracking error than the PID control method in terms of trajectory tracking. Combined with Figure 3 and Figure 8 It can be seen that the method of the embodiment has better behavior torque estimation ability and trajectory tracking performance than the traditional method.

[0142] The above only describes the preferred embodiments of the present application and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for learning and controlling action-behavior-critique neural networks in human-computer interaction processes, characterized in that, include: Based on the joint angle position, joint angular velocity, joint angular acceleration, inertia matrix, centripetal force and Coriolis force matrices, gravity matrix, and output torque of each joint in the joint space of the force-haptic feedback device during human-computer interaction, as well as the behavioral torque of the interactive force applied by the operator at the end of the force-haptic feedback device in each joint, a human-computer interaction dynamic model is constructed. An augmented system is used to represent the human-computer interaction dynamic model. The state vector of the augmented system is obtained based on the joint angle position, joint angular velocity, and desired trajectory of the force-haptic feedback device in the joint space during human-computer interaction. The expression of the augmented system is: In the formula, express The state vector of the system is augmented at any given time. This indicates the joint angle position tracking error of the system; express The state vector of the augmented system at any time Regarding the time derivative, Represents the Lipschitz continuous function Transpose; The derivative of the joint angle position tracking error of the system; This indicates the desired joint angle position for each joint; This represents the derivative of the expected joint angle position of each joint with respect to time. Represents a nonlinear dynamic function; Represents the gain function; This indicates the output torque of the force-haptic feedback device at each joint in the joint space during human-computer interaction; This indicates the torque of the interactive force applied by the operator at the end of the force-haptic feedback device at each joint; This represents the transpose of the system state function; This represents the transpose of the system output gain function; This indicates that the length of the string consisting of 0s is... The row vector; Gaussian function is used as radial basis function and behavior network is constructed based on multilayer RBF neural network. Based on system state observer, radial basis function and human-computer interaction dynamics model in human-computer interaction process, weight estimation vector update law of behavior network is obtained. Behavior network is updated based on weight estimation vector update law. Behavior torque of interactive behavior force applied by operator at force haptic feedback device end in human-computer interaction process is obtained at each joint. Based on the output torque of each joint of the force-tactile feedback device in the joint space during the human-computer interaction process, the behavioral torque of the interactive force applied by the operator at the end of the force-tactile feedback device on each joint, and the state vector of the augmented system, the performance index function of the system during the human-computer interaction process is obtained. Construct an action network and a critique network. Based on the performance index function, the activation vectors and weight estimation vectors of the action network and the critique network, as well as the behavioral torques output by the behavior network, obtain the update law of the weight estimation vectors of the action network and the critique network. Update the action network according to the update law of the weight estimation vectors of the action network and the critique network, and obtain the output torques of each joint of the force haptic feedback device based on the updated action network. Based on the updated behavior network, the behavioral torques of the interactive forces applied by the operator at the end of the force-haptic feedback device during the human-computer interaction process are obtained on each joint. The updated motion network is used to obtain the output torques of each joint of the force-haptic feedback device. The state observation error term and joint angular velocity observation values ​​of the state observer are updated, as well as the joint angular positions and joint angular velocities of the human-computer interaction dynamics model.

2. The action-behavior-critique neural network learning control method for human-computer interaction process according to claim 1, characterized in that, The expression for the human-computer interaction dynamics model is: In the formula, This indicates the joint angle position of the force-haptic feedback device in the joint space during human-computer interaction; This indicates the joint angular velocity of the force-haptic feedback device in the joint space during human-computer interaction; This indicates the joint angular acceleration of the force-haptic feedback device in the joint space during human-computer interaction; This represents the inertia matrix of the force-haptic feedback device in the joint space during human-computer interaction. This represents the centripetal and Coriolis force matrix of the force-haptic feedback device in the joint space during human-computer interaction. This represents the gravity matrix of the force-haptic feedback device in the joint space during human-computer interaction. This indicates the output torque of the force-haptic feedback device at each joint in the joint space during human-computer interaction; This represents the behavioral torque of the interactive force applied by the operator at the end of the force-haptic feedback device at each joint.

3. The action-behavior-critique neural network learning control method for human-computer interaction process according to claim 1, characterized in that, The expression for the system state observer during human-computer interaction is: In the formula, The derivative of the observed joint angular velocity; This represents the inverse matrix of the inertia matrix of the force-haptic feedback device in the joint space during human-computer interaction. This represents the centripetal and Coriolis force matrix of the force-haptic feedback device in the joint space during human-computer interaction. This represents the gravity matrix of the force-haptic feedback device in the joint space during human-computer interaction. This indicates the output torque of the force-haptic feedback device at each joint in the joint space during human-computer interaction; Indicates the estimated behavioral torque; This represents the observed joint angular velocity. Indicates state observation error; Represents a diagonal matrix; This represents the state observation error term; This represents the error correction matrix.

4. The action-behavior-critique neural network learning control method for human-computer interaction process according to claim 1, characterized in that, The expression for the weight estimation vector update law of the behavioral network is: In the formula, the behavior network includes the initial behavior network and Layered extended behavioral networks; This represents the update law for the weight estimation vector of the initial behavioral network. The first in the behavioral network Weight estimation vector update law for layer extended behavioral networks; The first in the behavioral network Weight estimation vector update law for layer extended behavioral networks; The first in the behavioral network Weight estimation vectors for layer-extended behavioral networks; This represents the initial weight estimation vector of the behavioral network in the behavioral network; The first in the behavioral network Weight estimation vectors for layer-extended behavioral networks; The first in the behavioral network Weight estimation vectors for layer-extended behavioral networks; This represents the weight estimation vector of the first layer extended behavioral network in the behavioral network; Represents the activation vector of the behavioral network; This represents the state observation error term; This represents the inverse matrix of the inertia matrix of the force-haptic feedback device in the joint space during human-computer interaction. This represents the initial learning rate of the behavioral network. The first in the behavioral network The learning rate of the layer-extended behavioral network. The first in the behavioral network The learning rate of the layer-extended behavioral network; These are adjustable parameters for the initial behavioral network. The first in the behavioral network Adjustable parameters of layer-extended behavioral networks. The first in the behavioral network Adjustable parameters for layer-extended behavioral networks.

5. The action-behavior-critique neural network learning control method for human-computer interaction process according to claim 1, characterized in that, The expression for the system's performance index function during human-computer interaction is: In the formula, A function representing the system's performance metrics during human-computer interaction; Indicates the discount factor; It is a positive semi-definite matrix. It is a positive definite matrix; The dimension is The set of real matrices; , All are positive definite matrices; This indicates the output torque of the force-haptic feedback device at each joint in the joint space during human-computer interaction; This indicates the torque of the interactive force applied by the operator at the end of the force-haptic feedback device at each joint; express The state vector of the augmented system at time t, where Time interval The integral variable on.

6. The action-behavior-critique neural network learning control method for human-computer interaction process according to claim 1, characterized in that, The expressions for constructing the action network and the critique network are as follows: The expression for the action network is: A critical expression of the internet: In the formula, This represents the estimated output torque of the force-haptic feedback device at each joint in the joint space during human-computer interaction. This represents the estimated value of the system's performance index function during human-computer interaction. Indicates the output peak limit; This represents the weight estimation vector of the action network; This represents the activation vector of the action network; This represents the weight estimation vector of the critical network; This represents the activation vector of the critical network.

7. The action-behavior-critique neural network learning control method for human-computer interaction process according to claim 1, characterized in that, The expressions for the weight estimation vector update law of action networks and critique networks are: Define the weight estimation vectors for the action network and the critique network. Then we have: In the formula, The update law for the weight estimation vectors of the action network and the critique network is represented. This represents the weight estimation vector of the action network; This represents the weight estimation vector of the critical network; This represents the learning rate of the action network and the critical network; This represents the residual error caused by the action network and the critique network in the performance index function part; This represents the residual error caused by the action network and the critique network in the remaining parts.

8. The action-behavior-critique neural network learning control method for human-computer interaction process according to claim 7, characterized in that, The expression is: In the formula, This represents the residual error caused by the action network and the critique network in the performance index function part; It is a positive semi-definite matrix. It is a positive definite matrix; superscript Represents the transpose of the output of a matrix or function; It is a positive definite matrix; Indicates an adjustable parameter; Indicates the output peak limit; This represents the weight estimation vector of the initial behavioral network; This represents the transpose of the activation vector of the behavioral network; express The state vector of the augmented system at time t, where Time interval Integral variables on; For output range The integral variable on.

9. The action-behavior-critique neural network learning control method for human-computer interaction process according to claim 7, characterized in that, The expression is: In the formula, This represents the residual error caused by the critique network; This represents the residual error introduced by the action network; This indicates the joint angle position tracking error of the system; It is a positive definite matrix; express The state vector of the augmented system at time t, where Time interval Integral variables on; Indicates the interval between operations; This represents the activation vector of the action network; This represents the activation vector of the critical network; Indicates an adjustable parameter; Indicates the output peak limit; Represents the matrix Kronecker product operator; express The state vector of the augmented system at time t; Let t represent the state vector of the augmented system at time t.

Citation Information

Patent Citations

  • Reconfigurable mechanical arm cooperative force / motion control system and method based on terminal task assignment

    CN113276114A

  • Remote operation-based bionic manipulator man-machine interaction system and method

    CN118238147A