A fuzzy approximate fractional order control method for human-machine interaction process

CN116165890BActive Publication Date: 2026-08-21NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211741417.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-31
Publication Date
2026-08-21
Estimated Expiration
2042-12-31

AI Technical Summary

Technical Problem

[0003]在该系统中对机械臂的轨迹跟踪控制是主要研究的对象,然而由于该系统具有复杂动力学耦合性、固有非线性以及模型参数不确定性等特点使得常规的控制方法并不能够达到较为理想的控制效果

Benefits of technology

[0014]The beneficial effects of this invention are: by using a reinforcement learning fractional order optimization strategy to solve the non-singular terminal sliding surface, this invention can optimize the order of the control system. With the principle of minimizing the long-term trajectory tracking error, the order of the control system is continuously optimized during system operation. This enables the control system to simultaneously take into account dynamic performance and stability performance during interaction, thereby reducing control error and improving control accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116165890B_ABST
    Figure CN116165890B_ABST
Patent Text Reader

Abstract

The application discloses a fuzzy approximate fractional order control method for human-computer interaction process, obtains external control force and a desired motion trajectory of a mechanical arm; generates a reference motion trajectory and a tracking error according to the external control force and the desired motion trajectory; constructs a non-singular terminal sliding mode surface based on the error; solves the non-singular terminal sliding mode surface by using a reinforcement learning fractional order suboptimal strategy, and obtains the order of the non-singular terminal sliding mode surface; and calculates an auxiliary control force according to the external control force, the reference motion trajectory, the tracking error and the order; by solving the non-singular terminal sliding mode surface by using the reinforcement learning fractional order suboptimal strategy, the order of the control system can be optimized, the order of the control system is continuously optimized in the system operation process on the principle of long-term minimum trajectory tracking error, the dynamic performance and the stability performance of the control system can be considered simultaneously in the interaction process, the control error is reduced, and the control precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human-machine collaborative interaction technology for robots, and particularly relates to a fuzzy approximate fractional-order control method for human-machine interaction processes. Background Technology

[0002] The human-computer interaction control method is based on human-computer collaborative robot systems. This system uses human experience in scene recognition and evaluation of highly dynamic inducing factors as the top-level link of control decision-making, reducing the probability of problems arising in handling complex or sudden events due to insufficient robot intelligence.

[0003] The trajectory tracking control of the robotic arm is the main research object in this system. However, due to the complex dynamic coupling, inherent nonlinearity and uncertainty of model parameters of this system, conventional control methods cannot achieve ideal control results. Summary of the Invention

[0004] The purpose of this invention is to provide a fuzzy approximate fractional-order control method for human-computer interaction processes, so as to improve the control accuracy of the robotic arm in a human-machine collaborative robot system and reduce control errors.

[0005] This invention adopts the following technical solution: a fuzzy approximate fractional-order control method for human-computer interaction processes, comprising the following steps: Obtain the external control force and the desired motion trajectory of the robotic arm; Generate a reference motion trajectory and tracking error based on the external control force and the desired motion trajectory; Based on constructing non-singular terminal sliding surfaces according to errors; The order of the non-singular terminal sliding surface is obtained by using a reinforcement learning fractional order optimization strategy. The auxiliary control force is calculated based on the external control force, reference motion trajectory, tracking error, and order.

[0006] Furthermore, the generation of a reference trajectory and tracking error based on the external control force and the desired motion trajectory includes: , in, The mass matrix of the desired impedance structure, For the acceleration of the reference motion trajectory, The acceleration of the desired trajectory. Here is the damping matrix of the desired impedance structure. For reference speed of the motion trajectory, The velocity of the desired trajectory, Let be the elastic matrix of the desired impedance structure. For reference of the motion trajectory, For the desired motion trajectory, External control force.

[0007] Furthermore, generating the reference trajectory and tracking error based on the external control force and the desired motion trajectory also includes: , in, To track errors, This represents the actual motion trajectory of the robotic arm's end effector.

[0008] Furthermore, the non-singular terminal sliding surface is: , in, It is a non-singular terminal sliding surface. To track speed errors, The first diagonal matrix is ​​positive definite. For calculus operators, The order of the sliding surface. It is a constant. , .

[0009] Furthermore, reinforcement learning score order optimization strategies include: To enhance the state space of the learning fractional order optimization strategy, The value is the action space of the reinforcement learning fractional order optimization strategy. To enhance the reward function of the learning score-order optimization strategy, To reinforce the Q-value function of the fractional order optimization strategy; The optimization strategy for the order of reinforcement learning scores is to maximize the Q-value function and find the corresponding action space. in, , for Each time step contains an extended-dimensional state vector that approximates the system state. for Tracking error at any time, for Auxiliary variables at time, , yes The state transition matrix of a fractional-order approximation system. yes The input matrix of the fractional-order approximation system. Let i be the i-th element in the tracking error. , For the control parameters of the reward, This represents the tracking error at time k. , This represents the optimal strategy. , Let k be the order at time k.

[0010] Furthermore, the auxiliary control force is calculated based on the external control force, reference motion trajectory, tracking error, and order, including: , in, To assist in control, Let the inertia matrix of the robotic arm's end effector in Cartesian space be denoted as . To achieve a fast terminal sliding mode approach rate, Let the centripetal and Coriolis force matrices of the robotic arm's end effector in Cartesian space be given. This refers to the actual speed of movement at the end of the robotic arm. Let be the gravity matrix of the robotic arm's end effector in Cartesian space.

[0011] Furthermore, the fast terminal sliding mode convergence rate is: , in, and It is a positive definite second diagonal matrix with all elements on the diagonal greater than zero. It is a constant. .

[0012] Another technical solution of the present invention: a fuzzy approximate fractional-order control device method for human-computer interaction processes, comprising: The acquisition module is used to acquire the external control force and the desired motion trajectory of the robotic arm; The generation module is used to generate a reference motion trajectory and tracking error based on the external control force and the desired motion trajectory. A module for constructing non-singular terminal sliding surfaces based on errors; The solution module is used to solve for the non-singular terminal sliding surface using a reinforcement learning fractional order optimization strategy, and obtain the order of the non-singular terminal sliding surface. The calculation module is used to calculate the auxiliary control force based on the external control force, reference motion trajectory, tracking error, and order.

[0013] Another technical solution of the present invention: a fuzzy approximate fractional-order control device for human-computer interaction process, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned fuzzy approximate fractional-order control method for human-computer interaction process.

[0014] The beneficial effects of this invention are: by using a reinforcement learning fractional order optimization strategy to solve the non-singular terminal sliding surface, this invention can optimize the order of the control system. With the principle of minimizing the long-term trajectory tracking error, the order of the control system is continuously optimized during system operation. This enables the control system to simultaneously take into account dynamic performance and stability performance during interaction, thereby reducing control error and improving control accuracy. Attached Figure Description

[0015] Figure 1 Auxiliary variables in the embodiments of the present invention A schematic diagram showing the convergence of each element over time; Figure 2 This is a schematic diagram illustrating the system order changes during the training process in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the changes in the reward function during training in an embodiment of the present invention; Figure 4 This is a schematic diagram of the trajectory tracking error of the robotic arm tracking a sine wave in the X-axis direction in an embodiment of the present invention; Figure 5 This is a schematic diagram of the trajectory tracking error of the robotic arm tracking a triangular wave in the X-axis direction in an embodiment of the present invention; Figure 6 This is a schematic diagram of the trajectory tracking error of the robotic arm tracking a square wave in the X-axis direction in an embodiment of the present invention. Detailed Implementation

[0016] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0017] Non-singular terminal sliding mode control based on fractional calculus theory can handle nonlinear systems such as human-robot collaborative robot systems well. However, in the system design process of applying this method, choosing a suitable sliding surface order is a major challenge. Although a larger order can bring a faster dynamic response speed, it sacrifices the steady-state tracking accuracy of the system and may even result in a large overshoot. Therefore, a fixed sliding surface order cannot meet the actual control requirements.

[0018] This invention provides a fractional-order human-machine interaction control method based on fuzzy approximation. It utilizes the Q-Learning algorithm from reinforcement learning to optimize the order of the control system, aiming to minimize the long-term trajectory tracking error. This ensures the control system's order continuously optimizes during system operation and tracks the motion of the robotic arm's end effector, which is driven by the proposed algorithm. Throughout the interaction process, the control system simultaneously considers dynamic and steady-state performance, ensuring the robotic arm's trajectory tracking error converges within a finite time.

[0019] This invention discloses a fuzzy approximate fractional-order control method for human-computer interaction processes, comprising the following steps: acquiring an external control force and the desired motion trajectory of a robotic arm; generating a reference motion trajectory and tracking error based on the external control force and the desired motion trajectory; constructing a non-singular terminal sliding surface based on the error; solving for the non-singular terminal sliding surface using a reinforcement learning fractional-order optimization strategy to obtain the order of the non-singular terminal sliding surface; and calculating an auxiliary control force based on the external control force, the reference motion trajectory, the tracking error, and the order.

[0020] This invention utilizes a reinforcement learning fractional order optimization strategy to solve for the non-singular terminal sliding surface, thereby optimizing the order of the control system. By minimizing the long-term trajectory tracking error, the order of the control system is continuously optimized during system operation. This allows the control system to simultaneously consider dynamic performance and stability during interaction, reducing control error and improving control accuracy.

[0021] In this embodiment of the invention, the desired trajectory is generally pre-designed in actual control. ,and , as well as The relationship between the three can be described using an impedance model: (1) in, The mass matrix of the desired impedance structure, , For the acceleration of the reference motion trajectory, The acceleration of the desired trajectory. Here is the damping matrix of the desired impedance structure. , For reference speed of the motion trajectory, The velocity of the desired trajectory, Let be the elastic matrix of the desired impedance structure. , For reference of the motion trajectory, For the desired motion trajectory, For external control force, By adjusting these parameters, the desired user experience can be achieved during human-computer interaction.

[0022] Specifically, the trajectory tracking error of the controller can be expressed as: (2) in, To track errors, This represents the actual motion trajectory of the robotic arm's end effector.

[0023] Next, a fractional-order non-singular terminal sliding surface is constructed based on the tracking error, specifically as follows: (3) in, It is a non-singular terminal sliding surface. To track speed errors, The first diagonal matrix is ​​positive definite. ,and All elements on the diagonal are greater than zero. For calculus operators, the following is adopted: Fractional calculus as defined, and since fractional calculus cannot be directly implemented in control systems, this invention employs... The method performs numerical approximation of fractional calculus operators. The order of the sliding surface. , It is a constant. , .

[0024] Specifically, in this embodiment, the reinforcement learning score order optimization strategy includes: the reinforcement learning score order optimization strategy is to find the corresponding action space by maximizing the Q-value function.

[0025] More specifically, consider the following dynamics on the sliding surface: (4) Theoretically, when the order satisfies In this case, the dynamic finite-time stability on the sliding surface can be guaranteed. Considering that different fixed orders have their own advantages for transient and steady-state performance, order optimization can be actively performed based on the dynamic characteristics of the surface to achieve optimal overall performance. The analytical solution of fractional calculus cannot be obtained directly. This method is one of the most commonly used methods for approximating fractional calculus using higher-order systems, for the error vector The element The mathematical description of the fractional numerical approximation is as follows: (5) in, , which is an auxiliary variable for calculating fractional numerical results. It actively determines the order of the auxiliary system used to approximate fractional values. yes The state transition matrix of a fractional-order approximation system. yes The input matrix of the fractional-order approximation system. yes The observation matrix of a fractional-order approximation system. yes The input-output matrix of a fractional-order approximation system. It is the system output, used to calculate the fractional order numerical result, which is the corresponding approximation fractional order value of the input.

[0026] Therefore, the dynamic response on the surface can be described using the following dynamic equation: (6) To track the i-th element of the error, the following Markov decision process is formed: (7) in, express Each time step contains an extended-dimensional state vector that approximates the system state. express The dynamic order at time step, It describes Each time step contains an extended-dimensional state vector that approximates the system state. for Tracking error at any time, for Auxiliary variables at time, .

[0027] Based on the principle of maximizing rewards, the reward function is designed according to the system trajectory tracking error as follows: (8) in For the control parameters of the reward, This represents the tracking error at time k. The order of the sliding surface is also considered. As the action space, the auxiliary state variables of the system trajectory tracking error and the fractional-order numerical approximation system are used as the state space. The optimal order selection can obtain the optimal Q-value function, which can be expressed as: (9) in, It is the optimal value function, which is usually difficult to obtain directly and requires the use of function approximation methods. The optimal strategy is defined by the following relationship: (10) To construct a near-optimal numerical approximation of the Q-value function, the following Q-value function transition description is constructed: (11) in, It is a transition mapping of the Q-value function, satisfying ,in, yes The Q-value function after the algorithm iteration. yes The Q-value function after this algorithm iteration. Define the mapping. ,in, It is a parameterized description of the state space, and a mapping. Then the following continuous mapping relation holds. , yes Parameter values ​​after the algorithm iteration yes After the algorithm iterations, the state values ​​are determined, and the problem is transformed into finding an approximate solution that satisfies the following expression: (12) At this point, the problem is transformed into the parameter... Finding the optimal approximate solution without loss of generality, using multidimensional parameters. There is a non-linear mapping relationship between the function and the Q-value as follows: (13) in, It is a description of state-space related parameters. Indicates the selected first The fractional order approximation of the parameter value corresponding to the minimum of the Q-value function is the first... One portion, It is a Q-value function with approximate parameters, and at this time... That is, the parameter value corresponding to the selected strategy, which can be marked as The near-optimal approximate parameter expression can be determined using the following strategy to generate a fractional-order parameter selection strategy: (14) in, It is a near-optimal order selection strategy.

[0028] The above completes the process of determining the order. Next, the required auxiliary control forces need to be generated. This embodiment of the invention describes the human-computer interaction process. The dynamic model of the robotic arm can be simplified using the Lagrange method to the following form: (15) in, and These represent the velocity and acceleration of the robotic arm's end effector in Cartesian space, respectively.

[0029] Then, the above equation is transformed, and reference motion trajectory information is added to obtain the control law of the trajectory tracking controller during human-computer interaction: (16) in, To assist in control, , Let the inertia matrix of the robotic arm's end effector in Cartesian space be denoted as . To achieve a fast terminal sliding mode approach rate, Let the centripetal and Coriolis force matrices of the robotic arm's end effector in Cartesian space be given. This refers to the actual speed of movement at the end of the robotic arm. Let be the gravity matrix of the robotic arm's end effector in Cartesian space.

[0030] Meanwhile, in this embodiment of the invention, in order to ensure that the system has a fast approach rate during the arrival motion phase of sliding mode control, a fast terminal sliding mode approach rate is adopted, which can be expressed as: (17) in, and It is a positive definite second diagonal matrix with all elements on the diagonal greater than zero. It is a constant. .

[0031] In summary, the control law of the trajectory tracking controller during human-computer interaction is: (18) In summary, the order optimization control method for a human-computer interaction process proposed in this invention utilizes... The algorithm optimizes the order of the sliding surface, ensuring that the control system simultaneously achieves both dynamic and steady-state performance. Furthermore, because the variable-order control method minimizes the long-term trajectory tracking error, the order of the control system continuously changes, thus reducing the operator's parameter tuning workload in practical engineering design. Compared to existing methods, this invention, for the first time, proposes a fractional-order non-singular terminal sliding mode control based on existing methods to simultaneously achieve high steady-state tracking accuracy and good dynamic response speed. The algorithm designs a variable-order control strategy. This method can ensure that the control system has both high trajectory tracking accuracy and good dynamic response speed.

[0032] To verify the effectiveness of the method of the present invention, verification experiments were also conducted. For example... Figure 1 As shown in the figure, to verify the convergence of relevant state variables during the training process in the experiment, the figure... , , , , and Auxiliary variables Elements in. For example, Figure 2 The diagram shown illustrates the change in order. Figure 3 The diagram shows the changes in the reward function. It is easy to see from the three diagrams above that the adopted... It is convergent, and the order of the control system changes continuously with the system trajectory tracking error, generating a local optimal order sequence.

[0033] In addition, from Figure 4 , Figure 5 and Figure 6 The tracking errors of the robotic arm tracking three different waveforms reveal that as the fixed order increases, the system's dynamic response speed improves, but this also leads to larger overshoot and lower steady-state tracking accuracy. However, the designed variable-order control strategy maintains high tracking accuracy while also exhibiting good dynamic response speed. Furthermore, compared to fixed-order control, variable-order control consistently demonstrates better control performance even when the trajectory waveform changes.

[0034] This invention also discloses a fuzzy approximate fractional-order control device method for human-computer interaction processes, comprising: an acquisition module for acquiring external control force and the desired motion trajectory of the robotic arm; a generation module for generating a reference motion trajectory and tracking error based on the external control force and the desired motion trajectory; a construction module for constructing a non-singular terminal sliding surface based on the error; a solution module for solving the non-singular terminal sliding surface using a reinforcement learning fractional-order optimization strategy to obtain the order of the non-singular terminal sliding surface; and a calculation module for calculating the auxiliary control force based on the external control force, the reference motion trajectory, the tracking error, and the order.

[0035] The present invention also discloses a fuzzy approximate fractional-order control device for human-computer interaction processes, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned fuzzy approximate fractional-order control method for human-computer interaction processes.

[0036] It should be noted that the information interaction and execution process between the above-mentioned devices are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0037] The device can be a computing device such as a desktop computer, laptop, handheld computer, radar, or cloud server. The device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that it may include more or fewer components, or a combination of certain components, or different components; for example, it may also include input / output devices, network access devices, etc.

[0038] The processor referred to can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0039] In some embodiments, the memory may be an internal storage unit of the extraction device, such as the hard drive or memory of the extraction device. In other embodiments, the memory may be an external storage device of the extraction device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the extraction device. Furthermore, the memory may include both internal storage units and external storage devices of the extraction device. The memory is used to store operating systems, applications, bootloaders, data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or will be output.

Claims

1. A fuzzy approximate fractional-order control method for human-computer interaction processes, characterized in that, Includes the following steps: Obtain the external control force and the desired motion trajectory of the robotic arm; A reference motion trajectory and tracking error are generated based on the external control force and the desired motion trajectory. Based on the aforementioned construction of a non-singular terminal sliding surface according to the error; The order of the non-singular terminal sliding surface is obtained by solving the non-singular terminal sliding surface using a reinforcement learning fractional order optimization strategy. The auxiliary control force is calculated based on the external control force, reference motion trajectory, tracking error, and order. The reinforcement learning score order optimization strategy is to maximize the Q-value function to obtain the corresponding action space; in, , To reinforce the Q-value function of the fractional order optimization strategy, For the control parameters of the reward, This represents the tracking error at time k. This represents the optimal strategy. To enhance the reward function of the learning score-order optimization strategy, , , To enhance the state space of the learning fractional order optimization strategy, Let k be the order at time k. , for Each time step contains an extended-dimensional state vector that approximates the system state. for Tracking error at any time, for Auxiliary variables at time, , yes The state transition matrix of a fractional-order approximation system. yes The input matrix of the fractional-order approximation system. The value is the action space of the reinforcement learning fractional-order optimization strategy. Let i be the i-th element in the tracking error. It is a constant. .

2. The fuzzy approximate fractional-order control method for human-computer interaction processes as described in claim 1, characterized in that, The generation of a reference trajectory and tracking error based on the external control force and the desired motion trajectory includes: , in, The mass matrix of the desired impedance structure, For the acceleration of the reference motion trajectory, The acceleration of the desired trajectory. Here is the damping matrix of the desired impedance structure. For reference speed of the motion trajectory, The velocity of the desired trajectory, Let be the elastic matrix of the desired impedance structure. For reference of the motion trajectory, For the desired motion trajectory, External control force.

3. The fuzzy approximate fractional-order control method for human-computer interaction processes as described in claim 2, characterized in that, The generation of the reference trajectory and tracking error based on the external control force and the desired motion trajectory also includes: , in, To track errors, This represents the actual motion trajectory of the robotic arm's end effector.

4. A fuzzy approximate fractional-order control method for human-computer interaction processes as described in claim 2 or 3, characterized in that, The non-singular terminal sliding surface is: , in, It is a non-singular terminal sliding surface. To track speed errors, The first diagonal matrix is ​​positive definite. For calculus operators, The order of the sliding surface. .

5. The fuzzy approximate fractional-order control method for human-computer interaction processes as described in claim 4, characterized in that, The auxiliary control force is calculated based on the external control force, reference motion trajectory, tracking error, and order, including: , in, To assist in control, Let the inertia matrix of the robotic arm's end effector in Cartesian space be denoted as . To achieve a fast terminal sliding mode approach rate, Let the centripetal and Coriolis force matrices of the robotic arm's end effector in Cartesian space be given. This refers to the actual speed of movement at the end of the robotic arm. Let be the gravity matrix of the robotic arm's end effector in Cartesian space.

6. The fuzzy approximate fractional-order control method for human-computer interaction processes as described in claim 5, characterized in that, The rapid terminal sliding mode convergence rate is: , in, and It is a positive definite second diagonal matrix with all elements on the diagonal greater than zero. It is a constant. .

7. A fuzzy approximate fractional-order control device method for human-computer interaction processes, characterized in that, include: The acquisition module is used to acquire the external control force and the desired motion trajectory of the robotic arm; The generation module is used to generate a reference motion trajectory and tracking error based on the external control force and the desired motion trajectory; A construction module is used to construct a non-singular terminal sliding surface based on the error. The solution module is used to solve the non-singular terminal sliding surface using a reinforcement learning score order optimization strategy to obtain the order of the non-singular terminal sliding surface. The calculation module is used to calculate the auxiliary control force based on the external control force, the reference motion trajectory, the tracking error, and the order. The reinforcement learning score order optimization strategy is to maximize the Q-value function to obtain the corresponding action space; in, , To reinforce the Q-value function of the fractional order optimization strategy, For the control parameters of the reward, This represents the tracking error at time k. This represents the optimal strategy. To enhance the reward function of the learning score-order optimization strategy, , , To enhance the state space of the learning fractional order optimization strategy, Let k be the order at time k. , for Each time step contains an extended-dimensional state vector that approximates the system state. for Tracking error at any time, for Auxiliary variables at time, , yes The state transition matrix of a fractional-order approximation system. yes The input matrix of the fractional-order approximation system. The value is the action space of the reinforcement learning fractional-order optimization strategy. Let i be the i-th element in the tracking error. It is a constant. .

8. A fuzzy approximate fractional-order control device for human-computer interaction processes, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes a computer program, it implements the fuzzy approximate fractional-order control method for human-computer interaction processes as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Mechanical arm trajectory tracking method based on fractional-order adaptive nonsingular terminal sliding mode

    CN107942684A

  • Fractional order sliding mode control method of flexible joint mechanical arm

    CN108181813A