Dynamic optimization fractional order control method for man-machine cooperation process
By designing a fractional sliding mode observer and using deep reinforcement learning to optimize control parameters, the problem of balancing steady-state and dynamic performance in human-machine collaborative interaction in teleoperation systems was solved, achieving high-precision behavior force estimation and improving the stability and interactive performance of the control system.
Patent Information
- Application Number
- CN202610090044.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, during the human-computer collaborative interaction process in teleoperation systems, the controller struggles to balance steady-state and dynamic performance, and force-haptic feedback devices lack force sensors, resulting in unsatisfactory performance of traditional control methods.
A fractional sliding mode observer is designed to estimate operator behavior forces. A reference trajectory is generated through an admittance mechanism, a fractional sliding mode surface is constructed, and a control law is designed. The control parameters are optimized by combining deep reinforcement learning, and a Markov decision process is established to achieve adaptive control.
It achieves a balance between steady-state performance and dynamic response characteristics in human-machine collaboration, improves the stability and interactive performance of the control system, and the fractional sliding mode observer can estimate the operator's behavior force with high accuracy.
Smart Images

Figure CN121857327A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of human-computer collaboration and interaction technology, and in particular to a dynamic optimization fractional-order control method for human-computer collaboration processes. Background Technology
[0002] In related technologies, the balance between steady-state and dynamic performance of the controller during human-computer interaction with a series force-haptic feedback device in a teleoperation system is studied, with the end-effector trajectory tracking control of the force-haptic feedback device under human intervention as the research object. The collaborative interaction scenario is described as follows: the force-haptic feedback device tracks a preset trajectory, and when necessary, the operator corrects the preset trajectory by applying a force to the end-effector of the force-haptic feedback device. During this process, to ensure the steady-state performance of trajectory tracking, the controller needs to have a low steady-state error to ensure that the system can achieve high-precision tracking of the preset trajectory without human intervention. When the operator intervenes, the controller should have a high dynamic response capability to quickly and smoothly track the corrected trajectory it guides.
[0003] In existing control methods, controller parameters are typically set to fixed values, making it difficult to meet performance requirements at different stages. This necessitates a trade-off between steady-state and dynamic performance. Furthermore, force-haptic feedback devices lack force sensors and exhibit nonlinearity, dynamic coupling, and model parameter uncertainties, thus traditional control methods cannot achieve ideal results.
[0004] Therefore, it is necessary to improve one or more of the problems existing in the above-mentioned related technical solutions.
[0005] It should be noted that this section is intended to provide background or context for the technical solutions of this disclosure as set forth in the claims. The description herein does not constitute an admission that it is prior art simply because it is included in this section. Summary of the Invention
[0006] The purpose of this disclosure is to provide a dynamic optimization fractional-order control method for human-machine collaboration processes, thereby overcoming, to at least to some extent, one or more problems caused by the limitations and defects of related technologies.
[0007] According to embodiments of this disclosure, a dynamic optimization fractional-order control method for a human-machine collaboration process is provided, comprising: Step S1: Based on the dynamic model of the force-haptic feedback device, design a fractional-order sliding mode observer and estimate the interactive force applied by the operator; Step S2: Based on the interactive behavior force, the preset expected trajectory is corrected through the admittance mechanism to generate a reference trajectory; Step S3: Based on the tracking error between the reference trajectory and the actual trajectory of the device, construct a fractional sliding surface, and design a fractional control law containing behavior estimation force according to the fractional sliding surface to calculate and generate the output control force of the haptic feedback device; Step S4: Using deep reinforcement learning, the control parameters in the fractional-order control law are used as optimization objects to construct and train a Markov decision process to obtain a parameter optimization strategy that can adaptively adjust the control parameters according to the system state; wherein, the system state includes at least trajectory tracking error, velocity error and behavior estimation force; Step S5: During operation, the control parameters are adjusted in real time according to the parameter optimization strategy, and steps S1 to S3 are executed cyclically to achieve adaptive control of the force tactile feedback device.
[0008] Furthermore, the expression for the fractional sliding mode observer is:
[0009] in, For force haptic feedback device end position The observed values, For force haptic feedback device end-effector velocity The observed values, For controller output, For behavioral estimation ability, The inertial matrix of the force-haptic feedback device, For the centripetal force and Coriolis force matrix of the force haptic feedback device, The gravity matrix for force-haptic feedback devices.
[0010] Furthermore, based on the interactive behavior force, the preset expected trajectory is corrected through an admittance mechanism to generate a reference trajectory:
[0011] in, The desired position of the preset trajectory, The desired speed for the preset trajectory, The desired acceleration for the preset trajectory, The reference position for the reference trajectory. The reference velocity is the reference trajectory. The reference acceleration is the reference trajectory. The inertia matrix, Here is the damping matrix. Here is the stiffness matrix.
[0012] Furthermore, the expression for a fractional-order sliding surface is:
[0013] in, It is a fractional-order sliding surface; This represents a vector consisting of 2s; It is the first positive definite diagonal matrix; It is the first calculus operator; It is a fractional term; To track errors, It is a natural constant.
[0014] Furthermore, the expression for the fractional-order control law is:
[0015] in, For the second calculus operator, For speed error, It is the second positive definite diagonal matrix;
[0016] in, This represents the error between the observed position and the actual position in a fractional sliding mode observer. This represents the error between the observed velocity and the actual velocity. For fractional-order observers, auxiliary variables; For adaptive update law, It is the third positive definite diagonal matrix. It is the fourth positive definite diagonal matrix. It is the fifth positive definite diagonal matrix. It is the sixth positive definite diagonal matrix.
[0017] Furthermore, step S4 specifically includes: The controller parameter optimization problem is modeled as a finite Markov decision process, which is defined as follows: ;in, For system state This constitutes the state space; For control parameters The space for action; The state transition function is defined as follows: ; For dense reward functions; State transition function for:
[0018] in, Indicates the system at the 1st The state at any given moment, Indicates the time step interval. To control the action, As an auxiliary variable; This represents the integral variable.
[0019] By employing deep reinforcement learning, the control parameters in the fractional-order control law are used as optimization objects to train a Markov decision process, thereby obtaining a parameter optimization strategy for adjusting the control parameters.
[0020] Furthermore, the deep reinforcement learning method is any one of the PPO algorithm, DDPG algorithm, or SAC algorithm.
[0021] Furthermore, the reward function is:
[0022] in, This is the first adjustable parameter. This is the second adjustable parameter; The maximum limiting force output by the controller. For the first Dimensional error.
[0023] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: In the embodiments of this disclosure, the dynamic optimization fractional-order control method for the human-machine collaboration process described above achieves the following: First, a new fractional-order sliding mode surface is constructed, and a fractional-order sliding mode controller is designed accordingly to simultaneously consider the steady-state performance and dynamic response characteristics of the system. Second, a fractional-order sliding mode observer is designed to estimate the force applied by the operator to the end of the tactile feedback device in real time. Finally, a Markov decision process is established based on the full-process closed-loop control system, and the control strategy is trained through deep reinforcement learning to achieve adaptive dynamic optimization of the control parameters, thereby effectively balancing the steady-state and dynamic performance of the controller during collaborative interaction. Furthermore, this method exhibits excellent performance in both steady-state and dynamic aspects. The designed fractional-order sliding mode observer can estimate the force applied by the operator with high accuracy, effectively ensuring the stability and interactive performance of the control system. Attached Figure Description
[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0025] Figure 1 A flowchart illustrating the steps of a dynamic optimization fractional-order control method for a human-machine collaboration process in an exemplary embodiment of this disclosure is shown. Figure 2A comparison diagram showing the trajectory tracking error of the end effector of the haptic feedback device in an exemplary embodiment of this disclosure is provided. Figure 3 A comparison diagram showing the output trajectory of the fractional-order sliding mode controller in an exemplary embodiment of this disclosure is provided. Figure 4 A comparison diagram of the reference trajectory in an exemplary embodiment of this disclosure is shown; Figure 5 A comparison diagram of the error trajectories of behavioral force and behavioral estimation force in an exemplary embodiment of this disclosure is shown. Detailed Implementation
[0026] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0027] Furthermore, the accompanying drawings are merely illustrative diagrams of embodiments of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0028] This example implementation provides a dynamic optimization fractional-order control method for human-machine collaboration processes. (See reference...) Figure 1 As shown, the dynamic optimization fractional-order control method for this human-machine collaboration process may include: Step S1: Based on the dynamic model of the force-haptic feedback device, design a fractional-order sliding mode observer and estimate the interactive force applied by the operator; Step S2: Based on the interactive behavior force, the preset expected trajectory is corrected through the admittance mechanism to generate a reference trajectory; Step S3: Based on the tracking error between the reference trajectory and the actual trajectory of the device, construct a fractional sliding surface, and design a fractional control law containing behavior estimation force according to the fractional sliding surface to calculate and generate the output control force of the haptic feedback device; Step S4: Using deep reinforcement learning, the control parameters in the fractional-order control law are used as optimization objects to construct and train a Markov decision process to obtain a parameter optimization strategy that can adaptively adjust the control parameters according to the system state; wherein, the system state includes at least trajectory tracking error, velocity error and behavior estimation force; Step S5: During operation, the control parameters are adjusted in real time according to the parameter optimization strategy, and steps S1 to S3 are executed cyclically to achieve adaptive control of the force tactile feedback device.
[0029] The aforementioned dynamic optimization fractional-order control method for human-machine collaboration firstly constructs a new fractional-order sliding surface and designs a fractional-order sliding mode controller accordingly, simultaneously considering the system's steady-state performance and dynamic response characteristics. Secondly, a fractional-order sliding mode observer is designed to estimate the force applied by the operator to the end of the tactile feedback device in real time. Finally, a Markov decision process is established based on the full-process closed-loop control system, and the control strategy is trained through deep reinforcement learning to achieve adaptive dynamic optimization of control parameters, thereby effectively balancing the controller's steady-state and dynamic performance during collaborative interaction. Furthermore, this method exhibits excellent performance in both steady-state and dynamic aspects. The designed fractional-order sliding mode observer can estimate the force applied by the operator with high accuracy, effectively ensuring the stability and interactive performance of the control system.
[0030] Below, we will refer to Figures 1 to 5 The steps of the dynamic optimization fractional-order control method for the human-machine collaboration process described in this example embodiment will be explained in more detail.
[0031] In steps S1 and S2, a fractional-order sliding mode observer is designed based on the dynamic model of the force haptic feedback device, and the interactive force applied by the operator is estimated; based on the interactive force, the preset expected trajectory is corrected through the admittance mechanism to generate a reference trajectory.
[0032] Specifically, the dynamic model of the force-haptic feedback device in the human-computer collaborative interaction process is represented in Cartesian space as follows: (1) in, , and These represent the position, velocity, and acceleration of the end effector of the force-haptic feedback device in Cartesian space, respectively. , and These represent the inertial matrix, centripetal force and Coriolis force matrix, and gravity matrix of the force-haptic feedback device in Cartesian space, respectively, which are determined by the structure of the force-haptic feedback device. Indicates the controller output; This refers to the interactive force applied by the operator at the end of the force haptic feedback device.
[0033] Since it is impossible to directly obtain the force of real interactive behavior through force sensors To estimate the operator's interactive behavior, a fractional sliding mode observer is designed as follows: (2) in, For force haptic feedback device end position Observations For force haptic feedback device end-effector velocity Observed values. This will be represented as behavioral estimation force, which will be discussed in a later design.
[0034] definition , and These represent the position, velocity, and acceleration of the preset trajectory, respectively, followed by the desired position, desired velocity, and desired acceleration. Definition , and These represent the position, velocity, and acceleration of the corrected trajectory, respectively, followed by a description of the reference position, reference velocity, and reference acceleration. Based on behavioral estimation capabilities. The corrected reference trajectory is described by the admittance mechanism: (3) in, , and These represent the artificially defined inertia matrix, damping matrix, and stiffness matrix, respectively.
[0035] In step S3, based on the tracking error between the reference trajectory and the actual trajectory of the device, a fractional sliding surface is constructed, and a fractional control law containing behavior estimation force is designed according to the fractional sliding surface to calculate and generate the output control force of the haptic feedback device.
[0036] Specifically, the controller trajectory tracking error is defined as... Based on this, a fractional-order sliding surface is constructed: (4) in, It is a fractional-order sliding surface; This represents a vector consisting of 2s; It is a positive definite diagonal matrix; It is a calculus operator; Let be a fractional-order term. Based on this fractional-order sliding surface, the fractional-order control law is designed as follows: (5) in, For velocity error, the behavior estimation force is designed as follows: (6) in, This represents the error between the observed position and the actual position in a fractional sliding mode observer. This represents the error between the observed velocity and the actual velocity. For fractional-order observers, auxiliary variables; The designed adaptive update law is used.
[0037] Traditional methods use controller parameters , and Designed with fixed values, it is difficult to adjust in real time according to the system state, thus making it difficult to balance steady-state performance and dynamic response performance. This application employs a deep reinforcement learning method to obtain parameter optimization strategies through training. To adaptively generate optimal control parameters based on the system state at each time step. , and This process can be represented by the following formula: (7) To obtain parameter optimization strategies This application first models the controller parameter optimization problem as a finite Markov decision process, which is defined as follows: .in, For system state This constitutes the state space; For control parameters The space for action; The state transition function is defined as follows: ; It is a dense reward function that returns a scalar reward value at each time step, providing continuous feedback based on the agent's actions and the resulting state changes.
[0038] definition Indicates the system at the 1st The state at any given moment, Indicates the time step interval. To control the action. (The following is a list of steps / sections, likely related to motion control.) The state at any given moment is described as follows: (8) Among them, the state transition function Represented as: (9) in, As an auxiliary variable, it is approximated by the Oustaloup filter; This represents the integral variable.
[0039] The reward function is designed in the following form: (10) in, and These are adjustable parameters; This indicates the maximum limiting force output by the controller.
[0040] In a specific embodiment, step one involves training the aforementioned Markov decision process using a deep reinforcement learning method. After training, the strategy is obtained. This application employs PPO, DDPG, and SAC methods respectively to... Training is performed, and the specific training process is shown in Table 1. After training, the following initialization operations are performed: Initialize the fractional sliding mode observer parameters. , and Initialize behavioral estimation power Initialize the adaptive update law Set the desired trajectory; set the initial position of the force haptic feedback device.
[0041] Step 2, based on formula (3) and the behavior estimation power of the previous moment. Calculate the current reference trajectory , and Calculate the behavioral estimation power based on formula (6). The control output is calculated based on the generated action using formula (5). .
[0042] Step 3: Control output based on calculation Input the system dynamics, i.e., obtain the updated position and velocity of the force haptic feedback device's end effector from formula (1). This completes one update calculation; return to step two.
[0043] Understandably, the parameter initialization in the training process is only for training purposes, and in actual implementation, some parameters need to be reinitialized according to the steps.
[0044] Table 1 Training Process
[0045] like Figures 2 to 5 As shown, the changes in trajectory tracking error, controller output, parameter variations, and behavioral force estimation error of the force-haptic feedback device are illustrated. To verify the effectiveness of the proposed method, two sets of fixed-parameter fractional-order sliding mode control methods were set as control groups. Control group 1's controller parameters aimed to improve steady-state performance, while control group 2's controller parameters aimed to enhance dynamic response. Simulation results show that the proposed dynamic optimization fractional-order control method outperforms the two fixed-parameter control methods in both steady-state and dynamic performance across the three different training algorithms. Furthermore, the designed fractional-order sliding mode observer can estimate the operator-applied behavioral force with high accuracy, effectively ensuring the stability and interactive performance of the control system.
[0046] The aforementioned dynamic optimization fractional-order control method for human-machine collaboration firstly constructs a new fractional-order sliding surface and designs a fractional-order sliding mode controller accordingly, simultaneously considering the system's steady-state performance and dynamic response characteristics. Secondly, a fractional-order sliding mode observer is designed to estimate the force applied by the operator to the end of the tactile feedback device in real time. Finally, a Markov decision process is established based on the full-process closed-loop control system, and the control strategy is trained through deep reinforcement learning to achieve adaptive dynamic optimization of control parameters, thereby effectively balancing the controller's steady-state and dynamic performance during collaborative interaction. Furthermore, this method exhibits excellent performance in both steady-state and dynamic aspects. The designed fractional-order sliding mode observer can estimate the force applied by the operator with high accuracy, effectively ensuring the stability and interactive performance of the control system.
[0047] It should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise" in the above description indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of this disclosure and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this disclosure.
[0048] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0049] In the embodiments of this disclosure, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this disclosure according to the specific circumstances.
[0050] In embodiments of this disclosure, unless otherwise expressly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature being directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature being directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.
[0051] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0052] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
Claims
1. A dynamic optimization fractional-order control method for human-machine collaboration processes, characterized in that, include: Step S1: Based on the dynamic model of the force-haptic feedback device, design a fractional-order sliding mode observer and estimate the interactive force applied by the operator; Step S2: Based on the interactive behavior force, the preset expected trajectory is corrected through the admittance mechanism to generate a reference trajectory; Step S3: Based on the tracking error between the reference trajectory and the actual trajectory of the device, construct a fractional sliding surface, and design a fractional control law containing behavior estimation force according to the fractional sliding surface to calculate and generate the output control force of the haptic feedback device; Step S4: Using deep reinforcement learning, the control parameters in the fractional-order control law are used as optimization objects to construct and train a Markov decision process to obtain a parameter optimization strategy that can adaptively adjust the control parameters according to the system state; wherein, the system state includes at least trajectory tracking error, velocity error and behavior estimation force; Step S5: During operation, the control parameters are adjusted in real time according to the parameter optimization strategy, and steps S1 to S3 are executed cyclically to achieve adaptive control of the force tactile feedback device.
2. The dynamic optimization fractional-order control method for human-machine collaboration process according to claim 1, characterized in that, The expression for the fractional sliding mode observer is: in, For force haptic feedback device end position The observed values, For force haptic feedback device end-effector velocity The observed values, For controller output, For behavioral estimation ability, The inertial matrix of the force-haptic feedback device, For the centripetal force and Coriolis force matrix of the force haptic feedback device, The gravity matrix for force-haptic feedback devices.
3. The dynamic optimization fractional-order control method for human-machine collaboration process according to claim 2, characterized in that, Based on interactive behavior, a reference trajectory is generated by correcting the preset expected trajectory through an admittance mechanism: in, The desired position of the preset trajectory, The desired speed for the preset trajectory, The desired acceleration for the preset trajectory, The reference position for the reference trajectory. The reference velocity is the reference trajectory. The reference acceleration is the reference trajectory. The inertia matrix, Here is the damping matrix. Here is the stiffness matrix.
4. The dynamic optimization fractional-order control method for human-machine collaboration process according to claim 3, characterized in that, The expression for a fractional-order sliding surface is: in, It is a fractional-order sliding surface; This represents a vector consisting of 2s; It is the first positive definite diagonal matrix; It is the first calculus operator; It is a fractional term; To track errors, It is a natural constant.
5. The dynamic optimization fractional-order control method for human-machine collaboration process according to claim 4, characterized in that, The expression for the fractional-order control law is: in, For the second calculus operator, For speed error, It is the second positive definite diagonal matrix; in, This represents the error between the observed position and the actual position in a fractional sliding mode observer. This represents the error between the observed velocity and the actual velocity. For fractional-order observers, auxiliary variables; For adaptive update law, It is the third positive definite diagonal matrix. It is the fourth positive definite diagonal matrix. It is the fifth positive definite diagonal matrix. It is the sixth positive definite diagonal matrix.
6. The dynamic optimization fractional-order control method for human-machine collaboration process according to claim 5, characterized in that, Step S4 specifically includes: The controller parameter optimization problem is modeled as a finite Markov decision process, which is defined as follows: ;in, For system state This constitutes the state space; For control parameters The space for action; The state transition function is defined as follows: ; For dense reward functions; State transition function for: in, Indicates the system at the 1st The state at any given moment, Indicates the time step interval. To control the action, As an auxiliary variable; This represents the integral variable. By employing deep reinforcement learning, the control parameters in the fractional-order control law are used as optimization objects to train a Markov decision process, thereby obtaining a parameter optimization strategy for adjusting the control parameters.
7. The dynamic optimization fractional-order control method for human-machine collaboration process according to claim 6, characterized in that, The deep reinforcement learning method is any one of the PPO algorithm, DDPG algorithm, or SAC algorithm.
8. The dynamic optimization fractional-order control method for human-machine collaboration process according to claim 6, characterized in that, The reward function is: in, This is the first adjustable parameter. This is the second adjustable parameter; The maximum limiting force output by the controller. For the first Dimensional error.