Soft stipulation performance reinforcement learning control method and system for rehabilitation exoskeleton

By employing a soft-stress performance reinforcement learning control method, combined with dynamic soft constraints and online learning, the safety and stability issues of a pneumatic artificial muscle-driven upper limb rehabilitation exoskeleton robot in complex environments were resolved. This approach enabled precise trajectory tracking and efficient disturbance handling, thereby enhancing the system's robustness and learning capabilities.

CN121364635APending Publication Date: 2026-01-20SHENZHEN RES INST OF NANKAI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511641424.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing technologies for controlling pneumatic artificial muscle-driven upper limb rehabilitation exoskeleton robots suffer from limitations such as the single nature of traditional control strategies, fixed performance boundaries leading to safety and stability issues, inability to effectively handle sudden disturbances in complex environments, and lack of dynamic adjustment mechanisms and disturbance handling efficiency.

Method used

A soft-constraint performance reinforcement learning control method is adopted, which combines dynamic soft constraint mechanism and online learning and compensation mechanism. Tunnel-type preset performance boundary and safety boundary are designed. The lumped disturbance is handled by reinforcement learning framework to achieve system safety and high performance coordination.

Benefits of technology

It achieves precise trajectory tracking of pneumatic artificial muscle-driven upper limb rehabilitation exoskeleton robot in complex environments, improves the robustness and safety of the system under extreme conditions, avoids the unlimited growth of control force, and ensures the stability and efficient learning ability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121364635A_ABST
    Figure CN121364635A_ABST
Patent Text Reader

Abstract

The invention discloses a soft stipulation performance reinforcement learning control method for rehabilitation exoskeleton, and the method is a reinforcement learning control method based on soft preset performance, and comprises the steps: S1, building a soft stipulation performance reinforcement learning control law based on a dynamic soft constraint mechanism and an online learning and compensation mechanism; wherein the dynamic soft constraint mechanism is used for intelligent coordination control of safety and high performance, and the online learning and compensation mechanism is used for intelligent coordination calculation of complexity, model dependence and control precision; and S2, controlling the rehabilitation exoskeleton based on the soft stipulation performance reinforcement learning control law. The invention further discloses a corresponding system, electronic equipment and a computer readable storage medium, a soft constraint dynamic adjustment mechanism is combined with an intelligent learning and decision-making mechanism, and therefore the balance problem of safety, accuracy and self-adaptability when the upper limb rehabilitation robot faces uncertainty is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical device control, in particular to a soft guaranteed performance reinforcement learning control method and system for a rehabilitation exoskeleton. BACKGROUND

[0002] Upper limb motor dysfunction is a common sequelae of diseases such as rotator cuff injury, which seriously affects the quality of life of patients. Upper limb rehabilitation exoskeleton robot as an important medical auxiliary equipment, through guiding or assisting patients to carry out repetitive and standardized training, can effectively promote the remodeling of neural function and the recovery of motor function, and prevent muscle atrophy. In such applications, safety and stability are the primary premise, and any control instability can cause secondary injury to patients. At the same time, the efficiency and comfort of training (i.e. the compliance of the machine) are also crucial, which requires the robot to accurately track the predetermined rehabilitation trajectory and provide standardized assisted movement.

[0003] To cope with the nonlinearity of the upper limb rehabilitation exoskeleton robot driven by pneumatic artificial muscle, model predictive control, adaptive control, neural network control and other advanced strategies are widely used. Although these methods perform well in dealing with uncertainties, their core design idea often focuses on a single index - minimizing the trajectory tracking error. A controller that only pursues minimum error may produce excessive overshoot, too long convergence time or become unstable under disturbance, which is absolutely unacceptable in rehabilitation training. Therefore, the design of the controller must shift from "single error index" to "comprehensive management of multiple performance indexes", that is, to simultaneously and explicitly constrain the tracking error, overshoot, convergence speed, etc.

[0004] The preset performance control is an advanced control idea to solve the above problems. It sets the convergence speed, steady-state accuracy and maximum allowed overshoot of the error in advance by designing a performance function. The preset performance control can ensure that the transient and steady-state performance of the system is constrained within the preset boundary, thereby providing strong performance guarantee. However, the traditional preset performance control has a key defect: its performance boundary is pre-set and static. This becomes very dangerous when dealing with sudden situations in rehabilitation training (such as the patient's involuntary shaking due to sudden pain). The fixed boundary may conflict with the sudden disturbance, forcing the controller to output a large control amount to maintain the constraint, which in turn triggers system oscillation or even instability, endangering patient safety.

[0005] In recent years, although a few studies have begun to focus on the dynamic adjustment of PPC boundary (such as dealing with input saturation, discontinuous trajectory, etc.), there is a significant research gap in how to handle the conflict between strong external disturbance and performance constraints, including: (1) Coupling problem: In the complex pneumatic artificial muscle driven upper limb rehabilitation exoskeleton robot system, unmodeled dynamics and external disturbances are tightly coupled, making it difficult to separate and accurately estimate disturbances alone.

[0006] (2) Lack of connection between disturbance and boundary: Existing methods fail to establish an effective connection mechanism between "external disturbance / performance degradation → dynamic adjustment of performance boundary".

[0007] (3) Lack of decision mechanism: How to determine when the system is at the critical point of "performance degradation"? How to adjust the boundary (timing, amplitude and rate of adjustment)? These problems lack theoretical guidance and solutions.

[0008] After searching the existing technology, the following existing technology implementation schemes are found: In Chinese patent application CN 120606376A (publication number), a double pneumatic artificial muscle control method and system with overshoot constraint are proposed. The state constraint in this method has no dynamic adjustment capability and does not consider the performance degradation of the system caused by strong external disturbance.

[0009] In Chinese patent application CN 116160442A (publication number), a mechanical arm system preset performance dynamic surface trajectory tracking control method is proposed. The preset performance is used to constrain the tracking performance of the system, and the use of preset performance will generate a large driving signal when the state reaches the system boundary, making the system recover to the preset boundary. However, this strong driving signal often leads to system jitter or even instability, and cannot guarantee safety when facing strong external disturbance.

[0010] In Chinese patent application CN 120161768A (publication number), a preset performance control method for ship dynamic positioning system under saturation constraint is proposed. By setting a safety factor and compensating for input saturation in the control law, the problem of simultaneously ensuring saturation constraint and preset performance is solved. Although it is a preset performance, its boundary is still a fixed boundary, and the influence of external disturbance is not considered.

[0011] In summary, the existing upper limb rehabilitation exoskeleton control technology based on pneumatic artificial muscle has the following limitations: 1. Single defect of traditional control strategy: Model prediction, adaptive control method can handle nonlinearity, but usually only aims to minimize tracking error, which may lead to overshoot, slow convergence or even instability, and cannot meet the comprehensive requirements of multiple performance indicators (such as convergence time and overshoot) in rehabilitation training.

[0012] 2. The boundary in traditional preset performance control is fixed: Although the conventional preset performance can constrain multiple performance indicators, the boundary is fixed after the steady state. When the system encounters a sudden situation (such as a sudden shaking of a patient), the fixed boundary cannot be adjusted adaptively, which may lead to constraint conflicts, endangering the stability and safety of the system.

[0013] 3. Conflict problem between external disturbance and performance constraint: When the system is subjected to strong external disturbance, the performance constraint may seriously conflict with the disturbance effect. Existing research is difficult to establish an effective connection between disturbance and performance constraint in a complex nonlinear system coupled with unmodeled dynamics and disturbance, and there is a lack of boundary dynamic adjustment mechanism when performance degrades.

[0014] 4. Efficiency challenge of disturbance processing and approximate learning: For the lumped disturbance (unmodeled dynamics + external disturbance) in the system, how to efficiently and accurately approximate and integrate into the control law while accelerating the learning convergence process of the control strategy is a key challenge.

[0015] Therefore, the existing technology in the control of upper limb rehabilitation robots faces a "triple dilemma": Conflict between precision and complexity: traditional backstepping method has high precision but explosive calculation; dynamic surface control is simple in calculation but sacrifices precision due to filtering error.

[0016] Conflict between performance and safety: fixed preset performance control either conserves performance for safety or risks high performance.

[0017] Conflict between model dependence and adaptability: model-based design (such as disturbance observer) has a sharp performance drop when the model is uncertain or the disturbance changes. SUMMARY

[0018] The purpose of the present application is to provide a soft specified performance reinforcement learning control method and system for rehabilitation exoskeleton, which combines "dynamic adjustment mechanism of soft constraint" with "intelligent learning and decision mechanism" to solve the balance problem of safety, accuracy and adaptability of upper limb rehabilitation robots in the face of uncertainty.

[0019] The first aspect of the present application is to provide a soft specified performance reinforcement learning control method for rehabilitation exoskeleton, which is a soft preset performance based reinforcement learning control method, comprising: S1, establishing a soft specified performance reinforcement learning control law based on a dynamic soft constraint mechanism and an online learning and compensation mechanism; wherein the dynamic soft constraint mechanism is used to intelligently coordinate control safety and high performance, and the online learning and compensation mechanism is used to intelligently coordinate calculation complexity, model dependence and control precision; S2, controlling the rehabilitation exoskeleton based on the soft specified performance reinforcement learning control law.

[0020] Preferably, the S1 comprises: S11, establishing a dynamics model for the upper limb rehabilitation exoskeleton robot system driven by pneumatic artificial muscle; S12, designing a tunnel type preset performance boundary for the dynamics model to constrain the tracking error of the upper limb rehabilitation exoskeleton robot system; the tunnel type preset performance boundary dynamically adjusts the boundary based on the soft preset performance function of error change, and has an intelligent decision mechanism for dynamically adjusting the performance boundary according to the system state, wherein the system state includes any one or more of the tracking error, the error derivative and the control input size; S13, designing a safety boundary and a soft boundary for the dynamics model; when the upper limb rehabilitation exoskeleton robot system is not subjected to strong external disturbance, the tracking error is constrained within the safety boundary; when the tracking error of the upper limb rehabilitation exoskeleton robot system exceeds the safety boundary, the soft boundary is dynamically adjusted to ensure the safety of the upper limb rehabilitation exoskeleton robot system, and at this time the tracking error is constrained within the soft boundary; S14, establishing a reinforcement learning framework based on the actor-critic strategy or an advanced estimator capable of online and adaptive estimation and compensation of unknown dynamics and disturbances, wherein the reinforcement learning framework includes an actor network and a critic network, the critic network is used to approximate the long-term cost function of the upper limb rehabilitation exoskeleton robot system, and the actor network is used to approximate the lumped disturbance of the upper limb rehabilitation exoskeleton robot system, the lumped disturbance includes the sum of all uncertainties acting on the upper limb rehabilitation exoskeleton robot system; S15, based on the tunnel type preset performance boundary, the safety boundary and the soft boundary, and the reinforcement learning framework or the advanced estimator capable of online and adaptive estimation and compensation of unknown dynamics and disturbances, the soft prescribed performance reinforcement learning control law is established by the command filter backstepping method with error compensation, the adaptive fuzzy / neural network backstepping control method, the model predictive control method or the sliding mode control method.

[0021] Preferably, the S11 comprises: The upper limb rehabilitation exoskeleton robot system driven by pneumatic artificial muscle is described as shown in the following formula (1): (1) ; Wherein, denotes the vector of each joint angle, denotes the vector of pneumatic artificial muscle input air pressure, denotes the external disturbance vector, denotes the matrix related to the three-element model, and I, C, g respectively represent an inertia matrix, a centrifugal coriolis matrix and a gravity vector; Simplifying the above equation (1) can obtain equation (2): (2) ; wherein, is a known nominal part in the middle, is an unknown part in the middle, is a known matrix, is a lumped disturbance vector containing unknown dynamics and external disturbances.

[0022] Preferably, the S12 comprises: The designed tunnel type preset performance is shown in the following equation (3): (3) ; wherein, is an adjustable parameter, is an initial tracking error, has the same form as and is designed as shown in equation (4): (4) ; wherein, is an adjustable parameter, is a convergence time, , is a positive adjustable parameter, has the same form as ; Or the tunnel type preset performance is designed based on the performance boundary regulator of fuzzy reasoning, the boundary adjustment of online optimization based on performance index or the preset performance control principle of multi-model switching.

[0023] Preferably, the S13 comprises: The designed safety boundary is shown in equation (5): (5) ; wherein, are all positive adjustable parameters; is a designed safety upper boundary, is a designed safety lower boundary; The designed soft boundary is shown in equation (6): (6) ; wherein, , is an output of the middle system; are all positive adjustable parameters; These are the soft lower boundary and the soft upper boundary. Even if the system experiences performance degradation and the tracking error exceeds the safety boundary, it is always constrained within the soft boundary. .

[0024] Preferably, S14 includes: Define the long-run cost function As shown in equation (7): (7); in, , It is a constant. It is a positive definite matrix. It is the compensated error vector; The long-run cost function is approximated using RBFNN, as shown in equation (8): (8); in, For the bounded approximation error, It is the actual weight vector; It is an activation function; It is the number of hidden layer nodes; Define the estimate of the long-run cost function As shown in equation (9): (9); in, It is the input state vector of the executor network. They are The estimated value; The estimation error of the long-run cost function is defined as shown in equation (10): (10); in, It is an adjustable parameter; The weight update rate of the executor network is defined as shown in Equation (11): (11); Case 1 is or Case 2 is , It is an adjustable parameter vector. It is the learning rate of the evaluator network. It is defined as shown in equation (12): (12); in, ,and, yes The gradient; The error transformation is introduced as shown in equation (13): (13); in, and It is the error after conversion. Defined in formula (2), The output of the following second-order filter is shown in equation (14): in, It is the virtual control law to be designed. It is a positive, adjustable parameter. It is an intermediate state; To compensate for the filtering error introduced by the filter, the following compensation mechanism is designed, as shown in equation (15): (15) in, All are adjustable normal values. , and It is the output state of the compensation mechanism, used to compensate for filtering errors; The compensated error is defined as shown in equation (16): (16); To construct soft boundaries, the following intermediate system is defined, as shown in equation (17): (17) in, , It is a positive, adjustable parameter. It is the output of the intermediate system. It is the input of the intermediate system. The form is shown in the following formula (18): (18); in, The two variables above are Components and All parameters are adjustable. ; The virtual control law is designed as shown in equation (19): (19); in, It is a virtual control law, in the above formula and The expression is defined as , , and is the output state of the intermediate system in (12), and are upper bounds of and respectively, and the value is an adjustable parameter.

[0025] Approximation of system with lumped disturbance using actor network As shown in equation (20): (20); where, is the real weight vector of the actor network, is the activation function, is the input vector, is the bounded approximation error, n ai is the number of hidden layer neurons; Definition of estimation error of actor network As shown in equation (21): (21); where, , K I is a positive adjustable parameter, is the ideal value of the cost function, is the estimated value of ; is the estimated value of the long-term cost function; Updating rate of actor network weight As shown in equation (22): (22); where, is the component of the updating rate , whose expression is defined as Case 1 is or , and Case 2 is , is a constant vector, is designed as shown in equation (23): (23); wherein the advanced estimator capable of estimating and compensating unknown dynamics and disturbances online and adaptively comprises an adaptive fuzzy / neural network observer, a high-order sliding mode observer or an extended state observer.

[0026] Preferably, the soft specified performance reinforcement learning control law of the S15 is established by the command filter backstepping method with band error compensation, and the method comprises the following steps: The soft specified performance reinforcement learning control law is shown in formula (24): (24) Wherein, is an adjustable parameter.

[0027] The second aspect of the present application provides a soft specified performance reinforcement learning control system for a rehabilitation exoskeleton, which is implemented by a soft preset performance based reinforcement learning control method, and is used to implement the method of the first aspect, and comprises: A control law construction module (101) is configured to establish a soft specified performance reinforcement learning control law based on a dynamic soft constraint mechanism and an online learning and compensation mechanism; wherein the dynamic soft constraint mechanism is used to intelligently coordinate control safety and control high performance, and the online learning and compensation mechanism is used to intelligently coordinate calculation complexity, model dependency and control accuracy. A control module (102) is configured to control the rehabilitation exoskeleton based on the soft specified performance reinforcement learning control law.

[0028] The third aspect of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores a plurality of instructions, and the processor is configured to read the instructions and execute the method of the first aspect.

[0029] The fourth aspect of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a plurality of instructions, and the plurality of instructions can be read and executed by a processor to execute the method of the first aspect.

[0030] The method and system of the present application have the following advantages: The application is to realize precise and safe trajectory tracking control of an upper limb rehabilitation exoskeleton robot driven by a pneumatic artificial muscle, and a soft preset performance based reinforcement learning control method is proposed; the tunnel type preset performance is used to improve the transient and steady state tracking performance of the system, and the convergence time of the system tracking error can be preset. Compared with the relatively loose funnel type preset performance, the tunnel type preset performance provides faster convergence speed and smaller overshoot; in order to ensure the safety of the system, especially for the problem of performance degradation caused by sudden conditions (such as strong external disturbance), a soft preset performance is designed; the proposed soft preset performance includes a safety boundary and a soft boundary. When the system is not subjected to strong external disturbance, the error is constrained within the safety boundary; when the system is subjected to external disturbance and the error exceeds the safety boundary, the soft boundary is adjusted dynamically to temporarily relax the performance constraint to ensure the safety of the system; the dynamic adjustment of the soft boundary is temporary, and when the external disturbance disappears, the part adjusted by the soft boundary will converge to zero exponentially; the evaluator-performer based reinforcement learning framework is used to handle the lumped disturbance of the system and improve the robustness of the system; the evaluator network is used to approximate the long-term cost function of the system, and the performer network is used to handle the lumped disturbance of the system; finally, the soft preset performance based reinforcement learning control law is designed, which introduces the command filtering technology in the derivation of the control law to avoid the differential explosion problem of the traditional backstepping method; the method and system finally realize accurate trajectory tracking of the exoskeleton robot while effectively improving the safety of the system.

[0031] Specifically: 1. A mechanical arm system preset performance dynamic surface trajectory tracking control method is proposed in the prior art, which uses preset performance to constrain the tracking error of the system. Compared with the soft preset performance designed in the application, the soft preset performance has dynamic adjustment capability and can temporarily relax the constraint to ensure the safety of the system when the system is subjected to strong external disturbance. The preset performance boundary used in the comparative method does not have the ability of dynamic adjustment, and when the tracking error of the system approaches the preset performance boundary, a large driving signal will be generated to force the system to return to the preset boundary; when the tracking error approaches the boundary, the driving signal will approach infinity; such a large driving signal often causes the system to lose stability directly, which cannot guarantee the safe operation of the system. The specific comparison advantages are as follows: (1) The contradiction between "safety" and "high performance" is solved, and intelligent trade-off is realized: The limitation of the prior art is that there is a fundamental contradiction in the fixed performance boundary. If the boundary is set too tight (pursue high performance), under strong disturbance, the controller may generate a huge control amount to maintain this rigid constraint, leading to actuator saturation, even causing system instability, and safety cannot be guaranteed. If the boundary is set too loose (ensure safety), the control accuracy and transient performance of the system will be greatly compromised under normal conditions. The beneficial effects of the present application are that the soft preset performance mechanism of the present application no longer pursues to meet the most stringent performance index at any time, but dynamically balances safety and performance according to the actual working condition. Under normal conditions, high accuracy is pursued, and under abnormal conditions, safety is prioritized. This flexible compromise mechanism enables the system to have higher intelligence and practicality in complex and variable environments.

[0032] (2) Significantly enhance the robustness and survivability of the system under extreme conditions: The limitation of the prior art is that when facing strong external disturbance with huge amplitude or dramatic dynamics, the fixed performance boundary may be forcibly broken, leading to the failure of the preset performance control law, and the loss of control of the system tracking error. The beneficial effects of the present application are that the "temporary relaxation constraint" mechanism of the present application provides a valuable "buffer interval" for the system. When strong disturbance occurs, the system allows the error to temporarily increase, but this is a controlled and intentional relaxation, not out of control. This effectively avoids the unlimited growth of control force and the deep saturation of the actuator, ensuring the stability basis of the control system, so that the system can "safely pass through" the peak period of disturbance, and then restore high-precision tracking when the disturbance subsides. This greatly improves the robustness and survivability of the system under extreme, unexpected conditions.

[0033] 2. In the prior art, a fixed-time disturbance observer is designed to handle the lumped disturbance of the system. In contrast, the present application uses reinforcement learning to handle the lumped disturbance, with the following comparative advantages: (1) Higher adaptability and robustness: from "accurate known" to "autonomous learning": The limitation of the prior art is that the performance of the fixed-time disturbance observer depends heavily on the accuracy of the system model. When there are unmodeled dynamics in the system, or the disturbance characteristics change dramatically (such as large-scale parameter jumps), the observer based on the fixed model may not be able to accurately estimate, leading to performance degradation or even instability. The beneficial effects of the present application are that the reinforcement learning framework of the present patent does not depend on the accurate mathematical model of the system. Two RBF neural networks continuously learn and adapt to changing disturbances and system dynamics by updating their weights in real time. Even in the face of unprecedented or time-varying disturbances, the system can autonomously adjust its strategy through interactive data, demonstrating strong learning robustness and environmental adaptability.

[0034] (2) From "disturbance rejection" to "performance optimization": global performance optimization is achieved: The limitation of the prior art is that the only goal of the disturbance observer is to accurately estimate and cancel the disturbance, which belongs to a kind of "local" compensation behavior. The beneficial effects of the present application are that the evaluator network is introduced to approximate the long-term cost function. This means that the goal of the system is not only to cancel the current disturbance, but also to minimize a long-term performance index. The executor network is updated under the "guidance" of the evaluator network, and the goal is to find an optimal disturbance compensation strategy to minimize the long-term comprehensive performance of the system. This realizes the leap from passive "disturbance rejection" to active "intelligent optimization control".

[0035] 3. The prior art designs a fixed-time dynamic surface control law, and the present application introduces command filtering in the derivation of the control law to avoid the problem of differential explosion, and designs a compensation mechanism for the filtering error caused by command filtering. The specific comparative analysis is as follows: (1) Higher control accuracy is achieved: The limitation of the prior art is that the first-order filter in the dynamic surface control replaces the analytical derivation, which introduces a filtering error (i.e. the difference between the input and output of the filter). This error is transmitted backward along the backstepping design steps and accumulates, eventually leading to a decrease in the tracking accuracy of the system and producing a steady-state error. This is an inherent theoretical defect of the dynamic surface control method. The beneficial effects of the present application are that the present application actively designs a compensation signal to construct an error compensation system. The compensator can observe and cancel the error introduced by the command filter in real time. This means that the system can enjoy the computational simplicity of command filtering without sacrificing the final tracking accuracy.

[0036] (2) The flexibility of the design and the robustness of the system are improved: The limitation of the prior art is that in the dynamic surface control, the time constant of the filter is a key but contradictory design parameter: a smaller constant (wide bandwidth) can reduce the error, but will result in a derivative estimate containing too much noise; a larger constant (narrow bandwidth) will smooth the signal, but will introduce more phase lag and error. The beneficial effects of the present application are that since the present application has an independent error compensation loop, it is no longer as sensitive to the selection of the bandwidth of the command filter as the dynamic surface control. A relatively conservative filter parameter can be selected to ensure smoothness, and the task of restoring accuracy is left to the high-performance compensator. This decoupled design brings greater flexibility. At the same time, the dynamic suppression ability of the compensator to the error also makes the system more robust to internal parameter perturbations and external disturbances. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the specific embodiments or the related art of the present application, the drawings needed to be used in the specific embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0038] Figure 1 A soft-guaranteed performance reinforcement learning control method flow chart for a rehabilitation exoskeleton is provided according to an embodiment of the present application. Figure 2 A physical principle diagram of a pneumatic artificial muscle driven upper limb rehabilitation exoskeleton robot is provided according to an embodiment of the present application. Figure 3 A control block diagram of a precise and safe trajectory tracking control method of a pneumatic artificial muscle driven upper limb rehabilitation exoskeleton robot corresponding to a soft-guaranteed performance reinforcement learning control law is provided according to an embodiment of the present application. Figure 4 A soft-guaranteed performance reinforcement learning control system architecture diagram for a rehabilitation exoskeleton is provided according to an embodiment of the present application. Figure 5 An electronic device structure diagram is provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0039] The technical solutions of the present application will be described in detail below with reference to the drawings. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0040] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance.

[0041] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connecting" should be understood in a broad sense, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be direct connection, or indirect connection through intermediate medium, or internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0042] The following embodiments aim to solve the following core technical problems: (1) Realize the cooperative optimization of multiple performance indicators: design a control method that can not only ensure high-precision trajectory tracking, but also strictly constrain and manage multiple performance indicators such as overshoot and convergence time, and comprehensively improve the efficiency and quality of rehabilitation training.

[0043] (2) Give dynamic adaptive ability to performance boundary: solve the safety problem that fixed performance boundary may cause in sudden situations, design a soft performance boundary that can be dynamically relaxed or tightened according to the system in real time, and balance the optimal performance under the premise of safety.

[0044] (3) Establish the association of disturbance and constraint and intelligently adjust the boundary: clearly propose the concept of "safety boundary" and build its linkage mechanism with performance boundary. When the system detects that the tracking error exceeds the safety tolerance due to disturbance, it can automatically trigger the adjustment strategy of temporarily relaxing the performance boundary, thereby resolving the conflict and ensuring the safe and stable operation of the system.

[0045] (4) Efficiently fuse intelligent approximation and constraint learning: use the Actor-Critic framework of reinforcement learning and the approximation ability of neural network to efficiently estimate and compensate for the lumped disturbance in the system; and integrate the preset performance constraint into the learning framework to significantly reduce the policy search space, accelerate the convergence speed of the control strategy, and improve the learning efficiency.

[0046] The professional terms and their meanings involved in the present embodiment are as follows: (1) Lumped disturbance: refers to the sum of all uncertainties acting on the controlled object. In the context of the present patent, it is a unified concept, which specifically includes unmodeled dynamics, parameter variations, etc. of the system and external disturbances received by the system.

[0047] (2) Soft preset performance: a dynamic performance constraint mechanism. It specifies the transient and steady-state behavior of the system tracking error (such as convergence speed, overshoot, and steady-state error band) through a time-varying, state-dependent performance boundary function.

[0048] (3) Reinforcement learning framework based on evaluator-performer: an adaptive optimization structure that uses two neural networks to learn in parallel.

[0049] (4) Command filtering technique: A numerical technique used in the backstepping design framework to generate the virtual control law and its derivative. It takes the virtual control law calculated in the last design step as the input to a filter (usually a first or second order linear filter), and the output of the filter is the smoothed virtual control command, whose analytical derivative can be directly obtained. The main purpose of this technique is to fundamentally avoid the "analytical differentiation explosion" problem in backstepping, greatly simplifying the controller structure, and introducing amplitude and rate limits by setting filter parameters.

[0050] (5) Filtered error compensation signal: A dynamic signal specially designed to offset the phase lag and amplitude attenuation introduced by the command filter dynamics.

[0051] Embodiment One As shown in Figure 1 , the embodiment provides a soft specified performance reinforcement learning control method for a rehabilitation exoskeleton, which is a soft preset performance based reinforcement learning control method, comprising: S1, establishing a soft specified performance reinforcement learning control law based on a dynamic soft constraint mechanism and an online learning and compensation mechanism; wherein the dynamic soft constraint mechanism is used to intelligently coordinate control safety and control high performance, and the online learning and compensation mechanism is used to intelligently coordinate calculation complexity, model dependency and control accuracy.

[0052] S2, controlling the rehabilitation exoskeleton based on the soft specified performance reinforcement learning control law.

[0053] As a preferred implementation, the S1 comprises: S11, establishing a dynamics model for a pneumatic artificial muscle driven upper limb rehabilitation exoskeleton robot system, wherein the pneumatic artificial muscle driven upper limb rehabilitation exoskeleton robot physical schematic diagram is as shown in Figure 2 .

[0054] In this embodiment, the pneumatic artificial muscle driven upper limb rehabilitation exoskeleton robot system is described as shown in the following formula (1): (1) ; wherein, denotes the vector of each joint angle, denotes the vector of pneumatic artificial muscle input air pressure, denotes the external disturbance vector, denotes the matrix related to the three-element model, and denote the inertia matrix, the centrifugal Coriolis matrix and the gravity vector, respectively.

[0055] Simplifying the above formula (1) can obtain formula (2): (2) ; wherein, is a known part in the middle, is a part unknown in the middle, is a known matrix, is a lumped disturbance vector containing unknown dynamics and external disturbances.

[0056] S12, a tunnel type preset performance boundary is designed for the dynamics model, to constrain the tracking error of the upper limb rehabilitation exoskeleton robot system; In this embodiment, the tunnel type preset performance is designed as shown in the following formula (3): (3) ; wherein, is an adjustable parameter, is an initial tracking error, has the same form as and is designed as shown in formula (4): (4) ; wherein, is an adjustable parameter, is a convergence time, , is a positive adjustable parameter, has the same form as .

[0057] This embodiment designs a soft preset performance function based on error change to dynamically adjust the boundary. As an alternative technical solution, the alternative solution for "performance constraint processing" is any intelligent decision mechanism that can dynamically adjust the performance boundary according to the system state (such as tracking error, error derivative, control input size, etc.), which should be considered as an equivalent alternative of the present application. For example: A. Performance boundary regulator based on fuzzy reasoning: design a fuzzy system, whose input is tracking error and error change rate, and output is the parameter of performance boundary function (such as convergence speed, steady state error band), to realize experience-based, smooth dynamic constraint adjustment.

[0058] B. Boundary adjustment based on online optimization of performance index: define an index reflecting the instantaneous performance of the system (such as control input saturation, error severity), when the index exceeds a certain threshold, trigger the optimization algorithm, and online solve a new set of more relaxed performance boundary parameters.

[0059] C. Prescribed performance control with multiple model switching: Prescribe multiple sets of performance boundary functions (e.g. "high performance mode", "safe mode", "robust mode"), and design a switching logic to switch between the performance boundaries of different modes based on the magnitude of the external disturbance estimate or the urgency of the system state.

[0060] S13, designing a safety boundary and a soft boundary for the dynamic model; wherein, when the upper limb rehabilitation exoskeleton robot system is not subjected to strong external disturbance, the tracking error is constrained within the safety boundary; when the tracking error of the upper limb rehabilitation exoskeleton robot system exceeds the safety boundary, the soft boundary is dynamically adjusted to ensure the safety of the upper limb rehabilitation exoskeleton robot system, and at this time the tracking error is constrained within the soft boundary; In this embodiment, then, on the basis of step S12, the safety boundary is designed as shown in formula (5): (5) ; wherein, are positive adjustable parameters, is the designed safety upper boundary, is the designed safety lower boundary.

[0061] The soft boundary is designed as shown in formula (6): (6) ; wherein, , is the output of the intermediate system, the specific form of which will be introduced later; are positive adjustable parameters; are the soft lower boundary and the soft upper boundary respectively, that is, even if the system performance degrades and the tracking error exceeds the safety boundary, it is always constrained within the soft boundary, that is .

[0062] S14, establishing a reinforcement learning framework based on the actor-critic strategy, wherein the reinforcement learning framework includes a critic network and an actor network, the critic network is used to approximate the long-term cost function of the upper limb rehabilitation exoskeleton robot system, and the actor network is used to approximate the lumped disturbance of the upper limb rehabilitation exoskeleton robot system, the lumped disturbance including the sum of all uncertainties acting on the upper limb rehabilitation exoskeleton robot system; In this embodiment, in order to handle the lumped uncertainty of the system, a reinforcement learning based on the actor-critic structure is used, wherein the critic network is used to approximate the long-term cost function, and the actor network is used to approximate the lumped uncertainty of the system.

[0063] Define the long-term cost function As shown in equation (7): (7); in, , It is a constant. It is a positive definite matrix. This is the compensated error vector, and its expression will be given later.

[0064] Since radial basis function neural networks (RBFNNs) have strong nonlinear approximation capabilities, RBFNNs are used to approximate the long-run cost function, as shown in equation (8): (8); in, For the bounded approximation error, It is the actual weight vector; It is an activation function; It represents the number of hidden layer nodes.

[0065] Furthermore, the estimated value of the long-run cost function is defined. As shown in equation (9): (9); in, It is the input state vector of the executor network. They are The estimated value.

[0066] The estimation error of the long-run cost function is defined as shown in equation (10): (10); in, It is an adjustable parameter.

[0067] Next, the weight update rate of the executor network is defined as shown in equation (11): (11); Case 1 is or Case 2 is , It is an adjustable parameter vector. It is the learning rate of the evaluator network. It is defined as shown in equation (12): (12); in, ,and, yes The gradient.

[0068] Next, the error transformation is introduced as shown in equation (13): (13); Among them, among them, and It is the error after conversion. Defined in formula (2), The output of the following second-order filter is shown in equation (14): in, It is the virtual control law to be designed. It is a positive, adjustable parameter. It is an intermediate state.

[0069] To compensate for the filtering error introduced by the filter, the following compensation mechanism is designed, as shown in equation (15): (15) in, All are adjustable normal values. , and The output state of the compensation mechanism is used to compensate for filtering errors.

[0070] The compensated error is defined as shown in equation (16): (16); To construct soft boundaries, the following intermediate system is defined, as shown in equation (17): (17) in, , It is a positive, adjustable parameter. It is the output of the intermediate system. It is the input of the intermediate system. The form is shown in the following formula (18): (18); in, The above two variables are Components and All parameters are adjustable. ; The virtual control law is designed as shown in equation (19): (19); in, It is a virtual control law, in the above formula and The expression is defined as , , and is the output state of the intermediate system in (12), and are upper bounds of and , respectively, and the value is an adjustable parameter.

[0071] Collective perturbation of an approximated system using an actor network as shown in equation (20): (20); where, is the true weight vector of the actor network, is the activation function, is the input vector, is the bounded approximation error, n ai is the number of hidden layer neurons.

[0072] Next, the estimation error of the actor network is defined as as shown in equation (21): (21); where, , K I is a positive adjustable parameter, is the ideal value of the cost function, is the estimation value of ; is the estimation value of the long-term cost function.

[0073] Updating rate of the actor network weights as shown in equation (22): (22); where, is the component of the updating rate , whose expression is defined as Case 1 is or , and Case 2 is , is a constant vector, is designed as shown in equation (23): (23); As an alternative to the "lumped disturbance handling" using the critic-actor based reinforcement learning framework, where the actor network is specialized for approximating the lumped disturbance of the system, any advanced estimator capable of estimating and compensating unknown dynamics and disturbances online and adaptively can be combined with the soft preset performance control of the present application to form an alternative. For example: A. Adaptive fuzzy / neural network observer: using the universal approximation property of fuzzy logic systems or neural networks (such as RBFNN, BPNN) to learn the lumped disturbance online, and update its parameters or weights in real time through adaptive law.

[0074] B. High-order sliding mode observer: using the exact robust differentiator property of high-order sliding mode to accurately estimate the disturbance in finite time, and use it for feedforward compensation. Its estimation accuracy and convergence speed can meet the high-performance control requirements.

[0075] C. Extended state observer: as the core component of active disturbance rejection control, ESO extends the lumped disturbance into a new state variable and observes it in real time. Combining it with soft preset performance control, active compensation of disturbance can also be achieved.

[0076] S15, based on the tunnel type preset performance boundary, safety boundary and soft boundary, and the reinforcement learning framework, the soft preset performance reinforcement learning control law is established by the command filter backstepping method with error compensation, and the soft preset performance reinforcement learning control law is shown as formula (24): (24) Wherein, is an adjustable parameter.

[0077] The soft preset performance reinforcement learning control law solves the problem of accurate and safe trajectory tracking control of the upper limb rehabilitation exoskeleton robot driven by the pneumatic artificial muscle. The control chart of the control method is as shown in Figure 3 .

[0078] The present application adopts the command filter backstepping method with error compensation to construct the final control law, as an alternative to "control law construction", any advanced control law design method that can handle system nonlinearity and integrate the above soft constraint and disturbance estimation information can be used as the implementation means of the present application. For example: A. Adaptive fuzzy / neural network backstepping control: using fuzzy systems or neural networks to directly approximate unknown nonlinear functions, and combining backstepping method to design control law, while integrating soft preset performance boundary as constraint condition.

[0079] B. Model predictive control: In the rolling optimization problem of MPC, the soft pre-specified performance bound is taken as the constraint of the optimization problem (which can be a hard constraint or a soft constraint), and the output of the disturbance observer is taken as the feedforward information or used for model correction, so as to realize optimal control with dynamic constraints.

[0080] C. Sliding mode control: An integral sliding mode surface is designed, and the error transformation function defined by the soft pre-specified performance bound is integrated into the design of the sliding mode surface, so as to ensure that the tracking error of the system is always constrained within the dynamically changing boundary.

[0081] The effectiveness of the control method is verified by stability analysis of the proposed control method.

[0082] Firstly, it is proved that the states of the intermediate system are bounded, and the proof process is as follows: From the definition of the input of the intermediate system, it is not difficult to see , that is, the input of the intermediate system is bounded. Next, consider , and , according to the definition of the intermediate system, it can be concluded that . Next, assume , then formula (25) is obtained: (25) This means , which is contrary to the conclusion obtained , so , and .

[0083] Similarly, it can be obtained that , so the states of the intermediate system are bounded.

[0084] Next, in order to prove that all signals of the closed-loop system are bounded, the first Lyapunov candidate function is designed as shown in formula (26): (26) Wherein, , the derivative of V1 is calculated as shown in formula (27): (27) According to the inequality , the above formula (27) can be processed to obtain formula (28): (28) Next, V a and V c are designed as shown in formula (29): (29) Wherein, When or and when .

[0085] Therefore, it can be concluded that, as long as is bounded, then is always satisfied.

[0086] Similarly, , are all normal numbers, so the weight estimates of the evaluator network and the executor network are both bounded.

[0087] Next, V2 and V3 are designed as shown in equation (30): (30); The derivative calculation of V2 is shown in equation (31): (31); where The expression of (32) where, and is bounded, .

[0088] The final derivative calculation formula of V2 is shown in equation (33): (33) The derivative calculation of V3 is shown in equation (34): (34) where, , , is bounded, and there is a constant .

[0089] The combined calculation of V1, V2, and V3 is shown in equation (35): (35) where, is bounded, and there is a constant satisfying and The above equation (35) can be rearranged as equation (36) by making the parameters satisfy (36)

[0090] Finally, the Lyapunov candidate function V4 is constructed as shown in equation (37): (37); Its derivative is shown in equation (38): (38) There exists a positive constant among them. Make The inequality holds true, and can be made true by choosing parameters. If this holds true, we can obtain equation (39): (39) in, .

[0091] In summary, the conclusion can be drawn as shown in equation (40): (40) Therefore, all signals in the closed-loop system are bounded.

[0092] Example 2 like Figure 4 As shown, this embodiment provides a soft-preset performance reinforcement learning control system for a rehabilitation exoskeleton. The control system is implemented using a reinforcement learning control method based on soft-preset performance, and is used to implement the method of Embodiment 1, including: The control law construction module 101 is used to establish a soft-constraint performance reinforcement learning control law based on a dynamic soft constraint mechanism and an online learning and compensation mechanism; wherein, the dynamic soft constraint mechanism is used to intelligently coordinate control security and control performance, and the online learning and compensation mechanism is used to intelligently coordinate computational complexity, model dependency and control accuracy.

[0093] The control module 102 is used to control the rehabilitation exoskeleton based on the soft-prescribed performance reinforcement learning control law.

[0094] The present invention also provides a memory that stores multiple instructions for implementing the method as described in Embodiment 1.

[0095] like Figure 5 As shown, the present invention also provides an electronic device, including a processor 301 and a memory 302 connected to the processor 301. The memory 302 stores a plurality of instructions, which can be loaded and executed by the processor to enable the processor to perform methods as described in Embodiments 2 and 3.

[0096] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A soft-guaranteed performance reinforcement learning control method for a rehabilitation exoskeleton, characterized by, The control method is a soft preset performance-based reinforcement learning control method, comprising: S1, establishing a soft specified performance reinforcement learning control law based on a dynamic soft constraint mechanism and an online learning and compensation mechanism; wherein the dynamic soft constraint mechanism is used for intelligently coordinating control safety and high performance, and the online learning and compensation mechanism is used for intelligently coordinating calculation complexity, model dependency and control accuracy; S2, controlling the rehabilitation exoskeleton based on the soft specified performance reinforcement learning control law.

2. The soft-guaranteed performance reinforcement learning control method for rehabilitation exoskeleton according to claim 1, wherein, The S1 comprises: S11, establishing a dynamics model for an upper limb rehabilitation exoskeleton robot system driven by a pneumatic artificial muscle; S12, designing a tunnel type preset performance boundary for the dynamics model to constrain a tracking error of the upper limb rehabilitation exoskeleton robot system; the tunnel type preset performance boundary dynamically adjusts the boundary based on a soft preset performance function of error change, and has an intelligent decision mechanism of dynamically adjusting the performance boundary according to system states, wherein the system states include any one or more of a tracking error, an error derivative and a control input size; S13, designing a safety boundary and a soft boundary for the dynamics model; when the upper limb rehabilitation exoskeleton robot system is not subjected to a strong external disturbance, the tracking error is constrained within the safety boundary; when the tracking error of the upper limb rehabilitation exoskeleton robot system exceeds the safety boundary, the soft boundary is dynamically adjusted to ensure safety of the upper limb rehabilitation exoskeleton robot system, and at this time the tracking error is constrained within the soft boundary; S14, establishing a reinforcement learning framework based on an actor-critic strategy or an advanced estimator capable of online and self-adaptive estimation and compensation of unknown dynamics and disturbances, wherein the reinforcement learning framework includes an actor network and a critic network, the actor network is used to approximate a long-term cost function of the upper limb rehabilitation exoskeleton robot system, and the critic network is used to approximate a lumped disturbance of the upper limb rehabilitation exoskeleton robot system, the lumped disturbance including a sum of all uncertainties acting on the upper limb rehabilitation exoskeleton robot system; S15, establishing the soft specified performance reinforcement learning control law by a command filter backstepping method with error compensation, an adaptive fuzzy / neural network backstepping control method, a model predictive control method or a sliding mode control method based on the tunnel type preset performance boundary, the safety boundary and the soft boundary and the reinforcement learning framework or the advanced estimator capable of online and self-adaptive estimation and compensation of unknown dynamics and disturbances.

3. The soft-guaranteed performance reinforcement learning control method for rehabilitation exoskeleton according to claim 2, wherein, The S11 comprises: An upper limb rehabilitation exoskeleton robot system driven by a pneumatic artificial muscle is described as shown in the following formula (1): (1); wherein, denotes a vector representing each joint angle, denotes a vector representing pneumatic artificial muscle input air pressure, denotes an external disturbance vector, denotes a matrix relating to a three-element model, and denote an inertia matrix, a centrifugal-Coriolis matrix, and a gravity vector, respectively; Simplifying the above formula (1) can obtain formula (2): (2); where is the known part of the nominal part of is the unknown part of is a known matrix, is a lumped disturbance vector containing unknown dynamics and external disturbances.

4. The soft-guaranteed performance reinforcement learning control method for rehabilitation exoskeleton according to claim 3, wherein, The S12 comprises: Designing a tunnel type preset performance as shown in the following formula (3): (3); wherein, is an adjustable parameter, is an initial tracking error, has the same form as and is designed as shown in equation (4): (4); wherein is an adjustable parameter, is a convergence time, , is a positive adjustable parameter, and have the same form; Or the tunnel type preset performance is based on a performance boundary adjuster of fuzzy reasoning, boundary adjustment based on online optimization of performance indicators or preset performance control principles of multi-model switching.

5. The soft-guaranteed performance reinforcement learning control method for rehabilitation exoskeleton according to claim 4, wherein, The S13 comprises: Designing a safety boundary as shown in formula (5): (5); wherein, are positive adjustable parameters; is a design safety upper bound, is a design safety lower bound; Designing a soft boundary as shown in formula (6): (6); wherein, , is the output of the intermediate system; are positive adjustable parameters; are respectively the soft lower and upper boundaries, i.e. even if the system performance degrades and the tracking error exceeds the safe boundaries, it is always constrained within the soft boundaries, i.e. .

6. The soft-guaranteed performance reinforcement learning control method for rehabilitation exoskeleton according to claim 5, wherein, The S14 comprises: Defining the long-term cost function As shown in equation (7): (7); wherein , is a constant, is a positive definite matrix, is the compensated error vector; An RBFNN is used to approximate the long-term cost function as shown in equation (8): (8); wherein, is a bounded approximation error, is a true weight vector; is an activation function; is the number of hidden layer nodes; defining an estimated value of a long-term cost function as shown in equation (9): (9); wherein, is an input state vector of the performer network, are, respectively, estimated values of The estimation error of the long-term cost function is defined as shown in equation (10): (10); wherein is an adjustable parameter; The weight update rate of the actor network is defined as shown in equation (11): (11); where Case 1 is or ; Case 2 is , is an adjustable parameter vector, is the learning rate of the evaluator network, is defined as shown in equation (12): (12); wherein and is the gradient of An error transformation is introduced as shown in equation (13): (13); where and is the converted error, is defined in equation (2), is the output of the following second order filter as shown in equation (14): wherein, is a virtual control law to be designed, is a positive adjustable parameter, is an intermediate state; To compensate for the filtering error introduced by the filter, the following compensation mechanism is designed as shown in equation (15): (15) wherein, are all adjustable positive constants, is the output state of the compensation mechanism used to compensate for filtering errors;​​ The compensated error is defined as shown in equation (16): (16); To construct a soft boundary, an intermediate system is defined as shown in equation (17): (17) wherein , is a positive adjustable parameter, is an output of the intermediate system, is an input of the intermediate system, is in the form of equation (18) as follows: (18); wherein, , the above two variables are components, and are adjustable parameters; is a design safety upper bound, is a design safety lower bound; The virtual control law is designed as shown in equation (19): (19); wherein is a virtual control law, and the expression is defined as , , and is the output state of the intermediate system in (12), and are respectively and upper bounds of the values, which are adjustable parameters; Collective perturbation using an approximating system of performers network As shown in equation (20): (20); wherein, is the true weight vector of the performer network, is an activation function, is an input vector, is a bounded approximation error, n ai is the number of hidden layer neurons; Defining an estimation error of an executor network As shown in equation (21): (21); wherein , K I is a positive adjustable parameter, is the ideal value of the cost function, is the estimate of ; and is the estimate of the long-term cost function; Update rate of performer network weights As shown in equation (22): (22); wherein is the update rate , whose expression is defined as , Case 1 is or , Case 2 is , is a constant vector, is designed as shown in equation (23): (23); The advanced estimator capable of online and adaptive estimation and compensation of unknown dynamics and disturbances includes an adaptive fuzzy / neural network observer, a high-order sliding mode observer, or an extended state observer.

7. The soft-guaranteed performance reinforcement learning control method for rehabilitation exoskeleton according to claim 6, wherein, The soft prescribed performance reinforcement learning control law established by the command filter backstepping method with error compensation of S15 includes: The soft prescribed performance reinforcement learning control law is shown in equation (24): (24) wherein is an adjustable parameter.

8. A soft-specified performance reinforcement learning control system for a rehabilitation exoskeleton, the control system being implemented by a soft-specified performance based reinforcement learning control method for implementing the method of any one of claims 1-7, characterized in that, It includes: A control law construction module (101) is configured to establish a soft prescribed performance reinforcement learning control law based on a dynamic soft constraint mechanism and an online learning and compensation mechanism; the dynamic soft constraint mechanism is used to intelligently coordinate control safety and control performance, and the online learning and compensation mechanism is used to intelligently coordinate computational complexity, model dependency, and control accuracy; A control module (102) is configured to control the rehabilitation exoskeleton based on the soft prescribed performance reinforcement learning control law.

9. An electronic device comprising a processor and a memory, the memory storing a plurality of instructions, the processor being configured to read the instructions and perform the method of any one of claims 1-7.

10. A computer-readable storage medium storing a plurality of instructions, the plurality of instructions being readable by a processor and executable to perform the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Mechanical arm system preset performance dynamic surface trajectory tracking control method

    CN116160442A

  • Preset performance control method of ship dynamic positioning system under saturation constraint

    CN120161768A

  • Double-pneumatic artificial muscle control method and system with overshoot constraint

    CN120606376A