Mechanical arm control method and system

By combining reinforcement learning and dynamic event triggering mechanism robot arm control methods, the problem that traditional methods are difficult to achieve high precision and low energy consumption under complex nonlinear systems and external perturbations is solved, and the efficient, robust and smooth control performance of robot arm systems is achieved.

CN120190818APending Publication Date: 2025-06-24ANHUI UNIV

Patent Information

Application Number
CN202510455852.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Traditional robotic arm control methods are difficult to achieve high-precision, low energy consumption and dynamic adaptive control under complex nonlinear systems and external disturbances, especially under full-state constraints, and are insufficient robustness and stability.

Method used

A robotic arm control method combining reinforcement learning and dynamic event triggering mechanism is proposed. Through the reviewer's neural network dynamic optimization control strategy, the trigger threshold and control update frequency are adaptively adjusted to ensure that the controller can respond in a timely manner when the system state changes violently or when external interference bursts.

Benefits of technology

It realizes high-precision trajectory tracking and optimal control of the robotic arm system in complex nonlinear systems and external disturbance environments, reduces the system's energy consumption and calculation burden, and enhances the system's robustness and smoothness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120190818A_ABST
    Figure CN120190818A_ABST
Patent Text Reader

Abstract

The invention discloses a mechanical arm control method and system. The method comprises the steps that a dynamic event triggering condition and a cost function are set according to error constraints of a mechanical arm system; estimating a cost function by using a commentator neural network to approach an optimal cost function, so that an optimal controller is generated when the Hamiltonian-Jacobi-Bellman estimation error is minimum; wherein the trigger threshold is adjusted according to the dynamic event trigger condition, and the update frequency of the reviewer neural network weight matrix update rule is controlled; by introducing a dynamic event triggering mechanism and a reinforcement learning method, the defects of the prior art in the aspects of self-adaptability, buffeting suppression, model errors, energy consumption optimization and the like are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robotic arm control, and particularly relates to a robotic arm control method and system. Background Art

[0002] As a key device in the fields of industrial automation, medical surgery, aerospace, etc., the control accuracy and stability of a robotic arm directly affect the task execution effect. With the complication of application scenarios, the robotic arm system needs to achieve high-precision control under high-dynamic and multi-constraint conditions. In existing robotic arm control methods, common control strategies include PID control based on periodic sampling, sliding mode control based on static event triggering, and model predictive control, etc. However, these methods all have certain limitations in complex task environments and nonlinear systems, and it is difficult to meet the control requirements of high precision, low energy consumption, and dynamic adaptability. The specific manifestations are as follows:

[0003] (1) The PID control method based on periodic sampling samples the system state and updates the control at a fixed time interval. Although a certain control stability is achieved, when the dynamic characteristics of the system change significantly or external disturbances are frequent, the fixed sampling period may cause the controller to fail to respond to system changes in a timely manner, resulting in problems such as an increase in tracking error or system oscillation. In addition, frequent control updates will also increase the computational burden and energy consumption of the system, reducing the overall operating efficiency of the system. Moreover, the PID gain parameters are usually fixed and lack self-adaptability, and it is difficult to maintain the best control performance in complex nonlinear systems. When there are state constraints or external disturbances, the fixed periodic sampling method is difficult to dynamically adjust the control strategy, which may lead to system instability or tracking failure, limiting the robustness and accuracy of the control system.

[0004] (2) The sliding mode control method based on static event triggering determines the update moment of the controller by setting a fixed triggering threshold. Although it effectively reduces the control update frequency, due to the static characteristics of the triggering condition, it is difficult to make adaptive adjustments according to the changes in the system state and external environment. When the system state changes violently or external disturbances are frequent, it may lead to frequent satisfaction of the triggering condition, resulting in overly frequent control updates and failure to effectively reduce the number of updates. If the threshold is set too large, it may cause the controller to fail to trigger an update in a timely manner when there is a large deviation in the system state, resulting in an increase in tracking error and even instability of the system. Therefore, when the dynamic range of the system changes greatly or external interference is strong, the fixed triggering threshold may lead to overly frequent triggering or sluggish response, affecting the response speed and tracking accuracy of the control system. At the same time, the chattering problem inherent in sliding mode control may cause mechanical losses to the actuator, affecting the smoothness and long-term stability of the system. Moreover, the sliding mode controller involves the adjustment of multiple hyperparameters such as sliding mode gain, saturation function boundary layer thickness, and triggering threshold parameters. There are complex coupling relationships between these parameters, making it difficult to optimize. It is difficult to accurately determine the balance between the triggering condition and the sliding mode gain, and a large number of experiments or simulations may be required for parameter optimization, increasing the design difficulty.

[0005] (3) The model predictive control method can predict and control the system state for a period of time in the future through an optimization algorithm when facing a complex multi-constrained system. However, the model predictive control method highly depends on the accuracy of the system model. In complex nonlinear systems and dynamic environments, modeling errors and unmodeled dynamic characteristics may lead to model mismatch of the model predictive control model, thereby affecting control accuracy and system stability. At the same time, the computational complexity of the model predictive control method is relatively high. In systems with high real-time requirements, control lag or instability may occur due to computational delay.

[0006] Therefore, traditional control methods such as PID control and sliding mode control often struggle to meet the requirements of modern applications when facing strong nonlinearity, external disturbances, and multi-state constraints. Especially under full-state constraints such as position, speed, and acceleration limitations, the performance of traditional methods is further restricted, making it difficult to achieve high-precision and strong-robustness control objectives.

[0007] With the development of technology, neural networks have been widely applied to the modeling and control of complex systems due to their powerful non-linear fitting ability and self-adaptability. However, traditional neural network control methods usually require high-frequency updates of control strategies, resulting in a large computational burden and difficulty in meeting real-time requirements. In addition, neural networks often lack effective constraint handling mechanisms when facing full-state constraints. Moreover, compared with static event-triggering mechanisms, dynamic event-triggering mechanisms can further adjust triggering conditions dynamically according to system states, and can better adapt to complex and changeable control environments. Therefore, in the published literature and materials, there have been studies on the combination of dynamic event-triggering mechanisms and reinforcement learning to achieve non-linear system control. For example, in the patent application document with the publication number CN118927260A, a combination of direct model-free adaptive control and reinforcement learning networks is proposed, and the event-triggering mechanism is used to solve the problems of limited system bandwidth and communication resources faced by the cloud-edge-end. This solution adopts a hierarchical control architecture, which pays more attention to the design of the system architecture, uses cloud-edge-end collaboration to improve resource utilization and real-time performance, optimizes communication resource allocation through a hierarchical computing architecture, and its control strategy mainly focuses on reducing network transmission load. The event-triggering mechanism adopted only optimizes data transmission and does not manage state constraints. In addition, although this solution realizes lightweight distributed control, its simplified control model (single-link system) and network optimization-oriented design idea make it neither require nor contain complex algorithms (such as obstacle functions or predictive constraint mechanisms) needed to handle the full-state constraints of multi-degree-of-freedom robotic arms. The event-triggering mechanism in this solution is mainly used to reduce the communication overhead between the cloud and the edge, and is suitable for lightweight and low-latency remote control scenarios, and is more suitable for distributed and resource-constrained industrial robotic arm control. Its advantage lies in reducing communication load and improving system scalability.

[0008] Another example is the research in the literature "Self-Learning and Optimal Control of Nonlinear Systems Based on Adaptive Dynamic Programming, Yu-Hong Tang, Master's Thesis of Tianjin University", which proposes to use a dynamic event-triggering mechanism and adaptive dynamic programming to design an event-triggered optimal controller for a saturated non-linear system. The event-triggered optimal control of the saturated non-linear system proposed in this solution focuses on solving the input saturation problem of general non-linear systems. It adopts an adaptive dynamic programming framework combined with an actor-critic structure, and alleviates the influence of limited control signals through a saturation compensation strategy, and is more suitable for various industrial control systems that need to handle the physical limitations of actuators. Similarly, although this solution adopts a more general adaptive dynamic programming method, its research scope is clearly limited to the input saturation problem of actuators, and it alleviates the saturation effect by optimizing the update frequency of event-triggered control, and does not involve the constraint handling of state variables such as joint angles and speeds during the movement of the robotic arm.

[0009] In summary, these two existing methods represent two different research directions of networked control and saturated system control respectively. There is neither an integrated state constraint modeling module nor a stability analysis to ensure constraint satisfaction in their control frameworks. Therefore, they cannot be directly applied to the robotic arm control scenario that requires strict state constraints. Thus, the application of the existing dynamic event-triggering mechanism in robotic arm control still faces challenges. Especially when dealing with full-state constraints and external disturbances, its robustness and stability need to be improved. In practical applications, robotic arm systems usually need to satisfy various state constraints to avoid overshoot, oscillation, or mechanical damage. Traditional constraint handling methods such as the penalty function method and the barrier function method often have difficulty in strictly satisfying the constraint conditions while ensuring control performance. Especially in a high-dynamic environment, the real-time performance and robustness of the constraint handling method become the key factors restricting system performance. Therefore, developing a control method that can effectively handle full-state constraints has become an important research direction for improving the performance of robotic arm systems.

[0010] In summary, there is an urgent need for a comprehensive control method in the field of robotic arm control that can combine the optimization ability of neural networks, the efficiency of event-triggering mechanisms, and the ability to handle full-state constraints, to overcome the deficiencies of existing technologies in aspects such as self-adaptation, chattering suppression, model error, and energy consumption optimization under complex nonlinear systems and external disturbances. Summary of the Invention

[0011] The present invention aims to solve the problems of low control accuracy and poor robustness of traditional methods under complex nonlinear systems and external disturbances.

[0012] The present invention solves the above technical problems through the following technical means:

[0013] A robotic arm control method is proposed, and the method includes:

[0014] Setting dynamic event-triggering conditions and a cost function according to the error constraints of the robotic arm system;

[0015] Using a critic neural network to estimate the cost function to approximate the optimal cost function, so as to generate an optimal controller when the Hamilton-Jacobi-Bellman estimation error is minimized;

[0016] Among them, the triggering threshold and the update frequency of the update rule of the weight matrix of the control critic neural network are adjusted according to the dynamic event-triggering conditions.

[0017] Further, the setting of the dynamic event-triggering conditions and the cost function according to the error constraints of the robotic arm system includes:

[0018] Setting a state constraint error vector according to the dynamic model of the robotic arm system;

[0019] The state constraint error vector is converted into an unconstrained error vector by using an error conversion function;

[0020] Based on the unconstrained error vector, the dynamic event-triggering condition and the cost function are set.

[0021] Further, the formula of the dynamic event-triggering condition is expressed as:

[0022]

[0023] In the formula, is the triggering error, η is an internal dynamic variable of the system, 0 < w < 1, γ and ν are positive design parameters, is the unconstrained error vector, Q and R are positive definite matrices, λ min (Q) is the minimum eigenvalue of matrix Q, θ is a positive design parameter, and |||| is the norm symbol.

[0024] Further, the formula of the triggering error is expressed as:

[0025]

[0026] In the formula, is the unconstrained error, is the unconstrained error at the triggering moment, is the derivative of the unconstrained error, is the constraint error vector, and are Lipschitz continuous and satisfy β M represents the upper bound of, β(0) = 0, and f1 > 0, f2 > 0.

[0027] Further, the formula of the cost function is expressed as:

[0028]

[0029] In the formula, is the cost function, is the unconstrained error vector, Q and R are positive definite matrices, τ is the input of the control system, ξ is the integration variable, and T is the transpose symbol.

[0030] Further, the formula of the Hamilton–Jacobi–Bellman estimation error is expressed as:

[0031]

[0032] In the formula, e H is the Hamilton–Jacobi–Bellman estimation error, is the Hamilton-Jacobi-Bellman estimation equation, is the Hamilton-Jacobi-Bellman equation.

[0033] Furthermore, the weight matrix update rule is constructed based on the Hamilton-Jacobi-Bellman estimation error and the auxiliary term, and the formula is expressed as:

[0034]

[0035] where W ∈ R m represents the ideal weight matrix vector, is the update rate of the critic neural network weight matrix under event triggering, λ1 is the neural network learning rate; is the gradient sign, represents the radial basis function vector of the critic neural network, and are Lipschitz continuous functions, is the optimal controller; is the auxiliary term, ρ is the design parameter, represents the continuously differentiable Lyapunov function, R is the positive definite matrix, λ2 is the auxiliary term learning rate, and T is the transpose sign.

[0036] Furthermore, the formula of the optimal controller is expressed as:

[0037]

[0038] where is the gradient sign, is the Lipschitz continuous function, R is the positive definite matrix, is the critic neural network weight matrix.

[0039] In addition, the present invention also proposes a robotic arm control system, and the system includes:

[0040] A dynamic event triggering condition setting module, which is used to set the dynamic event triggering condition and the cost function according to the error constraint of the robotic arm system;

[0041] An optimal cost function approximation module, which is used to estimate the cost function by using the critic neural network to approximate the optimal cost function, so as to generate an optimal controller when the Hamilton-Jacobi-Bellman estimation error is minimized;

[0042] An update module, which is used to adjust the triggering threshold and the update frequency of the update rule for controlling the critic neural network weight matrix according to the dynamic event triggering condition.

[0043] In addition, the present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the robotic arm control method as described above.

[0044] The advantages of the present invention are as follows:

[0045] (1) The present invention introduces the critic neural network in reinforcement learning, and learns the complex non-linear mapping relationship from the state space to the action space through the deep reinforcement learning method. The critic neural network dynamically optimizes the control strategy through the estimation of the cost function, enabling the system to adaptively adjust the control gain and trajectory tracking strategy under different working states; the continuous and smooth control input generated by reinforcement learning effectively suppresses the common chattering problem in sliding mode control, reduces mechanical wear and oscillation caused by frequent switching of the controller, and enhances the smoothness and tracking stability of the system. At the same time, a dynamic event-triggering mechanism is adopted, which adaptively adjusts the triggering threshold and control update frequency by real-time monitoring of the dynamic changes of the system state, ensuring that the controller can respond in a timely manner and maintain the tracking accuracy and system stability when the system state changes violently or external disturbances occur suddenly, with strong robustness.

[0046] (2) The optimal control method for the fully state-constrained robotic arm proposed by the present invention mainly aims at the problems of state overrun such as joint angles and speeds that may occur during the movement of the robotic arm. The dynamic event-triggering mechanism is used to reduce the computational burden while ensuring optimal control performance. By combining the critic neural network with the dynamic programming method, the obstacle Lyapunov function is used to handle state constraints, and the neural network update frequency is reduced through event triggering to achieve optimal control, which belongs to a refined control scheme for specific objects. The proposed control method improves the control performance in a constrained environment through the neural network and event triggering, and is more suitable for robotic arm systems with high precision and high dynamics requirements, and can explicitly handle complex constraints and optimize control performance. Compared with the related technologies, the general non-linear system scheme pays more attention to the adaptability of the control algorithm, and the event-triggering design needs to additionally consider the stability challenges brought by the saturation characteristics; while the present invention pays more attention to the optimization of specific dynamic models, and its event-triggering mechanism is mainly used to balance the computational cost of neural network updates, and is more suitable for the control of high-precision robotic arms such as surgical robots.

[0047] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0048] Figure 1 is a schematic flowchart of a robotic arm control method proposed in an embodiment of the present invention;

[0049] Figure 2It is a schematic diagram of the control principle of a robotic arm control system proposed in an embodiment of the present invention;

[0050] Figure 3 It is a schematic diagram of the structure of a robotic arm control system proposed in an embodiment of the present invention. Detailed implementation manners

[0051] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] As Figure 1 shown, an embodiment of the present invention proposes a robotic arm control method, and the method includes the following steps:

[0053] S10. Set dynamic event triggering conditions and a cost function according to the error constraints of the robotic arm system;

[0054] S20. Use a critic neural network to estimate the cost function to approximate the optimal cost function, so as to generate an optimal controller when the Hamilton-Jacobi-Bellman estimation error is minimized;

[0055] Among them, the triggering threshold and the update frequency of the update rule of the weight matrix of the control critic neural network are adjusted according to the dynamic event triggering conditions.

[0056] It should be noted that in this embodiment, the constraint error vector of the robotic arm system is defined according to the dynamic equation of the robotic arm system, and the dynamic event trigger condition and the cost function are set based on the constraint error vector. The critic neural network is used to estimate the cost function to approximate the optimal cost function, so as to generate an optimal controller. And a dynamic event trigger mechanism is adopted. By real-time monitoring the dynamic changes of the system state, the trigger threshold and the update frequency of the update rule of the control critic neural network weight matrix are adaptively adjusted. The advantages are as follows: 1) Improve the trajectory tracking accuracy and system adaptability. By adopting the dynamic event trigger mechanism, the trigger condition can be adaptively adjusted according to the robotic arm state and environmental changes, ensuring that the controller can respond in a timely manner when the dynamic characteristics of the system change or external disturbances occur, and effectively reducing the tracking error. At the same time, through the reinforcement learning method of the critic neural network, the system can dynamically learn and optimize the control strategy, improving the adaptive ability and tracking accuracy for complex nonlinear systems. 2) Suppress chattering and enhance the smoothness and stability of the system. Through the continuous and smooth control input generated by reinforcement learning, the present invention effectively suppresses the common chattering problem in sliding mode control, reducing mechanical wear and oscillation caused by frequent switching of the controller. Combined with the adaptive adjustment of the dynamic event trigger, the system can generate a smooth control trajectory while maintaining robustness, enhancing the overall stability and long-term stability of the system. 3) Reduce the system energy consumption and improve the control efficiency. Reinforcement learning dynamically balances the control accuracy and energy consumption during the learning process through a reward and punishment mechanism, generating an optimal control strategy, reducing unnecessary high-frequency switching operations, thereby reducing the overall energy consumption of the system. At the same time, the dynamic event trigger mechanism reduces unnecessary control updates, further reducing the computational burden and power consumption of the system and improving the operating efficiency of the system.

[0057] As a further preferred technical solution, in step S10: setting the dynamic event trigger condition and the cost function according to the error constraint of the robotic arm system specifically includes the following steps:

[0058] S11. Set the state constraint error vector according to the dynamic model of the robotic arm system;

[0059] Specifically, in this embodiment, according to the Lagrange equation, the dynamic equation of the n-degree-of-freedom robotic arm is deduced as:

[0060]

[0061] Where, respectively represent the joint position vector and the joint velocity vector, represents the joint acceleration vector, M(q) ∈ R n×n represents the symmetric positive definite inertia matrix, represents the Coriolis force, G(q) represents the gravity vector, τ represents the input of the control system, R n×nDenote the set of all real - valued, \(n\times n\) square matrices.

[0062] Define the system error as:

[0063]

[0064] where \(q\) d denotes that the desired target point is a constant vector, denotes the derivative of the desired target point, \(e\) q denotes the desired error, denotes the derivative of the desired error.

[0065] Define the state - constraint error vector whose derivative is:

[0066]

[0067] where denotes the derivative of \(M(q)\in\mathbb{R}\) n×n denotes a symmetric positive - definite inertia matrix, denotes the Coriolis force, \(G(q)\) denotes the gravity vector, 0 n denotes an \(n\) - dimensional vector with all elements being 0.

[0068] S12. Use the error - conversion function to convert the state - constraint error vector into an unconstrained error vector;

[0069] It should be noted that in practical applications, according to the actual working environment of the robotic arm, it is usually necessary to specify the state of the robotic arm to meet the actual working requirements. Therefore, the given constraint conditions are as follows. Given the system - state constraint \(x\) li \(<x\) i \(<x\) ui , \(x\) li and \(x\) ui respectively represent the lower and upper bounds of the state constraint. According to equation (2), the corresponding error constraint \(e\) li \(<e\) i \(<e\) ui , \(e\) li and \(e\) ui respectively represent the lower and upper bounds of the error constraint.

[0070] To find the optimal control strategy more effectively, this embodiment introduces a new error vector. By using the error - transformation function, the original constrained error vector is converted into a new unconstrained error vector as follows:

[0071]

[0072] Among them, represents the i-th element of the unconstrained error vector, Φ(e i , e li , e ui ) is an invertible strictly increasing smooth function that satisfies Φ(0, e li , e ui ) = 0.

[0073] Design the following error conversion function:

[0074]

[0075] where a > 1. According to equation (5), the converted system can be obtained as follows:

[0076]

[0077] Among them, and are Lipschitz continuous and satisfy β M represents 's upper bound, β(0) = 0, and f1 > 0, f2 > 0.

[0078] S13. Set the dynamic event trigger condition and the cost function based on the unconstrained error vector.

[0079] Specifically, this embodiment defines the trigger error and designs the dynamic event trigger condition as follows:

[0080]

[0081] Among them, is the trigger error, 0 < w < 1, γ and ν are positive design parameters, is the unconstrained error vector, Q and R are positive definite matrices, λ min (Q) is the minimum eigenvalue of matrix Q, θ is a positive design parameter, |||| is the norm symbol, η is the internal dynamic variable of the system, and its update rate is as follows:

[0082]

[0083] In the formula, represents the derivative of the internal dynamic variable, and μ represents a positive design constant.

[0084] Define the cost function as follows:

[0085]

[0086] Among them, Q and R are positive definite matrices, and ξ is the integration variable.

[0087] As a further preferred technical solution, in step S20, the cost function is estimated by using a critic neural network to approximate the optimal cost function, so that an optimal controller is generated when the Hamilton-Jacobi-Bellman estimation error is minimized, which specifically includes:

[0088] Set the cost function The corresponding Hamilton-Jacobi-Bellman equation is as follows:

[0089]

[0090] According to the Bellman optimality principle, it can be obtained that:

[0091]

[0092] Among them, τ * represents the optimal controller.

[0093] By solving The following optimal virtual controller is obtained:

[0094]

[0095] Among them, the optimal cost function is as follows:

[0096]

[0097] According to the universal approximation property, the optimal cost function can be approximated by a critic neural network:

[0098]

[0099] Among them, W ∈ R m represents the ideal weight matrix vector, represents the radial basis function vector of the critic neural network, m represents the number of neurons in the hidden layer, represents the approximation error.

[0100] According to the critic neural network, the Hamilton-Jacobi-Bellman estimation equation can be obtained:

[0101]

[0102] Therefore, according to equations (15) and (11), the Hamilton-Jacobi-Bellman estimation error can be defined as:

[0103]

[0104] To minimize the Hamilton-Jacobi-Bellman estimation error, the following error function is defined:

[0105]

[0106] Based on the Hamilton-Jacobi-Bellman estimation error and the auxiliary term, the update rule for the weight matrix of the critic neural network is as follows:

[0107]

[0108] where \(W\in\mathbb{R}^{}\) m represents the ideal weight matrix vector, \(\lambda_1\) is the learning rate of the neural network; is the gradient symbol, represents the radial basis function vector of the critic neural network, and are Lipschitz continuous, is the optimal controller; is the auxiliary term, \(\rho\) is the design parameter, \(\lambda_2\) is the learning rate of the auxiliary term, represents the continuously differentiable Lyapunov function, \(\chi\) represents the radial basis function vector of the critic neural network, \(\beta\) is the Lipschitz continuous function, \(R\) is the positive definite matrix, and \(T\) is the transpose symbol.

[0109] The optimal controller is generated when the Hamilton-Jacobi-Bellman estimation error is minimized:

[0110]

[0111] It should be noted that in this embodiment, by adjusting the trigger threshold and the update frequency of the update rule of the weight matrix of the critic neural network according to the dynamic event trigger condition, the trigger condition can be adaptively adjusted according to the state of the robotic arm and environmental changes, ensuring that the controller can respond in a timely manner when the dynamic characteristics of the system change or external disturbances occur, and effectively reducing the tracking error.

[0112] The present invention adopts a dynamic event-triggering mechanism. By real-time monitoring the dynamic changes of the system state, it adaptively adjusts the triggering threshold and control update frequency to ensure that when the system state changes violently or external disturbances occur suddenly, the controller can respond in a timely manner, maintaining the tracking accuracy and system stability. At the same time, when the system state is relatively stable, the dynamic event-triggering mechanism can reduce unnecessary control updates, lowering the system energy consumption and computational burden. In terms of optimizing the control strategy, the present invention introduces a critic neural network in reinforcement learning and learns the complex non-linear mapping relationship from the state space to the action space through the deep reinforcement learning method. The critic neural network dynamically optimizes the control strategy through the estimation of the value function, enabling the system to adaptively adjust the control gain and trajectory tracking strategy under different working states. Through continuous online learning and dynamic optimization, the system can maintain high-efficiency and stable control performance in long-term tasks and dynamic environmental changes. The continuous and smooth control input generated by reinforcement learning can effectively suppress the chattering problem caused by frequent control switching, enhancing the smoothness and tracking stability of the system. In addition, by introducing a reward and punishment mechanism, the reinforcement learning algorithm can balance the control accuracy and system energy consumption during the learning process, generating an optimal control strategy, reducing unnecessary high-frequency switching operations, lowering the overall energy consumption of the system, and improving the system operation efficiency.

[0113] In a complex non-linear system, the full-state constraints of the robotic arm, including position, velocity, and acceleration, pose higher requirements for the performance of the controller. By combining reinforcement learning with dynamic event triggering, the present invention can achieve efficient trajectory tracking and optimal control of the robotic arm while satisfying the full-state constraints. Through policy optimization, the system can minimize the tracking error and system oscillation while maintaining robustness, enhancing the overall control performance of the system.

[0114] In summary, by combining reinforcement learning with the dynamic event-triggering mechanism, the present invention constructs an optimal control method based on a critic neural network, solving the technical problems in aspects such as control accuracy, adaptability, chattering suppression, and energy consumption optimization in the prior art. Through the online learning and policy optimization of reinforcement learning, combined with the adaptive update of dynamic event triggering, the system can maintain high-efficiency and stable control performance in a complex dynamic environment, providing a better solution for robotic arm trajectory tracking and complex task execution.

[0115] In addition, as Figure 2 and Figure 3 shown, another embodiment of the present invention also proposes a robotic arm control system, including:

[0116] A dynamic event-triggering condition setting module 10, configured to set dynamic event-triggering conditions and a cost function according to the error constraints of the robotic arm system;

[0117] The optimal cost function approximation module 20 is used to estimate the cost function by using a critic neural network to approximate the optimal cost function, so as to generate an optimal controller when the Hamilton-Jacobi-Bellman estimation error is minimized;

[0118] The update module 30 is used to adjust the trigger threshold and the update frequency of the update rule for controlling the weight matrix of the critic neural network according to the dynamic event trigger condition.

[0119] As a further preferred technical solution, the dynamic event trigger condition setting module 10 specifically includes:

[0120] The constraint error vector setting unit is used to set the state constraint error vector according to the dynamic model of the robotic arm system;

[0121] The conversion unit is used to convert the state constraint error vector into an unconstrained error vector by using an error conversion function;

[0122] The condition setting unit is used to set the dynamic event trigger condition and the cost function based on the unconstrained error vector.

[0123] It should be noted that other embodiments or specific implementation methods of the robotic arm control system of the present invention can refer to the above method embodiments, which will not be elaborated here.

[0124] In addition, another embodiment of the present invention also proposes a computer program product, including a computer program, characterized in that the computer program is executed by a processor to implement the steps of the robotic arm control method as described in the above embodiments.

[0125] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise appropriately processing if necessary, and then storing it in a computer memory.

[0126] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0127] In the description of this specification, the description referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0128] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0129] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A robot arm control method, characterized in that: include: Setting dynamic event triggering conditions and cost functions according to the error constraints of the robot system; The cost function is estimated using a critic neural network to approximate the optimal cost function, so that the optimal controller is generated when the Hamilton-Jacobi-Bellman estimation error is minimized; Among them, the trigger threshold is adjusted according to the dynamic event triggering conditions and the update frequency of the critic neural network weight matrix update rule is controlled.

2. The robot arm control method according to claim 1, characterized in that: The step of setting the dynamic event triggering condition and the cost function according to the error constraint of the robotic arm system includes: Setting a state constraint error vector according to a dynamic model of the robot system; The state constraint error vector is converted into an unconstrained error vector using an error conversion function; The dynamic event triggering condition and the cost function are set based on an unconstrained error vector.

3. The robot arm control method according to claim 2, characterized in that: The formula of the dynamic event triggering condition is expressed as: wherein is the triggering error, η is the internal dynamic variable of the system, 0 < w < 1, γ and ν are positive design parameters, is the unconstrained error vector, Q and R are positive definite matrices, λ min (Q) is the minimum eigenvalue of matrix Q, θ is a positive design parameter, and ‖‖ is the norm symbol.

4. The robot arm control method according to claim 3, characterized in that: The trigger error is expressed as: In the formula, is the unconstrained error, The unconstrained error at the triggering time, is the derivative of the unconstrained error, is the constraint error vector, τ is the control input, and is Lipschitz continuous and satisfies β M express The upper bound of β(0) = 0, And f1>0, f2>0.

5. The robot arm control method according to claim 1, characterized in that: The formula of the cost function is expressed as: In the formula, is the cost function, is the unconstrained error vector, Q and R are positive definite matrices, τ is the control input, ξ is the integral variable, and T is the transposed sign.

6. The robot arm control method according to claim 1, characterized in that: The formula for the Hamilton-Jacobi-Bellman estimation error is expressed as: In the formula, e H is the Hamilton-Jacobi-Bellman estimate error, is the Hamilton-Jacobi-Bellman estimating equation, is the Hamilton-Jacobi-Bellman equation.

7. The robot arm control method according to claim 1, characterized in that: The weight matrix update rule is constructed based on the Hamilton-Jacobi-Bellman estimation error and auxiliary terms, and is publicly expressed as: Where W∈R m represents the ideal weight matrix vector, is the update rate of the critic neural network weight matrix, λ1 is the neural network learning rate; is the gradient symbol, represents the radial basis function vector of the critic neural network, and is Lipschitz continuous, is the optimal controller; is an auxiliary term, ρ is a design parameter, represents the continuously differentiable Lyapunov function, λ2 is the auxiliary term learning rate, χ represents the radial basis function vector of the critic neural network, β is the Lipschitz continuous function, R is a positive definite matrix, and T is the transposed sign.

8. The robot arm control method according to claim 1, characterized in that: The formula of the optimal controller is expressed as: In the formula, is the gradient symbol, is a Lipschitz continuous function, R is a positive definite matrix, χ is the radial basis function vector of the critic neural network, is the critic neural network weight matrix.

9. A robotic arm control system, characterized in that: include: A dynamic event trigger condition setting module, used to set dynamic event trigger conditions and cost functions according to the error constraints of the robotic arm system; An optimal cost function approximation module is used to estimate the cost function using a critic neural network to approximate the optimal cost function, so as to generate an optimal controller when the Hamilton-Jacobi-Bellman estimation error is minimized; The updating module is used to adjust the trigger threshold according to the dynamic event triggering conditions and control the updating frequency of the critic neural network weight matrix updating rule.

10. A computer program product, comprising a computer program, characterized in that The computer program is executed by a processor to implement the steps of the robot arm control method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Event-triggered single-link mechanical arm control method and system under cloud side-end cooperation

    CN118927260A

  • Robot trajectory tracking optimal control method based on an event trigger mechanism

    CN113093548A

  • Decentralized tracking control method for mechanical arm based on event triggering-neural dynamic programming

    CN113211446A

  • Modularized robot dispersion force / position optimal control method of event triggering mechanism

    CN115857353A

Cited By

  • Mechanical arm robust control method based on event triggering reinforcement learning

    CN122463188A