A control method and system for a dual-link robotic arm based on a neural network observer
By adopting a control method based on neural network observers, the problems of disturbance and attack on the robotic arm system in complex environments are solved, the control performance and stability are improved, and the precise tracking and state observation of the dual-link robotic arm are realized.
Patent Information
- Application Number
- CN202510445207.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Robotic arm systems are susceptible to external disturbances and attacks in complex environments, which can lead to a decline in control performance and may also be attacked during information transmission, affecting tracking accuracy and system stability.
A control method based on neural network observers is adopted to establish a system dynamics model of a double-link manipulator. By combining radial basis neural networks to identify unknown nonlinear dynamics, a neural network observer is designed, and a control scheme is constructed using reinforcement learning and adaptive update law. Adaptive safety control is achieved through an inverse stepping framework.
It improves the estimation accuracy of external attacks and disturbances, enables effective observation of the state of the robotic arm system, enhances control performance and stability, and reduces the tracking error of the actuator.
Smart Images

Figure CN120347734B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotic arm control technology, and in particular to a control method and system for a dual-link robotic arm based on a neural network observer. Background Technology
[0002] In recent years, robotic arm systems have unique advantages in improving production efficiency, improving working conditions, and reducing labor intensity, and their applications are becoming increasingly widespread. Especially in some complex operating environments, some repetitive or dangerous tasks are performed by robotic arms, which can reduce labor costs, improve productivity and manufacturing safety, and provide greater flexibility.
[0003] Practical robotic arm systems often face challenges from attacks and disturbances caused by complex environments such as turbulence, wind changes, mechanical resistance, unknown external conditions, and internal dynamic uncertainties. At present, in order to cope with these challenges, a variety of compensation strategies have emerged. The control scheme of robotic arms usually uses neural networks to act on the controlled system through the controller.
[0004] It is worth noting that although robotic arm systems have shown great potential in many fields, they are inherently complex and their algorithms and implementation processes are relatively cumbersome. In actual industrial production, they are easily affected by external disturbances. Therefore, by using a neural network observer to observe the disturbance state information of the dual-link robotic arm, the problem of the unavailability of the real state caused by external disturbances can be solved.
[0005] Furthermore, even though robotic arm systems have demonstrated great potential in many fields, they may encounter external attacks when sending data to controllers or actuators, thereby compromising the authenticity of the system data and rendering it unusable for control design; this has become a relatively difficult type of attack to handle; therefore, a neural network observer was designed to estimate unusable states for attacked robotic arm systems.
[0006] Furthermore, by combining control technology with neural network observers, a new control strategy and optimization path for robotic arms are generated. Specifically, a performance index function is constructed to handle optimization problems under attack and disturbance states, thereby further improving the control performance and stability of the robotic arm system and enabling the system to achieve good control performance.
[0007] Furthermore, position tracking is a fundamental problem when applying robotic arms in the engineering field. Neural network control methods have been widely used to solve the tracking problem of traditional rigid robotic arms. Due to the limitation of information transmission speed, there are often attacks and disturbances. If disturbances and attacks are ignored in the controller design, they may lead to the deterioration of system performance, especially in positioning applications, which will reduce the tracking accuracy of the actuator. Summary of the Invention
[0008] The purpose of this invention is to provide a control method and system for a dual-link robotic arm based on a neural network observer, so as to solve the problems existing in the prior art.
[0009] A control method and system for a dual-link manipulator based on a neural network observer, specifically including the following steps: Step 1. Establishing a system dynamics model of the dual-link manipulator under external attacks and disturbances; Step 2. Establishing a neural network observer based on the system dynamics model; Step 3. Establishing a control scheme for the dual-link manipulator based on reinforcement learning, the system dynamics model, and the neural network observer; Step 4. Performing adaptive safety control of the dual-link manipulator within the backstepping framework based on the control scheme.
[0010] Preferably, the system dynamics model of the dual-link robotic arm under external attacks and disturbances specifically includes:
[0011] use A system dynamics model of a dual-link robotic arm under external attacks and disturbances;
[0012] Where q = (q1, q2) T , These represent the angular position, angular velocity, and angular acceleration of a joint in a two-bar linkage robotic arm, which includes joint 1 and joint 2. q1 represents the angular position of joint 1, and q2 represents the angular position of joint 2. This represents the angular velocity of joint 1 of the robotic arm. This represents the angular velocity of joint 2 of the robotic arm. This represents the angular acceleration of joint 1 of the robotic arm. M(q) represents the angular acceleration of joint 2 of the robotic arm; M(q) represents the inertia matrix of the robotic arm. Represents the centrifugal Lippe force matrix; G(q)=(g 11 ,g 21 ) T Represents the gravitational vector, where g 11 =(m1+m2)gl1 cos(q1)+m2gl2 cos(q1+q2),g 21 =m2gl2 cos(q1+q2), where m1 and m2 represent the masses of robotic arm joint 1 and robotic arm joint 2, respectively, l1 and l2 represent the lengths of robotic arm joint 1 and robotic arm joint 2, respectively, and g represents the acceleration due to gravity. Represents the friction force matrix, τ d (t) represents the unknown external disturbance; τ(t) represents the input torque of the robotic arm joint; y = q represents the angular position of the robotic arm joint.
[0013] Preferably, after step 1, the method further includes transforming the system dynamics model of the dual-link robotic arm into a nonlinear system dynamic model with attack capability, wherein the nonlinear system dynamic model is:
[0014]
[0015] in, x i Represents the unmeasurable system state variables, i = 1, 2; u = τ(t)M -1 (q), where u represents the measurable system input, M -1 (q) represents the inverse of the robotic arm's inertia matrix; y represents the system output, and x1 represents the system output state; Represents an unknown nonlinear function. The derivatives of x1 and x2 are respectively, and D(t) = τ d (t)M -1 (q) represents an unknown external disturbance; a s (t,x i ) represents the state variable x i The attack state at time t It is x i The state after being attacked.
[0016] Preferably, a system dynamics model of a dual-link robotic arm under external attacks and disturbances is established, which further includes:
[0017] A radial basis function neural network is used to identify unknown, uncertain, and nonlinear dynamics in a two-link robotic arm, and the dynamic model is improved to determine the improved dynamic model; the improved dynamic model is as follows:
[0018]
[0019] in, Let θ1 represent the function obtained by approximation through a neural network. * Let T represent the weight vector, and T denote the transpose. Represents a basis function vector. Indicates the approximation error; A vector representing a subset of state variables. express The estimate; Let y represent the error term between the unknown system nonlinear function and the neural network approximation function, and let y represent the system output.
[0020] Preferably, based on a dynamic model, a neural network observer is established to observe the disturbance and attack state information of the dual-link robotic arm, and then the process further includes:
[0021] Based on the improved dynamic model, a neural network observer is established to observe the disturbance and attack state information of the dual-link robotic arm; the neural network observer is as follows:
[0022]
[0023] Among them, K i The values of i = 1 and 2 represent the observer gains of the design. When i = 1, K1 is the observer gain of robotic arm joint 1, and when i = 2, K2 is the observer gain of robotic arm joint 2. The estimation error is defined as... The state errors are respectively represented as follows: as well as These are state variables. The estimated value, The tables are respectively The state after differentiation, Represents the weight vector The estimate, is the basis function vector.
[0024] Preferably, a control scheme for the dual-link robot arm is established based on reinforcement learning, the dynamic model, and the neural network observer, specifically including:
[0025] Based on the gradient descent method and the stability determination method of Lyapunov function stability theory, an adaptive update law for the weights of the Actor-Critic neural network is designed.
[0026] Based on a neural network observer, an adaptive law is designed for the unknown nonlinear term according to the adaptive update law, thus forming a control scheme; the control scheme includes a virtual controller and an actual controller.
[0027] Preferably, based on a neural network observer, an adaptive law is designed for the unknown nonlinear term according to the adaptive update law to form a control scheme, specifically including:
[0028] use Construct the virtual controller;
[0029] use Construct the actual controller;
[0030] in, Represents the weights of the neural network. Representation of basis function vectors; Neural networks Used to approximate u * In the unknown dynamics, K1 and K2 represent the designed observer gains, and Γ2 represents a state variable. This indicates the pre-set reference signal y.r The derivative of θ * 1 represents the weight vector. Represents a basis function vector.
[0031] Preferably, based on the control scheme, within the backstepping framework, an adaptive update law for the uncertain parameters is designed based on the backstepping method and the neural network observer approximation design. According to the stability determination method based on gradient descent and Lyapunov function stability theory, an adaptive update law for the Actor-Critic neural network weights is designed. The adaptive update law is as follows:
[0032]
[0033] Where, r c1 ,r c2 ,r a1 ,r a2 ,r θ1 ,σ1,σ1,ω1 represent the design parameters, I m Describes an m×m identity matrix. Represents the weights of the neural network. This represents the adaptive update law for the weights of the designed neural network. For the basis function vector, Let i be a Gaussian function, i = 1, 2.
[0034] Preferably, the dual-link manipulator control system based on a neural network observer employs the aforementioned dual-link manipulator control method based on a neural network observer. Specifically, the dual-link manipulator control system based on a neural network observer includes: a system dynamics model establishment module for establishing a system dynamics model of the dual-link manipulator under external attacks and disturbances; a neural network observer establishment module for establishing a neural network observer based on the dynamics model; a control scheme establishment module for establishing a control scheme for the dual-link manipulator based on reinforcement learning, the dynamics model, and the neural network observer; and an adaptive control module for performing adaptive safety control of the dual-link manipulator within the backstepping framework based on the control scheme.
[0035] Preferably, the control scheme establishment module specifically includes: an adaptive update law design unit, used to design an adaptive update law for the weights of the Actor-Critic neural network based on the gradient descent method and the stability determination method of Lyapunov function stability theory; and a control scheme establishment unit, used to design an adaptive law for unknown nonlinear terms based on the neural network observer and the adaptive update law to form a control scheme; the control scheme includes a virtual controller and an actual controller.
[0036] Compared with the prior art, the present invention provides a dual-link robotic arm control method and system based on a neural network observer, which has the following beneficial effects:
[0037] 1. This invention takes into account that practical robotic arm systems often face challenges from attacks and disturbances caused by complex environments such as turbulence, wind changes, mechanical resistance, unknown external states, and internal dynamic uncertainties, which affect the control performance of the system. This invention designs a neural network observer for robotic arm systems. Because it takes into account the information of attacks and disturbances, it has better estimation accuracy for attacks and disturbances that change rapidly over time, thereby realizing the observation of the state, attacks, and external disturbances of the dual-link robotic arm system.
[0038] 2. The technical solution of this invention proposes an adaptive backstepping control strategy; by designing a performance index function, the designed controller enables the control of the dual-link robotic arm.
[0039] 3. The technical solution of this invention utilizes the learning ability of reinforcement learning algorithms and employs gradient descent to design an adaptive update law for the Actor-Critic neural network, enabling the designed approximate controller to approximate the actual controller under the premise of system stability. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a schematic diagram of the control method flow of the present invention;
[0042] Figure 2 This is a schematic diagram of the components of the robotic arm system of the present invention;
[0043] Figure 3 This is a schematic diagram of the dual-link robotic arm structure of the present invention;
[0044] Figure 4 This is a diagram showing the output and desired trajectory tracking of the dual-link robotic arm of the present invention;
[0045] Figure 5 This is a schematic diagram of the observation curve of the state of the dual-link robotic arm x1 based on the neural network observer without processing the external attack state according to the present invention;
[0046] Figure 6 This is an observation curve of the state of the dual-link robotic arm x1 based on the external attack state processed by the neural network observer according to the present invention;
[0047] Figure 7 This is an observation curve of the state of the double-link robotic arm x2 based on the neural network observer without processing the external attack state, according to the present invention.
[0048] Figure 8 This is an observation curve of the state of the double-link robotic arm x2 based on the external attack state processed by the neural network observer according to the present invention;
[0049] Figure 9 This is a virtual control input curve diagram of the dual-link robotic arm system based on neural networks in this invention;
[0050] Figure 10 This is a graph showing the actual control input curves of the dual-link robotic arm system based on neural networks in this invention.
[0051] Figure 11 This is a weight convergence curve of the neural network executed in this invention;
[0052] Figure 12 This is a weight convergence curve for evaluating the neural network in this invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0055] like Figure 1 As shown, a control method for a dual-link robotic arm based on a neural network observer specifically includes the following steps:
[0056] Step 1. Establish a system dynamics model of the dual-link robotic arm under external attacks and disturbances;
[0057] Step 2. Based on the dynamic model, establish a neural network observer to observe the disturbance and attack status information of the dual-link robotic arm;
[0058] Step 3. Based on reinforcement learning, the dynamic model, and the neural network observer, establish a control scheme for the dual-link robotic arm;
[0059] Step 4. Based on the control scheme, under the backstepping framework, establish a virtual controller and an actual controller to perform adaptive safety control on the double-link robotic arm;
[0060] As one example, combining a neural network observer with control addresses the problem of external attacks and disturbances affecting the control performance and stability of a robotic arm system. This optimizes the path, leading to a novel control strategy that further enhances the control performance and stability of the robotic arm system. Figure 1 The method shown is for, for example Figure 3 The dual-link robotic arm shown is used for backstepping control. The specific implementation process is as follows:
[0061] Step 1. Establish a system dynamics model of the dual-link robotic arm under external attacks and disturbances:
[0062]
[0063] in, These represent the angular position, angular velocity, and angular acceleration of a joint in a two-bar linkage robotic arm, which includes joint 1 and joint 2. q1 represents the angular position of joint 1, and q2 represents the angular position of joint 2. This represents the angular velocity of joint 1 of the robotic arm. This represents the angular velocity of joint 2 of the robotic arm. This represents the angular acceleration of joint 1 of the robotic arm. M(q) represents the angular acceleration of joint 2 of the robotic arm; M(q) represents the inertia matrix of the robotic arm. Represents the centrifugal Lippe force matrix; G(q)=(g 11 ,g 21 ) T Represents the gravitational vector, where g 11 =(m1+m2)gl1 cos(q1)+m2gl2 cos(q1+q2),g 21 =m2gl2 cos(q1+q2), where m1 and m2 represent the masses of robotic arm joint 1 and robotic arm joint 2, respectively, l1 and l2 represent the lengths of robotic arm joint 1 and robotic arm joint 2, respectively, and g represents the acceleration due to gravity. Represents the friction force matrix, τ d (t) represents the unknown external disturbance; τ(t) represents the input torque of the robotic arm joint; y = q represents the angular position of the robotic arm joint.
[0064] make u=τ(t)M -1 (q), D(t)=τ d (t)M -1 (q), u is the actual control input, which transforms the dynamic model of the double-link robotic arm into a state-space equation, i.e., a dynamic model:
[0065]
[0066] Where x1, x2, x i Indicates an unmeasurable system state. These are the states after differentiating x1 and x2, respectively, and a s (t,x i ) represents the state variable x i The attack state at time t x represents i The system state after being attacked, i = 1, 2; F(x) represents the unknown system nonlinear term, D(t) represents the unknown external disturbance, u represents the measurable system input, and y represents the system output.
[0067] A radial basis function neural network is used to identify unknown, uncertain, and nonlinear dynamics in a two-link robotic arm, and the dynamic model is improved accordingly. The improved dynamic model is as follows:
[0068]
[0069] in, This represents a function obtained through approximation using a neural network. This represents the weight vector, with the superscript T indicating transpose. For the basis function vector, A vector representing a subset of state variables. yes The estimate; The error term representing the unknown system nonlinear term function and the neural network approximation function. y represents the approximation error, and y represents the system output.
[0070] Step 2. Based on the dynamic model, establish a neural network observer to observe the disturbance and attack state information of the dual-link robotic arm:
[0071]
[0072] Where K1 and K2 are the observer gains of the design, and the estimation error is defined as... Where the state errors are respectively These are the estimated values of x1, x2, and θ1, respectively. The superscript T indicates transpose. This is the system state after x1 is attacked.
[0073] Based on a given reference signal, errors are generated using state variables estimated by a neural network observer to construct a consistent tracking error:
[0074]
[0075] Where z1 and z2 are the tracking errors of the dual-link robotic arm, y r α1 represents a pre-set reference signal, and α2 represents a virtual control signal. Represents virtual control signals. It is a virtual control signal α1 * The estimated value.
[0076] Step 3. Based on reinforcement learning, the dynamic model, and the neural network observer, establish a control scheme for the dual-link robot arm, specifically including:
[0077] Based on the stability criteria of gradient descent and Lyapunov function stability theory, we design an adaptive update law for the neural network weights of the Actor-Critic (AC) algorithm in Reinforcement Learning (RL).
[0078] Based on a neural network observer, an adaptive law is designed for the unknown nonlinear term according to the adaptive update law, thus forming a control scheme; the control scheme includes a virtual controller and an actual controller.
[0079] Step 4. Based on the control scheme, under the backstepping framework, establish a virtual controller and an actual controller to perform adaptive safety control on the double-link robotic arm;
[0080] Building a virtual controller specifically includes:
[0081] The optimal virtual controller is obtained by establishing the following performance index function and minimizing the cost on the admissible control set Ψ(Ω1).
[0082]
[0083] Where Ω1 is a compact set containing the origin, Ψ(Ω1) is the admissible control set, and α1 is the virtual controller. It is the optimal virtual controller. It is a cost function; The performance index function is represented by the minimum value obtained by integrating all possible virtual controls over the time interval [t,∞) under a given state z1; z1(s) represents the tracking error z1 of the double-link robotic arm, which is a function of time s.
[0084] α1(z1) indicates that this virtual control is a function of the given state z1; This represents the virtual control that minimizes the performance index function under a given state z1.
[0085] Differentiating both sides of equation (6), we obtain the Hamilton Jacobi Bellman (HJB) equation as follows:
[0086]
[0087] Among them, a s Change to a s (t,x1) represents the attack state that the state variable x1 is subjected to at time t. and These represent performance index functions with the same meaning.
[0088] By solving The virtual controller is obtained as shown below:
[0089]
[0090] Will Decomposed into:
[0091]
[0092] in, It is in response to The function is decomposed into a continuous function.
[0093] Substituting formula (9) into formula (8) yields:
[0094]
[0095] because It is an unknown continuous entity, so a neural network is used to approximate it. It can be rewritten as:
[0096]
[0097] in, This represents the ideal weight, and the superscript T indicates transpose. This represents a basis function vector, while Indicates the approximation error. The table shows the ideal weights. The transpose of .
[0098] Substituting formula (11) into formulas (9) and (10) yields:
[0099]
[0100] in,
[0101] Due to unknown items The existence of the controller in the above formula Unavailable; To achieve the control objective, reinforcement learning was employed, specifically through two types of neural networks: an evaluation network (critic, C) and an execution network (actor, A). The evaluation network is used to assess control performance, while the execution network is responsible for executing control actions.
[0102] To obtain an optimized controller, a critic neural network (Critic NN) was designed to evaluate control performance, and an actor neural network (Actor NN) was designed to execute control actions.
[0103]
[0104] in, This indicates the evaluation of neural network weights. Represents a basis function vector.
[0105]
[0106] in, yes The estimated value, yes The estimated value, i.e. the final virtual controller, is obtained from the neural network. Approaching Unknown dynamics within, The weights of an ideal evaluation neural network; neural network Approaching Unknown dynamics within, The K1 table represents the ideal execution neural network weights, and the K1 table represents the designed observer gain.
[0107] Under the control technology, the parameter adaptive update law of the evaluation neural network and the execution neural network weights of the designed virtual controller is as follows:
[0108]
[0109] Where, γ c1 It is a parameter for evaluating the design of neural networks, γ a1 These are the design parameters for executing the neural network.
[0110] Building the actual controller specifically includes:
[0111] Construct a performance index function and obtain the actual controller u by minimizing the cost function value on the admissible control set Ψ(Ω2). * :
[0112]
[0113] Where u is the actual controller, u * It is the actual controller. It is the cost function, z2(s) represents the tracking error of the dual-link robotic arm, and z2 is a function of time s.
[0114] The corresponding Hamilton-Jacobi-Bellman (HJB) equation is as follows:
[0115]
[0116] in, For virtual controllers The state after differentiation; and The expressions have the same meaning: Γ2 represents the performance index function, which is obtained by integrating all possible virtual controls over the time interval [t, ∞) and taking the minimum value under a given state z2; Γ2 represents a state variable that satisfies... Where α1 represents the input of the differential tracker, Γ1 and Γ2 represent a state variable, and ι represents the condition that satisfies and Positive parameters of the condition, This represents a constant threshold used to determine the output range of the saturation function. This indicates that the saturation function satisfies when hour, when hour, Where, represents the input variable of the saturation function, and sign(b) represents the sign function; For virtual controllers The state obtained after differentiation.
[0117] Similar to the virtual controller, we obtain the actual controller:
[0118]
[0119] in, For neural network weights, neural network Used to approximate u * Unknown dynamics within, Let K2 be the basis function vector, K2 be the designed observer gain, and Γ2 be a state variable. This indicates the pre-set reference signal y. r The derivative of θ * 1 represents the weight vector, and the superscript T indicates transpose. is the basis function vector.
[0120] Based on the aforementioned control scheme, adaptive safety control of the dual-link robotic arm is implemented within the backstepping framework, and the following steps are also included:
[0121] Based on the gradient descent method and the stability determination method of Lyapunov function stability theory, an adaptive update law for the weights of the Actor-Critic neural network is designed:
[0122]
[0123] in, It is the adaptive update law for the weights of the designed neural network. These are the weights of the neural network, I m It is an m×m identity matrix, r c1 ,r c2 ,r a1 ,r a2 ,r θ1 σ1, σ1, ω1 are the design parameters. A vector of basis functions; Let σ be a Gaussian function, i = 1, 2. In the design of the virtual controller, i.e., in the first step, i = 1; in the design of the actual controller, i.e., in the second step, i = 2. i This represents a design parameter that is added as a correction term to the adaptive law to ensure that the weights of each state are adjusted in the same way, thereby maintaining the consistency and balance of the system. Representation of neural network weights The adaptive update law.
[0124] Neural network weights represent the connection strength between different neurons in the neural network. The control strategy is approximated by designing an adaptive update law for the neural network weights. The virtual controller provides a safe testing environment to verify the effectiveness of these strategies. Finally, the actual controller applies these strategies to the actual system to achieve accurate and stable control. In the framework of the dual-link robotic arm system, the adaptive update law for neural network weights, the virtual controller, and the actual controller together constitute an important part of achieving high-performance control tasks and approximate the control strategy through continuous optimization and iteration.
[0125] This embodiment uses Lyapunov functions to evaluate the stability of a dual-link robotic arm system. Within the framework of Lyapunov functions, the convergence of the system is determined by analyzing the time derivative of the function. If the time derivative of the function approaches zero, it means that the system has reached a stable state. Therefore, by designing corresponding controllers and neural network weights, the influence of external attacks and disturbances is compensated, thereby enabling the time derivative of the Lyapunov function to converge. In other words, the control technology in this scheme can be reflected through the design of controllers and neural network weights.
[0126] To demonstrate the feasibility, effectiveness, and correctness of this example, the present invention conducts the following simulation experiments:
[0127] In this simulation experiment, an adaptive controller based on neural network states was designed for a dual-link robotic arm system under external disturbances and attacks. This controller ensures that the operating state of each dual-link robotic arm remains consistent with the desired state, meaning that each variable of the dual-link robotic arm tends to be consistent. Furthermore, while achieving consistency, this not only improves the accuracy of identifying unknown nonlinearities in the system but also significantly reduces the number of controller updates and mechanical wear in the dual-link robotic arm system.
[0128] During the controller design process, the system model parameters are set as follows:
[0129] The system's reference signal is y r =0.5sin(0.4t); the initial system state is x1(0) = x2(0) = -0.1; the system disturbance is τ. d =0.1cos(50t); the attack on the system is a s (t,x1)=0.01(0.1+0.1cos(t))x1; the relevant parameters of the designed neural network observer are K1=10, K2=40, The parameters related to the virtual controller and the actual controller are designed as σ1=σ2=0.5, ω1=0.3; the adaptive update law parameter is γ. a1 =γ a2 =1.5,γ c1 =γ c2 =1.1,γ θ1 =1.3.
[0130] The effectiveness of the simulation in this embodiment is further illustrated by referring to the accompanying drawings:
[0131] Simulation results, such as Figure 3 This diagram presents a typical two-link robotic arm, where q1 and q2 represent the angular positions of joint 1 and joint 2, l1 and l2 represent the lengths of joint 1 and joint 2, m1 and m2 represent the masses of joint 1 and joint 2, and g represents the acceleration due to gravity. Figures 4-12 , Figure 4 The output and desired trajectory tracking diagram of the dual-link robotic arm are presented. The dual-link robotic arm achieves accurate tracking of the reference signal under external disturbances and attacks. Figure 5 and 7 This demonstrates the observation performance of a neural network observer on the disturbances and states of a two-bar linkage robotic arm system that does not handle external attacks. Figure 6 and 8This demonstrates the observation performance of the neural network observer on external attacks and disturbances, as well as the state of the two-bar linkage robotic arm system; from Figure 5 , Figure 6 , Figure 7 as well as Figure 8 It can be seen from the data that the designed neural network observer has better observation results when dealing with attacks than when not dealing with attacks; Figure 9 The virtual control input curve of the dual-link robotic arm system is shown. Figure 10 The control input curves of the dual-link robotic arm system are shown. Figure 11 and Figure 12 The weight convergence curves of the execution neural network and the evaluation neural network are presented respectively. It can be seen from the figure that each weight in the execution neural network and the evaluation neural network eventually tends to stabilize and converge. Therefore, the designed reinforcement learning neural network control scheme can effectively achieve accurate approximation of the designed controller and has a good convergence effect.
[0132] In summary, all signals in the dual-link robotic arm system are consistent and eventually bounded, and simulation results demonstrate the effectiveness of the proposed control scheme.
[0133] This embodiment uses backstepping recursion and control as its design framework to address the tracking control problem of a dual-link manipulator system with unknown attacks and disturbances. Furthermore, this invention assumes that the dual-link manipulator system has an unmeasurable state and is affected by unknown attack and disturbance information, making the system more general. A neural network observer is designed, which can estimate not only the attacks and disturbances present in the system but also the system's state information, thereby achieving comprehensive observation of the system state, attacks, and disturbances. Subsequently, a reinforcement learning method with an Actor-Critic framework is used to construct a controller based on the neural network observer. Adaptive laws are used to adjust the weights of the neural network, enabling the designed approximate controller to closely approximate the actual controller. Finally, simulations verify that the proposed adaptive control strategy ensures that all signals are bounded. The widespread application of this invention in dual-link manipulator control is one of the important future research directions.
[0134] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0135] In the description of this invention, it should be understood that the terms "longitudinal", "lateral", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this invention, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0136] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the present invention.
Claims
1. A control method for a dual-link robotic arm based on a neural network observer, characterized in that, Specifically, the following steps are included: Step 1. Establish a system dynamics model of the dual-link robotic arm under external attacks and disturbances; Step 2. Based on the system dynamics model, establish a neural network observer; Step 3. Based on the system dynamics model and the neural network observer obtained through reinforcement learning, establish a control scheme for the dual-link robotic arm; Step 4. Based on the control scheme, adaptive safety control is performed on the dual-link robotic arm within the anti-stepping framework; The system dynamics model of the dual-link robotic arm under external attacks and disturbances includes: use A system dynamics model of a dual-link robotic arm under external attacks and disturbances; Where q = (q1, q2) T , These represent the angular position, angular velocity, and angular acceleration of a joint in a two-bar linkage robotic arm, which includes joint 1 and joint 2. q1 represents the angular position of joint 1, and q2 represents the angular position of joint 2. This represents the angular velocity of joint 1 of the robotic arm. This represents the angular velocity of joint 2 of the robotic arm. This represents the angular acceleration of joint 1 of the robotic arm. M(q) represents the angular acceleration of joint 2 of the robotic arm; M(q) represents the inertia matrix of the robotic arm. Represents the centrifugal Lippe force matrix; G(q)=(g 11 ,g 21 ) T Represents the gravity vector, where g 11 =(m1+m2)gl1cos(q1)+m2gl2cos(q1+q2),g 21 =m2gl2cos(q1+q2), where m1 and m2 represent the masses of robotic arm joint 1 and robotic arm joint 2, respectively, l1 and l2 represent the lengths of robotic arm joint 1 and robotic arm joint 2, respectively, and g represents the acceleration due to gravity. Represents the friction force matrix, τ d (t) represents the unknown external disturbance; τ(t) represents the input torque of the robotic arm joint; y = q represents the angular position of the robotic arm joint; Step 1 is followed by transforming the system dynamics model of the dual-link robotic arm into a nonlinear system dynamic model with attack capabilities. The nonlinear system dynamic model is as follows: Where x1=q, x i Represents the unmeasurable system state variables, i = 1, 2; u = τ(t)M -1 (q), where u represents the measurable system input, M -1 (q) represents the inverse of the robotic arm's inertia matrix; y represents the system output, and x1 represents the system output state; Represents an unknown nonlinear function. The derivatives of x1 and x2 are respectively, and D(t) = τ d (t)M -1 (q) represents an unknown external disturbance; a s (t,x i ) represents the state variable x i The attack state at time t It is x i The state after being attacked; A system dynamics model of a dual-link robotic arm under external attacks and disturbances is established, which then includes: A radial basis function neural network is used to identify unknown, uncertain, and nonlinear dynamics in a two-link robotic arm, and the dynamic model is improved to determine the improved dynamic model; the improved dynamic model is as follows: in, This represents a function obtained through approximation using a neural network. Let T represent the weight vector, and T denote the transpose. Represents a basis function vector. Indicates the approximation error; A vector representing a subset of state variables. express The estimate; Let y represent the error term between the unknown system nonlinear function and the neural network approximation function, and let y represent the system output.
2. The control method for a dual-link robotic arm based on a neural network observer according to claim 1, characterized in that, Based on the dynamic model, a neural network observer is established to observe the disturbance and attack state information of the dual-link robotic arm. This is followed by: Based on the improved dynamic model, a neural network observer is established to observe the disturbance and attack state information of the dual-link robotic arm; the neural network observer is as follows: Among them, K i The values of i = 1 and 2 represent the observer gains of the design. When i = 1, K1 is the observer gain of robotic arm joint 1, and when i = 2, K2 is the observer gain of robotic arm joint 2. The estimation error is defined as... The state errors are respectively represented as follows: x = (x1, x2) T as well as These are the state variables x1 and x2, respectively. The estimated value, The tables are respectively The state after differentiation, Represents the weight vector The estimate, is the basis function vector.
3. The method for controlling a dual-link robotic arm based on a neural network observer according to claim 2, characterized in that, Based on reinforcement learning, the aforementioned dynamic model, and the aforementioned neural network observer, a control scheme for a dual-link robot arm is established, specifically including: Based on the gradient descent method and the stability determination method of Lyapunov function stability theory, an adaptive update law for the weights of the Actor-Critic neural network is designed. Based on a neural network observer, an adaptive law is designed for the unknown nonlinear term according to the adaptive update law, thus forming a control scheme; the control scheme includes a virtual controller and an actual controller.
4. The control method for a dual-link robotic arm based on a neural network observer according to claim 3, characterized in that, Based on a neural network observer, an adaptive law is designed for the unknown nonlinear term according to the aforementioned adaptive update law, forming a control scheme, specifically including: use Construct the virtual controller; use Construct the actual controller; in, Represents the weights of the neural network. Representation of basis function vectors; Neural networks Used to approximate u * In the unknown dynamics, K1 and K2 represent the designed observer gains, and Γ2 represents a state variable. This indicates the pre-set reference signal y r The derivative of θ * 1 represents the weight vector. Represents a basis function vector.
5. The control method for a dual-link robotic arm based on a neural network observer according to claim 4, characterized in that, Based on the backstepping method and the neural network observer approximation design, an adaptive update law for the uncertain parameters is designed. According to the gradient descent method and the stability determination method of Lyapunov function stability theory, an adaptive update law for the weights of the Actor-Critic neural network is designed. The adaptive update law is as follows: Where, r c1 ,r c2 ,r a1 ,r a2 ,r θ1 ,σ1,σ1,ω1 represent the design parameters, I m Describes an m×m identity matrix. Represents the weights of the neural network. This represents the adaptive update law for the weights of the designed neural network. For the basis function vector, Let i be a Gaussian function, i = 1, 2.
6. A control system for a dual-link robotic arm based on a neural network observer, characterized in that, The dual-link robotic arm control system based on a neural network observer employs the dual-link robotic arm control method based on a neural network observer as described in any one of claims 1-5. The dual-link robotic arm control system based on a neural network observer specifically includes: The system dynamics model building module is used to build a system dynamics model of a dual-link robotic arm under external attacks and disturbances. The neural network observer building module is used to build neural network observers based on dynamic models. The control scheme establishment module is used to establish a control scheme for the dual-link robotic arm based on reinforcement learning, the dynamic model, and the neural network observer. An adaptive control module is used to perform adaptive safety control of the dual-link robotic arm based on the control scheme and within the anti-stepping framework.
7. A dual-link robotic arm control system based on a neural network observer according to claim 6, characterized in that, The control scheme establishment module specifically includes: The adaptive update law design unit is used to design the adaptive update law of the weights of the Actor-Critic neural network based on the stability determination method of gradient descent and Lyapunov function stability theory. The control scheme establishment unit is used to design an adaptive law for unknown nonlinear terms based on the neural network observer and the adaptive update law, thereby forming a control scheme; the control scheme includes a virtual controller and an actual controller.
Citation Information
Patent Citations
Self-adaptive neural network control method for hydraulic mechanical arm
CN117289612A
Fixed time optimal control method and system for single-connecting-rod mechanical arm
CN119501947A
Cited By
Robot anti-interference control method based on neural network interference observer
CN121979110A