Double-connecting-rod mechanical arm control method and system based on neural network observer

Through the dual-link robot arm control method based on neural network observers, the problem of the robot arm system being affected by disturbances and attacks in complex environments is solved, efficient adaptive safety control and accurate state observation are achieved, and the stability and tracking accuracy of the system are improved.

CN120347734AActive Publication Date: 2025-07-22BOHAI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510445207.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-22
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The robotic arm systems are susceptible to external perturbations and attacks in complex environments, resulting in reduced control performance, especially in positioning applications, and the prior art is difficult to effectively deal with these challenges.

Method used

Establish a dual-link robotic arm control method based on neural network observers. By establishing a system dynamics model, neural network observer and reinforcement learning, design an adaptive security control strategy, and use the Actor-Critic neural network adaptive update rate to approximate the actual controller to achieve accurate observation and control of perturbations and attacks.

Benefits of technology

It improves the control performance and stability of the robotic arm system in complex environments, achieves fast and accurate observation of attacks and disturbances caused by unknown factors, and reduces system wear and controller updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120347734A_ABST
    Figure CN120347734A_ABST
Patent Text Reader

Abstract

The invention discloses a double-connecting-rod mechanical arm control method and system based on a neural network observer. The double-connecting-rod mechanical arm control method comprises the following steps that 1, a double-connecting-rod mechanical arm dynamic model under external attack and disturbance is established; step 2, establishing a neural network observer based on the dynamic model; 3, based on reinforcement learning, a kinetic model and a neural network observer, a control scheme of the double-connecting-rod mechanical arm is established; 4, based on the control scheme, under a backstepping method framework, self-adaptive safety control is conducted on the double-connecting-rod mechanical arm; according to the method, attacks and disturbances caused by unknown factors and rapidly changing attacks and disturbances are observed; through the performance index function, the controller achieves control over the double-connecting-rod mechanical arm; according to the technical scheme of the invention, the adaptive update rate of the Actor-Critic neural network is designed by using the gradient descent method through the learning ability of the reinforcement learning algorithm, so that the approximation controller can be close to an actual controller.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robotic arm control, and particularly to a control method and system for a dual-link robotic arm based on a neural network observer. Background Art

[0002] In recent years, robotic arm systems have unique advantages in improving production efficiency, improving working conditions, and reducing labor intensity, and their applications are becoming more and more widespread; especially in some complex operating environments, some repetitive or dangerous tasks are performed by robotic arms, which can reduce labor costs, improve productivity and the safety of production manufacturing, and are more flexible.

[0003] Actual robotic arm systems often face challenges of attacks and disturbances caused by unknown factors such as complex environments like turbulence, wind changes, mechanical resistance, external unknown states, and internal dynamic uncertainties; at present, to address these challenges, various compensation strategies have emerged, and the control schemes of robotic arms usually use neural networks to act on the controlled system through controllers.

[0004] It should be noted that although robotic arm systems have demonstrated their powerful potential in many fields, the robotic arm systems themselves have high complexity, and their algorithms and implementation processes are relatively cumbersome, and they are easily affected by external disturbances in actual industrial production. Therefore, a neural network observer is used to observe the disturbance state information of the dual-link robotic arm, which solves the problem that the true state is unavailable due to external disturbances.

[0005] In addition, even though robotic arm systems have demonstrated their powerful potential in many fields, when sending data to controllers or actuators, they may encounter external attacks, which can damage the authenticity of system data and make it unusable for control design; this has become a relatively difficult type of attack to handle; therefore, for robotic arm systems under attack, a neural network observer is designed to estimate the unavailable state.

[0006] Moreover, by combining control technology with a neural network observer, a new control strategy and optimization path for robotic arms are generated, that is, a performance index function is constructed to handle the optimization problem under attack and disturbance states, to further improve the control performance and stability of the robotic arm system and enable the system to obtain good control performance.

[0007] In addition, when applying robotic arms in the engineering field, position tracking is also a basic problem; neural network control methods have been widely used to solve the tracking problems of traditional rigid robotic arms. Due to the limitation of information transmission speed, there are often attack and disturbance phenomena. If disturbances and attacks are ignored in controller design, disturbances and attacks may lead to deterioration of system performance, especially in positioning applications, which will reduce the tracking accuracy of actuators. Summary of the Invention

[0008] The objective of the present invention is to provide a control method and system for a two-link manipulator based on a neural network observer, so as to solve the problems existing in the above-mentioned prior art.

[0009] A control method and system for a two-link manipulator based on a neural network observer specifically include the following steps: Step 1. Establish a system dynamics model of the two-link manipulator under external attacks and disturbances; Step 2. Based on the system dynamics model, establish a neural network observer; Step 3. Based on reinforcement learning, the system dynamics model, and the neural network observer, establish a control scheme for the two-link manipulator; Step 4. Based on the control scheme, perform adaptive safety control on the two-link manipulator under the backstepping framework.

[0010] Preferably, establishing a system dynamics model of the two-link manipulator under external attacks and disturbances specifically includes:

[0011] Using To establish a system dynamics model of the two-link manipulator under external attacks and disturbances;

[0012] where q = (q1, q2) T , respectively represent the angular position, angular velocity, and angular acceleration of the joints of the two-link manipulator. The manipulator joints include manipulator joint 1 and manipulator joint 2. q1 represents the angular position of manipulator joint 1, q2 represents the angular position of manipulator joint 2, represents the angular velocity of manipulator joint 1, represents the angular velocity of manipulator joint 2, represents the angular acceleration of manipulator joint 1, represents the angular acceleration of manipulator joint 2; M(q) represents the manipulator inertia matrix; represents the centrifugal Coriolis force matrix; G(q) = (g 11 , g 21 ) T represents the gravity vector, where g 11 = (m1 + m2)gl1cos(q1) + m2gl2cos(q1 + q2), g 21 = m2gl2cos(q1 + q2), m1 and m2 respectively represent the masses of manipulator joint 1 and manipulator joint 2, l1 and l2 respectively represent the lengths of manipulator joint 1 and manipulator joint 2, and g represents the gravitational acceleration; represents the friction matrix, τ d (t) represents the unknown external disturbance; τ(t) represents the input torque of the manipulator joint; y = q represents the angular position of the manipulator joint.

[0013] Preferably, after step 1, it further includes converting the system dynamics model of the double-link manipulator into a nonlinear system dynamic model with attacks. The nonlinear system dynamic model is:

[0014]

[0015] where \(x_1 = q\), x i represents the unmeasurable system state variable, \(i = 1, 2\); \(u=\tau(t)M\) -1 (q), \(u\) represents the measurable system input, \(M\) -1 (q) represents the inverse matrix of the manipulator inertia matrix; \(y\) represents the system output, \(x_1\) represents the system output state; represents the unknown nonlinear function, are the derivatives of \(x_1\) and \(x_2\) respectively, \(D(t)=\tau\) d (t)M -1 (q) represents the unknown external disturbance; \(a\) s (t, x i ) represents the attack state of the state variable \(x\) i at time \(t\), is the state of \(x\) i after being attacked.

[0016] Preferably, after establishing the system dynamics model of the double-link manipulator under external attacks and disturbances, it further includes:

[0017] Using a radial basis neural network to identify the unknown and uncertain nonlinear dynamics in the double-link manipulator, improving the dynamic model, and determining the improved dynamic model; the improved dynamic model is:

[0018]

[0019] where, represents the function approximated by the neural network, represents the weight vector, \(T\) represents the transpose, represents the basis function vector, represents the approximation error; represents the subset vector of the state variables, represents the estimate of; represents the error term between the unknown system nonlinear term function and the neural network approximation function, \(y\) represents the system output.

[0020] Preferably, based on the dynamics model, a neural network observer is established to observe the disturbance and attack state information of the double-link manipulator, and then it further includes:

[0021] Based on the improved dynamic model, a neural network observer is established, and the neural network observer is used to observe the disturbance and attack state information of the double-link manipulator; the neural network observer is as follows:

[0022]

[0023] where K i , i = 1, 2 represent the designed observer gains. When i = 1, K1 is the observer gain of the manipulator joint 1, and when i = 2, K2 is the observer gain of the manipulator joint 2; the estimation error is defined as where the state errors are respectively expressed as x = (x1, x2) T and are the estimated values of the state variables x1, x2, respectively, are respectively expressed as the states after derivation, represents the weight vector of the estimation, is the basis function vector.

[0024] Preferably, based on reinforcement learning, the dynamic model, and the neural network observer, a control scheme for the double-link robot manipulator is established, specifically including:

[0025] According to the gradient descent method and the stability judgment method of the Lyapunov function theory, design the adaptive update law of the Actor–Critic neural network weights;

[0026] Based on the neural network observer, according to the adaptive update law, design the adaptive law for the unknown nonlinear terms to form a control scheme; the control scheme includes a virtual controller and an actual controller.

[0027] Preferably, based on the neural network observer, according to the adaptive update law, design the adaptive law for the unknown nonlinear terms to form a control scheme, specifically including:

[0028] Use to construct the virtual controller;

[0029] Use to construct the actual controller;

[0030] where represents the neural network weights, represents the basis function vector; the neural network is used to approximate u *Unknown dynamics in it, K1 and K2 represent the designed observer gains, Γ2 represents a state variable, represents the pre-set reference signal y r derivative of, θ * 1 represents the weight vector, represents the basis function vector.

[0031] Preferably, based on the control scheme, in the backstepping framework, an adaptive update rate for uncertain parameters approximated and designed based on backstepping and a neural network observer is used. According to the stability judgment method of the gradient descent method and the Lyapunov function stability theory, an adaptive update law for the weights of the Actor–Critic neural network is designed. The adaptive update rate is:

[0032]

[0033] where r c1 , r c2 , r a1 , r a2 , r θ1 , σ1, σ1, ω1 represent the designed parameters, I m represents the m×m order identity matrix, represents the neural network weights, represents the adaptive update rate of the designed neural network weights, is the basis function vector, is the Gaussian function, i = 1, 2.

[0034] Preferably, the double-link manipulator control system based on a neural network observer adopts a double-link manipulator control method according to any one of the above claims 1-8. The double-link manipulator control system based on a neural network observer specifically includes: a system dynamics model establishment module for establishing the system dynamics model of the double-link manipulator under external attacks and disturbances; a neural network observer establishment module for establishing a neural network observer based on the dynamics model; a control scheme establishment module for establishing a control scheme for the double-link manipulator based on reinforcement learning, the dynamics model, and the neural network observer; and an adaptive control module for adaptively and securely controlling the double-link manipulator based on the control scheme in the backstepping framework.

[0035] Preferably, the control scheme establishment module specifically includes: an adaptive update law design unit for designing an adaptive update law for the weights of the Actor–Critic neural network according to the gradient descent method and the stability judgment method of the Lyapunov function stability theory; a control scheme establishment unit for designing an adaptive law for unknown non-linear terms based on a neural network observer according to the adaptive update law to form a control scheme; the control scheme includes a virtual controller and an actual controller.

[0036] Compared with the prior art, the present invention provides a control method and system for a double-link manipulator based on a neural network observer, having the following beneficial effects:

[0037] 1. The present invention takes into account that the actual manipulator system often faces challenges of attacks and disturbances caused by unknown factors such as complex environments such as turbulence, wind force changes, mechanical resistance, external unknown states, and internal dynamic uncertainties, which affect the control performance of the system; the present invention designs a neural network observer for the manipulator system, and due to considering the information of attacks and disturbances, it has better estimation accuracy for attacks and disturbances that change rapidly with time, so as to realize the observation of the state, attacks, and external disturbances of the double-link manipulator system.

[0038] 2. The technical solution of the present invention proposes an adaptive backstepping control strategy; by designing a performance index function, the designed controller realizes the control of the double-link manipulator.

[0039] 3. The technical solution of the present invention designs an adaptive update rate of the Actor–Critic neural network by using the learning ability of the reinforcement learning algorithm and the gradient descent method, so that the designed approximate controller can approximate the actual controller on the premise of system stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0041] Figure 1 It is a schematic diagram of the control method flow of the present invention;

[0042] Figure 2 It is a schematic diagram of the components of the manipulator system of the present invention;

[0043] Figure 3 It is a schematic diagram of the structure of the double-link manipulator of the present invention;

[0044] Figure 4Output and desired trajectory tracking diagram of the double-link manipulator of the present invention;

[0045] Figure 5 Schematic diagram of the observation curve of the x1 state of the double-link manipulator based on the neural network observer without dealing with external attack states of the present invention;

[0046] Figure 6 Observation curve diagram of the x1 state of the double-link manipulator based on the neural network observer dealing with external attack states of the present invention;

[0047] Figure 7 Observation curve diagram of the x2 state of the double-link manipulator based on the neural network observer without dealing with external attack states of the present invention;

[0048] Figure 8 Observation curve diagram of the x2 state of the double-link manipulator based on the neural network observer dealing with external attack states of the present invention;

[0049] Figure 9 Virtual control input curve diagram of the double-link manipulator system based on the neural network of the present invention;

[0050] Figure 10 Actual control input curve diagram of the double-link manipulator system based on the neural network of the present invention;

[0051] Figure 11 Weight convergence curve diagram of the neural network executed by the present invention;

[0052] Figure 12 Weight convergence curve diagram for evaluating the neural network of the present invention. Detailed implementation manners

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0055] As Figure 1 shown, a control method for a double-link manipulator based on a neural network observer specifically includes the following steps:

[0056] Step 1. Establish a system dynamics model of the double-link manipulator under external attacks and disturbances;

[0057] Step 2. Based on the dynamic model, establish a neural network observer for observing the disturbance and attack state information of the double-link manipulator;

[0058] Step 3. Based on reinforcement learning, the dynamic model, and the neural network observer, establish a control scheme for the double-link manipulator;

[0059] Step 4. Based on the control scheme, in the framework of the backstepping method, establish a virtual controller and an actual controller for adaptive safety control of the double-link manipulator;

[0060] As an embodiment, combining the neural network observer with the control solves the problem that external attacks and disturbances affect the control performance and stability of the manipulator system, thereby optimizing the path, designing a new control strategy, further improving the control performance and stability of the manipulator system, and adopting Figure 1 the method shown in Figure 3 to perform backstepping control on the double-link manipulator shown in

[0061] Step 1. Establish the system dynamic model of the double-link manipulator under external attacks and disturbances:

[0062]

[0063] where q = (q1, q2) T , represent the angular position, angular velocity, and angular acceleration of the joints of the double-link manipulator respectively. The manipulator joints include manipulator joint 1 and manipulator joint 2. q1 represents the angular position of manipulator joint 1, q2 represents the angular position of manipulator joint 2, represents the angular velocity of manipulator joint 1, represents the angular velocity of manipulator joint 2, represents the angular acceleration of manipulator joint 1, represents the angular acceleration of manipulator joint 2; M(q) represents the manipulator inertia matrix; represents the centrifugal Coriolis force matrix; G(q) = (g 11 , g 21 ) T represents the gravity vector, where g 11 = (m1 + m2)gl1cos(q1) + m2gl2cos(q1 + q2), g 21 = m2gl2cos(q1 + q2), m1 and m2 represent the masses of manipulator joint 1 and manipulator joint 2 respectively, l1 and l2 represent the lengths of manipulator joint 1 and manipulator joint 2 respectively, and g represents the acceleration due to gravity; represents the friction matrix, τ d(t) represents an unknown external disturbance; τ(t) represents the input torque of the robotic arm joints; y = q represents the angular position of the robotic arm joints.

[0064] Let x1 = q, u = τ(t)M -1 (q), D(t) = τ d (t)M -1 (q), u is the actual control input. Transform the dynamic model of the two-link robotic arm into a state-space equation, i.e., the dynamic model:

[0065]

[0066] where, x1, x2, x i represent the unmeasurable system states, are the states after differentiating x1 and x2 respectively, a s (t, x i ) represents the attacked state that the state variable x i receives at time t, represents x i the system state after being attacked, i = 1, 2; F(x) represents the unknown system nonlinear term, D(t) represents the unknown external disturbance, u represents the measurable system input, and y represents the system output.

[0067] Utilize a radial basis neural network to identify the unknown uncertain nonlinear dynamics in the two-link robotic arm and improve the dynamic model. The improved dynamic model is:

[0068]

[0069] where, represents the function approximated by the neural network, represents the weight vector, the superscript T represents the transpose, is the basis function vector, represents the subset vector of the state variables, is the estimate of; represents the error term between the unknown system nonlinear term function and the neural network approximation function, is the approximation error, and y represents the system output.

[0070] Step 2. Based on the dynamic model, establish a neural network observer for observing the disturbance and attack state information of the two-link robotic arm:

[0071]

[0072] where K1 and K2 are the designed observer gains, and the estimation error is defined as where the state errors are respectively x = (x1, x2) T , are the estimated values of x1, x2, and θ1 respectively, and the superscript T represents the transpose, is the system state of x1 after being attacked.

[0073] Based on the state variables estimated by the neural network observer with respect to the given reference signal, an error is generated, and a consensus tracking error is constructed:

[0074]

[0075] where z1 and z2 are the tracking errors of the double-link manipulator, and y r represents the preset reference signal, and α1 represents the virtual control signal, represents the virtual control signal, is the estimated value of the virtual control signal α1 * .

[0076] Step 3. Based on reinforcement learning, the dynamic model, and the neural network observer, establish a control scheme for the double-link robot manipulator, specifically including:

[0077] According to the gradient descent method and the stability judgment method of the Lyapunov function theory, design the adaptive update law of the weights of the Actor–Critic (AC) neural network in reinforcement learning (RL);

[0078] Based on the neural network observer, design an adaptive law for the unknown nonlinear terms according to the adaptive update law to form a control scheme; the control scheme includes a virtual controller and an actual controller.

[0079] Step 4. Based on the control scheme, under the backstepping framework, establish a virtual controller and an actual controller to perform adaptive safety control on the double-link manipulator;

[0080] Construct a virtual controller, specifically including:

[0081] Establish the following performance index function, and obtain the virtual controller by minimizing the cost value on the admissible control set Ψ(Ω1)

[0082]

[0083] where Ω1 is a compact set containing the origin, Ψ(Ω1) is the admissible control set, and α1 is the virtual controller, is a virtual controller is the cost function; represents the performance index function, which is obtained by integrating all possible virtual controls over the time interval [t, ∞) and taking the minimum value given the state z1; z1(s) represents the tracking error z1 of the two-link manipulator, which is a function of time s; α1(z1) indicates that this virtual control is a function of the given state z1; represents the virtual control that can minimize the performance index function given the state z1.

[0084] Taking the derivative of both sides of Equation (6), the Hamilton Jacobi Bellman (HJB) equation is obtained as follows:

[0085]

[0086] where, changing a s to a s (t, x1) represents the attack state that the state variable x1 is subjected to at time t, and have the same meaning and represent the performance index function.

[0087] By solving the virtual controller is obtained as follows:

[0088]

[0089] Decompose into:

[0090]

[0091] where, is a continuous function generated by decomposing the function.

[0092] Substituting Equation (9) into Equation (8) gives:

[0093]

[0094] Since is unknown and continuous, a neural network is used to approximate it, which can be rewritten as:

[0095]

[0096] where, represents the ideal weight, and the superscript T represents the transpose, denotes the basis function vector, while denotes the approximation error, Table is the ideal weight transpose.

[0097] Substituting formula (11) into formulas (9) and (10) gives:

[0098]

[0099] where,

[0100] Due to the existence of the unknown term in the above formula, the controller cannot be used; to achieve the control goal, reinforcement learning is adopted, specifically implemented through two neural networks, the evaluation network (critic, C) and the execution network (actor, A). The evaluation network is used to evaluate the control performance, while the execution network is responsible for executing the control behavior.

[0101] To obtain an optimized controller, an evaluation neural network for evaluating control performance (critic neural network, Critic NN) and an execution neural network for executing control behavior (actor neural network, Actor NN) are designed:

[0102]

[0103] where, denotes the weights of the critic neural network, denotes the basis function vector.

[0104]

[0105] where, is the estimated value of, is the estimated value of, that is, the finally obtained virtual controller. The neural network approximates the unknown dynamics in, denotes the ideal weights of the critic neural network; the neural network approximates the unknown dynamics in, denotes the ideal weights of the actor neural network, and K1 represents the designed observer gain.

[0106] Under the control technology, the parameter adaptive update rates of the weights of the evaluation neural network and the execution neural network of the designed virtual controller are:

[0107]

[0108] Among them, γ c1 is a design parameter for evaluating the neural network, and γ a1 is a design parameter for implementing the neural network.

[0109] Construct an actual controller, specifically including:

[0110] Construct a performance index function, and obtain the actual controller u by minimizing the cost function value on the admissible control set Ψ(Ω2): * :

[0111]

[0112] Among them, u is the actual controller, and u * is the actual controller, is the cost function, and z2(s) represents the tracking error of the double-link manipulator. z2 is a function of time s.

[0113] The corresponding Hamilton-Jacobi-Bellman (HJB) equation is as follows:

[0114]

[0115] Among them, is the state after differentiating the virtual controller ; is the same as in meaning, representing the performance index function, which represents the function obtained by integrating all possible virtual controls over the time interval [t, ∞) and taking the minimum value under the given state z2; Γ2 represents a state variable that satisfies Among them, α1 represents the input of the differential tracker, Γ1 and Γ2 represent a state variable, and ι represents a positive parameter that satisfies and conditions, represents a constant threshold used to determine the output range of the saturation function, indicates that the saturation function satisfies when then When then Among them, represents the input variable of the saturation function, and sign(b) represents the sign function; is the state obtained by differentiating the virtual controller after differentiation.

[0116] Similar to the virtual controller, obtain the actual controller:

[0117]

[0118] Among them, is the neural network weight, and the neural network is used to approximate the unknown dynamics in * u, is the basis function vector, K2 is the designed observer gain, Γ2 represents a state variable, represents the preset reference signal y r derivative, θ * 1 represents the weight vector, and the superscript T represents the transpose, is the basis function vector.

[0119] Based on the above control scheme, under the backstepping framework, adaptive safety control is performed on the double-link manipulator, and it also includes:

[0120] According to the stability judgment method of the gradient descent method and the Lyapunov function stability theory, design the adaptive update rate of the Actor–Critic neural network weight:

[0121]

[0122] Among them, is the designed adaptive update rate of the neural network weight, is the neural network weight, I m is the m×m order identity matrix, r c1 , r c2 , r a1 , r a2 , r θ1 , σ1, σ1, ω1 are designed parameters, is the basis function vector; is the Gaussian function, i = 1, 2. In the process of designing the virtual controller, i.e., i = 1 in the first step, and in the process of designing the actual controller, i.e., i = 2 in the second step; σ i represents a designed parameter, which is added to the adaptive law as a correction term to ensure that the weights of each state are adjusted equally, thereby maintaining the consistency and balance of the system; represents the neural network weight adaptive update rate.

[0123] The neural network weight represents the connection strength between different neurons in the neural network. By designing the adaptive update rate of the neural network weight, the control strategy is approximated; the virtual controller provides a safe test environment to verify the effectiveness of these strategies; finally, the actual controller applies these strategies to the actual system to achieve precise and stable control; in the framework of the double-link manipulator system, the adaptive update rate of the neural network weight, the virtual controller, and the actual controller together constitute an important part of achieving high-performance control tasks, and approximate the control strategy through continuous optimization and iteration.

[0124] In this embodiment, the Lyapunov function is used to evaluate the stability of the double-link manipulator system. Under the framework of the Lyapunov function, the convergence of the system is judged by analyzing the time derivative of the function. If the time derivative of the function tends to zero, it means that the system has reached a stable state. Therefore, by designing the corresponding controller and neural network weights, the influence of external attacks and disturbances can be compensated, so that the time derivative of the Lyapunov function converges. That is to say, in this solution, the control technology can be reflected in the form of the design of the controller and neural network weights.

[0125] To prove the feasibility, effectiveness, and correctness of this example, the present invention conducts the following simulation experiments:

[0126] In this simulation experiment, for the double-link manipulator system under external disturbances and attacks, an adaptive controller based on the neural network state is designed to make the operating state of each double-link manipulator consistent with the desired double-link manipulator, that is, each variable of the double-link manipulator tends to be consistent. In addition, while achieving consistency, not only the identification accuracy of the unknown nonlinearity in the system is improved, but also the update times of the controller of the double-link manipulator system and mechanical wear are greatly reduced.

[0127] During the controller design process, the system model parameters are set as follows:

[0128] The reference signal of the system is y r = 0.5sin(0.4t); the initial values of the system states are x1(0) = x2(0) = -0.1; the system disturbance is τ d = 0.1cos(50t); the attack on the system is a s (t, x1) = 0.01(0.1 + 0.1cos(t))x1; the relevant parameters of the designed neural network observer are K1 = 10, K2 = 40, The parameters related to the virtual controller and the actual controller are designed as σ1 = σ2 = 0.5, ω1 = 0.3; the adaptive update rate parameter is γ a1 = γ a2 = 1.5, γ c1 = γ c2 = 1.1, γ θ1 = 1.3.

[0129] Combined with the attached drawings, the effectiveness of the simulation of this embodiment is further described:

[0130] The simulation results are as Figure 3Presents a schematic diagram of a typical two-link manipulator. q1 and q2 represent the angular positions of joint 1 and joint 2 of the manipulator, l1 and l2 represent the lengths of joint 1 and joint 2 of the manipulator, m1 and m2 represent the masses of joint 1 and joint 2 of the manipulator, and g represents the acceleration due to gravity; as Figures 4 - 12 , Figure 4 Presents the output of the two-link manipulator and the desired trajectory tracking diagram. The double-link manipulator achieves precise tracking of the reference signal under external disturbances and attacks; Figure 5 and 7 Shows the disturbance and state observation performance of the two-link manipulator system without the neural network observer dealing with external attacks; Figure 6 and 8 Shows the observation performance of the neural network observer for external attacks, disturbances, and states of the two-link manipulator system; From Figure 5 , Figure 6 , Figure 7 and Figure 8 It can be seen that the designed neural network observer has better observation effects when dealing with attacks than not dealing with attacks; Figure 9 Shows the virtual control input curve diagram of the two-link manipulator system; Figure 10 Shows the control input curve diagram of the two-link manipulator system; Figure 11 and Figure 12 Respectively present the weight convergence curve diagrams of the execution neural network and the evaluation neural network. It can be seen from the figure that each weight in the execution neural network and the evaluation neural network finally tends to be stable and converges. Therefore, the designed reinforcement learning neural network control scheme can effectively achieve precise approximation of the designed controller and has good convergence effects.

[0131] In summary, all signals in the two-link manipulator system are uniformly ultimately bounded, and the simulation results prove the effectiveness of the proposed control scheme.

[0132] This embodiment takes backstepping recursion and control as the design framework, aiming at the tracking control problem of a two-link manipulator system with unknown attacks and disturbances. In addition, the present invention assumes that the two-link manipulator system has unmeasurable states and is affected by unknown attack and disturbance information, making the system more general. A neural network observer is designed, which can not only estimate the attacks and disturbances existing in the system, but also estimate the state information of the system, so as to achieve a comprehensive observation of the system state, attacks and disturbances. Subsequently, a reinforcement learning method with an Actor-Critic framework is adopted to construct a controller based on the neural network observer. And the weights of the neural network are adjusted through an adaptive law, so that the designed approximate controller can approximate the actual controller. Finally, simulations verify that the proposed adaptive control strategy can ensure that all signals are bounded. The popularization and application of the present invention in two-link manipulator control is one of the important research directions in the future.

[0133] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0134] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.

[0135] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A control method for a double-link manipulator based on a neural network observer, characterized in that, Specifically, it includes the following steps: Step 1. Establish the system dynamics model of a two-link manipulator under external attacks and disturbances; Step 2. Based on the system dynamics model, establish a neural network observer; Step 3. Based on the system dynamics model and the neural network observer with reinforcement learning, establish a control scheme for the two-link manipulator; Step 4. Based on the control scheme, under the backstepping framework, perform adaptive safety control on the two-link manipulator.

2. The control method of a double-link manipulator based on a neural network observer according to claim 1, characterized in that To establish the system dynamics model of a two-link manipulator under external attacks and disturbances, it specifically includes: Utilize Establish a system dynamics model of a two-link manipulator under external attacks and disturbances; where \(q=(q_1,q_2)\) T , represent the angular position, angular velocity, and angular acceleration of the double-link manipulator joints respectively. The manipulator joints include manipulator joint 1 and manipulator joint 2. \(q_1\) represents the angular position of manipulator joint 1, and \(q_2\) represents the angular position of manipulator joint 2. represents the angular velocity of manipulator joint 1. represents the angular velocity of manipulator joint 2. represents the angular acceleration of manipulator joint 1. represents the angular acceleration of manipulator joint 2; \(M(q)\) represents the manipulator inertia matrix; represents the centrifugal Coriolis force matrix; \(G(q)=(g 11 ,g 21 ) T represents the gravity vector, where \(g 11 =(m1 + m2)gl1cos(q1)+m2gl2cos(q1 + q2)\), \(g 21 =m2gl2cos(q1 + q2)\), \(m1\) and \(m2\) represent the masses of manipulator joint 1 and manipulator joint 2 respectively, \(l1\) and \(l2\) represent the lengths of manipulator joint 1 and manipulator joint 2 respectively, and \(g\) represents the gravitational acceleration; represents the friction matrix, \(\tau d (t)\) represents the unknown external disturbance; \(\tau(t)\) represents the input torque of the manipulator joint; \(y = q\) represents the angular position of the manipulator joint.

3. A control method for a double-link manipulator based on a neural network observer according to claim 2, characterized in that, After Step 1, it further includes transforming the system dynamics model of the two-link manipulator into a nonlinear system dynamic model with attacks. The nonlinear system dynamic model is: where x1 = q, x i represents the immeasurable system state variable, i = 1, 2; u = τ(t)M -1 (q), u represents the measurable system input, M -1 (q) represents the inverse matrix of the manipulator inertia matrix; y represents the system output, and x1 represents the system output state; represents the unknown nonlinear function, are the derivatives of x1 and x2 respectively, D(t) = τ d (t)M -1 (q) represents the unknown external disturbance; a s (t, x i ) represents the attacked state of the state variable x i at time t, is the state of x i after being attacked.

4. A double-link manipulator control method based on a neural network observer according to claim 3, characterized in that After establishing the system dynamics model of a two-link manipulator under external attacks and disturbances, it further includes: Use a radial basis neural network to identify the unknown and uncertain nonlinear dynamics in the two-link manipulator, improve the dynamic model, and determine the improved dynamic model. The improved dynamic model is: Among them, represents the function approximated by the neural network, represents the weight vector, T represents the transpose, represents the basis function vector, represents the approximation error; represents the subset vector of the state variables, represents the estimate of; represents the error term between the unknown system nonlinear term function and the neural network approximation function, and y represents the output of the system.

5. A control method for a double-link manipulator based on a neural network observer according to claim 4, characterized in that, Based on the dynamics model, establish a neural network observer to observe the disturbance and attack state information of the two-link manipulator. After that, it further includes: Based on the improved dynamic model, establish a neural network observer. The neural network observer is used to observe the disturbance and attack state information of the two-link manipulator. The neural network observer is: Among them, K i , i = 1, 2 represent the designed observer gains. When i = 1, K1 is the observer gain of the robotic arm joint 1, and when i = 2, K2 is the observer gain of the robotic arm joint 2; the estimation error is defined as where the state errors are respectively expressed as x = (x1, x2) T and are respectively the estimated values of the state variables x1, x2, and are respectively expressed as the states after differentiation of represents the weight vector of the estimation of is the basis function vector.

6. A control method for a double-link manipulator based on a neural network observer according to claim 5, characterized in that, Based on reinforcement learning, the dynamics model, and the neural network observer, establish a control scheme for the two-link robotic manipulator. Specifically, it includes: According to the stability judgment method of the gradient descent method and the Lyapunov function stability theory, design the adaptive update law for the weights of the Actor–Critic neural network; Based on the neural network observer, design an adaptive law for the unknown nonlinear terms according to the adaptive update law to form a control scheme. The control scheme includes a virtual controller and an actual controller.

7. A control method for a double-link manipulator based on a neural network observer according to claim 6, characterized in that Based on the neural network observer, design an adaptive law for the unknown nonlinear terms according to the adaptive update law to form a control scheme. Specifically, it includes: Utilize Construct the virtual controller; Utilize Construct the actual controller; Among them, represents the neural network weights, represents the basis function vector; the neural network is used to approximate the unknown dynamics in u * , K1 and K2 represent the designed observer gains, Γ2 represents a state variable, represents the preset reference signal y r 's derivative, θ * 1 represents the weight vector, represents the basis function vector.

8. A control method for a double-link manipulator based on a neural network observer according to claim 7, characterized in that, Based on the adaptive update rate of the uncertain parameters designed by the backstepping method and the approximation of the neural network observer, according to the stability judgment method of the gradient descent method and the Lyapunov function stability theory, design the adaptive update law for the weights of the Actor–Critic neural network. The adaptive update rate is: where r c1 , r c2 , r a1 , r a2 , r θ1 , σ1, σ1, ω1 represent the designed parameters, I m represents the m×m identity matrix, represents the neural network weights, represents the adaptive update rate of the designed neural network weights, is the basis function vector, is the Gaussian function, i = 1, 2.

9. A double-link manipulator control system based on a neural network observer, characterized in that, The control system of the two-link manipulator based on the neural network observer adopts the control method of the two-link manipulator based on the neural network observer described in any one of the above claims 1-8. The control system of the two-link manipulator based on the neural network observer specifically includes: A system dynamics model establishment module, used to establish the system dynamics model of a two-link manipulator under external attacks and disturbances; A neural network observer establishment module, used to establish a neural network observer based on the dynamics model; A control scheme establishment module, used to establish a control scheme for the two-link manipulator based on reinforcement learning, the dynamics model, and the neural network observer; An adaptive control module, used to perform adaptive safety control on the two-link manipulator under the backstepping framework based on the control scheme.

10. A double-link manipulator control system based on a neural network observer according to claim 9, characterized in that, Control scheme establishment module, specifically including: Adaptive update law design unit, which is used to design the adaptive update law of the weights of the Actor–Critic neural network according to the gradient descent method and the stability judgment method of the Lyapunov function stability theory; Control scheme establishment unit, which is used to design an adaptive law for unknown nonlinear terms based on the neural network observer according to the adaptive update law to form a control scheme; the control scheme includes a virtual controller and an actual controller.

Citation Information

Patent Citations

  • Self-adaptive neural network control method for hydraulic mechanical arm

    CN117289612A

  • Flexible double-connecting-rod mechanical arm reinforcement learning control method based on disturbance observer

    CN117944050A

  • Fixed time optimal control method and system for single-connecting-rod mechanical arm

    CN119501947A

  • Neural network adaptive tracking control method for joint robots

    US20220152817A1

  • Robot external contact force estimation method based on artificial neural network

    WO2022257185A1