Zero-sum differential game-based modular mechanical arm actuator additive fault optimal fault-tolerant control method and equipment
By combining zero-sum differential game theory and neural networks, a dynamic model of a modular robotic arm is constructed. The Hamilton-Jacobi-Isax equation is approximately solved, which solves the dynamic unknown fault problem of the modular robotic arm, realizes real-time optimal fault-tolerant control, and reduces system energy consumption.
Patent Information
- Application Number
- CN202511173852.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-28
AI Technical Summary
The optimal fault-tolerant control of modular robotic arms is difficult to cope with dynamic random faults, the optimal control of nonlinear systems suffers from the 'curse of dimensionality' problem, and traditional methods are difficult to effectively handle the high energy consumption and stability problems caused by actuator failures.
By employing a zero-sum differential game approach, a dynamic model is constructed using joint friction torque and cross-linking coupling terms. Combining a radial basis function neural network identifier and a single-evaluation neural network, the Hamilton-Jacobi-Isax equation is approximately solved to obtain the optimal fault-tolerant control strategy, thereby reducing system energy consumption and achieving real-time control.
The accuracy of the dynamic model of the modular robotic arm system was improved, the system energy consumption was reduced, and the problem of dynamic unknown faults was effectively solved, achieving real-time optimal control.
Smart Images

Figure CN121018538A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of robot control algorithm, in particular to a modular manipulator actuator additive fault optimal fault-tolerant control method and device of zero-sum differential game. BACKGROUND
[0002] The modular manipulator is composed of a class of standardized modules and connectors, and each connection module integrates communication, sensing, driving and control units. This modular design enables it to freely add modules and flexibly reconfigure according to different task indicators and working environments, and has higher flexibility and more extensive applicability than traditional industrial robots. With these advantages, the modular manipulator has made great achievements in extreme environments such as deep space exploration, underwater exploration, rescue and disaster relief, and has been widely recognized by workers. However, due to the influence of extreme environment and use time, the probability of failure of internal components of the modular manipulator is also higher than that of traditional industrial robots. Once the actuator, which is the control command execution terminal of the modular manipulator, fails, it will not only reduce system performance, but also cause irreparable safety accidents. Therefore, the research on fault-tolerant control of modular manipulator actuator failure is of great significance.
[0003] Since the modular manipulator is usually deployed in extreme harsh environments, its control system not only needs to have strong fault-tolerant capability to ensure stable performance in the event of actuator failure, but also needs to pay attention to the energy optimization of the system to maximize the operation cycle, reduce operation and maintenance costs, and improve overall reliability, so it is particularly important to study the optimal fault-tolerant control scheme. As a major component of modern control theory, optimal control is widely used in aerospace, decision control, and industrial control. Dynamic programming, as a method for solving optimal control strategies, can solve linear quadratic optimal control problems by solving Riccati equations, but modular manipulator systems are similar to complex nonlinear systems, and the optimal control strategy needs to be obtained by solving the Hamilton-Jacobi-Bellman equation. However, this equation is a partial differential equation, and its analytical solution is not only difficult to obtain, but also often suffers from "curse of dimensionality" in the solving process. Fortunately, neural network approximators can be used to approximate the Hamilton-Jacobi-Bellman equation through adaptive dynamic programming algorithms to obtain the optimal control strategy while avoiding the "curse of dimensionality". Optimal fault-tolerant control combines the advantages of optimal control and fault-tolerant control, ensuring that the modular manipulator can not only operate stably when a fault occurs, but also reduce system energy consumption. However, the faults handled by optimal fault-tolerant control are usually static and known, but in reality, faults are random and variable.
[0004] In summary, in the optimal control process of the modular manipulator, there are problems such as the difficulty of optimal fault-tolerant control to cope with dynamic random faults, and the "curse of dimensionality" in the optimal control of nonlinear systems. SUMMARY
[0005] The purpose of the present application is to provide a zero-sum differential game modular mechanical arm actuator additive fault optimal fault-tolerant control method and device, which can reduce the system energy consumption and solve the dynamic unknown fault problem of the modular mechanical arm, and realize real-time optimal control of the modular mechanical arm.
[0006] To achieve the above purpose, the present application provides the following solutions.
[0007] In a first aspect, the present application provides a zero-sum differential game modular mechanical arm actuator additive fault optimal fault-tolerant control method, which comprises: using joint friction torque to represent joint nonlinear damping characteristics, and using cross-linking coupling terms to represent multi-joint dynamic interaction, to construct a dynamics model of a modular mechanical arm system with actuator additive fault; using a radial basis neural network identifier to identify uncertain terms in the dynamics model, to update the dynamics model; taking the position error of the modular mechanical arm system as the basis, taking the actuator additive fault and the controller input as two opposite participants in the zero-sum differential game, constructing a performance index function of the dynamics model, and defining a Hamilton function according to the performance index function, to obtain a Hamilton-Jacobi-Isaacs equation; using a single-judge neural network based on a radial basis function to approximate the performance index function to approximately solve the Hamilton-Jacobi-Isaacs equation, to obtain an approximate optimal fault-tolerant control strategy; using the approximate optimal fault-tolerant control strategy to perform real-time optimal control on the target modular mechanical arm.
[0008] In a second aspect, the present application provides a computer device, comprising: a memory, a processor to store a computer program on the memory and run the computer program on the processor, and the processor executes the computer program to implement the above-mentioned zero-sum differential game modular mechanical arm actuator additive fault optimal fault-tolerant control method.
[0009] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0010] The application represents the nonlinear damping characteristics by joint friction torque, describes the multi-joint dynamic interaction by cross-coupling terms, and constructs a dynamics model containing faults. The radial basis neural network identifier is used to update the uncertain terms in the model online, and the model accuracy is improved. The performance index function is constructed based on the position error, the actuator fault and the controller input are regarded as the opposite parties of the zero-sum differential game, the performance index function is approximated by a single-judge neural network, the Hamilton-Jacobi-Isaacs equation is approximately solved, and the optimal fault-tolerant control strategy is obtained. The application combines the game theory and the neural network, reduces the system energy consumption, effectively solves the dynamic unknown fault problem of the modular robot arm, and realizes the real-time optimal control of the modular robot arm. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the application or the related art, the drawings needed to be used in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor.
[0012] Figure 1 The flowchart of the optimal fault-tolerant control method for the modular robot arm actuator additive fault of the zero-sum differential game provided by the embodiment of the application Figure 1 .
[0013] Figure 2 The flowchart of the optimal fault-tolerant control method for the modular robot arm actuator additive fault of the zero-sum differential game provided by the embodiment of the application Figure 2 .
[0014] Figure 3 The principle diagram of the optimal fault-tolerant control method for the modular robot arm actuator additive fault of the zero-sum differential game provided by the embodiment of the application.
[0015] Figure 4 The structural schematic diagram of a computer device provided by the embodiment of the application. DETAILED DESCRIPTION
[0016] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.
[0017] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the application will be further described in detail below with reference to the drawings and specific embodiments.
[0018] Example 1, such as Figures 1-3 As shown, this embodiment provides a zero-sum differential game-based optimal fault-tolerant control method for additive faults in a modular robotic arm actuator. The optimal fault-tolerant control method for additive faults in a modular robotic arm actuator based on zero-sum differential game includes the following steps.
[0019] S1. The nonlinear damping characteristics of the joint are characterized by the joint friction torque, and the dynamic interaction of multiple joints is characterized by the cross-linking coupling term. A dynamic model of a modular robotic arm system with actuator additive faults is constructed.
[0020] Furthermore, the formula for calculating the joint friction torque is as follows.
[0021]
[0022] In the formula, f represents the joint friction torque; αi f βi f θi f εi Let q represent the viscous friction coefficient, Coulomb friction coefficient, static friction coefficient, and Stribeck friction coefficient of the i-th subsystem, respectively; i , Let γ represent the joint position vector, joint velocity vector, and joint acceleration vector of the i-th subsystem, respectively; i This represents the reduction ratio of the harmonic reducer. This represents the position-dependent friction term; sgn(·) is the sign function. These represent the nominal values of the corresponding friction coefficients. This represents the uncertainty of the coefficient of friction.
[0023] Optionally, the cross-linking coupling term can be calculated using the following formula:
[0024]
[0025] In the formula, z mi z represents the unit vector along the rotation direction of the i-th motor rotor. lj , z lk These represent unit vectors along the rotation axes of the k-th and j-th joints, respectively. These represent the dot product of the corresponding unit vectors. Represent and The estimated value, This represents the correction error.
[0026] In practical applications, the specific process of step S1 is as follows:
[0027] Firstly, the dynamics model of the ith subsystem of the modular robot manipulator with actuator additive fault is established by using joint feedback technology:
[0028]
[0029] where the subscript i represents the ith subsystem, q i , q mi , q i , respectively represent the joint position vector, joint velocity vector, joint acceleration vector of the ith subsystem, I si represents the inertia matrix of the actuator, γ i represents the reduction ratio of the harmonic reducer, represents the joint friction torque, represents the cross-coupling term between subsystems, τ fi represents the measured value of the joint torque sensor, τ αi represents the control torque output by the actuator, u βi represents the actuator additive fault.
[0030] (1) Joint friction torque
[0031] Joint friction torque The joint friction torque mainly comes from the harmonic reducer and DC motor of each joint module of the modular robot manipulator, and is expressed as a nonlinear function:
[0032]
[0033] Assuming that the nominal value of the friction parameter is close to its actual value, the joint friction torque can be approximately expressed according to the linearization criterion:
[0034]
[0035] In equation (4), f represents the uncertainty term of the friction coefficient, f θi , f εi , respectively represent the nominal value of f αi , f βi , f θi , f εi , and is defined as:
[0036]
[0037] (2) Cross-coupling term
[0038] The cross-coupling term is a nonlinear function of the coupling dynamics depending on the global joint position vector, joint velocity vector, and joint acceleration vector of the modular robot manipulator system:
[0039]
[0040] In the formula, z mi z represents the unit vector along the rotation direction of the i-th motor rotor. lj , z lk These represent unit vectors along the rotation axes of the k-th and j-th joints, respectively. The above equation can be further rewritten as:
[0041]
[0042] Define system state Control torque Rewriting the above equation as a state-space equation:
[0043]
[0044] In the formula, f i (x i ) represents the measurable part of the system, g i (x i ) represents the system control input matrix, ψ i (x) represents the system's uncertain functional term:
[0045]
[0046] S2. Use a radial basis function neural network identifier to identify uncertainties in the dynamic model and update the dynamic model.
[0047] Step S2 specifically includes:
[0048] S21. The expected value substitution method is adopted to replace the joint dynamic information in the multi-joint coupling term with the expected value, and at the same time, the substitution error term is introduced to simplify the dynamic model of a single joint.
[0049] S22. Construct a radial basis function neural network model, and establish the uncertainty term using the ideal weight vector and activation function of the radial basis function neural network model.
[0050] S23. Replace the ideal weight vector of the radial basis neural network with the actual estimated values, and approximate the uncertainty by updating the network parameters.
[0051] S24. Calculate the derivative of the identification error based on the uncertainty term and the radial basis function neural network identifier.
[0052] S25. Using the derivative of the identification error, the weights of the neural network are adjusted by adopting a weight update law based on gradient descent to update the dynamic model.
[0053] Furthermore, the calculation formula for the radial basis function neural network identifier is as follows.
[0054]
[0055] In the formula, These are system state observations; g is the estimated value of the uncertain term; i (x i ) represents the control input matrix; μ i K is the actual output of the actuator. i e is the positive definite observation gain parameter; oi To identify errors.
[0056] Optionally, an identifier based on a radial basis function neural network can be used to approximate the uncertainties and cross-coupling terms in the system dynamics model by using the system's position information as the input of the radial basis function neural network identifier, thereby improving the accuracy of the system dynamics model.
[0057] In practical applications, the specific implementation process of step S2 is as follows.
[0058] Using the substitution concept, the other joint information contained in the cross-linking coupling term is replaced with the corresponding expected value:
[0059] ψ i (x)=ψ i (x i ,x jd )+Δψ i (x,x jd ).
[0060] In the formula, x jd Δψ represents the expected value of other joint information. i (x,x jd ) represents substitution error.
[0061] Therefore, the i-th subsystem is rewritten as:
[0062]
[0063] In the formula, F i (x i ,x jd )=f i (x i )+ψ i (x,x jd ), μ i =u i +u fi This represents the actual output of the actuator.
[0064] Using radial basis function neural networks, the uncertainty term F is... i (x i ,x jdThe approximate estimate is:
[0065] F i (x i ,x jd ) = W fi Τ σ fi (x i ,x jd )+e fi .
[0066] In the formula, η represents the ideal weight vector of a radial basis function neural network. f σ represents the number of hidden neurons in a radial basis function neural network. fi (x i ,x jd ) represents the activation function of a radial basis function neural network, e fi Represents the uncertain term F i (x i ,x jd The approximate error of ).
[0067] Due to the ideal weight vector W of the radial basis function neural network fi Often unknown, therefore an estimated value is used. Replace the actual value W fi ,get:
[0068]
[0069] In the formula, Representing F respectively i (x i ,x jd ), σ fi (x i ,x jd The estimated value of ).
[0070] Define system state x i2 The observed value is The neural network identifier is designed as follows.
[0071]
[0072] In the formula, K i Represents the positive definite observation gain parameter. This represents the identification error.
[0073] Thus, the derivative of the identification error is obtained.
[0074]
[0075] In the formula, Derivative of identification error This represents the approximate error of the weights. This represents the activation function error.
[0076] The weight update law of the radial basis function neural network is designed as follows:
[0077]
[0078] In the formula, l fi This represents the positive learning rate.
[0079] S3. Based on the position error of the modular robotic arm system, the additive fault of the actuator and the input of the controller are regarded as two opposing participants in a zero-sum differential game. The performance index function of the dynamic model is constructed, and the Hamilton function is defined according to the performance index function to obtain the Hamilton-Jacobi-Isaks equation.
[0080] Furthermore, the formula for calculating the Hamilton-Jacobi-Isax equation is as follows:
[0081]
[0082] In the formula, e i For location information error, I mi γ is the moment of inertia of the motor. i For the reduction ratio of the harmonic reducer, ▽J i (e i ) represents the gradient of the performance index function, ρ i The attenuation coefficient is... For the desired velocity information, ψ i (x) is the system's uncertain function term, f i (x i () represents the measurable part of the system.
[0083] In practical applications, the specific process of step S3 is as follows.
[0084] First, define the performance metric function for infinite time as follows:
[0085]
[0086] In the formula, U i (e i (t),u i (t),u fi (t))=e i Τ Q i e i +u i Τ R i ui -ρ i 2 u fi Τ u fi Represents the utility function, e i =e i (t), e i =[e i1 ,e i2 ,…,e in ] Τ =x i -x id x represents the trajectory tracking error. id Q represents the desired position vector. i It is a positive semi-definite symmetric matrix, R i Represents a positive definite symmetric matrix, ρ i This represents the attenuation coefficient.
[0087] This leads to the optimal cost function:
[0088]
[0089] If the optimal cost function is continuously differentiable, its infinitesimal form can be expressed as:
[0090]
[0091] In the formula, Initial conditions are
[0092] Define the Hamiltonian function of a modular robotic arm system with additive actuator faults as follows:
[0093]
[0094] Based on the Nash-Pontryagin minimax theorem, we can further derive the Hamilton-Jacobi-Isax equation:
[0095]
[0096] In zero-sum differential games, by solving the Hamiltonian function for u... i u fi The partial derivatives yield the optimal solution with a value of one.
[0097]
[0098] In the formula, These represent the optimal control strategy and the "worst" failure law, respectively.
[0099] Therefore, the Hamilton-Jacobi-Isax equation can be rewritten as:
[0100]
[0101] S4. The Hamilton-Jacobi-Isax equation is approximated by a single-criteria neural network based on radial basis functions to obtain an approximate optimal fault-tolerant control strategy.
[0102] Furthermore, the formula for calculating the weight update law of a single-judgment neural network is as follows.
[0103]
[0104] In the formula, l ci β represents the adjustable parameter of the single-judge neural network. Represents the normalization factor. Represents robustness.
[0105] Optionally, a single-evaluation neural network based on radial basis functions can be used, with the system's position error information as the input of the single-evaluation neural network. The system performance index function is approximated by minimizing the squared residual formed by the difference between the ideal Hamiltonian function and the approximate Hamiltonian function, thus obtaining an approximately optimal fault-tolerant control strategy.
[0106] In practical applications, step S4 specifically includes the following steps.
[0107] First, the optimal cost function is obtained by using a single-judgment neural network. Redefining:
[0108]
[0109] In the formula, σ represents the ideal weight vector of a single-judge neural network. ci (e i η represents the activation function of a single-judge neural network. c ε represents the number of hidden neurons in a single-judge neural network. ci (e i The ) represents the approximation error. Taking the partial derivative of the above equation, we get:
[0110]
[0111] In the formula,
[0112] Therefore, we can conclude that:
[0113]
[0114]
[0115] Due to the ideal weight vector W of a single-judge neural network ci Often unknown, therefore adopt Replace W ci get:
[0116]
[0117] In the formula, This represents a near-optimal performance index function. The gradient represents the approximate optimal performance index function.
[0118] Therefore, we can conclude that:
[0119]
[0120] In the formula, These represent the approximate Hamiltonian function, the approximate optimal control strategy, and the approximate "worst-case" fault law, respectively.
[0121] Therefore, we can conclude that:
[0122]
[0123] In the formula, This represents the dynamics of the error in the weight vector. It requires minimizing the squared residuals. Force e ci →0 to achieve This is the objective. Therefore, the weight update law of the single-judge neural network is designed as follows:
[0124]
[0125] In the formula, l ci β represents the adjustable parameter of the single-judge neural network. Represents the normalization factor. Represents robustness, e ci This represents the approximate Hamiltonian function. Represents weight error,
[0126] Step S4 is followed by: constructing a modular robotic arm experimental platform, and using the modular robotic arm experimental platform to rationalize and verify the near-optimal fault-tolerant control strategy. The modular robotic arm experimental platform is a two-degree-of-freedom modular robotic arm experimental platform.
[0127] In practical applications, a modular robotic arm experimental platform is constructed to verify the rationality of the near-optimal fault-tolerant control strategy. Specifically, this includes:
[0128] First, given the desired trajectories of two joints:
[0129]
[0130] Secondly, add an additive actuator fault to joint 2:
[0131] u f2 = -0.02, t≥50s.
[0132] For the designed radial basis function neural network (RBF) identifier, a 1-5-1 structure (1 input neuron, 5 hidden neurons, and 1 output neuron) was selected. The initial weight vector of the identifier is... Activation function c1 = c2 = [2, 1, 1, 1, 2] T b j =1.5, j=1,2,3,4,5.
[0133] For the designed single-judge neural network, the Hamilton-Jacobi-Isax equation is approximated using a radial basis function neural network to solve the performance index function. A 1-5-1 structure is chosen for the single-judge neural network. The initial weight vector of the single-judge neural network is... Activation function c j1 = [-1, -0.5, 0, 0.5, 1] Τ c j2 =[-2,-1,0,1,2] Τ b j =1.5, j=1,2,3,4,5.
[0134] By utilizing the communication between the host computer and the data acquisition card, a controller is built on the host computer using SIMULINK software to control the two-degree-of-freedom modular robotic arm test platform.
[0135] S5. Utilize an approximate optimal fault-tolerant control strategy to perform real-time optimal control on the target modular robotic arm.
[0136] In practical applications, the specific implementation process of step S5 is as follows.
[0137] First, given the desired trajectories of two joints:
[0138]
[0139] Secondly, add an additive actuator fault to joint 2:
[0140] u f2 = -0.02, t≥50s.
[0141] For the designed radial basis function neural network (RBF) identifier, a 1-5-1 structure (1 input neuron, 5 hidden neurons, and 1 output neuron) was selected. The initial weight vector of the identifier is... Activation function c1 = c2 = [2, 1, 1, 1, 2] T b j =1.5, j=1,2,3,4,5.
[0142] For the designed single-judge neural network, the Hamilton-Jacobi-Isax equation is approximated using a radial basis function neural network to solve the performance index function. A 1-5-1 structure is chosen for the single-judge neural network. The initial weight vector of the single-judge neural network is... Activation function c j1 = [-1, -0.5, 0, 0.5, 1] Τ c j2 =[-2,-1,0,1,2] Τ b j =1.5, j=1,2,3,4,5.
[0143] By utilizing the communication between the host computer and the data acquisition card, a controller is built on the host computer using SIMULINK software to control the two-degree-of-freedom modular robotic arm test platform.
[0144] exist Figure 2 First, the actual joint position information is obtained using sensors in a modular robotic arm system with additive actuator faults, and the difference between this and the expected joint position information is used to obtain the joint position information error. Using the joint position information error, controller input, and unknown actuator faults, an optimal performance index function and the Hamilton-Jacobi-Isax equations are constructed. Second, a neural network based on radial basis functions is used to identify the uncertainties in the system model. Third, a single-criteria neural network is used to approximate the optimal performance index function and solve the Hamilton-Jacobi-Isax equations to obtain an approximately optimal fault-tolerant control strategy. Figure 3 This application first establishes a dynamic model of a modular robotic arm system with additive actuator faults using joint torque feedback technology. Second, leveraging the powerful approximation capability of neural networks, a neural network identifier based on radial basis functions is designed to identify uncertainties in the system's dynamic model, improving the model's accuracy. Third, a single-criteria neural network is used to approximate the performance index function to solve the Hamilton-Jacobi-Isax equations, yielding the optimal fault-tolerant control strategy for the system. Finally, the effectiveness of the proposed method is verified through experiments using a two-degree-of-freedom modular robotic arm experimental platform.
[0145] In summary, this application addresses modular robotic arm systems experiencing additive actuator failures by introducing the concept of zero-sum differential game theory into optimal fault-tolerant control. It treats the actuator failure and controller input as two completely opposing players in a game, transforming the optimal fault-tolerant control problem into a zero-sum differential game problem. First, this application establishes a dynamic model of the modular robotic arm experiencing additive actuator failures using joint torque technology and rewrites it in the form of state-space equations. Second, it designs a neural network observer based on radial basis functions to identify uncertainties and cross-coupling terms in the system, improving the accuracy of the system dynamic model. Third, it defines the position error of the modular robotic arm system, constructs a performance index function that includes the system position error, control strategy, and additive actuator failure, and further derives the Hamilton-Jacobi-Isax equations for the system. Then, it uses a single-criteria neural network to approximate the gradient of the performance index function to solve the Hamilton-Jacobi-Isax equations, obtaining the optimal control strategy and worst-case failure law. Finally, it uses a two-degree-of-freedom modular robotic arm experimental platform to experimentally verify the feasibility of the proposed method.
[0146] The technical effects of this application are as follows.
[0147] This application utilizes a radial basis function neural network to design an identifier that identifies uncertainties in a modular robotic arm system with actuator additive faults, including friction and cross-linking coupling terms. This improves the accuracy of the system's dynamic model and avoids the impact of uncertainties on the system's control performance. A single-judge neural network is used instead of a traditional actuator-judge neural network to approximate the performance index function, reducing the computational load during system operation by decreasing the number of neural networks. Regarding fault-tolerant control, two innovations are proposed. First, the performance index function constructed in this application includes not only the system's dynamic process but also its control process. By minimizing the performance index function, the optimal fault-tolerant control strategy is obtained, which, compared to traditional fault-tolerant control methods, not only ensures stable system operation but also minimizes the system's overall energy consumption. Second, by treating the actuator additive fault and the controller input as two completely opposing participants in a zero-sum differential game, the problem of optimal fault-tolerant control's difficulty in handling dynamic faults is addressed.
[0148] In summary, this application for modular robotic arm systems with additive actuator faults not only improves the accuracy of the system dynamics model and reduces the overall energy consumption of the system, but also solves the problem that traditional methods struggle to handle dynamic unknown faults.
[0149] Example 2: This application also provides a computer device, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 4As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores and processes data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements the methods described above.
[0150] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0151] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0152] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0153] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game theory, characterized in that, The modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game includes: The nonlinear damping characteristics of joints are characterized by joint friction torque, and the dynamic interaction of multiple joints is characterized by cross-linking coupling terms. A dynamic model of a modular robotic arm system with actuator additive faults is constructed. A radial basis function neural network (RBF) identifier is used to identify uncertainties in the dynamic model and update the dynamic model accordingly. Based on the position error of the modular robotic arm system, the additive fault of the actuator and the input of the controller are regarded as two opposing participants in a zero-sum differential game. The performance index function of the dynamic model is constructed, and the Hamilton function is defined according to the performance index function to obtain the Hamilton-Jacobi-Isaks equation. By using a single-criteria neural network based on radial basis functions to approximate the performance index function, the Hamilton-Jacobi-Isaac equation is solved to obtain an approximately optimal fault-tolerant control strategy. A near-optimal fault-tolerant control strategy is used to perform real-time optimal control on the target modular robotic arm.
2. The modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game theory according to claim 1, characterized in that, After approximating the performance index function using a single-criteria neural network based on radial basis functions to solve the Hamilton-Jacobi-Isax equations and obtain the approximately optimal fault-tolerant control strategy, the following steps are also included: A modular robotic arm experimental platform was constructed, and the near-optimal fault-tolerant control strategy was rationally verified using the modular robotic arm experimental platform.
3. The modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game theory according to claim 2, characterized in that, The modular robotic arm experimental platform is a two-degree-of-freedom modular robotic arm experimental platform.
4. The modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game theory according to claim 1, characterized in that, The formula for calculating the joint friction torque is as follows: In the formula, This represents the frictional torque of the joint; f αi f βi f θi f εi Let q represent the viscous friction coefficient, Coulomb friction coefficient, static friction coefficient, and Stribeck friction coefficient of the i-th subsystem, respectively; i , Let γ represent the joint position vector, joint velocity vector, and joint acceleration vector of the i-th subsystem, respectively; i This represents the reduction ratio of the harmonic reducer. This represents the position-dependent friction term; sgn(·) is the sign function. These represent the nominal values of the corresponding friction coefficients. This represents the uncertainty of the coefficient of friction.
5. The modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game theory according to claim 1, characterized in that, The calculation formula for the crosslinking coupling term is as follows: In the formula, z mi z represents the unit vector along the rotation direction of the i-th motor rotor. lj , z lk These represent unit vectors along the rotation axes of the k-th and j-th joints, respectively. These represent the dot product of the corresponding unit vectors. Represent and The estimated value, This represents the correction error.
6. The modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game theory according to claim 1, characterized in that, A radial basis function neural network (RBF) identifier is used to identify uncertainties in the dynamic model and update the dynamic model. Specifically, this includes: The expected value substitution method is adopted to replace the joint dynamic information in the multi-joint coupling term with the expected value, and at the same time, the substitution error term is introduced to simplify the dynamic model of a single joint. A radial basis function neural network model is constructed, and the uncertainty term is established by approximating the uncertainty term in the dynamic model through the ideal weight vector and activation function of the radial basis function neural network model. By replacing the ideal weight vector of the radial basis function neural network with actual estimated values, the uncertainty term is optimized by updating the network parameters. The derivative of the identification error is calculated based on the uncertainty term and the radial basis function neural network identifier; By utilizing the derivative of the identification error, a weight update law is designed using the gradient descent method. The weights of the neural network are adjusted through a positive learning rate to update the dynamic model.
7. The modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game theory according to claim 6, characterized in that, The calculation formula for the radial basis function neural network identifier is as follows: In the formula, These are system state observations; g is the estimated value of the uncertain term; i (x i ) represents the control input matrix; μ i K is the actual output of the actuator. i e is the positive definite observation gain parameter; oi To identify errors.
8. The modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game theory according to claim 1, characterized in that, The formula for calculating the Hamilton-Jacobi-Isax equation is as follows: In the formula, e i For location information error, I mi γ is the moment of inertia of the motor. i For the reduction ratio of the harmonic reducer, ▽J i (e i ) represents the gradient of the performance index function, ρ i The attenuation coefficient is... For the desired velocity information, ψ i (x) is the system's uncertain function term, f i (x i () represents the measurable part of the system.
9. The modular robotic arm actuator additive fault-tolerant control method based on zero-sum differential game theory according to claim 1, characterized in that, The formula for calculating the weight update law of the single-judge neural network is as follows: In the formula, l ci β represents the adjustable parameter of the single-judge neural network. Represents the normalization factor. Represents robustness, e ci This represents the approximate Hamiltonian function. Represents weight error, 10. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the optimal fault-tolerant control method for additive faults of a modular robotic arm actuator based on zero-sum differential game as described in any one of claims 1-9.
Citation Information
Cited By
Dual-arm robot cooperative control method and device based on differential game
CN121716088A
A differential game-based dual-arm robot cooperative control method and device
CN121716088B