A multi-modular robot optimal coordination control method, device and medium
By establishing the kinematic and dynamic models of modular robot collaborative handling tasks, and using neural network observers and two-level Nash game theory, the interaction strategy of the modular robot system is optimized, the accuracy and stability problems of the multi-level multi-agent collaborative system are solved, and optimal control is achieved.
Patent Information
- Application Number
- CN202411707310.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-11-27
AI Technical Summary
Existing control methods are difficult to effectively handle complex systems with multiple levels and multiple agents collaborating, and cannot guarantee the accuracy, optimality and stability of modular robotic systems.
A kinematic model of the collaborative handling task of modular robots is established. The cross-coupling terms between joint subsystems are observed through a neural network observer. A performance index function of a two-layer Nash game is constructed, and a radial basis function neural network is used for approximation to solve the Hamilton-Jacobi equation to obtain the optimal control law and optimal internal force.
In a complex multi-layer game system, the internal collaboration issues of modular robots are taken into consideration, and the interaction strategies between multiple modular robots are optimized to ensure the accuracy, optimality and stability of the system.
Smart Images

Figure CN119575814B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robot control, and in particular to a method, device and medium for optimal coordinated control of multi-modular robots. Background Art
[0002] With the implementation of Industry 5.0, robotic handling capabilities are finding widespread and in-depth application in a wide range of fields, including deep space exploration, disaster prevention and relief, and industrial production. However, as mission scenarios and working conditions become increasingly complex, the weaknesses of single robots, such as insufficient load capacity and limited working environments, are becoming increasingly apparent when handling complex, large-mass objects. Consequently, the need for multiple robots to collaboratively handle these complex tasks is growing. Using coordinated control algorithms to coordinate multiple robots to accomplish these tasks has become a hot research topic.
[0003] Each joint subsystem in a modular robot is independently equipped with sensing, actuation, communication, and control units. The subsystems can be reconfigured to create a modular robot system with different configurations to adapt to different mission scenarios. Compared to traditional fixed-configuration robots, modular robots offer numerous advantages, including greater scalability, maintainability, low cost, and portability.
[0004] Optimal control, as a control method that studies how to optimize the specific performance indicators of a given controlled system through appropriate control strategies, is an indispensable part of modern control theory and also one of the important research directions of modular robot system control.
[0005] Existing control methods only focus on differential games between modular robot subsystems, or analyze the game process between robots and human collaborators separately. This method works well when facing relatively simple systems, but for complex systems with multiple levels and multiple agents collaborating, it is difficult to effectively handle the interactions between different levels within the system, and cannot guarantee the accuracy, optimality and stability of the system. Summary of the Invention
[0006] The purpose of this application is to provide a multi-modular robot optimal coordination control method, equipment and medium, which can optimize the interaction strategies between multiple modular robots while taking into account the internal collaboration problems of modular robots in a complex multi-layer game system to ensure the accuracy, optimality and stability of the system.
[0007] To achieve the above objectives, this application provides the following solutions:
[0008] In a first aspect, the present application provides a multi-modular robot optimal coordination control method, which is applied to a multi-modular robot system; the multi-modular robot system includes a plurality of modular robots, each modular robot includes a plurality of joint subsystems; the multi-modular robot optimal coordination control method includes:
[0009] Establish a kinematic model for modular robot collaborative handling tasks;
[0010] Based on the kinematic model of the modular robot collaborative handling task, the overall dynamic model of the modular robot and the target object is established according to the robot joint torque feedback and Newton-Euler iteration;
[0011] Based on the overall dynamic model of the modular robot and the target object, a neural network observer is constructed to observe the cross-linked coupling terms between the joint subsystems of the modular robot. A radial basis function neural network is used to approximate the cross-linked coupling terms and obtain the update rate expression of the weights.
[0012] Construct a performance index function for the two-layer Nash game, and use a radial basis function neural network to approximate the performance index function of the two-layer Nash game. Determine the weight update rate of the two-layer Nash game based on the weight update rate expression.
[0013] The Hamilton-Jacobi equation corresponding to the performance index function of the double-level Nash game is solved according to the weight update rate of the double-level Nash game, and the optimal control law and optimal internal force are obtained to achieve coordinated control of the multi-modular robot system.
[0014] Optionally, based on the kinematic model of the modular robot collaborative handling task, an overall dynamic model of the modular robot and the target object is established according to the robot joint torque feedback and Newton-Euler iteration, including:
[0015] Establish a dynamic model of the modular robot based on the robot joint torque feedback;
[0016] Establish the dynamic model of the target object based on Newton-Euler iteration;
[0017] Based on the kinematic model of the modular robot's collaborative handling task, the dynamic model of the modular robot and the dynamic model of the target object, an overall dynamic model of the modular robot and the target object is established.
[0018] Optionally:
[0019] The kinematic model of the modular robot collaborative handling task is:
[0020] x=δ1(q1)=δ2(q2)=...=δ n (q n )=...=δN (q N );
[0021]
[0022] Among them, x represents the pose vector of the target object, represents the speed of the target object, represents the acceleration of the target object, q n represents the joint angle of the nth modular robot, represents the joint velocity of the nth modular robot, represents the joint acceleration of the nth modular robot, δ n () represents the relationship between the pose vector of the target object and the joint angle of the nth modular robot, Represents δ n The first-order derivative of (), Jcjn represents the Jacobian matrix from the center of mass of the target object to the joint of the n-th modular robot, Indicates J cjn The first derivative of , n = 1, 2, ..., N, where N represents the number of modular robots;
[0023] The dynamic model of the modular robot is:
[0024]
[0025] Among them, I n represents the moment of inertia vector of the motor of the nth modular robot, γ n represents the reduction ratio vector of the harmonic reducer of the nth modular robot, f n represents the friction torque of the nth modular robot, I cn represents the cross-coupling term between the joint subsystems of the nth modular robot, τ jn represents the torque measured by the joint torque sensor of the nth modular robot, f oen represents the force exerted by the target object on the end of the n-th modular robot, J ejn represents the Jacobian matrix from the end to the joint of the nth modular robot, Indicates J ejn The transpose of τ n represents the output torque of the motor of the nth modular robot;
[0026] The dynamic model of the target object is:
[0027]
[0028] Among them, M W (x) represents the inertia matrix of the target object, a Coriolis and centrifugal force matrix, G W (x) represents a gravity term acting on the target object, f o represents the resultant force of N modular robots acting on the target object;
[0029] The overall dynamics model of the modular robots and the target object is:
[0030]
[0031] where M zn represents the inertia matrix of the nth modular robot, C zn represents the Coriolis and centrifugal force term of the nth modular robot, G zn represents the gravity term of the nth modular robot, represents the transpose of J cjn , d n (t) represents the load distribution matrix of the nth modular robot at time t, f ni represents the internal force of the nth modular robot.
[0032] Optionally, the neural network observer is:
[0033]
[0034] where x n2 represents the second component of the state vector of the nth modular robot, represents the observer state vector of the nth modular robot, represents the first derivative of , represents the observation of the cross-coupling term between the joint subsystems of the nth modular robot, Ω n represents the positive definite observation gain matrix of the nth modular robot, u n represents the control input of the nth modular robot, M zn -1 represents the inverse of M zn .
[0035] Optionally, the performance index function of the double-layer Nash game includes: an inner-layer Nash game performance index function and an outer-layer Nash game performance index function;
[0036] The inner-layer Nash game performance index function is:
[0037]
[0038] where, represents the inner-layer Nash game performance index function, represents the utility function of the inner Nash game, represents the speed error, express The transpose of represents the inner error quadratic positive definite matrix, represents the inner interaction control input penalty matrix, u nm represents the control law of the mth joint subsystem of the nth modular robot, Indicates u nm The transpose of , m = 1, 2, ..., M, M represents the number of joint subsystems, τ represents the independent variable of time integration;
[0039] The outer Nash game performance index function is:
[0040]
[0041] in, represents the outer Nash game performance index function, represents the outer error quadratic positive definite matrix, represents the outer interaction control input penalty matrix, represents the internal force positive definite matrix, represents f ni The transpose of .
[0042] Optionally, the weight update rate of the double-layer Nash game includes: a weight update rate of the inner layer Nash game and a weight update rate of the outer layer Nash game;
[0043] The weight update rate of the inner Nash game is:
[0044]
[0045] in, represents the weight update rate of the inner Nash game, λ nm represents the learning rate of the inner Nash game radial basis neural network, e cnm represents the error function of the inner layer Nash game radial basis neural network, σ Inm represents the basis function of the inner Nash game radial basis neural network, represents the acceleration error, represents partial derivative;
[0046] The weight update rate of the outer Nash game is:
[0047]
[0048] in, represents the weight update rate of the outer Nash game, λ onrepresents the learning rate of the outer Nash game radial basis neural network, e con represents the error function of the outer Nash game radial basis neural network, Represents the basis functions of the outer Nash game radial basis neural network.
[0049] Optionally, the Hamilton-Jacobi equation is:
[0050]
[0051] Among them, J Inm represents the inner Nash game performance index function of the m-th joint subsystem of the n-th modular robot, Indicates J Inm The transpose of J On represents the outer Nash game performance index function of the nth modular robot, express The transpose of represents the dynamic characteristic matrix of the system, represents the control input matrix, represents the expected acceleration.
[0052] Optionally:
[0053] The optimal control law is:
[0054]
[0055] in, represents the optimal control law of the mth joint subsystem of the nth modular robot, represents the optimal performance index function of the inner Nash game of the m-th joint subsystem of the n-th modular robot, Representation matrix The inverse, Representation matrix The transpose of
[0056] The optimal internal force is:
[0057]
[0058] Among them, f ni * represents the optimal internal force of the nth modular robot, Represents the optimal performance indicator function of the outer Nash game of the nth modular robot.
[0059] In a second aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the optimal coordinated control method for multiple modular robots.
[0060] In a third aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the multi-modular robot optimal coordination control method.
[0061] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0062] The present application provides a method, device and medium for optimal coordinated control of multiple modular robots. First, a kinematic model of the modular robot collaborative handling task is established. Secondly, the overall dynamic model of the modular robot and the target object is established by combining the robot joint torque feedback and Newton-Euler iteration. The cross-linked coupling terms between the joint subsystems of the modular robot are observed by a neural network observer. Subsequently, the performance index functions of the inner and outer layers are constructed respectively. The radial basis function network is then used to approximate the performance index functions of the inner and outer layers respectively. The Hamilton-Jacobi equation is solved to obtain the optimal control law and the optimal internal force, thereby realizing coordinated control of the multi-modular robot system. The present application proposes a two-layer Nash game theory framework, which can optimize the interaction strategies between multiple modular robots while taking into account the internal collaboration issues of modular robots in a complex multi-layer game system, thereby ensuring the accuracy, optimality and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0064] Figure 1 Schematic diagram of a multi-modular robot system provided in this application performing a handling task;
[0065] Figure 2 Flowchart of the optimal coordinated control method for multiple modular robots provided in this application;
[0066] Figure 3 This is a diagram of the execution steps of the optimal coordinated control method for multiple modular robots provided in this application. DETAILED DESCRIPTION
[0067] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0068] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0069] In an exemplary embodiment, the present application provides a multi-modular robot optimal coordination control method, which is applied to a multi-modular robot system. Figure 1 As shown, the multi-modular robot system includes several modular robots, each of which includes several joint subsystems. Figure 2 and Figure 3 As shown, the multi-modular robot optimal coordination control method includes the following steps 101 to 105.
[0070] Step 101: Establish a kinematic model of the modular robot collaborative handling task.
[0071] The kinematic model of the modular robot collaborative handling task is as follows:
[0072] x=δ1(q1)=δ2(q2)=...=δ n (q n )=...=δ N (q N ), (1)
[0073] The above formula expresses the relationship between the robot's end position and posture and the variables of each joint. (p) T ,x (o) T ] T ∈R dt Indicates the position x of the target object (p) ∈R p and direction x (o) ∈R o The pose vector, d t (d t =p+o) represents the degree of freedom of the target object, where p represents the dimension of the target object position vector, o represents the dimension of the target object posture vector, and q n represents the joint angle of the nth modular robot, δ n() represents the relationship between the pose vector of the target object and the joint angle of the nth modular robot, n = 1, 2, ..., N, N represents the number of modular robots, and R represents a real number.
[0074] The first derivative of the pose vector represents the velocity of the target object:
[0075]
[0076] The second derivative of the pose vector represents the acceleration of the target object:
[0077]
[0078] in, represents the speed of the target object, represents the acceleration of the target object, represents the joint velocity of the nth modular robot, represents the joint acceleration of the nth modular robot, express The first derivative of J cjn represents the Jacobian matrix from the center of mass of the target object to the joints of the nth modular robot, Indicates J cjn The first derivative of .
[0079] Step 102: Based on the kinematic model of the modular robot collaborative handling task, an overall dynamic model of the modular robot and the target object is established according to the robot joint torque feedback and Newton-Euler iteration.
[0080] First, the dynamic model of the modular robot is established based on the robot joint torque feedback. The dynamic model of the nth modular robot is as follows:
[0081]
[0082] in, represents the moment of inertia vector of the motor of the nth modular robot, represents the reduction ratio vector of the harmonic reducer of the nth modular robot, represents the friction torque of the nth modular robot, represents the cross-coupling term between the joint subsystems of the n-th modular robot, represents the torque measured by the joint torque sensor of the nth modular robot, f oen ∈R M represents the force exerted by the target object on the end of the n-th modular robot, J ejn represents the Jacobian matrix from the end to the joint of the nth modular robot, represents the transpose of the end-to-joint Jacobian matrix of the nth modular robot, τ n ∈R M Represents the output torque of the motor of the nth modular robot.
[0083] Secondly, the dynamic model of the target object is established based on the Newton-Euler iteration. The dynamic model of the target object is as follows:
[0084]
[0085] Among them, M W (x) represents the inertia matrix of the target object, Represents the Coriolis force and centrifugal force matrix of the target object, G W (x) is the gravity term acting on the target object, f o represents the resultant force of N modular robots acting on the target object, which can be decomposed into:
[0086]
[0087] Among them, f ecn represents the force exerted by the nth modular robot on the target object. ecn and f oen The relationship can be expressed as:
[0088] f ecn =J ecn T f oen , (7)
[0089] Among them, J ecn The Jacobian matrix J represents the distance from the end of the nth modular robot to the center of mass of the target object. ecn T Indicates J ecn The transpose of , and f ecn can be decomposed into the internal force f ni With external force f ne The form of the sum:
[0090] f ecn =f ni +f ne , (8)
[0091] The sum of the internal forces is zero, that is:
[0092]
[0093] Therefore (5) can be written as:
[0094]
[0095] In other words, ne It can also be written as the resultant external force multiplied by the load distribution matrix:
[0096]
[0097] Among them, d n (t) is a time-dependent load distribution matrix, which represents the load distribution matrix of the n-th modular robot at time t, and thus we can get:
[0098]
[0099] in, Indicates J ecn The pseudo-inverse of , where the superscript T indicates the transpose of the matrix.
[0100] Finally, based on the kinematic model of the modular robot's collaborative handling task, the dynamic model of the modular robot and the dynamic model of the target object, an overall dynamic model of the modular robot and the target object is established.
[0101] The overall dynamic model of the modular robot and the target object is as follows:
[0102]
[0103] Among them, M zn represents the inertia matrix of the nth modular robot, C zn Denotes the Coriolis force and centrifugal force of the nth modular robot, G zn represents the gravity term of the nth modular robot, τ n Represents the output torque of the motor of the nth modular robot.
[0104] M zn It can be specifically expressed as:
[0105]
[0106] C zn It can be specifically expressed as:
[0107]
[0108] G zn It can be specifically expressed as:
[0109]
[0110] in, Indicates J cjn The transpose of .
[0111] The control input and state vector are defined as:
[0112] u n =τ n , (17)
[0113]
[0114] Among them, u n is the control input of the nth modular robot, x n is the state vector of the n-th modular robot, and the first component x of the state vector of the n-th modular robot n1 =q n Represents the joint position, the second component of the state vector of the nth modular robot Representing the joint velocity, the state space equation of the multi-modular robot collaborative handling task can be obtained as:
[0115]
[0116] Among them, the intermediate parameter l n and g n They are:
[0117]
[0118] g n =M zn -1 . (twenty one)
[0119] Among them, M zn -1 Indicates M zn The inverse of.
[0120] Step 103: Based on the overall dynamic model of the modular robot and the target object, a neural network observer is constructed to observe the cross-linked coupling terms between the joint subsystems of the modular robot, and a radial basis function neural network is used to approximate the cross-linked coupling terms to obtain an update rate expression for the weights.
[0121] The observation of unknown cross-coupling terms between joint subsystems of modular robots is completed through neural networks. The specific method is as follows:
[0122] First, substituting (16), (17), and (18) into the state space equation (19) yields:
[0123]
[0124] Construct a neural network observer to observe the unknown cross-coupling terms between the joint subsystems of the modular robot:
[0125]
[0126] in, represents the observer state vector of the nth modular robot, express The first derivative of represents the observation of the cross-coupling term between the joint subsystems of the nth modular robot, Ω n is a positive definite observation gain matrix of the nth modular robot. The error dynamics of the observer can be expressed as:
[0127]
[0128] The observation error can be expressed as: express The first-order derivative of . The cross-linked coupling term is approximated by a radial basis neural network:
[0129] I cn =W cn T σ cn +ε cn , (25)
[0130] Among them, W cn represents the ideal weight, σ cn represents the basis function, ε cn is the approximate error, from which we can get:
[0131]
[0132] and Respectively represent I cn and W cn The approximation of , and then the update rate expression of the weight is obtained:
[0133]
[0134] Among them, η cn is a positive constant.
[0135] Step 104: Construct a performance indicator function of the two-layer Nash game, and use a radial basis function neural network to approximate the performance indicator function of the two-layer Nash game, and determine the weight update rate of the two-layer Nash game according to the weight update rate expression.
[0136] Among them, the performance index function of the two-layer Nash game includes: the inner Nash game performance index function and the outer Nash game performance index function; the weight update rate of the two-layer Nash game includes: the weight update rate of the inner Nash game and the weight update rate of the outer Nash game.
[0137] The inner Nash game performance indicator function is constructed as follows:
[0138]
[0139] This performance indicator function represents the error as the optimization target of the control system, where represents the inner Nash game performance index function, represents the utility function of the inner Nash game, Is a function of time t, representing the speed error, x2=[x n12 ,...,x nm2 ,...,x nM2 ] T , represents the joint velocity, represents the expected speed, represents the inner error quadratic positive definite matrix, represents the inner interaction control input penalty matrix, u nm represents the control law of the mth joint subsystem in the nth modular robot, Indicates u nm The transpose of , m = 1, 2, ..., M, M represents the number of joint subsystems, τ represents the independent variable of time integration. From this, we can get the expression of the inner Nash game Hamiltonian function:
[0140]
[0141] The Hamiltonian function is the partial derivative of the performance index function with respect to time. The purpose is to find the extreme value of the performance index function to achieve optimal control. Indicates J Inm The partial derivative of the velocity error, represents the dynamic characteristic matrix of the system, represents the control input matrix. Therefore, the optimal performance index function of the mth joint subsystem in the nth modular robot can be expressed as:
[0142]
[0143] It means that by selecting a suitable control strategy, the integral of the utility function is minimized, thereby achieving the purpose of optimizing system performance. By taking the partial derivative of (29) with respect to the control law, we can obtain:
[0144]
[0145] According to (31), the optimal control law expression of the mth joint subsystem in the nth modular robot can be obtained as:
[0146]
[0147] It indicates that the optimal control law of the joint subsystem is determined by the gradient of the performance index function and the penalty matrix is input by the inner interactive control and the control input matrix Correction.
[0148] The outer Nash game performance indicator function is constructed as follows:
[0149]
[0150] in, Represents the outer Nash game performance index function, u n =[u n1 ,...,u nm ,...u nM ] T represents the control law vector of the nth modular robot, Q On represents the outer error quadratic positive definite matrix, represents the outer interaction control input penalty matrix, represents the positive definite matrix of internal forces, f ni Indicates internal force, represents f ni The transpose of . From this, we can get the expression of the outer Nash game Hamiltonian function:
[0151]
[0152] in, Indicates J On The partial derivative of , thus, the optimal performance index function of the outer Nash game can be expressed as:
[0153]
[0154] By taking the partial derivative of (34) with respect to the internal force, we can obtain:
[0155]
[0156] According to (36), the optimal internal force expression of the outer Nash game can be obtained:
[0157]
[0158] Among them, f ni * represents the optimal internal force of the nth modular robot, Represents the optimal performance indicator function of the outer Nash game of the nth modular robot.
[0159] Use radial basis neural network to approximate the inner Nash game performance index function:
[0160]
[0161] Among them, W Inm represents the ideal weight of the inner layer Nash game radial basis neural network, σ Inm represents the basis function of the inner Nash game radial basis neural network, ε Inm It represents the approximation error of the inner Nash game radial basis neural network.
[0162] According to (38), we can deduce that:
[0163]
[0164] in, represents the optimal control law of the mth joint subsystem of the nth modular robot, represents the optimal performance index function of the inner Nash game of the m-th joint subsystem of the n-th modular robot, Representation matrix The inverse, Representation matrix The transpose of represents the partial derivative of the basis function with respect to the error, It represents the gradient of the approximation error.
[0165] The approximate optimal performance indicator function can be expressed as:
[0166]
[0167] in, represents the estimated weights of the inner Nash game.
[0168] Ideal Hamiltonian function H Inm It can be expressed as:
[0169]
[0170] Approximate Hamiltonian function It can be expressed as:
[0171]
[0172] Define the error function expression:
[0173]
[0174] According to the gradient descent method, the weight update rate can be obtained as:
[0175]
[0176] in, represents the weight update rate of the inner Nash game, λnm represents the learning rate of the inner Nash game radial basis neural network, e cnm represents the error function of the inner layer Nash game radial basis neural network, σ Inm represents the basis function of the inner Nash game radial basis neural network, represents the acceleration error, represents the partial derivative, E cnm Represents the inner error gradient function.
[0177] Use radial basis neural network to approximate the outer Nash game performance index function:
[0178]
[0179] in, represents the ideal weights of the outer Nash game radial basis neural network, represents the basis function of the outer Nash game radial basis neural network, It represents the approximation error of the outer Nash game radial basis neural network.
[0180] According to (45), we can deduce that:
[0181]
[0182] in, represents the partial derivative of the basis function with respect to the error, It represents the gradient of the approximation error.
[0183] The approximate optimal performance indicator function can be expressed as:
[0184]
[0185] in, represents the estimated weights of the outer Nash game.
[0186] Ideal Hamiltonian function It can be expressed as:
[0187]
[0188] Approximate Hamiltonian function It can be expressed as:
[0189]
[0190] Define the error function expression:
[0191]
[0192] According to the gradient descent method, the weight update rate can be obtained as:
[0193]
[0194] in, represents the weight update rate of the outer Nash game, λ on represents the learning rate of the outer Nash game radial basis neural network, e con represents the error function of the outer Nash game radial basis neural network, Represents the basis function of the outer Nash game radial basis neural network, E con Represents the outer error gradient function.
[0195] Step 105: Solve the Hamilton-Jacobi equation corresponding to the performance index function of the double-layer Nash game according to the weight update rate of the double-layer Nash game to obtain the optimal control law and optimal internal force to achieve coordinated control of the multi-modular robot system.
[0196] For the inner Nash game, the control strategy is improved through strategy iteration. The specific process is as follows:
[0197] Step 1: Set variable k = 0 and select the initial admissible control strategy And choose a very small positive constant ε nm .
[0198] Step 2: When k>0, based on Approximate solution of the Hamilton-Jacobi equation based on the weight update rate And find the performance index function And there is
[0199] Step 3: Pass Update control strategy
[0200] Step 4: When When , the policy iteration process ends and the optimal control law is obtained, otherwise it returns to the second step and sets k = k + 1.
[0201] For the outer Nash game, the corresponding Hamilton-Jacobi equation is The iterative solution process is similar to that of the original one, and the optimal internal force can be obtained.
[0202] The derivation process of the above formula proves that compared with the traditional method of only focusing on the differential game between the joint subsystems of the modular robot, or only analyzing the game process between the modular robot and the human collaborator, the theoretical method proposed in this application takes into account the internal collaboration problems of the modular robot subsystems in a complex multi-layer game system, while optimizing the interaction strategies between multiple modular robot systems. It can obtain the optimal control law of each subsystem within the modular robot and obtain the optimal internal force.
[0203] In summary, this application proposes a two-layer Nash game theory framework to achieve the optimization of the interaction strategy between multiple modular robot systems while taking into account the internal collaboration problems of modular robot subsystems in a complex multi-layer game system. In addition, this application solves the problem of the Hamilton-Jacobi equation in the optimal control problem by applying the adaptive dynamic programming method. At the same time, it incorporates the idea of game theory and regards the multiple joint subsystems in the modular robot and the multiple modular robot systems when collaboratively transporting the target object as participants in the game theory. Through the inner and outer two-layer Nash games, the optimal control problem of the modular robot system is transformed into a multi-subsystem and multi-robot game problem.
[0204] In an exemplary embodiment, the present application further provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0205] In an exemplary embodiment, the present application further provides a computer-readable storage medium storing a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.
[0206] In an exemplary embodiment, the present application further provides a computer program product, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.
[0207] In this application, all actions to obtain signals, information, or data are performed in compliance with the relevant data protection laws and policies of the country in which they are located and with the authorization of the corresponding device owner. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with relevant laws and regulations.
[0208] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0209] The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0210] The technical features of the above embodiments can be combined arbitrarily. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.
[0211] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A multi-modular robot optimal coordination control method, characterized in that: Applied to a multi-modular robot system; the multi-modular robot system includes several modular robots, each modular robot includes several joint subsystems; the multi-modular robot optimal coordinated control method includes: Establish a kinematic model for modular robot collaborative handling tasks; Based on the kinematic model of the modular robot collaborative handling task, the overall dynamic model of the modular robot and the target object is established according to the robot joint torque feedback and Newton-Euler iteration; Based on the overall dynamic model of the modular robot and the target object, a neural network observer is constructed to observe the cross-linked coupling terms between the joint subsystems of the modular robot. A radial basis function neural network is used to approximate the cross-linked coupling terms and obtain the update rate expression of the weights. Construct a performance index function for the two-layer Nash game, and use a radial basis function neural network to approximate the performance index function of the two-layer Nash game. Determine the weight update rate of the two-layer Nash game based on the weight update rate expression. The Hamilton-Jacobi equation corresponding to the performance index function of the double-level Nash game is solved according to the weight update rate of the double-level Nash game, and the optimal control law and optimal internal force are obtained to achieve coordinated control of the multi-modular robot system.
2. The optimal coordination control method for multiple modular robots according to claim 1, characterized in that: Based on the kinematic model of the modular robot collaborative handling task, the overall dynamic model of the modular robot and the target object is established according to the robot joint torque feedback and Newton-Euler iteration, including: Establish a dynamic model of the modular robot based on the robot joint torque feedback; Establish the dynamic model of the target object based on Newton-Euler iteration; Based on the kinematic model of the modular robot's collaborative handling task, the dynamic model of the modular robot and the dynamic model of the target object, an overall dynamic model of the modular robot and the target object is established.
3. The optimal coordinated control method for multiple modular robots according to claim 2, characterized in that: The kinematic model of the modular robot collaborative handling task is: x=δ1(q1)=δ2(q2)=...=δ n (q n )=...=d N (q N ); Among them, x represents the pose vector of the target object, represents the speed of the target object, represents the acceleration of the target object, q n represents the joint angle of the nth modular robot, represents the joint velocity of the nth modular robot, represents the joint acceleration of the nth modular robot, δ n () represents the relationship between the pose vector of the target object and the joint angle of the nth modular robot, Represents δ n The first derivative of (), J cjn represents the Jacobian matrix from the center of mass of the target object to the joints of the nth modular robot, Indicates J cjn The first derivative of , n = 1, 2, ..., N, where N represents the number of modular robots; The dynamic model of the modular robot is: Among them, I n represents the moment of inertia vector of the motor of the nth modular robot, γ n represents the reduction ratio vector of the harmonic reducer of the nth modular robot, f n represents the friction torque of the nth modular robot, I cn represents the cross-coupling term between the joint subsystems of the nth modular robot, τ jn represents the torque measured by the joint torque sensor of the nth modular robot, f oen represents the force exerted by the target object on the end of the n-th modular robot, J ejn represents the Jacobian matrix from the end to the joint of the nth modular robot, Indicates J ejn The transpose of τ n represents the output torque of the motor of the nth modular robot; The dynamic model of the target object is: Among them, M W (x) represents the inertia matrix of the target object, Represents the Coriolis force and centrifugal force matrix of the target object, G W (x) represents the gravity acting on the target object, f o represents the resultant force exerted by N modular robots on the target object; The overall dynamic model of the modular robot and the target object is: Among them, M zn represents the inertia matrix of the nth modular robot, C zn Denotes the Coriolis force and centrifugal force of the nth modular robot, G zn represents the gravity term of the nth modular robot, Indicates J cjn The transpose of d n (t) represents the load distribution matrix of the nth modular robot at time t, f ni represents the internal force of the nth modular robot.
4. The optimal coordinated control method for multiple modular robots according to claim 3, characterized in that: The neural network observer is: Among them, x n2 represents the second component of the state vector of the n-th modular robot, represents the observer state vector of the nth modular robot, express The first derivative of represents the observation of the cross-coupling term between the joint subsystems of the nth modular robot, Ω n represents the positive definite observation gain matrix of the nth modular robot, u n represents the control input of the nth modular robot, M zn -1 Indicates M zn The inverse of.
5. The optimal coordinated control method for multiple modular robots according to claim 4, characterized in that: The performance index function of the double-layer Nash game includes: an inner layer Nash game performance index function and an outer layer Nash game performance index function; The inner Nash game performance index function is: in, represents the inner Nash game performance index function, represents the utility function of the inner Nash game, represents the speed error, express The transpose of represents the inner error quadratic positive definite matrix, represents the inner interaction control input penalty matrix, u nm represents the control law of the mth joint subsystem of the nth modular robot, Indicates u nm The transpose of , m = 1, 2, ..., M, M represents the number of joint subsystems, τ represents the independent variable of time integration; The outer Nash game performance index function is: in, represents the outer Nash game performance index function, represents the outer error quadratic positive definite matrix, represents the outer interaction control input penalty matrix, represents the internal force positive definite matrix, represents f ni The transpose of .
6. The optimal coordinated control method for multiple modular robots according to claim 5, characterized in that: The weight update rate of the double-layer Nash game includes: the weight update rate of the inner layer Nash game and the weight update rate of the outer layer Nash game; The weight update rate of the inner Nash game is: in, represents the weight update rate of the inner Nash game, λ nm represents the learning rate of the inner Nash game radial basis neural network, e cnm represents the error function of the inner layer Nash game radial basis neural network, σ Inm represents the basis function of the inner Nash game radial basis neural network, represents the acceleration error, represents partial derivative; The weight update rate of the outer Nash game is: in, represents the weight update rate of the outer Nash game, λ on represents the learning rate of the outer Nash game radial basis neural network, e con represents the error function of the outer Nash game radial basis neural network, Represents the basis functions of the outer Nash game radial basis neural network.
7. The optimal coordinated control method for multiple modular robots according to claim 6, characterized in that: The Hamilton-Jacobi equation is: Among them, J Inm represents the inner Nash game performance index function of the m-th joint subsystem of the n-th modular robot, Indicates J Inm The transpose of represents the outer Nash game performance index function of the nth modular robot, express The transpose of represents the dynamic characteristic matrix of the system, represents the control input matrix, represents the expected acceleration.
8. The optimal coordinated control method for multiple modular robots according to claim 7, characterized in that: The optimal control law is: in, represents the optimal control law of the mth joint subsystem of the nth modular robot, represents the optimal performance index function of the inner Nash game of the m-th joint subsystem of the n-th modular robot, Representation matrix The inverse, Representation matrix The transpose of The optimal internal force is: Among them, f ni * represents the optimal internal force of the nth modular robot, Represents the optimal performance indicator function of the outer Nash game of the nth modular robot.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-modular robot optimal coordination control method according to any one of claims 1 to 8 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the multi-modular robot optimal coordination control method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Multi-person zero-sum game reconfigurable robot optimal control method and system
CN113910241A
Coordination control method, system and equipment for double-arm reconfigurable robot and medium
CN116834016A