Humanoid robot servo constraint robust control method and system based on Stackelberg game

By constructing a dynamic model incorporating uncertainties and optimizing control parameters using Stackelberg game theory, the problem of decreased tracking performance caused by uncertainties in a humanoid robot dual-arm system was solved, achieving high-precision trajectory tracking and strong robustness, and optimizing the tuning process of controller parameters.

CN122008185APending Publication Date: 2026-05-12HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2025-12-25
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle uncertainties such as dynamic changes, sensor measurement errors, and external disturbances in humanoid robot dual-arm systems, leading to decreased tracking performance or even system instability. Furthermore, existing robust control methods lack systematic optimization theory, making it difficult to balance system performance and control costs.

Method used

A robust control method for humanoid robots based on Stackelberg game theory is constructed. By establishing a dynamic model containing uncertainties, a robust controller is designed, and the control parameters are optimized using Stackelberg game theory, so as to achieve high-precision trajectory tracking and strong robustness of the system in uncertain environments.

Benefits of technology

It achieves high-precision trajectory tracking capability, strong robustness and stable control input in complex environments, and optimizes controller parameters to enable the system to maintain high efficiency and stability under uncertainty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122008185A_ABST
    Figure CN122008185A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a humanoid robot servo constraint robust control method and system based on a Stackelberg game, and belongs to the technical field of robot servo control and robust control. The control method comprises the following steps: constructing a two-arm system of the humanoid robot; a humanoid robot dynamic model with uncertainty is constructed according to the humanoid robot two-arm system; determining a servo constraint equation of the kinetic model; constructing a robust controller; constructing a control parameter setting model based on a Stackelberg game, and performing optimization determination under engineering constraint on key control parameters in the robust controller according to the control parameter setting model to obtain an optimal parameter solution; and the optimal parameter solution is input into the humanoid robot double-arm system to obtain the optimal performance. According to the invention, optimization and setting of controller parameters are realized through the Stackelberg game, so that the system has high-precision trajectory tracking capability, high robustness and stable control input in an uncertain environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot servo control and robust control technology, specifically to a robust control method and system for humanoid robots based on Stackelberg game theory servo constraints. Background Technology

[0002] Humanoid robots have found widespread application in service, manufacturing, entertainment, and human-robot interaction. Dual-arm humanoid robots typically perform highly synchronized and force-sensitive operations in constrained, often unstructured, three-dimensional environments. Compared to single-arm tasks, these scenarios inherently involve closed-chain interactions and require the end effector to maintain precise posture, posing challenges in dynamic coupling, robust coordinated control, and other aspects of the system. High-precision trajectory tracking control is crucial for achieving these precise operations. However, dual-arm robots are inevitably affected by various uncertainties in actual operation, such as dynamic load changes, sensor measurement errors, unmodeled dynamics, and external disturbances. These uncertainties can significantly degrade the system's tracking performance and even lead to system instability.

[0003] In dynamic modeling, traditional energy-based Lagrangian methods have limitations when dealing with non-ideal constraints. Although the Udwadia-Kalaba (UK) equations provide a unified and accurate framework for modeling such constrained mechanical systems and transform the trajectory tracking problem into a servo-constraint following problem, effectively handling system uncertainties within this framework remains a key challenge. Existing research attempts to combine the UK method with robust control to suppress unknown disturbances, but their controller design often relies on experience or trial and error to tune parameters, lacking systematic optimization theory guidance.

[0004] Furthermore, existing robust control parameter optimization methods often focus on a single design objective and a single parameter, making it difficult to balance the complex trade-offs between system performance (such as transient response and steady-state accuracy) and control costs (such as energy consumption). Although game theory (such as Nash games) has been introduced into the field of control to solve multi-objective optimization problems, how to construct a non-cooperative game model and utilize the leader-follower decision sequence characteristics described by Stackelberg game theory to collaboratively optimize multiple key control parameters in the specific problem of servo constraint tracking of dual-arm robots has not yet been fully studied. Summary of the Invention

[0005] The purpose of this invention is to provide a robust control method and system for humanoid robots based on Stackelberg game theory, which solves the problems of low tracking accuracy, robustness and overall control efficiency of dual-arm humanoid robots in complex and uncertain environments.

[0006] To achieve the above objectives, embodiments of the present invention provide a robust control method for servo constraints of a humanoid robot based on Stackelberg game theory, the control method comprising: Constructing a humanoid robot dual-arm system; Based on the described humanoid robot dual-arm system, a dynamic model of the humanoid robot with uncertainties is constructed. Determine the servo constraint equations of the dynamic model; Build robust controllers; A control parameter tuning model based on Stackelberg game is constructed, and the key control parameters in the robust controller are optimized under engineering constraints according to the control parameter tuning model to obtain the optimal parameter solution. The optimal parameter solution is input into the humanoid robot dual-arm system to obtain optimal performance.

[0007] Optionally, a dynamic model of the humanoid robot with uncertainties is constructed based on the humanoid robot dual-arm system, including: Construct a dynamic model of the humanoid robot based on formula (1). (1) in, Let be the system's inertia matrix. For the Coriolis centrifugal force of the system, For the gravity term of the system, For system control torque, For the uncertainty parameters of the system, For joint angle, The joint angular velocity, Joint angular acceleration, For time, These are the uncertainty parameters of the system.

[0008] Optionally, determining the servo constraint equations of the dynamic model includes: The humanoid robot dual-arm system is simplified into a system of two planar two-degree-of-freedom robotic arms; The desired motion trajectories of the two planar two-degree-of-freedom robotic arms are determined according to formula (2). (2) The first-order constraints are obtained according to formula (3). (3) The second-order constraints are obtained according to formula (4). (4) in, For the desired motion trajectory, For the constraint matrix, For the angle of the joints in the system, Let be the angular velocity of the joints in the system. Let be the angular acceleration of the joints in the system. For the expected first-order constraint and , For the expected second-order constraint and , These are the dimensions of the expected first-order constraints and the dimensions of the expected second-order constraints, respectively. .

[0009] Optionally, a robust controller is constructed, including: The system control matrix is ​​determined according to formulas (5) to (8). (5) (6) (7) (8) in, For nominal binding terms, For correction items, For robust feedback items, It is a positive definite matrix. These are the weighting coefficients. For constant control parameters, , As the first intermediate parameter, It is a symmetrical term. The first control parameter, For smooth switching functions, This is the weighted constraint error vector. For the Coriolis centrifugal force of the system, This is the gravity term of the system.

[0010] Optionally, a control parameter tuning model based on Stackelberg game is constructed, and the key control parameters in the robust controller are optimized under engineering constraints according to the control parameter tuning model to obtain the optimal parameter solution, including: Construct a control parameter tuning model based on Stackelberg game theory; The first control parameter is determined as the leader of the game based on the robust controller; The second control parameter is determined to be the follower in the game based on the robust controller; The leader and followers engage in Stackelberg game based on the control parameter tuning model to optimize and tune the control parameters and obtain the optimal parameter solution.

[0011] Optionally, a control parameter tuning model based on Stackelberg game is constructed, including: The leader's cost function is determined according to formula (9). (9) The cost function of the followers is determined according to formula (10). (10) in, For the leader's cost function, For system transient performance indicators, The cost function for followers, For the steady-state performance indicators of the system, for Operations, The first control parameter, This is the second control parameter. It is a fuzzy number with uncertainty.

[0012] Optionally, the leader and followers perform a Stackelberg game based on the control parameter tuning model to optimize and tune the control parameters and obtain the optimal parameter solution, including: The leader selects parameters in the first decision set; Followers determine their strategy based on the parameters chosen by the leader in order to minimize the follower's cost function; The leader minimizes the leader's objective function according to the strategy to obtain the leader's optimal solution; The followers update their strategy based on the leader's optimal solution to obtain the followers' optimal solution in the second decision set.

[0013] Optionally, the leader and followers perform a Stackelberg game based on the control parameter tuning model to optimize and tune the control parameters and obtain the optimal parameter solution, further comprising: The optimal solution is then verified under sufficient conditions. If the verification passes, the optimal parameter solution will be output; If the verification fails, the Stackelberg game is repeated to optimize and tune the control parameters to obtain the optimal parameter solution.

[0014] On the other hand, the present invention also provides a robust control system for humanoid robot servo constraints based on Stackelberg game theory, the system including a processor for executing the control method as described above.

[0015] Through the above technical solution, this invention provides a robust control method and system for humanoid robots based on Stackelberg game theory. The trajectory tracking problem of the humanoid robot's dual-arm joints is transformed into a servo constraint problem. Servo constraint equations for the humanoid robot's dual-arm system are established, and considering the system's uncertainties, a dynamic model incorporating uncertainty terms is constructed. Based on this model, a robust controller is designed to ensure the consistent boundedness and consistent eventual boundedness of the constraint tracking error. Subsequently, Stackelberg game theory is introduced, with key controller parameters acting as game participants. A cost function model for both parties is constructed, and the optimal parameter solution is obtained by solving the Stackelberg equilibrium. This invention achieves optimized tuning of controller parameters through Stackelberg game theory, enabling the system to possess high-precision trajectory tracking capability, strong robustness, and stable control input under uncertain environments.

[0016] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a control method according to an embodiment of the present invention; Figure 2 This is a flowchart of constructing servo constraint equations according to one embodiment of the present invention; Figure 3 This is a flowchart illustrating the process of obtaining the optimal solution according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating the construction of a Stackelberg game model according to one embodiment of the present invention; Figure 5 This is a flowchart of a Stackelberg game according to one embodiment of the present invention; Figure 6 This is an error comparison chart of one embodiment of the present invention; Figure 7 This is a control input comparison diagram of one embodiment of the present invention; Figure 8 This is a simplified model diagram of one embodiment of the present invention. Detailed Implementation

[0018] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0019] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0020] Figure 1 This is a flowchart of a control method according to an embodiment of the present invention, in which the control method includes: In step S1, a humanoid robot dual-arm system is constructed.

[0021] In step S2, a dynamic model of the humanoid robot with uncertainties is constructed based on the humanoid robot's dual-arm system.

[0022] In step S3, the servo constraint equations of the dynamic model are determined.

[0023] In step S4, a robust controller is constructed.

[0024] In step S5, a control parameter tuning model based on Stackelberg game is constructed, and the key control parameters in the robust controller are optimized under engineering constraints according to the control parameter tuning model to obtain the optimal parameter solution.

[0025] In step S6, the optimal parameter solution is input into the humanoid robot dual-arm system to obtain optimal performance.

[0026] In steps S1 to S6, a dynamic model of a humanoid robot with uncertainty is constructed. The trajectory tracking problem of the humanoid robot's two arm joints is then transformed into a servo constraint problem. The servo constraint equation of the robot system is proposed. Based on this model, a robust controller is designed to ensure the consistent boundedness and consistent eventual boundedness of the constraint tracking error. Subsequently, Stackelberg game theory is introduced, and the two key parameters of the robust controller are used as game participants. The optimal parameter solution is obtained by solving the Stackelberg equilibrium.

[0027] In existing technologies, energy-based Lagrangian methods have limitations in handling non-ideal constraints. Although the Udwadia-Kalaba (UK) equations provide a unified and accurate framework for modeling such constrained mechanical systems and transform the trajectory tracking problem into a servo-constrained following problem, effectively handling system uncertainties within this framework remains a key challenge. While existing research attempts to combine the UK method with robust control to suppress unknown disturbances, controller design typically relies on experience or trial-and-error to tune parameters, lacking systematic optimization theoretical guidance. Furthermore, existing robust control parameter optimization methods often focus on single design objectives and single parameters, making it difficult to balance the complex trade-off between system performance and control costs. Compared to existing technologies, this invention transforms the trajectory tracking problem of a humanoid robot's dual-arm joints into a servo-constraint problem, establishes the servo-constraint equations for the humanoid robot's dual-arm system, and considers the system's uncertainties, constructing a dynamic model including uncertainty terms. Through Stackelberg game theory, the two key parameters of the controller are optimized and tuned, enabling the system to possess high-precision trajectory tracking capability, strong robustness, and stable control input under uncertain environments.

[0028] In step S2, the method for constructing the humanoid robot dynamics model can be varied and known to those skilled in the art. In one example of the present invention, the humanoid robot dynamics model can be constructed according to formula (1). (1) in, Let be the system's inertia matrix. For the Coriolis centrifugal force of the system, For the gravity term of the system, For system control torque, For the uncertainty parameters of the system, For joint angle, The joint angular velocity, Joint angular acceleration, For time, For the uncertainty parameters of the system, For the Coriolis centrifugal force of the system, The gravity term of the system is given above. All of the above systems are systems with two planar two-degree-of-freedom robotic arms.

[0029] For humanoid robots, the control objective of tracking the trajectory of their dual arms is achieved by constructing servo constraint equations. This transforms the desired joint motion trajectory into servo constraints that the system must follow, thus converting the complex trajectory tracking problem into a strict constraint following problem. In this implementation, the method for constructing the servo constraint equations can be as follows: Figure 2 The method shown. Figure 2The control method also includes: In step S31, the humanoid robot dual-arm system is simplified into a system of two planar two-degree-of-freedom robotic arms. For example... Figure 8 As shown, the humanoid robot dual-arm system is simplified into a system of two planar two-degree-of-freedom robotic arms. Each robotic arm has two joints, for a total of four joints, namely A, B, C, and D, which are controlled by four corresponding motors. Indicates the first The first robotic arm The mass of each link. Indicates the first The first robotic arm The length of each link Indicates the first The first robotic arm Moment of inertia of each link Indicates the first The first robotic arm The angle of each joint Indicates the first The first robotic arm angular velocity of each joint Indicates the first The first robotic arm angular acceleration of each joint Specifically, the dynamic model can be extended based on formulas (11) to (23). , (11) , (12) , (13) (14) (15) (16) (17) (18) (19) in, (20) ,(twenty one) in, ,(twenty two) ,(twenty three) in, The angle of the first joint of the first robotic arm. The angle of the second joint of the first robotic arm. The angle of the first joint of the second robotic arm. The angle of the second joint of the second robotic arm. The angular velocity of the first joint of the first robotic arm. The angular velocity of the second joint of the first robotic arm. The angular velocity of the first joint of the second robotic arm. The angular velocity of the second joint of the second robotic arm. Let be the angular acceleration of the first joint of the first robotic arm. The angular acceleration of the second joint of the first robotic arm. The angular acceleration of the first joint of the second robotic arm. The angular acceleration of the second joint of the second robotic arm. For the first robotic arm joint 1, This is the coupling inertia term between the first robotic arm joint 1 and joint 2. This is the coupling inertia term between the first robotic arm joint 1 and joint 2. For the principal inertia term of the first robotic arm joint 2, For the principal inertia term of the second robotic arm joint 1, For the coupling inertial term of the second robotic arm joint 1 and joint 2, For the coupling inertial term of the second robotic arm joint 1 and joint 2, For the principal inertial term of the second robotic arm joint 2, The control torque for the first robotic arm joint 1, The control torque for the first robotic arm joint 2, For the control torque of the second robotic arm joint 1, For the control torque of the second robotic arm joint 2, The mass of the first link of the first robotic arm. and The first The length of the first link of the robotic arm, the length of the first link... The length of the second link of the robotic arm. However, since the two robotic arms of the humanoid robot are symmetrical, with equal length, mass, and moment of inertia, it can be simplified to consider... , = , The mass of the second link of the first robotic arm. Let be the moment of inertia of the first link of the first robotic arm. Let be the moment of inertia of the second link of the first robotic arm. The mass of the first link of the second robotic arm. The mass of the second link of the second robotic arm. For the autocoupling Coriolis term of the first robotic arm joint 1, The Coriolis coupling term of the velocity of the second joint of the first robotic arm with respect to the first joint. Let the Coriolis coupling term be the velocity of the first joint of the first robotic arm with respect to the second joint. For the autocoupling Coriolis term of the second joint of the first robotic arm, For the autocoupling Coriolis term of the second robotic arm joint 1, The Coriolis coupling term of the velocity of the second joint of the second robotic arm with respect to the first joint. The Coriolis coupling term of the velocity of the first joint of the second robotic arm with respect to the second joint. For the autocoupling Coriolis term of the second joint of the second robotic arm, For the first joint of the first robotic arm, the gravity term is... For the gravity term of the second joint of the first robotic arm, For the gravity term of the first joint of the second robotic arm, This is the gravity term of the second joint of the second robotic arm.

[0030] In step S32, the desired motion trajectories of the two planar two-degree-of-freedom robotic arms are determined according to formula (2). (2) In step S33, the first-order constraints are obtained according to formula (3). (3) In step S34, the second-order constraints are obtained according to formula (4). (4) in, For the desired motion trajectory, For the constraint matrix, For the angle of the joints in the system, Let be the angular velocity of the joints in the system. Let be the angular acceleration of the joints in the system. For the expected first-order constraint and , For the expected second-order constraint and , These are the dimensions of the expected first-order constraints and the dimensions of the expected second-order constraints, respectively. , , , Formula (3) ensures that the velocity of each joint is equal to the desired velocity, and formula (4) ensures that the acceleration of each joint is equal to the desired acceleration.

[0031] In steps S31 to S34, the humanoid robot dual-arm system is simplified into two planar two-degree-of-freedom manipulators. The motion trajectories of each joint are converted into motion constraints and represented in matrix form. First-order and second-order constraints are obtained based on the matrix form of these motion trajectory constraints. By combining expected motion trajectory planning with first-order and second-order constraint equations, the computational complexity of motion planning is significantly reduced. The constraint equations ensure the stability and accuracy of the coordinated movement of the two arms, improving the real-time control performance of the system in complex tasks.

[0032] In this embodiment, the following assumptions are made regarding the above servo constraint equations: Assumption 1: For each Inertia matrix ,in In other words, only when the first assumption ensures the well-posedness and stability of the dynamic equations and the positive definiteness of the inertia matrix is ​​the study of the dynamics of the robotic arm system meaningful.

[0033] Assumption 2: For each of the first-order servo constraint equations , rank 1. All constraints are compatible. That is, it is guaranteed that at least one valid motion constraint and servo constraint exist, and the problem has a solution.

[0034] Assumption 3: For each , It is full rank, and It is invertible. That is, the full-rank condition guarantees the invertibility of the constraint matrix, providing a mathematical basis for subsequent control design.

[0035] Assumption 4: Under Assumption 3, for a given and ,make ,(twenty four) Assuming for all There exists a constant , making (25) This ensures that even with uncertainties, the controller... The fact that the device still functions effectively forms the basis for subsequent proofs of the robust controller's stability. Among these, This is the eleventh intermediate parameter.

[0036] Assumption 5: Assume the existence of an unknown vector of normal numbers. and known functions : ,for The following inequalities hold: (26) That is, to describe the upper bound of uncertainty in a fuzzy manner in order to construct This is more closely aligned with actual engineering practices. Among them, The uncertain part of the inverse inertial matrix. This represents the uncertain part of the Coriolis term. This represents the uncertain part of the gravity term. Based on the above assumptions, a robust controller for a humanoid robot dual-arm system is constructed, including: The system control matrix is ​​determined according to formulas (5) to (8). (5) (6) (7) (8) in, For nominal binding terms, For correction items, For robust feedback items, It is a positive definite matrix. These are the weighting coefficients. For constant control parameters, , The first intermediate parameter is the controller. item, It is a symmetrical term. The first control parameter, This is the second control parameter. For smooth switching functions and , This is the weighted constraint error vector. For fuzzy numbers with uncertainty, for The first derivative. A robust controller is designed based on the model, ensuring the consistent boundedness and consistent eventual boundedness of the constraint tracking error.

[0037] Existing Nash games involve simultaneous decision-making, failing to reflect the leader-follower relationship between parameters. This invention, however, employs the leader-follower decision sequence of a Stackelberg game, more effectively balancing transient and steady-state performance. In step S5, Stackelberg game theory is introduced, with key controller parameters treated as game participants. A cost function model for both sides is constructed, and the optimal parameter solution is obtained by solving the Stackelberg equilibrium. Specifically, as... Figure 3As shown, the following steps may be included: In step S51, a control parameter tuning model based on Stackelberg game is constructed. Specifically, this includes: The transient performance index of the system is determined according to formula (27). (27) The steady-state performance index of the system is determined according to formula (28). (28) in, , For servo constraint error, for time... The error energy accumulated by the system due to initial conditions, uncertainties, and other factors quantifies the degree to which the system deviates from the target constraint at that moment. for The largest eigenvalue, This represents the time elapsed since the controller started.

[0038] Because the cost function contains fuzzy numbers, it is necessary to handle the fuzzy parts of the system. One method for defuzzification is to perform D-operations on the steady-state and transient performance index functions of the system, thereby obtaining the cost functions for the leader and followers, respectively. Figure 4 As shown, it includes: In step S511, the leader's cost function is determined according to formula (9). (9) In step S512, the cost function of the follower is determined according to formula (10). (10) (29) (30) (31) (32) (33) (34) (35) (36) (37) in, For the leader's cost function, For system transient performance indicators, The cost function for followers, For the steady-state performance indicators of the system, for Operations, The first control parameter, This is the second control parameter. For fuzzy numbers with uncertainty, As the second intermediate parameter, As the third intermediate parameter, The fourth intermediate parameter, The fifth intermediate parameter, The sixth intermediate parameter, The seventh intermediate parameter, This is the eighth intermediate parameter. The ninth intermediate parameter, The tenth intermediate parameter and The leader's cost function is the cumulative error of the system from the initial moment to infinity, while the follower's cost function is the error level of the system at infinity.

[0039] In step S52, the first control parameter is determined as the leader of the game based on the robust controller. In the robust controller, Direct impact The control gain, and Influence The saturation region The control effect is far greater than Therefore, View it as the leader in the game. View them as followers in a game. Leaders Make the decision first, then the followers According to the leader Make a decision after that.

[0040] In step S53, the second control parameter is determined to be the follower in the game based on the robust controller.

[0041] In step S54, the leader and followers engage in a Stackelberg game based on the control parameter tuning model to optimize and tune the control parameters and obtain the optimal parameter solution.

[0042] In steps S51 to S54, a precise mathematical optimization model is established by constructing a cost function that includes transient and steady-state performance indices and using D-operations to handle fuzzy uncertainties. Secondly, the roles of leader and follower are reasonably allocated based on the degree of influence of parameters on system performance, making the parameter optimization process more in line with actual control logic. Finally, by solving the Stackelberg equilibrium solution, the optimal parameter solution can be obtained under the constraint of leader priority decision and follower optimal response, which not only ensures the robust stability of the control system, but also achieves the best balance between transient and steady-state performance.

[0043] In this embodiment, the specific steps for performing the Stackelberg game can be in various ways known to those skilled in the art. In a preferred example of the present invention, the specific steps for performing the Stackelberg game can be... Figure 5 The method shown. Figure 5 The control method also includes: In step S541, the leader selects parameters in the first decision set.

[0044] In step S542, the followers determine a strategy based on the parameters chosen by the leader to minimize the followers' cost function. If the leader chooses from the first decision set... Select one parameter To minimize the cost function Then the followers will move from the second decision set. Select a corresponding parameter

[0045] In step S543, the leader minimizes the leader's objective function according to the strategy to obtain the leader's optimal solution. Wherein, based on... Values ​​and related cost functions Leaders in the first decision set The optimal parameter solution is determined internally. The strategy is .

[0046] In step S544, the follower updates its policy based on the leader's optimal solution to obtain the follower's optimal solution in the second decision set. The follower updates its policy based on the leader's optimal parameter solution. Update its strategy. It changes from the second decision set. Select the best response .

[0047] Specifically, firstly, regarding about Taking the partial derivative, we get: (38) make ,get about Functions: . Next, Substitute into ,have to: (39) (40) make The optimal solution for the leader can be obtained. .

[0048] Finally, Substitution The optimal solution for the followers can be found. .

[0049] For the solution That is, it is the optimal solution to the Stackelberg game, and for any cost function... satisfy: , , That is .

[0050] In steps S541 to S544, the follower first obtains the optimal response function by solving the partial derivative based on the leader's current decision, thus establishing a dynamic coupling relationship between parameters. Secondly, the leader incorporates the follower's response function into its own optimization objective and obtains the optimal solution containing follower feedback by solving the derivative of the composite function, forming a closed loop of bidirectional optimization. The final Stackelberg equilibrium solution ensures that the cost functions of both parties can reach the minimum value when any parameter changes, enabling the control system to achieve the optimal balance between transient and steady-state performance, significantly improving the systematicness and coordination of parameter optimization.

[0051] Furthermore, the sufficiency of the obtained optimal solution needs to be verified. Specifically, this includes the following steps: In step S545, the sufficient conditions for the optimal solution are verified.

[0052] In step S546, if the verification passes, the optimal parameter solution is output.

[0053] In step S547, if the verification fails, the Stackelberg game is played again to optimize and tune the control parameters to obtain the optimal parameter solution.

[0054] Specifically, firstly, the cost function about Find the second-order partial derivative, and... Substituting into it, for the cost function about Find the second-order partial derivative, and... Substitute it into the context: (41) (42) Determine if the two second-order partial derivatives are greater than 0; if so, then confirm. The solution is the Stackelberg optimal solution, i.e., Stackelberg equilibrium is achieved; otherwise, the system reverts to the robust controller and the leader is readjusted. and followers The value is then verified again.

[0055] Figure 6 This is an error comparison diagram of one embodiment of the present invention, namely, an error comparison diagram between the actual motion trajectory of the joint and the expected motion trajectory. error1 to error4 represent the errors of the first joint of the first robotic arm, the second joint of the first robotic arm, the first joint of the second robotic arm, and the second joint of the second robotic arm, respectively. Figure 6 The difference between the left and right images is because it's a comparison of errors at different joints. Different joints have different dynamic models, and the initial errors for joint 1 of the first robotic arm and joint 1 of the second robotic arm are set to 0.1; the initial errors for joint 2 of the first robotic arm and joint 2 of the second robotic arm are set to 0.2. These different initial errors lead to different convergence rates, resulting in differences in the images. Figure 6 As shown, based on a simplified dual-arm model, different control methods, SMC (Sliding Mode Control) and LQR (Linear Quadratic Regulator), were used and plotted in the Matlab environment. Overall, LQR converged to 0 the slowest and had a larger error; while SMC had a better convergence speed, it exhibited some chattering during the convergence process. In the enlarged local plot, a simulation time of approximately 8-13 seconds was selected, during which the controller was more stable and had a certain duration, making it more meaningful for comparison. The red line represents the method proposed in this invention. It can be seen that, at the same order of magnitude, the red line always closely follows the 0 line, while the blue line (SMC) shows differences at joint 1 of the first robotic arm and joint 1 of the second robotic arm. , The error; there are at joint 2 of the first robotic arm and joint 2 of the second robotic arm respectively. , The error; the green line (LQR) has at joint 1 of the first robotic arm and joint 1 of the second robotic arm respectively. , The error; at joint 2 of the first robotic arm and joint 2 of the second robotic arm, there are respectively more than , The data shows that the Stackelberg game-based robust control method has superior performance compared to traditional control methods such as SMC and LQR, with greater stability, accuracy, and lower control cost. Compared to linear quadratic regulators (LQR) and sliding mode control (SMC), the Stackelberg game-based robust control method produces smaller joint errors.

[0056] Figure 7 This is a control input comparison diagram of one embodiment of the present invention. , , , These represent the control torques of the first robotic arm joint 1, the first robotic arm joint 2, the second robotic arm joint 1, and the second robotic arm joint 2, respectively. It can be seen that initially, the SMC exhibits severe chattering, with amplitudes reaching 100-120 Nm. While the LQR exhibits less severe chattering, its control torque is still larger than that of the robust control method based on Stackelberg game theory at the beginning. and More obviously, the values ​​are approximately 40 Nm and 10 Nm, respectively, while the robust control method based on Stackelberg game theory is approximately 3 Nm and 4 Nm. Figure 7 The large difference between the top and bottom is actually due to the effect of uncertainty. The uncertainty of the mass of the first robotic arm can be written as: The uncertainty in the mass of the second robotic arm can be written as: The first robotic arm exhibits greater uncertainty, which is reflected in the graph as a larger amplitude of control torque jitter. Compared to LQR and SMC, the robust control method based on Stackelberg game theory produces a lower and more stable control input at the beginning, without excessive jitter. The robust controller optimized by Stackelberg game theory demonstrates superior performance in trajectory tracking of humanoid robot dual-arm systems.

[0057] On the other hand, the present invention also provides a robust control system for humanoid robot servo constraints based on Stackelberg game theory, the system including a processor for executing any of the control methods described above.

[0058] Through the above technical solution, this invention provides a robust control method and system for humanoid robots based on Stackelberg game theory. The trajectory tracking problem of the humanoid robot's dual-arm joints is transformed into a servo constraint problem. Servo constraint equations for the humanoid robot's dual-arm system are established, and considering the system's uncertainties, a dynamic model incorporating uncertainty terms is constructed. Based on this model, a robust controller is designed to ensure the consistent boundedness and consistent eventual boundedness of the constraint tracking error. Subsequently, Stackelberg game theory is introduced, with key controller parameters as game participants. A cost function model for both parties is constructed, and the optimal parameter solution is obtained by solving the Stackelberg equilibrium. This invention achieves optimized tuning of controller parameters through Stackelberg game theory, enabling the system to possess high-precision trajectory tracking capability, strong robustness, and stable control input under uncertain environments. It should be noted that the Stackelberg game in this invention is only used for optimized tuning of control parameters in the robot's robust controller under engineering constraints. Its object of application is the physical control parameters of the robot's servo control system, and it does not involve abstract information processing or decision-making algorithms.

[0059] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0060] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0061] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0062] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0063] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0064] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0065] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0066] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0067] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A robust control method for servo constraints of a humanoid robot based on Stackelberg game theory, characterized in that, The control method includes: Constructing a humanoid robot dual-arm system; Based on the described humanoid robot dual-arm system, a dynamic model of the humanoid robot with uncertainties is constructed. Determine the servo constraint equations of the dynamic model; Build robust controllers; A control parameter tuning model based on Stackelberg game is constructed, and the key control parameters in the robust controller are optimized under engineering constraints according to the control parameter tuning model to obtain the optimal parameter solution. The optimal parameter solution is input into the humanoid robot dual-arm system to obtain optimal performance.

2. The control method according to claim 1, characterized in that, Based on the aforementioned humanoid robot dual-arm system, a dynamic model of the humanoid robot with uncertainties is constructed, including: Construct a dynamic model of the humanoid robot based on formula (1). ,(1) in, Let be the system's inertia matrix. For the Coriolis centrifugal force of the system, For the gravity term of the system, For system control torque, For the uncertainty parameters of the system, For joint angle, The joint angular velocity, Joint angular acceleration, For time, These are the uncertainty parameters of the system.

3. The control method according to claim 2, characterized in that, Determining the servo constraint equations of the dynamic model includes: The humanoid robot dual-arm system is simplified into a system of two planar two-degree-of-freedom robotic arms; The desired motion trajectories of the two planar two-degree-of-freedom robotic arms are determined according to formula (2). ,(2) The first-order constraints are obtained according to formula (3). ,(3) The second-order constraints are obtained according to formula (4). ,(4) in, For the desired motion trajectory, For the constraint matrix, For the angle of the joints in the system, Let be the angular velocity of the joints in the system. Let be the angular acceleration of the joints in the system. For the expected first-order constraint and , For the expected second-order constraint and , These are the dimensions of the expected first-order constraints and the dimensions of the expected second-order constraints, respectively. .

4. The control method according to claim 3, characterized in that, Building a robust controller includes: The system control matrix is ​​determined according to formulas (5) to (8). ,(5) ,(6) ,(7) ,(8) in, For nominal binding terms, For correction items, For robust feedback items, It is a positive definite matrix. These are the weighting coefficients. For constant control parameters, , As the first intermediate parameter, It is a symmetrical term. The first control parameter, For smooth switching functions, This is the weighted constraint error vector. For the Coriolis centrifugal force of the system, This is the gravity term of the system.

5. The control method according to claim 1, characterized in that, A control parameter tuning model based on Stackelberg game is constructed, and the key control parameters in the robust controller are optimized under engineering constraints according to the control parameter tuning model to obtain the optimal parameter solution, including: Construct a control parameter tuning model based on Stackelberg game theory; The first control parameter is determined as the leader of the game based on the robust controller; The second control parameter is determined to be the follower in the game based on the robust controller; The leader and followers engage in Stackelberg game based on the control parameter tuning model to optimize and tune the control parameters and obtain the optimal parameter solution.

6. The control method according to claim 5, characterized in that, Construct a control parameter tuning model based on Stackelberg game, including: The leader's cost function is determined according to formula (9). ,(9) The cost function of the followers is determined according to formula (10). ,(10) in, For the leader's cost function, For system transient performance indicators, The cost function for followers, For the steady-state performance indicators of the system, for Operations, The first control parameter, This is the second control parameter. It is a fuzzy number with uncertainty.

7. The control method according to claim 6, characterized in that, The leader and followers perform a Stackelberg game based on the control parameter tuning model to optimize and tune the control parameters and obtain the optimal parameter solution, including: The leader selects parameters in the first decision set; Followers determine their strategy based on the parameters chosen by the leader in order to minimize the follower's cost function; The leader minimizes the leader's objective function according to the strategy to obtain the leader's optimal solution; The followers update their strategy based on the leader's optimal solution to obtain the followers' optimal solution in the second decision set.

8. The control method according to claim 7, characterized in that, The leader and followers perform a Stackelberg game based on the control parameter tuning model to optimize and tune the control parameters and obtain the optimal parameter solution, which also includes: The optimal solution is then verified under sufficient conditions. If the verification passes, the optimal parameter solution will be output; If the verification fails, the Stackelberg game is repeated to optimize and tune the control parameters to obtain the optimal parameter solution.

9. A robust control system for servo constraints of a humanoid robot based on Stackelberg game theory, characterized in that, The system includes a processor for executing the control method as described in any one of claims 1 to 8.