A robot decision planning method and device based on a three-level player master-slave game

By constructing a three-level player master-slave game decision-making model and performing implicit differential calculations, the problem that existing technologies cannot fully characterize hierarchical structures in robot decision-making and planning is solved, achieving more rational and safer decision-making and planning results.

CN120287283BActive Publication Date: 2026-07-31TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TONGJI UNIV
Filing Date
2025-02-21
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing robot decision-making and planning methods cannot fully characterize the overall hierarchical structure of a robot system when simulating its hierarchical structure, resulting in poor decision-making performance.

Method used

We adopt a decision planning method for a three-level player master-slave game to construct a three-level player master-slave game decision model. We calculate the best response of lower-level players through implicit differentiation, and obtain the total differential of their own cost function by combining the policy mapping relationship. We then use gradient descent to obtain new policies.

Benefits of technology

It improves the effectiveness and safety of robot decision-making and planning by considering the hierarchical structure and optimal response of the robot system, thereby enhancing the rationality and safety of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120287283B_ABST
    Figure CN120287283B_ABST
Patent Text Reader

Abstract

This invention provides a robot decision-making and planning method and apparatus based on a three-level player master-slave game. The method includes: constructing a three-level player master-slave game decision model corresponding to the robot system based on the robot's decision-making and planning scenario; in three-level Stackelberg dynamic learning for the three-level player master-slave game decision model, higher-level players calculate the partial derivative of the lower-level players' optimal response based on implicit differentials, and, combined with the mapping relationship between the lower-level players' policies and their own policies, obtain the total differential of their own cost function; for all players, the total differential of the cost function given the higher-level player's policy is called, and a new policy is obtained through gradient descent; then, robot decision-making and planning are performed on the robot system based on this. In this way, the robot's decision-making and planning performance can be improved based on the three-level player master-slave game decision model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotics, and in particular to a robot decision-making and planning method and apparatus based on a three-level player master-slave game. Background Technology

[0002] In robotic decision-making and planning scenarios, game theory models are often used to guide decision-making and deployment. For example, the chase-escape game can be applied to robot path planning. Current research often seeks the Nash equilibrium of the game to update strategies and uses gradient descent optimization algorithms to find the Nash equilibrium. However, it is important to understand that in robotic decision-making and planning scenarios, participants also have differences in the order of their actions and their dominance, i.e., a hierarchical structure exists. In this case, Nash equilibrium and related algorithms cannot simulate the differences within the players. To address this, some researchers have used master-slave game theory to guide decision-making and deployment in robotic decision-making and planning scenarios. The stable state of the master-slave game is defined as the Stackelberg equilibrium.

[0003] Currently, in master-slave game theory, the player with higher dominance and first-mover advantage is called the leader, while the relatively weaker player is called the follower. However, due to the often complex and diverse internal structures of robotic systems, existing two-level master-slave game theory models cannot fully characterize the overall hierarchical structure of the system. Therefore, current robot decision-making and planning based on master-slave game theory suffers from poor performance. Summary of the Invention

[0004] In a first aspect, embodiments of the present invention provide a robot decision-making and planning method based on a three-level player master-slave game, the method comprising: Based on the robot decision-making and planning scenarios faced by the robot system, a three-level player master-slave game decision-making model is constructed for the robot system. In the three-level player master-slave game decision-making model, there are three levels of players. The highest level player is the senior leader who acts first and has a dominant position. The middle level players are the intermediate followers who make the best response after observing the actions of the senior leader. The lowest level players are the low-level followers who act last and make the best response after observing the actions of the senior leader and the intermediate followers. In the three-level Stackelberg dynamic learning for a three-level player master-slave game decision model, higher-level players calculate the partial derivative of the best response of lower-level players based on implicit differentials. Combining the mapping relationship between the lower-level players' policies and their own policies, they obtain the total differential of their own cost function. They then call the total differential of the cost function of each player after a given higher-level player's policy, and obtain a new policy through gradient descent. Based on this, they perform robot decision planning for the robot system.

[0005] Among some feasible approaches to the first aspect, a three-level player master-slave game decision-making model is constructed based on the robot system's decision-making and planning scenario, including: For robot decision-making and planning scenarios faced by robot systems, a three-level player master-slave game decision-making model is constructed based on the order of action and dominant position of different robots in the robot system, and combined with the cost function of each robot.

[0006] In some possible implementations of the first aspect, the three-level player master-slave game decision-making model corresponding to the robot system can be represented by the following formula: ; in, The robots involved in the robot decision-making and planning scenario are collectively referred to as players. This indicates the corresponding player strategy; This indicates the corresponding player strategy dimension; This represents the corresponding set of strategies; Indicates policy set constraints; For the player cost function; They represent senior leaders, mid-level followers, and junior followers, respectively. .

[0007] In some feasible ways of implementing the first aspect, in a three-level player master-slave game decision-making model, the lower-level follower... For senior leaders Intermediate Follower The best response to the action is represented as: ; Intermediate Follower For senior leaders The best response to the action is represented as: ; Based on this, the local Stackelberg equilibrium of the three-level player master-slave game decision model is defined as: If senior leader strategy satisfy: ; Then the strategy set Stackelberg equilibrium of a three-level player master-slave game decision-making model.

[0008] In some possible implementations of the first aspect, in the three-level Stackelberg dynamic learning for a three-level player master-slave game decision model, the cost function is first... right The partial derivative is denoted as ; Cost function right as well as The second-order partial derivatives obtained by successively calculating the gradients are denoted as follows: ,when When, the second-order partial derivative is denoted as The total differential is denoted as .

[0009] The gradient call vector for the three-level Stackelberg dynamic learning is denoted as: ; For each player, the total differential of the cost function after applying a higher-level player's policy is called. This total differential is the sum of the products of the partial derivatives of each unknown variable and the gradients of the variables themselves. Recorded as , Recorded as And remember right The partial derivative is From this, we can obtain for: ; Based on this, the solution is obtained using implicit differential computation. , , .

[0010] Secondly, embodiments of the present invention provide a robot decision-making and planning device based on a three-level player master-slave game, the device comprising: The module is used to construct a three-level player master-slave game decision-making model for the robot system based on the robot decision-making and planning scenarios it faces. In the three-level player master-slave game decision-making model, there are three levels of players: the highest level player is the senior leader who acts first and has a dominant position; the middle level player is the intermediate follower who makes the best response after observing the actions of the senior leader; and the lowest level player is the low-level follower who acts last and makes the best response after observing the actions of the senior leader and the intermediate follower. The planning module is used in the three-level Stackelberg dynamic learning of a three-level player master-slave game decision model. Higher-level players calculate the partial derivative of the best response of lower-level players based on implicit differentials. Combining the mapping relationship between the lower-level players' policies and their own policies, they obtain the total differential of their own cost function. They then call the total differential of the cost function of each player after a higher-level player's policy is given, and obtain a new policy through gradient descent. Based on this, the robot system performs robot decision planning.

[0011] Thirdly, embodiments of the present invention provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0012] Fourthly, embodiments of the present invention provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method described above.

[0013] Compared with the prior art, the present invention has at least the following technical effects: 1. This invention considers the hierarchical structure of the robot system, and by taking into account both the robot's motivation and optimal response, it obtains more rational and safer decisions. Furthermore, by considering the internal interactions of each participant during gradient invocation, it improves the efficiency of decision learning and ultimately enhances the robot's decision planning performance.

[0014] 2. For robot decision-making and planning scenarios, this invention improves the rationality of gradient invocation during dynamic learning of decision-making by taking into account the robot's best response information, which can effectively improve the safety and reliability of participants' decisions in robot decision-making and planning.

[0015] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0016] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A flowchart of a robot decision planning method based on a three-level player master-slave game provided in an embodiment of the present invention; Figure 2 This invention provides a synchronous dynamic learning three-player trajectory graph. Figure 3 This invention provides a three-level Stackelberg dynamic learning three-player trajectory graph. Figure 4 This is a schematic diagram illustrating the target-defender team utility under different dynamic learning scenarios provided in embodiments of the present invention. Figure 5 A structural diagram of a robot decision-making and planning device based on a three-level player master-slave game provided in an embodiment of the present invention; Figure 6 This is a structural diagram of an exemplary electronic device capable of implementing embodiments of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Furthermore, the term "and / or" in this invention is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this invention generally indicates that the preceding and following related objects have an "or" relationship.

[0019] To address the technical problems encountered in the background art, embodiments of the present invention provide a robot decision-making and planning method, apparatus, device, and storage medium based on a three-level player master-slave game. The following, in conjunction with the accompanying drawings, provides a detailed description of the robot decision-making and planning method, apparatus, device, and storage medium based on a three-level player master-slave game provided by the present invention through specific embodiments.

[0020] Figure 1 A flowchart of a robot decision-making and planning method based on a three-level player master-slave game is provided for an embodiment of the present invention, as follows: Figure 1 As shown, the robot decision-making and planning method 100 may include: S110. Based on the robot decision-making and planning scenarios faced by the robot system, a three-level player master-slave game decision-making model is constructed for the robot system. In the three-level player master-slave game decision-making model, there are three levels of players. The highest level player is the senior leader who acts first and has a dominant position. The middle level player is the intermediate follower who makes the best response after observing the senior leader's actions. The lowest level player is the low-level follower who acts last and makes the best response after observing the actions of the senior leader and the intermediate follower.

[0021] In S120, in the three-level Stackelberg dynamic learning for a three-level player master-slave game decision model, higher-level players calculate the partial derivative of the best response of lower-level players based on implicit differentials. Combining the mapping relationship between the lower-level players' policies and their own policies, they obtain the total differential of their own cost function. They then call the total differential of the cost function of each player after a higher-level player's policy is given, and obtain a new policy through gradient descent. Based on this, they perform robot decision planning for the robot system.

[0022] To facilitate further understanding, the above steps will be described in detail below with reference to specific embodiments: a) Construction of a three-level player master-slave game decision-making model For robot decision-making and planning scenarios, based on the order of actions and dominance of different robots in the system, and combined with the cost function of each robot, a three-level player master-slave game decision-making model is constructed for the robot system. This model can be expressed using the following formula: (1) in, The robots involved in the robot decision-making and planning scenario are collectively referred to as players. This indicates the corresponding player strategy; This indicates the corresponding player strategy dimension; This represents the corresponding set of strategies; Indicates policy set constraints; For the player cost function; They represent senior leaders, mid-level followers, and junior followers, respectively. .

[0023] In a three-level player master-follower game decision-making model, the lower-level follower... For senior leaders Intermediate Follower The best response to the action is represented as: (2) Intermediate Follower For senior leaders The best response to the action is represented as: (3) Based on this, the local Stackelberg equilibrium of the three-level player master-slave game decision model is defined as: If senior leader strategy satisfy: (4) Then the strategy set Stackelberg equilibrium of a three-level player master-slave game decision-making model.

[0024] Furthermore, certain assumptions need to be set for the three-player master-slave game decision model to ensure the existence of Stackelberg equilibrium: First, the cost functions of all players are q-th order continuously differentiable, i.e. First, q is greater than 2; second, all optimal response functions are continuously differentiable; and third, player policy constraints are all non-empty convex compact sets.

[0025] b) Three-level Stackelberg dynamic learning First, the cost function right The partial derivative is denoted as ; Cost function right as well as The second-order partial derivatives obtained by successively calculating the gradients are denoted as follows: ,when When, the second-order partial derivative is denoted as The total differential is denoted as .

[0026] The gradient call vector for the three-level Stackelberg dynamic learning is denoted as: (5) Unlike synchronous gradient descent in dynamic learning, where each player directly calls the cost function to apply the partial derivative of their own policy,... Following the strategy update approach, in three-level Stackelberg dynamic learning, players need to consider the optimal responses and behavioral motivations of lower-level players to establish an information advantage, gaining a more rational and stable strategy by understanding both themselves and their opponents. Therefore, given the determined behavior of higher-level players, each player will invoke the total differential of their own strategy and the strategies of lower-level players. The total differential consists of the sum of the partial derivatives of each unknown variable and the product of the gradients of the variables themselves. Recorded as , Recorded as And remember right The partial derivative is From this, we can obtain for: (6) Based on this, the solution is obtained using implicit differential computation. , , Specifically, based on the Stackelberg equilibrium definition and equations (2) and (3), we can obtain: (7) Based on the above, using implicit differential equations, the specific expressions for each optimal response gradient can be obtained by solving equations and performing matrix operations. For example, for Seeking the correct answer Taking the derivative and setting it to 0, we get... .

[0027] Therefore, the above three-level Stackelberg dynamic learning can be summarized into the algorithm described in Table 1: Table 1

[0028] Next, for the robot decision-making and planning task, the initial state of each robot is initialized, and the optimization objective function is then calculated. The robot then follows the policy update rules in the algorithm to guide its real-time actions in order to achieve Stackelberg equilibrium.

[0029] As an example, this section uses a three-robot path planning scenario to illustrate the robot decision-making and planning method based on a three-level player master-slave game, as shown below: Consider a three-robot path planning scenario, modeled as a three-player chase-escape zero-sum game. The three robots are divided into a target, a defender, and an attacker. The attacker pursues the target, while the defender and target form a team, with the defender intercepting and interfering with the attacker. The strategies of the target, attacker, and defender are all two-dimensional vectors, representing the horizontal and vertical coordinates in space, respectively denoted as... The cost functions for the three players are as follows: ,in The expression is: (8) In this simulation, all player strategies fall within the range of [-35, 35], and each player aims to minimize their respective cost function. The attacker is allowed to capture the target when the distance between them is less than 5, and the defender is allowed to intercept and interfere with the attacker when the distance between them is less than 15. In the simulation, this is manifested as a decay in the learning rate of the attacker's strategy update.

[0030] The three-player chase-escape zero-sum game was simulated using both synchronous dynamic learning (directly calling the self-cost function to differentiate its own policy) and the three-level Stackelberg dynamic learning of this invention. The resulting trajectory maps for each player are shown below. Figure 2-3 As shown.

[0031] Depend on Figure 2 It can be seen that when players update their strategies using synchronous dynamic learning, the attacker is able to capture the target in the 15th iteration, meaning the target fails to escape; while... Figure 3It can be seen that when players' strategy updates use three-level Stackelberg dynamic learning, the target successfully escapes within 200 iterations, and the defender is able to intercept and interfere with the attacker in the 40th iteration, slowing down its learning rate. The utility (cost function value) of the target-defender team within 200 iterations under different dynamic learning methods is as follows: Figure 4 As shown, in synchronous dynamic learning, the policy is no longer updated once the target is captured.

[0032] Depend on Figure 4 It can be seen that using three-level Stackelberg dynamic learning can cause the target-defender team utility to converge to a lower value (the final utility of synchronous dynamic learning is 263.3689, and the final utility of three-level Stackelberg dynamic learning is -988.5609). This result demonstrates the effectiveness of three-level Stackelberg dynamic learning for decision planning.

[0033] In summary, the present invention achieves at least the following technical effects: 1. This invention considers the hierarchical structure of the robot system, and by taking into account both the robot's motivation and optimal response, it obtains more rational and safer decisions. Furthermore, by considering the internal interactions of each participant during gradient invocation, it improves the efficiency of decision learning and ultimately enhances the robot's decision planning performance.

[0034] 2. For robot decision-making and planning scenarios, this invention improves the rationality of gradient invocation during dynamic learning of decision-making by taking into account the robot's best response information, which can effectively improve the safety and reliability of participants' decisions in robot decision-making and planning.

[0035] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0036] The above is an introduction to the method embodiments. The following describes the solution of the present invention further through device embodiments.

[0037] Figure 5 A structural diagram of a robot decision-making and planning device based on a three-level player master-slave game, provided in an embodiment of the present invention, is shown below. Figure 5 As shown, the robot decision-making and planning device 500 may include: Module 510 is used to construct a three-level player master-slave game decision-making model for the robot system based on the robot decision-making and planning scenarios it faces. In the three-level player master-slave game decision-making model, there are three levels of players: the highest level player is the senior leader who acts first and has a dominant position; the middle level player is the intermediate follower who makes the best response after observing the senior leader's actions; and the lowest level player is the low-level follower who acts last and makes the best response after observing the actions of the senior leader and the intermediate follower.

[0038] The planning module 520 is used in the three-level Stackelberg equilibrium dynamic learning of the three-level player master-slave game decision model. The higher-level player calculates the partial derivative of the lower-level player's best response based on implicit differential, and obtains the total differential of its own cost function by combining the mapping relationship between the lower-level player's policy and its own policy. The total differential of the cost function of each player after the higher-level player's policy is called, and then a new policy is obtained through gradient descent. Then, the robot system is used for robot decision planning based on this.

[0039] Understandable Figure 5 Each module / unit in the robot decision-making and planning device 500 shown has the ability to implement Figure 1 The functions of each step in the robot decision-making and planning method 100 shown, and the corresponding technical effects they achieve, will not be elaborated here for the sake of brevity.

[0040] Figure 6 This is a structural diagram of an exemplary electronic device capable of implementing embodiments of the present invention. Electronic device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 600 may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown in this invention, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0041] like Figure 6As shown, the electronic device 600 may include a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0042] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0043] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer program product, including a computer program tangibly contained in a computer-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform method 100 by any other suitable means (e.g., by means of firmware).

[0044] The various embodiments described above in this invention can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), payload programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0045] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0046] In the context of this invention, a computer-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of computer-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0047] It should be noted that the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute method 100 and achieve the corresponding technical effects achieved by the embodiments of the present invention in executing the method. For the sake of brevity, they will not be described in detail here.

[0048] In addition, the present invention also provides a computer program product, which includes a computer program that implements method 100 when executed by a processor.

[0049] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this invention does not impose any limitations on them.

[0050] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A robot decision-making and planning method based on a three-level player master-slave game, characterized in that, The method includes: Based on the robot decision-making and planning scenarios faced by the robot system, a three-level player master-slave game decision-making model is constructed for the robot system. In the three-level player master-slave game decision-making model, there are three levels of players. The highest level player is the senior leader who acts first and has a dominant position. The middle level players are the intermediate followers who make the best response after observing the actions of the senior leader. The lowest level players are the low-level followers who act last and make the best response after observing the actions of the senior leader and the intermediate followers. In the three-level Stackelberg dynamic learning for a three-level player master-slave game decision model, higher-level players calculate the partial derivative of the best response of lower-level players based on implicit differentials. Combining the mapping relationship between the lower-level players' policies and their own policies, they obtain the total differential of their own cost function. They then call the total differential of the cost function of each player after a higher-level player's policy is given, and obtain a new policy through gradient descent. Based on this, they perform robot decision planning on the robot system. The process of constructing a three-level player master-slave game decision-making model for the robot system based on the robot decision-making and planning scenarios it faces includes: For robot decision-making and planning scenarios faced by robot systems, a three-level player master-slave game decision-making model is constructed according to the order of action and dominant position of different robots in the robot system, and combined with the cost function of each robot. The three-level player master-slave game decision-making model corresponding to the robot system can be expressed by the following formula: ; in, The robots involved in the robot decision-making and planning scenario are collectively referred to as players. This indicates the corresponding player strategy; This indicates the corresponding player strategy dimension; This represents the corresponding set of strategies; Indicates policy set constraints; For the player cost function; They represent senior leaders, mid-level followers, and junior followers, respectively. ; In a three-level player master-follower game decision-making model, the lower-level follower... For senior leaders Intermediate Follower The best response to the action is represented as: ; Intermediate Follower For senior leaders The best response to the action is represented as: ; Based on this, the local Stackelberg equilibrium of the three-level player master-slave game decision model is defined as: If senior leader strategy satisfy: ; Then the strategy set The Stackelberg equilibrium that constitutes a three-level player master-slave game decision-making model; In the three-level Stackelberg dynamic learning for a three-level player master-slave game decision model, the cost function is first... right The partial derivative is denoted as ; Cost function right as well as The second-order partial derivatives obtained by successively calculating the gradients are denoted as follows: ,when When, the second-order partial derivative is denoted as The total differential is denoted as ; The gradient call vector for the three-level Stackelberg dynamic learning is denoted as: ; For each player, the total differential of the cost function after applying the higher-level player's policy is called. This total differential is the sum of the products of the partial derivatives of each unknown variable and the gradients of the variables themselves. Recorded as , Recorded as And remember right The partial derivative is From this, we can obtain for: ; Based on this, the solution is obtained using implicit differential computation. , , .

2. A robot decision-making and planning device based on a three-level player master-slave game, characterized in that, The apparatus is used to perform the method of claim 1, comprising: The module is used to construct a three-level player master-slave game decision-making model for the robot system based on the robot decision-making and planning scenarios it faces. In the three-level player master-slave game decision-making model, there are three levels of players: the highest level player is the senior leader who acts first and has a dominant position; the middle level player is the intermediate follower who makes the best response after observing the actions of the senior leader; and the lowest level player is the low-level follower who acts last and makes the best response after observing the actions of the senior leader and the intermediate follower. The planning module is used in the three-level Stackelberg dynamic learning of a three-level player master-slave game decision model. Higher-level players calculate the partial derivative of the best response of lower-level players based on implicit differentials. Combining the mapping relationship between the lower-level players' policies and their own policies, they obtain the total differential of their own cost function. They then call the total differential of the cost function of each player after a higher-level player's policy is given, and obtain a new policy through gradient descent. Based on this, the robot system performs robot decision planning.

3. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of claim 1.

4. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method of claim 1.