Robot decision planning method and device based on three-level player master-slave game
By constructing a three-level player master-slave game decision model and estimating the optimal response of low-level players using implicit differentials, the problem of poor robot decision planning in the existing technology is solved, and a more efficient and safe decision planning effect is achieved.
Patent Information
- Application Number
- CN202510195188.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing robot decision planning methods are not effective when simulating the internal hierarchy of the robot system, and cannot fully characterize the overall hierarchy of the system, resulting in unsatisfactory decision planning results.
The three-level player master-slave game method is adopted to build a three-level player master-slave game decision model of the robot system. The best response of low-level players is estimated through implicit differentialization, combined with the strategy mapping relationship, the full differentialization of its own cost function is obtained, and a new strategy is obtained for decision planning using gradient descent.
提高了机器人决策规划的效率和安全性,通过考虑机器人动机与最佳响应,实现更为理性和安全的决策,提升了决策学习的合理性和可靠性。
Smart Images

Figure CN120287283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and in particular, to a robot decision-making and planning method and device based on a three-level player master-slave game. Background Art
[0002] In robot decision-making and planning scenarios, game models are often used to guide decision-making deployments. For example, pursuit-evasion games can be applied to robot path planning. Currently, existing research often seeks the Nash equilibrium of the game to update strategies and uses gradient descent optimization algorithms to find the Nash equilibrium. However, it should be understood that in robot decision-making and planning scenarios, there are also differences in the order of actions and dominant positions among participants, that is, there is a hierarchical structure. At this time, the Nash equilibrium and its related algorithms cannot simulate the internal differences of players. In response to this, some researchers have used master-slave games to guide decision-making deployments in robot decision-making and planning scenarios, and the stable state of the master-slave game is defined as the Stackelberg equilibrium.
[0003] Currently, in the master-slave game, the player with a higher dominant position and a first-mover advantage in actions is called the leader, while the relatively disadvantaged player is called the follower. However, due to the often complex and diverse internal structures of robot systems, the existing two-level player master-slave game cannot fully characterize the overall hierarchical structure of the system. In view of this, there are currently defects in the robot decision-making and planning based on the master-slave game, resulting in poor effects. Summary of the Invention
[0004] In a first aspect, an embodiment of the present invention provides a robot decision-making and planning method based on a three-level player master-slave game. The method includes:
[0005] According to the robot decision-making and planning scenario faced by the robot system, a three-level player master-slave game decision model corresponding to the robot system is constructed. In the three-level player master-slave game decision model, there are three levels of players. The highest-level player is a senior leader who acts first and is in a dominant position. The middle-level player is an intermediate follower who makes the best response after observing the actions of the senior leader. The lowest-level player is a junior follower who acts last and makes the best response after observing the actions of the senior leader and the intermediate follower.
[0006] In the three-level Stackelberg dynamic learning for the three-level player master-slave game decision model, the higher-level player estimates the total derivative or partial derivative of the best response of the lower-level player based on implicit differentiation, combines the mapping relationship between the strategies of the lower-level player and its own strategy, obtains the total differential of its own cost function, calls the total differential of the cost function for all players after giving the strategy of the higher-level player, and then obtains a new strategy through gradient descent. After that, robot decision-making and planning for the robot system is carried out based on this.
[0007] In some realizable ways of the first aspect, according to the robot decision-making and planning scenarios faced by the robot system, a three-level player leader-follower game decision model corresponding to the robot system is constructed, including:
[0008] For the robot decision-making and planning scenarios faced by the robot system, according to the action sequence and dominant position of different robots in the robot system, and in combination with the goals or loss functions of each robot, a three-level player leader-follower game decision model corresponding to the robot system is constructed.
[0009] In some realizable ways of the first aspect, the three-level player leader-follower game decision model corresponding to the robot system can be expressed by the following formula:
[0010]
[0011] Among them, X i represents the robots corresponding to the robot decision-making and planning scenarios, collectively referred to as players; x i represents the corresponding player strategies; n i represents the corresponding player strategy dimension; represents the corresponding strategy set; U i represents the strategy set constraint; f i is the player cost function; X1, X2, X3 represent the high-level leader, middle-level follower, and low-level follower respectively; i = 1, 2, 3.
[0012] In some realizable ways of the first aspect, in the three-level player leader-follower game decision model, the best response of the low-level follower X3 to the actions of the high-level leader X1 and the middle-level follower X2 is expressed as:
[0013]
[0014] The best response of the middle-level follower X2 to the action of the high-level leader X1 is expressed as:
[0015]
[0016] Based on this, the local Stackelberg equilibrium of the three-level player leader-follower game decision model is defined as:
[0017] If the high-level leader strategy satisfies:
[0018]
[0019] Then the strategy set constitutes the Stackelberg equilibrium of the three-level player leader-follower game decision model.
[0020] In some realizable ways of the first aspect, in the three-level Stackelberg dynamic learning for the three-level player master-slave game decision model, first, the cost function f i (x1, x2, x3) (i = 1, 2, 3) with respect to x j (j = 1, 2, 3) of the partial derivative is denoted as The cost function f i (x1, x2, x3) (i = 1, 2, 3) with respect to x j (j = 1, 2, 3) and x k (k = 1, 2, 3) in turn to obtain the second-order partial derivative of the gradient is denoted as When j = k, the second-order partial derivative is denoted as The total differential is denoted as D(·).
[0021] The gradient call vector of the three-level Stackelberg dynamic learning is denoted as:
[0022] ω S (x) = [ω S,1 (x), ω S,2 (x), ω S,3 (x)];
[0023] For all players, the total differential of the cost function after calling the strategy of the higher-level player is called respectively. The total differential is composed of the sum of the partial derivatives of each unknown variable and the product of the variable's own gradient. Denote as Denote as And denote The partial derivative with respect to x2 is Thus, ω S (x) can be obtained as:
[0024]
[0025] On this basis, based on the implicit differential estimation, solve D mr , D br ,
[0026] In the second aspect, an embodiment of the present invention provides a robot decision-making and planning device based on a three-level player master-slave game. The device includes:
[0027] A construction module is used to construct a three - level player - based master - slave game decision model corresponding to a robot system according to the robot decision - making and planning scenarios faced by the robot system. In the three - level player - based master - slave game decision model, there are three levels of players. The highest - level player is a senior leader who acts first and is dominant. The middle - level player is an intermediate follower who makes the best response after observing the actions of the senior leader. The lowest - level player is a junior follower who makes the best response after observing the actions of the senior leader and the intermediate follower.
[0028] A planning module is used in the three - level Stackelberg dynamic learning for the three - level player - based master - slave game decision model. A higher - level player estimates the total derivative or partial derivative of the best response of a lower - level player based on implicit differentiation, combines the mapping relationship between the strategies of the lower - level player and its own strategy, obtains the total differential of its own cost function, calls the total differential of the cost function for all players respectively after a given higher - level player's strategy, and then obtains a new strategy through gradient descent. Based on this, robot decision - making and planning for the robot system are carried out.
[0029] Thirdly, an embodiment of the present invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0030] Fourthly, an embodiment of the present invention provides a non - transitory computer - readable storage medium storing computer instructions for causing a computer to execute the method as described above.
[0031] Compared with the prior art, the present invention has at least the following technical effects:
[0032] 1. The present invention takes into account the hierarchical structure in the robot system. By considering both the motivation and the best response of the robot, a more rational and safe decision is obtained, and the decision - making learning efficiency is improved by considering the internal interaction among participants during gradient calls, ultimately improving the robot decision - making and planning effect.
[0033] 2. For the robot decision - making and planning scenario, the present invention improves the rationality of gradient calls during decision - making dynamic learning by considering the best response information of the robot, and can effectively improve the safety and reliability of the decisions of the participants in robot decision - making and planning.
[0034] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] With reference to the accompanying drawings and the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. The drawings are used to better understand the present invention and do not limit the present invention. In the drawings, the same or similar reference numerals represent the same or similar elements, where:
[0036] Figure 1 is a flowchart of a robot decision-making and planning method based on a three-level player master-slave game provided by an embodiment of the present invention;
[0037] Figure 2 is a synchronous dynamic learning three-player trajectory diagram provided by an embodiment of the present invention;
[0038] Figure 3 is a three-level Stackelberg dynamic learning three-player trajectory diagram provided by an embodiment of the present invention;
[0039] Figure 4 is a schematic diagram of the target-defender team utility under different dynamic learning provided by an embodiment of the present invention;
[0040] Figure 5 is a structural diagram of a robot decision-making and planning device based on a three-level player master-slave game provided by an embodiment of the present invention;
[0041] Figure 6 is a structural diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. Detailed Embodiments
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] In addition, the term " / and" in the present invention is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.
[0044] To solve the technical problems in the background art, an embodiment of the present invention provides a robot decision-making and planning method, device, equipment, and storage medium based on a three-level player master-slave game. The following will combine the accompanying drawings to describe in detail a robot decision-making and planning method, device, equipment, and storage medium provided by an embodiment of the present invention through specific embodiments.
[0045] Figure 1 It is a flowchart of a robot decision-making and planning method provided by an embodiment of the present invention. As Figure 1 shown, the robot decision-making and planning method 100 may include:
[0046] S110. According to the robot decision-making and planning scenario faced by the robot system, construct a three-level player master-slave game decision model corresponding to the robot system; in the three-level player master-slave game decision model, there are three levels of players. The highest-level player is the senior leader who acts first and is dominant. The middle-level player is the intermediate follower who makes the best response after observing the actions of the senior leader. The lowest-level player is the junior follower who acts last and makes the best response after observing the actions of the senior leader and the intermediate follower.
[0047] S120. In the three-level Stackelberg dynamic learning for the three-level player master-slave game decision model, the higher-level player estimates the total derivative or partial derivative of the best response of the lower-level player based on implicit differentiation, combines the mapping relationship between the strategies of the lower-level player and its own strategy, obtains the total differential of its own cost function, calls the total differential of the cost function for all players respectively after giving the strategy of the higher-level player, and then obtains a new strategy through gradient descent. After that, based on this, robot decision-making and planning is performed on the robot system.
[0048] For the convenience of further understanding, the above steps will be described in detail below with specific embodiments:
[0049] a) Construction of the three-level player master-slave game decision model
[0050] For the robot decision-making and planning scenario faced by the robot system, according to the action sequence and dominant position of different robots in the robot system, and in combination with the goals or loss functions of each robot, construct a three-level player master-slave game decision model corresponding to the robot system. Specifically, it can be expressed by the following formula:
[0051]
[0052] Among them, X i represents the robot corresponding to the robot decision-making and planning scenario, collectively referred to as the player; x i represents the corresponding player strategy; n iDenote the corresponding player strategy dimension; Denote the corresponding strategy set; U i Denote the strategy set constraint; f i Is the player cost function; X1, X2, X3 represent the senior leader, intermediate follower, and junior follower respectively; i = 1, 2, 3.
[0053] In the three-level player leader-follower game decision model, the best response of the junior follower X3 to the actions of the senior leader X1 and the intermediate follower X2 is expressed as:
[0054]
[0055] The best response of the intermediate follower X2 to the action of the senior leader X1 is expressed as:
[0056]
[0057] Based on this, the local Stackelberg equilibrium of the three-level player leader-follower game decision model is defined as:
[0058] If the senior leader strategy Satisfies:
[0059]
[0060] Then the strategy set Constitutes the Stackelberg equilibrium of the three-level player leader-follower game decision model.
[0061] In addition, some assumptions need to be set for the three-level player leader-follower game decision model to ensure the existence of the Stackelberg equilibrium: First, the cost functions of all players are q-order continuously differentiable, that is, f i ∈C q , and q is greater than 2; Second, all best response functions are continuously differentiable; In addition, the player strategy constraints are all non-empty convex compact sets.
[0062] b) Three-level Stackelberg dynamic learning
[0063] First, denote the partial derivative of the cost function f i (x1, x2, x3) (i = 1, 2, 3) with respect to x j (j = 1, 2, 3) as Denote the cost function f i (x1, x2, x3) (i = 1, 2, 3) with respect to x j (j = 1, 2, 3) and x k (k = 1, 2, 3) and successively find the gradient to obtain the second-order partial derivative as When j = k, the second-order partial derivative is denoted as The total differential is denoted as D(·).
[0064] The gradient call vector of the three-level Stackelberg dynamic learning is denoted as:
[0065] ω S (x) = [ω S,1 (x), ω S,2 (x), ω S,3 (x)] (5)
[0066] Different from the synchronous gradient descent dynamic learning where each player directly calls the partial derivative of the cost function with respect to its own strategy D i f i In the three-level Stackelberg dynamic learning, in the way of updating the strategy, the player needs to consider the best response and behavioral motivation of the lower-level players, so as to establish an information difference advantage, know the opponent and oneself to obtain a more rational and stable strategy. Therefore, each player will call the total differential of its own and the lower-level players' strategies on the premise of the given determined behavior of the higher-level player. The total differential is composed of the sum of the partial derivatives of each unknown variable and the product of the variable's own gradient. Denote as Denote as And denote The partial derivative with respect to x2 is Thus, ω S (x) is:
[0067]
[0068] On this basis, based on the implicit differential estimation, solve D mr , D br , Specifically, first according to the definition of Stackelberg equilibrium and equations (2), (3), we can get:
[0069]
[0070] According to the above, using the implicit differentiation method, the specific expressions of the gradients of each best response can be obtained through solving equations and matrix operations. For example, for Taking the derivative with respect to x2 and setting it to 0, we can get
[0071] Thus, the above three-level Stackelberg dynamic learning can be summarized as the algorithm described in Table 1:
[0072] Table 1
[0073]
[0074] After that, for the robot decision-making and planning task, first initialize the initial states of each robot, and obtain ω according to their optimized objective functions. S Subsequently, guide the robots to act in real time according to the strategy update rules in the algorithm to pursue the Stackelberg equilibrium.
[0075] As an example, here, a three-robot path planning scenario is used to illustrate the robot decision-making and planning method based on the three-level player leader-follower game, which is specifically as follows:
[0076] Consider the three-robot path planning scenario and model it as a three-player pursuit-evasion zero-sum game. The three robots are divided into a target, a defender, and an attacker. The attacker chases the target, and the defender and the target are on the same team. The defender intercepts and interferes with the attacker. The strategies of the target, the attacker, and the defender are all two-dimensional vectors representing the horizontal and vertical coordinates in space, denoted as [x1; x2], [y1; y2]; [z1; z2] respectively. The cost functions of the three players are f(x, y, z), -f(x, y, z), f(x, y, z), where the expression of f is:
[0077]
[0078] Among them, all player strategies are in the interval [-35, 35], and each player endeavors to minimize the corresponding cost function. It is set that when the distance between the target and the attacker is less than 5, the attacker can capture the target, and when the distance between the attacker and the defender is less than 15, the defender can intercept and interfere with the attacker, which is manifested as the attenuation of the attacker's strategy update learning rate in the simulation experiment.
[0079] The three-player pursuit-evasion zero-sum game is respectively simulated by synchronous dynamic learning (directly calling the partial derivative of its own cost function with respect to its own strategy) and the three-level Stackelberg dynamic learning of the present invention. The trajectory diagrams of each player finally obtained are respectively as Figures 2-3 shown.
[0080] From Figure 2 it can be seen that when the strategy updates of each player adopt synchronous dynamic learning, the attacker can capture the target at the 15th iteration, that is, the target fails to escape; while from Figure 3 it can be seen that when the strategy updates of each player adopt the three-level Stackelberg dynamic learning, the target successfully escapes within 200 iterations, and at the 40th iteration, the defender can intercept and interfere with the attacker, slowing down its learning rate. The utility (cost function value) of the target-defender team within 200 iterations under different dynamic learnings is as Figure 4 shown. In synchronous dynamic learning, the strategy is no longer updated after the target is captured.
[0081] From Figure 4 It can be seen that the use of three - level Stackelberg dynamic learning can make the target - defender team utility converge to a lower value (the final utility of synchronous dynamic learning is 263.3689, and the final utility of three - level Stackelberg dynamic learning is - 988.5609). This result reflects the effectiveness of three - level Stackelberg dynamic learning for decision - making planning.
[0082] In summary, the present invention has at least achieved the following technical effects:
[0083] 1. The present invention takes into account the hierarchical structure in the robot system. By considering both the motivation and the best response of the robot, a more rational and safe decision is obtained, and the decision - making learning efficiency is improved by considering the interaction within each participant during gradient calls, ultimately improving the robot decision - making planning effect.
[0084] 2. For the robot decision - making planning scenario, the present invention improves the rationality of gradient calls during decision - making dynamic learning by considering the robot's best response information, and can effectively improve the security and reliability of the decisions of the participants in the robot decision - making planning.
[0085] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0086] The above is the introduction of the method embodiments. The following further illustrates the solution of the present invention through device embodiments.
[0087] Figure 5 Shown in the following is the structural diagram of a robot decision - making planning device based on three - level player - leader - follower game provided by an embodiment of the present invention. Figure 5 As shown, the robot decision - making planning device 500 may include:
[0088] A construction module 510, configured to construct a three - level player - leader - follower game decision model corresponding to the robot system according to the robot decision - making planning scenario faced by the robot system; in the three - level player - leader - follower game decision model, there are three levels of players. The highest - level player is a senior leader who acts first and is dominant. The middle - level player is an intermediate follower who makes the best response after observing the actions of the senior leader. The lowest - level player is a junior follower who makes the best response after observing the actions of the senior leader and the intermediate follower.
[0089] A planning module 520, in the process of the three - level Stackelberg equilibrium dynamic learning for the three - level player master - slave game decision - making model, is used for a higher - level player to estimate the total derivative or partial derivative of the best response of a lower - level player based on implicit differentiation, combine the mapping relationship between the strategies of the lower - level player and its own strategy, obtain the total differential of its own cost function, call the total differential of the cost function for all players after a given higher - level player strategy respectively, and then obtain a new strategy through gradient descent. After that, robot decision - making planning is carried out on the robot system based on this.
[0090] It can be understood that Figure 5 each module / unit in the robot decision - making planning device 500 shown has the function of implementing Figure 1 each step in the robot decision - making planning method 100 shown, and can achieve its corresponding technical effects. For the sake of brevity, it will not be elaborated here.
[0091] Figure 6 It is a structural diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. The electronic device 600 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 600 can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the present invention, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed in the present invention.
[0092] As Figure 6 shown, the electronic device 600 may include a computing unit 601, which can perform various appropriate actions and processes according to the computer program stored in the read - only memory (ROM) 602 or the computer program loaded from the storage unit 608 into the random - access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0093] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0094] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 can be implemented as a computer program product, including a computer program, which is tangibly contained in a computer-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method 100 described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute method 100 in any other suitable manner (e.g., by means of firmware).
[0095] The various embodiments described above in the present invention can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0096] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0097] In the context of the present invention, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a computer-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0098] It should be noted that the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute method 100 and achieve the corresponding technical effects achieved by the method of the embodiments of the present invention. For the sake of brevity of description, details are not repeated herein.
[0099] In addition, the present invention also provides a computer program product, which includes a computer program that implements method 100 when executed by a processor.
[0100] It should be understood that various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. The present invention is not limited herein.
[0101] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A robot decision-making and planning method based on a three-level master-slave game among players, characterized in that, The method includes: According to the robot decision-making and planning scenario faced by the robot system, a three-level player master-slave game decision-making model corresponding to the robot system is constructed; in the three-level player master-slave game decision-making model, there are three levels of players. The highest-level player is the senior leader who acts first and is dominant. The middle-level player is the intermediate follower who makes the best response after observing the actions of the senior leader. The lowest-level player is the junior follower who acts last and makes the best response after observing the actions of the senior leader and the intermediate follower. In the three-level Stackelberg dynamic learning for the three-level player master-slave game decision-making model, the higher-level player estimates the total derivative or partial derivative of the best response of the lower-level player based on implicit differentiation, combines the mapping relationship between the strategies of the lower-level player and its own strategy, obtains the total differential of its own cost function, calls the total differential of the cost function for all players after a given higher-level player's strategy respectively, and then obtains a new strategy through gradient descent. After that, robot decision-making and planning are carried out for the robot system based on this.
2. The method according to claim 1, characterized in that, The constructing of a three-level player master-slave game decision-making model corresponding to the robot system according to the robot decision-making and planning scenario faced by the robot system includes: For the robot decision-making and planning scenario faced by the robot system, according to the action sequence and dominant position of different robots in the robot system, and in combination with the objectives or loss functions of each robot, a three-level player master-slave game decision-making model corresponding to the robot system is constructed.
3. The method according to claim 2, characterized in that, The three-level player master-slave game decision-making model corresponding to the robot system can be expressed by the following formula: Among them, X i represents the robot corresponding to the robot decision-making and planning scenario, collectively referred to as the player; x i represents the corresponding player strategy; n i represents the dimension of the corresponding player strategy; represents the corresponding strategy set; U i represents the strategy set constraint; f i is the player cost function; X1, X2, X3 represent the senior leader, the intermediate follower, and the junior follower respectively; i = 1, 2, 3.
4. The method according to claim 3, characterized in that, In the three-level player master-slave game decision-making model, the best response of the junior follower X3 to the actions of the senior leader X1 and the intermediate follower X2 is expressed as: The best response of the intermediate follower X2 to the action of the senior leader X1 is expressed as: Based on this, the local Stackelberg equilibrium of the three-level player master-slave game decision-making model is defined as: If the senior leader strategy Satisfies: Then the strategy set constitutes the Stackelberg equilibrium of the three-level player principal-agent game decision model.
5. The method according to claim 4, wherein In the three-level Stackelberg dynamic learning for the three-level player master-slave game decision model, first, the partial derivative of the cost function f i (x1,x2,x3) (i = 1,2,3) with respect to x j (j = 1,2,3) is denoted as The second-order partial derivative obtained by successively taking the gradient of the cost function f i (x1,x2,x3) (i = 1,2,3) with respect to x j (j = 1,2,3) and x k (k = 1,2,3) is denoted as When j = k, the second-order partial derivative is denoted as The total differential is denoted as D(·); The gradient call vector of the three-level Stackelberg dynamic learning is denoted as: ω S (x) = [ω S,1 (x), ω S,2 (x), ω S,3 (x)]; Call the total differential of the cost function for each player separately after a given higher-level player strategy. The total differential is composed of the sum of the products of the partial derivatives of each unknown variable and the gradient of the variable itself. Denote it as Denote it as And denote The partial derivative with respect to x2 is From this, ω S (x) can be obtained as: Based on this, solve for D based on implicit differential estimation mr , D br , 6. A robot decision-making and planning device based on a three-level master-slave game of players, characterized in that The device includes: A construction module for constructing a three-level player master-slave game decision-making model corresponding to the robot system according to the robot decision-making and planning scenario faced by the robot system; in the three-level player master-slave game decision-making model, there are three levels of players. The highest-level player is the senior leader who acts first and is dominant. The middle-level player is the intermediate follower who makes the best response after observing the actions of the senior leader. The lowest-level player is the junior follower who acts last and makes the best response after observing the actions of the senior leader and the intermediate follower. A planning module for, in the three-level Stackelberg dynamic learning for the three-level player master-slave game decision-making model, the higher-level player estimating the total derivative or partial derivative of the best response of the lower-level player based on implicit differentiation, combining the mapping relationship between the strategies of the lower-level player and its own strategy, obtaining the total differential of its own cost function, calling the total differential of the cost function for all players after a given higher-level player's strategy respectively, and then obtaining a new strategy through gradient descent. After that, robot decision-making and planning are carried out for the robot system based on this.
7. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-5.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
Stackeberg game-based cognitive network resource allocation method
CN102076025A
Mobile target defense decision selection method, device and system based on Markov time game
CN110300106A
Multi-agent transaction method and system of integrated energy system
CN113592648A
Multi-level optimization control system and method based on Stackelberg-Nash differential game
CN114815598A
Random nonlinear predictive controller and method based on uncertainty propagation by means of Gaussian hypothesis density filter
CN117581166A