Intervention regulation and control method and device for promoting collaborative operation of multiple robots

Through non-cooperative game modeling and intervention signal design, the resource coordination problem in multi-robot collaborative operations is solved, and the overall optimization and efficient resource allocation of the robot system are achieved.

CN120066014APending Publication Date: 2025-05-30TONGJI UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510088472.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In multi-robot collaborative operations, it is difficult for the existing technology to effectively coordinate robot resources, resulting in uneven resource allocation and a decrease in overall efficiency.

Method used

By modeling the multi-robot collaborative operation scenarios of the robot system, a method for supervisors to impose intervention signals on the robot is designed to maximize the overall benefit function of the robot system. The method includes static intervention algorithm and dynamic intervention algorithm, and different types of interventions are carried out based on the acquired information.

Benefits of technology

The overall optimization of the robot system is achieved, the efficiency of resource allocation and system coordination are improved, and the stability and robustness of the robot behavior are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066014A_ABST
    Figure CN120066014A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intervention regulation and control method and device for promoting multi-robot collaborative operation, and the method comprises the steps: carrying out the modeling of a multi-robot collaborative operation scene of a robot system, and obtaining a non-cooperative game model; wherein the robot system comprises a plurality of robots and a supervisor, each robot has a corresponding revenue function, the objective of the robot is to maximize the corresponding revenue function by making a decision, and the objective of the supervisor is to maximize the overall revenue function of the robot system by applying an intervention signal to the robot to influence the decision of the robot; the position of the robot is subjected to spatial set constraint and coupling constraint; the intervention signal applied by the supervisor has budget constraint; and the supervisor applies an intervention signal to the robot by adopting a corresponding dry prediction algorithm based on the observed game information of the non-cooperative game model so as to maximize the overall revenue function of the robot system. In this way, the robot can be intervened, regulated and controlled, and efficient configuration of resources and efficient cooperation of the system are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robots, and in particular, to an intervention and regulation method and device for promoting multi-robot collaborative operation. Background Art

[0002] In the context of the rapid development of artificial intelligence technology, multi-robot collaborative operation has become a key force in promoting the intelligent transformation of cities. When multi-robots collaborate, if each robot only pursues the optimal completion of individual tasks and ignores the coordination of the overall system, it may lead to uneven resource allocation and a decline in overall efficiency. In view of this, there is an urgent need to explore an intervention and regulation method for promoting multi-robot collaborative operation to achieve efficient resource allocation and efficient system collaboration. Summary of the Invention

[0003] In a first aspect, an embodiment of the present invention provides an intervention and regulation method for promoting multi-robot collaborative operation, the method comprising:

[0004] Performing non-cooperative game modeling on the multi-robot collaborative operation scenario of the robot system to obtain a non-cooperative game model; in the non-cooperative game model, it includes multiple robots and a supervisor, each robot has a corresponding revenue function, and its goal is to maximize the corresponding revenue function by making decisions. The goal of the supervisor is to influence the decisions of the robots by applying intervention signals to the robots, thereby maximizing the overall revenue function of the robot system; the positions of the robots are subject to spatial set constraints and coupling constraints; the intervention signals applied by the supervisor are subject to budget constraints;

[0005] The supervisor applies an intervention signal to the robots based on the observed game information of the non-cooperative game model by using a corresponding intervention method to maximize the overall revenue function of the robot system.

[0006] In some realizable ways of the first aspect, the decision of the robot is the position of the robot, and maximizing the corresponding revenue function means that the energy loss of the robot when completing the expected task is minimized; the intervention signal applied by the supervisor refers to the disturbance signal applied by the supervisor to the robot.

[0007] In some realizable ways of the first aspect, the supervisor applies an intervention signal to the robots based on the observed game information of the non-cooperative game model by using a corresponding intervention method to maximize the overall revenue function of the robot system, including:

[0008] If the observed game information of the non-cooperative game model is complete information, the supervisor applies a static intervention signal to the robots based on the complete information by using a static intervention method to maximize the overall revenue function of the robot system.

[0009] In some realizable ways of the first aspect, the complete information includes updated gradient information, the activity space of the robot, the stable solution of the Lagrange multiplier, and the optimal position of the robot;

[0010] Based on the complete information, the supervisor uses a static intervention method to impose a static intervention signal on the robot to maximize the overall profit function of the robot system, including:

[0011] Based on the complete information, the supervisor designs a static intervention signal using a static intervention method;

[0012] d k = F(x opt ) + A T λ opt + o;

[0013] where d k represents the static intervention signal; F(·) represents the updated gradient information; x opt represents the optimal position of the robot; λ opt represents the stable solution of the Lagrange multiplier; A represents the coefficient matrix corresponding to the coupling constraint; o represents the projection of x opt on the normal cone; represents the activity space of the robot;

[0014] Imposing a static intervention signal on the robot so that the robot maximizes the overall profit function of the robot system according to its own update algorithm.

[0015] In some realizable ways of the first aspect, based on the game information of the observed non - cooperative game model, the supervisor uses the corresponding intervention method to impose an intervention signal on the robot to maximize the overall profit function of the robot system, including:

[0016] If the game information of the observed non - cooperative game model is restricted information, then based on the restricted information, the supervisor uses a dynamic intervention method to impose a dynamic intervention signal on the robot to maximize the overall profit function of the robot system.

[0017] In some realizable ways of the first aspect, the restricted information includes the optimal position of the robot;

[0018] Based on the restricted information, the supervisor uses a dynamic intervention method to impose a dynamic intervention signal on the robot to maximize the overall profit function of the robot system, including:

[0019] Based on the restricted information, the supervisor designs a dynamic intervention signal for the robot using a dynamic intervention method;

[0020]

[0021] where dk+1 represents the dynamic intervention signal, that is, the intervention signal at time k+1; represents the projection operator for the budget constraint, that is, this projection operator can project the updated intervention signal into the budget constraint of the intervention signal, ensuring that the intervention signal is always within the established budget constraint; d k represents the intervention signal at time k; β k represents the update step size; x opt represents the optimal position of the robot; x k represents the position of the robot at time k;

[0022] Apply a dynamic intervention signal to the robot so that the robot can maximize the overall revenue function of the robot system according to its own update algorithm.

[0023] In some implementable ways of the first aspect, the update algorithm of the robot is represented by the following formula:

[0024]

[0025]

[0026] where, x i,k+1 represents the position of robot i at time k+1; represents the projection operator for the activity space of robot i, that is, this projection operator projects the updated position of robot i into the activity space of robot i, ensuring that the position of robot i is always within its activity space; x i,k represents the position of robot i at time k; β k represents the update step size; u i,k represents the velocity drive signal of robot i at time k; represents the partial derivative of the revenue function of robot i with respect to its position, that is, it means moving in the direction that maximizes the revenue function of robot i; A i represents the coefficient matrix of the coupling constraint of robot i; λ k represents the Lagrange multiplier; d i,k represents the intervention signal applied to robot i at time k;

[0027] The update of the Lagrangian operator is represented by the following formula:

[0028]

[0029] where, λ i,k+1 represents the Lagrange multiplier related to robot i at time k+1; represents the projection operator of the Lagrange multiplier, that is, this operator projects the value of the updated Lagrange multiplier into the non-negative real number space, ensuring the non-negativity of the Lagrange multiplier; λ i,kDenote the Lagrange multiplier related to machine \(i\) at time \(k\); \(\beta\) k Denote the update step size; \(A\) i Denote the coefficient matrix of the coupling constraint of robot \(i\); \(x\) i,k Denote the position of robot \(i\) at time \(k\); \(b\) i Denote the constant matrix of the coupling constraint of robot \(i\).

[0030] In a second aspect, an intervention and regulation device for promoting multi-robot collaborative operation provided by an embodiment of the present invention includes:

[0031] A modeling module, configured to perform non-cooperative game modeling on the multi-robot collaborative operation scenario of a robot system to obtain a non-cooperative game model; in the non-cooperative game model, it includes multiple robots and a supervisor, each robot has a corresponding revenue function, and its goal is to maximize the corresponding revenue function by making decisions. The goal of the supervisor is to influence the decisions of the robots by applying intervention signals to the robots, thereby maximizing the overall revenue function of the robot system; the positions of the robots are subject to spatial set constraints and coupling constraints; there is a budget constraint on the intervention signals applied by the supervisor;

[0032] An intervention module, configured to enable the supervisor to apply an intervention signal to the robot based on the observed game information of the non-cooperative game model by using a corresponding intervention algorithm, so as to maximize the overall revenue function of the robot system.

[0033] In a third aspect, an embodiment of the present invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0034] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the method as described above.

[0035] In the embodiment of the present invention, the robots can be intervened and regulated to promote multi-robot collaborative operation, and then guide the entire robot system towards overall optimization, realizing efficient resource allocation and efficient system collaboration.

[0036] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. The accompanying drawings are used to better understand the present invention and do not limit the present invention. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0038] Figure 1 is a flowchart of an intervention regulation method for promoting multi-robot collaborative operation provided by an embodiment of the present invention;

[0039] Figure 2 is a schematic diagram of a non-cooperative game model provided by an embodiment of the present invention;

[0040] Figure 3 is a schematic diagram of an intervention algorithm provided by an embodiment of the present invention;

[0041] Figure 4 is a robot collaborative optimization task topology diagram provided by an embodiment of the present invention;

[0042] Figure 5 is a schematic diagram of error convergence provided by an embodiment of the present invention;

[0043] Figure 6 is a schematic diagram of the overall benefit under intervention provided by an embodiment of the present invention;

[0044] Figure 7 is a schematic diagram of the equilibrium change of the robot before and after intervention provided by an embodiment of the present invention;

[0045] Figure 8 is a structural diagram of an intervention regulation device for promoting multi-robot collaborative operation provided by an embodiment of the present invention;

[0046] Figure 9 is a structural diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. Detailed Embodiments

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] In addition, in the present invention, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the present invention, the character " / " generally represents an "or" relationship between the front and back associated objects.

[0049] To solve the technical problems appearing in the background art, the embodiments of the present invention provide an intervention regulation method, device, equipment, and storage medium for promoting multi-robot collaborative operation. The following will combine the accompanying drawings to describe in detail an intervention regulation method, device, equipment, and storage medium for promoting multi-robot collaborative operation provided by the embodiments of the present invention through specific embodiments.

[0050] Figure 1 The flowchart of an intervention regulation method for promoting multi-robot collaborative operation provided by the embodiments of the present invention is as Figure 1 shown. The intervention regulation method 100 may include:

[0051] S110, perform non-cooperative game modeling on the multi-robot collaborative operation scenario of the robot system to obtain a non-cooperative game model. In the non-cooperative game model, it includes multiple robots and a supervisor. Each robot has a corresponding payoff function, and its goal is to maximize the corresponding payoff function by making decisions. The goal of the supervisor is to influence the decisions of the robots by applying intervention signals to the robots, thereby maximizing the overall payoff function of the robot system. The positions of the robots are subject to spatial set constraints and coupling constraints. There is a budget constraint on the intervention signals applied by the supervisor.

[0052] S120, the supervisor applies an intervention signal to the robots based on the observed game information of the non-cooperative game model by using a corresponding intervention algorithm to maximize the overall payoff function of the robot system.

[0053] For the convenience of further understanding, the above steps will be described in detail below with specific embodiments:

[0054] The present invention studies an intervention regulation method for promoting multi-robot collaborative operation. First, non-cooperative game modeling is performed on the multi-robot collaborative operation scenario of the robot system to obtain a non-cooperative game model.

[0055] Among them, the non-cooperative game model may be as Figure 2 shown, including multiple robots and a supervisor. Each robot i has a corresponding payoff function U i (x, d i ), and its goal is to maximize the payoff function U by making decisions i (x, d i), where the decision of the robot is x i is the position of the robot, maximizing the profit function U i (x, d i ) means minimizing the energy loss when the robot i completes the expected task. The goal of the supervisor (i.e., the higher-level controller, the dominant party) is to influence the decision x of the robot by applying the intervention signal d i to the robot, and then maximizing the overall profit function of the robot system i Here, the intervention signal d is the disturbance signal applied by the supervisor to the robot. The position of the robot i is subject to spatial set constraints and coupling constraints, and there is a budget constraint on the intervention signal applied by the supervisor. Specifically, as follows: i The profit function (i.e., the energy profit function) of each robot i is:

[0056]

[0057]

[0058] Among them, the position of the robot is subject to spatial set constraints, that is the set usually represents the possible activity space of the robot; there is also a budget constraint on the intervention signal applied by the supervisor This budget constraint reflects the limitations of the supervisor's ability; to ensure that the distance between robots can maintain communication and avoid obstacles, the position of the robot also needs to satisfy the following coupling constraints:

[0059]

[0060] Among them, A := [A 1 , …, A N and represent the coefficient matrix and the constant matrix corresponding to the coupling constraint, and the feasible region of the robot strategy combination can be obtained as This feasible region is the activity space of the robot to ensure the constraint conditions.

[0061] Before applying the interference, that is, when d i ≡ 0, the profit function of each robot i is W i (x), and its goal is to maximize the profit function W i (x). This profit function is usually the quadratic term index of the robot, that is, related to the tracking error and energy loss; after the intervention, its profit function changes from W i (x) to U i (x), and its goal also changes from maximizing W i (x) to maximizing U i(x), the change in the revenue function is reflected in that after adding a perturbation signal to the robot, in order to overcome the interference, its revenue function increases by terms. Now consider the speed-driven robot.

[0062] The update algorithm for the position of robot i at time k is:

[0063]

[0064] where x i,k+1 represents the position of robot i at time k + 1; represents the projection operator of the activity space of robot i, that is, this projection operator projects the updated position of robot i into the activity space of robot i to ensure that the position of robot i is always within its activity space; x i,k represents the position of robot i at time k; β k represents the update step size, which determines the rate of robot position update; u i,k represents the speed drive signal of robot i at time k.

[0065] Speed drive signal:

[0066]

[0067] where u i,k represents the speed drive signal of robot i at time k; represents the partial derivative of the revenue function of robot i with respect to its position, that is, it means that the position moves in the direction that maximizes the revenue function of robot i, and its role is to update the position in the direction of gradient descent to maximize the revenue function; A i represents the coefficient matrix of the coupling constraint of robot i; λ k represents the Lagrange multiplier, and its role is to ensure that the robot can maintain communication and avoid obstacles during the update process; d i,k represents the intervention signal applied to robot i at time k.

[0068] In order to enable the robot to maintain communication and avoid obstacles, the update algorithm for the Lagrange multiplier is:

[0069]

[0070] where λ i,k+1 represents the Lagrange multiplier related to robot i at time k + 1; represents the projection operator of the Lagrange multiplier, that is, this operator projects the value of the updated Lagrange multiplier into the non-negative real number space to ensure the non-negativity of the Lagrange multiplier; λ i,k represents the Lagrange multiplier related to machine i at time k; β kdenotes the update step size; A i denotes the coefficient matrix of the coupling constraint of robot i; x i,k denotes the position of robot i at time k; b i denotes the constant matrix of the coupling constraint of robot i.

[0071] For the robot system, the position update algorithm of the robot system can be expressed as:

[0072]

[0073] The speed drive signal of the robot system:

[0074] u k = -F(x k ) - A T λ k + d k #(7)

[0075] The update algorithm for the Lagrange multiplier of the robot system is:

[0076]

[0077] where, denotes the value of the set of all robot positions at time k; denotes the value of the supervisor's intervention signal at time k; denotes the profit function W i (x k ) the negative of the gradient, and the Lagrange multiplier is updated in the direction that satisfies the coupling constraint during the update process; other parameters are similar to those in formulas (3)-(5) and will not be elaborated here.

[0078] In view of this, the position information of each robot when the robot system is stable can be solved using the following variational inequality:

[0079]

[0080] where, K denotes the activity space of the robot after satisfying the constraints; -F denotes the gradient information of the profit function; denotes the final intervention signal applied.

[0081] The goal of the supervisor is to maximize the overall profit function of the robot system, that is, to maximize Based on this, the desired speed drive signal u of robot i at time k can be obtained as: i,k is:

[0082]

[0083] Among them, It means that each robot updates in the gradient direction that maximizes the benefit of the robot system; A i represents the coefficient matrix of the coupling constraint of robot i; λ k represents the Lagrange multiplier.

[0084] The desired speed drive signal of the robot system is:

[0085] u k = -H(x k ) - A T λ k #(11)

[0086]

[0087] Among them, H(x k ) represents the negative of the gradient of the overall benefit function .

[0088] In view of this, the position information of each robot when the robot system is stable can be solved by using the following variational inequality, that is, the position information of the robot when the desired overall benefit function is maximized:

[0089]

[0090] Among them, K represents the activity space of the robot after satisfying the constraints; -H represents the gradient information of the overall benefit function of the robot system.

[0091] The interference methods used here are mainly divided into two categories. As Figure 3 shown, if the game information of the observed non - cooperative game model is complete information, the supervisor uses the static interference method to impose a static interference signal on the robot to maximize the overall benefit function of the robot system; if the game information of the observed non - cooperative game model is restricted information, the supervisor, based on the restricted information, uses the dynamic interference method to impose a dynamic interference signal on the robot to maximize the overall benefit function of the robot system. Specifically as follows:

[0092] a) Static interference method under complete information

[0093] Under complete information, the supervisor can obtain the update gradient information F(·) of the robot, the activity space of the robot the stable solution of the Lagrange multiplier λ opt and the optimal position x of the robot opt , so it can design a static interference signal according to this information:

[0094] d k = F(xopt ) + A T λ opt + o#(14)

[0095] where d k represents the static intervention signal; F(·) represents the updated gradient information; x opt represents the optimal position of the robot; λ o p t represents the stable solution of the Lagrange multiplier; A represents the coefficient matrix corresponding to the coupling constraint; o represents the projection of x opt on the normal cone, and is further represented as the set of all directions leaving the active space opt at the optimal position x ; represents the active space of the robot.

[0096] After that, a static intervention signal is applied to the robot so that the robot maximizes the overall revenue function of the robot system according to its own update algorithm.

[0097] Exemplarily, the static intervention algorithm and the robot update algorithm can be as shown in Table 1:

[0098] Table 1

[0099]

[0100] b) Dynamic intervention algorithm under limited information

[0101] Under limited information, the supervisor can only obtain the optimal position x opt of the robot, but cannot obtain other specific information. In this case, the intervention signal is no longer fixed, but can have the property of dynamically converging to the intervention stable point . And the integral controller based on error feedback has the above properties. Therefore, the supervisor can design a dynamic intervention signal using the dynamic intervention algorithm based on this information:

[0102]

[0103] where d k+1 represents the dynamic intervention signal, that is, the intervention signal at time k + 1; represents the projection operator of the budget constraint, that is, this projection operator can project the updated intervention signal into the budget constraint of the intervention signal to ensure that the intervention signal is always within the established budget constraint; d k represents the intervention signal at time k; β k represents the update step size; x opt represents the optimal position of the robot; x k represents the position of the robot at time k.

[0104] Subsequently, a dynamic intervention signal is applied to the robot so that the robot maximizes the overall benefit function of the robot system according to its own update algorithm.

[0105] Exemplarily, the dynamic intervention algorithm and the robot update algorithm can be as shown in Table 2:

[0106]

[0107]

[0108] In summary, in view of the information differences that the supervisor can obtain in different environments, the present invention designs different intervention algorithms. In an environment with relatively sufficient information, the supervisor can obtain the parameter information of the robot system. At this time, applying Algorithm 1 of the present invention can quickly motivate each robot to cooperate, and ultimately achieve the system optimum; while in a harsh environment with limited information, the supervisor cannot directly observe the parameters of the robot system. At this time, Algorithm 2 can be applied to the robot, so that the robot behavior can dynamically reach the system optimal strategy.

[0109] In view of this, the present invention has at least achieved the following technical effects:

[0110] 1. The introduction of the present invention constructs an innovative framework for dealing with the problem of robot cooperative regulation in non - cooperative games. This method comprehensively considers the limitations of the robot behavior set, the coupling constraints between behaviors, and the budget constraints of the supervisor in the regulation process, thus effectively achieving the system optimal goal. Compared with the prior art, the present invention exhibits outstanding characteristics such as low information dependence, dynamic regulation ability, excellent stability and robustness, and wide applicability, providing a novel perspective and means for solving the non - cooperative game regulation problem in robot cooperation, and having double values of theoretical innovation and practical application.

[0111] 2. The present invention explores a static intervention algorithm based on complete information. In this algorithm, the supervisor designs a static intervention signal based on known data by using the projected gradient descent algorithm and the KKT operator theory to guide the robot to gradually approach the overall optimal strategy, so as to maximize the overall benefit of the robot system. This method shows excellent performance in the regulation effect and has high pertinence.

[0112] 3. In addition, the present invention also develops a dynamic intervention algorithm based on limited information. The supervisor only needs to master the information related to the optimal behavior and dynamically adjusts the intervention signal through an integral controller based on error feedback. The robot updates its strategy in real - time according to these intervention signals, jointly forming a generalized dynamic system, and ultimately promoting the system benefit to reach the maximum. This algorithm demonstrates excellent adaptability and robustness, providing strong technical support for robot cooperative regulation.

[0113] Based on the above, the intervention and regulation method 100 provided by the present invention will be described in detail below in combination with a specific embodiment:

[0114] As Figure 4 shown, consider the collaborative optimization task of four robots. For each robot, they move on a plane, aiming to optimize some individual objectives related to their positions while maintaining communication connectivity. For each robot i ∈ {1, …, 4}, its profit function is:

[0115]

[0116] where p i = col(x i , y i ) is the coordinate position of the robot, is a local parameter usually related to the velocity term, is the intervention signal received by robot i. The activity space of the robot is 0.1 ≤ y i ≤ 0.5, that is, the activity position of the robot on the y-axis is [0.1, 0.5]. In order to keep all robots in communication with their neighbors, it is necessary to force the Chebyshev distance between any two adjacent robots to be less than 0.2 m. Therefore, the coupling constraint can be expressed as:

[0117]

[0118] where (i, j) ∈ E means that robot j can receive the position information of robot i, indicating that i and j are connected in the topological graph. This coupling constraint means that the maximum distance between robot i and robot j on the x or y axis cannot exceed 0.2.

[0119] In this simulation experiment, consider indicating that the magnitude of the intervention signal is between -1 and 1, and the update step size β k of the robot takes 1 / (1 + k) 0.6 . Through calculation, the position of the robot when the system profit is optimal is: p opt = [0.70; 0.40; 0.58; 0.38; 0.63; 0.40; 0.49; 0.32].

[0120] The error convergence of the robot behavior is as Figure 5 shown. Whether it is static intervention or dynamic intervention, the behavior of the robot gradually converges to the socially optimal state, and the convergence speed of dynamic intervention is slower than that of static intervention. The main reason is that dynamic intervention ultimately converges dynamically to static intervention, that is, its final point corresponds to the static budget constraint value, so its convergence speed is slower.

[0121] The overall benefit of the robot system is as Figure 6 shown. It can be seen that after the intervention, compared with before the intervention, the sum of its overall benefits increases. That is, before the intervention, even if all robots reach individual optimality, due to the mutual influence of the behaviors among the robots, the system benefit at this time is a sub-optimal state. However, after the intervention, the system benefit corresponding to the robot system is optimal.

[0122] The equilibrium change of the robots before and after the intervention is as Figure 7 shown. It can be seen that the essence of the intervention lies in adjusting the equilibrium state of the game. By introducing an intervention term into the robot benefit function, the equilibrium point of the game is transferred from the Nash equilibrium point to the socially optimal point.

[0123] In summary, the simulation experiments show that the proposed algorithm can achieve the maximization of the robot system benefit under complete and limited game information.

[0124] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0125] The above is the introduction of the method embodiments. The following further illustrates the solution of the present invention through device embodiments.

[0126] Figure 8 The following is a structural diagram of an intervention regulation device for promoting multi-robot collaborative operation provided by an embodiment of the present invention. As Figure 8 shown, the intervention regulation device 800 may include:

[0127] A modeling module 810, configured to perform non-cooperative game modeling on the multi-robot collaborative operation scenario of the robot system to obtain a non-cooperative game model. In the non-cooperative game model, it includes multiple robots and a supervisor. Each robot has a corresponding benefit function, and its goal is to maximize the corresponding benefit function by making decisions. The goal of the supervisor is to influence the decisions of the robots by applying intervention signals to the robots, thereby maximizing the overall benefit function of the robot system. The positions of the robots are subject to spatial set constraints and coupling constraints. There is a budget constraint on the intervention signals applied by the supervisor.

[0128] An intervention module 820, configured for a supervisor to apply an intervention signal to a robot by using a corresponding intervention algorithm based on the observed game information of the non - cooperative game model, so as to maximize the overall revenue function of the robot system.

[0129] It can be understood that Figure 8 each module / unit in the illustrated intervention and regulation device 800 has the function of implementing Figure 1 each step in the illustrated intervention and regulation method 100, and can achieve its corresponding technical effects. For the sake of brevity, they will not be described herein again.

[0130] Figure 9 It is a structural diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. The electronic device 900 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 900 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the present invention, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed in the present invention.

[0131] As Figure 9 shown, the electronic device 900 may include a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read - only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random - access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0132] A plurality of components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0133] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer program product, including a computer program that is tangibly embodied in a computer-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the method 100 described above may be executed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute method 100 in any other suitable manner (e.g., by means of firmware).

[0134] The various embodiments described above in the present invention can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0135] The program code for implementing the methods of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0136] In the context of the present invention, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a computer-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0137] It should be noted that the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute method 100 and achieve the corresponding technical effects achieved by the method of the embodiments of the present invention. For the sake of concise description, details are not repeated herein.

[0138] In addition, the present invention also provides a computer program product, which includes a computer program that implements method 100 when executed by a processor.

[0139] It should be understood that various forms of the flow shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. The present invention is not limited herein.

[0140] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intervention and control method for promoting multi-robot collaborative operation, characterized in that: The method comprises: The non-cooperative game modeling is carried out for the multi-robot collaborative operation scenario of the robot system to obtain a non-cooperative game model; in the non-cooperative game model, there are multiple robots and a supervisor, each robot has a corresponding profit function, and its goal is to maximize the corresponding profit function by making decisions, and the supervisor's goal is to influence the robot's decision by applying intervention signals to the robot, thereby maximizing the overall profit function of the robot system; the position of the robot is subject to spatial set constraints and coupling constraints; the intervention signal imposed by the supervisor has a budget constraint; Based on the observed game information of the non-cooperative game model, the regulator uses the corresponding intervention algorithm to apply intervention signals to the robot to maximize the overall benefit function of the robot system.

2. The method according to claim 1, characterized in that The decision of the robot is the position of the robot, and maximizing the corresponding benefit function means minimizing the energy loss when the robot completes the desired task; the intervention signal applied by the supervisor refers to the disturbance signal applied by the supervisor to the robot.

3. The method according to claim 1, characterized in that: The supervisor applies intervention signals to the robot using a corresponding intervention algorithm based on the observed game information of the non-cooperative game model to maximize the overall benefit function of the robot system, including: If the observed game information of the non-cooperative game model is complete information, the regulator uses a static intervention algorithm based on complete information to apply a static intervention signal to the robot to maximize the overall benefit function of the robot system.

4. The method according to claim 3, characterized in that The complete information includes updated gradient information, the robot's activity space, Lagrange multiplier stable solution and the robot's optimal position; The supervisor applies a static intervention signal to the robot using a static intervention algorithm based on complete information to maximize the overall benefit function of the robot system, including: Based on complete information, the regulator uses a static intervention algorithm to design a static intervention signal; d k =F(x opt )+A T l opt +o; Among them, d k represents the static intervention signal; F(·) represents the updated gradient information; x opt represents the optimal position of the robot; opt represents the Lagrange multiplier stable solution; A represents the coefficient matrix corresponding to the coupling constraint; o represents x opt exist Projection on the normal cone; Represents the robot's activity space; A static intervention signal is applied to the robot so that the robot maximizes the overall benefit function of the robot system according to its own update algorithm.

5. The method according to claim 1, characterized in that The supervisor applies intervention signals to the robot using a corresponding intervention algorithm based on the observed game information of the non-cooperative game model to maximize the overall benefit function of the robot system, including: If the observed game information of the non-cooperative game model is restricted information, the regulator uses a dynamic intervention algorithm to apply dynamic intervention signals to the robot based on the restricted information to maximize the overall benefit function of the robot system.

6. The method according to claim 5, characterized in that The restricted information includes the optimal position of the robot; The supervisor applies a dynamic intervention signal to the robot using a dynamic intervention algorithm based on the restricted information to maximize the overall benefit function of the robot system, including: Based on the restricted information, the supervisor uses a dynamic intervention algorithm to design dynamic intervention signals to the robot; Among them, d k+1 represents the dynamic intervention signal, that is, the intervention signal at time k+1; represents the projection operator of the budget constraint, that is, the projection operator can project the updated intervention signal into the budget constraint of the intervention signal to ensure that the intervention signal is always within the established budget constraint; d k represents the intervention signal at time k; β k represents the update step size; x opt represents the optimal position of the robot; x k represents the robot position at time k; A dynamic intervention signal is applied to the robot so that the robot maximizes the overall benefit function of the robot system according to its own update algorithm.

7. The method according to claim 4 or 6, characterized in that: The robot's update algorithm is expressed using the following formula: Among them, x i,k+1 represents the position of robot i at time k+1; The projection operator represents the activity space of robot i, that is, the projection operator projects the updated position of robot i into the activity space of robot i, ensuring that the position of robot i is always within its activity space; x i,k represents the position of robot i at time k; β k represents the update step size; u i,k represents the speed driving signal of robot i at time k; A represents the partial derivative of the profit function of robot i with respect to its position, that is, its position moves in the direction that maximizes the profit function of robot i; i The coefficient matrix representing the coupling constraints of robot i; λ k represents the Lagrange multiplier; d i,k represents the intervention signal applied to robot i at time k; The update of the Lagrangian operator is expressed as follows: Among them, λ i,k+1 represents the Lagrange multiplier associated with robot i at time k+1; represents the projection operator of the Lagrange multiplier, that is, the operator projects the updated value of the Lagrange multiplier into the non-negative real number space to ensure the non-negativity of the Lagrange multiplier; λ i,k represents the Lagrange multiplier associated with machine i at time k; β k represents the update step size; A i The coefficient matrix representing the coupling constraints of robot i; x i,k represents the position of robot i at time k; b i A constant matrix representing the coupling constraints of robot i.

8. An intervention and control device for promoting multi-robot collaborative operation, characterized in that: The device comprises: A modeling module is used to perform non-cooperative game modeling on a multi-robot collaborative operation scenario of a robot system to obtain a non-cooperative game model; in the non-cooperative game model, there are multiple robots and a supervisor, each robot has a corresponding profit function, and its goal is to maximize the corresponding profit function by making decisions, and the supervisor's goal is to influence the robot's decision by applying intervention signals to the robot, thereby maximizing the overall profit function of the robot system; the position of the robot is subject to spatial set constraints and coupling constraints; the intervention signal imposed by the supervisor has a budget constraint; The intervention module is used by the regulator to apply intervention signals to the robot based on the observed game information of the non-cooperative game model, using the corresponding intervention algorithm to maximize the overall benefit function of the robot system.

9. An electronic device, characterized in that: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Communication resource allocation method based on master-slave game

    CN120378224A

  • Business intervention method and device, equipment and storage medium

    CN121256406A

  • Business intervention methods, devices, equipment, and storage media

    CN121256406B