Intelligent Agent Scheduling Method and Device, Storage Medium and Program Product for Charging System
By building a two-layer topology and game model of intelligent parking lots, the decision-making conflict problems in multi-charging requests and heterogeneous multi-agent scheduling are solved, and efficient and responsive charging system scheduling is achieved.
Patent Information
- Application Number
- CN202510281593.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-11
AI Technical Summary
In smart parking lots, in the face of multi-charging requests and heterogeneous multi-agent control systems (charging robots and mobile charging piles), it is difficult for the existing technology to quickly and accurately schedule multi-agent systems to respond and solve the problem of decision-making conflicts.
A two-layer topology construction method is adopted, including static topology and dynamic topology, combined with cooperative game model and competition-cooperation game model, optimized through model prediction control to achieve Nash equilibrium scheduling decisions.
It realizes efficient scheduling in multi-charging requests and dynamic environments, ensures the response efficiency and resource utilization of the charging system, and avoids decision-making conflicts.
Smart Images

Figure CN119784115B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent agent scheduling method and scheduling device for a charging system, a storage medium, and a program product. Background Art
[0002] In recent years, as an environmentally friendly and sustainable means of transportation, electric vehicles have attracted extensive attention and achieved remarkable development globally. Electric vehicles are mainly divided into pure electric vehicles, plug-in hybrid vehicles, and hybrid vehicles, and the key lies in the progress of battery technology. With the development of lithium-ion battery technology, the driving range of electric vehicles has gradually increased, the charging time has gradually shortened, and the cost has gradually decreased. In addition, the motor efficiency and control system of electric vehicles are also constantly optimized, thereby improving the overall performance and driving experience.
[0003] In the past few years, the electric vehicle market has shown a rapid growth trend, mainly due to the support of government policies, especially the promotion of environmental protection and carbon emission reduction policies. Many countries have formulated subsidy policies and regulations to promote the popularization of electric vehicles. Especially in regions such as Europe and China, the sales volume of electric vehicles has increased significantly, and the market share has been continuously expanding. Automobile manufacturers have also launched new electric vehicle models one after another, and the market competition has become increasingly fierce. For example, as a pioneer in the field of electric vehicles, Tesla has continuously launched new models with high performance and long driving range, leading the development direction of the market.
[0004] However, as an important energy source for electric vehicles, the popularization and improvement of charging infrastructure are very necessary. For this reason, it is particularly important to build an autonomous charging system that can efficiently respond to charging requests and effectively schedule charging robots and mobile charging piles. Facing the increasing user charging demands, multiple charging requests often occur in a short period of time in a parking lot, and a single charging robot and mobile charging pile are difficult to meet the demands. Therefore, deploying multiple mobile charging devices can improve the overall charging demand response efficiency. In order to ensure the efficient operation of the multi-charging robot system and avoid conflicts, it is necessary to design an efficient and harmonious intelligent scheduling strategy.
[0005] Traditional strategies for scheduling heterogeneous multi-agents mostly design corresponding rules according to actual application problems, mainly including priority rule scheduling strategies (Fellir, Fadoua, et al. "A multi-Agent based model for task scheduling in cloud-fog computing platform." 2020 IEEE international conference on informatics, IoT, and enabling technologies (ICIoT). IEEE, 2020.) and multi-agent path finding strategies (Atzmon, Dor, et al. "Generalizing multi-agent pathfinding for heterogeneous agents." Proceedings of the International Symposium on Combinatorial Search. Vol. 11. No. 1. 2020.). The above rule-based strategies mainly rely on fixed and static environmental states, and for non-linear optimization problems, there are often disadvantages of large computational and communication overheads.
[0006] With the development of optimization techniques, advanced heuristic optimization algorithms have also been applied to the field of heterogeneous multi-agent scheduling. Zhou et al. combined the distributed multi-objective evolutionary algorithm and the greedy algorithm (Zhou, Jing, et al. "Task allocation for multi-agent systems based on distributed many-objective evolutionary algorithm and greedy algorithm." Ieee Access 8 (2020): 19306-19318.) to solve the optimal strategy for task allocation. However, the greedy algorithm lacks a global perspective during optimization and cannot handle the impacts brought about by dynamic environmental changes. Secondly, regarding the leader-follower formula problem, Wang et al. utilized particle swarm optimization as an effective evolutionary technique for distributed gain optimization in multi-agent networks (Wang, Xin, Dongsheng Yang, and Shuang Chen. "Particle swarm optimization based leader-follower cooperative control in multi-agent systems." Applied Soft Computing 151 (2024): 111130.), and ensured the stability of the closed-loop dynamics under a directed communication topology based on algebraic graph theory. Although particle swarm, as a heuristic algorithm with strong global search ability, is still prone to falling into local optima and incurring high computational costs when facing high-dimensional and complex heterogeneous multi-agent scheduling problems.
[0007] Facing the high-dimensional complex relationships and dynamic changing environment of heterogeneous multi-agent systems, the adaptive ability and reward mechanism of reinforcement learning can handle them well. Zhang et al. proposed an improved Q-value mixing matrix, combined with the multi-head attention mechanism of Qatten, which can adapt to heterogeneous robots operating asynchronously in different scenarios (Zhang, Han, et al. "Heterogeneous Multi-Robot Cooperation with Asynchronous Multi-Agent Reinforcement Learning." IEEE Robotics and Automation Letters (2023)). However, reinforcement learning requires a large amount of training data and time to converge, and although specific reward mechanisms can be designed to handle the heterogeneity of heterogeneous agents, it may still be difficult to achieve effective coordination and optimization. Summary of the Invention
[0008] Technical problem to be solved
[0009] In order to overcome the above problems in the prior art, the present invention aims to provide a scheduling method and device that can quickly and accurately schedule a multi-agent system to respond and solve its decision-making conflicts in the case of multiple charging requests in an intelligent parking lot and a heterogeneous multi-agent control system of charging robots and mobile charging piles.
[0010] Technical solution
[0011] According to a first aspect of the present invention, there is provided an agent scheduling method for a charging system, the charging system having a plurality of agents, the agents including charging robots and mobile charging piles, the scheduling method comprising: a charging system network topology construction step (S100) of constructing a network topology for a parking lot physical system and a heterogeneous multi-agent system of charging robots - mobile charging piles, the network topology including a static topology constructed based on parking lot spaces, robot rails, and charging vehicles, and a dynamic topology constructed based on charging robots, mobile charging piles, and communication links therebetween; a centralized task allocation step (S200) of, in response to a plurality of charging requests, modeling a task allocation problem as a cooperative game optimization model for the heterogeneous multi-agent system to perform centralized task allocation; a distributed path planning step (S300) of, with the allocated charging requests as the target, establishing a competition - cooperation game model based on a locally communicated fully connected graph constructed using the dynamic topology to perform distributed real-time path planning control; and a planning strategy design step (S400) of, based on game theory and using model predictive control, solving the game models in the centralized task allocation step and the distributed path planning step until convergence to a sequence of Nash equilibrium solutions satisfying conditions, and outputting a planning control sequence within a prediction horizon as a control vector.
[0012] As a preferred solution, in the agent scheduling method for a charging system according to the present invention, in the centralized task allocation step, the task allocation problem is modeled as the following cooperative game optimization model:
[0013]
[0014] Wherein, N r and N c respectively represent the numbers of charging robots and mobile charging piles in the system, N task represents the number of charging requests existing in the system at the current moment, X i,jIt represents the decision result of agent i being assigned to execute task j. The first-line constraint means that each charging request needs to be responded to by at most one heterogeneous agent formation of charging robots and mobile charging piles respectively. The second-line constraint means that each agent has exactly one charging task assignment at present. h(X) and g(X) respectively represent the equations or inequalities related to power storage, vehicle demand, and robotic arm structure, etc., u allocation represents the decision utility function.
[0015] As a preferred solution, for the agent scheduling method of the charging system according to the present invention, in the planning strategy design step, the cooperative game optimization model established in the centralized task assignment step is modified as follows to make the assigned tasks stable:
[0016]
[0017] h(X) = 0
[0018] g(X) ≤ 0,
[0019] where, N r and N c respectively represent the numbers of charging robots and mobile charging piles in the system, N task represents the number of charging requests existing in the system at the current moment, X i,j represents the decision result of agent i being assigned to execute task j. The first-line constraint means that each charging request needs to be responded to by at most one heterogeneous agent formation of charging robots and mobile charging piles respectively. The second-line constraint means that each agent has exactly one charging task assignment at present. h(X) and g(X) represent the equations or inequalities related to power storage, vehicle demand, and robotic arm structure, etc., represents the decision utility function at the next moment.
[0020] As a preferred solution, for the agent scheduling method of the charging system according to the present invention, the Nash equilibrium solution of the cooperation-competition game model in the distributed path planning step is calculated, and the expression of the optimization problem is obtained as:
[0021]
[0022] s.t. h i (π i ) = 0
[0023] g i (π i ) ≤ 0
[0024]
[0025] where, min Ui(π i, π -i ) represents the strategy π of agent i i and the strategies π of other agents -i that constitute the decision utility function, where α ij represents the correlation between agent i and agent j, s.t. h i (π i ) and g i (π i ) represent the equality and inequality constraints related only to the agent itself included in the planning strategy design step, γ ij (π i , π j ) represents the joint constraints of the generalized Nash equilibrium that depend on the decisions of other agents, which include collision constraints, task assignment constraints, etc. between two agents, and include the constraint conditions related to the decisions of the remaining agents in the optimization problem set up for agent i in the planning strategy design step,
[0026] And, iterate and optimize until converging to the Nash equilibrium solution that satisfies the KKT conditions.
[0027] As a preferred solution, for the agent scheduling method of the charging system according to the present invention, the charging system network topology construction step includes:
[0028] Static topology establishment step (S110): Construct a static physical topology map of the parking lot based on the relevant physical information of the parking lot,
[0029] Among them, in the static physical topology of the parking lot, the vertices have the following attributes:
[0030]
[0031] Among them represents the value of the parking space vertex in the static topology, r i represents whether there is an unresponded charging request for the current parking space, T i represents the waiting time after the electric vehicle owner places a charging request on the current parking space, t represents the current time, and t0 represents the initial time when the charging request appears. On this basis, the edge value of the static topology is represented by the guide rail length between two vertices;
[0032] Dynamic topology establishment step (S120), the dynamic topology is used to represent the positional relationship and communication connectivity of heterogeneous multi-agents of charging robots - mobile charging piles, and the information included in the vertices of the heterogeneous multi-agents includes the following attributes:
[0033]
[0034] Among them, (x i,t , y i,t) represents the physical position of the current agent in the parking lot coordinate system, task i represents the charging task assigned to the current agent; c i represents the energy consumption coefficient of the agent, b i represents the remaining power of the current agent; while s i represents the special attribute related to the vertex type, υ i represents the route along which the current robot moves forward,
[0035] Among them, the edge e i,j has the following communication intensity:
[0036]
[0037] Among them, d i,j represents the actual distance between two agents, d cs represents the distance threshold for communication. When the distance is greater than the threshold, the communication link between the two is disconnected. "!" and "&" are logical operators. Among them, "!" represents the NOT operation, and "&" represents the AND operation.
[0038] As a preferred solution, for the agent scheduling method of the charging system according to the present invention, the centralized task assignment construction steps include:
[0039] Task set definition step (S210), defining the following task set:
[0040]
[0041] Among them, the task has the following attributes:
[0042]
[0043] Among them, C i represents the power requirement for the current charging task, L i represents the position of the charging port of the current task vehicle. After the task appears, the camera directly above the parking space identifies the vehicle information to obtain C i , L i The specific information of; P i represents the priority of the current task. The higher the priority, the smaller P i ; PV i , PR i represents the pose of the vehicle in the parking space and the label of the vertex where the parking space is located in the static topology;
[0044] Task assignment matrix definition step (S220): Define the following task assignment matrix: Among them, X i,jIndicates the decision result that agent i is assigned to execute task j:
[0045]
[0046] In the utility function design step (S230), the utility function for each agent i is designed. The definition of the utility after agent i executes task j is:
[0047]
[0048] The definition of the joint utility function is:
[0049]
[0050] The decision utility function is:
[0051]
[0052] Constraint condition design step (S240): For the charging robot, constraint conditions are designed. The inverse kinematic equation is used to solve the end-effector target pose obtained by calculating based on the vehicle pose and the charging port position:
[0053]
[0054] p target = g(L j , PV j ),
[0055] where f -1 (·) represents the inverse kinematics of the carried robotic arm, and g(·) represents the position equation for calculating the end-effector based on the current pose of the vehicle and the prior charging port position information, so as to obtain the optimization model of the task assignment centralized cooperation game model.
[0056] As a preferred solution, according to the agent scheduling method of the charging system of the present invention, the distributed path planning step includes:
[0057] The definition step of participants and strategies (S310), the participants include charging robots and mobile charging piles, and the strategy space includes:
[0058] Π i =(a i ),
[0059] where a i is a vector, and there are two directions, positive and negative in the x direction and positive and negative in the y direction;
[0060] The utility function design step (S320), designing the utility function, which includes:
[0061]
[0062] Among them, respectively represent the competitive utility function and the cooperative utility function.
[0063] As a preferred solution, for the agent scheduling method of the charging system according to the present invention, the utility function includes a cooperative utility function, and the cooperative utility function is designed to penalize the situation where the distance between two agents is close or they are moving towards each other for the group of agents connected by the communication link constructed near agent i. The cooperative utility function is as follows:
[0064]
[0065] where N cs represents the agents having a communication link with i.
[0066] As a preferred solution, for the agent scheduling method of the charging system according to the present invention, the utility function includes a competitive utility function, and the competitive utility function is designed to stimulate each agent to complete the assigned task at the fastest speed and reduce the time difference in reaching the task location between heterogeneous agents for the same task. The competitive utility function is as follows:
[0067] And
[0068] The final system competitive utility function is as follows:
[0069] x i,k = 1, x j,k = 0.
[0070] According to a second aspect of the present invention, there is provided an agent scheduling device for a charging system, the charging system having a plurality of agents, the agents including a charging robot and a mobile charging pile, and the scheduling device including: a charging system network topology construction unit that constructs a network topology for a parking lot physical system and a charging robot-mobile charging pile heterogeneous multi-agent system, the network topology including a static topology constructed based on parking lot spaces, robot rails, and charging vehicles, and a dynamic topology constructed based on charging robots, mobile charging piles, and communication links therebetween; a centralized task allocation unit that, in response to a plurality of charging requests, models the task allocation problem as a cooperative game optimization model for the heterogeneous multi-agent system to perform centralized task allocation; a distributed path planning unit that, with the allocated charging requests as the target, establishes a competition-cooperation game model based on a locally communicated fully connected graph constructed using the dynamic topology to perform distributed real-time path planning control; and a strategy design unit that, based on game theory and using model predictive control, solves the game models established by the centralized task allocation unit and the distributed path planning unit until converging to a sequence of Nash equilibrium solutions that satisfy the conditions, and outputs a planned control sequence within the prediction horizon as a control vector.
[0071] According to a third aspect of the present invention, there is provided a non-transitory storage medium storing a computer program that, when executed by a processor, can implement the agent scheduling method for a charging system according to the first aspect of the present invention.
[0072] According to a fourth aspect of the present invention, there is provided a computer program product including computer instructions that, when executed by a processor, can implement the agent scheduling method for a charging system according to the first aspect of the present invention.
[0073] Beneficial technical effects
[0074] The scheduling method of the charging system of the present invention can, on the basis of constructing a two-layer topology of a static topology and a dynamic topology, better capture the relationships between multi-agents using a cooperative game model and a cooperation-competition game model, and solve for the Nash equilibrium of the optimal planned control decision in an optimized and predictive combination manner of model predictive control.
[0075] By the following description of exemplary embodiments with reference to the accompanying drawings, other features of the present invention will become clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is a flowchart showing the agent scheduling method of the charging system;
[0077] Figure 2 is a flowchart showing the charging system network topology construction steps;
[0078] Figure 3 is a flowchart showing the steps of centralized task allocation construction;
[0079] Figure 4 is a flowchart showing the steps of distributed path planning;
[0080] Figure 5 is a schematic diagram showing the establishment of a static topology;
[0081] Figure 6 is a schematic diagram showing the establishment of a dynamic topology;
[0082] Figure 7 is a schematic diagram showing task allocation and local trajectory planning control. Detailed implementation manners
[0083] Exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangements of components, numerical representations, and numerical values described in these embodiments do not limit the scope of the present invention.
[0084] The scheduling method of the automatic charging system of the present invention can be implemented by a processor in the automatic charging system executing a computer program stored on a memory in the automatic charging system. As an alternative solution, it can also be implemented by the automatic charging system communicating with a server, and a processor in the server executing a computer program stored on the server and real-time feedback the program execution result to the automatic charging system.
[0085] In the present invention, the term "unit" may refer to a software environment, a hardware environment, or a combination of a software and a hardware environment. In a software environment, the term "unit" refers to a functionality, an application, a software module, a function, a routine, a set of instructions, or a program that can be executed by a programmable processor, such as a microprocessor, a central processing unit (CPU), or a specially designed programmable device, or a controller. The memory contains instructions or programs that, when executed by the CPU, cause the CPU to perform operations corresponding to the unit or function. In a hardware environment, the term "unit" refers to a hardware element, a circuit, a component, a physical structure, a system, a module, or a subsystem. According to a specific embodiment, the term "unit" may include mechanical, optical, or electrical components, or any combination thereof. The term "unit" may include active (e.g., transistors) or passive (e.g., capacitors) components. The term "unit" may include a semiconductor device having a substrate and other material layers with various conductive concentrations. It may include a CPU or a programmable processor that can execute a program stored in the memory to perform a specified function. The term "unit" may include logic elements (e.g., AND, OR) implemented by transistor circuits or any other switching circuits. In a combination of a software and a hardware environment, the term "unit" or "circuit" refers to any combination of the software and hardware environments as described above. In addition, the terms "element", "component", "part", or "device" may also refer to a "circuit" integrated or not integrated with a packaging material.
[0086] Figure 1 is a flowchart for illustrating an agent scheduling method of a charging system according to an embodiment.
[0087] The following will refer to Figure 1 illustrate a flowchart of an agent scheduling method of a charging system according to an embodiment.
[0088] First, in step S100, a charging system network topology construction is performed. Specifically, a network topology is constructed for the parking lot physical system and the charging robot-mobile charging pile heterogeneous multi-agent system. The network topology includes a static topology constructed based on the parking lot spaces, the robot rails, and the charging vehicles, and a dynamic topology constructed based on the charging robots, the mobile charging piles, and the communication links between them.
[0089] The following will refer to Figure 2 the flowchart shown to illustrate the process of the charging system network topology construction step S100.
[0090] In step S110, a static topology is established. Figure 5 Schematically shows a schematic diagram of the static topology construction. In Figure 5 a solid circle represents the occupancy of a parking space. A static parking lot physical topology map is constructed based on the relevant physical information of the parking lot, including the existing N sThe vertex set composed of parking spaces The hanging charging robot and the guide rail of the mobile charging pile are used as edges A continuous section of the guide rail edge is segmented by the parking space vertices Indicates at the parking space vertex and The two-way passable guide rail part between them. All superscripts s represent the relevant structures or value parameters in the static network topology. Then the physical static topology of the parking lot can be expressed as G s =(V s , E s ).
[0091] In the physical static topology of the parking lot, the vertices mainly include the following attributes
[0092]
[0093] Among them, T i Represents the waiting time after the electric vehicle owner on the current parking space places an order for a charging request. On this basis, the edge value of the static topology is represented by the guide rail length between two vertices
[0094] In step S120, a dynamic topology is established Figure 6 Schematically shows the schematic diagram of the dynamic topology construction. The dynamic topology is used to represent the positional relationship and communication connection degree of the heterogeneous multi-agent of the charging robot - mobile charging pile. Then the dynamic topology G d mainly includes charging robots and the heterogeneous agent vertex set of mobile charging piles Among them, the type of the vertex will be described in the attributes. The edges of the dynamic topology G d are different from the static topology. It is represented by whether the communication line is connected, that is, it will change dynamically with the movement of the agent and there will be changes in the dynamic topology structure
[0095] The information contained in the heterogeneous agent vertex includes
[0096]
[0097] Among them, (x i,t , y i,t ) represents the physical position of the current agent in the parking lot coordinate system, task i represents the charging task assigned to the current agent. A charging task assigns a charging robot and a mobile charging pile formation using the method described in step (2), and is represented by the label of the corresponding parking space vertex; c i represents the energy consumption coefficient of the agent, indicating the electric energy consumed per unit time, b irepresents the remaining power of the current agent; and s i represents the special attributes related to the vertex type. When the agent is a charging robot, it represents the relevant parameters of the robotic arm, and when it is a mobile charging pile, it represents the maximum charging power that the charging pile can achieve. υ i represents the route that the current robot is moving forward.
[0098] Next, for the edge values, that is, the communication links, design is carried out. Define the vertices nearby d cs within which vertices can establish communication with it, that is, there is an edge between the two; on this basis, considering the heterogeneous characteristics of the dynamic topology vertices, a communication link will be constructed between the charging robot and the mobile charging pile that are assigned to the same vehicle charging task in step (2), and both communication links are undirected edges, and the situation of communication loss or delay is not considered. Finally, define the communication strength of the edge e i,j as:
[0099]
[0100] where d i,j represents the actual distance between the two agents, and d cs represents the distance threshold at which communication can be carried out. When the distance is greater than the threshold, the communication link between the two is disconnected. Overall, it means that the communication links in the dynamic topology only exist between isomorphic agents with a distance less than the threshold or heterogeneous agents with the same task.
[0101] Next, proceed to step S200. In step S200, centralized task allocation is performed. In response to multiple charging requests, for the heterogeneous multi-agent system, the task allocation problem is modeled as a cooperative game optimization model to perform centralized task allocation.
[0102] Next, the centralized task allocation steps will be described in detail with reference to Figure 3 Detailed description of the centralized task allocation steps.
[0103] In step S210, first define the N task task set:
[0104]
[0105] where the task has the following attributes:
[0106]
[0107] where C i represents the power requirement needed for the current charging task, and L iIndicates the position of the charging port of the current task vehicle. After the task appears (i.e., after the user initiates a charging request), the camera directly above the parking space identifies the vehicle information to obtain C i , L i specific information; P i Indicates the priority of the current task, which is mainly judged by the user level of the charging request initiator. The higher the priority, the smaller P i ; PV i , PR i Indicates the pose of the vehicle in the parking space and the label of the vertex where the parking space is located in the static topology.
[0108] Then, in step S220, define the task assignment matrix where X i,j represents the decision result of agent i being assigned to execute task j:
[0109]
[0110] Subsequently, in step S230, design the utility function. Specifically, design the utility function for each agent i. First, define the utility after agent i executes task j:
[0111]
[0112] Since each charging task requires a heterogeneous agent formation to complete, if there is a large time difference in the arrival times among heterogeneous agents, it will cause resource waste. Therefore, define the joint utility function:
[0113]
[0114] In the agent numbering, the front part represents the charging robot, and the back part represents the mobile charging pile. The finally obtained decision utility function is as follows:
[0115]
[0116] Finally, in step S240, design the corresponding constraint conditions according to the designed utility function and system topology above. First, it is necessary to ensure that the robot has sufficient power storage to travel to the task-assigned parking space, which is represented by the following formula:
[0117]
[0118] Design the constraint conditions for the mobile charging pile, then the maximum power supported by the charging equipment it carries needs to meet the requirements of the vehicle with the charging request, which is represented by the following formula:
[0119] x i,j ·(s i-C i ) ≥ 0,
[0120] Design the constraint conditions for the charging robot. The robotic arm it carries needs to be able to use the parking space coordinates as the base of the robotic arm and solve the inverse kinematic equation to reach the target pose of the end effector obtained by calculating the vehicle pose and the charging port position:
[0121]
[0122] p target = g(L j , PV j ),
[0123] where f -1 (·) represents the inverse kinematics of the robotic arm carried, and g(·) represents the position equation of the end effector calculated based on the current pose of the vehicle and the prior charging port position information. Obtain the optimization model of the task allocation centralized cooperative game model:
[0124]
[0125] Among them, constraint (a) ensures that each agent currently has at most one charging task assignment, constraint (b) ensures that each charging request currently has at most one charging robot - mobile charging pile formation assigned to respond, and constraints (c) and (d) mainly represent the three constraints defined above, which include power storage, charging requests, and the structural constraints of the robotic arm itself.
[0126] After completing the centralized task allocation step, the process enters step S300.
[0127] In step S300, distributed path planning is performed. Specifically, with the assigned charging request as the goal, based on the local communication fully connected graph constructed using dynamic topology, a competition - cooperation game model is established for distributed real - time path planning control.
[0128] Next, it will refer to Figure 4 Describe the distributed path planning steps in detail.
[0129] In step S310, the participants and strategies are defined. The participants mainly include charging robots and mobile charging piles. The strategy space mainly includes:
[0130] Π i = (a i ),
[0131] where a i is a vector. Due to the limitation of the robot guide rail, there are only two directions of positive and negative x and positive and negative y.
[0132] In step S320, a utility function is designed. Specifically, the utility function is represented by the following formula:
[0133]
[0134] where represent the competitive utility function and the cooperative utility function respectively.
[0135] For the cooperative utility function, an abstract concept of "congestion" is introduced. For the group of agents connected by the communication links constructed near agent i, the cases where the distance between two agents is too close or they are moving towards each other are punished, enabling heterogeneous multi-agent systems with small-range communication capabilities to cooperate to achieve local obstacle avoidance. The following cooperative utility function is designed:
[0136]
[0137] where N cs represents the agents having communication links with i, excluding heterogeneous agents for the same charging task. The design of the cooperative utility function is for local obstacle avoidance.
[0138] For the competitive utility function, based on the cooperative utility function, it is easy to obtain that all agents choose the stop strategy to optimize the cooperative utility function, which leads to the failure of task response. Therefore, a competitive utility function is designed to stimulate each agent to complete the assigned task at the fastest speed and, on this basis, be able to reduce the time difference in reaching the task location between heterogeneous agents for the same task. It is represented by the following formula:
[0139]
[0140] The final system competitive utility function is represented by the following formula:
[0141] x i,k = 1, x j,k = 0,
[0142] Subsequently, the process enters step S400. In step S400, a planning strategy is designed. Specifically, based on game theory and using model predictive control, the game models in the centralized task allocation step and the distributed path planning step are solved until convergence to a sequence of Nash equilibrium solutions that meet the conditions, and the planning control sequence within the prediction horizon is output as the control vector.
[0143] Next, step S400 will be described with reference to Figure 7 Figure 7 It is a schematic diagram showing task allocation and local trajectory planning control. Considering the stability of the allocated tasks, the following centralized task allocation cooperative game optimization model is obtained by modifying the cooperative game optimization model:
[0144]
[0145] h(X) = 0
[0146] g(X) ≤ 0,
[0147] For the Nash equilibrium solution of the cooperative-competitive game model, the following optimization problem is set up for agent i:
[0148]
[0149] x i,k = 1, x j,k = 1
[0150]
[0151] The establishment of the competitive and cooperative game model can help the agent avoid obstacles to possible collisions in the path on the premise of ensuring its own rapid arrival at the target location. Define the equality constraint as h i (π i ) = 0, the inequality constraint g i (π i ) ≤ 0, and the constraints involving other agents are On this basis, the optimal strategy of agent i is obtained as The optimal strategies of the remaining agents are defined as Thus, the generalized Nash equilibrium is defined as:
[0152]
[0153] In the optimization problem, the objective function of each agent only depends on its own strategy optimization. However, due to the existence of the competitive game model, the strategies π -i of other agents also have an impact on π i . To capture this characteristic, the optimal strategies of the remaining robots are added as sensitive terms to the objective function, resulting in the following formula:
[0154]
[0155] Among them, is the utility of the best strategy of agent j based on the strategy of agent i, and α i,j is the recognition parameter of the relationship between the robots:
[0156]
[0157] The established model is a static game, that is, when agent i makes an optimal decision, it cannot perceive the decisions of the other agents. However, due to the existence of the communication link, the final decision of the previous moment can be known. Therefore, while solving the problem of one agent, the strategies of other agents in the previous iteration are frozen and used, and the approximate Nash equilibrium is calculated using the Gauss-Seidel iteration best response method. In the k-th iteration optimization, Use The first-order Taylor expansion of the neighborhood is used for approximation:
[0158]
[0159] where is a constant term, Using the Karush Kuhn Tucker (KKT) conditions and simple derivation, the following formula is obtained:
[0160]
[0161] where represents the Lagrange multiplier related to the constraints involving other agents, and the final expression form of the optimization problem for distributed local path planning is:
[0162]
[0163] s.t. h i (π i ) = 0
[0164] g i (π i ) ≤ 0
[0165]
[0166] where, minUi(π i , π -i ) represents the decision utility function composed of the strategy π i of agent i and the strategy π -i of other agents. Among them, α ij represents the correlation between agent i and agent j. s.t.h i (π i ) and g i (π i ) represent the equality and inequality constraints related only to the agent itself included in step S400. γ ij (π i , π j) represents the joint constraints of the generalized Nash equilibrium that depend on other subject decisions, which include collision constraints, task assignment constraints, etc. between two agents, and include the constraint conditions related to the decisions of the remaining agents involved in the optimization problem set for agent i in step S400.
[0167] So far, iterative optimization is carried out until it converges to the Nash equilibrium solution that satisfies the KKT conditions.
[0168] The scheduling method of the charging system of the present invention can, on the basis of constructing a two-layer topology of a static topology and a dynamic topology, better capture the relationships between multiple agents by using a cooperative game model and a cooperative-competitive game model, and solve for the Nash equilibrium of the optimal planning control decision in an optimized and predictive combination manner of model predictive control.
[0169] Other embodiments
[0170] Embodiments of the present invention can also be implemented by a computer of a system or device that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium (which can also be more fully referred to as a "non-transitory computer-readable storage medium") to perform one or more of the functions in the above embodiments, and / or includes one or more circuits (e.g., an application-specific integrated circuit (ASIC)) for performing one or more of the functions in the above embodiments, and the embodiments of the present invention can be implemented by a method of, for example, the computer of the system or device reading and executing the computer-executable instructions from the storage medium to perform one or more of the functions in the above embodiments, and / or controlling the one or more circuits to perform one or more of the functions in the above embodiments. The computer can include one or more processors (e.g., a central processing unit (CPU), a microprocessing unit (MPU)), and can include a network of separate computers or separate processors to read and execute the computer-executable instructions. The computer-executable instructions can be provided to the computer, for example, from a network or the storage medium. The storage medium can include, for example, a hard disk, a random access memory (RAM), a read-only memory (ROM), a memory of a distributed computing system, an optical disc (such as a compact disc (CD), a digital versatile disc (DVD), or a Blu-ray disc (BD) TM )), a flash device, and a memory card, etc.
[0171] Embodiments of the present invention can also be implemented by the following method, that is, by providing software (program) that performs the functions of the above embodiments to a system or device through a network or various storage media, and the method of the computer or the central processing unit (CPU), the microprocessing unit (MPU) of the system or device reading and executing the program.
[0172] Although the present invention has been described with reference to exemplary embodiments, it should be understood that the present invention is not limited to the disclosed exemplary embodiments. The scope of the appended claims should be given the broadest interpretation so as to cover all such variations and equivalent structures and functions.
Claims
1. An intelligent agent scheduling method for a charging system, wherein the charging system has multiple intelligent agents, the intelligent agents include charging robots and mobile charging piles, and the scheduling method includes: The charging system network topology construction step (S100) constructs a network topology for the parking lot physical system and the charging robot-mobile charging pile heterogeneous multi-agent system, wherein the network topology includes a static topology based on parking spaces, robot rails, and charging cars, and a dynamic topology based on the charging robot, mobile charging pile, and the communication links between them; A centralized task allocation step (S200), in response to multiple charging requests, for a heterogeneous multi-agent system, modeling the task allocation problem as a cooperative game optimization model to perform centralized task allocation; The distributed path planning step (S300) takes the allocated charging request as the target, establishes a competition-cooperation game model based on the local communication fully connected graph constructed by using the dynamic topology, and performs distributed real-time path planning control; The planning strategy design step (S400) is based on game theory and uses model predictive control to solve the game models in the centralized task allocation step and the distributed path planning step until they converge to a Nash equilibrium solution sequence that minimizes the objective function in the corresponding model predictive control optimization model, and outputs the planning control sequence within the prediction field of view as a control vector. Among them, in the planning strategy design step, the cooperative game optimization model established in the centralized task allocation step is modified as follows to make the allocation task stable: Among them, N r and N c Respectively represent the number of charging robots and mobile charging piles in the system, N task Indicates the number of charging requests in the system at the current moment, X i,j represents the decision result of agent i being assigned to perform task j. The first row of constraints means that each charging request needs to be responded to by at most one and only one charging robot and mobile charging pile heterogeneous agent formation, respectively. The second row of constraints means that each agent currently has only one and only one charging task assigned. h(X) and g(X) represent equality or inequality constraints on power storage, vehicle requirements, and robot arm structure. Represents the decision utility function at the next moment.
2. The scheduling method according to claim 1, wherein: In the centralized task allocation step, the task allocation problem is modeled as the following cooperative game optimization model: Among them, N r and N c Respectively represent the number of charging robots and mobile charging piles in the system, N task Indicates the number of charging requests in the system at the current moment, X i,j represents the decision result of agent i being assigned to perform task j. The first row of constraints means that each charging request needs to be responded to by at most one and only one charging robot and mobile charging pile heterogeneous agent formation, respectively. The second row of constraints means that each agent currently has only one and only one charging task assigned. h(X) and g(X) represent the equality or inequality constraints on power storage, vehicle demand, and robot arm structure, respectively. u allocation represents the decision utility function.
3. The scheduling method according to claim 1, wherein: The Nash equilibrium solution of the cooperation-competition game model in the distributed path planning step is calculated, and the expression of the optimization problem is obtained as follows: s.t.h i (π i )=0 g i (π i )≤0 Among them, minUi(π i ,π -i ) represents the strategy π of agent i i and the strategies π of other agents -i The decision utility function is composed of ij represents the correlation between agent i and agent j, sth i (π i ) and g i (π i ) represents the equality and inequality constraints related only to the agent itself contained in the planning strategy design step, γ ij (π i ,π j ) represents the joint constraints of the generalized Nash equilibrium that depend on the decisions of other subjects, which include the collision constraints between the two agents, the task allocation constraints, and the constraints related to the decisions of the remaining agents in the optimization problem set up for agent i in the planning strategy design step. Furthermore, the optimization is iterated until it converges to a Nash equilibrium solution that satisfies the KKT condition.
4. The scheduling method according to claim 1, wherein: The steps to build the charging system network topology include: Static topology building step (S110): constructing a static parking lot physical topology map based on the relevant physical information of the parking lot. Among them, in the parking lot physical static topology, the vertex contains the following attributes: in Represents the value of the parking space vertex in the static topology, r i Indicates whether there is any uncompleted charging request in the current parking space, T i represents the waiting time after the electric car owner in the current parking space places a charging request, t represents the current time, and t0 represents the initial time when the charging request appears. On this basis, the edge value of the static topology is represented by the length of the rail between the two vertices; The dynamic topology establishment step (S120) is used to represent the positional relationship and communication connectivity of the heterogeneous multi-agents of the charging robot and the mobile charging pile. The information contained in the vertices of the heterogeneous multi-agents includes the following attributes: in Indicates the value contained in the vertex in the dynamic topology, type i Indicates that the current vertex belongs to a charging robot or a mobile charging pile, (x i,t ,y i,t ) represents the current physical position of the agent in the parking lot coordinate system, task i Indicates the charging task currently assigned to the agent; c i represents the energy consumption coefficient of the agent, b i represents the remaining power of the current agent; and s i Indicates special attributes related to the vertex type, v i Indicates the current route of the robot. Among them, edge e i,j The communication strength is: Among them, d i,j represents the actual distance between the two agents, d cs Indicates the distance threshold for communication. When the distance is greater than the threshold, the communication link between the two is disconnected.
5. The scheduling method according to claim 1, wherein: The steps of centralized task allocation include: The task set definition step (S210) defines the following task set: The tasks It has the following properties: Among them, C i Indicates the power requirement for the current charging task, L i Indicates the location of the charging port of the current mission vehicle. After the mission occurs, the camera directly above the parking space recognizes the vehicle information to obtain C i ,L i Specific information of P i Indicates the priority of the current task. The higher the priority, the higher the P i The smaller the PV i , P.R. i Indicates the position of the vehicle in the parking space and the number of the vertex where the parking space is located in the static topology; Task allocation matrix definition step (S220): define the following task allocation matrix: Where X i,j Indicates the decision result of agent i being assigned to perform task j: The utility function design step (S230) designs the utility function of each agent i. The utility of agent i after performing task j is defined as: The joint utility function is defined as: The decision utility function is: Among them, x prj and prj Indicates the coordinate position of task j, x i0 ,y i0 represents the coordinate position of the charging robot in the heterogeneous agent formation, x i1 ,y i1 represents the location of the mobile charging pile in the heterogeneous agent formation, represents the joint utility function of heterogeneous agent formation i when performing task j, i includes charging robot i0 and mobile charging pile i1, P i Indicates the priority of the current task, b i Indicates the remaining power of the current agent, c i represents the energy consumption coefficient of the agent, v i0 , v i1 represents the current speed of the charging robot i0 and the mobile charging pile i1 in the heterogeneous agent formation, Constraint design step (S240): Design constraints for the charging robot, and use the inverse kinematics equation to solve the end effector target posture calculated based on the vehicle posture and the charging port position: p target =g(L j ,PV j ), Among them, f -1 (·) represents the inverse kinematics of the carried robotic arm, g(·) represents the position equation of the end effector calculated according to the current posture of the vehicle and the prior information of the charging port position, to obtain the optimization model of the centralized cooperative game model of task allocation, p targrt Indicates the 6-dimensional position and attitude information of the charging port, s i Indicates special attributes related to the vertex type, x prj , y prj Represents the position information PV of the jth vertex in the overall topology i , P.R. i Represents the position of the vehicle in the parking space and the number of the vertex where the parking space is located in the static topology.
6. The scheduling method according to claim 1, wherein: The distributed path planning steps include: Participants and strategy definition step (S310), participants include A charging robot and Mobile charging stations, the strategy space includes: P i =(a i ), where a i It is a vector, and has two directions: positive and negative x and positive and negative y; The utility function design step (S320) designs a planning utility function, which includes: in, They represent the competitive utility function and the cooperative utility function respectively.
7. The scheduling method according to claim 5, wherein: The utility function includes a cooperative utility function, which is designed to punish the situation where two agents are close to each other or are moving in opposite directions for the group of agents connected by the communication link constructed near agent i. The cooperative utility function is as follows: Among them, N cs Represents the intelligent agent that has a communication link with i.
8. The scheduling method according to claim 5, wherein: The utility function includes a competitive utility function, which is designed to stimulate each agent to complete the assigned task as quickly as possible and reduce the time difference between heterogeneous agents of the same task and the time difference in arriving at the task location. The competitive utility function is as follows: and The final system competition utility function is as follows:
9. An intelligent agent scheduling device for a charging system, the charging system having a plurality of intelligent agents, the intelligent agents including a charging robot and a mobile charging pile, the scheduling device comprising: The charging system network topology construction unit builds the network topology for the parking lot physical system and the charging robot-mobile charging pile heterogeneous multi-agent system. The network topology includes a static topology based on parking spaces, robot rails and charging cars, and a dynamic topology based on charging robots, mobile charging piles and the communication links between them. A centralized task allocation unit, which responds to multiple charging requests, models the task allocation problem as a cooperative game optimization model for a heterogeneous multi-agent system to perform centralized task allocation; The distributed path planning unit takes the assigned charging request as the target and establishes a competition-cooperation game model based on the local communication fully connected graph constructed using dynamic topology to perform distributed real-time path planning control; as well as The planning strategy design unit, based on game theory and using model predictive control, solves the game model established by the centralized task allocation unit and the distributed path planning unit until it converges to a Nash equilibrium solution sequence that meets the conditions, and outputs the planning control sequence within the prediction field of view as the control vector. Among them, the cooperative game optimization model established in the planning strategy design unit for the centralized task allocation step is modified as follows to make the allocation task stable: Among them, N r and N c Respectively represent the number of charging robots and mobile charging piles in the system, N task Indicates the number of charging requests in the system at the current moment, X i,j represents the decision result of agent i being assigned to perform task j. The first row of constraints means that each charging request needs to be responded to by at most one and only one charging robot and mobile charging pile heterogeneous agent formation, respectively. The second row of constraints means that each agent currently has only one and only one charging task assigned. h(X) and g(X) represent equality or inequality constraints on power storage, vehicle requirements, and robot arm structure. Represents the decision utility function at the next moment.
10. A non-temporary storage medium storing a computer program, which, when executed by a processor, can implement the control method for heterogeneous multi-agent collaborative operations according to any one of claims 1-8.
11. A computer program product, comprising computer instructions, which, when executed by a processor, can implement the intelligent agent scheduling method for a charging system according to any one of claims 1-8.
Citation Information
Patent Citations
Cooperative game-based multi-park energy scheduling optimization method for comprehensive energy system
CN112465240A
Nash equilibrium specified time search method for intra-group decision consistent multi-group game
CN114488802A