Charging Robot Scheduling Method and Storage Medium

By introducing user priority rules and adaptive probability inference framework into the charging robot scheduling method, combined with model-based reinforcement learning, the problem of low scheduling efficiency is solved, and more robust and efficient charging request prediction and scheduling is achieved.

CN119443728BActive Publication Date: 2025-05-30ELU TECHNOLOGY HOLDINGS (ZHEJIANG)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510020235.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-30
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

In the face of multiple charging requests, it is difficult for the prior art to effectively schedule charging robots, resulting in low response efficiency. Traditional scheduling algorithms rely on expert knowledge and cannot cope with unknown situations.

Method used

A charging robot scheduling method based on user priority rules and sequence hierarchical adaptive probability inference framework was designed. Combined with model-based reinforcement learning, a world state derivative and reward model is established to improve training efficiency and enhance generalization ability.

Benefits of technology

It realizes that long-sequence decomposition prediction of future charging requests is carried out on the basis of taking into account the priority of each charging request, and captures the seasonal and trendy laws of charging requests, which improves the robustness of scheduling, and improves the training efficiency and testing effect of the overall scheduling strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119443728B_ABST
    Figure CN119443728B_ABST
Patent Text Reader

Abstract

The present invention provides a charging robot scheduling method and a storage medium. The charging robot scheduling method performs intelligent scheduling of charging robots according to the priorities of charging requests and the prediction of new charging requests for multiple charging requests in an intelligent parking lot, and includes the following steps: constructing a priority potential field of a parking space - charging system based on the multiple charging requests of users; modeling the multi - target tracking and priority intelligent scheduling process as a Markov decision process based on the constructed priority potential field of the parking space - charging system; designing a model - based reinforcement learning network based on the modeled Markov decision process, where the reinforcement learning network includes a dynamic network, a value - policy network, and a target network; and realizing parallel training of multiple networks by designing the interaction mode between the models of the dynamic network and the policy - value network and the collected data set based on the constructed reinforcement learning network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a scheduling method for multi-object single-agent, and particularly to a charging robot scheduling method and a storage medium. Background Art

[0002] With the popularization of electric vehicles, the demand for charging infrastructure is also growing rapidly. Urban and transportation planners are increasing their investment in the construction of charging piles to meet the growing charging needs of electric vehicles. With the gradual increase in the demand for charging equipment, intelligent charging systems can slow down the total demand for equipment while improving charging efficiency, and popularize the charging method of electric vehicles, enabling more electric vehicle owners without professional charging knowledge to achieve automatic charging with the help of intelligent systems. Therefore, mobile robots are used in the market to respond to charging requests, move to designated parking spaces, and operate charging equipment to complete the charging task of electric vehicles. However, when facing multiple charging requests, this charging mode requires an intelligent scheduling algorithm to determine the order of satisfaction of charging requests.

[0003] Traditional intelligent scheduling algorithms for robots to respond to multiple targets mainly include first-come, first-served (scheduling according to the order of arrival of charging requests), shortest job first (prioritizing tasks with the shortest execution time), priority scheduling (allocating resources according to the priority of tasks, and high-priority tasks are executed first), and preemptive scheduling (allowing higher-priority tasks to interrupt the currently executing task to ensure that urgent tasks can be responded to in a timely manner). However, traditional scheduling usually relies on expert knowledge to design scheduling rules for different situations to achieve scheduling, and cannot make correct countermeasures for unknown and unconsidered situations.

[0004] Advanced heuristic optimization algorithms have also been widely used in the field of intelligent scheduling. For example, genetic algorithms (see Non-Patent Document 1) and particle swarm algorithms (see Non-Patent Document 2) transform the scheduling problem into an optimization problem of designing an adaptation function. However, these algorithms have problems such as being prone to falling into local optimal solutions and being sensitive to algorithm parameters, and the adaptation function often only includes the waiting time of each current charging request, without rule design related to users, which often causes dissatisfaction among some users.

[0005] With the development of artificial intelligence technology, artificial neural networks are gradually used for intelligent scheduling. The most typical one is the reinforcement learning algorithm. The model-free reinforcement learning algorithm learns the optimal policy by interacting with the constructed environment to maximize the reward (see Non-Patent Document 3). However, this model-free reinforcement learning algorithm has the problem of low sample efficiency caused by the time cost and computing power cost of interacting with the environment, resulting in a long training time and high computing power, and is prone to instability during the training process, with performance fluctuations or difficulty in converging to the optimal solution for a long time.

[0006] Prior art documents:

[0007] Non-patent document 1: Samsuria, Erlianasha, et al. "Adaptive fuzzy-geneticalgorithm operators for solving mobile robot scheduling problem in job-shopFMS environment." Robotics and Autonomous Systems 176 (2024): 104683。

[0008] Non-patent document 2: Zhang, Junqi, et al. "Moving-distance-minimized PSO formobile robot swarm." IEEE Transactions on Cybernetics 52.9 (2021): 9871-9881。

[0009] Non-patent document 3: Zhu, Aiyu, et al. "Deep reinforcement learning for real-time assembly planning in robot-based prefabricated construction." IEEETransactions on Automation Science and Engineering 20.3 (2023): 1515-1526。 Summary of the invention

[0010] The present invention aims to address the above problems existing in the prior art, and provides a charging robot scheduling method. When facing multiple charging requests in an intelligent parking lot, the charging robot scheduling method schedules charging robots for quick response according to the analysis of the priorities of these charging requests and the prediction of future new charging requests.

[0011] The present invention mainly designs a priority rule based on the users who issue charging requests, combines a sequence-level adaptive probability inference framework to predict new charging requests, and uses model-based reinforcement learning to establish a world state evolution and reward model, so as to improve the training efficiency and make the optimal strategy closely combined with the environmental dynamics, thereby enhancing its generalization ability.

[0012] According to the first aspect of the present invention, there is provided a charging robot scheduling method, which, for multiple charging requests in an intelligent parking lot, intelligently schedules charging robots according to the priority of the charging requests and the prediction of new charging requests. The method includes the following steps: a step of constructing a priority potential field of a parking space-charging system, constructing a priority potential field of the parking space-charging system based on multiple charging requests of users; a step of converting a Markov decision process, modeling a multi-objective tracking and priority intelligent scheduling process as a Markov decision process based on the constructed priority potential field of the parking space-charging system; a step of constructing a reinforcement learning network, designing a model-based reinforcement learning network based on the modeled Markov decision process, where the reinforcement learning network includes a dynamic network, a value-policy network, and a target network; a step of training and testing the reinforcement learning network, based on the constructed reinforcement learning network, by designing the interaction mode between the models of the dynamic network and the policy-value network and the collected data set, to perform parallel training of multiple networks.

[0013] As a preferred solution, according to the charging robot scheduling method of the present invention, wherein the step of constructing the priority potential field of the parking space-charging system includes a step of establishing an undirected topological network and a step of designing a priority potential field of the charging request of the vehicle node based on the established undirected topological network.

[0014] As a preferred solution, according to the charging robot scheduling method of the present invention, wherein, in the step of establishing the undirected topological network, the topological network of the parking space-charging system is set including the existing number of parking space vertex sets and guide rail edge sets wherein represents a bidirectional edge connecting vertices and , and the topological network of the parking space-charging system .

[0015] As a preferred solution, according to the charging robot scheduling method of the present invention, wherein, in the step of designing the priority potential field of the charging request of the vehicle node, based on the occupancy matrix of the parking space system at the current time defined by using the topological network of the parking space-charging system , according to the following formula (1), the priority potential field intensity at the position is defined as:

[0016] (1)

[0017] where each element in formula (1) is represented as follows:

[0018] Vertex priority value : for vertex At The priority value of the charging request at the moment is represented by the time factor to describe the waiting time of the vehicle for charging.

[0019]

[0020] Mean matrix : Represents the center point of the priority potential field:

[0021] Covariance matrix :

[0022]

[0023] Among them, the user factor is obtained depending on the user's level and feedback, and the vehicle factor describes the charging request of the vehicle. is the weight coefficient.

[0024] As a preferred solution, according to the charging robot scheduling method of the present invention, wherein, in the Markov decision process conversion step, according to the action space of the charging robot describing the movement direction and speed of the charging robot, the state space at time k transfers to the state space at time k + 1 s k+1 The system state transition equation is as follows:

[0025]

[0026] Wherein respectively represent the vertex priority value of the parking space vertex, the occupancy matrix, the priority potential field intensity at the position where the charging robot is located, and the state transition matrix of the charging robot energy consumption.

[0027] As a preferred solution, according to the charging robot scheduling method of the present invention, wherein the step of constructing the reinforcement learning network includes the steps of building a mechanism and data-driven simulation environment and designing a model-based reinforcement learning network architecture based on the built simulation environment.

[0028] As a preferred solution, according to the charging robot scheduling method of the present invention, wherein the step of building the simulation environment includes the design of system state transition and the reward function, and the designed reward function is as follows:

[0029]

[0030] Among them, represents the magnitude of the potential field intensity at the position where the charging robot is located at the current time k, The reward value corresponding to the vertex priority value of the position where the charging robot is located at the current time k The reward value corresponding to the power consumption of the charging robot at the current time k 、 and are weight coefficients.

[0031] As a preferred solution, according to the charging robot scheduling method of the present invention, wherein

[0032] The steps of designing the model-based reinforcement learning network architecture include dynamic network model construction and policy-value network and corresponding target network model construction.

[0033] As a preferred solution, according to the charging robot scheduling method of the present invention, the training and testing steps of the reinforcement learning network include the following processes: self-play, data accumulation, joint training, and parameter saving. Among them, multiple game data are obtained and stored through self-play between the existing policy-value network model and the simulation environment. During joint training, data is randomly selected from the stored game data to train the parameters of the policy-value network and the dynamic network model, and parameter updates are performed in the target network.

[0034] According to the second aspect of the present invention, there is provided a non-transitory storage medium storing a computer program which, when executed by a processor, can implement the charging robot scheduling method according to the first aspect of the present invention.

[0035] The beneficial effects of the present invention are as follows: It can achieve long-sequence decomposition prediction of future charging requests on the basis of considering the priorities of various charging requests, better capture the seasonal and trend laws of charging requests, and achieve more robust scheduling. On this basis, the model-based reinforcement learning algorithm can improve the training efficiency and testing effect of the overall scheduling strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Illustrates the overall flowchart of the charging robot scheduling method according to the present invention.

[0037] Figure 2A Illustrates the flowchart of the sub-process of constructing the priority potential field of the parking space-charging system in the charging robot scheduling method according to the present invention.

[0038] Figure 2B Illustrates the schematic block diagram of the construction of the priority potential field of the parking space-charging system in the charging robot scheduling method according to the present invention.

[0039] Figure 3AThe flowchart of the sub - process of the Markov decision process transformation step in the charging robot scheduling method according to the present invention is illustrated.

[0040] Figure 3B The schematic diagram of the Markov process transformation of the charging robot scheduling method according to the present invention is illustrated.

[0041] Figure 4A The flowchart of the sub - process of the step of designing a model - based reinforcement learning network in the charging robot scheduling method according to the present invention is illustrated.

[0042] Figure 4B The schematic diagram of the long - sequence prediction adaptive probability inference framework in the charging robot scheduling method according to the present invention is illustrated.

[0043] Figure 5 The block diagram of the overall training process of the model - based reinforcement learning scheduling algorithm in the charging robot scheduling method according to the present invention is illustrated.

[0044] Figure 6 The schematic diagram of self - play through Monte Carlo tree search in the charging robot scheduling method according to the present invention is illustrated.

[0045] Figure 7 The schematic diagram of the joint training of network parameters in the charging robot scheduling method according to the present invention is illustrated.

[0046] Figure 8 The flowchart of the detailed process of the charging robot scheduling method according to the present invention is illustrated. Detailed implementation manners

[0047] The exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative configurations of components, numerical representations, and numerical values described in these embodiments do not limit the scope of the present invention.

[0048] The charging robot scheduling method of the present invention can be implemented by a processor in the charging robot executing a computer program stored in a memory in the charging robot. As an alternative solution, it can also be implemented by the charging robot communicating with a server, and a processor in the server executing a computer program stored in the server or the cloud and real - time feedback of the program execution result to the charging robot.

[0049] In the following description of the present invention, the scheduling method of the present invention is a charging robot scheduling method for scheduling charging robots to respond quickly in the face of multiple charging requests in an intelligent parking lot. In addition, the scheduling method of the present invention is not limited thereto, and it can also be applied to situations where multiple requests (such as multiple cleaning requests or multiple purchase requests, etc.) existing in other environments except the intelligent parking lot need to be processed.

[0050] In addition, it should be understood that the parking lot mentioned in the present invention may not only refer to the entire parking lot, but also refer to each area in the case where the parking lot is large and the entire parking lot is divided into multiple areas.

[0051] The core technology of the charging robot scheduling method of the present invention is to first establish an undirected topology graph of the parking space - charging system for the vehicle - charging system by using relevant theories of graph theory, and then design a priority calculation rule based on users to describe the priority level of each user's charging request. Then, combined with the established undirected topology graph of the parking space - charging system and the designed priority rule of users, a priority potential field of the parking space - charging system is established. On this basis, the charging robot and the priority intelligent scheduling process targeted by the present invention are transformed into a Markov decision process, and a dynamic network model is used to learn and reason about the vertex state transition model and the topology structure transition model of the undirected topology graph, where the latter uses an adaptive prediction probability framework based on sequence decomposition to predict new charging requests. Then, a "mechanism + data" - driven simulation environment is established. A mechanism environment is established by using the priority potential field evolution equation and the charging robot energy consumption equation, and a data environment is established by combining the previous charging request data. The value - policy network model interacts with the dynamic network model and the simulation environment, searches for the optimal policy in the way of Monte Carlo tree search, and obtains the corresponding value and reward, accumulates data in the way of self - play, and finally updates the network parameters by using the stored data to improve the training efficiency and test effect of the overall scheduling strategy.

[0052] Next, with reference to Figure 1 the charging robot scheduling method according to the present invention will be described. The method includes the following steps S100 to S400.

[0053] First, in step S100, a priority potential field of the parking space - charging system is constructed. Specifically, in step S100, based on the multiple charging requests of electric vehicle users in the intelligent parking lot, a priority potential field of the parking space - charging system is constructed.

[0054] Next, in step S200, based on the priority potential field of the parking space - charging system constructed in step S100, the multi - target tracking and priority intelligent scheduling process is modeled as a Markov decision process.

[0055] Then, in step S300, based on the transformation of the multi-objective tracking and priority intelligent scheduling created in step S200 into a Markov decision process, a model-based reinforcement learning network (algorithm) is designed to ensure the optimality of the policy and the efficiency of training. Specifically, the reinforcement learning network includes a dynamic network that describes the vertex state of the undirected topological network and the topological structure transition function, a value-policy network that explores the optimal scheduling policy, and a target network that updates the parameters.

[0056] Finally, in step S400, based on the undirected topological network established in step S100 and the reinforcement learning network designed in step S300, by designing the interaction method between the models of the dynamic network and the policy-value network and the collected data set, parallel training of multiple networks is achieved. Specifically, the dynamic network model is learned through the prior data set. Through interaction with the dynamic network model, the policy network maximizes the obtained reward, the value network trains the parameters to maximize the expected cumulative discounted reward, and soft updates of the parameters are performed in the target network for verification in the test set and actual operation.

[0057] The charging robot scheduling method of the present invention performs intelligent scheduling for multiple charging requests in an intelligent parking lot according to the priority of the charging requests and the prediction of new charging requests, so that the charging robot can respond quickly.

[0058] Next, with reference to Figure 2A and Figure 2B for Figure 1 the process of constructing the priority potential field of the parking space-charging system in step S100 of Figure 2A is described in detail. As

[0059] shown, the process of constructing the priority potential field of the parking space-charging system specifically includes step S101 of establishing an undirected topological network and step S102 of designing the priority potential field of the charging requests of the parking space vertices.

[0059] Specifically, with reference to Figure 2B the construction schematic diagram shown to describe in detail the specific process of constructing the priority potential field of the parking space-charging system of the intelligent parking lot.

[0060] In step S101 of establishing the undirected topological network, the topological graph of the parking space-charging system of the intelligent parking lot is set , including the existing set and the guide rail edge set , where represents the bidirectional edge connecting the vertices and , and the topological graph of the parking space-charging system (topological network) is denoted as Due to the characteristics of the charging robot guide rail, this topological graph is a bidirectional but non-connected graph, where the parking space vertices can also be called parking space nodes, corresponding to each parking space in the physical schematic diagram, and the edge value is set as the guide rail distance between adjacent parking spaces.

[0061] The steps S102 of the priority potential field design for the charging request of the parking space vertex include the step S1021 of calculating the vertex priority value and the step S1022 of calculating based on the covariance matrix of the user and the vehicle.

[0062] Specifically, in the step S102 of the priority potential field design for the charging request of the parking space vertex, according to the parking space - charging system topological graph established in the step S101, first define the occupancy matrix of the parking space system in the parking lot at the current moment : :

[0063]

[0064] where represents the currently occupied parking space vertex. On this basis, define the priority potential field intensity at the position as shown in the following formula (1):

[0065] (Formula 1)

[0066] The descriptions of each relevant element in formula (1) are as follows:

[0067] (1.1) Vertex priority (priority value) :

[0068]

[0069] where is the priority value of the charging request of vertex at time. The larger the vertex priority value, the more urgent the charging request. It is represented by the time factor , which mainly describes the waiting time of the vehicle for charging:

[0070]

[0071] where respectively represent the current world time and the initial time when the parking space represented by vertex is occupied, and represents the parameter characterized by the control time factor.

[0072] (1.2) Mean matrix , which represents the center point of the priority potential field:

[0073]

[0074] (1.3) Covariance Matrix , which is expressed as follows:

[0075]

[0076] where the user factor and the vehicle factor are defined respectively as follows, being the weights of two different factors to balance the influence on the overall priority value for different priority value pairs:

[0077] (1.3.1) Vehicle Factor : Describes the charging request for the vehicle. Specifically, after the vehicle parks in the parking space, the battery capacity of the vehicle is obtained by identifying the vehicle category through the camera on the guide rail , and then the vehicle factor is obtained through calculation:

[0078]

[0079] where is an adjustment parameter used to adjust the image of the vehicle factor function according to the actual situation.

[0080] (1.3.2) User Factor : Obtained depending on the user's level and feedback:

[0081]

[0082] where is the maximum tolerance time of the user deduced according to the user feedback, being the level of the user currently requesting charging:

[0083]

[0084] where is the current cumulative charging time of the user.

[0085] The above describes the construction process of the priority potential field of the parking space - charging system of the intelligent parking lot of the present invention. In addition, the reinforcement learning algorithm in the multi - target scheduling method for the charging robot of the present invention mainly depends on a Markov decision process (MDP). Specifically, based on the priority potential field of the parking space - charging system constructed in step S100, the multi - target tracking and priority intelligent scheduling process is modeled as a Markov decision process for subsequent model - based reinforcement learning algorithm design.

[0086] Next, referring toFigure 3A and Figure 3B Describe in detail the flowchart of the Markov decision process transformation step S200 in the charging robot scheduling method according to the present invention.

[0087] As Figure 3A shown, the Markov decision process transformation step S200 includes a step S201 of priority potential field state transition based on the priority potential field of the parking space - charging system constructed in step S100 and a step S202 of establishing a system state transition equation based on the energy consumption mechanism model of the charging robot.

[0088] As Figure 3B shown, the Markov decision process transformation mainly includes core elements such as a state space, an action space, and a reward function. Among them, the state space is mainly the information of the undirected topology graph of the parking space - charging system, the action space mainly contains the motion control instructions of the charging robot; the reward value is mainly obtained by representing the sum of the priority values of the vertices in the current undirected graph structure and the energy consumption of the charging robot. The following describes each core element in detail.

[0089] First, the state space is mainly the information of the undirected topology graph of the parking space - charging system, including the vertices V, edge values E, and occupancy matrix of the current undirected topology graph, the priority value of the vertices in the topological structure, and the priority potential field intensity as well as the energy consumption of the current charging robot, which is specifically represented as follows:

[0090]

[0091] where is the power of the charging robot at the current moment.

[0092] Next, describe the action space , which mainly includes the moving direction and speed of the charging robot:

[0093]

[0094] where respectively represent the current moving direction of the charging robot and the speed in this direction.

[0095] The reward function is obtained through the sum of the priority values of the vertices of the current undirected topology graph, the priority potential field intensity where the charging robot is located, and the energy consumption:

[0096]

[0097] Further, the specific setting of the reward function will be described in the reinforcement learning network design step S300 to be described below.

[0098] As Figure 3B shown, the left side represents the state space at the current time k, which consists of the vertices V of the undirected topological graph, the edge values E, and the occupancy matrix , the vertex priority value in the topological structure and the priority potential field strength as well as the energy consumption of the current charging robot is represented. According to the motion direction and speed of the charging robot described by the action space of the charging robot, the state space transfers to the state space at time k+1 s k+1 As Figure 3B shown on the right side of. Specifically, the system state transition equation is as follows:

[0099]

[0100] where respectively represent the state transition matrices of the vertex priority value of the parking space vertex, the occupancy matrix, the priority potential field strength at the location of the charging robot, and the energy consumption power of the charging robot. It is achieved through the priority potential field strength rule and the motion of the charging robot proposed in step S100. When the charging robot moves to an occupied parking space, the priority value of the current parking space will be cleared, and the action space will be locked to fix the position of the charging robot to achieve the charging operation. The network model is obtained through the long sequence prediction adaptive inference framework to be described below. The relationship between the motion actions of the charging robot and its energy consumption will be obtained through the learning to be described below.

[0101] The above-described Markov decision process transformation will be used for the model-based reinforcement learning algorithm design to be described below. The multi-object tracking and priority intelligent scheduling problem is transformed into a reinforcement learning algorithm to learn a scheduling policy network that can minimize the priority degree of each vertex in the overall undirected graph network.

[0102] Next, with reference to Figure 4A and Figure 4B describe the model-based reinforcement learning network design step S300 in the charging robot scheduling method according to the present invention.

[0103] The design and processing of the model-based reinforcement learning scheduling algorithm of the present invention is to design the model-based reinforcement learning algorithm on the basis of realizing the transformation design of the multi-objective tracking and intelligent scheduling Markov decision process of the charging robot in step S200, so as to ensure the optimality of the policy and the efficiency of training.

[0104] As Figure 4A shown, the steps S300 of the model-based reinforcement learning network design of the present invention include the following steps S301 to S302.

[0105] Step S301: Building a simulation environment driven by "mechanism + data". The construction of the simulation environment mainly includes the design of system state transition and reward function.

[0106] The system state transition is shown in an iterative form through the system state transition equation established in step S202. It mainly includes the state transition equations of the priority of the vertices of the undirected graph, the occupancy matrix, and the energy consumption power of the charging robot. Among them, since the occupancy matrix cannot be described by the mechanism model, it is established in a data-driven manner. The overall state transition equation consists of a "mechanism" state transition model composed of the designed vertex priority state transition and the energy consumption power state transition of the charging robot, and a "data" state transition model formed by collecting the daily parking space occupancy situation in the current parking lot.

[0107] Next, the design of the reward function is described. The real-time reward feedback is as described in the above step S200 regarding the reward function. The specific reward function is set as follows:

[0108]

[0109] Where represents the magnitude of the potential field intensity formed by the vertex priority at the position of the charging robot at the current time k, so as to prompt the charging robot to respond to relatively urgent charging requests; represents the reward value corresponding to the vertex priority value (vertex priority value) in the undirected topological graph at the current time k, so as to ensure the minimum charging request delay for each parking space; represents the reward value corresponding to the energy consumption power of the charging robot at the current time k; , and are weight coefficients used to measure the importance of these three reward values. The specific definitions of these parameters are shown in the following formulas:

[0110]

[0111] Step S302: Design of a model-based reinforcement learning network architecture. To enable the obtained optimal policy to be closely combined with the environmental dynamic model, a model-based reinforcement learning network architecture is designed, specifically including the construction of a dynamic network model and a policy-value network.

[0112] First, describe the construction of the dynamic network model. After transforming the intelligent scheduling problem of the charging robot into a reinforcement learning problem, effective reward value data needs to be obtained by testing the policy optimized by the policy network in the test set, which will lead to inefficient training results. To improve the training rate and closely combine the optimal policy with the environmental dynamics, the present invention constructs a dynamic network model to accurately describe the state transition probability and the obtained reward value, so that the obtained data not only comes from the real-world dataset, but also a part of it comes from a large amount of data jointly constructed by the dynamic network model and the policy model. Combining the strategy of establishing the simulation environment in step S301, the dynamic network model is defined as:

[0113]

[0114] where the network can be implemented by a multi-layer perceptron. The network does not belong to the state transition function of the Markov decision process, but predicts future new requests through a long-sequence prediction adaptive probability inference framework, enhancing the robustness of the model-based reinforcement learning scheduling algorithm to new requests. The overall framework diagram and its processing process of this long-sequence prediction adaptive probability inference framework are as Figure 4B shown.

[0115] Next, refer to Figure 4B to describe the details of the long-sequence prediction adaptive probability inference framework of the present invention. First, after analyzing the new charging request data, it is found that its long time series usually shows a long-term trend and short-term periodic fluctuations. Therefore, decomposing the time series can help the model effectively capture its internal complex time dynamics. Use a frequency-domain-based method to decompose the time series, map the sequence to the frequency domain, and then use frequency masking to take the high-frequency part as the seasonal component , and the low-frequency part as the trend component :

[0116]

[0117] where represents the Fourier transform, represents the inverse Fourier transform. Since is prone to noise, certain filtering is required to obtain it. In the present invention, the model parameters in the neural network are represented as:

[0118]

[0119] wherein respectively represent the seasonal encoder and the trend encoder, represents the seasonal trend global representation extractor, represents the decoder.

[0120] As Figure 4B shown, the long time series of the charging request dataset is decomposed into seasonal components and trend components , and the seasonal encoder and the trend encoder are used to encode the seasonal components and trend components respectively. Then, in the clustering matching mode, the sequences encoded by the seasonal trend global representation extractor are divided into meta-learning tasks. Based on the network parameters and corresponding to the meta-learning tasks respectively, the seasonal representation and trend representation of the time parameter are learned. Finally, the seasonal and trend components are fused and decoded in the decoder to complete the final prediction. Thus, the prediction of the new charging request is completed.

[0121] The construction of the policy-value network is described below. To better optimize the policy network parameters, a policy-value network and the corresponding target network model are established. The parameters in the two networks are trained using the stochastic gradient method, and the parameters in the target network are optimized by means of soft update. The established policy network is used to give the optimal policy according to the current system state:

[0122]

[0123] wherein represents the optimal action obtained by the charging robot at the current moment through the policy network according to the current state .

[0124] The value network is used to estimate the cumulative discounted reward to be generated according to the current state:

[0125]

[0126] wherein represents the expected value corresponding to the current state and action, Represents a neural network function for predicting the expected value. The policy network and the value network are co-optimized in a policy iteration manner. The goal of the policy network is to maximize the predicted cumulative discounted reward, while the goal of the value network is to match the predicted value with the actual reward value. The loss function of the value network is:

[0127]

[0128] Where Represents the true value of the expected value, which is obtained through the Markov property. The loss gradient of the policy network is as follows:

[0129]

[0130] Policy network Through n Number of explorations to obtain and store data. For the policy network The state-action value obtained in the neural network function For the actions taken is trained by backpropagating the gradient to update the model parameters of the policy network, so as to approximate the optimal through training the policy network.

[0131] The above describes the design process of the model-based reinforcement learning network. The following describes the steps S400 for training and testing the model-based reinforcement learning network.

[0132] Specifically, based on the reinforcement learning network constructed in step S300, by designing the interaction method between the dynamic network model and the policy-value network model established in step S302 and the collected data set, parallel training of multiple networks is achieved. The training and testing steps S400 of the model-based reinforcement learning network (scheduling algorithm) mainly include four processes: self-play, data accumulation (experience replay), parameter training (joint training), and parameter saving. Specifically as Figure 5 Shown in the flow chart of the overall training process of the model-based reinforcement learning scheduling algorithm.

[0133] As Figure 5 Shown, a large amount of game data is obtained through self-play between the established policy-value network model and the simulation environment. Then, the game data obtained through self-play is stored in the experience pool. During joint training, the game data (experience data) collected is randomly sampled from the experience pool for training the model parameters of the policy-value network and the dynamic network, and the trained model parameters are saved, and this is looped. It mainly includes the following processes 1 to 3:

[0134] Process 1 - Pre-training of the dynamic network model: First, collect the historical parking data of the parking lot, for Pre-train the world state transition model represented by the network to accurately represent the priority potential field of the undirected graph of the parking space - charging system and the occupancy matrix state transition model caused by new charging requests.

[0135] Process 2 - Generation of environment interaction data: Since the charging robot can only move on the guide rail, the action space can be discretized as:

[0136]

[0137] where is the action selected by the charging robot through the agent policy in the discretized action space, and the direction can be discretized into forward, backward, left turn, and right turn, and the speed can be discretized into specific gears. The specific speed gears depend on the charging robot used, and the present invention does not make specific limitations. To obtain the optimal policy, data is generated through the interaction between the self-play module and the real environment to provide data for subsequent model training. In the present invention, model-based Monte Carlo tree search (MCTs) is used to explore and select the action space, sample through the distribution of visit counts, and finally obtain the policy applied to the environment for interaction. The dynamic network and the policy - value network are used for search, and self-play is performed through Monte Carlo tree search to complete the process of data collection as Figure 6 shown.

[0138] As Figure 6 shown, every time a state is reached, the policy probability distribution is obtained through Monte Carlo tree search. After sampling and selection, the action at the current moment is obtained, and the reward at the current moment and the state at the next moment are obtained, and the data collection is completed in this cycle. The process of Monte Carlo tree search is to, after initializing the Monte Carlo tree, select the current action at each node according to the upper confidence bound algorithm, expand to the next node through the dynamic network model combined with the current node and the selected action, and update the selection count and reward value of the next node until the leaf node is reached. Then, the selection count and reward value are backtracked to reconstruct and optimize the Monte Carlo tree until the maximum search step is reached. The process of Monte Carlo tree search mainly includes the following subprocesses:

[0139] MCTs subprocess 1 - Initialize MCTs: The main role of MCTs is to determine the most effective method, obtain immediate rewards, and evaluate the prediction value, and initialize the node , and the information saved for each edge is:

[0140]

[0141] Among them represents the access count of the current node represents the optimal policy and predicted value obtained through the current parameters of the policy-value network is the predicted reward value and state transition equation obtained from the current parameters of the dynamic network

[0142] MCTs subprocess 2 - Policy selection: For each node, the upper confidence bound algorithm is combined with the value-policy network and the access count of the node to select the policy with the highest score. The specific formula is as follows

[0143]

[0144] Among them represents the number of times of selecting action a at the s-state node represents the number of times of selecting all actions other than a at the s-state node. The upper confidence bound algorithm makes the search more inclined to the action-state pairs with higher historical scores through the first term in the formula and makes the search tend to select the nodes that have not been explored through the latter term in the formula

[0145] Among them is a parameter used to adjust the balance between the efficiency and breadth of the upper confidence bound algorithm. If is larger, the search algorithm is more inclined to search the nodes with high current success rates, focusing on search efficiency; conversely, if is larger, it is more inclined to select the nodes that have not been explored, focusing on search breadth. The coordination between the two ensures the balance of the search tree's efficiency and breadth. Therefore, the search tree is more inclined to find and explore the high-value nodes that have not been discovered

[0146] MCTs subprocess 3 - Node expansion: New state nodes are expanded according to the policy selected in MCTs subprocess 2 through the dynamic network model, and the reward for the next state is predicted until the leaf node of the search tree is reached After that, the corresponding policy and value are predicted through the policy-value function as follows

[0147]

[0148] MCTs subprocess 4 - Backpropagation: Through the trajectory of the current search, the predicted value is backpropagated to all nodes in the search path, increasing the access count and updating the average value of the nodes

[0149]

[0150] Among them Indicates the number of times a node is visited. When the trajectory passes through this node, the visit count is incremented by 1. is the value estimate:

[0151]

[0152] where, represents the decay factor in the Markov process, represents the reward, represents the expected value, τ represents the number of steps of iteration, m is the last leaf node of the search tree, representing the total number of steps, k is the current moment of MCTS, representing the current number of steps.

[0153] Process 3 - Joint Training: Since the search time of MCTs is relatively long, in actual application, each network is trained to obtain parameters to represent the optimal strategy found. Since the state transition model in the dynamic network model has been pre-trained, the policy network, value network, and reward network are mainly jointly trained. The overall joint training process is as Figure 7 shown. In actual application, at every time interval, the original observation state is collected , and in the self-play process of interacting with the environmental data in Process 2 (generation of environmental interaction data) at the current moment, after using Monte Carlo tree search to search and sample for the optimal strategy, an action is executed in the environment , and the observation state at the next moment is obtained and the immediate reward obtained . The above data constitutes a set of moment data in the trajectory and is stored. After collecting a certain amount of data, the joint training process is started. A trajectory is randomly sampled from the numerous collected trajectories, and the first k moment data is taken as the initial state. Using the in the dynamic network model, the policy network and the value network According to the input at the current moment the corresponding immediate reward, policy action, and expected value are calculated using the existing network model parameters, and the loss functions are calculated respectively with the immediate reward fed back by the simulator, the optimal strategy obtained by MCTs search, and the corresponding expected value in the collected dataset. Then, the network model parameters are calculated and updated by means of backpropagation.

[0154]

[0155] The above formula represents the overall loss function of joint training. c represents the regularization term for network training to ensure the sparsity of the network, They respectively represent the loss functions of the dynamic network, the policy network, and the value network, which are expressed as minimizing the error between the predicted policy and the optimal policy searched by MCTs, the error between the predicted value and the target value of the value network, and minimizing the error between the predicted reward of the dynamic network model and the actually observed reward. Since the output state in the dynamic network will be the input of the policy-value network, its gradient can be backpropagated into the dynamic network model for parameter update. For the detailed process of the training and testing steps S400 of the model-based reinforcement learning scheduling algorithm of the present invention, refer to Figure 8 as shown.

[0156] As Figure 8 shown in the block diagram on the right, the training and testing steps S400 of the reinforcement learning scheduling algorithm of the present invention are carried out on the basis of the reinforcement learning network constructed in step S300. First, in step S401, the initial state of the self-play is selected. Then, in step S402, the optimal action is selected through Monte Carlo tree search. In step S403, the nodes of the Monte Carlo tree are expanded based on the selected optimal action. In step S404, it is determined whether the node has been expanded to a leaf node. If it is determined in step S404 that it has not been expanded to a leaf node (No in step S404), the process returns to step S402, and if it is determined that it has been expanded to a leaf node (Yes in step S404), the process proceeds to step S405, and the corresponding policy and value are predicted using the policy network until the self-play ends. Next, in S406, the rewards explored in this self-play are traced back along the path and backpropagated to all nodes in the search path. Then, in step S407, the data experience of this self-play is stored for data accumulation.

[0157] In step S408, it is determined whether the maximum batch of data accumulation has been reached, where the maximum batch of data accumulation can be set according to user requirements or actual situations. If it is determined in S408 that the maximum batch has not been reached (No in step S408), the process returns to step S401, and if it is determined in S408 that the maximum batch has been reached (Yes in step S408), the process proceeds to step S409 to start the joint training process of parameter training.

[0158] Specifically, in step S409, the data experience of the self-game is retrieved, and in step S410, the initial state of the experience is set. Next, in step S411, a policy-value network is used to make action selections to calculate the corresponding policy actions and expected values according to the existing network model parameters. In step S412, a dynamic network is used to generate rewards and perform state transitions, specifically calculating the corresponding immediate rewards according to the existing network model parameters. Then, in step S413, the policy actions, expected values, and immediate rewards calculated in steps S411 and S412 are compared with the experience data in the dataset to calculate the error. In step S414, the calculation results in step S413 are backpropagated into the dynamic network model for parameter update, thereby performing parameter training. In step S415, it is determined whether the training of this batch is completed. If it is determined in step S415 that the training of this batch is not completed, the process returns to step S409 to continue the joint training process. If it is determined in step S415 that the training of this batch is completed, the process returns to step S401 to start the self-game process again, and this cycle continues until the preset number of loop steps is reached.

[0159] The above reference Figure 8 describes the training and testing processes of the reinforcement learning scheduling algorithm of the present invention. For Figure 8 the other steps and processes in it have been described in the previous part and will not be elaborated here.

[0160] In the training and testing of the reinforcement learning scheduling algorithm of the present invention, based on the above construction of the undirected graph and the design of the reinforcement learning network, through interaction with the dynamic network, the policy network maximizes the obtained rewards, and the value network maximizes the predicted cumulative discounted rewards to train the parameters, and performs soft parameter updates in the target network. Finally, it is verified in the test set and actual operation, improving the training efficiency and testing effect of the overall scheduling strategy.

[0161] In summary, the charging robot scheduling method of the present invention constructs the priority potential field of the parking space-charging system by using the formulated priority rules, combines deep learning to predict charging requests, transforms the intelligent scheduling process into a Markov decision process, designs a model-based reinforcement learning intelligent scheduling algorithm on this basis, and uses Monte Carlo tree search to train and test the optimal strategy. Thus, the present invention realizes the long-sequence decomposition prediction of future charging requests on the basis of considering the priorities of each charging request, can better capture the seasonal and trend laws of charging requests, schedules the charging robot for rapid response, and realizes a more robust scheduling. In addition, on this basis, the model-based reinforcement learning algorithm can improve the training efficiency and testing effect of the overall scheduling strategy.

[0162] [Other Embodiments]

[0163] Embodiments of the present invention can also be implemented by a computer of a system or apparatus that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium (which may also be more fully referred to as a "non-transitory computer-readable storage medium") to perform one or more of the functions in the above embodiments, and / or includes one or more circuits (e.g., an application specific integrated circuit (ASIC)) for performing one or more of the functions in the above embodiments. Moreover, embodiments of the present invention can be implemented by a method of, for example, reading and executing the computer-executable instructions from the storage medium by the computer of the system or apparatus to perform one or more of the functions in the above embodiments, and / or controlling the one or more circuits to perform one or more of the functions in the above embodiments. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessing unit (MPU)), and may include a network of separate computers or separate processors to read and execute the computer-executable instructions. The computer-executable instructions may be provided to the computer, for example, from a network or the storage medium. The storage medium may include, for example, one or more of a hard disk, a random access memory (RAM), a read-only memory (ROM), a memory of a distributed computing system, an optical disc (such as a compact disc (CD), a digital versatile disc (DVD), or a Blu-ray Disc (BD)™), a flash device, and a memory card, etc.

[0164] Although the present invention has been described above with reference to exemplary embodiments, the above embodiments are only for illustrating the technical concept and features of the present invention and cannot be used to limit the protection scope of the present invention. Any equivalent variation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.

Claims

1. A charging robot scheduling method, characterized in that: For multiple charging requests in a smart parking lot, a charging robot is intelligently scheduled according to the priority of the charging requests and the prediction of new charging requests. The method includes the following steps: A priority potential field construction step S100 of the parking space-charging system is to construct a priority potential field of the parking space-charging system based on the user's multiple charging requests; A Markov decision process conversion step S200, based on the priority potential field of the parking space-charging system constructed in the priority potential field construction step S100 of the parking space-charging system, the multi-target tracking and priority intelligent scheduling process is modeled as a Markov decision process; A reinforcement learning network construction step S300, based on the Markov decision process modeled in the Markov decision process transformation step S200, to design a model-based reinforcement learning network, wherein the reinforcement learning network includes a dynamic network, a value-strategy network and a target network; The reinforcement learning network training and testing step S400 is based on the reinforcement learning network constructed in the reinforcement learning network construction step S300, and performs parallel training of multiple networks by designing the interaction between the dynamic network model and the strategy-value network model and the collected data set. The priority potential field construction step S100 of the parking space-charging system includes a step S101 of establishing an undirected topological network and a step S102 of designing a priority potential field of vehicle node charging requests based on the established undirected topological network. In the step of establishing an undirected topological network, the parking space-charging system topological network G is set to include the existing M parking space vertex sets v m ∈V, m∈(1, 2, ..., M) and the guide edge set e i,j ∈E, i, j∈(1, 2, ..., M), where e i,j Represents the connected vertex v i and v j The bidirectional edge of the parking space-charging system topology network G = (V, E), According to the parking space-charging system topology established in step S101, first define the occupancy matrix O of the parking space system of the parking lot at the current time k. k : The k ={o k,i },i∈V Where V occupy Represents the currently occupied parking space vertex, based on which the priority potential field strength g at position X = (x, y) is defined k As shown in the following formula (1): The description of each relevant element in formula (1) is as follows: Vertex priority value j k,i : where j k,i is the priority value of the charging request of vertex i at time k. The larger the vertex priority value, the more urgent the charging request. It mainly describes the time the vehicle waits for charging: where t k , t 0,i Respectively represent the current world time and the initial time when the parking space represented by vertex i is occupied, θ time represents the parameter characterizing the control time factor, The mean matrix u represents the center point of the priority potential field: The covariance matrix ∑ is expressed as follows: ∑ 11 =∑ 22 =λq k,i The user factor Depends on the user's rating and feedback, the vehicle factor Describe the charging request of the vehicle, λ is the weight coefficient, Define the dynamic network model as: o k,i =d η (o k-1,i ,t k ) r k =d θ (s k ) Among them, r k is the reward function, s k is the state space, a k is the action space, P k is the current energy consumption of the charging robot, t k is the current world time, Among them, the network d θ Through the multi-layer perceptron, the network d η It does not belong to the state transition function of the Markov decision process, but predicts future new requests through a long sequence prediction adaptive probabilistic reasoning framework.

2. The charging robot scheduling method according to claim 1, wherein: In the Markov decision process transformation step S200, according to the action space a of the charging robot k The movement direction and speed of the charging robot described by k Transfer to the state space s at time k+1 k+1 The system state transfer equation is as follows: s k+1 ~q W (s k+1 |s k ,a k ) =(q j (j k+1 |j k ),q o (o k+1 |o k ),q g (g k+1 ||g k ,a k ),q P (P k+1 |P k ,a k )) where q j (·), q o (·), q g (·), q P (·) respectively represent the vertex priority value of the parking space vertex, the occupancy matrix, the priority potential field strength of the charging robot’s location, and the state transition matrix of the charging robot’s energy consumption.

3. The charging robot scheduling method according to claim 1, wherein: The reinforcement learning network construction step S300 includes a step S301 of constructing a mechanism and data-driven simulation environment and a step S302 of designing a model-based reinforcement learning network architecture based on the constructed simulation environment.

4. The charging robot scheduling method according to claim 3, wherein: Step S301 of setting up the simulation environment includes the system state transfer and the design of the reward function, wherein the designed reward function is as follows: r k =λ g r g,k +λ j r j,k +λ P r P,k Among them, r g,k represents the potential field strength of the charging robot at the current time k, r j,k represents the reward value corresponding to the vertex priority value of the charging robot at the current time k, r P,k represents the reward value corresponding to the energy consumption of the charging robot at the current time k, λ g , j and λ P is the weight coefficient.

5. The charging robot scheduling method according to claim 3, wherein: Step S302 of designing a model-based reinforcement learning network architecture includes constructing a dynamic network model and a strategy-value network and a corresponding target network model.

6. The charging robot scheduling method according to claim 1, wherein: The training and testing step S400 of the reinforcement learning network includes the following processes: self-play, data accumulation, joint training and parameter preservation, Among them, multiple game data are obtained and stored by playing self-games with the existing strategy-value network model and the simulation environment. During joint training, data are randomly extracted from the stored game data to train the parameters of the strategy-value network and the dynamic network model, and the parameters are updated in the target network.

7. A non-temporary storage medium storing a computer program, which, when executed by a processor, can implement the charging robot scheduling method according to any one of claims 1 to 6.