Automatic charging system scheduling method and scheduling device, storage medium and computer program product
By adopting the scheduling method of multi-intelligent reinforcement learning in the electric vehicle charging parking lot, the scheduling strategies of charging piles and robotic arms are optimized, and the problems of low utilization rate of charging piles and long charging waiting time are solved, and efficient and safe charging services are achieved.
Patent Information
- Application Number
- CN202510095984.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-21
AI Technical Summary
In the existing electric vehicle charging parking lot, the charging pile utilization rate is low, the charging waiting time is long, and the scheduling mechanism is inflexible. Especially when the parking lot size becomes larger, automatic charging robots are prone to collision and congestion problems.
The scheduling method based on multi-agent reinforcement learning is adopted. By constructing the strategy value function and action value function, using neural network modeling and training, the scheduling strategies of the charging pile and the robotic arm are optimized, so that they can learn the optimal strategy to maximize common rewards under a fully cooperative relationship, and insert a security control module into the algorithm to limit the actions to satisfy physical constraints.
It significantly improves the utilization rate and charging efficiency of charging piles, shortens the charging waiting time, ensures the rationality and safety of system operation, and has high flexibility and scalability, adapts to parking lots of different sizes and configurations.
Smart Images

Figure CN119539438B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automatic charging of electric vehicles, and in particular to a scheduling method and scheduling device, a storage medium and a computer program product of an automatic charging system. Background Art
[0002] With the rapid growth of the electric vehicle market, the demand for charging infrastructure has also increased. At present, most charging stations use fixed charging piles. In this mode, electric vehicles fail to leave in time after charging and continue to occupy idle charging piles, resulting in low utilization of charging piles. This will cause serious charging waiting problems during peak hours, affecting user experience.
[0003] In order to solve the above problems, charging parking lots have begun to try to introduce automation technology. In existing automation technologies (for example, see Chinese patent applications with publication numbers CN119142178A and CN112776624A), an automatic charging robot is used to carry a mobile charging pile and autonomously move to the electric vehicle that needs to be charged, and the charging gun is inserted into the electric vehicle to complete automatic charging, thereby releasing the strong binding relationship between the pile and the parking space, and changing "car looking for pile" to "pile looking for car".
[0004] However, the above technology does not take into account the collision and congestion problems that may occur when the parking lot becomes larger and the number of automatic charging robots increases, nor does it take into account how to efficiently dispatch a large number of mobile automatic charging robots to minimize the charging waiting time.
[0005] To solve the above problems, it is necessary to design a corresponding scheduling algorithm for a specific parking lot. The scheduling problem of the automatic charging robot is a complex multi-objective optimization problem. First, the automatic charging robot works in a dynamically changing environment, and the charging demand changes with time and location; second, the scheduling algorithm must not only consider the charging waiting time but also the robot's energy consumption; in addition, when there are multiple charging robots, the coordination between them is also crucial. Traditional scheduling algorithms are difficult to cope with the above difficulties.
[0006] In recent years, reinforcement learning has demonstrated excellent performance in many fields, including games, robot control, and autonomous driving. Reinforcement learning provides a potential solution to the above problems by learning the optimal strategy to maximize long-term rewards through the interaction between the agent and the environment. Summary of the invention
[0007] The present invention aims to solve the problems of low charging pile utilization, long charging waiting time and inflexible scheduling mechanism in existing electric vehicle charging parking lots in the prior art. By introducing a scheduling method based on multi-agent reinforcement learning, the scheduling process of the parking lot suspension rail automatic charging system is optimized, the parking lot charging efficiency is improved, and the user experience is improved.
[0008] To this end, the present invention is made in view of the above-mentioned technical background, and it applies reinforcement learning to the system scheduling problem of the suspended rail automatic charging parking lot. Specifically, a multi-agent reinforcement learning algorithm is used to enable the charging piles and the robotic arms to work closely and coordinate with each other in a fully cooperative relationship, learn their own optimal strategies to maximize the common rewards, and at the same time use reinforcement learning with a safety control module to limit the behavior of the scheduling system to comply with physical constraints.
[0009] According to a first aspect of the present invention, there is provided a scheduling method for an automatic charging system, wherein the automatic charging system comprises a charging pile and a robotic arm, wherein the robotic arm performs plugging and unplugging operations on a charging gun on the charging pile so that the charging gun performs charging operations on an electric vehicle, and the scheduling method comprises the following steps: a construction step, constructing a policy value function and an action value function for the charging pile and the robotic arm respectively; a modeling step, using a neural network to model the action value function and the policy value function of the charging pile and the robotic arm constructed in the construction step, so as to establish an action value network and a policy network; a training step, using a temporal difference algorithm to train the action value network of the charging pile and the robotic arm established in the modeling step, and training the policy network of the charging pile and the robotic arm established in the modeling step with the goal of maximizing the cumulative reward; a parameter updating step, updating the neural network parameters after the training step is completed, so that the policies of the charging pile and the robotic arm are the optimal scheduling policies.
[0010] As a preferred solution, according to the scheduling method of the automatic charging system of the present invention, the objective function of the strategy value function constructed for the charging pile and the robotic arm in the construction step is defined as follows:
[0011]
[0012]
[0013] in and Represent the strategy value functions for charging piles and robotic arms respectively and The objective function is Respectively represent the actions of the charging pile and the robotic arm, Indicates the status of the entire environment of the automatic charging system, is the expected function,
[0014] is the reward function of the charging pile and the robotic arm, and They are and The entropy indicates that the charging pile strategy and the robotic arm strategy are in the state The degree of randomness under is the regularization coefficient.
[0015] As a preferred solution, according to the scheduling method of the automatic charging system of the present invention, the reward function of the charging pile and the mechanical arm is The definition is as follows:
[0016]
[0017] in, Indicates in status Next, the charging pile performs an action And the robot performs the action , No. The waiting time for a charging request; Indicates that the charging pile performs an action The total energy consumption, and Indicates that the robot arm performs an action Total energy consumption; are all weight coefficients.
[0018] As a preferred solution, according to the scheduling method of the automatic charging system of the present invention, in the construction step, based on the objective function of the constructed strategy value function, the action value functions are constructed for the charging pile and the robotic arm respectively as follows:
[0019]
[0020]
[0021] in, are the action value functions of the charging pile and the robotic arm, are the state value functions of the charging pile and the robotic arm, is the reward function of the charging pile and the robotic arm, E is the expected function, is the discount factor.
[0022] As a preferred solution, according to the scheduling method of the automatic charging system of the present invention, in the modeling step, the neural network parameters are used. and Model the action value functions of the charging pile and the robotic arm respectively to establish an action value network and , and use the neural network parameters and Model the policy value functions of the charging pile and the robotic arm respectively to establish a policy network and .
[0023] As a preferred solution, according to the scheduling method of the automatic charging system of the present invention, in the training step, the loss functions used for training the action value networks of the charging pile and the robotic arm constructed by the neural network in the modeling step using the temporal difference algorithm are defined as follows:
[0024]
[0025] Where E is the expected function, R is the data set collected by the strategy in the past, represents the data sample taken from the data set R, t represents the current moment, t+1 represents the next moment, and According to the strategy Collect actions, is the discount factor, is the regularization coefficient, and Indicates the target network.
[0026] As a preferred solution, according to the scheduling method of the automatic charging system of the present invention, in the training step, the loss functions used to train the strategy network of the charging pile and the robotic arm established by the neural network in the modeling step with the goal of maximizing the cumulative reward are defined as follows:
[0027]
[0028]
[0029] Among them, E is the expected function, R is the data set collected by the strategy in the past, Represents the data sample obtained from the data set R , and Indicates that the charging pile and the robotic arm collect actions according to the strategy. is the regularization coefficient.
[0030] As a preferred solution, according to the scheduling method of the automatic charging system of the present invention, the scheduling method also includes: a safety control step of limiting the actions of the charging pile and the robotic arm so that the actions of the charging pile and the robotic arm meet the constraints.
[0031] As a preferred solution, according to the scheduling method of the automatic charging system of the present invention, in the training step, data samples are sampled from the data set R, and the loss functions of the action value functions of the charging pile and the mechanical arm are calculated respectively. and To train the network and update the target network.
[0032] According to a second aspect of the present invention, there is provided a scheduling device for an automatic charging system, wherein the automatic charging system comprises a charging pile and a robotic arm, wherein the robotic arm performs plugging and unplugging operations on a charging gun on the charging pile so that the charging gun performs charging operations on an electric vehicle, and the scheduling device comprises: a construction unit, which constructs a policy value function and an action value function for the charging pile and the robotic arm, respectively; a modeling unit, which uses a neural network to model the action value function and the policy value function of the charging pile and the robotic arm constructed by the construction unit, so as to establish an action value network and a policy network; a training unit, which uses a temporal difference algorithm to train the action value network of the charging pile and the robotic arm established in the modeling step, and trains the policy network of the charging pile and the robotic arm established in the modeling step with the goal of maximizing the cumulative reward; and a parameter updating unit, which updates the neural network parameters after the training performed by the training unit is completed, so that the policies of the charging pile and the robotic arm are the optimal scheduling policies.
[0033] According to a third aspect of the present invention, there is provided a non-transitory storage medium storing a computer program, which, when executed by a processor, can implement the scheduling method of the automatic charging system according to the first aspect of the present invention.
[0034] According to a fourth aspect of the present invention, a computer program product is provided, which includes computer instructions, and when the computer instructions are executed by a processor, the scheduling method of the automatic charging system according to the first aspect of the present invention can be implemented.
[0035] The beneficial effects of the present invention are as follows: the scheduling method of the present invention transforms the scheduling problem of the suspended rail charging system into the problem of finding the optimal strategy through reinforcement learning through the combination of a multi-agent system and reinforcement learning. The strategy neural network of the suspended rail charging pile and the suspended rail manipulator continuously learns from the parking lot operation process data, realizing the intelligent scheduling of charging scheduling, significantly improving the utilization rate and charging efficiency of the charging pile, and shortening the charging waiting time. In addition, the scheduling method of the automatic charging system of the present invention inserts a safety control module designed according to constraints into the algorithm, thereby ensuring the rationality and safety of the system operation. In addition, due to the diversity and flexibility of the reward function design, the method is highly flexible and scalable, and can adapt to suspended rail automatic charging parking lots of different sizes and configurations, providing electric vehicle users with more efficient, convenient and safe charging services. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A schematic diagram illustrating an exemplary structure of an automatic charging system according to the present invention.
[0037] Figure 2 A flow chart illustrating a charging operation process of the automatic charging system according to the present invention.
[0038] Figure 3 The overall flow chart of the scheduling method of the automatic charging system according to the present invention is illustrated.
[0039] Figure 4 The flowchart illustrates the training and execution process of the scheduling method of the automatic charging system according to the present invention. DETAILED DESCRIPTION
[0040] Exemplary embodiments of the present invention are described in detail below with reference to the accompanying drawings. It should be noted that the relative arrangement of components, numerical expressions and numerical values described in these embodiments do not limit the scope of the present invention unless otherwise specifically stated.
[0041] The scheduling method of the automatic charging system of the present invention can be implemented by a processor in the automatic charging system executing a computer program stored in a memory in the automatic charging system. As an alternative, the automatic charging system can also communicate with a server, and the processor in the server executes a computer program stored in the server or the cloud and feeds back the program execution results to the automatic charging system in real time.
[0042] In the present invention, the term "unit" may refer to a software environment, a hardware environment, or a combination of software and hardware environments. In a software environment, the term "unit" refers to functionality, an application, a software module, a function, a routine, a set of instructions or a program that can be executed by a programmable processor (such as a microprocessor, a central processing unit (CPU) or a specially designed programmable device) or a controller. The memory contains instructions or programs, and when the CPU executes the instructions or programs, the CPU performs operations corresponding to the unit or function. In a hardware environment, the term "unit" refers to a hardware element, circuit, component, physical structure, system, module or subsystem. According to a specific embodiment, the term "unit" may include mechanical, optical or electrical components, or any combination thereof. The term "unit" may include active (such as transistors) or passive (such as capacitors) components. The term "unit" may include a semiconductor device having a substrate and other material layers with various conductive concentrations. It may include a CPU or a programmable processor that can execute a program stored in a memory to perform a specified function. The term "unit" may include a logic element (such as AND, OR) implemented by a transistor circuit or any other switching circuit. In the combination of software and hardware environment, the term "unit" or "circuit" refers to any combination of software and hardware environment as described above. In addition, the term "element", "component", "part" or "device" may also refer to a "circuit" integrated or not integrated with packaging materials.
[0043] The following first describes the system architecture of the automatic charging system of the present invention in conjunction with the accompanying drawings.
[0044] The automatic charging system of the present invention designs a network consisting of a main track and a branch track to ensure that the intelligent body component can move to any designated parking space efficiently and safely. The main track is in a closed state and runs through the entire parking lot, while the branch tracks extending from the main track lead to each parking space, allowing the intelligent body component to stay next to the parking space without affecting the passage of the main track.
[0045] The intelligent components in the automatic charging system of the present invention include charging piles and robotic arms, both of which have the ability to move autonomously on the main track and branch track. The charging pile can move to a specific parking space according to the dispatching instructions to provide power replenishment for the electric vehicle; the robotic arm is responsible for the plug-in and unplugging operations of the charging gun into the charging pile and the charging gun from the charging pile, and also has the ability to move to any parking space to ensure the automation and seamless connection of the charging process.
[0046] An exemplary structure of the automatic charging system of the present invention can be found in Figure 1 shown. Figure 1The shape of the track of the automatic charging system and its positional relationship with the parking spaces in the parking lot are shown. The automatic charging system of the present invention comprises a main track 1 and a branch track 2, as well as a charging pile and a mechanical arm that can move autonomously on the two tracks. Figure 1 As shown, the main track 1 is in a closed loop and is arranged to cover the entire parking lot. Branch tracks 2 extend from the main track 1, and parking spaces 3 for charging are distributed on both sides of each branch track. The branch tracks 2 allow charging piles and robotic arms to stay and operate at the parking spaces without affecting the smooth flow of the main track 1.
[0047] Although Figure 1 The main track 1 shown in the figure is in the shape of a closed ring, but the present invention is not limited thereto. The shape of the main track 1 may also be other closed shapes such as a circle or an ellipse.
[0048] In addition, although in the following description, the main track and branch track in the automatic charging system of the present invention are preferably suspended to facilitate passage and operation while saving space, the present invention is not limited to this. Depending on the height and layout of the parking lot, the automatic charging system of the present invention may also be applicable to ground-type tracks in which the tracks are laid on the ground.
[0049] The charging operation process of the automatic charging system of the present invention will be described in detail below.
[0050] Figure 2 A flow chart illustrating the charging operation process of the automatic charging system according to the present invention is shown. Figure 2 As shown, first, in step S101, the electric vehicle enters a parking space in the automatic charging system. After the electric vehicle stops, it sends a charging request to the automatic charging system in step S102. In step S103, the automatic charging system selects a suitable suspension track charging pile and suspension track manipulator. Specifically, after receiving the charging request from the electric vehicle, the automatic charging system selects the most suitable suspension track charging pile and suspension track manipulator to undertake the charging task of the parking space according to the scheduling method described later.
[0051] In steps S104A and S104B, the selected suspended track charging pile and suspended track mechanical arm are moved to the designated parking space, respectively. It should be understood that the present invention does not specifically limit the order in which the suspended track charging pile and the suspended track mechanical arm arrive at the designated parking space, and the two may arrive successively or simultaneously.
[0052] In step S105, after the suspended track charging pile and the suspended track manipulator move along the track to the corresponding branch track of the designated parking space, the suspended track manipulator will automatically grab the charging gun on the suspended track charging pile and insert it into the electric vehicle. In step S106, it is determined whether the insertion and docking of the charging gun to the electric vehicle by the suspended track manipulator is completed. If it is determined in step S106 that the insertion and docking is not completed (No in step S106), the process returns to step S105 to continue the insertion operation of the charging gun into the electric vehicle. If it is determined in step S106 that the insertion and docking is completed (Yes in step S106), the process proceeds to steps S107A and S107B.
[0053] In step S107A, the suspended track charging pile starts charging the electric vehicle, and in step S107B the suspended track manipulator can leave to perform the task of plugging and unplugging charging guns in other parking spaces. Alternatively, the suspended track manipulator can also stay in place and wait for the charging to end without leaving to perform the task of plugging and unplugging charging guns in other parking spaces. Here, it should be understood that the charging operation of the suspended track charging pile and the departure of the suspended track manipulator can be performed in parallel or one after another, and the present invention does not specifically limit this.
[0054] Next, in step S108, it is determined whether the charging of the electric vehicle by the suspended track charging pile is completed. If it is determined in step S108 that the charging is not completed (no in step S108), the process returns to step S107A to continue charging. If it is determined in step S108 that the charging is completed (yes in step S108), the process proceeds to step S109. In step S109, the automatic charging system uses the scheduling method to reselect a suitable suspended track robot arm to reach the current parking space. Then, in step S120, after the selected suspended track robot arm moves along the track to the current parking space, it pulls out the charging gun and puts it back on the charging pile. Thus, the charging operation is completed, and the charging pile and the robot arm leave the current parking space and perform other charging tasks.
[0055] As an alternative, the suspension track robot arm can also stay in place and wait in step S107B, and after charging is completed, the operation of pulling out the charging gun and putting it back on the charging pile in step S120 is performed without leaving to perform the task of plugging and unplugging charging guns in other parking spaces. In this case, the operation of reselecting the suspension track robot arm in step S109 will be omitted.
[0056] The scheduling method used to select a suitable charging pile and a robotic arm in step S103 will be described in detail below.
[0057] In order to further give full play to the advantages of the entire automatic charging system and improve the utilization rate and charging efficiency of parking lot charging piles, the present invention provides a scheduling method based on multi-agent reinforcement learning for selecting appropriate charging piles and robotic arms for operations according to charging requests, thereby realizing intelligent scheduling of charging scheduling.
[0058] Reinforcement learning usually refers to an agent learning the optimal strategy to maximize the cumulative reward in a Markov process by interacting with the environment, performing actions and receiving feedback (usually rewards or penalties). The concepts involved include state sets, action sets and rewards.
[0059] The state set defined in the present invention includes the position and state of each charging pile, the position and state of each robotic arm, and the state of each parking space; the action set includes three actions of the charging pile and the robotic arm, namely, going to the parking space, working, and resting; the reward is the sum of the waiting time of all charging requests, the moving energy consumption of all robots, and the weighted average of the waiting time of vehicles in the earlier queue.
[0060] The two intelligent agents involved in the present invention, the suspended track charging pile and the suspended track robotic arm, independently interact with the environment in the common environment of the automatic charging system created above, and use the rewards of environmental feedback to improve their own strategies to obtain higher cumulative rewards. The suspended track charging pile and the suspended track robotic arm have the same goals and are in a fully cooperative relationship, so they obtain the same rewards and have the same reward function, and use the feedback rewards to continuously improve their respective strategies.
[0061] In the present invention, the scheduling method of the automatic charging system mainly includes a function construction step, a network modeling step, a network training step and a parameter updating step. Figure 3 The scheduling method of the automatic charging system of the present invention is described.
[0062] like Figure 3 As shown, first, in the function construction step S201, a strategy value function and an action value function are constructed for the charging pile and the robotic arm respectively, and the constructed strategy value function and action value function will be described in detail below.
[0063] Next, in the network modeling step S202, the action value function and strategy value function of the charging pile and the robot arm constructed in step S201 are modeled using a neural network to establish an action value network and a strategy network. Specifically, two neural networks are used: The action value functions of the suspended rail charging pile and the suspended rail manipulator are modeled respectively, and two neural networks are used The strategy value functions of the suspended track charging pile and the suspended track robotic arm are modeled separately.
[0064] Then in the network training step S203, the action value network of the charging pile and the robotic arm established in step S202 is trained using the temporal difference algorithm, and the strategy network of the charging pile and the robotic arm established in step S202 is trained with the goal of maximizing the cumulative reward.
[0065] After the network training in step S203 is completed, the neural network parameters are updated in step S204 so that the strategies of the charging pile and the robotic arm are the optimal scheduling strategies. Specifically, in step S204, the strategies of the suspended track charging pile and the suspended track robotic arm are gradually made the optimal scheduling strategies through the soft update of the neural network parameters.
[0066] The following is a detailed description of each step in the scheduling method of the present invention.
[0067] The suspended track charging pile and the suspended track manipulator in the automatic charging system of the present invention are two groups of intelligent agents in reinforcement learning, each using different strategies, and their respective states are composed of positions and progress of tasks undertaken. , state set It consists of the status of all suspended track charging piles and suspended track robotic arms and the status of all parking spaces. Actions of suspended track charging piles and suspended track robotic arms , action set The two groups of agents, the suspended track charging pile and the suspended track robotic arm, are in a fully cooperative relationship with the same goal, so they have the same reward function. , the reward function It is preferably expressed by the following formula (1):
[0068] Formula (1)
[0069] in, Indicates the current state Next, the suspended rail charging pile performs an action And the suspension track robot performs the action , No. The waiting time for a charging request; Indicates that the hanging rail charging pile performs an action Total energy consumption; Indicates that the suspension track robot arm performs an action Total energy consumption; are all weight coefficients.
[0070] As a preferred solution, the design of the reward function in the present invention also takes into account the waiting time of all charging requests and the system energy consumption, and a fourth term is added to the formula to prevent the situation where the waiting time of a single request is too long. The waiting time in the reward function can be calculated based on the moving speed and relative position of the suspended track charging pile and the suspended track manipulator, and the energy consumption can be calculated based on the unit energy consumption of the suspended track charging pile and the suspended track manipulator moving a certain distance.
[0071] In order to achieve the goal of maximizing the cumulative reward, the suspended track charging pile and the suspended track manipulator hope to learn a strategy for selecting actions from the current state to achieve this goal. In the present invention, the strategies (strategy functions) of the suspended track charging pile and the suspended track manipulator are respectively expressed as Indicates that it is used to represent the state S to the action and 's mapping.
[0072] Based on the design of the above-mentioned reward function, for the suspended track charging pile and the suspended track manipulator, the objectives of the multi-agent reinforcement learning of the present invention can be defined as the objective functions shown in the following formulas (2) and (3):
[0073] (Formula 2)
[0074] (Formula 3)
[0075] The objective function above and It represents the goal to be achieved by training the policy functions of the charging pile and the robotic arm, that is, the optimal policy. is the expectation function, which means calculating the average value of the contents in the brackets ([]) following it. for The entropy of represents the measure of the randomness of a random variable. In this multi-agent reinforcement learning, and Indicates that the suspended track charging pile strategy and the suspended track manipulator strategy are in state The degree of randomness. is the regularization coefficient, which is used to control the importance of entropy.
[0076] The scheduling method of the present invention uses the offline strategy algorithm SAC to train the strategies of the suspended track charging pile and the suspended track manipulator. In the SAC algorithm, the idea of maximum entropy reinforcement learning is used. Therefore, the goal of the multi-agent reinforcement learning of the present invention is to maximize the cumulative reward by introducing entropy and regularization coefficients, and to make the strategy more random and more exploratory, thereby accelerating the strategy learning of the suspended track charging pile and the suspended track manipulator, and reducing the possibility of the strategy falling into a poor local optimum.
[0077] Based on the objective function designed above, since the strategy is made more random and more exploratory, the Bellman equation in the reinforcement learning of the present invention becomes a soft form. Specifically, the Bellman equation of the action-value function constructed in the present invention is shown in the following formula 4 and formula 5:
[0078] (Formula 4)
[0079] (Formula 5)
[0080] in, They are the action value functions of the suspended track charging pile and the suspended track manipulator respectively; It is the expectation, which has the same meaning as the expectation function above, and is used to calculate the average value of the content in the brackets ([]) that follow it; The discount factor is a hyperparameter set according to the situation (manually set), which is mainly used to balance immediate rewards and future rewards to avoid infinite rewards; They are the state value functions of the suspended track charging pile and the suspended track robotic arm, respectively, which are used to measure the expected rewards that the suspended track charging pile and the suspended track robotic arm can obtain from a certain state when the current strategy is adopted. The state value function is expressed as:
[0081]
[0082]
[0083] In order to utilize the powerful high-dimensional data learning ability and nonlinear function approximation ability of neural networks, two neural networks are used in the present invention. The action value functions of the suspended track charging pile and the suspended track manipulator constructed above are modeled respectively, and two neural networks are used The strategy value functions of the suspended track charging pile and the suspended track robotic arm are modeled separately.
[0084] Specifically, using neural networks and The action value function of the suspended track charging pile and the suspended track manipulator (parameters are ) and the policy function (with parameters ) to represent (model) to establish the action value network and strategy network. The action value network is specifically represented as and , the policy network is represented as and , they can also be called trained networks for use in subsequent training steps.
[0085] In the scheduling method of the automatic charging system of the present invention, according to the temporal difference algorithm, the loss function of the action value network used to train the action value functions of the charging pile and the robot arm modeled by the neural network is defined as the following formula 6 and formula 7:
[0086] (Formula 6)
[0087] (Formula 7)
[0088] Where E is the expected function, which means calculating the average value of the content in brackets ([]); R is the data set collected by the strategy in the past; represents the data sample taken from the data set R, where t represents the current moment (also called the current time step), and t+1 represents the next moment after the current moment (also called the next time step after the current time step); and Indicates collecting actions according to the strategy, that is, determining the actions taken by the charging pile and the robot arm at the next moment according to the strategy; the same as above is the discount factor, which is an artificially set hyperparameter; is the regularization coefficient mentioned above. Because SAC is a discrete strategy algorithm, in order to make the training more stable, the present invention preferably uses the target network and , used to represent the trained network and The objective to be achieved is that the network parameters in the present invention can be updated by a soft update method, which is a slow update method because it only updates a part or a certain proportion of the network parameters, and is more conducive to the stability of neural network training.
[0089] In addition, in the scheduling method of the automatic charging system of the present invention, the loss functions of the strategies used to train the strategy networks of the charging pile and the robotic arm modeled by the neural network with the goal of maximizing the cumulative reward are defined as shown in the following formulas 8 and 9 respectively:
[0090] (Formula 8)
[0091] (Formula 9)
[0092] Where E is the expected function, which is the same as above. It means calculating the average value of the content in brackets ([]), and R is the data set collected by the strategy in the past. Represents a data sample taken from the data set R (state ), and Indicates that the charging pile and the robotic arm collect actions according to the strategy, the same as above is the regularization coefficient mentioned above.
[0093] On the other hand, as a preferred embodiment, the scheduling method of the present invention also includes a safety control step for limiting the actions of the suspended track charging pile and the suspended track robotic arm so as to satisfy the constraints (constraint conditions). Specifically, the scheduling method of the present invention additionally inserts a safety control module (unit) to limit the actions of the suspended track charging pile and the suspended track robotic arm to satisfy the constraints. In the present invention, the main constraints are, for example, that the suspended track charging pile and the suspended track robotic arm moving on the main track cannot collide, and that the suspended track charging pile is in the charging state and the suspended track robotic arm is in the plugging and unplugging state, and cannot be interrupted to perform other tasks. The specific operation method is that the safety control module performs real-time control of the action reward function Check, if there is a violation of the above constraints, the current reward function Make corrections, such as changing it to a very small negative value, and impose a huge penalty on the action, so that the strategy avoids violating the constraints in subsequent learning.
[0094] In addition, it should be understood that in the present invention, the safety control module can be implemented in the form of hardware installed in the automatic charging system, or can be implemented in the form of software of method steps executed by a processor in the automatic charging system.
[0095] The specific training and execution process of the scheduling method of the automatic charging system in the present invention is described in detail below.
[0096] Reference Figure 4 The training and execution process of the scheduling algorithm based on multi-agent reinforcement learning of the automatic charging system of the present invention is described.
[0097] like Figure 4 As shown, first in step S301, the automatic charging system is initialized. Specifically, the automatic charging system is initialized in the already arranged site, using random network parameters Initialize the action value network of the suspended track charging pile and the suspended track manipulator and , and use random network parameters Initialize the policy network of the suspended track charging pile and the suspended track manipulator and .
[0098] Next, in step S302, the target network of the action value network is initialized. Specifically, the same parameters are copied arrive , respectively initialize the target network of the action value network and .
[0099] In step S303, the initial state of the current system environment is obtained In step S304, the time step Set to 1. Then, in step S305, determine whether t is less than or equal to the set time threshold T. If it is determined in step S305 that t is less than or equal to T (yes in step S305), the process proceeds to step S306. In step S306, the suspended track charging pile and the suspended track manipulator select their respective actions according to the current strategy. and Next, in step S307, the suspended track charging pile and the suspended track manipulator perform the action selected in step S306 and obtain a reward. , and changes the environment state to .
[0100] Then, in step S308, the safety control module checks the constraints. Specifically, the actions and environmental states of the suspended track charging pile and the suspended track robot are transmitted to the safety control module to check whether the constraints are violated. If the constraints are violated, the rewards obtained are corrected and penalties are given. Then, in step S309, the collected data samples are Put into the data playback pool (dataset).
[0101] Next, in step S310, data is sampled from the data playback pool R, and the loss function of the action value function is calculated respectively. and And perform network training, and calculate the loss function of the strategy separately and And perform network training to update the target network. Preferably, in the present invention, network training refers to using the data continuously collected in the above process (these data are stored in R), calculating the above-mentioned loss function, and then using conventional methods for training neural networks such as gradient descent for training.
[0102] After the processing of step S310 is completed, the time step t is increased by 1, that is, t=t+1 is set, and the processing returns to step S305 to determine whether the time step t is less than or equal to the time threshold T. From 1 to Steps S306 to S310 are executed in a loop until t is greater than T.
[0103] If it is determined in step S305 that t is greater than T (step S305 is no), the process proceeds to step S311. In step S311, it is determined whether e is less than E, where e represents the cumulative reward obtained within this T time, and E represents the expected (satisfactory) cumulative reward, for example, E is a score set by the system (for example, given subjectively by humans). If it is determined in step S311 that e is less than E (step S311 is yes), it means that the score of the cumulative reward is not high enough, that is, the cumulative reward obtained within this T time does not meet the expectations of the system, then the process returns to step S303, and the process is started again until a satisfactory scheduling strategy for the suspension track charging pile and the suspension track manipulator is obtained. If it is determined in step S311 that e is greater than or equal to E (step S311 is no), it means that the cumulative reward obtained within this T time meets the expectations of the system, and the process ends. A satisfactory scheduling strategy for the suspension track charging pile and the suspension track manipulator is obtained through training.
[0104] Finally, the strategy of the suspended track charging pile and the suspended track manipulator gradually becomes the optimal scheduling strategy through the soft update of the neural network parameters. The soft update of the neural network parameters adopted in the present invention is more conducive to the stability of the neural network training.
[0105] Through the above process, the scheduling strategy in the current suspended rail automatic charging parking lot can be pre-trained, and data can be continuously collected during the subsequent actual operation to conduct real-time training of the scheduling strategy, so as to dynamically respond to different situations, reduce the waiting time of charging requests, and improve the charging efficiency of the parking lot.
[0106] In summary, the scheduling method of the automatic charging system of the present invention transforms the scheduling problem of the suspended rail charging system into the problem of finding the optimal strategy through reinforcement learning through the combination of multi-agent system and reinforcement learning. The strategy neural network of the suspended rail charging pile and the suspended rail manipulator continuously learns from the parking lot operation process data, realizing the intelligent scheduling of charging scheduling, significantly improving the utilization rate and charging efficiency of the charging pile, and shortening the charging waiting time. In addition, the scheduling method of the automatic charging system of the present invention inserts a safety control module designed according to constraints into the algorithm, thereby ensuring the rationality and safety of the system operation. In addition, due to the diversity and flexibility of the reward function design, the method is highly flexible and scalable, and can adapt to suspended rail automatic charging parking lots of different sizes and configurations, providing electric vehicle users with more efficient, convenient and safe charging services.
[0107] Other Implementations
[0108] The embodiments of the present invention may also be implemented by reading out and executing computer executable instructions (e.g., one or more programs) recorded on a storage medium (which may also be more completely referred to as a "non-transitory computer-readable storage medium") to perform one or more functions in the above-mentioned embodiments, and / or a computer of a system or device including one or more circuits (e.g., an application-specific integrated circuit (ASIC)) for performing one or more functions in the above-mentioned embodiments, and the embodiments of the present invention may be implemented by, for example, reading out and executing the computer executable instructions from the storage medium by the computer of the system or device to perform one or more functions in the above-mentioned embodiments, and / or controlling the one or more circuits to perform one or more functions in the above-mentioned embodiments. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessing unit (MPU)), and may include a network of separate computers or separate processors to read out and execute the computer executable instructions. The computer executable instructions may be provided to the computer, for example, from a network or the storage medium. The storage medium may include, for example, one or more of a hard disk, a random access memory (RAM), a read-only memory (ROM), a memory of a distributed computing system, an optical disk (such as a compact disc (CD), a digital versatile disc (DVD), or a Blu-ray Disc (BD)™), a flash memory device, a memory card, and the like.
[0109] Although the present invention has been described above with reference to exemplary embodiments, the above embodiments are only for illustrating the technical concept and features of the present invention and cannot be used to limit the protection scope of the present invention. Any equivalent variations or modifications made according to the spirit and essence of the present invention should be included in the protection scope of the present invention.
Claims
1. A scheduling method for an automatic charging system, characterized in that: The automatic charging system includes a charging pile and a mechanical arm, and the mechanical arm performs plugging and unplugging operations on the charging gun on the charging pile so that the charging gun performs charging operations on the electric vehicle. The scheduling method includes the following steps: A construction step (S201) constructs a strategy value function and an action value function for each charging pile and a robotic arm, and constructs the same reward function for the charging pile and the robotic arm; A modeling step (S202), using a neural network to model the action value function and the strategy value function of the charging pile and the robotic arm constructed in the construction step, so as to establish an action value network and a strategy network; A training step (S203), using a temporal difference algorithm to train the action value network of the charging pile and the robotic arm established in the modeling step, and training the strategy network of the charging pile and the robotic arm established in the modeling step with the goal of maximizing the cumulative reward; A parameter updating step (S204), in which the neural network parameters are updated after the training step is completed, so that the strategies of the charging pile and the robotic arm are the optimal scheduling strategies; Among them, the objective function of the strategy value function constructed for the charging pile and the robotic arm in the construction step is defined as follows: in and Represent the strategy value functions for charging piles and robotic arms respectively and The objective function is Respectively represent the actions of the charging pile and the robotic arm, Indicates the status of the entire environment of the automatic charging system, The state set consists of the states of all charging piles and robotic arms and the states of all parking spaces. is the expected function, is the reward function of the charging pile and the robotic arm, and They are and The entropy indicates that the charging pile strategy and the robotic arm strategy are in the state The degree of randomness under is the regularization coefficient.
2. The scheduling method of the automatic charging system according to claim 1, characterized in that: The reward functions of the charging pile and the robotic arm are defined as follows: in, Indicates in status Next, the charging pile performs an action And the robot performs the action , No. The waiting time for a charging request; Indicates that the charging pile performs an action The total energy consumption, and Indicates that the robot arm performs an action Total energy consumption; are all weight coefficients.
3. The scheduling method of the automatic charging system according to claim 1, characterized in that: In the construction step, based on the objective function of the constructed strategy value function, the action value functions are constructed for the charging pile and the robotic arm respectively as follows: in, are the action value functions of the charging pile and the robotic arm, are the state value functions of the charging pile and the robotic arm, is the reward function of the charging pile and the robotic arm, E is the expected function, is the discount factor.
4. The scheduling method of the automatic charging system according to claim 1, characterized in that: In the modeling step, the neural network parameters are used and Model the action value functions of the charging pile and the robotic arm respectively to establish an action value network and , and use the neural network parameters and Model the policy value functions of the charging pile and the robotic arm respectively to establish a policy network and .
5. The scheduling method of the automatic charging system according to claim 4, characterized in that: In the training step, the loss functions used to train the action value networks of the charging pile and the robotic arm constructed by the neural network in the modeling step using the temporal difference algorithm are defined as follows: Where E is the expected function, R is the data set collected by the strategy in the past, represents the data sample taken from the data set R, t represents the current moment, t+1 represents the next moment, and Indicates that actions are collected according to the strategy. is the discount factor, is the regularization coefficient, and Indicates the target network.
6. The scheduling method of the automatic charging system according to claim 4, characterized in that: In the training step, the loss functions used to train the policy networks of the charging pile and the robotic arm established by the neural network in the modeling step with the goal of maximizing the cumulative reward are defined as follows: Among them, E is the expected function, R is the data set collected by the strategy in the past, Represents the data sample obtained from the data set R , and Indicates that the charging pile and the robotic arm collect actions according to the strategy. is the regularization coefficient.
7. The scheduling method of the automatic charging system according to claim 1, characterized in that: The scheduling method further includes: The safety control step limits the movements of the charging pile and the robotic arm so that the movements of the charging pile and the robotic arm meet the constraints.
8. The scheduling method of the automatic charging system according to claim 5, characterized in that: In the training step, data samples are sampled from the data set R, and the loss functions of the action value functions of the charging pile and the robotic arm are calculated respectively. and To train the network and update the target network.
9. A dispatching device for an automatic charging system, characterized in that: The automatic charging system includes a charging pile and a mechanical arm, the mechanical arm performs plugging and unplugging operations on the charging gun on the charging pile so that the charging gun performs charging operations on the electric vehicle, and the scheduling device includes: The construction unit constructs the policy value function and action value function for the charging pile and the robotic arm respectively, and constructs the same reward function for the charging pile and the robotic arm; A modeling unit, which uses a neural network to model the action value function and the strategy value function of the charging pile and the robotic arm constructed by the construction unit, so as to establish an action value network and a strategy network; A training unit, which uses a temporal difference algorithm to train the action value network of the charging pile and the robotic arm established in the modeling step, and trains the policy network of the charging pile and the robotic arm established in the modeling step with the goal of maximizing the cumulative reward; A parameter updating unit, which updates the neural network parameters after the training performed by the training unit is completed, so that the strategies of the charging pile and the robotic arm are the optimal scheduling strategies; Among them, the objective function of the strategy value function constructed by the construction unit for the charging pile and the robotic arm is defined as follows: in and Represent the strategy value functions for charging piles and robotic arms respectively and The objective function is Respectively represent the actions of the charging pile and the robotic arm, Indicates the status of the entire environment of the automatic charging system, The state set consists of the states of all charging piles and robotic arms and the states of all parking spaces. is the expected function, is the reward function of the charging pile and the robotic arm, and They are and The entropy indicates that the charging pile strategy and the robotic arm strategy are in the state The degree of randomness under is the regularization coefficient.
10. A non-transitory storage medium storing a computer program, wherein the computer program, when executed by a processor, can implement the scheduling method of the automatic charging system according to any one of claims 1 to 8.
11. A computer program product, comprising computer instructions, which when executed by a processor can implement the scheduling method of the automatic charging system according to any one of claims 1 to 8.
Citation Information
Patent Citations
Vehicle-to-vehicle charging system and method and electric vehicle
CN112776624A
Track mobile charging pile
CN119142178A
Charging pile scheduling method and system based on deep reinforcement learning
CN117332954A
Charging control method and charging control device of automobile charging equipment
CN119142176A