New energy vehicle charging infrastructure and edge computing gateway energy storage cooperative scheduling method

Through the combination of deep reinforcement learning and multi-head attention mechanism model, the scheduling problems of new energy vehicle charging infrastructure and edge computing gateway energy storage system are solved, intelligent resource optimization and maximum benefits are achieved, and the operation efficiency and economics of charging infrastructure and energy storage system are improved.

CN120280966AActive Publication Date: 2025-07-08CHONGQING ARCHITECTURAL DESIGN INST CO LTD

Patent Information

Application Number
CN202510764305.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The scheduling methods of the existing charging infrastructure and edge computing gateway energy storage systems of new energy vehicles rely on manual experience or simple rules, making it difficult to adapt to complex and changing charging needs and energy storage states, resulting in low resource allocation efficiency and unable to maximize the overall system benefits.

Method used

The deep reinforcement learning algorithm and the multi-head attention mechanism model are used to construct state space and action space, and the dependence between parameters is captured through multi-head parallel calculation, multi-dimensional feature information is formed, scheduling action probability distribution is output, and stable coordinated scheduling strategies are obtained through real-time reward and training optimization.

Benefits of technology

It has realized intelligent coordinated scheduling of new energy vehicle charging infrastructure and edge computing gateway energy storage systems, optimized resource allocation, reduced overload of charging piles and idle energy storage nodes, improved power resource utilization efficiency, controlled operating costs, and increased energy storage system benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280966A_ABST
    Figure CN120280966A_ABST
Patent Text Reader

Abstract

The invention discloses a new energy vehicle charging capital construction and edge computing gateway energy storage cooperative scheduling method, which comprises the following steps of: constructing a state space and an action space containing various parameters of a charging capital construction and edge computing gateway energy storage system, extracting parameter characteristics by utilizing a multi-head attention mechanism model, and analyzing an association relationship; and then scheduling action probability distribution is output by means of a strategy network in a deep reinforcement learning algorithm so as to select a cooperative scheduling action, after the action is executed, an instant reward including operation cost, charging and discharging income, charging waiting time variation and the like is obtained, a state is updated, and experience is stored for iteratively training the strategy network and a value network. According to the method, reward calculation is completed by optimizing training processes such as priority experience playback and a gating mechanism and comprehensively considering equipment loss cost, a stable collaborative scheduling strategy is finally obtained, efficient collaboration of new energy vehicle charging capital construction and an edge computing gateway energy storage system is achieved, and the overall resource scheduling benefit and the system operation performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of charging infrastructure and energy storage supervision of edge computing gateways, and particularly to a collaborative scheduling method for new energy vehicle charging infrastructure and energy storage of edge computing gateways. Background Art

[0002] In recent years, the global new energy vehicle industry has developed rapidly, and the market ownership has continued to climb. The construction of supporting charging infrastructure has been accelerating continuously. However, due to the random and volatile characteristics of the charging demand of new energy vehicles, traditional charging infrastructure often has problems of low efficiency in power resource allocation. At the same time, the combination of edge computing gateway technology and energy storage systems has brought new possibilities for the efficient management of power resources. Through the decentralized and tamper-proof characteristics of edge computing gateways, secure and trusted interaction and collaboration between nodes in the energy storage system can be achieved. Therefore, researching the collaborative scheduling method for charging infrastructure and energy storage of edge computing gateways has important practical significance and application value.

[0003] Currently, the scheduling methods of most new energy vehicle charging infrastructure and energy storage systems are relatively traditional, relying on manual experience or simple rule settings, and it is difficult to adapt to complex and changing charging demands and energy storage states. When the existing technology deals with the collaborative scheduling of charging infrastructure and energy storage of edge computing gateways, it does not sufficiently explore the complex correlation relationships between various parameters. Parameters such as the real-time power and demand quantity of charging infrastructure, and the stored electricity and charge-discharge power of the energy storage system of the edge computing gateway affect each other. However, traditional methods often view these parameters in isolation and do not fully utilize the internal connections between them, resulting in the scheduling strategy being unable to maximize the overall benefits of the system, and there are obvious defects in aspects such as operation cost control and revenue improvement. Summary of the Invention

[0004] In order to overcome the disadvantages and deficiencies existing in the prior art, the present invention provides a collaborative scheduling method for new energy vehicle charging infrastructure and energy storage of edge computing gateways.

[0005] The technical solution adopted by the present invention is a collaborative scheduling method for new energy vehicle charging infrastructure and energy storage of edge computing gateways, including the following steps:

[0006] Step S1: Construct a state space including the real-time charging power, charging demand quantity of each charging pile of new energy vehicle charging infrastructure, and the stored electricity and charge-discharge power of each energy storage node of the edge computing gateway energy storage system, and define an action space including the probability of selecting scheduling actions based on the deep reinforcement learning algorithm. The scheduling actions of the action space include the charging power allocation for each charging pile in the charging infrastructure and the charge-discharge decisions of each energy storage node of the edge computing gateway energy storage system;

[0007] Step S2: Use the multi-head attention mechanism model to perform feature extraction and correlation analysis on the parameters of the collaborative scheduling of charging infrastructure and edge computing gateway energy storage in the state space. Capture the dependencies between different parameters through multi-head parallel computing to form a set of feature vectors including multi-dimensional feature information;

[0008] Step S3: Based on the policy network in the deep reinforcement learning algorithm, use the set of feature vectors as input, output the probability distribution of scheduling actions in the action space, and select the collaborative scheduling actions for the charging infrastructure and edge computing gateway energy storage system according to the probability distribution;

[0009] Step S4: Execute the selected collaborative scheduling actions, obtain the immediate rewards including the operating cost of the charging infrastructure, the charge and discharge benefits of the edge computing gateway energy storage system, and the change in the charging waiting time of new energy vehicle users, and update the parameters in the state space;

[0010] Step S5: Store the state space, the selected scheduling actions, the immediate rewards, and the updated state space in the experience replay pool. Randomly sample data from the experience replay pool based on the deep reinforcement learning algorithm to construct a training data set;

[0011] Step S6: Use the training data set to iteratively train and optimize the policy network and the value network in the deep reinforcement learning algorithm, return to Step S3 and repeat the execution until the preset training termination condition is met, and obtain a stable collaborative scheduling policy for the new energy vehicle charging infrastructure and edge computing gateway energy storage.

[0012] Further, the multi-head attention mechanism model used in Step S2 satisfies the following formula in its calculation process:

[0013]

[0014] Among them, represents the query matrix composed of the parameters of the collaborative scheduling of charging infrastructure and edge computing gateway energy storage in the state space, including the real-time charging power matrix of charging piles and the stored power matrix of energy storage nodes; represents the key matrix, corresponding to the query matrix ; represents the value matrix; are the weight matrices of the query, key, and value of different heads respectively; is the weight matrix of the multi-head output; is the dimension of the key matrix; is the number of heads of the multi-head attention mechanism; Concat is the concatenation operation.

[0015] Further, the policy network of the deep reinforcement learning algorithm based on in Step S3 satisfies the following formula in the calculation of the output probability distribution of scheduling actions:

[0016]

[0017] Among them, represents the probability that, under the policy network with parameter , the state takes the action . is the state in the state space constructed for step S1, including various parameters of the charging infrastructure and the energy storage system of the edge computing gateway; is the scheduling action in the action space defined for step S1; is the action space; is the output score function of the policy network, and its parameter is optimized through subsequent training.

[0018] Furthermore, the immediate reward obtained in step S4 is calculated to satisfy the following formula:

[0019]

[0020] Among them, is the immediate reward; is the operating cost of the charging infrastructure, including the energy consumption cost of the charging pile; is the charging and discharging income of the energy storage system of the edge computing gateway; is the change in the charging waiting time of new energy vehicle users; are the weight coefficients of the operating cost of the charging infrastructure, the charging and discharging income of the energy storage system of the edge computing gateway, and the change in the charging waiting time of new energy vehicle users, respectively.

[0021] Furthermore, when the deep reinforcement learning algorithm extracts sample data from the experience replay pool to construct a training data set in step S5, a prioritized experience replay mechanism is adopted, and the sample extraction probability satisfies the following formula:

[0022]

[0023] Among them, represents the extraction probability of the th sample in the experience replay pool; is the temporal difference error of the th sample, reflecting the importance of the sample to learning; is the total number of samples in the experience replay pool; is the parameter for controlling the sampling bias.

[0024] Furthermore, when training and optimizing the value network in step S6, the following loss function formula is adopted:

[0025]

[0026]

[0027] Among them, is the loss function of the value network, are the value network parameters; represents the expectation; are respectively the state, action, immediate reward, and updated state of the sample extracted in step S5; is the valuation of the value network in state ; is the valuation of the target value network in state ; is the discount factor, used to balance the immediate reward and future rewards.

[0028] Furthermore, after using the multi-head attention mechanism model to extract features from the parameters in step S2, the feature fusion formula is further adopted:

[0029]

[0030] Among them, is the fused feature vector; is the feature vector extracted from the charging infrastructure parameters by the multi-head attention mechanism; is the feature vector extracted from the edge computing gateway energy storage system parameters by the multi-head attention mechanism; is the fusion weight coefficient.

[0031] Furthermore, before the policy network outputs the scheduling action probability distribution in step S3, a gating mechanism is introduced, and its calculation satisfies the following formula:

[0032]

[0033] Among them, is the gating signal; is the activation function; are the weight matrix and bias vector of the gating mechanism; is the feature extraction result of the policy network for state ; is the scheduling action probability distribution output by the original policy network; is the prior probability distribution; is the scheduling action probability distribution adjusted by the gating mechanism.

[0034] Furthermore, when calculating the immediate reward in step S4, combining the equipment loss costs of the charging infrastructure and the edge computing gateway energy storage system, the immediate reward calculation formula is adjusted to:

[0035]

[0036] Among them, is the equipment loss cost of the charging infrastructure and the energy storage system of the edge computing gateway; is the weight coefficient of the equipment loss cost.

[0037] Beneficial effects: The present invention proposes a collaborative scheduling method for a new energy vehicle charging infrastructure and an edge computing gateway energy storage. This method effectively overcomes the defects of the prior art through the organic combination of a deep reinforcement learning algorithm and a multi-head attention mechanism model. In terms of intelligence, it abandons the traditional scheduling method that relies on manual experience or simple rules. The deep reinforcement learning algorithm can automatically learn and optimize the scheduling strategy based on the real-time state of the charging infrastructure and the edge computing gateway energy storage system. In the face of the charging peak of new energy vehicles, it can intelligently coordinate the power distribution between the charging piles and the energy storage nodes, avoid overloading of the charging piles and idleness of the energy storage nodes, reduce the waiting time of users, and improve the utilization efficiency of power resources. In terms of parameter correlation analysis, the multi-head attention mechanism model can deeply explore the complex dependence relationships between the parameters of the charging infrastructure and the edge computing gateway energy storage system, and no longer view parameters such as real-time power, demand quantity, and stored power in isolation. By extracting features and performing correlation analysis on these parameters, multi-dimensional feature information is formed, providing a comprehensive basis for scheduling decisions, thereby maximizing the overall benefits of the system, effectively controlling the operating costs of the charging infrastructure, increasing the charging and discharging benefits of the edge computing gateway energy storage system, and at the same time making the scheduling strategy more comprehensive and scientific by considering the equipment loss cost, etc., and giving full play to the advantages of the collaborative scheduling of the charging infrastructure and the edge computing gateway energy storage. Description of the Drawings

[0038] Figure 1 is the flowchart of the method operation steps of the present invention;

[0039] Figure 2 is the implementation diagram of the method operation unit of the present invention. Detailed Embodiments

[0040] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following further describes this application in detail with reference to the drawings and specific embodiments.

[0041] As Figure 1 shown, a collaborative scheduling method for a new energy vehicle charging infrastructure and an edge computing gateway energy storage includes the following steps:

[0042] Step S1: Construct a state space including the real-time charging power of each charging pile in the new energy vehicle charging infrastructure, the number of charging demands, the stored power of each energy storage node in the edge computing gateway energy storage system, and the charge and discharge power. Define an action space including the scheduling action selection probability based on the deep reinforcement learning algorithm. The scheduling actions in the action space include the charging power allocation for each charging pile in the charging infrastructure and the charge and discharge decisions for each energy storage node in the edge computing gateway energy storage system;

[0043] Specifically, in the process of collaborative scheduling of the new energy vehicle charging infrastructure and the edge computing gateway energy storage, step S1 is the basic construction link of the entire scheduling method. First, construct a state space, which covers the key information of the operation of the new energy vehicle charging infrastructure and the edge computing gateway energy storage system. For the new energy vehicle charging infrastructure, it includes the real-time charging power of each charging pile, which directly reflects the current workload of the charging pile. Different powers correspond to different charging speeds and energy consumptions; the number of charging demands reflects the number of vehicles waiting to be charged at a specific moment and is an important indicator to measure the urgency of charging demands. In the aspect of the edge computing gateway energy storage system, the stored power of each energy storage node determines its schedulable energy reserve. When the stored power is sufficient, it can be used to support the charging peak, and when it is insufficient, the charging needs to be reasonably arranged; the charge and discharge power parameter affects the energy interaction efficiency between the energy storage node and the external power grid or the charging infrastructure. Define an action space based on the deep reinforcement learning algorithm. The scheduling actions in the action space focus on the direct control of the charging infrastructure and the edge computing gateway energy storage system. For the charging infrastructure, it involves the allocation of the charging power of each charging pile. For example, during the charging demand peak, according to the real-time status of the charging pile and the situation of the energy storage node, the power is reasonably allocated to avoid overloading some charging piles while some are idle; in the edge computing gateway energy storage system, the scheduling action is manifested as the charge and discharge decisions for each energy storage node, such as determining when to charge the energy storage node using low-cost electricity and when to discharge during the peak electricity consumption period to obtain benefits. Each action corresponds to a certain selection probability, providing a basis for subsequent scheduling decisions.

[0044] Step S2: Use the multi-head attention mechanism model to perform feature extraction and correlation analysis on the parameters of the collaborative scheduling of the charging infrastructure and the edge computing gateway energy storage in the state space, capture the complex dependence relationships between different parameters through the multi-head parallel computing method, and form a set of feature vectors including multi-dimensional feature information;

[0045] Specifically, the core of step S2 is to deeply process the parameters in the state space using the multi-head attention mechanism model. In the scenario of coordinated scheduling of new energy vehicle charging infrastructure and edge computing gateway energy storage, there are numerous and interrelated parameters in the state space, and it is difficult for traditional methods to fully explore their internal connections. The multi-head attention mechanism model captures the complex dependence relationships between different parameters through the unique method of multi-head parallel computing. Taking the two parameters of the real-time charging power of the charging pile and the stored power of the energy storage node as an example, multiple "heads" of the model will analyze them from different perspectives. One "head" may focus on studying the impact of high-power charging of the charging pile on the overall cost when the energy storage node has sufficient power; another "head" may focus on the feasibility and benefits of the energy storage node discharging to supplement when the charging demand surges and the charging pile power is insufficient. Through this multi-head parallel computing, the model comprehensively extracts the features of the parameters for the coordinated scheduling of the charging infrastructure and the edge computing gateway energy storage, converts the information contained in each parameter into a feature vector, and further analyzes the associations between these feature vectors. Finally, a set of feature vectors containing multi-dimensional feature information is formed. These sets of feature vectors are no longer isolated parameter information but deep information that integrates the complex relationships between parameters, providing a more valuable basis for subsequent scheduling decisions and enabling the system to select scheduling actions based on a more comprehensive and in-depth understanding.

[0046] Step S3: Based on the policy network in the deep reinforcement learning algorithm, use the set of feature vectors as input, output the probability distribution of the scheduling actions in the action space, and select the coordinated scheduling actions for the charging infrastructure and the edge computing gateway energy storage system according to the probability distribution;

[0047] Specifically, step S3 relies on the policy network in the deep reinforcement learning algorithm, takes the set of feature vectors generated in step S2 as input, and makes decisions on scheduling actions. The policy network plays the role of a "decision-making brain" throughout the collaborative scheduling process. It analyzes the comprehensive state of the current new energy vehicle charging infrastructure and the edge computing gateway energy storage system based on the input multi-dimensional feature information, and then outputs the probability distribution of each scheduling action in the action space. For example, when the set of feature vectors shows that the current charging demand is at a peak, some charging piles are approaching full load, and some nodes in the edge computing gateway energy storage system have sufficient power, the policy network will calculate the probabilities of scheduling actions such as increasing the discharge of energy storage nodes and adjusting the power of some charging piles. Based on this probability distribution, the system will select the corresponding collaborative scheduling action. This selection process does not simply choose the action with the highest probability, but performs random sampling according to the probability distribution. This method not only considers the optimal scheduling direction in the current state but also introduces a certain degree of exploration to avoid falling into local optimal solutions. By continuously outputting probability distributions and selecting actions according to different states, the policy network gradually learns the optimal collaborative scheduling strategy for the charging infrastructure and the edge computing gateway energy storage system in various complex situations. As the training progresses, the selected actions will increasingly tend to maximize the overall benefits of the system.

[0048] Step S4: Execute the selected collaborative scheduling action, obtain the immediate rewards including the operating cost of the charging infrastructure, the charging and discharging benefits of the edge computing gateway energy storage system, and the change in the charging waiting time of new energy vehicle users, and update the parameters in the state space;

[0049] Specifically, step S4 is the link where the collaborative scheduling actions selected in step S3 are put into practice, and the execution effects of the actions are evaluated and the status is updated. After the system determines and executes a certain scheduling action, immediate rewards are obtained from multiple dimensions. In terms of charging infrastructure, the changes in operating costs after the action execution are considered, including the energy consumption cost of charging piles, equipment loss cost, etc. If reasonable power allocation reduces energy consumption, it will be reflected in the rewards; in the edge computing gateway energy storage system, the charging and discharging benefits are concerned. Successful discharging during high electricity price periods or charging during low electricity price periods will increase the benefits and be reflected in the rewards; at the same time, the change in the charging waiting time of new energy vehicle users is also an important evaluation index. Shortening the waiting time will obtain positive rewards, and vice versa will result in negative rewards. These rewards combined comprehensively reflect the impact of this scheduling action on the entire system. While obtaining immediate rewards, the system will update the parameters in the state space according to the actual situation after the action execution. For example, after executing the action of the energy storage node discharging to support the charging of the charging pile, the stored power of the energy storage node will decrease accordingly, and the charging power and the number of charging demands of the charging pile may also change due to some vehicles completing charging. The updated state space will be used as the basis for the next scheduling decision, enabling the system to continuously optimize the scheduling strategy based on the latest system state.

[0050] Step S5: Store the state space, the selected scheduling actions, the immediate rewards, and the updated state space in the experience replay pool, and randomly extract sample data from the experience replay pool based on the deep reinforcement learning algorithm to construct a training data set;

[0051] Specifically, the main task of step S5 is to store the key information in each scheduling process and construct a data set for training to continuously optimize the scheduling strategy. After each execution of the scheduling action and update of the state, the system will store a series of information such as the current state space, the selected scheduling actions, the obtained immediate rewards, and the updated state space in the experience replay pool. The experience replay pool is like an "experience warehouse", accumulating the scheduling experiences of the system at different times. Based on the deep reinforcement learning algorithm, the system will randomly extract sample data from the experience replay pool to construct a training data set. This random extraction method breaks the temporal correlation of the data and avoids the overfitting problem caused by continuously learning similar data. The extracted sample data includes various different system states, scheduling actions, and their corresponding reward feedbacks. These diverse data provide rich materials for subsequent model training. By continuously storing new scheduling experiences in the experience replay pool and extracting samples to construct a training data set, the system can make full use of historical experiences, learn more effective scheduling strategies from a large amount of practical data, and gradually improve the accuracy and efficiency of the collaborative scheduling of new energy vehicle charging infrastructure and edge computing gateway energy storage systems.

[0052] Step S6: Use the training dataset to iteratively train and optimize the policy network and the value network in the deep reinforcement learning algorithm, return to step S3 and repeat the execution until the preset training termination condition is met, and obtain a stable collaborative scheduling strategy for the new energy vehicle charging infrastructure and the edge computing gateway energy storage.

[0053] Specifically, step S6 is the optimization core of the entire collaborative scheduling method. By using the training dataset constructed in step S5, the policy network and the value network in the deep reinforcement learning algorithm are iteratively trained and optimized. The value network is used to evaluate the future expected benefits of the system after taking different actions in a given state, and it provides a reference standard for the optimization of the policy network; the policy network adjusts its own parameters according to the evaluation feedback of the value network to output a more optimal probability distribution of scheduling actions. During the training process, the system will input the samples in the training dataset into the network in sequence, calculate the error between the prediction result and the actual result, and through the backpropagation algorithm, transmit the error back to each layer of the network to update the parameters in the network and continuously reduce the error. After multiple iterative trainings, the performance of the policy network and the value network gradually improves, and the scheduling decision-making ability for the new energy vehicle charging infrastructure and the edge computing gateway energy storage system is continuously enhanced. When the preset training termination condition is met, such as the error converges to a certain range, the number of training times reaches the set value, etc., the system stops training. At this time, the obtained scheduling strategy is the stable collaborative scheduling strategy for the new energy vehicle charging infrastructure and the edge computing gateway energy storage. This strategy can continuously and efficiently coordinate the operation of the charging infrastructure and the edge computing gateway energy storage system under different system states, achieving the optimal allocation of resources and the maximization of the overall system benefits.

[0054] Preferably, for the multi-head attention mechanism model used in step S2, its calculation process satisfies the following formula:

[0055]

[0056] where represents the query matrix composed of the collaborative scheduling parameters of the charging infrastructure and the edge computing gateway energy storage in the state space, including the real-time charging power matrix of charging piles, the stored power matrix of energy storage nodes, etc.; represents the key matrix, corresponding to the query matrix ; represents the value matrix; are the weight matrices of the query, key, and value of different heads respectively; is the weight matrix of the multi-head output; is the dimension of the key matrix; is the number of heads of the multi-head attention mechanism; Concat is the concatenation operation; through the above formula, multi-head parallel feature extraction and fusion are performed on the parameters of the charging infrastructure and the edge computing gateway energy storage system to generate a set of feature vectors.

[0057] Specifically, the multi-head attention mechanism model analyzes the parameters of the charging infrastructure and the energy storage system of the edge computing gateway through multi-head parallel computing from different dimensions. For example, one "head" focuses on analyzing the dynamic relationship between the real-time charging power of the charging pile and the charging and discharging power of the energy storage node. During the charging peak, it can identify the alleviating effect of the discharging power of the energy storage node on the charging pile load; another "head" focuses on the matching degree between the number of charging demands and the stored power of the energy storage node to determine whether the existing energy storage can meet the expected charging demands. Through this multi-dimensional analysis, the model can capture the complex dependency relationships between the parameters, generate a more comprehensive set of feature vectors, provide a richer information basis for subsequent scheduling decisions, and improve the system's adaptability to complex scenarios.

[0058] Preferably, for the policy network of the deep reinforcement learning algorithm based in step S3, the calculation of the output scheduling action probability distribution satisfies the following formula:

[0059]

[0060] Wherein, represents the probability that under the policy network with parameters , the state takes the action ; is the state in the state space constructed in step S1, including the parameters of the charging infrastructure and the energy storage system of the edge computing gateway; is the scheduling action in the action space defined in step S1; is the action space; is the output scoring function of the policy network, and its parameter is optimized through subsequent training. This formula realizes the output of the probability distribution of the collaborative scheduling action for the charging infrastructure and the edge computing gateway energy storage according to the state space information.

[0061] Specifically, the policy network in step S3 outputs the scheduling action probability distribution based on the state space information. In practical applications, when the state space shows that the charging demand of the charging piles in a certain area surges and the energy storage node has sufficient power, the policy network will increase the probability of the action of "increasing the discharging support of the energy storage node to the charging piles". At the same time, by adjusting the probabilities of each action in different states, the policy network can learn the optimal scheduling strategy, such as preferentially charging the energy storage node during the low electricity price period and discharging to obtain benefits during the peak period, so as to achieve the balanced optimization of the operating cost of the charging infrastructure and the benefits of the edge computing gateway energy storage system.

[0062] Preferably, for the immediate reward obtained in step S4, its calculation satisfies the following formula:

[0063]

[0064] Wherein, is the immediate reward; is the operation cost of the charging infrastructure, including the energy consumption cost of charging piles, etc.; is the charging and discharging income of the edge computing gateway energy storage system; is the change in the charging waiting time of new energy vehicle users; α, β, and γ are the weight coefficients of the operation cost of the charging infrastructure, the charging and discharging income of the edge computing gateway energy storage system, and the change in the charging waiting time of new energy vehicle users respectively. The immediate reward after performing the scheduling action is comprehensively calculated through this formula to reflect the collaborative scheduling effect of the charging infrastructure and the edge computing gateway energy storage.

[0065] Specifically, the calculation of the immediate reward comprehensively considers various factors. Among the operation costs of the charging infrastructure, the energy consumption cost of the charging pile accounts for the main part, and the energy consumption can be reduced by reasonably allocating the power; the charging and discharging income of the edge computing gateway energy storage system is closely related to the peak-valley electricity price of the power grid. Charging during the valley period and discharging during the peak period can obtain the price difference income; the change in the charging waiting time of new energy vehicle users reflects the service quality, and reducing the waiting time can improve user satisfaction. The weight coefficients α, β, and γ can be adjusted according to actual needs. For example, when the power grid load is tight, increasing the value of α can give priority to controlling the operation cost of the charging infrastructure to achieve a comprehensive evaluation of the collaborative scheduling effect.

[0066] Preferably, when the deep reinforcement learning algorithm extracts sample data from the experience replay pool to construct a training data set in step S5, a prioritized experience replay mechanism is adopted, and the sample extraction probability satisfies the following formula:

[0067]

[0068] where, represents the extraction probability of the th sample in the experience replay pool; is the time difference error of the th sample, reflecting the importance of the sample for learning; is the total number of samples in the experience replay pool; is a parameter for controlling the sampling bias. This formula realizes the priority extraction of important experience samples in the collaborative scheduling process of the charging infrastructure and the edge computing gateway energy storage, and improves the training efficiency.

[0069] Specifically, the prioritized experience replay mechanism improves the training efficiency. In the collaborative scheduling of charging infrastructure and edge computing gateway energy storage, some key experience samples are crucial for learning the optimal strategy, such as the scheduling process when a charging pile fails in extreme weather. The temporal difference error reflects the importance of the sample. A sample with a large error indicates that the current strategy performs poorly in this state and needs to be focused on for learning. By preferentially sampling such samples, the model can improve the strategy faster, reduce the training time, and especially can adapt and optimize the scheduling strategy more quickly when dealing with complex electricity market price fluctuations and sudden changes in charging demand.

[0070] Preferably, when optimizing the training of the value network in step S6, the following loss function formula is used:

[0071]

[0072]

[0073] Where, is the loss function of the value network, are the parameters of the value network; represents the expectation; are respectively the state, action, immediate reward, and updated state of the sample extracted in step S5; is the valuation of the value network in state ; is the valuation of the target value network in state ; is the discount factor, which is used to balance the immediate reward and future rewards. By minimizing this loss function, the value network is optimized to accurately evaluate the value of the collaborative scheduling state of the charging infrastructure and edge computing gateway energy storage.

[0074] Specifically, the training optimization of the value network ensures the accurate evaluation of the state value. During peak charging demand periods, the value network needs to accurately evaluate the impact of each scheduling action in the current state on future benefits. For example, choosing to increase the discharge of energy storage nodes can relieve the current charging pressure, but may affect the discharge benefits during subsequent high-price periods. The discount factor γ balances the immediate reward and future rewards. A larger γ value makes the model pay more attention to long-term benefits and is suitable for the long-term stable operation of the power grid; a smaller γ value focuses on immediate benefits and is applicable to short-term emergency scheduling. By minimizing the loss function, the value network can accurately predict the long-term effects of different scheduling strategies and provide a reliable basis for the optimization of the policy network.

[0075] Preferably, after extracting the features of the parameters using the multi-head attention mechanism model in step S2, the following feature fusion formula is further used:

[0076]

[0077] Among them, is the fused feature vector; is the feature vector extracted from the charging infrastructure parameters through the multi-head attention mechanism; is the feature vector extracted from the energy storage system parameters of the edge computing gateway through the multi-head attention mechanism; is the fusion weight coefficient, and the feature vectors of the charging infrastructure and the energy storage system of the edge computing gateway are weighted and fused through this formula to enhance the feature representation ability.

[0078] Specifically, feature fusion enhances the representation ability. The charging infrastructure parameters and the energy storage system parameters of the edge computing gateway have different characteristics. The number of charging demands reflects the dynamics on the user side, while the stored electricity in the energy storage nodes reflects the resource status on the system side. By weighted fusing these two types of feature vectors, the respective advantages can be fully utilized. During the peak charging period, the feature vector of the charging infrastructure may show that some charging piles are overloaded, while the feature vector of the energy storage system indicates that the nearby energy storage nodes have sufficient electricity. After fusion, more accurate scheduling suggestions can be generated, such as transferring some demands of the overloaded charging piles to the energy storage for support, improving the overall performance of the system.

[0079] Preferably, before the policy network outputs the scheduling action probability distribution in step S3, a gating mechanism is introduced, and its calculation satisfies the following formula:

[0080]

[0081] Among them, is the gating signal; is the activation function; is the weight matrix and bias vector of the gating mechanism; is the feature extraction result of the policy network for the state ; is the scheduling action probability distribution output by the original policy network; is the prior probability distribution; is the scheduling action probability distribution adjusted by the gating mechanism. Through this gating mechanism, the output of the policy network is adjusted by combining prior knowledge to optimize the selection of the charging infrastructure and the coordinated scheduling actions of the energy storage system of the edge computing gateway.

[0082] Specifically, the gating mechanism optimizes the action selection. Prior knowledge plays an important role in the coordinated scheduling of the charging infrastructure and the energy storage system of the edge computing gateway. For example, according to historical data, weekends and afternoons are usually peak charging demand periods. The gating mechanism combines the prior probability distribution and the output of the original policy network, and automatically increases the action probability of "starting the energy storage node charging in advance" during the peak period, reducing manual intervention. The gating signal dynamically adjusts the fusion ratio according to the current state. In extreme cases, such as a power grid failure, it can rely entirely on prior knowledge to respond quickly, improving the robustness of the system and its ability to cope with emergencies.

[0083] Preferably, in step S4, when calculating the immediate reward , combining the equipment loss cost of the charging infrastructure and the edge computing gateway energy storage system, the immediate reward calculation formula is adjusted as follows:

[0084]

[0085] where is the equipment loss cost of the charging infrastructure and the edge computing gateway energy storage system; is the weight coefficient of the equipment loss cost. By this formula, the immediate reward after executing the scheduling action is calculated more comprehensively to optimize the collaborative scheduling strategy of the charging infrastructure and the edge computing gateway energy storage.

[0086] Specifically, the equipment loss cost is incorporated into the reward calculation. The long-term operation of the equipment of the charging infrastructure and the edge computing gateway energy storage system will cause losses, and frequent charging and discharging will shorten the equipment life. Incorporating the equipment loss cost into the immediate reward calculation can prompt the scheduling strategy to use the equipment more reasonably. For example, when the price difference between peak and valley electricity prices is small, avoid frequent charging and discharging of energy storage nodes and choose to operate in the power range with less equipment loss. The weight coefficient can be adjusted according to factors such as equipment depreciation rate to balance short-term benefits and long-term equipment maintenance costs and achieve the sustainable operation of the system.

[0087] As Figure 2 shown, a collaborative scheduling method for a new energy vehicle charging infrastructure and an edge computing gateway energy storage is realized through the following six units:

[0088] A parameter acquisition and state construction unit, which is used to perform the operations of constructing the state space and defining the action space in step S1, and transmit state information and action space information to subsequent units;

[0089] A multi-head feature processing unit, connected to the parameter acquisition and state construction unit, which is used to perform the operations of feature extraction and correlation analysis on the parameters using the multi-head attention mechanism model in step S2, and transmit the processed set of feature vectors to the policy decision-making unit;

[0090] A policy decision-making unit, connected to the multi-head feature processing unit and the action execution unit respectively, which is used to perform the operations of outputting the scheduling action probability distribution based on the policy network and selecting the scheduling action in step S3;

[0091] An action execution and reward acquisition unit, connected to the policy decision-making unit, which is used to perform the operations of executing the scheduling action, obtaining the immediate reward and updating the state space parameters in step S4, and transmit the immediate reward and the updated state space parameters to the experience storage and training unit;

[0092] The experience storage and training unit, which is respectively connected to the action execution and reward acquisition unit and the policy decision-making unit, is used to execute the operations of storing experience in the experience replay pool, extracting samples to construct a training data set in step S5, and training and optimizing the policy network and value network in step S6, and feeding back the optimized network parameters to the policy decision-making unit;

[0093] The loop control unit, which is connected to the policy decision-making unit and the experience storage and training unit, is used to judge whether the preset training termination condition is satisfied, and control each unit to execute the loop operation of steps S3 - S6 until a stable cooperative scheduling policy is obtained.

[0094] A cooperative scheduling method for a new energy vehicle charging infrastructure and an edge computing gateway energy storage proposed by the present invention solves the problems of lack of intelligence and insufficient parameter correlation analysis in traditional technologies, and realizes the optimization and upgrading of resource scheduling by means of a deep reinforcement learning algorithm and a multi-head attention mechanism model.

[0095] Improving scheduling intelligence to cope with complex scenarios: Traditional scheduling relies on manual experience and simple rules and is difficult to cope with the complex changes in the charging demands of new energy vehicles. This cooperative scheduling method endows the system with the ability of autonomous learning and decision-making through a deep reinforcement learning algorithm. It can perceive the dynamic states of the charging infrastructure and the edge computing gateway energy storage system in real time, such as the real-time charging power of charging piles, the stored power of energy storage nodes, etc. When there is a charging peak for new energy vehicles, instead of overloading the charging piles or leaving the energy storage idle as in the traditional way, the system automatically learns and optimizes the scheduling strategy, intelligently distributes the power of charging piles, reasonably arranges the charging and discharging of energy storage nodes, effectively reduces the user waiting time, significantly improves the utilization efficiency of power resources, and makes the operation of the charging infrastructure and the energy storage system more efficient and flexible.

[0096] Deeply mining parameter correlations to optimize the overall benefit: Previous technologies regarded the parameters of the charging infrastructure and the edge computing gateway energy storage system in isolation and could not maximize the overall benefit. This method uses a multi-head attention mechanism model to deeply mine the complex dependence relationships between various parameters. Whether it is the number of charging demands of the charging infrastructure or the charging and discharging power of the edge computing gateway energy storage system, this model can perform feature extraction and correlation analysis on these parameters through multi-head parallel computing to form comprehensive multi-dimensional feature information. The scheduling strategy formulated based on this information fully considers the mutual influence between various parameters, can effectively control the operating cost of the charging infrastructure, improve the charging and discharging income of the edge computing gateway energy storage system, and realize the optimization of the overall benefit of the system.

[0097] Multi-link collaborative optimization to ensure long-term stability: In addition to the above two major core breakthroughs, the collaborative scheduling method has also been optimized in multiple links. In the reward calculation link, factors such as the operating cost of charging infrastructure, energy storage revenue, user waiting time, and equipment loss are fully considered to comprehensively evaluate the effect of scheduling actions; through the experience replay pool and the prioritized experience replay mechanism, historical experience is effectively utilized to improve the training efficiency; the policy network and value network are optimized by means of cyclic training to ensure the continuous evolution of the scheduling strategy, and finally a stable and reliable collaborative scheduling scheme is obtained to ensure the long-term stable and efficient operation of the charging infrastructure and the energy storage system of the edge computing gateway.

[0098] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "set", "install", "connected", "connected", "fixed" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0099] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various equivalent changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalent scope.

Claims

1. A collaborative scheduling method for new energy vehicle charging infrastructure and edge computing gateway energy storage, characterized in that, Including the following steps: Step S1: Construct a state space including the real-time charging power of each charging pile in the new energy vehicle charging infrastructure, the number of charging demands, the stored power of each energy storage node in the edge computing gateway energy storage system, and the charging and discharging power, and define an action space for the selection probability of scheduling actions; Step S2: Use the multi-head attention mechanism model to perform feature extraction and correlation analysis on the parameters of the collaborative scheduling of the charging infrastructure and the edge computing gateway energy storage in the state space, capture the dependence relationship between different parameters through the multi-head parallel computing method, and form a set of feature vectors including multi-dimensional feature information; Step S3: Based on the policy network, use the set of feature vectors as input, output the scheduling action probability distribution in the action space, and select the collaborative scheduling action for the charging infrastructure and the edge computing gateway energy storage system according to the probability distribution; Step S4: Execute the selected collaborative scheduling action, obtain the immediate reward including the operating cost of the charging infrastructure, the charging and discharging income of the edge computing gateway energy storage system, and the change in the charging waiting time of new energy vehicle users, and update the parameters in the state space; Step S5: Store the state space, the selected scheduling action, the immediate reward, and the updated state space in the experience replay pool, randomly extract sample data from the experience replay pool, and construct a training data set; Step S6: Iteratively train and optimize the policy network and the value network in the deep reinforcement learning algorithm, return to Step S3 and repeat the execution until the preset training termination condition is met.

2. A coordinated scheduling method for a new energy vehicle charging infrastructure and an edge computing gateway energy storage according to claim 1, characterized in that, In Step S2, for the multi-head attention mechanism model used, the calculation process satisfies the following formula: , Among them, represents the query matrix composed of the collaborative scheduling parameters of the charging infrastructure and the energy storage of the edge computing gateway in the state space, including the real-time charging power matrix of the charging piles and the stored electricity matrix of the energy storage nodes; represents the key matrix, corresponding to the query matrix ; represents the value matrix; are the weight matrices of the queries, keys, and values with different heads respectively; is the weight matrix for the multi-head output; is the dimension of the key matrix; is the number of heads of the multi-head attention mechanism; Concat is the concatenation operation.

3. A coordinated scheduling method for new energy vehicle charging infrastructure and edge computing gateway energy storage according to claim 1, characterized in that, In Step S3, the calculation of the scheduling action probability distribution output by the policy network satisfies the following formula: , Among them, represents the probability that, under the policy network with parameter , the state takes the action ; is the state in the state space constructed for step S1, including various parameters of the charging infrastructure and the energy storage system of the edge computing gateway; is the scheduling action in the action space defined for step S1; is the action space; is the output scoring function of the policy network, and the parameter is optimized through subsequent training.

4. A coordinated scheduling method for new energy vehicle charging infrastructure and edge computing gateway energy storage according to claim 1, characterized in that, In Step S4, for the immediate reward obtained, the calculation satisfies the following formula: , Among them, is the instant reward; is the charging infrastructure operation cost, including the energy consumption cost of charging piles; is the charging and discharging income of the edge computing gateway energy storage system; is the change in the charging waiting time of new energy vehicle users; are the weight coefficients of the charging infrastructure operation cost, the charging and discharging income of the edge computing gateway energy storage system, and the change in the charging waiting time of new energy vehicle users, respectively.

5. A method for collaborative scheduling of new energy vehicle charging infrastructure and edge computing gateway energy storage according to claim 1, characterized in that, In Step S5, when extracting sample data from the experience replay pool to construct a training data set, a prioritized experience replay mechanism is adopted, and the sample extraction probability satisfies the following formula: , Among them, represents the sampling probability of the th sample in the experience replay pool; is the temporal difference error of the th sample, reflecting the importance of the sample for learning; is the total number of samples in the experience replay pool; is a parameter for controlling sampling bias.

6. A collaborative scheduling method for a new energy vehicle charging infrastructure and an edge computing gateway energy storage according to claim 1, characterized in that, In Step S6, when training and optimizing the value network, the following loss function formula is adopted: , , Among them, is the loss function of the value network, are the value network parameters; denotes expectation; are respectively the state, action, immediate reward, and updated state of the sample extracted in step S5; is the valuation of the value network at state ; is the valuation of the target value network at state ; is the discount factor, which is used to balance immediate rewards and future rewards.

7. A coordinated scheduling method for a new energy vehicle charging infrastructure and an edge computing gateway energy storage according to claim 1, characterized in that, In Step S3, before the policy network outputs the scheduling action probability distribution, a gating mechanism is introduced, and the calculation satisfies the following formula: , Among them, is the gating signal; is the activation function; is the weight matrix and bias vector of the gating mechanism; is the feature extraction result of the policy network for the state ; is the scheduling action probability distribution output by the original policy network; is the prior probability distribution; is the scheduling action probability distribution adjusted by the gating mechanism.

8. A coordinated scheduling method for a new energy vehicle charging infrastructure and an edge computing gateway energy storage according to claim 4, characterized in that, In step S4, when calculating the immediate reward , considering the equipment loss cost of the charging infrastructure and the edge computing gateway energy storage system, the immediate reward calculation formula is adjusted to: , Among them, is the equipment loss cost of the charging infrastructure and the energy storage system of the edge computing gateway; is the weight coefficient of the equipment loss cost.

Citation Information

Patent Citations

  • Vehicle-mounted task cooperative migration method for Internet of Vehicles resource fusion

    CN111885155A

  • New energy automobile charging station billing data transmission method based on mobile edge computing and block chain technology

    CN112115505A

  • Vehicle-road cooperation method and system based on communication-traffic-power integration

    CN118783445A

  • Charging pile intelligent scheduling system and method based on edge calculation

    CN119886656A

  • Method and apparatus for constructing dispatching model of integrated energy system, medium, and electronic device

    WO2022160705A1

Cited By

  • Direct current screen switching method and system based on reinforcement learning

    CN120511860A

  • A DC screen switching method and system based on reinforcement learning

    CN120511860B