A multi-charging station line-level collaborative scheduling method and system based on reinforcement learning and safety constraint projection
Patent Information
- Application Number
- CN202611118409.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-09-29
AI Technical Summary
然而,现有技术多聚焦于单站V2G控制或简单聚合策略,未能充分考虑多站协同下的复杂约束与全局优化目标
本发明通过将强化学习策略输出与安全约束投影相结合,既保留了智能体对复杂多站调度场景的学习能力,又保证了输出指令满足工程可执行边界;通过剩余需求迭代再分配,提高安全投影后的需求满足率;通过历史负担、参与频率和响应率的闭环更新,提升站点调用公平性和长期调度稳定性;通过数据库批次管理和后台服务部署策略,支持在线运行、离线回放、模型升级和结果追溯,具有良好的工程应用价值。
Smart Images

Figure CN122844318A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power and energy system technology, and relates to a multi-charging station line-level collaborative scheduling method and system based on reinforcement learning and security constraint projection. Background Technology
[0002] With the increasing prevalence of electric vehicles (EVs) and the development of V2G (Vehicle-to-Grid) technology, multiple charging stations connected to the same regional power distribution network can coordinate and utilize on-site online bidirectional charging and discharging equipment and transformers to feed on-demand battery power back to the power distribution network, thereby participating in grid regulation. This technology aims to improve the power distribution network's capacity to absorb distributed energy, alleviate peak load pressure, and optimize the benefits for EV users. However, existing technologies mostly focus on single-station V2G control or simple aggregation strategies, failing to fully consider the complex constraints and global optimization objectives under multi-station coordination.
[0003] Currently, existing multi-charging station load control methods suffer from problems such as inaccurate assessment of control capacity, difficulty in directly meeting engineering safety constraints with reinforcement learning actions, easy overstepping of control task allocation, difficulty in closed-loop utilization of execution feedback, and difficulty in connecting online deployment with offline training.
[0004] Therefore, there is an urgent need for a new and efficient collaborative scheduling method. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a multi-charging station line-level collaborative scheduling method and system based on reinforcement learning and safety constraint projection. The method takes the distribution line as the scheduling object and the multiple charging stations connected to the line as flexible resource aggregation units. It combines the output of reinforcement learning strategy with engineering safety constraint projection, so that the intelligent scheduling strategy can meet the requirements of station-level up and down adjustment capability, transformer safety boundary, online status, ramping constraint and scheduling command executability while retaining the learning ability.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection, the method specifically includes the following steps: S1. Data Acquisition and Scheduling Cycle Initialization of Multi-Source Operation Data of Distribution Lines: At the beginning of each scheduling cycle, the following data are acquired: distribution line adjustment demand, predicted power of the line, target power of the line, current aggregated power of each charging station, online status of each charging station, current load rate of distribution transformer of each charging station, upward and downward adjustment capacity of each charging station, historical cumulative adjustment burden of each charging station, recent participation frequency of each charging station, and line tracking error of the previous scheduling cycle, forming a distribution line dataset. S2. Reinforcement learning state space construction: Construct global state features based on the line adjustment requirements, construct station-level state features based on the state of each charging station, and concatenate the global state features with all the station-level state features to form a reinforcement learning state vector. S3. Station-level initial action generation based on reinforcement learning strategy: The reinforcement learning state vector is input into the trained reinforcement learning strategy network, the strategy network outputs the continuous action vector corresponding to each charging station, and the continuous action vector is converted into the initial adjustment power proposal of each charging station. S4. Station-level safety constraint modeling and feasible interval determination: Based on the current operating status of each charging station, adjustable resources within the station, transformer safety constraints, station-level power boundaries, and equipment online status, determine the feasible adjustment interval for each charging station within the current scheduling cycle; S5. Safety constraint projection of reinforcement learning actions: Project the initial adjustment power suggestion into the feasible adjustment range to obtain a safe and executable station-level adjustment power, and calculate the remaining deviation between the allocated adjustment amount and the line adjustment demand. S6. Iterative redistribution based on remaining capacity and historical fairness: When the remaining deviation exceeds the preset tolerance, iterative redistribution is carried out according to the remaining adjustable capacity and historical fairness weight of each charging station until the remaining deviation meets the tolerance or there are no resources to be allocated, and the final station-level regulating power is output. S7. Reinforcement learning training reward function design: During the offline training phase, a reward function is constructed with line demand tracking error, safety projection correction error, smoothness of adjacent cycle adjustment power, and fairness of station cumulative adjustment burden as the core, to guide the policy network to learn the collaborative allocation rules among multiple charging stations; S8. Execution Feedback, Closed-Loop Correction and Historical Status Update: Obtain the actual execution adjustment amount of each charging station, calculate the line adjustment tracking error for this cycle, and update the historical cumulative adjustment burden, normalized historical adjustment burden and recent participation frequency of each charging station. Use the updated status as the input for the next scheduling cycle.
[0007] Furthermore, in step S1, the power distribution line dataset is represented as follows:
[0008]
[0009]
[0010] in, Indicates power distribution lines During the scheduling period The line-side dataset; Indicates the target power or warning boundary power of the line; This indicates the original predicted power of the line; This represents the predicted power of the line after correction by the prediction correction model; This indicates the measured power of the line; Indicates the need for line adjustment; This represents the set of status characteristics of all charging stations along the line; Indicates the first Each charging station cycle Station-level status characteristics; This indicates that the charging station is online. This indicates the current aggregate charging power of the charging station; and They represent the first The charging station's capacity can be adjusted upwards and downwards; This indicates the current load rate of the station's distribution transformer; This indicates the historical cumulative adjustment burden; Indicates recent participation frequency; This indicates the adjustment tracking error of the station or line in the previous cycle; The line adjustment requirement is expressed as follows:
[0011] when When, it indicates that the current line needs to reduce the charging load; when When, it indicates that the current line allows or needs to increase the charging load; the absolute value of the scheduling demand is expressed as:
[0012] The data collection and scheduling batch initialization described above provide a unified data foundation for subsequent reinforcement learning state construction, action generation, safe projection, and instruction writing.
[0013] Furthermore, in step S2, the global state feature is represented as:
[0014] No. The station-level state characteristics of a charging station are represented as follows:
[0015] The final reinforcement learning state vector is represented as:
[0016] in, Indicates the need for line adjustment; This represents the absolute value of demand; Indicates the adjustment direction factor; This represents the power normalization reference value; This indicates the number of charging stations with adjustable direction capabilities at the current location; This indicates the total number of charging stations under the same power distribution line; This indicates the line tracking error in the previous scheduling cycle; Represents global state characteristics; Indicates the first Station-level status characteristics of a charging station; Indicates the first The actual adjustment amount executed by each charging station in the previous cycle; This indicates the burden of normalized historical adjustment; Indicates recent participation frequency; This represents the complete state vector of the input reinforcement learning policy model.
[0017] Furthermore, in step S3, the state vector obtained in step S2 is... Input a reinforcement learning policy network based on the proximal policy optimization algorithm, and output a continuous action vector corresponding to each charging station from the policy network:
[0018] in, For parameters The policy network, This is the initial action vector at the station level; the action values are normalized and mapped to convert them into initial regulation power recommendations for each station.
[0019] in, This indicates that the reinforcement learning strategy applies to the site. The generated unprojected adjustment amount reflects the reinforcement learning strategy's initial allocation preference for line adjustment tasks based on historical training experience, but this result does not necessarily meet the engineering safety boundary, so it should not be directly issued for execution.
[0020] Furthermore, in step S4, based on the current operating status of each charging station, adjustable resources within the station, transformer safety constraints, station-level power boundaries, and equipment online status, the feasible adjustment range for each charging station within the current scheduling cycle is determined, specifically including: When the line needs to reduce the charging load, the station The maximum acceptable reduction is When the line needs to increase the charging load, the station The maximum allowable upward adjustment is Station-level regulation volume The feasible interval can be represented as:
[0021] in, Indicates site Execute the command to reduce the charging load. Indicates site Execute the command to increase the charging load; If a site is offline, experiencing communication failure, has transformer out of bounds, has zero adjustable capacity, or has failed critical data, then both its up-adjustment and down-adjustment capabilities will be set to zero.
[0022] By modeling the feasible intervals described above, the reinforcement learning action space and the actual engineering safety boundary can be described in a unified way.
[0023] Furthermore, in step S5, to ensure that the output of the reinforcement learning policy meets engineering safety requirements, the unprojected actions obtained in step S3 are subjected to safety constraint projection; the initial adjustment amount for each station is truncated to obtain a safe and executable station-level adjustment amount.
[0024] in, This is the station-level adjustment amount after safety constraint projection; this step explicitly decouples the reinforcement learning policy preference from the engineering constraints: the policy network is responsible for giving the adjustment tendency, and the safety projection layer is responsible for ensuring that the adjustment result does not exceed the station capacity boundary, transformer boundary, and power adjustment limits; After safety projection, calculate the remaining deviation between the currently allocated adjustment amount and the line adjustment demand:
[0025] like If the requirement does not exceed the set tolerance, the instruction generation process can proceed directly; otherwise, the remaining demand will be reallocated.
[0026] Furthermore, in step S6, when there is still remaining demand after the adjustment results of the safe projection, the system iteratively redistributes the demand based on the remaining adjustable capacity of each site, historical adjustment burden, and response reliability; for sites that still have remaining capacity, redistribution weights are constructed:
[0027] in, For the site Remaining adjustable capacity in the current direction, This refers to the site's historical response rate. To adjust the burden of history, For participation frequency, The participation frequency penalty coefficient; Distribute the remaining demand to available sites based on their weights:
[0028] in, The set of sites that still have remaining adjustable capacity; the redistribution results still need to be truncated by the feasible range at the site level and the remaining demand is updated; this process is executed iteratively until the remaining demand is less than the tolerance, the set of available sites is empty, or the maximum number of iterations is reached; Through this mechanism, the system can perform engineering-executable compensation allocation for incomplete adjustment quantities after the initial actions of reinforcement learning are safely projected, thereby improving the line adjustment demand satisfaction rate and avoiding long-term centralized calls to a few stations.
[0029] Furthermore, in step S7, during the offline training phase, a reward function is constructed with demand tracking, constraint safety, allocation smoothness, site fairness, and action executability as its core principles to guide the policy network in learning the collaborative allocation patterns among multiple charging stations; the single-cycle reward function is expressed as:
[0030] in, To constrain violations and penalties, , , , The reward weighting coefficients are as follows: the first term is used to reduce line adjustment residuals; the second term is used to penalize violations of safety constraints; the third term is used to suppress sudden changes in instructions between adjacent scheduling cycles; and the fourth term is used to reduce the probability that sites with high historical loads will continue to bear too many adjustment tasks.
[0031] Through the aforementioned reward function, the reinforcement learning model can learn scheduling strategies that "meet line adjustment requirements, comply with capacity boundaries, balance site fairness, and maintain instruction smoothness" during historical simulations and offline training.
[0032] Furthermore, in step S8, the specific steps include: after the station-side control system executes the scheduling command, it writes the actual execution power, execution status, failure reason and feedback time into the database feedback table; the active control service calculates the actual completed amount, line residual and station historical status based on the feedback data. Site The actual amount completed is expressed as follows:
[0033] in, For the site The collection of charging stations within the area For charging piles The actual power change; the line execution residual is expressed as:
[0034] When the line residual exceeds the set threshold and there is still a usable closed-loop time window in the current scheduling cycle, the system rereads the remaining capacity of each station and... The system performs a limited number of rescheduling operations as new adjustment demands; after the loop is closed, the system updates the historical adjustment load, participation frequency, and response rate of each site.
[0035]
[0036] in, and These are the attenuation coefficients for historical burden and response rate, respectively. To prevent extremely small positive numbers from being divided by zero, the system forms a closed-loop scheduling process of "state acquisition - reinforcement learning decision-making - security projection - instruction execution - feedback correction - historical update" by implementing feedback and updating historical status. This improves the reliability and sustainable operation capability of active load balancing scheduling for multi-charging station power distribution lines.
[0037] The present invention also provides a multi-charging station line-level collaborative scheduling system based on reinforcement learning and security constraint projection, which adopts the method described above.
[0038] The beneficial effects of this invention are as follows: This invention combines reinforcement learning strategy output with security constraint projection, preserving the agent's learning ability in complex multi-station scheduling scenarios while ensuring that output instructions meet engineering executable boundaries. It improves the demand satisfaction rate after security projection through iterative redistribution of remaining demands. Closed-loop updates of historical burden, participation frequency, and response rate enhance the fairness of station calls and long-term scheduling stability. Database batch management and backend service deployment strategies support online operation, offline playback, model upgrades, and result traceability, demonstrating significant engineering application value.
[0039] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating the overall process of the method of the present invention. Figure 2 This is a schematic diagram of the active load balancing and dispatching system for power distribution lines of multiple charging stations. Detailed Implementation
[0041] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0042] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0043] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0044] This application proposes a multi-charging station line-level collaborative scheduling method and system based on reinforcement learning and security constraint projection. The aim is to leverage the charging behavior changes of numerous electric vehicles (EVs) across multiple charging stations to influence the state of the distribution network, thereby improving the power grid. When EVs connect to the grid, the charging behavior and power changes of numerous EVs are used to influence the grid state through transformers within the charging stations. The adjustable power of numerous charging stations is aggregated, serving as energy storage capacity for an adjustable grid. This capacity is flexibly utilized, and the charging stations monitor grid state changes in real time, issuing dispatch commands promptly based on the grid status to improve the power grid.
[0045] Figure 1 The above is a flowchart of the overall process of the present invention. The multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection provided by the present invention specifically includes: S1. Data Acquisition and Scheduling Cycle Initialization for Multi-Source Operation of Power Distribution Lines: At the beginning of each scheduling cycle, the proactive scheduling service determines the current scheduling time. The system retrieves data such as line number, unified data cutoff time, and scheduling batch number from the database or offline playback data source, including predicted load of distribution lines, target power or warning power of lines, operating status of each charging station, transformer load status, station-level historical scheduling status, and adjustable capacity data of each charging station in the current direction.
[0046] The multi-source operational data includes at least: power distribution line correction and predicted load. Line target power , site Current power Site upgrade capabilities Site downgrade capability Site response rate Historical adjustment burden Frequency of participation and the previous cycle adjustment instructions .
[0047] The line adjustment demand can be expressed as:
[0048] when When, it indicates that the current line needs to reduce the charging load; when When, it indicates that the current line allows or needs to increase the charging load; the absolute value of the scheduling demand is expressed as:
[0049] The data collection and scheduling batch initialization described above provide a unified data foundation for subsequent reinforcement learning state construction, action generation, safe projection, and instruction writing.
[0050] S2. Reinforcement Learning State Space Construction: Based on the current scheduling cycle's line demand, the adjustable capacity of each charging station, historical scheduling status, and execution reliability indicators, a state vector is constructed for the reinforcement learning agent. This state vector characterizes the current line adjustment pressure, station resource availability, station adjustment history, and execution reliability, enabling the agent to generate initial adjustment actions with global coordination characteristics in multi-station collaborative scenarios.
[0051] In a preferred embodiment, the state vector is constructed as follows:
[0052] in, As a power normalization reference, This refers to the number of charging stations participating in the dispatch. For the site Available capabilities in the current adjustment direction:
[0053] Normalization can reduce the impact of different power levels on the stability of reinforcement learning training and enhance the generalization ability under different lines and site sizes.
[0054] S3. Station-level initial action generation based on PPO reinforcement learning strategy: The state vector obtained in step S2 Input a reinforcement learning policy network based on the proximal policy optimization algorithm, and output a continuous action vector corresponding to each charging station from the policy network:
[0055] in, For parameters The policy network, This is the initial action vector at the station level. The action values, after normalization and mapping, are converted into initial regulation power recommendations for each station.
[0056] in, This indicates that the reinforcement learning strategy applies to the site. The generated unprojected adjustment amount. This adjustment amount reflects the reinforcement learning policy's initial allocation preference for line adjustment tasks based on historical training experience, but the result does not necessarily meet the engineering safety boundary, so it should not be directly issued for execution.
[0057] S4. Station-level safety constraint modeling and feasible interval determination: Based on the current operating status of each charging station, the adjustable resources within the station, the transformer safety constraints, the station-level power boundary, and the online status of the equipment, the feasible adjustment range for each charging station within the current scheduling cycle is determined.
[0058] When the line needs to reduce the charging load, the station The maximum acceptable reduction is When the line needs to increase the charging load, the station The maximum allowable upward adjustment is Station-level regulation volume The feasible interval can be represented as:
[0059] in, Indicates site Execute the command to reduce the charging load. Indicates site Execute the command to increase the charging load.
[0060] If a site is offline, experiencing communication failure, has transformer out of bounds, has zero adjustable capacity, or has failed critical data, then both its up-adjustment and down-adjustment capabilities will be set to zero.
[0061] By modeling the feasible intervals described above, the reinforcement learning action space and the actual engineering safety boundary can be described in a unified way.
[0062] S5. Safety constraint projection of reinforcement learning actions: To ensure that the reinforcement learning policy output meets engineering safety requirements, safety-constrained projections are applied to the unprojected actions obtained in step S3. Interval truncation is performed on the initial adjustment amount for each station to obtain safe and executable station-level adjustment amounts:
[0063] in, This represents the station-level adjustment after safety constraint projection. This step explicitly decouples the reinforcement learning policy preferences from the engineering constraints: the policy network is responsible for providing the adjustment tendency, while the safety projection layer is responsible for ensuring that the adjustment result does not exceed the station capacity boundary, transformer boundary, and power adjustment limits.
[0064] After safety projection, calculate the remaining deviation between the currently allocated adjustment amount and the line adjustment demand:
[0065] like If the requirement does not exceed the set tolerance, the instruction generation process can proceed directly; otherwise, the remaining demand will be reallocated.
[0066] S6. Iterative redistribution based on residual capacity and historical fairness: When there is still residual demand after the adjustment results are projected safely, the system iteratively redistributes the demand based on the remaining adjustable capacity of each site, historical adjustment burden, and response reliability. For sites that still have residual capacity, redistribution weights are constructed as follows:
[0067] in, For the site Remaining adjustable capacity in the current direction, This refers to the site's historical response rate. To adjust the burden of history, For participation frequency, This is the participation frequency penalty coefficient.
[0068] Distribute the remaining demand to available sites based on their weights:
[0069] in, This represents the set of sites that still have remaining adjustable capacity. The reallocation results still need to be truncated according to the feasible range at the site level, and the remaining demand needs to be updated. This process is executed iteratively until the remaining demand is less than the tolerance, the set of available sites is empty, or the maximum number of iterations is reached.
[0070] Through this mechanism, the system can perform engineering-executable compensation allocation for incomplete adjustment quantities after the initial actions of reinforcement learning are safely projected, thereby improving the line adjustment demand satisfaction rate and avoiding long-term centralized calls to a few stations.
[0071] S7. Reinforcement Learning Training Reward Function Design: During the offline training phase, a reward function is constructed with demand tracking, constraint safety, allocation smoothness, site fairness, and action executability as its core, to guide the policy network to learn the collaborative allocation rules among multiple charging stations.
[0072] In a preferred embodiment, the single-cycle reward function is expressed as:
[0073] in, To constrain violations and penalties, , , , These are the reward weighting coefficients. The first term is used to reduce line regulation residuals, the second term is used to penalize violations of safety constraints, the third term is used to suppress sudden changes in instructions between adjacent scheduling cycles, and the fourth term is used to reduce the probability that historically overloaded stations will continue to bear too many regulation tasks.
[0074] Through the aforementioned reward function, the reinforcement learning model can learn scheduling strategies that "meet line adjustment requirements, comply with capacity boundaries, balance site fairness, and maintain instruction smoothness" during historical simulations and offline training.
[0075] S8. Execution feedback, closed-loop correction, and historical state update: After executing the scheduling command, the station-side control system writes the actual executed power, execution status, failure reason, and feedback time into the database feedback table. The proactive control service calculates the actual completed output, line residual, and station historical status based on the feedback data.
[0076] Site The actual amount completed is expressed as follows:
[0077] in, For the site The collection of charging stations within the area For charging piles The actual power change. The line execution residual is expressed as:
[0078] When the line residual exceeds the set threshold and there is still a usable closed-loop time window in the current scheduling cycle, the system rereads the remaining capacity of each station and... A limited number of reschedules are performed as new adjustment demands. After the loop is closed, the system updates the site's historical adjustment load, participation frequency, and response rate.
[0079]
[0080] in, and These are the attenuation coefficients for historical burden and response rate, respectively. To prevent extremely small positive numbers from being divided by zero, the system forms a closed-loop scheduling process of "state acquisition - reinforcement learning decision-making - security projection - instruction execution - feedback correction - historical update" by implementing feedback and updating historical states. This improves the reliability and sustainable operation capability of active load balancing scheduling for multi-charging station power distribution lines.
[0081] S9. Scheduling instruction generation, database transaction writing, and backend service deployment: After completing the safety constraint projection and remaining demand reallocation, the system generates scheduling instructions based on the final station-level adjustment results. These scheduling instructions include at least the scheduling batch number, line number, scheduling time, station number, adjustment direction, station-level target adjustment amount, pre-projection actions, post-projection actions, station-level upward adjustment capacity, station-level downward adjustment capacity, remaining demand, model version, and algorithm version.
[0082] In terms of software deployment, the system adopts a backend service architecture that decouples training, inference, and scheduling. The reinforcement learning training service is used to train the PPO policy model offline and output the model file; the active control service is triggered every 15 minutes to read the latest running status and model version from the database and perform a single-cycle scheduling; the result storage module writes to the scheduling batch table, station-level instruction table, stub-level instruction table, and running indicator table in a transactional manner.
[0083] For the same line and the same scheduling time, the system sets a unique service key:
[0084] Unique key constraints, database row locks, or task locks are used to ensure that the same scheduling cycle is not executed repeatedly by multiple processes. If an exception occurs during the write process, the current transaction is rolled back, and the batch status is marked as failed to avoid a partial batch state where some instructions have been written but others are missing.
[0085] This deployment strategy decouples reinforcement learning model training, online scheduling inference, and scheduling result storage, facilitating model updates, anomaly localization, offline playback, and engineering maintenance.
[0086] Figure 2 This is a schematic diagram of a multi-charging-station distribution line load active balancing dispatching system. The program described in this invention can run on a distribution network dispatching center server, a charging station edge computing gateway, an industrial controller, or a cloud-based EMS energy management platform. It supports real-time data acquisition, reinforcement learning model inference, security constraint projection calculation, database transaction writing, and multi-threaded parallel scheduling. Through the execution of this program, intelligent collaborative scheduling of flexible charging loads from multiple charging stations under the same distribution line can be achieved, providing safe, stable, and interpretable control support for active load balancing of the distribution network.
[0087] This program aims to achieve coordinated control between the charging load of multiple charging stations and the power distribution lines. Without changing the original charging structure of vehicles and charging piles, and without relying on vehicles to discharge back into the grid, the system constructs a unified load active balancing scheduling platform to acquire data in real time, such as line regulation demand, charging station operating status, adjustable capacity of charging piles, transformer load rate, user charging status, and historical regulation burden. It then calls the trained reinforcement learning strategy model to generate initial regulation power suggestions at the station level.
[0088] To ensure the engineering feasibility of the scheduling results, the program further introduces a safety constraint projection mechanism after the reinforcement learning strategy output. Based on the charging station's up-and-down adjustment capabilities, online status, transformer safety boundaries, single-cycle ramp limits, and historical fairness constraints, the initial adjustment power is truncated and the remaining demand is iteratively redistributed, ultimately generating a station-level scheduling instruction that meets the safety constraints. This instruction can be written to the database or sent to the charging station control system, where it is further decomposed at the station level for execution by specific charging piles.
[0089] During system operation, information such as line tracking error, cumulative station regulation load, previous cycle regulation power, and response status can be continuously updated based on execution feedback. This information is then used as part of the reinforcement learning state for the next scheduling cycle, achieving closed-loop operation of "state acquisition—strategy reasoning—safety projection—command output—feedback update". Its core objective is to achieve orderly dispatch of flexible loads from multiple charging stations, fair allocation of regulation tasks, and active balancing of distribution line loads, while ensuring the safe operation of the distribution network, the boundaries of charging station equipment, and the normal charging needs of users.
[0090] Example 1: A line arrival scheduling method combining reinforcement learning and security constraint projection.
[0091] In this embodiment, the system performs line-to-station power allocation for multiple charging stations under the same power distribution line. At the start of each scheduling cycle, the system acquires the line adjustment requirements. And the current status of each charging station. Line adjustment needs. The power values are signed. This indicates a need to increase the charging power of the charging station. This indicates that the charging power of the charging station needs to be reduced.
[0092] The system calculates the current directional adjustable capacity for each charging station. When At that time, read or calculate the maximum charging power that the station can increase; when At that time, the maximum charging power that the station can reduce is read or calculated. If the charging station is offline, or the transformer load rate has reached the corresponding directional safety boundary, the station's current directional adjustment capability is set to zero.
[0093] The system encodes the line demand and the status of each station into reinforcement learning state vectors and inputs them into the PPO policy network. The PPO policy network outputs a continuous action vector of length equal to the number of charging stations, with each action component in a specific state. The system converts motion components into... Adjust the ratio and multiply it by the station-level adjustable capacity in the current direction to obtain the initial adjustment power recommendation.
[0094] To prevent the initial reinforcement learning proposals from exceeding the safety boundary, the system inputs the initial adjustment power proposals into the safety constraint projection layer. This layer performs station-by-station truncation based on the feasible interval and calculates the remaining adjustment demand. If the remaining demand still exists, iterative redistribution is performed according to the remaining capacity and fairness weights until the remaining demand meets the tolerance or there is no remaining adjustable capacity.
[0095] Ultimately, the system outputs the station-level regulating power that satisfies the constraints. If the total adjustable capacity is insufficient, it outputs the maximum feasible result within the current safety boundary and records the unmet demand, rather than forcibly exceeding the safety constraints of the charging station or transformer.
[0096] Example 2: Construction of a scheduling system.
[0097] This embodiment provides a multi-charging station power distribution line load active balancing dispatching system, including a data acquisition module, a station-level capability assessment module, a reinforcement learning decision-making module, a security constraint projection module, a dispatching command output module, and a state update module. This embodiment provides a multi-charging station power distribution line load active balancing dispatching system to implement the aforementioned dispatching method based on reinforcement learning and security constraint projection. The system includes a data acquisition module, a station-level capability assessment module, a reinforcement learning decision-making module, a security constraint projection module, a dispatching command output module, and a state update module. These modules work together to form a closed-loop dispatching process of "data acquisition—capability assessment—strategy decision—security correction—command output—feedback update".
[0098] The data acquisition module is used to acquire data such as distribution line regulation demand, line predicted correction power, line target power, online status of each charging station, power of charging piles in the station, SOC, charging status, transformer load rate, historical regulation burden and execution error of the previous cycle, and to verify the integrity and timeliness of the data.
[0099] The station-level capability assessment module is used to calculate the upward and downward adjustment capabilities of each charging station. Based on the charging pile's online status, charging status, current power, power upper and lower limits, SOC constraints, and ramping limitations, this module calculates the adjustable margin of a single pile and, combined with user profile coefficients and transformer safety boundaries, obtains the final adjustable capability at the station level.
[0100] The reinforcement learning decision module is used to construct reinforcement learning state vectors and input information such as line adjustment demand, station-level adjustable capability, online status, historical burden and previous cycle error into the trained PPO policy model, and output the initial adjustment power suggestions for each charging station.
[0101] The safety constraint projection module is used to make engineering safety corrections to the initial regulation power recommendation. This module constructs a feasible range based on the station-level up and down regulation capacity, transformer boundaries, online status, and single-cycle ramping limits, and truncates the initial action; if there is still remaining regulation demand, it iterative redistribution is performed according to the remaining adjustable capacity and historical fairness weights to obtain the final executable station-level regulation power.
[0102] The dispatch instruction output module generates dispatch batches, station-level regulation power, target power, and safety verification results, and writes the dispatch instructions to the database or sends them to the charging station control system. The status update module updates the actual response, line tracking error, cumulative regulation load, and previous cycle regulation power based on execution feedback, and uses the updated results as input for the next dispatch cycle.
[0103] Through the above system structure, this embodiment can achieve coordinated adjustment of charging load of multiple charging stations without relying on vehicle reverse discharge. It also ensures that the reinforcement learning output results meet the station-level capability, transformer safety, online status and ramping constraints through safety constraint projection, thereby improving the safety, adaptability and interpretability of the scheduling results.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection, characterized in that, The method specifically includes the following steps: S1. Data Acquisition and Scheduling Cycle Initialization of Multi-Source Operation Data of Distribution Lines: At the beginning of each scheduling cycle, the following data are acquired: distribution line adjustment demand, predicted power of the line, target power of the line, current aggregated power of each charging station, online status of each charging station, current load rate of distribution transformer of each charging station, upward and downward adjustment capacity of each charging station, historical cumulative adjustment burden of each charging station, recent participation frequency of each charging station, and line tracking error of the previous scheduling cycle, forming a distribution line dataset. S2. Reinforcement learning state space construction: Construct global state features based on the line adjustment requirements, construct station-level state features based on the state of each charging station, and concatenate the global state features with all the station-level state features to form a reinforcement learning state vector. S3. Station-level initial action generation based on reinforcement learning strategy: The reinforcement learning state vector is input into the trained reinforcement learning strategy network, the strategy network outputs the continuous action vector corresponding to each charging station, and the continuous action vector is converted into the initial adjustment power proposal of each charging station. S4. Station-level safety constraint modeling and feasible interval determination: Based on the current operating status of each charging station, adjustable resources within the station, transformer safety constraints, station-level power boundaries, and equipment online status, determine the feasible adjustment interval for each charging station within the current scheduling cycle; S5. Safety constraint projection of reinforcement learning actions: Project the initial adjustment power suggestion into the feasible adjustment range to obtain a safe and executable station-level adjustment power, and calculate the remaining deviation between the allocated adjustment amount and the line adjustment demand. S6. Iterative redistribution based on remaining capacity and historical fairness: When the remaining deviation exceeds the preset tolerance, iterative redistribution is carried out according to the remaining adjustable capacity and historical fairness weight of each charging station until the remaining deviation meets the tolerance or there are no resources to be allocated, and the final station-level regulating power is output. S7. Reinforcement learning training reward function design: During the offline training phase, a reward function is constructed with line demand tracking error, safety projection correction error, smoothness of adjacent cycle adjustment power, and fairness of station cumulative adjustment burden as the core, to guide the policy network to learn the collaborative allocation rules among multiple charging stations; S8. Execution Feedback, Closed-Loop Correction and Historical Status Update: Obtain the actual execution adjustment amount of each charging station, calculate the line adjustment tracking error for this cycle, and update the historical cumulative adjustment burden, normalized historical adjustment burden and recent participation frequency of each charging station. Use the updated status as the input for the next scheduling cycle.
2. The multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection according to claim 1, characterized in that, In step S1, the power distribution line dataset is represented as follows: in, Indicates power distribution lines During the scheduling period The line-side dataset; Indicates the target power or warning boundary power of the line; This indicates the original predicted power of the line; This represents the predicted power of the line after correction by the prediction correction model; This indicates the measured power of the line; Indicates the need for line adjustment; This represents the set of status characteristics of all charging stations along the line; Indicates the first Each charging station cycle Station-level status characteristics; This indicates that the charging station is online. This indicates the current aggregate charging power of the charging station; and They represent the first The charging station's capacity can be adjusted upwards and downwards; This indicates the current load rate of the station's distribution transformer; This indicates the historical cumulative adjustment burden; Indicates recent participation frequency; This indicates the adjustment tracking error of the station or line in the previous cycle; The line adjustment requirement is expressed as follows: when When, it indicates that the current line needs to reduce the charging load; when When, it indicates that the current line allows or needs to increase the charging load; the absolute value of the scheduling demand is expressed as: The data collection and scheduling batch initialization described above provide a unified data foundation for subsequent reinforcement learning state construction, action generation, safe projection, and instruction writing.
3. The multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection according to claim 2, characterized in that, In step S2, the global state feature is represented as: No. The station-level state characteristics of a charging station are represented as follows: The final reinforcement learning state vector is represented as: in, Indicates the need for line adjustment; This represents the absolute value of demand; Indicates the adjustment direction factor; This represents the power normalization reference value; This indicates the number of charging stations with adjustable direction capabilities at the current location; This indicates the total number of charging stations under the same power distribution line; This indicates the line tracking error in the previous scheduling cycle; Represents global state characteristics; Indicates the first Station-level status characteristics of a charging station; Indicates the first The actual adjustment amount executed by each charging station in the previous cycle; This indicates the burden of normalized historical adjustment; Indicates recent participation frequency; This represents the complete state vector of the input reinforcement learning policy model.
4. The multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection according to claim 3, characterized in that, In step S3, the state vector obtained in step S2 is... Input a reinforcement learning policy network based on the proximal policy optimization algorithm, and output a continuous action vector corresponding to each charging station from the policy network: in, For parameters The policy network, This is the initial action vector at the station level; the action values are normalized and mapped to convert them into initial regulation power recommendations for each station. in, This indicates that the reinforcement learning strategy applies to the site. The generated unprojected adjustment amount reflects the reinforcement learning strategy's initial allocation preference for line adjustment tasks based on historical training experience, but this result does not necessarily meet the engineering safety boundary, so it should not be directly issued for execution.
5. The multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection according to claim 4, characterized in that, In step S4, based on the current operating status of each charging station, adjustable resources within the station, transformer safety constraints, station-level power boundaries, and equipment online status, the feasible adjustment range for each charging station within the current scheduling cycle is determined, specifically including: When the line needs to reduce the charging load, the station The maximum acceptable reduction is When the line needs to increase the charging load, the station The maximum allowable upward adjustment is Station-level regulation volume The feasible interval can be represented as: in, Indicates site Execute the command to reduce the charging load. Indicates site Execute the command to increase the charging load; If a site is offline, experiencing communication failure, has transformer out of bounds, has zero adjustable capacity, or has failed critical data, then both its up-adjustment and down-adjustment capabilities will be set to zero. By modeling the feasible intervals described above, the reinforcement learning action space and the actual engineering safety boundary can be described in a unified way.
6. The multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection according to claim 5, characterized in that, In step S5, to ensure that the reinforcement learning policy output meets engineering safety requirements, the unprojected actions obtained in step S3 are subjected to safety constraint projection; the initial adjustment amount for each station is truncated to obtain a safe and executable station-level adjustment amount. in, This is the station-level adjustment amount after safety constraint projection; this step explicitly decouples the reinforcement learning policy preference from the engineering constraints: the policy network is responsible for giving the adjustment tendency, and the safety projection layer is responsible for ensuring that the adjustment result does not exceed the station capacity boundary, transformer boundary, and power adjustment limits; After safety projection, calculate the remaining deviation between the currently allocated adjustment amount and the line adjustment demand: like If the requirement does not exceed the set tolerance, the instruction generation process can proceed directly; otherwise, the remaining demand will be reallocated.
7. The multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection according to claim 6, characterized in that, In step S6, when there is still remaining demand after the adjustment results of the safe projection, the system iteratively redistributes the demand based on the remaining adjustable capacity of each site, historical adjustment burden, and response reliability; for sites that still have remaining capacity, redistribution weights are constructed: in, For the site Remaining adjustable capacity in the current direction, This refers to the site's historical response rate. To adjust the burden of history, For participation frequency, The participation frequency penalty coefficient; Distribute the remaining demand to available sites based on their weights: in, The set of sites that still have remaining adjustable capacity is selected; the redistribution results still need to be truncated at the site level and the remaining demand is updated; this process is executed iteratively until the remaining demand is less than the tolerance, the set of available sites is empty, or the maximum number of iterations is reached.
8. The multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection according to claim 7, characterized in that, In step S7, during the offline training phase, a reward function is constructed with demand tracking, constraint safety, allocation smoothness, site fairness, and action executability as its core principles to guide the policy network in learning the collaborative allocation patterns among multiple charging stations; the single-cycle reward function is expressed as: in, To constrain violations and penalties, , , , The reward weighting coefficients are as follows: the first term is used to reduce line adjustment residuals; the second term is used to penalize violations of safety constraints; the third term is used to suppress sudden changes in instructions between adjacent scheduling cycles; and the fourth term is used to reduce the probability that sites with high historical loads will continue to bear too many adjustment tasks.
9. The multi-charging station line-level cooperative scheduling method based on reinforcement learning and security constraint projection according to claim 8, characterized in that, Step S8 specifically includes: after the station-side control system executes the scheduling command, it writes the actual execution power, execution status, failure reason and feedback time into the database feedback table; the active control service calculates the actual completed amount, line residual and station historical status based on the feedback data; Site The actual amount completed is expressed as follows: in, For the site The collection of charging stations within the area For charging piles The actual power change; the line execution residual is expressed as: When the line residual exceeds the set threshold and there is still a usable closed-loop time window in the current scheduling cycle, the system rereads the remaining capacity of each station and... The system performs a limited number of rescheduling operations as new adjustment demands; after the loop is closed, the system updates the historical adjustment load, participation frequency, and response rate of each site. in, and These are the attenuation coefficients for historical burden and response rate, respectively. To prevent extremely small positive numbers from being divided by zero, the system forms a closed-loop scheduling process of "state acquisition - reinforcement learning decision-making - security projection - instruction execution - feedback correction - historical update" by implementing feedback and updating historical status. This improves the reliability and sustainable operation capability of active load balancing scheduling for multi-charging station power distribution lines.
10. A multi-charging station line-level collaborative scheduling system based on reinforcement learning and security constraint projection, characterized in that, The system employs the method as described in any one of claims 1 to 9.