Multi-energy collaborative control and management system for automobile energy storage charging pile

By using real-time grid data perception and reinforcement learning model optimization strategies, multi-energy collaborative control of vehicle energy storage charging piles was achieved, solving the problems of low energy utilization efficiency and unreasonable load management in traditional charging piles, and improving the operating efficiency and safety of the power grid and users.

CN121822210BActive Publication Date: 2026-05-15HANGZHOU SUNWELL TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU SUNWELL TECH
Filing Date
2026-03-10
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional car energy storage charging piles are difficult to integrate multiple energy sources efficiently, resulting in low utilization efficiency of clean energy, unstable grid load, and unreasonable load management, leading to uneven distribution of charging resources and affecting the safety of residents' electricity use.

Method used

The grid sensing module collects grid status data in real time, combines the status of energy storage units and user needs, uses a pre-trained reinforcement learning model to calculate multi-objective optimization strategies, generates a charging and discharging strategy instruction set, and achieves precise control and evaluation calibration through the execution control module. It also adaptively adjusts the configuration parameters of the decision engine to achieve multi-energy collaborative control.

Benefits of technology

It has improved the frequency stability and operational efficiency of the power grid, reduced charging costs for users, optimized the utilization of power grid peak-shaving resources, ensured stable power grid load, and improved the overall operational efficiency of the power grid and users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121822210B_ABST
    Figure CN121822210B_ABST
Patent Text Reader

Abstract

The application provides a multi-energy collaborative control and management system for automobile energy storage charging piles, relates to the technical field of automobile charging, and comprises: a power grid sensing module, which is used for communicating with a power grid dispatching device through a charging pile side sensor, collecting power grid frequency, voltage and electricity price signals in real time, and constructing a power grid real-time state data set; a strategy decision module, which is used for performing multi-objective real-time optimization strategy calculation by using a pre-trained reinforcement learning model based on the power grid real-time state data set, in combination with the current state of charge of an energy storage unit, predicted user charging demand and time-of-use electricity price information, and obtaining a final charging and discharging strategy instruction set; and an execution control module, which is used for controlling the energy storage unit to execute corresponding instructions according to the charging and discharging strategy instruction set and synchronously collecting instruction execution characteristic data. The application realizes efficient collaborative scheduling of automobile energy storage charging piles and various energies, improves energy utilization efficiency, and guarantees power supply stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle charging technology, and in particular to a multi-energy collaborative control and management system for vehicle energy storage charging piles. Background Technology

[0002] In residential communities, traditional car energy storage charging pile technology has some limitations. In terms of energy coordination, it is difficult to efficiently integrate multiple energy sources. When the photovoltaic power is sufficient at noon, due to the lack of intelligent linkage with the photovoltaic system, traditional charging piles cannot prioritize the absorption of clean energy. Some photovoltaic power can only be consumed by grid connection. When the grid load is low, it may even be curtailed, resulting in poor utilization efficiency of clean energy. During the evening peak electricity consumption period, the concentrated charging of electric vehicles exacerbates the pressure on the grid. Traditional charging piles cannot call on energy storage to share the power supply, which may lead to short-term overload of the distribution lines and affect residents' electricity use.

[0003] In addition, the load management capability is somewhat insufficient. When multiple vehicles are charging at the same time, most traditional charging piles adopt fixed power allocation, which cannot be dynamically adjusted according to the vehicle battery status, the urgency of charging and the grid load. For example, if five vehicles are charging at the same time in a community, three vehicles have only 10% range left and urgently need to be charged, while two vehicles have 50% range left and do not need to be charged. However, the power is evenly allocated, resulting in slow charging for vehicles in urgent need and resources being occupied by vehicles in non-urgent need, which is an unreasonable allocation. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a multi-energy collaborative control and management system for vehicle energy storage charging piles, which realizes stable regulation of grid load and reduces charging costs by sensing the grid status in real time, dynamically allocating multiple energy sources and charging and discharging strategies.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] The first aspect is the multi-energy collaborative control and management system for vehicle energy storage charging piles, including:

[0007] The power grid sensing module is used to communicate with the power grid dispatching device through sensors on the charging pile side, collect power grid frequency, voltage and electricity price signals in real time, and build a real-time power grid status dataset.

[0008] The strategy decision module is used to perform multi-objective real-time optimization strategy calculation based on the real-time power grid status dataset, combined with the current state of charge of the energy storage unit, the predicted user charging demand and time-of-use electricity price information, and using a pre-trained reinforcement learning model to obtain the final charging and discharging strategy instruction set.

[0009] The execution control module is used to control the energy storage unit to execute corresponding instructions according to the charging and discharging strategy instruction set, and to synchronously collect instruction execution characteristic data;

[0010] The benefit assessment module is used to calculate the user's charging cost savings, grid frequency stability contribution, and peak-shaving resource savings based on instruction execution characteristic data, and generate a set of basic assessment indicators.

[0011] The coupling analysis module is used to extract key feature parameters from the dynamic changes of the basic evaluation index set, and to analyze the nonlinear coupling relationship by establishing a correlation model between the parameters, thereby generating dynamic compensation coefficients.

[0012] The evaluation and calibration module is used to calibrate the basic evaluation index set in real time using dynamic compensation coefficients to obtain the final evaluation index set.

[0013] The adaptive module is used to adaptively adjust the configuration parameters of the decision engine based on the final evaluation index set, the updated grid status data and user behavior characteristics, and realize the coordinated control of multiple energy sources in energy storage.

[0014] In a second aspect, a computing device includes:

[0015] One or more processors;

[0016] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0017] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0018] The above-described solution of the present invention has at least the following beneficial effects:

[0019] By combining time-of-use electricity pricing information and user charging demand forecasts with strategic decision-making, a final charging and discharging strategy is formulated to guide users to charge during off-peak electricity periods, reducing their charging expenses. Simultaneously, precise control of the charging and discharging process reduces unnecessary energy loss, saving costs for users. For the power grid, its stability and operational efficiency are improved. Grid sensing collects real-time status data such as grid frequency and voltage, and the charging and discharging strategy formulated based on this data can smooth grid frequency fluctuations and increase the grid's contribution to frequency stability. Furthermore, through coordinated control of energy storage units, it can effectively address grid peak-shaving demands, reduce the investment of peak-shaving resources, save peak-shaving costs, and improve the overall operational efficiency of the grid. In terms of its own performance, it possesses strong adaptive capabilities and a precise evaluation and calibration mechanism. Coupled analysis extracts key feature parameters and analyzes their coupling relationships to generate dynamic compensation coefficients. Evaluation and calibration use these coefficients to perform real-time calibration of the basic evaluation index set, making the evaluation results more accurate. Based on the final evaluation index set, the configuration parameters of the decision engine are continuously adjusted to ensure continuous adaptation to changes in grid status and user behavior, achieving efficient coordinated control of multiple energy sources. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the multi-energy collaborative control and management system for vehicle energy storage charging piles provided in an embodiment of the present invention.

[0021] Figure 2 This is a schematic diagram provided by an embodiment of the present invention, which controls the energy storage unit to execute corresponding instructions according to the charging and discharging strategy instruction set, and synchronously collects instruction execution feature data. Detailed Implementation

[0022] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0023] like Figure 1 As shown, embodiments of the present invention propose a multi-energy collaborative control and management system for vehicle energy storage charging piles, including:

[0024] The power grid sensing module is used to communicate with the power grid dispatching device through sensors on the charging pile side, collect power grid frequency, voltage and electricity price signals in real time, and build a real-time power grid status dataset.

[0025] The strategy decision module is used to perform multi-objective real-time optimization strategy calculation based on the real-time power grid status dataset, combined with the current state of charge of the energy storage unit, the predicted user charging demand and time-of-use electricity price information, and using a pre-trained reinforcement learning model to obtain the final charging and discharging strategy instruction set.

[0026] The execution control module is used to control the energy storage unit to execute corresponding instructions according to the charging and discharging strategy instruction set, and to synchronously collect instruction execution characteristic data;

[0027] The benefit assessment module is used to calculate the user's charging cost savings, grid frequency stability contribution, and peak-shaving resource savings based on instruction execution characteristic data, and generate a set of basic assessment indicators.

[0028] The coupling analysis module is used to extract key feature parameters from the dynamic changes of the basic evaluation index set, and to analyze the nonlinear coupling relationship by establishing a correlation model between the parameters, thereby generating dynamic compensation coefficients.

[0029] The evaluation and calibration module is used to calibrate the basic evaluation index set in real time using dynamic compensation coefficients to obtain the final evaluation index set.

[0030] The adaptive module is used to adaptively adjust the configuration parameters of the decision engine based on the final evaluation index set, the updated grid status data and user behavior characteristics, and realize the coordinated control of multiple energy sources in energy storage.

[0031] In this embodiment of the invention, a final charging and discharging strategy is formulated by combining time-of-use electricity price information and user charging demand prediction through strategic decision-making. This guides users to charge during off-peak electricity periods, reducing their charging expenses. Simultaneously, precise control of the charging and discharging process reduces unnecessary energy loss, saving costs for users. For the power grid, its stability and operational efficiency are improved. The grid sensing system collects real-time status data such as grid frequency and voltage, and the charging and discharging strategy formulated based on this data can smooth grid frequency fluctuations and increase the grid's contribution to frequency stability. Furthermore, through the coordinated control of energy storage units, it can effectively address grid peak-shaving demands, reduce the investment in peak-shaving resources, save peak-shaving costs, and improve the overall operational efficiency of the power grid. In terms of its own performance, it possesses strong adaptive capabilities and a precise evaluation and calibration mechanism. Coupled analysis extracts key feature parameters and analyzes their coupling relationships to generate dynamic compensation coefficients. Evaluation and calibration use these coefficients to perform real-time calibration of the basic evaluation index set, making the evaluation results more accurate. Based on the final evaluation index set, the configuration parameters of the decision engine are continuously adjusted to ensure continuous adaptation to changes in grid status and user behavior, achieving efficient coordinated control of multiple energy sources.

[0032] In a preferred embodiment of the present invention, communication is established between the charging pile-side sensor and the power grid dispatching device to collect real-time power grid frequency, voltage, and electricity price signals, thereby constructing a real-time power grid status dataset, which may include:

[0033] In this embodiment of the invention, high-precision grid frequency sensors and voltage sensors are first deployed at the incoming end of each car energy storage charging pile in the community and at key nodes of the power distribution line. The frequency sensors are AC frequency detection devices with millisecond-level response capabilities, and the voltage sensors are high-precision detection devices capable of simultaneously capturing line voltage and phase voltage. All sensors are calibrated and debugged to ensure that the measurement range matches the rated parameters of the community power distribution network and can accurately capture instantaneous fluctuation signals during grid operation. Next, a two-way communication link is established between the charging pile and the community power grid dispatching device. The communication method prioritizes stable and reliable wired communication. For scenarios where wiring is inconvenient, an industrial-grade wireless communication mode is used as a supplement. All communication links follow the common communication protocol of the power grid dispatching system to ensure the compatibility and security of data transmission. In specific connection, the sensor output port on the side of each charging pile is connected to the communication gateway of the community power grid dispatching device through communication cables or wireless mode to complete the physical connection and protocol adaptation between the sensor and the dispatching device. At the same time, a communication link monitoring mechanism is set up to investigate problems such as disconnection and signal interference in real time to ensure communication continuity.

[0034] Then, the real-time acquisition process is initiated. The frequency sensor continuously monitors the frequency changes of the AC power in the distribution network at a sampling interval of once every 10 milliseconds, recording the frequency value at each sampling moment in real time, including the stable value during normal grid operation and the fluctuation value caused by load changes. The voltage sensor synchronously collects the line voltage and phase voltage values ​​of the distribution lines at a sampling interval of once every 10 milliseconds, accurately recording the instantaneous changes and continuous status of the voltage amplitude. The electricity price signal is obtained in real time from the grid dispatching device through the communication link. The grid dispatching device receives the time-of-use electricity price plan issued by the power operation department in advance and updates it dynamically according to the adjustment of the electricity market. The charging pile side keeps synchronized with the dispatching device through the communication link, automatically checking the electricity price information once an hour. If there is an electricity price adjustment, the collected electricity price signal is updated immediately to ensure the timeliness of the electricity price data. After that, the collected raw signals are preprocessed. First, the raw sampling data of frequency and voltage are filtered to remove abnormal values ​​caused by electromagnetic interference and sensor instantaneous errors, retaining those that conform to the actual operation of the distribution network. The system first obtains valid data on the power grid status. Then, it calibrates and corrects this data by comparing the sensor-collected values ​​with the standard reference values ​​built into the power grid dispatching device to correct system errors in the sensors and ensure data accuracy. Simultaneously, frequency and voltage data are converted into fixed-format values, and electricity price signals are converted into standardized data corresponding to specific time periods. Finally, a real-time power grid status dataset is constructed. Using timestamps as the core index, the preprocessed frequency, voltage, and electricity price data are structurally integrated. Each timestamp corresponds to a complete data record, including the collection time, real-time power grid frequency value, line voltage value, phase voltage value, current electricity price standard, sensor number, corresponding charging pile number, and the node identifier of the distribution line. This forms a standardized real-time power grid status dataset. The dataset uses a real-time update mechanism, appending newly collected and preprocessed valid data to the dataset in chronological order every 10 milliseconds, while automatically removing historical data older than 24 hours to ensure the dataset always reflects the latest operating status of the power grid.

[0035] In a preferred embodiment of the present invention, based on a real-time power grid status dataset, combined with the current state of charge of the energy storage unit, predicted user charging demand, and time-of-use pricing information, a pre-trained reinforcement learning model is used to perform multi-objective real-time optimization strategy calculations to obtain the final charging and discharging strategy instruction set, which may include:

[0036] The real-time status dataset of the power grid, the current state of charge of energy storage units, the predicted user charging demand, and time-of-use electricity price information are integrated to form a data combination;

[0037] Based on data combination, a multi-objective benefit evaluation system is established in a pre-trained reinforcement learning model to simultaneously calculate the expected effects on multiple aspects such as user charging costs, grid frequency stability support, and peak-shaving resource utilization.

[0038] Based on the expected results from multiple aspects, a preliminary set of charging and discharging operation schemes is generated using the action generation unit in the reinforcement learning model.

[0039] A feasibility check was performed on the initial set of charging and discharging operation schemes, and all feasible operation schemes that meet the requirements were selected.

[0040] The future sustainable benefits of feasible operational schemes are predicted and compared, and the scheme with the best overall benefits is selected to obtain the final charging and discharging strategy instruction set, which specifically includes: the charging and discharging strategy instruction set includes the power control instructions of the charging pile, the charging and discharging power instructions of the energy storage unit, and the execution time nodes.

[0041] In this embodiment of the invention, firstly, a real-time power grid status dataset is acquired. This dataset originates from multi-point monitoring of the community's distribution network and includes real-time power grid frequency values ​​every 5 milliseconds, phase and line voltage values ​​of the power grid, current total load power, load change rate of the distribution network, and information indicating whether the power grid is in a state of alert. This data is collected in real-time through sensors on the charging pile side and the power grid dispatching device, and then preprocessed. Next, the current state of charge (SOC) of the energy storage unit is acquired. The actual stored capacity is read using the energy acquisition device built into the energy storage unit, and the rated total capacity parameter of the energy storage unit is queried. The current actual stored capacity is divided by the rated total capacity to calculate the current SOC of the energy storage unit, presented as a percentage. This calculation result directly reflects the remaining chargeable and dischargeable capacity of the energy storage unit. Then, predicted user charging demand is acquired by collecting historical charging data of electric vehicles in the community over the past 12 months, including the number of charging vehicles at different times each day, the initial remaining capacity per vehicle per charge, the charging duration, the charging date to the target capacity, and the charging date. Information such as the weather conditions for the corresponding weekday or holiday type is collected. Simultaneously, the reservation information submitted through the community charging reservation system is analyzed, including the vehicle model, estimated arrival time at the charging station, current remaining battery level, and target charging level. This historical and real-time reservation data is then aggregated and statistically analyzed to identify charging demand patterns under different time periods and scenarios. Combined with the current time, date, type, and weather conditions, the system predicts the number of vehicles needing charging in each 5-minute sub-period within the next 24 hours, the initial remaining battery level for each vehicle, the estimated charging time, and the target charging level, forming structured predicted user charging demand data. Furthermore, time-of-use electricity price information is obtained. Through a dedicated communication link with the power operator, the system acquires the peak, flat, and valley segment identifiers for each 5-minute sub-period within the next 24 hours, as well as the unit electricity price standard for each time period, clarifying the cost differences for charging at different times. Simultaneously, the electricity price information is verified hourly with the power grid dispatching device to ensure consistency between the electricity price data and that of the power operator. Finally, the data is integrated, using 5 minutes as a unified time index. The real-time frequency value, voltage value, total load power, load change rate, current state of charge percentage of energy storage units, predicted number of charging vehicles in the sub-period, expected charging amount per vehicle, expected charging start and end time, and time-of-use electricity price standard corresponding to the sub-period are matched one by one in chronological order. Each time index contains complete four types of data information, ultimately forming a structured data combination.

[0042] First, a pre-trained reinforcement learning model is constructed. The network structure design of the reinforcement learning model fully meets the needs of the multi-energy collaborative scenario in the community, adopting a deep neural network architecture. The input layer has 32 neurons, each corresponding to one of the 32 feature data in the data combination, ensuring comprehensive reception of various input information. Four hidden layers are set: the first layer contains 96 neurons for initial extraction of basic features; the second layer contains 192 neurons for in-depth mining of the correlation between features; the third layer contains 144 neurons for optimizing feature representation; and the fourth layer contains 80 neurons for integrating higher-order features. All hidden layers use the ReLU activation function because this function can effectively alleviate the gradient vanishing problem and adapt to the nonlinear transformation requirements of multi-dimensional data. The output layer has 3 neurons, corresponding to the evaluation results of three benefit dimensions: user charging cost saving, grid frequency stability, and support for peak shaving and resource utilization, ensuring that the output corresponds one-to-one with the target. The reward function of the reinforcement learning model is designed as a multi-dimensional weighted summation form, fully taking into account the needs of the community, users, and grid peak shaving. The reward consists of three parts: user cost savings, grid frequency stability, and peak-shaving resource utilization. The user cost savings reward is calculated by multiplying the current user charging cost savings by a first reward coefficient of 0.35. The grid frequency stability reward is calculated by multiplying the reduction in grid frequency deviation by a second reward coefficient of 0.35. The peak-shaving resource utilization reward is calculated by multiplying the amount of traditional peak-shaving resource substitution by a third reward coefficient of 0.3. The total reward value of the reinforcement learning model is obtained by adding the reward results of the three parts. The coefficient setting refers to the balanced priority of the three parties' needs in the daily operation of the community. During the training of the reinforcement learning model, historical data of the community over the past 18 months are collected, including historical data of grid status, historical data of energy storage unit charging and discharging, historical data of user charging, historical data of time-of-use electricity prices, and historical data of photovoltaic power output. Valid data is selected and eliminated to form 180,000 training samples, which are then divided into training set and validation set in an 8:2 ratio. The training set is used to adjust the parameters of the reinforcement learning model, and the validation set is used to monitor the generalization ability of the reinforcement learning model.

[0043] The training process employs the stochastic gradient descent algorithm, with the initial learning rate set to 0.001, decaying to 0.9 every 200 rounds. The batch size is set to 128 samples. After each iteration, the mean squared error between the reinforcement learning model's predicted results and the actual returns is calculated. Based on error backpropagation, the weights and bias parameters of each layer of the reinforcement learning model are adjusted. After 2000 rounds of iterative training, when the mean squared error on the validation set decreases to below 0.005 and remains stable for 30 consecutive rounds, the reinforcement learning model training is complete, iteration is stopped, and the trained reinforcement learning model parameters are saved to complete pre-training. Then, a multi-objective return evaluation system is established. This system corresponds one-to-one with the output layer of the reinforcement learning model and consists of three independent and complete evaluation processes. When calculating the expected charging cost effect for users, the baseline total charging cost is first calculated. For each vehicle predicted in the data set, its expected charging capacity and the time-of-use electricity price corresponding to the expected charging period are extracted. The baseline charging cost for each vehicle is obtained by multiplying its expected charging capacity by the time-of-use electricity price for the expected charging period. For example, if a vehicle is expected to charge 50 kWh during a peak period with a price of 1.2 yuan per kWh, then the baseline charging cost for that vehicle is 50 kWh multiplied by 1.2 yuan per kWh. Then, the baseline charging costs for all predicted charging vehicles are added together to obtain the baseline total charging cost for the entire community. The optimized total charging cost is calculated by analyzing the data combination using a pre-trained reinforcement learning model. The reinforcement learning model matches each vehicle with an optimized charging period that has a lower charging cost based on time-of-use (TOU) electricity prices. The TOU price during this period is lower than the base charging period price. The optimized charging cost for each vehicle is obtained by multiplying the expected charging amount for each vehicle by the TOU price during the optimized charging period. For example, if the vehicle is matched with an off-peak charging price of 0.3 yuan per kilowatt-hour, the optimized charging cost is 50 kilowatt-hours multiplied by 0.3 yuan per kilowatt-hour. The optimized charging costs for all vehicles are then added together to obtain the optimized total charging cost. Finally, the optimized total charging cost is subtracted from the base total charging cost to obtain the expected cost savings for the user.

[0044] When calculating the expected effect of grid frequency stability support, first define the standard operating range of grid frequency as 50Hz ± 0.2Hz, with a maximum allowable deviation of 0.2Hz. Then, obtain the frequency response coefficient of the community distribution network. This coefficient is obtained by statistically analyzing the correspondence between load changes and frequency changes in the community distribution network over the past 12 months. Specifically, for every 100kW increase in load, the grid frequency decreases by an average of 0.05Hz, i.e., the frequency response coefficient is 0.05Hz. For every 100kW, based on the current grid frequency and the predicted user charging load increment in the data combination, calculate the expected frequency deviation value without using an optimization strategy. Multiply the predicted user charging load increment by the frequency response coefficient to obtain the frequency change. Add this frequency change value to the current grid frequency. If the result exceeds the standard range of 50Hz ± 0.2Hz, the excess part is the expected frequency deviation value. Simultaneously, calculate the duration of this expected frequency deviation value based on the predicted charging load duration. Analyze the data combination model using a pre-trained reinforcement learning model. Based on the current grid frequency and load increment, determine the optimized charging and discharging power of the energy storage unit. This optimized charging and discharging power is inversely superimposed on the predicted user charging load increment. For example, if the predicted charging load increment is... A 200 kW frequency drop of 0.1 Hz will trigger a reinforcement learning model that instructs the energy storage unit to discharge 200 kW to offset the load increase. The optimized actual frequency deviation is calculated, and the reduction in frequency deviation is obtained by subtracting the actual frequency deviation from the maximum allowable deviation within the standard operating range. Simultaneously, the duration of the optimized actual frequency deviation is calculated, and the duration reduction is obtained by subtracting the actual duration from the expected duration. The frequency deviation reduction and duration reduction together constitute the expected effect of grid frequency stability support. When calculating the expected effect of peak-shaving resource utilization, the current total load and future load of the grid in the data combination are first considered. The total peak-shaving capacity required by the power grid is calculated using 2-hour load change forecast data. This is obtained by subtracting the current total load from the maximum predicted load value for the next 2 hours, and this load increment is the total peak-shaving capacity required by the power grid. For example, if the current total load is 800 kW and the maximum predicted load for the next 2 hours is 1200 kW, then the required total peak-shaving capacity is 1200 kW minus 800 kW. If no optimization strategy is adopted, this total peak-shaving capacity must be entirely provided by traditional peak-shaving resources. The unit cost of traditional peak-shaving resources is determined to be 0, referencing the local electricity market's peak-shaving service price.The total cost of traditional peak-shaving resources is calculated by multiplying the total peak-shaving resource capacity by the unit cost of traditional peak-shaving resources. A pre-trained reinforcement learning model determines the discharge power curve of the energy storage unit over the next two hours. This curve is dynamically adjusted according to the grid load change trend; for example, increasing the discharge power during periods of rapid load increase and maintaining a stable discharge power during periods of stable load. Integrating this curve over a two-hour time axis yields the total discharge capacity of the energy storage unit. This total discharge capacity represents the capacity of traditional peak-shaving resources that the energy storage unit can replace. Multiplying the replaced traditional peak-shaving resource capacity by the unit cost of traditional peak-shaving resources gives the value of energy storage replacing peak-shaving resources. Subtracting the value of energy storage replacing peak-shaving resources from the total cost of traditional peak-shaving resources gives the expected effect of peak-shaving resource utilization.

[0045] The action generation unit in a reinforcement learning model is the core functional part responsible for outputting specific operation instructions after the model has been trained. It includes an action space definition and an action selection strategy, both designed based on the actual equipment configuration and operational requirements of the community. The action space definition comprehensively covers all operational dimensions of the energy storage unit and charging pile. The action dimensions of the energy storage unit include charging / discharging state, charging / discharging power, and charging / discharging time interval. The charging / discharging state is divided into three explicit options: charging, discharging, and stopping. The charging / discharging power is divided into five fixed levels based on the rated power of the community's energy storage units: 10 kW, 20 kW, 30 kW, 40 kW, and 50 kW. The charging / discharging time interval is set in 5-minute increments. The clock unit is consistent with the time index of the data combination, and any continuous or non-continuous 5-minute sub-periods within the next 24 hours can be selected. The action dimensions of the charging pile include power source, power supply, and power supply time interval. The power source is divided into three options: grid power supply, energy storage power supply, and direct photovoltaic power supply. The power supply is divided into four fixed levels according to the type of charging pile in the community: 3 kW, 7 kW, 11 kW, and 20 kW. Among them, 3 kW and 7 kW are suitable for AC slow charging piles, and 11 kW and 20 kW are suitable for DC fast charging piles. The power supply time interval is completely synchronized with the charging and discharging time interval of the energy storage unit to ensure supply and demand matching. The action selection strategy adopts the ε-greedy strategy, and the ε value is set to 0.2. This numerical setting ensures an 80% probability of selecting the final action combination learned during the reinforcement learning model training process, guaranteeing the overall effectiveness of the solution, while retaining a 20% probability of random action selection to avoid the reinforcement learning model getting stuck in local finalities and ensure the diversity of generated solutions. In specific implementation, the calculated expected effects of user charging cost savings, grid frequency stability support, and peak-shaving resource utilization, along with the original data in the data combination, are simultaneously input into the action generation unit. The action generation unit first queries the final action mapping table stored during the reinforcement learning model training process based on this input information. This mapping table records the high-yield action combinations corresponding to different input conditions, filters out candidate action combinations with a matching degree higher than 85% with the current input information, and then selects candidate action combinations through an ε-greedy strategy to generate multiple different preliminary charging and discharging operation schemes. Each preliminary charging and discharging operation scheme contains complete and detailed operational details. For example, during the midday period of 11:00-14:00 when photovoltaic output is sufficient, the scheme specifies that the energy storage unit charges at a power of 30 kilowatts, and the charging piles prioritize direct photovoltaic power supply. For three non-urgent charging vehicles, two will be allocated 7 kW of power, and the vehicles scheduled for charging will be allocated 11 kW. Two fast-charging interfaces will also be reserved for temporary charging vehicles. During the evening peak electricity consumption period from 5:00 PM to 9:00 PM, the plan specifies that the energy storage unit will discharge at 40 kW, the charging piles will switch to energy storage power supply, 20 kW of fast charging will be allocated to five vehicles urgently needing charging, and 7 kW of slow charging will be allocated to three non-urgent vehicles. Adding new temporary charging vehicles will be restricted. During the nighttime off-peak period from 11:00 PM to 5:00 AM the next day, the plan specifies… The energy storage unit is charged at full power (50 kW), and all charging piles are switched to grid power. 3 kW or 11 kW of power is allocated to each reserved vehicle based on demand, maximizing the use of low-priced electricity. During cloudy days when solar power output is insufficient, the plan specifies that the energy storage unit maintains its current state of charge, with grid power as the primary source and energy storage power as a secondary source, prioritizing charging for vehicles in urgent need. Through these methods, a preliminary set of charging and discharging operation plans with 20 to 25 different operational details is generated, each focusing on different revenue dimensions, comprehensively covering different operating scenarios in the community.

[0046] First, the constraints of the feasibility check are clearly defined. These standards are based on the actual hardware performance and safe operation requirements of the community power distribution network, energy storage unit, and charging pile to ensure their enforceability. Specifically, the constraints for the energy storage unit are: maximum charging power of 50 kW, maximum discharging power of 40 kW, allowable charging / discharging terminal voltage range of 380V±10V, allowable operating current range of 0A-120A, minimum state of charge of 10% and maximum of 90%, internal resistance of 0.5 ohms, and no continuous overcurrent or overvoltage conditions exceeding 3 seconds during charging / discharging (i.e., operating current exceeding 120A). The duration of a voltage exceeding 380V±10V at the terminal cannot exceed 3 seconds; the specific constraints for charging piles are as follows: the maximum power supply of a single AC slow charging pile is 7 kW, the maximum power supply of a single DC fast charging pile is 60 kW, the maximum carrying current of the community power distribution line is 400A, the line impedance is 0.1 ohms, the total power supply current obtained by dividing the sum of the power supply of all charging piles at the same time by the line voltage of 380V cannot exceed 400A, and the power supply of a single pile cannot exceed its corresponding maximum power supply. The specific constraints for the safe operation of the power grid are that the power grid frequency must be maintained at 50Hz±0.Within a 2Hz range, the grid line voltage must be maintained within 380V±7%, and the harmonic distortion rate generated by charging and discharging operations must not exceed 5%. The harmonic distortion rate is calculated based on the harmonic emission standards of the energy storage units and charging piles within the community, using the Total Harmonic Distortion (THD) calculation method. Then, each scheme in the preliminary charging and discharging operation scheme set is thoroughly and meticulously checked, proceeding sequentially from energy storage unit to charging pile and grid safety to ensure no omissions. When checking the energy storage unit operation, the charging and discharging status, charging and discharging power, and start and end times of the energy storage unit in the scheme are extracted. First, it is determined whether the charging and discharging power is between the maximum charging power of 50 kW and the maximum discharging power of 40 kW. If the charging power exceeds 50 kW or the discharging power exceeds 40 kW, the scheme is directly deemed infeasible. Then, the charging and discharging terminal voltage is calculated based on the charging and discharging power and the internal resistance of the energy storage unit. The terminal voltage during charging... The terminal voltage during discharge is calculated as follows: 380V plus the charging / discharging power divided by the operating current and then multiplied by the internal resistance. During discharge, the terminal voltage is equal to 380V minus the charging / discharging power divided by the operating current and then multiplied by the internal resistance. The operating current is equal to the charging / discharging power divided by the terminal voltage. The actual values ​​of the terminal voltage and operating current are obtained through iterative calculation. It is determined whether the terminal voltage is within the range of 380V ± 10V and the operating current is within the range of 0A-120A. Finally, the state of charge (SOC) at the end of charging / discharging is calculated by adding the product of the charging power and the charging time to the current SOC, or subtracting the product of the discharging power and the discharging time. The total charge at the end is then divided by the rated total charge to obtain the SOC at the end. It is determined whether this SOC is between 10% and 90%. Simultaneously, it is checked whether there are overcurrent or overvoltage conditions lasting more than 3 seconds during the charging / discharging process. If any of these conditions are not met, the energy storage unit is not feasible to operate.

[0047] When checking the operation of charging piles, extract the type, power supply, and power supply time of each charging pile in the plan. First, determine whether the power supply of each charging pile does not exceed its corresponding maximum power supply: AC slow charging piles should not exceed 7 kW, and DC fast charging piles should not exceed 60 kW. Then, add up the power supply of all charging piles at the same time point to obtain the total power supply. Divide the total power supply by 380V to obtain the total power supply current. Determine whether the total power supply current does not exceed 400A. If any of these conditions are not met, the charging pile operation is not feasible. When checking the grid safety status, combine the charging load data predicted by the charging and discharging power in the plan with the current grid status data to simulate and calculate the frequency and voltage changes of the grid after the plan is implemented. The frequency change is equal to the difference between the predicted charging load increment and the energy storage charging and discharging power, multiplied by the frequency response coefficient of 0.05Hz per 100 kW. The current grid frequency is then added to the frequency change. The following checks are performed: Quantity is measured to determine if it falls within the range of 50Hz ± 0.2Hz; the voltage change is equal to the total supply current multiplied by the distribution line impedance of 0.1 ohms, and the current line voltage is subtracted from the voltage change to determine if it falls within the range of 380V ± 7%; simultaneously, based on the charging and discharging power and the harmonic emission parameters of the charging pile in the plan, the total harmonic distortion rate is calculated to determine if it does not exceed 5%. If any one of these conditions is not met, the grid safety is not up to standard. If a plan meets all the constraints in the three checks of energy storage unit, charging pile, and grid safety, it is considered a feasible plan; if any one of the check items fails to meet the constraints, it is considered an infeasible plan and is eliminated. The reasons for elimination are recorded in detail, including the name of the unmet constraint, the actual value in the plan, the constraint standard value, and the degree of difference. After all plans are checked, the remaining feasible plans are arranged in the order of generation to form a set of feasible operational plans.

[0048] When predicting the future sustainable benefits of feasible operational schemes, the prediction period is first defined. Considering the temporal distribution characteristics of community charging demand, the prediction period is set to the next two hours. This period covers the key stages of a single charging session and the dynamic changes in grid load. During the prediction process, based on the generated data combination, combined with the community grid load change patterns, photovoltaic power generation fluctuation patterns, and user charging behavior trends, the dynamic changes in the real-time grid status, energy storage unit state of charge, user charging demand, and time-of-use pricing within the prediction period are predicted. Then, each feasible operational scheme is substituted into the above dynamic change prediction data, and the user charging cost, grid frequency stability support, and peak-shaving resource utilization benefits of the scheme at each time point within the prediction period are recalculated. Finally, these benefits are summed to obtain the total benefits of each feasible scheme throughout the entire prediction period. The total sustainable benefits within the community are calculated. When comparing sustainable benefits, a comprehensive benefit assessment logic is adopted. Combining the current core needs of the community, reasonable weights are assigned to the three types of benefits. For example, during the evening peak hours, the weight of grid frequency stability and peak-shaving resource utilization is emphasized; during periods of large electricity price differences, the weight of user charging costs is emphasized; and during the midday solar peak hours, the weight of clean energy utilization is emphasized. Then, the comprehensive benefit value of each feasible solution is calculated based on the weights. By ranking the comprehensive benefit values ​​of all feasible solutions, the solution with the highest comprehensive benefit value is selected as the optimal solution. Finally, the charging and discharging operation parameters in the final solution are converted into specific instructions, including power control instructions for each vehicle from the charging pile, charging and discharging power instructions for the energy storage unit, and clear execution time node instructions. These instructions together constitute the final charging and discharging strategy instruction set.

[0049] By regulating the charging and discharging of energy storage units, the grid load is effectively shared, avoiding short-term overload of distribution lines and ensuring the stability of residents' daily electricity use. The diverse solutions generated by the action generation unit cover different operating scenarios in the community. After rigorous feasibility checks that meet the requirements of equipment performance and grid safety, the safety and feasibility of the final solution are ensured, avoiding equipment damage or grid safety hazards caused by improper operation and improving the reliability of system operation.

[0050] like Figure 2 As shown, in a preferred embodiment of the present invention, controlling the energy storage unit to execute corresponding instructions according to the charging and discharging strategy instruction set, and synchronously collecting instruction execution characteristic data, may include:

[0051] The charging and discharging strategy instruction set is analyzed, the specified power threshold parameters and time window parameters are extracted, and dynamic charging and discharging power allocation instructions for each energy storage unit are generated.

[0052] Execute dynamic power allocation commands for charging and discharging to control each energy storage unit to perform charging and discharging operations according to the target power.

[0053] During command execution, the actual charging and discharging rate of each energy storage unit, the real-time deviation data between the terminal voltage and the target voltage, and the command response delay data from command issuance to power response are collected in real time.

[0054] Based on real-time voltage deviation data, the cumulative voltage offset of each energy storage unit within the complete time window is calculated.

[0055] Multi-dimensional feature fusion is performed on the actual charge / discharge rate, command response delay data, and accumulated voltage offset to generate a dynamic feature vector characterizing the current command execution state.

[0056] In this embodiment of the invention, each instruction text in the charge / discharge strategy instruction set is analyzed one by one to identify descriptions related to power limits, such as "charging power must not exceed 50kW" and "discharge power is minimum 10kW". 50kW is extracted as the upper limit threshold for charging power, and 10kW as the lower limit threshold for discharging power. The instructions are then searched for statements regarding time ranges, such as "charge instructions are executed daily from 8:00 to 18:00". The time difference between 18:00 and 8:00 is calculated to obtain 36,000 seconds (10 hours × 3600 seconds / hour), which is determined as the time window parameter for this instruction. The total number of energy storage units participating in the scheduling is counted. Set up 3 units, labeled A, B, and C respectively. Obtain the rated power of each unit, for example, A is 30kW, B is 40kW, and C is 30kW, with a total rated power of 100kW. If the current total charging and discharging power demand is 80kW, allocate power according to the rated power ratio of each unit. The allocated power of A is 80×(30 / 100)=24kW, B is 80×(40 / 100)=32kW, and C is 80×(30 / 100)=24kW. At the same time, check whether each allocated power is within its power threshold range. If the charging upper limit of A is 30kW and 24kW is not exceeded, then retain the value and finally generate the dynamic allocation instruction for each unit.

[0057] Each energy storage unit receives a corresponding dynamic allocation command. For example, if the target charging power of unit A is 24kW, the current output power of unit A is monitored in real time. Assuming the initial output power is 15kW, the difference between the current and the target power is calculated to be 24-15=9kW. To achieve the target power, the input current of unit A is gradually adjusted, increasing the current every 30 seconds. After each adjustment, the power is measured again until the power stabilizes within the range of 24kW±1kW. If the power reaches 25kW after a certain adjustment, exceeding the allowable fluctuation range, the current is appropriately reduced to bring the power back down to around 24kW. For discharge operations, for example, if the target discharge power of unit B is 32kW and the initial actual power is 35kW, the difference is calculated to be 32-35=-3kW. The power is reduced by decreasing the output current, adjusting every 30 seconds until the power stabilizes within the range of 32kW±1kW.

[0058] Connect a power meter to unit A and record the amount of electricity charged within 60 seconds. If 0.4 kWh is charged within 60 seconds, the actual charging rate is calculated as 0.4 kWh ÷ (60 / 3600 hours) = 24 kW (the unit of rate, "kW," is consistent with the unit of power, and no additional time dimension is needed). For unit B, if 0.533 kWh is discharged within 60 seconds, the actual discharge rate is 0.533 kWh ÷ (60 / 3600 hours) = 32 kW. The target voltage for Unit A during charging is 500V. The terminal voltage is measured every 5 seconds by a voltage sensor. If a measurement value is 502V, the deviation is calculated as 502-500=2V; if another measurement value is 498V, the deviation is 498-500=-2V. The time when the charging and discharging command of Unit A is issued is recorded as 8:00:00, and the time when the power starts to change is monitored as 8:00:02. The time difference between the two is calculated to be 2 seconds, which is the command response delay data of this unit.

[0059] Given that the time window for Unit A is 36,000 seconds (10 hours), voltage deviation data is collected every 5 seconds within this window. 36,000 seconds contains 36,000 ÷ 5 = 7,200 data points. The absolute value of the deviation for each data point is taken. For example, if the deviation is 2V in a certain 5-second period, the absolute value is 2V; if the deviation is -2V in a certain 5-second period, the absolute value is 2V. Each absolute value is multiplied by the time interval (5 seconds) to obtain the voltage offset for that period, such as 2V × 5 seconds = 10V·second. The voltage offsets for the 7,200 periods are added together, for example, the sum is 72,000V·second, which is the cumulative voltage offset of Unit A within this time window.

[0060] The actual charge / discharge rate range is 0-50kW, the command response delay range is 0-10 seconds, and the cumulative voltage offset range is 0-180000V·s (calculated based on a 10-hour window and a maximum deviation of 5V: 5V×36000 seconds=180000V·s). The actual data of Unit A is standardized. The actual charging rate is 24kW, and the standardized value is 24÷50=0.48; the command response delay is 2 seconds, and the standardized value is 2÷10=0.2; the cumulative voltage offset is 72000V·s, and the standardized value is 72000÷180000=0.4. The three values ​​are combined in the order of "actual charge / discharge rate standardized value - command response delay standardized value - cumulative voltage offset standardized value" to form (0.48, 0.2, 0.4), which forms the dynamic feature vector of the current command execution state of Unit A.

[0061] By extracting detailed power thresholds and time window parameters and precisely allocating power, it ensures that the operation of each energy storage unit is strictly limited to a safe and efficient range, avoiding equipment damage or energy waste caused by improper power allocation. This makes the overall charging and discharging process more orderly. Real-time monitoring and adjustment of power to match target values ​​ensures that the charging and discharging state of the energy storage unit is always highly consistent with the command requirements, reducing the impact of power fluctuations on the power grid, ensuring the stability of grid voltage and frequency, and providing users with a continuous and reliable energy supply. Detailed collection of data such as actual charging and discharging rates, voltage deviations, and response delays can comprehensively capture the real-time operating details of the energy storage unit, promptly detecting problems such as abnormal rates, excessive voltage fluctuations, and slow response. Calculating the cumulative voltage offset within the complete time window can comprehensively reflect the voltage stability of the energy storage unit throughout the entire operating cycle. Compared with deviation data at a single time point, it better reflects the voltage control effect in long-term operation, providing a comprehensive reference for evaluating unit performance. Multi-dimensional feature fusion generates dynamic feature vectors, which can integrate scattered operating data into an intuitive comprehensive indicator, improving the intelligent management level of the entire energy storage system.

[0062] In a preferred embodiment of the present invention, based on instruction execution characteristic data, the user's charging cost savings, the contribution to grid frequency stability, and the peak-shaving resource savings are calculated respectively to generate a basic evaluation index set, which may include:

[0063] Based on the actual charging and discharging rate and time-of-use electricity price information in the instruction execution feature data, the amount of charging cost savings for users is calculated and generated.

[0064] The amount of savings in user charging costs is coupled with the cumulative voltage offset and grid frequency fluctuation data to generate the contribution to grid frequency stability.

[0065] The contribution of power grid frequency stability is matched and calculated with the characteristics of command response delay and the gap in power grid peak-shaving demand to generate peak-shaving resource saving benefits.

[0066] The system integrates multiple dimensions of user charging cost savings, grid frequency stability contribution, and peak-shaving resource savings to generate a set of basic evaluation indicators.

[0067] In this embodiment of the invention, the actual charging and discharging rates and durations (unit of time is uniformly seconds) of a user's associated energy storage unit are extracted from the instruction execution feature data for each time period of the day. For example, during the off-peak period (0:00-6:00, a total of 21600 seconds, electricity price 0.5 yuan / kWh), the actual charging rate is 20kW, and the charging continues for 10800 seconds (3 hours), so the charging amount during this period is calculated as 20kW × (10800 ÷ 3600) hours = 60kWh; during the flat period (6:00-8:00, a total of 7200 seconds, electricity price 0.8 yuan / kWh), the actual discharging rate is 10kW, and the discharging continues for 3600 seconds (1 hour), so the discharging amount is 10kW × (3600 ÷ 3600) hours = 10kWh; during the peak period (8:00-22:00, a total of 50400 seconds, ... Within an electricity price of 1.2 yuan / kWh, the actual discharge rate is 15kW, with continuous discharge for 7200 seconds (2 hours). The discharge amount is 15kW × (7200 ÷ 3600) hours = 30kWh. The user's total daily charging demand is 60kWh (equal to the energy storage charging amount). If distributed according to the actual electricity consumption time period (10kWh in the flat period corresponds to 3600 seconds, 30kWh in the peak period corresponds to 7200 seconds, and the remaining 20kWh in the valley period corresponds to 7200 seconds), the traditional cost is 10kWh × 0.8 yuan / kWh + 30kWh × 1.2 yuan / kWh + 20kWh × 0.5 yuan / kWh = 8 yuan + 36 yuan + 10 yuan = 54 yuan. Only the cost of charging 60kWh in the valley period needs to be borne, that is, 60kWh × 0.5 yuan / kWh = 30 yuan. No additional electricity cost is required for discharge during the flat and peak periods.

[0068] The cumulative voltage deviation of the energy storage unit during this period is extracted as 800 V·s (time unit is unified to seconds, consistent with the time dimension of the voltage deviation). The maximum allowable cumulative voltage deviation benchmark value of the power grid is set at 2000 V·s. The voltage stability correction coefficient is calculated as 1 - (actual cumulative deviation ÷ benchmark value) = 1 - (800 ÷ 2000) = 0.6 (the higher the value, the more stable the voltage). A total of 3600 instantaneous values ​​with a frequency deviation of 50 Hz are recorded (once per second, consistent with the time unit). The sum of the absolute values ​​of all instantaneous deviations is calculated to be 720 Hz. The average deviation is 720 Hz ÷ 3600 = 0.2 Hz. The average deviation of the frequency fluctuation benchmark is set at 0.5 Hz. The frequency stability correction coefficient is calculated. Positive coefficient = 1 - (actual average deviation ÷ benchmark deviation) = 1 - (0.2 ÷ 0.5) = 0.6; Taking the maximum possible savings of 50 yuan as the benchmark, the standardized value of the current savings of 24 yuan = 24 ÷ 50 = 0.48 (eliminating the influence of currency units and unifying it to a relative value of 0-1); set the weights of the three parameters (the sum is 1), the standardized value of cost savings accounts for 30%, the voltage stability correction coefficient accounts for 30%, and the frequency stability correction coefficient accounts for 40%, and couple the calculation of the contribution of power grid frequency stability = 0.48 × 0.3 + 0.6 × 0.3 + 0.6 × 0.4 = 0.144 + 0.18 + 0.24 = 0.564 (the result is rounded to three decimal places, ranging from 0 to 1, the higher the value, the greater the contribution).

[0069] Extracting the command response delay characteristics, the energy storage unit received 10 charge / discharge commands during this period. The response delays (unit: seconds) for each command were 1 second, 2 seconds, 1.5 seconds, 3 seconds, 2.5 seconds, 1 second, 2 seconds, 3.5 seconds, 2 seconds, and 1.5 seconds, respectively. The average response delay was calculated as (1+2+1.5+3+2.5+1+2+3.5+2+1.5)÷10=20÷10=2 seconds. The baseline response delay was set to 5 seconds. The response efficiency coefficient was calculated as 1-(average response delay ÷ baseline delay)=1-(2÷5)=0. 0.4 (unified as a relative value of 0-1), collect data on the peak-shaving demand gap of the power grid. During this period, the peak-shaving demand gap of the power grid during the morning peak (8:00-10:00, a total of 7200 seconds) is 800kW, and the actual discharge contribution of the energy storage unit is 320kW; the gap during the evening peak (18:00-20:00, a total of 7200 seconds) is 1000kW, and the actual contribution is 400kW. Calculate the total peak-shaving coverage ratio = (320+400)÷(800+1000) = 720÷1800 = 0.4 (unified as a relative value of 0-1).

[0070] Assuming a weighting of 40% for grid frequency stability contribution, 20% for response efficiency coefficient, and 40% for peak-shaving coverage ratio (totaling 1), the comprehensive peak-shaving coefficient is calculated as: 0.564 × 0.4 + 0.62 × 0.2 + 0.4 × 0.4 = 0.2256 + 0.124 + 0.16 = 0.5096 (consistent with relative values ​​between 0 and 1). Given that the unit cost of purchasing peak-shaving resources for the grid during this period is 1.2 yuan / kW, and the total peak-shaving demand gap is 1800kW, the peak-shaving resource saving benefit is calculated as: Comprehensive Peak-Shaving Coefficient × Total Gap × Unit Cost = 0.5096 × 1800 × 1.2 ≈ 1100 yuan. Standardization rules for each indicator are determined, and users... The upper limit for charging cost savings is set at 100 yuan, the contribution of grid frequency stability is set at 0-1, and the upper limit for peak-shaving resource savings is set at 2000 yuan. Each indicator is standardized: the standardized value for user charging cost savings is 24 ÷ 100 = 0.24; the standardized value for peak-shaving resource savings is 1100 ÷ 2000 = 0.55. Integrating the original and standardized data, a basic evaluation indicator set is generated according to the structure: "User charging cost savings (original value / standardized value) - grid frequency stability contribution - peak-shaving resource savings (original value / standardized value)", specifically (24 yuan / 0.24, 0.564, 1100 yuan / 0.55).

[0071] By refining the matching calculations of charging and discharging rates with electricity prices at different times, the cost savings for users in energy storage systems can be accurately quantified, allowing users to intuitively perceive the economic benefits and increasing their acceptance and enthusiasm for using the system. The cost savings are coupled with voltage offset and frequency fluctuation data to calculate the contribution to grid frequency stability, taking into account both economic factors and grid operation stability indicators. This makes the contribution assessment more comprehensive and objectively reflects the actual role of energy storage systems in grid frequency stability. Combining grid frequency stability contribution, command response delay, and peak-shaving gap data to calculate peak-shaving resource savings quantifies the specific value of energy storage systems in alleviating grid peak-shaving pressure. This helps grid operators clarify the actual contribution of energy storage systems to reducing peak-shaving resource input and provides data support for resource allocation optimization.

[0072] In a preferred embodiment of the present invention, key feature parameters are extracted from the dynamic changes of the basic evaluation index set, and nonlinear coupling relationships are analyzed by establishing a correlation model between the parameters to generate dynamic compensation coefficients. This may include:

[0073] Time series analysis was performed on the basic evaluation indicator set to identify the dynamic change patterns of each indicator;

[0074] Based on the dynamic change pattern, the fluctuation range of user charging cost savings, the rate of change of the contribution of grid frequency stability, and the decay trend of peak-shaving resource savings are extracted as key feature parameters.

[0075] By inputting key feature parameters into a pre-defined correlation model, the synergistic effect between cost savings and frequency stability contribution, as well as the constraint relationship between peak shaving benefits and charging demand fluctuations, can be analyzed.

[0076] Based on the synergy and constraint analysis results obtained from the correlation model, the nonlinear coupling strength between each parameter is quantified, and a comprehensive dynamic compensation coefficient is calculated based on the nonlinear coupling strength.

[0077] In this embodiment of the invention, firstly, complete operational data of three core indicators in the basic evaluation index set are comprehensively and accurately acquired over a continuous time dimension. The three indicators are user charging cost savings, grid frequency stability contribution, and peak-shaving resource savings. All data are sourced from the real-time monitoring unit of the community's energy storage charging piles, the grid dispatch interface, and the user charging record system. The acquisition frequency is set to once per minute to ensure data continuity and timeliness. The acquisition period fully covers various typical operating scenarios, such as the noon photovoltaic power supply period in the community, the peak period of concentrated electric vehicle charging in the evening, the low grid load period at night, and the period of insufficient photovoltaic output on cloudy or rainy days. At the same time, the environmental parameters of light intensity are recorded for each period. The system collects data on temperature, grid operation parameters, line current and voltage deviations, user charging behavior, number of charging vehicles, and initial battery range to ensure that the data comprehensively reflects the true changing patterns of indicators under different scenarios. Next, the time series data for each type of indicator undergoes systematic preprocessing using a fixed-window sliding average method. The sliding window is set to 5 minutes, meaning that the indicator values ​​at five consecutive time points are summed and divided by 5 to obtain the smoothed value at the midpoint of the window. This method gradually covers all time points, eliminating abnormal interference data caused by factors such as sudden charging starts by individual vehicles in the community, instantaneous power output fluctuations due to cloud cover on photovoltaic panels, and short-term voltage surges in the grid, making the data changes more closely reflect actual operating trends.

[0078] After preprocessing, the time series data of each indicator is regarded as a continuous and complete data change trajectory. The trajectory is scanned and analyzed node by node to accurately identify key feature points. The criteria for judging each feature point are clear: peak point is judged as the value of three consecutive adjacent time nodes satisfying the condition that the value of the previous node is less than the value of the middle node is greater than the value of the next node, and the increase of the value of the middle node compared with the value of the previous node exceeds a preset peak threshold of 10%; valley point is judged as the value of three consecutive adjacent time nodes satisfying the condition that the value of the previous node is greater than the value of the middle node is less than the value of the next node, and the decrease of the value of the middle node compared with the value of the previous node exceeds a preset valley threshold of 10%; inflection point is judged as the direction of change of the value of two adjacent time nodes reverses from rising to falling or from falling to rising, and the absolute value of the change exceeds a preset inflection point threshold of 8%; the starting point of the stable segment is judged as the absolute value of the change of the value of any two adjacent nodes within five consecutive time nodes is less than a preset stable threshold of 5%; the ending point of the stable segment is judged as the first time node after the starting point of the stable segment where the absolute value of the change of the value of adjacent nodes exceeds 5%.

[0079] Finally, based on the precise minute-by-minute distribution of these key feature points on the time axis, the specific differences and percentages calculated from the numerical changes between adjacent feature points, the connection methods between feature points (such as peak points, stable segments, inflection points, valley points, rising segments, and peak points), and the time intervals corresponding to each feature point, the specific number of minutes is calculated. This allows for the summarization of the dynamic change patterns of each type of indicator. For example, user charging cost savings, when solar power is abundant and sunlight intensity is stable at noon, will rapidly increase from an initial value with a stable increase until reaching a peak point and then entering a stable segment, exhibiting a rapid increase, peak, and stable change pattern. The contribution of grid frequency stability will rapidly increase with the increase in charging load during the evening charging peak, reaching a peak and then slowly decreasing as the grid load gradually eases, exhibiting a rapid increase and slow decrease change pattern. Peak-shaving resource savings will gradually decrease as the duration of the load increases during periods of sustained high charging load, exhibiting a continuous and slow decay change pattern.

[0080] Based on the accurately identified dynamic change patterns of each indicator, key feature parameters are extracted in greater depth and detail. The extraction process for each parameter includes a complete operation procedure and calculation logic to ensure accurate and reproducible results. For the fluctuation range of user charging cost savings, the time series data change trajectory of this indicator is first scanned throughout the entire period. The scanning step size is consistent with the data collection frequency, which is once per minute. All peak points and valley points that meet the judgment criteria on the trajectory are accurately identified. These peak points and valley points are numbered sequentially according to time sequence to ensure that each peak point forms a unique matching pair with the next adjacent valley point, avoiding mismatches across time periods. The specific cost savings value corresponding to each peak point and the specific cost savings value corresponding to each valley point are recorded one by one. Then, for each pair of matching peak and valley points, the value corresponding to the peak point is subtracted from the value corresponding to the next adjacent valley point to calculate the numerical difference between each pair of peak and valley points. All calculated differences are sorted in descending order, and the largest difference ranked first is selected as the fluctuation range of user charging cost savings. This value directly reflects the maximum fluctuation range of cost savings during the monitoring period.

[0081] To determine the rate of change of the power grid frequency stability contribution, a second in-depth analysis of the dynamic change pattern of this indicator is first performed. During the analysis, the focus is on the inflection points in the direction of numerical change. All inflection points that meet the judgment criteria are selected and numbered in chronological order. Starting from the first inflection point, two adjacent inflection points are selected as a group for calculation. The frequency stability contribution value corresponding to the later inflection point is subtracted from the frequency stability contribution value corresponding to the earlier inflection point. The resulting difference is then divided by the time interval between the two inflection points, calculated to the second based on the difference in time points. This yields the instantaneous rate of change of the power grid frequency stability contribution within that time period. Following the same calculation method, the rate of change calculation is performed for all time periods consisting of adjacent inflection points. All the obtained instantaneous rate of change values ​​are accumulated, and the sum is divided by the total number of instantaneous rates of change to obtain the average rate of change of the power grid frequency stability contribution. This average rate of change is used as a key characteristic parameter characterizing the speed of change of this indicator.

[0082] To assess the decay trend of peak-shaving resource saving benefits, multiple consecutive feature points are extracted from the time-series data of this indicator at fixed time intervals. The fixed time interval is set to 5 minutes to ensure a comprehensive and uniform capture of the indicator's changing trend. The extracted feature points are sorted chronologically. Starting from the first feature point, the difference between the peak-shaving resource saving benefit value corresponding to each subsequent feature point and the value corresponding to the previous feature point is calculated. If the calculated difference is negative, it indicates that the peak-shaving resource saving benefit is decaying during that period, and this negative difference is recorded and stored. If the difference is positive or zero, it indicates no decay during that period and is not included in the calculation. All stored negative differences are summed to obtain the total decay of peak-shaving resource saving benefits. The total decay is then divided by the total time length corresponding to these negative differences. The total time length is obtained by summing the time intervals of each decay period to a value accurate to the minute, ultimately yielding the average decay rate of peak-shaving resource saving benefits. This average decay rate is used as a key feature parameter characterizing the decay trend of this indicator.

[0083] First, a pre-defined correlation model is constructed. The entire model construction process revolves around the actual operational scenarios of energy storage charging piles in the community. Initially, a comprehensive and systematic collection of historical data generated by the long-term operation of energy storage charging piles within the community is undertaken. The collection period is set at 12 months, covering operational data under different weather conditions (spring, summer, autumn, winter, sunny, cloudy, rainy, snowy) to ensure the comprehensiveness and representativeness of the data. The collected historical data specifically includes data on different photovoltaic output levels, divided into five gradient intervals based on rated output, recording the real-time output power and light intensity matching data of the photovoltaic panels in each interval; data on the distribution of different user charging needs; scenarios with multiple vehicles charging simultaneously (no less than 80% of the charging pile capacity); and scenarios with distributed charging. The system records the number of vehicles charging at any given time, ranging from 30% to 50% of the charging pile capacity. It also records the battery capacity and initial driving range of each vehicle in mixed emergency and non-emergency scenarios, indicating the urgency level of charging and the required charging time. Furthermore, it tracks different grid load conditions: during off-peak hours, the grid load does not exceed 40% of the rated capacity; during normal hours, the grid load is 40% to 70% of the rated capacity; and during peak hours, the grid load is not less than 70% of the rated capacity. Real-time current, voltage, frequency, and dispatch instructions for distribution lines are also recorded. Additionally, it includes three types of core indicator data and quantitative data on the interrelationships between these indicators, such as the change in frequency stability contribution when cost savings increase by a certain percentage, and the attenuation of peak-shaving benefits when charging demand fluctuations increase by a certain percentage.

[0084] The collected historical data was then categorized, cleaned, and segmented. Based on a combination of photovoltaic power output, charging demand distribution, and grid load status, the data was divided into multiple data groups. The 3σ criterion was used to screen effective data: the mean and standard deviation of each group were calculated, and outliers exceeding the mean plus or minus three times the standard deviation were removed to ensure data reliability. The effective data was then divided into training and validation datasets in a 7:3 ratio. The training dataset was used for iterative optimization of model parameters, while the validation dataset was used for final testing of model performance. During the segmentation process, the scene distribution ratio of the two data groups was kept consistent to avoid scene bias. The training dataset was then input into the initially constructed association model for training. The core structure of the association model includes a feature input layer, a relationship analysis layer, and a result output layer. The functions and implementation logic of each layer are clearly defined. The feature input layer has three independent input ports, each corresponding to a key feature parameter. The ports support data format standardization, converting the input parameter values ​​into normalized values ​​from 0 to 1 for easier calculation by the association model. The relationship analysis layer is the core structure of the association model... The core computing unit employs the Delaunay triangulation algorithm to deeply analyze the correlation relationships of key feature parameters. Specifically, the value of each key feature parameter at different time points is considered as a discrete point in three-dimensional space. The x-axis represents the time dimension, the y-axis represents the fluctuation range of user charging cost savings, and the z-axis represents the rate of change of grid frequency stability contribution or the attenuation trend of peak-shaving resource savings. These discrete points are grouped into a point set and input into the Delaunay triangulation algorithm. The algorithm automatically connects adjacent points to form a non-overlapping triangular mesh that covers all points. By analyzing the side lengths, angles, and adjacency relationships of the triangles in the triangular mesh, the spatial correlation strength between points corresponding to different parameters is determined. For example, when the triangle formed by the point representing the fluctuation range of cost savings and the point representing the rate of change of frequency stability contribution has a short side length and a stable angle, it indicates a strong cooperative correlation between the two within that time interval. When the triangle formed by the point representing the attenuation trend of peak-shaving resource savings and the point representing the fluctuation of charging demand shows obvious stretching deformation, it indicates a forced constraint relationship between the two.

[0085] During the training of the correlation model, the core objective is to improve the model's accuracy in identifying synergistic effects and constraints. An initial learning rate of 0.001 and 1000 training iterations are set. After each iteration, the model's accuracy on the training dataset is calculated. If the accuracy improves, the current parameters are retained; if the accuracy decreases, the learning rate is halved and the model reverts to the previous parameters. Through continuous iteration, the correlation weight parameters within the model are adjusted to ensure accurate identification of the synergistic effect between cost savings and frequency stability contribution. Specifically, when the photovoltaic power in the community is sufficient, the fluctuation range of cost savings increases, and the rate of change of the grid frequency stability contribution increases simultaneously, with both mutually reinforcing each other. Simultaneously, the model is also ensured to accurately identify the constraint relationship between peak-shaving benefits and charging demand fluctuations. Specifically, when multiple vehicles charge simultaneously in the community, leading to increased charging demand fluctuations, the decline in peak-shaving resource savings intensifies, with both mutually constraining each other.

[0086] After training, the validation dataset is input into the association model for performance testing. The accuracy of the association model in identifying synergistic effects and constraints in the validation dataset is calculated by dividing the number of correctly identified synergistic or constraint cases by the total number of cases in the validation dataset and then multiplying by 100%. If the accuracy does not reach the preset 95% standard, the association weight parameters and learning rate of the association model are adjusted, and the testing process is repeated until the analysis accuracy of the association model meets the preset requirements, thus completing the construction of the association model. Then, the extracted key feature parameters are input into the constructed association model according to their corresponding ports. The association model uses its internal relationship analysis logic to parse the input parameter values ​​in real time and time period by time, and analyzes in detail the strength of the synergistic effect between the fluctuation range of user charging cost savings and the rate of change of the contribution of grid frequency stability, clarifying the specific degree of mutual promotion between the two in different operating scenarios. At the same time, it deeply analyzes the degree of constraint between the decay trend of peak-shaving resource savings and the fluctuation of charging demand, accurately determining the magnitude of the impact of charging demand fluctuations on the decay of peak-shaving benefits, providing a precise and detailed basis for the subsequent quantification of nonlinear coupling strength.

[0087] Based on the synergistic effect analysis and constraint relationship analysis results output by the correlation model, and combined with the grid analysis data from the Delaunay triangulation algorithm, the nonlinear coupling strength between various parameters is accurately quantified. Each step of the quantification process has clearly defined calculation logic and parameter basis to ensure the results are traceable and verifiable. For the synergistic coupling strength between user charging cost savings and grid frequency stability contribution, the calculation logic is as follows: First, obtain the synergistic effect strength value output by the correlation model. This value ranges from 0 to 1.0 and directly reflects the strength of the synergistic relationship between the two. Second, perform Delaunay triangulation... The algorithm generates a triangular mesh, and calculates the ratio of the mean side length of the triangle formed by the corresponding point sets of the two to the mean side length of all triangles to obtain the coefficient of synergistic association. The third step is to multiply the synergistic effect strength value by the coefficient of synergistic association to obtain the basic value of synergistic association. The fourth step is to superimpose the basic synergistic coefficient of the community's energy storage charging pile under standard operating conditions. The basic synergistic coefficient is calculated based on historical data of the community under standard operating conditions of stable photovoltaic output, balanced charging demand, and stable grid load over the past year, and is set as a fixed value. Finally, by calculating the basic value of synergistic association plus the basic synergistic coefficient, the synergistic coupling strength value of this set of parameters is obtained.

[0088] The calculation logic for the constraint coupling strength between peak-shaving resource saving benefits and charging demand fluctuations is as follows: First, obtain the constraint strength value output by the correlation model. This value ranges from 0 to 1.0 and directly reflects the strength of the constraint relationship between the two. Second, calculate the deformation coefficient of the triangle formed by the corresponding point sets of the two through the triangular mesh generated by the Delaunay triangulation algorithm. The deformation coefficient is the ratio of the actual area of ​​the triangle to the area of ​​the equilateral triangle. The smaller the ratio, the greater the deformation and the stronger the constraint. Third, multiply the constraint strength value by the triangle deformation coefficient to obtain the basic constraint correlation value. Fourth, superimpose the basic constraint coefficient of the two under standard operating conditions. The basic constraint coefficient is calculated based on the historical data of the community under standard operating conditions over the past year and is set as a fixed value. Finally, by calculating the basic constraint correlation value plus the basic constraint coefficient, the constraint coupling strength value of this set of parameters is obtained.

[0089] Then, the calculation of the comprehensive dynamic compensation coefficient begins. This step fully considers the actual energy configuration of the community and the grid operation requirements. The calculation process is progressive. The first step determines the priority weights for clean energy utilization and grid load regulation within the community. The weights are set based on the proportion of photovoltaic installed capacity and the grid load pressure level. If the photovoltaic installed capacity accounts for more than 30% of the total power supply capacity of the charging piles, and the grid load pressure is high during peak hours, then the priority weight for clean energy utilization is set to 0.6, and the priority weight for grid load regulation is set to 0.4, with the sum of the two weights always being 1.0. The second step multiplies the previously obtained synergistic coupling strength value by the priority weight for clean energy utilization to obtain the weighted value for synergistic coupling. The third step multiplies the constraint coupling strength value by the priority weight for grid load regulation to obtain the weighted value for synergistic coupling. The fourth step involves adding the weighted values ​​of the cooperative coupling and the restrictive coupling to obtain the total value of the comprehensive coupling strength. The fifth step calculates the benchmark coupling strength value for the operation of the community's energy storage charging piles. This value is the weighted sum of the average cooperative coupling strength and the average restrictive coupling strength of the community's charging piles under standard operating conditions over the past three months, with equal weights, and is set as a fixed benchmark value. The sixth step divides the total value of the comprehensive coupling strength by the benchmark coupling strength value to obtain the coupling strength ratio. The seventh step multiplies this ratio by a preset basic compensation coefficient, which is set according to relevant industry standards. If the community has special power grid requirements, it can be adjusted within a reasonable range. Through the above series of complete calculations, a dynamic compensation coefficient that can comprehensively reflect the nonlinear coupling relationship between various parameters is finally obtained.

[0090] By meticulously processing the time series data of the indicators multiple times, the subtle dynamic changes of each evaluation indicator under different operating scenarios in the community are captured. This ensures that the extraction process of key feature parameters fully conforms to the actual operating rules. The construction of the correlation model is based on the long-term historical data of the community. After incorporating the Delaunay triangulation algorithm, it can explore the spatial correlation between different parameters. This allows the identification of collaborative and restrictive relationships to no longer rely on single numerical comparisons, but rather combine geometric spatial structural features to conform to the actual operating conditions of the community's energy storage and charging piles.

[0091] In a preferred embodiment of the present invention, the basic evaluation index set is calibrated in real time using a dynamic compensation coefficient to obtain the final evaluation index set, which may include:

[0092] Using the dynamic compensation coefficient as a weighting parameter, the basic evaluation index set is weighted and calibrated to generate a weighted calibration index set that includes the weighted cost saving index.

[0093] The weighted cost-saving index is extracted from the weighted calibration index set. The compensation value is calculated based on the degree of decay of the peak-shaving resource saving benefits. The compensation value is then superimposed on the weighted cost-saving index to generate the compensated and corrected cost-saving index.

[0094] All indicators in the weighted calibration indicator set, except for the weighted cost savings indicator, are integrated with the compensated and corrected cost savings indicator to obtain the final evaluation indicator set.

[0095] In this embodiment of the invention, firstly, a weighted calibration index set is generated. The first step is to determine the dynamic compensation coefficient as a weight parameter. This requires collecting various real-time data that affect the basic evaluation index at the current moment, such as real-time resource supply, market price fluctuation data, and equipment operating status parameters. Then, the degree of influence of these data on each basic evaluation index is analyzed. Factors with a high degree of influence are assigned a higher weight. The dynamic compensation coefficient is obtained by comprehensively calculating these degrees of influence. Next, each index in the basic evaluation index set is weighted with this dynamic compensation coefficient. For example, if the basic evaluation index set includes raw material cost index, labor cost index, and energy consumption index, and the basic raw material cost index is M and the dynamic compensation coefficient is N, then the weighted raw material cost index is the result of multiplying M and N; if the basic labor cost index is P, the weighted labor cost index is the result of multiplying P and N, and so on. This weighted calculation is performed on all indicators in the basic evaluation index set to generate a weighted calibration index set that includes the weighted cost saving index.

[0096] Secondly, a compensated and corrected cost-saving indicator is generated. The weighted cost-saving indicator, denoted as Q, is precisely extracted from the weighted calibration indicator set. Then, the degree of attenuation of the peak-shaving resource savings is calculated. First, the savings data of peak-shaving resources in the initial stage is collected, for example, the savings brought by peak-shaving resources in the first month of project operation is R; then, the savings data of the current stage is collected, for example, the savings brought by peak-shaving resources in the third month of project operation is S. The attenuation of benefits is obtained by calculating (RS). The attenuation is divided by the savings R in the initial stage to obtain the attenuation ratio T, i.e., (RS) / R=T, which reflects the degree of attenuation. Then, the compensation value is calculated based on this attenuation ratio T. If the attenuation ratio T is between 0-20%, the compensation value is 5% of the weighted cost-saving indicator Q according to the preset rules; if the attenuation ratio T is between 21%-50%, the compensation value is 10% of Q, and so on to determine the compensation value U. Finally, the compensation value U is superimposed on the weighted cost-saving indicator Q, i.e., Q plus U, to obtain the compensated and corrected cost-saving indicator V.

[0097] Finally, the final evaluation index set is obtained by integrating all the indicators in the weighted calibration index set except for the weighted cost savings index, such as the weighted raw material cost index, the weighted labor cost index, the weighted energy consumption index, etc. Then, these indicators are put together with the compensated and corrected cost savings index V to form a complete index set, which is the final evaluation index set.

[0098] In terms of assessment accuracy, since the dynamic compensation coefficient is calculated based on real-time data, it accurately reflects the impact of various factors on the indicators, avoiding the problem of not being able to cope with real-time changes when using fixed weights. Furthermore, the calculation and summarization of compensation values ​​for the attenuation of peak-shaving resource savings further refines the deviation of cost savings indicators caused by changes in benefits, making each indicator more closely reflect the actual situation and improving the accuracy of the assessment. From a real-time perspective, the entire calibration process revolves around real-time data; the dynamic compensation coefficient is updated promptly as the real-time data changes, and the weighted calibration of the indicators is also performed in real time. The calculation of compensation values ​​is based on the current degree of attenuation, ensuring that the indicators keep pace with changes in the actual situation and making the evaluation results highly real-time, providing timely reference for decision-making. In terms of adaptability, this method can flexibly cope with various complex situations. When influencing factors change, the dynamic compensation coefficient will be adjusted accordingly to ensure that the weighted indicators can adapt to new situations. When the degree of attenuation of peak-shaving resource saving benefits is different, the calculated compensation values ​​will also be different, so that the cost saving indicators can adapt to different attenuation conditions. This flexibility allows the evaluation method to be effectively applied in various scenarios, improving the applicability of the evaluation system.

[0099] In a preferred embodiment of the present invention, based on the final evaluation index set, the updated grid status data and user behavior characteristics are integrated to adaptively adjust the configuration parameters of the decision engine, thereby achieving coordinated control of multiple energy sources in energy storage. This may include:

[0100] The final evaluation index set is input into the parameter update unit of the decision engine, and the grid frequency fluctuation data and user charging behavior deviation feature data are collected and updated in real time.

[0101] The final evaluation index set, real-time power grid frequency fluctuation data, and user charging behavior deviation feature data are integrated to form a dynamically adjusted input dataset.

[0102] Based on dynamically adjusted input datasets, the change values ​​of weight coefficients used for multi-objective processing in the decision engine are calculated using gradient descent algorithm;

[0103] Based on the changes in weight coefficients, the weight coefficients currently used in multi-objective processing are dynamically updated, an updated set of weight coefficients is generated, and the result is output to the instruction generation unit of the decision engine.

[0104] The instruction generation unit applies the updated weight coefficient set and combines it with real-time status information to generate multi-energy collaborative scheduling instructions for the next control cycle, thereby realizing the collaborative control of multiple energy storage sources.

[0105] In this embodiment of the invention, firstly, a dynamically adjusted input dataset is prepared. Each indicator in the final evaluation index set is input into the parameter update unit of the decision engine. Simultaneously, power grid frequency fluctuation data is collected in real time. The actual power grid frequency is recorded every 0.1 seconds through the power grid monitoring terminal. After 100 consecutive records, a set of frequency data is obtained. The difference between each value in this set of data and the standard frequency of the power grid is calculated. Then, the maximum, minimum, and average values ​​of these differences are statistically analyzed to fully present the amplitude and trend of frequency fluctuations. When collecting user charging behavior deviation feature data, the user's preset charging time period and preset charging amount are obtained first. Then, the user's actual charging start time, end time, and actual charging amount are recorded. The overlap ratio between the actual charging time period and the preset time period (overlap length ÷ preset length) and the deviation ratio between the actual charging amount and the preset charging amount ((actual amount - preset amount) ÷ preset amount) are calculated. These ratio data are used as deviation features. Finally, they are integrated in the order of "final evaluation index → ​​power grid frequency fluctuation data → user behavior deviation data" to form a dynamically adjusted input dataset.

[0106] Secondly, the weight coefficient changes are calculated. Based on the dynamically adjusted input dataset, the gradient descent algorithm is used to calculate the weight coefficient changes for multi-objective processing. The first step determines the multi-objective evaluation criteria, including three objectives: lowest energy cost, highest power supply stability, and ultimate user satisfaction. Each objective corresponds to an initial weight coefficient. The second step substitutes the dynamically adjusted input data into the multi-objective evaluation criteria. For example, cost index data is used to calculate the current score for the lowest energy cost objective, frequency fluctuation data is used to calculate the current score for the highest power supply stability objective, and behavioral offset data is used to calculate the current score for the ultimate user satisfaction objective. Then, the three scores are weighted and summed according to the initial weight coefficients to obtain the final score. The third step is to calculate the gradient of the target total score by slightly adjusting each weight coefficient and recalculating the target total score. This yields the magnitude of the change in the total score with each weight coefficient (i.e., the gradient). The fourth step is to adjust the weights along the gradient descent direction (the direction that leads to the final total score). If the gradient of the target with the lowest energy cost is positive (i.e., increasing the weight will improve the total score), then the change value of that weight is increased by the preset learning rate. If the gradient is negative, then the change value is decreased by the learning rate. This process of substituting data, calculating scores, finding gradients, and adjusting change values ​​is repeated until the target total score no longer improves significantly. The recorded change in each weight coefficient at this point is the final change value of the weight coefficient.

[0107] Then, an updated set of weight coefficients is generated. Each weight coefficient currently used by the decision engine is extracted and added to its corresponding change value. For example, if the current weight of the lowest energy cost target is 0.3 and the change value is +0.05, then the updated weight is 0.35; if the current weight of the highest power supply stability target is 0.4 and the change value is -0.02, then the updated weight is 0.38. This process is repeated to update all weight coefficients. All updated weight coefficients are then arranged in the order of the targets to form an updated set of weight coefficients, which is then sent to the instruction generation unit.

[0108] Finally, multi-energy coordinated dispatch instructions are generated. After receiving the updated weight coefficient set, the instruction generation unit collects real-time status information, including the current power of each energy storage device, the input power of each energy source, and the real-time load of the power grid. This information is combined with the weight coefficient set, and the priority of each target is calculated according to the updated weights. If the energy cost is the lowest and the weight is the highest, then the use of low-cost energy is given priority. If the power supply stability is the highest and the weight increases, then the energy storage devices are given priority to ensure sufficient reserve capacity. Based on the priority, specific dispatch schemes are formulated. For example, in the next 15-minute control cycle, photovoltaic full-power output of 200kW, wind power output of 150kW, and lithium battery energy storage discharge of 100kW are arranged to make up for the load gap, while the non-emergency charging power of users is limited to below 50kW. These schemes are converted into specific instructions as dispatch instructions for the next control cycle, realizing the coordinated control of multiple energy storage sources.

[0109] In terms of control precision, the system can capture minute state changes through grid frequency acquisition and detailed user behavior offset calculation. The gradient descent algorithm, after multiple iterations to adjust weight coefficients, can accurately find parameter configurations suitable for the current state, keeping the decision engine's response error within a minimal range. Regarding dynamic adaptability, the real-time integrated dataset ensures the system responds to changes in grid state and user behavior without delay. When users suddenly increase their charging demand, the weight coefficients are adjusted within one control cycle to prioritize user power supply. Conversely, when grid frequency fluctuations intensify, the weight of the power supply stability objective is rapidly increased to ensure timely power supply. By utilizing energy storage devices to smooth out fluctuations, the system can maintain efficient operation in complex and ever-changing scenarios. In terms of improving energy utilization efficiency, the dynamic optimization of multi-objective weights enables refined energy allocation. Regarding improving user experience, real-time response to user charging behavior deviations can reduce power consumption restrictions. In terms of ensuring system stability, the dynamic updating of weight coefficients allows multi-energy collaborative control of energy storage to proactively address potential risks. For example, when the grid frequency shows a downward trend, the system will increase the energy storage discharge weight in advance to replenish power before the frequency falls below the threshold, reducing the probability of grid instability and improving the reliability of energy supply.

[0110] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0111] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0112] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multi-energy collaborative control and management system for vehicle energy storage charging piles, characterized in that, include: The power grid sensing module is used to communicate with the power grid dispatching device through sensors on the charging pile side, collect power grid frequency, voltage and electricity price signals in real time, and build a real-time power grid status dataset. The strategy decision module is used to perform multi-objective real-time optimization strategy calculation based on the real-time power grid status dataset, combined with the current state of charge of the energy storage unit, the predicted user charging demand and time-of-use electricity price information, and using a pre-trained reinforcement learning model to obtain the final charging and discharging strategy instruction set. The execution control module is used to control the energy storage unit to execute corresponding instructions according to the charging and discharging strategy instruction set, and to synchronously collect instruction execution characteristic data; The benefit assessment module is used to calculate user charging cost savings, grid frequency stability contribution, and peak-shaving resource savings based on instruction execution characteristic data, generating a basic assessment index set. Specifically, this includes: calculating user charging cost savings based on actual charging and discharging rates and time-of-use pricing information in the instruction execution characteristic data; coupling user charging cost savings with cumulative voltage offset and grid frequency fluctuation data to generate grid frequency stability contribution; matching grid frequency stability contribution with instruction response delay characteristics and grid peak-shaving demand gap data to generate peak-shaving resource savings; and integrating user charging cost savings, grid frequency stability contribution, and peak-shaving resource savings in multiple dimensions to generate the basic assessment index set. The coupling analysis module is used to extract key feature parameters from the dynamic changes of the basic evaluation index set, and to analyze the nonlinear coupling relationship by establishing a correlation model between the parameters, thereby generating dynamic compensation coefficients. The evaluation and calibration module is used to calibrate the basic evaluation index set in real time using dynamic compensation coefficients to obtain the final evaluation index set. The adaptive module is used to adaptively adjust the configuration parameters of the decision engine based on the final evaluation index set, the updated grid status data and user behavior characteristics, and realize the coordinated control of multiple energy sources in energy storage.

2. The multi-energy collaborative control and management system for vehicle energy storage charging piles according to claim 1, characterized in that, Based on the real-time power grid status dataset, combined with the current state of charge of energy storage units, predicted user charging demand, and time-of-use pricing information, a pre-trained reinforcement learning model is used to calculate multi-objective real-time optimization strategies, resulting in the final charging and discharging strategy instruction set, including: The real-time status dataset of the power grid, the current state of charge of energy storage units, the predicted user charging demand, and time-of-use electricity price information are integrated to form a data combination; Based on data combination, a multi-objective benefit evaluation system is established through a pre-trained reinforcement learning model to simultaneously calculate the expected effects on multiple aspects, including user charging costs, grid frequency stability support, and peak-shaving resource utilization. Based on the expected results from multiple aspects, a preliminary set of charging and discharging operation schemes is generated using the action generation unit in the reinforcement learning model. A feasibility check was performed on the initial set of charging and discharging operation schemes, and all feasible operation schemes that meet the requirements were selected. The future sustainable benefits of feasible operational schemes are predicted and compared, and the scheme with the best overall benefits is selected to obtain the final charging and discharging strategy instruction set.

3. The multi-energy collaborative control and management system for vehicle energy storage charging piles according to claim 2, characterized in that, The charging and discharging strategy instruction set includes power control instructions for charging piles, charging and discharging power instructions for energy storage units, and execution time nodes.

4. The multi-energy collaborative control and management system for vehicle energy storage charging piles according to claim 3, characterized in that, According to the charging and discharging strategy instruction set, the energy storage unit is controlled to execute corresponding instructions, and instruction execution characteristic data is collected synchronously, including: The charging and discharging strategy instruction set is analyzed, the specified power threshold parameters and time window parameters are extracted, and dynamic charging and discharging power allocation instructions for each energy storage unit are generated. Execute dynamic power allocation commands for charging and discharging to control each energy storage unit to perform charging and discharging operations according to the target power. During command execution, the actual charging and discharging rate of each energy storage unit, the real-time deviation data between the terminal voltage and the target voltage, and the command response delay data from command issuance to power response are collected in real time. Based on real-time voltage deviation data, the cumulative voltage offset of each energy storage unit within the complete time window is calculated. Multi-dimensional feature fusion is performed on the actual charge / discharge rate, command response delay data, and accumulated voltage offset to generate a dynamic feature vector characterizing the current command execution state.

5. The multi-energy collaborative control and management system for vehicle energy storage charging piles according to claim 4, characterized in that, Key feature parameters are extracted from the dynamic changes of the basic evaluation index set, and nonlinear coupling relationships are analyzed by establishing a correlation model between the parameters to generate dynamic compensation coefficients, including: Time series analysis was performed on the basic evaluation indicator set to identify the dynamic change patterns of each indicator; Based on the dynamic change pattern, the fluctuation range of user charging cost savings, the rate of change of the contribution of grid frequency stability, and the decay trend of peak-shaving resource savings are extracted as key feature parameters. By inputting key feature parameters into a pre-defined correlation model, the synergistic effect between cost savings and frequency stability contribution, as well as the constraint relationship between peak shaving benefits and charging demand fluctuations, can be analyzed. Based on the synergy and constraint analysis results obtained from the correlation model, the nonlinear coupling strength between each parameter is quantified, and a comprehensive dynamic compensation coefficient is calculated based on the nonlinear coupling strength.

6. The multi-energy collaborative control and management system for vehicle energy storage charging piles according to claim 5, characterized in that, The basic evaluation index set is calibrated in real time using dynamic compensation coefficients to obtain the final evaluation index set, including: Using the dynamic compensation coefficient as a weighting parameter, the basic evaluation index set is weighted and calibrated to generate a weighted calibration index set that includes the weighted cost saving index. The weighted cost-saving index is extracted from the weighted calibration index set. The compensation value is calculated based on the degree of decay of the peak-shaving resource saving benefits. The compensation value is then superimposed on the weighted cost-saving index to generate the compensated and corrected cost-saving index. All indicators in the weighted calibration indicator set, except for the weighted cost savings indicator, are integrated with the compensated and corrected cost savings indicator to obtain the final evaluation indicator set.

7. The multi-energy collaborative control and management system for vehicle energy storage charging piles according to claim 6, characterized in that, Based on the final evaluation index set, and by integrating updated grid status data and user behavior characteristics, the configuration parameters of the decision engine are adaptively adjusted to achieve coordinated control of multiple energy sources, including: The final evaluation index set is input into the parameter update unit of the decision engine, and the grid frequency fluctuation data and user charging behavior deviation feature data are collected and updated in real time. The final evaluation index set, real-time power grid frequency fluctuation data, and user charging behavior deviation feature data are integrated to form a dynamically adjusted input dataset. Based on dynamically adjusted input datasets, the change values ​​of weight coefficients used for multi-objective processing in the decision engine are calculated using gradient descent algorithm; Based on the changes in weight coefficients, the weight coefficients currently used in multi-objective processing are dynamically updated, an updated set of weight coefficients is generated, and the result is output to the instruction generation unit of the decision engine. The instruction generation unit applies the updated weight coefficient set and combines it with real-time status information to generate multi-energy collaborative scheduling instructions for the next control cycle, thereby realizing the collaborative control of multiple energy storage sources.

8. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the system as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the system as described in any one of claims 1 to 7.