A Vehicle-to-Grid Charging and Discharging Optimization Method and System Based on Dynamic Electricity Price

By constructing a prediction model of the degree of battery polarization and optimizing charging and discharging strategies, the problems of limited battery life and economic benefits in the existing technology are solved, and the effects of extended battery life and stable returns are achieved.

CN120156382BActive Publication Date: 2025-07-08SICHUAN SHUXING YOUCHUANG SAFETY TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510637058.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-08
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing charging and discharging strategy optimization methods are difficult to adapt to the real-time changes in dynamic electricity prices and user needs, and fail to accurately predict the degree of battery polarization, resulting in limited battery life and economic benefits.

Method used

By constructing a prediction model of the degree of battery polarization, combining dynamic electricity price and user demand data, a reinforcement learning algorithm is used to optimize the charging and discharging strategy, a genetic algorithm and electrochemical simulation model are used to extend the battery life, and economic benefits are maximized through linear planning algorithms.

Benefits of technology

It achieves the extension of battery life and maximizes economic benefits, improves energy utilization efficiency, and maintains profit stability in a dynamic environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120156382B_ABST
    Figure CN120156382B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle-to-grid charging and discharging optimization method and system based on dynamic electricity prices, which relates to the technical fields of smart grids and electric vehicles. The method includes: S1, constructing a prediction model of battery polarization degree through an electrochemical model and historical operation data to obtain a quantization result of the battery polarization degree; S2, according to the quantization result of the battery polarization degree, combining dynamic electricity price data and user demand data, and adopting a reinforcement learning algorithm to construct a charging and discharging strategy optimization framework, outputting a preliminary charging and discharging strategy, determining the adaptability of the strategy in different scenarios, and obtaining a smooth charging and discharging curve. The vehicle-to-grid charging and discharging optimization method and system based on dynamic electricity prices realize the intelligent management of the battery charging and discharging process, effectively extend the battery life, improve the energy utilization efficiency, and maximize the economic benefits, providing technical support for the efficient operation of the battery energy storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of smart grid and electric vehicle, and particularly relates to a vehicle-grid charging and discharging optimization method and system based on dynamic electricity price. Background Art

[0002] Battery energy storage systems play a key role in energy transformation and smart grids, and their performance optimization directly affects energy utilization efficiency and economic benefits. With the rapid development of renewable energy and the popularization of dynamic electricity price mechanisms, battery energy storage systems need to operate efficiently in complex usage scenarios while taking into account battery life and economic returns.

[0003] However, existing optimization methods for charging and discharging strategies have significant limitations. Many traditional methods rely on static models, making it difficult to adapt to the real-time changes in dynamic electricity prices and user demands, and often ignoring the impact of battery polarization effects on battery life, resulting in limited effectiveness of the strategies in practical applications. In addition, some optimization algorithms overly simplify the charging and discharging curves, failing to fully consider the complexity of the internal electrochemical characteristics of the battery, causing losses in energy efficiency and economic benefits. In this field, the core challenges focus on the following technical factors: First, it is difficult to accurately predict the degree of battery polarization under different charging and discharging strategies, which directly affects the charging and discharging efficiency and battery life; Second, optimizing the charging and discharging curves requires finding a balance between dynamic electricity prices and user demands, but existing methods are difficult to achieve real-time adaptation; Finally, how to maximize economic benefits while extending battery life involves complex trade-offs in multi-objective optimization. These technical factors have not been effectively solved, resulting in limited performance of battery energy storage systems in dynamic scenarios and the economic potential not being fully released. Summary of the Invention

[0004] The purpose of the present invention is to provide a vehicle-grid charging and discharging optimization method and system based on dynamic electricity price, which accurately predicts the degree of battery polarization under different charging and discharging strategies, optimizes the charging and discharging curves to adapt to dynamic electricity prices and user demands, and maximizes economic benefits while extending battery life.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A vehicle-grid charging and discharging optimization method based on dynamic electricity price, the method comprising:

[0006] S1. Construct a prediction model of the degree of battery polarization through an electrochemical model and historical operation data, and obtain a quantitative result of the degree of battery polarization including dynamic adaptation;

[0007] S2. According to the quantitative result of the degree of battery polarization, combined with dynamic electricity price data and user demand data, adopt a reinforcement learning algorithm to construct a charging and discharging strategy optimization framework, output a preliminary charging and discharging strategy, determine the adaptability of the strategy in different scenarios, and obtain a smooth charging and discharging curve;

[0008] S3. If the energy efficiency of the optimized charge-discharge curve during simulated operation is lower than the preset threshold, then by adjusting the mutation rate and crossover rate of the genetic algorithm, regenerate the charge-discharge curve and output the updated charge-discharge curve;

[0009] S4. According to the updated charge-discharge curve, calculate the capacity attenuation rate of the battery at different cycle numbers. Using an electrochemical simulation model, input including curve parameters and polarization degree, output the predicted battery life value, and obtain the quantitative result of life extension;

[0010] S5. Extract the capacity attenuation trend from the predicted battery life value, combine with dynamic electricity price data, use the linear programming algorithm to optimize the economic benefit objective function, output the optimal charge-discharge scheduling plan, and determine the result of maximizing economic benefits;

[0011] S6. If the revenue fluctuation of the optimal charge-discharge scheduling plan during real-time operation exceeds the preset threshold, then by real-time collecting electricity price and demand data, adjust the reward function of the reinforcement learning algorithm, regenerate the scheduling plan, output the dynamically adapted scheduling plan, and obtain the result of improved revenue stability.

[0012] Preferably, after constructing the prediction model of the battery polarization degree in S1, use the recurrent neural network algorithm. Input including charge-discharge rate, temperature, and cycle number, output the predicted polarization degree value, and obtain the quantitative result of the battery polarization degree.

[0013] Preferably, obtaining the smooth charge-discharge curve in S2 specifically includes extracting key parameters from the preliminary charge-discharge strategy, using the genetic algorithm to optimize the charge-discharge curve. Input including the predicted polarization degree value and strategy parameters, output the optimized charge-discharge curve, and obtain a smooth and efficient charge-discharge curve.

[0014] Preferably, it further includes updating the battery operation parameters according to the dynamically adapted scheduling plan. The specific steps are as follows:

[0015] Collect real-time scheduling execution data and environmental variable data. Among them, the real-time scheduling execution data includes charge-discharge rate, charge-discharge start time, and charge-discharge end time, and the environmental variable data includes temperature and user load;

[0016] After each scheduling cycle is completed, feedback the operation state data of the battery. The operation state data includes the current polarization voltage, actual SOC change rate, and estimated life attenuation rate;

[0017] Store the structured operation log. The structured operation log includes various input data within the time window, and records each scheduling decision and its feedback polarization response, energy efficiency, and economic benefit evaluation;

[0018] After the cumulative operation reaches a certain time, update the historical dataset of the model, append the new data to the historical dataset, and construct a time window rolling training sample;

[0019] When the observation error exceeds the set threshold, trigger incremental training and retrain using the new samples.

[0020] Preferably, it further includes regularly training a recurrent neural network according to the updated historical dataset and adjusting the model parameters. The specific steps are as follows:

[0021] Input the updated historical dataset, where the historical dataset contains structured time series samples, and the time series samples include input features and true polarization value labels;

[0022] The system triggers a model training task every time it runs full of the specified period;

[0023] Use the latest dataset to replace or append to the original training set, adopt an adaptive learning rate strategy to improve the convergence efficiency, and the training objective is to minimize the prediction error;

[0024] Update the model parameters by evaluating the performance of the new model on the validation set.

[0025] Preferably, the prediction model of the battery polarization degree constructed in S1 uses a controlled voltage experiment to obtain the change curve of the polarization voltage and time, and combines the equivalent circuit model parameter fitting method to realize the real-time modeling of the polarization state.

[0026] Preferably, the reinforcement learning algorithm in S2 is the double deep Q network in deep reinforcement learning. Its state space includes the current SOC, electricity price gradient, polarization degree, and user travel time window, and the action space includes the charge and discharge power values and durations.

[0027] Preferably, the adaptability evaluation indicators of the preliminary charge and discharge strategy in S2 include three items: energy efficiency indicator, regulation frequency smoothness indicator, and user demand satisfaction rate, which are used as the criteria for judging the quality of the strategy.

[0028] Preferably, the mutation rate of the genetic algorithm in S3 is increased from the initial value of 0.05 to 0.1 when the simulation runs fail, and the crossover rate is dynamically adjusted between 0.7 and 0.9 to improve the diversity of the search space and avoid local optima.

[0029] A vehicle-to-grid charge and discharge optimization system based on dynamic electricity price is used to implement the steps of the vehicle-to-grid charge and discharge optimization method based on dynamic electricity price. The system includes:

[0030] A polarization prediction module, which is used to construct a prediction model of the battery polarization degree through an electrochemical model and historical operation data, and output the quantization result of the battery polarization degree;

[0031] A strategy optimization module, which is used to construct a charging and discharging strategy optimization framework by using a reinforcement learning algorithm according to the battery polarization degree quantization result output by the polarization prediction module, in combination with dynamic electricity price data and user demand data, generate a preliminary charging and discharging strategy, and determine the adaptability of the strategy in different scenarios, so as to output a smooth charging and discharging curve;

[0032] An energy efficiency evaluation and genetic optimization module, which is used to judge whether the energy efficiency of the charging and discharging curve during simulated operation is lower than a preset threshold. If it is lower than the threshold, the charging and discharging curve is regenerated by adjusting the mutation rate and crossover rate of the genetic algorithm, and the updated charging and discharging curve is output;

[0033] A battery life evaluation module, which is used to calculate the capacity attenuation rate of the battery at different cycle times in combination with an electrochemical simulation model according to the updated charging and discharging curve, and takes the charging and discharging curve parameters and the polarization degree as inputs, outputs a battery life prediction value, and obtains a quantified result of life extension;

[0034] A scheduling optimization module, which is used to extract the capacity attenuation trend according to the life prediction value, in combination with dynamic electricity price data, optimize the economic benefit objective function by using a linear programming algorithm, output an optimal charging and discharging scheduling plan, and determine the result of maximizing economic benefits;

[0035] A dynamic scheduling update module, which is used to adjust the reward function of the reinforcement learning algorithm in the strategy optimization module and regenerate a scheduling plan when the revenue fluctuation of the optimal charging and discharging scheduling plan during real-time operation exceeds a preset threshold, so as to output a dynamically adapted scheduling plan and improve revenue stability.

[0036] As can be seen from the above technical solutions, the present invention has the following beneficial effects:

[0037] The vehicle-to-grid charging and discharging optimization method and system based on dynamic electricity price constructs a battery polarization degree prediction model through an electrochemical model and historical data, optimizes the charging and discharging strategy in combination with dynamic electricity price and user demand, and uses a genetic algorithm to optimize the charging and discharging curve. The present invention also predicts the battery life according to the optimized charging and discharging curve and formulates an optimal charging and discharging scheduling plan by using a linear programming algorithm. During real-time operation, the present invention can dynamically adjust the scheduling plan to improve revenue stability and continuously update the historical data set to optimize the prediction model. This method realizes the intelligent management of the battery charging and discharging process, effectively extends the battery life, improves the energy utilization efficiency, and maximizes the economic benefits, providing technical support for the efficient operation of the battery energy storage system. Description of the Drawings

[0038] Figure 1 It is a flow chart of the method of the present invention;

[0039] Figure 2 This is the connection diagram of the system modules of the present invention. Specific embodiments

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0041] As Figure 1 shown, the present invention provides a technical solution: a method for optimizing the charging and discharging of a vehicle-to-grid based on dynamic electricity prices, the method comprising:

[0042] S1. Construct a prediction model for the degree of battery polarization through an electrochemical model and historical operation data, and obtain a quantitative result of the degree of battery polarization;

[0043] S2. According to the quantitative result of the degree of battery polarization, combine dynamic electricity price data and user demand data, and adopt a reinforcement learning algorithm to construct an optimization framework for charging and discharging strategies, output a preliminary charging and discharging strategy, determine the adaptability of the strategy in different scenarios, and obtain a smooth charging and discharging curve;

[0044] S3. If the energy efficiency of the optimized charging and discharging curve in the simulated operation is lower than a preset threshold, then by adjusting the mutation rate and crossover rate of the genetic algorithm, regenerate the charging and discharging curve, and output the updated charging and discharging curve;

[0045] S4. According to the updated charging and discharging curve, calculate the capacity attenuation rate of the battery at different cycle numbers, adopt an electrochemical simulation model, input including curve parameters and polarization degree, output a battery life prediction value, and obtain a quantitative result of life extension;

[0046] S5. Extract the capacity attenuation trend from the battery life prediction value, combine dynamic electricity price data, and adopt a linear programming algorithm to optimize the economic benefit objective function, output an optimal charging and discharging scheduling plan, and determine the result of maximizing economic benefits;

[0047] S6. If the revenue fluctuation of the optimal charging and discharging scheduling plan in the real-time operation exceeds a preset threshold, then by real-time collecting electricity price and demand data, adjust the reward function of the reinforcement learning algorithm, regenerate the scheduling plan, output a dynamically adapted scheduling plan, and obtain a result of improving revenue stability.

[0048] In this embodiment, first, a model capable of predicting the degree of battery polarization is constructed through electrochemistry modeling technology and historical data analysis. This model takes parameters such as temperature, current, and voltage during battery operation as inputs and outputs a quantitative polarization degree index through machine learning algorithms. Further, the system uses this prediction result as a key input parameter for charge-discharge strategy design, and combines dynamic electricity prices and user electricity demand, and applies reinforcement learning algorithms to output a preliminary charge-discharge strategy. The adaptability of this strategy in various operating scenarios is verified through simulation to ensure its smoothness and feasibility. If the energy efficiency shown by the curve generated by this preliminary strategy does not reach the threshold set by the system during simulation, a genetic algorithm is enabled to adjust the mutation rate and crossover rate to improve the optimization ability and generate a new charge-discharge strategy. The newly generated strategy is used to analyze the capacity decay trend of the battery under multiple cycle charge-discharge conditions. By taking the curve parameters and the degree of polarization as inputs through an electrochemistry simulation model, a predicted value of the battery life is obtained, thereby quantifying the positive effect of the strategy on extending the battery life. Then, based on the life prediction trend and electricity price change information, a linear programming model is used to optimize the economic objective function, and an optimal charge-discharge scheduling plan under current conditions is output. If the actual revenue fluctuates by more than the preset threshold during the real-time implementation of this plan, the system dynamically adjusts the reward function weight of reinforcement learning by collecting the latest electricity price and user load demand in real time to respond to market changes, and finally realizes the dynamic adaptive optimization of the scheduling scheme.

[0049] This embodiment uses an electrochemistry modeling method to predict the degree of battery polarization. The polarization voltage As a key quantitative indicator of battery health, the calculation formula is as follows:

[0050] ;

[0051] Where is the open-circuit voltage, measured through a static experiment; is the terminal voltage, from vehicle operation data or real-time collection by sensors.

[0052] The degree of battery polarization Can be normalized as:

[0053] ;

[0054] In step S2, reinforcement learning uses the Q-learning algorithm based on temporal difference, and the policy value function is:

[0055] ;

[0056] Where is the current state, including battery SOC, electricity price, and user load, is an action, i.e., a charging or discharging decision, is an immediate reward, considering both economic benefits and battery life, is the learning rate, usually set between 0.01 - 0.1, is the discount factor, reflecting the importance of future rewards, often taking values between 0.9 - 0.99, is to execute the action and reach the new state, is the new state and all possible actions in it have the maximum Q - value.

[0057] This Q - learning algorithm based on temporal difference is a traditional Q - learning, based on a tabular or simple approximate Q - function, used to calculate the value corresponding to the state and action. Its update method is that the value of the current state and action is updated according to the immediate reward plus the future optimal expected return. The Q - learning formula represents the update method of the theoretical algorithm, indicating that the principle of Q - learning based on temporal difference is used.

[0058] In step S3, the genetic algorithm is used, and its fitness function is:

[0059] ;

[0060] where f(x) represents the fitness function, and x represents the set of charge - discharge decision parameters formulated for the battery S0C, electricity price, and user load status; is the energy - efficiency function, measuring the unit - energy revenue, is the life - extension rate, is the weight factor, set according to the actual operation goal, usually = 1;

[0061] The capacity attenuation rate is modeled as:

[0062] ;

[0063] where C loss (n) is the capacity attenuation rate, is the initial capacity, is the number of charge - discharge cycles, is the capacity attenuation coefficient, obtained by experimental regression.

[0064] The economic - benefit function adopts the form of linear - programming optimization:

[0065] ;

[0066] where, is the electricity price at time , is the charge - discharge electricity quantity, The loss cost caused by battery usage, T represents the total number of time periods within the entire optimization cycle, and t represents the index of the specific time step (moment) within the optimization cycle.

[0067] The real-time adaptability adjusts the reward function as follows:

[0068] ;

[0069] Among them, r t represents the value of the reward function after real-time adjustment, is the total revenue at time t, is the revenue fluctuation, is the weight adjustment parameter, which is set according to user preferences or operation strategies.

[0070] This method integrates multi-source data and intelligent optimization algorithms, enhancing the intelligence and response ability of the vehicle-grid interaction system. By introducing a battery polarization prediction and life simulation model, the battery service life is significantly extended; while dynamically optimizing the charge and discharge strategy and combining with a linear programming objective function, the economy of the system under fluctuating electricity prices is effectively enhanced. The reinforcement learning mechanism ensures the stability of the scheduling scheme and maximizes the revenue through real-time adjustment of the reward function, and has high adaptability and robustness.

[0071] Deploy the system of the present invention at a large commercial vehicle charging station, and establish a model using the time-of-use electricity price provided by the State Grid and taxi operation data. Through the polarization prediction model and reinforcement learning strategy trained by historical data, charging is automatically started during the night valley period and intelligent discharging is performed during the day peak period, and the strategy is fine-tuned using the genetic algorithm. Within three months, the system optimized the charge and discharge cycle 240 times, the battery capacity attenuation rate was reduced by about 18%, the monthly energy cost was saved by 21%, and the battery life was extended by more than 15% as expected. This system can still maintain stable revenue after adapting to the load change in sudden high-temperature weather, reflecting its high practicality and commercial value.

[0072] After constructing the prediction model of the battery polarization degree in S1, the recurrent neural network algorithm is adopted. The input includes the charge and discharge rate, temperature, and number of cycles, and the output is the predicted value of the polarization degree, obtaining the quantization result of the battery polarization degree.

[0073] In this embodiment, the polarization degree prediction model takes the recurrent neural network (RNN) as the core and is specifically used to process time series data. The model structure includes an input layer, multiple hidden layers, and an output layer. The input data dimensions include:

[0074] The charge and discharge rate , (unit: C-rate);

[0075] The battery operating temperature , (unit: °C);

[0076] Cumulative number of cycles (unit: times).

[0077] The input vector is represented as: ;

[0078] The state update formula of the RNN is:

[0079] ;

[0080] The output layer predicts the polarization degree value at the current time point :

[0081] ;

[0082] Among them, is the hidden state at the moment, are the input weight matrix, state transition matrix, and output matrix, is the bias term.

[0083] The model training uses the mean squared error loss function:

[0084] ;

[0085] L represents the mean squared error loss function. The training samples come from historical charge-discharge cycle records and measured polarization voltage data, and the data volume is not less than 5000 groups to ensure the generalization ability and prediction accuracy of the model.

[0086] In terms of parameter settings, it is recommended that the hidden layer size be 64 - 128 units, the optimization algorithm use Adam, the initial learning rate be set to 0.001, the number of training epochs be not less than 200 epochs, and the training set and validation set be divided in a ratio of 8:2.

[0087] The polarization degree quantization value output by this model is directly used for loss prediction and life assessment in the charge-discharge strategy, significantly improving the technical support accuracy of the scheduling strategy.

[0088] By introducing the RNN structure to process the time series data of charge-discharge rate, temperature, and number of cycles, the nonlinear evolution law of the battery state can be effectively captured. Compared with traditional linear modeling methods, the prediction accuracy of this method is significantly improved, which can early warn the deterioration trend of battery polarization and enhance the battery health management ability. At the same time, due to the strong learning ability of RNN for time series features, it is more adaptable in a dynamic environment, providing high-precision input support for subsequent optimization strategies, effectively extending the battery life and ensuring the stable operation of the system.

[0089] Obtaining a smooth charge-discharge curve in S2 specifically includes extracting key parameters from the preliminary charge-discharge strategy, optimizing the charge-discharge curve using a genetic algorithm. The input includes the predicted polarization degree and strategy parameters, and the output is the optimized charge-discharge curve, obtaining a smooth and efficient charge-discharge curve.

[0090] In this embodiment, the generation of the smooth charge-discharge curve is based on a genetic algorithm for non-linear multi-objective optimization. Its input includes the predicted polarization degree output by the RNN model in step S1 , and the key charge-discharge parameter set determined by the preliminary reinforcement learning strategy , where:

[0091] is the charge-discharge rate at time , is the total daily demand energy, is the allowed charge-discharge time window.

[0092] The form of the optimization objective function is:

[0093] ;

[0094] where: the first term is the curve smoothness term (the sum of the squares of the first derivatives of the charge-discharge rate), and the second term is the polarization suppression term (to avoid high-rate operations when the polarization degree is high), are the weights of smoothness and energy consumption penalty respectively, and the typical value range is 0.3 - 0.7.

[0095] The genetic algorithm performs optimization through the following steps:

[0096] Initialize the population, and each chromosome represents a set of charge-discharge rate sequences ;

[0097] Calculate the fitness of each individual according to the above objective function;

[0098] Generate a new generation through selection (roulette wheel), crossover (single-point / multi-point) and mutation (Gaussian perturbation);

[0099] Repeat the iteration until the target converges or reaches the maximum number of generations.

[0100] The constraint conditions include: satisfying the charge balance constraint ; r t represents the charge-discharge power at time step t, Δt represents the length of each time step t, and dt represents the cumulative charge change;

[0101] Ensure the rate boundary ;

[0102] r t represents the reward function value at time t, rmax Denote the maximum value of the reward function, \(r\). min Denote the minimum value of the reward function. The output charge and discharge curve, while meeting the total power demand, maximally smooths the curve shape and avoids the polarization deterioration region, improving the overall energy efficiency.

[0103] By taking the preliminary strategy parameters and the polarization prediction results as the optimization inputs of the genetic algorithm, this method realizes the dual optimization of the charge and discharge curve at the physical level and the strategy level. On the premise of ensuring the energy delivery target, it greatly improves the smoothness of the charge and discharge rate, reduces the peak power, slows down the battery polarization effect, and extends its service life. This optimization mechanism has strong adaptability and can be deployed in different vehicle models, battery types, and user scenarios, with broad engineering application value.

[0104] It also includes updating the battery operation parameters according to the dynamically adapted scheduling scheme. The specific steps are as follows:

[0105] Collect real-time scheduling execution data and environmental variable data. Among them, the real-time scheduling execution data includes the charge and discharge rate, the start time of charge and discharge, and the end time of charge and discharge, and the environmental variable data includes temperature and user load;

[0106] After each scheduling cycle is completed, feedback the battery operation status data, and the operation status data includes the current polarization voltage, the actual SOC change rate, and the estimated value of the life attenuation rate;

[0107] Store the structured operation log, and the structured operation log includes various input data within the time window, and records each scheduling decision and its feedback polarization response, energy efficiency, and economic benefit evaluation;

[0108] After the cumulative operation reaches a certain time, update the historical data set of the model, append the new data to the historical data set, and construct a time window rolling training sample;

[0109] When the observation error exceeds the set threshold, trigger incremental training and retrain with the new samples.

[0110] In this embodiment, the system adds a data closed-loop feedback module, which is mainly composed of two parts: a "data recording module" and a "model update trigger", and is used to support the dynamic continuous optimization of the prediction model. The system framework is as follows:

[0111] Data acquisition input: including real-time scheduling execution data (such as real-time charge and discharge rate , the start time of charge and discharge , the end time ) and environmental variables (such as temperature , user load );

[0112] Operating status feedback: After each scheduling cycle is completed, the BMS feeds back the recorded battery operating status data, including:

[0113] Current polarization voltage ;

[0114] Actual SOC change rate ;

[0115] Estimated value of life attenuation rate ;

[0116] Function of data recording module:

[0117] Store structured operation logs , where are various input data within the time window, H represents storing structured operation logs, and X t represents the set of input data at time t of the time window (window);

[0118] Record each scheduling decision and its feedback polarization response, energy efficiency, and economic benefit evaluation;

[0119] Data update mechanism: The system performs model dataset update every 24 hours or after the cumulative operation reaches a certain mileage; new data is appended to the historical dataset to construct a time window rolling training sample;

[0120] Predictive model adaptive update mechanism: If the observation error exceeds the threshold , then incremental training is triggered;

[0121] Use the new sample for retraining to maintain the model accuracy and adaptability.

[0122] Among them, represents the estimated value of the predictive strategy parameters at time t, ΔD represents the newly added dataset, represents the actual strategy value at time t.

[0123] The system ensures that through the data recording and feedback mechanism, an intelligent closed-loop optimization system is constructed to dynamically adapt to the long-term changes of vehicles, batteries, and the external environment.

[0124] By establishing a data-driven continuous optimization mechanism, the present invention not only has high accuracy in the initial stage of the prediction model, but also can continuously absorb actual feedback during the long-term operation process to correct the model accuracy. As the core component of the intelligent scheduling system, the data recording module not only improves the prediction accuracy, but also supports advanced functions such as long-term maintenance plans, fault prediction, and operation optimization. The dynamic evolution ability of the system greatly enhances its scalability and reliability, adapting to different operation cycles and operation strategy adjustment requirements.

[0125] It also includes regularly training a recurrent neural network according to the updated historical data set and adjusting the model parameters. The specific steps are as follows:

[0126] Input the updated historical data set, where the historical data set contains structured time series samples, and the time series samples include input features and true polarization value labels;

[0127] The system triggers a model training task every time it runs for a specified period.

[0128] Use the latest data set to replace or append to the original training set, adopt an adaptive learning rate strategy to improve the convergence efficiency, and the training objective is to minimize the prediction error;

[0129] By evaluating the performance of the new model on the validation set, update the model parameters. Based on the completion of the closed-loop of scheduling and data recording, this embodiment introduces a regular model update mechanism to maintain the accuracy and adaptability of the battery polarization degree prediction model. This mechanism mainly includes the following technical processes:

[0130] 1. Historical data set construction: The historical data set generated by the data recording module , contains structured time series samples:

[0131] ;

[0132] It includes input features and true polarization value labels.

[0133] 2. Regular scheduling mechanism: The system triggers a model training task every time it runs for a specified period (such as daily, every 100 schedules, or a cumulative of 1GB of new data).

[0134] 3. RNN model retraining: The model structure remains the same, using a multi-layer RNN or LSTM; use the latest data set to replace or append to the original training set;

[0135] Adopt an adaptive learning rate strategy (such as AdaGrad, Adam) to improve the convergence efficiency;

[0136] The model training objective is to minimize the prediction error:

[0137] ;

[0138] 4. Model Evaluation and Selection: Evaluate the performance of the new model on the validation set through cross - validation or the hold - out method; only replace the currently deployed model if the mean squared error (MSE) of the new model is better than that of the current model; update the model parameters including input weights , state transition matrix , output weights and bias terms .

[0139] 5. Deployment and Iteration: The optimized model is automatically deployed to the vehicle's local BMS or the cloud platform of the dispatching center to ensure that the prediction model is always in the optimal state.

[0140] By regularly iteratively training the RNN model, the method of the present invention can adapt to the long - term changes in the battery operating environment, grid dynamics, and user behavior, and achieve dynamic self - adaptation of battery polarization prediction. This mechanism significantly reduces the prediction deviation problem caused by model aging, improves the strategy accuracy, reduces the system misoperation rate, and ensures the scientificity and robustness of dispatching. Compared with the one - time modeling method, it has better continuous and stable performance and is suitable for long - term intelligent energy management systems.

[0141] The prediction model of the battery polarization degree constructed in S1 uses a controlled voltage experiment to obtain the change curve of polarization voltage and time, and combines the equivalent circuit model parameter fitting method to achieve real - time modeling of the polarization state.

[0142] In this embodiment, the system uses a controlled voltage charge - discharge experiment to obtain battery polarization characteristic data, and combines an equivalent circuit model for parameter fitting to achieve real - time prediction modeling of the battery polarization state. Specifically, it includes the following steps:

[0143] 1. Controlled Voltage Experiment Design:

[0144] The experiment is set to discharge the battery at multiple different constant voltage points (such as 3.6V, 3.8V, 4.0V);

[0145] During the experiment, collect current , terminal voltage , and the internal temperature of the battery ;

[0146] The polarization voltage is defined as: , represents the polarization voltage at time , represents the set voltage;

[0147] 2. Construction of Polarization Change Curve:

[0148] Integrate the data at different voltage points to generate a three-dimensional change surface of voltage-polarization-time; Integrate the data at different voltage points to generate a three-dimensional change surface of voltage-polarization-time;

[0149] Interpolate and smooth each charge-discharge process to remove high-frequency noise.

[0150] 3. Equivalent circuit model construction:

[0151] Adopt a typical first-order or second-order model of parallel R-C:

[0152] ;

[0153] where , is the polarization time constant;

[0154] Use the least squares method for parameter fitting to obtain the optimal estimated value of;

[0155] Update the above parameters in real time, and dynamic modeling of the polarization state can be carried out during online operation.

[0156] 4. Online application of the model:

[0157] During actual operation, input the current collected in real time and the current parameters into the model for calculation ;

[0158] The system combines this result as a polarization prediction index and provides it to the scheduling optimization algorithm for decision support.

[0159] By introducing an electrochemistry physical model constructed based on experimental data, the present invention enhances the modeling accuracy of the battery polarization state. The equivalent circuit model can effectively characterize the time dependence of the polarization process and support dynamic update, with strong real-time performance and generalization ability. Compared with the data-driven black-box prediction model, this method has higher physical interpretability and model transparency, and is suitable for application in the vehicle-grid interaction system with high safety requirements. The parameter fitting method supported by experiments also enables the model to have good personalized expansion ability and can adapt to different types of battery systems.

[0160] In S2, the reinforcement learning algorithm is the double deep Q network in deep reinforcement learning. Its state space includes the current SOC, electricity price gradient, polarization degree, and user travel time window, and the action space includes the charge-discharge power value and duration.

[0161] The double deep Q-network is a deep reinforcement learning method that approximates the Q-function with a neural network, solving the problem that traditional Q-learning cannot be represented by a table in a high-dimensional space (when the state space and action space are large). The double deep Q-network is an improved version of Q-learning that uses two neural networks to prevent overestimation problems. The double deep Q-network is a specific implementation method that uses a deep neural network to replace the table to represent and update Q-values.

[0162] In this embodiment, to achieve efficient and stable policy learning, the double deep Q-network algorithm is introduced. This method effectively alleviates the overestimation problem existing in traditional DQN and improves the decision-making quality and policy robustness.

[0163] 1. Definition of state space

[0164] ;

[0165] is the current state of charge of the battery, is the electricity price gradient, reflecting the price change trend; is the current degree of polarization; is the travel time window available to the user.

[0166] 2. Definition of action space

[0167] ;

[0168] is the charge / discharge power corresponding to the current action; is the charge / discharge duration;

[0169] The action can be discretized into several gears. For example, the power is divided into 5 gears (-50kW to +50kW), and the duration is divided into 4 gears (15 - 60 minutes).

[0170] 3. Double DQN structure

[0171] Two neural networks: the main network and the target network ;

[0172] Update strategy: ;

[0173] ;

[0174] represents the target value at the current step t, is the action a executed at the t-th step t and the next state reached, represents in the new state and the optimal action a is found by the main network ​ For immediate reward, it combines electricity price revenue, battery loss, and polarization penalty; is the discount factor, usually set to 0.95; is the learning rate, generally taken as 0.001, θ is the parameter of the main network, and θ′ is the parameter of the target network.

[0175] 4. Reward function design

[0176] ;

[0177] is the reward value at the th moment, w1, w2, and w3 are weight parameters, is the th moment's revenue related to the electricity price, is the th moment's battery degradation loss, is the actual power currently executed;

[0178] Both profit drive and battery health protection are emphasized;

[0179] = 5:3:2 is a typical weight setting.

[0180] 5. Training mechanism

[0181] The experience replay mechanism is adopted to improve the sample utilization rate;

[0182] Soft update the target network weights every K steps, , where = 0.005.

[0183] This method comprehensively perceives the current status of the battery, electricity price changes, and user behavior characteristics by constructing a multi-dimensional state space, and effectively controls the stability and accuracy of the policy output by introducing a dual DQN structure. Compared with conventional Q-learning or DQN, dual DQN performs better in preventing overfitting and reducing policy fluctuations. Using the action space jointly composed of power and time can better adapt to the flexible needs of multiple types of vehicles and different charging station configurations.

[0184] The adaptability evaluation indicators of the preliminary charge and discharge strategy in S2 include three items: energy efficiency index, regulation frequency smoothness index, and user demand satisfaction rate, which are used as the criteria for judging the quality of the strategy.

[0185] In this embodiment, to achieve the scientific evaluation and feedback optimization of the preliminary charge and discharge strategy, three core adaptability evaluation indicators are introduced to quantitatively measure the comprehensive performance of the strategy in different scenarios:

[0186] 1. Energy efficiency index

[0187] ;

[0188] is the total energy output by the battery during discharge;

[0189] is the total corresponding charging input energy;

[0190] This index is used to measure the energy utilization efficiency of the system under the current strategy and reflects the loss situation during the conversion process;

[0191] Typical threshold: More than 85% is an efficient strategy.

[0192] 2. Regulation frequency smoothness index

[0193] Smoothness index ;

[0194] is the charging and discharging power at the t-th moment; This index measures the smoothness of power regulation. The smaller the value, the smoother the change, which is beneficial to extending the battery life and reducing the grid impact; In the evaluation criteria, usually less than 5kW / min is the ideal value.

[0195] 3. User demand satisfaction rate

[0196] Satisfaction rate ;

[0197] is the number of times the charging task is successfully completed within the user's reserved travel time window;

[0198] is the total number of times involving user requests in all scheduling cycles;

[0199] This index directly reflects the ability of the scheduling strategy to respond to the needs of end users and evaluates its acceptability from the user's perspective;

[0200] Usually, it needs to be maintained above 95%.

[0201] The above three indexes comprehensively evaluate the charging and discharging strategy from three dimensions: electric energy conversion, strategy dynamic performance and user experience, form a strategy scoring mechanism, and serve as an important reference for subsequent genetic optimization and reinforcement learning fine-tuning.

[0202] By introducing a scientific and comprehensive three-dimensional evaluation system, this method constructs a feedback closed-loop at the strategy generation stage, improving the convergence efficiency of the optimization path and the engineering reliability of the scheduling output. The energy efficiency index ensures the economy of the strategy, the regulation frequency smoothness restricts the dynamic stability of the strategy, and the user demand satisfaction rate guarantees the service quality, thus achieving a balance among performance, safety, and user satisfaction. The system can dynamically eliminate inferior strategies through strategy scoring, continuously promoting the level of intelligent operation.

[0203] In S3, the mutation rate of the genetic algorithm is increased from the initial value of 0.05 to 0.1 when the simulation runs fail, and the crossover rate is dynamically adjusted between 0.7 and 0.9 to increase the diversity of the search space and avoid local optima.

[0204] In this embodiment, to address the failure situations such as low energy efficiency in the simulation operation of the initial optimization strategy, the genetic algorithm introduces a dynamic parameter adjustment mechanism to enhance the algorithm's search ability and improve the globality and stability of the optimization strategy.

[0205] 1. Initial parameter setting:

[0206] Mutation rate = 0.05, used to introduce individual random perturbations and maintain population diversity;

[0207] Crossover rate = 0.8, controlling the proportion of the combination of information of two parent individuals to improve the local search ability.

[0208] 2. Judgment of simulation operation failure:

[0209] If the optimized charging and discharging strategy does not meet the set threshold in the evaluation indicators (such as energy efficiency, user satisfaction rate);

[0210] Or the individual fitness improvement is slow within N generations, indicating being trapped in a local optimum;

[0211] Trigger the parameter adjustment mechanism.

[0212] 3. Mutation rate adjustment mechanism:

[0213] Increase the mutation rate: ;

[0214] Initial is 0.05, increased to 0.1 when failed;

[0215] = 0.01 is the step size.

[0216] 4. Crossover rate dynamic adjustment mechanism:

[0217] According to the current population diversity or fitness variance Make adjustments:

[0218] , is expressed as the crossover probability;

[0219] is the fitness variance of the current population, is the maximum fitness variance;

[0220] The state adjustment range is controlled between [0.7, 0.9];

[0221] The purpose is to increase the crossover frequency when the population converges too quickly and enhance the ability to explore combined paths.

[0222] 5. Adjustment effect feedback mechanism

[0223] If the fitness is significantly improved within N generations after adjustment, restore the initial parameters to maintain convergence stability;

[0224] If there is no improvement, further increase it or switch to other optimization methods such as simulated annealing.

[0225] This method effectively expands the search range and avoids the problem of premature convergence of the algorithm by dynamically adjusting the key parameters of the genetic algorithm, especially when the optimization falls into a local optimum or the strategy quality does not meet the standard. The increase in the mutation rate enhances the random exploration ability of individuals, while the adjustment of the crossover rate optimizes the recombination efficiency of genetic information, enabling the search to achieve a better balance between globality and locality and significantly increasing the probability of the global optimal solution of the charge-discharge strategy.

[0226] Such as Figure 2 As shown, a vehicle-to-grid charge-discharge optimization system based on dynamic electricity prices is also provided for implementing the steps of the vehicle-to-grid charge-discharge optimization method based on dynamic electricity prices. The system includes:

[0227] A polarization prediction module for constructing a prediction model of the battery polarization degree through an electrochemical model and historical operation data and outputting a quantization result of the battery polarization degree;

[0228] A strategy optimization module for constructing a charge-discharge strategy optimization framework using a reinforcement learning algorithm, generating a preliminary charge-discharge strategy, and determining the adaptability of the strategy in different scenarios based on the quantization result of the battery polarization degree output by the polarization prediction module, in combination with dynamic electricity price data and user demand data, so as to output a smooth charge-discharge curve;

[0229] An energy efficiency evaluation and genetic optimization module for determining whether the energy efficiency of the charge-discharge curve in the simulated operation is lower than a preset threshold. If it is lower than the threshold, regenerate the charge-discharge curve by adjusting the mutation rate and crossover rate of the genetic algorithm and output the updated charge-discharge curve;

[0230] A battery life evaluation module, which is used to calculate the capacity attenuation rate of the battery at different cycle numbers according to the updated charge-discharge curve in combination with an electrochemical simulation model, take the charge-discharge curve parameters and the degree of polarization as inputs, output a battery life prediction value, and obtain a quantitative result of life extension;

[0231] A scheduling optimization module, which is used to extract the capacity attenuation trend according to the life prediction value, combine dynamic electricity price data, optimize the economic benefit objective function by using a linear programming algorithm, output an optimal charge-discharge scheduling plan, and determine the result of maximizing economic benefits;

[0232] A dynamic scheduling update module, which is used to, when the revenue fluctuation of the optimal charge-discharge scheduling plan during real-time operation exceeds a preset threshold, collect electricity price and user demand data in real time, adjust the reward function of the reinforcement learning algorithm in the strategy optimization module, and regenerate a scheduling plan, so as to output a dynamically adapted scheduling plan and improve revenue stability.

[0233] In this embodiment, the optimization system is constructed by the collaboration of multiple functional modules to form a closed-loop intelligent scheduling platform. The system operation process starts from polarization state perception and gradually progresses to policy generation, performance evaluation, life prediction, and economic scheduling optimization, and finally realizes real-time feedback regulation to ensure the operation efficiency and stability of the system. First, the polarization prediction module generates a quantization value of the current polarization degree of the battery based on the operation data collected by the vehicle and the results of battery modeling experiments. This result is used as the key basis for formulating the charge and discharge strategy and is provided to the policy optimization module. Next, based on the input of the current battery state, the trend of electricity price changes, and the user's travel demand, the policy optimization module uses the reinforcement learning algorithm to generate a preliminary charge and discharge strategy. The system verifies the adaptability performance of this strategy in multiple typical operation scenarios through the scheduling simulation platform and outputs the initial charge and discharge curve. The generated curve is then evaluated for performance by the energy efficiency evaluation module, which mainly focuses on its energy utilization efficiency, power change smoothness, and the ability to meet user needs. If the strategy fails to pass the evaluation, the genetic optimization module is activated to explore a better strategy combination in a larger search space by adjusting the core parameters of the genetic algorithm, namely the mutation rate and the crossover rate, and then outputs an improved charge and discharge curve. The battery life evaluation module uses the updated charge and discharge curve and combines it with the battery simulation model to evaluate the capacity attenuation trend of the battery under different charge and discharge cycle conditions. At the same time, this module also outputs a prediction result of the impact on the battery service life based on inputs such as the polarization state and power parameters to measure the positive effect of the current strategy on battery health management. Subsequently, the scheduling optimization module constructs an economic optimization framework based on the battery life prediction trend and future electricity price information, with the goal of formulating a scheduling plan that maximizes economic benefits under the premise of meeting energy and life conditions. This scheduling scheme is used as the main operation basis of the system and will be applied to actual charge and discharge control. During the actual operation process, if it is detected that the system revenue fluctuation exceeds the preset threshold, the dynamic scheduling update module will be triggered. This module will re-collect the latest electricity price and user demand information and dynamically adjust the reward mechanism in the policy optimization module to generate a new charge and discharge strategy to improve the response flexibility and revenue stability of the system. The entire system constructs a complete closed-loop process from data collection, policy learning to real-time optimization feedback. Through the close cooperation of each functional module, it realizes the refined management and intelligent control of the vehicle-to-grid charge and discharge process, and has strong adaptability, high reliability, and good engineering practicability.

[0234] Through modular design and information linkage mechanism, this system realizes the whole-process intelligent management from battery state perception, strategy formulation, performance evaluation, life prediction to economic benefit optimization. It has remarkable flexibility, adaptability and scalability. Its closed-loop feedback control ability ensures operation stability, and improves the scheduling effect through multi-algorithm integration (reinforcement learning + genetic algorithm + linear programming), ultimately significantly improving system energy efficiency, user satisfaction and equipment service life.

[0235] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A vehicle-to-grid charging and discharging optimization method based on dynamic electricity prices, characterized in that The method includes: S1. Construct a prediction model for the degree of battery polarization through an electrochemical model and historical operation data, and obtain a quantitative result of the degree of battery polarization; S2. According to the quantitative result of the degree of battery polarization, combined with dynamic electricity price data and user demand data, adopt a reinforcement learning algorithm to construct a charging and discharging strategy optimization framework, output a preliminary charging and discharging strategy, determine the adaptability of the strategy in different scenarios, and obtain a smooth charging and discharging curve; S3. If the energy efficiency of the optimized charging and discharging curve in the simulated operation is lower than the preset threshold, then by adjusting the mutation rate and crossover rate of the genetic algorithm, regenerate the charging and discharging curve, and output the updated charging and discharging curve; S4. According to the updated charging and discharging curve, calculate the capacity attenuation rate of the battery at different cycle numbers, adopt an electrochemical simulation model, input including curve parameters and the degree of polarization, output a battery life prediction value, and obtain a quantitative result of life extension; S5. Extract the capacity attenuation trend from the battery life prediction value, combined with dynamic electricity price data, adopt a linear programming algorithm to optimize the economic benefit objective function, output an optimal charging and discharging scheduling plan, and determine the result of maximizing economic benefits; S6. If the revenue fluctuation of the optimal charging and discharging scheduling plan in the real-time operation exceeds the preset threshold, then by collecting real-time electricity price and demand data, adjust the reward function of the reinforcement learning algorithm, regenerate the scheduling plan, output a dynamically adapted scheduling plan, and obtain a result of improved revenue stability.

2. The vehicle-to-grid charging and discharging optimization method based on dynamic electricity price according to claim 1, wherein: After constructing the prediction model for the degree of battery polarization in S1, adopt a recurrent neural network algorithm, input including charging and discharging rate, temperature, and cycle number, output a predicted value of the degree of polarization, and obtain a quantitative result of the degree of battery polarization.

3. The vehicle-to-grid charging and discharging optimization method based on dynamic electricity price according to claim 1, wherein: The obtaining of the smooth charging and discharging curve in S2 specifically includes extracting key parameters from the preliminary charging and discharging strategy, optimizing the charging and discharging curve using a genetic algorithm, input including the predicted value of the degree of polarization and strategy parameters, outputting the optimized charging and discharging curve, and obtaining a smooth and efficient charging and discharging curve.

4. The vehicle-to-grid charging and discharging optimization method based on dynamic electricity price according to claim 1, wherein It also includes updating the battery operation parameters according to the dynamically adapted scheduling plan, and the specific steps are as follows: Collect real-time scheduling execution data and environmental variable data, where the real-time scheduling execution data includes charging and discharging rate, charging and discharging start time, and charging and discharging end time, and the environmental variable data includes temperature and user load; After each scheduling cycle is completed, feedback the operation state data of the battery, and the operation state data includes the current polarization voltage, actual SOC change rate, and estimated value of the life attenuation rate; Store structured operation logs, and the structured operation logs include various input data within a time window, and record each scheduling decision and its feedback polarization response, energy efficiency, and economic benefit evaluation; After the cumulative operation reaches a certain time, update the historical data set of the model, append the new data to the historical data set, and construct a time window rolling training sample; When the observation error exceeds the set threshold, trigger incremental training and retrain using the new samples.

5. The vehicle-to-grid charging and discharging optimization method based on dynamic electricity price according to claim 4, characterized in that It also includes regularly training the recurrent neural network according to the updated historical data set and adjusting the model parameters, and the specific steps are as follows: Input the updated historical dataset, where the historical dataset contains structured time series samples, and the time series samples include input features and true polarization value labels; The system triggers a model training task every time a specified cycle is completed; Replace or append the latest dataset to the original training set, adopt an adaptive learning rate strategy to improve the convergence efficiency, and the training objective is to minimize the prediction error; Update the model parameters by evaluating the performance of the new model on the validation set.

6. The vehicle-to-grid charging and discharging optimization method based on dynamic electricity price according to claim 1, wherein: The prediction model for the degree of battery polarization constructed in S1 uses a controlled voltage experiment to obtain the change curve of polarization voltage and time, and combines the equivalent circuit model parameter fitting method to realize real-time modeling of the polarization state.

7. A vehicle-to-grid charging and discharging optimization method based on dynamic electricity prices according to claim 1, characterized in that: The reinforcement learning algorithm in S2 is the double deep Q network in deep reinforcement learning. Its state space includes the current SOC, electricity price gradient, degree of polarization, and user travel time window, and the action space includes the charge and discharge power values and durations.

8. A vehicle-to-grid charging and discharging optimization method based on dynamic electricity price according to claim 1, characterized in that: The adaptability evaluation indicators of the preliminary charge and discharge strategy in S2 include three items: energy efficiency indicator, regulation frequency smoothness indicator, and user demand satisfaction rate, which are used as the judgment criteria for the quality of the strategy.

9. A vehicle-to-grid charging and discharging optimization method based on dynamic electricity price according to claim 1, characterized in that: In S3, the mutation rate of the genetic algorithm is increased from the initial value of 0.05 to 0.1 when the simulation runs fail, and the crossover rate is dynamically adjusted between 0.7 and 0.9 to improve the diversity of the search space and avoid local optima.

10. A vehicle-to-grid charging and discharging optimization system based on dynamic electricity prices, which is used to implement the steps of the vehicle-to-grid charging and discharging optimization method according to any one of claims 1-9, and is characterized in that, The system includes: A polarization prediction module, which is used to construct a prediction model for the degree of battery polarization through an electrochemical model and historical operation data, and output the quantization result of the degree of battery polarization; A strategy optimization module, which is used to combine the quantization result of the degree of battery polarization output by the polarization prediction module, dynamic electricity price data and user demand data, and adopt a reinforcement learning algorithm to construct a charge and discharge strategy optimization framework, generate a preliminary charge and discharge strategy, and determine the adaptability of the strategy in different scenarios to output a smooth charge and discharge curve; An energy efficiency evaluation and genetic optimization module, which is used to judge whether the energy efficiency of the charge and discharge curve in the simulation operation is lower than a preset threshold. If it is lower than the threshold, regenerate the charge and discharge curve by adjusting the mutation rate and crossover rate of the genetic algorithm, and output the updated charge and discharge curve; A battery life evaluation module, which is used to calculate the capacity attenuation rate of the battery at different cycle times according to the updated charge and discharge curve, combined with an electrochemical simulation model, and take the charge and discharge curve parameters and the degree of polarization as inputs to output the battery life prediction value and obtain the quantization result of life extension; A scheduling optimization module, which is used to extract the capacity attenuation trend according to the life prediction value, combine dynamic electricity price data, and adopt a linear programming algorithm to optimize the economic benefit objective function, output the optimal charge and discharge scheduling plan, and determine the result of maximizing economic benefits; A dynamic scheduling update module, which is used to adjust the reward function of the reinforcement learning algorithm in the strategy optimization module and regenerate the scheduling plan when the revenue fluctuation of the optimal charge and discharge scheduling plan in real-time operation exceeds a preset threshold, so as to output a dynamically adapted scheduling plan and improve the revenue stability.

Citation Information

Patent Citations

  • Lithium ion battery quick charging method considering battery polarization degree

    CN113410530A

  • Intermittent charging control method and system for lithium battery pack

    CN119093540A