Super-computing center battery charging and discharging strategy optimization method and system based on reinforcement learning
Through the method based on reinforcement learning, the battery charging and discharging strategy of supercomputing centers is optimized in real time, and the problems of strategy lag and insufficient adaptability in the existing technology are solved, and the stability and life of the battery system are improved.
Patent Information
- Application Number
- CN202510837230.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The battery charging and discharging strategies of existing supercomputing centers cannot analyze the multi-parameter coupling effect in real time, resulting in overcharging or power distribution mismatch, affecting battery life and system stability. The fixed strategy library is difficult to adapt to load fluctuations, which easily triggers the system protection mechanism to interrupt power supply.
Using reinforcement learning-based method, real-time operation data of the battery system is obtained in real time, and charging and discharging strategies are optimized through load feature extraction and reinforcement learning strategy network, optimized charging and discharging strategies are generated, and the strategy parameters are iteratively updated through feedback data to achieve in-depth analysis and strategy optimization of current changes, voltage fluctuations and temperature responses.
It improves the generalization ability of charging and discharging strategies, reduces the risk of battery performance attenuation, improves the stability of system power supply, avoids overcharge and discharge and power oscillation problems, and achieves independent decision-making with multi-objective optimization.
Smart Images

Figure CN120357594A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning, and in particular, to a method and system for optimizing the battery charging and discharging strategy of a supercomputer center based on reinforcement learning. Background Art
[0002] With the exponential growth of the computing power demand of supercomputer centers, the optimization of the battery charging and discharging strategy of its power supply system has become a key link to ensure the stable operation of the system. Existing technologies usually adopt a charging and discharging trigger mechanism based on fixed thresholds, control the battery charging and discharging behavior by presetting current and voltage threshold values, or perform policy library matching selection under the drive of the load fluctuation amplitude. However, the battery system of the supercomputer center is in a scenario of the coupling effect of load and complex environment for a long time. There is a strong non-linear correlation between current transient impact, voltage fluctuation mode and temperature response characteristics. Traditional methods cannot resolve the multi-parameter coupling effect in real time, resulting in lag in policy generation and frequently causing problems such as overcharging / discharging or power distribution mismatch; the predefined rules of the fixed policy library are difficult to adapt to the evolution law of the load fluctuation characteristics, and there is a lack of feedback correction ability for the battery performance decay path after the policy is executed, resulting in an increase in battery cycle life loss; at the same time, the existing solutions rigidly bind the charging and discharging trigger conditions to power parameters, resulting in limited policy flexibility. When dealing with sudden load increase or environmental parameter mutation, it is easy to trigger the system protection mechanism to interrupt the power supply, seriously affecting the continuity of supercomputer tasks. Summary of the Invention
[0003] The present invention provides a method and system for optimizing the battery charging and discharging strategy of a supercomputer center based on reinforcement learning.
[0004] In a first aspect, an embodiment of the present invention provides a method for optimizing the battery charging and discharging strategy of a supercomputer center based on reinforcement learning, including: Obtaining a real-time operation data set of the battery system of the supercomputer center in real time, where the real-time operation data set includes a battery state parameter set, a load fluctuation parameter set, and an environmental monitoring parameter set; Performing load feature extraction processing on the real-time operation data set to generate a load feature set of the battery system; the load feature set includes a current change trend feature, a voltage fluctuation mode feature, and a temperature correlation response feature; Inputting the load feature set into a pre-trained reinforcement learning policy network, and performing charging and discharging strategy priority matching processing on the load feature set through the reinforcement learning policy network to generate a priority ranking result of the target charging and discharging strategy; Adjusting the current charging and discharging strategy set according to the priority ranking result to generate an optimized charging and discharging strategy set; each charging and discharging strategy in the optimized charging and discharging strategy set includes a charging and discharging trigger condition parameter and a power distribution parameter; Execute the charge and discharge operations based on the optimized charge and discharge strategy set, and collect the battery performance feedback data after the charge and discharge operations to iteratively update the policy matching parameters of the reinforcement learning policy network.
[0005] In a second aspect, an embodiment of the present invention provides a computer system, including: A memory in which a computer program is stored; A processor for loading the computer program to implement the method for optimizing the battery charge and discharge strategy of a supercomputer center based on reinforcement learning as described above.
[0006] The method for optimizing the battery charge and discharge strategy of a supercomputer center based on reinforcement learning provided by the present invention generates a load feature set by real-time fusing battery state parameters, load fluctuation parameters, and environmental monitoring parameters, comprehensively captures the current change law, voltage fluctuation characteristics, and temperature response mode during the operation of the battery, enables the reinforcement learning policy network to deeply analyze the electro-thermal-load coupling relationship and generate a policy priority ranking; optimizes the trigger condition parameters and power distribution parameters of the decoupled design in the optimized charge and discharge strategy set, realizes the collaborative optimization of the charge and discharge timing and execution intensity, effectively balances the load demand matching accuracy and the battery performance protection intensity; drives the iterative update of the policy network parameters by real-time collecting the performance feedback data after the policy execution, establishes an effective adaptation mechanism between policy generation and battery state changes, significantly improves the generalization ability of the charge and discharge strategy for complex load scenarios in the supercomputer center, reduces the risk of battery performance degradation while improving the power supply stability of the system, overcomes the overcharge / discharge and power oscillation problems caused by the lagged response of traditional fixed policies to environmental parameters, and can achieve autonomous decision-making for multi-objective optimization without relying on artificial experience rules. Description of the Drawings
[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0008] Figure 1 It is a flowchart of a method for optimizing the battery charge and discharge strategy of a supercomputer center based on reinforcement learning provided by an embodiment of the present invention.
[0009] Figure 2 It is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. Detailed Embodiments
[0010] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0011] Please refer to Figure 1 , Figure 1 which is a flowchart of an optimization method for the battery charge and discharge strategy of a supercomputer center based on reinforcement learning. This optimization method for the battery charge and discharge strategy of a supercomputer center based on reinforcement learning can be executed by a computer system, and this optimization method for the battery charge and discharge strategy of a supercomputer center based on reinforcement learning may include the following steps: Step S100: Real-time obtain the real-time operation data set of the supercomputer center battery system. The real-time operation data set includes a battery state parameter set, a load fluctuation parameter set, and an environmental monitoring parameter set.
[0012] The real-time operation data set refers to a series of data collected in real time during the operation of the supercomputer center battery system, which reflects the operation state of the system. The battery state parameter set is data describing the state of the battery itself, covering parameters such as the voltage, current, temperature, and state of charge of the battery. These parameters can reflect the current performance and health status of the battery. The load fluctuation parameter set is data on the load change situation of the supercomputer center, such as information on the power change and current fluctuation of the load, which reflects the dynamic change of the workload of the supercomputer center at different times. The environmental monitoring parameter set is data obtained by monitoring the environment where the battery system is located, including parameters such as environmental temperature, humidity, and air pressure. Environmental factors have an important impact on the performance and life of the battery.
[0013] In order to obtain these data in real time, a variety of sensors and monitoring devices can be used. For the battery state parameter set, a voltage sensor can be used to measure the voltage of the battery in real time, a current sensor to collect the current of the battery, a temperature sensor to monitor the temperature of the battery, and a state-of-charge monitoring device to obtain the state of charge of the battery. For the load fluctuation parameter set, the load power of the supercomputer center can be measured by a power sensor, and the change in load current can be collected by a current transformer. For the environmental monitoring parameter set, a temperature sensor can be used to measure the environmental temperature, a humidity sensor to obtain the environmental humidity, and a barometric pressure sensor to monitor the environmental air pressure. For example, in a supercomputer center, corresponding sensors can be installed at key parts of the battery system, and data can be collected at regular intervals and transmitted to the data processing center for storage and analysis.
[0014] Step S200: Perform load feature extraction processing on the real-time operation data set to generate a load feature set of the battery system; the load feature set includes current change trend features, voltage fluctuation pattern features, and temperature correlation response features.
[0015] The load feature extraction processing is to analyze and process the real-time operation data set to extract key features that can reflect the load characteristics of the battery system. The load feature set is a set used to describe the load characteristics of the battery system, where the current change trend feature reflects the change trend of the battery current over time, the voltage fluctuation pattern feature reflects the fluctuation law of the battery voltage, and the temperature correlation response feature represents the correlation between temperature changes and load fluctuations.
[0016] As an implementation, step S200, performing load feature extraction processing on the real-time operation data set to generate a load feature set of the battery system, may specifically include the following steps S210 to S240: Step S210: Perform time-domain trend analysis on the current parameter sequence in the battery state parameter set, extract the fluctuation period feature and peak interval distribution feature of the current parameter sequence, and perform linear superposition processing on the fluctuation period feature and the peak interval distribution feature to generate a current change trend feature.
[0017] The time-domain trend analysis is to analyze the current parameter sequence in the time domain to find its change law and trend. The fluctuation period feature refers to the periodic characteristics of the fluctuations in the current parameter sequence, reflecting the periodic changes in the current fluctuations. The peak interval distribution feature is the distribution of the time intervals between adjacent peaks in the current parameter sequence, reflecting the regularity of the occurrence of current peaks. The linear superposition processing is to combine the fluctuation period feature and the peak interval distribution feature according to a set linear relationship to obtain a comprehensive current change trend feature.
[0018] As an implementation, step S210, performing time-domain trend analysis on the current parameter sequence in the battery state parameter set, extracting the fluctuation period feature and peak interval distribution feature of the current parameter sequence, may specifically include the following steps S211 to S215: Step S211: Perform sliding window segmentation processing on the current parameter sequence to generate multiple current parameter subsequences.
[0019] The sliding window segmentation process divides the current parameter sequence according to the set window size and sliding step length to obtain multiple subsequences. The window size refers to the number of current data points contained in each subsequence, and the sliding step length refers to the distance the window moves each time. Through the sliding window segmentation process, a long current parameter sequence can be divided into multiple shorter subsequences, facilitating subsequent analysis and processing. For example, if the current parameter sequence has 1000 data points, the window size is set to 100, and the sliding step length is set to 10, then (1000 - 100) / 10 + 1 = 91 current parameter subsequences can be obtained.
[0020] Step S212: Perform local extreme value detection on each current parameter subsequence, and extract the peak point positions and the time interval data between adjacent peak points.
[0021] Local extreme value detection is to find the local maximum and minimum points in each current parameter subsequence. A peak point refers to a point where the current value reaches the maximum within a local range. Through local extreme value detection, the peak point positions in the current parameter subsequence can be determined, and the time interval data between adjacent peak points can be calculated. Specifically, the difference method can be used to calculate the difference between adjacent data points. When the difference changes from positive to negative, the corresponding point is the peak point. For example, for a current parameter subsequence [1, 3, 5, 4, 2, 6, 5, 3], the peak points 5 and 6 can be detected by the difference method, and the time interval between adjacent peak points is 3 data points.
[0022] Step S213: Statistically analyze the distribution histogram of all time interval data, and calculate the kurtosis coefficient and skewness coefficient of the distribution histogram.
[0023] The distribution histogram is used to display the data distribution. Group the time interval data of all adjacent peak points, count the frequency of data in each group, and then draw the histogram. The kurtosis coefficient is a statistic that describes the kurtosis of the data distribution, reflecting the steepness of the data distribution. The skewness coefficient is a statistic that describes the skewness of the data distribution, reflecting the asymmetry of the data distribution. The kurtosis coefficient and skewness coefficient can be calculated using relevant functions in statistical software or programming languages. For example, in Python, the kurtosis and skew functions in the scipy.stats library can be used to calculate the kurtosis coefficient and skewness coefficient.
[0024] Step S214: Generate the peak interval distribution characteristics based on the kurtosis coefficient and skewness coefficient.
[0025] The peak interval distribution feature is a feature vector constructed based on the kurtosis coefficient and skewness coefficient, used to describe the distribution characteristics of current peak intervals. The kurtosis coefficient and skewness coefficient can be used as two components of the feature vector to form the peak interval distribution feature. For example, if the kurtosis coefficient is 2.5 and the skewness coefficient is 0.8, the peak interval distribution feature can be expressed as [2.5, 0.8].
[0026] Step S215: Perform Fourier transform processing on the current parameter sequence, extract the main frequency components and their corresponding amplitudes, and generate a fluctuation period feature based on the reciprocals of the main frequency components.
[0027] Fourier transform processing converts the current parameter sequence from the time domain to the frequency domain to analyze its frequency components. The main frequency components refer to the frequency components with larger amplitudes in the frequency domain, corresponding to the main periods of current fluctuations. Through Fourier transform, the frequency spectrum of the current parameter sequence can be obtained, and the main frequency components and their corresponding amplitudes can be extracted from it. The fluctuation period feature is calculated based on the reciprocals of the main frequency components because frequency and period are reciprocal to each other. For example, if the main frequency component is 5 Hz, the fluctuation period feature is 1 / 5 = 0.2 seconds.
[0028] Step S220: Perform frequency domain conversion processing on the voltage parameter sequences in the voltage fluctuation parameter set, extract the energy ratio feature of the voltage fluctuation components with voltages lower than the first preset frequency and the noise components with voltages higher than the second preset frequency, and calculate the correlation degree between the energy ratio feature and the mean variance feature of the voltage parameter sequences to generate a voltage fluctuation pattern feature.
[0029] Frequency domain conversion processing converts the voltage parameter sequence from the time domain to the frequency domain, such as using Fourier transform. The first preset frequency and the second preset frequency are preset frequency thresholds used to divide the voltage fluctuation components and noise components (such as the frequency band division standard based on battery characteristics). The voltage fluctuation components refer to the part of the voltage lower than the first preset frequency, reflecting the low-frequency fluctuation characteristics of the voltage. The noise components refer to the part of the voltage higher than the second preset frequency, reflecting the high-frequency noise characteristics of the voltage. The energy ratio feature is the ratio of the energy of the voltage fluctuation components to the energy of the noise components, reflecting the relative magnitudes of the voltage fluctuations and noise. The mean variance feature is the mean and variance of the voltage parameter sequence, describing the average level and fluctuation degree of the voltage respectively. Correlation degree calculation is to perform correlation analysis on the energy ratio feature and the mean variance feature to generate a comprehensive voltage fluctuation pattern feature.
[0030] In specific implementation, first, perform Fourier transform on the voltage parameter sequence to obtain its spectrum. Then, according to the first preset frequency and the second preset frequency, divide the spectrum into a voltage fluctuation component and a noise component, and calculate their energies. Next, calculate the mean and variance of the voltage parameter sequence. Finally, use methods such as the correlation coefficient to calculate the correlation degree between the energy ratio feature and the mean-variance feature, and obtain the voltage fluctuation pattern feature. For example, if the first preset frequency is 10 Hz, the second preset frequency is 100 Hz, the energy of the voltage fluctuation component obtained through Fourier transform is 50 J, the energy of the noise component is 10 J, the mean of the voltage parameter sequence is 220 V, and the variance is 100, then the energy ratio feature is 50 / 10 = 5. If the correlation degree calculated through the correlation coefficient is 0.8, then the voltage fluctuation pattern feature can be represented as a comprehensive feature vector.
[0031] Step S230: Based on the temporal correlation between the temperature parameter sequence in the environmental monitoring parameter set and the load fluctuation parameter set, construct a response delay function of temperature change to load fluctuation, and generate a temperature correlation response feature through the integral operation of the response delay function.
[0032] Temporal correlation refers to the correlation relationship between the temperature parameter sequence and the load fluctuation parameter set in the time series, reflecting the synchronization and delay between temperature change and load fluctuation. The response delay function is a function that describes the response delay characteristics of temperature change to load fluctuation, which indicates how long it takes for the temperature to change correspondingly after the load fluctuation occurs. The integral operation is to integrate the response delay function within a set time interval to obtain the cumulative response effect of temperature change to load fluctuation. The temperature correlation response feature is a feature constructed based on the integral operation result, used to describe the correlation degree between temperature change and load fluctuation.
[0033] As an implementation manner, in step S230, based on the temporal correlation between the temperature parameter sequence in the environmental monitoring parameter set and the load fluctuation parameter set, construct a response delay function of temperature change to load fluctuation, and generate a temperature correlation response feature through the integral operation of the response delay function, which may specifically include the following steps S231 to S235: Step S231: Align the time stamps of the temperature data points in the temperature parameter sequence and the load data points in the load fluctuation parameter set to generate a set of temperature-load time series pairs within the synchronous time window.
[0034] Timestamp alignment processing is to match and align the data points in the temperature parameter sequence and the load fluctuation parameter set according to the timestamps, ensuring that they are synchronized in time. The synchronous time window refers to a specific time interval within which the temperature and load data are processed. The temperature-load time series pair set is a set of pairs composed of the aligned temperature data points and load data points, and each pair contains a temperature value and the corresponding load value. For example, if the temperature parameter sequence and the load fluctuation parameter set are [(t1, T1), (t2, T2),...] and [(t1', L1), (t2', L2),...] respectively, through timestamp alignment processing, the data points with similar timestamps are paired to obtain the temperature-load time series pair set [(t1, T1, L1), (t2, T2, L2),...] within the synchronous time window.
[0035] Step S232: Determine the delay time corresponding to the maximum cross-correlation value of the temperature change relative to the load fluctuation according to the cross-correlation analysis of the temperature change rate and the load change rate in the temperature-load time series pair set.
[0036] Cross-correlation analysis is a method for analyzing the correlation between two time series. By calculating their cross-correlation function, the delay relationship between them is determined. The temperature change rate refers to the change amount of temperature per unit time, and the load change rate refers to the change amount of load per unit time. The delay time corresponding to the maximum cross-correlation value refers to the time delay corresponding to the maximum cross-correlation value in the cross-correlation function. This delay time represents the response delay of the temperature change relative to the load fluctuation. In specific implementation, the numpy.correlate function in Python can be used to calculate the cross-correlation function of the temperature change rate and the load change rate, and then find the index corresponding to the maximum value of the cross-correlation function, and convert this index into a time delay.
[0037] Step S233: Based on the delay time and the proportional relationship between the temperature and load change amplitudes in the temperature-load time series pair set, construct a piecewise linear response delay function of the temperature change to the load fluctuation; the piecewise linear response delay function includes the temperature response slope parameters and delay compensation parameters corresponding to different load intervals.
[0038] The piecewise linear response delay function is a piecewise-defined linear function used to describe the response relationship between temperature changes and load fluctuations. The temperature response slope parameters corresponding to different load intervals represent the rate at which temperature changes with load within different load ranges. The delay compensation parameter is determined based on the delay time and is used to compensate for the delay of temperature changes relative to load fluctuations. When specifically constructing it, first divide the load interval into multiple sub-intervals according to the size of the load. Then, within each sub-interval, calculate the temperature response slope parameter based on the proportional relationship between the temperature and the load change amplitude in the temperature-load time series pair set. Finally, combine the delay time to determine the delay compensation parameter and construct the piecewise linear response delay function. For example, if the load interval is divided into three sub-intervals [0, 100], [100, 200], and [200, 300], and the temperature response slope parameters calculated within each sub-interval are k1, k2, and k3 respectively, and the delay compensation parameter is d, then the piecewise linear response delay function can be expressed as a piecewise function.
[0039] Step S234: Integrate the cumulative effect of the piecewise linear response delay function within the synchronous time window to generate the cumulative response energy distribution feature of temperature changes with respect to load fluctuations.
[0040] The integration process is to integrate the piecewise linear response delay function within the synchronous time window to calculate the cumulative response effect of temperature changes with respect to load fluctuations. The cumulative response energy distribution feature is a feature constructed based on the integration result and is used to describe the temporal distribution of the cumulative response of temperature changes with respect to load fluctuations. When specifically implementing it, numerical integration methods such as the trapezoidal integration method or Simpson's integration method can be used to integrate the piecewise linear response delay function. For example, using the trapezoidal integration method, divide the synchronous time window into multiple small intervals, approximate the piecewise linear response delay function as a linear function within each small interval, then calculate the integral value of each small interval, and finally sum up the integral values of all small intervals to obtain the cumulative response energy distribution feature.
[0041] Step S235: Couple and calculate the cumulative response energy distribution feature with the mean offset of the temperature parameter sequence and output the temperature-related response feature.
[0042] The mean offset refers to the difference between the mean of the temperature parameter sequence and a certain reference value, reflecting the overall offset of the temperature. The coupling calculation is to comprehensively calculate the cumulative response energy distribution feature and the mean offset to generate a comprehensive temperature-related response feature. When specifically implementing it, methods such as weighted summation can be used to couple the cumulative response energy distribution feature and the mean offset. For example, if the cumulative response energy distribution feature is E, the mean offset is ΔT, and the weights are w1 and w2 respectively, then the temperature-related response feature can be expressed as w1×E + w2×ΔT.
[0043] Step S240: Perform multi-dimensional normalization fusion processing on the current change trend feature, voltage fluctuation pattern feature, and temperature correlation response feature, and output a load feature set.
[0044] Multi-dimensional normalization fusion processing is to perform normalization processing on feature data of different dimensions to make them have the same dimension and range, and then fuse these features to obtain a comprehensive load feature set. Various methods can be used for normalization processing, such as min-max normalization, Z-score normalization, etc. Vector splicing, weighted summation, etc. can be used for fusion processing. For example, if the current change trend feature is [1, 2, 3], the voltage fluctuation pattern feature is [4, 5, 6], and the temperature correlation response feature is [7, 8, 9], use min-max normalization to normalize them to the range of [0, 1], and then use the method of vector splicing to fuse them to obtain the load feature set [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9].
[0045] Step S300: Input the load feature set into a pre-trained reinforcement learning policy network, and perform charge and discharge policy priority matching processing on the load feature set through the reinforcement learning policy network to generate a priority ranking result of the target charge and discharge policy.
[0046] The pre-trained reinforcement learning policy network is a neural network model that has been pre-trained and is used to learn the optimal charge and discharge policies under different load characteristics. Charge and discharge policy priority matching processing is to match the load feature set with historical charge and discharge policies and evaluate the priority of each charge and discharge policy. The target charge and discharge policy refers to the charge and discharge policy that is suitable for the current load characteristics obtained after matching. The priority ranking result is the result obtained by sorting the target charge and discharge policies from high to low according to the priority.
[0047] As an implementation, in step S300, perform charge and discharge policy priority matching processing on the load feature set through the reinforcement learning policy network to generate a priority ranking result of the target charge and discharge policy, which can specifically include the following steps S310 to S350: Step S310: Obtain the policy execution records in the historical charge and discharge policy set, and the policy execution records include the load matching degree parameter, battery loss parameter, and response timeliness parameter of each historical charge and discharge policy.
[0048] The set of historical charge-discharge strategies refers to the set of all charge-discharge strategies used in the past. The strategy execution record is the data recording the execution of each historical charge-discharge strategy, including the load matching degree parameter, the battery loss parameter, and the response timeliness parameter. The load matching degree parameter refers to the matching degree between the historical charge-discharge strategy and the load characteristics, reflecting the applicability of the strategy to the current load. The battery loss parameter refers to the battery loss situation during the execution of the historical charge-discharge strategy, such as capacity attenuation and internal resistance increase. The response timeliness parameter refers to the response speed of the historical charge-discharge strategy to load changes, reflecting the timeliness of the strategy. The strategy execution record can be obtained by querying relevant data from the database. For example, in the battery management system of a supercomputer center, the execution record of each charge-discharge strategy is stored in the database, and the strategy execution record in the set of historical charge-discharge strategies can be obtained by querying the database.
[0049] Step S320: Construct a strategy matching weight coefficient according to the load matching degree parameter, and perform inverse normalization processing on the battery loss parameter and the response timeliness parameter to generate a strategy optimization constraint coefficient.
[0050] The strategy matching weight coefficient is a coefficient constructed according to the load matching degree parameter, used to measure the importance of each historical charge-discharge strategy in priority matching. The inverse normalization processing is to process the battery loss parameter and the response timeliness parameter so that their value ranges are consistent with the load matching degree parameter, facilitating subsequent calculations. The strategy optimization constraint coefficient is a coefficient constructed based on the inversely normalized battery loss parameter and response timeliness parameter, used to constrain the selection of charge-discharge strategies and avoid selecting strategies with excessive battery loss or long response timeliness.
[0051] In specific implementation, first, according to the magnitude of the load matching degree parameter, a weight is assigned to each historical charge and discharge strategy. The larger the weight, the higher the matching degree of the strategy with the current load. Then, reverse normalization processing is performed on the battery loss parameter and the response timeliness parameter. For example, the maximum-minimum normalization method is used to normalize them to the range of [0, 1]. Finally, the reverse-normalized battery loss parameter and response timeliness parameter are weighted and summed to obtain the policy optimization constraint coefficient. For example, if the load matching degree parameters are [0.8, 0.6, 0.4], the battery loss parameters are [0.2, 0.3, 0.4], and the response timeliness parameters are [0.1, 0.2, 0.3], after reverse normalization processing, the reverse-normalized battery loss parameters are [0.8, 0.7, 0.6], and the response timeliness parameters are [0.9, 0.8, 0.7]. If the weights are 0.5 and 0.5 respectively, the policy optimization constraint coefficient is [0.8×0.5 + 0.9×0.5, 0.7×0.5 + 0.8×0.5, 0.6×0.5 + 0.7×0.5] = [0.85, 0.75, 0.65].
[0052] Step S330: Use the policy evaluation module in the reinforcement learning policy network to predict the policy matching degree for the load feature set, and output the initial priority score of each candidate charge and discharge strategy.
[0053] The policy evaluation module is a sub-module in the reinforcement learning policy network, used to evaluate the matching degree of each candidate charge and discharge strategy with the load feature set. Policy matching degree prediction is to predict the matching degree of each candidate charge and discharge strategy according to the load feature set. The initial priority score is the preliminary evaluation score of each candidate charge and discharge strategy, reflecting the priority of the strategy under the current load characteristics.
[0054] As an implementation method, in step S330, using the policy evaluation module in the reinforcement learning policy network to predict the policy matching degree for the load feature set and output the initial priority score of each candidate charge and discharge strategy can specifically include the following steps S331 to S337: Step S331: Perform steady-state and transient component separation processing on the current change trend feature in the load feature set, and extract the steady-state current parameter sequence and transient current impact feature.
[0055] The separation and processing of steady-state and transient components decompose the current change trend characteristics into steady-state and transient components. The steady-state current parameter sequence is the steady-state part of the current change trend characteristics, reflecting the long-term stable change trend of the current. The transient current impact characteristics are the transient part of the current change trend characteristics, reflecting the sudden change of the current. In specific implementation, filtering methods such as low-pass filters and high-pass filters can be used to separate the current change trend characteristics. The low-pass filter can filter out high-frequency components to obtain the steady-state current parameter sequence; the high-pass filter can filter out low-frequency components to obtain the transient current impact characteristics. For example, using Butterworth low-pass and high-pass filters and setting appropriate cut-off frequencies to filter the current change trend characteristics to obtain the steady-state current parameter sequence and the transient current impact characteristics.
[0056] Step S332: Perform weight assignment processing on the main energy distribution parameter and the secondary fluctuation component parameter in the voltage fluctuation mode characteristics to generate a voltage fluctuation sensitivity coefficient.
[0057] The main energy distribution parameter is the part with a relatively large energy proportion in the voltage fluctuation mode characteristics, reflecting the main characteristics of the voltage fluctuation. The secondary fluctuation component parameter is the part with a relatively small energy proportion in the voltage fluctuation mode characteristics, reflecting the secondary characteristics of the voltage fluctuation. The weight assignment processing is to assign different weights to the main energy distribution parameter and the secondary fluctuation component parameter, and then perform weighted summation on them to obtain the voltage fluctuation sensitivity coefficient. In specific implementation, different weights can be assigned to the two parameters according to the actual situation. For example, if the main energy distribution parameter is 0.8, the secondary fluctuation component parameter is 0.2, and the weights are 0.7 and 0.3 respectively, then the voltage fluctuation sensitivity coefficient is 0.8×0.7 + 0.2×0.3 = 0.62.
[0058] Step S333: Based on the cumulative response energy distribution characteristics of the temperature correlation response characteristics, perform thermal effect compensation processing on the steady-state current parameter sequence to generate a temperature-calibrated steady-state current sequence.
[0059] The thermal effect compensation process takes into account the influence of temperature on current and corrects the steady-state current parameter sequence. The cumulative response energy distribution characteristic reflects the cumulative response of temperature changes to load fluctuations. Based on this characteristic, the degree of influence of temperature on current can be determined. The steady-state current sequence after temperature calibration is the steady-state current sequence obtained after thermal effect compensation processing, which more accurately reflects the current change situation considering the influence of temperature. In specific implementation, a thermal effect compensation model can be established. According to the cumulative response energy distribution characteristic and the steady-state current parameter sequence, the compensation amount of temperature on current is calculated, and then the compensation amount is added to the steady-state current parameter sequence to obtain the steady-state current sequence after temperature calibration. For example, if the thermal effect compensation model is I_calibrated = I_steady + k×E, where I_calibrated is the steady-state current sequence after temperature calibration, I_steady is the steady-state current parameter sequence, k is the compensation coefficient, and E is the cumulative response energy distribution characteristic.
[0060] Step S334: Perform time-series superposition processing on the transient current impact characteristic and the steady-state current sequence after temperature calibration to generate a composite current feature vector.
[0061] Time-series superposition processing is to superimpose the transient current impact characteristic and the steady-state current sequence after temperature calibration in the time series to obtain a comprehensive composite current feature vector. The composite current feature vector contains both the steady-state and transient information of the current, and more comprehensively reflects the change situation of the current. In specific implementation, the data at the corresponding time points of the transient current impact characteristic and the steady-state current sequence after temperature calibration can be added to obtain the composite current feature vector. For example, if the transient current impact characteristic is [1, 2, 3] and the steady-state current sequence after temperature calibration is [4, 5, 6], then the composite current feature vector is [1 + 4, 2 + 5, 3 + 6] = [5, 7, 9].
[0062] Step S335: Perform cross-dimensional correlation analysis on the composite current feature vector and the voltage fluctuation sensitivity coefficient to construct a current-voltage joint fluctuation feature matrix.
[0063] Cross - dimensional correlation analysis is to analyze the correlation between the composite current feature vector and the voltage fluctuation sensitivity coefficient, considering their interactions in different dimensions. The current - voltage joint fluctuation feature matrix is a matrix constructed based on the results of cross - dimensional correlation analysis, which is used to describe the joint fluctuation characteristics of current and voltage. In specific implementation, correlation analysis methods such as the Pearson correlation coefficient can be used to calculate the correlation between the composite current feature vector and the voltage fluctuation sensitivity coefficient, and then the correlation results are formed into a matrix. For example, if the composite current feature vector is [5, 7, 9] and the voltage fluctuation sensitivity coefficient is 0.62, calculate the Pearson correlation coefficient between them to obtain the correlation coefficient matrix, which is used as the current - voltage joint fluctuation feature matrix.
[0064] Step S336: Input the current - voltage joint fluctuation feature matrix into the feature matching layer of the policy evaluation module, and compare it with the load adaptation feature template in the historical policy execution record to generate the initial matching degree score of the candidate charge - discharge policy.
[0065] The feature matching layer is a level in the policy evaluation module for feature matching. The load adaptation feature template is a template related to the load characteristics in the historical policy execution record, which reflects the characteristics of different charge - discharge policies under different load conditions. The similarity comparison is to compare the current - voltage joint fluctuation feature matrix with the load adaptation feature template and calculate the similarity between them. The initial matching degree score is the score obtained according to the similarity comparison result, which reflects the matching degree of the candidate charge - discharge policy and the historical policy in terms of load characteristics. In specific implementation, methods such as Euclidean distance and cosine similarity can be used to calculate the similarity between the current - voltage joint fluctuation feature matrix and the load adaptation feature template. For example, use the Euclidean distance to calculate the distance between two matrices. The smaller the distance, the higher the similarity. Convert the similarity into a score to obtain the initial matching degree score.
[0066] Step S337: Perform policy conflict detection and processing on the initial matching degree score, eliminate the policy matching results that are incompatible with the current battery state parameter set, and output the initial priority score.
[0067] Strategy conflict detection and handling is to check whether the candidate charge and discharge strategies corresponding to the initial matching degree scores are compatible with the current set of battery state parameters. If a certain candidate charge and discharge strategy may cause overcharging, over-discharging or other abnormal conditions of the battery, it is considered that the strategy is not compatible with the current set of battery state parameters. Eliminating the incompatible strategy matching results is to set the initial matching degree scores of the candidate charge and discharge strategies that are not compatible with the current set of battery state parameters to a lower value or directly remove them. The initial priority score is the score output after the strategy conflict detection and handling, which more accurately reflects the priority of each candidate charge and discharge strategy under the current battery state. In specific implementation, some constraint conditions can be set according to the set of battery state parameters, such as the state of charge range and voltage range of the battery, to check whether each candidate charge and discharge strategy meets these constraint conditions and handle the strategies that do not meet the constraint conditions.
[0068] Step S340: Perform a weighted operation on the initial priority score, the strategy matching weight coefficient, and the strategy optimization constraint coefficient to generate a corrected strategy priority score.
[0069] The weighted operation is to perform a weighted sum of the initial priority score, the strategy matching weight coefficient, and the strategy optimization constraint coefficient according to the set weights to obtain the corrected strategy priority score. The corrected strategy priority score comprehensively considers factors such as load matching degree, battery loss, and response timeliness, and more comprehensively evaluates the priority of each candidate charge and discharge strategy. In specific implementation, if the initial priority score, the strategy matching weight coefficient, and the strategy optimization constraint coefficient are S_init, W_match, and C_constraint respectively, and the weights are w1, w2, and w3, then the corrected strategy priority score S_corrected = w1×S_init + w2×W_match + w3×C_constraint.
[0070] Step S350: Sort the candidate charge and discharge strategies in descending order according to the corrected strategy priority score to generate a priority sorting result.
[0071] Sorting in descending order is to sort the candidate charge and discharge strategies from high to low according to the corrected strategy priority score. The priority sorting result is the order of the candidate charge and discharge strategies obtained after sorting in descending order, which reflects the priority of each strategy under the current load characteristics and battery state. In specific implementation, sorting algorithms such as quicksort and mergesort can be used to sort the candidate charge and discharge strategies. For example, using the sorted function in Python to sort the candidate charge and discharge strategies in descending order according to the corrected strategy priority score to obtain the priority sorting result.
[0072] Step S400: Adjust the current charge-discharge strategy set according to the priority sorting result to generate an optimized charge-discharge strategy set; each charge-discharge strategy in the optimized charge-discharge strategy set includes charge-discharge trigger condition parameters and power distribution parameters.
[0073] The current charge-discharge strategy set refers to the set of charge-discharge strategies currently in use. The adjustment is to select and modify the strategies in the current charge-discharge strategy set according to the priority sorting result to improve the performance of the strategies. The optimized charge-discharge strategy set is the charge-discharge strategy set obtained after adjustment, and each charge-discharge strategy in it includes charge-discharge trigger condition parameters and power distribution parameters. The charge-discharge trigger condition parameters refer to the conditions for triggering charge-discharge operations, such as the state of charge and voltage of the battery. The power distribution parameters refer to the magnitude of the power allocated to the battery during the charge-discharge process.
[0074] As an implementation, in step S400, adjusting the current charge-discharge strategy set according to the priority sorting result to generate an optimized charge-discharge strategy set may specifically include the following steps S410 to S440: Step S410: Screen out a target strategy subset with a priority score higher than a preset threshold from the priority sorting result, and extract the charge-discharge trigger condition parameters in the target strategy subset.
[0075] The preset threshold is a pre-set priority score threshold for screening out strategies with higher priorities. The target strategy subset is a set of strategies screened out from the priority sorting result with a priority score higher than the preset threshold. The charge-discharge trigger condition parameters are the charge-discharge trigger conditions of each strategy in the target strategy subset, such as the state of charge threshold and voltage threshold of the battery. Specifically, when implementing, traverse the priority sorting result, add the strategies with a priority score higher than the preset threshold to the target strategy subset, and then extract the charge-discharge trigger condition parameters from the target strategy subset. For example, if the preset threshold is 0.8 and the priority sorting result is [(Strategy 1, 0.9), (Strategy 2, 0.7), (Strategy 3, 0.85)], then the target strategy subset is [(Strategy 1, 0.9), (Strategy 3, 0.85)], and the extracted charge-discharge trigger condition parameters are the charge-discharge trigger conditions of Strategy 1 and Strategy 3.
[0076] Step S420: Calibrate the charge-discharge trigger condition parameters according to the current change trend feature in the load feature set to generate a calibrated trigger condition parameter set.
[0077] Calibration is to adjust the charge-discharge trigger condition parameters according to the current change trend feature to make them more suitable for the current load situation. The calibrated trigger condition parameter set is the set of charge-discharge trigger condition parameters obtained after calibration, which more accurately reflects the charge-discharge trigger conditions under the current load characteristics.
[0078] As an implementation manner, in step S420, the charge and discharge trigger condition parameters are calibrated according to the current change trend feature in the load feature set to generate a calibrated trigger condition parameter set, which may specifically include the following steps S421 to S426: Step S421: Extract the fluctuation period feature and the peak interval distribution feature in the current change trend feature, and generate a current stability parameter based on the period stability coefficient of the fluctuation period feature and the dispersion degree of the peak interval distribution feature.
[0079] The period stability coefficient of the fluctuation period feature is an index to measure the stability of the fluctuation period, reflecting the change degree of the fluctuation period. The dispersion degree of the peak interval distribution feature is an index to measure the dispersion degree of the peak interval distribution, reflecting the regularity of the peak appearance. The current stability parameter is a parameter constructed based on the period stability coefficient and the dispersion degree, used to describe the stability of the current. In specific implementation, the period stability coefficient and the dispersion degree can be weighted and summed to obtain the current stability parameter. For example, if the period stability coefficient is 0.8, the dispersion degree is 0.2, and the weights are 0.7 and 0.3 respectively, then the current stability parameter is 0.8×0.7 + 0.2×0.3 = 0.62.
[0080] Step S422: Determine the adjustment ratio of the charge and discharge trigger threshold according to the difference between the current stability parameter and the preset stability reference value, and perform a linear scaling process on the current threshold in the charge and discharge trigger condition parameters based on the adjustment ratio to generate a calibrated current trigger threshold.
[0081] The preset stability reference value is a pre-set current stability reference value. The adjustment ratio is determined according to the difference between the current stability parameter and the preset stability reference value, used to adjust the charge and discharge trigger threshold. The linear scaling process is to scale the current threshold in the charge and discharge trigger condition parameters according to the adjustment ratio to obtain a calibrated current trigger threshold. In specific implementation, first calculate the difference between the current stability parameter and the preset stability reference value, then determine the adjustment ratio according to the difference. For example, the larger the difference, the larger the adjustment ratio. Finally, multiply the current threshold in the charge and discharge trigger condition parameters by the adjustment ratio to obtain the calibrated current trigger threshold. For example, if the preset stability reference value is S_base, the current stability parameter is S_current, and the adjustment ratio is α, then ΔS = |S_current - S_base|, α = 1 - kΔS (k is the attenuation coefficient).
[0082] Step S423: Extract the energy ratio relationship between the main energy distribution parameter and the secondary fluctuation component parameter in the voltage fluctuation mode feature, and perform a range expansion on the voltage fluctuation tolerance in the charge and discharge trigger condition parameters based on the energy ratio relationship to generate an extended voltage tolerance parameter.
[0083] The energy ratio relationship between the main energy distribution parameter and the secondary fluctuation component parameter reflects the ratio between the main characteristics and the secondary characteristics of the voltage fluctuation. The voltage fluctuation tolerance is the allowable voltage fluctuation range in the charge and discharge trigger condition parameters. Range extension is to expand the voltage fluctuation tolerance according to the energy ratio relationship to adapt to different voltage fluctuation situations. The extended voltage tolerance parameter is the voltage fluctuation tolerance parameter obtained after range extension, which can more flexibly adapt to the changes in voltage fluctuations. In specific implementation, according to the energy ratio relationship between the main energy distribution parameter and the secondary fluctuation component parameter, the expansion multiple is determined, and then the voltage fluctuation tolerance is multiplied by the expansion multiple to obtain the extended voltage tolerance parameter. For example, if the energy ratio relationship between the main energy distribution parameter and the secondary fluctuation component parameter is 2:1, the voltage fluctuation tolerance is ±5V, and the expansion multiple is 1.2, then the extended voltage tolerance parameter is ±5×1.2 = ±6V.
[0084] Step S424: Perform parameter synchronization verification processing on the calibrated current trigger threshold and the extended voltage tolerance parameter, detect whether the joint effective interval of the current threshold and the voltage tolerance covers the fluctuation range of the load fluctuation parameter set, and generate a conflict-free trigger parameter combination.
[0085] The parameter synchronization verification processing is to check whether the calibrated current trigger threshold and the extended voltage tolerance parameter can take effect synchronously to avoid conflicts. The joint effective interval refers to the interval where the calibrated current trigger threshold and the extended voltage tolerance parameter are both satisfied. The fluctuation range of the load fluctuation parameter set refers to the change range of the current and voltage of the load during operation. The conflict-free trigger parameter combination is the combination of the calibrated current trigger threshold and the extended voltage tolerance parameter obtained after verification, which can effectively trigger the charge and discharge operation and will not have conflicts. In specific implementation, the calibrated current trigger threshold and the extended voltage tolerance parameter are compared within the fluctuation range of the load fluctuation parameter set, check whether their joint effective interval covers the fluctuation range, and adjust the parameters that do not meet the conditions until a conflict-free trigger parameter combination is obtained.
[0086] Step S425: Perform thermal inertia compensation processing on the conflict-free trigger parameter combination based on the cumulative response energy distribution characteristic of the temperature correlation response characteristic, and correct the influence of the temperature delay on the effective timing of the trigger condition to generate a temperature-corrected trigger parameter set.
[0087] The thermal inertia compensation process takes into account the influence of temperature on the effective timing of trigger conditions and corrects the trigger parameter combinations without conflicts. The cumulative response energy distribution characteristic reflects the cumulative response of temperature changes to load fluctuations. Based on this characteristic, the influence degree of temperature delay on the effective timing of trigger conditions can be determined. The temperature-corrected trigger parameter set is the trigger parameter set obtained after the thermal inertia compensation process, which more accurately reflects the trigger conditions considering the influence of temperature. In specific implementation, a thermal inertia compensation model can be established. According to the cumulative response energy distribution characteristic and the trigger parameter combinations without conflicts, the compensation amount of temperature delay on the effective timing of trigger conditions is calculated, and then the compensation amount is added to the trigger conditions to obtain the temperature-corrected trigger parameter set. For example, if the thermal inertia compensation model is T_corrected = T_uncorrected + k × E, where T_corrected is the temperature-corrected trigger parameter set, T_uncorrected is the trigger parameter combination without conflicts, k is the compensation coefficient, and E is the cumulative response energy distribution characteristic.
[0088] Step S426: For the temperature-corrected trigger parameter set, screen out the trigger condition parameters that match the fluctuation frequency and amplitude in the load fluctuation parameter set to generate a calibrated trigger condition parameter set.
[0089] Screening is to select the trigger condition parameters that match the fluctuation frequency and amplitude in the load fluctuation parameter set from the temperature-corrected trigger parameter set. The fluctuation frequency and amplitude are important characteristics of load fluctuations. Selecting the trigger condition parameters that match them can improve the effectiveness of the charge and discharge strategy. The calibrated trigger condition parameter set is the trigger condition parameter set obtained after screening, which is more in line with the current load fluctuation situation. In specific implementation, the trigger condition parameters in the temperature-corrected trigger parameter set are compared with the fluctuation frequency and amplitude in the load fluctuation parameter set, and the trigger condition parameters with higher matching degrees are selected to form the calibrated trigger condition parameter set.
[0090] Step S430: Verify the load balance degree of the power distribution parameters. If it is detected that the matching degree of the power distribution parameters and the load fluctuation parameter set is lower than the preset standard, perform non-linear compensation processing on the power distribution parameters based on the voltage fluctuation mode characteristic.
[0091] Load balancing degree verification is to check whether the power distribution parameters can evenly distribute the load over different time periods, avoiding situations where the load is too large or too small. The matching degree refers to the similarity between the power distribution parameters and the set of load fluctuation parameters, reflecting the rationality of power distribution. The preset standard is a preset matching degree threshold used to determine whether the power distribution is reasonable. Nonlinear compensation processing is to adjust the power distribution parameters according to the characteristics of the voltage fluctuation pattern to improve the matching degree between power distribution and load fluctuation. In specific implementation, the matching degree between the power distribution parameters and the set of load fluctuation parameters is calculated. If the matching degree is lower than the preset standard, the power distribution parameters are nonlinearly adjusted according to the main energy distribution parameters and secondary fluctuation component parameters in the voltage fluctuation pattern characteristics, such as using a nonlinear function for correction.
[0092] Step S440: Combine and optimize the calibrated trigger condition parameter set and the compensated power distribution parameters to generate an optimized charge-discharge strategy set; the combination optimization includes parameter synchronization verification and conflict elimination processing.
[0093] Combination optimization is to comprehensively process the calibrated trigger condition parameter set and the compensated power distribution parameters to generate an optimal charge-discharge strategy set. Parameter synchronization verification is to check whether the calibrated trigger condition parameter set and the compensated power distribution parameters can take effect synchronously, avoiding conflicts. Conflict elimination processing is to adjust the conflicting parameters to ensure the effectiveness of the optimized charge-discharge strategy set. In specific implementation, the calibrated trigger condition parameter set and the compensated power distribution parameters are combined, and then parameter synchronization verification and conflict elimination processing are performed. The parameters that do not meet the conditions are adjusted until an optimized charge-discharge strategy set is obtained.
[0094] Step S500: Perform charge-discharge operations based on the optimized charge-discharge strategy set, and collect the battery performance feedback data after the charge-discharge operations to iteratively update the strategy matching parameters of the reinforcement learning policy network.
[0095] The optimized charge-discharge strategy set is the charge-discharge strategy set obtained after adjustment and optimization. Performing charge-discharge operations is to charge and discharge the battery according to the strategies in the optimized charge-discharge strategy set. The battery performance feedback data is the data reflecting the changes in battery performance collected after the charge-discharge operations, such as the capacity, internal resistance, temperature of the battery, etc. Iterative update is to update the strategy matching parameters of the reinforcement learning policy network according to the battery performance feedback data to improve the performance of the policy network.
[0096] As an implementation manner, in step S500, performing charge-discharge operations based on the optimized charge-discharge strategy set, and collecting the battery performance feedback data after the charge-discharge operations to iteratively update the strategy matching parameters of the reinforcement learning policy network may specifically include the following steps S510 to S540: Step S510: When executing each charge-discharge strategy in the optimized charge-discharge strategy set, the deviation degree between the battery output power curve and the strategy expected power curve is monitored in real time.
[0097] The battery output power curve is the curve of the actual output power of the battery changing with time during the execution of the charge-discharge strategy. The strategy expected power curve is the curve of the power expected by each strategy in the optimized charge-discharge strategy set changing with time. The deviation degree refers to the degree of difference between the battery output power curve and the strategy expected power curve, reflecting the degree of conformity between the actual charge-discharge situation and the expected situation. Real-time monitoring is to continuously collect battery output power data during the charge-discharge process and calculate its deviation degree from the strategy expected power curve. In specific implementation, a power sensor can be used to collect battery output power data in real time, and then a curve fitting method can be used to obtain the battery output power curve, calculate the difference between it and the strategy expected power curve, and obtain the deviation degree.
[0098] Step S520: If the deviation degree exceeds the preset tolerance range, trigger an abnormal mark for strategy execution, and record the abnormal timestamp corresponding to the deviation degree and the snapshot of environmental parameters.
[0099] The preset tolerance range is a preset deviation degree threshold used to judge whether the charge-discharge operation is normal. The abnormal mark for strategy execution is to mark the current charge-discharge strategy execution situation when the deviation degree exceeds the preset tolerance range, indicating that the strategy execution is abnormal. The abnormal timestamp is the time point when the deviation degree exceeds the preset tolerance range. Recording the abnormal timestamp can facilitate subsequent analysis and processing. The snapshot of environmental parameters is the environmental parameter data collected at the abnormal time point, such as temperature, humidity, air pressure, etc. Environmental factors may affect the performance of the battery. Recording the snapshot of environmental parameters helps to analyze the cause of the abnormality. In specific implementation, compare the size of the deviation degree with the preset tolerance range. If the deviation degree exceeds the preset tolerance range, trigger an abnormal mark for strategy execution, and record the abnormal timestamp and the snapshot of environmental parameters.
[0100] Step S530: Conduct an attenuation coefficient analysis on the battery performance feedback data, extract the capacity attenuation rate and internal resistance change rate after the strategy execution, and input the capacity attenuation rate and internal resistance change rate into the strategy effectiveness evaluation model to generate a strategy effectiveness score.
[0101] The attenuation coefficient analysis is to analyze the battery performance feedback data to determine the capacity attenuation and internal resistance change of the battery. The capacity attenuation rate refers to the proportion of the battery capacity decrease after implementing the charge-discharge strategy, reflecting the degree of battery capacity attenuation. The internal resistance change rate refers to the proportion of the battery internal resistance change after implementing the charge-discharge strategy, reflecting the change of the battery internal resistance. The strategy effect evaluation model is a pre-trained model used to evaluate the effectiveness of the charge-discharge strategy. The strategy effectiveness score is the score obtained through the strategy effect evaluation model based on the capacity attenuation rate and the internal resistance change rate, reflecting the long-term impact degree of the charge-discharge strategy on the battery performance.
[0102] As an implementation manner, in step S530, perform attenuation coefficient analysis on the battery performance feedback data, and extract the capacity attenuation rate and the internal resistance change rate after the strategy execution. Specifically, it may include the following steps S531 to S535: Step S531: After each charge-discharge strategy is executed, collect the battery charge-discharge cycle count and the capacity test data sequence.
[0103] The battery charge-discharge cycle count refers to the number of charge-discharge cycles experienced by the battery during the execution of the charge-discharge strategy. The capacity test data sequence is the data sequence obtained by testing the battery capacity after each charge-discharge cycle, reflecting the change of the battery capacity with the charge-discharge cycle count. These data can be collected using a battery test device, such as a battery charge-discharge tester. After each charge-discharge cycle ends, measure the battery capacity, and record the charge-discharge cycle count and the capacity test data.
[0104] Step S532: Perform linear regression analysis on the capacity test data sequence, and calculate the capacity decrease slope corresponding to the unit cycle count as the capacity attenuation rate.
[0105] Linear regression analysis is a statistical method used to analyze the linear relationship between two variables. In this step, take the capacity test data sequence as the dependent variable and the charge-discharge cycle count as the independent variable, perform linear regression analysis, and obtain the linear equation of the capacity change with the charge-discharge cycle count. The capacity decrease slope corresponding to the unit cycle count is the slope of the linear equation, reflecting the proportion of the battery capacity decrease in each charge-discharge cycle, that is, the capacity attenuation rate. Specifically, when implementing, the `numpy.polyfit` function in Python can be used to perform linear regression analysis to obtain the coefficients of the linear equation, so as to calculate the capacity attenuation rate.
[0106] Step S533: Measure the DC internal resistance value of the battery at different state of charge, and construct the internal resistance-state of charge relationship curve.
[0107] The DC internal resistance value is the internal resistance of the battery in a DC circuit. Different states of charge refer to the states of the battery under different degrees of charging or discharging. The internal resistance-state of charge relationship curve is a curve that describes the change of the battery's internal resistance with the state of charge, reflecting the relationship between the battery's internal resistance and the state of charge. To measure the DC internal resistance value of the battery under different states of charge, an internal resistance tester can be used to measure the internal resistance of the battery under different states of charge, record the state of charge and the corresponding internal resistance values. Then, using a curve fitting method, such as polynomial fitting, the internal resistance-state of charge relationship curve is constructed.
[0108] Step S534: Calculate the change in curvature of the internal resistance-state of charge relationship curve, and generate the internal resistance change rate based on the integral result of the change in curvature.
[0109] The change in curvature refers to the change in the curvature of the internal resistance-state of charge relationship curve at different points, reflecting the change in the degree of curvature of the curve. The integral result is the result obtained by integrating the change in curvature within a set range of the state of charge, reflecting the cumulative effect of the change in the battery's internal resistance with the state of charge. The internal resistance change rate is calculated based on the integral result, reflecting the change ratio of the battery's internal resistance during the entire charge and discharge process. In specific implementation, first calculate the change in curvature of the internal resistance-state of charge relationship curve, and a numerical differentiation method can be used. Then, integrate the change in curvature, using a numerical integration method, such as the trapezoidal integration method or Simpson's integration method. Finally, calculate the internal resistance change rate based on the integral result.
[0110] Step S535: Standardize the capacity attenuation rate and the internal resistance change rate to generate an input vector for evaluating the strategy effect.
[0111] Standardization processing is to process the capacity attenuation rate and the internal resistance change rate so that they have the same dimension and range, facilitating subsequent analysis and processing. The input vector for evaluating the strategy effect is a vector composed of the capacity attenuation rate and the internal resistance change rate obtained after standardization processing, serving as the input to the strategy effect evaluation model. In specific implementation, the minimum-maximum normalization or Z-score normalization method can be used to standardize the capacity attenuation rate and the internal resistance change rate. For example, using the minimum-maximum normalization method, the capacity attenuation rate and the internal resistance change rate are normalized to the range [0, 1] to obtain the input vector for evaluating the strategy effect.
[0112] As an implementation manner, in step S530, input the capacity attenuation rate and the internal resistance change rate into the strategy effect evaluation model to generate a strategy effectiveness score, which specifically may include the following steps S536 to S5311: Step S536: Perform unit-time attenuation amount standardization processing on the capacity attenuation rate to generate a standardized capacity attenuation parameter; perform state-of-charge interval normalization processing on the internal resistance change rate to generate a standardized internal resistance change parameter.
[0113] The normalization processing of the decay amount per unit time is to convert the capacity decay rate into the decay amount per unit time and perform normalization processing to make it have the same dimension and range. The normalized capacity decay parameter is the capacity decay parameter obtained after the normalization processing of the decay amount per unit time. The normalization processing of the state of charge interval is to normalize the internal resistance change rate within the state of charge interval to make it comparable under different states of charge. The normalized internal resistance change parameter is the internal resistance change parameter obtained after the normalization processing of the state of charge interval. In specific implementation, for the capacity decay rate, divide it by the charge-discharge time to obtain the decay amount per unit time, and then use the min-max normalization or Z-score normalization method for normalization processing. For the internal resistance change rate, normalize it within the state of charge interval so that its value range is between [0, 1].
[0114] Step S537: Perform time-series feature splicing processing on the normalized capacity decay parameter and the normalized internal resistance change parameter to generate a capacity-internal resistance joint evaluation vector.
[0115] The time-series feature splicing processing is to splice the normalized capacity decay parameter and the normalized internal resistance change parameter in the time series to obtain a comprehensive vector. The capacity-internal resistance joint evaluation vector contains information on capacity decay and internal resistance change, and more comprehensively reflects the change of battery performance. In specific implementation, splice the normalized capacity decay parameter and the normalized internal resistance change parameter in chronological order to obtain the capacity-internal resistance joint evaluation vector. For example, if the normalized capacity decay parameter is [0.1, 0.2, 0.3] and the normalized internal resistance change parameter is [0.4, 0.5, 0.6], then the capacity-internal resistance joint evaluation vector is [0.1, 0.2, 0.3, 0.4, 0.5, 0.6].
[0116] Step S538: Input the capacity-internal resistance joint evaluation vector into the recurrent neural network layer in the policy effect evaluation model for time-series dependence feature extraction, and output the decay mode feature vector.
[0117] The recurrent neural network layer is a layer in the policy effect evaluation model, which is used to process sequential data and can capture the temporal dependencies in the data. Temporal dependency feature extraction is to extract time-related features from the capacity-internal resistance joint evaluation vector, which reflects the trends and patterns of battery performance changes. The decay mode feature vector is the feature vector output after being processed by the recurrent neural network layer, which contains the temporal pattern information of battery capacity decay and internal resistance change. Specifically, when implemented, the recurrent neural network layer can adopt structures such as long short-term memory network (LSTM) or gated recurrent unit (GRU). The capacity-internal resistance joint evaluation vector is input into the recurrent neural network layer, and after the calculation and processing of the network, the decay mode feature vector is output.
[0118] Step S539: Input the capacity-internal resistance joint evaluation vector into the convolutional layer in the policy effect evaluation model to capture the local fluctuation pattern, and generate the internal resistance fluctuation feature vector.
[0119] The convolutional layer is a layer in the policy effect evaluation model, which is used to extract local features from the data. Local fluctuation pattern capture is to extract the local fluctuation patterns from the capacity-internal resistance joint evaluation vector, which reflects the changes in battery performance within a local range. The internal resistance fluctuation feature vector is the feature vector output after being processed by the convolutional layer, which contains the fluctuation information of the battery internal resistance within a local range. Specifically, when implemented, the convolutional layer can adopt a one-dimensional convolutional kernel to perform a convolutional operation on the capacity-internal resistance joint evaluation vector, extract local features, and then perform dimensionality reduction processing through a pooling layer to obtain the internal resistance fluctuation feature vector.
[0120] Step S5310: Perform cross-channel attention weight assignment processing on the decay mode feature vector and the internal resistance fluctuation feature vector to generate the policy joint evaluation feature matrix.
[0121] Cross-channel attention weight assignment processing is to perform weighted combination on the decay mode feature vector and the internal resistance fluctuation feature vector to highlight important feature information. The policy joint evaluation feature matrix is the matrix obtained after cross-channel attention weight assignment processing, which synthesizes the decay mode features and the internal resistance fluctuation features, and more comprehensively evaluates the effect of the charge and discharge policy. Specifically, when implemented, an attention mechanism can be used to assign different weights to the decay mode feature vector and the internal resistance fluctuation feature vector, and then perform weighted summation on them to obtain the policy joint evaluation feature matrix. For example, using the multi-head attention mechanism to calculate the attention weights of each feature vector, and then performing weighted combination of the feature vectors according to the weights.
[0122] Step S5311: Input the policy joint evaluation feature matrix into the fully connected layer in the policy effect evaluation model to perform linear transformation processing to generate the policy effectiveness score; the policy effectiveness score is used to quantify the long-term impact degree of the charge and discharge policy on battery performance.
[0123] The fully connected layer is the last layer in the policy effect evaluation model, which is used to perform a linear transformation on the policy joint evaluation feature matrix and map it to a scalar value, that is, the policy effectiveness score. The linear transformation process is to perform matrix multiplication on the policy joint evaluation feature matrix and the weight matrix of the fully connected layer, and then add a bias term to obtain the policy effectiveness score. The policy effectiveness score is a numerical value used to quantify the long-term impact of the charge-discharge policy on the battery performance. The higher the score, the smaller the impact of the policy on the battery performance, and the more effective the policy.
[0124] Step S540: Adjust the policy matching parameters of the reinforcement learning policy network according to the comparison result between the policy effectiveness score and the preset score threshold; the adjustment includes updating the policy priority weight and correcting the policy matching rule.
[0125] The preset score threshold is a preset scoring criterion used to judge the effectiveness of the charge-discharge policy. The comparison result is the magnitude relationship between the policy effectiveness score and the preset score threshold. The adjustment is to modify the policy matching parameters of the reinforcement learning policy network according to the comparison result to improve the performance of the policy network. Updating the policy priority weight is to adjust the priority weights of different charge-discharge policies so that more effective policies have higher priorities. Correcting the policy matching rule is to modify the policy matching rule to improve the accuracy of policy matching. Specifically, when the policy effectiveness score is higher than the preset score threshold, it is considered that the policy is effective, and the priority weight of this policy can be appropriately increased; when the policy effectiveness score is lower than the preset score threshold, it is considered that the policy effect is not good, and the policy matching rule needs to be corrected to avoid selecting this policy. By continuously adjusting the policy matching parameters, the reinforcement learning policy network can learn better charge-discharge policies and improve the performance and service life of the battery.
[0126] It can be understood that in the above introductions of the embodiments of the present invention, various algorithms involved can be obtained from relevant contents in the prior art. For the sake of saving space, they are not elaborated in the embodiments of the present invention. In addition, those skilled in the art can make detailed supplements according to the common general knowledge in the art when implementing the solution of the present invention. For example, according to the general knowledge in the art, normalization can be used to eliminate the dimension conflict before feature fusion, interpolation can be used to eliminate the dimension difference, the threshold can be reasonably set by combining historical data, experience or business scenario requirements, and the model can be trained based on the general model training method, etc. The present invention will no longer introduce the redundant implementation process in too much detail.
[0127] Please refer to Figure 2 , Figure 2Schematic structural diagram of a computer system provided by an embodiment of the present invention. The computer system at least includes a processor 101, a communication interface 102, and a memory 103. Among them, the processor 101, the communication interface 102, and the memory 103 can be connected through a bus or other means. Among them, the processor 101 (or central processing unit (CPU)) is the computing core and control core of the computer system, which can parse various instructions in the computer system and process various data in the computer system. The communication interface 102 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and under the control of the processor 101, it can be used to send and receive data; the communication interface 102 can also be used for the transmission and interaction of internal data in the computer system. The memory 103 (Memory) is a memory device in the computer system, used to store programs and data. It can be understood that the memory 103 here can include both the built-in memory of the computer system and, of course, the extended memory supported by the computer system. The memory 103 provides a storage space, and the operating system of the computer system is stored in this storage space, which is not limited in the present invention.
[0128] In one embodiment, the processor 101 executes the optimization method for the battery charge and discharge strategy of the supercomputer center based on reinforcement learning provided above in the embodiment of the present invention by running the computer program in the memory 103.
Claims
1. An optimization method for the charging and discharging strategy of the battery in a supercomputing center based on reinforcement learning, characterized in that, Including: Obtaining in real time a set of real-time operation data of the battery system of the supercomputer center, where the set of real-time operation data includes a set of battery state parameters, a set of load fluctuation parameters, and a set of environmental monitoring parameters; Performing load feature extraction processing on the set of real-time operation data to generate a set of load features of the battery system; the set of load features includes a current change trend feature, a voltage fluctuation pattern feature, and a temperature correlation response feature; Inputting the set of load features into a pre-trained reinforcement learning policy network, and performing charge and discharge policy priority matching processing on the set of load features through the reinforcement learning policy network to generate a priority ranking result of the target charge and discharge policy; Adjusting the current set of charge and discharge policies according to the priority ranking result to generate an optimized set of charge and discharge policies; each charge and discharge policy in the optimized set of charge and discharge policies includes charge and discharge trigger condition parameters and power distribution parameters; Performing charge and discharge operations based on the optimized set of charge and discharge policies, and collecting battery performance feedback data after the charge and discharge operations to iteratively update the policy matching parameters of the reinforcement learning policy network.
2. The method according to claim 1, wherein The performing load feature extraction processing on the set of real-time operation data to generate a set of load features of the battery system includes: Performing time-domain trend analysis on the current parameter sequence in the set of battery state parameters, extracting the fluctuation period feature and peak interval distribution feature of the current parameter sequence, and performing linear superposition processing on the fluctuation period feature and peak interval distribution feature to generate the current change trend feature; Performing frequency-domain conversion processing on the voltage parameter sequence in the set of voltage fluctuation parameters, extracting the energy ratio feature of the voltage fluctuation component below the first preset frequency and the noise component above the second preset frequency, and calculating the correlation degree between the energy ratio feature and the mean variance feature of the voltage parameter sequence to generate the voltage fluctuation pattern feature; According to the time-series correlation between the temperature parameter sequence in the set of environmental monitoring parameters and the set of load fluctuation parameters, constructing a response delay function of temperature change to load fluctuation, and generating the temperature correlation response feature through the integral operation of the response delay function; Performing multi-dimensional normalization fusion processing on the current change trend feature, the voltage fluctuation pattern feature, and the temperature correlation response feature, and outputting the set of load features.
3. The method according to claim 1, characterized in that, The performing charge and discharge policy priority matching processing on the set of load features through the reinforcement learning policy network to generate a priority ranking result of the target charge and discharge policy includes: Obtaining the policy execution records in the set of historical charge and discharge policies, where the policy execution records include the load matching degree parameters, battery loss parameters, and response timeliness parameters of each historical charge and discharge policy; Constructing a policy matching weight coefficient according to the load matching degree parameters, and performing inverse normalization processing on the battery loss parameters and response timeliness parameters to generate a policy optimization constraint coefficient; Performing policy matching degree prediction on the set of load features through the policy evaluation module in the reinforcement learning policy network, and outputting the initial priority scores of each candidate charge and discharge policy; Perform a weighted operation on the initial priority score, the policy matching weight coefficient, and the policy optimization constraint coefficient to generate a corrected policy priority score; Arrange the candidate charge and discharge strategies in descending order according to the corrected policy priority score to generate the priority ranking result.
4. The method according to claim 1, characterized in that Adjust the current charge and discharge strategy set according to the priority ranking result to generate an optimized charge and discharge strategy set, including: Screen a target strategy subset with a priority score higher than a preset threshold from the priority ranking result, and extract the charge and discharge trigger condition parameters in the target strategy subset; Calibrate the charge and discharge trigger condition parameters according to the current change trend feature in the load feature set to generate a calibrated trigger condition parameter set; Verify the load balance degree of the power distribution parameters. If it is detected that the matching degree between the power distribution parameters and the load fluctuation parameter set is lower than the preset standard, perform a non-linear compensation process on the power distribution parameters based on the voltage fluctuation mode feature; Perform combined optimization on the calibrated trigger condition parameter set and the compensated power distribution parameters to generate the optimized charge and discharge strategy set; the combined optimization includes parameter synchronization verification and conflict elimination processing.
5. The method according to claim 1, wherein Execute the charge and discharge operation based on the optimized charge and discharge strategy set, and collect the battery performance feedback data after the charge and discharge operation to iteratively update the policy matching parameters of the reinforcement learning policy network, including: When executing each charge and discharge strategy in the optimized charge and discharge strategy set, monitor the deviation degree between the battery output power curve and the policy expected power curve in real time; If the deviation degree exceeds the preset tolerance range, trigger a policy execution anomaly mark, and record the anomaly timestamp and environmental parameter snapshot corresponding to the deviation degree; Conduct an attenuation coefficient analysis on the battery performance feedback data, extract the capacity attenuation rate and internal resistance change rate after the policy execution, and input the capacity attenuation rate and internal resistance change rate into the policy effectiveness evaluation model to generate a policy effectiveness score; Adjust the policy matching parameters of the reinforcement learning policy network according to the comparison result between the policy effectiveness score and the preset score threshold; the adjustment includes updating the policy priority weight and correcting the policy matching rule.
6. The method according to claim 2, wherein Perform a time-domain trend analysis on the current parameter sequence in the battery state parameter set, and extract the fluctuation period feature and peak interval distribution feature of the current parameter sequence, including: Perform a sliding window segmentation process on the current parameter sequence to generate multiple current parameter subsequences; Perform local extreme value detection on each current parameter subsequence, and extract the peak point position and the time interval data between adjacent peak points; Statistically calculate the distribution histogram of all time interval data, and calculate the kurtosis coefficient and skewness coefficient of the distribution histogram; Generate the peak interval distribution feature according to the kurtosis coefficient and skewness coefficient; Perform a Fourier transform process on the current parameter sequence, extract the main frequency components and their corresponding amplitudes, and generate the fluctuation period feature according to the reciprocal of the main frequency components.
7. The method according to claim 3, characterized in that, Performing policy matching degree prediction on the load feature set through the policy evaluation module in the reinforcement learning policy network, and outputting the initial priority scores of each candidate charge and discharge policy, including: Separating the steady-state and transient components of the current change trend feature in the load feature set, and extracting the steady-state current parameter sequence and the transient current impact feature; Performing weight allocation processing on the main energy distribution parameter and the secondary fluctuation component parameter in the voltage fluctuation mode feature to generate a voltage fluctuation sensitivity coefficient; Based on the cumulative response energy distribution feature of the temperature correlation response feature, performing thermal effect compensation processing on the steady-state current parameter sequence to generate a temperature-calibrated steady-state current sequence; Performing time-series superposition processing on the transient current impact feature and the temperature-calibrated steady-state current sequence to generate a composite current feature vector; Performing cross-dimensional correlation analysis on the composite current feature vector and the voltage fluctuation sensitivity coefficient to construct a current-voltage joint fluctuation feature matrix; Inputting the current-voltage joint fluctuation feature matrix into the feature matching layer of the policy evaluation module, and comparing the similarity with the load adaptation feature template in the historical policy execution record to generate the initial matching degree score of the candidate charge and discharge policy; Performing policy conflict detection processing on the initial matching degree score, eliminating the policy matching results that are incompatible with the current battery state parameter set, and outputting the initial priority score.
8. The method according to claim 4, characterized in that Calibrating the charge and discharge trigger condition parameters according to the current change trend feature in the load feature set to generate a calibrated trigger condition parameter set, including: Extracting the fluctuation period feature and the peak interval distribution feature in the current change trend feature, and generating a current stability parameter based on the period stability coefficient of the fluctuation period feature and the dispersion degree of the peak interval distribution feature; Determining the adjustment ratio of the charge and discharge trigger threshold according to the difference between the current stability parameter and the preset stability reference value, and performing linear scaling processing on the current threshold in the charge and discharge trigger condition parameters based on the adjustment ratio to generate a calibrated current trigger threshold; Extracting the energy ratio relationship between the main energy distribution parameter and the secondary fluctuation component parameter in the voltage fluctuation mode feature, and performing range expansion on the voltage fluctuation tolerance in the charge and discharge trigger condition parameters based on the energy ratio relationship to generate an extended voltage tolerance parameter; Performing parameter synchronization verification processing on the calibrated current trigger threshold and the extended voltage tolerance parameter, and detecting whether the joint effective interval of the current threshold and the voltage tolerance covers the fluctuation range of the load fluctuation parameter set to generate a conflict-free trigger parameter combination; Performing thermal inertia compensation processing on the conflict-free trigger parameter combination based on the cumulative response energy distribution feature of the temperature correlation response feature, and correcting the influence of temperature delay on the effective timing of the trigger condition to generate a temperature-corrected trigger parameter set; For the temperature-corrected trigger parameter set, screening out the trigger condition parameters that match the fluctuation frequency and amplitude in the load fluctuation parameter set to generate the calibrated trigger condition parameter set.
9. The method according to claim 5, characterized in that, Performing attenuation coefficient analysis on the battery performance feedback data, and extracting the capacity attenuation rate and internal resistance change rate after the execution of the strategy, including: After each charge-discharge strategy is executed, collecting the battery charge-discharge cycle times and the capacity test data sequence; Performing linear regression analysis on the capacity test data sequence, and calculating the capacity decline slope corresponding to the unit cycle times as the capacity attenuation rate; Measuring the DC internal resistance value of the battery under different state of charge, and constructing an internal resistance-state of charge relationship curve; Calculating the curvature change amount of the internal resistance-state of charge relationship curve, and generating the internal resistance change rate according to the integral result of the curvature change amount; Performing standardization processing on the capacity attenuation rate and the internal resistance change rate, and generating an input vector for evaluating the strategy effect.
10. A computer system, characterized in that, Including: A memory, in which a computer program is stored; A processor, configured to load the computer program to implement the method for optimizing the battery charge-discharge strategy of the supercomputer center based on reinforcement learning according to any one of claims 1-9.
Citation Information
Patent Citations
Data center battery charging and discharging optimization control method and device
CN113131584A
Equalization method of energy storage battery pack management system based on neural network and medium
CN117613421A
Battery charging and discharging control method and device, equipment and storage medium
CN117691713A
AI-based big data distributed computing task automatic optimization method and system
CN119576507A
Battery state monitoring analysis method and system based on machine learning
CN119881669A
Cited By
Direct current screen switching method and system based on reinforcement learning
CN120511860A
Intelligent energy scheduling method and system of flow battery
CN120565729A
Intelligent management method and system for optimizing cycle life of storage and charging equipment
CN120655067A
Motor load adaptive adjustment method based on analog circuit
CN122092742A