Charging and discharging strategy intelligent optimization method and system applied to energy storage power supply

By performing time-series segmentation and feature extraction on the charging and discharging process of energy storage power sources, and combining pattern clustering and reinforcement learning, a highly adaptive charging and discharging strategy is generated. This solves the problems of optimization lag and insufficient accuracy of energy storage power sources under dynamic load environments, and achieves more efficient strategy adaptation.

CN120855435AInactive Publication Date: 2025-10-28GUANGDONG BEIBEI ELECTRIC TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511121241.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-28
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture the dynamic changes during the charging and discharging process of energy storage power sources, resulting in insufficient adaptability of charging and discharging strategies to grid load demands, especially in scenarios with frequent load fluctuations where optimization lags or lacks accuracy.

Method used

By performing time-series segmentation processing on the historical charging and discharging process of the energy storage power source, a data sequence of charging and discharging periods with continuous time stamps is generated. Charging and discharging features are extracted, pattern clustering analysis is performed, and a charging and discharging power allocation scheme adapted to different charging and discharging pattern clusters is generated by combining reinforcement learning algorithms, and the power allocation parameters are adjusted in real time.

Benefits of technology

It improves the optimization accuracy and environmental adaptability of energy storage power supply charging and discharging strategies, enhances the ability to adapt to complex behavior patterns and dynamic load environments, and reduces the interference of noise on strategy optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120855435A_ABST
    Figure CN120855435A_ABST
Patent Text Reader

Abstract

The invention provides a charging and discharging strategy intelligent optimization method and system applied to an energy storage power supply, and the method comprises the steps: carrying out the time sequence segmentation processing of a historical charging and discharging process of the energy storage power supply, employing a sliding time window to traverse historical charging and discharging data, and generating a charging and discharging time period data sequence with a continuous time stamp; performing charging and discharging feature extraction on the charging and discharging time period data sequence to obtain a charging and discharging feature set corresponding to each charging and discharging time period, performing mode clustering analysis based on the charging and discharging feature sets, dividing the charging and discharging time periods with similar feature distribution into a plurality of charging and discharging mode clusters through time sequence clustering, and performing mode clustering analysis on the charging and discharging mode clusters; and carrying out association matching on the charging and discharging mode clusters and power grid load prediction information, carrying out strategy iteration optimization processing on an association matching result through reinforcement learning, and generating a charging and discharging power distribution scheme adapted to different charging and discharging mode clusters. According to the invention, the optimization accuracy and environmental adaptability of the charging and discharging strategy of the energy storage power supply can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of charging control, and more specifically, to a method and system for intelligent optimization of charging and discharging strategies for energy storage power supplies. Background Technology

[0002] With the increasing penetration of renewable energy and the development of smart grid technology, energy storage power supplies, as key devices for balancing grid load and improving energy efficiency, have seen increasing attention paid to their charging and discharging strategy optimization. By regulating the charging and discharging process of energy storage power supplies, grid operating efficiency can be improved and energy storage resources can be utilized more rationally. Currently, methods typically extract features from historical charging and discharging data and then allocate charging and discharging power using mathematical models or preset rules to adapt to changes in grid load. However, these methods often struggle to fully capture the dynamic changes during the charging and discharging process, leading to insufficient adaptability between the charging and discharging strategy and the actual load demand of the grid. This is especially true in scenarios with diverse charging and discharging behavior patterns and frequent grid load fluctuations, where traditional methods are prone to optimization lag or insufficient accuracy. Therefore, improving the adaptability of energy storage power supply charging and discharging strategies to complex behavior patterns and dynamic load environments has become a pressing technical problem in the field of energy storage system optimization. Summary of the Invention

[0003] This invention provides a method and system for intelligent optimization of charging and discharging strategies for energy storage power supplies.

[0004] In a first aspect, embodiments of the present invention provide an intelligent optimization method for charging and discharging strategies applied to energy storage power supplies. The method includes: performing time-series segmentation processing on the historical charging and discharging process of the energy storage power supply; traversing the historical charging and discharging data using a sliding time window to generate a charging and discharging time period data sequence with continuous time markers; performing charging and discharging feature extraction on the charging and discharging time period data sequence to obtain a charging and discharging feature set corresponding to each charging and discharging time period, wherein the charging and discharging feature set contains feature information reflecting the changing patterns of the charging and discharging process; performing pattern clustering analysis based on the charging and discharging feature set, dividing charging and discharging time periods with similar feature distributions into multiple charging and discharging pattern clusters through time-series clustering, wherein each charging and discharging pattern cluster corresponds to a type of charging and discharging behavior pattern; associating and matching the charging and discharging pattern clusters with grid load prediction information, and performing strategy iterative optimization processing on the association matching results through reinforcement learning to generate a charging and discharging power allocation scheme adapted to different charging and discharging pattern clusters, wherein the charging and discharging power allocation scheme is used to dynamically adjust the real-time charging and discharging execution process of the energy storage power supply.

[0005] Secondly, embodiments of the present invention provide a computer system, including: a memory storing a computer program; and a processor for loading the computer program to implement the intelligent optimization method for charging and discharging strategies applied to energy storage power sources as described above.

[0006] The present invention provides an intelligent optimization method for charging and discharging strategies of energy storage power supplies. This method performs time-series segmentation processing on the historical charging and discharging processes of the energy storage power supply, using a sliding time window to traverse historical charging and discharging data to generate a data sequence of charging and discharging periods with continuous time markers. By preserving the temporal order correlation of data in each period, it provides a structured time-series data foundation for subsequent feature extraction, avoiding the loss of temporal correlation caused by traditional fixed-period segmentation. Charging and discharging feature extraction is performed on the charging and discharging period data sequence to obtain a set of charging and discharging features reflecting the changing patterns of the charging and discharging process. Based on the charging and discharging feature set, pattern clustering analysis is performed. Through time-series clustering, charging and discharging periods with similar feature distributions are divided into multiple charging and discharging pattern clusters. This achieves an abstract representation from raw data to charging and discharging behavior patterns, elevating the charging and discharging process from the data level to the pattern level, facilitating the capture of dynamic changes in charging and discharging behavior, and reducing the interference of noise from single data points on strategy optimization. By associating and matching charging and discharging mode clusters with grid load forecasting information, and using reinforcement learning algorithms to iteratively optimize the association and matching results, charging and discharging power allocation schemes adapted to different charging and discharging mode clusters are generated. By establishing a high-order mapping relationship between charging and discharging behavior patterns and grid load demand, the adaptability of charging and discharging strategies to grid load change trends is improved. At the same time, by utilizing the dynamic iteration mechanism of reinforcement learning, the power allocation parameters corresponding to different charging and discharging mode clusters are adjusted in real time, avoiding the limitations of traditional fixed models that are difficult to adapt to charging and discharging behaviors in multiple scenarios, and improving the optimization accuracy and environmental adaptability of energy storage power supply charging and discharging strategies. Attached Figure Description

[0007] Figure 1 This is a flowchart of an intelligent optimization method for charging and discharging strategies of energy storage power sources provided in an embodiment of the present invention.

[0008] Figure 2 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present invention. Detailed Implementation

[0009] Please see Figure 1 The flowchart below illustrates an intelligent optimization method for charging and discharging strategies of energy storage power supplies, provided by an embodiment of the present invention. This method can be executed by a computer system and may include the following steps:

[0010] Step S100: Perform time series segmentation processing on the historical charging and discharging process of the energy storage power supply, use a sliding time window to traverse the historical charging and discharging data, and generate a charging and discharging period data sequence with continuous time markers. The data of each period in the charging and discharging period data sequence maintains the time sequence correlation.

[0011] Time series segmentation is the process of dividing continuous time series data into multiple time periods. Its purpose is to better analyze and process the data, reduce data complexity, and highlight the characteristics of the data in different time periods. A sliding window is a window with a fixed length that slides across the time series data, starting from the beginning and sliding sequentially with a certain step size. After each slide, data within the window is extracted for processing. Historical charge / discharge data is the charging and discharging data of the energy storage power source over a past period, including information such as charging and discharging time, capacity, and power. This data records the charging and discharging behavior history of the energy storage power source. The charge / discharge period data sequence is a data sequence generated after time series segmentation and sliding window traversal. Each period of data has a continuous time stamp, and the data from each period are interconnected in chronological order.

[0012] For example, first, the length and sliding step of the sliding time window are determined. Then, starting from the beginning of the historical charge / discharge data, the data covered by the sliding time window is treated as a time period, and the time range and charge / discharge related information, such as battery level and power, are recorded for that time period. Next, the sliding time window is slid to the right according to the set step, and the data within the window is captured again as the next time period. This process is repeated until the sliding time window has traversed all the historical charge / discharge data. During the traversal, continuous time stamps are added to each time period to ensure that the data of each time period maintains a temporal sequence.

[0013] Step S200: Perform charge and discharge feature extraction on the data sequence of charge and discharge periods to obtain the charge and discharge feature set corresponding to each charge and discharge period. The charge and discharge feature set contains feature information that reflects the change law of the charge and discharge process.

[0014] Charging and discharging feature extraction is the process of extracting key features that reflect the changing patterns of the charging and discharging process from a data sequence of charging and discharging periods. The charging and discharging feature set is a collection of various extracted charging and discharging features, which includes characteristic information reflecting the changing patterns of the charging and discharging process, such as waveform features and trend features.

[0015] In one implementation, step S200 may specifically include the following steps S210 to S240:

[0016] Step S210: Perform waveform morphology analysis on the data of each time period in the charging and discharging time period data sequence, and extract the waveform features of the charging and discharging curves in each time period. The waveform features include the distribution position of the extreme points of the charging and discharging curves and the order in which the curve inflection points appear.

[0017] Waveform morphology analysis is the process of analyzing the shape and characteristics of charge-discharge curves. By analyzing the waveform morphology of the curve, we can understand the changes in the charge-discharge process. A charge-discharge curve is a curve showing the change in charge or power over time during the charge-discharge process of an energy storage power source. Waveform characteristics are parameters reflecting the shape and characteristics of the charge-discharge curve, including the distribution location of extreme points and the order inflection points. The distribution location of extreme points refers to the location of the maximum and minimum values ​​on the time axis of the charge-discharge curve; these extreme points reflect the peak and trough values ​​during the charge-discharge process. The order inflection points refers to the chronological order in which the curve's concavity / convexity changes (i.e., inflection points). The appearance of an inflection point may indicate a change in the trend of the charge-discharge process.

[0018] For example, the data for each time period in the charge / discharge time series is first preprocessed to ensure data quality and reliability. Then, waveform morphology analysis is performed on the charge / discharge curve for each time period. Next, extreme points and inflection points on the charge / discharge curve are identified. For example, extreme points can be found by using the method of the first derivative being zero. When the first derivative of the charge / discharge curve is zero, the point may be a maximum or minimum point. The sign of the second derivative is then used to determine whether it is a maximum or minimum point. For inflection points, the second derivative of the charge / discharge curve can be calculated. When the second derivative is zero and its sign changes, the point is an inflection point. The position of each extreme point and inflection point on the time axis is recorded, and the inflection points are arranged in chronological order to obtain the waveform characteristics of the charge / discharge curve.

[0019] In one implementation, step S210 may specifically include the following steps S211 to S215:

[0020] Step S211: Perform curve smoothing on the data of each time period in the charging and discharging time period data sequence to eliminate high-frequency noise interference in the charging and discharging curve and obtain the smoothed charging and discharging curve.

[0021] The purpose of curve smoothing is to reduce noise interference in the curve, making it smoother. High-frequency noise interference refers to high-frequency noise signals present in the charge-discharge curve. These noise signals may be caused by factors such as measurement errors and environmental interference, and can affect the accurate extraction of charge-discharge curve features. The smoothed charge-discharge curve is obtained after curve smoothing, removing most of the high-frequency noise interference and better reflecting the true changes in the charge-discharge process.

[0022] Exemplarily, a variety of curve smoothing methods can be adopted, such as the moving average method, the Gaussian filtering method, etc. Taking the moving average method as an example, assuming that a period of data in the charge-discharge period data sequence contains n data points, and the window size of the moving average is set to k (k < n). Starting from the first data point, calculate the average value of k data points within the window, and take this average value as the smoothed value of the data point at the center of the window. Then, move the window one data point to the right, calculate the average value of k data points within the window again, and repeat this process until the window traverses all the data points. In this way, the smoothed charge-discharge curve can be obtained.

[0023] Step S212: Perform an extreme point detection operation on the smoothed charge-discharge curve to identify the maximum and minimum points on the charge-discharge curve, and record the distribution positions of each extreme point on the time axis.

[0024] The extreme point detection operation is a process of finding the maximum and minimum points on the charge-discharge curve. The maximum point refers to a point on the charge-discharge curve where the value of this point is greater than the values of its adjacent points; the minimum point refers to a point on the charge-discharge curve where the value of this point is less than the values of its adjacent points. Recording the distribution positions of each extreme point on the time axis is for subsequent analysis of the peak and valley situations in the charge-discharge process, as well as the time rules when they appear.

[0025] Exemplarily, perform extreme point detection based on the smoothed charge-discharge curve. The method of the first derivative being zero can be used to find the extreme points. First, perform numerical differentiation on the smoothed charge-discharge curve to calculate its first derivative. When the first derivative is zero, this point may be a maximum point or a minimum point. Then, judge whether it is a maximum point or a minimum point by calculating the second derivative of this point. If the second derivative is less than zero, this point is a maximum point; if the second derivative is greater than zero, this point is a minimum point. For example, for a smoothed charge-discharge curve, use the difference method to calculate its first derivative and second derivative. Assume that the function of the curve is y = f(x), where x represents time and y represents the charge-discharge power or energy. For two adjacent data points (x1, y1) and (x2, y2), the first derivative can be approximately expressed as (y2 - y1) / (x2 - x1). When the calculated first derivative is zero, then calculate the second derivative of this point, and judge whether it is a maximum point or a minimum point by the positive or negative of the second derivative. Finally, record the position of each extreme point on the time axis for subsequent analysis.

[0026] Step S213: Calculate the second derivative of the smoothed charge-discharge curve, determine the inflection point positions of the curve according to the positive and negative changes of the second derivative, and arrange the inflection point positions in chronological order to obtain the order of appearance of the curve inflection points.

[0027] The second derivative calculation involves taking the second derivative of the smoothed charge-discharge curve. The second derivative reflects the concavity / convexity of the curve; a positive second derivative indicates a concave curve, while a negative second derivative indicates a convex curve. An inflection point is a point on the charge-discharge curve where the curve's concavity / convexity changes; that is, the point where the second derivative is zero and its sign changes. The order in which inflection points appear refers to the chronological order in which these points occur, reflecting the turning points in the trend of the charge-discharge process.

[0028] For example, firstly, the second derivative of the smoothed charge-discharge curve is calculated. The second derivative can be calculated using numerical differentiation. Then, the values ​​of the second derivative are iterated through; when the second derivative is zero and its sign changes, that point is the inflection point of the curve. The position of each inflection point on the time axis is recorded, and these inflection point positions are arranged in chronological order to obtain the order in which the curve's inflection points appear.

[0029] Step S214: Associate and store the distribution locations of extreme points and the order in which curve inflection points appear, and generate a waveform feature descriptor containing location coordinates and sequence markers.

[0030] Associative storage combines the locations of extreme points and the order in which curve inflection points appear. A waveform feature descriptor is a set of information describing the waveform characteristics of a charge / discharge curve, containing the coordinates of the extreme point locations and a marker indicating the order in which curve inflection points appear. Location coordinates refer to the specific position of the extreme point on the time axis, expressed in units of time. The sequence marker is the number of the curve inflection points according to their chronological order, used to identify the order in which the inflection points appear. For example, the recorded locations of extreme points and the order in which curve inflection points appear are integrated. This associative storage can be achieved using a database or data structure. For instance, a table can be created with columns for "Extreme Point Location Coordinates," "Curve Inflection Point Location Coordinates," and "Curve Inflection Point Sequence Marker." The location coordinates of each extreme point are entered into the "Extreme Point Location Coordinates" column, the location coordinates of each curve inflection point are entered into the "Curve Inflection Point Location Coordinates" column, and a sequence marker is assigned to each curve inflection point according to its order of appearance and entered into the "Curve Inflection Point Sequence Marker" column. This generates a waveform feature descriptor containing location coordinates and sequence markers.

[0031] Step S215: Convert the waveform feature descriptor into a standardized feature vector form as the waveform feature of the charge-discharge curve.

[0032] Standardized feature vector form is a vector representation of waveform feature descriptors after a unified format conversion, making the waveform features of different charge-discharge curves comparable and computable. Converting waveform feature descriptors into standardized feature vector form facilitates subsequent processing and analysis by machine learning algorithms.

[0033] For example, a normalization method can be used to standardize the position coordinates and sequence markers in the waveform feature descriptor. For instance, the position coordinates of extreme points and curve inflection points can be normalized to the interval [0,1]. Assuming the range of extreme point position coordinates is from t1 to t2, for a specific extreme point position coordinate t, normalization can be performed using the formula (t-t1) / (t2-t1). For the curve inflection point sequence markers, they can be converted into binary or one-hot encoding. Then, the standardized position coordinates and sequence markers are combined into a vector, which is the standardized feature vector form of the waveform feature.

[0034] Step S220: Perform trend change analysis on the data of each time period in the charging and discharging time period data sequence, and calculate the trend characteristics of the charging and discharging rate in each time period. The trend characteristics include the duration of the rising phase and the duration of the falling phase of the charging and discharging rate.

[0035] Trend analysis involves analyzing data from different time periods within a charge / discharge time series to understand the trend of charge / discharge rate changes over time. Charge / discharge rate refers to the change in the amount of charge / discharge by the energy storage power source per unit time, reflecting the speed of charging and discharging. Trend characteristics are feature parameters describing the trend of charge / discharge rate changes, including the duration of the rising and falling phases of the charge / discharge rate. The duration of the rising phase refers to the length of the period from a lower value to a higher value; the duration of the falling phase refers to the length of the period from a higher value to a lower value.

[0036] For example, firstly, the charge / discharge rate is calculated for each time period in the charge / discharge time series. Then, trend analysis is performed on the calculated charge / discharge rate series to determine the rising and falling phases of the charge / discharge rate, and their durations are calculated. For example, for a 24-hour charge / discharge time series, the hourly charge / discharge rate is calculated. Next, by analyzing the changes in the charge / discharge rate series, the start and end times of the rising and falling phases are determined, thereby calculating the durations of the rising and falling phases.

[0037] In one implementation, step S220 may specifically include the following steps S221 to S226:

[0038] Step S221: Calculate the charge and discharge rate for each time period in the charge and discharge time period data sequence, and obtain the change value of charge and discharge amount per unit time as the charge and discharge rate. Perform three exponential smoothing processes on the charge and discharge rate to eliminate short-term fluctuation interference and obtain a smoothed charge and discharge rate sequence.

[0039] Charge / discharge rate calculation involves determining the change in charge / discharge amount per unit time based on data from different time periods within a charge / discharge time series. Triple exponential smoothing is a method for smoothing time series data. By weighted averaging of historical data, it effectively eliminates short-term fluctuations and highlights the long-term trend of the data. The smoothed charge / discharge rate series, obtained after triple exponential smoothing, removes most of the short-term fluctuations and better reflects the long-term trend of charge / discharge rates.

[0040] For example, the charge / discharge rate is first calculated for each time period in the charge / discharge time period data sequence. Assuming the data sequence is recorded at hourly intervals, for two adjacent time periods, the charge / discharge rate is obtained by subtracting the charge / discharge amount of the previous period from the amount of the charge / discharge in the later period, and then dividing by the time interval (1 hour). For example, for a 24-hour charge / discharge time period, the hourly charge / discharge rate is calculated. Then, the calculated charge / discharge rate sequence is subjected to triple exponential smoothing. The formula for triple exponential smoothing is as follows:

[0041] Let the original charge / discharge rate sequence be {y1, y2, ..., yn}, and the first exponential smoothing value... The calculation formula is:

[0042] Where t = 1, 2, ..., n, α is the smoothing coefficient, taking values ​​in the range [0, 1], and y t These are the observations of the original time series at time 10:00. This is the first exponentially smoothed value from the previous time step. First exponential smoothing can be used to eliminate random fluctuations in time series data and provide a preliminary reflection of the data's trend.

[0043] Quadratic exponential smoothing value This value is obtained by performing exponential smoothing again on the value obtained from the first exponential smoothing. The calculation formula is as follows:

[0044]

[0045] Triple exponential smoothing value It is the value obtained by performing a third exponential smoothing process on the second exponentially smoothed value. The calculation formula is as follows:

[0046]

[0047] Through iterative calculations, a smoothed charge / discharge rate sequence is finally obtained.

[0048] Step S222: Perform multi-scale trend decomposition on the smoothed charge and discharge rate sequence. Decompose the rate sequence into multiple intrinsic mode function components and residual trend terms using the ensemble empirical mode decomposition algorithm. Extract the residual trend terms that reflect the long-term trend as the principal components for trend analysis.

[0049] Multi-scale trend decomposition is the process of decomposing a smoothed charge / discharge rate sequence at different scales to separate different frequency components. Ensemble Empirical Mode Decomposition (EEMD) is an adaptive signal decomposition method that can decompose complex time-series signals into multiple Intrinsic Mode Function (IMF) components and a residual trend term. IMF components are oscillating components with different frequencies and amplitudes, reflecting the local characteristics of the signal at different time scales. The residual trend term is the signal component remaining after multiple decompositions, reflecting the long-term trend of the signal. The principal components for trend analysis are the main components extracted from the decomposition results for trend analysis; for example, the residual trend term can be selected as the principal components for trend analysis.

[0050] In one implementation, step S222 may specifically include the following steps S2221 to S2225:

[0051] Step S2221: Add white noise perturbation with a preset amplitude to the smoothed charge and discharge rate sequence to generate an extended rate sequence set containing different noise implementations. The preset amplitude is adaptively adjusted according to the dynamic range of the rate sequence.

[0052] White noise perturbation is a noise signal with random characteristics, whose amplitude is uniformly distributed across the entire frequency range. The purpose of adding white noise perturbation is to overcome the mode aliasing problem that may occur when the Empirical Mode Decomposition (EMD) algorithm processes non-stationary signals. The preset amplitude refers to the pre-defined magnitude of the white noise perturbation, which is adaptively adjusted based on the dynamic range of the smoothed charge / discharge rate sequence to ensure that the added white noise perturbation effectively improves the decomposition effect without excessively affecting the original signal. The extended rate sequence set is a collection of multiple charge / discharge rate sequences containing different noise implementations generated after adding white noise perturbation. For example, firstly, the dynamic range of the smoothed charge / discharge rate sequence is calculated, i.e., the difference between the maximum and minimum values. Then, the preset amplitude is adaptively determined based on the dynamic range. For example, the preset amplitude can be set to a certain proportion of the dynamic range, such as 10%. Next, white noise perturbation with the preset amplitude is added to the smoothed charge / discharge rate sequence. A white noise sequence following a normal distribution with a mean of 0 and a standard deviation equal to the preset amplitude can be generated using a random number generator. The generated white noise sequence is added to the smoothed charge / discharge rate sequence to obtain a noisy extended rate sequence. This process is repeated multiple times, generating a different white noise sequence each time, thus producing a set of extended rate sequences with different noise implementations.

[0053] Step S2222: Perform empirical mode decomposition on each sequence in the extended rate sequence set, generate upper and lower envelopes by cubic spline interpolation, calculate the difference between the mean of the envelope and the original sequence, iteratively filter the components that satisfy the intrinsic mode function conditions until the remaining sequence is a monotonic function or a constant, and obtain multiple intrinsic mode function components and corresponding residual trend terms.

[0054] Empirical Mode Decomposition (EMD) is the process of decomposing each sequence in a set of expanding rate sequences into multiple intrinsic mode function (IMF) components and a residual trend term. Cubic spline interpolation is an interpolation method used to generate upper and lower envelopes. It fits the curve by constructing a cubic polynomial between data points, resulting in a curve with good smoothness and approximation. The upper and lower envelopes refer to the upper envelope of the maximum point and the lower envelope of the minimum point, respectively, fitted using cubic spline interpolation at each local extremum point of the expanding rate sequence. The envelope mean is the average value of the corresponding points on the upper and lower envelopes. Iterative selection involves continuously calculating the difference between the envelope mean and the original sequence, repeating this process until components satisfying the IMF conditions are selected. The IMF conditions require a component to satisfy two conditions: first, the number of extrema and the number of zero-crossing points in the entire data sequence are equal or differ by at most 1; second, at any point, the average of the upper envelope defined by the local maximum point and the lower envelope defined by the local minimum point is zero. The residual trend term is the sequence remaining after multiple iterations of filtering. It is a monotonic function or a constant that reflects the long-term trend of the expansion rate sequence.

[0055] For example, for each sequence in the set of expanding rate sequences, its upper and lower envelopes are first generated using cubic spline interpolation. For instance, for an expanding rate sequence, all its local maxima and minima are identified, and then the upper and lower envelopes are fitted using cubic spline interpolation. Next, the mean of the upper and lower envelopes is calculated, and the mean of the envelopes is subtracted from the original sequence to obtain a new sequence. It is then determined whether this new sequence satisfies the intrinsic mode function (IMF) condition. If it does, it is used as an IMF component; otherwise, it is used as a new original sequence, and the above steps are repeated until a component satisfying the IMF condition is obtained. This process iterates until the remaining sequence is a monotonic function or a constant; this remaining sequence is the residual trend term. Ultimately, for each expanding rate sequence, multiple IMF components and one residual trend term can be obtained.

[0056] Step S2223: Calculate the statistical average of the intrinsic mode function components of the same order under different noise levels, eliminate the influence of noise interference, and obtain the set of intrinsic mode function components after noise reduction. The statistical average is calculated by the arithmetic mean method or the median mean method.

[0057] The statistical average is the value obtained by averaging the intrinsic mode function components of the same order under different noise levels. It can eliminate the noise interference caused by added white noise perturbation. The set of intrinsic mode function components after noise reduction is the set of intrinsic mode function components that have been eliminated from noise interference after statistical averaging.

[0058] For example, for the intrinsic mode function components of the same order obtained from the decomposition of each sequence in the extended rate sequence set, their statistical average is calculated using the arithmetic mean method or the median mean method. This process is repeated to calculate the statistical average for the intrinsic mode function components of each order, finally obtaining the denoised intrinsic mode function component set.

[0059] Step S2224: Perform Fourier transform on the set of intrinsic mode function components after noise reduction to obtain the amplitude-frequency characteristic curves of each component, and extract the characteristic frequency range of the curves and the energy proportion of the corresponding frequency bands.

[0060] The Fourier transform converts the set of intrinsic mode function components after noise reduction from the time domain to the frequency domain, thus obtaining the amplitude-frequency response curves of each component. The amplitude-frequency response curve describes the amplitude of the signal at different frequencies, reflecting the frequency components and energy distribution of the signal. The characteristic frequency range refers to the frequency intervals with distinct characteristics on the amplitude-frequency response curve; these characteristic frequencies usually correspond to the main frequency components of the signal. The energy proportion of the corresponding frequency band refers to the proportion of signal energy to total energy within the characteristic frequency range, which helps to understand the contribution of different frequency components to the signal.

[0061] For example, a Fourier transform is performed on each component in the denoised intrinsic mode function component set. A Fast Fourier Transform (FFT) algorithm can be used to improve computational efficiency. Then, the amplitude-frequency response curves of each component are plotted. Next, the characteristic frequency range is determined by analyzing the amplitude-frequency response curves. The characteristic frequency range can be determined based on the peaks and troughs of the curves; typically, the frequency range near the peak is chosen as the characteristic frequency range. Finally, the energy percentage within the characteristic frequency range is calculated. This can be obtained by integrating the power spectral density within the characteristic frequency range and then dividing by the integral of the total power spectral density.

[0062] Step S2225: Divide the intrinsic mode function components into high-frequency component group, mid-frequency component group and low-frequency component group according to the characteristic frequency range. Calculate the correlation coefficient between the low-frequency component group and the residual trend term. Perform weighted superposition processing on the low-frequency component with the highest correlation coefficient and the residual trend term. The weights are dynamically allocated according to the energy proportion to generate the main component of trend analysis.

[0063] The high-frequency, mid-frequency, and low-frequency component groups are three groups of components divided according to the characteristic frequency range of the intrinsic mode functions (IMFs). The high-frequency component group contains IMFs with higher characteristic frequencies, the mid-frequency component group contains IMFs with characteristic frequencies in the middle range, and the low-frequency component group contains IMFs with lower characteristic frequencies. The division can be achieved by setting different thresholds and dividing the data into intervals according to these thresholds. The correlation coefficient is an indicator that measures the degree of linear correlation between two variables, with a value range of [-1, 1]. The closer the absolute value of the correlation coefficient is to 1, the stronger the linear correlation between the two variables. Weighted superposition processing involves adding the low-frequency component with the highest correlation coefficient to the residual trend term according to certain weights. The weights are dynamically allocated based on energy proportions to highlight the components that contribute significantly to trend analysis. The principal components for trend analysis are the main components generated after weighted superposition processing and used for trend analysis.

[0064] For example, the denoised intrinsic mode function component set is first divided into high-frequency component group, mid-frequency component group, and low-frequency component group according to the characteristic frequency range. For instance, components with characteristic frequencies greater than a certain threshold f are... high The components are divided into high-frequency component groups, and the characteristic frequencies are set at a certain threshold f. low The components are divided into low-frequency component groups, and the characteristic frequencies in [f low ,f high The components between [values] are divided into mid-frequency component groups. Then, the correlation coefficient between each component in the low-frequency component group and the residual trend term is calculated. The Pearson correlation coefficient can be used to calculate the correlation coefficient. Next, the low-frequency component with the highest correlation coefficient is identified. Weights are dynamically assigned based on the energy proportion of this low-frequency component and the residual trend term within the characteristic frequency range. Finally, the low-frequency component with the highest correlation coefficient and the residual trend term are weighted and superimposed according to the dynamically assigned weights to generate the principal components for trend analysis.

[0065] Step S223: Perform first-order and second-order difference processing on the principal components of the trend analysis to obtain the rate of change sequence and acceleration change sequence, respectively. Based on the positive and negative values ​​of the rate of change sequence, the charging and discharging rate is initially divided into the rising stage, the falling stage and the stable stage.

[0066] First-order differencing is the process of calculating the difference between adjacent data points of the principal component in trend analysis, yielding a rate of change sequence that reflects the speed of change in charge / discharge rate. Second-order differencing is the process of calculating the difference between adjacent data points of the rate of change sequence, yielding an acceleration sequence that reflects the change in the rate of change of charge / discharge rate. The rate of change sequence, obtained after first-order differencing, indicates whether the charge / discharge rate is increasing or decreasing. The acceleration sequence, obtained after second-order differencing, helps in further analyzing the stability of charge / discharge rate changes. The increasing, decreasing, and stable phases are preliminary divisions of the charge / discharge rate change stages based on the positive or negative values ​​of the rate of change sequence. When the rate of change sequence value is positive, the charge / discharge rate is in the increasing phase; when the rate of change sequence value is negative, the charge / discharge rate is in the decreasing phase; and when the rate of change sequence value is close to zero, the charge / discharge rate is in the stable phase.

[0067] Step S224: Dynamically adjust the initially divided stage boundaries by combining the acceleration change sequence. When the absolute value of the acceleration change sequence exceeds the preset trend threshold, the stage boundary correction mechanism is triggered, and the final stage boundary is determined by verifying the trend consistency of adjacent time periods.

[0068] The stage boundary refers to the dividing point between the rising, falling, and stable phases of the charge / discharge rate. Dynamic adjustment refers to the real-time modification and optimization of the initially defined stage boundaries based on the acceleration change sequence. The preset trend threshold is a pre-set threshold used to determine whether the acceleration change sequence exceeds the normal range. The stage boundary correction mechanism is a mechanism that corrects the stage boundaries when the absolute value of the acceleration change sequence exceeds the preset trend threshold. The trend consistency verification between adjacent time periods refers to the method of determining the final stage boundary by checking whether the charge / discharge rate change trends of adjacent time periods are consistent.

[0069] For example, a preset trend threshold is first set. For instance, the preset trend threshold can be set to a certain multiple of the standard deviation of the acceleration change sequence, such as twice. Then, the acceleration change sequence is traversed, and when the absolute value of the acceleration change sequence exceeds the preset trend threshold, a stage boundary correction mechanism is triggered. When correcting the stage boundary, the charging and discharging rate change trends of adjacent time periods are checked. For example, for a preliminary boundary point between an ascending and descending stage, if the acceleration change sequence near that boundary point exceeds the preset trend threshold, the rate change sequence of adjacent time periods is checked. If the rate change sequence of adjacent time periods shows a consistent trend (e.g., both ascending or both descending), the stage boundary is not modified; if the rate change sequence of adjacent time periods shows an inconsistent trend, the stage boundary is dynamically adjusted based on the trend consistency check result to ensure the accuracy and rationality of the stage division.

[0070] Step S225: During the rising phase, start timing from the time when the rate of change first turns from negative to positive, and end timing from the time when the rate of change turns from positive to negative again or approaches zero, to obtain the duration of the rising phase; during the falling phase, start timing from the time when the rate of change first turns from positive to negative, and end timing from the time when the rate of change turns from negative to positive again or approaches zero, to obtain the duration of the falling phase.

[0071] The duration of the rising phase refers to the length of the period during which the charge / discharge rate rises from a lower value to a higher value. The duration of the falling phase refers to the length of the period during which the charge / discharge rate falls from a higher value to a lower value. The start and end times of the rising and falling phases are determined by recording the positive and negative changes in the rate of change, thus calculating the duration.

[0072] For example, firstly, the rising phase and the falling phase are determined based on the rate of change sequence and the dynamically adjusted phase boundaries. For the rising phase, timing begins from the point when the rate of change first turns from negative to positive. Then, the rate of change sequence is continuously monitored, and timing ends when the rate of change turns from positive to negative again or approaches zero. This time point is recorded, and the duration of the rising phase is obtained by subtracting the start time from the end time. For the falling phase, timing begins from the point when the rate of change first turns from positive to negative, and the end time is recorded in the same way to calculate the duration of the falling phase.

[0073] Step S226: Perform time granularity normalization on the duration of the rising phase and the duration of the falling phase, convert them into relative duration indicators based on the time interval of the time period data sequence, and generate trend characteristics with the same time dimension.

[0074] Time granularity normalization is the process of converting the duration of the rising and falling phases into relative duration indicators based on the time intervals of the time-segment data sequence. This makes the durations of the rising and falling phases comparable and calculable across different charging and discharging periods. The time interval of the time-segment data sequence refers to the time interval between two adjacent data points in the charging and discharging period data sequence. The relative duration indicator is the duration indicator obtained after time granularity normalization, using the time interval of the time-segment data sequence as the unit, ensuring that the changing trend characteristics of different charging and discharging periods have the same time dimension.

[0075] For example, the time interval of the period data sequence is determined. Assume the period data sequence is recorded at hourly intervals. Then, the duration of the rising phase and the duration of the falling phase are divided by the time interval of the period data sequence to obtain a relative duration index in hours. Similarly, the duration of the falling phase is processed to obtain a relative duration index of the falling phase.

[0076] Step S230: Align the waveform features and trend features by feature dimension to ensure that the time dimension labels of the two features are consistent, and generate a preliminary feature set.

[0077] Feature dimension alignment is the process of unifying and matching waveform features and trend features along the time dimension, ensuring a temporal correspondence between the two features. Time dimension markers refer to the labels on the feature data on the time axis, used to represent the corresponding time points or time periods. The preliminary feature set is the feature set generated after aligning the waveform and trend features along their respective dimensions. It contains information from both types of features, and the time dimension markers remain consistent, facilitating further processing and analysis.

[0078] For example, first check the time dimension labels of the waveform features and the trend features. If the time dimension labels of the two features are inconsistent, adjustments are needed. For instance, if the waveform features are recorded in hourly intervals, while the trend features are recorded in half-hourly intervals, the trend features need to be interpolated or aggregated to make their time dimension labels consistent with the waveform features. Linear interpolation can be used to interpolate the trend features to make them also use hourly intervals. Then, combine the aligned waveform features and trend features to generate a preliminary feature set. For example, for charging and discharging period data containing waveform features and trend features, arrange the waveform features and trend features in chronological order to form a new feature vector; this feature vector is the preliminary feature set.

[0079] Step S240: Perform feature redundancy removal processing on the preliminary feature set. Redundant features that repeatedly represent the changing patterns of the charging and discharging process are eliminated through feature correlation analysis to obtain a simplified charging and discharging feature set.

[0080] Feature redundancy removal is the process of processing the initial feature set to remove duplicate or highly correlated features, thereby reducing the dimensionality and complexity of the features and improving the efficiency of subsequent analysis and processing. Feature correlation analysis is a method used to evaluate the correlation between features, judging the degree of correlation by calculating indicators such as correlation coefficients. Redundant features refer to those features in the initial feature set that repeatedly characterize the changes in the charging and discharging process; removing these redundant features will not affect the understanding and analysis of the charging and discharging process. The simplified charging and discharging feature set is the feature set obtained after feature redundancy removal, containing only the key features that can effectively characterize the changes in the charging and discharging process.

[0081] For example, a feature correlation analysis is first performed on the preliminary feature set. The Pearson correlation coefficient can be used to calculate the correlation between features. For instance, for each pair of features in the preliminary feature set, the Pearson correlation coefficient is calculated between them. When the absolute value of the correlation coefficient exceeds a certain preset threshold, the two features are considered highly correlated and redundant. Then, based on the results of the feature correlation analysis, one feature is selected to be retained, and the other redundant feature is removed. The decision to retain a feature can be based on factors such as its importance and stability. This process is repeated until no redundant features remain in the preliminary feature set, resulting in a streamlined charge / discharge feature set.

[0082] Step S300: Perform pattern clustering analysis based on the charging and discharging feature set. Divide charging and discharging periods with similar feature distributions into multiple charging and discharging pattern clusters through time series clustering. Each charging and discharging pattern cluster corresponds to a type of charging and discharging behavior pattern.

[0083] Pattern clustering analysis is an analytical method that groups objects with similar characteristics into the same category. In the method provided in this invention embodiment, the charging and discharging feature set is analyzed, and charging and discharging periods with similar feature distributions are divided into different categories. Time series clustering is a clustering method specifically designed for processing time series data, considering the sequential and dynamic characteristics of time series data. Charging and discharging pattern clusters are different categories obtained after time series clustering. Each charging and discharging pattern cluster contains charging and discharging periods with similar feature distributions, corresponding to a type of charging and discharging behavior pattern. Charging and discharging behavior patterns refer to the behavioral patterns and characteristic patterns exhibited by energy storage power sources during the charging and discharging process.

[0084] In one implementation, step S300 may specifically include the following steps S310 to S350:

[0085] Step S310: Perform time series standardization on each charge-discharge feature in the charge-discharge feature set to ensure that the numerical ranges of different features are consistent, thereby obtaining a standardized charge-discharge feature set.

[0086] Time series standardization is a process of standardizing the scale and range of various charge-discharge features in a charge-discharge feature set. Its purpose is to eliminate differences in numerical ranges between different features, making them comparable. The standardized charge-discharge feature set is the set of charge-discharge features obtained after time series standardization, where the numerical ranges of different features remain consistent.

[0087] For example, the standardization method is first determined. A normalization method can be used to standardize each charge / discharge feature in the charge / discharge feature set to the interval [0,1]. For each charge / discharge feature, its maximum and minimum values ​​are calculated. Assuming a charge / discharge feature ranges from x1 to x2, for each data point x of that feature, normalization can be performed using the formula (x-x1) / (x2-x1). This process is applied to all charge / discharge features in the charge / discharge feature set to obtain the standardized charge / discharge feature set. For example, for a charge / discharge feature set containing multiple features, one of which is charge / discharge power and the other is charge / discharge time, both features are normalized so that their values ​​are within the range [0,1], resulting in the standardized charge / discharge feature set.

[0088] Step S320: Calculate the dynamic time curvature distance between any two charging / discharging time period features in the standardized charging / discharging feature set, as an indicator to measure the similarity of feature distributions.

[0089] Dynamic time warping distance (VTW) is a distance metric used to measure the similarity between two time series. It allows the two time series to be warped and stretched over time to find the optimal matching path, thus more accurately measuring their similarity. In this method, it is used to calculate the similarity between any two charge / discharge period features in a standardized charge / discharge feature set. Feature distribution similarity refers to the degree of similarity in the distribution of two charge / discharge period features in the feature space. The smaller the VTW, the more similar the distributions of the two charge / discharge period features.

[0090] In one implementation, step S320 may specifically include the following steps S321 to S325:

[0091] Step S321: The two charge-discharge period features in the standardized charge-discharge feature set are respectively designated as the first charge-discharge period feature and the second charge-discharge period feature, wherein the first charge-discharge period feature contains multiple first feature points arranged in chronological order, and the second charge-discharge period feature contains multiple second feature points arranged in chronological order.

[0092] The first charge / discharge period feature and the second charge / discharge period feature are two charge / discharge period features selected from the standardized charge / discharge feature set for calculating the dynamic time curvature distance. The first feature point is the feature data points in the first charge / discharge period feature arranged in chronological order, and the second feature point is the feature data points in the second charge / discharge period feature arranged in chronological order.

[0093] For example, two charge / discharge period features are randomly selected from a standardized charge / discharge feature set and labeled as the first charge / discharge period feature and the second charge / discharge period feature, respectively. For instance, for a standardized charge / discharge feature set containing 100 charge / discharge period features, the 10th and 20th charge / discharge period features are selected as the first charge / discharge period feature and the second charge / discharge period feature, respectively. Each feature data point in the first charge / discharge period feature is sequentially labeled as the first feature point in chronological order, and each feature data point in the second charge / discharge period feature is sequentially labeled as the second feature point in chronological order.

[0094] Step S322: Construct a time curvature path matrix between the first charge / discharge period feature and the second charge / discharge period feature. Each element in the time curvature path matrix represents the feature point distance between the first feature point and the second feature point.

[0095] The time-curved path matrix is ​​a two-dimensional matrix used to store the feature point distances of all possible matching paths between the first and second charge-discharge period features. The feature point distance refers to the distance between the first and second feature points in the feature space, and can be calculated using methods such as Euclidean distance.

[0096] For example, the lengths of the first and second charge-discharge period features are first determined. Assume the first charge-discharge period feature contains m first feature points, and the second charge-discharge period feature contains n second feature points. Then, an m×n matrix is ​​constructed as the time-curved path matrix. For each element (i,j) in the matrix, the feature point distance between the i-th first feature point and the j-th second feature point is calculated, such as the Euclidean distance.

[0097] The calculated feature point distances are then filled into the corresponding positions in the time curvature path matrix. For example, for a charge / discharge period feature pair containing 5 first feature points and 6 second feature points, a 5×6 time curvature path matrix is ​​constructed, and the feature point distance for each element is calculated.

[0098] Step S323: Determine the path with minimum cost in the time-curved path matrix through heuristic search. The path with minimum cost meets the time order constraint and has the minimum cumulative path distance.

[0099] Heuristic search is a method for finding the optimal path in a time-curved path matrix. It uses heuristic rules to guide the search process and improve efficiency. The minimum-cost path is a path from the top-left corner to the bottom-right corner of the time-curved path matrix that satisfies the temporal order constraint (i.e., the path can only move right, down, or down-right), and minimizes the sum of the distances between the feature points of all elements on the path. The cumulative path distance is the sum of the distances between the feature points of all elements on the minimum-cost path.

[0100] For example, a heuristic search can be performed using dynamic programming. First, initialize the first row and first column of the time-curved path matrix. For the element (1,j) in the first row, its cumulative distance is the sum of the cumulative distance of the previous element (1,j-1) and the feature point distance of the current element; for the element (i,1) in the first column, its cumulative distance is the sum of the cumulative distance of the previous element (i-1,1) and the feature point distance of the current element. Then, for the other elements (i,j) in the matrix, calculate the cumulative distances from the three possible previous elements (i-1,j), (i,j-1), and (i-1,j-1) to that element, and select the path with the smallest cumulative distance as the path to the current element. Finally, backtrack from the bottom right corner of the matrix to find the path with the minimum cost. For example, for a 5×6 time-curved path matrix, the path with the minimum cost can be found using dynamic programming.

[0101] Step S324: Calculate the sum of the feature point distances of all elements on the path with the minimum cost, and use it as the dynamic time curvature distance between the features of the first charging and discharging period and the features of the second charging and discharging period.

[0102] After determining the minimum-cost path in the time-curvature path matrix, the feature point distances of all elements on that path are summed. The sum is the dynamic time-curvature distance between the features of the first and second charge / discharge periods. This distance reflects the similarity between the two charge / discharge period features.

[0103] For example, all elements on the path with the minimum cost are traversed, and the feature point distances of each element are summed. In this way, the dynamic time curvature distance between the features of the first charge / discharge period and the features of the second charge / discharge period is obtained.

[0104] Step S325: Normalize the dynamic time bending distance to obtain a similarity index with a value range within a set interval. The smaller the similarity index value, the higher the similarity of the distribution of features of the two charging and discharging periods.

[0105] Normalization is the process of adjusting the dynamic time curvature distance to ensure its value falls within a set range, typically normalized to the interval [0,1]. The similarity index, obtained after normalization, measures the similarity of the characteristic distributions of two charging / discharging periods; a smaller value indicates greater similarity in the distributions of the characteristics of the two charging / discharging periods.

[0106] For example, the range of values ​​for the dynamic time bending distance is first determined. Assume the maximum value of the dynamic time bending distance is D. max The minimum value is D min For a specific dynamic time bending distance D, the formula (DD) can be used.min ) / (D max -D min The data is then normalized. The normalized dynamic time curvature distance is used as a similarity index. The smaller the similarity index value, the more similar the distribution of characteristics between the two charging and discharging periods.

[0107] Step S330: Construct a charge-discharge feature similarity matrix based on the dynamic time bending distance. The element values ​​in the similarity matrix represent the degree of similarity between the features of two corresponding charge-discharge periods.

[0108] The charge-discharge feature similarity matrix is ​​a two-dimensional matrix used to store the similarity between any two charge-discharge period features in the standardized charge-discharge feature set. The number of rows and columns of the matrix is ​​equal to the number of charge-discharge period features. The element (i,j) in the matrix represents the similarity between the i-th charge-discharge period feature and the j-th charge-discharge period feature. This similarity is represented by the normalized value of the dynamic time curvature distance (i.e., the similarity index).

[0109] For example, for each pair of charge / discharge period features in the standardized charge / discharge feature set, the dynamic time curvature distance between them is calculated and normalized to obtain a similarity index. Then, the similarity index is filled into the corresponding positions in the charge / discharge feature similarity matrix. This process is repeated until all elements in the matrix are filled, resulting in the charge / discharge feature similarity matrix.

[0110] Step S340: Cluster the charging and discharging feature similarity matrix using a time series clustering algorithm, and divide the charging and discharging periods with similarity higher than a preset threshold into the same clustering unit.

[0111] Time series clustering algorithms are used to cluster time series data, dividing the data into different categories based on the similarity between them. In this method, a time series clustering algorithm is used to cluster the charge-discharge feature similarity matrix, grouping charge-discharge periods with a similarity higher than a preset threshold into the same cluster unit. The preset threshold is a pre-defined threshold used to determine whether the features of two charge-discharge periods are similar; only when the similarity of the features of two charge-discharge periods exceeds this threshold are they grouped into the same cluster unit.

[0112] For example, a suitable time series clustering algorithm is selected, such as hierarchical clustering or K-means clustering. Taking hierarchical clustering as an example, each charging / discharging period feature is first treated as a separate clustering unit. Then, based on the charging / discharging feature similarity matrix, the similarity between clustering units is calculated. Each time, the two clustering units with the highest similarity exceeding a preset threshold are merged into a new clustering unit. This process is repeated until all clustering units meeting the conditions have been merged.

[0113] Step S350: Extract pattern features for each cluster unit, generate pattern center features that characterize the charging and discharging behavior pattern of the cluster unit, and define the cluster unit containing the pattern center features and the corresponding charging and discharging time period as a charging and discharging pattern cluster.

[0114] Pattern feature extraction is the process of processing each cluster unit to extract key features that represent the charging and discharging behavior pattern of that cluster unit. The pattern center feature is a feature vector obtained after pattern feature extraction, used to characterize the charging and discharging behavior pattern of that cluster unit, reflecting the overall characteristics and trends of the charging and discharging time periods within that cluster unit. A charging and discharging pattern cluster is a set formed by combining the pattern center features and the corresponding charging and discharging time periods, representing a class of charging and discharging behavior patterns.

[0115] In one implementation, step S350 may specifically include the following steps S351 to S356:

[0116] Step S351: Extract the standardized charge and discharge feature set of all charge and discharge periods in the cluster unit, take the standardized charge and discharge features of each charge and discharge period as feature samples, and perform time-series alignment processing on the feature samples to make the time axis labels of all feature samples consistent.

[0117] The standardized charge-discharge feature set is a collection of standardized charge-discharge features from all charge-discharge periods within a clustering unit. Feature samples are those where the standardized charge-discharge features of each charge-discharge period are treated as a single sample. Time-series alignment is the process of processing these feature samples to ensure their timeline labels are consistent, facilitating subsequent feature analysis and comparison.

[0118] For example, firstly, a standardized set of charge-discharge features for all charge-discharge periods is extracted from the clustering unit. Then, the standardized charge-discharge features for each charge-discharge period are treated as a feature sample. The timeline markings of the feature samples are checked, and if inconsistent, time-series alignment is performed. For example, some feature samples are recorded at hourly intervals, while others are recorded at half-hourly intervals, requiring the time intervals to be unified. Time-series alignment can be performed using interpolation or aggregation methods. If linear interpolation is used, for feature samples recorded at half-hourly intervals, they are interpolated to hourly intervals to ensure their timeline markings are consistent with other feature samples. Finally, the time-series aligned feature samples are arranged in chronological order to ensure that the timeline markings of all feature samples remain consistent.

[0119] Step S352: Perform anomaly detection on feature samples using the isolated forest algorithm, identify and remove abnormal feature samples that deviate from the cluster center, and obtain a set of normal feature samples.

[0120] The Isolation Forest algorithm identifies outliers in data by constructing random decision trees. Outlier samples are those that deviate significantly from the cluster centers, potentially due to measurement errors, equipment malfunctions, or other reasons, and can negatively impact subsequent feature analysis and processing. The normal feature sample set is the set of features remaining after anomaly detection and removal of outlier samples.

[0121] For example, the parameters of the Isolation Forest algorithm are first initialized, such as the number of trees and the number of samples. Then, the time-aligned feature samples are input into the Isolation Forest algorithm for anomaly detection. The Isolation Forest algorithm calculates an anomaly score for each feature sample; the higher the score, the more likely the sample is to be an anomaly. An anomaly score threshold is set; when the anomaly score of a feature sample exceeds this threshold, it is identified as an anomalous feature sample. Finally, all identified anomalous feature samples are removed, resulting in a set of normal feature samples.

[0122] Step S353: Perform temporal feature alignment on each feature sample in the normal feature sample set, calculate the feature alignment offset between samples, and perform position calibration on the sample feature points based on the feature alignment offset to obtain the calibrated feature samples.

[0123] Temporal feature alignment is a process that further processes each feature sample based on a normal feature sample set to make their features more aligned and consistent. Feature alignment offset refers to the degree of offset between two feature samples in the feature dimension; calculating the feature alignment offset reveals the differences between feature samples. Position calibration is the process of adjusting the position of sample feature points according to the feature alignment offset to eliminate the offset and differences between feature samples, resulting in calibrated feature samples.

[0124] For example, first, a feature sample is selected as a reference sample. For other feature samples in the normal feature sample set, the feature alignment offset between them and the reference sample is calculated. A dynamic time warping algorithm can be used to calculate the feature alignment offset. For each feature sample, the sample feature points are positionally calibrated based on the calculated feature alignment offset. This process is repeated for each feature sample in the normal feature sample set to obtain the calibrated feature sample.

[0125] Step S354: Perform feature fusion on the time dimension of the calibrated feature samples to generate a fused feature sequence containing the correlation of features in multiple time periods. Perform feature dimensionality reduction on the fused feature sequence to obtain the core feature dimension reflecting the charging and discharging behavior pattern.

[0126] Feature fusion is the process of merging and integrating calibrated feature samples along the time dimension to generate a fused feature sequence that contains correlations between features from multiple time periods. The fused feature sequence is a new feature sequence obtained after feature fusion, containing correlation information between features from multiple time periods. Feature dimensionality reduction is the process of processing the fused feature sequence to reduce the dimensionality of the features and extract the core feature dimensions that reflect the charging and discharging behavior patterns.

[0127] In one implementation, step S354 may specifically include the following steps S3541 to S3545:

[0128] Step S3541: Calculate the variance inflation factor of each feature dimension in the fused feature sequence. When the variance inflation factor is greater than a preset threshold, it is marked as a feature dimension with multicollinearity and assigned to the same feature dimension group.

[0129] The variance inflation factor (VIF) is an indicator that measures the degree of multicollinearity among independent variables in a multiple linear regression model. In a fused feature sequence, each feature dimension may be correlated with other feature dimensions. When there is a high degree of correlation among some feature dimensions, multicollinearity occurs, which can affect the stability and accuracy of the model. The VIF is calculated based on the correlation matrix between feature dimensions. For each feature dimension in the fused feature sequence, a linear regression model is constructed with it as the dependent variable and the remaining feature dimensions as independent variables. Then, the coefficient of determination R of this regression model is calculated. 2 The formula for calculating the variance inflation factor is VIF = 1 / (1-R). 2) .

[0130] For example, first, correlation analysis is performed on the fused feature sequences to calculate the correlation coefficients between each feature dimension, constructing a correlation matrix. Then, for each feature dimension, its variance inflation factor is calculated according to the formula above. The preset threshold is a standard pre-set based on the actual situation to determine whether multicollinearity exists in the feature dimensions; for example, the preset threshold can be set to 5 or 10. When the variance inflation factor of a certain feature dimension is greater than the preset threshold, the feature dimension is marked as having multicollinearity and is grouped into the same feature dimension group.

[0131] Step S3542: Perform principal component analysis on each feature dimension group, extract eigenvalues ​​and eigenvectors through eigenvalue decomposition of the covariance matrix, and generate an orthogonal feature space containing multiple principal component components. The principal component components are sorted from largest to smallest variance explained.

[0132] Principal Component Analysis (PCA) is a data dimensionality reduction method that transforms the original features into a set of uncorrelated principal components through linear transformation. These principal components are arranged in descending order of their explained variance, which represents the contribution of each principal component to the variance of the original data. The covariance matrix describes the covariance relationship between feature dimensions. By performing eigenvalue decomposition on the covariance matrix, eigenvalues ​​and eigenvectors can be obtained. Eigenvalues ​​represent the magnitude of the variance of each principal component, while eigenvectors define the orientation of the principal components.

[0133] For example, for each feature dimension group, its covariance matrix is ​​first calculated. Assuming a feature dimension group contains n feature dimensions, then the covariance matrix is ​​an n×n matrix with elements C. ij Let represent the covariance between the i-th and j-th feature dimensions. Then, perform eigenvalue decomposition on the covariance matrix to solve for its eigenvalues ​​and eigenvectors. Numerical methods such as singular value decomposition (SVD) can be used to implement eigenvalue decomposition. After obtaining the eigenvalues ​​and eigenvectors, sort the eigenvectors according to their corresponding eigenvalues ​​in descending order. Each eigenvector corresponds to a principal component. These principal components form an orthogonal eigenspace, meaning that the principal components are independent of each other.

[0134] Step S3543: Calculate the variance explained rate of the principal component components cumulatively. When the cumulative variance explained rate reaches the preset threshold, stop feature selection and retain the corresponding principal component components as the initial dimensionality reduction features. The preset threshold is dynamically adjusted according to the clustering accuracy requirements of the charging and discharging mode.

[0135] The cumulative variance explained rate refers to the value obtained by summing the variance explained rates of each principal component, starting from the first principal component. The preset threshold is a standard pre-set based on the accuracy requirements of charge / discharge pattern clustering, used to determine how many principal components should be retained as initial dimensionality reduction features. By retaining principal components with higher variance explained rates, we can reduce feature dimensionality while preserving as much information as possible from the original data.

[0136] For example, first calculate the variance explained rate for each principal component, which is the proportion of each principal component's eigenvalue to the sum of all eigenvalues. Then, starting with the principal component with the highest variance explained rate, accumulate the variance explained rates sequentially. When the accumulated variance explained rate reaches a preset threshold, feature selection stops, and the corresponding principal component is retained as the initial dimensionality reduction feature. The preset threshold can be dynamically adjusted according to the accuracy requirements of charge / discharge pattern clustering.

[0137] Step S3544: Construct a feature importance evaluation model based on random forest, with the charging and discharging mode cluster identifier as the target variable and the initial dimensionality reduction features as the input variable, and calculate the importance score of each feature dimension through out-of-bag data error.

[0138] Random forests consist of multiple decision trees, and predictions and evaluations are made by combining the results of these decision trees. When constructing a feature importance evaluation model based on random forests, the charge / discharge mode cluster identifier is used as the target variable; that is, the model output is the charge / discharge mode cluster to which each sample belongs. The initial dimensionality-reduced features are used as input variables; that is, the model input consists of features obtained after dimensionality reduction through principal component analysis. Out-of-bag error is a method used in random forests to evaluate feature importance. It calculates the error using samples not used during the training of each decision tree (i.e., out-of-bag data), and evaluates the importance of a feature dimension by comparing the change in out-of-bag error after shuffling a certain feature dimension.

[0139] For example, the data first needs to be partitioned into a training set and a test set. Then, a random forest model is built using the training set, setting parameters such as the number of decision trees and the maximum depth of each tree. During training, each decision tree is trained using a subset of samples, while another subset serves as out-of-bag data. For each feature dimension, after training, the values ​​of that feature dimension are shuffled, and the out-of-bag error is recalculated. If shuffling a feature dimension significantly increases the out-of-bag error, it indicates that that feature dimension has a significant impact on the model's prediction results, and its importance score is high; conversely, if shuffling a feature dimension does not significantly change the out-of-bag error, it indicates that the feature dimension has low importance. In this way, the importance score of each initial dimensionality-reduced feature is calculated.

[0140] Step S3545: Use a recursive feature elimination algorithm to perform a second screening of the initial dimensionality reduction features. Remove feature dimensions in descending order of importance score and evaluate the model classification accuracy. Stop removal when the classification accuracy decreases beyond a preset threshold. Retain the current set of feature dimensions as the core feature dimensions that reflect the charging and discharging behavior pattern.

[0141] Recursive Feature Elimination (RFE) is a feature selection method that progressively removes features. It continuously removes less important features while evaluating the model's classification accuracy until a stopping condition is met. In this step, RFE is used to perform a secondary screening of the initial dimensionality-reducing features to further reduce feature dimensions and improve the model's efficiency and accuracy. For example, the initial dimensionality-reducing features are first sorted from highest to lowest importance based on the importance score obtained from the feature importance evaluation model using a random forest. Then, starting with the feature dimension with the lowest importance, feature dimensions are removed sequentially, and the model is retrained using the remaining feature dimensions, and the model's classification accuracy is evaluated. The preset threshold is a pre-defined standard for determining whether to stop feature removal, representing the maximum allowable decrease in model classification accuracy. When the decrease in model classification accuracy exceeds the preset threshold, feature removal stops, and the current set of feature dimensions is retained as the core feature dimensions reflecting the charging and discharging behavior pattern.

[0142] Step S355: Calculate the weighted average of the fusion feature sequence after dimensionality reduction on each core feature dimension. The weight is positively correlated with the feature contribution of the corresponding feature dimension within the clustering unit. The feature contribution is determined by the feature importance evaluation algorithm.

[0143] The weighted average is used to calculate the overall performance of the fused feature sequence after dimensionality reduction across each core feature dimension. Feature contribution refers to the degree to which each core feature dimension contributes to the data features within a cluster unit, determined through a feature importance evaluation algorithm, such as the previously mentioned random forest-based feature importance evaluation model. Weights are positively correlated with feature contribution; that is, the higher the feature contribution, the greater the corresponding weight.

[0144] Step S356: Combine the weighted average values ​​into an eigenvector form, which serves as the pattern center feature characterizing the charging and discharging behavior pattern of the cluster unit.

[0145] The pattern center feature is a feature vector used to characterize the charging and discharging behavior pattern of cluster units. It integrates information from the dimensionality-reduced fused feature sequence across each core feature dimension. By combining weighted averages into a feature vector form, the charging and discharging behavior pattern within a cluster unit can be quantitatively represented, facilitating subsequent analysis and applications.

[0146] Step S400: The charging and discharging mode clusters are correlated and matched with the grid load prediction information. The correlation and matching results are iteratively optimized through reinforcement learning to generate charging and discharging power allocation schemes that are adapted to different charging and discharging mode clusters. The charging and discharging power allocation schemes are used to dynamically adjust the real-time charging and discharging execution process of the energy storage power supply.

[0147] Association matching is the process of mapping and combining charging / discharging mode clusters with grid load forecast information. Through association matching, the applicability of different charging / discharging mode clusters under different grid load conditions can be understood. Grid load forecast information is prediction data of grid load changes over a future period, including the grid's load demand at different time points. Reinforcement learning is a machine learning method that uses an agent to interact with the environment and continuously learn the optimal strategy. In this step, reinforcement learning is used to iteratively optimize the association matching results to generate charging / discharging power allocation schemes adapted to different charging / discharging mode clusters. The charging / discharging power allocation scheme is a plan for allocating the charging / discharging power of the energy storage power source based on different charging / discharging mode clusters and grid load conditions. It can dynamically adjust the real-time charging / discharging process of the energy storage power source to improve the grid's operating efficiency and stability.

[0148] In one implementation, step S400 may specifically include the following steps S410 to S460:

[0149] Step S410: Obtain grid load forecast information of the power grid system, which includes the trend of grid load changes in the future period.

[0150] Grid load forecasting information is predictive data on the grid load over a future period, which is crucial for rationally scheduling the charging and discharging of energy storage power sources. Grid load forecasting information can be obtained in various ways, such as using historical grid load data for time series analysis and forecasting, or combining external factors such as meteorological data and holiday information with machine learning models for prediction.

[0151] For example, power grid load forecasting information can be obtained from power grid system monitoring equipment, data acquisition systems, or professional power grid load forecasting agencies. This information is typically presented in time series format, containing power grid load forecasts for different points in time within a future period.

[0152] Step S420: Align the mode center features of the charge / discharge mode cluster with the grid load prediction information on the time axis to keep their time stamps synchronized.

[0153] Time axis alignment is the process of unifying and matching the mode center characteristics of the charging and discharging mode cluster with the grid load forecast information in the time dimension, ensuring that the two have a corresponding relationship in time, so as to facilitate subsequent correlation matching and analysis.

[0154] For example, first, the time stamps of the mode center features of the charge / discharge mode cluster and the grid load forecast information are checked. If their time resolutions are inconsistent—for example, the mode center features are recorded at hourly intervals while the grid load forecast information is recorded at half-hourly intervals—adjustment is required. Time axis alignment can be achieved using interpolation or aggregation methods. If interpolation is used, for the longer time interval, data at intermediate time points can be generated using linear interpolation or similar methods to match the time resolution of the other. If aggregation is used, for the shorter time interval, multiple adjacent data points can be averaged or summed to obtain data with longer time intervals. Then, the aligned mode center features and grid load forecast information are arranged chronologically to ensure their time stamps remain synchronized. For example, the mode center features with hourly intervals and the grid load forecast information with half-hourly intervals can be interpolated to make both hourly intervals, and then they can be matched chronologically.

[0155] Step S430: Calculate the cross-correlation coefficient between the central features of the model and the power grid load forecast information to determine the correlation strength between the charging and discharging behavior pattern and the power grid load change trend.

[0156] The cross-correlation coefficient is an indicator used to measure the correlation between two time series, reflecting the degree of association between charging and discharging behavior patterns and the trend of grid load changes. By calculating the cross-correlation coefficient, we can understand the performance of different charging and discharging pattern clusters under different grid load changes, providing a basis for subsequent correlation matching.

[0157] For example, the aligned pattern center features and grid load forecast information are first converted into time series data. Then, cross-correlation analysis is used to calculate the cross-correlation coefficient between the two. Cross-correlation analysis can be achieved by calculating the correlation between the two time series at different time delays, such as calculating the Pearson correlation coefficient. By calculating the cross-correlation coefficient at different time delays, the time delay with the highest correlation coefficient is found. This highest correlation coefficient represents the strength of the association between the charging and discharging behavior pattern and the trend of grid load change.

[0158] Step S440: Construct the association matching results between charging and discharging modes and grid load based on the association strength. The association matching results include the grid load adaptation range corresponding to different charging and discharging mode clusters.

[0159] The correlation matching results are a correspondence constructed based on the correlation strength between charging and discharging behavior patterns and grid load change trends, clarifying the applicability of different charging and discharging mode clusters under different grid load conditions. The grid load adaptation range refers to the range of grid loads that each charging and discharging mode cluster can adapt to. By determining the grid load adaptation range, a reference can be provided for subsequent charging and discharging power allocation.

[0160] For example, charging / discharging mode clusters are matched with grid loads based on the calculated cross-correlation coefficients and association strengths. For each charging / discharging mode cluster, its performance under different grid loads is analyzed to determine the range of grid loads it can adapt to, i.e., the grid load adaptation interval. The grid load adaptation intervals of all charging / discharging mode clusters are then compiled to form the association matching results.

[0161] Step S450: Construct a reinforcement learning environment, using the association matching results as the environment state input, the charging and discharging power allocation scheme as the action output, and the power grid operation efficiency index as the reward signal.

[0162] A reinforcement learning environment is the context in which an agent learns and makes decisions, consisting of a state space, an action space, and a reward function. In this step, the association matching result is used as the environment state input, allowing the agent to select a suitable charging and discharging power allocation scheme based on the current environment state. The charging and discharging power allocation scheme is used as the action output, allowing the agent to interact with the environment by executing different actions. The power grid operating efficiency index is used as the reward signal, and the agent's goal is to maximize the reward signal by continuously adjusting its actions, thereby learning the optimal charging and discharging power allocation strategy.

[0163] In one implementation, step S450 may specifically include the following steps S451 to S456:

[0164] Step S451: When constructing the state space of the reinforcement learning environment, the charging and discharging mode cluster identifier, grid load adaptation range, current charging and discharging state parameters and environmental interference factors in the association matching results are combined into a multi-dimensional state vector as the environmental state input. The environmental interference factors include grid voltage fluctuation characteristics and energy storage power supply temperature change characteristics.

[0165] The state space is the set of all possible states in a reinforcement learning environment. In this step, multiple relevant factors are combined into a multi-dimensional state vector to represent the environment state. The charge / discharge mode cluster identifier is used to distinguish different charge / discharge mode clusters. The grid load adaptation range represents the grid load range that each charge / discharge mode cluster can adapt to. The current charge / discharge state parameters reflect the current charge / discharge status of the energy storage power supply, such as charge / discharge power and energy level. Environmental interference factors include grid voltage fluctuation characteristics and energy storage power supply temperature change characteristics, which affect the charge / discharge process of the energy storage power supply.

[0166] For example, the representation of each factor is first determined. The charge / discharge mode cluster identifier can be represented using integer encoding or one-hot encoding; the grid load adaptation range can be represented using the upper and lower limits of the range; the current charge / discharge state parameters can be directly represented using actual measured values; the grid voltage fluctuation characteristics among environmental interference factors can be represented using parameters such as voltage fluctuation range and fluctuation frequency; and the energy storage power supply temperature change characteristics can be represented using parameters such as temperature change rate and current temperature value. Then, these factors are combined into a multi-dimensional state vector in a certain order. For example, a reinforcement learning environment containing 5 charge / discharge mode clusters may have the following state vector: [mode cluster identifier, lower limit of grid load adaptation range, upper limit of grid load adaptation range, current charge / discharge power, current energy level, grid voltage fluctuation range, grid voltage fluctuation frequency, energy storage power supply temperature change rate, energy storage power supply current temperature}].

[0167] Step S452: Normalize each component in the multidimensional state vector, convert the charge / discharge mode cluster identifier into a one-hot encoded vector, convert the grid load adaptation interval into a combination feature of the interval midpoint value and the interval width, and convert the current charge / discharge state parameters and environmental interference factors into standardized features with zero mean and unit variance.

[0168] Normalization aims to eliminate differences in numerical ranges between different components, improving the training efficiency and stability of reinforcement learning algorithms. By converting charge / discharge mode cluster identifiers into one-hot encoded vectors, different charge / discharge mode clusters can have the same representation in the state vector. Converting the grid load adaptation range into a combination of the midpoint value and the range width provides a more concise representation of the grid load adaptation range information. Converting the current charge / discharge state parameters and environmental disturbance factors into standardized features with zero mean and unit variance makes these parameters comparable.

[0169] Step S453: When constructing the action space of the reinforcement learning environment, a continuous action space is used to represent the charging and discharging power allocation scheme. The dimension of the action space corresponds to the number of charging and discharging mode clusters, and the value range of each dimension is determined according to the maximum charging and discharging power capacity of the energy storage power source.

[0170] The action space is the set of all possible actions an agent can perform in a reinforcement learning environment. In this step, a continuous action space is used to represent the charging and discharging power allocation scheme. A continuous action space means that the agent can continuously select actions within a certain range, i.e., it can choose different charging and discharging power values. The dimensions of the action space correspond to the number of charging and discharging mode clusters, with each dimension representing the charging and discharging power allocation value for one charging and discharging mode cluster. The value range of each dimension is determined based on the maximum charging and discharging power capacity of the energy storage power source, ensuring that the actions selected by the agent are within the safe operating range of the energy storage power source.

[0171] For example, first determine the number n of the charging / discharging mode clusters, then the dimension of the action space is n. For each dimension, based on the maximum charging / discharging power capacity P of the energy storage power source... max and minimum charge / discharge power capacity P min Its value range is determined to be [P] min , P max ].

[0172] Step S454: When constructing the reward function for the reinforcement learning environment, the power grid operation efficiency index is used as the core reward item. The power grid operation efficiency index includes the proportion of power grid loss reduction, the proportion of renewable energy consumption, and the proportion of voltage stability improvement. The comprehensive efficiency index is calculated by weighted summation.

[0173] The reward function is used in a reinforcement learning environment to evaluate the quality of an agent's actions. It assigns a reward value based on the agent's actions and the environment's state. The agent's goal is to maximize this reward value by continuously adjusting its actions. In this step, the power grid operation efficiency index is used as the core reward item. This index comprises multiple aspects, which are combined into a comprehensive efficiency index through weighted summation, serving as the main part of the reward function.

[0174] For example, the calculation methods for the grid loss reduction ratio, renewable energy consumption ratio, and voltage stability improvement ratio are first determined. The grid loss reduction ratio can be calculated by comparing grid losses before and after adopting the current charging and discharging power allocation scheme; the renewable energy consumption ratio can be determined by calculating the proportion of renewable energy in total energy consumption; and the voltage stability improvement ratio can be evaluated by comparing voltage fluctuation range and stability indicators. Then, a weight is assigned to each indicator, with the weight set according to the actual situation and importance. For example, the weight of the grid loss reduction ratio is set to 0.4, the weight of the renewable energy consumption ratio is set to 0.3, and the weight of the voltage stability improvement ratio is set to 0.3. Finally, the comprehensive efficiency index is calculated by weighted summation: R = 0.4 × grid loss reduction ratio + 0.3 × renewable energy consumption ratio + 0.3 × voltage stability improvement ratio. The comprehensive efficiency index is used as the core reward item in the reward function.

[0175] Step S455: Introduce a dynamic penalty mechanism into the reward function. When the charging and discharging power allocation scheme exceeds the safe operating range of the energy storage power supply, calculate the penalty value based on the degree and duration of the excess. The degree of excess is measured by the deviation ratio between the current value and the safety threshold, and the duration is calculated by accumulating the number of consecutive sampling cycles.

[0176] The dynamic penalty mechanism is introduced to ensure that the charging and discharging power allocation scheme chosen by the agent is within the safe operating range of the energy storage power source. When the charging and discharging power allocation scheme exceeds the safe operating range of the energy storage power source, it will cause damage to the energy storage power source, so this behavior needs to be penalized. The penalty value is calculated based on the degree of exceedance and the duration; the greater the degree of exceedance and the longer the duration, the higher the penalty value.

[0177] For example, firstly, the safe operating range of the energy storage power source is determined, including safety thresholds such as maximum and minimum charge / discharge power. When the charge / discharge power allocation scheme selected by the agent exceeds the safety threshold, the deviation ratio between the current value and the safety threshold is calculated. Simultaneously, the duration is accumulated through consecutive sampling periods, i.e., the number of consecutive sampling periods exceeding the safe operating range is recorded. Then, a penalty value is calculated based on the degree and duration of the deviation. A linear or nonlinear function can be used to calculate the penalty value. Finally, the penalty value is subtracted from the reward function to obtain the final reward value.

[0178] Step S456: The reward signal is discounted and accumulated using the time difference algorithm, and the instantaneous reward fluctuations are smoothed using the exponential moving average method to generate the optimized target signal for the reinforcement learning algorithm policy iteration.

[0179] The temporal difference algorithm is used to address the delayed reward problem in reinforcement learning. By discounting and accumulating future rewards, it allows the agent to consider long-term rewards. The exponential moving average method is used to smooth data fluctuations by weighting historical data to make it smoother and more stable. In this step, the temporal difference algorithm is used to discount and accumulate the reward signal, while the exponential moving average method is used to smooth fluctuations in immediate rewards, generating the optimization target signal for the reinforcement learning algorithm's policy iteration.

[0180] Step S460: Iteratively optimize the policy in the reinforcement learning environment using a reinforcement learning algorithm, dynamically adjust the charging and discharging power allocation parameters corresponding to different charging and discharging mode clusters until the reward signal reaches a stable state, and generate the final charging and discharging power allocation scheme.

[0181] Policy iterative optimization involves continuously interacting with the environment and adjusting the agent's policy based on reward signals to find the optimal charging and discharging power allocation scheme. In this step, reinforcement learning algorithms are used to perform policy iterative optimization in a constructed reinforcement learning environment, dynamically adjusting the charging and discharging power allocation parameters corresponding to different charging and discharging mode clusters until the reward signal reaches a stable state, i.e., the reward signal no longer fluctuates or changes significantly. The charging and discharging power allocation scheme generated at this point is the final charging and discharging power allocation scheme.

[0182] In one implementation, step S460 may specifically include the following steps S461 to S466:

[0183] Step S461: Initialize the policy network parameters of the reinforcement learning algorithm, and set the maximum number of policy iterations and the reward signal stability threshold.

[0184] The policy network is a neural network used in reinforcement learning algorithms to generate actions; its parameters determine the agent's policy. Before starting policy iteration optimization, the policy network parameters need to be initialized. Random initialization can be used to initialize the policy network parameters to random values. The maximum number of policy iterations is a pre-set upper limit used to control the number of iterations and prevent the algorithm from getting stuck in an infinite loop. The reward signal stability threshold is a standard used to determine whether the reward signal has reached a stable state; when the fluctuation amplitude of the reward signal is less than this threshold, the reward signal is considered to have reached a stable state.

[0185] For example, the architecture and number of parameters of the policy network are first determined. For instance, the policy network could be a multilayer perceptron containing an input layer, hidden layers, and an output layer. Then, the parameters of the policy network are initialized using a random initialization method. Simultaneously, the maximum number of policy iterations is set according to the actual situation, such as 1000 times. The reward signal stability threshold can be set according to the fluctuation range and accuracy requirements of the reward signal; for example, the reward signal stability threshold can be set to 0.01.

[0186] Step S462: In each iteration, select the charge / discharge power allocation parameters from the action space based on the current policy network parameters and apply them to the reinforcement learning environment.

[0187] In each policy iteration, the agent selects a charge / discharge power allocation parameter from the action space based on the current policy network parameters; that is, a specific charge / discharge power allocation scheme. This scheme is then applied to the reinforcement learning environment to interact with it and receive feedback, including the next state and reward signal.

[0188] For example, the current multidimensional state vector is input into a policy network, which calculates an action vector based on its parameters. This action vector corresponds to the charge / discharge power allocation parameters. These charge / discharge power allocation values ​​are then applied to a reinforcement learning environment to update the charge / discharge state of the energy storage power source and the operating state of the power grid.

[0189] Step S463: Obtain the comprehensive reward signal and the next state vector from the reinforcement learning environment feedback, and store the current state vector, the charge / discharge power allocation parameters, the comprehensive reward signal, and the next state vector into the experience replay pool.

[0190] The experience replay pool is a buffer used in reinforcement learning to store the agent's interaction experience with the environment, which can improve the algorithm's sample utilization and training stability. In each iteration, after the agent interacts with the environment, it receives a comprehensive reward signal from the environment and the next state vector. The current state vector, the charge / discharge power allocation parameters, the comprehensive reward signal, and the next state vector are stored as a set of experience data in the experience replay pool.

[0191] For example, the comprehensive reward signal and the next state vector are obtained from the reinforcement learning environment. For instance, the comprehensive reward signal is R, and the next state vector is S. t+1 The current state vector S t Charging and discharging power allocation parameter A t The combined reward signal R and the next state vector S t+1 Combined into a quadruple (S t A t ,R,S t+1The data is stored in the experience replay pool. The experience replay pool can be implemented using data structures such as queues or lists. When the experience replay pool reaches a certain capacity, a first-in, first-out (FIFO) strategy can be used to update the data.

[0192] Step S464: When the number of samples in the experience replay pool reaches a preset threshold, a batch of samples is randomly sampled from the experience replay pool, and the policy network parameters are updated using the gradient descent algorithm.

[0193] When the number of samples in the experience replay pool reaches a preset threshold, it indicates that sufficient empirical data has been accumulated, and the parameters of the policy network can be updated. Randomly sampling a batch of samples from the experience replay pool increases sample diversity and avoids overfitting. The gradient descent algorithm calculates the gradient of the loss function with respect to the policy network parameters and then updates the parameters in the opposite direction of the gradient, gradually reducing the loss function.

[0194] Step S465: Calculate the moving average of the comprehensive reward signal over multiple consecutive iterations. When the fluctuation of the moving average is less than the stability threshold of the reward signal, stop the strategy iteration.

[0195] The moving average is a statistic used to smooth out data fluctuations. It is calculated by averaging the comprehensive reward signal over multiple consecutive iterations. When the fluctuation range of the moving average is less than the stability threshold of the reward signal, it indicates that the reward signal has reached a stable state, and strategy iteration can be stopped at this point.

[0196] Step S466: Extract the charge and discharge power allocation parameters output by the final strategy network, associate and bind them with the corresponding charge and discharge mode clusters, and generate a charge and discharge power allocation scheme containing the adaptation parameters of each charge and discharge mode cluster.

[0197] When the strategy iteration stops, it means that the optimal strategy network parameters have been found. At this point, the output charge / discharge power allocation parameters are extracted from the final strategy network, and these parameters are associated and bound with the corresponding charge / discharge mode clusters to form a complete charge / discharge power allocation scheme.

[0198] For example, the current multidimensional state vector is input into the final policy network to obtain the charge / discharge power allocation parameters. For instance, the charge / discharge power allocation parameters output by the policy network are [P1, P2, ..., Pn], corresponding to n charge / discharge mode clusters. These charge / discharge power allocation parameters are associated and bound to their corresponding charge / discharge mode clusters, recording the adapted charge / discharge power for each cluster. The adapted parameters for all charge / discharge mode clusters are then compiled to generate a charge / discharge power allocation scheme containing the adapted parameters for each cluster.

[0199] Please see Figure 2 , Figure 2This is a schematic diagram of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 can be connected via a bus or other means. The processor 101 (or Central Processing Unit, CPU) is the computing and control core of the computer system, capable of parsing various instructions and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for data transmission and interaction within the computer system. The memory 103 is a storage device in the computer system used to store programs and data. It is understood that the memory 103 here can include the computer system's built-in memory, or it can include extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system; this invention does not limit this storage space.

[0200] In one embodiment, the processor 101 executes the intelligent optimization method for charging and discharging strategies of energy storage power supplies provided above in the embodiments of the present invention by running a computer program in the memory 103.

Claims

1. A smart optimization method for charging and discharging strategies applied to energy storage power supplies, characterized in that, The method includes: The historical charging and discharging process of the energy storage power supply is segmented into time series, and a sliding time window is used to traverse the historical charging and discharging data to generate a data sequence of charging and discharging periods with continuous time markers. Perform charge and discharge feature extraction on the charge and discharge period data sequence to obtain the charge and discharge feature set corresponding to each charge and discharge period. The charge and discharge feature set contains feature information that reflects the change law of the charge and discharge process. Based on the charging and discharging feature set, pattern clustering analysis is performed. By using time series clustering, charging and discharging periods with similar feature distributions are divided into multiple charging and discharging pattern clusters, and each charging and discharging pattern cluster corresponds to a type of charging and discharging behavior pattern. The charging and discharging mode clusters are correlated and matched with grid load prediction information. The correlation and matching results are then iteratively optimized using reinforcement learning to generate charging and discharging power allocation schemes that are adapted to different charging and discharging mode clusters.

2. The method according to claim 1, characterized in that, The step of performing charge-discharge feature extraction on the charge-discharge period data sequence yields a charge-discharge feature set corresponding to each charge-discharge period, including: Waveform morphology analysis is performed on the data of each time period in the charging and discharging time period data sequence to extract the waveform features of the charging and discharging curves in each time period. The waveform features include the distribution position of the extreme points of the charging and discharging curves and the order of the appearance of the curve inflection points. Trend change analysis is performed on the data of each time period in the charging and discharging time period data sequence to calculate the trend characteristics of the charging and discharging rate in each time period. The trend characteristics include the duration of the rising phase and the duration of the falling phase of the charging and discharging rate. Align the waveform features and the trend features by feature dimensions to ensure that the time dimension labels of the two features are consistent, and generate a preliminary feature set. The preliminary feature set is subjected to feature redundancy removal processing. Redundant features that repeatedly represent the changes in the charging and discharging process are eliminated through feature correlation analysis, resulting in a simplified charging and discharging feature set.

3. The method according to claim 2, characterized in that, The step of performing waveform morphology analysis on the data of each time period in the charging and discharging time period data sequence and extracting the waveform features of the charging and discharging curves in each time period includes: The data for each time period in the charge / discharge time period data sequence are subjected to curve smoothing to obtain the smoothed charge / discharge curve; Perform extreme point detection on the smoothed charge-discharge curve to identify the maximum and minimum points on the charge-discharge curve and record the distribution position of each extreme point on the time axis. The second derivative of the smoothed charge-discharge curve is calculated, and the inflection point position of the curve is determined according to the positive and negative changes of the second derivative. The inflection point positions are arranged in chronological order to obtain the order in which the curve inflection points appear. The distribution locations of the extreme points and the order in which the curve inflection points appear are associated and stored to generate a waveform feature descriptor containing position coordinates and order markers; The waveform feature descriptor is converted into a standardized feature vector form, which serves as the waveform feature of the charge-discharge curve.

4. The method according to claim 2, characterized in that, The step of performing trend change analysis on the data of each time period in the charging and discharging time period data sequence, and calculating the trend characteristics of the charging and discharging rate changes in each time period, includes: The charging and discharging rate is calculated for each time period in the charging and discharging time period data sequence. The change in charging and discharging amount per unit time is used as the charging and discharging rate. The charging and discharging rate is then subjected to three exponential smoothing processes to obtain a smoothed charging and discharging rate sequence. The smoothed charge and discharge rate sequence is subjected to multi-scale trend decomposition, which decomposes the rate sequence into multiple intrinsic mode function components and residual trend terms. The residual trend terms that reflect the long-term change trend are extracted as the principal components of trend analysis. The principal components of the trend analysis are subjected to first-order and second-order difference processing to obtain the rate of change sequence and acceleration change sequence, respectively. Based on the positive and negative values ​​of the rate of change sequence, the charging and discharging rate is initially divided into the rising stage, the falling stage and the stable stage. The initial stage boundaries are dynamically adjusted by combining the acceleration change sequence. When the absolute value of the acceleration change sequence exceeds the preset trend threshold, the stage boundary correction mechanism is triggered, and the final stage boundary is determined by verifying the trend consistency of adjacent time periods. During the rising phase, the duration of the rising phase is determined by starting the time when the rate of change first turns from negative to positive and ending the time when the rate of change turns from positive to negative again or approaches zero. During the falling phase, the duration of the falling phase is determined by starting the time when the rate of change first turns from positive to negative and ending the time when the rate of change turns from negative to positive again or approaches zero. The duration of the rising phase and the duration of the falling phase are normalized to a time granularity and converted into a relative duration index based on the time interval of the time period data sequence, generating a trend feature with the same time dimension.

5. The method according to claim 1, characterized in that, The pattern clustering analysis based on the charging and discharging feature set divides charging and discharging periods with similar feature distributions into multiple charging and discharging pattern clusters through time series clustering. Each charging and discharging pattern cluster corresponds to a type of charging and discharging behavior pattern, including: Time series standardization is performed on each charge-discharge feature in the charge-discharge feature set to ensure that the numerical range of different features is consistent, thus obtaining a standardized charge-discharge feature set. Calculate the dynamic time curvature distance between any two charge / discharge period features in the standardized charge / discharge feature set, and use it as an indicator to measure the similarity of feature distributions; A charge-discharge feature similarity matrix is ​​constructed based on the dynamic time bending distance, and the element values ​​in the similarity matrix represent the degree of similarity between the features of two corresponding charge-discharge periods. The charging and discharging feature similarity matrix is ​​clustered using a time series clustering algorithm, and charging and discharging periods with similarity levels higher than a preset threshold are divided into the same clustering unit. For each cluster unit, pattern features are extracted to generate pattern center features that characterize the charging and discharging behavior pattern of the cluster unit. Cluster units containing pattern center features and corresponding charging and discharging time periods are defined as charging and discharging pattern clusters.

6. The method according to claim 5, characterized in that, The calculation of the dynamic time curvature distance between any two charge / discharge time period features in the standardized charge / discharge feature set, used as an indicator to measure the similarity of feature distributions, includes: The two charge-discharge period features in the standardized charge-discharge feature set are respectively denoted as the first charge-discharge period feature and the second charge-discharge period feature. The first charge-discharge period feature contains multiple first feature points arranged in chronological order, and the second charge-discharge period feature contains multiple second feature points arranged in chronological order. Construct a time curvature path matrix between the first and second charge / discharge period features, where each element of the time curvature path matrix represents the feature point distance between the first and second feature points; The minimum cost path in the time-curved path matrix is ​​determined through heuristic search. The minimum cost path meets the time order constraint and has the minimum cumulative path distance. Calculate the sum of feature point distances for all elements on the path with the minimum cost, and use it as the dynamic time curvature distance between the features of the first and second charge / discharge periods. The dynamic time bending distance is normalized to obtain a similarity index with a value range within a set interval. The smaller the similarity index value, the higher the similarity of the distribution of features of the two charging and discharging periods.

7. The method according to claim 5, characterized in that, The step of extracting pattern features for each cluster unit to generate pattern center features characterizing the charging and discharging behavior pattern of that cluster unit includes: Extract the standardized charge-discharge feature set of all charge-discharge periods in the clustering unit, take the standardized charge-discharge features of each charge-discharge period as feature samples, and perform time-series alignment processing on the feature samples to make the time axis labels of all feature samples consistent. Anomaly detection is performed on the feature samples using the isolated forest algorithm to identify and remove abnormal feature samples that deviate from the cluster center, thus obtaining a set of normal feature samples. Perform temporal feature alignment on each feature sample in the normal feature sample set, calculate the feature alignment offset between samples, and perform position calibration on the sample feature points based on the feature alignment offset to obtain the calibrated feature sample. The calibrated feature samples are fused along the time dimension to generate a fused feature sequence containing the correlation of features across multiple time periods. The fused feature sequence is then subjected to feature dimensionality reduction to obtain the core feature dimension reflecting the charging and discharging behavior pattern. Calculate the weighted average of the fusion feature sequence after dimensionality reduction on each core feature dimension. The weight is positively correlated with the feature contribution of the corresponding feature dimension within the clustering unit. The feature contribution is determined by the feature importance evaluation algorithm. The weighted average values ​​are combined into a feature vector form, which serves as the pattern center feature characterizing the charging and discharging behavior pattern of the cluster unit.

8. The method according to claim 1, characterized in that, The step of associating and matching the charging and discharging mode clusters with grid load forecast information, and then using reinforcement learning to iteratively optimize the association and matching results to generate charging and discharging power allocation schemes adapted to different charging and discharging mode clusters includes: Obtain grid load forecast information of the power grid system, wherein the grid load forecast information includes the trend of grid load changes in future time periods; The mode center features of the charging and discharging mode cluster are aligned with the power grid load prediction information on the time axis to keep their time stamps synchronized. The cross-correlation coefficient between the central characteristics of the calculation pattern and the power grid load forecast information is used to determine the correlation strength between the charging and discharging behavior pattern and the power grid load change trend. Based on the correlation strength, a correlation matching result between the charging and discharging modes and the grid load is constructed. The correlation matching result includes the grid load adaptation range corresponding to different charging and discharging mode clusters. A reinforcement learning environment is constructed, with the association matching results as the environment state input, the charging and discharging power allocation scheme as the action output, and the power grid operation efficiency index as the reward signal. The policy is iteratively optimized in the reinforcement learning environment using a reinforcement learning algorithm. The charging and discharging power allocation parameters corresponding to different charging and discharging mode clusters are dynamically adjusted until the reward signal reaches a stable state, thus generating the final charging and discharging power allocation scheme.

9. The method according to claim 8, characterized in that, The construction of the reinforcement learning environment, using the association matching results as the environment state input, the charging and discharging power allocation scheme as the action output, and the power grid operation efficiency index as the reward signal, includes: When constructing the state space of the reinforcement learning environment, the charging and discharging mode cluster identifier, grid load adaptation range, current charging and discharging state parameters and environmental interference factors in the association matching results are combined into a multi-dimensional state vector as the environmental state input. The environmental interference factors include grid voltage fluctuation characteristics and energy storage power temperature change characteristics. Normalize each component in the multidimensional state vector, convert the charge / discharge mode cluster identifier into a one-hot encoded vector, convert the grid load adaptation interval into a combination feature of the interval midpoint value and the interval width, and convert the current charge / discharge state parameters and environmental interference factors into standardized features with zero mean and unit variance. When constructing the action space of the reinforcement learning environment, a continuous action space is used to represent the charging and discharging power allocation scheme. The dimension of the action space corresponds to the number of charging and discharging mode clusters, and the value range of each dimension is determined according to the maximum charging and discharging power capacity of the energy storage power supply. When constructing the reward function for the reinforcement learning environment, the grid operation efficiency index is used as the core reward item. The grid operation efficiency index includes the reduction ratio of grid losses, the renewable energy consumption ratio, and the voltage stability improvement ratio. The comprehensive efficiency index is calculated by weighted summation. A dynamic penalty mechanism is introduced into the reward function. When the charging and discharging power allocation scheme exceeds the safe operating range of the energy storage power supply, the penalty value is calculated based on the degree and duration of the excess. The degree of excess is measured by the deviation ratio between the current value and the safety threshold, and the duration is calculated by accumulating the number of consecutive sampling cycles. The reward signal is discounted and accumulated using a time difference algorithm, and the instantaneous reward fluctuations are smoothed using an exponential moving average method to generate an optimized target signal for reinforcement learning algorithm policy iteration.

10. A computer system, characterized in that, include: A memory, wherein a computer program is stored; A processor is configured to load the computer program to implement the intelligent optimization method for charging and discharging strategies applied to energy storage power sources as described in any one of claims 1-9.

Citation Information

Cited By

  • Rare earth lithium battery high-speed fast charging intelligent management system based on rare earth composite material

    CN121404071A

  • Adaptability analysis method and system for overcharge load and multi-element energy storage equipment

    CN121440720A