Operation pressure assessment method and system based on energy storage and new energy power station
By constructing an operational stress assessment method based on energy storage and new energy power plants, and utilizing data cleaning, active learning, and sparse generalized linear models, the complexity and sparse correlation problems of the assessment models for new energy power plants and energy storage systems are solved, achieving efficient and accurate operational stress assessment and early warning support.
Patent Information
- Application Number
- CN202511581511.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-02-17
Smart Images

Figure CN121543053A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of new energy power plant management technology, and in particular to a method and system for assessing the operational pressure of energy storage and new energy power plants. Background Technology
[0002] The synergistic development of new energy power plants and energy storage systems has become an important trend in the energy industry. As the proportion of renewable energy continues to increase, energy storage systems are playing an increasingly important role in the power system, and their operational stress assessment is crucial for the sustainable development of power plants.
[0003] Currently, common methods for assessing the operational stress of new energy power plants are mainly based on traditional financial indicator analysis, such as single dimensions like payback period and net present value. These methods typically employ deterministic linear models, which fail to fully reflect the complexity of the coordinated operation of new energy power plants and energy storage systems, thus limiting the accuracy and practicality of the assessment results.
[0004] More advanced evaluation techniques employ a multi-dimensional indicator system, combining factors such as the charging and discharging characteristics of energy storage systems, equipment lifespan, and market prices to construct evaluation models. This technology uses complex nonlinear models to describe the relationships between various factors, enabling a better characterization of the operational features of new energy and energy storage systems and improving the accuracy of the evaluation.
[0005] However, existing technologies still have some problems: First, the models are highly complex and computationally inefficient, making it difficult to meet real-time assessment needs; second, they struggle to effectively handle the sparse correlation between assessment indicators, easily leading to overfitting and affecting the model's application in real-world environments. These issues limit the effectiveness and practicality of assessing the operational stress of new energy power plants and energy storage systems. Summary of the Invention
[0006] In view of this, this application provides a method and system for assessing the operational stress of energy storage and new energy power plants, which solves the problems of high model complexity, low computational efficiency and difficulty in effectively handling the sparse correlation characteristics between assessment indicators in the prior art.
[0007] This application provides a method for assessing the operational stress of energy storage and new energy power plants, including:
[0008] Collect energy storage system operation data, new energy power generation data, and electricity market data; clean and standardize the energy storage system operation data, the new energy power generation data, and the electricity market data to obtain a standardized evaluation index dataset.
[0009] Using the standardized evaluation index dataset, an index correlation matrix is constructed and an active learning strategy is used to select training samples. Through iterative optimization, a near-optimal regression model for a single index is obtained.
[0010] A sparse feature matrix is constructed based on the output of the single-index approximate optimal regression model. The sparse feature matrix is then optimized using an iterative hard threshold learning method to form an optimized sparse generalized linear model.
[0011] Collect real-time operational data, input the real-time operational data into the optimized sparse generalized linear model, and output operational pressure assessment results and early warning information.
[0012] The energy storage system operation data, the new energy power generation data, and the electricity market data are cleaned and standardized to obtain the standardized evaluation index dataset, including:
[0013] The energy storage system operation data, the new energy power generation data, and the electricity market data are used to detect and process outliers, resulting in cleaned data.
[0014] Based on the cleaned data, the utilization rate, charging and discharging efficiency, equipment depreciation rate, power generation efficiency, and electricity price fluctuation characteristics of the energy storage system are extracted to obtain characteristic data;
[0015] The feature data is subjected to Z-score standardization to generate the standardized evaluation index dataset.
[0016] Using the standardized evaluation index dataset, an index correlation matrix is constructed, and an active learning strategy is employed to select training samples, including:
[0017] The correlation coefficients between the evaluation indicators are calculated using the standardized evaluation indicator dataset to generate the indicator correlation matrix;
[0018] The sample information entropy is calculated based on the correlation matrix of the indicators, and training samples are selected using an uncertainty sampling strategy to form evaluation indicator data to be trained.
[0019] A near-optimal regression model for a single index is obtained through iterative optimization, including:
[0020] An objective function containing a prediction error term and a regularization term is established using the evaluation index data to be trained;
[0021] The gradient of the model parameters is calculated based on the objective function, and the model parameters are updated using the first learning rate to generate the single-index approximate optimal regression model.
[0022] Based on the output of the single-index approximate optimal regression model, a sparse feature matrix is constructed, including:
[0023] The L1 regularization process is performed on the output of the single-index approximate optimal regression model to generate the filtered features.
[0024] Principal component analysis is performed on the filtered features to construct the sparse feature matrix.
[0025] The sparse feature matrix is optimized using an iterative hard thresholding learning method, including:
[0026] A first hard threshold function is created using the data distribution characteristics of the sparse feature matrix to form a parameter threshold index;
[0027] Based on the parameter threshold index, the model parameters are iteratively optimized using the gradient descent method, and parameter sparsification is performed using the first hard threshold function to generate an optimized sparse generalized linear model containing model performance indicators.
[0028] It also includes: validating and tuning the optimized sparse generalized linear model containing model performance metrics, including:
[0029] Cross-validation is performed using the optimized sparse generalized linear model containing model performance metrics and the first hard threshold function to obtain the validated model performance metrics.
[0030] Based on the verified model performance metrics, the model hyperparameters are optimized using a grid search method to form the final optimized model parameters.
[0031] Inputting the real-time running data into the optimized sparse generalized linear model includes:
[0032] The real-time operational data is used to perform intraday, monthly, and annual forecasts, generating multi-scale operational pressure forecast results.
[0033] The multi-scale operational stress prediction results are weighted and fused to form a fused operational stress assessment result.
[0034] It also includes: generating early warning information based on the integrated operational pressure assessment results, including:
[0035] A multi-level early warning threshold system is created using the fused operational stress assessment results;
[0036] The operational pressure level is graded and assessed according to the multi-level early warning threshold system, and the early warning information is output.
[0037] This application also provides an operational stress assessment system based on energy storage and new energy power plants, including:
[0038] The data acquisition module is used to collect energy storage system operation data, new energy power generation data and electricity market data, and to clean and standardize the energy storage system operation data, the new energy power generation data and the electricity market data to obtain a standardized evaluation index dataset.
[0039] The model building module is used to construct an index correlation matrix using the standardized evaluation index dataset and select training samples using an active learning strategy, and obtain a single index approximate optimal regression model through iterative optimization.
[0040] The model optimization module is used to construct a sparse feature matrix based on the output of the single-index approximate optimal regression model, and to optimize the sparse feature matrix using an iterative hard threshold learning method to form an optimized sparse generalized linear model.
[0041] The evaluation module is used to collect real-time operating data, input the real-time operating data into the optimized sparse generalized linear model, and output the operational pressure evaluation results and early warning information.
[0042] This application embodiment also provides a computer device, the computer device comprising:
[0043] At least one processor; and a memory communicatively connected to said at least one processor;
[0044] The memory stores instructions that can be executed by the at least one processor, which are then executed by the at least one processor to enable the at least one processor to perform the above-mentioned operational stress assessment method based on energy storage and new energy power plants.
[0045] This application also provides a computer-readable storage medium that stores computer instructions for causing a computer to execute the above-described method for assessing the operational stress of energy storage and new energy power plants.
[0046] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above-described method for assessing the operational pressure of energy storage and new energy power plants.
[0047] This application has the following technical effects:
[0048] By using an active learning strategy to create a single-index approximate optimal regression model, computational complexity is reduced and model training efficiency is improved.
[0049] An iterative hard threshold learning method is used to construct a sparse generalized linear model, which effectively handles the sparse correlation characteristics between evaluation indicators and improves the model's generalization ability.
[0050] Multi-level data preprocessing and feature extraction methods improve data quality and enhance model accuracy;
[0051] Real-time assessment and early warning mechanisms support dynamic operational stress assessments, providing timely decision support for the operation of new energy power plants and energy storage systems. Attached Figure Description
[0052] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0053] Figure 1 A flowchart illustrating the operational stress assessment method for energy storage and new energy power plants provided in this application embodiment;
[0054] Figure 2 A detailed flowchart of the data acquisition and preprocessing steps provided for embodiments of this application;
[0055] Figure 3 A detailed flowchart illustrating the steps for constructing a single-index near-optimal active regression model, as provided in the embodiments of this application;
[0056] Figure 4 A detailed flowchart illustrating the steps for constructing and optimizing a sparse generalized linear model as provided in the embodiments of this application;
[0057] Figure 5 A detailed flowchart of the comprehensive assessment steps for operational stress provided in this application embodiment. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0059] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0060] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0061] like Figure 1 As shown in the embodiments of this application, a method for assessing the operational stress of energy storage and new energy power plants is provided, including:
[0062] S1: Collect operational data of energy storage embodiments of this application, new energy power generation data, and electricity market data; perform cleaning and standardization processing on the operational data of energy storage embodiments of this application, the new energy power generation data, and the electricity market data to obtain a standardized evaluation index dataset;
[0063] In the operational stress assessment methodology for energy storage and new energy power plants, data acquisition and preprocessing are fundamental to the entire assessment process. This stage first acquires raw data from multiple sources, including energy storage devices, new energy power generation units, and the electricity trading market, through a distributed acquisition architecture. The operational data for energy storage in this embodiment mainly includes parameters such as charge / discharge power curves, SOC changes, cycle count, and device temperature; new energy power generation data covers real-time power generation, environmental parameters (such as solar intensity, wind speed and direction), and device status information; electricity market data includes market information such as spot price fluctuations, ancillary service prices, and electricity demand forecasts. This multi-source heterogeneous data is acquired in real-time through standardized interfaces and initially cached locally to handle anomalies such as network outages.
[0064] After the raw data is collected, this embodiment of the application requires comprehensive data cleaning. First, the data integrity is checked, missing values are identified and processed using methods such as time-series interpolation, similar day substitution, or statistical mean. Second, outlier detection is performed, using statistical methods (such as the 3σ rule or box plot analysis) or machine learning algorithms (such as clustering or anomaly detection models) to identify and process data anomalies. Finally, smoothing filtering algorithms are used to eliminate data noise and improve data quality. After basic cleaning, this embodiment of the application extracts key features from the processed data, including the utilization rate of energy storage, charging and discharging efficiency and degradation trend, equipment depreciation rate, new energy power generation efficiency, and electricity price fluctuation patterns. These features are transformed to a unified dimension range using standardization methods such as Z-score or Min-Max, forming a standardized evaluation index dataset, laying the foundation for subsequent modeling and analysis.
[0065] S2: Using the standardized evaluation index dataset, construct an index correlation matrix and use an active learning strategy to select training samples. Through iterative optimization, obtain a single-index approximate optimal regression model.
[0066] In the single-index near-optimal active regression model construction stage, intelligent algorithms are used to mine index relationships from standardized data and build an efficient model. This stage first analyzes the intrinsic correlations between various evaluation indicators, constructing an index correlation matrix by calculating Pearson correlation coefficients or Spearman rank correlation coefficients to visually demonstrate the degree of influence between each index. Based on this matrix, this embodiment sets a correlation threshold (e.g., |ρ|>0.7) to identify significantly correlated index combinations, providing a basis for subsequent modeling. In the sample selection stage, this embodiment no longer simply uses the full dataset but adopts an active learning strategy, selecting the most informative sample subset based on the principle of information entropy. Specifically, this embodiment calculates the uncertainty index (e.g., information entropy or prediction variance) for each sample point, prioritizing samples with high uncertainty, strong representativeness, and coverage of the feature space for labeling and training. This strategy significantly reduces the number of training samples required while maintaining the model's learning effect and greatly improving computational efficiency.
[0067] In terms of model optimization, this embodiment constructs an objective function containing a prediction error term and a regularization term based on selected training samples. The prediction error term typically uses mean squared error (MSE) to measure the deviation between the model's predicted value and the actual value; the regularization term introduces L1 or L2 norm constraints to prevent overfitting. This embodiment uses gradient descent to iteratively optimize the model parameters, calculating the gradient of the objective function with respect to the parameters in each iteration and updating the parameter values according to an adaptive learning rate strategy. To improve convergence efficiency, this embodiment also employs a dynamic learning rate adjustment mechanism, automatically adjusting the learning step size based on changes in the objective function during the optimization process. When the parameter change over multiple consecutive iterations is less than a preset threshold or the maximum number of iterations is reached, this embodiment considers a local optimum to have been found, completing the construction of the single-index model. This approximate optimization strategy significantly reduces the solution complexity while ensuring model performance.
[0068] S3: Construct a sparse feature matrix based on the output of the single-index approximate optimal regression model, and optimize the sparse feature matrix using an iterative hard threshold learning method to form an optimized sparse generalized linear model.
[0069] The sparse generalized linear model construction and optimization stage is the core innovation of the evaluation method, aiming to address the sparse correlation characteristics between evaluation indicators of energy storage and new energy power plants. This stage first performs L1 regularization on the output of the single-indicator model. By solving an optimization problem with L1 norm constraints, some model parameters are forced to zero, achieving feature selection. This process not only removes redundant features but also significantly improves the model's interpretability and generalization ability. The embodiments of this application then apply Principal Component Analysis (PCA) to the selected feature set, transforming the original features into a set of linearly uncorrelated principal components. Dimensions are then selected for retention based on the cumulative explained variance (typically a 95% threshold), further reducing model complexity and constructing a structurally optimized sparse feature matrix.
[0070] In the model optimization stage, this application innovatively introduces a hard threshold function to achieve parameter sparsity. This function employs a piecewise design: when the absolute value of a parameter is less than the threshold λ, it is set to zero; when it is greater than or equal to λ, it remains unchanged. The setting of the threshold λ is crucial. This application adaptively determines the initial threshold based on the data distribution characteristics of the feature matrix, typically a fixed proportion of the feature standard deviation (e.g., 0.1 times). The optimization process uses an alternating iterative strategy: first, the hard threshold function parameters are fixed, and the model parameters are optimized through gradient descent; then, the hard threshold function is applied to perform parameter sparsity; finally, the hard threshold parameters are updated, and the next iteration begins. This alternating optimization method ensures both the model's prediction accuracy and maintains the high sparsity of the parameters. This application also evaluates model performance through cross-validation and optimizes hyperparameters, including regularization coefficients and hard threshold function parameters, using methods such as grid search, ultimately forming a sparse generalized linear model with good generalization ability.
[0071] S4: Collect real-time operating data, input the real-time operating data into the optimized sparse generalized linear model, and output the operational pressure assessment results and early warning information.
[0072] In the comprehensive assessment phase of operational pressure, the optimized model is applied to actual business scenarios to support management decisions. This application embodiment first establishes a real-time data acquisition mechanism, obtaining the latest operational data of the energy storage embodiment and the new energy power plant through standardized interfaces. Unlike traditional methods, this method simultaneously performs predictive analysis across multiple time scales: short-term (intraday) forecasts focus on immediate operational efficiency and market price fluctuations, employing a sliding window strategy for frequent updates; medium-term (monthly) forecasts consider equipment maintenance cycles and seasonal factors; and long-term (annual) forecasts focus on strategic factors such as equipment depreciation and technological iteration. This application embodiment integrates the forecast results from different time scales through a weighted fusion method, with weight coefficients determined based on historical data learning, forming a comprehensive assessment result that fully reflects the operational status.
[0073] Regarding the early warning mechanism, this application embodiment establishes a multi-level early warning threshold system, typically classifying operational pressure into four levels: normal, attention, warning, and danger. The early warning thresholds comprehensively consider historical statistical patterns, expert experience, and operational objectives, and are dynamically adjusted based on external factors such as seasonal changes and policy adjustments. When this application embodiment detects that operational pressure exceeds a specific threshold, it automatically triggers the corresponding level of early warning information and initiates cause diagnosis analysis to identify key factors leading to increased pressure (such as decreased equipment efficiency, market price fluctuations, or policy changes), and provides targeted handling suggestions. This application embodiment also provides multi-dimensional visualization, including real-time monitoring panels, trend analysis charts, and detailed assessment reports, allowing managers to view assessment results anytime via desktop terminals or mobile devices and respond promptly to operational risks. This comprehensive assessment and early warning mechanism provides a powerful decision support tool for the refined management of energy storage and new energy power plants.
[0074] Among them, such as Figure 2 As shown, S1 specifically includes:
[0075] S1.1: Using the operating data of the energy storage embodiment of this application, the new energy power generation data, and the electricity market data, outlier detection and processing are performed to form cleaned data;
[0076] Firstly, outlier detection and processing are the primary steps in data preprocessing and are crucial to ensuring the accuracy of subsequent evaluation models. In practical implementation, this application embodiment comprehensively examines the collected operational data of energy storage, new energy power generation data, and electricity market data. Firstly, this application embodiment employs statistical methods for outlier identification, including applying the 3σ principle to perform boundary detection on numerical data, marking data exceeding the mean ± 3 standard deviations as potential outliers. Simultaneously, this application embodiment also uses box plots to identify outliers, marking data points falling below Q1-1.5IQR or above Q3+1.5IQR as outliers (where Q1 and Q3 are the first and third quartiles of the data, respectively, and IQR is the interquartile range). For the data in this application's embodiments regarding energy storage, this application's embodiments pay particular attention to abnormal SOC values (such as exceeding the 0-100% range), sudden changes in charging and discharging power, and abnormal equipment temperatures; for new energy power generation data, the focus is on detecting situations where power generation does not match environmental parameters, such as data points where the solar irradiance is high but the photovoltaic power generation is abnormally low; for electricity market data, the main focus is on identifying price points with drastic price fluctuations and price points that significantly deviate from historical patterns.
[0077] After identifying anomalous data, this embodiment adopts corresponding processing strategies based on the type and severity of the anomaly. For isolated anomalies, if determined to be measurement errors, the average or median of the nearest valid values is used for replacement; for continuous anomalous data segments, time series interpolation methods (such as linear interpolation, spline interpolation, or ARIMA model prediction) are used for correction; for data with a qualitative shift, a regression correction model is used for adjustment. Simultaneously, this embodiment performs imputation processing on missing values. For missing data with short time intervals, linear or polynomial interpolation methods are used; for data missing for long periods, historical data from the same period or data under similar operating conditions are used for replacement. Through this series of processing steps, this embodiment effectively removes noise and anomalies from the original data, forming a reliable cleaned dataset, laying the foundation for subsequent feature extraction.
[0078] S1.2: Extract the utilization rate, charging and discharging efficiency, equipment depreciation rate, power generation efficiency, and electricity price fluctuation characteristics of the energy storage embodiment of this application based on the cleaned data to obtain feature data;
[0079] This application embodiment extracts key features based on the cleaned dataset. These features are core indicators for assessing the operational pressure of energy storage and new energy power plants. The utilization rate of energy storage in this application embodiment is an important indicator for measuring the efficiency of energy storage equipment. This application embodiment quantifies this by calculating the ratio of actual charge / discharge capacity to rated capacity, specifically expressed as: Utilization rate = (Σ|P(t)|·Δt) / (2·C·T), where P(t) is the charge / discharge power at time t, C is the rated capacity of this application embodiment, and T is the statistical period. Charge / discharge efficiency reflects the energy loss during the energy conversion process of this application embodiment. The calculation method is: Efficiency = Total discharge capacity / Total charge capacity. This application embodiment obtains this indicator through the integral charge / discharge power curve. Equipment depreciation rate considers the value reduction of energy storage equipment over time and cycle count, and is usually calculated based on a cycle life model: Depreciation rate = f(N,DOD,T), where N is the number of cycles, DOD is the depth of discharge, and T is the usage time.
[0080] For new energy power generation efficiency, this application uses different calculation methods based on the power generation type. For example, photovoltaic power generation efficiency = actual power generation / (irradiance intensity · panel area · theoretical conversion efficiency), and wind power efficiency = actual power generation / (wind energy resources · wind turbine swept area · theoretical power coefficient). Electricity price fluctuation characteristics are obtained through statistical analysis of electricity market price data, including indicators such as peak-valley price difference, price volatility, and arbitrage opportunity frequency. Furthermore, this application also extracts some composite features, such as the peak-valley arbitrage index for energy storage (considering peak-valley price difference and charging / discharging efficiency) and the new energy absorption coefficient (considering wind and solar curtailment rates and grid dispatch restrictions). These features together constitute a multi-dimensional indicator system for assessing operational pressure, comprehensively reflecting the operational status and external environment of energy storage and new energy power plants.
[0081] S1.3: Perform Z-score standardization on the feature data to generate the standardized evaluation index dataset.
[0082] Z-score standardization addresses the issue of inconsistent dimensions among different features, ensuring that each indicator has equal weight during model training. Specifically, for each feature x, this embodiment calculates its standardized value z = (x - μ) / σ, where μ is the mean of the feature and σ is the standard deviation. This process transforms each feature into a standard normal distribution with a mean of 0 and a standard deviation of 1, eliminating the influence of dimensional differences. For example, the utilization rate of energy storage in this embodiment is typically a ratio between 0 and 1, while electricity price fluctuations can range from tens to hundreds of yuan. Without standardization, electricity price factors would unreasonably dominate the model results in subsequent modeling. This embodiment calculates the mean and standard deviation of each feature separately during standardization and then applies the transformation to the entire dataset. It is important to note that this embodiment saves these statistical parameters to maintain a consistent transformation standard when processing new data. Through standardization, this embodiment generates a balanced and comparable dataset of evaluation indicators, providing a high-quality data foundation for building an accurate operational stress assessment model.
[0083] like Figure 3 As shown, S2 specifically includes:
[0084] S2.1: Calculate the correlation coefficients between the evaluation indicators using the standardized evaluation indicator dataset to generate the indicator correlation matrix;
[0085] This application embodiment conducts an in-depth analysis of the standardized evaluation index dataset, constructing an index correlation matrix to reveal the intrinsic relationships between the evaluation indicators. Specifically, this application embodiment calculates the correlation coefficient between any two indicators Xi and Xj, forming an n×n symmetric matrix (n being the number of indicators). Two main methods are used to calculate the correlation coefficient: the Pearson correlation coefficient is suitable for detecting linear relationships, calculated as ρ(Xi,Xj)=cov(Xi,Xj) / (σXi·σXj), where cov represents the covariance and σ represents the standard deviation; the Spearman rank correlation coefficient is suitable for detecting nonlinear monotonic relationships, assessing the strength of the correlation by calculating the correlation between indicator rankings. For energy storage, this application embodiment pays particular attention to the correlation between indicators such as charge / discharge efficiency and cycle count, and SOC management strategy and battery life; for new energy power generation, the focus is on analyzing the relationship between power generation efficiency and environmental factors, and power generation volatility and ancillary service demand.
[0086] After correlation analysis, this embodiment identifies significantly correlated indicator pairs by setting thresholds (typically |ρ|>0.7 indicates strong correlation, and 0.4<|ρ|<0.7 indicates moderate correlation), and further distinguishes between positive and negative correlations. These analytical results not only help understand the influence mechanisms between indicators but also provide important basis for subsequent feature selection and model construction. For example, when two indicators are found to be highly correlated, it may be advisable to retain only one to reduce model complexity; when an indicator is found to be significantly correlated with multiple other indicators, it may indicate that the indicator is a key driving factor and should be given special attention. Visualization of the correlation matrix (such as heatmaps) also provides analysts with an intuitive tool for understanding data structure, helping to discover potential data patterns and regularities.
[0087] S2.2: Calculate the sample information entropy based on the index correlation matrix, and use an uncertainty sampling strategy to screen training samples to form evaluation index data to be trained;
[0088] Based on information theory principles, an efficient sample selection strategy is implemented. By calculating sample information entropy and applying an uncertainty sampling method, the most informative training samples are selected from massive historical data. First, this embodiment evaluates the information entropy of each sample point based on the index correlation matrix, calculated as H(x) = -Σp(xi)log p(xi), where p(xi) represents the probability of sample x in the i-th partition (obtained through kernel density estimation or histogram methods). Higher information entropy indicates greater uncertainty in the sample and a greater potential contribution to model training. In actual calculations, this embodiment divides the sample space into multiple regions and calculates the probability distribution of samples falling into each region, thereby estimating the information entropy value.
[0089] In the sample selection stage, this embodiment employs an uncertain sampling strategy, comprehensively considering three key factors: the information content of the samples (prioritizing samples with high information entropy), the representativeness of the samples (ensuring that the selected samples can well cover the entire feature space), and the diversity of the samples (avoiding information redundancy caused by selecting overly similar samples). Specifically, this embodiment first sorts the samples according to their information entropy values, then uses clustering methods (such as K-means or density clustering) to group the samples, selecting the sample with the highest information entropy from each group, while ensuring that samples in different groups maintain sufficient distance. This strategy not only significantly reduces the number of samples that need to be labeled and used (typically by 60-80%), but also ensures the representativeness and diversity of the selected samples, providing a high-quality, streamlined dataset for subsequent model training.
[0090] S2.3: Establish an objective function containing a prediction error term and a regularization term using the evaluation index data to be trained;
[0091] This application embodiment utilizes high-quality training samples obtained through screening to construct an objective function containing a prediction error term and a regularization term, laying the foundation for the optimization of the single-index regression model. The design of the objective function is the core of model training, defining the direction of optimization and the evaluation criteria. For the prediction error term, this application embodiment mainly uses the mean squared error (MSE) to quantify the deviation between the model's predicted value and the actual value, expressed as: MSE=(1 / n)Σ(yi-f(xi)) 2 , where yi is the actual value of sample i, f(xi) is the model's prediction of sample i, and n is the number of samples. MSE has good mathematical properties (continuous and differentiable), which facilitates subsequent gradient calculation and optimization.
[0092] To prevent overfitting and improve generalization ability, this application introduces a regularization term into the objective function. Different types of regularization can be selected depending on the specific scenario: L1 regularization (Lasso) promotes model sparsity by adding the sum of the absolute values of the parameters (λΣ|θj|), which is helpful for feature selection; L2 regularization (Ridge) adds the sum of squared values of the parameters (λΣθj|). 2 Limiting parameter values prevents parameter inflation; Elastic Net combines the advantages of both, using α·λΣ|θj|+(1-α)·λΣθj 2 As a regularization term, λ is the regularization strength parameter, controlling the degree of regularization; α is a mixing parameter, controlling the ratio of L1 and L2 regularization. In embodiments of this application, these parameters are typically preset to initial values based on data characteristics and problem requirements, and adjusted during subsequent optimization.
[0093] Taking all the above into consideration, the complete objective function of the single-index regression model can be expressed as: J(θ)=(1 / n)Σ(yi-f(xi,θ)) 2 +R(θ), where R(θ) is the selected regularization term. This objective function considers both the model's prediction accuracy and its complexity and generalization ability, providing a clear optimization objective for subsequent parameter optimization.
[0094] S2.4: Calculate the gradient of the model parameters according to the objective function, update the model parameters using the first learning rate, and generate the single-index approximate optimal regression model.
[0095] An iterative optimization process for the model parameters was implemented. By calculating the gradient of the objective function and applying gradient descent, the model parameters were gradually adjusted, ultimately obtaining a near-optimal regression model for a single index. First, in this embodiment, the gradient of the objective function J(θ) with respect to each parameter θj is calculated, expressed as... Where p is the number of parameters. For mean squared error and common regularization terms, these gradients usually have analytical solutions and can be calculated directly; for complex models, numerical differentiation methods may be needed to approximate the gradients.
[0096] After obtaining the gradient, this embodiment of the application uses gradient descent to update the parameters: Here, α is the learning rate, which controls the step size of each iteration. A suitable learning rate is crucial for optimization performance: an excessively large learning rate may lead to instability or divergence in the optimization process, while an excessively small learning rate will make the convergence process too slow. To address this issue, this embodiment employs an adaptive learning rate strategy, dynamically adjusting the learning rate based on changes in the objective function during the optimization process. Specifically, when the objective function value continuously decreases over multiple iterations, this embodiment appropriately increases the learning rate to accelerate convergence; when the objective function fluctuates or increases, the learning rate is decreased to ensure stability.
[0097] This application also introduces an early stopping strategy to avoid excessive iteration and overfitting. Specifically, this application divides the data into a training set and a validation set, and evaluates the model's performance on the validation set after each iteration. When the validation set performance no longer improves for several consecutive iterations, or the improvement is less than a preset threshold, this application stops the iteration process. Furthermore, to avoid getting trapped in local optima, this application employs a strategy of multiple random initializations, starting optimization from different starting points, and selecting the model with the best final performance as the output. Through this series of optimization techniques, this application efficiently finds approximately optimal model parameters, constructing an accurate and efficient single-index regression model, laying a solid foundation for the subsequent construction of sparse generalized linear models.
[0098] like Figure 4 As shown, S3 specifically includes:
[0099] S3.1: Perform L1 regularization processing using the output of the single-index approximate optimal regression model to generate the filtered features;
[0100] This step is the starting point for constructing the sparse generalized linear model. In this embodiment, the L1 regularization process is performed using the output of the single-index approximate optimal regression model obtained in the previous stage to achieve effective feature selection. Specifically, L1 regularization (also known as Lasso regularization) introduces the sum of the absolute values of the parameters into the objective function as a penalty term, causing the model parameters to tend towards sparsity. The optimization problem can be expressed as: min(||y-Xβ|| 2 +λ||β||1), where y represents the target variable, X is the feature matrix, β is the parameter vector to be optimized, and λ is a coefficient controlling the strength of regularization. The unique feature of L1 regularization is that it forces certain parameters to become exactly zero, rather than merely becoming very small, making it a powerful tool for feature selection. In practical implementation, embodiments of this application employ efficient algorithms such as Coordinate Descent or Least Angle Regression (LARS) to solve this type of optimization problem, ensuring computational feasibility on large-scale datasets.
[0101] The choice of regularization strength parameter λ is crucial, as it directly affects the rigor of feature selection: a small λ value retains more features, potentially leading to an overly complex model; a large λ value may oversimplify the model, resulting in the loss of important information. This embodiment uses cross-validation to determine the optimal λ value. By trying a series of λ values and evaluating the performance of the corresponding models on the validation set, the best-performing λ value is selected. Furthermore, this embodiment considers the domain significance of features, setting lower regularization strengths for certain features deemed important based on professional knowledge (such as the charging and discharging efficiency and peak-valley electricity price differences in this embodiment of energy storage) to ensure they are not incorrectly removed during feature selection. Through this process, this embodiment effectively removes redundant or irrelevant features, retaining the most predictive subset of features, laying the foundation for building an efficient operational stress assessment model.
[0102] S3.2: Perform principal component analysis on the filtered features to construct the sparse feature matrix;
[0103] This application applies Principal Component Analysis (PCA) to the filtered feature set to further reduce feature dimensionality and handle correlations between features, constructing a sparse feature matrix with optimized structure. PCA is a classic unsupervised dimensionality reduction technique that transforms the original feature space into a set of orthogonal new feature spaces through linear transformation. Each principal component is a linear combination of the original features, ordered by variance. In specific implementation, this application first calculates the feature covariance matrix, then solves for its eigenvalues and eigenvectors. The magnitude of the eigenvalue corresponding to the eigenvector represents the variance explained by the corresponding principal component. This application arranges the eigenvectors in descending order of eigenvalues to form a transformation matrix. By multiplying the original features by this matrix, the representation in the principal component space is obtained.
[0104] When selecting the number of principal components to retain, this embodiment uses the cumulative explained variance ratio as a metric. Specifically, it selects the top k principal components, ensuring that their cumulative explained variance ratio reaches a preset threshold (typically 95%). This method guarantees the preservation of key information in the data during dimensionality reduction. For applications such as energy storage and new energy operation pressure assessment, this embodiment typically extracts 8-12 principal components from 20-30 original features, significantly reducing model complexity. Notably, this embodiment also considers the interpretability of principal components when applying PCA. By analyzing the load coefficients (i.e., elements of the eigenvectors) of the principal components, the physical meaning represented by each principal component is identified, such as the "charge / discharge efficiency-lifetime" principal component and the "market price fluctuation" principal component, enhancing the model's interpretability. Through PCA processing, this embodiment constructs a sparse feature matrix that retains key information and possesses good mathematical properties, providing an ideal data foundation for subsequent model optimization.
[0105] S3.3: Create a first hard threshold function using the data distribution characteristics of the sparse feature matrix to form a parameter threshold index;
[0106] An innovative hard thresholding mechanism is introduced, a key tool for maintaining model sparsity. The hard thresholding function is a non-linear operation used to force parameter sparsity during parameter updates. Specifically, the function is defined piecewise: when the absolute value of a parameter is less than the threshold λ, the parameter is set to 0; when the absolute value of a parameter is greater than or equal to λ, the parameter value remains unchanged. Mathematically, it is expressed as: HardThreshold(θ,λ)=θ·I(|θ|≥λ), where I(·) is an indicator function that takes a value of 1 when the condition is met, and 0 otherwise. This approach differs from the soft thresholding function (implicitly used in Lasso), which shrinks all parameter values, while the hard thresholding function only processes parameters smaller than the threshold, preserving the original values of larger parameters, which helps reduce estimation bias.
[0107] This application embodiment intelligently sets the parameters of the hard threshold function based on the data distribution characteristics of the sparse feature matrix. First, the statistical properties of the feature matrix are analyzed, including the variance and covariance structure of each dimension and the eigenvalue distribution. The initial threshold λ is typically set as a certain proportion of the feature standard deviation, with the specific proportion determined according to the required sparsity level; a commonly used initial setting is 0.1 times the feature standard deviation. This application embodiment also considers the differences in importance of different features, setting lower thresholds for important features (such as principal components that contribute significantly in principal component analysis) and higher thresholds for less important features, achieving adaptive feature selection. In this way, this application embodiment forms a set of parameter threshold indices, providing specific standards and implementation schemes for sparsity reduction in subsequent model optimization.
[0108] S3.4: Based on the parameter threshold index, iteratively optimize the model parameters using the gradient descent method and apply the first hard threshold function to perform parameter sparsification, generating an optimized sparse generalized linear model containing model performance indicators;
[0109] The core optimization process of the sparse generalized linear model is implemented. Through the synergistic effect of gradient descent and a hard thresholding function, the model's predictive performance is optimized while maintaining the high sparsity of the parameters. This application's embodiment employs an improved gradient descent algorithm for optimization, performing two key steps in each iteration: First, the parameter gradient is calculated based on the objective function, and the parameter values are updated using the gradient information. Where α is the learning rate; then, a hard threshold function is applied to the updated parameters to perform sparsification: θnew = HardThreshold(θ',λ). This alternating optimization strategy is called Iterative Hard Thresholding (IHT), which applies sparsification immediately after each gradient update to ensure that the sparse structure of the parameters is maintained throughout the optimization process.
[0110] To improve the stability and convergence of the algorithm, this application implements a series of optimization techniques. First, an adaptive learning rate strategy is adopted, dynamically adjusting the learning step size based on changes in the objective function value. Second, a momentum term is introduced to accelerate convergence, that is, the update direction of the previous iteration is considered when calculating parameter updates. Where β is the momentum coefficient. Furthermore, this embodiment employs a mini-batch gradient descent method combining batch gradient descent and stochastic gradient descent, using a subset of data to calculate the gradient in each iteration, balancing computational efficiency and optimization stability. During optimization, this embodiment continuously monitors the model's performance metrics, including prediction error (typically measured by mean squared error (MSE) or mean absolute error (MAE), model sparsity (proportion of non-zero parameters), and computational complexity. When the performance improvement over multiple iterations falls below a preset threshold, or when the maximum number of iterations is reached, this embodiment terminates the optimization process and outputs an optimized sparse generalized linear model incorporating these performance metrics.
[0111] S3.5: Perform cross-validation using the optimized sparse generalized linear model containing model performance metrics and the first hard threshold function to obtain the validated model performance metrics;
[0112] The model's performance is evaluated through a rigorous cross-validation procedure to ensure good generalization ability. This embodiment employs a k-fold cross-validation method (typically k = 5 or 10), randomly dividing the dataset into k equal-sized subsets. Each time, k-1 subsets are used as training data, and the remaining subset is used as validation data. This process is repeated k times, ensuring each subset serves as the validation set once. The average result is then used as an estimate of the model's performance. This method makes full use of limited data resources and provides a reliable assessment of the model's generalization ability. During validation, this embodiment simultaneously uses an optimized sparse generalized linear model and a hard thresholding function to ensure that the evaluation reflects the model's performance in real-world application environments.
[0113] This application evaluates various performance metrics to comprehensively understand the model's characteristics. Prediction accuracy metrics include mean squared error (MSE), mean absolute error (MAE), and coefficient of determination (R²). 2 Model structure metrics include parameter sparsity (proportion of non-zero parameters) and model complexity (number of effective parameters); computational efficiency metrics include training time and prediction time. Specifically, in the application scenario of this energy storage embodiment, this embodiment also evaluates the model's prediction stability at different time scales, such as the trend of prediction error changes within a day, week, and month, as well as its response capability to extreme situations (such as drastic market price fluctuations and abnormal equipment operation). Through these comprehensive verifications, this embodiment gains a deep understanding of model performance, providing a reliable basis for subsequent hyperparameter optimization.
[0114] S3.6: Based on the verified model performance indicators, the model hyperparameters are optimized using a grid search method to form the final optimized model parameters.
[0115] In the final stage of model optimization, this embodiment uses a grid search method to specifically optimize the model hyperparameters based on the validation results, ultimately forming the optimal model configuration. Grid search is an exhaustive hyperparameter optimization method that specifically tries all possible parameter combinations in a predefined parameter space to find the optimal parameter settings. In this embodiment, the main hyperparameters optimized include: regularization strength coefficient λ, which controls the strength of L1 regularization; hard threshold function parameter, which controls the degree of parameter sparsity; learning rate α, which controls the step size of gradient descent; and momentum coefficient β, which affects the stability and convergence speed of optimization, etc.
[0116] To improve search efficiency, this application employs a two-stage grid search strategy: first, a coarse-grained grid is used for an initial search over a wide parameter range to find parameter regions with good performance; then, a fine-grained grid is used around these regions for a refined search to determine the optimal parameters. For example, for the regularization coefficient λ, the initial search might be within the range of [0.001, 0.01, 0.1, 1, 10], and then the search is further refined within the range of [0.05, 0.08, 0.1, 0.12, 0.15] around the best-performing λ value (e.g., λ = 0.1). This application uses the same cross-validation scheme to evaluate the performance of each set of parameters during the search process, ensuring the consistency and comparability of the evaluation results.
[0117] In addition to prediction accuracy, this application embodiment also considers model complexity and computational efficiency when selecting the final parameters. For example, when the prediction performance of two sets of parameter configurations is similar, this application embodiment tends to choose the configuration with sparser parameters (fewer non-zero parameters) or lower computational complexity. Ultimately, this application embodiment outputs fully optimized model parameters, including the weight coefficients, regularization parameters, and hard threshold function parameters of the linear model. These parameters together constitute an efficient sparse generalized linear model for assessing the operational pressure of energy storage and new energy power plants. Through this series of meticulous optimizations, this application embodiment ensures that the model has excellent performance and high applicability in real-world application environments.
[0118] like Figure 5 As shown, S4 specifically includes:
[0119] S4.1: Utilize the real-time operational data to perform intraday, monthly, and annual forecasts, generating multi-scale operational pressure forecast results;
[0120] This application achieves multi-timescale operational pressure prediction based on an optimization model, providing a multi-dimensional perspective for comprehensively assessing the operational status of energy storage and new energy power plants. The embodiment first establishes a real-time data acquisition mechanism, obtaining the latest operational data from energy storage devices and new energy power generation units through a distributed data acquisition architecture. In the data acquisition stage, this embodiment employs a dual-caching mechanism to ensure data real-time performance and continuity. The primary cache is used for real-time data updates, while the backup cache stores historical data to handle abnormal situations such as network interruptions. The acquired real-time data includes operational parameters of the energy storage unit such as charging and discharging power, SOC status, cycle count, and equipment temperature; output power, environmental parameters, and status information of the new energy power generation unit; and market data such as real-time electricity prices and ancillary service prices. These data, after undergoing a standardized process consistent with previous processing, are used as input variables for the model.
[0121] This application's embodiments are based on the same sparse generalized linear model, but employ different forecasting strategies and parameter configurations to simultaneously forecast operational pressure at three time scales. Intraday forecasting (within 24 hours) primarily focuses on changes in operational pressure caused by immediate operational efficiency and short-term market price fluctuations. This application's embodiments use a sliding time window strategy, updating the forecast results at fixed time intervals (e.g., 15 minutes), and combining real-time meteorological data and power load forecasts to improve the accuracy of short-term forecasts. Monthly forecasting (1 week to 1 month) considers more medium-term factors such as equipment maintenance cycles, seasonal fluctuations in the power market, and monthly settlements, using a seasonally adjusted model based on historical data for forecasting, paying particular attention to the impact of cyclical events such as equipment maintenance and market policy adjustments on operational pressure. Annual forecasting (quarterly and annually) focuses on long-term strategic factors such as equipment depreciation, technological iteration, and energy policies. This application's embodiments combine historical data from the same period and industry development trends, using a time series model to conduct long-term trend forecasting, providing a reference for strategic planning and investment decisions.
[0122] S4.2: Perform weighted fusion calculation on the multi-scale operational pressure prediction results to form a fused operational pressure assessment result;
[0123] By using weighted fusion calculation, the prediction results from different time scales are integrated into a comprehensive operational stress assessment result, achieving comprehensiveness and balance in the assessment. This application embodiment employs a dynamic weighted fusion method, determining weight coefficients based on the historical accuracy of the prediction results at different time scales, current data quality, and business focus. Mathematically, this is expressed as: Pcomprehensive = wshort·Pshort-term + wmedium-term·Pmedium-term + wlong-term·Plong-term, where P represents the prediction result at each time scale, and w is the corresponding weight coefficient satisfying wshort + wmedium-term + wlong-term = 1. The determination of the weight coefficients employs an adaptive mechanism. This application embodiment periodically evaluates the historical performance of the prediction models at each time scale, finding the optimal weight configuration by minimizing the weighted combination of prediction errors.
[0124] In practical implementation, the embodiments of this application adjust the weighting emphasis according to the power plant type and operation mode. For example, for energy storage power plants that mainly provide intraday peak-shaving services, the short-term forecast weight is higher (wshort≈0.6-0.7); for energy storage power plants that cooperate with seasonal renewable energy power generation, the medium-term forecast weight is increased (wmedium≈0.4-0.5); for power plants that have just been put into operation or are considering expansion, the long-term forecast weight is relatively increased (wlong≈0.3-0.4). In addition, the embodiments of this application also consider weighting adjustments during special periods, such as increasing the short-term forecast weight during periods of sharp fluctuations in electricity market prices, and increasing the medium- and long-term forecast weights during periods of policy adjustment. Through this intelligent weighted fusion mechanism, the operating pressure assessment results generated by the embodiments of this application reflect both the immediate operating conditions and the medium- and long-term development trends, providing comprehensive support for management decisions.
[0125] S4.3: Utilize the fused operational stress assessment results to create a multi-level early warning threshold system;
[0126] Based on the integrated operational stress assessment results, a multi-level early warning threshold system was constructed to provide refined quantitative standards for risk management. This application embodiment designs four early warning levels: Normal (green), Attention (yellow), Warning (orange), and Danger (red), corresponding to different levels of operational stress. The thresholds are determined using a comprehensive method, combining historical statistical data analysis, expert judgment, and operational target considerations. Specifically, this application embodiment first analyzes the distribution characteristics of historical operational stress indicators and determines preliminary thresholds using statistical methods (such as the percentile method). For example, stress values below the 50th percentile are defined as Normal, 50-75th percentiles as Attention, 75-90th percentiles as Warning, and above the 90th percentile as Danger. Then, this application embodiment adjusts these preliminary thresholds based on industry expert experience, particularly considering the differences in characteristics and risk resistance capabilities of different types of power plants.
[0127] The system also incorporates a dynamic threshold adjustment mechanism, adjusting early warning thresholds in real time based on external factors such as seasonal changes, market conditions, and policy adjustments. For example, during peak electricity demand seasons, the system appropriately increases the threshold tolerance to avoid over-warning; during periods of abnormal market price fluctuations, it correspondingly reduces the threshold sensitivity to identify potential risks in advance. Furthermore, the system sets differentiated thresholds for different types of operational pressures, employing different evaluation criteria for different pressure sources such as equipment efficiency decline, market price fluctuations, and policy changes. This multi-dimensional, differentiated, and dynamically adjusted early warning threshold system enables the system to provide accurate risk identification for complex and ever-changing operating environments, significantly improving the practicality and effectiveness of early warnings.
[0128] S4.4: The operational pressure level is graded and assessed according to the multi-level early warning threshold system, and the early warning information is output.
[0129] This system implements a multi-level early warning threshold system for classifying and assessing operational stress and generating early warning information, providing managers with timely and clear decision support. When the system detects that operational stress exceeds a specific threshold, it triggers the corresponding level of early warning mechanism. The specific classification and assessment process is as follows: the system compares the integrated operational stress assessment results with the early warning thresholds to determine the current stress level; then, it judges the urgency of the early warning based on the stress change trend (such as continuous increase or increased fluctuation); finally, it comprehensively considers the stress level and urgency to generate the final early warning level. The early warning information includes not only the stress level but also an analysis of the main factors leading to the increase in stress, an assessment of potential impacts, and response suggestions, providing decision-makers with a comprehensive understanding of the situation.
[0130] The system employs a decision tree algorithm to trace and analyze key factors leading to increased pressure. For example, when it detects increased operational pressure due to decreased charging and discharging efficiency, the system automatically analyzes the specific cause, such as equipment aging, improper operation, or changes in the external environment, and provides corresponding handling suggestions. Early warning information is delivered through multiple channels, including system console display, email notifications, SMS alerts, and mobile application push notifications, ensuring that critical information is delivered to responsible personnel in a timely manner. For different levels of warnings, the system has set differentiated information dissemination strategies: alerts at the "attention" level are periodically summarized and sent; alerts at the "warning" level are immediately pushed to the operations team; and alerts at the "danger" level simultaneously notify management and the operations team, triggering emergency response procedures.
[0131] The system also provides multi-dimensional visualization capabilities to help users intuitively understand early warning information. The real-time monitoring panel displays the current operational pressure level and key indicator status in a dashboard format, using color coding to intuitively represent the warning level. Trend analysis charts display the historical trends and predicted trends of each indicator through an interactive graphical interface, helping to identify abnormal patterns. The comprehensive assessment report automatically generates a document containing detailed analysis conclusions and improvement suggestions, supporting in-depth research and decision-making reference. Users can customize the displayed content and report format according to their focus, meeting the information needs of different management levels. Through this refined and intelligent early warning mechanism, the system provides strong technical support for the refined management and risk control of energy storage and new energy power plants, significantly improving the operational safety and profitability stability of the power plants.
[0132] This application also provides an operational stress assessment system based on energy storage and new energy power plants. The system includes a data acquisition module, a model building module, a model optimization module, and an assessment module, corresponding to the four main steps in the aforementioned method. The data acquisition module is responsible for collecting multi-source heterogeneous data and preprocessing it; the model building module is responsible for constructing a single-index approximate optimal regression model; the model optimization module is responsible for constructing and optimizing a sparse generalized linear model; and the assessment module is responsible for assessing and issuing early warnings based on real-time data. The entire system adopts a modular design, with data exchange between various functional modules through standardized interfaces, ensuring system maintainability and facilitating future functional expansion.
[0133] This application addresses the problems of high model complexity, low computational efficiency, and difficulty in effectively handling the sparse correlation characteristics between evaluation indicators in existing technologies. Compared with existing technologies, this application has the following advantages: 1) A single-indicator approximate optimal regression model using an active learning strategy obtains local optima through iterative optimization, reducing computational complexity and improving model training efficiency; 2) A sparse generalized linear model based on a hard threshold function maintains model sparsity through iterative learning, effectively handling the sparse correlation characteristics between evaluation indicators and improving the model's generalization ability; 3) Multi-level data preprocessing and feature extraction methods improve data quality and model accuracy; 4) A real-time evaluation and early warning mechanism supports dynamic operational pressure assessment, providing timely and accurate decision support for the operation of new energy power plants and energy storage systems.
[0134] This application also provides an operational stress assessment system based on energy storage and new energy power plants, including:
[0135] The data acquisition module is used to collect energy storage system operation data, new energy power generation data and electricity market data, and to clean and standardize the energy storage system operation data, the new energy power generation data and the electricity market data to obtain a standardized evaluation index dataset.
[0136] The model building module is used to construct an index correlation matrix using the standardized evaluation index dataset and select training samples using an active learning strategy, and obtain a single index approximate optimal regression model through iterative optimization.
[0137] The model optimization module is used to construct a sparse feature matrix based on the output of the single-index approximate optimal regression model, and to optimize the sparse feature matrix using an iterative hard threshold learning method to form an optimized sparse generalized linear model.
[0138] The evaluation module is used to collect real-time operating data, input the real-time operating data into the optimized sparse generalized linear model, and output the operational pressure evaluation results and early warning information.
[0139] This application embodiment also provides a computer device, the computer device comprising:
[0140] At least one processor; and a memory communicatively connected to said at least one processor;
[0141] The memory stores instructions that can be executed by the at least one processor, which are then executed by the at least one processor to enable the at least one processor to perform the above-mentioned operational stress assessment method based on energy storage and new energy power plants.
[0142] This application also provides a computer-readable storage medium that stores computer instructions for causing a computer to execute the above-described method for assessing the operational stress of energy storage and new energy power plants.
[0143] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above-described method for assessing the operational pressure of energy storage and new energy power plants.
[0144] This application has the following technical effects:
[0145] By using an active learning strategy to create a single-index approximate optimal regression model, computational complexity is reduced and model training efficiency is improved.
[0146] An iterative hard threshold learning method is used to construct a sparse generalized linear model, which effectively handles the sparse correlation characteristics between evaluation indicators and improves the model's generalization ability.
[0147] Multi-level data preprocessing and feature extraction methods improve data quality and enhance model accuracy;
[0148] Real-time assessment and early warning mechanisms support dynamic operational stress assessments, providing timely decision support for the operation of new energy power plants and energy storage systems.
[0149] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0150] This disclosure also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program performs the steps of the operational stress assessment method and system based on energy storage and new energy power plants described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0151] In addition, this disclosure also provides a computer program product, which stores a computer program. When the computer program is run by a processor, it executes the steps of the operation pressure assessment method and system based on energy storage and new energy power plants provided in any of the above embodiments of this disclosure. For details, please refer to the above method embodiments, which will not be repeated here.
[0152] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0153] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices and apparatuses described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0154] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0155] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0156] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0157] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A method for assessing the operational pressure of energy storage and new energy power plants, characterized in that, include: Collect energy storage system operation data, new energy power generation data, and electricity market data; clean and standardize the energy storage system operation data, the new energy power generation data, and the electricity market data to obtain a standardized evaluation index dataset. Using the standardized evaluation index dataset, an index correlation matrix is constructed and an active learning strategy is used to select training samples. Through iterative optimization, a near-optimal regression model for a single index is obtained. A sparse feature matrix is constructed based on the output of the single-index approximate optimal regression model. The sparse feature matrix is then optimized using an iterative hard threshold learning method to form an optimized sparse generalized linear model. Collect real-time operational data, input the real-time operational data into the optimized sparse generalized linear model, and output operational pressure assessment results and early warning information.
2. The method according to claim 1, characterized in that, The energy storage system operation data, the new energy power generation data, and the electricity market data are cleaned and standardized to obtain the standardized evaluation index dataset, including: The energy storage system operation data, the new energy power generation data, and the electricity market data are used to detect and process outliers, resulting in cleaned data. Based on the cleaned data, the utilization rate, charging and discharging efficiency, equipment depreciation rate, power generation efficiency, and electricity price fluctuation characteristics of the energy storage system are extracted to obtain characteristic data; The feature data is subjected to Z-score standardization to generate the standardized evaluation index dataset.
3. The method according to claim 1, characterized in that, Using the standardized evaluation index dataset, an index correlation matrix is constructed, and an active learning strategy is employed to select training samples, including: The correlation coefficients between the evaluation indicators are calculated using the standardized evaluation indicator dataset to generate the indicator correlation matrix; The sample information entropy is calculated based on the correlation matrix of the indicators, and training samples are selected using an uncertainty sampling strategy to form evaluation indicator data to be trained.
4. The method according to claim 3, characterized in that, A near-optimal regression model for a single index is obtained through iterative optimization, including: An objective function containing a prediction error term and a regularization term is established using the evaluation index data to be trained; The gradient of the model parameters is calculated based on the objective function, and the model parameters are updated using the first learning rate to generate the single-index approximate optimal regression model.
5. The method according to claim 1, characterized in that, Based on the output of the single-index approximate optimal regression model, a sparse feature matrix is constructed, including: The L1 regularization process is performed on the output of the single-index approximate optimal regression model to generate the filtered features. Principal component analysis is performed on the filtered features to construct the sparse feature matrix.
6. The method according to claim 1, characterized in that, The sparse feature matrix is optimized using an iterative hard thresholding learning method, including: A first hard threshold function is created using the data distribution characteristics of the sparse feature matrix to form a parameter threshold index; Based on the parameter threshold index, the model parameters are iteratively optimized using the gradient descent method, and parameter sparsification is performed using the first hard threshold function to generate an optimized sparse generalized linear model containing model performance indicators.
7. The method according to claim 6, characterized in that, Also includes: The optimized sparse generalized linear model, which includes model performance metrics, is validated and tuned, including: Cross-validation is performed using the optimized sparse generalized linear model containing model performance metrics and the first hard threshold function to obtain the validated model performance metrics. Based on the verified model performance metrics, the model hyperparameters are optimized using a grid search method to form the final optimized model parameters.
8. The method according to claim 1, characterized in that, Inputting the real-time running data into the optimized sparse generalized linear model includes: The real-time operational data is used to perform intraday, monthly, and annual forecasts, generating multi-scale operational pressure forecast results. The multi-scale operational stress prediction results are weighted and fused to form a fused operational stress assessment result.
9. The method according to claim 8, characterized in that, Also includes: Early warning information is generated based on the integrated operational pressure assessment results, including: A multi-level early warning threshold system is created using the fused operational stress assessment results; The operational pressure level is graded and assessed according to the multi-level early warning threshold system, and the early warning information is output.
10. An operational stress assessment system based on energy storage and new energy power plants, characterized in that, include: The data acquisition module is used to collect energy storage system operation data, new energy power generation data and electricity market data, and to clean and standardize the energy storage system operation data, the new energy power generation data and the electricity market data to obtain a standardized evaluation index dataset. The model building module is used to construct an index correlation matrix using the standardized evaluation index dataset and select training samples using an active learning strategy, and obtain a single index approximate optimal regression model through iterative optimization. The model optimization module is used to construct a sparse feature matrix based on the output of the single-index approximate optimal regression model, and to optimize the sparse feature matrix using an iterative hard threshold learning method to form an optimized sparse generalized linear model. The evaluation module is used to collect real-time operating data, input the real-time operating data into the optimized sparse generalized linear model, and output the operational pressure evaluation results and early warning information.