Industrial load prediction method, system and equipment fusing standard mutual information and improved bidirectional LSTM (Long Short Term Memory)
Through improved fuzzy C-mean clustering and standard mutual information combined with bidirectional LSTM neural network and Transformer multi-head attention mechanism, the problem of insufficient modeling of external nonlinear factors in the prior art is solved, and a higher precision medium- and long-term load prediction is achieved.
Patent Information
- Application Number
- CN202511081774.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-08-04
AI Technical Summary
The prior art over-reliance on the timing continuity of load historical data in medium and long-term load prediction, making it difficult to effectively model the nonlinear correlation of external multi-source heterogeneous factors, and traditional correlation analysis methods are difficult to adapt to the dynamic response characteristics of power loads, resulting in insufficient prediction accuracy.
The improved fuzzy C-mean clustering algorithm is used to combine kernel density estimation for data preprocessing, and nonlinear correlation is quantified through standard mutual information, and a load prediction model is constructed using a bidirectional LSTM neural network and a Transformer multi-head attention mechanism, and feature fusion is performed by combining convolutional neural networks.
It significantly improves the accuracy and robustness of load prediction, can better adapt to complex and changeable power system environment, and improves the adaptability and accuracy of the prediction model.
Smart Images

Figure CN120562841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and in particular to an industry load forecasting method, system, and device integrating standard mutual information and improved bidirectional LSTM. Background Art
[0002] Accurate medium- and long-term load forecasting is the core technical support for ensuring stable power system operation and the effectiveness of grid planning. Its results directly impact the progress of new power system construction and the optimal allocation of electric energy resources. With the development of smart grid construction, improving forecast accuracy in complex and changing operating environments has become a key research direction in the industry.
[0003] Currently, mainstream monthly load forecasting technologies in the industry focus on two main approaches: extrapolation forecasting methods based on time series characteristics and intelligent forecasting methods based on machine learning. Time series analysis methods primarily include exponential smoothing, gray forecasting, and the similar day method, while the latter utilizes machine learning algorithms such as artificial neural networks and support vector regression. Existing technologies generally suffer from two limitations: First, they overly rely on the temporal continuity of historical load data and lack effective modeling of the nonlinear correlations of heterogeneous external factors, such as temperature fluctuations and industrial policy adjustments. Second, traditional correlation analysis methods struggle to adapt to the dynamic response characteristics of power load. For example, linear correlation tests based on the Pearson coefficient cannot capture the high-order nonlinear coupling relationships found in real-world scenarios. While Copula functions can construct nonlinear correlations, they overly rely on a priori parameterization assumptions of the correlation function, resulting in generalization issues in the feature correlation model. Existing technologies are particularly prone to systematic biases when sudden external disturbances trigger sudden changes in power demand. Furthermore, the correlation analysis results of the aforementioned studies only select a subset of external factors with strong correlations with load, failing to fully quantify the impact of each factor on power load.
[0004] Therefore, current technology faces two key areas for improvement: 1. A multimodal data fusion architecture is urgently needed to overcome the limitations of unidirectional time series analysis; 2. Nonparametric estimation methods that can quantify nonlinear relationships are urgently needed. By effectively combining temporal and spatial characteristics with diverse data, while accurately capturing load variations and constructing effective coupling models for dynamic characteristics such as meteorological changes and economic fluctuations, this will become a key research direction for breaking through the bottleneck of monthly load forecasting technology in the industry. Summary of the Invention
[0005] The technical problem to be solved and the technical task to be addressed by this invention are to improve and enhance existing technical solutions by providing an industry load forecasting method, system, and device that integrates standard mutual information with an improved bidirectional LSTM, with the goal of improving the accuracy of monthly industry load forecasts. To this end, this invention adopts the following technical solution.
[0006] The first technical solution of the present invention is an industry load forecasting method integrating standard mutual information and improved bidirectional LSTM, comprising the following steps: 1) An improved fuzzy C-means clustering algorithm taking into account kernel density estimation is used to cluster information including load series, consumption levels, temperature, and vacations; 2) Based on the clustering results, standard mutual information is calculated to quantitatively analyze the correlation between information factors such as consumption level, temperature, and vacation time and the monthly load of the industry; 3) Based on the correlation analysis results, weights are assigned and a bidirectional LSTM neural network is constructed to capture the temporal variation patterns of industry loads. The Transformer multi-head attention mechanism is used to analyze the impact of external features, and the monthly industry load forecast results are output through convolutional neural network connections.
[0007] This technical solution addresses the key bottlenecks of traditional industry monthly load forecasting, including over-reliance on linear time series, difficulty in effectively integrating multi-source heterogeneous nonlinear external factors, and poor adaptability to complex dynamic environments, through "clustering preprocessing to reduce noise and improve efficiency → standard mutual information to accurately quantify nonlinear correlations → dynamic integration of correlation weights → bidirectional LSTM + multi-head attention spatiotemporal modeling." This significantly improves the robustness and accuracy of the forecasting model, as well as its adaptability to complex real-world scenarios (such as outliers, nonlinear relationships, and external mutations), providing more reliable data support for energy planning and scheduling. Specifically: An improved fuzzy C-means clustering algorithm that incorporates kernel density estimation is used to jointly cluster load series and external factors (consumption levels, temperature, and vacations), effectively improving the efficiency and accuracy of subsequent standard mutual information calculations and enhancing robustness to outliers. By identifying the primary patterns and inherent structure in the data, the algorithm reduces noise, redundancy, seasonal fluctuations, and interference from mixed patterns in the raw data, providing a more reliable and homogeneous data foundation for mutual information calculations. Direct mutual information calculations on raw high-dimensional data are prone to distortion and computationally intensive. Clustering preprocessing significantly reduces computational complexity, especially in large-scale industrial load data scenarios. Kernel density estimation is used to optimize the selection of initial cluster centers, overcoming the traditional fuzzy C-means algorithm's sensitivity to initial values and reducing the risk of falling into local optimal solutions. This improves clustering quality and lays a more solid foundation for subsequent analysis.
[0008] Standardized mutual information (SMI) effectively captures and quantifies complex nonlinear dependencies between variables. It addresses the fundamental limitation of traditional linear correlation coefficients (such as the Pearson coefficient) in industry load analysis—linear methods can fail when load exhibits complex nonlinear relationships with factors such as temperature and consumption levels. Based on information theory, SMI does not rely on specific data distribution assumptions, enabling a more scientific and comprehensive assessment of the true impact of various factors on load.
[0009] The bidirectional LSTM model simultaneously leverages forward and backward information from historical sequences to more comprehensively understand and capture the inherent long- and short-term temporal dependencies and changing trends of the industry's monthly load. The Transformer multi-head attention mechanism analyzes the impact of external features (consumption levels, temperature, and vacations). This allows for the dynamic and parallel capture of complex cross-dimensional relationships between these features and between them and the forecast target, making it particularly adept at identifying the heterogeneous impact of key events (such as sudden temperature changes and holidays). Finally, a convolutional neural network connects the bidirectional LSTM (temporal features) and the Transformer (external feature influences) to produce a comprehensive forecast. This architecture, which deeply integrates temporal patterns with external dynamic features, significantly improves the model's adaptability and forecast accuracy in complex and changing environments.
[0010] An improved bidirectional LSTM neural network is used for load forecasting. Compared to traditional unidirectional LSTM or other forecasting models, the bidirectional LSTM can simultaneously consider both forward and backward information in time series data, providing a more comprehensive understanding of the data's temporal dependencies and improving the accuracy of monthly load forecasts. Furthermore, a multi-head attention mechanism is employed to incorporate cluster-enhanced features into the forecasting model to account for external factors. This approach ensures that forecasts better reflect actual conditions, resulting in higher forecast accuracy, especially in complex and changing external environments.
[0011] This method combines clustering algorithms, mutual information analysis, and deep learning models, making it applicable to a wide range of industries and load types, demonstrating strong versatility and adaptability. By combining kernel density estimation, standard mutual information, and an improved bidirectional LSTM, it effectively improves the accuracy of monthly industry load forecasts, reduces forecast errors, and provides more reliable data support for energy planning and scheduling.
[0012] As a preferred technical means: the clustering process in step 1) further includes: 1.1) Determine the membership matrix: Based on the improved fuzzy C-means clustering algorithm, set the number of clusters, select the initial cluster centers based on the kernel density estimation method, and calculate the membership between the monthly load data and the cluster centers; 1.2) Constructing the objective function: The objective function is set to minimize the Euclidean distance between the monthly load data and the cluster center; 1.3) Update cluster centers: Update the cluster center matrix by iteratively optimizing the objective function until the objective function converges; 1.4) Obtaining the optimal clustering result: Obtain the clustering result based on the minimized objective function value; 1.5) Determine the optimal number of clusters: Select the optimal number of clusters using the Xie & Ben index and determine the final clustering result.
[0013] The fuzzy C-means clustering algorithm allows data points to belong to multiple clusters simultaneously with different degrees of membership, which improves the flexibility of data processing and can effectively identify the complex change patterns and inherent structures in monthly load data. The algorithm adopts the objective function of minimizing the Euclidean distance between data points and cluster centers, and gradually updates the cluster centers through iterative optimization to ensure that the results converge to the global optimum. The Xie & Ben indicator is used to dynamically determine the optimal number of clusters, avoiding reliance on preset parameters and improving the adaptability and intelligence of clustering. The cluster centers calculated through multiple iterative calculations can accurately reflect the distribution characteristics of load data and reduce analytical bias. The clustered monthly load data provides an excellent foundation for correlation analysis and predictive modeling, and can more accurately identify the underlying patterns in the data. The iterative optimization process reduces the computational burden of full-space search, effectively reducing the complexity of the algorithm and improving operational efficiency.
[0014] As a preferred technical means: Step 1.1) Determine the membership matrix and set the number of clusters based on the number of months , using the kernel density estimation method from the electricity consumption data matrix Y Before the selection N l The data with the largest row density is used as the initial cluster center matrix V =[ V 1,…, V l ,…, V Nl ] T , and then calculate any monthly data Y m Belong to l The membership degree of each cluster center.
[0015] Step 1.2) Construct the objective function, the objective function G The sum of the squares of the weighted distances of membership between all historical months and all cluster centers; Step 1.3) Update the cluster center, if the objective function G If the minimum value is not reached, the cluster center matrix is updated according to the given update strategy. V Update and recalculate the membership matrix until the objective function G Reach minimum; Step 1.4) Obtain the optimal clustering result from the membership matrix and cluster center matrix V The number of clusters obtained is N l The optimal clustering result corresponding to the minimum value of the objective function when ; Step 1.5) In determining the optimal number of clusters, the Xie & Ben index is used to reflect the degree of compactness within the cluster and the degree of separation between clusters. The clustering quality index corresponding to different numbers of clusters is compared. I XB , filter out the best number of clusters , and obtain the best clustering result.
[0016] By clustering historical monthly load data and external influencing factor data, the data is categorized before correlation analysis. This allows the data to more clearly demonstrate their underlying patterns, improving the accuracy of subsequent correlation analysis and ensuring more reliable results. The membership matrix is flexible and accounts for data uncertainty, allowing data points to belong to multiple cluster centers simultaneously. The fuzzy processing approach better captures complex data characteristics, reduces errors, and increases the robustness of the algorithm. The Xie & Ben index dynamically evaluates clustering effectiveness at different numbers of clusters based on intra-cluster compactness and inter-cluster separation, enabling the optimal number of clusters to be selected. This reduces manual intervention and ensures a scientific and rational selection of cluster numbers. Cluster center updates are iteratively optimized based on the objective function of minimizing the Euclidean distance. During each iteration, cluster centers are adjusted through rapid calculations to ensure the objective function reaches its minimum value in a short period of time, reducing computational complexity and improving algorithm efficiency. By optimizing the membership matrix and cluster center matrix, each data point is optimally assigned to its corresponding cluster. This not only improves the quality of the clustering results but also provides a more accurate foundation for subsequent data analysis based on the clustering results. This method can effectively process load data and multidimensional external influencing factor data, fully utilizing the characteristics of the data for clustering and analysis, and ensuring that the clustering results accurately reflect the inherent structure and similarity of the data.
[0017] As a preferred technical means: step 2) further comprises the steps of: 2.1) Data Fuzzification Based on the clustering results obtained by the improved fuzzy C-means (DFCM) clustering algorithm in step 1), the original external influencing factor value x(m) and monthly load value y(m) are replaced by the cluster center value corresponding to each sampling point to obtain the fuzzy processed sequences U and V; 2.2) Standard mutual information calculation 2.2.1) Mutual Information Calculation: The mutual information I(U;V) between the fuzzified data U and V is calculated to quantitatively describe the correlation between the influencing factors and the monthly load. 2.2.2) Information entropy calculation: Calculate the information entropy H(U) and H(V) of sequences U and V; 2.2.3) Standard mutual information calculation: Calculate the standard mutual information value J(U;V) between the influencing factors and the monthly load by using the ratio of mutual information I(U;V) to information entropy H(U) and H(V); 2.3) Quantitative evaluation of correlation 2.3.1) Result Analysis: Based on the size of the standard mutual information value J(U;V), the correlation between external influencing factors and the industry's monthly load is quantitatively evaluated. The larger the value, the stronger the correlation. 2.3.2) Parameter input: The correlation analysis results based on standard mutual information are used as network hyperparameters and input into the monthly load forecasting network to optimize the model's forecasting ability.
[0018] A fuzzy clustering algorithm is used to preprocess raw data, mapping high-dimensional data to the feature space of cluster centers. This achieves data noise reduction and feature enhancement, establishing a stable data foundation for correlation analysis. Based on the standard mutual information theory, a quantitative model of the dependency relationship between external factors and industry loads is constructed, accurately characterizing this dependency relationship. Compared to traditional correlation analysis, standard mutual information can deeply explore nonlinear correlation patterns and complex interactions between variables. Correlation analysis results are directly used as model hyperparameters, dynamically assigning weights based on the importance of each influencing factor. This enhances the model's responsiveness to key factors, improving forecasting accuracy and stability. Data uncertainty is measured by calculating sequence information entropy, and the standard mutual information metric is constructed based on mutual information to comprehensively assess the relative importance of influencing factors. This method balances the direct correlation between data and information complexity, ensuring the scientific and accurate analysis results. The standard mutual information value intuitively reflects the strength of the correlation between factors and loads, helping to select key predictive variables and appropriately assign model weights. This data-driven parameter optimization strategy can effectively improve model performance and forecasting results. By accurately identifying and quantitatively analyzing the influencing mechanisms of external factors, the model possesses enhanced environmental adaptability and scenario generalization capabilities, making it widely applicable to various load forecasting tasks. Clustering-optimized data makes the calculation of standard mutual information more efficient and accurate, avoiding the complexity of processing high-dimensional raw data and improving overall analysis efficiency.
[0019] As a preferred technical means: In step 2.1) data fuzzification processing, for the monthly load of the industry to be analyzed and the external influencing factors, it is known that within a time period N m The influencing factor value of each sampling pointx ( m ) and monthly load values y ( m ), m ∈{1,2,…, N m},all x ( m ) constitutes a sequence X ,all y ( m ) constitutes a sequence Y ; Use the data after DFCM clustering to replace the original data, calculate the standard mutual information, and record the sampling value x ( m The cluster center value of the cluster where u ( m ),all u ( m ) constitutes a sequence U , sampling value y ( m The cluster center value of the cluster where v ( m ),all v ( m ) constitutes a sequence V ; u ( m )Total F Possible values, F ≤ N m , recorded as , f ∈{1,2,…, F}, v ( m )Total G Possible values, G ≤ N m , recorded as , ∈{1,2,…, G}; In step 2.2) of the standard mutual information calculation, the standard mutual information value between the influencing factors and the monthly load is J ( U ; V ) is calculated as follows:
[0020] Where,
[0021]
[0022]
[0023] Where, I ( U ; V )for u ( m )and v ( m )’s mutual information value; H ( U )and H ( V ) are respectively U and V Information entropy; To satisfy simultaneously u ( m )= and v ( m )= The proportion of sampling points to all sampling points; To satisfy u ( m )= The proportion of sampling points to all sampling points; To satisfy v ( m )= The proportion of sampling moments to all sampling moments; J ( U ; V ) ranges from [0,1]. The larger the value, the stronger the correlation between the influencing factor and the monthly load of the industry.
[0024] Standardized mutual information (SMI) quantitatively assesses the nonlinear relationship between monthly industry load and external influencing factors, avoiding complex correlations that traditional linear correlation analysis methods may overlook, thereby more comprehensively reflecting the actual connection between the two. By fuzzifying the original data and replacing the original numerical values with fuzzy C-means clustering results, this method reduces errors caused by subtle numerical differences, allowing the analysis to focus more on the overall trends and key features of the data rather than being distracted by minor changes, thereby improving the accuracy and robustness of the correlation analysis. Correlation analysis results based on SMI can be used as network hyperparameter inputs for the prediction model, helping to optimize the model structure and better adapt it to different external influencing factors and monthly load fluctuation scenarios, thereby improving prediction accuracy and generalization. By calculating the mutual information value and information entropy and further normalizing the resulting SMI, it is possible to quantitatively assess the strength of the correlation between the influencing factors and monthly load. Since the SMI value ranges from [0 to 1], this metric is intuitive and has a clear physical meaning, facilitating more informed model decisions in practical applications. By mapping complex raw data to fewer cluster center values, computational complexity is reduced, alleviating the model's burden in processing large-scale, high-dimensional data and thereby improving model efficiency. Using fuzzy C-means clustering to process data makes the analysis results more robust to noise and outliers, enhancing the stability and reliability of the load forecasting model in practical applications. Using the results of standard mutual information analysis as a network hyperparameter enables the forecasting model to more specifically learn and adapt to key external influencing factors, thereby optimizing network performance and improving forecast accuracy.
[0025] As a preferred technical means: the monthly load forecasting of step 3) includes the following steps: 3.1) Correlation weight calculation: Normalize the standard mutual information value between the industry monthly load and each external influencing factor to determine the characteristic weight of the external influencing factor; 3.2) Eigenvector Weighting: Based on the calculated eigenweights, the eigenvectors of all external influencing factors are weighted and summed to obtain the enhanced comprehensive eigenvector. 3.3) Bidirectional LSTM Network Modeling: Using a bidirectional long short-term memory network, the time series feature vector of historical monthly load is used as input to capture the time series trend of monthly load in the industry. 3.4) Modeling the Transformer Multi-Head Attention Mechanism: Applying the Transformer multi-head attention mechanism, we take multi-dimensional external feature vectors and integrated feature vectors as input and capture dynamic correlations across feature dimensions through parallel attention heads. 3.5) Network Optimization and Output: The time series feature vectors extracted by the bidirectional LSTM and the multi-head attention features encoded by the Transformer are combined through a convolutional neural network to construct a feature interaction layer. Finally, the fully connected layer outputs the environmentally adaptable monthly industry load forecast value.
[0026] By calculating the correlation weights in step 3.1 and normalizing the standard mutual information values, we obtain the feature weights of the external influencing factors, enabling a more accurate reflection of the impact of each factor on the monthly load forecast. The weighted feature vectors are further processed in step 3.2, enhancing the feature representation and thus deepening the model's understanding of the data. Using the combined feature vectors of external influencing factors and the time series feature vectors of historical monthly loads as input to the improved bidirectional LSTM network helps the model better capture the complex dynamics of monthly load variations. The weighted summation of the combined feature vectors makes the input features more representative and important, improving the model's forecasting accuracy. The bidirectional LSTM network combines the advantages of both forward and backward layers, leveraging both past and future information in time series data to more comprehensively capture the regularity of monthly load variations. This bidirectional processing approach better leverages historical data and external influencing factors, resulting in more accurate forecasts. Through the modeling and optimization of the bidirectional LSTM network and the Transformer multi-head attention mechanism in steps 3.3 and 3.4, the final monthly load forecast output fully reflects the complex temporal dependencies and the influence of external influencing factors. Bidirectional networks can more effectively process long-sequence data, reducing information loss caused by large time spans. The multi-head attention mechanism, through the parallel application of multiple self-attention heads, dynamically captures global feature dependencies from different dimensions, enhancing the model's adaptability to multi-scale temporal patterns and compensating for the limitations of the bidirectional LSTM model in external feature input. By comprehensively considering the feature weights of various external influencing factors and processing information bidirectionally, the model can more effectively utilize meaningful data features, reducing the risk of overfitting due to redundant information or noise, making the model more robust to new data and demonstrating good generalization. Combining the forward and backward layers of the bidirectional LSTM network allows for flexible adjustment of the network structure and optimization strategy to achieve optimal prediction results, ensuring the model remains efficient and accurate in handling complex prediction tasks.
[0027] As a preferred technical means: the bidirectional LSTM neural network model captures the long-term and short-term dependencies of the load sequence based on the bidirectional time series feature extraction module, and controls the transmission of information through the forget gate, input gate, and output gate. The formulas for calculating the forward and backward monthly load forecast outputs are as follows:
[0028]
[0029]
[0030] Where, Represents the forward calculation value; Represents the backward calculated value; o t Represents the final output value calculated by the forward and backward values; ω 1 is the weight matrix mapping the input layer to the forward layer; ω 3 is the weight matrix mapping the input layer to the backward layer; ω 2 is the weight matrix mapping the output of the forward layer at the previous moment to the current calculation moment; ω 5 is the weight matrix that maps the output of the backward layer at the previous moment to the current calculation moment; ω 4 is the weight matrix mapping the output of the forward layer to the output layer; ω 6 is the weight matrix mapping the output of the backward layer to the output layer; The Transformer multi-head attention mechanism maps input features to different subspaces by calculating multiple sets of attention weights in parallel. It captures the dynamic correlation between data in each subspace, effectively identifying the heterogeneous impact of key external events on load, such as sudden temperature changes and holiday effects. Single-head attention is calculated using the following formula:
[0031] Where, Q 、 K 、 V is the query key-value matrix; is the scaling factor; The model is divided into multiple attention heads to form multiple subspaces, allowing the model to focus on different aspects of information. The specific expression is:
[0032]
[0033]
[0034] Where, Q i 、 K i 、 V i For the i The query key-value matrix of the attention heads; 、 、 For the i The weight matrix of the attention head; is the output projection matrix; h is the number of attention heads.
[0035] Although bidirectional LSTMs can process sequential time series information, they only process data sequentially, making it difficult to focus on relationships between global time points and unable to leverage future weather forecasts. The Transformer multi-head attention mechanism calculates multiple sets of attention weights in parallel, mapping input features into different subspaces. This mechanism captures dynamic correlations between data in each subspace, effectively identifying the heterogeneous impact of key external events such as sudden temperature changes and holiday effects on load. This multi-head attention mechanism effectively addresses the shortcomings of bidirectional LSTMs.
[0036] The second technical solution of the present invention is to provide an industry load forecasting system that integrates standard mutual information and improved bidirectional LSTM, and the system is used to implement the aforementioned industry monthly load forecasting method, and the system includes: Data receiving module: used to receive historical industry monthly load data and external influencing factor data; Clustering processing module: clustering the received data using the improved fuzzy C-means clustering algorithm based on kernel density estimation; Correlation analysis module: performs standard mutual information calculation based on clustering results to quantitatively analyze the correlation between external influencing factors and the industry's monthly load; Monthly load forecasting module: uses a bidirectional LSTM neural network and combines the results of correlation analysis to predict the industry's monthly load.
[0037] The data receiving module effectively integrates historical monthly industry load data and data on external influencing factors, providing comprehensive data support for subsequent analysis and forecasting. Data clustering using the fuzzy C-means clustering algorithm can better identify patterns and structures within the data, improving data understanding. Standardized mutual information is used to quantify the relationship between external influencing factors and monthly industry load, revealing key factors influencing monthly load fluctuations and enhancing forecast accuracy. A bidirectional LSTM (bidirectional long short-term memory) network captures temporal dependencies in the data and incorporates bidirectional information from historical data, thereby improving the accuracy and reliability of monthly load forecasts. The system integrates cluster analysis, correlation analysis, and deep learning methods to fully leverage the information in the data and provide a more comprehensive and accurate monthly load forecast. The system can adjust to changes in monthly industry load and external influencing factors, making forecast results more realistic and adaptable to actual conditions, thereby enhancing system adaptability.
[0038] As a preferred technical means: the clustering processing module includes a membership determination unit, an objective function construction unit, a cluster center update unit, an optimal clustering result acquisition unit, and an optimal cluster number determination unit; the membership determination unit sets the number of clusters according to the improved fuzzy C-means clustering algorithm, selects the initial cluster centers based on the kernel density estimation method, and calculates the membership between the monthly load data and the cluster centers; the optimal cluster number determination unit selects the optimal number of clusters according to the Xie & Ben index and determines the final clustering result; The correlation analysis module includes a data fuzzification processing unit, a standard mutual information calculation unit, and a correlation quantitative evaluation unit; the data fuzzification processing unit replaces the original external influencing factor values and monthly load values with cluster center values to obtain a fuzzified sequence; the standard mutual information calculation unit calculates the standard mutual information value between the fuzzified data to quantitatively describe the correlation between the external influencing factors and the monthly load; The monthly load forecasting module includes a correlation weight calculation unit, a feature vector weighting unit, an improved bidirectional LSTM modeling unit, and a network optimization and output unit; the improved bidirectional LSTM modeling unit constructs a bidirectional LSTM neural network, adopts the Transformer multi-head attention mechanism to analyze the impact of external features, and connects the Transformer and bidirectional LSTM models through a convolutional neural network.
[0039] The membership determination unit uses an improved fuzzy C-means clustering algorithm to effectively determine the number of clusters and calculate the membership between monthly load data and cluster centers, improving understanding of data patterns and clustering effectiveness. The cluster center update unit iteratively optimizes the objective function to ensure that the cluster centers fully reflect the data distribution, thereby obtaining more accurate clustering results. The data fuzzification unit fuzzifies external influencing factors and monthly load data based on the clustering results, enhancing the analysis of data correlations. The standard mutual information calculation unit quantifies the correlation between external influencing factors and monthly load, providing more refined data relationship analysis. The correlation weight calculation unit determines the weight of external influencing factors by normalizing the standard mutual information value, helping to accurately identify the key factors influencing monthly load fluctuations. The improved bidirectional LSTM modeling unit utilizes a bidirectional long short-term memory network and the Transformer multi-head attention mechanism, combining the comprehensive feature vector with the time series feature vectors of historical monthly load data. This effectively captures the complex patterns and trends in time series data, improving the accuracy and robustness of monthly load forecasts. The network optimization and output unit, combined with a bidirectional LSTM network and the Transformer multi-head attention mechanism, generates robust and high-quality monthly load forecasts, meeting the industry's medium- and long-term planning and decision-making needs. The overall technical solution, which integrates cluster analysis, correlation analysis, and deep learning methods, is capable of processing complex data and addressing practical forecasting challenges, providing the industry with reliable forecasting tools and decision-making support.
[0040] The improved bidirectional LSTM modeling unit realizes bidirectional processing of sequence data through the combination of forward long short-term memory network and backward long short-term memory network; in each monthly load forecasting process, the transmission of information is controlled by forget gate, input gate and output gate; the Transformer multi-head attention mechanism establishes a direct correlation between any two time steps in the sequence through global recognition of enhanced comprehensive feature vectors and external feature vectors, effectively capturing long-distance dependency features and the impact of sudden events; the convolutional neural network layer splices the output features of the bidirectional LSTM and Transformer in spatial dimensions, extracts local pattern features through multi-scale convolution kernels, and realizes deep fusion of heterogeneous features; the fully connected layer maps the fused high-dimensional feature vector to the target output space to complete the final load forecast value calculation; this hybrid architecture gives full play to the technical advantages of each component. The bidirectional LSTM ensures the continuity of time series modeling, the Transformer enhances the global correlation perception capability, the convolutional neural network realizes feature fusion, and the fully connected layer completes nonlinear mapping, forming an end-to-end high-precision monthly load forecasting solution.
[0041] The third technical solution of the present invention is to provide a computer device, which includes one or more processors and one or more memories, wherein at least one program code is stored in the one or more memories, and when the program code is executed by the one or more processors, the aforementioned industry load forecasting method that integrates standard mutual information and improved bidirectional LSTM is implemented.
[0042] The fourth technical solution of the present invention is to provide a storage medium, in which at least one program code is stored. When the program code is executed by a processor, the steps of the aforementioned industry load forecasting method that integrates standard mutual information and improved bidirectional LSTM are implemented.
[0043] Beneficial effects: The present invention effectively integrates the industry's monthly load series and multiple external influencing factors such as consumption level, temperature, and vacation to achieve accurate prediction of the industry's monthly load. Specifically, the first step is to cluster the load and external factor information using the improved fuzzy C-means clustering algorithm taking into account kernel density estimation, which can capture the similar patterns and internal structures hidden in the data in a targeted manner, laying a data foundation for subsequent correlation analysis; the second step is to calculate the standard mutual information based on the clustering results, which can quantitatively characterize the nonlinear correlation between external influencing factors and the industry's monthly load, breaking through the limitations of traditional linear correlation analysis methods and more in line with the complex interaction mechanism between multiple factors and loads in actual scenarios; the third step is to assign weights based on the correlation analysis results, capture the load time series change law through a bidirectional LSTM neural network, and at the same time introduce the Transformer multi-head attention mechanism to analyze the impact of external features and realize feature fusion through a convolutional neural network, effectively integrating the time series features and external factor features, overcoming the problem of traditional methods over-relying on the time series continuity of load historical data and insufficient modeling of external dynamic influences, and ultimately improving the accuracy of the industry's monthly load forecast and its adaptability to complex operating environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is the basic model architecture of the present invention.
[0045] Figure 2 It is the LSTM network unit of the present invention.
[0046] Figure 3 It is the bidirectional LSTM network structure of the present invention.
[0047] Figure 4 These are the monthly load forecast results for the steel industry using different methods. DETAILED DESCRIPTION
[0048] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings.
[0049] Example 1: This paper proposes a method for monthly industry load forecasting that integrates kernel density estimation, standard mutual information, and improved bidirectional LSTM. The implementation process includes the following detailed steps: Step 1: Use the improved fuzzy C-means clustering algorithm with kernel density estimation to cluster the load series and information including consumption level, temperature, and vacation; Step 2: Based on the clustering results, standard mutual information is calculated to quantitatively analyze the correlation between information factors such as consumption level, temperature, and vacation time and the monthly load of the industry; Step 3: Assign weights based on the correlation analysis results, build a bidirectional LSTM neural network to capture the temporal variation pattern of industry load, use the Transformer multi-head attention mechanism to analyze the impact of external features, and output the monthly industry load forecast results through convolutional neural network connection.
[0050] The basic model architecture of this method is as follows Figure 1 shown.
[0051] Step 1: Based on the improved fuzzy C-means clustering algorithm, the monthly load series is clustered with factors such as consumption level, temperature, and vacation information. The details are as follows: Before analyzing the correlation between external factors and monthly industry load, the data was clustered to improve the accuracy of the correlation analysis. The improved fuzzy C-means (DFCM) clustering algorithm and the Xie & Ben index were used to cluster the historical monthly load data and external factor data, respectively. The clustering process based on the DFCM algorithm is as follows: Step 1.1: Determine the membership matrix: Set the number of clusters , N m Based on the kernel density estimation method, the electricity consumption data matrix Y Before the selection N l The data with the largest row density is used as the initial cluster center matrix V =[ V 1,…, V l ,…, V Nl ] T , the density calculation method is:
[0052] Where, Representative History Month m The density around the monthly load data; Y m andY j History Month m and j The corresponding monthly load data vector includes the load value and external influencing factors (such as temperature, consumption level, etc.);|| Y m - Y j ||Represents Y m and Y j The Euclidean distance between d c is the distance threshold parameter, specifically the bandwidth parameter of the kernel density estimation, which is adaptively selected according to the data distribution or determined by cross-validation.
[0053] Y m Belong to l The membership degree of a cluster center is expressed as:
[0054] Where, Q m,l Representative History Month m The monthly load data belongs to the cluster center l The membership degree is 0≤ Q m,l ≤1; q is the membership factor, Specifically, it is a fuzzy parameter that controls the fuzziness of the membership degree, usually q>1 To enhance the fuzziness of clustering results; V l is the cluster center l Corresponding monthly load data;|| Y m - V l ||Represents Y m and V l The Euclidean distance between .
[0055] Step 1.2: Construct the objective function G :
[0056] Where, The membership factor is q The historical month calculated in the case of m Belong to the cluster center l The degree of membership.
[0057] Step 1.3: Update the cluster center matrix V :If the objective function obtained in step 2 G If the minimum value is not reached, the cluster center matrix is calculated according to the following formula: V Update and recalculate the membership matrix until the objective function J Reach minimum.
[0058]
[0059] Step 1.4: Get the number of clusters N l The optimal clustering result under: from the membership matrix and cluster center matrix V The number of clusters obtained is N l The optimal clustering result corresponding to the minimum value of the objective function when .
[0060] Step 1.5: Determine the optimal number of clusters N l *: The Xie&Ben index is selected to evaluate the quality of clustering results under different cluster numbers. The index can be expressed as:
[0061] In the formula, the numerator reflects the degree of compactness within the clustering result, while the denominator reflects the degree of separation between the clustering results; I XB The smaller the value, the better the clustering effect. Therefore, by comparing the clustering quality indicators corresponding to different cluster numbers I XB , the optimal number of clusters can be screened out N l *, get the clustering result under the optimal number of clusters.
[0062] Furthermore, in step 2, standard mutual information is calculated based on the clustering results to quantitatively analyze the correlation between factors such as consumption level, temperature, and vacation information and the monthly load of the industry, as follows: Standardized mutual information (SMI) was used to analyze the correlation between monthly industry load and external influencing factors, fully accounting for their nonlinear relationship. Directly using SMI from raw data can affect correlation analysis due to subtle numerical differences, failing to capture the key trends in the variables. Therefore, based on SMI, the clustering results obtained using the DFCM algorithm were fuzzified to account for subtle numerical differences, improving the accuracy of the results.
[0063] For the monthly load and external influencing factors of the industry to be analyzed, it is known that within a period of time N m The influencing factor value of each sampling point x( m ) and monthly load values y ( m ), m ∈{1,2,…, N m},all x ( m ) constitutes a sequence X ,all y ( m ) constitutes a sequence Y The data after FCM clustering is used to replace the original data and calculate the standard mutual information. x ( m The cluster center value of the cluster where u ( m ),all u ( m ) constitutes a sequence U , sampling value y ( m The cluster center value of the cluster where v ( m ),all v ( m ) constitutes a sequence V . u ( m )Total F Possible values ( F ≤ N m ), denoted as u 1( f ), f ∈{1,2,…, F}, v ( m )Total G Possible values ( G ≤ N m ), denoted as v 1( g ), g ∈{1,2,…, G}, then the standard mutual information value of the influencing factors and monthly load is J ( U ; V ) is calculated as follows:
[0064] Where,
[0065]
[0066]
[0067] Where, I ( U ; V )for u ( m )and v ( m )’s mutual information value; H ( U )and H ( V ) are respectively U and V Information entropy; P ( u 1( f ), v 1( g )) satisfies both u ( m )= u 1( f )and v ( m )= v 1( g ) to all sampling points; P ( u 1( f ))To satisfy u ( m )= u 1( f ) to all sampling points; P ( v 1( g ))To satisfy v ( m )= v 1( g )’s sampling time to all sampling times; J ( U ; V ) ranges from [0,1]. The larger the value, the stronger the correlation between the influencing factor and the monthly load.
[0068] It should be noted that the correlation analysis results obtained based on standard mutual information will be input into the monthly load forecasting network as network hyperparameters.
[0069] Furthermore, in step 3, the external feature data and the comprehensive feature vector are input into the improved bidirectional LSTM neural network to perform monthly industry load forecasting, as follows: Different external factors have varying degrees of influence on the monthly load of each industry. Therefore, when conducting monthly load forecasts for each industry, the correlation between the monthly load of the industry and each external influencing factor should be fully considered. To this end, the correlation results obtained by the improved correlation analysis based on standard mutual information are used as hyperparameters of the monthly load forecasting network and integrated into the monthly load forecasting network structure. The specific approach is as follows: Combine the monthly industry load with the r External factors ( r ∈{1,2,.., R}, R is the number of external influencing factors) is recorded as J r , normalize the standard mutual information value and get the monthly load of the industry and the r The normalized standard mutual information value of external influencing factors J 0,r ,Right now:
[0070] Then, J 0,r For the r The characteristic vector of the external influencing factors x r The weight of all external influencing factors is weighted and summed to obtain the enhanced comprehensive feature vector x o :
[0071] Finally, the basic external feature vector, the enhanced comprehensive feature vector and the historical monthly load feature vector are all used as inputs of the improved bidirectional LSTM monthly load forecasting network.
[0072] The improved bidirectional LSTM monthly load forecasting network consists of a bidirectional LSTM module, a Transformer multi-head attention module, and a convolutional neural network module. The modules are described as follows: Bidirectional LSTM is a forward-backward combination of LSTMs. It combines forward LSTMs with backward LSTMs. It incorporates both past and future information and possesses the memory capabilities of long short-term memory networks (LSTMs). It combines the characteristics of both bidirectional recurrent neural networks and LSTMs.
[0073] The single LSTM unit in the bidirectional LSTM network transmits information through three basic "gate" structures, namely the forget gate, input gate and output gate, as shown in Figure 2 shown.
[0074] 1) Forget Gate The forget gate in LSTM is used to determine the degree of feature forgetting of the state memory unit. The forget gate takes the monthly load forecast value from the previous moment into account. h t-1 and the current input features x t At the same time, the sigmoid function is passed in and the sigmoid function outputs As the forget weight of the state memory unit. f t It can be expressed as:
[0075] Where, W IF Output of the forget gate at the moment h t-1 The weight matrix of W HF Current input feature x t The weight matrix of b IF and b HF is the forget gate bias term; h t-1 It represents the monthly load forecast output at the previous moment; x t Represents the monthly load forecast characteristic factors currently input, including external characteristics such as temperature.
[0076] 2) Input Gate The function of the input gate is to generate the intermediate value of the state memory unit. First, the input gate converts the predicted output value of the previous moment into h t-1 and the newly entered x t Pass in the sigmoid function together and output i t As input node state g t Update weight of . Input node state g t Depend on h t-1 and x t Joint decision.
[0077] The sigmoid function output of the input gate i t The expression is as follows:
[0078] Where, WII and W HI The sigmoid function is passed into the input gate h t-1 and input features x t The multiplied weight matrix; b HI and b II is the bias term of the input gate sigmoid function.
[0079] Tanh function output of the input gate g t The expression is as follows:
[0080] Where, W IG and W HG The input gate is passed into the tanh function h t-1 and input features x t The multiplied weight matrix; b IG and b HG is the bias term of the input gate tanh function.
[0081] 3) Update state memory unit The third step of the LSTM workflow is to update the state memory unit, that is, to update the C t First, the state of the state memory unit at the previous moment is combined with the output of the forget gate f t Multiply and output the selective memory result. Then add this value to the intermediate value of the state memory unit point by point to update the state memory unit together. C t , C t The expression is as follows:
[0082] 4) Output gate The function of the output gate is to determine the monthly load forecast value at the current moment h t The output gate first converts the monthly load forecast value of the previous moment into h t-1 and input features x t Pass in the sigmoid function to determine the weight o t. Secondly, the output gate updates the state memory unit C t Passed to the tanh function, the output is combined with the weight o t Multiply them to get the monthly load forecast value at the current moment h t Among them, the output weight o t The expression is as follows:
[0083] Where, W IO and W HO for h t-1 and input features x t The multiplied weight matrix; b IO and b HO is the output gate bias term.
[0084] Monthly load forecast value at the current moment h t The expression is as follows:
[0085] Where, C t is the updated state memory unit.
[0086] Bidirectional LSTM uses two independent hidden layers to process the sequence data forward and backward based on LSTM, connects the two connection layers to the same output layer, and uses the previous information and the next information as the current time basis of the time series data, such as Figure 3 As shown in the figure, a bidirectional LSTM network is used for prediction, which has good expressive power for continuous time series. The reuse of weight parameters makes it less demanding on data. The relevant calculation formula is as follows:
[0087]
[0088]
[0089] Where, Represents the forward calculation value; Represents the backward calculated value; o t Represents the final output value calculated by the forward and backward values; ω1 is the weight matrix mapping the input layer to the forward layer; ω 3 is the weight matrix mapping the input layer to the backward layer; ω 2 is the weight matrix mapping the output of the forward layer at the previous moment to the current calculation moment; ω 5 is the weight matrix that maps the output of the backward layer at the previous moment to the current calculation moment; ω 4 is the weight matrix mapping the output of the forward layer to the output layer; ω 6 is the weight matrix that maps the output of the backward layer to the output layer.
[0090] Because the bidirectional LSTM model cannot utilize future weather forecast data and only has the capability of analyzing sequential trends, the Transformer multi-head attention module is used to capture the influence of multidimensional external features. By computing multiple sets of attention weights in parallel, the input features are mapped into different subspaces. Dynamic correlations between data in each subspace are captured separately, effectively identifying the heterogeneous impact of key external events such as sudden temperature changes and holiday effects on load. This multi-head attention mechanism effectively complements the shortcomings of the bidirectional LSTM. Its single-head attention is calculated using the following formula:
[0091] Where, Q 、 K 、 V is the query key-value matrix; is the scaling factor.
[0092] The multi-head attention mechanism divides the input feature space into multiple independent sub-representation spaces. Each attention head performs attention calculations in parallel in its corresponding subspace, thereby achieving parallel learning and representation of the multi-dimensional features of the input sequence. This mechanism enables the model to simultaneously capture multiple potential correlation patterns and global information of the data. Its specific expression is:
[0093]
[0094]
[0095] Where, Q i 、 K i 、 V i For the i The query key-value matrix of the attention heads; 、 、 For the i The weight matrix of the attention head; is the output projection matrix;h is the number of attention heads.
[0096] Convolutional neural networks serve as a bridge between bidirectional LSTM and Transformer models. They are primarily composed of convolutional layers, pooling layers, and fully connected layers. Convolutional layers convolve input data using convolution kernels (filters) to extract local features. Pooling layers downsample feature maps to reduce the number of parameters while retaining key information. Fully connected layers map the extracted features to the final output space.
[0097] The convolution operation is the core of the convolutional neural network. It performs bit-by-bit product sum operations by sliding the convolution kernel over the input data. The weight sharing mechanism allows the same convolution kernel to share the same weight parameters across the entire input region, significantly reducing the number of model parameters. The design of the local receptive field ensures that each neuron is connected only to a local region of the input, effectively capturing the spatial locality of the data. The specific output feature calculation formula is:
[0098] Where, is the output of the convolutional layer; is the activation function; is the convolution kernel weight; For input data; For bias.
[0099] In order to verify the effectiveness and accuracy of the proposed industry monthly load forecasting method that integrates kernel density estimation, standard mutual information and bidirectional LSTM, the steel, shipbuilding and textile industries of a certain city are taken as examples.
[0100] The monthly load data of the industry adopts the monthly load data of the city from January 2016 to December 2020. The corresponding external influencing factor data include residents' consumption data, meteorological data, and holiday data. First, the improved clustering algorithm and standard mutual information are used to analyze the correlation between each external influencing factor and the monthly load of each industry; then, the above data are divided into training set and test set in a ratio of 5:1 and input into the constructed monthly load forecasting network to verify the effectiveness of the proposed algorithm. Among them, the training set is used to train the model, and the test set is used to evaluate the prediction effect of the model. The network parameters of the bidirectional LSTM monthly load forecasting model constructed by the present invention are selected as follows: the time step is 1 month, the maximum number of iterations is 1000, the initialization learning rate is 0.005, and the number of hidden layer units of the forward LSTM and the backward LSTM are both 128.
[0101] Based on standard mutual information, the correlations between the steel, shipbuilding, and textile industries and various external factors were analyzed. These external factors included household consumption levels, monthly average maximum and minimum temperatures, and the number of monthly vacation days. The correlation analysis results between each external factor and the monthly load of each industry, obtained using standard mutual information, are shown in Table 1.
[0102] Table 1 Correlation analysis results based on standard mutual information
[0103] Table 1 shows that among the external factors, temperature has the greatest impact on the monthly load of various industries, while consumption levels and the number of monthly holidays have relatively little impact. This is primarily due to the fact that temperature affects both heating and cooling air conditioning loads in industry. When the monthly average maximum temperature is high, cooling air conditioning loads are higher, while when the monthly average temperature is low, heating air conditioning loads are higher. Nevertheless, the relative strength of the correlation between load and monthly average temperature varies across different industries, with the shipbuilding and textile industries showing a significantly higher correlation than the steel industry. This is primarily due to the fact that, due to the specific nature of the steel industry, air conditioning demand is significantly lower than in the shipbuilding and textile industries. Furthermore, it is not difficult to see that the correlation between the number of monthly holidays and the textile industry is significantly higher than in the other two industries. This is because the shipbuilding and steel industries are heavy industries, while the textile industry is light industry, which is generally unaffected by holidays. Consumption levels are national economic data, representing macroeconomic indicators, and therefore have a lower correlation with monthly industry load data.
[0104] Comparing the correlation analysis results based on standard mutual information before and after clustering reveals that, before clustering, the data in the table show a low correlation between consumption levels and the monthly load of the three industries (steel, shipping, and textiles). After clustering, the correlation between consumption levels and monthly load further decreased, indicating that clustering effectively mitigated the influence of non-critical factors. However, for temperature, a factor with a strong correlation with monthly industry load, both the monthly average maximum and minimum temperatures showed a significant increase in correlation after clustering. For example, the monthly average maximum temperature for the shipping industry increased from 0.52 to 0.58, and for the textile industry from 0.49 to 0.59. This trend indicates that clustering effectively reduces the influence of outliers in the data, reducing the impact of random errors and making important features more prominent in the data.
[0105] The correlation analysis results between each external influencing factor and the load of each industry were obtained. On this basis, the external feature data was weighted and input into the bidirectional LSTM prediction network as external features together with the load features to obtain the prediction results of the load of each industry. In order to verify the performance of the proposed bidirectional LSTM neural network prediction model taking into account external influencing factors, the prediction results were compared with the results of the load prediction model based on GRU, CNN, LSTM, bidirectional LSTM, exponential smoothing and extreme gradient boosting decision tree (XGBoost). Among them, the prediction error comparison results of the ship industry load are shown in Table 2, and the prediction result curve is shown in Figure 4 shown.
[0106] Table 2 Prediction error results under different methods
[0107] From Table 2 and Figure 4 It can be seen that the LSTM prediction method based on the traditional deep learning model and the exponential smoothing method based on the traditional time series prediction model have poor prediction effects. The main reason is that they only consider the load time series data and ignore the impact of external factors on the industry load. The performance of GRU and XGBoost is at an intermediate level, inferior to CNN and the method proposed in this invention. This is because the time span of the industry monthly load forecast is long, GRU has more parameters, and it is easy to overfit when there are relatively few monthly data samples. XGBoost is essentially a tree-based integration method with limited ability to model the continuity and temporal dependency of time series. The time span of the industry monthly load forecast is long, but the key features may be concentrated in the recent months. The local receptive field of CNN is suitable for capturing this short-term and medium-term correlation. However, the monthly load is also optimized with the same month in history. The global vision of the method proposed in this invention can effectively capture this long-term and medium-term correlation, and combined with the temporal correlation that the bidirectional LSTM is good at learning, it achieves the lowest prediction error.
[0108] Example 2: This embodiment provides an industry monthly load forecasting system that integrates kernel density estimation, standard mutual information, and bidirectional LSTM. The system includes: Data receiving module: used to receive historical industry monthly load data and external influencing factor data; Clustering processing module: clustering the received data using the improved fuzzy C-means clustering algorithm based on kernel density estimation; Correlation analysis module: performs standard mutual information calculation based on clustering results to quantitatively analyze the correlation between external influencing factors and the industry's monthly load; Monthly load forecasting module: uses a bidirectional LSTM neural network and combines the results of correlation analysis to predict the industry's monthly load.
[0109] The data receiving module effectively integrates historical monthly industry load data and data on external influencing factors, providing comprehensive data support for subsequent analysis and forecasting. Data clustering using the fuzzy C-means clustering algorithm can better identify patterns and structures within the data, improving data understanding. Standardized mutual information is used to quantify the relationship between external influencing factors and monthly industry load, revealing key factors influencing monthly load fluctuations and enhancing forecast accuracy. A bidirectional LSTM (bidirectional long short-term memory) network captures temporal dependencies in the data and incorporates bidirectional information from historical data, thereby improving the accuracy and reliability of monthly load forecasts. The system integrates cluster analysis, correlation analysis, and deep learning methods to fully leverage the information in the data and provide a more comprehensive and accurate monthly load forecast. The system can adjust to changes in monthly industry load and external influencing factors, making forecast results more realistic and adaptable to actual conditions, thereby enhancing system adaptability.
[0110] Membership determination unit: used to set the number of clusters according to the improved fuzzy C-means clustering algorithm, select the initial cluster center based on the kernel density estimation method, and calculate the membership between the monthly load data and the cluster center; Objective function construction unit: used to set the objective function to minimize the Euclidean distance between the monthly load data and the cluster center; Cluster center update unit: updates the cluster center matrix by iteratively optimizing the objective function until the objective function converges; Optimal clustering result acquisition unit: used to obtain clustering results according to the minimized objective function value; Optimal cluster number determination unit: selects the optimal number of clusters through Xie&Ben index and determines the final clustering result; The correlation analysis module further includes: Data fuzzification processing unit: used to replace the original external influencing factor values and monthly load values with the cluster center values corresponding to each sampling point based on the clustering results of the clustering processing module to obtain the fuzzy processed sequence; Standard mutual information calculation unit: used to calculate the standard mutual information value between the fuzzy processed data to quantitatively describe the correlation between external influencing factors and monthly load; Correlation quantitative evaluation unit: used to evaluate the correlation between external influencing factors and the industry's monthly load based on the size of the standard mutual information value, and input the results into the monthly load forecast module; The monthly load forecasting module further includes: Correlation weight calculation unit: used to normalize the standard mutual information value between the industry monthly load and each external influencing factor to determine the characteristic weight of the external influencing factor; Eigenvector weighting unit: used to perform weighted summation of the eigenvectors of all external influencing factors according to the calculated eigenweights to obtain a comprehensive eigenvector; Improved bidirectional LSTM modeling unit: A bidirectional LSTM neural network is constructed to capture the temporal variation of industry loads. The Transformer multi-head attention mechanism is used to analyze the impact of external features. The Transformer and bidirectional LSTM models are connected through a convolutional neural network. Network optimization and output unit: The fully connected layer adjusts the final output data shape to form the monthly load forecast result.
[0111] Through the detailed description of the industry monthly load forecasting method that integrates kernel density estimation, standard mutual information and bidirectional LSTM in the above specification, those skilled in the art can clearly understand the industry monthly load forecasting system that integrates kernel density estimation, standard mutual information and bidirectional LSTM in this embodiment. For the system disclosed in Example 2, since it corresponds to the method disclosed in Example 1 and has corresponding functional modules and beneficial effects, the relevant details can be referred to the method section.
[0112] Example 3 This embodiment provides an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a current limiting method for a grid-type flexible direct current system based on power matching as described in any embodiment of the present invention is implemented.
[0113] Example 4 This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the current limiting method for a grid-type flexible direct current system based on power matching as described in any embodiment of the present invention is implemented.
[0114] Those skilled in the art will appreciate that the various units and algorithm steps described in the disclosed embodiments can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0115] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0116] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), magnetic disk or optical disk, and other media that can store program codes.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A method for industry load forecasting that integrates standard mutual information and improved bidirectional LSTM, characterized in that: The following steps are involved: 1) An improved fuzzy C-means clustering algorithm taking into account kernel density estimation is used to cluster information including load series, consumption levels, temperature, and vacations; 2) Based on the clustering results, standard mutual information is calculated to quantitatively analyze the correlation between information factors such as consumption level, temperature, and vacation time and the monthly load of the industry; 3) Based on the correlation analysis results, weights are assigned and a bidirectional LSTM neural network is constructed to capture the temporal variation patterns of industry loads. The Transformer multi-head attention mechanism is used to analyze the impact of external features, and the monthly industry load forecast results are output through convolutional neural network connections.
2. The industry load forecasting method integrating standard mutual information and improved bidirectional LSTM according to claim 1 is characterized by: The clustering process of step 1) further includes: 1.1) Determine the membership matrix: Based on the improved fuzzy C-means clustering algorithm, set the number of clusters, select the initial cluster centers based on the kernel density estimation method, and calculate the membership between the monthly load data and the cluster centers; 1.2) Constructing the objective function: The objective function is set to minimize the Euclidean distance between the monthly load data and the cluster center; 1.3) Update cluster centers: Update the cluster center matrix by iteratively optimizing the objective function until the objective function converges; 1.4) Obtaining the optimal clustering result: Obtain the clustering result based on the minimized objective function value; 1.5) Determine the optimal number of clusters: Select the optimal number of clusters using the Xie & Ben index and determine the final clustering result.
3. The industry load forecasting method integrating standard mutual information and improved bidirectional LSTM according to claim 2 is characterized by: Step 1.1) Determine the membership matrix and set the number of clusters based on the number of months , using the kernel density estimation method from the electricity consumption data matrix Y Before the selection N l The data with the largest row density is used as the initial cluster center matrix V =[ V 1,…, V l ,…, V Nl ] T , and then calculate any monthly data Y m Belong to l The membership degree of each cluster center. Step 1.2) Construct the objective function, the objective function G The sum of the squares of the weighted distances of membership between all historical months and all cluster centers; Step 1.3) Update the cluster center, if the objective function G If the minimum value is not reached, the cluster center matrix is updated according to the given update strategy. V Update and recalculate the membership matrix until the objective function G Reach minimum; Step 1.4) Obtain the optimal clustering result from the membership matrix and cluster center matrix V The number of clusters obtained is N l The optimal clustering result corresponding to the minimum value of the objective function when ; Step 1.5) In determining the optimal number of clusters, the Xie & Ben index is used to reflect the degree of compactness within the cluster and the degree of separation between clusters. The clustering quality index corresponding to different numbers of clusters is compared. I XB , filter out the best number of clusters , and obtain the best clustering result.
4. The industry load forecasting method integrating standard mutual information and improved bidirectional LSTM according to claim 3 is characterized by: Step 2) further comprises the steps of: 2.1) Data Fuzzification Based on the clustering results obtained by the improved fuzzy C-means (DFCM) clustering algorithm in step 1), the original external influencing factor value x(m) and monthly load value y(m) are replaced by the cluster center value corresponding to each sampling point to obtain the fuzzy processed sequences U and V; 2.2) Standard mutual information calculation 2.2.1) Mutual Information Calculation: The mutual information I(U;V) between the fuzzified data U and V is calculated to quantitatively describe the correlation between the influencing factors and the monthly load. 2.2.2) Information entropy calculation: Calculate the information entropy H(U) and H(V) of sequences U and V; 2.2.3) Standard mutual information calculation: Calculate the standard mutual information value J(U;V) between the influencing factors and the monthly load by using the ratio of mutual information I(U;V) to information entropy H(U) and H(V); 2.3) Quantitative evaluation of correlation 2.3.1) Results Analysis: Based on the standard mutual information value J(U;V), the correlation between information factors such as consumption level, temperature, and vacation time and the industry's monthly load was quantitatively evaluated. A larger value indicates a stronger correlation. 2.3.2) Parameter input: The correlation analysis results based on standard mutual information are used as network hyperparameters and input into the monthly load forecasting network to optimize the model's forecasting ability.
5. The industry load forecasting method integrating standard mutual information and improved bidirectional LSTM according to claim 4 is characterized by: In step 2.1) data fuzzification processing, for the monthly load of the industry to be analyzed and the external influencing factors, it is known that within a period of time N m The influencing factor value of each sampling point x ( m ) and monthly load values y ( m ), m ∈{1,2,…, N m },all x ( m ) constitutes a sequence X ,all y ( m ) constitutes a sequence Y ; Use the data after DFCM clustering to replace the original data, calculate the standard mutual information, and record the sampling value x ( m The cluster center value of the cluster where u ( m ),all u ( m ) constitutes a sequence U , sampling value y ( m The cluster center value of the cluster where v ( m ),all v ( m ) constitutes a sequence V ; u ( m )Total F Possible values, F ≤ N m , recorded as , f ∈{1,2,…, F }, v ( m )Total G Possible values, G ≤ N m , recorded as , ∈{1,2,…, G }; In step 2.2) of the standard mutual information calculation, the standard mutual information value between the influencing factors and the monthly load is J ( U ; V ) is calculated as follows: Where, Where, I ( U ; V )for u ( m )and v ( m )’s mutual information value; H ( U )and H ( V ) are respectively U and V Information entropy of To satisfy both u ( m )= and v ( m )= The proportion of sampling points to all sampling points; To satisfy u ( m )= The proportion of sampling points to all sampling points; To satisfy v ( m )= The proportion of sampling moments to all sampling moments; J ( U ; V ) ranges from [0,1]. The larger the value, the stronger the correlation between the influencing factor and the monthly load of the industry.
6. The industry load forecasting method integrating standard mutual information and improved bidirectional LSTM according to claim 5 is characterized by: The monthly load forecast in step 3) includes the following steps: 3.1) Correlation Weight Calculation: Normalize the standard mutual information value between the industry's monthly load and each external influencing factor to determine the characteristic weights of information factors including consumption level, temperature, and vacation time; 3.2) Eigenvector Weighting: Based on the calculated eigenweights, the eigenvectors of all external influencing factors are weighted and summed to obtain the enhanced comprehensive eigenvector. 3.3) Bidirectional LSTM Network Modeling: Using a bidirectional long short-term memory network, the time series feature vector of historical monthly load is used as input to capture the time series trend of monthly load in the industry. 3.4) Modeling the Transformer Multi-Head Attention Mechanism: Applying the Transformer multi-head attention mechanism, we take multi-dimensional external feature vectors and integrated feature vectors as input and capture dynamic correlations across feature dimensions through parallel attention heads. 3.5) Network Optimization and Output: The time series feature vectors extracted by the bidirectional LSTM and the multi-head attention features encoded by the Transformer are combined through a convolutional neural network to construct a feature interaction layer. Finally, the fully connected layer outputs the environmentally adaptable monthly industry load forecast value.
7. The industry load forecasting method integrating standard mutual information and improved bidirectional LSTM according to claim 6 is characterized by: The bidirectional LSTM neural network model captures the long-term and short-term dependencies of the load sequence based on the bidirectional time series feature extraction module. It controls the transmission of information through the forget gate, input gate, and output gate. The formulas for calculating the forward and backward monthly load forecast outputs are as follows: Where, Represents the forward calculation value; Represents the backward calculated value; o t Represents the final output value calculated by the forward and backward values; ω 1 is the weight matrix mapping the input layer to the forward layer; ω 3 is the weight matrix mapping the input layer to the backward layer; ω 2 is the weight matrix mapping the output of the forward layer at the previous moment to the current calculation moment; ω 5 is the weight matrix that maps the output of the backward layer at the previous moment to the current calculation moment; ω 4 is the weight matrix mapping the output of the forward layer to the output layer; ω 6 is the weight matrix mapping the output of the backward layer to the output layer; The Transformer multi-head attention mechanism maps input features to different subspaces by calculating multiple sets of attention weights in parallel. It captures the dynamic correlation between data in each subspace, effectively identifying the heterogeneous impact of key external events on load, such as sudden temperature changes and holiday effects. Single-head attention is calculated using the following formula: Where, Q 、 K 、 V is the query key-value matrix; is the scaling factor; The model is divided into multiple attention heads to form multiple subspaces, allowing the model to focus on different aspects of information. The specific expression is: Where, Q i 、 K i 、 V i For the i The query key-value matrix of the attention heads; 、 、 For the i The weight matrix of the attention head; is the output projection matrix; h is the number of attention heads.
8. An industry load forecasting system integrating standard mutual information and improved bidirectional LSTM, characterized by: The system is used to perform the method according to any one of claims 1 to 7, and the system comprises: Data receiving module: used to receive historical industry monthly load data and external influencing factor data; Clustering processing module: clustering the received data using the improved fuzzy C-means clustering algorithm based on kernel density estimation; Correlation analysis module: performs standard mutual information calculation based on clustering results to quantitatively analyze the correlation between external influencing factors and the industry's monthly load; Monthly load forecasting module: uses a bidirectional LSTM neural network and combines the results of correlation analysis to predict the industry's monthly load.
9. The industry load forecasting system integrating standard mutual information and improved bidirectional LSTM according to claim 8, characterized in that: The cluster processing module includes a membership determination unit, an objective function construction unit, a cluster center update unit, an optimal clustering result acquisition unit, and an optimal cluster number determination unit; the membership determination unit sets the number of clusters according to the improved fuzzy C-means clustering algorithm, selects the initial cluster center based on the kernel density estimation method, and calculates the membership between the monthly load data and the cluster center; The optimal cluster number determination unit selects the optimal cluster number through the Xie&Ben index and determines the final clustering result; The correlation analysis module includes a data fuzzification processing unit, a standard mutual information calculation unit, and a correlation quantitative evaluation unit; the data fuzzification processing unit replaces the original external influencing factor values and monthly load values with cluster center values to obtain a fuzzified sequence; the standard mutual information calculation unit calculates the standard mutual information value between the fuzzified data to quantitatively describe the correlation between the external influencing factors and the monthly load; The monthly load forecasting module includes a correlation weight calculation unit, a feature vector weighting unit, an improved bidirectional LSTM modeling unit, and a network optimization and output unit; the improved bidirectional LSTM modeling unit constructs a bidirectional LSTM neural network, adopts the Transformer multi-head attention mechanism to analyze the impact of external features, and connects the Transformer and bidirectional LSTM models through a convolutional neural network.
10. A computer device, characterized in that: The device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories. When the program code is executed by the one or more processors, an industry load forecasting method that integrates standard mutual information and improved bidirectional LSTM as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-dimensional short-term power load prediction method based on improved K-means clustering CCA-BiLSTM
CN114358185A
Transform-based multi-factor load prediction method
CN119151040A
Energy prediction device, energy prediction method, and energy prediction program
US20240195174A1
Cited By
Load prediction method and system based on virtual power plant operation
CN120893637A
Power business data prediction method and device, terminal equipment and storage medium
CN121211028A
Multi-dimensional data processing method and system corresponding to power load prediction
CN121233943A
Electric quantity prediction method and device based on production condition, equipment, storage medium and program product
CN121724224A