A method for predicting distribution transformer load based on multi-scale wave mode adaptive decomposition
By using an improved multi-scale fluctuation mode adaptive decomposition method, the accuracy and adaptability issues in distribution transformer load forecasting are solved, achieving efficient load forecasting and power grid safety monitoring.
Patent Information
- Application Number
- CN202411809293.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-12-10
AI Technical Summary
Existing technologies for load forecasting of distribution transformers suffer from insufficient forecasting accuracy, poor adaptability, mode aliasing, and endpoint effects, and are not effective enough in handling abnormal fluctuations.
An improved maximum overlap discrete wavelet packet transform is used for multi-scale decomposition. Combined with adaptive density estimation, multi-level clustering, and improved dynamic time warping, an efficient index structure is constructed for prediction through adaptive matching and anomaly detection optimization.
It improves the accuracy and adaptability of load forecasting, reduces forecasting errors, enhances the robustness and stability of the system, enables real-time monitoring and early warning, and ensures the safe operation of the power grid.
Smart Images

Figure CN119740016B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system load forecasting and grid safe operation technology, specifically to a distribution transformer load forecasting method based on multi-scale fluctuation mode adaptive decomposition. Background Technology
[0002] In the field of power system load forecasting, related technologies have developed rapidly, and traditional forecasting methods such as time series analysis and regression analysis have been widely used. However, with the increasing complexity of distribution network structures and the diversification of load characteristics, these traditional methods have shown certain limitations in terms of forecasting accuracy and adaptability. In recent years, data-driven forecasting technologies, especially those combining signal processing and machine learning, have gradually become a research hotspot, providing a new perspective for load forecasting.
[0003] Despite advancements in some areas, existing technologies still fall short in terms of accuracy and adaptability in distribution transformer load forecasting. First, traditional methods often struggle to handle the nonlinear characteristics and fluctuations in load data, leading to inaccurate forecasts. Second, existing technologies have room for improvement in feature extraction and characterization of load data, particularly when dealing with multi-timescale fluctuation patterns. Furthermore, outliers and noise significantly impact forecast results, and existing technologies have limited means to address these issues.
[0004] Currently, the main technical shortcomings of distribution transformer load forecasting are as follows:
[0005] Traditional time series forecasting methods (such as ARIMA and SARIMA) cannot effectively handle the multi-scale characteristics of load data, thus limiting forecast accuracy.
[0006] While deep learning methods (such as LSTM and Transformer) have strong nonlinear fitting capabilities, they are poorly adapted to small sample scenarios and lack model interpretability.
[0007] Existing decomposition prediction methods (such as EMD, VMD, etc.) suffer from problems such as mode aliasing and endpoint effects, and lack the ability to adaptively identify fluctuation characteristics at different scales.
[0008] Existing methods generally neglect the handling of abnormal fluctuations in load data, resulting in poor stability of prediction results.
[0009] In summary, this invention has significant advantages in improving the accuracy, adaptability, and robustness of distribution transformer load forecasting, and is expected to provide strong support for the safe operation and optimized dispatch of power systems. Summary of the Invention
[0010] In view of the aforementioned existing problems, this invention aims to solve the following problems: how to achieve adaptive multi-scale decomposition of load data to avoid pattern aliasing and endpoint effects; how to construct a scientific fluctuation characteristic characterization system to accurately depict fluctuation patterns at different scales; how to establish a dynamic pattern matching mechanism to improve the adaptability of the prediction model to load changes; and how to effectively handle abnormal fluctuations in load data to improve the reliability of prediction results.
[0011] To address the aforementioned technical issues, a distribution transformer load forecasting method based on multi-scale fluctuation mode adaptive decomposition is proposed, including:
[0012] The system takes raw data as input and employs an improved maximum overlap discrete wavelet packet transform to decompose the load signal at multiple scales, optimize energy distribution, extract and optimize features, and construct a feature fluctuation library to store features. Through adaptive density estimation and an improved multi-level clustering strategy, it refines the fluctuation patterns at different time scales, constructs an efficient index structure, and performs prediction through adaptive matching. Improved dynamic time warping optimizes distance metrics and search strategies, and an improved particle swarm optimization algorithm dynamically adjusts weights to construct a combined prediction framework. Anomaly detection is optimized by combining ensemble isolated forests and local anomaly factors, and the online correction mechanism is optimized through adaptive windowing and robust estimation.
[0013] As a preferred embodiment of the distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition described in this invention, the improved maximum overlap discrete wavelet packet transform for multi-scale decomposition of the load signal includes: using the symmetric symmetric symmetric symmetric symmetric symmetric wavelet basis function as the basic decomposition tool, setting the support to 16, selecting the optimal wavelet basis by minimizing the reconstruction error, and the error function... for:
[0014]
[0015] Where x is the original signal, For reconstructing the signal;
[0016] Set the reconstruction error threshold to When the reconstruction error is less than At that time, it was considered that the wavelet basis selection was appropriate;
[0017] At the same time, the wavelet basis function must satisfy three key properties: the first key property is orthogonality, the second key property is compact support, and the third key property is vanishing moment.
[0018] An adaptive extension algorithm is used to address the endpoint effect problem in load data. An AR(p) autoregressive model is established to describe the temporal correlation of the load data. The model order p is dynamically determined by minimizing the AIC criterion. The extension length L is adaptively determined based on the data length N and the signal-to-noise ratio (SNR). Denoising is performed using an improved soft thresholding method. The threshold parameter... A scale-dependent adaptive approach is adopted, which uses the median absolute deviation method for robust estimation, optimizes the threshold selection using the SURE criterion, and determines the optimal threshold by minimizing the unbiased risk estimate.
[0019] The optimized energy distribution includes, for the wavelet decomposition coefficients of the j-th layer, calculating the wavelet energy proportion of the j-th layer, setting an adaptive energy threshold, and calculating the mutual information between adjacent decomposition layers through a constraint mechanism based on mutual information. When the mutual information value exceeds the threshold... When the information gain is 0.1, it indicates information redundancy, requiring hierarchical adjustment. An improved information gain calculation method dynamically determines the optimal hierarchy, and a distance function is defined using a hierarchy merging strategy based on dynamic programming. ,when Merging at that time The cross-validation method was used to determine that, among which, This represents the distance between adjacent floors.
[0020] As a preferred embodiment of the distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition described in this invention, the feature extraction and characterization optimization includes: performing deep mining and enhancement processing on distribution transformer load data to construct a feature fluctuation library;
[0021] The load data includes an amplitude feature family and a shape feature family. The amplitude feature family includes maximum amplitude, root mean square amplitude, peak factor, margin factor, and waveform factor. The shape feature family includes skewness, kurtosis, waveform index, and impulse index.
[0022] The frequency domain characteristics are enhanced by improving the power spectral density estimation method. The multi-window spectrum estimation technique is adopted, and the final power spectral density is obtained by weighted superposition of multiple independent spectrum estimation results.
[0023] The frequency band energy features extracted by energy distribution characterization include sub-band energy ratio, energy entropy, spectral centroid, and spectral broadening. The sub-band energy ratio includes calculating the proportion of energy in each sub-band to the total energy. The spectral centroid is obtained by weighted averaging of energy to obtain the center position of the energy distribution. The spectral broadening is obtained by calculating the second moment of the frequency relative to the centroid.
[0024] As a preferred embodiment of the distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition described in this invention, the method of adaptive density estimation and improved multi-level clustering strategy includes: characterizing data distribution features through improved kernel density estimation, selecting the Epanechnikov kernel function, selecting bandwidth using an adaptive method, and determining the density threshold based on the dynamic determination method of the density threshold using the silhouette coefficient.
[0025] After characterization, an improved multi-level clustering strategy is used to identify fluctuation patterns at different time scales. Based on the inherent laws of distribution network load changes, the time scale is divided into four levels: intraday, daily, weekly, and monthly.
[0026] The intraday layer covers fluctuations from 15 minutes to 4 hours, reflecting short-term electricity consumption behavior; the daily layer covers changes from 4 hours to 24 hours, reflecting the characteristics of the daily load curve; the weekly layer spans from 1 day to 7 days, depicting the differences between weekdays and weekends; and the monthly layer ranges from 7 days to 30 days, capturing seasonal variation characteristics.
[0027] Within each timescale layer, an improved density clustering algorithm is used for wave pattern identification, and a density-sensitive factor is introduced during the boundary point absorption process. The criteria for determining the ownership of boundary points are as follows:
[0028]
[0029] in, For point With point The distance between them and These are the corresponding density values. Let be the neighborhood radius, where ,when When the value is less than 0.35, a relatively dense cluster is formed; when... When the value is greater than 0.35, points with large density differences are allowed to be grouped into the same category;
[0030] The refined characterization of the fluctuation pattern includes probabilistic distribution modeling using an improved Gaussian mixture model and parameter estimation using an optimized expectation-maximization algorithm.
[0031]
[0032] in, For conditional expectation, For model parameters, The parameters for the current iteration. Let be the likelihood function containing latent variables. For expectation operators;
[0033] In each iteration, the parameters are dynamically adjusted and the step size is updated based on the gradient change of the likelihood function through an adaptive learning rate adjustment mechanism. Simultaneously, the optimal number of components K is adaptively determined using the Bayesian information criterion.
[0034]
[0035]
[0036] Among them, BIC stands for Bayesian Information Criterion. The log-likelihood value is M, where M is the number of samples.
[0037] Meanwhile, the nonlinear correlation characteristics in the load data of power distribution equipment are characterized using the copula function framework:
[0038]
[0039] in, This represents the cumulative distribution function value of the d-th random variable. Denotes the Copula joint distribution function. The cumulative distribution function value represents the marginal distribution. Represents a random variable. The dimension is represented; the joint distribution of the multidimensional random variables is decomposed into two parts, marginal distribution and dependency structure, by using the copula function. Based on the characteristics of the load data, t-copula is selected as the basic dependency structure, and the relevant parameters are estimated by the maximum likelihood method. At the same time, an adaptive strategy is adopted for the selection of the marginal distribution.
[0040] Multidimensional features are extracted by characterizing the nonlinear correlation features in the load data of power distribution equipment, including the construction of statistical feature systems and morphological feature systems.
[0041] The statistical moment feature system includes mean, variance, skewness, and kurtosis. The mean is calculated... The variance reflects the central tendency of the load level through Describe the dispersion of load fluctuations and calculate the skewness. and the kurtosis Characterizes the asymmetry and peak characteristics of load distribution;
[0042] The morphological feature system includes spikes in the opening operation suppression signal, valleys in the closing operation filling signal, and abrupt changes in the morphological gradient detection load curve.
[0043] As a preferred embodiment of the distribution transformer load prediction method based on multi-scale fluctuation pattern adaptive decomposition described in this invention, the construction of an efficient index structure includes constructing a multi-layer hybrid index and retrieval mechanism and supporting fast fluctuation pattern retrieval, including a multi-layer hybrid index, a feature dimension index, and support for fast fluctuation pattern retrieval.
[0044] The multi-layer hybrid index constructs a time-dimensional index structure based on a B+ tree, which organizes the timestamps of the load data as keys. Each non-leaf node stores the boundary values of the time interval, and the leaf nodes are linked in chronological order. At the same time, a caching mechanism is built in the B+ tree to cache the index information of hot time intervals in memory, and a pre-read strategy is used to preload data of adjacent time periods.
[0045] The feature dimension index employs a hybrid indexing strategy combining locality-sensitive hashing and KD-trees to index the feature vectors. Construct a hash function:
[0046]
[0047] Where a is a random vector and b is a random offset. Using the bucket width parameter, a KD tree is used to partition the feature space. At each node, the dimension with the largest variance is selected as the partition axis to perform adaptive segmentation of the feature space.
[0048] The fast retrieval of the fluctuation pattern is achieved through an index structure based on an inverted linked list. Each fluctuation pattern is assigned a unique ID, and a mapping relationship between the pattern ID and the feature vector is established through the inverted linked list. Considering the high-dimensionality of the feature vector, a compression storage strategy is adopted, including feature quantization and sparse representation, to reduce storage overhead. At the same time, the inverted linked list adopts a block storage method to store patterns with similar features in adjacent physical spaces.
[0049] Post-indexing retrieval mechanism optimization includes multi-path parallel retrieval strategies and multi-level pruning strategies;
[0050] The multi-path parallel retrieval strategy filters data in the time dimension, locates data within the target time window, performs similarity matching in the feature space, finds candidate similar fluctuation patterns through the combination of LSH and KD trees for pattern verification, and determines the final matching result by calculating a precise similarity metric.
[0051] The multi-level pruning strategy includes: distance-based pruning, which quickly eliminates regions that cannot contain the target pattern by calculating the minimum distance between the query vector and the feature space partition; density-based pruning, which prioritizes high-density regions and skips sparse regions by utilizing the density distribution characteristics of fluctuating patterns; and similarity-threshold-based pruning, which sets a dynamic similarity threshold to terminate the verification of candidate patterns that do not meet the requirements.
[0052] As a preferred embodiment of the distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition described in this invention, the improved dynamic time warping includes: constructing an efficient index structure and performing adaptive matching prediction; and optimizing distance metric and search strategy through an improved dynamic time warping algorithm.
[0053] The optimized distance metric includes constructing a feature distance vector using a multi-feature fusion distance calculation method. , where each component Representing distance metrics across different feature dimensions, an adaptive weight allocation mechanism based on entropy weighting is used for features. and j, information entropy Calculated as
[0054]
[0055] in, Calculate the feature weights for the normalized feature values:
[0056]
[0057] in, Let i be the weight of the i-th indicator. Let i be the entropy value of the i-th index, while ensuring
[0058]
[0059] The final fusion distance is expressed as ;
[0060] Simultaneously, the path constraints of DTW are optimized, and bandwidth parameters are adjusted. Dynamically adjust based on the periodicity of the sequence:
[0061]
[0062] in, For sequence length, The periodicity intensity of the sequence, and To adjust the parameters, and simultaneously through slope constraints The degree of inclination of the control path, among which, Through cross-validation Optimize selection within the range;
[0063] The search strategy employs an improved A* heuristic algorithm for optimal path search, defining nodes. Evaluation function of ode :
[0064]
[0065]
[0066] in, Indicates the distance from the starting point to the node. The actual cumulative distance, For the node The estimated distance to the destination is determined by a heuristic function based on a lower bound estimate:
[0067]
[0068] in, For heuristic distance estimation from node n to the target, For the k-th element of the target sequence, For the i-th element of the reference sequence, For the search window width, Sequence length;
[0069] Furthermore, a multi-level pruning strategy is employed to utilize the lower bound of LB_Keogh on the sequence. Constructing the envelope and Perform fast filtering when candidate sequences If the distance to the envelope exceeds the current optimal value, the branch is pruned directly using the Early abandoning strategy. During the cumulative distance calculation, the calculation is stopped immediately when the threshold is exceeded. The lower bound estimate is gradually tightened through the Cascadinglower bounds mechanism, and multi-level filtering from coarse to fine is performed.
[0070] As a preferred embodiment of the distribution transformer load forecasting method based on multi-scale fluctuation mode adaptive decomposition described in this invention, the anomaly detection optimization includes: employing an improved ensemble isolated forest algorithm for anomaly detection; and calculating the information gain value of each feature for each node by optimizing the feature space partitioning strategy combined with an information gain-based feature selection mechanism. :
[0071]
[0072] in, Let the information entropy of dataset D be , Given the conditional entropy under feature f, the feature with the maximum information gain is selected as the splitting feature to capture abnormal patterns in the load data of power distribution equipment.
[0073] The split threshold is selected by a dynamic split threshold determination method. Based on the distribution characteristics of the current node data, the probability density function of the data distribution is obtained by kernel density estimation method, and the optimal split point is selected from the local minimum points of the density function.
[0074] The computational accuracy is improved by using an improved local anomaly factor algorithm, and the distance metric is optimized by using generalized Minkowski distance.
[0075] In the calculation of local anomaly factors, the reach-distance is optimized to take into account the local density information of sample points. In the density estimation stage, an adaptive bandwidth selection mechanism is combined, and the bandwidth parameter is determined by minimizing the mean square error criterion. The improved isolation forest and local anomaly factor algorithms are integrated, and the detection results of the two algorithms are fused by weighted voting. The weight values are dynamically updated through online learning.
[0076] The optimized online correction mechanism includes adaptive window design and robust estimation optimization;
[0077] The adaptive window design establishes a base window length of 24 hours × 4 samples per hour × 7 days, i.e., L = 672 sampling points. The window length is adjusted based on prediction error feedback. When the continuous prediction error exceeds a set threshold, the window length is adjusted accordingly. Adjustments are made, where δ is the dynamic adjustment coefficient, which is positively correlated with the prediction error;
[0078] Optimize window data weighting using an exponential decay weighting strategy:
[0079]
[0080] Where α is the forgetting factor and i is the time sequence number of the data point; the forgetting factor α adopts an adaptive adjustment mechanism, with an initial value of 0.1, and is dynamically adjusted according to the changes in prediction error; when the prediction error increases, the value of α is increased to accelerate the forgetting of historical data; when the prediction error decreases, the value of α is decreased to maintain model stability, and the adjustment range of α is limited to [0.05, 0.3].
[0081] The robust estimation optimization includes improving the M-estimation method and enhancing the S-estimation method. Outliers are handled by improving the M-estimation method and combining it with the Huber loss function.
[0082] The enhanced S-estimation method optimizes the upper limit of the outlier ratio tolerated by the estimator. Considering the characteristics of distribution network load data, the breakpoint is set as:
[0083]
[0084] Where n is the number of samples and p is the number of parameters, an improved weight function is used during the weight iteration process. ,in, Let be the residual, s be the scaling estimate, and during the iteration process, when the Euclidean distance between two adjacent parameter estimates is less than 1... Convergence is determined at that time.
[0085] Another objective of this invention is to provide a distribution transformer load forecasting system based on multi-scale fluctuation mode adaptive decomposition. This invention captures the fluctuation characteristics of load data more accurately through multi-scale fluctuation decomposition and feature extraction; it achieves the adaptability and robustness of load forecasting: through adaptive matching forecasting and anomaly handling, the system can adapt to fluctuation modes at different time scales and handle abnormal data; by monitoring and forecasting distribution transformer load in real time, it provides safety early warning for power grid operation and ensures the stable operation of the power grid.
[0086] As a preferred embodiment of the distribution transformer load prediction system based on multi-scale fluctuation mode adaptive decomposition according to the present invention, it is characterized by including a multi-scale fluctuation decomposition module, a characteristic fluctuation library construction module, an adaptive matching prediction module, and an anomaly handling module.
[0087] The multi-scale fluctuation decomposition module takes the original load data as input, uses the improved maximum overlap discrete wavelet packet transform to decompose the load signal at multiple scales, optimizes the energy distribution, and transmits the characteristics of the decomposed load signal to the feature fluctuation library construction module.
[0088] The feature fluctuation library construction module takes the load signal features after multi-scale decomposition as input, performs feature extraction and characterization optimization, and constructs a feature fluctuation library to provide data support for the adaptive matching prediction module.
[0089] The adaptive matching prediction module performs fine characterization of fluctuation patterns through adaptive density estimation and improved multi-level clustering strategy, constructs an efficient index structure, and makes predictions, while passing the prediction results to the anomaly handling module.
[0090] The anomaly handling module inputs the prediction results and combines them with integrated isolated forest and local anomaly factors for anomaly detection optimization. It optimizes the online correction mechanism through adaptive window and robust estimation. The corrected prediction results are used for online monitoring of power grid safety margin.
[0091] A computer device includes a memory and a processor, the memory storing a computer program, characterized in that the processor executes the computer program to implement the steps of the distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition.
[0092] A computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of the distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition.
[0093] The beneficial effects of this invention are as follows: By inputting raw data and performing multi-scale decomposition, this invention optimizes the distribution of load signal energy and extracts key features, improving the accuracy and efficiency of load forecasting. Subsequently, by employing adaptive density estimation and an improved multi-level clustering strategy, we refine the fluctuation patterns, construct an efficient index structure, and perform forecasting through adaptive matching. This not only identifies the time-scale features in the load data but also improves the accuracy and adaptability of the forecast. Improved dynamic time warping and particle swarm optimization algorithms further optimize the distance metric and search strategy, constructing a combined forecasting framework that enhances the flexibility and robustness of the forecast. Deep mining and enhancement processing of feature extraction and characterization optimization, along with the construction of a multi-level hybrid index and retrieval mechanism, comprehensively capture the intrinsic features of the load data, improving the accuracy of the forecast and system performance. Finally, the introduction of anomaly detection optimization and online correction mechanisms, through the integration of isolated forest and local anomaly factor algorithms, as well as adaptive window design and robust estimation optimization, reduces forecast errors, improves the adaptability and stability of the forecast model, and achieves the beneficial effect of improving overall forecast efficiency and accuracy. Attached Figure Description
[0094] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:
[0095] Figure 1 The above is a flowchart of a distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition, which is provided as an embodiment of the present invention.
[0096] Figure 2 The diagram below shows a system scheme block diagram of a distribution transformer load prediction system based on multi-scale fluctuation mode adaptive decomposition, as provided in one embodiment of the present invention. Detailed Implementation
[0097] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0098] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0099] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is mutually exclusive, either alone or selectively, with other embodiments.
[0100] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.
[0101] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0102] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0103] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides a method for predicting distribution transformer load based on adaptive decomposition of multi-scale fluctuation modes, including:
[0104] S1: Input the original data, use the improved maximum overlap discrete wavelet packet transform to perform multi-scale decomposition of the load signal, optimize the energy distribution, perform feature extraction and characterization optimization, and construct a feature fluctuation library to store the features.
[0105] Furthermore, the improved maximum overlap discrete wavelet packet transform for multi-scale decomposition of load signals includes using the symmetric sym8 wavelet basis function as the basic decomposition tool, with a tight support characteristic of set support of 16, which can effectively capture the multi-scale features in the distribution network load data.
[0106] The optimal wavelet basis is selected by minimizing the reconstruction error; the error function is... for:
[0107]
[0108] Where x is the original signal, For reconstructing the signal;
[0109] Set the reconstruction error threshold to When the reconstruction error is less than At that time, it was considered that the wavelet basis selection was appropriate;
[0110] Wavelet basis functions must simultaneously satisfy three key properties, the first of which is orthogonality:
[0111]
[0112] The second key property is tight support:
[0113]
[0114] The third key property is the vanishing moment:
[0115]
[0116] in, Let the wavelet basis functions be translated at scale j and k. For the Kronecker delta function, For time variables, Let ψ be the support set of the wavelet basis function. The length parameter of the wavelet basis function. The k-th power of time, For wavelet basis functions, Let be the order of the vanishing moment;
[0117] To ensure the stability and reliability of wavelet transform, the selected wavelet basis functions must simultaneously satisfy three key properties. First, orthogonality: wavelet functions at different scales and translations must remain orthogonal. This can be verified through integration, ensuring that there is no redundant information among the coefficients obtained from the decomposition. Second, compact support: the wavelet functions must be non-zero in the finite interval [0, 2N-1]. This helps improve computational efficiency and reduce the impact of boundary effects. Third, a sufficiently high order of vanishing moments: the integral of the wavelet function with the polynomial is zero. This guarantees a good approximation of the polynomial trend in the load data.
[0118] To address the common endpoint effect problem in load data processing, this invention employs an adaptive extension algorithm to handle the endpoint effect in load data, establishing an AR(p) autoregressive model to describe the temporal correlation of load data:
[0119]
[0120] in, The time series value at time t, These are the autoregressive coefficients. This is the white noise term;
[0121] The model order p is dynamically determined by minimizing the AIC criterion:
[0122]
[0123] The AIC criterion takes into account both the model's goodness of fit (characterized by the sum of squared residuals, RSS) and complexity (characterized by the number of model parameters, p), achieving a balance between the two.
[0124] The extension length L is adaptively determined based on the data length N and the signal-to-noise ratio (SNR):
[0125]
[0126] This ensures the sufficiency of the extension while avoiding the computational burden caused by excessive extension.
[0127] Denoising is achieved using an improved soft thresholding shrinkage method, with the threshold parameter... Employing a scale-dependent adaptive form:
[0128]
[0129] in, Let be the threshold parameter for the j-th layer, and α be a scaling factor to balance the noise levels at different decomposition scales. ; The standard deviation of noise. Given the data length, a robust estimate is made using the median absolute deviation method:
[0130]
[0131] Threshold selection is optimized using the SURE criterion, and the optimal threshold is determined by minimizing the unbiased risk estimate.
[0132]
[0133] in, For the sample size, These are wavelet coefficients. is the threshold parameter, RSS is the sum of squared residuals, and SNR is the signal-to-noise ratio.
[0134] It can achieve optimal noise suppression while preserving signal characteristics.
[0135] The above technical solutions are particularly suitable for load forecasting scenarios at the equipment level in power distribution networks.
[0136] Because distribution transformer loads exhibit strong randomness and multi-scale characteristics, the optimized wavelet transform can effectively separate load fluctuation components at different scales. Meanwhile, considering potential abnormal fluctuations and measurement noise in the load data, the endpoint effect suppression and threshold optimization schemes of this invention ensure the reliability of the preprocessing results, laying the foundation for subsequent feature extraction and predictive modeling.
[0137] Specifically, optimizing the energy distribution includes, for the wavelet decomposition coefficients of the j-th layer... Calculate the wavelet energy proportion of the j-th layer.
[0138]
[0139] in, Calculate the relative energy ratio for the signal length:
[0140]
[0141] in, The optimal decomposition level is dynamically determined by summing the energy values of all layers.
[0142] Set adaptive energy threshold :
[0143]
[0144] This setup can effectively adapt to the energy distribution characteristics of different types of load data. Here, J represents the total number of decomposition layers; the mutual information between adjacent decomposition layers is calculated through a constraint mechanism based on mutual information. :
[0145]
[0146] in, Describes the joint probability distribution. and These represent the marginal probability distributions; when the mutual information value exceeds the threshold... When the value is 0.1, it indicates that there is information redundancy and hierarchical adjustment is required. This constraint mechanism is particularly suitable for handling complex load situations in distribution networks where different user types (such as residential, commercial, and industrial) are mixed.
[0147] The optimal level is dynamically determined using an improved information gain calculation method.
[0148]
[0149] in, and For the j-th layer and the j-th layer Shannon entropy value, To balance the model complexity for regularization terms, As a balance factor, This design can effectively avoid over-decomposition and improve the model's ability to adapt to sudden load changes.
[0150] A distance function is defined using a hierarchical merging strategy based on dynamic programming:
[0151]
[0152] when Merging at that time This was determined through cross-validation; this merging strategy is particularly suitable for handling multi-cycle superposition phenomena in the equipment-level loads of distribution networks.
[0153] in, This represents the distance between adjacent floors. To balance the weighting coefficients of the importance of the two distance metrics, For the i-th layer signal, For the j-th layer signal, The signal is the Euclidean distance. For probability distribution, Let i be the probability distribution of the i-th layer. Let be the probability distribution of the j-th layer.
[0154] The energy distribution optimization decomposition hierarchical method of this invention fully considers the characteristics of distribution network load data. Through adaptive energy analysis and dynamic hierarchical optimization, it can effectively extract load change characteristics at different time scales, providing a reliable basis for subsequent load forecasting. Practice shows that this method exhibits good adaptability and stability when dealing with different types of distribution transformer loads, such as residential areas, commercial areas, and industrial parks.
[0155] It should be noted that in-depth mining and enhancement processing are performed on the load data of distribution transformers to construct a characteristic fluctuation library;
[0156] The load data includes an amplitude feature family and a shape feature family. The amplitude feature family includes maximum amplitude, root mean square amplitude, peak factor, margin factor, and waveform factor. The shape feature family includes skewness, kurtosis, waveform index, and impulse index.
[0157] The maximum amplitude is This reflects the extreme characteristics of load fluctuations;
[0158] The root mean square amplitude is It can effectively characterize the effective value level of the load;
[0159] The peak factor is This can reflect the peak characteristics of the load;
[0160] The margin factor is ;
[0161] The waveform factor is It can depict the morphological characteristics of load fluctuations from different perspectives;
[0162] The skewness is This can characterize the asymmetry of load distribution;
[0163] The kurtosis is , used to describe the steepness of load distribution;
[0164] The waveform index is This can effectively reflect the degree of drastic change in load;
[0165] The pulse index is The pulse characteristics of the load are measured by the ratio of the maximum amplitude to the average absolute value.
[0166] in, The square of the first derivative of the reconstructed load signal within time t. For random variables, , This is the mean of skewness and kurtosis. , These are the standard deviations of skewness and kurtosis, respectively.
[0167] The frequency domain characteristics are enhanced by improving the power spectral density estimation method. A multi-window spectral estimation technique is employed, which involves weighted superposition of multiple independent spectral estimation results. The final power spectral density is obtained. :
[0168]
[0169] Adopting an improved Hamming window :
[0170]
[0171] It ensures good frequency resolution while effectively suppressing spectral leakage.
[0172] Dynamic overlap rate adjustment mechanism:
[0173]
[0174] in, The overlap rate, The weight coefficients of the k-th window function are... Let be the complex representation of the k-th frequency component. The power spectrum of the frequency components. 'Index to the sampling points of the window function, u is the total length of the window function, which ensures the stability of the spectral estimation;
[0175] Frequency band energy features extracted by energy distribution characterization include sub-band energy ratio, energy entropy, spectral centroid, and spectral broadening. The sub-band energy ratio includes calculating the proportion of each sub-band energy to the total energy. This reflects the distribution characteristics of energy in the frequency domain;
[0176] The energy entropy is It can measure the uniformity of energy distribution.
[0177] The spectral centroid was obtained by energy-weighted averaging. This allows us to determine the center location of the energy distribution.
[0178] Spectral broadening is obtained by calculating the second moment of the frequency relative to the centroid. ,in, Let i be the energy probability of the i-th layer frequency band. For frequency values, The power spectral density corresponding to the frequency reflects the degree of dispersion of energy distribution. These characteristics together constitute a comprehensive characterization of the load's frequency domain properties.
[0179] The aforementioned extraction process of time-domain and frequency-domain features fully considers the characteristics of distribution network load data, effectively capturing the load characteristics of different types of users (such as residential, commercial, and industrial users). These features retain important indicators from traditional load analysis while introducing new feature representation methods, providing rich feature support for subsequent load forecasting. Through this multi-dimensional feature extraction and representation optimization, the model's ability to perceive load change patterns is significantly improved.
[0180] S2: Through adaptive density estimation and improved multi-level clustering strategy, fluctuation patterns are finely characterized from different time scales, an efficient index structure is constructed, and prediction is performed through adaptive matching.
[0181] Furthermore, the data distribution characteristics are characterized by an improved kernel density estimation, using the Epanechnikov kernel function:
[0182]
[0183] It can effectively reduce computational complexity and ensure the smoothness of the estimation results.
[0184] Adaptive method is used to select bandwidth:
[0185]
[0186] This selection method takes into account both the dispersion of the data and can be dynamically adjusted with the sample size, ensuring the accuracy of the density estimation.
[0187] A dynamic density threshold determination method based on silhouette coefficients is used to determine the density threshold.
[0188]
[0189] in, The standard deviation of the data. , These are the density mean and standard deviation, respectively. To optimize through contour coefficients, This ensures the compactness and separability of the clustering results; For indicator functions, For sample size;
[0190] After characterization, an improved multi-level clustering strategy is used to identify fluctuation patterns at different time scales. Based on the inherent laws of distribution network load changes, the time scale is divided into four levels: intraday, daily, weekly, and monthly.
[0191] The intraday layer covers fluctuations from 15 minutes to 4 hours, reflecting short-term electricity consumption behavior; the daily layer covers changes from 4 hours to 24 hours, reflecting the characteristics of the daily load curve; the weekly layer spans from 1 day to 7 days, depicting the differences between weekdays and weekends; and the monthly layer ranges from 7 days to 30 days, capturing seasonal variation characteristics.
[0192] Within each timescale layer, an improved density clustering algorithm is used for fluctuation pattern identification, where the determination of core points adopts an adaptive approach.
[0193]
[0194] in, To account for the current layer's data volume, a density-sensitive factor is introduced during the boundary point absorption process. The criteria for determining the ownership of boundary points are as follows:
[0195]
[0196] in, For point With point The distance between them and These are the corresponding density values. The neighborhood radius;
[0197] in, ,when When the value is less than 0.35, a relatively dense cluster is formed; when... When the value is greater than 0.35, points with large density differences are allowed to be grouped into the same category;
[0198] By employing the improved hierarchical density clustering method described above, this invention can accurately capture the fluctuation characteristics of distribution transformer load data across different time scales, providing a reliable pattern matching basis for subsequent load forecasting. Practice has shown that this method exhibits good adaptability and stability when processing different types of distribution transformer load data, including residential areas, commercial areas, and industrial parks.
[0199] The refined characterization of the fluctuation pattern includes probabilistic distribution modeling using an improved Gaussian mixture model and parameter estimation using an optimized expectation-maximization algorithm.
[0200]
[0201] in, For conditional expectation, For model parameters, The parameters for the current iteration. Let be the likelihood function containing latent variables. For expectation operators;
[0202] In each iteration, the parameters are dynamically adjusted and the step size is updated based on the gradient change of the likelihood function through an adaptive learning rate adjustment mechanism. Simultaneously, the optimal number of components K is adaptively determined using the Bayesian information criterion.
[0203]
[0204]
[0205] Among them, BIC stands for Bayesian Information Criterion. Here, M represents the log-likelihood value, and M represents the number of samples. This method effectively balances the relationship between model complexity and fitting accuracy.
[0206] Meanwhile, the nonlinear correlation characteristics in the load data of power distribution equipment are characterized using the copula function framework:
[0207]
[0208] in, This represents the cumulative distribution function value of the d-th random variable. Denotes the Copula joint distribution function. The cumulative distribution function value represents the marginal distribution. Represents a random variable. The dimension is represented by the copula function, which decomposes the joint distribution of multidimensional random variables into two parts: marginal distribution and dependency structure. Based on the characteristics of the load data, t-copula is selected as the basic dependency structure, and the relevant parameters are estimated by the maximum likelihood method. At the same time, the selection of the marginal distribution adopts an adaptive strategy, which can flexibly select normal distribution, log-normal distribution or generalized Pareto distribution, etc., according to the data characteristics.
[0209] Multidimensional features are extracted by characterizing the nonlinear correlation features in the load data of power distribution equipment, including the construction of statistical feature systems and morphological feature systems.
[0210] The statistical moment feature system includes mean, variance, skewness, and kurtosis. The mean is calculated... The variance reflects the central tendency of the load level through Describe the dispersion of load fluctuations and calculate the skewness. and the kurtosis These statistical moment features characterize the asymmetry and peak characteristics of load distribution; they can reflect the probability distribution characteristics of power distribution equipment load from different perspectives, providing an important basis for subsequent load forecasting.
[0211] The morphological feature system includes opening operations, closing operations, and morphological gradients;
[0212] The opening operation is: Suppress spikes in the signal;
[0213] The closing operation is: Fill in the valleys in the signal;
[0214] The morphological gradient is Detect abrupt changes in the load curve;
[0215] Where B is the structuring element and X is the signal. and These represent expansion and erosion operations, respectively.
[0216] In practical applications, the refined characterization method for fluctuation patterns of this invention can effectively handle the diversity and complexity of load data in power distribution network equipment. For example, for distribution transformers in commercial areas, this method can accurately capture the load fluctuation patterns during business hours; for power distribution equipment in industrial parks, it can effectively identify the characteristics of load abrupt changes during the transition between different production shifts. This refined characterization provides a reliable data foundation for subsequent load forecasting, significantly improving the accuracy and robustness of the forecasting model.
[0217] It should also be noted that, in order to meet the retrieval requirements of fluctuation patterns in load forecasting at the equipment level of distribution networks, the construction of an efficient index structure includes building a multi-layer hybrid index and retrieval mechanism and supporting fast fluctuation pattern retrieval, including multi-layer hybrid index, feature dimension index, and support for fast fluctuation pattern retrieval.
[0218] The multi-layer hybrid index constructs a time-dimensional index structure based on a B+ tree, which organizes the timestamps of the load data as keys. Each non-leaf node stores the boundary values of the time interval, and the leaf nodes are linked in chronological order. At the same time, a caching mechanism is built in the B+ tree to cache the index information of hot time intervals in memory, and a pre-read strategy is used to preload data of adjacent time periods.
[0219] The feature dimension index employs a hybrid indexing strategy combining locality-sensitive hashing and KD-trees to index the feature vectors. Construct a hash function:
[0220]
[0221] Where a is a random vector and b is a random offset. Using the bucket width parameter, a KD tree is used to partition the feature space. At each node, the dimension with the largest variance is selected as the partition axis to perform adaptive segmentation of the feature space.
[0222] To support fast retrieval of fluctuation patterns, fast retrieval is achieved through an index structure based on an inverted linked list. Each fluctuation pattern is assigned a unique ID, and a mapping relationship between the pattern ID and the feature vector is established through the inverted linked list. Considering the high-dimensionality of the feature vector, a compression storage strategy is adopted, including feature quantization and sparse representation, to reduce storage overhead. At the same time, the inverted linked list adopts a block storage method, storing patterns with similar features in adjacent physical spaces, which improves access locality.
[0223] Post-indexing retrieval mechanism optimization includes multi-path parallel retrieval and multi-level pruning strategies;
[0224] The multi-path parallel retrieval strategy filters data in the time dimension, locates data within the target time window, performs similarity matching in the feature space, finds candidate similar fluctuation patterns through the combination of LSH and KD trees for pattern verification, and determines the final matching result by calculating a precise similarity metric. The entire retrieval process adopts multi-threaded parallel processing, making full use of the parallel computing capabilities of modern processors.
[0225] The multi-level pruning strategies include: distance-based pruning, which quickly eliminates regions unlikely to contain the target pattern by calculating the minimum distance between the query vector and the feature space partition; density-based pruning, which prioritizes high-density regions and skips sparse regions by utilizing the density distribution characteristics of fluctuating patterns; and similarity-threshold-based pruning, which sets a dynamic similarity threshold to terminate the verification of candidate patterns that do not meet the requirements. The combined use of these pruning strategies significantly reduces the computational complexity of the retrieval process.
[0226] The aforementioned index structure and retrieval mechanism design fully considers the characteristics of distribution network load data and can effectively support the retrieval needs of large-scale load data fluctuation patterns. Experiments show that this scheme achieves good results in terms of retrieval efficiency and accuracy, and is of great significance for improving the performance of load forecasting.
[0227] S3: By improving dynamic time warping, distance metrics and search strategies are optimized, and weights are dynamically adjusted through an improved particle swarm optimization algorithm to construct a combined prediction framework.
[0228] Furthermore, we construct an efficient index structure and perform adaptive matching prediction, and optimize distance metrics and search strategies through an improved dynamic time warping algorithm.
[0229] The optimized distance metric includes constructing a feature distance vector using a multi-feature fusion distance calculation method. , where each component Representing distance metrics across different feature dimensions, an adaptive weight allocation mechanism based on entropy weighting is used for features. and j, information entropy Calculated as
[0230]
[0231] in, Calculate the feature weights for the normalized feature values:
[0232]
[0233] in, Let i be the weight of the i-th indicator. Let i be the entropy value of the i-th index, while ensuring that:
[0234]
[0235] The final fusion distance is expressed as ;
[0236] Simultaneously, the path constraints of DTW are optimized, and bandwidth parameters are adjusted. Dynamically adjust based on the periodicity of the sequence:
[0237]
[0238] in, For sequence length, The periodicity intensity of the sequence, and To adjust the parameters, and simultaneously through slope constraints The degree of inclination of the control path, among which, Through cross-validation Optimize selection within the range;
[0239] The search strategy employs an improved A* heuristic algorithm for optimal path search, defining nodes. Evaluation function of ode :
[0240]
[0241]
[0242] in, Indicates the distance from the starting point to the node. The actual cumulative distance, For the node The estimated distance to the destination is determined by a heuristic function based on a lower bound estimate:
[0243] ,
[0244] in, For heuristic distance estimation from node n to the target, For the k-th element of the target sequence, For the i-th element of the reference sequence, For the search window width, Sequence length.
[0245] Furthermore, a multi-level pruning strategy is employed to utilize the lower bound of LB_Keogh on the sequence. Constructing the envelope and Perform fast filtering when candidate sequences If the distance to the envelope exceeds the current optimal value, the branch is pruned directly. The Early abandoning strategy is adopted. During the cumulative distance calculation, the calculation is stopped immediately when the threshold is exceeded. The lower bound estimate is gradually tightened through the Cascadinglower bounds mechanism to perform multi-level filtering from coarse to fine.
[0246] Experiments show that in a case study of load forecasting for a distribution transformer in a residential area, the improved DTW algorithm increases matching accuracy by 15.3% compared to the traditional method, while also improving computational efficiency by approximately 3.5 times. This algorithm exhibits good adaptability to time-dependent elastic deformation in load sequences and can effectively identify similar load patterns, providing a reliable pattern matching basis for subsequent forecasting.
[0247] It should be noted that the dynamic weight adjustment includes, firstly, weight optimization using an improved particle swarm optimization algorithm. This algorithm introduces a non-linearly decreasing inertial weight strategy, enabling particles to possess strong global exploration capabilities in the early stages of the search. The particle velocity update uses an improved velocity v update formula.
[0248]
[0249]
[0250]
[0251] in, The inertial weights are non-linearly decreasing. and These are the maximum and minimum values of the inertia weight, respectively. This represents the current iteration number. The maximum number of iterations, and As a learning factor, and These are the upper and lower limits of the learning factor, respectively. For the individual's optimal position, The optimal position is determined by this non-linear decreasing strategy, which allows the inertial weight to decrease slowly in the early stages of optimization and then decrease rapidly in the later stages, thus better balancing global search and local development.
[0252] Simultaneously through the Sigmoid mapping function Constraining particle positions through weight normalization Ensure the sum of the weights is 1:
[0253]
[0254]
[0255] The load characteristics are adapted to dynamic changes through a time decay update mechanism. Calculate the time decay factor, where, , The half-life is used to determine the weight of samples. This mechanism allows newer samples to have a larger weight, while older samples have a gradually decreasing weight, thus reflecting changes in load characteristics in a timely manner.
[0256] Incremental learning is performed using a sliding window mechanism, with a window length of... Set as This refers to the number of sampling points in one week. For samples within the window, their contribution to model updates is measured by sample importance weights.
[0257]
[0258] in, The distance between the sample and the current predicted target. For scale parameters;
[0259] Through an abnormal punishment mechanism:
[0260]
[0261] in, For the outlier scores of the sample, This is a penalty coefficient to reduce the impact of outlier samples on model updates;
[0262] This adaptive weighting mechanism fully considers the characteristics of distribution network load data, effectively handles the differences in load characteristics among different types of users (such as residential, commercial, and industrial users), and can adapt to dynamic changes in load over time and seasons. Experimental verification shows that this mechanism exhibits excellent predictive performance in multiple real-world scenarios, with a prediction accuracy significantly superior to traditional methods.
[0263] It should be noted that the constructed combined prediction framework includes a deep learning architecture with multi-head attention mechanism, a deep residual network structure, and a multi-level feature fusion mechanism;
[0264] The deep learning architecture of the multi-head attention mechanism includes two key components: a self-attention layer and a temporal attention layer.
[0265] The self-attention layer is transformed by a linear transformation matrix. , and Map the input features X to the feature space of the query vector Q, the key vector U, and the value vector V:
[0266]
[0267] Among them, the query vector Q is used to capture the load characteristics at the current moment, while the key vector U and the value vector V are used to store key information and specific values of historical load patterns, respectively.
[0268] By calculating the dot product of Q and U and normalizing it using the scaling factor sf, the attention weight matrix is obtained, which reflects the correlation strength between load characteristics at different times:
[0269]
[0270] The model's ability to express multi-dimensional features in the load sequence is enhanced through a multi-head attention mechanism, i.e., parallel operation. Each attention head has an independent parameter matrix and captures feature associations in different subspaces. These associations are then processed using a concat operation. The outputs of each head are concatenated and then fused using a linear transformation matrix WO to obtain the final multi-head attention output:
[0271]
[0272] This design allows the model to simultaneously focus on the dependencies of the load sequence across different feature dimensions, improving the comprehensiveness of feature extraction.
[0273] In the time attention layer, the temporal dependencies in the load sequence are characterized by a relative position encoding mechanism. The temporal information is modeled by calculating the relative distance between any two moments. Unlike the traditional absolute position encoding, the relative position encoding models the temporal information by calculating the relative distance between any two moments. This design is more in line with the periodic characteristics of load changes.
[0274] Meanwhile, a local-global dual attention structure is added. Local attention captures short-term load fluctuation features, while global attention captures long-term trends. The combination of the two enhances the model's ability to capture multi-scale time features.
[0275] To enhance the model's ability to learn periodic patterns in load sequences, a periodicity enhancement module is integrated into the time attention layer. This module decomposes the input sequence into periodic components at multiple scales, including daily, weekly, and monthly cycles. Through a defined periodicity attention sublayer, the correlation between different periodic components is established, thereby improving the model's ability to identify and predict periodic load patterns.
[0276] The deep residual network structure establishes a direct connection between shallow and deep features through identity mapping. For each residual unit, the output can be expressed as... F(x) is the residual mapping and x is the identity mapping. After outputting the residual mapping, a pre-activation design is used to optimize the residual learning effect. The batch normalization layer and the activation function layer are placed before the convolutional layer. This design not only improves the convergence performance of the model, but also enhances the stability of feature propagation.
[0277] By using an identity mapping protection mechanism, residual information is transmitted in deep networks to reduce loss. This structural design effectively alleviates the gradient vanishing problem during deep network training while maintaining the model's sensitivity to the original features.
[0278] Regarding feature fusion, this invention constructs a multi-layered feature fusion mechanism. First, skip connections enable direct transfer of features from different depth layers, avoiding information loss in deep networks. Second, a dense connection strategy is employed to cumulatively fuse features from each layer, fully utilizing feature information at different levels of abstraction. Finally, a multi-scale feature fusion module enables adaptive fusion of features at different time scales, improving the model's ability to perceive load changes.
[0279] S4: Combine integrated isolated forest and local anomaly factors to optimize anomaly detection, and optimize the online correction mechanism through adaptive windowing and robust estimation.
[0280] Furthermore, an improved ensemble isolated forest algorithm is employed for anomaly detection. By optimizing the feature space partitioning strategy and combining it with an information gain-based feature selection mechanism, the information gain value of each feature is calculated for each node's feature selection. :
[0281]
[0282] in, Let the information entropy of dataset D be , Given the conditional entropy under feature f, the feature with the maximum information gain is selected as the splitting feature to capture abnormal patterns in the load data of power distribution equipment.
[0283] A dynamic splitting threshold determination method is used to select the splitting threshold. Based on the distribution characteristics of the current node data, a kernel density estimation method is used to obtain the probability density function of the data distribution. The optimal splitting point is then selected from the local minima of the density function. Simultaneously, the tree depth... Determined through an adaptive approach:
[0284]
[0285] in, This setting, based on the number of samples, can effectively balance computational efficiency and model expressive power.
[0286] Improvements to anomaly calculation:
[0287]
[0288] Where z(x) represents the sample Path length, For the desired path length, As the normalization factor, , This is the harmonic number; this calculation method takes into account the degree of isolation of the sample in the feature space, and can more accurately identify outliers in the power distribution load data.
[0289] The computational accuracy is improved by refining the local anomaly factor algorithm. An adaptive k-value selection strategy is adopted for determining the k-nearest neighbors.
[0290]
[0291] Where n is the number of samples, and μ and σ are the mean and standard deviation of the historical best k value, respectively. This dynamic adjustment mechanism can better adapt to the data distribution with different load characteristics.
[0292] For distance measurement, the generalized Minkowski distance is used:
[0293]
[0294] in, , Let x and y be the i-th components. The norm parameter is used to better characterize the similarity relationships of distribution load data in a high-dimensional feature space. For different types of load characteristics (such as electricity consumption, power factor, etc.), a feature weighting strategy is adopted, and the weights are optimized and determined through a data-driven approach.
[0295] In the calculation of the local anomaly factor, the reach-distance is optimized to consider the local density information of the sample points. In the density estimation stage, an adaptive bandwidth selection mechanism is used, and the bandwidth parameter is determined by minimizing the mean square error criterion. This invention integrates the improved isolated forest and local anomaly factor algorithms, and uses a weighted voting method to fuse the detection results of the two algorithms. The weight values are dynamically updated through online learning, which can adaptively adjust the importance of different algorithms in different scenarios, thereby achieving a more robust anomaly detection effect. This integration strategy is particularly suitable for the complex anomaly pattern recognition problem in distribution network load data.
[0296] It should be noted that the online correction mechanism optimization is a key module in this invention, mainly comprising two parts: adaptive window design and robust estimation optimization. This mechanism can effectively handle abnormal fluctuations in load forecasting at the equipment level of the distribution network, improving forecast accuracy and stability.
[0297] The adaptive window design establishes a base window length of 24 hours × 4 (number of samples per hour) × 7 days, i.e., L = 672 sampling points. The window length is adjusted based on prediction error feedback. When the continuous prediction error exceeds a set threshold, the window length is adjusted accordingly. Adjustments are made, where δ is the dynamic adjustment coefficient, which is positively correlated with the prediction error;
[0298] For distribution transformers in residential areas, the value of δ ranges from [0.1, 0.3]; for distribution transformers in commercial areas, the value of δ ranges from [0.2, 0.4]; and for distribution transformers in industrial areas, the value of δ ranges from [0.3, 0.5].
[0299] Optimize window data weighting using an exponential decay weighting strategy:
[0300]
[0301] Where α is the forgetting factor and i is the time index of the data point; the forgetting factor α adopts an adaptive adjustment mechanism, with an initial value set to 0.1, and is dynamically adjusted according to changes in the prediction error. When the prediction error increases, the value of α is increased to accelerate the forgetting of historical data; when the prediction error decreases, the value of α is decreased to maintain model stability, and the adjustment range of α is limited to [0.05, 0.3].
[0302] The robust estimation optimization includes improving the M-estimation method and enhancing the S-estimation method. The improved M-estimation method, combined with the Huber loss function, handles outliers. When the absolute value of the residual does not exceed a threshold k, squared loss is used; otherwise, linear loss is used.
[0303]
[0304] The threshold k is estimated using the median absolute deviation (MAD) method, k = 1.345 × MAD, where MAD = median(|x i -median(x)|);
[0305] The influence function is designed to use a Huber-type function. ;
[0306] To further improve the robustness of the estimation, the enhanced S-estimation method optimizes the upper limit of the outlier proportion that the estimator can tolerate. Considering the characteristics of distribution network load data, the breakpoint is set as:
[0307]
[0308] Where n is the number of samples and p is the number of parameters, an improved weight function is used during the weight iteration process. ,in, Let be the residual, s be the scaling estimate, and during the iteration process, when the Euclidean distance between two adjacent parameter estimates is less than 1... Convergence is determined at that time.
[0309] The online correction mechanism of this invention has demonstrated excellent performance in practical applications. Taking a distribution transformer in a commercial area as an example, after adopting this mechanism, the mean absolute percentage error (MAPE) of load forecasting decreased from 3.8% to 2.8%, and the root mean square error (RMSE) decreased from 8.5kW to 6.2kW. Simultaneously, the reliability of the forecast results was significantly improved, with a 96.2% coverage rate for the 95% confidence interval. This fully demonstrates the significant advantages of this mechanism in handling abnormal fluctuations in distribution network load forecasting.
[0310] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
[0311] Example 2 is the second embodiment of the present invention, which provides a method for predicting distribution transformer load based on adaptive decomposition of multi-scale fluctuation modes. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0312] Example 1: Load prediction of a distribution transformer in a residential area, details are shown in the table below:
[0313] Table 1 Input Data
[0314]
[0315] Table 2 Parameter Configuration
[0316] MODWPT parameters Decomposition level: 5; Wavelet basis: sym8; Threshold: λ=0.05 DBSCAN parameters Eps=0.5; MinPts=5; Density factor: α=0.3 DTW parameters bandwidth: =48; Step mode: Symmetric; Distance metric: Euclidean distance PSO parameters Population size: 50; Maximum iterations: 100; W: 0.9 → 0.4; c1 = c2 = 2.0
[0317] Table 3 Data Preprocessing
[0318] Missing value handling Short-time: cubic spline interpolation; Medium-time: similar day substitution; Long-time: model estimation Outlier handling 3σ rule initial screening; LOF algorithm fine screening; expert rule verification Standardized processing Z-score standardization; range: [0,1]
[0319] Feature engineering includes time feature construction, meteorological feature extraction, and load feature calculation;
[0320] Model training includes parameter optimization, model ensemble, and prediction validation;
[0321] Table 4 Prediction Results
[0322] MAPE 3.2% RMSE 4.5kW Coverage 95.8%
[0323] Example 2: Load forecasting of a distribution transformer in a commercial area
[0324] Table 5 Input Data
[0325]
[0326] Table 6 Parameter Configuration
[0327] MODWPT parameters Number of decomposition levels: 6; Wavelet basis: sym20; Threshold: λ=0.06 Clustering parameters Eps=0.6; MinPts=6; Density factor: α=0.4 DTW parameters bandwidth: =72; Step mode: Symmetric; Distance metric: Euclidean distance Optimize parameters Population size: 60; Maximum iterations: 120; W: 0.9 → 0.3; c1 = c2 = 2.5
[0328] Table 7 Prediction Results
[0329] MAPE 2.8% RMSE 6.2kW Coverage 96.2%
[0330] Example 3: Load forecasting of distribution transformers in an industrial park
[0331] Table 8 Input Data
[0332]
[0333] Table 9 Parameter Configuration
[0334] MODWPT parameters Number of decomposition levels: 7; Wavelet basis: sym2; Threshold: λ=0.04 Clustering parameters Eps=0.4; MinPts=8; Density factor: α=0.5 DTW parameters bandwidth: =96; Step size mode: Symmetric; Distance metric: Mahalanobis distance Optimize parameters Population size: 80; Maximum iterations: 150; W: 0.9 → 0.35; c1 = c2 = 2.2
[0335] Table 10 Prediction Results
[0336] MAPE 2.5% RMSE 8.3kW Coverage 96.5%
[0337] Example 4: Load Prediction of a Distribution Transformer in a Large Shopping Mall
[0338] Table 11 Input Data
[0339]
[0340] Table 12 Parameter Configuration
[0341] MODWPT parameters Number of decomposition levels: 6; Wavelet basis: sym9; Threshold: λ=0.055 Clustering parameters Eps=0.55; MinPts=7; Density factor: α=0.35 DTW parameters bandwidth: =60; Step size mode: Asymmetric; Distance metric: Cosine distance Optimize parameters Population size: 70; Maximum iterations: 130; W: 0.85 → 0.3; c1 = c2 = 2.3
[0342] Table 13 Prediction Results
[0343] MAPE 3.0% RMSE 5.8kW Coverage 95.9%
[0344] Example 3, the third embodiment of the present invention, differs from the previous two embodiments in that:
[0345] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0346] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0347] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0348] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0349] Example 4, refer to Figure 2 This is the fourth embodiment of the present invention. This embodiment provides a distribution transformer load prediction system based on multi-scale fluctuation mode adaptive decomposition, including a multi-scale fluctuation decomposition module 10, a feature fluctuation library construction module 20, an adaptive matching prediction module 30, and an anomaly handling module 40.
[0350] The multi-scale fluctuation decomposition module 10 takes the original load data as input, uses the improved maximum overlap discrete wavelet packet transform to decompose the load signal at multiple scales, optimizes the energy distribution, and transmits the characteristics of the decomposed load signal to the feature fluctuation library construction module 10.
[0351] The feature fluctuation library construction module 20 takes the load signal features after multi-scale decomposition as input, performs feature extraction and characterization optimization, and constructs a feature fluctuation library to provide data support for the adaptive matching prediction module 30.
[0352] The adaptive matching prediction module 30 performs fine characterization of fluctuation patterns through adaptive density estimation and improved multi-level clustering strategy, constructs an efficient index structure, and makes predictions, while passing the prediction results to the anomaly handling module 40.
[0353] The anomaly handling module 40 inputs the prediction results and combines them with integrated isolated forest and local anomaly factors to optimize anomaly detection. It optimizes the online correction mechanism through adaptive window and robust estimation. The corrected prediction results are used for online monitoring of power grid safety margin.
[0354] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for predicting distribution transformer load based on adaptive decomposition of multi-scale fluctuation modes, characterized in that: include, Input the raw data, use the improved maximum overlap discrete wavelet packet transform to decompose the load signal at multiple scales, optimize the energy distribution, perform feature extraction and characterization optimization, and build a feature fluctuation library to store the features; By using adaptive density estimation and an improved multi-level clustering strategy, fluctuation patterns are finely characterized at different time scales, an efficient index structure is constructed, and prediction is performed through adaptive matching. By improving dynamic time warping, optimizing distance metrics and search strategies, and by dynamically adjusting weights through an improved particle swarm optimization algorithm, a combined prediction framework is constructed. By combining integrated isolated forests and local anomaly factors, anomaly detection is optimized, and the online correction mechanism is optimized through adaptive windowing and robust estimation. The improved maximum overlap discrete wavelet packet transform for multi-scale decomposition of the load signal includes: using the symmetric symmetric sym8 wavelet basis function as the basic decomposition tool; setting the support to 16; selecting the optimal wavelet basis by minimizing the reconstruction error; and using the error function... for: , Where x is the original signal, For reconstructing the signal; Set the reconstruction error threshold to When the reconstruction error is less than At that time, it was considered that the wavelet basis selection was appropriate; At the same time, the wavelet basis function must satisfy three key properties: the first key property is orthogonality, the second key property is compact support, and the third key property is vanishing moment. An adaptive extension algorithm is used to address the endpoint effect problem in load data. An AR(p) autoregressive model is established to describe the temporal correlation of the load data. The model order p is dynamically determined by minimizing the AIC criterion. The extension length L is adaptively determined based on the data length N and the signal-to-noise ratio (SNR). Denoising is performed using an improved soft-thresholding contraction method. The threshold parameter... A scale-dependent adaptive approach is adopted, which uses the median absolute deviation method for robust estimation, optimizes the threshold selection using the SURE criterion, and determines the optimal threshold by minimizing the unbiased risk estimate. The optimized energy distribution includes, for the wavelet decomposition coefficients of the j-th layer, calculating the wavelet energy proportion of the j-th layer, setting an adaptive energy threshold, and calculating the mutual information between adjacent decomposition layers through a constraint mechanism based on mutual information. When the mutual information value exceeds the threshold... When the information gain is 0.1, it indicates information redundancy, requiring hierarchical adjustment. An improved information gain calculation method dynamically determines the optimal hierarchy, and a distance function is defined using a hierarchy merging strategy based on dynamic programming. ,when Merging at that time The cross-validation method was used to determine that, among which, This represents the distance between adjacent floors.
2. The distribution transformer load forecasting method based on multi-scale fluctuation mode adaptive decomposition as described in claim 1, characterized in that: The feature extraction and characterization optimization includes performing in-depth mining and enhancement processing on distribution transformer load data to construct a feature fluctuation library; The load data includes an amplitude feature family and a shape feature family. The amplitude feature family includes maximum amplitude, root mean square amplitude, peak factor, margin factor, and waveform factor. The shape feature family includes skewness, kurtosis, waveform index, and impulse index. The frequency domain characteristics are enhanced by improving the power spectral density estimation method. The multi-window spectrum estimation technique is adopted, and the final power spectral density is obtained by weighted superposition of multiple independent spectrum estimation results. The frequency band energy features extracted by energy distribution characterization include sub-band energy ratio, energy entropy, spectral centroid, and spectral broadening. The sub-band energy ratio includes calculating the proportion of energy in each sub-band to the total energy. The spectral centroid is obtained by weighted averaging of energy to obtain the center position of the energy distribution. The spectral broadening is obtained by calculating the second moment of the frequency relative to the centroid.
3. The distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition as described in claim 2, characterized in that: The improved multi-level clustering strategy using adaptive density estimation includes characterizing data distribution features through improved kernel density estimation, selecting the Epanechnikov kernel function, using an adaptive method to select bandwidth, and determining the density threshold using a dynamic method based on the silhouette coefficient. After characterization, an improved multi-level clustering strategy is used to identify fluctuation patterns at different time scales. Based on the inherent laws of distribution network load changes, the time scale is divided into four levels: intraday, daily, weekly, and monthly. The intraday layer covers fluctuation characteristics from 15 minutes to 4 hours, reflecting users' short-term electricity consumption behavior; the intraday layer covers variation patterns from 4 hours to 24 hours, reflecting the characteristics of the daily load curve. The weekly span is 1 to 7 days, depicting the difference between weekdays and weekends; the monthly range is 7 to 30 days, capturing seasonal variation characteristics. Within each timescale layer, an improved density clustering algorithm is used for wave pattern identification, and a density-sensitive factor is introduced during the boundary point absorption process. The criteria for determining the ownership of boundary points are as follows: , in, For point With point The distance between them and These are the corresponding density values. Let be the neighborhood radius, where ,when When the value is less than 0.35, a relatively dense cluster is formed; when... When the value is greater than 0.35, points with large density differences are allowed to be grouped into the same category; The refined characterization of the fluctuation pattern includes probabilistic distribution modeling using an improved Gaussian mixture model and parameter estimation using an optimized expectation-maximization algorithm. , in, For conditional expectation, For model parameters, The parameters for the current iteration. Let be the likelihood function containing latent variables. For expectation operators; In each iteration, the parameters are dynamically adjusted and the step size is updated based on the gradient change of the likelihood function through an adaptive learning rate adjustment mechanism. Simultaneously, the optimal number of components K is adaptively determined using the Bayesian information criterion. , , Among them, BIC stands for Bayesian Information Criterion. Here, M represents the log-likelihood value, and M represents the sample size. Meanwhile, the nonlinear correlation characteristics in the load data of power distribution equipment are characterized using the copula function framework: , in, This represents the cumulative distribution function value of the d-th random variable. Denotes the Copula joint distribution function. The cumulative distribution function value represents the marginal distribution. Represents a random variable. The dimension is represented; the joint distribution of the multidimensional random variables is decomposed into two parts, marginal distribution and dependency structure, by using the copula function. Based on the characteristics of the load data, t-copula is selected as the basic dependency structure, and the relevant parameters are estimated by the maximum likelihood method. At the same time, an adaptive strategy is adopted for the selection of the marginal distribution. Multidimensional features are extracted by characterizing the nonlinear correlation features in the load data of power distribution equipment, including the construction of statistical feature systems and morphological feature systems. The statistical feature system includes mean, variance, skewness, and kurtosis. The mean is calculated... The variance reflects the central tendency of the load level through Describe the dispersion of load fluctuations and calculate the skewness. and the kurtosis Characterizes the asymmetry and peak characteristics of load distribution; The morphological feature system includes spikes in the opening operation suppression signal, valleys in the closing operation filling signal, and abrupt changes in the morphological gradient detection load curve.
4. The distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition as described in claim 3, characterized in that: The construction of an efficient index structure includes building a multi-layered hybrid index and retrieval mechanism and supporting fast fluctuation pattern retrieval, including a multi-layered hybrid index, a feature dimension index, and fast retrieval supporting fluctuation patterns; The multi-layer hybrid index constructs a time-dimensional index structure based on a B+ tree, which organizes the timestamps of the load data as keys. Each non-leaf node stores the boundary values of the time interval, and the leaf nodes are linked in chronological order. At the same time, a caching mechanism is built in the B+ tree to cache the index information of hot time intervals in memory, and a pre-read strategy is used to preload data of adjacent time periods. The feature dimension index employs a hybrid indexing strategy combining locality-sensitive hashing and KD-trees to index the feature vectors. Construct a hash function: , Where a is a random vector and b is a random offset. Using the bucket width parameter, a KD tree is used to partition the feature space. At each node, the dimension with the largest variance is selected as the partition axis to perform adaptive segmentation of the feature space. The fast retrieval of the fluctuation pattern is achieved through an index structure based on an inverted linked list. Each fluctuation pattern is assigned a unique ID, and a mapping relationship between the pattern ID and the feature vector is established through the inverted linked list. Considering the high-dimensionality of the feature vector, a compression storage strategy is adopted, including feature quantization and sparse representation, to reduce storage overhead. At the same time, the inverted linked list adopts a block storage method to store patterns with similar features in adjacent physical spaces. Post-indexing retrieval mechanism optimization includes multi-path parallel retrieval strategies and multi-level pruning strategies; The multi-path parallel retrieval strategy filters data in the time dimension, locates data within the target time window, performs similarity matching in the feature space, finds candidate similar fluctuation patterns through the combination of LSH and KD trees for pattern verification, and determines the final matching result by calculating a precise similarity metric. The multi-level pruning strategy includes: distance-based pruning, which quickly eliminates regions that cannot contain the target pattern by calculating the minimum distance between the query vector and the feature space partition; density-based pruning, which prioritizes high-density regions and skips sparse regions by utilizing the density distribution characteristics of fluctuating patterns; and similarity-threshold-based pruning, which sets a dynamic similarity threshold to terminate the verification of candidate patterns that do not meet the requirements.
5. The distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition as described in claim 4, characterized in that: The improved dynamic time warping includes constructing an efficient index structure and performing adaptive matching prediction, and optimizing distance metrics and search strategies through an improved dynamic time warping algorithm. The optimized distance metric includes constructing a feature distance vector using a multi-feature fusion distance calculation method. , where each component Representing distance metrics across different feature dimensions, an adaptive weight allocation mechanism based on entropy weighting is used for features. and j, information entropy Calculated as , in, Calculate the feature weights for the normalized feature values: , in, Let i be the weight of the i-th indicator. Let i be the entropy value of the i-th index, while ensuring , The final fusion distance is expressed as ; Simultaneously, the path constraints of DTW are optimized, and bandwidth parameters are adjusted. Dynamically adjust based on the periodicity of the sequence: , in, For sequence length, The periodicity intensity of the sequence, and To adjust the parameters, and simultaneously through slope constraints The degree of inclination of the control path, among which, Through cross-validation Optimize selection within the range; The search strategy employs an improved A* heuristic algorithm for optimal path search, defining nodes. Evaluation function of ode : , , in, Indicates the distance from the starting point to the node. The actual cumulative distance, For the node The estimated distance to the destination is determined by a heuristic function based on a lower bound estimate: , in, For heuristic distance estimation from node n to the target, For the k-th element of the target sequence, For the i-th element of the reference sequence, For the search window width, Sequence length; Furthermore, a multi-level pruning strategy is employed to utilize the lower bound of LB_Keogh on the sequence. Constructing the envelope and Perform fast filtering when candidate sequences If the distance to the envelope exceeds the current optimal value, the branch is pruned directly using the Early Abandoning strategy. During the cumulative distance calculation, the calculation is stopped immediately when the threshold is exceeded. The lower bound estimate is gradually tightened through the Cascading lower bounds mechanism, and multi-level filtering from coarse to fine is performed.
6. The distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition as described in claim 5, characterized in that: The anomaly detection optimization includes employing an improved ensemble isolated forest algorithm for anomaly detection. This is achieved by optimizing the feature space partitioning strategy and combining it with an information gain-based feature selection mechanism. For each node's feature selection, the information gain value of each feature is calculated. : , in, Let the information entropy of dataset D be , Given the conditional entropy under feature f, the feature with the maximum information gain is selected as the splitting feature to capture abnormal patterns in the load data of power distribution equipment. The split threshold is selected by a dynamic split threshold determination method. Based on the distribution characteristics of the current node data, the probability density function of the data distribution is obtained by kernel density estimation method, and the optimal split point is selected from the local minimum points of the density function. The computational accuracy is improved by using an improved local anomaly factor algorithm, and the distance metric is optimized by using generalized Minkowski distance. In the calculation of local anomaly factors, the reach-distance is optimized to take into account the local density information of sample points. In the density estimation stage, an adaptive bandwidth selection mechanism is combined, and the bandwidth parameter is determined by minimizing the mean square error criterion. The improved isolation forest and local anomaly factor algorithms are integrated, and the detection results of the two algorithms are fused by weighted voting. The weight values are dynamically updated through online learning. The optimized online correction mechanism includes adaptive window design and robust estimation optimization; The adaptive window design establishes a base window length of 24 hours × 4 samples per hour × 7 days, i.e., L = 672 sampling points. The window length is adjusted based on prediction error feedback. When the continuous prediction error exceeds a set threshold, the window length is adjusted accordingly. Adjustments are made, where δ is the dynamic adjustment coefficient, which is positively correlated with the prediction error; Optimize window data weighting using an exponential decay weighting strategy: , Where α is the forgetting factor and i is the time sequence number of the data point; the forgetting factor α adopts an adaptive adjustment mechanism, with an initial value of 0.1, and is dynamically adjusted according to the changes in prediction error; when the prediction error increases, the value of α is increased to accelerate the forgetting of historical data; when the prediction error decreases, the value of α is decreased to maintain model stability, and the adjustment range of α is limited to [0.05, 0.3]. The robust estimation optimization includes improving the M-estimation method and enhancing the S-estimation method. Outliers are handled by improving the M-estimation method and combining it with the Huber loss function. The enhanced S-estimation method optimizes the upper limit of the outlier ratio tolerated by the estimator. Considering the characteristics of distribution network load data, the breakpoint is set as: , Where n is the number of samples and p is the number of parameters, an improved weight function is used during the weight iteration process. ,in, Let be the residual, and s be the scaling estimate. Also, during the iteration process, when the Euclidean distance between two adjacent parameter estimates is less than... Convergence is determined at that time.
7. A system employing a distribution transformer load forecasting method based on multi-scale fluctuation mode adaptive decomposition as described in any one of claims 1 to 6, characterized in that: It includes a multi-scale fluctuation decomposition module, a feature fluctuation library construction module, an adaptive matching prediction module, and an anomaly handling module; The multi-scale fluctuation decomposition module takes the original load data as input, uses the improved maximum overlap discrete wavelet packet transform to decompose the load signal at multiple scales, optimizes the energy distribution, and transmits the characteristics of the decomposed load signal to the feature fluctuation library construction module. The feature fluctuation library construction module takes the load signal features after multi-scale decomposition as input, performs feature extraction and characterization optimization, and constructs a feature fluctuation library to provide data support for the adaptive matching prediction module. The adaptive matching prediction module performs fine characterization of fluctuation patterns through adaptive density estimation and improved multi-level clustering strategy, constructs an efficient index structure, and makes predictions, while passing the prediction results to the anomaly handling module. The anomaly handling module inputs the prediction results and combines them with integrated isolated forest and local anomaly factors for anomaly detection optimization. It optimizes the online correction mechanism through adaptive window and robust estimation. The corrected prediction results are used for online monitoring of power grid safety margin.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the distribution transformer load prediction method based on multi-scale fluctuation mode adaptive decomposition as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Load prediction method based on discrete wavelet transformation and optimization of a minimum support vector machine
CN109871977A
Hierarchical fault positioning method and system for distributed power distribution network
CN118707246A