Transformer load rate prediction method based on data mining

By collecting real-time electrical parameters from transformers and users, classifying user nodes, generating a list of flexible loads, and combining this with overload risk prediction, the problem of indiscriminate response in transformer load management is solved, achieving precise demand response and optimized equipment utilization.

CN121834465AInactive Publication Date: 2026-04-10无锡则安电力科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
无锡则安电力科技有限公司
Filing Date
2026-01-12
Publication Date
2026-04-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing industrial data mining technologies struggle to differentiate between users' willingness and ability to respond and adjust during power load analysis. This results in indiscriminate power rationing or broad-based invitation responses during peak load periods, making it impossible to effectively manage transformer loads.

Method used

By collecting real-time electrical parameters and equipment status of smart meters on the low-voltage side of the transformer and at the user end, a time-series electrical status parameter set is generated, user electricity consumption behavior characteristics are calculated, adjustable and rigid nodes are classified, a list of flexible load users is generated, and a demand response dynamic pricing strategy is generated in combination with transformer overload risk prediction.

Benefits of technology

It enables precise management of transformer load, uses price leverage to guide highly sensitive users to stagger their electricity consumption during off-peak hours, ensures grid security, maximizes equipment capacity utilization, and reduces equipment overload risk and economic costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834465A_ABST
    Figure CN121834465A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial data mining, in particular to a transformer load rate prediction method based on data mining, which comprises the following steps of: acquiring real-time voltage, current and active power readings according to sensing signals of intelligent electric meters at a low-voltage side of a transformer and at each user side in a power supply area; synchronously obtaining state values indicated by the transformer winding temperature and the oil conservator oil level, and generating a time sequence electrical state parameter set. According to the method, the expected pressure of a transformer insulation system and a conductive part under a specific working condition can be restored more truly, a matched demand response dynamic pricing strategy can be generated based on an overload risk prediction value, and a price lever is utilized to guide a high-sensitivity user to actively shift peaks; the capacity of existing equipment is utilized to the maximum extent while the operation safety of a power grid is guaranteed, and the economic cost caused by blind capacity expansion and the thermal aging risk caused by long-term overload operation of the equipment are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial data mining, and in particular to a transformer load rate prediction method based on data mining. BACKGROUND

[0002] Industrial data mining technology is to use statistics, artificial intelligence, machine learning and database technology to automatically extract hidden, unknown and potentially valuable knowledge and patterns from massive, high-dimensional and heterogeneous industrial production data.

[0003] The existing industrial data mining technology in the field of power load analysis often focuses on the historical trend extrapolation of macro total quantity or the statistical analysis of single dimension electrical values. All users in the transformer power supply area are usually regarded as a homogeneous load as a whole, and it is difficult to distinguish which users have the willingness and ability to respond to the adjustment instructions, resulting in indiscriminate power cut or extensive invitation response when facing peak load pressure. Therefore, improvement is needed. SUMMARY

[0004] The purpose of the present application is to solve the shortcomings in the prior art, and a transformer load rate prediction method based on data mining is proposed.

[0005] In order to achieve the above-mentioned purpose, the present application adopts the following technical scheme, a transformer load rate prediction method based on data mining, comprising the following steps: According to the sensing signals of the transformer low-voltage side and the intelligent electric meters of each user terminal in the power supply area, the real-time voltage, current and active power readings are collected, the state values of the transformer winding temperature and the oil level of the oil pillow are synchronously obtained, the time sequence electrical state parameter set is generated, the power usage standard deviation and the peak-valley difference ratio of each user node in the set period are calculated according to the time sequence electrical state parameter set, and the user electricity behavior feature data set is quantitatively generated; Each data point in the user electricity behavior feature data set is mapped to a multi-dimensional feature vector coordinate system, the distance between each feature vector point in the coordinate system is calculated, the behavior cluster is filtered according to the distance, the feature vector space distribution set is generated, the adjustable node and the rigid node are classified according to the feature vector space distribution set, the device ID information of the adjustable node is extracted, and the elastic load user identity recognition list is generated; The current real-time load amplitude of each user in the elastic load user identity recognition list is extracted, the user expected load value is calculated and obtained, the user expected load value is combined with the transformer total expected load value, and the transformer overload risk prediction value is generated; According to the transformer overload risk prediction value, a demand response dynamic pricing strategy file is generated.

[0006] Preferably, the user electricity behavior feature data set acquisition step is: Based on the collected voltage, current and active power values, they are matched with the transformer winding temperature and oil conservator level indication values ​​according to a unified timestamp. The power readings of each user node within a set period are serialized, the power usage standard deviation and peak-to-valley ratio are calculated, and the data are bound to the user identifier and time identifier to generate a time-series electrical state parameter set. Calculate the load sensitivity based on the set of time-series electrical state parameters; Based on the load sensitivity, the load sensitivity of each user node is aggregated with the user identifier and time identifier, and then arranged into continuous records in chronological order to generate a user electricity consumption behavior feature dataset.

[0007] Preferably, the step of obtaining the feature vector space distribution set is as follows: Based on the user electricity consumption behavior feature dataset, each data point is mapped to a multi-dimensional feature vector coordinate system according to a unified time index. Covariance analysis is performed on the feature vectors of each pair of user nodes and the correlation between features is corrected. The weighted Mahalanobis distance is calculated based on the linear correlation structure between features. Based on the weighted Mahalanobis distance, node pairs with a density less than a preset threshold are selected to establish a connection relationship. All sets of interconnected nodes are defined as behavioral clusters. The center vector, dispersion, and number of internal members of each behavioral cluster are counted to generate a set of feature vector spatial distributions.

[0008] Preferably, the step of obtaining the resilient load user identification list is as follows: Based on the set of feature vector spatial distributions, calculate the adjustable potential index for each behavioral cluster; Based on the comparison between the adjustable potential index and the preset threshold, when the adjustable potential index is greater than the preset threshold, the cluster is marked as an adjustable node, and when the adjustable potential index is less than the preset threshold, it is marked as a rigid node. The device ID information corresponding to all adjustable nodes is extracted to generate a list of elastic load user identities.

[0009] Preferably, the step of obtaining the user's expected load value is as follows: Based on the flexible load user identification list, the real-time load amplitude of each user is read one by one. The current ambient temperature and the operating status of the transformer cooling system are synchronized according to the timestamp. The temperature reading is extracted, the operating status indication value is extracted from the cooling control unit, and input into the adjustment coefficient mapping table to retrieve the corresponding adjustment coefficient. The real-time load amplitude of each user is multiplied by the adjustment coefficient to obtain the corrected load value, and recorded with user identifier and time index to generate the user's expected load value.

[0010] Preferably, the step of obtaining the transformer overload risk prediction value is as follows: Based on the user's expected load value, the user group is divided according to the future time period. The expected load values ​​of users in the flexible load user identification list are summed by time period to obtain the total load value of the flexible user time period. At the same time, the baseline load values ​​of non-listed users are summed by the same time period to obtain the total load value of the non-listed user time period. The total load value of the flexible user time period and the total load value of the non-listed user time period are added together to obtain the total expected load value of the transformer in the current time period. Based on the total expected load value of the transformer in each time period, iterate through all the total expected load value sequences in the future time period, select the maximum value and define it as the peak load prediction value of the transformer, compare the peak load prediction value with the rated capacity on the transformer nameplate, calculate the difference ratio, and look up the corresponding risk level in the overload risk judgment table according to the difference ratio to generate the transformer overload risk prediction value.

[0011] Preferably, the steps for obtaining the demand response dynamic pricing strategy file are as follows: Based on the predicted value of transformer overload risk, extract the difference ratio and time index for each future period, merge adjacent risk periods according to the operation and maintenance calendar, look up the table to map the difference ratio to the response level, mark the triggering method and applicable power supply area, and generate a list of demand response triggering periods. Based on the demand response triggering period list, the current time-of-use electricity price table is read for each triggering period, and the price adjustment level, start and end time, user scope and settlement method are bound according to the response level. The billing accuracy and minimum billing duration limits are supplemented, and a time-of-use price adjustment list is generated.

[0012] Preferably, the step of obtaining the demand response dynamic pricing strategy file further includes: Based on the time-of-use pricing adjustment list, compile the strategy number, version number, generation time, applicable power supply area, validity period, rollback rules, notification template, approver list, and execution monitoring items, write them into structured text in the order of fields, and attach the effective conditions to generate a demand response dynamic pricing strategy file.

[0013] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This invention integrates physical sensing signals from the low-voltage side of the transformer with smart meter readings at the user end, simultaneously collecting real-time electrical parameters and key status indicators such as equipment winding temperature and oil level. It constructs a time-series electrical status parameter set including the standard deviation of power usage and the peak-to-valley ratio, enabling a deep correlation between load fluctuations and equipment health status from the data level. Multidimensional features are mapped to a vector coordinate system and spatial distances are calculated to generate behavioral clusters. Based on the cluster distribution characteristics, the dispersion of user electricity consumption behavior is quantified, thus separating flexible load nodes with adjustment potential from rigid nodes that cannot be adjusted, avoiding a "one-size-fits-all" load management approach. Furthermore, by introducing adjustment coefficients based on current ambient temperature and cooling system operating status, the real-time load amplitude of flexible users is dynamically corrected and superimposed. This not only more accurately reflects the expected pressure on the transformer insulation system and conductive components under specific operating conditions but also generates a matching demand response dynamic pricing strategy based on overload risk prediction values. This leverages price to guide highly sensitive users to proactively stagger peak loads, maximizing the utilization of existing equipment capacity while ensuring grid operation safety, reducing the economic costs of blind expansion and the risk of thermal aging caused by long-term overload operation. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the steps of the present invention. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0016] Please see Figure 1 This invention provides a technical solution: a transformer load rate prediction method based on data mining, comprising the following steps: Based on the sensing signals of smart meters at each user end in the low-voltage side of the transformer and the power supply area, real-time voltage, current and active power readings are collected, and the status values ​​of transformer winding temperature and oil level indication are obtained synchronously to generate a time-series electrical state parameter set. Based on the time-series electrical state parameter set, the standard deviation of power usage and peak-valley ratio of each user node in the set period are calculated, and a user electricity consumption behavior feature dataset is generated in a quantitative manner. Map each data point in the user electricity consumption behavior feature dataset to a multi-dimensional feature vector coordinate system, calculate the distance between each feature vector point in the coordinate system, filter behavior clusters based on the distance, generate a feature vector spatial distribution set, classify adjustable nodes and rigid nodes based on the feature vector spatial distribution set, extract the device ID information of adjustable nodes, and generate a list of resilient load user identification. Extract the current real-time load amplitude of each user in the flexible load user identification list, calculate the expected load value of the user, and combine the expected load value of the user with the total expected load value of the transformer to generate the transformer overload risk prediction value. Based on the predicted values ​​of transformer overload risk, a dynamic demand response pricing strategy document is generated.

[0017] The steps to obtain the user electricity consumption behavior feature dataset are as follows: Based on the collected voltage, current and active power values, they are matched with the transformer winding temperature and oil conservator level indication values ​​according to a unified timestamp. The power readings of each user node within a set period are serialized, the power usage standard deviation and peak-to-valley ratio are calculated, and the data are bound to the user identifier and time identifier to generate a time-series electrical state parameter set. Based on the time-series electrical state parameter set, the load sensitivity is calculated using the following formula: ; in, For the first Load sensitivity index for individual users For the first Standard deviation of power usage within a user-defined period For the first The average active power within a user-defined period is obtained by averaging the power values ​​for the same period. For the first Maximum active power within a user cycle For the first Minimum active power per user cycle This is a fluctuation sensitivity index used to measure the relative impact of power fluctuation terms. This is the peak-valley sensitivity index, used to measure the relative impact of peak-valley differences. Based on load sensitivity, the load sensitivity of each user node is aggregated with user identifier and time identifier, and then arranged into continuous records in chronological order to generate a user electricity consumption behavior feature dataset.

[0018] Specifically, based on the collected voltage, current, and active power values, as well as transformer winding temperature and oil level data, a unified time reference axis is first established, and the sampling period is set. The sampling interval is 24 hours. For a 15-minute interval, high-frequency electrical data from low-voltage side smart meters and low-frequency status data from transformer monitoring units are time-domain aligned. Missing time points are filled in using cubic spline interpolation to ensure all variables have corresponding values ​​at the same timestamp. Then, the power readings of each user node within a set period are serialized to form a 96-bit time series vector. For this power series, the standard deviation of power usage is obtained by calculating the square root of the sum of the squares of the differences between each point in the series and the mean. Simultaneously, the maximum value is found by iterating through the power series. and minimum value The difference between the two values ​​is calculated and divided by the average power within the period to obtain the peak-to-valley ratio. During this process, if the average power in the denominator is zero, it is set to the minimum non-zero measurement value of 0.001kW allowed by the equipment to avoid calculation errors. Then, the calculated standard deviation and peak-to-valley ratio are associated and bound with the corresponding user unique identification code (UID) and the current processing time window ID to construct a structured record containing electrical characteristics and identity information. Finally, all processed records are summarized to generate a time-series electrical state parameter set.

[0019] In the load sensitivity calculation formula, by introducing the fluctuation sensitivity index and the peak-valley sensitivity index, the random fluctuation of user load (characterized by the standard deviation term) and the extreme variation range (characterized by the peak-valley difference term) are nonlinearly weighted and combined to construct a dimensionless load sensitivity index. This comprehensively quantifies the impact of user electricity consumption behavior on transformer stability, making users with different power bases comparable under the same evaluation system.

[0020] For the first The standard deviation of power usage within a user's defined period is a parameter that reflects the dispersion of a user's power consumption from its average level within a statistical period, and is a core indicator for measuring load fluctuation stability. The steps for obtaining this parameter are as follows: First, data is collected from the smart meter... Individual users within a set period (e.g., 24 hours) Data sequence of power sampling points ,in (Sampling at 15-minute intervals); then calculate the average value of the sequence. Then calculate each sampling point. Compared with the average The square of the difference, for this Sum the differences of squares and divide by . (or Finally, take the square root of the result to get... The calculation formula is: For example, a user's power value over 5 sampling periods is... If the power is kW and the average power is 4.0 kW, then the variance is... Therefore, the standard deviation kW.

[0021] For the first The average active power within a user-defined period is used to normalize fluctuations and eliminate the impact of differences in users' basic electricity consumption on sensitivity evaluation. The acquisition steps are as follows: Read the... The active power values ​​of each user at all valid sampling points within a set period are summed and then divided by the total number of sampling points. If data is missing, only the valid data points are averaged. Continuing with the example above, the sampling points are... ,but kW.

[0022] For the first Maximum active power within a user cycle For the first The minimum active power within each user cycle, along with the two parameters, constitutes the extreme range of load variation, reflecting the maximum range of transformer capacity occupancy by users. The steps to obtain this are: traversing the [number]th [user cycle] using bubble sort or the built-in extreme value search algorithm... Power sequence of individual users within a set period Select the largest value as Select the smallest value as In the example above, kW, kW.

[0023] This is a fluctuation sensitivity index used to measure the relative weight of the power fluctuation term (standard deviation) in the overall sensitivity evaluation, reflecting the grid's tolerance to high-frequency random fluctuations. The acquisition steps are as follows: Based on historical operating data of the power supply area, extract the standard deviation data of all user loads corresponding to periods of transformer overheating or voltage flicker, establish a regression analysis model of "load standard deviation - transformer loss rate," and obtain the exponential part of the power-law relationship through least squares fitting. The specific empirical method setting process is as follows: During full-load operation testing of the transformer, monitor the winding temperature rise rate at different standard deviation levels, and fit the curve using multiple sets of experimental data. (in (where is the coefficient of variation), and the best-fit index is calculated. In this example, based on the age of the transformers in the area and their sensitivity to fluctuations, the optimal fit index is set after calculation. This indicates that the fluctuation term has a high nonlinear effect on the final sensitivity.

[0024] The peak-valley sensitivity index measures the relative weight of the peak-valley difference in the overall sensitivity evaluation, reflecting the power grid's level of attention to significant load peak shaving and valley filling demands. Its acquisition steps are as follows: based on the overload capacity curve in the transformer design specifications, analyze the impact weight of the peak-valley difference on insulation life loss, and determine this in conjunction with the incentive orientation for peak shaving and valley filling in the demand response strategy. Specifically, referring to historical peak-shaving records, calculate the contribution of the peak-valley difference change rate to the transformer load rate variance, and determine its weighting factor through principal component analysis (PCA). In this example, to balance the impact of the fluctuation term and considering the current seasonal peak-shaving demand, the index is calculated and set... .

[0025] Calculations based on parameters: Select a specific user Calculations were performed based on electricity consumption data for a certain day, with the following parameters known: kW (calculated from the example above); kW; kW; kW; ; .

[0026] The first step is to calculate the coefficient of variation (fluctuation term): ; ; The second step is to calculate the relative peak-to-valley difference term: ; ; The third step is to calculate the final load sensitivity index. : ; This user The load sensitivity index is 0.536. Because... This is a normalized index that integrates volatility and peak-valley difference. This value quantifies the potential stress on the transformer caused by the user's current electricity consumption pattern. If the standard baseline for this area is set at 0.5, then the user's value of 0.536 is slightly higher than the baseline value, which means that its load fluctuation is relatively large compared to its average electricity consumption, or its peak-valley difference is more obvious. It belongs to a "sensitive" user that needs to be closely monitored or has certain adjustment potential. In the subsequent demand response pricing strategy, this user may be allocated a more aggressive dynamic electricity price to guide it to smooth out the volatility.

[0027] Based on load sensitivity, a hash mapping table structure with timestamps as keys is first established to temporarily store the calculation results. Then, the load sensitivity value of each user node after calculation is retrieved. Extract the corresponding user unique identifier (UID) and the time window identifier (TID) from which the data was generated. A triplet data packet is formed by UID and TID. To ensure the temporal continuity of the data, all triplet data packets are traversed to check for time breakpoints. For any detected missing time points, if data exists before and after the time point, linear interpolation is used to fill in the missing records. If more than three consecutive time points are missing, the time period is marked as invalid and not filled. Subsequently, all triplets are quickly sorted in ascending order of time identifier TID. Based on the time dimension alignment, a secondary sort is performed by user identifier UID, so that the behavioral characteristics of the same user at different time points are continuously distributed in the physical storage space. Then, the sorted data is formatted and transformed, mapping each triplet to a row in a feature matrix, where the columns correspond to time, user ID, and sensitivity index, respectively. Finally, metadata header information is added to the dataset, including the data generation time, the total number of users covered, and the time span. The organized memory data blocks are written to persistent storage media to generate a user electricity consumption behavior feature dataset.

[0028] The steps to obtain the feature vector space distribution set are as follows: Based on the user electricity consumption behavior feature dataset, each data point is mapped to a multi-dimensional feature vector coordinate system according to a unified time index. Covariance analysis is performed on the feature vectors of each pair of user nodes, and the correlation between features is corrected. The weighted Mahalanobis distance is calculated based on the linear correlation structure between features. The calculation formula is as follows: ; in, Let be the weighted Mahalanobis distance between user node a and user node b. Let be the K-dimensional feature vector of the a-th user node, where each component corresponds to the quantized value of the node in each feature dimension. Let be the K-dimensional feature vector of the b-th user node. This is the inverse of the covariance matrix of the feature vectors of all user nodes, used to eliminate linear correlation between features. Given a K×K diagonal weight matrix, its k-th diagonal element... This represents the importance coefficient of the k-th feature dimension. Let K be the transpose of the eigenvector differences, where K is the total number of feature dimensions. Based on the weighted Mahalanobis distance, node pairs with a density less than a preset threshold are selected to establish connections. All interconnected node sets are defined as behavioral clusters. The center vector, dispersion, and number of internal members of each behavioral cluster are statistically analyzed to generate a set of feature vector spatial distributions.

[0029] Specifically, the weighted Mahalanobis distance calculation formula introduces a weight matrix. This allows the model to adjust the contribution of different feature dimensions according to business needs, thereby quantifying the similarity of electricity consumption behavior among users.

[0030] Let be the K-dimensional feature vector of the a-th user node. Each component corresponds sequentially to the "load sensitivity index" obtained in the preceding steps. "Standard deviation of power usage" "and peak-to-valley ratio" The acquisition steps are as follows: Read the profile data of the a-th user and extract the dimensionless load sensitivity. Standard deviation of power in kW and dimensionless peak-to-valley ratio , forming column vectors For example, if user a's data is extracted with the following parameters: sensitivity 0.536, standard deviation 1.414kW, peak-to-valley ratio 1.0, then... User B's data: sensitivity 0.600, standard deviation 1.500kW, peak-to-valley ratio 1.1, then... .

[0031] This is the inverse of the covariance matrix of the feature vectors of all user nodes, used to eliminate linear correlations between features and perform dimension normalization. The steps to obtain it are: select the feature vectors of all users within the power supply area to construct the matrix. Calculate the covariance matrix between features ,in The main diagonal elements correspond to , , The variance, with units of , are respectively , , .right Inverse is obtained Its main diagonal element unit becomes , , In this example, the covariance matrix is ​​calculated. After inverting the equation, we get... .

[0032] The weight matrix is ​​a K×K diagonal matrix, obtained by using the Analytic Hierarchy Process (AHP) to determine the weights and constructing a judgment matrix to compare the relative importance of features. In this example, load sensitivity is considered the most important, and the weight vector is set as follows: Construct a diagonal matrix .

[0033] Calculations based on parameters: The first step is to calculate the difference between eigenvectors. (Note the units): ; The second step is to calculate the weighted difference vector. : ; The third step is to calculate the quadratic form. : First calculate the intermediate vector : ; ; (unit ) ; Recalculate : ; (The result is a dimensionless numerical value).

[0034] Fourth step: Take the square root to get the distance: ; The weighted Mahalanobis distance between user a and user b is 0.318. Compressing the multidimensional feature differences into a standardized scalar, a smaller value indicates that the two users are very close in the weighted feature space, belonging to the same category of users with highly similar electricity consumption behaviors, and can be grouped into the same behavioral cluster.

[0035] Based on the weighted Mahalanobis distance, a symmetric matrix containing the distances between all user nodes is constructed. To automatically identify reasonable cluster density boundaries from massive amounts of data, the K-nearest neighbor distance graph method is used for adaptive parameter optimization. For each user node, the weighted Mahalanobis distance to its k-th nearest neighbor is calculated, where k is typically twice the feature dimension (k=6). All calculated k-distance values ​​are sorted in descending order and plotted as a two-dimensional curve. The difference method is used to identify "elbow points" where the curve slope changes abruptly, and the distance value corresponding to this point is read as a preset density threshold. For example, if the algorithm automatically identifies a distance abrupt change point as 0.45, then the density threshold is set to 0.45. Subsequently, a density-based DBSCAN clustering process is executed, starting from any unvisited node... Starting from a node, retrieve all its neighboring nodes within a radius of 0.45. If the number of nodes in the neighborhood meets the minimum cluster size requirement (e.g., set to 3), create a new cluster and mark these nodes as core members. Expand the cluster range by continuously absorbing densely connected nodes through breadth-first search until it can no longer be expanded. Define all interconnected node sets as behavioral clusters. Remove isolated points that cannot be classified as noise data. For each generated cluster, traverse the feature vectors of its members, calculate the average to obtain the center vector, calculate the average Euclidean distance of each member to the center as the dispersion, and count the total number of user IDs in the cluster. Encapsulate these statistical indicators in a structured way to generate a feature vector spatial distribution set.

[0036] The steps for obtaining the user identification list for elastic loads are as follows: Based on the feature vector space distribution set, the adjustable potential index of each behavioral cluster is calculated using the following formula: ; in, Let C be the adjustable potential index of the Cth behavioral cluster. The load sensitivity index of all user nodes within the Cth cluster. The standard deviation represents the degree of fluctuation in user response behavior within a cluster. The load sensitivity index of all user nodes within the Cth cluster. The average value, Let C be the number of user nodes contained in the Cth behavior cluster. This is a logarithmic penalty term for cluster size, used to balance the impact of cluster size on the regulation potential; Based on the comparison between the adjustable potential index and the preset threshold, when the adjustable potential index is greater than the preset threshold, the cluster is marked as an adjustable node, and when the adjustable potential index is less than the preset threshold, it is marked as a rigid node. The device ID information corresponding to all adjustable nodes is extracted to generate a list of elastic load user identities.

[0037] Specifically, in the formula for calculating the adjustable potential index, the coefficient of variation term is used... The dispersion of user load sensitivity within a cluster is quantified. Higher dispersion indicates stronger complementarity among users in terms of response time and capability. This is further analyzed using the natural logarithm term. Nonlinear rewards are applied to cluster size to avoid excessive weighting of ultra-large clusters, thereby selecting high-quality elastic load resources that have both a certain scale and internal adjustment complementarity.

[0038] The load sensitivity index of all user nodes within the Cth cluster. The standard deviation of the index is used to measure the variability of user behavior within a cluster. The steps to obtain it are as follows: extract all user IDs contained in the C-th cluster from the feature vector space distribution set, and then backtrack based on the ID index to obtain the load sensitivity index for each user. , forming a sequence Calculate the variance of the sequence and then take the square root. For example, a cluster has 5 users. The values ​​are respectively The mean is 0.55. Calculate the standard deviation. .

[0039] The load sensitivity index of all user nodes within the Cth cluster. The average value is used as a normalization factor. The steps to obtain this are: averaging the extracted sensitivity sequences. Here... .

[0040] This represents the number of user nodes contained in the Cth cluster, reflecting the absolute capacity of the aggregated resources. The steps to obtain this number are: directly count the number of user IDs in the list of the Cth cluster. In this example... .

[0041] This is a logarithmic penalty term for the cluster size, used to smooth out the marginal benefits of size growth. Its acquisition steps are: calculate... The natural logarithm of .

[0042] Calculations based on parameters: Using the above example data: ; ; ; The first step is to calculate the coefficient of variation (diversity factor): ; The second step is to calculate the scale factor: ; The third step is to calculate the potential index: ; The adjustability potential index of the Cth behavioral cluster is 0.550. This result comprehensively reflects the value of this 5-person group in participating in demand response. Although small in size, the significant differences in load sensitivity among its members (high CV values, with extreme differences between 0.4 and 0.85) indicate that they may have complementary peak-shifting habits in their electricity consumption, resulting in good adjustment flexibility after aggregation. If this index exceeds a preset threshold, the cluster will be identified as a resilient adjustment unit with high development value.

[0043] Based on the comparison between the adjustable potential index and the preset threshold, a reasonable preset threshold is first determined through backtracking analysis of historical operating data. The adjustable potential index of all behavioral clusters generated in the past quarter is collected, and a frequency distribution histogram of the potential index is constructed. Following the Pareto principle, the top 20% of the high-resolution segments of the cumulative frequency distribution curve are selected as the judgment threshold. For example, the preset threshold is calculated to be 0.50. Then, the classification and discrimination process is initiated, traversing all behavioral clusters generated at the current moment and reading the data for each cluster. The value is compared with 0.50. This indicates that the cluster has high aggregation response value, so it is marked as an "adjustable node" and given a high-priority label. If a node is marked as a "rigid node," it is included in the basic load management. Then, for all clusters marked as adjustable nodes, the user member list is parsed, the smart terminal physical address and device ID of each user are extracted, duplicate records are removed by deduplication algorithm, and these device IDs are associated and bound with the corresponding cluster ID, potential index and power supply area information. The clusters are sorted from high to low according to the potential index to generate a structured list of flexible load user identification.

[0044] The steps to obtain the user's expected load value are as follows: Based on the flexible load user identification list, the real-time load amplitude of each user is read one by one. The current ambient temperature and the operating status of the transformer cooling system are synchronized according to the timestamp. The temperature reading is extracted, the operating status indication value is extracted from the cooling control unit, and input into the adjustment coefficient mapping table to retrieve the corresponding adjustment coefficient. The real-time load amplitude of each user is multiplied by the adjustment coefficient to obtain the corrected load value, and recorded with user identifier and time index to generate the expected load value of the user.

[0045] Specifically, based on the flexible load user identification list, a data buffer channel for parallel processing is established. Real-time load amplitude data for each user node in the list is read sequentially. A time alignment algorithm is used to synchronously acquire the ambient temperature reading of the transformer's area and the operating status indicator value of the transformer's cooling system. The cooling status indicator value includes three status codes: natural cooling mode, low-speed air cooling mode, and high-speed forced oil circulation mode. Subsequently, a pre-built adjustment coefficient mapping table is accessed. This mapping table is generated by collecting historical operating data from the past three years for the transformer area, removing abnormal records from power outage periods, dividing the temperature into 2-degree Celsius intervals, and combining different cooling statuses to create multiple operating scenarios. Statistics for each scenario are then compiled. The average rate of change of user load in the next hour is used as the adjustment coefficient. For example, under the condition of 35 degrees Celsius and the cooling system is running at high speed, historical data shows that the user load is amplified by an average of 1.1 times. Therefore, the adjustment coefficient under this condition is set to 1.1. Based on the currently extracted real-time temperature and cooling status, the corresponding adjustment coefficient is extracted by index matching in the mapping table. The real-time load amplitude of the user is multiplied by the adjustment coefficient to obtain the corrected load value. This process effectively corrects the user load prediction deviation caused by environmental thermal effects and equipment heat dissipation capacity constraints. Finally, the corrected load value is bound to the corresponding user unique identifier and the current timestamp and appended to the result sequence in time sequence to generate the expected user load value.

[0046] The steps for obtaining the transformer overload risk prediction value are as follows: Based on the expected load value of users, users are divided into groups according to future time periods. The expected load values ​​of users in the flexible load user identification list are summed by time period to obtain the total load value of flexible users in the time period. At the same time, the baseline load values ​​of non-listed users are summed by the same time period to obtain the total load value of non-listed users in the time period. The total load value of flexible users in the time period and the total load value of non-listed users in the time period are added together to obtain the total expected load value of transformers in the current time period. Based on the total expected load value of the transformer in each time period, iterate through all the total expected load value sequences in the future time period, select the maximum value and define it as the peak load prediction value of the transformer. Compare the peak load prediction value with the rated capacity on the transformer nameplate, calculate the difference ratio, and look up the corresponding risk level in the overload risk judgment table according to the difference ratio to generate the transformer overload risk prediction value.

[0047] Specifically, based on the expected load values ​​of users, a 4-hour time window for future forecasting is set, and this window is divided into continuous discrete time periods with a 15-minute time step. For each future time period, firstly, all records belonging to the flexible load user identification list are selected from the generated expected load dataset. The modified load values ​​of these users are accumulated within the time period to obtain the total load value of the flexible user time period. At the same time, the remaining users in the power supply area not included in the identification list are identified. For these non-listed users, the feature matching method is used to select five sample days from the historical database that are similar to the current day's weather type and belong to the same workday type. The load records of these sample days in the same time period are extracted. After removing outliers exceeding twice the standard deviation of the mean, the average value of the remaining sample data is calculated as the baseline load of the non-listed user. The baseline load values ​​of all non-listed users are summed within the time period to obtain the total load value of the non-listed user time period. Finally, the total load value of the flexible user time period and the total load value of the non-listed user time period are algebraically added and merged to obtain the total load forecast result of the entire distribution area within the time period, thus obtaining the total expected load value of the transformer in the current time period.

[0048] Based on the total expected transformer load for each time period, an extreme value search algorithm is used to traverse the load data for all time periods within the future prediction window. The load point with the largest value is compared and extracted, and this is defined as the transformer peak load prediction value. Subsequently, the rated capacity parameter on the transformer's nameplate is read, and the difference between the peak load prediction value and the rated capacity is calculated. This difference is then divided by the rated capacity to obtain the difference ratio. Based on this difference ratio, the overload risk is classified and judged. A preset overload risk judgment table is called, and the threshold settings of this table are derived from the thermal aging life model of transformer insulation paper and relevant IEEE load guidelines. Then, by experimentally determining the inflection point of insulation life loss rate under different overload rates, a difference ratio less than or equal to 0.05 is set as the first-level safety threshold, representing normal operation within the measurement error range; a difference ratio greater than 0.05 and less than or equal to 0.20 is set as the second-level warning threshold, representing the existence of short-term recoverable overheating risk; a difference ratio greater than 0.20 is set as the third-level danger threshold, representing irreversible thermal damage to the insulation material. Based on the specific interval in which the calculated difference ratio value falls, the corresponding risk level number is searched and matched in the judgment table to generate the transformer overload risk prediction value.

[0049] The steps to obtain the demand response dynamic pricing strategy file are as follows: Based on the predicted values ​​of transformer overload risk, the difference ratio and time index for each future period are extracted, adjacent risk periods are merged according to the operation and maintenance calendar, the difference ratio is mapped to the response level by looking up the table, the triggering method and applicable power supply area are marked, and a list of demand response triggering periods is generated. Based on the list of demand response trigger periods, the current time-of-use electricity price table is read for each trigger period. The price adjustment level, start and end time, user scope and settlement method are bound according to the response level. The billing accuracy and minimum billing duration limits are supplemented to generate a time-of-use price adjustment list. Based on the time-of-use pricing adjustment list, compile the strategy number, version number, generation time, applicable power supply area, validity period, rollback rules, notification template, approver list, and execution monitoring items, write them into structured text in the order of fields, and attach the effective conditions to generate a demand response dynamic pricing strategy file.

[0050] Specifically, based on the transformer overload risk prediction values, the process first iterates through all risk prediction data points within the future time period, extracting the difference ratio value for each time step and the corresponding UTC timestamp index to construct a time series set of risk events. A time window merging algorithm is then used to aggregate the discrete risk points. A time continuity threshold of 30 minutes is set; that is, if the interval between two adjacent high-risk time points is less than or equal to 30 minutes, they are considered as the same continuous overload event and merged to form several independent risk periods. For each merged risk period, the weighted average of all difference ratios within that period is calculated as the intensity index of the event. Subsequently, a pre-set response level mapping table is called, whose threshold setting references the short-term overload of the transformer. The load capacity curve and insulation aging model are used to define, for example, a weighted difference ratio between 0.05 and 0.15 corresponds to a light overload, mapped to a level 3 response; a ratio between 0.15 and 0.25 corresponds to a moderate overload, mapped to a level 2 response; and a ratio exceeding 0.25 corresponds to a severe overload, mapped to a level 1 response. The triggering method is further labeled according to the determined response level. The level 3 response is configured as a "flexible invitation" mode, which pushes suggestions through the APP, while the level 1 and level 2 responses are configured as an "automatic command" mode, which directly issues control signals. At the same time, the power supply area code and administrative street information of the transformer are extracted from the geographic information system, and these attribute information are bound to the risk period to generate a list of demand response triggering periods.

[0051] Based on the list of demand response trigger periods, the system accesses the basic database of the electricity marketing system and retrieves the current time-of-use tariff table applicable to the current power supply area. This table contains the basic electricity price rates for peak, flat, and valley periods. For each trigger period in the list, a corresponding price adjustment strategy is matched according to its response level. The adjustment coefficient in this strategy is obtained through historical demand price elasticity analysis. For example, the price increase coefficient for a level 3 response is set to 1.2, meaning a 20% increase in electricity price; the increase coefficient for a level 2 response is 1.5; and the increase coefficient for a level 1 response is 2.0. Alternatively, a specific peak-shaving subsidy unit price can be set, and a multiplication operation is performed to calculate the target execution price for that period. The billing process is then refined to two decimal places for billing precision. The start and end times for price adjustments are then determined. To ensure metering accuracy and equipment stability, a minimum billing duration constraint check is performed, setting the minimum execution window to one hour. If the trigger period is less than one hour, the timeframe is extended forward and backward from the center of that period to a full hour. The start and end times are aligned to the nearest 15-minute mark. The user range applicable to this pricing strategy is defined, and the previously generated flexible load user identification list is directly loaded. The settlement method prioritizes high-frequency frozen data from smart meters. Rates, times, users, and settlement rules are encapsulated to generate a time-of-use price adjustment list.

[0052] Based on the time-of-use pricing adjustment list, initialize the structured container of the strategy file, create a globally unique strategy number using a UUID generation algorithm combined with the current server's millisecond-level timestamp, read the number of strategies generated that day and increment it to determine the version number, call the system clock to record the specific time of strategy generation, extract the logical node ID and physical address of the applicable power supply area from the power grid topology model, set the validity period of the strategy based on the predicted risk end time and add a 2-hour observation margin, configure the dynamic rollback rules of the strategy, set that if the real-time monitored transformer load rate is lower than 85% of the rated capacity for three consecutive sampling cycles (15 minutes per cycle), an automatic stop command will be triggered and the price will roll back to the benchmark price, load a standardized notification message template, fill the placeholders in the template with the calculated dynamic price, execution period and response reason, query the permission configuration table of the dispatch management system to obtain the list of approval personnel and their contact information, define key monitoring items during the execution process, including user response execution rate, load voltage drop value and voltage deviation rate, and finally serialize and encode all the above fields according to the JSON data exchange format and attach a digital signature to ensure the integrity and non-repudiation of the file, generating a demand response dynamic pricing strategy file.

[0053] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A transformer load rate prediction method based on data mining, characterized in that, Includes the following steps: Based on the sensing signals of smart meters at each user terminal on the low-voltage side of the transformer and within the power supply area, real-time voltage, current and active power readings are collected, and the status values ​​of transformer winding temperature and oil level indication are obtained synchronously to generate a time-series electrical state parameter set. Based on the time-series electrical state parameter set, the standard deviation of power usage and peak-valley ratio of each user node within a set period are calculated to quantify and generate a user electricity consumption behavior feature dataset. Map each data point in the user electricity consumption behavior feature dataset to a multi-dimensional feature vector coordinate system, calculate the distance between each feature vector point in the coordinate system, filter behavior clusters based on the distance, generate a feature vector spatial distribution set, classify adjustable nodes and rigid nodes based on the feature vector spatial distribution set, extract the device ID information of the adjustable nodes, and generate a flexible load user identification list. Extract the current real-time load amplitude of each user in the flexible load user identification list, calculate the expected load value of the user, and combine the expected load value of the user with the total expected load value of the transformer to generate the transformer overload risk prediction value. Based on the predicted transformer overload risk values, a dynamic demand response pricing strategy file is generated.

2. The transformer load rate prediction method based on data mining according to claim 1, characterized in that, The steps for obtaining the user electricity consumption behavior feature dataset are as follows: Based on the collected voltage, current and active power values, they are matched with the transformer winding temperature and oil conservator level indication values ​​according to a unified timestamp. The power readings of each user node within a set period are serialized, the power usage standard deviation and peak-to-valley ratio are calculated, and the data are bound to the user identifier and time identifier to generate a time-series electrical state parameter set. Calculate the load sensitivity based on the set of time-series electrical state parameters; Based on the load sensitivity, the load sensitivity of each user node is aggregated with the user identifier and time identifier, and then arranged into continuous records in chronological order to generate a user electricity consumption behavior feature dataset.

3. The transformer load rate prediction method based on data mining according to claim 1, characterized in that, The steps for obtaining the feature vector spatial distribution set are as follows: Based on the user electricity consumption behavior feature dataset, each data point is mapped to a multi-dimensional feature vector coordinate system according to a unified time index. Covariance analysis is performed on the feature vectors of each pair of user nodes and the correlation between features is corrected. The weighted Mahalanobis distance is calculated based on the linear correlation structure between features. Based on the weighted Mahalanobis distance, node pairs with a density less than a preset threshold are selected to establish a connection relationship. All sets of interconnected nodes are defined as behavioral clusters. The center vector, dispersion, and number of internal members of each behavioral cluster are counted to generate a set of feature vector spatial distributions.

4. The transformer load rate prediction method based on data mining according to claim 1, characterized in that, The steps for obtaining the elastic load user identification list are as follows: Based on the set of feature vector spatial distributions, calculate the adjustable potential index for each behavioral cluster; Based on the comparison between the adjustable potential index and the preset threshold, when the adjustable potential index is greater than the preset threshold, the cluster is marked as an adjustable node, and when the adjustable potential index is less than the preset threshold, it is marked as a rigid node. The device ID information corresponding to all adjustable nodes is extracted to generate a list of elastic load user identities.

5. The transformer load rate prediction method based on data mining according to claim 1, characterized in that, The steps for obtaining the user's expected load value are as follows: Based on the flexible load user identification list, the real-time load amplitude of each user is read one by one. The current ambient temperature and the operating status of the transformer cooling system are synchronized according to the timestamp. The temperature reading is extracted, the operating status indication value is extracted from the cooling control unit, and input into the adjustment coefficient mapping table to retrieve the corresponding adjustment coefficient. The real-time load amplitude of each user is multiplied by the adjustment coefficient to obtain the corrected load value, and recorded with user identifier and time index to generate the user's expected load value.

6. The transformer load rate prediction method based on data mining according to claim 1, characterized in that, The steps for obtaining the transformer overload risk prediction value are as follows: Based on the user's expected load value, the user group is divided according to the future time period. The expected load values ​​of users in the flexible load user identification list are summed by time period to obtain the total load value of the flexible user time period. At the same time, the baseline load values ​​of non-listed users are summed by the same time period to obtain the total load value of the non-listed user time period. The total load value of the flexible user time period and the total load value of the non-listed user time period are added together to obtain the total expected load value of the transformer in the current time period. Based on the total expected load value of the transformer in each time period, iterate through all the total expected load value sequences in the future time period, select the maximum value and define it as the peak load prediction value of the transformer, compare the peak load prediction value with the rated capacity on the transformer nameplate, calculate the difference ratio, and look up the corresponding risk level in the overload risk judgment table according to the difference ratio to generate the transformer overload risk prediction value.

7. The transformer load rate prediction method based on data mining according to claim 1, characterized in that, The steps for obtaining the demand response dynamic pricing strategy file are as follows: Based on the predicted value of transformer overload risk, extract the difference ratio and time index for each future period, merge adjacent risk periods according to the operation and maintenance calendar, look up the table to map the difference ratio to the response level, mark the triggering method and applicable power supply area, and generate a list of demand response triggering periods. Based on the demand response triggering period list, the current time-of-use electricity price table is read for each triggering period, and the price adjustment level, start and end time, user scope and settlement method are bound according to the response level. The billing accuracy and minimum billing duration limits are supplemented, and a time-of-use price adjustment list is generated.

8. The transformer load rate prediction method based on data mining according to claim 7, characterized in that, The steps for obtaining the demand response dynamic pricing strategy file also include: Based on the time-of-use pricing adjustment list, compile the strategy number, version number, generation time, applicable power supply area, validity period, rollback rules, notification template, approver list, and execution monitoring items, write them into structured text in the order of fields, and attach the effective conditions to generate a demand response dynamic pricing strategy file.