Power distribution area cooling and heating load identification method and system based on multi-dimensional feature fusion
Through the multi-dimensional feature fusion method, sliding window and Pearson correlation coefficient are used to capture temperature mutations. Combined with the load spectrum entropy characteristics, the XGboost model is used to classify cold and warm loads. This solves the problems of load response lag and insufficient feature representation in the existing technology and improves the classification accuracy.
Patent Information
- Application Number
- CN202510699828.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-12
AI Technical Summary
Existing temperature-controlled load identification technologies have problems such as temperature response lag, insufficient feature representation, sample imbalance, and timeliness bottlenecks, resulting in low load classification accuracy.
By acquiring time-series load data and ambient temperature data, the sliding window algorithm is used to capture temperature mutation events. The load response characteristics are verified by combining the Pearson correlation coefficient, and the load spectrum entropy features are extracted. The classification model of the XGboost framework is used to perform multi-dimensional feature fusion, dimensionality reduction, and classification prediction.
The accuracy of cooling and heating load classification was increased to 92.3%, an increase of 23.5% compared with traditional methods. This solved the problems of load response lag and insufficient feature representation and reduced the complexity of the model.
Smart Images

Figure CN120632652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of temperature control load identification, and in particular to a method and system for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion. Background Art
[0002] The power system is currently undergoing a profound transformation from traditional centralized power supply to a high proportion of renewable energy integration. Against this backdrop, load characteristic analysis technology faces unprecedented challenges and opportunities. However, existing technologies still have shortcomings in the field of temperature control load identification, as shown in the following aspects:
[0003] 1. Temperature response hysteresis: Traditional temperature correlation coefficient methods, such as the short-term combined load forecasting method disclosed in Chinese patent CN103617467A, directly make predictions based on the original load data of the power grid, ignoring the temporal and spatial delay characteristics of the load response.
[0004] 2. Insufficient feature representation: Existing methods, such as the building energy consumption and power load forecasting method and system based on the CNN-LSTM hybrid network model disclosed in Chinese patent CN113537571A, rely on a single statistical feature in the proposed CNN-LSTM model and are unable to capture the complexity of load fluctuations.
[0005] 3. Sample imbalance: In actual substations, normal loads account for over 70%, and cold / warm load samples are scarce, causing the classifier to favor the majority class.
[0006] 4. Timeliness bottleneck: When processing massive amounts of data, traditional models have high computational complexity, resulting in prediction delays of more than 2 hours. Summary of the Invention
[0007] The purpose of the present invention is to overcome the defects of the above-mentioned existing temperature control load identification technology, which ignores the spatiotemporal delay characteristics of load response, relies on a single statistical feature, and cannot capture the complexity of load fluctuations, and to provide a distribution station area heating and cooling load identification method and system based on multi-dimensional feature fusion.
[0008] The purpose of the present invention can be achieved by the following technical solutions:
[0009] A method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion includes the following steps:
[0010] Obtain the time series load data and ambient temperature data of the target substation and perform data preprocessing;
[0011] For the pre-processed data, a sliding window algorithm is used to capture the temperature mutation time and verify the load response correlation to select valid events and generate a dynamic thermal response event set;
[0012] extracting a multidimensional feature vector including load spectrum entropy based on the dynamic thermal response event set;
[0013] The classification model based on the XGboost framework is used to classify and predict the extracted multidimensional feature vectors, and the classification results of cooling load, heating load and normal load are obtained;
[0014] An energy efficiency optimization strategy for the target area is generated according to the classification result.
[0015] Furthermore, the method of capturing the temperature mutation time by the sliding window algorithm is as follows:
[0016] A sliding window of the first time scale is used to perform sliding calculations on the entire time period of the preprocessed data, and the standard deviation of the temperature within each sliding window is calculated;
[0017] If the calculated temperature standard deviation within the window exceeds the first fraction of the historical temperature, it is marked as a candidate event;
[0018] The calculation expression of the temperature standard deviation in the window is:
[0019]
[0020] Where σ T is the temperature standard deviation in the window, N is the number of data points in the sliding window, T i is the temperature sampling value, is the window mean.
[0021] Furthermore, the load response correlation verification process includes the following steps:
[0022] The load average of the same type of time period 7 days before the candidate event occurs is selected as the baseline;
[0023] intercepting the load data of the second time scale after the candidate event occurs, and calculating the difference series between the load data and the baseline;
[0024] The Pearson correlation coefficient between the difference sequence and the time series of the sliding window corresponding to the candidate event is calculated, and the candidate events are screened according to the Pearson correlation coefficient to obtain valid events.
[0025] Furthermore, the calculation expression of the baseline is:
[0026]
[0027] Where, P base As the baseline, P historical (t d ) is the load mean d days before the candidate event occurs;
[0028] The calculation expression of the difference sequence is:
[0029] ΔP=P response -P base
[0030] Where ΔP is the difference sequence, P response is the load data of the second time scale after the candidate event occurs;
[0031] The calculation expression of the Pearson correlation coefficient is:
[0032]
[0033] Where r is the Pearson correlation coefficient, t i is the time point i in the time series, ΔP i is the load change value at time point i.
[0034] Furthermore, the first time scale is 3 hours, and the second time scale is 6 hours.
[0035] Furthermore, the calculation process of the load spectrum entropy is specifically as follows:
[0036] The load data in the dynamic thermal response event set is divided into days, and a sequence of 24 sampling points is generated every day;
[0037] Perform approximate entropy calculation on each sampling point sequence to obtain the calculation result of load spectrum entropy;
[0038] In the calculation process of the approximate entropy calculation, the pattern dimension m=3 and the tolerance coefficient r=0.15*σ are defined. P ,σ P is the daily load standard deviation, and the calculation expression of the load spectrum entropy is:
[0039] ApEn(m,r)=φ m (r)-φ m+1 (r)
[0040] Where, ApEn is the load spectrum entropy, φ m (r) is the logarithmic mean of the pattern similarity probability in the m-dimensional space and under the tolerance coefficient r; φ m+1 (r) is the logarithmic mean of the pattern similarity probability in the m+1 dimensional space and under the tolerance coefficient r.
[0041] Furthermore, the classification model based on the XGboost framework compresses the multidimensional feature vector through PCA dimensionality reduction to perform classification prediction.
[0042] Furthermore, the method further includes performing seasonal correlation verification and confidence level classification prediction on the classification results output by the classification model;
[0043] The seasonal correlation verification includes:
[0044] The negative correlation between temperature and load was verified from June to September;
[0045] The positive correlation between temperature and load was verified from November to March;
[0046] The confidence level prediction includes:
[0047] When the predicted probability output by the classification model is greater than 0.6, the classification result is directly output;
[0048] When the predicted probability output by the classification model is between 0.4 and 0.6, it is marked as a suspected load and triggers manual review;
[0049] When the predicted probability output by the classification model is less than 0.4, it is classified as a normal load.
[0050] Furthermore, the data preprocessing specifically includes: data integrity check and outlier repair;
[0051] The data integrity check includes: marking abnormal labels for time periods with a missing rate exceeding 5%, and filling missing values using a third-order polynomial interpolation method;
[0052] The outlier repair includes: using similar day data from neighboring stations to fill in data segments that are missing for more than 3 hours; and cross-validating and correcting the obtained ambient temperature data using meteorological station observations.
[0053] The present invention also provides a distribution station area heating and cooling load identification system based on multi-dimensional feature fusion, including a memory and a processor, the memory stores a computer program, and the processor calls the computer program to execute the steps of the distribution station area heating and cooling load identification method based on multi-dimensional feature fusion as described above.
[0054] Compared with the prior art, the present invention has the following advantages:
[0055] (1) The present invention first extracts time-series load data and ambient temperature data simultaneously, captures temperature mutation events through a sliding window standard deviation algorithm, and compares the temperature response lag with the data before and after the temperature mutation event in combination with the Pearson correlation coefficient to verify the load response characteristics and realize the extraction of spatiotemporal correlation features; secondly, a load spectrum entropy feature system is designed, in which an approximate entropy algorithm is used to quantify the complexity of load fluctuations to solve the defect of insufficient representation power of traditional statistical features; finally, the XGboost classification model is used to reduce the dimensionality of the multi-dimensional feature vector, reduce the complexity of the model, and realize the final classification prediction; overall, the cold / warm load classification accuracy reaches 92.3% (F1-score 0.87), which is 23.5% higher than the traditional temperature correlation coefficient method. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A schematic flow chart of a method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion according to an embodiment of the present invention;
[0057] Figure 2 The present invention provides a flowchart of obtaining and preprocessing time series load data and ambient temperature data of a target substation in an embodiment of the present invention. DETAILED DESCRIPTION
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0059] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0060] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0061] In the description of the present invention, it should be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, or are the orientation or position relationship in which the product of the invention is usually placed when in use. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the system or component referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they should not be understood as limiting the present invention.
[0062] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0063] Furthermore, terms such as "horizontal" and "vertical" do not necessarily mean that a component must be absolutely horizontal or overhanging, but rather that it can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal than "vertical," and does not mean that the structure must be completely horizontal, but rather that it can be slightly tilted.
[0064] Example 1
[0065] like Figure 1 As shown, this embodiment provides a method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion, comprising the following steps:
[0066] S1: Obtain the time series load data and ambient temperature data of the target substation and perform data preprocessing;
[0067] S2: For the pre-processed data, a sliding window algorithm is used to capture the temperature mutation time and verify the load response correlation to select valid events and generate a dynamic thermal response event set;
[0068] S3: Extracting a multidimensional feature vector containing load spectrum entropy based on the dynamic thermal response event set;
[0069] S4: A classification model based on the XGboost framework is used to classify and predict the extracted multidimensional feature vectors, and the classification results of cooling load, heating load and normal load are obtained;
[0070] S5: Generate an energy efficiency optimization strategy for the target area based on the classification results.
[0071] Specifically, the data and processing process in step S1 includes: data integrity check and outlier repair;
[0072] Data integrity checks include:
[0073] (1) Mark the time period where the missing rate exceeds 5% with an abnormal mark;
[0074] (2) Third-order polynomial interpolation is used to fill missing values.
[0075] Outlier fixes include:
[0076] (1) For data segments missing for more than 3 consecutive hours, similar day data from neighboring stations are used to fill in the gaps;
[0077] (2) The temperature data were cross-validated and corrected using meteorological station observations.
[0078] In step S2, the sliding window algorithm is used to capture the temperature mutation time:
[0079] A sliding window of the first time scale is used to perform sliding calculations on the entire time period of the preprocessed data, and the standard deviation of the temperature within each sliding window is calculated;
[0080] The calculation expression of the temperature standard deviation within the window is:
[0081]
[0082] Where σ T is the temperature standard deviation in the window, N is the number of data points in the sliding window, T i is the temperature sampling value, is the window mean.
[0083] If the calculated temperature standard deviation within the window exceeds the first quantile before the historical temperature, it is marked as a candidate event. In this embodiment, when the standard deviation exceeds the first 10% quantile of the historical temperature, it is marked as a candidate event.
[0084] The verification process for load response relevance includes the following steps:
[0085] Baseline load calculation: The load average of the same type of time period (weekdays / holidays) 7 days before the candidate event occurs is selected as the baseline;
[0086]
[0087] Where, P base As the baseline, P historical (t d ) is the load mean d days before the candidate event occurs;
[0088] Data interference from extreme weather days (such as temperature >35℃ or <0℃) was excluded.
[0089] Response window extraction: intercept the load data of the second time scale after the candidate event occurs (Presponse ), calculate the difference series between the load data and the baseline;
[0090] ΔP=P response -P base
[0091] Where ΔP is the difference sequence, P response is the load data of the second time scale after the candidate event occurs;
[0092] Pearson correlation coefficient verification: Calculate the Pearson correlation coefficient of the difference sequence ΔP and the time series of the sliding window corresponding to the candidate event, and screen the candidate events based on the Pearson correlation coefficient to obtain valid events.
[0093]
[0094] Where r is the Pearson correlation coefficient, t i is the time point i in the time series, ΔP i is the load change value at time point i; the threshold |r| is set to ≥ 0.4, and only significant related events are retained.
[0095] In this embodiment, the first time scale is 3 hours (covering a typical temperature control device response cycle), and the second time scale is 6 hours.
[0096] In step S3, the calculation process of the load spectrum entropy is specifically as follows:
[0097] First, the time series is segmented: the load data in the dynamic thermal response event set is divided by day, and a sequence of 24 sampling points is generated every day;
[0098] Then perform approximate entropy calculation: perform approximate entropy calculation on each sampling point sequence to obtain the calculation result of load spectrum entropy;
[0099] In the calculation process of approximate entropy calculation, the model dimension m = 3 and the tolerance coefficient r = 0.15*σ are defined. P ,σ P is the daily load standard deviation, and the calculation expression of load spectrum entropy is:
[0100] ApEn(m,r)=φ m (r)-φ m+1 (r)
[0101] Where, ApEn is the load spectrum entropy, φ m (r) is the logarithmic mean of the pattern similarity probability in the m-dimensional space and under the tolerance coefficient r; φ m+1 (r) is the logarithmic mean of the pattern similarity probability in the m+1 dimensional space and under the tolerance coefficient r.
[0102] When the time series length N ≤ m + 2, the default entropy value 0.0 is returned.
[0103] In step S4, the classification model based on the XGboost framework performs dimensionality reduction through PCA to compress the multidimensional feature vector for classification prediction; and optimizes the classification performance in combination with dynamic category weights.
[0104] For sample balance and dimensionality reduction: PCA dimensionality reduction, as follows:
[0105] Retaining 95% of the principal components of the variance, compressing the 12-dimensional features to 7 dimensions, and calculating the contribution rate:
[0106]
[0107] Among them, λ i are the eigenvalues of the covariance matrix.
[0108] That is, the principal component features with 95% variance are retained through PCA dimensionality reduction;
[0109] In addition, set the category weights: cooling load 1.5, heating load 1.2, and normal load 0.8.
[0110] Preferably, the method further comprises performing seasonal correlation verification and confidence grading prediction on the classification results output by the classification model, that is, implementing result credibility grading based on the prediction probability threshold, and improving classification reliability by combining the seasonal correlation verification mechanism;
[0111] Seasonal correlation verification includes:
[0112] In summer (June-September), the negative correlation between temperature and load is verified;
[0113] In winter (January 1-March), it is mandatory to verify the positive correlation between temperature and load;
[0114] Confidence-graded predictions include:
[0115] When the predicted probability output by the classification model is greater than 0.6, the classification result is directly output;
[0116] When the predicted probability output by the classification model is between 0.4 and 0.6, it is marked as a suspected load and triggers manual review;
[0117] When the predicted probability output by the classification model is less than 0.4, it is classified as a normal load.
[0118] Specific implementation process:
[0119] In this implementation, the cooling and heating load identification method of the distribution station area based on multi-dimensional feature fusion includes the following steps:
[0120] S1. Obtain massive historical load data of distribution transformers and temperature data of corresponding dates, and perform preprocessing
[0121] Obtain massive historical load data of distribution transformers in substations and preprocess the data. The preprocessing process includes mean correction for abnormal values and outliers in the load data and filling in null values using polynomial interpolation method.
[0122] For n+1 different sample points (x i ,y i ), there exists a unique polynomial:
[0123] L n (x) = a0 + a1x + a2x 2 +...+a n x n
[0124] Make L n (x j )=y j (j=0,1,2,…,n)
[0125] Therefore, the null value load ratio under any date can be filled according to the unique polynomial.
[0126] At the same time, the corresponding temperature data and load data are combined to obtain complete sample data. Figure 2 shown.
[0127] S2. Establish a multi-feature cooling and heating load recognition model based on XGboost
[0128] First, the system calculates each feature, using a sliding window algorithm to capture sudden temperature changes. A three-hour window (covering the typical response cycle of temperature control equipment) is set, and the temperature standard deviation within the window is calculated in real time. A dynamic threshold strategy is employed: the threshold is set to the top 10% of the historical standard deviation (e.g., 1.5°C) in summer and lowered to 0.8°C in winter to accommodate seasonal temperature fluctuations. For detected events, the load response correlation is further verified: the mean load value for the same period of time over the seven days prior to the event is used as a baseline. The difference between the load and the baseline for the six hours after the event is calculated, and the Pearson correlation coefficient is used to verify the strength of the correlation with temporal variation. Only valid events with an absolute value of 0.4 or higher are retained. Next, features are extracted from valid events. For load spectrum entropy, the full-day load data is divided into three-hour segments, the frequency of similar fluctuation patterns is counted, and the entropy value is calculated using an approximate entropy algorithm. A fault-tolerance mechanism is introduced to automatically return to a default value when insufficient data points are available to prevent calculation errors.
[0129] S3. Feature optimization and dimensionality reduction
[0130] First, to address the scarcity of cold and warm load samples, we randomly selected neighbors of the cold load in the feature space and generated new data through linear interpolation, increasing the number of cold and warm samples to 80% of the normal load, effectively alleviating category imbalance. Then, based on principal component analysis (PCA), we eliminated redundancy and screened the top seven principal components with a cumulative variance contribution rate ≥ 95%, compressing the feature dimension to 58% of the original data while retaining 96.3% of the information, significantly reducing computational complexity.
[0131] S4. Model parameter setting and training
[0132] The model was trained using the XGboost framework. The maximum depth of the tree model and the initial learning rate were set to 0.5 to balance the model complexity and generalization ability. A gradient boosting framework was constructed by combining 1000 decision trees. An L2 regularization of 0.2 was used to suppress overfitting, and a sub-sampling rate of 0.8 was used to enhance data perturbation robustness. To address the problem of class imbalance, weights of 1.5, 1.2, and 0.8 were assigned to the cooling load, heating load, and normal load, respectively. 5-fold cross-validation was used during training, and 80% of the data was randomly sampled from each fold as the training set. An early stopping mechanism was implemented by monitoring the validation set loss. The batch size was set to 256 to adapt to the GPU memory limitation. Feature importance evaluation showed that temperature difference sensitivity, load spectral entropy, and 2-hour lag correlation coefficient ranked in the top three, with the cumulative gain accounting for 67%, confirming that the response characteristics of the temperature-controlled load dominated the classification decision. The test set covers data from 1,133 substations. The model achieved 85% recall and 89% precision for cooling load identification, 88% recall and 86% precision for heating load identification, and 94% accuracy for general load classification. The overall macro-average F1-score reached 89.3%. Confusion matrix analysis showed a misclassification rate of only 4.2% for cooling / warming loads. The classification results are shown in Table 1.
[0133] Table 1
[0134] Load type Cooling load Heating load Normal load Number of stations 265 833 35
[0135] Example 2
[0136] This embodiment provides a distribution station area heating and cooling load identification system based on multi-dimensional feature fusion, including a memory and a processor. The memory stores a computer program, and the processor calls the computer program to execute the steps of the distribution station area heating and cooling load identification method based on multi-dimensional feature fusion as in Example 1.
[0137] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0138] It will be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented in various computer languages, for example, the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0139] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A system that specifies the functions of a box or boxes.
[0140] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction system that is implemented in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0142] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A method for identifying cooling and heating loads in distribution station areas based on multi-dimensional feature fusion, characterized in that: The following steps are involved: Obtain the time series load data and ambient temperature data of the target substation and perform data preprocessing; For the pre-processed data, a sliding window algorithm is used to capture the temperature mutation time and verify the load response correlation to select valid events and generate a dynamic thermal response event set; extracting a multidimensional feature vector including load spectrum entropy based on the dynamic thermal response event set; The classification model based on the XGboost framework is used to classify and predict the extracted multidimensional feature vectors, and the classification results of cooling load, heating load and normal load are obtained; An energy efficiency optimization strategy for the target area is generated according to the classification result.
2. The method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion according to claim 1 is characterized in that: The specific method of capturing the temperature mutation time by the sliding window algorithm is as follows: A sliding window of the first time scale is used to perform sliding calculations on the entire time period of the preprocessed data, and the standard deviation of the temperature within each sliding window is calculated; If the calculated temperature standard deviation within the window exceeds the first fraction of the historical temperature, it is marked as a candidate event; The calculation expression of the temperature standard deviation in the window is: Where, σ T is the temperature standard deviation in the window, N is the number of data points in the sliding window, T i is the temperature sampling value, is the window mean.
3. The method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion according to claim 2 is characterized in that: The load response correlation verification process includes the following steps: The load average of the same type of time period 7 days before the candidate event occurs is selected as the baseline; intercepting the load data of the second time scale after the candidate event occurs, and calculating the difference series between the load data and the baseline; The Pearson correlation coefficient between the difference sequence and the time series of the sliding window corresponding to the candidate event is calculated, and the candidate events are screened according to the Pearson correlation coefficient to obtain valid events.
4. The method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion according to claim 3 is characterized in that: The calculation expression of the baseline is: Where, P base As the baseline, P historical (t d ) is the load mean d days before the candidate event occurs; The calculation expression of the difference sequence is: ΔP=P response -P base Where ΔP is the difference sequence, P response is the load data of the second time scale after the candidate event occurs; The calculation expression of the Pearson correlation coefficient is: Where r is the Pearson correlation coefficient, t i is the time point i in the time series, ΔP i is the load change value at time point i.
5. The method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion according to claim 4 is characterized in that: The first time scale is 3 hours, and the second time scale is 6 hours.
6. The method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion according to claim 1 is characterized in that: The calculation process of the load spectrum entropy is specifically as follows: The load data in the dynamic thermal response event set is divided into days, and a sequence of 24 sampling points is generated every day; Perform approximate entropy calculation on each sampling point sequence to obtain the calculation result of load spectrum entropy; In the calculation process of the approximate entropy calculation, the pattern dimension m=3 and the tolerance coefficient r=0.15*σ are defined. P ,σ P is the daily load standard deviation, and the calculation expression of the load spectrum entropy is: ApEn(m,r)=φ m (r)-φ m+1 (r) Where, ApEn is the load spectrum entropy, φ m (r) is the logarithmic mean of the pattern similarity probability in the m-dimensional space and under the tolerance coefficient r; φ m+1 (r) is the logarithmic mean of the pattern similarity probability in the m+1 dimensional space and under the tolerance coefficient r.
7. The method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion according to claim 1 is characterized in that: The classification model based on the XGboost framework compresses the multidimensional feature vector through PCA dimensionality reduction to perform classification prediction.
8. The method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion according to claim 1 is characterized in that: The method further comprises performing seasonal correlation verification and confidence grading prediction on the classification results output by the classification model; The seasonal correlation verification includes: The negative correlation between temperature and load was verified from June to September; The positive correlation between temperature and load was verified from November to March; The confidence level prediction includes: When the predicted probability output by the classification model is greater than 0.6, the classification result is directly output; When the predicted probability output by the classification model is between 0.4 and 0.6, it is marked as a suspected load and triggers manual review; When the predicted probability output by the classification model is less than 0.4, it is classified as a normal load.
9. The method for identifying cooling and heating loads in a distribution station area based on multi-dimensional feature fusion according to claim 1 is characterized in that: The data preprocessing specifically includes: data integrity check and outlier repair; The data integrity check includes: marking abnormal labels for time periods with a missing rate exceeding 5%, and filling missing values using a third-order polynomial interpolation method; The outlier repair includes: using similar day data from neighboring stations to fill in data segments that are missing for more than 3 hours; and cross-validating and correcting the obtained ambient temperature data using meteorological station observations.
10. A cooling and heating load identification system for distribution station area based on multi-dimensional feature fusion, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor calls the computer program to execute the steps of any one of the methods according to claims 1 to 9.
Citation Information
Patent Citations
Short-period combined load prediction method
CN103617467A
Building energy consumption electric load prediction method and device based on CNN-LSTM hybrid network model
CN113537571A