Data Mining Method and Device for Energy Consumption Behavior of Typical Factor Integrated Energy System

The K-means algorithm corrects the damaged data and performs standardized processing, extracts the characteristics of user electricity use behavior, solves the problem of low data mining accuracy in the existing technology, and achieves higher data mining accuracy.

CN114529330BActive Publication Date: 2025-06-10SICHUAN ENERGY INTERNET RES INST TSINGHUA UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111669628.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-06-10
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The prior art is difficult to accurately analyze and mine user's electricity consumption behavior data, especially in the case of damage or interference in the data, which affects the accuracy of load analysis.

Method used

The K-means algorithm is used to correct the damaged data, and the data is quantified through standardized processing, forming a sequence of user electricity consumption behavior characteristics, calculating the correlation between numerical relationships of adjacent periods, and extracting feature indicators to realize data mining.

Benefits of technology

It effectively avoids interference from typical factors, improves the accuracy of data mining, and can accurately extract user's power consumption behavior characteristics, and its accuracy rate is higher than that of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The present invention discloses a method and device for mining energy consumption behavior data of an integrated energy system considering typical factors. The method includes using the K-means algorithm to correct damaged data, and mining and marginalizing some disturbed data in the data. Therefore, a standardized processing method is adopted to quantify the data. After filtering, the electricity consumption behavior characteristics of users can be accurately extracted. Experimental results show that traditional data mining methods are easily affected by typical factors such as nature, and the accuracy of data mining is not ideal. The present invention can effectively extract correct behavior data, and the accuracy rate is higher than that of traditional methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data mining, and particularly relates to a data mining method and device for the energy consumption behavior of an integrated energy system considering typical factors. Background Art

[0002] With the advent of the big data era, in the context of the smart grid, a large amount of power consumption data has been accumulated in the power grid information collection system and the customer service information system, hiding a large amount of power consumption information. As industrial loads are major power consumers, it is of great significance to use electricity in an orderly, efficient, energy-saving and environmental-friendly manner. Therefore, the future smart grid should ensure the safe and reliable power consumption while providing higher-quality, more targeted services and scientific suggestions for different users. Therefore, it is of great significance to analyze the growth laws and characteristics of the power consumption of users and power supply enterprises.

[0003] The massive historical load data of users accumulated in the power consumption information collection system contains the power consumption behaviors and habits of users. It can not only improve the accuracy of load forecasting and the dispatching management level, but also provide support for electricity price setting, economic dispatching and demand response. With the new round of power system reform, large-power and stable-power-consuming industrial and commercial users will directly participate in bilateral transactions, the power spot market and demand-side response, and bear clean energy quotas, which will have a significant impact on power generation dispatching plans, power grid operation modes, power grid peak shaving capabilities, new energy consumption, etc.

[0004] At present, there are many analysis methods for energy usage behaviors, such as using duration algorithms, analyzing energy usage anomalies according to community characteristics, using fuzzy clustering algorithms to study the behavior patterns of power customers, etc. However, it is difficult to quantitatively estimate the current user behavior data of energy usage, and the mining accuracy of user behavior data needs to be further improved. In order to further improve the accuracy of data mining, this paper proposes a data mining method for the energy consumption behavior of an integrated energy system considering typical factors. Summary of the Invention

[0005] Under the power reform, the opening of the power selling end makes it very important for power companies, consumers and the entire power market to obtain the power consumption behaviors of different users. Users with different living habits and backgrounds have different power consumption behaviors. From the perspective of power companies, the analysis of user behaviors can help power retailers formulate more efficient and satisfactory marketing and demand plans, and contribute to accumulating a large number of customers. For residential electricity users, through interaction with power companies, they can better understand their consumption patterns, and even adjust and optimize their power consumption habits, cut peaks and fill valleys under the incentive of electricity packages, reduce household electricity costs, and also contribute to the overall stability of the power grid.

[0006] Single - user load pattern extraction is the basis for studying user classification, which helps to achieve demand response, enabling power suppliers to implement effective energy control, flexible pricing, and demand management based on load patterns and consumption categories, and enabling power users to understand their load patterns and cope with fluctuating prices, thereby reducing electricity bills. There are two challenges in extracting load patterns from daily load curve clustering. One is dimensionality reduction to reduce information loss, and the other is to improve the performance of daily load curve clustering.

[0007] To achieve the above - mentioned purpose, the present invention provides the following technical solutions:

[0008] A method for data mining of energy - using behavior data in an integrated energy system considering typical factors, the method is carried out according to the following steps:

[0009] Obtain the load characteristic curve of behavior data through the K - means algorithm;

[0010] Correct the bad data on the load characteristic curve to obtain an electricity - using correction curve;

[0011] Perform unified quantization on the correction curve by using standardization processing to obtain the normalized value of electricity - using data;

[0012] Form a user electricity - using behavior feature sequence from the normalized values of the electricity - using data;

[0013] Calculate the correlation between the numerical relationships of adjacent periods in the feature sequence according to covariance to obtain a correlation numerical relationship;

[0014] Extract feature indicators according to the correlation numerical relationship to achieve data mining.

[0015] Preferably, the bad data includes data loss or damage under the influence of the data acquisition system and other external factors; the bad data will affect the accuracy of load analysis.

[0016] Preferably, obtaining the load characteristic curve of behavior data through the K - means algorithm includes using the K - means algorithm to cluster the daily load curves of each user. Among them, the load curves of each user on weekdays, weekends, and holidays are quite different, and the weekdays, weekends, and holidays of users are clustered separately.

[0017] Preferably, the correction of the bad data includes that the bad - data correction equation is:

[0018]

[0019] In formula (1), T d is the curve to be modified, T cIt is a characteristic curve, p and q are points on Td and Tc respectively, T(n) is the load characteristic curve, Q(n) is the electricity consumption correction curve, and n is a point on the user curve.

[0020] Preferably, the correction curve is uniformly quantified by using standardization processing, and the transformation process is expressed as:

[0021] g * = Q(n) × (g'+g min ) / g max (2)

[0022] In formula (2), g * represents the normalized value of electricity consumption data, g' represents the electricity consumption of the user within the selected time period, g min represents the minimum electricity consumption of the user within the selected time period, and g max represents the maximum value of electricity consumption behavior, and Q(n) is the electricity consumption correction curve.

[0023] Preferably, the numerical relationship of the user's electricity consumption behavior characteristic sequence can be expressed as:

[0024] E m ={e d , d = 1, 2,..., M} (3)

[0025] In formula (3), E m represents the user's electricity consumption behavior characteristic sequence in the mth measurement period, m represents the number of periods, M is the number of days in the period, and e d is the maximum value of the periodic load, and d is the characteristic day.

[0026] Preferably, the correlation between the numerical relationships of adjacent periods in the characteristic sequence is expressed by the correlation numerical relationship as:

[0027]

[0028] In formula (4), E(E m ) represents the maximum screening coefficient of abnormal electricity consumption data, n is the total number of periods, and x j represents the real-time discrimination vector element of abnormal electricity consumption data.

[0029] Preferably, the chi-square test process is used to screen the interference data of the characteristic sequence, which can be expressed as:

[0030]

[0031] In formula (5), β is the chi-square test result, and the larger this result is, the greater the deviation between the actual value and the theoretical value; m is the mth measurement period; n is the total number of periods; A m represents the standard calculation term coefficient of the missing value of electricity consumption data.

[0032] Preferably, the process of extracting the feature index, including the user's electricity consumption behavior characteristics, can be expressed as:

[0033] α = P m-1 +(P m+1 -P m-1 ) 2 (6)

[0034] In formula (6), P m-1 represents the user's electricity consumption behavior data set in the previous period, and P m+1 represents the user's electricity consumption behavior data set in the next period.

[0035] Preferably, the P m+1 and P m-1 are the user's electricity consumption behavior feature sequences after removing the interference data through the chi-square test process. m

[0036] An energy consumption behavior data mining device for an integrated energy system considering typical factors, the mining device includes:

[0037] A correction unit, configured to obtain the load characteristic curve of the behavior data through the K-means algorithm, correct the bad data on the load characteristic curve to obtain the electricity consumption correction curve, and perform unified quantization on the correction curve by using standardization processing to obtain the normalized value of the electricity consumption data;

[0038] A correlation determination unit, configured to form a user's electricity consumption behavior feature sequence from the normalized value of the electricity consumption data, and calculate the correlation between the numerical relationships of adjacent periods in the feature sequence according to the covariance to obtain the correlation numerical relationship;

[0039] A data mining unit, configured to extract feature indexes according to the correlation numerical relationship to achieve data mining.

[0040] The technical effects and advantages of the present invention:

[0041] In order to avoid the interference of typical factors and improve the accuracy of data mining, the present invention proposes a method and device for mining energy consumption behavior data of an integrated energy system considering typical factors.

[0042] 1. The K-means algorithm is used to correct the damaged data, and some of the disturbed data in the data is mined and marginalized;

[0043] 2. There are some interference data in the processed user's electricity consumption behavior data. In order to avoid the influence of typical factors, the standardization processing method is used to quantify the behavior data, and the electricity consumption behavior characteristics of the user can be accurately extracted.

[0044] In summary, to avoid the interference of typical factors and improve the accuracy of data mining, this paper proposes a data mining method for the energy consumption behavior of integrated energy systems. The K-means algorithm is used to correct the damaged data. There are some interfering data in the processed user electricity consumption behavior data. To avoid the influence of typical factors and effectively extract electricity consumption data, a standardization processing method is used to quantify the behavior data, which improves the accuracy of data mining. At the same time, the extraction accuracy of the present invention is compared with the long and short time algorithm (traditional method) and the fuzzy clustering algorithm (traditional method). The method proposed by the present invention has obvious...

[0045] Other features and advantages of the present invention will be described in the following specification, and in part will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures pointed out in the specification and claims. Detailed implementation manners

[0046] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0047] Under the power reform, the opening of the power selling end makes it very important for power companies, consumers and the entire power market to obtain the electricity consumption behaviors of different users. Users with different living habits and backgrounds have different electricity consumption behaviors. From the perspective of power companies, the analysis of user behaviors can help power retailers formulate more efficient and satisfactory marketing and demand plans, and contribute to accumulating a large number of customers. For residential electricity users, through interaction with power companies, they can better understand their consumption patterns, and even adjust and optimize their electricity consumption habits, cut peaks and fill valleys under the incentive of electricity packages, reduce household electricity costs, and also contribute to the overall stability of the power grid.

[0048] The extraction of single-user load patterns is the basis for studying user classification, which helps to achieve demand response, enables power suppliers to achieve effective energy control, flexible pricing and demand management based on load patterns and consumption categories, enables power users to understand their load patterns, and cope with fluctuating prices, thereby reducing electricity bills. There are two challenges in extracting load patterns from daily load curve clustering. One is to reduce dimensionality and avoid information loss, and the other is to improve the performance of daily load curve clustering. Due to the influence of data acquisition systems and other external factors, data is often lost or damaged. The existence of such bad data will affect the accuracy of load analysis and must be excluded and repaired.

[0049] To address the deficiencies of the prior art, the present invention discloses a method for mining energy consumption behavior data of an integrated energy system considering typical factors. The method includes: obtaining the load characteristic curve of the behavior data through the K-means algorithm; correcting the bad data on the load characteristic curve to obtain the corrected power consumption curve; performing unified quantization on the corrected curve using standardization processing to obtain the normalized value of the power consumption data; forming a user power consumption behavior feature sequence based on the normalized value of the power consumption data according to the measurement period; calculating the correlation between the numerical relationships of adjacent periods in the feature sequence based on covariance to obtain the correlation numerical relationship; and extracting feature indicators according to the correlation numerical relationship.

[0050] Traditional data mining methods, such as the usage duration algorithm, analyzing energy usage anomalies based on community characteristics, and using fuzzy clustering algorithms to study the behavior patterns of electricity customers, are easily affected by typical factors such as nature, and the accuracy of data mining is not ideal. To solve the above problems, the present invention proposes a method and device for mining energy consumption behavior data of an integrated energy system considering typical factors. The K-means algorithm is used to correct the damaged data and mine and marginalize some of the disturbed data in the data. Therefore, after filtering by using the standardization processing method to quantify the data, the power consumption behavior characteristics of users can be accurately extracted. The method of the present invention can effectively extract the correct behavior data, and the accuracy rate is higher than that of the traditional method.

[0051] Among them, a device for mining energy consumption behavior data of an integrated energy system considering typical factors, the mining device includes:

[0052] A correction unit for obtaining the load characteristic curve of the behavior data through the K-means algorithm, correcting the bad data on the load characteristic curve to obtain the corrected power consumption curve, and performing unified quantization on the corrected curve using standardization processing to obtain the normalized value of the power consumption data;

[0053] A correlation determination unit for forming a user power consumption behavior feature sequence based on the normalized value of the power consumption data, and calculating the correlation between the numerical relationships of adjacent periods in the feature sequence based on covariance to obtain the correlation numerical relationship;

[0054] A data mining unit for extracting feature indicators according to the correlation numerical relationship to achieve data mining.

[0055] A data mining method for the energy consumption behavior of an integrated energy system considering typical factors, the method comprising the following steps: obtaining a load characteristic curve of behavior data through the K-means algorithm; correcting bad data on the load characteristic curve to obtain a corrected power consumption curve; performing unified quantization on the corrected curve by means of standardization processing to obtain a normalized value of power consumption data; forming a user power consumption behavior characteristic sequence from the normalized value of the power consumption data; calculating the correlation between the numerical relationships of adjacent periods in the characteristic sequence according to covariance to obtain a correlation numerical relationship; and extracting characteristic indicators according to the correlation numerical relationship to achieve data mining.

[0056] Further, for the horizontal similarity of load data, the K-means algorithm is used to cluster the daily load curves of each user (since the load curves of users on weekdays, weekends, and holidays vary greatly among users, separate clustering is required on weekdays, weekends, and holidays for users), and the clustering center, i.e., the load characteristic curve, is obtained through the K-means algorithm, and bad data is identified and processed. The bad data correction equation is: The bad data correction equation is:

[0057]

[0058] In equation (1), T d is the curve to be modified; T c is the characteristic curve; p and q are points on Td and Tc respectively; T(n) is the load characteristic curve; Q is the corrected power consumption curve; and n is a point on the user curve.

[0059] After the user power consumption behavior data is mined and marginalized, there are some interference data. Therefore, standardization processing is used to perform unified quantization on the behavior data, and the transformation process can be expressed as:

[0060] g* = Q(n) × (g’ + g min ) / g max (2)

[0061] In equation (2), g * represents the normalized value of power consumption data, g’ represents the power consumption of the user during this period, g min represents the minimum power consumption of the user during this period, and g max represents the maximum value of power consumption behavior.

[0062] Further, after the dimensions of the same transformed behavior data, the within-class difference matrix is used to process the separated state after transformation.

[0063] Given that users of different orders of magnitude may have the same load curve pattern, the data needs to be normalized. After normalizing the user's power consumption data, the data is processed as a feature sequence, and its numerical relationship can be expressed as:

[0064] E m ={e d ,d=1,2,...,M} (3)

[0065] In formula (3), E m represents the characteristic sequence of user electricity consumption behavior in the mth metering cycle, where m represents the cycle number, M represents the cycle days, and e represents the cycle length. d is the maximum value of the periodic load and d is the characteristic day.

[0066] Furthermore, corresponding to the numerical relationship formed in adjacent periods, the correlation between adjacent period variables is calculated using covariance, and the numerical relationship of the correlation is expressed as:

[0067]

[0068] In formula (4), E(E m ) represents the maximum screening coefficient of abnormal power consumption data, n is the total number of cycles, x j Represents the real-time discriminant vector elements of abnormal power consumption data.

[0069] Furthermore, the chi-square test process is used to screen the interference data of the feature sequence, which can be expressed as:

[0070]

[0071] In formula (5), β is the chi-square test result. The larger the result, the greater the deviation between the actual value and the theoretical value. m is the mth measurement cycle. n is the total number of cycles. A m The standard calculated term coefficient representing missing values ​​in electricity usage data.

[0072] Furthermore, E m It is a user electricity consumption behavior feature sequence containing interference data. Because there is interference data, it will affect the extraction of electricity consumption behavior features and electricity consumption behavior prediction, resulting in inaccurate extraction of user electricity consumption features. Therefore, it is necessary to remove the interference data. The chi-square test method in the original handover material is used to remove the interference data, so E m Change to P m , and then according to P m Perform feature extraction.

[0073] Furthermore, in order to improve the accuracy of feature extraction indicators, the chi-square test process is used to filter the interference data. After filtering, the process of extracting the user's electricity consumption behavior features can be expressed as:

[0074] α = P m-1 +(P m+1 -P m-1 ) 2 (6)

[0075] In formula (6), P m-1 represents the user electricity consumption behavior dataset of the previous period, and P m+1 represents the user electricity consumption behavior dataset of the next period. Under the above numerical processing, after deducting the characteristic data with drastic trend changes, the research on the method for intelligent feature extraction is finally completed.

[0076] Furthermore, the P m+1 and P m-1 are the user electricity consumption behavior feature sequences after removing interference data through the chi-square test process of Em.

[0077] For a clearer understanding of the technical features, objectives, and effects of the present invention, the technical solution of the present invention will be described in detail below, but it should not be construed as a limitation on the scope of implementation of the present invention.

[0078] Currently, there are many analysis methods for energy usage behavior, such as using duration algorithms, analyzing energy usage anomalies based on community characteristics, using fuzzy clustering algorithms to study the behavior patterns of electricity customers, etc. However, it is difficult to quantitatively estimate the user behavior data of current energy usage, and the mining accuracy of user behavior data needs to be further improved. In order to avoid the interference of typical factors and improve the accuracy of data mining, the present invention proposes a data mining method for the energy consumption behavior of an integrated energy system. The K-means algorithm is used to correct the damaged data, and there are some interference data in the processed user electricity consumption behavior data. In order to avoid the influence of typical factors and effectively extract electricity consumption data, a standardization processing method is used to quantify the behavior data.

[0079] The three multivariate datasets used in this embodiment are the annual electricity consumption, gas consumption, and climate data, presented in the form of 24:00 daily data. Among them, the electricity consumption data includes 6 variables of facility, fan, refrigeration, heating, indoor lighting, and indoor equipment electricity load data. The gas and weather data contain four variables. The former is the natural gas usage data of facilities, heating, indoor equipment, and water heaters, and the latter is the temperature, dew point temperature, and wind speed of the climate data. Twenty groups of user electricity consumption data are collected and sorted as the processing object, and the actual daily electricity consumption data is used as the processing object. The user electricity consumption data is compiled as shown in Table 1.

[0080] Table 1 Comprehensive user electricity consumption data

[0081]

[0082] Taking the data in Table 1 as the standard values, the extraction accuracy of this method is compared with the long short - term algorithm (traditional method 1) and the fuzzy clustering algorithm (traditional method 2). Based on the preparations in this embodiment, after calibrating and extracting the same behavioral characteristics, three methods for defining the numerical relationship of extraction accuracy are as follows:

[0083]

[0084] In formula (6), K is the accuracy value, which is equivalent to normalizing the results of various extraction methods to facilitate comparison between different methods. The closer the processing result is to 1, the more accurate the feature extraction is; n(o) represents the data processing object, ω o represents the index data set, and W represents the behavioral feature parameter. The comparison of the extraction accuracies corresponding to the three methods is shown in Table 2.

[0085] Table 2. Three methods for extraction accuracy results

[0086]

[0087] The closer the accuracy value is to 1, the higher the accuracy of the intelligent extraction method. As shown in Table 2, the precision of traditional method 1 is 0.22 - 0.28, and the average value is about 0.26; the accuracy of traditional method 2 is 0.34 - 0.47, and the average value is about 0.42; the accuracy of the method of the present invention is 0.82 - 0.96, and the average value is about 0.86. Comparing the three methods, the precision of traditional method 1 is the worst, and the precision of the method of the present invention is significantly higher than that of the two traditional methods. Therefore, it can be clearly concluded from Table 2 that compared with the two traditional methods, the method proposed by the present invention has the best accuracy.

[0088] In summary, the present invention proposes a comprehensive energy system energy consumption behavior data mining method considering typical factors. The K - means algorithm is used to correct damaged data, and some disturbed data in the data is mined and marginalized. Therefore, a standardized processing method is used to quantify the data. After filtering, the user's electricity consumption behavior characteristics can be accurately extracted. Combining the experimental results of this embodiment shows that the method of the present invention can effectively extract correct behavior data, and the accuracy rate is higher than that of traditional methods.

[0089] It should be noted later that the above - mentioned are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data mining method for the energy consumption behavior of an integrated energy system considering typical factors, characterized in that, the method is carried out according to the following steps, obtaining the load characteristic curve of the behavior data through the K-means algorithm; correcting the bad data on the load characteristic curve to obtain the electricity consumption correction curve; wherein, the bad data correction equation is: In formula (1), T d is the curve to be modified, T c is the characteristic curve, p and q are respectively points on T d and T c ; T(n) is the load characteristic curve, Q(n) is the power consumption correction curve, and n is a point on the user curve. using standardization processing to uniformly quantify the electricity consumption correction curve to obtain the normalized value of the electricity consumption data; forming the normalized values of the electricity consumption data into a user electricity consumption behavior feature sequence; calculating the correlation between the numerical relationships of adjacent periods in the feature sequence according to the covariance to obtain the correlation numerical relationship; extracting feature indicators according to the correlation numerical relationship to achieve data mining.

2. A data mining method for the energy consumption behavior of an integrated energy system considering typical factors according to claim 1, characterized in that, the bad data includes the loss or damage of data under the influence of the data acquisition system and other external factors.

3. A data mining method for the energy consumption behavior of an integrated energy system considering typical factors according to claim 1, characterized in that, obtaining the load characteristic curve of the behavior data through the K-means algorithm includes clustering the daily load curves of each user by using the K-means algorithm, wherein the load curves of each user on weekdays, weekends and holidays are quite different from each other, and the weekdays, weekends and holidays of the users are clustered separately.

4. A data mining method for the energy consumption behavior of an integrated energy system considering typical factors according to claim 1, characterized in that, using standardization processing to uniformly quantify the correction curve, and the transformation process is expressed as: In formula (2), g * represents the normalized value of electricity consumption data, g' represents the electricity consumption of the user within the selected time period, and g min represents the minimum electricity consumption of the user within the selected time period, and g max represents the maximum value of electricity consumption behavior, and Q(n) is the electricity consumption correction curve.

5. A data mining method for the energy consumption behavior of an integrated energy system considering typical factors according to claim 1, characterized in that, the numerical relationship of the user electricity consumption behavior feature sequence is expressed as: E m = {e d , d = 1, 2, ……, M} (3) In formula (3), E m represents the user's electricity consumption behavior characteristic sequence in the m-th measurement period, where m represents the number of periods, M is the number of days in a period, and e d is the maximum load in a period, and d is the characteristic number of days.

6. A data mining method for the energy consumption behavior of an integrated energy system considering typical factors according to claim 1, characterized in that, the correlation between the numerical relationships of adjacent periods in the feature sequence, and the correlation numerical relationship is expressed as: In formula (4), E(E m ) represents the maximum screening coefficient of abnormal power consumption data, n is the total number of periods, and x j represents the real-time discrimination vector element of abnormal power consumption data.

7. A data mining method for the energy consumption behavior of an integrated energy system considering typical factors according to claim 5 or 6, characterized in that, using the chi-square test process to screen the interference data of the feature sequence, which is expressed as: In formula (5), β is the chi-square test result. The larger this result is, the greater the deviation between the actual value and the theoretical value; E m represents the user's electricity consumption behavior characteristic sequence in the m-th measurement period, where m represents the number of periods; n is the total number of periods; A m represents the standard calculation term coefficient of the missing value of the electricity consumption data.

8. A data mining method for the energy consumption behavior of an integrated energy system considering typical factors according to claim 1, characterized in that, the extraction of the feature indicators includes the process of extracting the user electricity consumption behavior features, which is expressed as: α = P m-1 +(P m+1 -P m-1 ) 2 (6) In formula (6), P m-1 represents the user's electricity consumption behavior dataset of the previous period, and P m+1 represents the user's electricity consumption behavior dataset of the next period.

9. A data mining method for the energy consumption behavior of an integrated energy system considering typical factors according to claim 8, characterized in that, The described P m+1 and P m-1 are E m The user's electricity consumption behavior feature sequence after removing interference data through the chi-square test process.

10. A data mining device for the energy consumption behavior of an integrated energy system considering typical factors, characterized in that, the mining device includes: A correction unit, which is used to obtain the load characteristic curve of behavior data through the K-means algorithm, correct the bad data on the load characteristic curve to obtain an electricity consumption correction curve, and perform unified quantization on the electricity consumption correction curve by using standardization processing to obtain a normalized value of electricity consumption data; wherein, The bad data correction equation is: In formula (1), T d is the curve to be modified, T c is the characteristic curve, p and q are points on T d and T c respectively, T(n) is the load characteristic curve, Q(n) is the power consumption correction curve, and n is a point on the user curve; A correlation determination unit, which is used to form a user electricity consumption behavior feature sequence with the normalized value of the electricity consumption data, calculate the correlation between the numerical relationships of adjacent periods in the feature sequence according to the covariance, and obtain a correlation numerical relationship; A data mining unit, which is used to extract feature indicators according to the correlation numerical relationship to realize data mining.

Citation Information

Patent Citations

  • Abnormal electricity consumption client detection method based on PSO algorithm

    CN103678766A