A residential space state recognition method based on power big data
Patent Information
- Application Number
- CN202610955910.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-22
AI Technical Summary
①缺乏日级别的精细化判定能力:现有方法主要基于月度或统计周期内的累计指标进行判定,无法实现日级别的精细化空置状态判定
[0057]优点一:日级别精细化判定,动态捕捉空置状态变化。现有方法主要基于月度或统计周期内的累计指标进行判定,无法实现日级别的精细化判定。本发明以“日”为最小尺度单元,通过机器学习模型对每个用户每一天的空置状态进行判定,能够动态捕捉空置状态的变化过程,为实时监控和预警提供数据支撑。
Smart Images

Figure CN122796512A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of resource planning technology, and specifically to a method for identifying the status of residential spaces based on big data on electricity. Background Technology
[0002] Accurate estimation of urban residential vacancy rates is a crucial basis for decision-making in real estate market regulation, urban planning, and public resource allocation. For a long time, this indicator has relied primarily on traditional methods such as population censuses, door-to-door surveys by community grid workers, or sampling surveys. These methods are extremely costly in terms of manpower and time, with survey cycles typically measured in years. They suffer from low monitoring frequency and poor timeliness, and are constrained by factors such as the separation of residents from their registered address and low resident cooperation, making it difficult to achieve routine, broad-coverage, and continuous dynamic monitoring.
[0003] In recent years, the industry has begun to explore using big data to address the shortcomings of traditional methods. One mainstream approach relies on spatiotemporal trajectory data such as mobile phone signaling and internet location tracking to identify population activity patterns and infer vacancy status from the perspective of "whether anyone lives there." However, this method essentially measures "human behavior," while vacancy rates focus on the utilization status of "housing space." There is a natural conceptual misalignment between "person" and "household." Complex situations such as multiple households, frequent travel, and differences in behavior among family members can easily lead to misjudgments from a population perspective, making it difficult to accurately reflect the true vacancy status of physical space.
[0004] Another approach directly utilizes electricity data to characterize housing usage from the energy consumption perspective. Electricity data possesses advantages such as objectivity, high frequency, comprehensive coverage, and natural binding to physical space, effectively avoiding the conceptual misalignment problem from a population perspective. However, current vacancy identification methods based on electricity data are generally quite crude. Most technical solutions simply use single judgment rules such as "zero electricity consumption for several consecutive days" or "monthly electricity consumption below a certain absolute threshold," remaining at the level of identifying "absolutely vacant households" with completely no power. These methods do not fully utilize the minute-level high-frequency data collection capabilities of the now widely adopted smart meters, nor do they delve into the deep features hidden in intraday high-frequency sequences and massive horizontal data, such as the peak-valley shape of daily power curves, the fluctuation of electricity consumption patterns across multiple time scales, and climate sensitivity responses. Consequently, they lack the ability to effectively identify complex vacancy patterns such as "hidden vacancy" where only a few appliances like refrigerators and security systems maintain automatic operation, and "periodic vacancy" like migratory birds, leading to significantly distorted vacancy rate estimates.
[0005] Existing technologies for determining urban housing vacancy rates based on electricity big data have the following main shortcomings: ① Lack of daily-level refined judgment capability: Existing methods mainly rely on cumulative indicators within a month or statistical period for judgment, failing to achieve daily-level refined vacancy status determination. This results in the inability to capture the dynamic changes in vacancy status and to distinguish different vacancy scenarios. ② Insufficient utilization of intraday high-frequency time-series features: Existing methods do not fully utilize high-frequency (e.g., hourly) electricity consumption data collected by smart meters. The morphological characteristics of intraday electricity consumption curves (e.g., peak-to-valley ratio, morning and evening peak distribution, etc.) can effectively reflect residential behavior patterns, but existing methods do not incorporate them into their feature systems. ③ Lack of effective sample construction methods: Supervised machine learning methods require labeled samples, but accurate labeling of vacancy status requires on-site verification, which is costly. Existing technologies have not proposed effective weak-label sample construction methods to address the scarcity of labeled data. ④ Lack of a hierarchical judgment mechanism for continuous vacancy status: Existing methods directly output a binary classification result indicating whether a user is vacant, without establishing a hierarchical judgment mechanism based on the number of consecutive vacancy days. Different durations of vacancy (such as seasonal vacations or long-term idleness) have different business implications and need to be handled in a tiered manner. Summary of the Invention
[0006] The technical problem this invention aims to solve is that existing technologies lack the ability to make refined daily-level judgments, do not fully utilize intraday high-frequency time-series features, lack effective sample construction methods, and lack a hierarchical judgment mechanism for continuous vacancy status. The goal is to provide a residential space status identification method based on electricity big data. Using the day as the smallest unit, a four-dimensional feature system is constructed, including electricity consumption level features, intraday time-series features, interday correlation features, and date type features. Weakly labeled training samples are constructed using a multi-rule fusion method to train a machine learning classification model to achieve daily-level vacancy status judgment. Hierarchical judgment of user idle status is achieved based on the number of consecutive vacant days. This method enables refined daily-level judgment, dynamically captures changes in vacancy status, fully utilizes intraday high-frequency time-series features to improve judgment accuracy, solves the problem of scarce labeled data through weakly labeled samples, and supports differentiated vacancy status management based on a two-stage judgment mechanism.
[0007] This invention is achieved through the following technical solution:
[0008] This invention provides a method for identifying the status of residential spaces based on big data of electricity, comprising the following specific steps:
[0009] Acquire electricity consumption time-series data of the target living space collected at sampling time intervals within the statistical period;
[0010] Based on the electricity consumption time-series data, extract electricity consumption category features, intraday time-series features, interday correlation features, and date type features;
[0011] Based on electricity consumption characteristics and intraday time series characteristics, a multi-dimensional judgment threshold is obtained to determine whether the state is vacant or not.
[0012] Daily samples that simultaneously satisfy all rules in the first set of judgment rules are marked as positive samples, and the positive samples represent empty samples.
[0013] Daily samples that simultaneously satisfy all rules in the second set of judgment rules are marked as negative samples, and the negative samples represent non-empty samples.
[0014] A weakly labeled training set is constructed based on the labeled positive and negative samples;
[0015] A machine learning classification model is trained using the weakly labeled training set, and the trained machine learning classification model is used to identify the status of the electricity consumption time series data of the target residential space, and output the vacancy status identifier of the target residential space within the statistical period.
[0016] Furthermore, the process of extracting electricity consumption-related features:
[0017] The system acquires hourly electricity consumption time series data of users, and uses the end of a period as the cutoff boundary to divide the hourly electricity consumption time series data of each user into daily segments. The electricity consumption within the same period after segmentation is accumulated to obtain the cumulative daily electricity consumption value of the user. The cumulative daily electricity consumption value is used as the electricity consumption category feature of the user on the corresponding date.
[0018] Furthermore, the intraday time-series feature extraction process:
[0019] Obtain the user's electricity consumption data within a period of time to construct the daily electricity consumption curve;
[0020] Calculate the interquartile range of electricity consumption over the 24 hours of the day, and set an adaptive amplitude threshold based on the interquartile range;
[0021] Calculate the first-order difference for the internal points within the 24-hour period of the day, and identify candidate peak points and candidate valley points based on the sign relationship between the forward and backward differences. The forward difference is the electricity consumption of the next point minus the electricity consumption of the current point, and the backward difference is the electricity consumption of the current point minus the electricity consumption of the previous point.
[0022] When there are at least three consecutive identical electricity consumption values, the midpoint of the interval of consecutive equal values is added as a candidate peak point or candidate valley point.
[0023] Calculate the peak height for candidate peaks and the valley depth for candidate valleys. Retain candidate points whose peak height or valley depth is greater than or equal to the adaptive amplitude threshold and remove candidate points that do not meet the conditions. If the time interval between two peaks is less than a preset hour threshold, retain the peak with the larger peak height. If the time interval between two valleys is less than the preset hour threshold, retain the valley with the larger valley depth.
[0024] The merged peaks and valleys are sorted by time and forced to alternate, forming a peak-valley alternation sequence;
[0025] Based on the peak-valley alternation sequence and the daily 24-hour electricity consumption, calculate one or more of the following parameters: maximum peak-valley ratio, load factor, proportion of morning and evening peak electricity consumption, proportion of midnight low value, curve flatness index, the period of occurrence of maximum peak value and / or the period of occurrence of minimum valley value, as intraday time series characteristics.
[0026] Furthermore, the diurnal correlation feature extraction process:
[0027] Obtain the user's daily electricity consumption time series, the series including the current day's electricity consumption and historical daily electricity consumption data;
[0028] Obtain the electricity consumption data of the same day last week corresponding to the current day's electricity consumption, and calculate the ratio of the current day's electricity consumption to the electricity consumption of the same day last week, as a comparative feature of daily electricity consumption in the same period.
[0029] Obtain a total dataset consisting of electricity consumption data for the current day and historical days, and calculate the percentile of the electricity consumption for the current day in the total dataset as the global percentile feature of the daily electricity consumption.
[0030] Calculate the user's average daily electricity consumption over the past year, and calculate the ratio of the current day's electricity consumption to the average daily electricity consumption over the past year, as a characteristic of the daily electricity consumption deviation rate;
[0031] Output at least one of the following: daily electricity consumption comparison feature, daily electricity consumption global percentile feature, and daily electricity consumption deviation rate feature, as an inter-day correlation feature.
[0032] Furthermore, the date type feature extraction process:
[0033] Obtain the target date for the feature to be extracted and the corresponding date information of the electricity consumption time-series data of the user to which the target date belongs;
[0034] Based on the weekday attribute of the target date and the preset holiday database, determine whether the target date is a weekday. If it is, generate a binary feature value of 1; otherwise, generate a binary feature value of 0 as the feature of whether it is a weekday.
[0035] Based on the month to which the target date belongs, and according to the preset seasonal classification rules, the target date is classified into one of spring, summer, autumn or winter, and a corresponding seasonal category label is generated as the seasonal feature.
[0036] Output at least one of the "whether it is a weekday" feature and the "to which season" feature as the date type feature.
[0037] Furthermore, the first set of determination rules includes:
[0038] Rule A: Daily electricity consumption is less than the first electricity threshold;
[0039] Rule B: The maximum peak-to-valley ratio of the daily electricity consumption curve is less than the first peak-to-valley ratio threshold;
[0040] Rule C: The flatness index of the daily electricity consumption curve is less than the first flatness threshold.
[0041] Furthermore, the second set of determination rules includes:
[0042] Rule D: Daily electricity consumption exceeds the second electricity threshold, and the second electricity threshold is greater than the first electricity threshold;
[0043] Rule E: The maximum peak-to-valley ratio of the daily electricity consumption curve is greater than the second peak-to-valley ratio threshold;
[0044] Rule F: The proportion of electricity consumption during morning and evening peak hours to the total daily electricity consumption is greater than the threshold for the first peak period.
[0045] Furthermore, the step of using the trained machine learning classification model to perform state identification on the electricity consumption time-series data of the target living space includes:
[0046] Obtain the daily electricity consumption feature vector, intraday time series feature vector, interday correlation feature vector, and date type feature vector for each user. Input the electricity consumption feature vector, intraday time series feature vector, interday correlation feature vector, and date type feature vector into the trained machine learning classification model to obtain the daily empty label of the machine learning classification model. Perform the above prediction process on all days of a single user within a preset time window and output the daily judgment result sequence.
[0047] Based on the daily determination result sequence, starting from the most recent day and going back, count the number of consecutive days that each user has been determined to be idle, and output the sequence of consecutive idle days.
[0048] The idle status level of the corresponding user is determined based on the preset grading threshold range into which the number of consecutive idle days falls.
[0049] The vacancy rate of a region is calculated based on the number of consecutive vacant days or the level of idle status of each user within the target region.
[0050] Furthermore, after outputting the vacancy status indicator of the target living space within the statistical period, the method further includes:
[0051] Based on the vacancy status identifiers of multiple consecutive first statistical periods, the number of consecutive vacancy days is counted using a sliding window.
[0052] Based on a preset time threshold, the idle status of the target residential space is divided into short-term idle, medium-term idle, or long-term idle according to the number of consecutive vacant days.
[0053] Furthermore, after classifying the vacancy status of the target residential space into short-term, medium-term, or long-term vacancy based on the number of consecutive vacancy days according to a preset time threshold, the method further includes:
[0054] Acquire geographic information data across multiple spatial scales;
[0055] Based on the aforementioned geographic information data, the number of vacant users and the total number of users within each spatial range are statistically analyzed according to different vacancy levels, and the regional vacancy rate is calculated.
[0056] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0057] Advantage 1: Daily-level refined judgment, dynamically capturing changes in vacancy status. Existing methods mainly rely on cumulative indicators within a month or statistical period for judgment, failing to achieve daily-level refined judgment. This invention uses the "day" as the smallest unit of measurement, employing a machine learning model to determine the vacancy status of each user every day. This dynamically captures the changes in vacancy status, providing data support for real-time monitoring and early warning.
[0058] Advantage 2: Fully utilizes intraday high-frequency time-series characteristics to improve judgment accuracy. Existing methods do not fully utilize the high-frequency electricity consumption data collected by smart meters. This invention proposes a time-series feature extraction method based on intraday high-frequency data. Through morphological features such as peak-to-valley ratio, load factor, and the proportion of morning and evening peak hours, it effectively distinguishes between "low-energy-consuming resident users" and "vacant residences," significantly improving judgment accuracy.
[0059] Advantage 3: Weakly labeled sample construction method to solve the problem of scarce labeled data. Existing technologies have not proposed an effective sample construction method to solve the problem of scarce labeled data. This invention proposes a multi-rule fusion weakly labeled sample construction method, which can automatically construct high-quality training samples without manual annotation, enabling the application of supervised machine learning methods.
[0060] Advantage 4: A two-stage determination mechanism supports differentiated management of vacancy status. Existing methods directly output binary classification results without distinguishing between different types of vacancy status. This invention proposes a two-stage determination mechanism, from level-based determination to continuous day-based determination, which can distinguish between different scenarios such as short-term vacancy and long-term vacancy, providing more accurate data support for business decisions such as urban planning and real estate regulation. Attached Figure Description
[0061] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:
[0062] Figure 1 This is a flowchart of the residential space status identification method in an embodiment of the present invention. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.
[0064] As one possible implementation method, such as Figure 1 As shown, this embodiment provides a method for identifying the status of residential spaces based on electricity big data. It constructs a four-dimensional feature system including electricity consumption category features, intraday time series features, interday correlation features, and date type features. Weakly labeled training samples are built using a multi-rule fusion method to train a machine learning classification model to achieve daily-level vacancy status determination. The method further classifies user idle status based on the number of consecutive vacant days. The specific implementation process is as follows:
[0065] Step 1: Acquisition and processing of high-frequency electricity consumption time series data.
[0066] Electricity consumption data: High-frequency electricity consumption time series data for each residential user over the past year (365 days) is obtained from the power system, mainly hourly high-frequency electricity consumption sampling data. This data, along with user addresses and spatial location coordinates, provides a basis for subsequent regional vacancy rate statistics.
[0067] The quality of data preprocessing directly affects the accuracy of subsequent feature extraction. The acquisition and processing process of high-frequency electricity consumption time series data is as follows:
[0068] Missing value handling: Obtain the electricity consumption data for each time period of the original high-frequency electricity consumption time series data, and check for missing data in each user's time series. For single-point missing data (such as missing data for a certain 1 hour), linear interpolation is used to fill the missing data; for time periods with continuous missing data exceeding 6 hours, the time period is marked as invalid and will not be included in subsequent feature calculations.
[0069] Outlier detection and correction: Based on box plot method or 3σ principle, identify and eliminate obviously unreasonable electricity consumption values (e.g., negative values, sudden increases exceeding 10 times the user's transformer capacity). For isolated outliers, use the average of adjacent time points for correction.
[0070] Timestamp standardization: Align all users' electricity consumption data to a standard timeline (e.g., Beijing time, hourly time) to eliminate time drift that may exist between different data collection terminals.
[0071] Resampling at a uniform frequency: For high-frequency time-series data, check if the sampling interval is uniform. If it is not uniform, resample to an evenly spaced 1-hour sequence. If the sampling interval is less than 1 hour, aggregate it into an hourly sequence. The final result is continuous, evenly spaced electricity consumption time-series data for each user over 24 hours per day, as shown in Table 1.
[0072] Table 1. Example of hourly high-frequency electricity consumption time series data.
[0073] 1 3 6 7 …… Road No. 104.334 30.552 2 88 69 81 …… Road No. 104.312 30.556
[0074] Step 2: Acquisition of multi-scale spatial range data.
[0075] Multi-scale spatial data: Spatial GIS data covering different statistical ranges such as residential land parcels, communities, streets, districts (cities) and counties. Primarily used for spatial statistics and visualization in subsequent vacancy rate calculations.
[0076] Step 3: Construction of multi-dimensional feature engineering.
[0077] After data preprocessing, features of four dimensions are extracted from the data of each user obtained in step (1): electricity consumption category features, intraday time series features, interday correlation features, and date type features. The details are as follows:
[0078] ① Electricity consumption characteristics:
[0079] Daily electricity consumption Q_day: The cumulative electricity consumption over the 24 hours of the day. The hourly electricity consumption time series data for each user is truncated at 24:00 each day, and then accumulated on a daily basis to obtain daily electricity consumption data, forming a continuous daily electricity consumption characteristic, as shown in Table 2:
[0080] Table 2 Examples of Daily Electricity Consumption
[0081] 1 36 62 …… Road No. 104.334 30.552 2 888 695 …… Road No. 104.312 30.556
[0082] ② Intraday Time Series Characteristics: Time series analysis is performed on the daily electricity consumption curve. An optimized peak-valley detection algorithm is used to identify local maxima (peaks) and local minima (valleys) in the 24-hour electricity consumption curve. The optimized peak-valley detection algorithm mainly considers the large differences in electricity consumption base among different users (such as between elderly people living alone and large households), and fixed thresholds may lead to missed detections or false detections; in order to avoid being too sensitive to users with low electricity consumption, the minimum peak-valley amplitude threshold is not set to a fixed value, but an adaptive threshold is adopted, as detailed below.
[0083] First, calculate the adaptive amplitude threshold: calculate the interquartile range (IQR) of the user's 24 points on the same day, and set the adaptive amplitude threshold θself = 0.25 × IQR.
[0084] Second, calculate the first-order difference to identify candidate extreme points: For the 22 internal points each day, perform backward difference (electricity consumption at this point minus the electricity consumption at the previous point) and forward difference (electricity consumption at the next point minus the electricity consumption at this point). If the forward difference at a point is less than or equal to 0 and the backward difference is greater than 0, it is listed as a candidate peak; if the forward difference at a point is greater than or equal to 0 and the backward difference is less than 0, it is listed as a candidate valley. For consecutive equal values (such as multiple points with zero electricity consumption), to avoid missed detections, if ≥3 consecutive identical values appear, the midpoint of that interval is added as a candidate extreme point.
[0085] Thirdly, amplitude filtering (applying adaptive thresholds):
[0086] Calculate peak height: For each candidate peak point i, if there are identified candidate valley points or sequence endpoints on the left and right sides, the peak height is as follows: ,in Let i be the value of the i-th valley point. The value is the nearest valley point on the left; if none is found, then p1 (the left endpoint) is used. Similarly... The value is the nearest valley point on the right; if none exists, then p24 (the right endpoint) is used. If there are no valley points on either side, then... .
[0087] Calculate valley depth: For each candidate valley point j, the valley depth is calculated as follows: ,in The value of the nearest peak on the left. This is the value of the nearest peak to the right. If there are no adjacent peaks, the difference between this value and the endpoint is used.
[0088] Retention condition: For the calculation results, if the peak height If the value of a candidate point is greater than or equal to the adaptive amplitude threshold θself and the valley value is greater than or equal to the adaptive amplitude threshold θself, then the corresponding candidate point is retained; otherwise, it is discarded.
[0089] Fourth, merge adjacent extreme points: if the time interval between two peaks is less than 2 hours, retain the peak with the larger peak height and remove the other. If the time interval between two valleys is less than 2 hours, retain the valley with the larger valley depth (i.e., the lower valley).
[0090] Fifth, form an alternating peak-valley sequence: Sort the retained peaks and valleys by time, and then force an alternation. If the first one is a peak, the sequence pattern is "peak-valley-peak-valley..."; if the first one is a valley, the pattern is "valley-peak-valley-peak..."; delete redundant points that violate the alternation rule (e.g., if there is no valley between two consecutive peaks, retain the peak with the larger amplitude). The final output is a set of paired or nearly paired peak-valley points.
[0091] Finally, intraday time-series features are calculated based on the previously identified peak and trough data. The specific features are as follows.
[0092] Maximum peak-to-valley ratio: The ratio of the maximum peak value (the largest of the peak values) to the minimum valley value (the smallest of the valley values) of the daily electricity consumption curve.
[0093] Load factor: The ratio of the average electricity consumption per hour to the maximum peak value.
[0094] Peak electricity consumption ratio: The ratio of the sum of electricity consumption during the morning (6:00-9:00) and evening (17:00-23:00) periods to the total daily electricity consumption.
[0095] Midnight low value ratio: The ratio of the sum of electricity consumption during the early morning period (0:00-6:00) to the total daily electricity consumption.
[0096] Curve flatness index: the ratio of the standard deviation to the mean of daily electricity consumption.
[0097] The time periods of the maximum peak and minimum trough are encoded as categorical features (morning peak / evening peak / other time periods).
[0098] ③ Daytime correlation characteristics: These reflect the correlation between daily electricity consumption and electricity consumption on other days.
[0099] Daily electricity consumption comparison: The ratio of daily electricity consumption to the electricity consumption on the same day last week.
[0100] Global percentile of daily electricity consumption: The percentile of the total daily electricity consumption among all selected days.
[0101] Daily electricity consumption deviation rate: The ratio of daily electricity consumption to the average daily electricity consumption over the past year.
[0102] ④ Date type characteristics:
[0103] Whether it is a working day: a binary feature (1 = working day, 0 = non-working day (weekend / holiday));
[0104] Seasonal characteristics: Category features (spring / summer / autumn / winter);
[0105] Finally, the constructed multi-dimensional feature engineering is shown in Table 3:
[0106] Table 3 Examples of Feature Engineering
[0107]
[0108] Step 4: Construct weakly labeled training samples.
[0109] This step addresses the reality of a lack of manually labeled data by proposing an adaptive weak label construction process based on multi-rule fusion. The thresholds involved are calculated based on the statistical distribution characteristics of the user electricity consumption dataset, rather than being subjectively set by humans, ensuring the objectivity and technical consistency of the rules. The specific rules are as follows:
[0110] ① Rules for constructing positive samples (empty samples):
[0111] Rule A: Daily electricity consumption < threshold T_low. T_low is calculated by statistically analyzing the distribution of daily electricity consumption of all residential users over a recent period (e.g., the last 30 days) and taking the 5th percentile; or for each user, taking the 10th percentile of their historical daily electricity consumption during non-holiday periods.
[0112] Rule B: Maximum peak-to-valley ratio < threshold PR_th indicates a flat daily electricity consumption curve. PR_th is calculated by taking the peak-to-valley ratio of the daily electricity consumption curves of the same user over a recent period (e.g., the last 30 days), taking the 25th percentile of these values as a benchmark, and then multiplying it by a coefficient k (this coefficient is determined by experimental verification of sensitivity to vacancy conditions).
[0113] Rule C: Curve flatness index < threshold FI_th indicates minimal fluctuation in electricity consumption. FI_th is calculated by adding 0.5 times the standard deviation to the mean of the coefficient of variation (CV), where CV = σ / μ (where σ is the standard deviation and μ is the mean). The statistical window is the CV distribution of non-vacant suspected days in the user's recent period (e.g., the last 30 days).
[0114] A high-confidence positive sample must simultaneously satisfy rules A, B, and C, and is then marked as "empty" (label=1).
[0115] ② Rules for constructing negative samples (non-empty samples):
[0116] Rule D: Daily electricity consumption > threshold T_high. T_high is calculated by taking the 95th percentile of the daily electricity consumption of all users in a recent period (e.g., the last 30 days), or the 90th percentile of the user's historical daily electricity consumption.
[0117] Rule E: Maximum peak-to-valley ratio > threshold PR_high indicates the presence of significant peak-to-valley characteristics. PR_high is calculated by taking the peak-to-valley ratio from the daily electricity consumption curves of the same user over a recent period (e.g., the last 30 days), taking the 75th percentile of these values, and then multiplying it by a coefficient k (this coefficient is determined by experimental verification of sensitivity to vacancy conditions).
[0118] Rule F: Peak electricity consumption ratio > threshold R_th indicates the existence of a typical residential pattern. R_th is calculated by manually surveying a small number (e.g., 100) of non-vacant days during predefined peak hours (6:00-9:00) and evening peak hours (17:00-23:00), calculating the average and standard deviation of the peak electricity consumption ratio for the selected sample. Then, R_th = μ - 0.5σ, ensuring that the ratio for users with normal residential patterns is typically higher than this threshold.
[0119] High-confidence negative samples must simultaneously satisfy rules D, E, and F, and are then marked as "non-empty" (label=0).
[0120] ③ Other sample processing: Samples that do not meet the above conditions will not be included in the training set for the time being.
[0121] Step 5: Train the residential daily electricity usage vacancy classification model.
[0122] This step uses the feature vector constructed in the previous steps as input and the weak label samples generated in step (4) as training data to train a binary classification model to determine whether any user is in an idle state on any day.
[0123] This invention does not limit the specific type of classification model and can employ mainstream supervised learning classifiers such as Support Vector Machines, Random Forests, and Gradient Boosting Trees (e.g., XGBoost, LightGBM). The following explanation uses a Random Forest model as an example, but the scope of this invention is not limited thereto.
[0124] ① Random forest model parameter settings (initial values, which can be optimized through grid search):
[0125] Number of decision trees n_estimators: 100-500;
[0126] Maximum depth (max_depth): 10-30;
[0127] Minimum number of sample splits (min_samples_split): 2-10;
[0128] Minimum number of leaf node samples (min_samples_leaf): 1-5;
[0129] Feature sampling ratio max_features: sqrt or log2;
[0130] ② Training process:
[0131] The weakly labeled training sample set is randomly divided into a training set and a validation set in a 7:3 ratio (or cross-validation is used).
[0132] The selected classification model is trained using the training set. Hyperparameters (such as the number of trees, maximum depth, regularization coefficient, etc.) are determined by grid search or Bayesian optimization on the validation set, with F1 score or AUC as the optimization objective.
[0133] After training, evaluate the model performance on the validation set and calculate metrics such as accuracy, precision, recall, and F1 score to ensure that the model has good generalization ability.
[0134] ③ Model output:
[0135] Binary classification result: empty (1) or non-empty (0);
[0136] Step 6: Determine the daily-scale vacancy status;
[0137] For each user's feature vector for each day, a pre-trained random forest model is used for prediction. The daily decision sequence is output as: {d1,d2,d3,...,dn}, where di∈{0,1}. The specific prediction process is as follows:
[0138] ①Predictive input: For any user's electricity consumption data on a certain day, extract its feature vector according to step (3) and input it into the model.
[0139] ② Model Prediction: Load the classification model trained in step 5. Predict the daily empty label for each user based on the daily input feature vector.
[0140] ③ Output sequence: For a single user, within the time window [1,n], output the daily judgment result sequence: {d1,d2,d3,...,dn}, where di∈{0,1}.
[0141] Step 7: Count the number of consecutive vacant days.
[0142] Based on the daily determination results, the number of consecutive idle days for each user is counted.
[0143] (1) Sliding window statistics: Backtracking from the most recent day, count the number of consecutive days that are judged as "vacant".
[0144] (2) Output the sequence of consecutive idle days: {T1,T2,...,Tm}, where Ti represents the current consecutive idle days of the i-th user.
[0145] Step 8: Determine the level of idle status.
[0146] Based on the number of consecutive idle days, the user's idle status level is determined according to the tiered threshold, as shown in Table 4:
[0147] Table 4. Determination of User Idle Status Level
[0148] Short-term idle 15-30 days Business trips or travel Mid-term idle 30-90 days May be seasonally vacant Long-term idle ≥90 days High probability of long-term vacancy
[0149] Step 9: Regional vacancy rate statistics
[0150] The calculation of regional vacancy rate mainly includes: overall regional vacancy rate = number of vacant users / total number of users in the region × 100%. Different idle levels of vacancy rate can be calculated according to actual needs.
[0151] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for identifying the status of residential spaces based on big data of electricity, characterized in that, The specific steps include the following: Acquire electricity consumption time-series data of the target living space collected at sampling time intervals within the statistical period; Based on the electricity consumption time-series data, extract electricity consumption category features, intraday time-series features, interday correlation features, and date type features; Based on electricity consumption characteristics and intraday time series characteristics, a multi-dimensional judgment threshold is obtained to determine whether the state is vacant or not. Daily samples that simultaneously satisfy all rules in the first set of judgment rules are marked as positive samples, and the positive samples represent empty samples. Daily samples that simultaneously satisfy all rules in the second set of judgment rules are marked as negative samples, and the negative samples represent non-empty samples. A weakly labeled training set is constructed based on the labeled positive and negative samples; A machine learning classification model is trained using the weakly labeled training set, and the trained machine learning classification model is used to identify the status of the electricity consumption time series data of the target residential space, and output the vacancy status identifier of the target residential space within the statistical period.
2. The method for identifying the status of residential spaces based on power big data according to claim 1, characterized in that, The process of extracting electricity consumption features: The system acquires hourly electricity consumption time series data of users, and uses the end of a period as the cutoff boundary to divide the hourly electricity consumption time series data of each user into daily segments. The electricity consumption within the same period after segmentation is accumulated to obtain the cumulative daily electricity consumption value of the user. The cumulative daily electricity consumption value is used as the electricity consumption category feature of the user on the corresponding date.
3. The method for identifying the status of residential spaces based on power big data according to claim 1, characterized in that, The intraday time-series feature extraction process: Obtain the user's electricity consumption data within a period of time to construct the daily electricity consumption curve; Calculate the interquartile range of electricity consumption over the 24 hours of the day, and set an adaptive amplitude threshold based on the interquartile range; Calculate the first-order difference for the internal points within the 24-hour period of the day, and identify candidate peak points and candidate valley points based on the sign relationship between the forward and backward differences. The forward difference is the electricity consumption of the next point minus the electricity consumption of the current point, and the backward difference is the electricity consumption of the current point minus the electricity consumption of the previous point. When there are at least three consecutive identical electricity consumption values, the midpoint of the interval of consecutive equal values is added as a candidate peak point or candidate valley point. Calculate the peak height for candidate peaks and the valley depth for candidate valleys. Retain candidate points whose peak height or valley depth is greater than or equal to the adaptive amplitude threshold and remove candidate points that do not meet the conditions. If the time interval between two peaks is less than a preset hour threshold, retain the peak with the larger peak height. If the time interval between two valleys is less than the preset hour threshold, retain the valley with the larger valley depth. The merged peaks and valleys are sorted by time and forced to alternate, forming a peak-valley alternation sequence; Based on the peak-valley alternation sequence and the daily 24-hour electricity consumption, calculate one or more of the following parameters: maximum peak-valley ratio, load factor, proportion of morning and evening peak electricity consumption, proportion of midnight low value, curve flatness index, the period of occurrence of maximum peak value and / or the period of occurrence of minimum valley value, as intraday time series characteristics.
4. The method for identifying the status of residential spaces based on power big data according to claim 1, characterized in that, The process of extracting daytime correlation features: Obtain the user's daily electricity consumption time series, the series including the current day's electricity consumption and historical daily electricity consumption data; Obtain the electricity consumption data of the same day last week corresponding to the current day's electricity consumption, and calculate the ratio of the current day's electricity consumption to the electricity consumption of the same day last week, as a comparative feature of daily electricity consumption in the same period. Obtain a total dataset consisting of electricity consumption data for the current day and historical days, and calculate the percentile of the electricity consumption for the current day in the total dataset as the global percentile feature of the daily electricity consumption. Calculate the user's average daily electricity consumption over the past year, and calculate the ratio of the current day's electricity consumption to the average daily electricity consumption over the past year, as a characteristic of the daily electricity consumption deviation rate; Output at least one of the following: daily electricity consumption comparison feature, daily electricity consumption global percentile feature, and daily electricity consumption deviation rate feature, as an inter-day correlation feature.
5. The method for identifying the status of residential spaces based on power big data according to claim 1, characterized in that, The date type feature extraction process: Obtain the target date for the feature to be extracted and the corresponding date information of the electricity consumption time-series data of the user to which the target date belongs; Based on the weekday attribute of the target date and the preset holiday database, determine whether the target date is a weekday. If it is, generate a binary feature value of 1; otherwise, generate a binary feature value of 0 as the feature of whether it is a weekday. Based on the month to which the target date belongs, and according to the preset seasonal classification rules, the target date is classified into one of spring, summer, autumn or winter, and a corresponding seasonal category label is generated as the seasonal feature. Output at least one of the "whether it is a weekday" feature and the "to which season" feature as the date type feature.
6. The method for identifying the status of residential spaces based on power big data according to claim 1, characterized in that, The first set of determination rules includes: Rule A: Daily electricity consumption is less than the first electricity threshold; Rule B: The maximum peak-to-valley ratio of the daily electricity consumption curve is less than the first peak-to-valley ratio threshold; Rule C: The flatness index of the daily electricity consumption curve is less than the first flatness threshold.
7. The method for identifying the status of residential spaces based on power big data according to claim 1, characterized in that, The second set of determination rules includes: Rule D: Daily electricity consumption exceeds the second electricity threshold, and the second electricity threshold is greater than the first electricity threshold; Rule E: The maximum peak-to-valley ratio of the daily electricity consumption curve is greater than the second peak-to-valley ratio threshold; Rule F: The proportion of electricity consumption during morning and evening peak hours to the total daily electricity consumption is greater than the threshold for the first peak period.
8. The method for identifying the status of residential spaces based on power big data according to claim 1, characterized in that, The process of identifying the state of the electricity consumption time-series data of the target residential space using the machine learning classification model trained through training includes: Obtain the daily electricity consumption feature vector, intraday time series feature vector, interday correlation feature vector, and date type feature vector for each user. Input the electricity consumption feature vector, intraday time series feature vector, interday correlation feature vector, and date type feature vector into the trained machine learning classification model to obtain the daily empty label of the machine learning classification model. Perform the above prediction process on all days of a single user within a preset time window and output the daily judgment result sequence. Based on the daily determination result sequence, starting from the most recent day and going back, count the number of consecutive days that each user has been determined to be idle, and output the sequence of consecutive idle days. The idle status level of the corresponding user is determined based on the preset grading threshold range into which the number of consecutive idle days falls. The vacancy rate of a region is calculated based on the number of consecutive vacant days or the level of idle status of each user within the target region.
9. The method for identifying the status of residential spaces based on power big data according to claim 8, characterized in that, After outputting the vacancy status indicator of the target residential space within the statistical period, the method further includes: Based on the vacancy status identifiers of multiple consecutive first statistical periods, the number of consecutive vacancy days is counted using a sliding window. Based on a preset time threshold, the idle status of the target residential space is divided into short-term idle, medium-term idle, or long-term idle according to the number of consecutive vacant days.
10. The method for identifying the status of residential spaces based on power big data according to claim 9, characterized in that, After classifying the vacancy status of the target residential space into short-term, medium-term, or long-term vacancy based on the number of consecutive vacancy days according to a preset time threshold, the method further includes: Acquire geographic information data across multiple spatial scales; Based on the aforementioned geographic information data, the number of vacant users and the total number of users within each spatial range are statistically analyzed according to different vacancy levels, and the regional vacancy rate is calculated.