Multi-source data driven homestay electricity consumption behavior intelligent identification method and system
Through the intelligent identification method driven by multi-source data, the homestay electricity consumption behavior characteristic data is obtained, and multi-layer model modeling and weighting fusion is carried out, which solves the problem of inefficient traditional manual investigation and achieves efficient and accurate power consumption prediction and improvement of power supply service quality.
Patent Information
- Application Number
- CN202510310898.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-08-12
AI Technical Summary
The existing methods for identifying electricity consumption behavior in homestays rely on manual investigation, which is inefficient and difficult to grasp the changes in electricity consumption categories in a timely and accurate manner, affecting the efficiency and quality of power supply services.
Using a multi-source data-driven intelligent identification method, by obtaining the multi-source electricity consumption behavior characteristic data of the homestay, using multiple algorithms for preliminary modeling, forming a high-dimensional feature matrix, and using sub-level models for weighted fusion to achieve power consumption prediction.
It improves the accuracy and stability of homestay electricity consumption behavior recognition, reduces the workload of manual verification, and improves the intelligent operation and maintenance capabilities of the power supply system.
Smart Images

Figure CN120470296A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data recognition technology, and in particular to a method and system for intelligently identifying electricity usage behavior in homestays driven by multi-source data. Background Art
[0002] After in-depth analysis and research on the technical fields related to the present invention and utility model, we found that the existing solutions have the following major shortcomings:
[0003] 1. Limitations of traditional methods: Traditional user electricity category identification relies on manual full-area surveys. This method is time-consuming and labor-intensive, inefficient, and difficult to accurately and timely identify situations where users have privately changed their electricity category, which seriously affects the efficiency and quality of power supply services.
[0004] 2. Necessity of intelligent transformation: Due to problems with traditional methods, the work of identifying user electricity usage categories must be transformed from "manual" to "intelligent" to improve the efficiency and quality of power supply services.
[0005] Applying machine learning to traditional user electricity classification, we've developed an intelligent homestay identification method to improve the efficiency and quality of power supply services. Accurately understanding the distribution and load characteristics of homestay users in each substation supports marketing operations such as business expansion management and electricity price and fee management, helping to better meet the electricity needs of customers. Summary of the Invention
[0006] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of this application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.
[0007] In view of the above existing problems, the present invention is proposed.
[0008] Therefore, the present invention provides a multi-source data-driven intelligent identification method and system for homestay electricity usage behavior to solve the problem of poor recognition accuracy and stability of homestay electricity usage behavior in the existing technology.
[0009] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0010] In a first aspect, the present invention provides a multi-source data-driven intelligent identification method for electricity usage behavior in homestays, comprising:
[0011] Obtain the multi-source electricity consumption behavior feature data of the homestay and extract the first feature;
[0012] Use multiple algorithms to preliminarily model the first feature and obtain the prediction results of the first-level model;
[0013] The output of the first-layer model is integrated with the original input data to form the first high-dimensional feature matrix;
[0014] Input the first high-dimensional feature matrix into the sub-layer model for training to obtain the prediction result of the sub-layer model;
[0015] The prediction results of the sub-layer model are weighted and fused to obtain the power consumption prediction results.
[0016] As a preferred solution of the multi-source data-driven intelligent identification method for electricity consumption behavior of homestays described in the present invention, the following is a solution:
[0017] The first feature is preliminarily modeled using multiple algorithms, including:
[0018] Select appropriate model algorithms based on the characteristics of different data sources;
[0019] Each model is trained independently and generates predictions for the first-level model.
[0020] As a preferred solution of the multi-source data-driven intelligent identification method for electricity consumption behavior of homestays described in the present invention, the following is a solution:
[0021] The step of fusing the output of the first-layer model with the original input data includes:
[0022] By drawing the power change curve covering different dimensions, key features are identified;
[0023] Use various machine learning algorithms to evaluate feature importance and identify and extract the most critical features;
[0024] The key features are concatenated with the original data set to form the first high-dimensional feature matrix.
[0025] As a preferred solution of the multi-source data-driven intelligent identification method for electricity consumption behavior of homestays described in the present invention, the following is a solution:
[0026] Inputting the first high-dimensional feature matrix into the sub-layer model for training includes:
[0027] Use the same model algorithm as the first-layer model, combined with cross-validation and hyperparameter optimization strategies for training.
[0028] As a preferred solution of the multi-source data-driven intelligent identification method for electricity consumption behavior of homestays described in the present invention, the following is a solution:
[0029] The weighted fusion using the prediction results of the sub-layer model includes:
[0030] By evaluating the performance of each sub-layer model on the validation set, a weight reflecting the importance is calculated for each model;
[0031] According to the calculated weights, clarify the specific weight value of each model;
[0032] The weighted average is used to combine the prediction results of different models and their corresponding weights to output the final power consumption prediction result.
[0033] As a preferred solution of the multi-source data-driven intelligent identification method for electricity consumption behavior of homestays described in the present invention, the following is a solution:
[0034] The first features include time series features, address information features, special date features and historical load features.
[0035] As a preferred solution of the multi-source data-driven intelligent identification method for electricity consumption behavior of homestays described in the present invention, the following is a solution:
[0036] The time series features include capturing the periodic variation pattern of the user's electricity consumption;
[0037] The address information features include converting the geographical location influence into a numerical feature by a first algorithm, and generating a numerical feature column by address vectorization;
[0038] The special date features include extracting typical electricity consumption fluctuation patterns by comparing holidays and tourist peak periods with daily electricity consumption data;
[0039] The historical load characteristics include seasonal decomposition and trend analysis of historical load data.
[0040] In a second aspect, the present invention provides a multi-source data-driven intelligent identification system for electricity usage behavior in homestays, comprising:
[0041] A data acquisition module is used to obtain the multi-source electricity consumption behavior feature data of the homestay and extract the first feature;
[0042] The first-layer model training module is used to use multiple algorithms to perform preliminary modeling on the first feature and obtain the prediction results of the first-layer model;
[0043] The feature fusion module is used to fuse the output results of the first-layer model with the original input data to form the first high-dimensional feature matrix;
[0044] A sub-layer model training module is used to input the first high-dimensional feature matrix into the sub-layer model for training to obtain the prediction result of the sub-layer model;
[0045] The weighted fusion module is used to perform weighted fusion using the prediction results of the sub-layer model to obtain the power consumption prediction results.
[0046] In a third aspect, the present invention provides a computing device, comprising:
[0047] Memory, used to store programs;
[0048] A processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the multi-source data-driven intelligent identification method for homestay electricity usage behavior.
[0049] In a fourth aspect, the present invention provides a computer-readable storage medium, comprising: when the program is executed by a processor, the steps of implementing the multi-source data-driven intelligent identification method for homestay electricity usage behavior.
[0050] The present invention integrates multi-source data, enabling model information maintenance, basic data management, and abnormal user identification. By visualizing data analysis, accurately identifying homestay users, and presenting identification results in multiple dimensions, the application can quickly identify potential homestay users, significantly reducing manual verification workload. In the future, the application is expected to be expanded to other regions, further enhancing the intelligent operation and maintenance capabilities of the power supply system. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:
[0052] Figure 1 A schematic diagram of the basic process of a multi-source data-driven intelligent identification method for electricity consumption behavior in homestays provided by one embodiment of the present invention;
[0053] Figure 2 A model architecture diagram of a multi-source data-driven intelligent identification method for electricity usage behavior in homestays provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0054] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0055] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0056] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0057] The present invention is described in detail with reference to schematic diagrams. For ease of illustration, cross-sectional views of device structures may be partially enlarged and not to scale when describing embodiments of the present invention. Furthermore, the schematic diagrams are merely illustrative and should not limit the scope of the present invention. Furthermore, in actual production, the three-dimensional dimensions of length, width, and depth should be included.
[0058] In the description of the present invention, it should be noted that the terms "upper, lower, inner, and outer" and other references to orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first, second, or third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0059] In this disclosure, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be interpreted broadly. For example, they may refer to fixed, removable, or integral connections. They may also refer to mechanical, electrical, or direct connections, indirect connections through an intermediary, or internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in this disclosure.
[0060] Example 1
[0061] Reference Figure 1 , as one embodiment of the present invention, provides a multi-source data-driven intelligent identification method for electricity usage behavior in homestays, comprising:
[0062] S1: Obtain the multi-source electricity consumption behavior feature data of the homestay and extract the first feature;
[0063] S2: Use multiple algorithms to preliminarily model the first feature and obtain the prediction results of the first-level model;
[0064] S3: Fuse the output of the first-layer model with the original input data to form the first high-dimensional feature matrix;
[0065] S4: Input the first high-dimensional feature matrix into the sub-layer model for training to obtain the prediction result of the sub-layer model;
[0066] S5: Use the prediction results of the sub-layer model for weighted fusion to obtain the power consumption prediction result.
[0067] It should be noted that the present invention solves the accuracy and stability problems in the identification of electricity consumption behavior in homestays by combining prediction with weighted optimization through a hierarchical model. Figure 1 As shown in Figure 3, this integration strategy can fully utilize the advantages of different models, significantly improve the prediction performance, and provide stronger robustness and generalization capabilities.
[0068] Example 2
[0069] Reference Figure 2 , which is an embodiment of the present invention, provides a multi-source data-driven intelligent identification method for electricity consumption behavior of homestays based on the previous embodiment, including:
[0070] In an embodiment of the present application, in step S1, characteristic data of multi-source electricity consumption behavior of a homestay is obtained, and a first feature is extracted, including time series features, address information features, special date features, and historical load features.
[0071] In the embodiment of the present application, the time series features include periodic features and trend features;
[0072] In the embodiment of the present application, the periodic characteristics include daily cycles, weekly cycles, etc. For example, the electricity consumption of a certain homestay increases significantly between 7 pm and 10 pm every day.
[0073] In the embodiment of the present application, the trend feature includes a trend of change in electricity consumption over time, such as seasonal fluctuations or long-term growth / decline trends.
[0074] In the embodiment of the present application, the address information features include geographic location vectorization and spatial correlation;
[0075] In an embodiment of the present application, geographic location vectorization includes converting the specific address of the homestay (such as "No. xx, xx Community, xx Street, xx City") into a numerical feature vector (for example, through Word2Vec or other vectorization technologies) so that the model can understand and utilize the impact of geographic location on electricity consumption behavior.
[0076] In the embodiment of the present application, the spatial correlation includes analyzing the spatial distribution of B&B addresses and finding that B&Bs in certain areas may have similar electricity usage patterns.
[0077] In this embodiment, the special date feature includes comparing electricity consumption changes between weekdays, holidays, and peak tourist seasons to extract typical electricity consumption fluctuation patterns. For example, during the May Day holiday, electricity consumption in homestays increased significantly compared to normal times. The impact of specific events (such as festivals, local events, etc.) on electricity consumption is identified and converted into quantifiable features.
[0078] In this embodiment of the present application, historical load characteristics include analyzing electricity consumption fluctuations in different time periods and seasons. For example, summer electricity consumption is 20% higher than winter, or winter electricity consumption is 10% higher than summer. By performing seasonal decomposition and trend analysis on historical load data, the cyclical and long-term trends of electricity consumption are revealed.
[0079] In the embodiment of the present application, the address information feature includes converting the geographical location influence into a numerical feature by a first algorithm, and generating a numerical feature column by address vectorization;
[0080] In the embodiment of the present application, the special date features include extracting typical electricity consumption fluctuation patterns by comparing holidays and peak travel periods with daily electricity consumption data;
[0081] In an embodiment of the present application, the historical load characteristics include seasonal decomposition and trend analysis of historical load data.
[0082] In an optional embodiment, the first algorithm may be Word2vec technology, GloVe technology, or FastText technology;
[0083] In an optional embodiment, the Word2vec technology includes converting the homestay address information "No. xx, xx Community, xx Street, xx City" into a numerical feature vector, thereby integrating it into a multi-source data feature set to improve the model's spatial correlation analysis and prediction accuracy of electricity consumption behavior.
[0084] In an optional embodiment, GloVe technology includes converting the homestay address information "No. xx, xx Community, xx Street, xx City" into a numerical feature vector based on global word co-occurrence statistics, and integrating it into a multi-source data feature set to enhance the model's capture of spatial correlation and the prediction accuracy of electricity consumption behavior.
[0085] In an optional embodiment, FastText technology includes converting the homestay address information "No. xx, xx Community, xx Street, xx City" into a numerical feature vector based on subword information, effectively processing unregistered words and enhancing the capture of spatial correlation, thereby integrating it into a multi-source data feature set and improving the model's prediction accuracy for electricity consumption behavior.
[0086] It should be noted that B&B address information is typically standardized and does not require complex subword processing (such as that provided by FastText) or global statistical information (such as that provided by GloVe). Therefore, Word2Vec technology was chosen in this solution based on its advantages in processing local context information, efficiency, and flexibility, as well as its applicability to the current application scenario.
[0087] In this embodiment of the present application, step S1 acquires multi-source electricity usage behavior feature data for homestays. Extracting the first feature involves effectively integrating information from different data sources through multi-source data fusion and feature engineering, rather than processing a single data source in isolation. Specifically, user electricity usage data is analyzed through time series analysis to capture its periodic variations, such as daily and weekly cycles. Address information is vectorized to convert the influence of the user's geographic location into numerical features, enabling collaborative modeling with other features. For example, the electricity usage addresses of homestay users may be spatially clustered. Address vectorization generates numerical feature columns, facilitating subsequent model learning of spatial correlations. To account for electricity usage fluctuations during holidays and peak travel periods, the present invention compares specific dates with daily electricity usage data to extract typical patterns of electricity usage fluctuations. These patterns are further converted into features to identify homestay users' electricity usage behavior during specific time periods. Historical load data is analyzed through seasonal decomposition and trend analysis to reveal seasonal differences and long-term trends in electricity usage. This information is then incorporated into the model through feature engineering to enhance the ability to predict future electricity demand. Through deep fusion and feature extraction of multi-source data, combined with data standardization and cleaning, the model's prediction accuracy and robustness were significantly improved. This comprehensive data processing approach not only enhanced the feature representation capabilities but also provided the model with richer input information, resulting in excellent performance in the task of identifying electricity usage behavior in homestays.
[0088] In the embodiment of the present application, in step S2, multiple algorithms are used to perform preliminary modeling on the first feature to obtain the prediction results of the first-level model, including:
[0089] During the first-level model training phase, multiple base models (see Table 1) are used to perform parallel modeling and prediction based on the characteristics of different data sources. This approach leverages the strengths of each model and improves overall prediction performance. Specifically, the first-level model selection is optimized based on the characteristics of the data, ensuring that each data type receives the most appropriate modeling. During training, each model independently models and generates corresponding predictions. These predictions serve not only as the output of the first-level model but are also fused with the original input data through feature concatenation to form a new high-dimensional feature matrix, providing richer input information for the second-level models.
[0090] Table 1 Base model table
[0091] Model Name Model Features Decision Tree Simple and intuitive, computationally efficient, and easy to understand. Random Forest By integrating multiple decision trees, the risk of overfitting is reduced and the generalization ability is strong. Gradient Boosting Higher prediction accuracy, improving performance by gradually improving weak models. XGBoost Efficient, strong regularization mechanism, suitable for large-scale data processing. LightGBM Fast, suitable for large-scale data, and good at processing sparse data. CatBoost Excellent category feature processing capabilities and fast training speed. Support Vector Machine (SVM) It is strong in processing high-dimensional data and suitable for classification tasks. K-nearest neighbor (KNN) Simple and intuitive, no training required, instance-based classification method. Neural Network It is suitable for processing complex nonlinear relationships and can mine deeper features.
[0092] In the embodiments of the present application, for region A, which has a large amount of data and high complexity, we give priority to using algorithms such as neural networks, XGBoost and LightGBM. These algorithms are considered ideal for data modeling in this region because of their significant advantages in mining nonlinear relationships in data and efficiently processing large-scale data. For region B, which has a smaller amount of data but obvious seasonal characteristics, we tend to use models such as random forests, gradient boosting trees and support vector machines. These models have unique advantages in capturing the periodic characteristics of data, while helping to reduce the risk of overfitting due to limited data volume. When processing data from region C, which has a small amount of data and relatively stable features, we choose models such as CatBoost, K-nearest neighbor algorithms and decision trees. These models have excellent performance in processing categorical features and are computationally efficient, so they are suitable for data modeling in this region. For region D, which has a small amount of data but may contain complex relationships, we comprehensively considered efficiency and complex pattern modeling capabilities and decided to use algorithms such as XGBoost, LightGBM and neural networks. These algorithms can effectively capture the complex relationships in the data while maintaining efficiency, thereby achieving accurate prediction of data in region D.
[0093] In the embodiment of the present application, a single model training step:
[0094] 1. Data Collection:
[0095] User electricity consumption data: including hourly, daily, and monthly electricity consumption (unit: kWh), reflecting the electricity consumption behavior characteristics of homestay users.
[0096] Address data: Using Word2vec technology, the homestay address information "No. xx, xx Community, xx Street, xx City" is transformed into a vector [-0.14852104 -0.02127483 -0.01589894 -0.05125818 0.19365442].
[0097] Holiday and peak tourist season data: Data analysis tools are used to analyze changes in electricity consumption during weekdays, holidays, and peak tourist seasons. For example, during the May Day holiday in a certain region, electricity consumption at homestays increased significantly compared to weekdays.
[0098] Historical load data: Analyze the fluctuations in electricity consumption among B&B users over different time periods and seasons. Summer electricity consumption in region xx is 20% higher than winter, while winter electricity consumption in region xx is 10% higher than summer.
[0099] 2. Data cleaning:
[0100] Eliminate redundant or abnormal data, such as filling missing values and removing noise data. For example, if a homestay user's electricity consumption is 0 kWh on a certain day, this data is considered abnormal and removed.
[0101] Use interpolation to fill missing values:
[0102]
[0103] Among them, y i is a missing value, y i -1 and y i +1 for the preceding and following data points.
[0104] 3. Feature extraction and standardization:
[0105] Extract statistical characteristics of electricity consumption behavior, such as the daily and monthly average differences and average values.
[0106] Normalize the eigenvalues:
[0107]
[0108] Among them, x norm is the normalized value, x min and x max are the minimum and maximum values of the feature, respectively.
[0109] 4. Data Segmentation: The preprocessed dataset was randomly divided into training and test sets with a ratio of 8:2. Furthermore, the ratio of the training and test sets was set to 4:1 based on the real-life distribution of homestay users and ordinary users to ensure the representativeness and generalization ability of the dataset.
[0110] 5. Model training and evaluation:
[0111] The processed data sources and extracted feature values are trained on several specific models, and the precision, recall, and F1 value of the models are statistically evaluated on the test set.
[0112] In an embodiment of the present application, in step S3, the output result of the first-layer model is fused with the original input data to form a first high-dimensional feature matrix, including the fusion of the first-layer output result and the original input data through feature splicing technology. The core advantage of this feature fusion strategy is that it can completely retain the basic information of the original data, thereby ensuring the comprehensiveness and integrity of the data. During this fusion process, we implemented in-depth feature engineering modeling on the original data, aiming to analyze and quantify the influence of key features on the prediction results. Through this process, we can extract multiple key features that have a significant impact on the prediction results, and accordingly eliminate those features that contribute less to the prediction, thereby providing a more expressive and predictive feature set for the next layer model. Specifically, the key operations of this layer include: carrying out detailed feature engineering modeling work, and deeply exploring potential key features by drawing power change curves covering different dimensions; using the feature importance evaluation method of each algorithm to conduct in-depth analysis of all the features selected by the first layer, so as to identify and extract the most critical features; finally, merging these key features with the original data set to form the input data set for training the next layer model.
[0113] In the embodiment of the present application, in step S4, the first high-dimensional feature matrix is input into the sub-layer model for training, and obtaining the prediction result of the sub-layer model includes the following steps:
[0114] By drawing the power change curve covering different dimensions, key features are identified;
[0115] Use various machine learning algorithms to evaluate feature importance and identify and extract the most critical features;
[0116] The key features are concatenated with the original data set to form the first high-dimensional feature matrix.
[0117] In an embodiment of the present application, in step S4, the first high-dimensional feature matrix is input into the sub-layer model for training, and the prediction results of the sub-layer model obtained include the algorithm integrating the new feature matrix (including the original features and the prediction results of the first-layer model) on the basis of the output of the first-layer model to further optimize the model performance. Specifically, this layer continues to use the same model algorithm as the first layer for iterative training, but the difference is that it combines cross-validation, hyperparameter optimization strategy and new features after feature selection to ensure the performance consistency of the model on the training set and the validation set, thereby improving the generalization ability.
[0118] In this embodiment of the present application, step S5 utilizes the prediction results of the sub-layer models for weighted fusion to obtain the power consumption forecast result. To maximize the accurate prediction capabilities of the sub-layer models, a weighted fusion strategy is introduced when constructing the model's final output. This strategy aims to integrate the advantages of each model by rationally assigning weights to the different sub-layer models, thereby improving overall forecasting performance and stability.
[0119] 1. Weight calculation:
[0120] Assign weights to each model based on validation set performance:
[0121]
[0122] Among them, W i is the weight of the i-th model, Error i is the validation error of the model.
[0123] 2. Weight distribution:
[0124] In the prediction of homestays in area A, the weight of LightGBM is 0.5, XGBoost is 0.3, and Random Forest is 0.2.
[0125] 3. Final prediction value:
[0126] Fusion of multiple models to output the final prediction value:
[0127]
[0128] Among them, y predict is the final predicted value, w i is the weight, y i is the predicted value of the ith model.
[0129] This embodiment also provides a multi-source data-driven intelligent identification system for electricity usage behavior in homestays, including:
[0130] A data acquisition module is used to obtain the multi-source electricity consumption behavior feature data of the homestay and extract the first feature;
[0131] The first-layer model training module is used to use multiple algorithms to perform preliminary modeling on the first feature and obtain the prediction results of the first-layer model;
[0132] The feature fusion module is used to fuse the output results of the first-layer model with the original input data to form the first high-dimensional feature matrix;
[0133] A sub-layer model training module is used to input the first high-dimensional feature matrix into the sub-layer model for training to obtain the prediction result of the sub-layer model;
[0134] The weighted fusion module is used to perform weighted fusion using the prediction results of the sub-layer model to obtain the power consumption prediction results.
[0135] Furthermore, it also includes:
[0136] Memory, used to store programs;
[0137] A processor is used to load the program to execute the multi-source data-driven intelligent identification method for homestay electricity usage behavior.
[0138] This embodiment also provides a computer-readable storage medium storing a program. When the program is executed by a processor, the method for intelligently identifying electricity usage behavior of homestays driven by multi-source data is implemented.
[0139] The storage medium proposed in this embodiment and the multi-source data-driven intelligent identification method for electricity consumption behavior of homestays proposed in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0140] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer's floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0141] Example 3
[0142] This is an embodiment of the present invention, which provides a deep brain stimulation surgery and treatment auxiliary method and system based on digital twin technology. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through specific implementation methods and implementation effects.
[0143] The details of this embodiment are as follows:
[0144] 1. Identification of electricity usage behavior of homestays in area A
[0145] Input data: electricity consumption data, address data, and holiday data of 395 households in area A.
[0146] First-layer model: multiple base models (LightGBM, XGBoost, random forest, etc.).
[0147] Sub-layer model: a specific base model.
[0148] Weight distribution: LightGBM (0.5), XGBoost (0.3), Random Forest (0.2).
[0149] Results: 124 homestay users were accurately identified, with an accuracy of 91.21%, a recall rate of 89.10%, and an F1 of 90.32%.
[0150] 2. Identification of electricity usage behavior of B&Bs in Region B
[0151] Input data: electricity consumption data, address data, and holiday data of 500 users in area B.
[0152] First-layer model: multiple base models (LightGBM, XGBoost, random forest, etc.).
[0153] Sub-layer model: specific base model.
[0154] Weight distribution: LightGBM (0.4), CatBoost (0.4), Random Forest (0.2).
[0155] Results: 289 homestay users were accurately identified, with an accuracy of 92%, a recall rate of 88.7%, and an F1 of 90.32%.
[0156] Currently, mainstream modeling paradigms are mostly limited to the use of a single data source and a single model, following standard classification and regression steps that are highly similar to the traditional process of single-model training. However, this paper aims to go beyond this traditional paradigm by deeply exploring the multi-source nature of data and the complementary potential between models, innovatively combining multiple base models with a layered architecture.
[0157] Specifically, the present invention constructs a multi-model parallel hierarchical architecture that not only fully leverages the richness of multi-source data but also enhances the model's adaptability to data complexity through multiple iterative training. During the prediction phase, the present invention employs an innovative weighting strategy to effectively integrate the outputs of multiple base models, resulting in a significant improvement in prediction performance.
[0158] Compared with the single model training of traditional methods, although it can apply different algorithms for modeling, it still has certain limitations in processing highly complex data sets and capturing potential nonlinear relationships in the data. The multi-model hierarchical architecture of the present invention shows greater flexibility and robustness, and can more effectively address the challenges of data diversity and model complementarity. Therefore, the method proposed in this patent is superior to traditional methods in terms of overall performance, adaptability and prediction accuracy, providing new ideas and technical support for research and application in related fields.
[0159] Example 4
[0160] This is an embodiment of the present invention, which provides a multi-source data-driven intelligent identification system for electricity consumption behavior in homestays, including a data acquisition module, a first-level model training module, a feature fusion module, a second-level model training module, and a weighted fusion module;
[0161] In this embodiment, the data acquisition module acquires multi-source electricity usage behavior data from homestays and extracts the first feature. It also collects user electricity usage data, address information, and holiday and peak travel period data. Highly discriminative features are extracted through methods such as time series analysis, address vectorization, holiday feature extraction, and historical load analysis.
[0162] In an embodiment of the present application, the first-layer model training module includes using multiple algorithms to preliminarily model the first feature to obtain the prediction results of the first-layer model; selecting a suitable base model, independently training each model, and generating corresponding prediction results.
[0163] In an embodiment of the present application, the feature fusion module includes fusing the output results of the first-layer model with the original input data to form a first high-dimensional feature matrix; and performing feature splicing on the prediction results of the first-layer model and the original input data to form a new high-dimensional feature matrix.
[0164] In this embodiment of the present application, the sub-layer model training module includes inputting the first high-dimensional feature matrix into the sub-layer model for training to obtain the sub-layer model's prediction results; and then continuing to iteratively train using the same model algorithm as the first layer. This module combines cross-validation, hyperparameter optimization strategies, and new features from feature selection to ensure consistent performance across both the training and validation sets.
[0165] In this embodiment, the weighted fusion module combines the prediction results of the sub-level models to generate a power consumption forecast. The module calculates the weight of each sub-level model and assigns weights based on the performance of the validation set. The prediction results of multiple models are combined using a weighted average method to generate a final power consumption forecast.
[0166] In the embodiment of the present application, the above modules work together to achieve accurate identification and prediction of B&B electricity consumption behavior through multi-source data-driven methods and multi-level model combination prediction and weighted optimization.
[0167] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A multi-source data-driven intelligent identification method for electricity consumption behavior in homestays, characterized by: include: Obtain the multi-source electricity consumption behavior feature data of the homestay and extract the first feature; Use multiple algorithms to model the first feature in parallel to obtain the prediction results of the first-level model; The output of the first-layer model is integrated with the original input data to form the first high-dimensional feature matrix; Input the first high-dimensional feature matrix into the sub-layer model for training to obtain the prediction result of the sub-layer model; The prediction results of the sub-layer model are weighted and fused to obtain the power consumption prediction results.
2. The multi-source data-driven intelligent identification method for electricity consumption behavior in homestays according to claim 1, characterized in that: The first feature is preliminarily modeled using multiple algorithms, including: Select appropriate model algorithms for training based on the characteristics of different data sources; Each model is trained independently and generates predictions for the first-level model.
3. The multi-source data-driven intelligent identification method for electricity usage behavior in homestays according to claim 1 or 2, characterized in that: The step of fusing the output of the first-layer model with the original input data includes: By drawing the power change curve covering different dimensions, key features are identified; Use feature evaluation algorithms to evaluate feature importance, identify and extract the most critical features; The key features are concatenated with the original data set to form the first high-dimensional feature matrix.
4. The multi-source data-driven intelligent identification method for electricity consumption behavior in homestays according to claim 3, characterized in that: Inputting the first high-dimensional feature matrix into the sub-layer model for training includes: Set the objective function of the sub-level model; Identify hyperparameters that need to be optimized; Use cross-validation to evaluate and select the best hyperparameter combination; The objective function of the sub-layer model is expressed as: Among them, y i and are the actual value and the predicted value respectively, R(θ) is the regularization term used to control the complexity of the model, S(θ) is the spatial correlation penalty term, H(θ) is the holiday effect penalty term, and w ij Represents a weight matrix defined based on geographic distance, and Holidays represents a dataset of holidays.
5. The multi-source data-driven intelligent identification method for electricity consumption behavior in homestays according to claim 4 is characterized by: The weighted fusion using the prediction results of the sub-layer model includes: By evaluating the performance of each sub-layer model on the validation set, a weight reflecting the importance is calculated for each model; According to the calculated weights, clarify the specific weight value of each model; The weighted average is used to combine the prediction results of different models and their corresponding weights to output the final power consumption prediction result.
6. The multi-source data-driven intelligent identification method for electricity consumption behavior in homestays according to claim 5, characterized in that: The first features include time series features, address information features, special date features and historical load features.
7. The multi-source data-driven intelligent identification method for electricity consumption behavior in homestays according to claim 6, characterized in that: The time series features include capturing the periodic variation pattern of the user's electricity consumption; The address information features include converting geographic location influence into numerical features through a first algorithm, and generating numerical feature columns through address vectorization; The special date features include extracting typical electricity consumption fluctuation patterns by comparing holidays and tourist peak periods with daily electricity consumption data; The historical load characteristics include seasonal decomposition and trend analysis of historical load data.
8. A system based on the multi-source data-driven intelligent identification method for electricity usage behavior in homestays according to claim 1, characterized in that: A data acquisition module is used to obtain the multi-source electricity consumption behavior feature data of the homestay and extract the first feature; The first-layer model training module is used to perform preliminary modeling on the first feature using multiple algorithms to obtain the prediction results of the first-layer model; The feature fusion module is used to fuse the output results of the first-layer model with the original input data to form the first high-dimensional feature matrix; A sub-layer model training module is used to input the first high-dimensional feature matrix into the sub-layer model for training to obtain the prediction result of the sub-layer model; The weighted fusion module is used to perform weighted fusion using the prediction results of the sub-layer model to obtain the power consumption prediction results.
9. A computing device, characterized in that include: Memory, used to store programs; A processor is used to load the program to execute the steps of the multi-source data-driven intelligent identification method for electricity consumption behavior of homestays as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by the processor, the steps of the multi-source data-driven intelligent identification method for electricity consumption behavior of homestays as described in any one of claims 1 to 7 are implemented.