Electric vehicle individual user charging demand prediction method and device

Through HDBSCAN and K-means clustering combined with LightGBM algorithm, a multi-source data set is constructed, which solves the data heterogeneity and exogenous variable interference problems in the charging demand prediction of individual electric vehicle users, realizes high-precision charging demand prediction, and supports the optimization planning of the power grid and charging piles.

CN120542618APending Publication Date: 2025-08-26TIANJIN UNIV +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510556653.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-23
Filing Date
2025-04-29
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the charging needs of individual users of electric vehicles. Especially when considering exogenous variable interference, there are problems of insufficient data representation and universality, which leads to an increase in the peak load of the power grid and affects the safety of the power grid.

Method used

Data layer clustering based on HDBSCAN algorithm and user layer clustering of K-means algorithm are used, and time series coding and exogenous variables are combined to build a multi-source data set, and prediction is carried out through the LightGBM algorithm to improve prediction accuracy and scene adaptability.

Benefits of technology

It significantly improves the accuracy and stability of charging demand prediction for individual electric vehicle users, reduces the cost of manual parameter adjustment, provides high-precision charging demand prediction support, and promotes grid scheduling and charging pile layout planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542618A_ABST
    Figure CN120542618A_ABST
Patent Text Reader

Abstract

The invention discloses an electric vehicle individual user charging demand prediction method and device, and the method comprises the steps: constructing a multi-source data set based on the charging behavior data, weather data and holiday and festival data of EV users in a community; cleaning and standardizing the multi-source data set through a multi-source data preprocessing model to obtain a processed multi-source data set; performing clustering processing on the processed multi-source data set through a double-layer clustering model to obtain charging type features and user type features; extracting input features from the processed multi-source data set; and inputting the charging type characteristics, the user type characteristics and the input characteristics into a charging demand prediction model, and performing prediction processing through a LightGBM algorithm to obtain a user charging demand. According to the method, interference of exogenous variables can be reasonably incorporated, the prediction precision is improved, and the problems in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric vehicle charging demand prediction, and in particular to a method and device for predicting the charging demand of individual electric vehicle users. Background Art

[0002] Energy shortages and environmental pressures are driving the rapid development of the global electric vehicle (EV) industry. By 2024, China's new energy vehicle sales will exceed 12 million, accounting for over 70% of the global total. With the widespread adoption of EVs, the uncontrolled surge in charging demand has led to increasing peak loads on the power grid, creating a "peak upon peak" phenomenon, significantly widening the peak-to-valley difference and seriously threatening the operational safety of the power grid. According to forecasts, by 2030, China's new energy vehicle ownership is expected to exceed 100 million. Under extreme conditions, the EV charging load will approach 25% of the country's total installed capacity. Communities are the primary charging locations for EVs. According to statistics, 65% of EV users prefer to recharge in their communities. However, peak EV charging periods overlap significantly with peak community electricity consumption periods, resulting in severe overload risks for distribution transformers and reduced power quality. This has become a key bottleneck hindering the large-scale deployment of EVs.

[0003] With the nationwide development and improvement of the National New Energy Vehicle Regulatory Platform, the Smart Internet of Vehicles Platform, and the New Energy Vehicle Three-Level Regulatory Platform, a large amount of charging behavior data has been accumulated, laying a solid foundation for predicting EV user charging demand. Currently, research primarily focuses on predicting charging demand in public charging stations and EV clusters. However, the charging behavior revealed in these scenarios differs significantly from that of individual EV users in the community, necessitating in-depth research on predicting charging demand for individual users within the community. Existing literature has initially explored methods for predicting individual user charging demand, but the charging behavior data used suffers from limited time span and scale. Furthermore, this data primarily originates from manual records or limited EV collection, limiting its representativeness and generalizability. Furthermore, these studies focus on the impact of prediction results on EV charging scheduling and battery health, but have yet to systematically explore the impact of individual user charging behavior variability on prediction results. Research has shown that predicting charging demand for individual EV users is not only influenced by the highly personalized nature of their charging behavior but also requires consideration of exogenous variables such as weather and holidays. The interaction of these complex factors significantly increases the difficulty of accurately predicting individual user charging demand.

[0004] Therefore, how to invent a method to predict the charging demand of individual electric vehicle users, reasonably incorporate the interference of exogenous variables, and improve the accuracy of prediction has become an urgent problem to be solved. Summary of the Invention

[0005] To this end, the present invention provides a method and device for predicting the charging demand of individual electric vehicle users. This method uses a three-level cleaning process to construct a highly consistent data source, solving the noise interference problem caused by the heterogeneity of multi-source data. It innovatively adopts a "data layer-user layer" two-layer clustering architecture, first mining the fine-grained features of single charging behavior (such as charging duration and power intensity) based on the HDBSCAN algorithm, and then dynamically clustering user groups using the K-means algorithm to effectively capture the differentiated patterns of user charging habits. Furthermore, by integrating time series coding (sine / cosine transform) with exogenous variables such as weather and holidays, the model's ability to model spatiotemporal dependencies and external environmental factors is enhanced, solving the defects of feature redundancy and insufficient scenario generalization in traditional methods.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for predicting charging demand of individual electric vehicle users, comprising:

[0007] Constructing a multi-source dataset based on charging behavior data, weather data, and holiday data of EV users in the community; cleaning and standardizing the multi-source dataset using a multi-source data preprocessing model to obtain a processed multi-source dataset;

[0008] Clustering the processed multi-source dataset using a two-layer clustering model to obtain charging type features and user type features;

[0009] Extract input features from the processed multi-source data set; input the charging type features, the user type features and the input features into a charging demand prediction model, perform prediction processing through the LightGBM algorithm, and obtain user charging demand.

[0010] As a preferred solution for a method for predicting charging demand of individual electric vehicle users, the charging behavior data of EV users in the community includes: EV starting charging time, charging amount and charging duration data.

[0011] As a preferred solution of a method for predicting charging demand of individual electric vehicle users, in the process of cleaning and standardizing the multi-source data set using the multi-source data preprocessing model, the multi-source data set is cleaned using a three-level data cleaning strategy; the expression of the processed multi-source data set is:

[0012] A=[A1,A2,…,A i ,…,A N ]

[0013] Where A i is the i-th data, i = 1, 2,…, N; N is the total number of data.

[0014] As a preferred solution for a method for predicting charging demand for individual electric vehicle users, the two-layer clustering model includes a data layer and a user layer. The data layer performs cluster analysis on the processed multi-source dataset using a charging type clustering strategy based on the HDBSCAN algorithm to extract the charging type characteristics. The user layer characterizes the user's long-term charging behavior based on the charging type characteristics using a user type clustering strategy based on the K-means algorithm to obtain the user type characteristics.

[0015] The quality of the clustering results of the charging type clustering strategy based on the HDBSCAN algorithm is evaluated by the Davies-Bouldin index; the calculation formula of the Davies-Bouldin index is:

[0016]

[0017] Where DB is the Davies-Bouldin index; H is the number of mutually exclusive clusters; and are the average distances from all samples in the rth cluster and the tth cluster to the center of the class to which they belong, respectively; ω rt is the class center distance of different categories;

[0018] The user's charging behavior is characterized by the user's proportional combination feature vector, and the expression of the proportional combination feature is:

[0019]

[0020] Where x u is the proportional combination feature vector; u is the individual EV user; U is the total number of EV users; N is the number of samples belonging to the hth charging type among user u; u is the total number of samples of user u; The number of samples of the hth charging type for user u accounts for N u proportion.

[0021] As a preferred solution for a method for predicting charging demand of individual electric vehicle users, during the prediction process using the LightGBM algorithm, the objective function expression of the LightGBM algorithm is:

[0022]

[0023] Where, δ i is the true value of the i-th sample in D; is the predicted value of the i-th sample; L is the regression loss function, which is set to the mean absolute error; Reg1 and Reg2 are the L1 and L2 regularization terms respectively; λ1 and λ2 are the coefficients of the L1 and L2 regularization terms respectively; N is the total number of data.

[0024] The present invention further provides a device for predicting charging demand of individual electric vehicle users, based on the above method for predicting charging demand of individual electric vehicle users, comprising:

[0025] A multi-source dataset construction and preprocessing module is used to construct a multi-source dataset based on charging behavior data, weather data, and holiday data of EV users in the community; the multi-source dataset is cleaned and standardized using a multi-source data preprocessing model to obtain a processed multi-source dataset;

[0026] A two-layer clustering processing module, configured to perform clustering processing on the processed multi-source data set using a two-layer clustering model to obtain charging type features and user type features;

[0027] A user charging demand prediction module is used to extract input features from the processed multi-source data set; input the charging type features, the user type features and the input features into a charging demand prediction model, perform prediction processing through the LightGBM algorithm, and obtain user charging demand.

[0028] As a preferred solution for a device for predicting charging demand of individual electric vehicle users, in the multi-source dataset construction and preprocessing module, the charging behavior data of EV users in the community includes: the starting charging time, charging amount and charging duration data of the EV.

[0029] As a preferred embodiment of a device for predicting charging demand of individual electric vehicle users, in the multi-source data set construction and preprocessing module, in the process of cleaning and standardizing the multi-source data set using the multi-source data preprocessing model, the multi-source data set is cleaned using a three-level data cleaning strategy; the expression of the processed multi-source data set is:

[0030] A=[A1,A2,…,A i ,…,A N ]

[0031] Where A i is the i-th data, i = 1, 2,…, N; N is the total number of data.

[0032] As a preferred solution for an electric vehicle individual user charging demand prediction device, in the two-layer clustering processing module, the two-layer clustering model includes a data layer and a user layer; the data layer performs cluster analysis on the processed multi-source data set using a charging type clustering strategy based on the HDBSCAN algorithm to extract the charging type characteristics; the user layer uses a user type clustering strategy based on the K-means algorithm to characterize the user's long-term charging behavior based on the charging type characteristics to obtain the user type characteristics;

[0033] The quality of the clustering results of the charging type clustering strategy based on the HDBSCAN algorithm is evaluated by the Davies-Bouldin index; the calculation formula of the Davies-Bouldin index is:

[0034]

[0035] Where DB is the Davies-Bouldin index; H is the number of mutually exclusive clusters; and are the average distances from all samples in the rth cluster and the tth cluster to the center of the class to which they belong, respectively; ω rt is the class center distance of different categories;

[0036] The user's charging behavior is characterized by the user's proportional combination feature vector, and the expression of the proportional combination feature is:

[0037]

[0038] Where x u is the proportional combination feature vector; u is the individual EV user; U is the total number of EV users; N is the number of samples belonging to the hth charging type among user u; u is the total number of samples of user u; The number of samples of the hth charging type for user u accounts for N u proportion.

[0039] As a preferred solution of a device for predicting charging demand of individual electric vehicle users, in the user charging demand prediction module, during the prediction process using the LightGBM algorithm, the objective function expression of the LightGBM algorithm is:

[0040]

[0041] Where, δ i is the true value of the i-th sample in D; is the predicted value of the i-th sample; L is the regression loss function, which is set to the mean absolute error; Reg1 and Reg2 are the L1 and L2 regularization terms respectively; λ1 and λ2 are the coefficients of the L1 and L2 regularization terms respectively; N is the total number of data.

[0042] The present invention has the following advantages: It constructs a multi-source dataset based on charging behavior data, weather data, and holiday data from EV users within a community; cleans and standardizes the multi-source dataset using a multi-source data preprocessing model to obtain a processed multi-source dataset; clusters the processed multi-source dataset using a two-layer clustering model to obtain charging type features and user type features; extracts input features from the processed multi-source dataset; and inputs the charging type features, user type features, and input features into a charging demand prediction model, performing prediction processing using the LightGBM algorithm to obtain user charging demand. The present invention significantly improves prediction accuracy and scenario adaptability through a data-driven hierarchical feature extraction and model optimization mechanism. Specifically, a highly consistent data source is constructed through a three-level cleaning process, addressing the noise interference caused by heterogeneous multi-source data. An innovative "data layer-user layer" two-layer clustering architecture is employed. The HDBSCAN algorithm is first used to mine fine-grained features of individual charging behaviors (such as charging duration and charge intensity), and then combined with the K-means algorithm to dynamically cluster user groups, effectively capturing the differentiated patterns of user charging habits. Furthermore, by integrating time series coding (sine / cosine transform) with exogenous variables such as weather and holidays, the model's ability to model spatiotemporal dependencies and external environmental factors is enhanced, solving the defects of feature redundancy and insufficient scenario generalization in traditional methods. Finally, the Optuna library is used to achieve adaptive optimization of model hyperparameters, significantly reducing the cost of manual parameter adjustment and improving prediction stability. Compared with the existing technology, the present invention has significant advantages in data utilization, feature expression capabilities and model generalization performance. It can provide high-precision, real-time charging demand forecasting support for scenarios such as power grid scheduling and charging pile layout planning, and promote the coordinated development of new energy vehicles and smart energy systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can, without inventive effort, derive other implementation drawings based on the provided drawings.

[0044] The structures, proportions, sizes, etc. illustrated in this specification are intended solely to complement the contents disclosed herein and to facilitate understanding and reading by persons skilled in the art. They are not intended to limit the conditions under which the present invention may be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportions, or adjustments in sizes, without affecting the efficacy and objectives of the present invention, shall remain within the scope of the technical contents disclosed herein.

[0045] Figure 1 This is a flow chart of a method for predicting charging demand of individual electric vehicle users provided in Example 1 of the present invention;

[0046] Figure 2 This is a schematic diagram of a specific implementation architecture of a method for predicting charging demand of individual electric vehicle users provided in Example 1 of the present invention;

[0047] Figure 3 Schematic diagram of data availability and feature analysis for all EVs in a possible embodiment provided in Example 1 of the present invention; (a) is the number of charging days; (b) is the number of charging times; (c) is the weekly charging frequency; and (d) is the average charging duration.

[0048] Figure 4 Schematic diagram of the joint distribution of charging duration and charge amount versus start charging time in a possible embodiment provided in Example 1 of the present invention; wherein the X-axis histogram represents the marginal distribution of the start charging time; and the Y-axis histogram represents the marginal distribution of the dependent variable; wherein (a) is the joint distribution diagram of the start charging time and charging duration; and (b) is the joint distribution diagram of the start charging time and charge amount;

[0049] Figure 5 This is a schematic diagram of the probability density distribution of different charging types in three dimensions: starting charging time, charging duration, and charging amount, in a possible embodiment provided in Example 1 of the present invention;

[0050] Figure 6 This is a schematic diagram of comparison between different user types in a possible embodiment provided in Example 1 of the present invention;

[0051] Figure 7 A schematic diagram of the MAE and SMAPE distribution of each model in different prediction tasks in a possible embodiment provided in Example 1 of the present invention;

[0052] Figure 8 A schematic diagram of the MAE and SMAPE distribution of the prediction results of each model in a possible embodiment provided in Example 1 of the present invention;

[0053] Figure 9Schematic diagram of MAE comparison of each user under each model in a possible embodiment provided in Example 1 of the present invention; wherein (a) is a comparison between Model 3 and Model 1; (b) is a comparison between Model 3 and Model 2;

[0054] Figure 10 A schematic diagram of the MAE and SMAPE distribution of the prediction results of each model in a possible embodiment provided in Example 1 of the present invention;

[0055] Figure 11 Schematic diagram of the MAE of each user under each model in a possible embodiment provided in Example 1 of the present invention; wherein (a) is a comparison between Model 3 and Model 1; (b) is a comparison between Model 3 and Model 2;

[0056] Figure 12 This is a schematic diagram of the architecture of a device for predicting charging demand of individual electric vehicle users provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0057] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0058] Example 1

[0059] See also Figure 1 and Figure 2 Embodiment 1 of the present invention provides a method for predicting charging demand of individual electric vehicle users, comprising the following steps:

[0060] S1. Constructing a multi-source dataset based on charging behavior data, weather data, and holiday data of EV users in the community; and cleaning and standardizing the multi-source dataset using a multi-source data preprocessing model to obtain a processed multi-source dataset.

[0061] S2. performing clustering processing on the processed multi-source dataset using a two-layer clustering model to obtain charging type features and user type features;

[0062] S3. Extract input features from the processed multi-source dataset; input the charging type features, the user type features, and the input features into a charging demand prediction model, perform prediction processing using the LightGBM algorithm, and obtain user charging demand.

[0063] In this embodiment, in step S1, a multi-source dataset is constructed based on the charging behavior data, weather data, and holiday data of EV users in the community; the multi-source dataset is cleaned and standardized using a multi-source data preprocessing model to obtain a processed multi-source dataset;

[0064] Specifically, the multi-source dataset contains charging behavior data, weather data, and holiday data of EV users in the community. The charging behavior data includes metadata such as the EV's start charging time, charge amount, and charging duration.

[0065] To address the outlier issues in the original data, a three-level data cleaning process was established. First, time threshold filtering was implemented to delete abnormal data with charging time less than five minutes or exceeding seven days. Then, the charging behavior exclusion principle was applied. Based on the physical constraint that "the same electric vehicle can only perform a single charging behavior at the same time node", abnormal data with overlapping time was removed. Then, a user activity screening mechanism was established. Each EV user must have sufficient data to analyze their charging behavior. Therefore, low-activity EV users with an average weekly charging frequency of less than 1 time were eliminated. Finally, a standardized charging data set A was obtained, which is in the form of:

[0066] A=[A1,A2,…,A i ,…,A N ]

[0067] Where A i is the i-th data, i = 1, 2,…, N; N is the total number of data.

[0068] This dataset fully preserves the effective charging behavior data of all EV users in the community, where each data A i It contains structured metadata such as the starting charging time (ts), charging duration (d) and charging capacity (c).

[0069] In this embodiment, in step S2, clustering is performed on the processed multi-source dataset using a two-layer clustering model to obtain charging type features and user type features;

[0070] Specifically, analyzing charging types and EV user types is a key prerequisite for accurately predicting user charging needs. However, individual users' charging types often exhibit multimodal characteristics (e.g., long-term charging at night coexisting with short-term charging during the afternoon), and charging behavior also varies among different user groups. Directly performing single-layer clustering on user or charging behavior data is susceptible to interference from the randomness of user charging behavior. It is difficult to simultaneously capture micro-level charging event characteristics (e.g., time, duration, and power consumption) and macro-level user behavior preferences (e.g., charging type distribution). It ignores the multi-level characteristics of charging behavior, resulting in inaccurate classification of charging types and user types.

[0071] Therefore, the present invention adopts a two-layer clustering model to solve the above problems by performing hierarchical clustering from the data layer and the user layer. At the data layer, the charging data is first clustered to identify different charging types, and the stability of the user layer clustering is improved by filtering abnormal data; at the user layer, based on the results of the data layer clustering, the proportion of each user in different charging types is calculated, and different user types are clustered and identified accordingly. This model can effectively filter out random noise in the charging data, thereby ensuring the robustness of user classification. Specifically, if a user has abnormal charging data, using only a single-layer clustering for the user may cause it to be misclassified. By first identifying its main charging type (such as long-term charging at night) through data layer clustering, and then clustering the feature distribution at the user layer, it can effectively reduce the interference of abnormal data and more accurately classify users into the correct user type.

[0072] In this embodiment, the data layer is based on HDBSCAN charging type clustering:

[0073] Before implementing data layer clustering, select t from the dataset A. s , d, and c are three metadata items that reflect EV users' charging needs: t reveals the user's preference for the starting time of charging; d reflects the intensity of the demand for charging duration; and c reflects the energy intensity required for a single charge. Based on the theory of feature subset selection (FSS), this three-dimensional feature combination maximizes the preservation of the core characteristics of user charging behavior while ensuring computational efficiency. Furthermore, after min-max normalization, a new sample set C is constructed, which is in the form of:

[0074] C=[C1,C2,…,C i ,…,C N ]

[0075] C i =[c i,1 ,c i,2 ,c i,3 ]

[0076] Where C i is the i-th sample; c i,1 , c i,1 and c i,3 C i Respectively in t s , d and c in three dimensions.

[0077] Samples C1 to C in C N Respectively with A1 to A in A N The final three-dimensional feature matrix C will be used as the input data of the HDBSCAN clustering algorithm.

[0078] In this embodiment, the HDBSCAN algorithm is used as the clustering method for the data layer. Compared with traditional clustering algorithms, HDBSCAN can adapt to the shape and size of clusters, avoiding the limitation of a preset number of clusters, thereby achieving effective identification of different charging types.

[0079] The clustering results of the HDBSCAN clustering algorithm depend on the settings of hyperparameters. To achieve the best clustering results, the parameter optimization algorithm in the Optuna library is used to search for the best combination of hyperparameters. The hyperparameter search range is as follows:

[0080] As shown in Table 1:

[0081] parameter meaning scope min_simples Neighborhood contains the minimum number of samples 6-60 min_size The cluster contains the minimum number of samples 4-3000 alpha Sensitivity Control 0.1-3.0

[0082] Table 1 Hyperparameter search range of HDBSCAN

[0083] The quality of the clustering results of the charging type clustering strategy based on the HDBSCAN algorithm is evaluated by the Davies-Bouldin index; the calculation formula of the Davies-Bouldin index is:

[0084]

[0085] Where DB is the Davies-Bouldin index; H is the number of mutually exclusive clusters; and are the average distances from all samples in the rth cluster and the tth cluster to the center of the class to which they belong, respectively; ω rt is the class center distance of different categories;

[0086] The DB index evaluates the clustering effect by quantifying the balance between intra-class cohesion and inter-class separation. The smaller the DB value, the better the compactness and separability of the clustering structure.

[0087] The data layer clustering process follows the following steps: Take C as the input sample set, and use the HDBSCAN algorithm to divide C into H mutually exclusive clusters, where each cluster corresponds to a specific charging type. i When sample C is assigned to cluster h, i Belonging to the hth charging type, its classification feature is marked as h∈{1,2,…,H}, which is used to represent its category.

[0088] In this embodiment, user type clustering is based on the K-means algorithm:

[0089] The charging data of each EV user is composed of different charging types in different proportions. Therefore, after extracting H types of charging types at the data layer, the proportion combination feature vector x of individual user u is calculated. u , to characterize the charging behavior of user u.

[0090] Wherein, the expression of the proportion combination feature is:

[0091]

[0092] Where x u is the proportional combination feature vector; u is the individual EV user; U is the total number of EV users; N is the number of samples belonging to the hth charging type among user u; u is the total number of samples of user u; The number of samples of the hth charging type for user u accounts for N u proportion.

[0093] The K-Means algorithm has significant application advantages in user-level clustering analysis. The input data for the user layer is the proportion of charging types, which usually presents a convex structure in low-dimensional space (for example, the user's charging data is concentrated in a few types of charging types). The K-means algorithm can effectively identify this structure and then divide user groups into different user types. In addition, the data layer has completed noise filtering and charging type identification through the HDBSCAN algorithm. Therefore, the core goal of user-level clustering shifts to mining the long-term charging behavior patterns of user groups, rather than processing cluster structures with complex geometric shapes.

[0094] In this embodiment, at the user level, the process of using the K-means algorithm to perform cluster analysis on U EV users to identify EV user types is as follows: First, calculate the x u ; Then, the x of the U user u Using the vector as input, the K-means algorithm is applied to segment user groups. Finally, a clustering algorithm divides U EV users into L user types with differentiated charging behaviors. If user u is assigned to the lth user type, then user u's classification feature l∈{1,2,…,L}. This clustering method achieves high similarity in charging behavior for users within a cluster and significant differences between clusters.

[0095] Based on the one-to-one mapping relationship between the samples of dataset C and A, C i The classification feature h obtained by clustering is assigned to A in A i, allowing A to add the charging type metadata attribute while retaining the original features. Since all EV user charging data originates from A, user-level clustering analysis allows the user type feature l to be simultaneously mapped to all charging behavior data corresponding to user u in A. Thus, the two-layer clustering model, through the expansion of the feature space, adds the charging type and user type features to each charging data entry in A's original metadata structure, ultimately forming dataset B, which provides richer and more structured input features for subsequent prediction models.

[0096] In this embodiment, in step S3, input features are extracted from the processed multi-source data set; the charging type features, the user type features and the input features are input into a charging demand prediction model, and prediction processing is performed using the LightGBM algorithm to obtain user charging demand.

[0097] Specifically, based on dataset B and related work, features related to user charging needs can be extracted from B, as shown in Table 2:

[0098] Feature Name describe <![CDATA[t sin ]]> Sine value of the EV starting charging time <![CDATA[t cos ]]> Cosine value of the EV starting charging time <![CDATA[t m ]]> Month when EV charging started <![CDATA[t l ]]> The length of time since the EV was last charged <![CDATA[l d ]]> How long did it take for the EV to be last charged? <![CDATA[l c ]]> The last charge level of the EV m EV charging methods are divided into ordered charging and disordered charging <![CDATA[t max ]]> The highest temperature of the day when EV is charging <![CDATA[t min ]]> Lowest temperature of the day when EV is charging v Boolean value, 1 means holiday, 0 means non-holiday l Classification features characterizing charging types h Classification features that characterize user types

[0099] Table 2 Input features of the prediction model

[0100] where t max , t min and v represent two exogenous variables, weather and holidays. The remaining features in the table are significantly correlated with user charging demand. Based on the above features, a dataset D is formed and used as the input data for the subsequent prediction model. s The time series features are expressed in a cyclic coding manner, that is, t s Convert to sine value (t sin ) and cosine value (t cos ). This encoding strategy effectively captures t s The periodic characteristics of the time interval ensure the continuity and uniformity of the time interval.

[0101] The present invention defines the user charging demand as two dimensions: charging time demand and charging amount demand. The input features of LightGBM are divided into two categories: numerical features and categorical features, as shown in Table 2, where m, v, h and l are categorical features, and the rest are numerical features. The key reason for choosing LightGBM as the prediction model is its ability to efficiently process categorical features. LightGBM can directly and efficiently process categorical features without the need for One-Hot encoding, avoiding problems such as feature dimension expansion and improving computational efficiency. At the same time, the gradient histogram algorithm used by LightGBM can effectively reduce the amount of calculation, making it faster to train on large-scale data sets. By introducing h and l as classification features, the model can effectively distinguish different charging types and user types, thereby improving the prediction accuracy and generalization ability of LightGBM.

[0102] The LightGBM model aims to minimize the predicted value The regression loss between the actual value δ is the optimization target. In order to effectively suppress the overfitting phenomenon, the objective function f loss Combining the regression loss function and the regularization term, its expression is:

[0103]

[0104] Where, δ i is the true value of the i-th sample in D; is the predicted value of the i-th sample; L is the regression loss function, which is set to the mean absolute error; Reg1 and Reg2 are the L1 and L2 regularization terms respectively; λ1 and λ2 are the coefficients of the L1 and L2 regularization terms respectively; N is the total number of data.

[0105] This model involves the tuning of multiple hyperparameters. To improve the model prediction effect, the present invention uses the Optuna library to optimize the hyperparameters to obtain the best hyperparameter combination. The search range of the hyperparameters is shown in Table 3:

[0106] parameter meaning scope num_leaves The maximum number of leaf nodes in a single tree 10-1000 learning_rate Learning rate 0.001-0.1 max_depth Maximum depth of the tree 2-13 min_data The minimum amount of data contained in a leaf node 1-1000 lambda_l1 L1 regularization coefficient 0.001-100 lambda_l2 L2 regularization coefficient 0.001-100

[0107] Table 3 Hyperparameter search range of LightGBM

[0108] In this embodiment, combining the multi-source data preprocessing model and the two-layer clustering model, the user charging demand prediction steps based on LightGBM are as follows:

[0109] T1. Use the multi-source data preprocessing model to clean the multi-source data set to obtain data set A;

[0110] T2. Extract key parameters h and l based on the two-layer clustering model, fuse h and l with features of A to construct enhanced dataset B, and derive dataset D based on B;

[0111] T3. Use the Optuna library to search for the optimal hyperparameter combination, divide D into a training set and a test set, and use the optimization objective as the training criterion for the model, and use the training set to train the model;

[0112] T4. Implement model performance verification on the test set and conduct quantitative evaluation of user charging demand prediction results.

[0113] In a possible embodiment, a simulation prediction example is provided as follows:

[0114] The charging data in this example is derived from a car network platform, which contains complete charging records for 300 EVs in a community in northern China between November 1, 2021, and November 8, 2022. Weather data and holiday data are acquired through network collection, and a multi-source dataset is constructed based on these multi-dimensional data sources. The simulation experiment is based on the Python 3.10 environment, the computer operating system is Windows 11, and the CPU is an Intel(R) Core(TM) i7-12700KF 3.60GHz.

[0115] 1. Multi-source data preprocessing and analysis

[0116] The multi-source dataset A is preprocessed, which contains 59,911 data items from 281 EV users. Figure 3 (a) and Figure 3 (b) shows that the amount of available data for each EV user is different, resulting in a wide distribution of weekly charging frequencies for each EV user, e.g. Figure 3 As shown in (c), the mean and variance of the weekly charging frequency of the overall EV users are 4.6±2.6; Figure 3 (d) shows that the average charging time for EVs is distributed narrowly (6-8 hours), but there are still a few extreme values ​​exceeding 20 hours. These empirical results indicate that there are significant differences in charging behavior among EV users, which provides a foundation for subsequent work.

[0117] Figure 4 The distribution of EV charging time and charging amount relative to the starting time of charging is shown. The distribution histogram shows that user charging behavior presents a significant bimodal feature, mainly concentrated in the two traffic peak hours of noon (11:00-13:00) and evening (18:00-22:00), and there is almost no charging behavior in the late night period (0:00-5:00). Figure 4The heat map in (a) shows that short charging times (2-4 hours) often occur during the day (11:00-1:00 PM), while long charging times (8-15 hours) tend to occur at night (6:00-10:00 PM). It is worth noting that EVs in this community rarely charge in the early morning, and charging times exceeding 20 hours are extremely rare, confirming that charging times vary significantly depending on the starting time. Figure 4 The charge capacity distribution analysis in (b) shows that the charge capacity during the evening period (18:00-24:00) is distributed in the range of 20kWh-40kWh, while the charge capacity during the daytime period (10:00-14:00) is significantly lower, concentrated in the range of 10kWh-20kWh. Extreme cases where the charge capacity exceeds 50kWh are very rare.

[0118] 2. Analysis of the two-layer clustering results Extract t from A s , d, c and normalized to obtain the data set C. Based on the HDBSCAN algorithm, C is divided into three types of data, namely, charging types 1 to 3. The results are as follows Figure 5 As shown in Table 4:

[0119]

[0120] Table 4 Data layer clustering results

[0121] based on Figure 5As analyzed in Table 4, each type of charging is analyzed based on three dimensions: charging start time, charge amount, and charging duration. In terms of charging start time distribution, the three charging types are concentrated in the early morning, midday, and nighttime periods, respectively, demonstrating different time preferences. In terms of charging duration distribution, the third type of charging type has the shortest average charging duration, at only 2.9 hours, indicating that midday charging is primarily a short-term top-up. The second type of charging type has the longest average charging duration, at 9.3 hours, indicating that evening charging tends to be more prolonged. The first type of charging type has a charging duration between the second and third types, being both shorter than the long evening charging and longer than the short midday charging, thus belonging to the medium evening charging type. In terms of charge capacity distribution, the third type of charging has the lowest average charge capacity, only 14.0kWh, confirming the short-term nature of midday charging, which is usually used for small-scale power replenishment. The second type of charging has an average charge capacity of 29.9kWh, indicating that this type of charging has a higher charge capacity demand based on long-term charging. The first type of charging has the highest average charge capacity, reaching 32.0kWh. Although its charging time is shorter than the second type, its overall power demand is larger, indicating that this type of charging may be related to the power replenishment needs of high-capacity batteries. In addition, the second type of charging begins in the evening and continues into the next day, belonging to the long-term overnight charging type. In summary, the first type is a medium-term, high-capacity charging type in the early morning; the second type is a long-term, high-capacity charging type overnight; and the third type is a short, low-capacity charging type in the afternoon.

[0122] Based on the data layer clustering, x is calculated u , the K-means algorithm is used to cluster users into five types of users. The number of users and charging times of each type are shown in Table 5 and Figure 6 As shown:

[0123] User Category Number of users Cumulative charging times / times Average starting charging time Average charging time / h Average charge capacity 1 81 15873 19:41 7.0 24.63 2 69 13314 18:02 5.6 21.2 3 76 14082 20:22 8.8 29.2 4 22 4622 23:22 6.2 26.2 5 33 4178 12:34 3.7 13.8

[0124] Table 5 User layer clustering results

[0125] From Table 5 and Figure 6As can be seen, different user categories exhibit significant differences in terms of number of charges, start time, and duration, reflecting the varying charging preferences of community users. Categories 1, 3, and 4 all charge at night. Category 3 has the latest start time (20:22) and the longest average charging duration (8.8 hours), representing a user type that charges for extended periods at night. This user type likely primarily charges at home, taking advantage of low nighttime electricity prices to complete a full charge. Their average charge capacity is the highest (29.2 kWh), indicating a high demand for charging. Their cumulative number of charges is also significantly higher than that of other categories, indicating frequent and continuous charging, possibly favoring a single full charge. Category 1 (19:41, 7.0 hours), while also preferring evening charging, has a shorter charging duration than Category 3, likely representing community users who commute daily and charge regularly. Their average charge capacity is 24.6 kWh, between Category 3 and Category 4, indicating moderate charging demand. This category also represents the largest number of users, suggesting that the charging behavior represented by this user type is more prevalent in the community. Category 4 has a later average charging start time (23:22) and an average charging duration of 6.2 hours. This may be influenced by work habits or electricity pricing policies, representing night shift workers or users who prefer off-peak electricity prices. Their average charge capacity is 26.2 kWh, similar to category 3, indicating that this category also tends to charge for longer periods. Although their number is relatively small, their charging demand is still significant. Category 5 (12:34, 3.7 hours) primarily charges during the afternoon, with shorter charging times. This may correspond to users with regular daytime travel needs, such as residents commuting or running errands. Their average charge capacity is the lowest (13.8 kWh), indicating that this category has relatively low charging needs and is relatively small in number, suggesting that their charging behavior is more sporadic and infrequent. Category 2 (18:02, 5.6 hours) charges more towards the evening, with a moderate charging duration and an average charge capacity of 21.2 kWh. This type of user may have regular daily charging needs, with a high frequency of charging, especially in the evening hours on weekdays. This suggests that their charging needs are widespread within the community, but slightly lower than those of other categories. A two-layer clustering model was used to obtain classification features h and l. Here, h∈{1,2,…,H}, where H is the number of charging types obtained by data-level clustering (i.e., H=3), and l∈{1,2,…,L}, where l is the number of user types obtained by user-level clustering (i.e., L=5).

[0126] 3. User Charging Demand Forecast Results

[0127] In order to verify the effectiveness of the LightGBM user charging demand prediction method combined with the two-layer clustering model proposed in this paper, LSTM and ARIMA are selected as the comparison methods for the user charging demand prediction part. Each method uses the same data set. A is combined with the classification features h and l to obtain set B, and finally the data set D is obtained. D is divided into training set D according to date. t and the test set D v , D t Contains all data from November 1, 2021 to September 30, 2022; D v Contains all data from October 1, 2022 to November 8, 2022. t Train three models and v The prediction performance of each model is evaluated and the evaluation indicators of each model are calculated.

[0128] MAE directly reflects absolute error and clearly demonstrates the magnitude of the prediction error, making it suitable for assessing the actual scale of the prediction error. SMAPE, on the other hand, measures relative error and allows for horizontal comparison of the prediction performance of different models across different prediction tasks. Table 6 lists the average values ​​of various evaluation metrics for the three models when predicting overall EV user charging demand.

[0129]

[0130] Table 6 Prediction results of each model in the test set

[0131] As can be seen from Table 6, LightGBM has shown significantly better prediction performance than other models in both charging time and charging capacity prediction tasks. Specifically, in the charging time prediction task, the MAE of LightGBM is 1.1 hours, which is much lower than LSTM's 3.0 hours and ARIMA's 3.5 hours, a decrease of 68.6% and 63.3% respectively. This shows that LightGBM is more accurate in capturing the changing trend of charging time. At the same time, in the charging capacity prediction task, the MAE of LightGBM is 6.2kWh, which is 45.6% and 50.8% lower than LSTM's 11.4kWh and ARIMA's 12.6kWh, respectively, which also shows that LightGBM can predict the charging capacity more accurately.

[0132] From the perspective of other indicators, the MSE and MAPE of LightGBM also show its advantages in accuracy. For example, in the charging time prediction task, the MSE of LightGBM is 2.2h2, which is significantly lower than LSTM (13.2h2) and ARIMA (17.2h2). This shows that LightGBM not only performs well in average error, but also has a smaller fluctuation range of prediction error, indicating that the model has better stability. In terms of SMAPE, LightGBM performs equally well. Its SMAPE for predicting charging time is 23.9%, and its SMAPE for predicting charging amount is 31.8%, which is significantly lower than the SMAPE values ​​of LSTM and ARIMA. This shows that LightGBM can maintain consistent and smaller relative errors between different tasks, and has a stronger horizontal comparison advantage. Comprehensive analysis shows that the present invention has better prediction accuracy.

[0133] Figure 7 The MAE and SMAPE distributions of all users for each model in different prediction tasks are shown. The violin plots intuitively reflect the differences in prediction errors of individual users, and also include the mean value (Mean) of the overall EV user evaluation index. Table 7 summarizes the extreme values ​​of MAE and SMAPE for each model in different prediction tasks to further compare the performance of each model in individual user prediction accuracy.

[0134]

[0135]

[0136] Table 7 Extreme values ​​of MAE and SMAPE for each model in different prediction tasks

[0137] The results show that the average MAE of the LightGBM model is roughly consistent with the median value, and the MAE gap between the best and worst individual users is the smallest, indicating that the model's prediction error distribution is concentrated and exhibits good stability. In contrast, the LSTM and ARIMA models have a wider distribution range, especially in charging time prediction, where their MAEs reach a maximum of 7.74h and 8.55h, respectively, and their SMAPEs reach 142.02% and 126.71%, respectively, significantly higher than LightGBM, showing volatility. In addition, the SMAPE range of LightGBM in charging capacity prediction is wider than that in charging time prediction, indicating that task characteristics have an impact on the model's prediction performance.

[0138] As shown in Table 7, the LightGBM model achieves the best prediction results compared to the ARIMA and LSTM models, with a minimum MAE of 0.16h and a maximum MAE of 3.7h, both lower than those of the other models. Specifically, the LightGBM model's prediction results outperform those in the Phipps study. The MAE and MSE for the best individual user were 0.16h and 0.076h², respectively, significantly lower than the MSE value (0.75h²) for the best individual user in the Chen et al. study.

[0139] Although the LightGBM model performed the best, the individual user with the worst prediction effect had a MAE of 3.7h, which is more than three times the average value (1.1h). This shows that the prediction effect of the model varies greatly among different individual users. This difference may be due to the significant difference in the charging behavior of individual users from other users, or the presence of extreme outliers in their charging behavior data. Since the charging behavior of individual users varies significantly, and the charging data of some users may contain extremely irregular patterns or noise, it is difficult for a simple model to completely solve this problem. Therefore, although LightGBM performs well for most individual users, it is still necessary to pay attention to the prediction error of individual users to further improve the prediction accuracy of the model.

[0140] 4. The impact of two-layer clustering and weather and holidays on LightGBM prediction of individual user charging demand

[0141] The following analysis focuses on the impact of weather, holidays, and the classification features generated by the two-layer clustering model on the prediction of individual user charging demand. The control experiment model is shown in Table 8:

[0142]

[0143]

[0144] Table 8 Control experimental group settings

[0145] Analyze the impact on charging time demand forecast:

[0146] Figure 8 The prediction results of each model are visualized. Table 9 shows the individual users with the best and worst prediction results in each model.

[0147]

[0148] Table 9 Extreme values ​​of MAE and SMAPE of each model in charging time prediction

[0149] from Figure 8As can be seen from Table 9, compared with Model 3, the MAE value distribution of Model 1 is wider, and the MAE maximum value is significantly higher than that of Model 3. This indicates that when the classification features generated by the two-layer clustering model are not included, the volatility of the prediction error is greater; the error of Model 3 is smaller and the distribution is concentrated, which proves the importance of the two-layer clustering model in improving prediction accuracy; the maximum values ​​of MAE and SMAPE in Model 2 are greater than the maximum values ​​of Model 3, indicating that weather and holidays also have an improving effect on prediction accuracy. Although the MAE minimum values ​​of the three models are similar, LightGBM using the two-layer clustering model can effectively reduce the MAE value of the entire EV user and improve the overall prediction accuracy. The following will analyze in detail the effect of the two-layer clustering model and weather and holidays in improving the accuracy of predicting individual user charging demand.

[0150] Figure 9 The MAE of each user of Model 3 and other models is compared in the form of a scatter plot. Figure 9 As shown in (a), the MAE values ​​of many users in Model 3 are lower than those in Model 1. The statistical results show that Model 3 reduces the MAE values ​​of users by 85.1%, which fully demonstrates the superiority of the two-layer clustering model in improving the accuracy of charging time prediction. Figure 9 Figure (b) shows that Model 3 reduces the MAE for some individual users, indicating that exogenous variables such as weather and holidays also play a positive role in improving charging time prediction accuracy. Statistical results show that Model 3 reduces the MAE for 57.9% of individual EV users. Comprehensive analysis shows that incorporating weather and holidays, as well as using a two-layer clustering model, significantly improves charging time demand prediction accuracy, with the two-layer clustering model showing a more significant improvement.

[0151] Analyze the impact on charging demand forecast:

[0152] Depend on Figure 10 Table 10 shows the MAE and SMAPE distribution of different models in charging demand prediction.

[0153]

[0154] Table 10 Extreme values ​​of MAE and SMAPE of each model in charge capacity prediction

[0155] As can be observed, Model 1 has higher maximum and minimum MAE values ​​than Model 3, and Model 2 also has higher maximum and minimum MAE and SMAPE values ​​than Model 3. This indicates that the use of a two-layer clustering model and weather and holiday factors can improve the accuracy of charging demand forecasts. In the case of extreme forecast errors (i.e., maximum and minimum values), Model 3 has lower volatility than other models, demonstrating its greater stability.

[0156] like Figure 11 As shown in the figure, by comparing the MAE values ​​of each EV individual user in the charging demand forecast of each model, the advantage of Model 3 is further demonstrated. The statistical results show that compared with Model 1, Model 3 reduces the MAE values ​​of individual users by 69.7%, indicating that the two-layer clustering model has significant advantages in handling individual differences and improving accuracy. Compared with Model 2, Model 3 reduces the MAE values ​​of EV users by 57.9%. Figure 11 Figure (b) shows that weather and holidays can also improve the accuracy of charging demand forecasts. Comprehensive analysis shows that the two-layer clustering model, including weather, holidays, and the two-layer clustering model, can all improve charging demand forecast accuracy. This analysis further understands the role of various features in improving forecast accuracy, helping to optimize and personalize future models.

[0157] In summary, the present invention has the following advantages: It constructs a multi-source dataset based on charging behavior data, weather data, and holiday data from EV users within a community; cleans and standardizes the multi-source dataset using a multi-source data preprocessing model to obtain a processed multi-source dataset; clusters the processed multi-source dataset using a two-layer clustering model to obtain charging type features and user type features; extracts input features from the processed multi-source dataset; and inputs the charging type features, user type features, and input features into a charging demand prediction model, performing prediction processing using the LightGBM algorithm to obtain user charging demand. The present invention significantly improves prediction accuracy and scenario adaptability through a data-driven hierarchical feature extraction and model optimization mechanism. Specifically, a highly consistent data source is constructed through a three-level cleaning process, addressing the noise interference caused by heterogeneous multi-source data. An innovative "data layer-user layer" two-layer clustering architecture is employed. The HDBSCAN algorithm is first used to mine fine-grained features of individual charging behaviors (such as charging duration and charge intensity), and then combined with the K-means algorithm to dynamically cluster user groups, effectively capturing the differentiated patterns of user charging habits. Furthermore, by integrating time series coding (sine / cosine transform) with exogenous variables such as weather and holidays, the model's ability to model spatiotemporal dependencies and external environmental factors is enhanced, solving the defects of feature redundancy and insufficient scenario generalization in traditional methods. Finally, the Optuna library is used to achieve adaptive optimization of model hyperparameters, significantly reducing the cost of manual parameter adjustment and improving prediction stability. Compared with the existing technology, the present invention has significant advantages in data utilization, feature expression capabilities and model generalization performance. It can provide high-precision, real-time charging demand forecasting support for scenarios such as power grid scheduling and charging pile layout planning, and promote the coordinated development of new energy vehicles and smart energy systems.

[0158] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.

[0159] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0160] Example 2

[0161] See also Figure 12 Embodiment 2 of the present invention further provides a device for predicting charging demand of individual electric vehicle users, comprising:

[0162] The multi-source dataset construction and preprocessing module 001 is used to construct a multi-source dataset based on the charging behavior data, weather data, and holiday data of EV users in the community; clean and standardize the multi-source dataset using a multi-source data preprocessing model to obtain a processed multi-source dataset;

[0163] A two-layer clustering processing module 002 is used to perform clustering processing on the processed multi-source data set using a two-layer clustering model to obtain charging type features and user type features;

[0164] The user charging demand prediction module 003 is used to extract input features from the processed multi-source data set; input the charging type features, the user type features and the input features into the charging demand prediction model, perform prediction processing through the LightGBM algorithm, and obtain the user charging demand.

[0165] In this embodiment, in the multi-source data set construction and pre-processing module 001, the charging behavior data of EV users in the community includes: the starting charging time, charging amount and charging duration data of the EV.

[0166] In this embodiment, in the multi-source data set construction and preprocessing module 001, in the process of cleaning and standardizing the multi-source data set using the multi-source data preprocessing model, the multi-source data set is cleaned using a three-level data cleaning strategy; the expression of the processed multi-source data set is:

[0167] A=[A1,A2,…,A i ,…,A N ]

[0168] Where A i is the i-th data, i = 1, 2,…, N; N is the total number of data.

[0169] In this embodiment, in the two-layer clustering processing module 002, the two-layer clustering model includes a data layer and a user layer; the data layer performs cluster analysis on the processed multi-source data set using a charging type clustering strategy based on the HDBSCAN algorithm to extract the charging type features; the user layer uses a user type clustering strategy based on the K-means algorithm to characterize the user's long-term charging behavior according to the charging type features to obtain the user type features;

[0170] The quality of the clustering results of the charging type clustering strategy based on the HDBSCAN algorithm is evaluated by the Davies-Bouldin index; the calculation formula of the Davies-Bouldin index is:

[0171]

[0172] Where DB is the Davies-Bouldin index; H is the number of mutually exclusive clusters; and are the average distances from all samples in the rth cluster and the tth cluster to the center of the class to which they belong, respectively; ω rt is the class center distance of different categories;

[0173] The user's charging behavior is characterized by the user's proportional combination feature vector, and the expression of the proportional combination feature is:

[0174]

[0175] Where x u is the proportional combination feature vector; u is the individual EV user; U is the total number of EV users; N is the number of samples belonging to the hth charging type among user u; u is the total number of samples of user u; The number of samples of the hth charging type for user u accounts for N u proportion.

[0176] In this embodiment, in the user charging demand prediction module 003, during the prediction process using the LightGBM algorithm, the objective function expression of the LightGBM algorithm is:

[0177]

[0178] Where, δ i is the true value of the i-th sample in D; is the predicted value of the i-th sample; L is the regression loss function, which is set to the mean absolute error; Reg1 and Reg2 are the L1 and L2 regularization terms respectively; λ1 and λ2 are the coefficients of the L1 and L2 regularization terms respectively; N is the total number of data.

[0179] It should be noted that the information interaction, execution process, etc. between the modules of the above-mentioned system are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and no further details will be given here.

[0180] Example 3

[0181] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code for a method for predicting the charging demand of individual users of electric vehicles is stored. The program code includes instructions for executing embodiment 1 or any possible implementation thereof. A method for predicting the charging demand of individual users of electric vehicles.

[0182] Computer-readable storage media can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0183] Example 4

[0184] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0185] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute a method for predicting charging demand of individual electric vehicle users in embodiment 1 or any possible implementation thereof.

[0186] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in a memory. The memory can be integrated into the processor or located outside the processor and exist independently.

[0187] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode.

[0188] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing system. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Alternatively, they can be implemented using program code executable by a computing system, and thus, they can be stored in a storage system and executed by the computing system. In some cases, the steps shown or described herein can be performed in a different order than that shown, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0189] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.

Claims

1. A method for predicting charging demand of individual electric vehicle users, characterized in that: include: Construct a multi-source dataset based on charging behavior data, weather data, and holiday data of EV users in the community; Cleaning and standardizing the multi-source data set using a multi-source data preprocessing model to obtain a processed multi-source data set; Clustering the processed multi-source dataset using a two-layer clustering model to obtain charging type features and user type features; Extract input features from the processed multi-source data set; input the charging type features, the user type features and the input features into a charging demand prediction model, perform prediction processing through the LightGBM algorithm, and obtain user charging demand.

2. The method for predicting charging demand of individual electric vehicle users according to claim 1, characterized in that: The charging behavior data of EV users in the community includes: the starting charging time, charging amount and charging duration data of the EV.

3. The method for predicting charging demand of individual electric vehicle users according to claim 2, characterized in that: In the process of cleaning and standardizing the multi-source data set by the multi-source data preprocessing model, the multi-source data set is cleaned by a three-level data cleaning strategy; the expression of the processed multi-source data set is: A=[A1,A2,…,A i ,…,A N ] Where A i is the i-th data, i = 1, 2,…, N; N is the total number of data.

4. The method for predicting charging demand of individual electric vehicle users according to claim 3, characterized in that: The two-layer clustering model includes a data layer and a user layer; the data layer performs cluster analysis on the processed multi-source data set using a charging type clustering strategy based on the HDBSCAN algorithm to extract the charging type features; The user layer is a user type clustering strategy based on the K-means algorithm, which characterizes the user's long-term charging behavior according to the charging type characteristics to obtain the user type characteristics; The quality of the clustering results of the charging type clustering strategy based on the HDBSCAN algorithm is evaluated by the Davies-Bouldin index; the calculation formula of the Davies-Bouldin index is: Where DB is the Davies-Bouldin index; H is the number of mutually exclusive clusters; and are the average distances from all samples in the rth cluster and the tth cluster to the center of the class to which they belong, respectively; ω rt is the class center distance of different categories; The user's charging behavior is characterized by the user's proportional combination feature vector, and the expression of the proportional combination feature is: Where x u is the proportional combination feature vector; u is the individual EV user; U is the total number of EV users; N is the number of samples belonging to the hth charging type among user u; u is the total number of samples of user u; The number of samples of the hth charging type for user u accounts for N u proportion.

5. The method for predicting charging demand of individual electric vehicle users according to claim 4, characterized in that: In the process of prediction processing by the LightGBM algorithm, the objective function expression of the LightGBM algorithm is: Where, δ i is the true value of the i-th sample in D; is the predicted value of the i-th sample; L is the regression loss function, which is set to the mean absolute error; Reg1 and Reg2 are the L1 and L2 regularization terms respectively; λ1 and λ2 are the coefficients of the L1 and L2 regularization terms respectively; N is the total number of data.

6. A device for predicting charging demand of individual electric vehicle users, using the method for predicting charging demand of individual electric vehicle users according to any one of claims 1 to 5, characterized in that: include: A multi-source dataset construction and preprocessing module is used to construct a multi-source dataset based on the charging behavior data, weather data, and holiday data of EV users in the community; Cleaning and standardizing the multi-source data set using a multi-source data preprocessing model to obtain a processed multi-source data set; A two-layer clustering processing module, configured to perform clustering processing on the processed multi-source data set using a two-layer clustering model to obtain charging type features and user type features; A user charging demand prediction module is used to extract input features from the processed multi-source data set; input the charging type features, the user type features and the input features into a charging demand prediction model, perform prediction processing through the LightGBM algorithm, and obtain user charging demand.

7. The device for predicting charging demand of individual electric vehicle users according to claim 6, characterized in that: In the multi-source dataset construction and preprocessing module, the charging behavior data of EV users in the community includes: EV starting charging time, charging amount and charging duration data.

8. The device for predicting charging demand of individual electric vehicle users according to claim 7, characterized in that: In the multi-source data set construction and preprocessing module, in the process of cleaning and standardizing the multi-source data set using the multi-source data preprocessing model, the multi-source data set is cleaned using a three-level data cleaning strategy; the expression of the processed multi-source data set is: A=[A1,A2,…,A i ,…,A N ] Where A i is the i-th data, i = 1, 2,…, N; N is the total number of data.

9. The device for predicting charging demand of individual electric vehicle users according to claim 8, characterized in that: In the two-layer clustering processing module, the two-layer clustering model includes a data layer and a user layer; the data layer performs cluster analysis on the processed multi-source data set using a charging type clustering strategy based on the HDBSCAN algorithm to extract the charging type features; The user layer is a user type clustering strategy based on the K-means algorithm, which characterizes the user's long-term charging behavior according to the charging type characteristics to obtain the user type characteristics; The quality of the clustering results of the charging type clustering strategy based on the HDBSCAN algorithm is evaluated by the Davies-Bouldin index; the calculation formula of the Davies-Bouldin index is: Where DB is the Davies-Bouldin index; H is the number of mutually exclusive clusters; and are the average distances from all samples in the rth cluster and the tth cluster to the center of the class to which they belong, respectively; ω rt is the class center distance of different categories; The user's charging behavior is characterized by the user's proportional combination feature vector, and the expression of the proportional combination feature is: Where x u is the proportional combination feature vector; u is the individual EV user; U is the total number of EV users; N is the number of samples belonging to the hth charging type among user u; u is the total number of samples of user u; The number of samples of the hth charging type for user u accounts for N u proportion.

10. The device for predicting charging demand of individual electric vehicle users according to claim 9, characterized in that: In the user charging demand prediction module, during the prediction process using the LightGBM algorithm, the objective function expression of the LightGBM algorithm is: Where, δ i is the true value of the i-th sample in D; is the predicted value of the i-th sample; L is the regression loss function, which is set to the mean absolute error; Reg1 and Reg2 are the L1 and L2 regularization terms respectively; λ1 and λ2 are the coefficients of the L1 and L2 regularization terms respectively; N is the total number of data.

Citation Information

Cited By

  • Water resource scheduling demand prediction method based on multi-source data analysis

    CN122047650A