EV Charging Load Prediction Method Based on Hierarchical Quantum Clustering and User Portrait
Through a method based on hierarchical quantum clustering and user portrait, combined with battery degradation model and Monte Carlo simulation, the problem of difficulty in accurately predicting EV charging load in the prior art is solved, and more accurate load prediction and grid management benefits are achieved.
Patent Information
- Application Number
- CN202410621626.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-05-17
AI Technical Summary
The prior art is difficult to accurately predict the charging load of electric vehicles (EVs), especially considering the variability of user behavior and degradation of battery capacity.
Using a method based on hierarchical quantum clustering and user portrait, the charging load of EV is predicted by obtaining user feature data, applying an improved quantum clustering algorithm, and combining battery degradation model and Monte Carlo simulation.
This method can more accurately predict the charging behavior of people with different characteristics, reduce the cost of grid operation, reduce the load pressure of grid, and more effectively predict the actual situation of regional charging stations.
Smart Images

Figure CN118539420B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of load forecasting, and in particular to an EV charging load forecasting method based on hierarchical quantum clustering and user portraits. Background Art
[0002] Under the background of global warming, increasing proportion of new energy power generation, and shortage of traditional fossil energy, electric vehicles have been well developed. Their low pollution, smooth driving experience, and relatively low energy replenishment price have attracted more and more people's favor. However, the concentrated charging of a large number of electric vehicles (EVs) at a certain time and place will accumulate a non-negligible load on the power grid. Therefore, effective prediction and planning of the charging behavior of electric vehicles can bring good benefits to society, the power grid, and users.
[0003] The factors affecting the charging of electric vehicles are complex and variable. The charging behavior of users is comprehensively affected by many uncertain factors, and there is also a greater or lesser correlation between various factors, which further increases the difficulty of load forecasting. The main influencing factors include but are not limited to the battery capacity of electric vehicles, charging methods, charging duration, cruising range, temperature of the weather, and psychological conditions of users.
[0004] The current research objects mainly include three types: private electric vehicles, electric taxis, and electric buses. The load forecasting of private electric vehicles is the most difficult. The previous stage of research mainly focused on using stochastic mathematics methods to overall forecast the spatio-temporal distribution of the load. With the enrichment of research methods, the research began to adopt artificial intelligence, big data, etc., and use methods such as intelligent agent multi-agent and simulating user behavior to refine and simulate the model. Summary of the Invention
[0005] Technical Objective: Aiming at the defects in the prior art, the present invention discloses an EV charging load forecasting method based on hierarchical quantum clustering and user portraits.
[0006] Technical Solution: To achieve the above technical objective, the present invention adopts the following technical solutions.
[0007] An EV charging load forecasting method based on hierarchical quantum clustering and user portraits includes the following steps:
[0008] Step A: Obtain user feature data, calculate the weight of each feature after removing strongly correlated features; wherein, obtain user feature data, perform correlation analysis on user features and remove strongly correlated features; after screening, perform normalization processing on the remaining feature data, and apply the entropy weight method to determine the weight of each feature;
[0009] Step B: Apply the improved quantum clustering algorithm to the weight values of each feature for clustering, and adjust the hyperparameters to obtain a series of user groups and group labels with different numbers of clusters. Then, screen the number of clusters through the internal evaluation index of clustering, and use the kernel density estimation method to obtain the probability density function for each feature of the user data in each group;
[0010] Step C: For the user feature data after removing strongly correlated features in Step A, divide each user feature into several intervals and number them, and arrange them in combination as the user ID. Based on the distribution of the number of people in different intervals, extract the user ID as a virtual user. Determine which group the user ID belongs to by judging the distance to the group center in Step B. After determining the user group, extract the probability density function of this group to determine the user features of the virtual user;
[0011] Step D: Obtain the original vehicle parameters of the electric vehicle: Score each vehicle model according to the different user features of the virtual user to determine the appropriate vehicle model for each user, thereby obtaining the vehicle parameters. The obtained vehicle parameters and the user feature parameters together form the user portrait unique to the virtual user;
[0012] Step E: Based on the driving mileage and charging times of the virtual user, use the battery degradation model to estimate the change in the battery capacity of the virtual user, and correct the battery capacity and cruising range based on the losses during the charging and discharging processes. Considering the change in the internal resistance of the electric vehicle during the constant current and constant voltage charging process, the charging stage is decomposed into a constant current charging stage with constant power and a constant voltage charging stage with power attenuation;
[0013] Step F: Use the obtained virtual user travel data and vehicle data as inputs, design the charging behavior of the virtual user in a day, and on the basis of Step E, use the Monte Carlo method to simulate a large number of randomly selected virtual users to obtain the load prediction data.
[0014] Beneficial effects:
[0015] 1. The present invention discloses an EV charging load prediction method based on hierarchical quantum clustering and user portrait, which can effectively extract the characteristics of user travel data, effectively express the influence of each variable reflected by the data on the load, and through clustering, obtain the distribution characteristics of vehicles and charging when people with different characteristics drive electric vehicles, which is more in line with the actual situation.
[0016] 2. The present invention considers that people with different characteristics have different performances in battery capacity, battery degradation rate, daily driving mileage, vehicle driving speed, charging time, etc., and can more accurately predict the load situation, thereby reducing the operating cost of the power grid and the load pressure on the power grid.
[0017] 3. The present invention takes into account the charging behaviors of users of electric vehicles in constant current constant voltage charging and fast / slow charging modes, and can more effectively predict the actual situation of regional charging stations. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of a method for predicting EV charging load based on hierarchical quantum clustering and user profiling according to the present invention;
[0019] Figure 2 It is a heat map of correlation coefficients of user characteristics;
[0020] Figure 3 It is a graph of the proportion of user characteristic weights;
[0021] Figure 4 It is a flowchart of the hierarchical quantum clustering algorithm;
[0022] Figure 5 It is a graph of the change in the number of clusters in hierarchical quantum clustering;
[0023] Figure 6 It is a graph of the change in internal evaluation indicators of clustering;
[0024] Figure 7 It is a schematic diagram of an excerpt of the probability distribution function of user characteristics for each group;
[0025] Figure 8 It is a probability graph of user ID distribution;
[0026] Figure 9 It is a schematic diagram of an excerpt of the heat map of scores for each vehicle model of some users under the auxiliary decision-making model;
[0027] Figure 10 It is a schematic diagram of an excerpt of the heat map of scores for each vehicle model of some users under the linear logistic regression method;
[0028] Figure 11 It is a charging curve graph of an electric vehicle battery;
[0029] Figure 12 It is a charging load prediction curve graph of an electric vehicle under the auxiliary decision-making model;
[0030] Figure 13 It is a charging load prediction curve graph of an electric vehicle under the linear logistic regression method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The following further describes and explains a method for predicting EV charging load based on hierarchical quantum clustering and user profiling according to the present invention with reference to the drawings and embodiments. It should be noted that the method of the present invention is applicable to charging piles or charging stations for predicting the charging load of electric vehicles (EVs) that may have charging requirements at the present charging pile or charging station.
[0032] Embodiment
[0033] As shown in the Figure 1 accompanying drawings, an EV charging load prediction method based on hierarchical quantum clustering and user portraits in this embodiment includes the following steps:
[0034] Step A: Obtain user feature data, calculate the weight of each feature after removing strongly correlated features; among them, obtain user feature data, perform correlation analysis on user features and remove strongly correlated features. After screening, perform normalization processing on the remaining feature data, and apply the entropy weight method to determine the weight of each feature; it includes the following steps:
[0035] Step A1: After obtaining user feature data, preprocess the user feature data, remove strongly correlated user feature data, and obtain processed user feature data;
[0036] In this embodiment, the data of several users in a certain place is used as a sample, and a total of nine original user feature data are obtained, including vehicle usage years, trip start time, trip end time, trip duration, driving mileage (that is, travel mileage), end point stay time, annual income, age, and education. Calculate the correlation between pairwise features using the Pearson correlation coefficient, and then screen and remove the user features with higher correlation. The formula for the Pearson correlation coefficient is:
[0037]
[0038] In the formula, X i and Y i represent the numerical values of user feature X and user feature Y of the i-th user, 1 ≤ i ≤ n, where n represents the number of users, that is, the number of samples, represents the numerical mean of all users corresponding to user feature X, represents the numerical mean of all users of user feature Y, Pearson represents the Pearson correlation coefficient, and the correlation coefficient value between user feature X and user feature Y is obtained. User feature X and user feature Y are selected from the above nine features.
[0039] After calculation, the Pearson correlation coefficient matrix results of 9 user features are as shown in the Figure 2As shown. According to the Pearson correlation coefficient matrix, it can be seen that there is a strong correlation between the trip start time and the trip end time, and between the trip duration and the driving mileage. Therefore, one variable is selected from each pair of the two pairs of variables for elimination, so that the original 9 variables become 7 variables for subsequent analysis, and 7 user characteristic data after processing are obtained. The 7 user characteristics are the vehicle usage years, the trip start time, the driving mileage (i.e., the travel mileage), the end point stay time, the annual income, the age, and the education.
[0040] Step A2: After normalizing the processed user characteristic data, calculate the weight value of each characteristic according to the entropy weight method;
[0041] For the data after eliminating strongly correlated characteristics, that is, the processed user characteristic data, perform normalization, and then apply the entropy weight method to determine the weights of each characteristic, and the sum of the weights of all characteristics is 1; the normalization formula is:
[0042]
[0043] In the formula, N ij represents the user characteristic data after normalization of the jth user characteristic of the ith user, and org ij represents the original value before normalization of the jth user characteristic of the ith user. min(org j ) represents the minimum value of all user data corresponding to the jth user characteristic, and max(org j ) represents the maximum value of all user data corresponding to the jth user characteristic.
[0044] Subsequently, calculate the weight value of each user characteristic according to the normalized data using the entropy weight method. The formula is as follows:
[0045]
[0046]
[0047] In the formula, K ij represents the proportion of the ith user in all users of the jth user characteristic. ω j represents the proportion of the jth user characteristic in all user characteristics. n represents the number of samples, that is, the number of users, and J represents the total number of user characteristics. In this embodiment, J takes the value of 7.
[0048] In this embodiment, the processed user characteristic data is used as the input matrix, with rows representing each user and columns representing the characteristics of the user. The weights of each characteristic calculated after normalization are as shown in the appendix Figure 3 shown.
[0049] Step B: Apply the improved quantum clustering algorithm to the weight values of each feature for clustering, and adjust the hyperparameters to obtain a series of user groups and group labels with different numbers of clusters. Then, screen the number of clusters through the internal evaluation index of clustering, and use the kernel density estimation method to obtain the probability density function for each feature of the user data in each group.
[0050] Step B1: Construct a data set {x} based on the weight values of each feature as the input matrix of the hierarchical quantum clustering algorithm; multiply the weights of each feature by the corresponding numerical values of the feature to obtain a new data set {x} as the input matrix of the hierarchical quantum clustering algorithm; the data set {x} is user feature data with n rows and J columns.
[0051] Step B2: Construct the potential function and wave function in the hierarchical quantum clustering algorithm.
[0052] Use the wave function to characterize the distribution of the processed input matrix data points in space, introduce the concept of quantum potential energy, and use the Schrödinger equation to characterize the quantity value when the data points tend to change to a stable solution. Find the class relationship in the data set by looking for the minimum value of the qubit. Among them, the input matrix data points represented by the wave function do not have an actual definite position, but appear at different positions in space with different probabilities. The quantum potential represents the energy when a quantum particle moves. Through the particle motion characteristics described by the Schrödinger equation, the quantum potential represents the "difference" quantity value between different data points in the clustering algorithm. Use the potential energy surface to characterize the concavity and convexity of the data set. The low points of the depression represent high-density regions. Apply the method of gradient descent, and the points tend to gather from high places to the depression, thus forming clusters. Here, the calculation formula of the potential function is:
[0053]
[0054] where V(x) is the potential function, ψ(x) is the wave function, x is the data in the data set {x}, is a row vector with dimension J, and x i is the i-th data in the data set {x}, and is a row vector with dimension J. 1 ≤ i ≤ n, n represents the number of users, that is, the number of samples, σ is the Gaussian width estimated by the Parzen window, and is the hyperparameter of the algorithm. The value of σ will affect the clustering result. The larger σ is, the more data points a cluster can contain. Among them, the calculation formula of the wave function is:
[0055]
[0056] where cst is a constant factor, and here it is taken as 1;
[0057] Step B3: Use the potential function and the improved quantum clustering algorithm to cluster the data set {x} to obtain the clustering result; use the internal evaluation indexes of clustering, such as SSE, DB index, and silhouette coefficient, to evaluate the selection of the number of clusters.
[0058] Among them, when the hyperparameter σ selects different step sizes, the change in the number of clusters is as follows Figure 5 shown. The initial value of the hyperparameter σ is 0.0001, and the step sizes are selected as 0.00001, 0.00005, and 0.0001. As the hyperparameter σ increases, the change in the number of clusters under different step sizes is different. The internal evaluation indexes of clustering, namely the sum of squared errors (SSE), the DB index, and the silhouette coefficient, are used to evaluate the selection of the number of clusters. The calculation formula of the sum of squared errors (SSE) is:
[0059]
[0060] In the formula, C represents the total number of groups in the cluster, c represents the index of the group in the cluster, where 1 ≤ c ≤ C, and grp c represents the c-th group containing multiple x, x is the data in the dataset {x}, and is a row vector with dimension J. l c represents the group center of group c.
[0061]
[0062] S c is the average distance from each data in group c to the group center, and M c,cj represents the distance between the center of group c and the center of group cj, where S c and M c,cj are calculated according to the following formulas:
[0063]
[0064] M c,cj = ||l c - l cj ||2
[0065] In the formula, u represents the sample index in the group, 1 ≤ u ≤ U c , and U c represents the total number of samples in group c. GX u represents the value of the u-th sample in group c and is a row vector with dimension J. l c represents the group center of group c, and l cj represents the group center of group cj, and ||||2 represents the two-norm. The silhouette coefficient S c of the clustering is calculated using the matlab toolbox.
[0066] The index results of each cluster are calculated, and the curves of the three indexes changing with the decrease in the number of clusters are as shown in the appendix Figure 6 shown. According to the characteristics of the indexes, select the point where there is a jump in the SSE curve, the DB index is at a low point, and the silhouette coefficient is as large as possible. Thus, multiple clusters are screened out.
[0067] Step B3 includes the following steps:
[0068] Step B31: Set the initial value of the hyperparameter σ to 0.0001, and use the dataset {x} normalized and weighted in Step B1 as the input matrix data; here, data is an observation matrix, and the dataset {x} is user feature data with n rows and J columns; each row represents a user, and each column represents a user feature;
[0069] Step B32: Set the objective function as an unconstrained minimization problem, select to use the optimization algorithm toolbox built in matlab, and select the quasi-Newton method to provide gradient information for the subsequent objective function f min , where the gradient information is automatically supplemented by the toolbox, which can accelerate the convergence speed of the function;
[0070] Step B33: Define the objective function f min , that is, the potential function V, a function used to calculate the potential energy and potential energy gradient of the replication points, which requires the input of data and σ. Among them, V has the same number of rows as x, and the potential energy value corresponds to the point represented by x. The expression of the potential function V is as follows:
[0071]
[0072] Among them, x is the data in the dataset {x}, which is a row vector with dimension J, x i is the i-th data in the dataset {x}, which is a row vector with dimension J. 1 ≤ i ≤ n, n represents the number of users, that is, the number of samples, and p(x) is a probability density function, which is the normalization of the Gaussian kernel function generated by the data points, expressed as follows:
[0073]
[0074] Among them, Z is the normalization factor, which is the sum of the Gaussian kernel function values of all data points, and its calculation formula is:
[0075]
[0076] Step B34: Initialize the index c of the group in the clustering, that is, the cluster number c is 1, and i is also used as the data point index, initialized to 1, and the cluster number matrix TC is 0. Among them, the cluster number c is used to assign which category the data point belongs to. When the data points belong to the same cluster, the cluster numbers of these points are the same. AL is used to record the position of the data point in the matrix to determine which data points have been assigned cluster numbers. The cluster number matrix is used to record the current cluster numbers of all data points. When the data point has not been assigned a cluster number, its number value is 0, and when it has been assigned a number, its number is c.
[0077] Step B35: Determine whether all data points have been processed completely, that is, whether all data points have been assigned cluster numbers. The judgment condition is whether AL is an empty set. When AL is an empty set, it means that the cluster numbers of all data points in the cluster number matrix TC are not 0. Therefore, there is no data point index with a value of 0 in AL, and the sub-loop ends at this time. If AL is not an empty set, go to Step B36; if AL is an empty set, go to Step B38;
[0078] Step B36: Determine the Euclidean distance between points. Calculate the distance between the points according to the indices of the unnumbered data points in the current cluster number matrix TC. When the distance is less than dr, it means that these points can be assigned to the same cluster. Set the cluster number of this point set to c and go to Step B37. Here, the distance dr is defined as one-tenth of the hyperparameter σ. When the distance does not satisfy being less than dr, do not assign a cluster number and directly go to Step B37;
[0079] Step B37: Increment the cluster number c by 1, find the next point index in the cluster number matrix TC that is not 0 and assign it to i, that is, sequentially find the data points that have not been assigned to cluster classes from the cluster number matrix TC and obtain their indices i, and then return to Step B35;
[0080] Step B38: Determine whether the maximum value of the cluster numbers in the current cluster number matrix TC is small enough. When the maximum value of the cluster numbers is 1, it means that all data points have been assigned to one category, and at this time, further clustering cannot be performed again, and the overall loop ends. When the maximum value of the cluster numbers is not 1, slightly increase σ. Here, the step size is selected as 0.00001 each time to make the change in the number of clusters relatively stable. Finally, return to Step B32. As σ increases, the number of clusters gradually decreases, that is, the maximum value of the cluster numbers gradually decreases. Before each increase in σ, store the current number of clusters and the clustering labels of the data. The overall algorithm flow chart is as Figure 4 shown.
[0081] Step B4: Obtain groups with different numbers of clusters according to the clustering results, and at the same time obtain the user data and group labels under different groups for this number of clusters. Apply the kernel density estimation method to the user data of each group to obtain the probability density function of each feature of each group. The calculation formula for the kernel density estimator or probability density function is:
[0082]
[0083] In the formula, x is the data in the data set {x}, is a row vector with dimension J, x iThe \(i\)-th data in the data set \(\{x\}\) is a row vector with dimension \(J\). \(1\leq i\leq n\), where \(n\) represents the number of users, i.e., the number of samples, \(\sigma\) is a hyperparameter, \(h\) is a smoothing coefficient or bandwidth, and \(c\) is a constant; \(L\) is a kernel function, which adopts a normal distribution form here. The probability distribution diagrams of some grouped data obtained by fitting are shown in the appendix Figure 7 As shown, where each row represents a group and the columns represent 7 user characteristics. Since some groups have fewer individuals, the KMEANS method is used to cluster the individuals with fewer groups again, and the kernel density estimation method is also used to estimate their probability distribution
[0084] Step C: For the user characteristic data after removing strongly correlated features in Step A, divide each user characteristic into several intervals and number them, and arrange and combine them as user IDs; based on the distribution of the number of people in different intervals, extract user IDs as virtual users. Through the distance from the group center in Step B, determine which group the user ID belongs to. After determining the user group, extract the probability density function of this group to determine the user characteristics of this virtual user
[0085] In some embodiments of the present invention, for the user characteristic data after removing strongly correlated features in Step A, each user characteristic of the user is divided into 3 to 6 intervals according to the size value, and each interval is numbered. The number combinations corresponding to each feature are combined as user IDs, and the number of users in different intervals of each feature is counted, and the probability of the user being in this interval is set according to the number of people
[0086] Among them, each user characteristic is grouped according to the numerical value as shown in the feature coding table of Table 1. The first row represents the number represented by this interval. A seven-digit code is extracted from a user individual with 7 characteristics as the user ID. Among them, according to the different numbers of people in each interval, the probability that a number falls into a certain interval is determined. The probabilities of each feature are shown in the appendix Figure 8 . Among them, Education 1 - 3 represents a high school education or below, 4 - 6 represents a bachelor's degree, and 7 - 8 represents an education above a bachelor's degree. By dividing intervals, users are first broadly defined to match the generality and privacy of the statistical and analysis data when users select vehicles in existing survey materials. The specific implementation steps are described in Step D
[0087] Table 1
[0088]
[0089] Step D: Obtain the original vehicle parameters of the electric vehicle: Combine the different user characteristics of the virtual user to score each vehicle model to determine the appropriate vehicle model for each user, thereby obtaining vehicle parameters; the obtained vehicle parameters and the user characteristic parameters together form a unique user portrait of the virtual user; use an auxiliary decision-making model or linear logistic regression method to calculate the score. The process of calculating the score by the auxiliary decision-making model includes:
[0090] Obtain the original vehicle parameters of the electric vehicle, including obtaining five vehicle characteristic data of the electric vehicle, and use two objective weight assignment methods to obtain two weight values of the same vehicle characteristic, forming a weight interval of large and small values; combine the user ID in step C, and according to the different intervals where the user characteristics are located, select different weights of vehicle characteristics, and combine the user characteristics of the virtual user to score each vehicle model. According to the positive or negative correlation correspondence between the user characteristics and the vehicle characteristics, divide the weights of the user characteristics and the vehicle characteristics into the same number of intervals, and according to the user ID in step C, select the maximum value in the corresponding interval of the vehicle characteristics as the weight. Among them, according to the method that the vehicle price is positively correlated with the user's annual income, the acceleration is negatively correlated with the age, the cruising range is positively correlated with the user's driving mileage, the charging time is positively correlated with the end-point stay time, and the energy consumption is positively correlated with the education, divide the weights of the user characteristics and the vehicle characteristics into the same number of intervals, and according to the user ID in step C, select the maximum value in the corresponding interval of the vehicle characteristics as the weight. Specifically:
[0091] Use two objective weight assignment methods: the entropy weight method and the CRITIC method to determine the weight interval, and select the large and small values of the interval according to the user's characteristics to determine the scoring weight, score the vehicle models according to the weight, obtain the vehicle model with the highest score for this user, extract the parameters of this vehicle, and jointly use them with the user characteristic parameters for Monte Carlo simulation. The mathematical model of the entropy weight method is the same as that in step A2:
[0092]
[0093]
[0094] In the formula, K io represents the proportion of the i-th user in the o-th vehicle characteristic among all users of this vehicle characteristic, N io represents the normalized user characteristic data of the i-th user in the o-th vehicle characteristic. ω o represents the proportion of the o-th vehicle characteristic among all vehicle characteristics, n represents the number of samples, that is, the number of users, O represents the total number of user characteristics. In this embodiment, O takes the value of 5, that is, the vehicle characteristics include 5 characteristics of the vehicle price, acceleration, cruising range, charging time, and energy consumption.
[0095] The mathematical formula of the CRITIC method is as follows, and thus establish its auxiliary decision-making model for assisting users to select electric vehicles. The vehicle characteristics participating in the evaluation include 5 characteristics of the vehicle price, acceleration, cruising range, charging time (the fast charging power is simplified as the charging time), and energy consumption.
[0096]
[0097] Among them, W o represents the weight of the o-th vehicle feature, SD o represents the standard deviation of the o-th vehicle feature, and R o represents the correlation coefficient matrix of the o-th vehicle feature with other vehicle features;
[0098] According to two weight assignment methods, two weight values ω o and W o of the same vehicle feature are obtained, forming a value range. According to the user ID in step C, different weights are selected according to the different intervals where the user is located. Among them, according to the method that the vehicle price is positively correlated with the user's annual income, the acceleration is negatively correlated with the age, the cruising range is positively correlated with the user's driving mileage, the charging time is positively correlated with the end stay time, and the energy consumption is positively correlated with the education, the weights of the user features and the vehicle features are divided into the same number of intervals, and the weights are selected. For example, when the higher the annual income interval where the user is located, the greater the proportion of the vehicle price. According to the four intervals of the annual income, the weight interval of the vehicle price is also divided into four sub-intervals. When the user's annual income is in the first interval, the weight of the vehicle price is selected as the maximum value of the first interval. The same method is used for other vehicle features. It should be noted that the older the age, the lower the weight corresponding to the acceleration, the shorter the stay time, the higher the weight corresponding to the charging time, the higher the education level, the higher the weight corresponding to the energy consumption, and the cruising range and the user's driving mileage also have a positive impact. Multiply the obtained weights by the five normalized vehicle parameters respectively, and then sum them to obtain the scoring value. Among them, the heat map of the scoring distribution of different users for 96 vehicles is as attached Figure 9 where different gray colors represent the scoring levels. The closer to black, the higher the score, and the lighter the color, the lower the score.
[0099] To avoid the singularity of the auxiliary decision-making model, the linear logistic regression method is used as the alternative scoring method. The expression of the linear logistic regression method formula is:
[0100]
[0101] Among them, a0, a1, a2, a3,..., a q are regression coefficients, which are selected or set according to the actual situation. In some embodiments of the present invention, they are obtained based on the questionnaire survey data, relevant paper datasets and inferences. Among them, ft1, ft2, ft3,......, ft q correspond to user features and vehicle features respectively. For user features, it is the interval number corresponding to the user feature, that is, the first row of the feature coding table. For vehicle features, it is the specific parameter after normalization. pb is the probability of making this choice. The higher the probability, the higher the score of this user for this vehicle model;
[0102] The estimated value of the regression coefficient is obtained by using a linear logistic regression model, where the selected features f include user features and vehicle features, including age, education, annual income, vehicle price, energy consumption, fast charging time, acceleration, and cruising range. Other features are not included in the linear logistic regression model for calculation because they have little impact on vehicle scoring. Multiply the obtained regression coefficient by the corresponding feature value to get the logarithm value on the right side of the equation, and then calculate the pb value based on this logarithm value, which represents the satisfaction degree of users choosing different models, that is, the score. Select the largest value as the model chosen by the user, and output the vehicle parameters of this vehicle. Among them, the heat map of the score distribution of 96 vehicles by different users is attached Figure 10 , different shades of gray represent the high and low scores. The closer to black, the higher the score, and the lighter the color, the lower the score.
[0103] The scoring algorithm generates a new user id (interval number) according to the user intervals divided in step C and Figure 8 the probability of the feature interval. Reverse lookup the intervals where each user feature is located according to the obtained user id and get the median of the interval. Find the group with the smallest Euclidean distance based on the median as the group to which the user belongs. Look up the probability density function of this group in step B4. Generate specific values of 7 user features, User_data, for the user according to the interval to which the id belongs. Use it as the input value of the user features in the scoring system. The scoring part calculates the scores of all vehicles for a user and selects the vehicle with the highest score as the vehicle feature of this user, and outputs the user feature parameter User_data and the vehicle parameters VEH_data of the obtained vehicle. When using the weighted method for scoring, different weights are selected according to the size of the interval number of the user id, and scores are calculated only for different models according to the weights. When using the linear logistic regression for scoring, the user feature parameters of the same user and the vehicle parameters of different models are jointly calculated to obtain the score value. The following are the specific implementation steps:
[0104] Step D1: Read the probability distribution function Kernel of the user feature j in group c of step B j,c , set the interval boundaries of the vehicle usage years as 1, 4, 7, 9, 11, 15; the interval boundaries of the trip start time as 400, 800, 1200, 1500, 1900, 2300. A total of 7 user features are set in sequence, which is Group j ;
[0105] Step D2: Input the interval boundaries of the set user features into the histcounts function to obtain the number of users Num of user feature j and interval et c,et , and the total number of users Tnum in interval et j,et , take the ratio of the two to get the interval probability Ratio j,et , Derive the median Central of the interval based on the interval set in step D1 j,et ;
[0106] Step D3: Use the randsample function to generate a random number. The interval of this random number is from a minimum of 1 to a maximum of the feature interval limit minus 1, and input the interval probability Ratio j,et Ensure the generation probability of different values, and obtain the random number Rdnum of user feature j j This is the user ID;
[0107] Step D4: Based on the random number Rdnum of user feature j j Use the value as an index to determine the interval et, and obtain the median Central of interval et j,et Output the value as User j Assume the user is at the midpoint of the interval, which is used to find the nearest group center;
[0108] Step D5: Set the number of loop iterations to the number of groups ωnum obtained by hierarchical quantum clustering. Calculate the mean of all data in each group to obtain the clustering center Cmean of this group j,ωnum ;
[0109] Step D6: Set the number of loop iterations to the number of groups ωnum obtained by hierarchical quantum clustering again. Calculate the Euclidean distance between User j and the clustering centers Cmean of each group j,ωnum Compare to obtain the minimum distance value, and extract the grouping ωnum c ;
[0110] Step D7: Based on the grouping ωnum c Extract the probability distribution function Kernel of seven user features c ;
[0111] Step D8: Find the index Idx of the first interval boundary larger than the user feature value j to determine the interval where the user feature value is located;
[0112] Step D9: Set a while loop to randomly generate a random number that conforms to the probability distribution function Kernel c Only when the generated user feature value is within the first interval corresponding to the index Idx of Group c minus 1 can the value be output (break), and the output is the user feature data User_data in one row with seven columns; j Step D10: Calculate the weight interval Weight of five vehicle features according to the entropy weight method and CRITIC method formulas
[0113] j ;
[0114] Step D11: Divide Weight j into et intervals. According to the positive or negative correlation correspondence between user characteristics and vehicle characteristics, map them to the Ldx j -th interval of Weight j . Among them, the reverse correspondence such as acceleration and age is the (et - Idx) j -th interval of Weight j to obtain the weight Wg of each user characteristic j ;
[0115] Step D12: Preprocess the weight Wg of each user characteristic j . If the sum of the weights is greater than 1, divide it by 0.999 until the sum of the weights is approximately equal to 1 after rounding to two decimal places. If the sum of the weights is less than 1, multiply it by 1.001 until the sum of the weights is approximately equal to 1 after rounding to two decimal places, and output the final weight value Wg j ;
[0116] Step D13: Normalize the vehicle characteristic data. Multiply the vehicle parameters of each row by the weight Wg j , and then sum each row to obtain the final score value Score of vehicle model v v . Find the vehicle model v with the highest score value and output its vehicle data VEH_data;
[0117] Step D14: Perform linear logistic regression fitting on the original data to obtain regression coefficients a0, a1, a2, a3,..., a q , calculate the probability pb according to the linear logistic regression model formula, so as to obtain the probability pb that this user selects vehicle model v v , find the vehicle model v with the highest score and output its vehicle data VEH_data.
[0118] Step E: Based on the driving mileage and charging times of the virtual user, use the battery degradation model to estimate the change in the battery capacity of the virtual user, and correct the battery capacity and cruising range based on the losses during the charging and discharging processes. Considering the change in the internal resistance of the electric vehicle during the constant current and constant voltage charging process, the charging stage is decomposed into a constant current charging stage with constant power and a constant voltage charging stage with power attenuation:
[0119] As the number of charging times of the electric vehicle increases, its battery capacity will gradually decrease, and the cruising range will also decrease. Among them, the battery capacity attenuation rate D Tp,bj is defined as:
[0120]
[0121] Among them, γ Tp,bj represents a variable related to the battery type bj and the temperature Tp. Taking the ternary lithium battery at 25°C, γ Tp,bj is taken as a constant. Cb j,r is the sum of the charge and discharge electricity during the charging and discharging interval between the rth charge and the (r - 1)th charge for the battery type bj, and K Tp,bj represents a variable related to the battery type bj and the temperature Tp. The battery type and temperature are the same as above and are taken as constants;
[0122] The capacity attenuation rate D of the obtained electric vehicle Tp,bj is in percentage. Therefore, the battery capacity after the rth charge can be obtained as:
[0123] A r = A r-1 (1 - D Tp,bj )
[0124] Among them, A r is the battery capacity after the rth charge, and A r-1 is the battery capacity after the (r - 1)th charge.
[0125] The charging process is divided into two categories: fast charging and slow charging (conventional charging). Among them, in the first stage, constant current charging is adopted, and the voltage rises slowly. Considering that the internal resistance of the battery is small when the state of charge is low, the charging power at this time is the constant rated charging power. In the second stage, when the battery capacity reaches 80%, the internal resistance of the battery increases, and the constant voltage charging method is adopted, and the current gradually decreases. At this time, the power decays exponentially and finally decays to 10% of the rated power. Among them, the power P cv in the constant voltage stage can be calculated by the formula:
[0126]
[0127] In the formula, T cc is the end time of constant current charging, T max is the time when the state of charge of the battery is 100%, P cc is the charging power of constant current charging, g is the charging efficiency, which also represents the loss constant of the charged electricity, and T x is the upper limit of integration of the variable limit integral. The discharge loss constant of the battery during discharge is ε. When calculating the cruising range, multiplying by the discharge loss constant represents the actual cruising range. The schematic diagram of two-stage charging is as shown in the appendix Figure 11 . P max represents the rated power and is equal to P cc . P min represents the minimum power after attenuation and is 10% of the rated power.
[0128] Taking the user characteristics and vehicle characteristics of the virtual user obtained in step D as inputs, calculate the attenuated battery capacity to replace the original battery capacity. Since the estimated battery capacity attenuation may exceed the standard regulations, at this time, calculate according to the minimum battery capacity that meets the travel requirements. When charging is required in step F, it needs to be processed in step E. After step E receives the charging demand from step F (charging required, current remaining battery power of the vehicle, required battery power of the vehicle), it determines whether fast charging is needed, calculates the charging time, further determines whether there is a constant voltage charging stage according to the charged power or SOC value, calculates the charging power according to the above judgment, and outputs the charging power and charging time as the load of a single charging behavior of this user, which is used for the accumulation in the load matrix in step F.
[0129] The following gives the specific implementation steps of step E:
[0130] Step E1: Extract User_data obtained in step D9, obtain the vehicle usage years and mileage of this user, simplify the annual mileage by multiplying the mileage by 365 and calculate the power consumption, calculate the charging times according to the power consumption and the cruising range (or battery capacity) in VEH_data, and calculate the battery capacity N under the usage years according to the battery capacity attenuation formula cpa , and use N cpa to replace the original battery capacity;
[0131] Step E2: Determine whether the total mileage exceeds the vehicle standard. If it exceeds, reset it according to the vehicle standard; otherwise, according to the assumed mileage. Determine whether the battery capacity after attenuation exceeds the travel requirement. If it exceeds, reset it to the battery capacity that meets the travel requirement; otherwise, calculate according to the attenuated battery capacity;
[0132] Step E3: Determine the SOC of the current vehicle and the required SOC value. If the sum of the two exceeds 0.8, a constant voltage charging mode is required, and go to step E5; otherwise, only constant power charging is required, and go to step E4;
[0133] Step E4: Determine whether to use the fast charging method according to step F. If the fast charging method is used, divide the required power by the fast charging power to obtain the charging time. If the slow charging method is used, divide the required power by the slow charging power to obtain the charging time, and output the charging power P and charging time T, where the fast charging power is obtained from VEH_data, and the slow charging power is uniformly set to 11 kwh;
[0134] Step E5: Set the constant current charging stage to be from 0% to 80% of SOC, and determine whether to use fast charging according to step F. If the fast charging method is used, calculate T according to the fast charging power ccf, that is, the required power divided by the fast charging power, where the required power is calculated based on the driving mileage, with the unit of kilowatt-hour, and the fast charging power unit is kilowatt per hour, which varies according to different vehicle models. According to the variable limit integral P cv Calculate the constant voltage charging time T cvf , that is, the difference between the charging completion time and the moment when the constant current switches to constant voltage charging. Among them, the charging completion time is the upper limit value obtained by reverse calculation based on the stored power and the variable limit integral. If the slow charging method is adopted, then the calculated T ccs and T cvs ;
[0135] Step E6: Whether it is fast charging or slow charging, according to the fitted P cv curve, sample P cv at equal intervals within the T cv time interval to obtain a series of discrete constant voltage charging powers as the output power P, and output the charging time T cv .
[0136] Step F: Use the obtained user travel data and vehicle data of the virtual user as inputs, namely User_data and VEH_data. The charging behavior of the virtual user in a day is designed in a loop. Based on Step E, the Monte Carlo method is used to simulate a large number of randomly selected virtual users to obtain load prediction data. Each loop of the Monte Carlo method simulates the behavior of a virtual user in a day. First, it judges when the virtual user has a charging demand according to the current battery level and the distance to the destination, and further judges whether to use fast charging or slow charging in combination with the stay time. If a charging demand occurs, it calls the charging calculation method in Step E to obtain the charging time and charging power. For behaviors with time exceeding 24 o'clock, they are counted after 0 o'clock. Finally, the charging loads of the virtual users obtained in each loop are accumulated to obtain the total charging load prediction result. Specifically, it includes the following steps:
[0137] Step F1: Use Monte Carlo simulation, set the total number of users to be simulated as mtn, set the individual unit as mti, and start the loop; extract the parameter values of the vehicle usage years, trip start time, driving mileage, end stay time, annual income, age, and education level of the user characteristics, calculate the score of user mti for each vehicle model using Step D, select the vehicle model with the highest score, and obtain parameters such as battery capacity, fast charging power, and energy consumption. Finally, obtain User_data and VEH_data. Set the charging identifier nc = 0 to indicate that the current user has no charging demand;
[0138] Step F2: Divide a day into 2400 sub-intervals, initialize the SOC, initialize the load matrix P, the rows of P represent time, with 2400 rows, and the columns represent the individual mti, and the values represent power;
[0139] Step F3: Calculate the power required for the trip at the start of the trip, and determine whether the SOC meets the travel demand. If it meets the demand, proceed to Step F4; otherwise, set the charging identifier nc = 1, start fast charging, and proceed to Step F7;
[0140] Step F4: Calculate the time to reach the destination and the remaining power, and determine whether the departure time plus the travel time exceeds 2399. If it exceeds, subtract 2399 from the arrival time. Determine whether the remaining power supports the consumption for the return trip. If it meets the return trip requirement, again determine whether the remaining SOC is lower than 5%. If the return trip power is sufficient and the SOC is greater than 5%, end the trip and the loop ends. Otherwise, set the charging identifier nc = 1 and proceed to Step F5;
[0141] Step F5: Statistically calculate the required power, and determine whether the stay time is sufficient when using slow charging. If the time is not enough, proceed to Step F6; if the time is sufficient, use the slow charging power and proceed to Step F7;
[0142] Step F6: Determine whether the stay time is sufficient when using fast charging. If it is sufficient, use the fast charging power and proceed to Step F7. If the time is not enough, extend the data in User data: the end point stay time until the power is satisfied. Use the power that supports the consumption for the return trip as the stored power and input it into Step E. Output the load matrix P of the charging time and charging power, and determine whether the charging completion time exceeds 2399. If it exceeds, proceed to Step F8; otherwise, end the loop;
[0143] Step F7: Determine which is longer between the stay time and the time required to fully charge. If the end point stay time is longer, the charging time is based on the fully charged power as the required power; if the fully charged time is longer, the power that can be charged according to the end point stay time is used as the required power. Input the current required power and the remaining power into Step E, statistically calculate the time interval in which the power is located, and record it in the load matrix P. Determine whether the charging completion time exceeds 2399. If it exceeds, proceed to Step F8; otherwise, determine whether charging is carried out at the start of the trip. If no trip operation has been carried out yet, proceed to Step F4; if there has been a trip, end the loop;
[0144] Step F8: Accumulate the part of the charging time that does not reach 2399 to the rows from the charging start time to 2399 in the load matrix P, and accumulate the excess part to the rows from 0 to the charging end time minus 2399 in the load matrix P; output the load matrix P and end the loop.
[0145] According to the load matrix P, output the electric vehicle charging load prediction result. The result graph includes the load prediction results of two scoring methods under different clusters, and the load prediction results brought by the hierarchical quantum clustering results of different numbers of clusters are as shown in Attachment Figure 12 ,13 Figure 12The predicted curve graph of the electric vehicle charging load for scoring method 1; Figure 13 The predicted curve graph of the electric vehicle charging load for scoring method 2.
Claims
1. A method for predicting EV charging load based on hierarchical quantum clustering and user portrait, characterized in that: The following steps are involved: Step A, obtaining user feature data, and calculating the weight of each feature after removing strongly correlated features; wherein, obtaining user feature data, performing correlation analysis on user features and removing strongly correlated features; normalizing the remaining feature data after screening, and applying the entropy weight method to determine the weight of each feature; Step B: Apply the improved quantum clustering algorithm to cluster the weight value of each feature, and adjust the hyperparameters to obtain a series of user groups and group labels with different cluster numbers, and select the number of clusters through the cluster internal evaluation index, and use the kernel density estimation method to obtain the probability density function of each feature of the user data of each group; Step B includes: According to the weight value of each feature, a data set {x} is constructed as the input matrix of the hierarchical quantum clustering algorithm; the weight of each feature is multiplied by the value corresponding to the feature to obtain a new data set {x} as the input matrix of the hierarchical quantum clustering algorithm; the data set {x} is user feature data with n rows and J columns; Construct the potential function and wave function in the hierarchical quantum clustering algorithm; the potential function calculation formula is: Where V(x) is the potential function, ψ(x) is the wave function, x is the data in the data set {x}, and x i is the i-th data in the data set {x}, 1≤i≤n, n represents the number of users, i.e. the number of samples, σ is the Gaussian width estimated by the Parzen window, and is a hyperparameter of the algorithm; the wave function calculation formula is: Among them, cst is a constant factor; Step C: for the user feature data after the strong correlation features are removed in step A, each user feature is divided into several intervals and numbered, and the permutations and combinations are used as user IDs; based on the distribution of the number of people in different intervals, the user ID is extracted as a virtual user, and the user ID is judged in which group by the distance to the group center in step B. After the user group is determined, the probability density function of the group is extracted to determine the user characteristics of the virtual user; Step D, obtaining original vehicle parameters of the electric vehicle: scoring each vehicle model in combination with different user characteristics of the virtual user to determine the vehicle model suitable for each user and thereby obtain vehicle parameters; the obtained vehicle parameters and the parameters of the user characteristics together constitute a user profile unique to the virtual user; Step E: Based on the driving mileage and charging times of the virtual user, a battery degradation model is used to estimate the battery capacity change of the virtual user, and the battery capacity and driving range are corrected based on the loss during the charging and discharging processes; considering the change of the internal resistance of the electric vehicle during the constant current and constant voltage charging process, the charging stage is decomposed into a constant current charging stage with constant power and a constant voltage charging stage with power attenuation; Step F: Using the obtained virtual user travel data and vehicle data as input, the charging behavior of the virtual user in one day is designed. Based on step E, a large number of randomly selected virtual users are simulated using the Monte Carlo method to obtain load forecast data.
2. The EV charging load prediction method based on hierarchical quantum clustering and user portrait according to claim 1 is characterized by: The step A comprises the following steps: Step A1: After obtaining the user feature data, pre-process the user feature data, remove strongly correlated user feature data, and obtain processed user feature data; Step A2: normalize the processed user feature data and calculate the weight value of each feature according to the entropy weight method.
3. The EV charging load prediction method based on hierarchical quantum clustering and user portrait according to claim 2 is characterized by: In step A2, the normalization formula is: Where N ij represents the user feature data of the i-th user after the j-th user feature is normalized, org ij represents the original value of the jth user feature of the ith user before normalization; min(org j ) represents the minimum value of all user data corresponding to the jth user feature, max(org j ) represents the maximum value of all user data corresponding to the jth user feature; The entropy weight method is used to calculate the weight value of each user feature. The formula is as follows: Where K ij represents the proportion of the i-th user in the j-th user feature among all the users of the user feature; ω j It represents the proportion of the jth user feature in all user features, n represents the number of samples, that is, the number of users, and J represents the total number of user features.
4. The EV charging load prediction method based on hierarchical quantum clustering and user portrait according to claim 1 is characterized by: The step B comprises the following steps: Step B1, constructing a data set {x} as an input matrix of the hierarchical quantum clustering algorithm according to the weight value of each feature; multiplying the weight of each feature by the numerical value corresponding to the feature to obtain a new data set {x} as the input matrix of the hierarchical quantum clustering algorithm; the data set {x} is user feature data with n rows and J columns; Step B2, constructing potential function and wave function in hierarchical quantum clustering algorithm; Step B3, clustering the data set {x} using the potential function and the improved quantum clustering algorithm to obtain a clustering result; using the internal evaluation indicators of clustering, SSE, DB index and silhouette coefficient to evaluate the selection of the number of clusters; Step B4: obtain groups with different numbers of clusters according to the clustering results, and obtain user data and group labels under different groups with the number of clusters; apply the kernel density estimation method to the user data of each group to obtain the probability density function of each feature of each group.
5. The EV charging load prediction method based on hierarchical quantum clustering and user portrait according to claim 4 is characterized by: The calculation formula of SSE in step B3 is: Where C represents the total number of groups in the cluster, c represents the index of the group in the cluster, where 1≤c≤C, grp c represents the cth group containing multiple x, where x is the data in the data set {x} and is a row vector with dimension J; l c represents the group center of group c; S c is the average distance from each data in group c to the group center, M c,cj represents the distance between the center of group c and the center of group cj, where S c With M c,cj Calculated according to the following formula: M c,cj =||l c -l cj ||2 Where u represents the sample index in the group, 1≤u≤U c , U c represents the total number of samples in group c; GX u represents the u-th sample value in group c, which is a row vector of dimension J; l c represents the group center of group c, l cj represents the group center of group cj, and || ||2 represents the two-norm.
6. The EV charging load prediction method based on hierarchical quantum clustering and user portrait according to claim 4 is characterized by: The calculation formula of the probability density function in step B4 is: In the formula, x is the data in the data set {x}, is a row vector with dimension J, and x i is the i-th data in the data set {x}, which is a row vector of dimension J; 1≤i≤n, n represents the number of users, i.e. the number of samples, σ is a hyperparameter, h is the smoothing coefficient or bandwidth, which is a constant; L is the kernel function, which adopts the normal distribution form.
7. The EV charging load prediction method based on hierarchical quantum clustering and user portrait according to claim 1 is characterized by: In step D, the score is calculated by using an auxiliary decision model or a linear logistic regression method. The process of calculating the score by the auxiliary decision model includes: obtaining the original vehicle parameters of the electric vehicle, including obtaining five vehicle feature data of the electric vehicle, using two objective weighting methods to obtain two weight values of the same vehicle feature, and forming a weight interval of large and small values; combining the user ID in step C, selecting different weights of the vehicle feature according to the different intervals of the user feature, and scoring each vehicle model in combination with the user features of the virtual user; The linear logistic regression formula is expressed as: where a0, a1, a2, a3, ..., a q is the regression coefficient, which is selected or set according to the actual situation, ft1, ft2, ft3, ..., ft q Corresponding to user characteristics and vehicle characteristics, for user characteristics, it is the interval number corresponding to the user characteristics, and for vehicle characteristics, it is the specific parameter after normalization; pb is the probability of making this choice, and the higher the probability, the higher the user's score for the model.
8. The EV charging load prediction method based on hierarchical quantum clustering and user portrait according to claim 7 is characterized by: Different weights of vehicle features are selected according to different intervals where user features are located, including: dividing the weights of user features and vehicle features into the same number of intervals according to the positive or negative correlation between the user features and the vehicle features, and selecting the maximum value in the interval corresponding to the vehicle features as the weight according to the user ID in step C.
9. The EV charging load prediction method based on hierarchical quantum clustering and user portrait according to claim 1 is characterized by: In step E, the electric vehicle battery capacity decay rate D Tp,bj Defined as: Among them, γ Tp,bj Represents variables related to battery type bj and temperature Tp, taken as a constant; C bj,r K is the battery type bj, the sum of the charge and discharge between the rth charge and the r-1th charge, Tp,bj represents the variable related to the battery type bj and temperature Tp, which is taken as a constant; The capacity decay rate D of the electric vehicle is obtained Tp,bj As a percentage, the battery capacity after the rth charge can be obtained as: A r =A r-1 (1-D Tp,bj ) Among them, A r is the battery capacity after the rth charge, A r-1 is the battery capacity after the r-1th charge.
10. The EV charging load prediction method based on hierarchical quantum clustering and user portrait according to claim 1 is characterized by: In step F, the Monte Carlo method is used to simulate a large number of randomly selected virtual users to obtain load forecast data: each cycle of the Monte Carlo method simulates the behavior of a virtual user for one day. First, the virtual user's charging demand is determined based on the current power and destination distance, and the fast charging or slow charging method is further determined based on the stay time. If a charging demand is generated, the charging calculation method of step E is called to obtain the charging time and charging power. Behaviors that exceed 24 o'clock are counted after 0 o'clock. Finally, the virtual user charging loads obtained in each cycle are accumulated to obtain the total charging load prediction result.
Citation Information
Patent Citations
Electric vehicle charging load forecasting method and system based on data mining
CN109325631A
Electric vehicle charging load prediction method based on Monte Carlo and deep learning
CN110570014A