Electric vehicle charging duration prediction method based on multi-factor influence fusion model

By using a multi-factor influence fusion model in the charging time prediction of electric vehicles and combining MLP and RF models for fusion prediction, the problems of prediction results deviation and overfitting in the prior art are solved, and high-precision and widely applicable charging time prediction are achieved.

CN120067683APending Publication Date: 2025-05-30KUNMING UNIV OF SCI & TECH

Patent Information

Application Number
CN202510126073.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing charging time prediction methods are difficult to achieve high-precision prediction under the influence of complex and multi-factors, and are prone to result deviations and overfitting.

Method used

The charging time prediction method for electric vehicles based on the multi-factor influence fusion model is adopted. By recording and cleaning new energy vehicle big data, key features are extracted and fusion prediction is combined with MLP and RF models, and model training and prediction are used to use backpropagation algorithm and regression decision tree algorithm.

Benefits of technology

It realizes high-precision prediction of charging time in different models and charging modes, avoids result deviation and overfitting of a single model, and has good generalization ability and applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067683A_ABST
    Figure CN120067683A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of new energy electric vehicles, and discloses an electric vehicle charging duration prediction method based on a multi-factor influence fusion model, and the method comprises the steps: carrying out the cleaning and segmentation of vehicle operation data collected by a big data platform; the charging power of each charging section is calculated, and the charging data obtained through segmentation is divided into a fast charging data set and a slow charging data set according to the power. The method comprises the following steps: respectively training a random forest model (RF) by using a fast charging data set and a slow charging data set, obtaining leaf node features according to an RF division rule, newly establishing a rule layer in a multi-layer perceptron model (MLP) for receiving and processing the RF leaf node features and deep features of an MLP hidden layer, realizing structural fusion of the model, and finally obtaining a charging duration prediction result in an output layer. According to the method, the influence of multiple factors including the charging power, the state of charge (SOC) and the temperature is comprehensively considered, the average error rate in the fast charging mode and the slow charging mode can be controlled within 5%, and accurate prediction of the charging duration of the electric vehicle is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of new energy electric vehicles, and particularly relates to a method for predicting the charging duration of electric vehicles based on a multi-factor influence fusion model. Background Art

[0002] The charging technology of new energy vehicles is one of the key areas to promote the popularization of electric vehicles. An efficient charging solution can solve the anxiety of electric vehicle users about the cruising range and charging time, and improve the driving experience of consumers. At the technical level, accurate prediction of the charging time can provide reliable data support, improve charging efficiency, optimize the allocation of charging station resources, extend battery life, and at the same time promote the in-depth application of big data and intelligent charging technologies in the energy field, accelerate the optimization and research and development of charging technologies, which is the key to promoting the development of new energy vehicle technologies; at the market level, accurate prediction of the charging duration directly improves the charging experience and satisfaction of users, and is an important means for new energy vehicle brands to enhance their market competitiveness. At the same time, accurate prediction of the charging duration can also optimize the operation efficiency of charging stations, promote shared charging and value-added services, create greater economic benefits for the market, and promote the development of the future market of intelligent vehicles. However, the accurate prediction of the charging duration is affected by multiple complex factors, including battery state (such as initial SOC, state of health SOH, temperature, etc.), charging pile power limitation, and energy consumption differences of different vehicles. These factors not only interact with each other, but also have the characteristics of dynamic change, resulting in the prediction of the charging duration becoming a highly non-linear and multi-variable complex problem.

[0003] At present, most prediction methods still rely on traditional empirical models or algorithms based on simple linear regression. Such methods have certain applicability when dealing with single variables or a small number of parameters, but it is difficult to ensure the accuracy of predictions in actual complex scenarios. Moreover, many studies only develop models for fixed scenarios or specific vehicle models, resulting in obvious deficiencies in generalization in cross-vehicle and cross-scenario applications. At the same time, due to equipment differences among different charging piles and anomalies in data recording, the quality of data collection will be affected, which also restricts the further development of charging duration prediction technology. Additionally, existing technologies generally lack in-depth mining of the time series characteristics during the charging process, making it difficult to capture the non-linear variation law of the charging curve, posing a significant challenge to the accurate prediction of charging duration. Some scholars have conducted corresponding research on charging duration prediction. The patent with the patent number 202410393292.3 discloses a method, device, equipment, and storage medium for predicting the remaining charging time of a vehicle. This method collects the vehicle condition information of the target vehicle, including the current SOC, target SOC, and battery parameters, selects vehicle charging data similar to the vehicle condition information of the target vehicle from historical data according to the vehicle condition information, and determines the reference remaining charging time of the vehicle in combination with the current output information of the charging pile. Based on the reference remaining time and the estimated remaining time of the vehicle, the remaining charging time of the target vehicle is predicted. By considering complex environmental conditions, accurate prediction of the remaining charging time of various vehicles is achieved. However, this method overly relies on the selection and comparison of historical data, which will greatly increase the workload. Moreover, this method needs to find vehicles similar to the vehicle condition information of the target vehicle as reference data. Due to the existence of differences in vehicle types and in-vehicle battery types, it is difficult to accurately predict the remaining charging time of different vehicle models. The patent with the patent number 202310605878.7 provides a method, system, and electronic device for predicting the remaining charging duration of new energy vehicles. This method determines the charging time consumption per unit battery power according to historical charging data, combines the current SOC and target SOC of the battery, estimates the empirical charging duration of the current battery, and uses an improved CatBoost model to achieve accurate prediction of the remaining charging duration of the battery. However, this method uses a single prediction model, and when dealing with a large amount of data, it is prone to result deviation or overfitting phenomena. In summary, it is urgent to consider using more advanced machine learning methods and prediction models that integrate neural networks to achieve high-precision prediction of the charging duration required by different electric vehicles under various complex charging conditions, reduce the result deviation of a single model, and have strong adaptability to meet the needs of complex and diverse charging scenarios. Summary of the Invention

[0004] Aiming at the problems of insufficient consideration of influencing factors, inability to meet the requirements of complex charging environments, and easy deviation of results in the existing charging duration prediction methods, the present invention provides an electric vehicle charging duration prediction method based on a multi-factor influence fusion model.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] An electric vehicle charging duration prediction method based on a multi-factor influence fusion model includes the following steps:

[0007] Step 1: Record the vehicle operation data collected by the new energy vehicle big data platform and unify the data format to obtain an initial data set;

[0008] Step 2: Cleaning and segmentation of the initial data set: Perform visual inspection on the initial data set, process the outliers, missing values, and duplicate values in the data set, and segment the initial data set into charge and discharge cycle segment data;

[0009] Step 3: Calculate the charging power in the charging cycle segment data, and divide the charging cycle segment data into a fast charging data set and a slow charging data set according to the power;

[0010] Step 4: Conduct correlation analysis on the fast charging data set and the slow charging data set, extract the key features for predicting the charging duration, and combine the extracted features and the charging duration to obtain a fast charging training set and a slow charging training set respectively;

[0011] Step 5: Perform model fusion based on the backpropagation algorithm (BP) inside the MLP model and the regression decision tree algorithm (CART) inside the RF model. Input the fast charging and slow charging training sets into the RF respectively, obtain the leaf node features according to the CART splitting rules, input the leaf node features and the deep features processed by the MLP hidden layer into the newly built "rule layer" of the MLP for fusion processing. Finally, according to whether the input training set is for fast charging or slow charging, the charging duration prediction result of the corresponding charging mode can be obtained at the output layer.

[0012] The present invention inputs the processed fast charging training set and slow charging training set into the RF model for training respectively, obtains the leaf node number index of each feature according to the splitting rules, converts the number index into a rule feature vector using one-hot encoding, and jointly inputs it into the "rule layer" together with the deep features obtained by the MLP hidden layer for further processing. By learning the complex feature relationships and the structural fusion of the MLP-RF model, accurate prediction of the fast and slow charging durations is achieved.

[0013] As a preferred embodiment of the present invention, the data collected by the new energy vehicle big data platform includes collection time, vehicle charge and discharge status, total voltage, total current, SOC, temperature, etc.; the data format is unified, including: converting the collection time into a time series format, and converting the total voltage, total current, and temperature into the smallest measurement unit to obtain the initial data set.

[0014] As a preferred embodiment of the present invention, the cleaning and segmentation of the initial data set specifically includes:

[0015] (1) Visualize the initial data set using a box plot according to the data valid range standard provided by the new energy big data platform. The data in the box plot are valid values, and the data outside the box plot are outside the valid range and are regarded as outliers and removed. The linear interpolation method is used to fill in the missing values, and by observing the time sampling points, the data with repeated time points are deleted;

[0016] (2) After processing the outliers, missing values, and duplicate values in the initial data set, use the moving average method to smooth and denoise the initial data set. The calculation formula is as follows:

[0017]

[0018] In the formula, MA t is the moving average value at time t, l is the window size, indicating the l nearest data points on the left used in the calculation, and x i is the i-th to t-th data points in the time series to be processed;

[0019] (3) After smoothing and denoising, classify the complex vehicle states in the initial data set. The parking charging and driving charging states are classified into the vehicle charging segment, and the non-charging state is classified into the vehicle discharging segment, and the initial data set is segmented into charging and discharging cycle segment data.

[0020] As a preferred embodiment of the present invention, the specific method for dividing the charging cycle segment data into a fast charging data set and a slow charging data set according to the charging power is as follows: Extract the voltage U and current I of each charging segment, calculate the power P according to the formula P = UI, and use the power as the basis for dividing the charging mode. By analyzing the power scatter plot, it is found that the charging power distribution is basically in the intervals [3, 10] kw and [30, 50] kw. Since the amount of data within [10, 30] kw is very small and cannot meet the data volume required for model training, the charging piles are divided into two modes: fast charging and slow charging with a charging power of 10 kw as the classification node, and the corresponding data sets are the fast charging data set and the slow charging data set.

[0021] As a preferred embodiment of the present invention, the steps of extracting key features, obtaining fast-charging training sets, and slow-charging training sets are as follows: Calculate the charging duration by subtracting the charging start time from the charging end time of each charging segment, and use the charging duration as the prediction target. Perform correlation analysis on a total of eight features, namely the initial state of charge (first SOC), final state of charge (final SCO), initial total voltage (Vt), charging current (I), initial maximum and minimum voltages of the single-cell battery (Vmax, Vmin), initial maximum temperature and minimum temperature of the single-cell battery (Tmax, Tmin), in the fast-charging data set and the slow-charging data set respectively. Extract six features with higher correlation with the charging duration, namely the initial SOC, final SOC, initial total voltage, initial minimum temperature, initial maximum and minimum voltages of the single-cell battery, as input features. Combine the extracted input features and the calculated charging duration to obtain fast-charging training data sets and slow-charging training data sets respectively, and divide the fast-charging and slow-charging training data sets according to the ratio of 7:2:1 for model training, testing, and verification. The expression of the training data set is as follows:

[0022]

[0023] In the formula, Train represents the training data set, the Time column represents the charging duration of each charging segment from charging segment 1 to charging segment N, and from first SOC to V min A total of six columns of data are included, and each column represents the extracted input features respectively.

[0024] As a preferred embodiment of the present invention, the core method used in the MLP model is the backpropagation algorithm (BP). The specific implementation steps of the algorithm are as follows:

[0025] (1) Propagate the extracted input features forward from the input layer to the hidden layer. In the hidden layer, the input features are calculated by summing the weights and biases, and are processed by the Sigmoid activation function for non-linear mapping of the features, and finally represented in the form of intermediate features and output from the hidden layer. The formula is as follows:

[0026]

[0027] In the formula, i is the i-th feature of the input layer, Y j is the output value of the hidden layer, m is the number of input features, w ij is the weight matrix from the input layer to the hidden layer, used to learn the relationship between the input features and the intermediate features, x i is the i-th input feature, θ jis the bias vector of the hidden layer, which is used to add an additional linear offset after the weighted sum, allowing the model to make a translational adjustment before activation and increasing the flexibility of the model. f(x) is the Sigmoid activation function used in the hidden layer, which introduces non-linearity and improves the model's ability to fit complex data. In f(x), x is the output value Y input to the hidden layer. j , a is a constant that affects the slope of the Sigmoid activation function, changes the learning rate and performance of the neural network. Too large an a will cause the gradient vanishing problem, and too small an a is prone to overfitting. Therefore, the selection of a is very important. After manual tuning, the selected a is 5;

[0028] (2) After the input features are output from the hidden layer in the form of intermediate features, they continue to be passed to the output layer. The output layer directly obtains the actual predicted output value using a linear activation. The formula is as follows:

[0029]

[0030] In the formula, C k is the predicted output value of the output layer, n is the number of intermediate features in the hidden layer, w jk is the weight matrix from the hidden layer to the output layer, which is used to learn the mapping relationship from intermediate features to the output. θ k is the bias vector of the output layer, which allows the model to make a linear translational adjustment after the weighted sum processing and is used to improve the accuracy of the prediction result;

[0031] (3) After obtaining the predicted output value, the mean squared error (MSE) loss function is calculated to measure the error between the predicted value and the true value. If the MSE converges to 0.05 or less, the training can be stopped in advance. Otherwise, it is necessary to enter the backpropagation process to adjust the weight matrix and bias to improve the prediction performance. The MSE calculation formula is as follows:

[0032]

[0033] In the formula, M is the total number of true value samples used, C k is the predicted output value of the output layer, y i is the true value;

[0034] (4) If the MSE does not converge to 0.05, it enters the backpropagation process. The system will transmit the mean squared error signal MSE in the backpropagation stage and use the gradient descent method to iteratively update the weights and biases of each layer to improve the prediction performance of the model. The gradient calculation formulas for the weights and biases of the output layer and the hidden layer are as follows:

[0035]

[0036] In the formula, The partial derivative of the mean squared error with respect to the weights from the output layer to the hidden layer, Y j The transpose of the output value of the hidden layer, The partial derivative of the mean squared error with respect to the bias from the output layer to the hidden layer, δ Y The error term of the hidden layer, w jk The transpose of the weight matrix from the hidden layer to the output layer, f′(x) is the derivative of the activation function of the hidden layer, The partial derivative of the mean squared error with respect to the weights from the hidden layer to the input layer, x i The transpose of the i-th feature of the input, The partial derivative of the mean squared error with respect to the bias from the hidden layer to the input layer, δ Y,i The error term corresponding to the i-th layer.

[0037] (5) By calculating the gradient, the weights and biases of each layer are continuously updated, gradually reducing the error between the predicted output and the true value. Until after several rounds of iteration, the MSE no longer decreases or even starts to increase, it is considered that the model reaches the best prediction state, and the backpropagation can be stopped. The output layer of the model in the best prediction state outputs the final prediction result. The calculation formulas for updating the weights and biases of each layer in backpropagation are as follows:

[0038]

[0039] In the formula, Δw ij Is the weight update value from the hidden layer to the input layer, Δθ j Is the bias update value from the hidden layer to the input layer, Δw jk Is the weight update value from the output layer to the hidden layer, Δθ k Is the bias update value from the output layer to the hidden layer, η represents the learning rate, which is used to control the step size of parameter update. When the learning rate is high, the model will be adjusted quickly. On the contrary, each step of learning is relatively small, so more iterations are required to make the model converge.

[0040] As a preferred embodiment of the present invention, the core method used in the RF model is the regression decision tree algorithm (CART). The specific implementation steps of the algorithm are as follows:

[0041] (1) Each decision tree performs bootstrap sampling from the training dataset, randomly extracting a Bootstrap sample, which has the same size as the training dataset and consists of input features and the training charging duration target. Each tree is trained with the samples randomly extracted from the training dataset. The expression of the extracted samples is as follows:

[0042] S t = Bootstrap(X,y)

[0043] In the formula, S trepresents the number of training data samples of the t-th tree, X is the feature matrix input to the random forest, and y is the true value of the charging duration target corresponding to each row of features;

[0044] (2) Each tree is recursively split using the splitting rules of a regression decision tree. First, the splitting point corresponding to the best feature needs to be found. After dividing the data into two subsets, by calculating the mean squared error (MSE) of the two subsets respectively, the region with a smaller MSE is selected for further splitting. For the selection of the best splitting point, all features and splitting points need to be traversed to select the best splitting point (j*, s*) that can minimize the total error. According to the best splitting point, the subsets are divided, and the search for feature splitting points and the division of subsets are continuously carried out until the minimum splitting samples or the depth of the tree in the hyperparameters are met, and then the splitting stops. The minimum splitting samples set is 4, and the depth of the tree is set to 10:

[0045]

[0046] In the formula, X i is the selected feature sample, y i is the true charging duration corresponding to the feature sample, X i,j is the j-th feature value of sample i, s represents the splitting point selected according to feature j, S left and S right are the numbers of training samples less than or equal to s and greater than s divided according to the splitting point s, called the left subset and the right subset respectively, S t is the total number of training samples of the current t-th tree, Total Error(j, s) is the total error, MSE(S left ) and MSE(S right ) are the mean squared errors of the left subset and the right subset respectively;

[0047] (3) For each time of starting to select different features and the corresponding charging duration targets, until the splitting ends, by calculating the MSE of the splitting regions respectively, the subset with a smaller MSE is defined as the leaf node region, and the mean value of the charging duration targets corresponding to all feature values in the leaf node region is the predicted output result of this regression tree. The calculation formula is as follows:

[0048]

[0049] In the formula, is the predicted result of the t-th regression tree, is the mean value of the charging duration targets corresponding to the remaining feature values in the leaf node region;

[0050] (4) After obtaining the predicted output results independently generated by each regression decision tree, the final random forest prediction value is the average of the predicted values of all regression trees. The calculation formula is as follows:

[0051]

[0052] In the formula, is the final prediction result of the random forest, T represents the number of trees in the random forest, is the prediction result of the t-th regression tree.

[0053] As a preferred embodiment of the present invention, the internal algorithm that combines the MLP and RF models is used to implement model fusion for predicting the charging duration required for the vehicle under fast charging and slow charging. The specific implementation steps are as follows:

[0054] (1) Input the fast charging or slow charging training data set into the RF model. The model internally uses bootstrap sampling to randomly extract data sets from the training data set for training, and repeatedly and randomly selects data sets for combined learning of the complex relationship information between the input features and the charging duration. An example expression of the extracted training data set is as follows:

[0055]

[0056] In the formula, Train1 represents the first training data set, and the number of samples taken is from 1 to 200. Train2 represents the second training data set, and the number of samples taken is from 301 to 500. The selection of training samples can be repeated, as shown in Train3. The RF model will traverse all feature combinations to find the complex relationship between the input features and the prediction results.

[0057] (2) The RF will divide the input training features into the corresponding leaf node regions based on the regression decision tree algorithm. The trained features will be numbered and indexed according to the leaf node regions they fall into after the splitting ends. Finally, all the training features will form a leaf node index matrix through numbering and indexing. The expressions of the leaf node numbering index and the index matrix are as follows:

[0058] L RF,i,t = x i,t

[0059] In the formula, L RF,i,t is the leaf node numbering index of the i-th training feature xi on the t-th tree;

[0060]

[0061] In the formula, L RF is the leaf node index matrix, n represents the total number of training features, each row of the matrix corresponds to a training feature, each column corresponds to a tree, and each element in the matrix is the leaf node numbering index where the sample falls after the splitting ends in the corresponding regression tree;

[0062] (3) Use one-hot encoding to transform the leaf node index matrix into a higher-dimensional and sparse feature representation, thereby forming a feature vector that implies rule information, which is used to capture the rule information of the feature partitioning in the random forest. When the leaf node number index is 3 and the total number of leaf nodes is 5, the rule feature vector after one-hot encoding will be converted into OneHot(L RF ) = [0, 0, 1, 0, 0], which is used as the feature vector that implies the random forest splitting rule for subsequent model prediction. The encoding formula is as follows:

[0063] Z RF = OneHot(L RF )

[0064] In the formula, Z RF is the feature vector matrix transformed from the leaf node encoding index of each sample;

[0065] (4) Through the leaf node number index and one-hot encoding, the rule information hidden in the RF is transformed into feature inputs for the MLP model. After the MLP hidden layer, a newly defined "rule layer" is added to receive the rule features from the RF and the hidden features obtained through weighted summation and non-linear mapping of the activation function in the MLP hidden layer, that is, the output value Y j of the hidden layer. The fused rule features and the output of the hidden layer are further processed through the newly established "rule layer". The fused features are linearly transformed and non-linearly mapped by the activation function in the "rule layer", and the fused features processed by the "rule layer" are input to the output layer to obtain the final prediction result. The processing of features inside the "rule layer" is similar to that of the hidden layer, but the "rule layer" uses different fused features and the ReLU activation function, introducing stronger non-linear processing of the fused feature relationship and avoiding the problem of gradient disappearance that is prone to occur in the Sigmoid function. The calculation formulas for feature fusion and the "rule layer" are as follows:

[0066]

[0067] In the formula, Z fusion represents the fused features, Combine represents feature combination, Z RF represents the rule features extracted from the random forest, G rule represents the "rule layer", f ReLU represents the ReLU activation function used by the rule layer, and the output range is [0, \infty), where infty is infinity, indicating that the predicted charging duration result can vary according to the time required for fast charging or slow charging, without limitation and the time cannot be negative. w rule represents the weight matrix of the "rule layer", and b rule represents the bias of the "rule layer".

[0068] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0069] (1) When the present invention performs feature extraction, it combines multiple parameters such as SOC change, temperature, and voltage, and divides them into fast charging and slow charging modes by calculating the charging power. Considering various influencing factors comprehensively, it ensures that the model can accurately predict the charging duration under actual complex conditions;

[0070] (2) The present invention transforms the RF rule features into rule feature vectors by using one-hot encoding, and performs non-linear processing by fusing them with the deep features obtained from the MLP hidden layer, and uses more complex deep features for model training;

[0071] (3) The present invention can accurately predict the charging duration required under different vehicle models and different charging modes, avoiding the situation that the model can only predict the charging duration under certain vehicle conditions. The prediction model has good generalization ability and strong applicability;

[0072] (4) By fusing the MLP-RF model, the present invention avoids the situation of deviation and overfitting that are prone to occur in a single model, and makes full use of the characteristics of the MLP model to capture complex non-linear relationships and the RF model to reduce the risk of overfitting through ensemble learning, realizing the accurate prediction of the charging duration of different vehicles by the fusion model;

[0073] (5) By using the BP neural network and the CART regression tree integration algorithm, and comparing with the prediction results of other model algorithms, it is found that when using the present invention for charging duration prediction, whether it is fast charging or slow charging, the average prediction error can be guaranteed to be within 360 seconds (6 minutes), which is 2 - 3 minutes less than the error predicted by other algorithms, and the error rates are only 4.325% and 3.862%, realizing the high-precision prediction of the charging duration of different vehicle models under different power charging piles. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0075] Figure 1 It is a schematic flowchart of the implementation method of the present invention;

[0076] Figure 2 It is a scatter plot of charging pile power classification;

[0077] Figure 3These are the implementation processes of two algorithms used in the present invention. Among them, part (a) shows the implementation process of the BP algorithm of the MLP model, and part (b) shows the implementation process of the CART regression tree of the RF model; part (c) shows the implementation process of model fusion prediction.

[0078] Figure 4 This is a schematic diagram for comparing the prediction results of the charging duration of different vehicles under different charging piles. Among them, part (a) shows the results of fast charging for vehicle type 1, and part (b) shows the results of slow charging for vehicle type 2.

[0079] Figure 5 This is a schematic diagram for comparing the prediction of the charging duration by the present invention and other different algorithms. Among them, part (a) is the comparison result based on fast charging piles, and part (b) is the comparison result based on slow charging piles. Detailed implementation manners

[0080] To better illustrate the purpose, technical solution, and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments.

[0081] Embodiment 1

[0082] A method for predicting the charging duration of electric vehicles based on a multi-factor influence fusion model can be applied to the field of electric vehicle charging research. The flow of predicting the charging duration of electric vehicles provided in this embodiment is as Figure 1 shown, and specifically includes the following:

[0083] Step 1: Record the data collected by the new energy vehicle big data platform, including collection time, vehicle status, total voltage, total current, SOC, temperature, etc.; the data format is unified as follows: convert the collection time into a time series format, and convert the total voltage, total current, and temperature into the smallest measurement unit to obtain the initial data set.

[0084] In this embodiment, the data of the real operation conditions of 5 different electric vehicles from January to March in a year collected by the platform is used, including driving speed, driving mileage, start-stop status, and voltage, current, and SOC change data, etc. Due to the large amount of data, only a part of the initial data set of vehicle type 2 is shown in Table 1. Among them, in the vehicle status column, "1" indicates that the vehicle is started, and "2" indicates that the vehicle is turned off; in the charging status column, "1" indicates charging while parked, "2" indicates charging while driving, "3" indicates not charging, and if there are "FE" or "FF" in the data, it means the data is abnormal or invalid.

[0085] Table 1 Part of the initial data set extracted for vehicle type 2

[0086]

[0087] Step 2: Screen and remove the data in the initial dataset that is outside the reasonable range, and perform corresponding smoothing processing. The cleaning and segmentation of the initial dataset specifically include:

[0088] (1) According to the data valid range standards provided by the new energy big data platform, the valid ranges of SOC, total voltage, total current, single-cell battery voltage, and single-cell battery temperature are [0, 100], [50V, 500V], [-1000A, 1000A], [0V, 15V], and [-40°C, 210°C] respectively. Visualize the initial dataset using a box plot according to the valid range standards. The data in the box plot are valid values, and the data outside the box are outside the valid range and are regarded as outliers and removed. The linear interpolation method is used to fill in the missing values, and by observing the time sampling points, the data with repeated time points are deleted;

[0089] (2) After processing the outliers, missing values, and duplicate values in the initial dataset, use the moving average method to smooth and denoise the initial dataset. The calculation formula is as follows:

[0090]

[0091] In the formula, MA t is the moving average value at time t, l is the window size, indicating the l nearest data points on the left used in the calculation, x i is the i-th to t-th data points in the time series to be processed. For some of the total voltage data when I need to process: {360 362.1 363.5 364 365.3 450 367 368 368.5 370}, it can be clearly seen that 450 is a fluctuating value. Select a window size of 5 and use the moving average method to process this value, and then the short-term data fluctuations are smoothed out.

[0092] (3) After smoothing and denoising, classify the complex vehicle states in the initial dataset. The parking charging and driving charging states are classified into the vehicle charging segment, and the non-charging state is classified into the vehicle discharging segment. The initial dataset is segmented into charging and discharging cycle segment data. Table 2 shows the results of cleaning a part of the initial dataset extracted from vehicle model 2.

[0093] Table 2 Initial dataset of vehicle model 2 after cleaning

[0094]

[0095]

[0096] Step 3: Divide the charging cycle segment data into fast charging data set and slow charging data set according to the charging power. The method is as follows: Extract the voltage U and current I of each charging segment, calculate the power P according to the formula P = UI, and use the power as the basis for dividing the charging mode. For example, Figure 2 It is found from the analysis of the power scatter plot that the charging power distribution is basically between the intervals [3, 10] kW and [30, 50] kW. Since the amount of data within [10, 30] kW is very small and cannot meet the data volume required for model training, the charging piles are divided into two modes: fast charging and slow charging with 10 kW charging power as the classification node. The corresponding data sets are the fast charging data set and the slow charging data set. Table 3 shows the power calculated by multiplying the initial total voltage and the initial total current. From the power calculated in the table, it can be seen that charging segments 1-3 are all slow charging, and charging segments 4-5 are fast charging.

[0097] Table 3 Division of fast and slow charging according to charging power calculation

[0098]

[0099]

[0100] Step 4: The steps to extract key features and obtain the fast charging training set and the slow charging training set are as follows: Calculate the charging duration by subtracting the charging start time from the charging end time of each charging segment, and use the charging duration as the prediction target. Perform correlation analysis on the initial SOC (first SOC), end SOC (final SCO), initial total voltage (Vt), charging current (I), initial maximum and minimum voltages of single cells (Vmax, Vmin), initial maximum temperature and minimum temperature of single cells (Tmax, Tmin) in the fast charging data set and the slow charging data set respectively. Extract six features with higher correlation with the charging duration among the initial SOC, end SOC, initial total voltage, initial minimum temperature, initial maximum and minimum voltages of single cells as input features. Combine the extracted input features and the calculated charging duration to obtain the fast charging training data set and the slow charging training data set respectively, and divide the fast charging and slow charging training data sets according to the ratio of 7:2:1 for model training, testing and verification. The expression of the training data set is as follows:

[0101]

[0102] In the formula, Train represents the training data set, and the Time column represents the charging duration of each charging segment from charging segment 1 to charging segment N. From first SOC to V min A total of six columns of data are included. Each column represents the extracted input features respectively. Train* shows the slow charging training data set using the actual extracted features and the corresponding charging duration combination.

[0103] Step 5: The core method used by the MLP model is the backpropagation algorithm (BP). As shown in part (a) of Figure 3 , the specific implementation steps of the algorithm are as follows:

[0104] (1) Propagate the extracted input features forward from the input layer to the hidden layer. In the hidden layer, the input features are summed through weights and biases, and non-linearly mapped through the Sigmoid activation function. Finally, they are represented in the form of intermediate features and output from the hidden layer. The formula is as follows:

[0105]

[0106] In the formula, i is the i-th feature of the input layer, Y j is the output value of the hidden layer, m is the number of input features, w ij is the weight matrix from the input layer to the hidden layer, used to learn the relationship between input features and intermediate features, x i is the i-th input feature, θ j is the bias vector of the hidden layer, used to add an additional linear offset after the weighted sum, allowing the model to perform a translation adjustment before activation and increasing the flexibility of the model. f(x) is the Sigmoid activation function used in the hidden layer, introducing non-linearity and improving the model's fitting ability for complex data. In f(x), x is the output value Y j input to the hidden layer. a is a constant that affects the slope of the Sigmoid activation function, changes the learning rate and performance of the neural network. Too large an a will cause the gradient vanishing problem, and too small an a is prone to overfitting. Therefore, the selection of a is very important. After manual parameter tuning, the selected a is 5;

[0107] Specifically, for the training feature matrix input to the present invention and the target value when Feature is input into the input layer of the MLP and processed by the hidden layer, the MLP will give the initial hidden layer weights and biases according to the size of the input feature matrix. θ j = [θ 1 θ 2 θ 3 θ 4 θ 5 θ 6 θ 7 . The hidden layer will perform weighted summation according to the weights and biases and perform non-linear processing according to the Sigmoid activation function. The calculation process is as follows:

[0108]

[0109] (2) The input features are processed by the hidden layer and output a hidden layer output matrix of size 700×700, which is then passed as input to the output layer. The output layer directly obtains the actual predicted output value using linear activation. The formula is as follows:

[0110]

[0111] In the formula, C k is the predicted output value of the output layer, n is the number of intermediate features in the hidden layer, w jk is the weight matrix from the hidden layer to the output layer, which is used to learn the mapping relationship from intermediate features to the output. θ k is the bias vector of the output layer, which allows the model to linearly translate and adjust the prediction result after weighted sum processing, and is used to improve the accuracy of the prediction result;

[0112] Specifically, the output layer continues to provide the weights and biases of this layer θ k =[θ 8 , and the output layer obtains the predicted output value through linear activation:

[0113] (3) After obtaining the predicted output value, the mean squared error (MSE) loss function is calculated to measure the error between the predicted value and the true value. If the MSE converges to 0.05 or less, training can be stopped early. Otherwise, backpropagation is required to adjust the weight matrix and bias to improve the prediction performance. The MSE calculation formula is as follows:

[0114]

[0115] In the formula, M is the total number of true value samples used, C k is the predicted output value of the output layer, and y i is the true value;

[0116] Calculate the mean squared error through the formula

[0117] (4) If the MSE does not converge to 0.05, the system enters the backpropagation process. The system will transmit the mean squared error signal MSE during the backpropagation stage and use the gradient descent method to iteratively update the weights and biases of each layer to improve the prediction performance of the model. The gradient calculation formulas for the weights and biases of the output layer and the hidden layer are as follows:

[0118]

[0119] In the formula, is the partial derivative of the mean squared error with respect to the weights from the output layer to the hidden layer, Y j is the transpose of the output value of the hidden layer, The partial derivative of the mean squared error with respect to the bias from the output layer to the hidden layer, δ Y is the error term of the hidden layer, w jk is the transpose of the weight matrix from the hidden layer to the output layer, f′(x) is the derivative of the activation function of the hidden layer, The partial derivative of the mean squared error with respect to the weight from the hidden layer to the input layer, x i is the transpose of the i-th feature of the input, The partial derivative of the mean squared error with respect to the bias from the hidden layer to the input layer, δ Y,i is the error term corresponding to the i-th layer.

[0120] (5) By calculating the gradient, the weights and biases of each layer are continuously updated, gradually reducing the error between the predicted output and the true value. Until after several rounds of iteration, the MSE no longer decreases or even starts to rise, it is considered that the model reaches the optimal prediction state, and the backpropagation can be stopped. The output layer of the model in the optimal prediction state outputs the final prediction result. The calculation formulas for updating the weights and biases of each layer in backpropagation are as follows:

[0121]

[0122] In the formula, Δw ij is the weight update value from the hidden layer to the input layer, Δθ j is the bias update value from the hidden layer to the input layer, Δw jk is the weight update value from the output layer to the hidden layer, Δθ k is the bias update value from the output layer to the hidden layer. η represents the learning rate, which is used to control the step size of parameter update. When the learning rate is high, the model will be adjusted quickly. On the contrary, each step of learning is relatively small, so more iterations are required to make the model converge.

[0123] Step 6: The core method used by the RF model is the regression decision tree algorithm (CART), as shown in part (b) of Figure 3 The specific implementation steps of the algorithm are as follows:

[0124] (1) Each decision tree performs bootstrap sampling from the training dataset, randomly extracting a Bootstrap sample, which has the same size as the training dataset and is composed of input features and the training charging duration target. Each tree is trained using the samples randomly extracted from the training dataset. The expression for the extracted sample is as follows:

[0125] S t = Bootstrap(X,y)

[0126] In the formula, S t represents the number of training data samples for the t-th tree, X is the feature matrix input to the random forest, and y is the true value of the charging duration target corresponding to each row of features;

[0127] Specifically, the extracted samples and targets are X and

[0128]

[0129] (2) Each tree is recursively split using the splitting rules of a regression decision tree. First, the splitting point corresponding to the best feature needs to be found. After dividing the data into two subsets, by calculating the mean squared error (MSE) of the two subsets respectively, the region with a smaller MSE is selected for further splitting. For the selection of the best splitting point, all features and splitting points need to be traversed to select the best splitting point (j*, s*) that can minimize the total error. According to the best splitting point, the subsets are divided, and the search for feature splitting points and the division of subsets are continuously carried out until the minimum splitting samples or the depth of the tree in the hyperparameters are met and the splitting stops. The minimum splitting samples set is 4, and the depth of the tree is set to 10:

[0130]

[0131] In the formula, X i is the selected feature sample, y i is the true charging duration corresponding to the feature sample, X i,j is the j-th feature value of sample i, s represents the splitting point selected according to feature j, S left and S right are the number of training samples less than or equal to s and greater than s divided according to the splitting point s, called the left subset and the right subset respectively, S t is the total number of training samples of the current t-th tree, Total Error(j, s) is the total error, MSE(S left ) and MSE(S right ) are the mean squared errors of the left subset and the right subset respectively;

[0132] Specifically, the data in the left subset and right subset regions after dividing according to the splitting point s are Combined with the MSE(S t ) formula to calculate the total errors Total Error(S left ) and Total Error(S right ), S t is 700, S left is 400, S right is 300. If Total Error(S left ) is less than Total Error(S right ), then continue to split in S left , otherwise split in S right until the stopping condition is met.

[0133] (3) For each time of starting to select different features and the corresponding charging duration targets until the splitting ends, by calculating the MSE of the splitting regions respectively, the subset with a smaller MSE is defined as the leaf node region, and the mean value of the charging duration targets corresponding to all feature values in the leaf node region is the predicted output result of this regression tree. The calculation formula is as follows:

[0134]

[0135] In the formula, is the predicted result of the t-th regression tree, is the mean value of the charging duration targets corresponding to the remaining feature values in the leaf node region;

[0136] Specifically, after the splitting stops, the remaining leaf node data is as follows: Its corresponding target value is Then the predicted output value of this regression decision tree is

[0137] (4) After obtaining the predicted output results independently generated by each regression decision tree, the final random forest predicted value is the average value of the predicted values of all regression trees. The calculation formula is as follows:

[0138]

[0139] In the formula, is the final predicted result of the random forest, T represents the number of trees in the random forest, is the predicted result of the t-th regression tree.

[0140] Step 7: Combine the internal algorithms used by the MLP and RF models to achieve model fusion for predicting the charging duration results required for the vehicle under fast charging and slow charging, as shown in part (c) of Figure 3 . The specific implementation steps are as follows:

[0141] (1) Input the fast charging or slow charging training dataset into the RF model. The model internally uses bootstrap sampling to randomly extract a dataset from the training dataset for training, and repeatedly and randomly selects the dataset for combined learning of the complex relationship information between the input features and the charging duration. An example expression of the extracted training dataset is as follows:

[0142]

[0143]

[0144] In the formula, Train1 represents the first training data set, and the number of samples taken ranges from 1 to 200. Train2 represents the second training data set, and the number of samples taken ranges from 301 to 500. The selection of training samples can be repeated, as shown in Train3. The RF model will traverse all feature combinations to find the complex relationship between the input features and the prediction results.

[0145] (2) RF will divide the input training features into corresponding leaf node regions based on the regression decision tree algorithm. The training features will be numbered and indexed according to the leaf node regions they fall into after the splitting is completed. Finally, all the training features will form a leaf node index matrix through numbered indexing. The expressions for the leaf node number indexing and the index matrix are as follows:

[0146] L RF,i,t =x i,t

[0147] In the formula, L RF,i,t is the leaf node number index of the i-th training feature xi on the t-th tree;

[0148]

[0149] In the formula, L RF is the leaf node index matrix. n represents the total number of training features. Each row of the matrix corresponds to a training feature, and each column corresponds to a tree. Each element in the matrix is the leaf node number index where the sample falls after the splitting of the corresponding regression tree ends;

[0150] (3) Using one-hot encoding, the leaf node index matrix is transformed into a higher-dimensional and sparse feature representation, thereby forming a feature vector containing rule information, which is used to capture the rule information of the random forest's feature partitioning. As a feature vector implicitly representing the random forest splitting rules, it is used for subsequent model prediction. The encoding formula is as follows:

[0151] Z RF =OneHot(L RF )

[0152] In the formula, Z RF is the feature vector matrix formed by transforming the leaf node encoding index of each sample;

[0153] Based on the above-mentioned numbered indexing and one-hot encoding implementation steps, an example is given below to simulate the numbering and encoding processes, as follows:

[0154] (3.1) Numbering process: Assume the input feature is The target value is Suppose the random forest consists of 2 decision trees, i.e., T = 2. Tree 1 and Tree 2 correspond to the first column and the second column of feature x respectively. Tree 1 numbers the leaf nodes according to the following partitioning rule: if x ≤ 2, enter leaf node 1; otherwise, enter leaf node 2. Then the leaf node numbers are Tree 2 numbers the leaf nodes according to the following partitioning rule: if x > 3, enter leaf node 1; otherwise, enter leaf node 2. Then the leaf node numbers are The index matrix formed by combining the leaf node numbers of all trees is

[0155] (3.2) Encoding process: The decision tree has two leaf nodes. When performing one-hot encoding, there will be two columns of data. According to the index matrix L RF It can be known that the leaf node numbers of sample 1 are [1 2]. The leaf node number in Tree 1 is 1, which becomes [1 0] after one-hot encoding. The leaf node number in Tree 2 is 2, which becomes [0 1] after one-hot encoding. After splicing, it becomes [1 0 0 1]. This is the regular feature vector obtained after one-hot encoding of the leaf node encoding of sample 1. Similarly, the leaf node numbers of sample 2 are [2 1], and the regular feature vector obtained after one-hot encoding is [0 1 1 0]. The regular feature vector matrix obtained after splicing the feature vectors is

[0156] (4) Through the leaf node number index and one-hot encoding, the implicit rule information in the RF is converted into features and input into the MLP model. After the MLP hidden layer, a newly defined "rule layer" is added to receive the rule features from the RF, as well as the hidden features obtained through weighted summation and non-linear mapping of the activation function in the MLP hidden layer, that is, the output value Y of the hidden layer j , and the fused rule features and the output of the hidden layer are further processed through the newly established "rule layer". The fused features are linearly transformed and non-linearly mapped by the activation function in the "rule layer", and the fused features processed by the "rule layer" are input into the output layer to obtain the final prediction result. The calculation formulas for feature fusion and the "rule layer" are as follows:

[0157]

[0158] In the formula, Z fusion represents the fused feature, Combine represents feature combination, Z RF represents the rule feature extracted from the random forest, G rule represents the "rule layer", f ReLU represents the ReLU activation function used by the rule layer, and the output range is [0, ∞), where ∞ is infinity, indicating that the predicted charging duration result can vary according to the time required for fast charging or slow charging, is unrestricted and the time cannot be negative, w ruleThe weight matrix representing the "rule layer", b rule represents the bias of the "rule layer".

[0159] Specifically, from step 6, it can be seen that the Y obtained by the MLP hidden layer j is a 700×700 depth feature matrix. After one-hot encoding by the RF model After feature fusion, in step 5, the implementation process of calculating the weights and biases of each layer of the MLP has been specifically described, and this step will not be described again.

[0160] The implementation mode of the "rule layer" is similar to that of the hidden layer. According to the specific weight matrix w of the "rule layer" rule and bias b rule , combined with the ReLU activation function to process the fused features. For the features with negative values in the output after processing, the ReLU activation function will output the negative values as 0 to avoid the influence of negative value output on the subsequent prediction of the charging duration result. After the "rule layer" processes the feature output, the charging duration prediction result is obtained at the output layer of the MLP.

[0161] Based on the above MLP-RF fusion model to predict the charging duration, the charging duration results of different vehicle models using different charging modes for charging prediction obtained in this embodiment are as Figure 4 shown. From Figure 4 (a) and (b) in it, it can be seen that the charging duration prediction method proposed by the present invention can achieve accurate prediction for different vehicle models and different charging modes. And through the comparison calculation of the predicted value and the true value, it can be known that whether it is the slow charging mode with a long charging time or the fast charging mode with a short charging time, the prediction results can all perform well. Among them, according to the prediction result data in Table 4, calculate the error compared with the true value. Error = charging real time - charging prediction time, such as 24440 - 24071.387 = 368.613, error rate = error / charging real time, such as 368.613 / 24440 * 100% = 1.508%. Using the above calculation method, traverse the prediction results of the entire fast and slow charging training set. Through calculation, it can be known that the average error of the fast and slow charging prediction durations can be controlled within 420 seconds (7 minutes), and the average error rate can be controlled within 5%. The overall prediction effect is better under different vehicle models and different charging modes.

[0162] Table 4 Comparison of slow charging prediction results

[0163]

[0164]

[0165] Figure 5Shown is the comparison of the prediction results of the MLP-RF fusion prediction model used in the present invention and two other models, including RF and BiLSTM-RF. From Figure 5 It is not difficult to see from part (a) in Figure 5 that the BiLSTM-RF model has the worst prediction effect and fails to follow the change trend of the true value well. In the 300-400th training samples, the prediction effect of RF is significantly inferior to that of MLP-RF; from

[0166] part (b) in

[0167] it can be seen that the RF model has the worst prediction effect, and the other two models perform well. However, the BiLSTM-RF model shows large fluctuations in the predicted values within the range of the 200th and 500th samples, proving that the prediction performance of this model is not stable enough.

[0166] In summary, the prediction model used in the present invention has stronger generalization ability than other models. While ensuring the prediction accuracy of the charging duration, it can also be applied to complex and changeable charging environment changes.

[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for predicting the charging time of electric vehicles based on a multi-factor fusion model, characterized in that: The following steps are involved: Step 1: Record the vehicle operation data collected by the new energy vehicle big data platform and unify the data format to obtain the initial data set; Step 2: Cleaning and segmenting the initial data set: Visually inspect the initial data set, process the outliers, missing values, and duplicate values, and segment the initial data set into charge and discharge cycle fragment data; Step 3: Calculate the charging power in the charging cycle segment data, and divide the charging cycle segment data into a fast charging data set and a slow charging data set according to the charging power; Step 4: Perform correlation analysis on the fast charging data set and the slow charging data set to extract key features for predicting charging time, and combine the extracted key features with the charging time to obtain the fast charging training set and the slow charging training set respectively; Step 5: Based on the back propagation algorithm (BP) in the MLP model and the regression decision tree algorithm (CART) in the RF model, the model is fused. The fast charging training set and the slow charging training set are input into RF respectively. The leaf node features are obtained according to the CART splitting rule. The leaf node features and the deep features processed by the MLP hidden layer are input into the newly created "rule layer" of the MLP for fusion processing. Finally, according to whether the input training set is fast charging or slow charging, the charging time prediction result of the corresponding charging mode is obtained in the output layer.

2. The method for predicting the charging time of an electric vehicle based on a multi-factor fusion model according to claim 1 is characterized in that: The data format unification in step 1 includes: converting the time data into a time series format; converting the voltage, current and temperature data into the minimum measurement unit.

3. The method for predicting the charging time of an electric vehicle based on a multi-factor fusion model according to claim 1 is characterized in that: The initial dataset cleaning and segmentation method described in step 2 is as follows: (1) According to the data valid range standard provided by the new energy vehicle big data platform, the initial data set is visualized using a box plot. The data in the box plot is a valid value, and the data outside the box plot is outside the valid range. It is considered as an outlier and removed, and the missing values ​​are filled by linear interpolation. By observing the time sampling points, the data recorded at duplicate time points are deleted; (2) After processing the outliers, missing values, and duplicate values ​​in the initial data set, the moving average method is used to smooth and reduce noise on the initial data set; (3) The initial data set after smoothing and denoising is divided into charging and discharging cycle segment data.

4. The method for predicting the charging time of an electric vehicle based on a multi-factor fusion model according to claim 1 is characterized in that: The method described in step 3 for dividing the charging cycle segment data into a fast charging data set and a slow charging data set according to the charging power is as follows: extract the voltage and current data of each charging segment, and calculate the power as the basis for dividing the charging piles, and draw the calculated power into a scatter plot for analysis. According to the power data distribution, the charging piles are divided into fast charging and slow charging modes with 10kw charging power as the classification node. The corresponding data sets are the fast charging data set and the slow charging data set.

5. The method for predicting the charging time of an electric vehicle based on a multi-factor fusion model according to claim 1 is characterized in that: The steps of extracting key features and obtaining fast charging training sets and slow charging training sets described in step 4 are as follows: calculate the charging time of each charging segment as the prediction target, and perform correlation analysis on the eight features of initial SOC, end SOC, initial total voltage, charging current, initial maximum and minimum voltage of the single cell, initial maximum temperature, and initial minimum temperature in the fast charging data set and the slow charging data set, respectively, and extract the six features of initial SOC, end SOC, initial total voltage, initial minimum temperature, initial maximum and minimum voltage of the single cell that are more correlated with the charging time as input features, combine the extracted input features with the calculated charging time, and obtain the fast charging training set and the slow charging training set, respectively, and divide the fast charging and slow charging training data sets into a ratio of 7:2:1 for model training, testing, and verification.

6. The method for predicting the charging time of an electric vehicle based on a multi-factor fusion model according to claim 1 is characterized in that: The core method used by the MLP model is the back propagation algorithm (BP), and the specific implementation steps of the algorithm are as follows: (1) The extracted input features are forward propagated from the input layer to the hidden layer. In the hidden layer, the input features are calculated and transmitted through weights and biases, represented in the form of intermediate features, and output from the hidden layer. The formula is as follows: In the formula, i is the i-th feature of the input layer, Y j is the output value of the hidden layer, m is the number of input features, and w ij is the weight matrix from the input layer to the hidden layer, which is used to learn the relationship between the input features and the intermediate features, x i is the i-th feature of the input, θ j is the bias vector of the hidden layer, allowing the model to be translated and adjusted before the activation function. f(x) is the Sigmoid activation function used by the hidden layer, which is used to introduce nonlinear features and improve the model's ability to fit complex data. In f(x), x is the output value Y of the input hidden layer. j , a is a constant that affects the slope of the Sigmoid activation function and changes the learning rate and performance of the neural network. Too large a will lead to the gradient vanishing problem, and too small a will easily lead to overfitting. Therefore, the selection of a is very important. After manual parameter adjustment, the selected a is 5; (2) After the input features are output from the hidden layer in the form of intermediate features, they are passed to the output layer. The output layer uses linear activation to directly obtain the actual predicted output value. The formula is as follows: In the formula, C k is the predicted output value of the output layer, n is the number of intermediate features of the hidden layer, and w jk is the weight matrix from the hidden layer to the output layer, which is used to learn the mapping relationship from the intermediate features to the output, θ k is the bias vector of the output layer, allowing the model to introduce translation adjustments during prediction; (3) After obtaining the predicted output value, the mean square error (MSE) loss function is calculated to measure the error between the predicted value and the true value. If the MSE converges to 0.05 or below, the training can be stopped early. Otherwise, it is necessary to enter back propagation and adjust the weight matrix and bias to improve the prediction performance. The MSE calculation formula is as follows: Where M is the total number of true value samples used, C k is the predicted output value of the output layer, y i is the true value; (4) If the MSE does not converge to 0.05, it will enter the back propagation phase. The system will transmit the mean square error signal MSE in the back propagation phase and iteratively update the weights and biases of each layer using the gradient descent method. The gradient calculation formulas for the output layer, hidden layer weights and biases are as follows: In the formula, The partial derivative of the mean square error with respect to the weights from the output layer to the hidden layer, Y j Transpose the output value of the hidden layer, The partial derivative of the mean square error with respect to the bias from the output layer to the hidden layer, δ Y is the error term of the hidden layer, w jk is the transpose of the weight matrix from the hidden layer to the output layer, f′(x) is the derivative of the hidden layer activation function, The partial derivative of the mean square error with respect to the weights from the hidden layer to the input layer, x i is the transpose of the i-th feature of the input, The partial derivative of the mean square error with respect to the bias from the hidden layer to the input layer, δ Y,i is the error term corresponding to the i-th layer. (5) The gradient is calculated to continuously update the weights and biases of each layer, gradually reducing the error between the predicted output and the true value. After several rounds of iterations, the MSE no longer decreases, or even starts to increase. The model is considered to have reached the best prediction state, and back propagation can be stopped. The output layer of the model in the best prediction state outputs the final prediction result. The calculation formula for updating the weights and biases of each layer of back propagation is as follows: In the formula, Δw ij is the weight update value from the hidden layer to the input layer, Δθ j is the bias update value from the hidden layer to the input layer, Δw jk is the weight update value from the output layer to the hidden layer, Δθ k is the bias update value from the output layer to the hidden layer, η represents the learning rate, which is used to control the step size of parameter update. When the learning rate is high, the model will adjust quickly. On the contrary, each learning step is relatively small, so more iterations are required to make the model converge.

7. The method for predicting the charging time of an electric vehicle based on a multi-factor fusion model according to claim 1 is characterized in that: The core method used by the RF model is the regression decision tree algorithm (CART), and the specific implementation steps of the algorithm are as follows: (1) Each decision tree performs bootstrap sampling from the training data set and randomly extracts a Bootstrap sample. The sample has the same size as the training data set and is composed of input features and the training charging time target. Each tree is trained by randomly extracting samples from the training data set. The extracted sample expression is as follows: S t =Bootstrap(X,y) In the formula, S t represents the number of training data samples of the tth tree, X is the feature matrix of the random forest input, and y is the true value of the charging time target corresponding to each row of features; (2) Each tree uses the splitting rule of the regression decision tree for recursive splitting. First, it is necessary to find the splitting point corresponding to the best feature. After dividing the data into two subsets, the mean square error (MSE) of the two subsets is calculated respectively, and the area with the smallest MSE is selected to continue splitting. For the selection of the best splitting point, it is necessary to traverse all features and splitting points, select the best splitting point (j*, s*) that can minimize the total error, and divide the subsets according to the best splitting point. The search for feature splitting points and the division of subsets are continuously performed until the minimum splitting sample or the depth of the tree in the hyperparameter is met. The minimum splitting sample is set to 4, and the depth of the tree is set to 10: Where, X i is the selected feature sample, y i is the actual charging time corresponding to the feature sample, X i,j is the jth eigenvalue of sample i, s represents the split point selected according to feature j, S left and S right The number of training samples less than or equal to s and greater than s divided according to the split point s are called the left subset and the right subset, respectively. t is the total number of training samples of the current t-th tree, TotalError(j,s) is the total error, MSE(S left ) and MSE(S right ) are the mean square errors of the left subset and the right subset respectively; (3) For each time different features and corresponding charging time targets are selected, until the end of the split, the MSE of the split regions is calculated respectively. The subset with smaller MSE is defined as the leaf node region. The mean of the charging time targets corresponding to all feature values ​​in the leaf node region is the predicted output result of the regression tree. The calculation formula is as follows: In the formula, is the prediction result of the tth regression tree, is the mean of the charging time targets corresponding to the remaining eigenvalues ​​in the leaf node area; (4) After obtaining the prediction output results generated independently by each regression decision tree, the final random forest prediction value is the average of all regression tree prediction values, calculated as follows: In the formula, is the final prediction result of the random forest, T represents the number of trees in the random forest, is the prediction result of the tth regression tree.

8. The method for predicting the charging time of an electric vehicle based on a multi-factor fusion model according to claim 1 is characterized in that: The internal algorithm used in the combination of MLP and RF models implements model fusion to predict the charging time required for the vehicle under fast charging and slow charging. The specific implementation steps are as follows: (1) The fast charging or slow charging training data set is input into the RF model. The model will repeatedly randomly extract data from the training data set and combine them, traversing the entire input training data set to learn the complex relationship between input features and charging duration; (2) RF will map the input training features to the corresponding leaf node regions based on the regression decision tree algorithm. The training features will be numbered and indexed according to the leaf node regions they fall into after the split. Finally, all the training features will form a leaf node index matrix through number indexing. The expressions of leaf node number index and index matrix are as follows: L RF,i,t =x i,t Where, L RF,i,t The leaf node number index of the i-th training feature xi in the t-th tree; Where, L RF is the leaf node index matrix, n represents the total number of training features, each row of the matrix corresponds to a training feature, each column corresponds to a tree, and each element in the matrix is ​​the leaf node number index where the sample falls after the split in the corresponding regression tree; (3) Using one-hot encoding, the leaf node index matrix is ​​converted into a higher-dimensional, sparse feature representation, thereby forming a feature vector that contains implicit rule information, which is used to capture the rule information of random forest feature partitioning. When the leaf node number index is 3 and the total number of leaf nodes is 5, the rule feature vector after one-hot encoding will be converted into OneHot (L RF )=[0,0,1,0,0], as the feature vector of the implicit random forest split rule for subsequent model prediction, the encoding formula is as follows: From RF =OneHot(L RF ) In the formula, Z RF Transform the leaf node encoding index of each sample into a feature vector matrix; (4) Through leaf node number indexing and one-hot encoding, the implicit rule information in RF is converted into features for input into the MLP model. After the MLP hidden layer, a newly defined "rule layer" is added to receive the rule features from RF and the hidden features obtained in the MLP hidden layer through weighted summation and nonlinear mapping of the activation function, that is, the output value Y of the hidden layer j , the fused rule features and the output of the hidden layer are further processed through the newly established "rule layer". The fused features are linearly transformed and nonlinearly mapped with the activation function in the "rule layer", and the fused features processed by the "rule layer" are input into the output layer to obtain the final prediction results in the output layer. The calculation formulas for feature fusion and the "rule layer" are as follows: In the formula, Z fusion Indicates fusion features, Combine indicates feature combination, Z RF represents the regular features extracted from the random forest, G rule stands for "rule layer", f ReLU represents the ReLU activation function used in the rule layer. The output range is [0,\inf ty), where infty is infinite, indicating that the predicted charging time can vary with the time required for fast charging or slow charging, without restriction and the time cannot be negative. w rule represents the weight matrix of the "rule layer", b rule Indicates the bias of the "rule layer".

Citation Information

Patent Citations

  • Method, device and equipment for predicting remaining charging time of vehicle and storage medium

    CN118342996A

  • New energy automobile residual charging duration prediction method and system and electronic device

    CN119026717A

Cited By

  • Charging station recommendation method and recommendation system based on comprehensive cost

    CN121258127A