Method for predicting charging behavior of electric vehicle driven by multi-mode large model
Through multimodal data fusion and pre-trained large models, the charging behavior prediction method of single data and insufficient real-time response in the existing technology is solved, and high-precision and real-time electric vehicle charging behavior prediction is achieved, the model adaptability and system scalability are improved, and the electric vehicle charging management and power grid scheduling are optimized.
Patent Information
- Application Number
- CN202510397533.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
In the existing charging behavior prediction methods for electric vehicles, the data mode is single, the model generalization ability is poor, the lack of real-time response ability, and the cross-modal feature fusion are insufficient, resulting in insufficient prediction accuracy and insufficient adaptability.
Multimodal data fusion technology is adopted, pre-trained large models such as Transformer, ViT, BERT and GNN are used to extract features, and charging behavior prediction is performed through cross-modal attention mechanism and fully connected neural network, combining real-time data flow and dynamic scene perception mechanism to achieve high-precision and real-time prediction.
It improves the comprehensiveness, accuracy, generalization ability and adaptability of charging behavior prediction, realizes real-time response and system scalability to complex scenarios, and optimizes charging management and grid scheduling.
Smart Images

Figure FT_1 
Figure SMS_10 
Figure SMS_23
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of smart grid and electric vehicle charging management, and particularly relates to a method for predicting electric vehicle charging behavior driven by a multimodal large model. This method realizes high-precision, real-time, and scenario-aware prediction of electric vehicle charging behavior by fusing multimodal data and using a pre-trained large model for feature extraction and behavior prediction, and is applicable to the management optimization of electric vehicle charging stations, grid dispatching optimization, and the improvement of user charging experience. Background Art
[0002] With the transformation of the global energy structure and the enhancement of environmental awareness, the popularization speed of electric vehicles (EVs) has increased significantly. The rapid growth of electric vehicles has brought new challenges to the power grid, especially in terms of charging infrastructure and power load management. The charging behavior of electric vehicles is highly random and diverse, affected by various factors such as users' charging habits, geographical locations, weather conditions, traffic flow, and the distribution of charging stations. Effectively predicting the charging behavior of electric vehicles is of great significance for the balance of the smart grid, the rational layout of charging facilities, and the optimization of user charging experience.
[0003] Traditional charging behavior prediction methods mainly rely on single-modal data for modeling, such as charging load data based on time series or user historical behavior data. These methods may have certain prediction effects in some specific scenarios, but due to the failure to fully consider the complex relationships between multimodal data, they often show the following deficiencies in practical applications:
[0004] 1. Single data modality and incomplete information: Traditional methods mainly rely on time series data, ignoring the potential value of multimodal data such as geographical location, environmental factors, user behavior characteristics, and images and videos, resulting in insufficient comprehensiveness and accuracy of prediction results.
[0005] 2. Poor model generalization ability and insufficient adaptability: Existing models are mostly customized for specific regions or specific user groups, and it is difficult to maintain high prediction accuracy in different geographical regions or changing environmental conditions, lacking good generalization ability.
[0006] 3. Lack of real-time dynamic response ability: Most existing models use static data for offline training and are difficult to respond in a timely manner to real-time changes in charging scenarios, such as sudden weather, temporary traffic congestion, or changes in user charging preferences, and cannot meet the real-time prediction requirements.
[0007] 4. Contradiction between model complexity and scalability: To improve prediction accuracy, traditional models continuously increase model complexity, resulting in high training costs, difficult deployment, and poor scalability of the model in different scenarios.
[0008] 5. Ignoring cross-modal feature relationships leads to limited prediction accuracy: Traditional methods fail to effectively utilize the correlation characteristics of cross-modal data, resulting in poor performance of the model when dealing with complex scenarios. In particular, there is a lack of systematic solutions in the fusion of images, text, and structured data.
[0009] Based on the above problems, a new type of multi-modal data-driven efficient prediction method is needed to address the deficiencies of the existing technology and improve prediction accuracy, real-time performance, and system adaptability. Summary of the Invention
[0010] Aiming at the deficiencies of the existing technology, the present invention proposes a method for predicting electric vehicle charging behavior driven by a multi-modal large model, aiming to solve problems such as single data modality, poor model generalization ability, lack of real-time response ability, and insufficient cross-modal feature fusion in the existing methods. By introducing multi-modal data fusion technology, pre-trained large models, and dynamic scene perception mechanisms, the present invention realizes high-precision and real-time prediction of electric vehicle charging behavior, and improves the intelligent level of the charging management system and the user charging experience.
[0011] The specific technical solutions are as follows:
[0012] A method for predicting electric vehicle charging behavior driven by a multi-modal large model:
[0013] Data collection and preprocessing: Collect multi-modal data, including time series data, geographical and environmental data, user behavior data, and image and video data. Standardize, extract features, and reduce the dimension of various types of data to construct a charging behavior prediction data set.
[0014] Multi-modal feature extraction: Use pre-trained large models to extract multi-modal features, specifically including time series feature extraction based on the Transformer architecture, image feature extraction using the Vision Transformer (ViT) model, text feature extraction using the BERT model, and feature modeling of geospatial data using graph neural networks (GNNs).
[0015] Multi-modal feature fusion: Adopt feature concatenation and cross-modal attention mechanisms to fuse features of different modalities, generate a unified feature representation, and provide input for subsequent charging behavior prediction.
[0016] Construction of the charging behavior prediction model: Based on the fused features, construct a fully connected neural network (FCNN) for charging behavior prediction. Adopt a multi-task learning strategy to optimize the time prediction error and load prediction error to improve the overall performance of the prediction model.
[0017] Real-time Prediction and Dynamic Scene Awareness: Access real-time data streams, use the attention mechanism to re-weight and fuse features, dynamically update the charging behavior prediction results, and achieve scene awareness and adaptive prediction.
[0018] Preferably, in the data collection and preprocessing step, the time series data includes charging load and historical electricity consumption data, the geographical and environmental data includes charging station locations, road traffic flow, and weather information, the user behavior data includes user reservation records, consumption preferences, and charging time preferences, and the image and video data includes parking lot camera images and vehicle queuing situations.
[0019] Preferably, in the multi-modal feature extraction step, the time series features are extracted by a model based on the Transformer architecture and are described by the following formula:
[0020]
[0021] In the formula, represents the input time series data, represents the extracted time series features.
[0022] The image features are extracted by the ViT model, and the formula is as follows:
[0023]
[0024] In the formula, is the input image, is the extracted image feature vector.
[0025] The text features are extracted by the BERT model, and the formula is:
[0026]
[0027] In the formula, is the input text data, is the extracted text feature representation.
[0028] The geographical and environmental features are modeled by a graph neural network (GNN), and the formula is:
[0029]
[0030] In the formula, represents the feature representation of node in the -th layer, is the set of adjacent nodes of node , is the weight matrix of the -th layer, is the activation function.
[0031] Preferably, in the multi-modal feature fusion step, feature splicing and cross-modal attention mechanism are adopted for fusion. The feature splicing formula is:
[0032]
[0033] In the formula, is the image feature, is the text feature, is the geographical feature.
[0034] The cross-modal attention mechanism formula is:
[0035]
[0036] In the formula, and are the weight matrices of the query and the key respectively, and are the feature representations of different modalities, is the attention weight, which is used to align the features of different modalities.
[0037] Preferably, the charging behavior prediction model is a fully connected neural network (FCNN), and the model output includes the prediction results of the charging time and the charging load. The loss function of the model is a multi-task learning loss, which is defined as:
[0038]
[0039] In the formula, is the charging time prediction error, which is defined as:
[0040]
[0041] In the formula, is the load prediction error, and are the loss weight parameters, is the number of samples.
[0042] Compared with the prior art, the beneficial effects of the present invention are:
[0043] 1. The prediction accuracy and comprehensiveness are improved: By fusing multi-modal data such as time series, geographical information, user behavior, and image and video, the comprehensiveness and accuracy of charging behavior prediction are significantly improved.
[0044] 2. The generalization ability and adaptability of the model are enhanced: By introducing pre-trained large models (such as ViT, BERT, etc.) for feature extraction and combining with the cross-modal attention mechanism, the adaptability and generalization ability of the model under different regions and user characteristics are effectively improved.
[0045] 3. Real-time dynamic response and scenario awareness are achieved: By accessing real-time data streams and dynamic scenario awareness mechanisms, the model can update prediction results in real time according to environmental changes, meeting the charging behavior prediction requirements in complex scenarios.
[0046] 4. The scalability and flexibility of the system are improved: The proposed method has good scalability and can adapt to different regions, different scales of charging stations and diverse user needs, facilitating its popularization and application in large-scale charging networks.
[0047] 5. Charging management and grid scheduling are optimized: The present invention can not only provide users with personalized charging time prediction and optimization suggestions, but also combine with the regional grid scheduling system to optimize the charging load distribution, improving the operation efficiency and reliability of the grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] To more clearly describe the technical solutions and implementation steps of the present invention, a drawing is provided to show the overall framework of the present invention. The drawing is intended to assist in understanding the content of the invention and is not used to limit its protection scope. Similar elements in the drawing are marked consistently, and the scales and dimensions in the drawing may differ from the actual situation.
[0049] Figure 1 : The overall process framework of the method for predicting electric vehicle charging behavior driven by the multimodal large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The embodiments of the present invention will be described in detail with reference to the drawings. The following embodiments are intended to show the technical details of the present invention, but do not constitute a limitation on the protection scope of the present invention. Any modification or replacement based on the technical solution of the present invention, if it does not exceed the substantial content of the present invention, shall be regarded as within the protection scope of the present invention.
[0051] As a preferred embodiment, step S1 is specifically as follows:
[0052] In the present invention, data collection and preprocessing are the basis for ensuring the accuracy of the model. By integrating multiple data sources, multi-dimensional information covering electric vehicle charging behavior is obtained, providing high-quality data for subsequent model training. Data collection covers four types of information: time series data, geographical and environmental data, user behavior data, and image and video data. These data are sourced from smart meters, geographic information systems (GIS), user interaction platforms, and charging station monitoring devices.
[0053] 1. Processing of time series data
[0054] Time series data mainly includes charging load and historical electricity consumption records. The preprocessing of this type of data includes detrending, standardization, and periodic analysis.
[0055] Detrending: The difference method is used to eliminate the long-term trend of the data to highlight the periodic changes. The difference formula is:
[0056]
[0057] In the formula, is the load data at time , and is the data after differencing.
[0058] Normalization: To eliminate the dimensional difference, the Z-score normalization method is used to process the data. The formula is:
[0059]
[0060] In the formula, is the mean of the data, and is the standard deviation.
[0061] Periodic analysis: The period characteristics in the load data are extracted through the fast Fourier transform (FFT). The formula is:
[0062]
[0063] In the formula, is the total number of samples, and is the frequency.
[0064] 2. Processing of Geographic and Environmental Data
[0065] The geographic data includes the location information of charging stations, and the environmental data covers weather, traffic flow, etc. These data are modeled through a graph neural network (GNN) to capture spatial relationships and environmental characteristics.
[0066] Adjacency matrix construction: The spatial relationship between charging stations is represented as an adjacency matrix A. The formula:
[0067]
[0068] Normalization of environmental characteristics: Weather data such as temperature and humidity are processed through min-max normalization. The formula:
[0069]
[0070] In the formula: is the normalized value; is the original data value; is the minimum value in the data; is the maximum value in the data.
[0071] 3. Processing of User Behavior Data
[0072] User behavior data includes reservation records, consumption preferences, etc. Natural Language Processing (NLP) technology is used to process the text data to extract user charging preference features. Text vectorization: The TF-IDF method is used to convert the text data into numerical features:
[0073]
[0074] In the formula, is the word in the document the occurrence frequency of, is the total number of documents, is the document containing the word the number of documents.
[0075] 4. Image and Video Data Processing
[0076] Image and video data are processed through Convolutional Neural Network (CNN) and Vision Transformer (ViT) models to extract features such as parking lot occupancy rate and queue length. Key frame extraction: The frame difference method is used to detect key frames for video data. The formula is:
[0077]
[0078] In the formula, and are the pixel matrices of the current frame and the previous frame respectively, is the frame difference.
[0079] As a preferred embodiment, step S2 is specifically as follows:
[0080] Multi-modal feature extraction. After data preprocessing, it enters the feature extraction stage. Multi-modal feature extraction aims to obtain key features from different types of data and provide rich information input for the prediction model.
[0081] 1. Time Series Feature Extraction
[0082] A model based on the Transformer architecture is used to extract time series features, which can capture long-term dependencies and complex dynamic changes.
[0083] Feature extraction formula:
[0084]
[0085] In the formula, is the input time series data, is the extracted time feature.
[0086] Calculation of load volatility:
[0087]
[0088] In the formula, is the average load, is the load value at a single time point, is the number of time points.
[0089] 2. Image feature extraction
[0090] Use a pre-trained Vision Transformer (ViT) model to process image data and extract the occupancy rate and queue length features of the charging station.
[0091] ViT feature extraction formula:
[0092]
[0093] In the formula, is the input image, is the extracted image feature vector.
[0094] Calculation of parking lot occupancy rate:
[0095]
[0096] 3. Text feature extraction
[0097] Use the BERT model to extract the text features in the user reservation information and feedback, and capture the user's charging preferences.
[0098] BERT feature extraction formula:
[0099]
[0100] In the formula, is the input text, is the extracted text feature vector.
[0101] 4. Geographic and environmental feature extraction
[0102] Geographic and environmental features are modeled through a graph neural network (GNN) to capture the impact of spatial relationships and environmental changes on charging behavior. GNN feature extraction formula:
[0103]
[0104] In the formula, is the node at the layer's feature representation, is the set of adjacent nodes, is the weight matrix, is the activation function.
[0105] As a preferred embodiment, step S3 is specifically as follows:
[0106] Multi-modal feature fusion: After feature extraction is completed, it is necessary to fuse features of different modalities to generate a unified feature representation. The present invention adopts a method combining feature concatenation and cross-modal attention mechanism to ensure the integrity and consistency of the fused features.
[0107] 1. Feature concatenation, features of different modalities are concatenated to form a comprehensive feature representation, and the formula is:
[0108]
[0109] In the formula, is the time series feature, is the image feature, is the text feature, is the geographical feature.
[0110] 2. Cross-modal attention mechanism: On the basis of feature concatenation, a cross-modal attention mechanism is introduced to strengthen important features and suppress irrelevant information through a weighting method. The attention mechanism formula:
[0111]
[0112] In the formula, and are the weight matrices of the query and the key respectively, and are the feature representations of different modalities, is the attention weight.
[0113] 3. Output of the fused features: The fused features are used as the input of the charging behavior prediction model, providing multi-dimensional comprehensive information for the model.
[0114] As a preferred embodiment, step S4 is specifically as follows:
[0115] After multi-modal feature extraction and fusion are completed, it enters the construction stage of the charging behavior prediction model. The present invention uses a fully connected neural network (FCNN) to train the fused features to predict the charging behavior of electric vehicles, including key parameters such as charging time, charging load, and user arrival time. This model can not only process complex multi-modal data, but also has strong generalization ability, and is suitable for prediction requirements of different user groups and geographical regions.
[0116] 1. Model structure design
[0117] The fully connected neural network consists of an input layer, multiple hidden layers, and an output layer. The input layer receives multi-modal fusion features, the hidden layers perform feature transformation and deep learning through non-linear activation functions, and the output layer generates the final prediction results of the charging behavior.
[0118] Input layer: The main function of the input layer is to receive the feature vector from the multi-modal feature fusion step. . This feature vector combines time series features, image features, text features, and geographical features, and can comprehensively represent various influencing factors of the electric vehicle charging behavior. The input dimension depends on the number and dimension of the fusion features.
[0119]
[0120] In the formula, is the time series feature, is the image feature, is the text feature, is the geographical feature.
[0121] Hidden layer: The hidden layer is responsible for learning the non-linear relationships from the input features to capture complex patterns and rules. The neurons in each layer are fully connected to all neurons in the previous layer, and the non-linear activation function ReLU (Rectified Linear Unit) is widely used in the hidden layer to introduce non-linearity and improve the expressive power of the model. The output formula of the hidden layer is:
[0122]
[0123] In the formula, represents the output of the -th layer, is the weight matrix, is the bias vector, is the ReLU activation function. Through the multi-layer stacking method, the model can gradually extract the high-level features in the data and finally obtain the prediction results in the output layer.
[0124] Output layer: The design of the output layer depends on the specific prediction task. In the present invention, the output layer is responsible for generating multiple prediction results, including the charging time, charging load, and user arrival time. To achieve this multi-task prediction, the output layer adopts a multi-head output structure, and each output head corresponds to a specific prediction task. The calculation formula of the output layer is:
[0125]
[0126] In the formula, and are the weight matrix and bias term of the output layer, respectively, is the output of the last hidden layer, is the final prediction result.
[0127] 2. Loss Function Design
[0128] To improve the accuracy and robustness of the prediction model, the present invention adopts a multi-task learning loss function. Multi-task learning can optimize multiple related tasks simultaneously in the same model, improving the overall performance and generalization ability of the model.
[0129] Time prediction error: The prediction error of the charging time is measured using the mean squared error (MSE). MSE can effectively penalize large prediction errors and ensure the accuracy of the model in the time prediction task. The formula is:
[0130]
[0131] In the formula, is the number of samples, is the actual charging time, is the charging time predicted by the model.
[0132] Load prediction error: The prediction of the charging load is also calculated using MSE to ensure the accuracy of the load prediction. The formula is:
[0133]
[0134] In the formula, is the actual charging load, is the predicted load.
[0135] Multi-task learning loss function: To optimize the time and load predictions simultaneously, the present invention adopts a weighted combination multi-task loss function. This loss function balances the importance of different tasks by adjusting the weight parameters and The formula is:
[0136]
[0137] This loss function design enables the model to find the optimal balance point between different tasks and improve the overall prediction performance.
[0138] 3. Model Training and Optimization
[0139] Model training is the process of adjusting network parameters by optimizing the loss function. The present invention uses the Adam optimizer for training and introduces regularization and early stopping strategies to prevent overfitting and improve the generalization ability of the model.
[0140] Optimization Algorithm: The Adam optimizer combines the methods of momentum and adaptive learning rate, which can adaptively adjust the learning rate during training, improving the convergence speed and stability. The parameter update formula is as follows:
[0141]
[0142] In the formula, is the model parameter, is the learning rate, and are the first-order and second-order moment estimates respectively, is a small constant to prevent division by zero.
[0143] Regularization: To prevent the model from overfitting during training, an L2 regularization term is added to constrain the weight parameters of the model. The formula for the regularization term is:
[0144]
[0145] In the formula, is the regularization coefficient, are the weight parameters in the model.
[0146] Early Stopping Strategy: During the model training process, continuously monitor the loss function value of the validation set. If the error of the validation set does not decrease after several consecutive rounds of training, stop the training in advance. This strategy effectively prevents the model from overfitting on the training data and improves the generalization ability in practical applications.
[0147] As a preferred embodiment, step S5 is specifically as follows:
[0148] Another core innovation of the present invention lies in introducing real-time data streams to achieve dynamic scene perception and real-time prediction of charging behavior. By updating the model input in real time, the adaptability and prediction accuracy of the model in complex scenarios are improved.
[0149] 1. Access to Real-Time Data Streams: The introduction of real-time data streams is the key difference between the present invention and traditional static models. By accessing real-time data such as traffic flow, weather changes, and charging station queue situations, the model can dynamically adjust the prediction results, significantly enhancing the response ability to emergencies and environmental changes.
[0150] • Traffic Flow Data: Traffic conditions directly affect the time when users arrive at the charging station. By obtaining road condition information through a real-time traffic monitoring system, the model can predict the arrival probability of users at a specific time period.
[0151] • Weather data: Weather changes have a significant impact on the charging demand of electric vehicles. For example, extreme high or low temperatures increase the energy consumption demand of the battery, thus changing users' charging behavior. The present invention dynamically adjusts the charging demand prediction by obtaining information such as the current temperature, humidity, and precipitation in real time.
[0152] • Charging station status data: The real-time occupancy rate and queue length of charging stations are important factors affecting users' choice of charging stations. By monitoring the real-time status of charging stations, the model can provide more targeted charging suggestions.
[0153] 2. Dynamic feature weighting: To make full use of real-time data, the present invention introduces a dynamic feature weighting mechanism. Through the attention mechanism, real-time data and historical data are weighted and fused to highlight the contribution of real-time data to the prediction result. Attention mechanism formula:
[0154]
[0155] In the formula, is the fused feature, is the real-time data feature. The attention mechanism can dynamically adjust the importance of features according to the changes in real-time data, ensuring that the model can quickly respond in the face of emergencies.
[0156] 3. Dynamic scenario awareness and adaptive adjustment: Dynamic scenario awareness is another innovation point of the present invention. After receiving real-time data, the model can automatically identify environmental changes and adjust the prediction strategy to adapt to the changing charging scenarios.
[0157] • Adaptive adjustment strategy: When significant environmental changes are detected, the model recalculates the feature weights and adjusts the prediction result.
[0158] • Real-time feedback mechanism: The prediction result of the model is fed back to the user in real time through the API interface, providing personalized charging suggestions for the user. These suggestions include the best charging time, recommended charging stations, and estimated queue time, etc.
Claims
1. A method for predicting electric vehicle charging behavior driven by a multimodal large model, characterized in that, It includes the following steps: Data collection and preprocessing: Collect multi-modal data related to electric vehicle charging behaviors, including time series data, geographical and environmental data, user behavior data, as well as image and video data. Standardize, extract features, and perform dimensionality reduction on each type of data respectively; Multi-modal feature extraction: Use pre-trained large models to extract features of each modality, including time series models based on Transformer, image feature extraction using Vision Transformer (ViT) models, text feature extraction using BERT models, and feature modeling of geographical and environmental data using graph neural networks (GNNs); Multi-modal feature fusion: Concatenate and cross-modal align features of different modalities, and adopt a cross-modal attention mechanism to generate a unified fused feature representation; Build a charging behavior prediction model: Based on the fused features, build a fully connected neural network (FCNN) model for charging behavior prediction, and optimize the time prediction error and load prediction error through multi-task learning; Real-time prediction and dynamic scenario perception: Introduce real-time data streams, re-weight features using the attention mechanism, and dynamically update the charging behavior prediction results.
2. The method according to claim 1, wherein The time series data includes charging load and historical electricity consumption data, the geographical and environmental data includes charging station locations, road traffic flow, and weather information, the user behavior data includes user reservation records, consumption preferences, and charging time preferences, and the image and video data includes parking lot camera images and vehicle queuing situations.
3. The method according to claim 1, characterized in that In the multi-modal feature extraction step, time series features are extracted through a time series model based on the Transformer architecture, image features are extracted through ViT models, text features are extracted through BERT models, and geographical and environmental features are modeled using graph neural networks (GNNs).
4. The method according to claim 1, characterized in that, In the multi-modal feature fusion step, feature concatenation and a cross-modal attention mechanism are adopted for feature fusion, and the cross-modal attention mechanism realizes the alignment of different modality features through query weight matrices and key weight matrices.
5. The method according to claim 1, wherein The charging behavior prediction model is a fully connected neural network (FCNN), and its output includes prediction results of charging time and charging load. The loss function of the model includes time prediction error and load prediction error.
6. The method according to claim 1, wherein In the real-time prediction and dynamic scenario perception step, by accessing real-time data streams, the attention mechanism is used to re-weight the fused features.
Citation Information
Cited By
Electric vehicle charging behavior characteristic description method and system fused with multi-modal data
CN121365218A