Charging pile load balancing method and system based on neural network
By constructing a unified charging pile load balancing framework based on neural networks, the problems of knowledge sharing and stability of the charging pile load balancing system in multiple scenarios are solved, realizing multi-scenario adaptation and smooth switching, and improving the system's adaptability and response speed.
Patent Information
- Application Number
- CN202511555686.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2025-11-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing charging pile load balancing technologies lack a unified framework for handling diverse scenarios, and there is a lack of knowledge sharing mechanisms between different scenario systems, which causes system performance to fluctuate during scenario transitions, affecting user experience and system efficiency.
A neural network-based approach is adopted, which analyzes charging pile status, environmental and user behavior data through a multimodal autoencoder network to generate a unified scene vector representation. The variational autoencoder and density clustering algorithm are combined to identify scene categories, and a knowledge decomposition engine is built to achieve cross-scene knowledge sharing. Finally, a load balancing decision with smooth transition is generated through an adaptive neural network architecture and meta-reinforcement learning algorithm.
It achieves multi-scenario adaptability, improves the system's generalization ability and response speed, reduces resource redundancy, ensures stable operation of the system during scenario switching, and enhances user experience and system efficiency.
Smart Images

Figure CN121032147A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent management of electric vehicle charging infrastructure, more specifically, it relates to a charging pile load balancing method and system based on neural network. BACKGROUND
[0002] With the rapid popularization of electric vehicles, the construction and management of charging infrastructure are facing unprecedented challenges. As the key facility for electric vehicle energy supply, the load balancing of charging piles directly affects user experience, power grid stability and overall operation efficiency. Traditional charging pile load balancing techniques mainly use static planning or simple dynamic scheduling algorithms. These methods can achieve certain results in a single scenario, but they are difficult to cope with complex and variable actual operating environments.
[0003] There are three main problems in current charging pile load balancing technology: existing technologies are usually designed for specific scenarios, such as daily operation, peak hours or special weather conditions, lack of unified framework to handle diversified scenarios, resulting in system redundancy and resource waste; load balancing systems in different scenarios are often independent of each other, lack of effective knowledge sharing mechanism, cannot fully utilize cross-scene experience data, reduces the overall learning efficiency of the system; during scenario transition, system performance often fluctuates significantly or even drops sharply, affecting service quality and user experience.
[0004] With the development of artificial intelligence technology, neural networks have shown great potential in solving complex decision-making problems. However, how to build a unified charging pile load balancing framework that can adapt to multiple scenarios, realize knowledge sharing and smooth switching, is still a technical problem to be solved. The present application is aimed at this challenge, and proposes a multi-modal fusion charging load balancing framework based on scene adaptability computing theory and knowledge decomposition and reorganization paradigm, aiming to achieve the goals of multi-scene adaptation, knowledge sharing and smooth switching. SUMMARY
[0005] The present application provides a charging pile load balancing method and system based on neural network, which solves the technical problems of charging pile load balancing system in related technology lacking unified framework to handle diversified scenarios, lacking knowledge sharing mechanism between systems in different scenarios, and system performance fluctuating greatly during scenario transition.
[0006] The present application provides a charging pile load balancing method based on neural network, comprising the following steps: Scene modeling and representation learning, using a multi-modal autoencoder network to analyze charging pile state, environment, user behavior multi-source heterogeneous data, fusing features of each modality through attention mechanism, generating a unified scene vector representation; Scene recognition and classification, using variational autoencoder to reduce the dimension of scene vector and apply density clustering algorithm to identify typical scene categories, and establish a scene classification system; Knowledge decomposition and representation, build knowledge decomposition engine, decompose the experience knowledge of charging load balancing into cross-scene shared knowledge and scene-specific knowledge, realize effective separation and organization of knowledge; Adaptive neural network architecture generation, based on the similarity between the current scene and the known scene, combined with shared knowledge and scene-specific knowledge, automatically generate a neural network architecture that adapts to the current scene; Meta-strategy learning and scene smooth transition, train the meta-strategy network to generate load balancing decisions through meta-reinforcement learning algorithm, and realize smooth transition through gradual model switching when the scene changes.
[0007] In a preferred embodiment, the multi-modal autoencoder network comprises: Convolutional neural network for processing spatial-related data; Long short-term memory network for processing time series data; Multilayer perceptron for processing static feature data; Attention mechanism for fusing features of each modality.
[0008] In a preferred embodiment, the knowledge decomposition algorithm further comprises: L1 regularization constraint is imposed on the parameters of all scene-specific knowledge models to make the scene-specific knowledge models as sparse as possible; Structured regularization constraints are imposed on shared knowledge models and scene-specific knowledge models to make knowledge have better organization and interpretability.
[0009] In a preferred embodiment, the step of calculating the similarity between the current scene and the known scene comprises: Calculate the dot product of the current scene vector and the known scene vector, then divide by the product of the magnitudes of the two vectors to get the similarity value; Select the K scenes with the highest similarity as the similar scenes of the current scene.
[0010] In a preferred embodiment, the step of automatically generating a neural network architecture that adapts to the current scene comprises: Build a search space containing multiple neural network layer types, connection methods and activation functions; Apply reinforcement learning controller to select the most suitable network structure and hyperparameters for the current scene; Fuse shared knowledge models and related scene-specific knowledge models according to similarity weights to initialize the parameters of the newly generated neural network.
[0011] In a preferred embodiment, the meta-reinforcement learning algorithm comprises: sampling a batch of scenes from the set of scenes; quickly adapting a meta-policy network using a small number of samples per scene; generating decision trajectories in each scene using the adapted policy; computing policy and value function gradients based on the collected trajectories; updating the meta-policy network parameters using an optimizer.
[0012] In a preferred embodiment, the scene change detection includes: abrupt change detection, identifying sudden changes by computing cosine distance of consecutive scene vectors; trend detection, monitoring the trend of key dimensions of scene vectors by exponential moving average method.
[0013] In a preferred embodiment, the progressive model switching includes: designing a dynamic smoothing coefficient to control the fusion ratio of old and new model decisions; smoothly adjusting the smoothing coefficient over time, gradually reducing the weight of the old model and gradually increasing the weight of the new model, to achieve a smooth transition.
[0014] In a preferred embodiment, a neural network-based charging pile load balancing method further includes a scene prediction and preparation step: using a time series prediction model to analyze the evolution trend of the scene and predict possible new scenes; preparing model parameters and resources in advance for high-probability scenes to reduce response delay during actual switching.
[0015] In a preferred embodiment, a neural network-based charging pile load balancing system for executing a neural network-based charging pile load balancing method includes: a scene modeling and representation learning unit for analyzing charging pile status, environment, and user behavior multi-source heterogeneous data using a multi-modal autoencoder network, and generating a unified scene vector representation by fusing features of each modality through an attention mechanism; a scene recognition and classification unit for dimensionality reduction of the scene vector using a variational autoencoder and identifying typical scene categories by applying a density clustering algorithm, establishing a scene classification system; a knowledge decomposition and representation unit for constructing a knowledge decomposition engine to decompose the experience knowledge of charging load balancing into cross-scene shared knowledge and scene-specific knowledge, achieving effective separation and organization of knowledge; an adaptive neural network architecture generation unit for automatically generating a neural network architecture adapted to the current scene based on the similarity of the current scene and known scenes, combining shared knowledge and scene-specific knowledge; The meta-policy learning and scene smooth transition unit is used to train the meta-policy network to generate load balancing decisions through meta-reinforcement learning algorithms, and to achieve a smooth transition through progressive model switching when the scene changes.
[0016] The beneficial effects of this invention are as follows: The constructed unified charging pile load balancing framework can adapt to various scenarios, including daily operations, special events, extreme weather, and emergency situations, effectively solving the resource redundancy problem caused by designing separate systems for different scenarios in traditional methods. By fusing various types of data through a multimodal autoencoder network and attention mechanism, the system can accurately identify and represent the current scenario, providing a reliable basis for subsequent decision-making.
[0017] The knowledge decomposition and representation mechanism enables cross-scenario knowledge sharing and transfer, decomposing the experiential knowledge of charging load balancing into cross-scenario shared knowledge and scenario-specific knowledge, significantly improving the system's generalization ability. When faced with new scenarios, the system can quickly construct initial decision-making strategies based on existing knowledge without the need for redevelopment or significant adjustments, thus significantly improving the system's adaptability and response speed.
[0018] By extracting high-order decision principles from multi-scenario decision-making experience through a meta-policy learning mechanism, the load balancing decision for the current scenario is guided, achieving a higher level of knowledge transfer. Meanwhile, the scenario smooth transition controller ensures stable system operation during scenario transitions through gradual model switching and buffered decision generation, avoiding the problem of precipitous performance degradation.
[0019] The adaptive neural network architecture generation mechanism can automatically generate the most suitable neural network architecture for the current or new scenario based on scene characteristics and decomposed knowledge, realizing dynamic adaptation of the model. This method not only improves the system's ability to handle various scenarios, but also reduces the need for manual intervention, reduces maintenance costs, and provides strong technical support for the intelligent operation of charging infrastructure. Attached Figure Description
[0020] Figure 1 This is a flowchart of a charging pile load balancing method based on neural networks according to the present invention; Figure 2 This is a radar chart comparing the system performance of the present invention in different scenarios; Figure 3 This is a line graph showing the changes in system performance during scene switching according to the present invention; Figure 4 This is a bar chart comparing the different knowledge decomposition methods of the present invention; Figure 5 This is an area graph showing the long-term performance change trend of the system according to the present invention. Figure 6This is a scatter plot of the zero-sample scene adaptability test results of the present invention; Figure 7 This is the Sankey diagram for multi-scenario knowledge sharing efficiency analysis in this invention. Detailed Implementation
[0021] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0022] At least one embodiment of the present invention discloses a charging pile load balancing method based on neural networks, such as... Figure 1 As shown, it includes the following steps: Step 1, Scene Modeling and Representation Learning: Utilize a multimodal autoencoder network to analyze multi-source heterogeneous data on charging pile status, environment, and user behavior. Integrate features from various modalities through an attention mechanism to generate a unified scene vector representation. Specifically, it includes the following sub-steps: Step 1.1, multimodal data preprocessing; The charging pile status data (such as usage rate and waiting queue length), environmental data (such as meteorological data and power grid load), and user behavior data (such as historical charging records and location information) are cleaned, normalized, and time-series aligned to form a structured multimodal input dataset.
[0023] Step 1.2, Scene Feature Extraction; Specialized neural network modules are used to process various types of data. Specifically, convolutional neural networks are used to process spatially relevant data (such as the geographical distribution of charging piles), long short-term memory networks are used to process time-series data (such as historical load changes), and multilayer perceptrons are used to process static feature data (such as charging pile configuration information), thereby extracting feature representations for each modality.
[0024] In one embodiment of this application, the implementation of the multimodal autoencoder network specifically includes the following structure: The convolutional neural network for spatial data processing consists of three convolutional layers, each using 64, 128, and 256 3×3 convolutional kernels respectively. Each layer is followed by a batch normalization layer and a ReLU activation function. Finally, a global average pooling layer converts the feature map into a fixed-length feature vector. This module is specifically designed for processing spatially relevant data such as the geographical distribution of charging stations and traffic flow.
[0025] It should be noted that the ReLU activation function is a widely used non-linear activation function in neural networks, and its definition is: ; in, Represents the ReLU function. Indicates the input value. This means taking the larger value between 0 and x, when The function outputs 0 when Time function output This piecewise linear design allows ReLU to maintain a linear mapping relationship on the positive half-axis, while completely suppressing signal transmission on the negative half-axis, thus introducing nonlinear transformation capabilities into the neural network while maintaining computational efficiency.
[0026] That is, the output is 0 when the input is negative, and the output is equal to the input value when the input is positive. This function design can effectively solve the gradient vanishing problem in deep neural networks, and it is also simple to compute and converges quickly, making it suitable for use in convolutional neural networks for feature extraction.
[0027] The Long Short-Term Memory (LSTM) network for time-series data processing comprises a two-layer LSTM structure, each layer containing 128 hidden units. A temporal embedding layer is added between the input layer and the LSTM layer to capture periodic patterns (such as weekday / weekend, morning / evening rush hour, etc.). This module is specifically designed to process time-series data such as historical load changes and user arrival frequency.
[0028] The multilayer perceptron for static features consists of three fully connected layers with 256, 128, and 64 neurons respectively. Each layer is followed by a batch normalization layer, a LeakyReLU activation function, and a Dropout layer with a probability of 0.2 to prevent overfitting. This module is used to process static data such as charging station configuration information and fixed user group characteristics.
[0029] In scenarios with high-dimensional features, self-attention mechanisms can be used to enhance feature extraction capabilities within each modality. Specifically, this is implemented as a multi-head self-attention layer with 8 heads and an attention dimension of 64, preserving sequence order information through positional encoding.
[0030] In some embodiments, the spatial data processing module may also employ a graph convolutional network structure to better capture the spatial topological relationships between charging piles. Furthermore, temporal data processing may use gated recurrent units (GRUs) or temporal convolutional networks (TCNs) instead of LSTM structures to obtain different temporal modeling characteristics.
[0031] like Figure 2As shown, the key performance indicators of the multimodal fusion charging load balancing framework driven by scenario-adaptive computing theory are demonstrated in different scenarios (daily operation, special events, extreme weather, and emergency situations), and compared with traditional methods, including five dimensions: balancing efficiency, response time, resource utilization, user satisfaction, and system stability.
[0032] Step 1.3, Scene Representation Fusion; By fusing features from various modalities through an attention mechanism, a unified scene vector representation is generated. .
[0033] In this process, the system first acquires multiple modal feature vectors, each of which represents the feature representation of different types of data.
[0034] Then, the system calculates an importance weight for each modality feature vector, which reflects the relevance and importance of the modality in the current scenario.
[0035] Next, the system performs a weighted sum of all modal feature vectors according to their corresponding weight values to obtain a vector that comprehensively represents the current scene. The weight values are determined by the data relevance and the characteristics of the current scene, enabling the system to automatically adjust the fusion ratio of each modal information according to different scenarios.
[0036] In one embodiment of this application, the calculation of attention weights is implemented using a soft attention mechanism.
[0037] First, the system applies three different linear transformation functions to each modal feature vector to generate a query vector, a key vector, and a value vector.
[0038] Then, by calculating the inner product of the query vector and the key vector and scaling it, a score representing the similarity is obtained.
[0039] Next, the system applies the Softmax function to normalize these scores to obtain the final attention weights.
[0040] Finally, the system uses these weights to perform a weighted summation of the value vectors, generating a fused scene representation vector. The scaling factor, the square root of the key vector dimension, is used to adjust the numerical range of the inner product result and prevent the gradient vanishing problem. The transformation function is implemented using different learnable weight matrices and is continuously optimized during training via backpropagation to adapt to different data features.
[0041] It should be noted that, in another embodiment, a gating fusion mechanism can be used instead of an attention mechanism, which automatically adjusts the fusion ratio of each modality by learning the interaction relationship between different modal features.
[0042] Step 1.4, Scene Clustering and Recognition; Based on the generated scene vector representations, a scene space is constructed using a variational autoencoder, and a density clustering algorithm is applied to automatically identify typical scene categories. Each scene is represented by a scene vector, forming a scene set.
[0043] It should be noted that the variational autoencoder (VAE) is a generative model whose internal implementation includes two core components: an encoder and a decoder.
[0044] The encoder maps the input scene vector to distribution parameters (mean vector and variance vector) in the latent space, rather than deterministic points; The decoder then samples the latent vectors from this distribution and reconstructs them back into the original space.
[0045] Specifically, the encoder outputs not a single latent vector, but rather the probability distribution parameters that define the latent vector, typically the mean and variance of a Gaussian distribution.
[0046] This design enables VAEs to capture the uncertainty in scene representation and generate a smoother latent space, which is beneficial for subsequent scene clustering and interpolation operations.
[0047] Optionally, in some embodiments, in addition to variational autoencoders, other combinations of dimensionality reduction and clustering methods can be used, such as a combination of t-SNE dimensionality reduction and K-means clustering, or a combination of UMAP dimensionality reduction and HDBSCAN clustering, to obtain scene representations and clustering results with different characteristics.
[0048] like Figure 3 As shown, the performance change curves of the system when switching from one scene to another are displayed under the scene smooth transition control technology, and the performance difference between having and not having the smooth transition control technology is compared.
[0049] Step 2, Scene Recognition and Classification: Use variational autoencoders to reduce the dimensionality of scene vectors and apply density clustering algorithms to identify typical scene categories, thus establishing a scene classification system; Specifically, it includes the following sub-steps: Step 2.1, Neural Network Model Training; For each identified typical scenario Supervised learning methods are used to train specific load balancing neural network models. Each model receives the current charging network state as input and outputs the optimal load distribution strategy.
[0050] Step 2.2, Application of knowledge decomposition algorithm; By applying knowledge distillation and decomposition algorithms, we analyze the weight parameters of all scenario-specific models to identify common and dissimilar features.
[0051] In this process, the system first calculates the degree of difference between the neural network model for each scene and the sum of the shared knowledge model and the scene-specific knowledge model, and uses the Frobenius norm of the matrix (i.e. the square root of the sum of squares of all elements) to measure this difference.
[0052] At the same time, the system applies L1 regularization constraints to the parameters of all scenario-specific knowledge models, making the scenario-specific knowledge models as sparse as possible and retaining only the truly scenario-specific knowledge.
[0053] The system optimizes the shared knowledge model and the scenario-specific knowledge model by minimizing the sum of these two parameters, where the regularization parameter... The sparsity of the decomposition is controlled; the larger the value, the sparser the scene-specific knowledge model and the higher the proportion of shared knowledge.
[0054] It should be noted that L1 regularization is a commonly used regularization technique in machine learning. It makes the model parameters sparse by adding a penalty term to the sum of the absolute values of the model parameters in the optimization objective.
[0055] The specific calculation method involves summing the absolute values of all parameters in the model and then multiplying the sum by a regularization coefficient as a penalty term added to the optimization objective. This regularization method is characterized by precisely compressing unimportant parameters to zero, thereby achieving feature selection and effectively identifying and retaining scenario-specific key knowledge during knowledge decomposition.
[0056] In some embodiments, the objective function described above can be extended by adding a knowledge structure constraint term to give shared knowledge and scenario-specific knowledge better structured characteristics: In this extended decomposition objective, the system adds a structured regularization term to the original decomposition objective. This structured regularization term imposes structural constraints on the shared knowledge model and the scenario-specific knowledge model, resulting in better organization and interpretability of the decomposed knowledge.
[0057] The weighting coefficients of structural constraints are used to balance the importance of the original decomposition objective and the structural constraints. Structured regularization functions can be designed based on graph Laplacian regularization to ensure coherence between related knowledge points; or based on hierarchical constraints to present a clear hierarchical relationship in the knowledge organization.
[0058] It should be noted that the structured regularization function The specific implementation can take many forms: When using the graph Laplacian regularization form, the function first constructs a graph relating the parameters, where nodes represent parameters and edges represent the degree of association between parameters. Then, it calculates the Laplacian matrix of the graph and finally calculates the structural constraint values through the quadratic form of the parameter vectors and the Laplacian matrix. When using a hierarchical constraint form, this function applies different strengths of regularization to the parameters of different layers according to the hierarchical structure of the neural network. Generally, the regularization strength of the parameters of deeper layers is higher, so as to promote knowledge sharing in shallow layers and specialization in deep layers.
[0059] like Figure 4 As shown, the performance of different knowledge decomposition methods (the method of this patent, traditional model fusion, single-scene model, and transfer learning) in four aspects are compared: efficiency of shared knowledge extraction, retention rate of scene-specific knowledge, knowledge transfer capability, and consumption of computing resources.
[0060] Step 2.3, shared knowledge extraction; Based on the decomposition results, a shared knowledge base is constructed. It contains load balancing principles, modes, and parameters applicable across various scenarios. This knowledge base uses a graph neural network structure for storage, enabling it to express the topological relationships and dynamic characteristics of charging networks.
[0061] Step 2.4, Scenario-specific knowledge organization; For each scenario Build its specific knowledge base It contains load balancing strategies and parameters unique to this scenario. It adopts a key-value storage structure, with scenario features as keys and corresponding specific strategy parameters as values.
[0062] It should be understood that, in another embodiment, the shared knowledge base may also be stored in the form of tensor decomposition or neural symbolic rules to adapt to different application needs. Furthermore, a scenario-specific knowledge base may also be structured as a decision tree or rule set to improve the interpretability of the knowledge.
[0063] like Figure 5 As shown, the performance improvement trend of the system over time is demonstrated, verifying the technical effect that "the system performance improves by 35% over time". At the same time, the trend of parameter reduction is shown, verifying the technical effect that "the system can achieve near-dedicated system performance with 20% fewer parameters".
[0064] Step 3, Knowledge Decomposition and Representation: Construct a knowledge decomposition engine to decompose the experience knowledge of charging load balancing into cross-scenario shared knowledge and scenario-specific knowledge, thereby achieving effective separation and organization of knowledge. Specifically, it includes the following sub-steps: Step 3.1, Define the network structure search space; We construct a search space that includes various neural network layer types (such as convolutional layers, attention layers, graph convolutional layers, etc.), connection methods, and activation functions, providing a basic component library for the automatic generation of model architectures.
[0065] Step 3.2, scene similarity calculation; For the input current scene vector, calculate its similarity with all known typical scenes, and find the most similar one. A scenario.
[0066] In this process, the system calculates the similarity between the current scene vector and the known scene vectors. Specifically, it first calculates the dot product of the two scene vectors, and then divides it by the product of the magnitudes (i.e., the Euclidean norms) of the two vectors.
[0067] This calculation method measures the similarity between the directions of two vectors, with similarity values ranging from -1 to 1. Values closer to 1 indicate greater similarity between the two scenes. This similarity calculation method is unaffected by the absolute magnitude of the vectors, focusing only on their direction, making it suitable for comparing the relative importance of features across different dimensions.
[0068] It should be noted that the Euclidean norm (i.e. the magnitude of a vector) is calculated by summing the squares of each element in the vector and then taking the square root, which represents the length of the vector in Euclidean space.
[0069] In the context of scene vectors, this value reflects the overall strength of scene features. By normalizing it by dividing by the product of the modulus, the influence of differences in feature strength between different scenes can be eliminated, allowing similarity calculation to focus on the comparison of feature patterns.
[0070] In some embodiments, in addition to cosine distance, other distance metrics such as Euclidean distance, Mahalanobis distance, or Earth movement distance can be used to obtain similarity assessments from different perspectives.
[0071] also, The value can be dynamically adjusted according to the specific application scenario. For example, it can be appropriately increased for less common scenarios. It is worthwhile to learn more about related scenarios.
[0072] Step 3.3, Adaptive network architecture generation; Based on scene similarity and a knowledge base, a reinforcement learning controller automatically generates the most suitable neural network architecture for the current scene. The controller selects appropriate network structures, connection methods, and hyperparameters according to scene characteristics to maximize the expected reward.
[0073] In this process, the reward function designed by the system consists of two parts: Performance metrics of the model in specific scenarios, such as load balancing efficiency and user waiting time; The complexity of a model is measured by factors such as the number of parameters and the amount of computation.
[0074] The system seeks a network architecture with good performance and moderate complexity by balancing these two components. The balancing parameters... This is used to adjust the relative importance of the two parts; the larger the value, the higher the requirement for model simplicity.
[0075] Optionally, in some embodiments, the reinforcement learning controller can be replaced with a gradient-based differentiable architecture search method or evolutionary algorithm to adapt to different architecture search requirements. Simultaneously, the reward function can also be extended to a multi-objective form, simultaneously considering multiple metrics such as performance, complexity, latency, and energy consumption.
[0076] Step 3.4, Knowledge Reorganization and Parameter Initialization; The shared knowledge base and the specific knowledge base of the relevant scenario are merged according to similarity weights to initialize the parameters of the newly generated neural network, forming an initial model for the current scenario.
[0077] In this process, the system first uses the parameters of the shared knowledge model as a basis, and then weights and fuses the specific knowledge model parameters of multiple similar scenarios and adds them to the base model.
[0078] The weight coefficient of each similar scene is calculated based on its similarity to the current scene. The higher the similarity, the greater the weight, ensuring that more relevant scene knowledge has a greater impact on the current model.
[0079] In another embodiment of this application, the initialization process can also be combined with knowledge distillation technology, allowing multiple models of similar scenarios to jointly guide the training of the current scenario model, thereby further improving knowledge transfer efficiency.
[0080] like Figure 6 As shown, the system's adaptability to new scenarios under different sample conditions (including zero samples) is demonstrated. The horizontal axis represents the number of training samples, and the vertical axis represents the performance score, verifying the technical effect that "the system can still achieve 90% of the performance level of a dedicated system in new scenarios with only a small amount or zero sample data".
[0081] Step 4: Adaptive neural network architecture generation. Based on the similarity between the current scene and known scenes, and combining shared knowledge and scene-specific knowledge, an adaptive neural network architecture is automatically generated to adapt to the current scene. Specifically, it includes the following sub-steps: Step 4.1, Multi-scenario decision data collection; Collect historical data on load balancing decisions under different scenarios, including state-action pairs, reward signals, and scenario information.
[0082] Construct a structured decision database as the foundation for meta-policy learning.
[0083] Step 4.2, Meta-policy network construction; A meta-policy neural network is constructed. This network takes the current state and scenario representation as input and outputs a load balancing strategy suitable for the current scenario. The network structure includes a scenario encoding layer, a policy generation layer, and a value evaluation layer, enabling it to handle decision knowledge across scenarios.
[0084] In one embodiment of this application, the meta-policy neural network is specifically implemented as follows: The overall network architecture adopts a context-based meta-reinforcement learning framework, which includes three main components: a scene encoding network, a policy network, and a value network.
[0085] Scene encoding networks convert scene vectors Mapped to the context embedding space, this is implemented using a three-layer fully connected network with 256, 128, and 64 neurons per layer, employing the ELU activation function and LayerNorm normalization layers. The output of the scene encoding network serves as the modulation signal for the policy and value networks, guiding them to adapt to different scenarios.
[0086] The policy network adopts a dual-stream architecture: one stream processes charging network state information, and the other stream processes scenario encoding information. The two streams are fused through an attention mechanism.
[0087] Specifically, the state flow consists of two graph convolutional layers (handling the topological relationships of charging piles) and one self-attention layer (handling the dependencies between charging piles). The scene flow includes a multilayer perceptron that converts scene encodings into policy-relevant representations; The fused representation generates a policy distribution, i.e., the load adjustment probability of each charging pile, through a multilayer perceptron.
[0088] The value network evaluates the expected reward of the current state and policy in a given scenario, employing a similar two-stream architecture, but its output is a scalar value instead of a policy distribution. The value network and policy network share parameters from early layers to improve learning efficiency.
[0089] To enhance the generalization ability of the meta-policy network, a self-regularization mechanism can be optionally introduced, including: weight decay (coefficient of 1e-4), feature normalization (L2 normalization of each intermediate layer representation) and information bottleneck constraint (limiting the information capacity of scene encoding).
[0090] Optionally, during the inference phase, a few-sample adaptation mechanism can be used for new scenarios to quickly fine-tune key subsets of network parameters through gradient descent, thereby further improving the adaptability of the meta-policy network.
[0091] It should be noted that, in another embodiment, the meta-policy network can also adopt a prototype-based meta-learning architecture, achieving rapid scene adaptation by learning scene prototypes and adaptation rules. Alternatively, model-independent meta-learning methods, such as optimized MAML (Model-Independent Meta-Learning Algorithm), can be used, enabling the model to quickly adapt to new scenes with a small number of gradient steps.
[0092] Step 4.3, Application of meta-reinforcement learning algorithm; Meta-policy networks are trained using meta-reinforcement learning algorithms, enabling them to quickly adapt to different scenarios and generate optimal policies. The training process includes three stages: policy pre-training, policy distillation, and policy fine-tuning, progressively extracting and fusing decision-making knowledge from various scenarios.
[0093] In this optimization process, the system performs expectation calculations on data from different scenarios, including two parts: The expected cumulative reward of the trajectory generated by the strategy in each scenario, where the discount factor reduces the future reward compared to the immediate reward; The KL divergence between the new policy and the prior policy is used to regularize the learning process and prevent the new policy from deviating too much from the prior knowledge.
[0094] By optimizing this objective function, the system can learn a meta-policy that performs well in various scenarios and conforms to prior knowledge.
[0095] It's important to note that KL divergence (KL divergence) is an asymmetric measure of the difference between two probability distributions. Specifically, it's calculated by taking the logarithm of the ratio of the probability of the new policy choosing that action to the probability of the prior policy choosing that action for each possible action in the policy distribution. Then, these logarithmic values are weighted and averaged using the probability of the new policy. This method measures the degree of change in the policy distribution, preventing the new policy from deviating too far from prior knowledge, thus maintaining the stability of the learning process.
[0096] In another embodiment of this application, the meta-reinforcement learning algorithm can be trained using the Model Independent Policy Optimization (MAPO) method, which combines the advantages of policy gradients and actor-critic architecture.
[0097] The training process includes the following steps: Scene sampling: Sample a batch of scenes from the scene set.
[0098] Inner loop adaptation: For each scenario, the meta-policy network is quickly adapted using a small number of samples.
[0099] Trajectory generation: Generate decision trajectories in various scenarios using an adapted strategy.
[0100] Gradient calculation: Calculate the policy gradient and value function gradient based on the collected trajectories.
[0101] Parameter update: Update the meta-policy network parameters using the Adam optimizer.
[0102] Furthermore, it is understood that in some embodiments, additional regularization terms may be added to the objective function, such as entropy regularization to encourage policy exploration or temporal consistency regularization to promote policy stability. In this extended objective function, the system adds two regularization terms: Policy entropy is used to encourage the policy to explore more possible actions and avoid premature convergence to a suboptimal policy. Its weight coefficient controls the degree of exploration. The temporal consistency regularization term is used to ensure that policies produce similar decisions under similar conditions, thereby enhancing the stability and predictability of the system. Its weight coefficient controls the importance of stability.
[0103] It should be noted that policy entropy The calculation method is to take the logarithm of the probability of all possible actions in the strategy distribution, and then use the negative value of the weighted sum of these probability values.
[0104] This calculation can measure the uncertainty or "diffusion" of the policy distribution. The higher the entropy value, the more uncertain the policy is and the more inclined it is to explore multiple possibilities.
[0105] The time consistency regularization term The implementation involves calculating the degree of difference in policy distributions under similar states, typically using KL divergence or L2 distance as metrics, and then taking a weighted average of the differences across all pairs of similar states.
[0106] This design ensures that strategies produce similar decisions in similar situations, enhancing the predictability and consistency of system behavior.
[0107] Step 4.4, Adaptive decision generation; For the current scenario, the meta-policy network combines scenario characteristics and the current charging network status to generate the optimal load balancing decision.
[0108] The decision-making process includes load allocation targets for each charging station, charging price adjustment strategies, and user charging guidance schemes, forming a complete load balancing execution plan.
[0109] It should be noted that the load allocation target is generated using a hierarchical decision-making method: First, the high-level decision-making module determines the overall load distribution strategy, such as centralized, decentralized, or hybrid. Then, the mid-level planning module refines the overall strategy into regional objectives, balancing the load of each region based on grid capacity and geographical location; Finally, the lower-level execution module calculates the specific load value for each charging station, taking into account the physical limitations and current state of the charging station.
[0110] This hierarchical decision-making structure enables the system to simultaneously meet the needs of global optimization and local constraints.
[0111] like Figure 7 As shown, the visualization shows the flow and sharing of knowledge in different scenarios, demonstrating the working mechanism of the knowledge decomposition and reorganization paradigm, and verifying the technical effect of "data utilization efficiency increased by 75% and model training speed increased by 65%".
[0112] Step 5: Meta-policy learning and smooth scene transition. The meta-policy network is trained using a meta-reinforcement learning algorithm to generate load balancing decisions, and a smooth transition is achieved through progressive model switching when the scene changes.
[0113] Specifically, it includes the following sub-steps: Step 5.1, Scene change detection; Continuously monitor environmental data and system status, calculate the difference between the current scene vector and the scene vector at the previous moment, and assess the degree of scene change. When the change exceeds a threshold, trigger the scene switching process.
[0114] In one embodiment of this application, the scene smooth transition controller is specifically implemented as follows: The scene change detection module employs a dual detection mechanism: Mutation detection and trend detection. Mutation detection identifies sudden changes by calculating the cosine distance between continuous scene vectors, with a threshold of 0.85 (an alarm is triggered when the cosine similarity is below this value). Trend detection monitors the changing trends of key dimensions of the scene vector using the exponential moving average method. When the direction of change is consistent for five consecutive time steps and the cumulative change exceeds a preset threshold, an alert is issued.
[0115] The combination of these two mechanisms can both capture sudden changes in the scene and detect gradual scene transitions in advance.
[0116] It should be noted that the exponential moving average method is a time series analysis technique that calculates the current trend by assigning different weights to historical data. More recent data is given higher weights, while the weight of earlier data decays exponentially over time.
[0117] The specific calculation formula is as follows:
[0118] ; in, Indicates the current time The exponential moving average, i.e., the smoothed value obtained after calculation; Indicates the current time The actual measured value, i.e., the original input data; Indicates the previous moment The exponential moving average, which is the result calculated in the previous step; It is the smoothing coefficient, with a value ranging from 0 to 1. It is used to control the degree of influence of the current measurement on the average value. The larger the value, the higher the weight of the current data and the more sensitive the system is to the response of new data; the smaller the value, the higher the weight of historical data and the stronger the smoothing effect.
[0119] The progressive model switching module is implemented as an adaptive smoothing function, and the rate of change of the fusion coefficients is dynamically adjusted according to the urgency of the scene change and the system stability requirements.
[0120] In this implementation, the system updates the fusion coefficients at each time step, and the calculation method is as follows: The fusion coefficient value from the previous time step is incremented by an increment determined by both the base handover rate and the system stability metric. The base handover rate controls the basic speed of handover, while the system stability metric reflects the current stability of the system; a larger value indicates greater system instability, a smaller increment, and a slower handover process.
[0121] In emergency scenarios, the system can accelerate the switching process by increasing the base switching rate or reducing the impact of stability metrics. The entire formula ensures that the fusion coefficient never exceeds 1, thus achieving a smooth transition from the old model to the new model.
[0122] The buffer decision generation module is implemented based on a constraint satisfaction optimization framework. During scenario transitions, the system transforms the current load state, resource constraints, and system stability requirements into constraints for an optimization problem. The objective function balances system stability with the transition speed to the optimal strategy for the new scenario.
[0123] Key constraints include: limits on the rate of change of charging pile load, limits on the range of grid pressure fluctuations, and guarantees of user service quality. This module generates transitional decisions that satisfy all constraints in real time using a quadratic programming solver.
[0124] The scene prediction module employs a hybrid architecture combining a Temporal Convolutional Network (TCN) and an attention mechanism. The network input is past data. The scene vector sequence at each time step is used to capture multi-scale temporal dependencies through dilated convolutional layers, while an attention mechanism is used to identify key time points.
[0125] The predicted output is the scene probability distribution and transition time estimate for future time moments. When the transition probability of a certain scene exceeds 70%, the system prepares the model parameters in advance to reduce the loading delay during actual switching.
[0126] It should be understood that, in some embodiments, scene change detection can also be combined with external event information (such as weather forecasts and large-scale event arrangements) to predict scene changes in advance using a hybrid detection method.
[0127] Furthermore, in actual deployments, the threshold for mutation detection and the sensitivity for trend detection can be adjusted according to specific application requirements.
[0128] Step 5.2, Progressive model switching; A progressive model switching mechanism is designed so that, during scenario changes, the decisions of the old and new models are merged according to dynamic weights to avoid abrupt decision changes. The fusion weights change smoothly over time.
[0129] In this process, the system's model decision at the current moment is formed by a weighted combination of the old model decision and the new model decision.
[0130] In this model, the weight coefficients of the old model are the complement of the current smoothing coefficient (i.e., 1 minus the smoothing coefficient), and the weight coefficients of the new model are the current smoothing coefficient. As time progresses, the smoothing coefficient gradually increases from 0 to 1, corresponding to a gradual decrease in the weights of the old model and a gradual increase in the weights of the new model, thus achieving a smooth switching of model decisions.
[0131] The rate of change of the smoothing coefficient will be adaptively adjusted according to the urgency of the scene switch, with faster switching in emergency situations and smoother switching in normal situations.
[0132] In another implementation, the model switching process can adopt a hierarchical and gradual strategy, switching shallow network parameters first and then deep network parameters, or switching non-critical decision components first and then critical decision components, to further enhance the stability of the switching process.
[0133] Step 5.3, Buffer Decision Generation; During scene transitions, the controller considers system stability constraints and generates buffer decisions.
[0134] These decisions ensure the basic operation of the system while reserving resources and time for the complete takeover of new scenarios, thus reducing the risk of switching.
[0135] Step 5.4, Scenario Prediction and Advance Preparation; Utilize time-series forecasting models to analyze scenario evolution trends and predict potential new scenarios. Prepare model parameters and resources in advance for high-probability scenarios to reduce response latency during actual switchovers.
[0136] In this process, the system uses a prediction function based on a recurrent neural network to calculate the conditional probability distribution of the possible scenarios at the next moment, given the current and historical scenario sequences. This prediction function analyzes the patterns and regularities in the scenario time series, learns the dynamic characteristics of scenario changes, and thus can effectively predict future scenarios.
[0137] It should be noted that the prediction function The specific implementation is as follows: First, the historical scene sequence is converted into a fixed-length vector representation through an embedding layer. Then, it is input into a bidirectional LSTM network to capture temporal dependencies. Next, a self-attention mechanism is used to identify key time points and their influencing weights. Finally, after passing through fully connected layers and softmax layers, the output is a vector representing the probability of different scene categories. This structure considers both the long-term dependencies of the scene sequence and the importance of different time points, thereby improving the accuracy of scene prediction.
[0138] It should be noted that, in some embodiments, in addition to the prediction method based on recurrent neural networks described above, Bayesian inference, Markov decision processes, or causal inference-based methods can also be used for scene transition prediction to adapt to different scene change patterns and prediction needs.
[0139] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. A load balancing method for charging piles based on neural networks, characterized in that, Includes the following steps: Scene modeling and representation learning utilizes a multimodal autoencoder network to analyze multi-source heterogeneous data on charging pile status, environment, and user behavior. It then fuses features from various modalities through an attention mechanism to generate a unified scene vector representation. Scene recognition and classification: A variational autoencoder is used to reduce the dimensionality of scene vectors and a density clustering algorithm is applied to identify typical scene categories, thus establishing a scene classification system. Knowledge decomposition and representation: Construct a knowledge decomposition engine to decompose the experience knowledge of charging load balancing into cross-scenario shared knowledge and scenario-specific knowledge, thereby achieving effective separation and organization of knowledge. Adaptive neural network architecture generation automatically generates a neural network architecture that adapts to the current scene based on the similarity between the current scene and known scenes, combined with shared knowledge and scene-specific knowledge. Meta-policy learning and smooth scene transition: The meta-policy network is trained using a meta-reinforcement learning algorithm to generate load balancing decisions, and a smooth transition is achieved through progressive model switching when the scene changes.
2. The charging pile load balancing method based on neural networks according to claim 1, characterized in that, Multimodal autoencoder networks include: Convolutional neural networks used for processing spatially related data; Long Short-Term Memory (LSTM) networks used for processing time-series data; Multilayer perceptrons are used to process static feature data; An attention mechanism for fusing features from different modalities.
3. The charging pile load balancing method based on neural networks according to claim 1, characterized in that, Knowledge decomposition algorithms further include: Apply L1 regularization constraints to the parameters of all scenario-specific knowledge models to make the scenario-specific knowledge models as sparse as possible. Applying structured regularization constraints to shared knowledge models and scenario-specific knowledge models improves the organization and interpretability of knowledge.
4. The charging pile load balancing method based on neural networks according to claim 1, characterized in that, The steps for calculating the similarity between the current scene and known scenes include: Calculate the dot product of the current scene vector and the known scene vector, and then divide by the product of the magnitudes of the two vectors to obtain the similarity value; Select the K scenes with the highest similarity as similar scenes of the current scene.
5. The charging pile load balancing method based on neural networks according to claim 1, characterized in that, The steps for automatically generating a neural network architecture adapted to the current scenario include: Construct a search space that includes various neural network layer types, connection methods, and activation functions; The reinforcement learning controller is used to select the most suitable network structure and hyperparameters for the current scenario. The shared knowledge model and the specific knowledge model of the relevant scenario are fused according to similarity weights to initialize the parameters of the newly generated neural network.
6. The charging pile load balancing method based on neural networks according to claim 1, characterized in that, Meta-reinforcement learning algorithms include: Sample batches of scenes from the scene collection; The meta-policy network is quickly adapted using a small number of samples for each scenario; Use the adapted strategy to generate decision trajectories in each scenario; The policy gradient and value function gradient are calculated based on the collected trajectories; Update the meta-policy network parameters using the optimizer.
7. The charging pile load balancing method based on neural networks according to claim 1, characterized in that, Scene change detection includes: Mutation detection identifies sudden changes by calculating the cosine distance between continuous scene vectors; Trend detection monitors the changing trends of key dimensions of scene vectors using the exponential moving average method.
8. The charging pile load balancing method based on neural networks according to claim 1, characterized in that, Progressive model switching includes: Design a dynamically changing smoothing coefficient to control the fusion ratio of decisions from the old and new models; The smoothing coefficient is adjusted over time to gradually decrease the weights of the old model and gradually increase the weights of the new model, thus achieving a smooth transition.
9. The charging pile load balancing method based on neural networks according to claim 1, characterized in that, It also includes scenario prediction and advance preparation steps: Utilize time-series prediction models to analyze scenario evolution trends and predict potential new scenarios; Prepare model parameters and resources in advance for high-probability scenarios to reduce response latency during actual switching.
10. A charging pile load balancing system based on a neural network, used to execute the charging pile load balancing method based on a neural network as described in any one of claims 1-9, characterized in that, include: The scene modeling and representation learning unit is used to analyze multi-source heterogeneous data on charging pile status, environment, and user behavior using a multimodal autoencoder network, and to fuse features from various modalities through an attention mechanism to generate a unified scene vector representation. The scene recognition and classification unit is used to reduce the dimensionality of scene vectors using a variational autoencoder and apply a density clustering algorithm to identify typical scene categories and establish a scene classification system. The knowledge decomposition and representation unit is used to build a knowledge decomposition engine, which decomposes the experience knowledge of charging load balancing into cross-scenario shared knowledge and scenario-specific knowledge, thereby achieving effective separation and organization of knowledge. The adaptive neural network architecture generation unit automatically generates a neural network architecture that adapts to the current scene based on the similarity between the current scene and known scenes, combined with shared knowledge and scene-specific knowledge. The meta-policy learning and scene smooth transition unit is used to train the meta-policy network to generate load balancing decisions through meta-reinforcement learning algorithms, and to achieve a smooth transition through progressive model switching when the scene changes.
Citation Information
Cited By
Charging station comprehensive operation management method and system based on multi-dimensional user grouping
CN122264473A