A battery charging optimization method and system based on edge-cloud collaboration
Patent Information
- Application Number
- CN202610216554.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-14
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-02-14
AI Technical Summary
但直接将强化学习应用于真实电池进行在线训练风险高、成本大;而在仿真环境中训练则严重依赖于模型的准确性,即“模拟器与实物的差距”
本发明解决了模型精度与计算资源的矛盾,通过“云侧高保真模型+边侧轻量化模型”的协同架构,将高精度建模与复杂优化计算放在云端,将需快速响应的验证与执行放在边端,实现了全局最优与局部实时的高效统一。
Smart Images

Figure CN122159423B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of battery charging optimization technology, specifically relating to a battery charging optimization method and system based on edge-cloud collaboration. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] As the core power source for electric vehicles and portable electronic devices, the optimization of the charging process of lithium-ion batteries is crucial. An excellent charging strategy needs to minimize charging time while ensuring safety, and simultaneously slow down battery aging and extend its cycle life. However, the complex electrochemical characteristics of batteries make their charging process highly nonlinear, and the optimal charging current trajectory is dynamically influenced by multiple factors such as the current state of the battery, the environment, historical operating conditions, and individual differences. This makes finding a globally optimal or near-optimal charging strategy extremely challenging.
[0004] Traditional battery charging methods can be broadly categorized into two types: mechanistic model-based methods and data-driven methods. Model-based optimization methods, such as using electrochemical models for optimal control, can theoretically generate high-performance charging curves. However, these methods face an inherent contradiction between "model accuracy" and "computational complexity." While high-fidelity mechanistic models can accurately simulate the internal state of the battery, their massive computational demands make real-time operation difficult on the resource-constrained battery management system. Simplified models, on the other hand, neglect key physical processes or nonlinear dynamics, leading to optimization results that may deviate from the optimal solution in practical applications, potentially even causing safety issues. Furthermore, battery model parameters drift with aging; if the model cannot be updated online, its long-term applicability will be significantly reduced.
[0005] Data-driven approaches, particularly reinforcement learning, have shown promise in optimizing complex systems. However, directly applying reinforcement learning to real batteries for online training is risky and costly; training in simulation environments heavily relies on model accuracy, i.e., the "difference between simulators and real devices." Furthermore, the complex policy networks trained also face computational bottlenecks in edge deployment. While existing cloud-based optimization solutions can handle complex calculations, uploading large amounts of raw data to the cloud introduces communication latency and bandwidth pressure, and struggles to meet the real-time requirements of policy execution during charging. Summary of the Invention
[0006] To address the aforementioned problems, this invention proposes a battery charging optimization method and system based on edge-cloud collaboration, which can balance accuracy, computational efficiency, real-time performance, and adaptability.
[0007] According to some embodiments, the present invention adopts the following technical solution: A battery charging optimization method based on edge-cloud collaboration includes the following steps: A digital twin model integrating multi-physics coupling and unknown nonlinear dynamic compensation is constructed on the cloud side. The multi-physics coupling part is used to establish a basic white-box model based on the battery's electro-thermal-aging coupling mechanism. The unknown nonlinear dynamic compensation part is used to introduce a black-box model to learn the unknown nonlinear dynamics in the residuals of the basic white-box model. On the edge side, a lightweight digital twin model is obtained through knowledge distillation based on the aforementioned digital twin model; Using the digital twin model as a simulation environment, the optimal charging strategy is generated through reinforcement learning for policy training. The optimal charging strategy is then deployed to the edge and verified using the lightweight digital twin model. The verified and feasible charging strategy will be sent to the end-side battery management system for execution. The charging strategy is executed on the device side and battery operation data is collected. The battery operation data is then fed back to the cloud to drive the dynamic update and strategy iteration of the digital twin model.
[0008] As an alternative implementation method, the process of constructing a digital twin model integrating multi-physics coupling and unknown nonlinear dynamic compensation on the cloud side includes: establishing a basic white-box model based on the electrochemical-thermal-aging multi-physics coupling mechanism, considering lithium-ion concentration distribution, electrode reaction kinetics, heat conduction, and capacity decay mechanism; simultaneously, introducing a deep neural network to construct a black-box compensation model, which adopts a long short-term memory network structure to learn the unknown nonlinear dynamics in the residuals of the basic model; and obtaining the digital twin model by weighted fusion of the mechanism output of the white-box model and the compensation output of the black-box compensation model.
[0009] As a further defined implementation, the basic white-box model includes a coupled second-order equivalent circuit model, a coupled thermal model, a coupled aging model, and a coupled ampere-hour integral model. The electrical model adopts a second-order equivalent circuit, and its voltage output is affected by temperature. The calculation of the thermal model depends on the heat generation power of the electrical model. The decay rate of the aging model is driven by the current and voltage of the electrical model and the temperature of the thermal model. Each model is bidirectionally coupled through state variables and parameters.
[0010] As a further defined implementation, the black-box compensation model adopts a long short-period memory neural network model, whose input includes the model input and the internal state of the system. Through a gating mechanism, it captures the dynamic memory effect and historical dependence of the battery, and outputs residual compensation terms for voltage, temperature and SOC.
[0011] As an alternative implementation, after constructing a digital twin model that integrates multi-physics coupling and unknown nonlinear dynamic compensation on the cloud side, the method further includes: using a dual-timescale extended Kalman filter algorithm to perform online correction of the model. Specifically, this is achieved through two coordinated estimation loops, one of which is a fast-changing loop used to quickly track the dynamic state of SOC and voltage; and the other is a slow-changing loop used to identify and correct physical parameters that change slowly with aging.
[0012] As an alternative implementation, after constructing a digital twin model that integrates multi-physics coupling and unknown nonlinear dynamic compensation on the cloud side, the method also includes iteratively optimizing the weight parameters of the unknown nonlinear dynamic compensation part using a stochastic gradient descent algorithm. The training process is based on the battery operation data uploaded from the edge side, and the data is divided into multiple batches for training. A weighted loss function is defined to simultaneously consider the prediction errors of voltage, temperature and SOC. By minimizing the final loss function, the optimized network structure and parameters are obtained.
[0013] As an alternative implementation, the process of obtaining a lightweight digital twin model by knowledge distillation based on the digital twin model on the edge side includes: inheriting the structure and parameters of the electrochemical-thermal-aging multiphysics physical model in the cloud-side model, performing structured pruning on the unknown nonlinear dynamic compensation part in the model, and removing redundant neuron connections and hidden layer units. A teacher-student distillation framework is constructed, with the cloud-side network of the complete unknown nonlinear dynamic compensation part as the teacher model and the edge-side network of the lightweight unknown nonlinear dynamic compensation part as the student model. A temperature regulation mechanism is used to soften the output distribution of the teacher model. An alternating training strategy is adopted, with fixed teacher model parameters and only updated student model parameters. The stochastic gradient descent algorithm is used to optimize the student model. The student model is trained until its performance on the validation set meets the set requirements. The lightweight unknown nonlinear dynamic compensation part obtained by distillation is integrated with the distributed physical model parameters on the edge side to form an edge-side lightweight digital twin model.
[0014] As an alternative implementation, the process of generating the optimal charging strategy through reinforcement learning using the digital twin model as the simulation environment includes: modeling the battery charging process as a Markov decision process, defining its state space as the battery's state of charge, terminal voltage, current, and surface temperature; the action space as the charging current command; and the state transition probability being dynamically described by the digital twin model. Design a comprehensive reward function, which is used to guide the agent to learn the optimal strategy that balances charging speed, aging reduction and temperature rise suppression; The policy is trained using a deep deterministic policy gradient algorithm until the predetermined requirements are met. The trained policy network parameters are then saved, thus obtaining the optimal charging policy.
[0015] As a further defined implementation, during the policy training process using the deep deterministic policy gradient algorithm, the policy network is responsible for outputting the determined charging current action based on the current state; while the value network is responsible for evaluating the long-term expected cumulative reward of the state-action pair. Both the policy network and the value network use deep neural networks. The policy network takes a state vector as its input, which is processed by several fully connected layers and the ReLU activation function. Finally, it outputs a normalized charging current command through a Tanh activation function. The value network takes the state vector and action vector as joint inputs and outputs a scalar Q value after network calculation. During training, the agent interacts with the digital twin simulation environment, collecting experience data and storing it in an experience replay buffer. Batch data is randomly sampled from the buffer to update the network parameters; the goal of updating the value network is to minimize the temporal difference error. A soft update mechanism is introduced to update the target network parameters in order to improve training stability.
[0016] A battery charging optimization system based on edge-cloud collaboration includes: A cloud-based device is used to construct a digital twin model that integrates multi-physics coupling and unknown nonlinear dynamic compensation. The multi-physics coupling part is used to establish a basic white-box model based on the battery's electro-thermal-aging coupling mechanism. The unknown nonlinear dynamic compensation part is used to introduce a black-box model to learn the unknown nonlinear dynamics in the residuals of the basic white-box model. Using the digital twin model as a simulation environment, the optimal charging strategy is generated through reinforcement learning. The optimal charging strategy is then distributed to the edge side and verified using the lightweight digital twin model. Finally, the verified feasible charging strategy is distributed to the edge battery management system for execution. An edge-side device is used to obtain a lightweight digital twin model through knowledge distillation based on the digital twin model. The battery management system, located on the edge, is used to execute the charging strategy and collect battery operation data, and feed the battery operation data back to the cloud to drive the dynamic update and strategy iteration of the digital twin model.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention resolves the contradiction between model accuracy and computing resources. Through a collaborative architecture of "cloud-side high-fidelity model + edge-side lightweight model", high-precision modeling and complex optimization calculations are placed in the cloud, while verification and execution that require rapid response are placed at the edge, achieving an efficient unification of global optimization and local real-time performance.
[0018] This invention enhances the sophistication and security of the strategy. It utilizes a high-fidelity model in the cloud for risk-free reinforcement learning training, and the strategy is rapidly verified by the edge model before execution, providing "double insurance" for safety and avoiding the risk of unsafe strategies directly affecting the physical battery. It also achieves adaptive optimization throughout the entire life cycle, driving continuous updates of the cloud model and iterative evolution of the strategy through edge data feedback. This enables the system to track the dynamic aging characteristics of the battery and always maintain the optimality and applicability of the charging strategy.
[0019] This invention constructs an efficient collaborative computing paradigm. By utilizing an edge-cloud collaborative architecture, it rationally allocates computing and communication loads, which not only overcomes the bottleneck of insufficient computing power on the pure edge side, but also avoids the problems of high communication latency and poor real-time performance in pure cloud solutions.
[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0022] Figure 1 This is a flowchart illustrating a battery charging optimization method based on edge-cloud collaboration provided in one embodiment. Figure 2 This is a schematic diagram of the complete process and architecture of a battery charging optimization method based on edge-cloud collaboration in one embodiment; Figure 3 This is a comparison curve of the output and the actual value of a cloud-side high-fidelity digital twin model and an edge-side lightweight digital twin model in one embodiment; Figure 4 This is a convergence curve and optimized current graph of the reinforcement learning training process in one embodiment; Figure 5 This is a performance comparison chart of one embodiment and a traditional charging method; Figure 6 This is a diagram of an edge-cloud collaborative battery charging optimization system architecture provided in one embodiment. Detailed Implementation
[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0024] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0025] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0026] Where there is no conflict, the embodiments and features described in this application may be combined with each other.
[0027] Example 1 A battery charging optimization method based on edge-cloud collaboration, such as Figure 1 As shown, it includes the following steps: Step S101: Construct a high-fidelity digital twin model of the battery on the cloud side. Specifically, a basic white-box model of the battery is first established based on the electrochemical-thermal-aging multiphysics coupling mechanism. Then, a black-box model based on a deep neural network is introduced to compensate for unknown nonlinear dynamic characteristics not considered in the basic model. Through a fusion of data-driven and mechanism-driven approaches, a high-precision gray-box model is constructed.
[0028] Step S102: Based on the high-fidelity model in the cloud, a lightweight digital twin model is generated for the edge computing nodes using knowledge distillation technology. Specifically, the high-fidelity model is used as the teacher model, while a student model with 70%-80% fewer parameters is designed. Knowledge from the teacher model is transferred to the student model through output distillation, feature distillation, and relation distillation. After training, the lightweight model is deployed to edge computing nodes.
[0029] Step S103: Using the high-fidelity battery digital twin model as the simulation environment, a charging strategy is trained using a reinforcement learning algorithm to generate a globally optimal charging strategy. Specifically, the battery charging process is modeled as a Markov decision process, where the state space includes the battery's state of charge, terminal voltage, current, and surface temperature; the action space is the charging current command. A deep deterministic policy gradient algorithm is used for large-scale, risk-free training in the digital twin simulation environment, generating the optimal charging current curve through the policy network.
[0030] Step S104: The globally optimal charging strategy is distributed to the edge side, and rapid verification and security assessment are performed using a lightweight digital twin model on the edge side. In specific implementation, the charging strategy distributed from the cloud is received at the edge computing node, and it is input into the lightweight model for millisecond-level rapid simulation. Key parameters such as voltage, current, and temperature are monitored in real time during the simulation process. The strategy is marked as verified as successful when all parameters are within the safe range, and is ready to be distributed to the edge side for execution.
[0031] Step S105 involves sending the verified charging strategy to the edge side for execution and collecting battery operation data in real time. In specific implementation, the edge-side battery management system receives the charging current curve sent from the edge side, controls the charging process through the battery charging and discharging equipment, and simultaneously collects battery voltage, current, and surface temperature data.
[0032] Step S106 involves feeding back the battery operation data collected at the device side to the cloud to drive model updates and policy iterations. Specifically, preprocessed operation data is uploaded to the cloud via a 5G communication network. When the voltage prediction error of the high-fidelity digital twin model continuously exceeds 3%, a model parameter update mechanism is triggered. Simultaneously, a user-defined policy generation threshold is set; when this threshold is exceeded, the system automatically initiates a new round of reinforcement learning policy training to generate a new generation of charging strategies adapted to the current battery state.
[0033] By cyclically executing the above steps, a closed-loop optimization mechanism for edge-cloud collaboration is established. An integrated architecture centered on "cloud decision-making - edge verification - edge execution" is constructed: the cloud is responsible for high-precision modeling and policy generation, the edge handles policy verification and real-time monitoring, and the edge focuses on policy execution and data acquisition. Through this three-level collaboration, optimized allocation of computing resources is achieved. The cloud processes model updates and policy training on an hourly basis, the edge completes policy verification on a second-level basis, and the edge achieves policy execution and data acquisition on a millisecond-level basis.
[0034] The specific structure of the battery optimization charging algorithm based on edge-cloud collaboration in this embodiment is as follows: Figure 2 As shown, this includes a battery digital twin modeling method based on edge-cloud collaboration and a battery charging optimization method combining digital twins and reinforcement learning. The figure systematically illustrates the entire closed-loop optimization process from high-fidelity modeling and decision-making on the cloud side, rapid verification on the edge side, to execution feedback on the edge side: On the cloud side, a high-fidelity digital twin model integrating the electrochemical-thermal-aging multi-physics coupling mechanism and deep neural network compensation is first constructed. Then, using this model as a simulation environment, the optimal charging strategy is generated through reinforcement learning training, and simultaneously, a lightweight edge model is obtained through knowledge distillation. The strategy and lightweight model generated in the cloud are distributed to the edge side, where the lightweight model performs rapid simulation and safety verification of the charging strategy. The verified strategy is finally distributed to the edge side for execution by the battery management system. Simultaneously, real-time operational data collected on the edge side is fed back to the cloud, driving parameter updates and structural self-correction of the high-fidelity model, thus forming a continuously optimized integrated closed loop of "perception-decision-verification-execution".
[0035] The battery digital twin modeling method based on edge-cloud collaboration specifically includes: In the cloud, a basic white-box model is established based on the electrochemical-thermal-aging multiphysics coupling mechanism, considering lithium-ion concentration distribution, electrode reaction kinetics, heat conduction, and capacity decay mechanisms; simultaneously, a deep neural network is introduced to construct a black-box compensation model, which adopts a long short-term memory network structure to learn the unknown nonlinear dynamics in the residuals of the basic model; by weighted fusion of the mechanism output of the white-box model and the compensation output of the black-box model, a high-fidelity gray-box digital twin model is constructed, the mathematical model of which can be expressed as: (1) in, For the multivariate output of the digital twin model, This represents the output of the multiphysics coupling model. This represents the compensation part of the unknown nonlinear neural network, where Input to the model, This refers to the internal state of the system. For physical parameter vectors, These are the weight parameters of the neural network.
[0036] The multiphysics coupling model is established based on the battery's electro-thermal-aging state coupling mechanism. Specifically, it includes a coupled second-order equivalent circuit model, a coupled thermal model, a coupled aging model, and a coupled ampere-hour integral model. The electrical model uses a second-order equivalent circuit, and its voltage output is affected by temperature. The calculation of the thermal model depends on the heat generation power of the electrical model. The decay rate of the aging model is driven by the current and voltage of the electrical model and the temperature of the thermal model. Through bidirectional coupling of state variables and parameters, each model achieves high-fidelity synchronous simulation of the battery's external characteristics and internal state.
[0037] A dual-time-scale extended Kalman filter algorithm is employed for online model calibration. This algorithm comprises two cooperative estimation loops: a fast-changing loop responsible for rapidly tracking dynamic states such as SOC and voltage; and a slow-changing loop responsible for identifying and correcting physical parameters that change slowly with aging, such as ohmic resistance and thermal resistance. This mechanism effectively solves the problem of time-varying drift of model parameters, ensuring the prediction accuracy of the digital twin throughout its entire lifecycle.
[0038] The neural network compensation The network architecture employs Long Short Period Memory (LSTM), and the network input includes both the model input and the system's internal state. (2) The network input at the previous moment is represented as: The external input of the model at the previous time step is represented as: The model's internal input at the previous time step is represented as: This 8-dimensional input vector comprehensively covers the electrical, thermal, and state characteristics of the battery, enabling LSTM to effectively capture the battery's dynamic memory effect and historical dependence through its internal control mechanism, thereby accurately outputting residual compensation terms for voltage, temperature, and SOC.
[0039] Specifically, these compensation values are added to the prediction values of the multiphysics coupling model to form the final output of the digital twin model, as shown in Equation 1.
[0040] For network training, a stochastic gradient descent algorithm is used to iteratively optimize the weight parameters of the LSTM neural network. The training process is based on battery operation data uploaded from the device, and the data is divided into multiple mini-batches for efficient training. A weighted loss function is defined to simultaneously consider the prediction errors of voltage, temperature, and SOC. By minimizing the above loss function, the optimized network structure and parameters are obtained.
[0041] The construction of a lightweight digital twin model on the edge side inherits the electrochemical-thermal-aging multiphysics physical model structure and parameters from the high-fidelity model on the cloud side, ensuring that the edge-side model follows the physical laws of battery operation. Simultaneously, the complete LSTM network compensation module on the cloud side is structurally pruned, removing redundant neuron connections and hidden layer units to obtain a lightweight LSTM student model while maintaining time-series prediction capabilities. Specifically, this includes: By using width pruning and depth pruning, a method can be obtained with a depth range from... Layer reduction to The number of layers and nodes per layer from Reduce to The lightweight student model structure is significantly less complex than the cloud-based teacher model. This structure fundamentally reduces computational complexity and memory consumption. Knowledge distillation training is then performed based on this structure.
[0042] A teacher-student distillation framework is constructed, using a cloud-side complete LSTM network as the teacher model and an edge-side lightweight LSTM network as the student model. A temperature regulation mechanism is employed to soften the output distribution of the teacher model. Let the outputs of the teacher model and the student model be respectively... and Temperature parameters are The distribution of soft targets is as follows: (3) An alternating training strategy is employed, fixing the teacher model parameters and updating only the student model parameters. Stochastic gradient descent is used to optimize the student model. After training until the student model's performance on the validation set reaches over 95% of the teacher model's, the distilled lightweight LSTM network is integrated with the distributed physical model parameters at the edge, forming a lightweight digital twin model on the edge side. Figure 3As shown in the figure, this comparison illustrates the relationship between the predicted outputs and actual measured values of the two models for key state parameters of battery voltage, temperature, and SOC under the same input conditions. The comparison curves clearly demonstrate that the cloud-side high-fidelity model has higher prediction accuracy, while the edge-side lightweight model achieves a significant improvement in computational efficiency while maintaining relative accuracy.
[0043] In a preferred embodiment, the effectiveness of the method is verified using a 26650 lithium-ion battery. A high-fidelity digital twin model integrating the electrochemical-thermal-aging multiphysics coupling mechanism with LSTM neural network compensation is constructed, and a lightweight edge-side model is obtained based on this model through knowledge distillation. Figure 4 As shown, the convergence curve plot illustrates the trend of reward value changes with training rounds during the training process, reflecting the complete learning process of the algorithm from exploration to convergence; the optimized current plot shows the optimal charging current curve obtained through reinforcement learning training, which achieves the optimal balance between charging speed and battery life while ensuring battery safety. The two plots together intuitively demonstrate the effectiveness and superiority of the proposed method in charging strategy optimization.
[0044] It can be seen that, under the same test conditions, the voltage, temperature, and SOC prediction outputs of both the high-fidelity model and the lightweight edge-side model maintained good consistency with the actual measured values of the battery. Specifically, the high-fidelity model's average absolute error for voltage prediction was less than 6.3mV, its SOC prediction error was less than 0.5%, and its temperature prediction error was less than 0.3℃. The lightweight model obtained through knowledge distillation, while maintaining similar accuracy, reduced the number of parameters and improved inference speed, fully verifying that the proposed method significantly improves the computational efficiency of the edge side while ensuring model accuracy.
[0045] Using a high-fidelity battery digital twin model as the simulation environment, the optimal charging strategy is generated through reinforcement learning training, specifically including: The battery charging process is modeled as a Markov decision process. Its state space is defined as the battery's state of charge, terminal voltage, current, and surface temperature; the action space is the charging current command; and the state transition probabilities are dynamically described by the high-fidelity digital twin model.
[0046] Design a comprehensive reward function to guide the agent in learning an optimal strategy that balances charging speed, aging mitigation, and temperature rise suppression. The reward function design is as follows: (4) in, This represents the increment of the state of charge (SOC) per unit time. This incentive motivates the reinforcement learning agent to seek faster charging strategies. The measured voltage of the end-side battery management system; The optimal reference voltage, calculated by a high-fidelity digital twin model in the cloud, satisfies the internal electrochemical safety boundaries (such as preventing lithium plating) under the current state and current. This penalty is designed to prevent any charging behavior that leads to excessive battery polarization, and is central to ensuring a healthy battery lifespan. To measure the surface temperature of the battery, This is the preset maximum safe temperature threshold. This feature serves as a safety redundancy, ensuring that the battery operates within an absolutely safe temperature range.
[0047] The Deep Deterministic Policy Gradient (DDPG) algorithm is used for policy training. DDPG, an Actor-Critic algorithm applicable to continuous action spaces, consists of a policy network (Actor) and a value network (Critic). The policy network is responsible for determining the charging current action based on the current state; the value network is responsible for evaluating the long-term expected cumulative reward of the state-action pair.
[0048] Both the policy network and the value network employ deep neural networks. The policy network takes a state vector as input, which is processed through several fully connected layers and the ReLU activation function, ultimately outputting a normalized charging current command via a Tanh activation function. The value network takes the state vector and action vector as joint input, and calculates and outputs a scalar Q-value.
[0049] During training, the agent interacts with a high-fidelity digital twin simulation environment, collecting experience data and storing it in an experience replay buffer. Subsequently, small batches of data are randomly sampled from the buffer to update the network parameters. The update objective of the value network is to minimize the temporal difference error, and its loss function is defined as: (5) in, It is the value network parameter, θ - These are the target network parameters. It is a state-action value estimation. It's an instant reward. It is a discount factor. It is the policy action for the next state. This loss function provides an accurate gradient direction for policy optimization by minimizing the difference between the current value estimate and the target value.
[0050] A soft update mechanism is introduced to update the target network parameters in order to improve training stability: (6) in, This is the soft update coefficient, which is usually taken as a small value close to 1.
[0051] After sufficient training, when the policy performance converges and the reward value stabilizes at a high level, the trained policy network parameters are saved, thus obtaining the optimal charging policy. This policy can adaptively output the optimal charging current based on the real-time state of the battery.
[0052] The step of distributing the optimal charging strategy to the edge side and using a lightweight digital twin model on the edge side for rapid verification specifically includes: The cloud sends the optimal charging strategy, i.e., the charging current curve sequence, generated during training, to the edge computing nodes. After receiving the strategy, the edge nodes load it locally.
[0053] On the edge side, a lightweight policy execution simulation environment is built. The core of this environment is a lightweight digital twin model deployed on the edge side, which takes the charging policy to be verified as the action command input model.
[0054] Initiate rapid simulation and extrapolation. The lightweight model extrapolates forward in millisecond-level time steps according to policy instructions, predicting the evolution trajectory of key states such as battery voltage, current, temperature, and SOC within a future charging cycle, for example, from the current SOC to the target SOC.
[0055] Safety, effectiveness, and aging impact assessments are conducted based on the simulation results. A series of safety thresholds are set, such as upper limits for voltage, current, and temperature. Throughout the entire simulation time series, all state variables are monitored in real time to ensure they remain within the safety threshold ranges. Simultaneously, the effectiveness of the strategy is evaluated.
[0056] Generate verification conclusions. If all state parameters remain within limits and performance indicators meet expectations during the simulation, the charging strategy is marked as "verification passed." If any parameter exceeds limits or performance fails to meet standards, it is marked as "verification failed." The edge side only distributes verified charging strategies to the end-side battery management system for execution.
[0057] The validated and feasible strategies will be further distributed to the edge battery management system for execution; the charging strategy will be executed on the edge and battery operation data will be collected, and the data will be fed back to the cloud to drive the dynamic update and strategy iteration of the digital twin model, specifically including: The edge-side battery management system receives validated charging strategies from the edge side. This strategy is typically represented as a charging current profile or an executable strategy network.
[0058] The battery management system analyzes the charging strategy and precisely executes charging current commands by controlling the battery charging and discharging equipment to achieve optimized charging control of the battery.
[0059] While executing the charging strategy, the device uses built-in sensors to collect real-time battery operating data. The collected data includes voltage, current, and surface temperature. Data is collected at a fixed high sampling frequency.
[0060] The edge device preprocesses and transmits the collected high-frequency raw data: the collected voltage, current and temperature raw data are preprocessed by filtering, noise reduction and timestamp alignment; then, the processed data is transmitted in two paths through the communication module: one path is uploaded to the cloud database for permanent storage, used for long-term training and iterative optimization of the high-fidelity digital twin model; the other path is sent to the edge device in real time, used for the instant state update of the lightweight digital twin model and the initial state setting for the next charging strategy verification.
[0061] After receiving real-time operational data from the edge, the cloud drives the following two core iterative processes: Dynamic updates of high-fidelity digital twin models: Using newly acquired data, the physical parameters and neural network weights in the high-fidelity model are dynamically calibrated through online learning algorithms, so that the model always keeps up with the actual aging state and characteristic changes of the battery.
[0062] Iterative optimization of reinforcement learning strategies: When a certain amount of new data is accumulated, or when a significant decrease in model accuracy due to battery aging is detected, or when battery state parameters exceed a set threshold, a new round of reinforcement learning training is triggered. In the new training cycle, the agent will learn in an updated and more accurate digital twin model environment, thereby generating a new generation of optimal charging strategies that are more adapted to the current health state of the battery.
[0063] Through the cyclical execution of the above process, a complete integrated closed-loop optimization system of "perception-decision-verification-execution" is constructed. This system enables the charging strategy to continuously evolve as the battery ages, always providing a personalized, adaptive, safe, and efficient optimal charging solution.
[0064] This embodiment experimentally verifies the superiority of the proposed method. For example... Figure 4 As shown, the reward value tends to converge after approximately 300 iterations in the reinforcement learning training process, indicating that the policy has become stable and optimal. The optimal charging current curve obtained from the training exhibits a complex shape, and compared with the traditional constant current and constant voltage method, it significantly improves charging efficiency while ensuring safety.
[0065] Figure 5This paper presents the charging experiment results of the optimized charging method proposed in this application under the faster, balanced, and more conservative modes, as well as the performance comparison results with the constant current charging method with average current under the three modes on the same battery. Experiments show that in the test of charging the battery from 0% SOC to 80% SOC, the method of this application shortens the charging time by approximately 1.33%, reduces the peak temperature by approximately 24.07%, and reduces the energy consumption per unit SOC by approximately 6.35%. This fully demonstrates that the method provided by this invention can effectively extend battery life and ensure thermal safety while achieving fast charging.
[0066] Example 2 A battery charging optimization system based on edge-cloud collaboration, such as Figure 6 As shown in the figure, the three-level collaborative architecture of cloud-side processing center, edge-side service node and terminal-side execution unit is clearly displayed to implement the above method, as well as the bidirectional data flow of "model / policy distribution" and "data feedback upload", which reflects the closed-loop optimization mechanism of the system.
[0067] The system includes: The cloud-side processing center is configured to: construct and maintain the high-fidelity battery digital twin model that integrates multi-physics coupling and unknown nonlinear dynamic compensation; dynamically update the parameters and perform online self-correction of the network structure of the high-fidelity model based on battery operation data fed back from the edge side; run reinforcement learning algorithms in the high-fidelity model as a simulation environment to generate the optimal charging strategy; and generate a lightweight edge-side digital twin model based on the high-fidelity model using knowledge distillation technology.
[0068] The edge service node is configured to: receive and store the edge lightweight digital twin model and the optimal charging strategy from the cloud-side processing center; use the edge lightweight digital twin model to perform rapid simulation verification and security assessment on the received charging strategy; and distribute the verified feasible charging strategy to the edge.
[0069] The edge execution unit, embedded in the battery management system, is configured to: receive and execute charging strategies from the edge service node; collect real-time operating data of the battery during the charging process; preprocess the collected data; and upload the preprocessed battery operating data to the cloud processing center.
[0070] The system achieves continuous optimization and self-evolution of battery charging strategies through the collaborative cooperation and data closed loop of the terminal, edge, and cloud levels.
[0071] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of one or more computer-usable storage media (including, but not limited to, disk storage, etc.) containing computer-usable program code. CD - ROM It takes the form of a computer program product implemented on (such as optical memory, etc.).
[0072] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0073] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0074] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0075] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made by those skilled in the art without creative effort within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A battery charging optimization method based on edge-cloud collaboration, characterized in that, Includes the following steps: A digital twin model integrating multi-physics coupling and unknown nonlinear dynamic compensation is constructed on the cloud side. The multi-physics coupling part is used to establish a basic white-box model based on the battery's electro-thermal-aging coupling mechanism. The unknown nonlinear dynamic compensation part is used to introduce a black-box model to learn the unknown nonlinear dynamics in the residuals of the basic white-box model. On the edge side, a lightweight digital twin model is obtained through knowledge distillation based on the aforementioned digital twin model; The process of constructing a digital twin model on the cloud side that integrates multi-physics coupling and unknown nonlinear dynamic compensation includes: establishing a basic white-box model based on the electrochemical-thermal-aging multi-physics coupling mechanism, considering lithium-ion concentration distribution, electrode reaction kinetics, heat conduction, and capacity decay mechanism; simultaneously, introducing a deep neural network to construct a black-box compensation model, which adopts a long short-term memory network structure to learn the unknown nonlinear dynamics in the residuals of the basic model; and obtaining the digital twin model by weighted fusion of the mechanism output of the white-box model and the compensation output of the black-box compensation model. The black-box compensation model adopts a long short-term memory network model. Its input includes the model input and the internal state of the system. Through a gating mechanism, it captures the dynamic memory effect and historical dependence of the battery and outputs residual compensation terms for voltage, temperature and SOC. Using the digital twin model as a simulation environment, the optimal charging strategy is generated through reinforcement learning for policy training. Specifically, this includes: modeling the battery charging process as a Markov decision process, defining its state space as the battery's state of charge, terminal voltage, current, and surface temperature; the action space as the charging current command; and the state transition probability being dynamically described by the digital twin model. Design a comprehensive reward function, which is used to guide the agent to learn the optimal strategy that balances charging speed, aging reduction and temperature rise suppression; The policy is trained using a deep deterministic policy gradient algorithm until the predetermined requirements are met. The trained policy network parameters are then saved, thus obtaining the optimal charging policy. The optimal charging strategy is then deployed to the edge and verified using the lightweight digital twin model. The verified and feasible charging strategy will be sent to the end-side battery management system for execution. The charging strategy is executed on the device side and battery operation data is collected. The battery operation data is then fed back to the cloud to drive the dynamic update and strategy iteration of the digital twin model.
2. The battery charging optimization method based on edge-cloud collaboration as described in claim 1, characterized in that, The basic white-box model includes a coupled second-order equivalent circuit model, a coupled thermal model, a coupled aging model, and a coupled ampere-hour integral model. The electrical model adopts a second-order equivalent circuit, and its voltage output is affected by temperature. The calculation of the thermal model depends on the heat generation power of the electrical model. The decay rate of the aging model is driven by the current and voltage of the electrical model and the temperature of the thermal model. Each model is bidirectionally coupled through state variables and parameters.
3. The battery charging optimization method based on edge-cloud collaboration as described in claim 1, characterized in that, After constructing a digital twin model on the cloud side that integrates multi-physics coupling and unknown nonlinear dynamic compensation, the method also includes: using a dual-timescale extended Kalman filter algorithm to perform online correction of the model. Specifically, this is achieved through two coordinated estimation loops: a fast-changing loop for quickly tracking the dynamic state of SOC and voltage; and a slow-changing loop for identifying and correcting physical parameters that change slowly with aging.
4. The battery charging optimization method based on edge-cloud collaboration as described in claim 1, characterized in that, After constructing a digital twin model that integrates multi-physics coupling and unknown nonlinear dynamic compensation on the cloud side, the training process also includes iteratively optimizing the weight parameters of the unknown nonlinear dynamic compensation part using the stochastic gradient descent algorithm. The training process is based on the battery operation data uploaded from the edge side, and the data is divided into multiple batches for training. A weighted loss function is defined to simultaneously consider the prediction errors of voltage, temperature and SOC. By minimizing the final loss function, the optimized network structure and parameters are obtained.
5. The battery charging optimization method based on edge-cloud collaboration as described in claim 1, characterized in that, The process of obtaining a lightweight digital twin model by knowledge distillation based on the digital twin model on the edge side includes: inheriting the structure and parameters of the electrochemical-thermal-aging multiphysics physical model in the cloud side model, performing structured pruning on the unknown nonlinear dynamic compensation part in the model, and removing redundant neuron connections and hidden layer units. A teacher-student distillation framework is constructed, with the cloud-side network of the complete unknown nonlinear dynamic compensation part as the teacher model and the edge-side network of the lightweight unknown nonlinear dynamic compensation part as the student model. A temperature regulation mechanism is used to soften the output distribution of the teacher model. An alternating training strategy is adopted, with fixed teacher model parameters and only updated student model parameters. The stochastic gradient descent algorithm is used to optimize the student model. The student model is trained until its performance on the validation set meets the set requirements. The lightweight unknown nonlinear dynamic compensation part obtained by distillation is integrated with the distributed physical model parameters on the edge side to form an edge-side lightweight digital twin model.
6. The battery charging optimization method based on edge-cloud collaboration as described in claim 1, characterized in that, During policy training using the deep deterministic policy gradient algorithm, the policy network is responsible for outputting the determined charging current action based on the current state; the value network is responsible for evaluating the long-term expected cumulative reward of the state-action pair. Both the policy network and the value network use deep neural networks. The policy network takes a state vector as its input, which is processed by several fully connected layers and the ReLU activation function. Finally, it outputs a normalized charging current command through a Tanh activation function. The value network takes the state vector and action vector as joint inputs and outputs a scalar Q value after network calculation. During training, the agent interacts with the digital twin simulation environment, collects experience data and stores it in the experience replay buffer; batch data is randomly sampled from the buffer to update the network parameters, and the update objective of the value network is to minimize the temporal difference error. A soft update mechanism is introduced to update the target network parameters in order to improve training stability.
7. A battery charging optimization system based on edge-cloud collaboration, employing a battery charging optimization method based on edge-cloud collaboration as described in any one of claims 1-6, characterized in that, include: A cloud-based device is used to construct a digital twin model that integrates multi-physics coupling and unknown nonlinear dynamic compensation. The multi-physics coupling part is used to establish a basic white-box model based on the battery's electro-thermal-aging coupling mechanism. The unknown nonlinear dynamic compensation part is used to introduce a black-box model to learn the unknown nonlinear dynamics in the residuals of the basic white-box model. Using the digital twin model as a simulation environment, the optimal charging strategy is generated through reinforcement learning. The optimal charging strategy is then distributed to the edge side and verified using a lightweight digital twin model. Finally, the verified feasible charging strategy is distributed to the edge battery management system for execution. An edge-side device is used to obtain a lightweight digital twin model through knowledge distillation based on the digital twin model. The battery management system, located on the edge, is used to execute the charging strategy and collect battery operation data, and feed the battery operation data back to the cloud to drive the dynamic update and strategy iteration of the digital twin model.
Citation Information
Patent Citations
Management method of power battery module based on cloud control technology
CN110492186A
Distributed photovoltaic energy intelligent group dispatching and group control system based on machine learning
CN118554625A