Electrode charging variable flow thermal safety control method and system based on AI multi-objective optimization
By constructing a digital twin model and using an AI multi-objective optimization method based on deep reinforcement learning, the problems of lag and accuracy in the thermal safety control of the extreme charging converter were solved, achieving intelligent, precise, and adaptive thermal safety control, and improving the system's efficiency and safety.
Patent Information
- Application Number
- CN202511606115.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2025-12-05
AI Technical Summary
Traditional extreme charge converter thermal safety control methods suffer from problems such as control lag, excessive conservatism, and lack of high-precision prediction capabilities. They are unable to effectively address model mismatch caused by equipment aging and environmental changes, resulting in compromised system efficiency and safety. Existing technologies are unable to effectively solve the technical challenges that existing technologies cannot address.
An AI-based multi-objective optimization-based extreme-charge converter thermal safety control method is adopted. By constructing a digital twin model and deep reinforcement learning (DRL), electrical and thermodynamic state variables are collected in real time, an intelligent agent optimization strategy is constructed, and a safety protection strategy is set to achieve intelligent, precise, and adaptive control of the converter.
It achieves intelligent, precise, and adaptive control of the thermal safety of the extreme charge converter, improving the system's thermal management level, operating efficiency, reliability, and long-term adaptability to equipment aging and environmental changes, ensuring absolute operational safety.
Smart Images

Figure CN121070102A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of extreme charge and variable current thermal safety control technology, and in particular to an extreme charge and variable current thermal safety control method and system based on AI multi-objective optimization. Background Technology
[0002] With the increasing popularity of electric two-wheelers, shortening charging time is a key requirement for improving user experience, leading to the development of "ultra-fast charging" (Ultra-Charging) technology. Ultra-Charging technology aims to replenish a vehicle's battery in a very short time by significantly increasing charging power. The core equipment for achieving Ultra-Charging is a high-power charging pile, and its core power conversion unit is the Ultra-Charging converter. Ultra-Charging converters handle power levels of hundreds of kilowatts or even megawatts during operation. The power devices inside, such as insulated-gate bipolar transistors (IGBTs), generate significant power losses during high-speed switching. These losses accumulate as heat, causing a sharp rise in junction temperature. If this heat cannot be dissipated promptly and effectively, the junction temperature of the power module may exceed its safe limit, leading to material aging, performance degradation, and even permanent thermal breakdown, resulting in equipment damage and safety accidents. Therefore, thermal safety control is a core technological challenge for ensuring the high reliability and long lifespan of Ultra-Charging converters. Traditional thermal management strategies mainly rely on simple control of the heat dissipation system or protection mechanisms based on fixed temperature thresholds, which face severe challenges in terms of control accuracy and dynamic response speed. Against this backdrop, developing smarter and more precise thermal safety control methods is crucial for promoting the commercial application of extreme charging technology.
[0003] Typical extreme charge converter thermal safety control usually relies on simple rules based on fixed thresholds or traditional PID control methods, which suffer from control lag and over-conservatism, making it difficult to achieve coordinated optimization of multiple objectives such as system efficiency while ensuring absolute safety. At the same time, due to the lack of high-precision digital twin models and self-learning capabilities, existing methods cannot accurately predict dynamic changes in electrothermal energy, and are also difficult to adapt to model mismatch caused by equipment aging and environmental changes, resulting in decreased control performance and insufficient overall system efficiency, safety margin, and long-term adaptability.
[0004] To address the shortcomings of the existing technologies, this technical solution proposes an extreme charge converter thermal safety control method and system based on AI multi-objective optimization. Summary of the Invention
[0005] This invention provides an extreme charge converter thermal safety control method and system based on AI multi-objective optimization to address the deficiencies in the prior art.
[0006] On the one hand, this invention provides an extreme charge converter thermal safety control method based on AI multi-objective optimization, including: S1: Real-time acquisition of electrical and thermodynamic state variables of the polarization converter, outputting state-space data; S2: Construct a digital twin model of the extreme-charge converter based on state-space data; S3: Based on the digital twin model, construct a deep reinforcement learning DRL training environment, and train the DRL agent according to the training environment to output the optimal optimization strategy. S4: Set a security protection strategy independent of the DRL agent. When the system's critical parameters exceed the preset security hard threshold, the output action of the DRL agent will be unconditionally overridden. S5: Deploy the optimal optimization strategy to the converter control system, generate control actions based on state space data, and work in conjunction with the safety protection strategy to perform real-time control of the converter and cooling system; S6: Continuously compare the actual operating data with the predicted output of the digital twin model, and calibrate the key parameters of the digital twin model online.
[0007] According to the extreme charge converter thermal safety control method based on AI multi-objective optimization provided by the present invention, step S1, the step of outputting state space data includes: S11: Through the converter controller local area network bus and sensors, electrical status quantities are collected in real time at a sampling frequency of not less than 1Hz; electrical status quantities include DC side voltage and current, AC side three-phase voltage and current, switching frequency and power device drive signals; S12: Thermodynamic state quantities are collected in real time through temperature sensors and heat flow meters at a sampling frequency of not less than 0.5Hz. The thermodynamic state quantities include power module junction temperature, radiator inlet and outlet water temperature, coolant flow rate and cooling fan speed. S13: Preprocess the electrical and thermodynamic state variables and generate state-space data in time series format.
[0008] According to the extreme charge converter thermal safety control method based on AI multi-objective optimization provided by the present invention, step S2, the step of constructing a digital twin model includes: S21: Based on the converter topology, power device physical characteristics and cooling system structure, establish a set of multi-physics field coupled differential equations to describe its electro-thermal dynamic characteristics, which serves as the model kernel of the digital twin model. S22: Divide the state space data into training and validation sets, use the system identification algorithm to identify and fit the unknown parameters in the model kernel, and output the parameterized model kernel. S23: Encapsulate the parameterized model kernel into an executable digital twin to complete the construction of the digital twin model.
[0009] According to the extreme charge converter thermal safety control method based on AI multi-objective optimization provided by the present invention, step S3, the step of outputting the optimal optimization strategy includes: S31: Using a digital twin model as the environment, define the state space, action space, and reward function of the DRL agent; S32: Select proximal policy optimization as the core algorithm of the DRL agent, and initialize the policy network and value network of the DRL agent; S33: In the DRL training environment, the DRL agent and the digital twin model engage in extensive temporal interactions to collect empirical data. The empirical data is then used to iteratively update the policy network and value network, outputting the optimal optimization policy.
[0010] According to the extreme charge converter thermal safety control method based on AI multi-objective optimization provided by the present invention, in step S31, the reward function is expressed as follows: In the formula, R t For the reward function, To predict the junction temperature, A j,max To the upper limit of junction temperature safety, P is the out-of-bounds penalty function. loss P represents the total power loss of the system. loss,rated B represents the system's rated power loss. t Let be the action vector at time t, where t is the time node. Let B be the norm of the change in motion. range Let α be the range of action variation, β be the weight of temperature variation, β be the weight of system power loss, and γ be the weight of the action to be performed, and α + β + γ = 1.
[0011] According to the extreme charge converter thermal safety control method based on AI multi-objective optimization provided by the present invention, step S4, the step of setting the safety protection strategy includes: S41: Set the preset safety hard threshold for key system parameters, including power module junction temperature, radiator inlet and outlet water temperature, DC side voltage, and AC side current. S42: Construct a supervisory controller that is independent of the DRL agent's policy, and the supervisory controller continuously monitors the key parameters of the system; S43: When any critical system parameter exceeds its corresponding preset safety hard threshold, the supervisory controller is immediately triggered and outputs a predefined safety action; the predefined safety actions include frequency reduction, power reduction operation, and forced maximum cooling.
[0012] According to the extreme charge converter thermal safety control method based on AI multi-objective optimization provided by the present invention, step S6, the step of online calibration of key parameters of the digital twin model includes: S61: Continuously compare the actual state quantity data collected during actual operation with the corresponding state quantity data predicted by the digital twin model under the same input, and calculate their root mean square error; S62: Set the deviation tolerance threshold. If the error continues to exceed the preset tolerance for a period of time that reaches the preset time window, the model is determined to be mismatched and the model parameter self-learning algorithm is triggered. S63: The self-learning algorithm uses the recursive least squares method to update and calibrate the predefined calibrable parameter set in the digital twin model online based on the latest collected real-time running data.
[0013] According to the extreme charge converter thermal safety control method based on AI multi-objective optimization provided by the present invention, step S62, the step of determining model mismatch includes: S621: For different state variables predicted by the digital twin model, set their corresponding deviation tolerance thresholds, including: the deviation tolerance threshold for the power module junction temperature is set to ±2°C, the deviation tolerance threshold for the radiator inlet and outlet water temperature is set to ±1°C, the deviation tolerance threshold for the DC side voltage is set to ±5V, and the deviation tolerance threshold for the AC side current is set to ±2A. S622: Set a uniform decision time window with a length of 60 seconds; S623: The triggering condition is defined as the root mean square error between the actual value and the predicted value of any state variable continuously exceeding its corresponding deviation tolerance threshold within a continuous time window. S624: When the triggering condition is met, a model mismatch alarm signal is generated and the model parameter self-learning algorithm is started.
[0014] According to the extreme-charge converter thermal safety control method based on AI multi-objective optimization provided by the present invention, in step S63, the calibrable parameter set includes the on-state resistance of power devices, switching energy loss, heat sink thermal resistance, heat sink thermal capacity, and equivalent heat capacity of coolant; online updating and calibration adopt a recursive least squares method with a forgetting factor, the forgetting factor ranges from 0.95 to 0.99, and the parameters are updated at an update frequency of not less than 0.5Hz.
[0015] This invention also provides an AI-based multi-objective optimization-based extreme charge converter thermal safety control system, comprising: The data acquisition module is used to acquire the electrical and thermodynamic state variables of the polarization converter in real time and output state space data. The electro-thermal digital twin model module is used to construct a digital twin model of the polar-charge converter based on state-space data; The agent optimization decision module is used to build a deep reinforcement learning DRL training environment based on the digital twin model, and to train the DRL agent according to the training environment and output the optimal optimization strategy. The security protection and execution module is used to set a security protection policy independent of the DRL agent. When the system's critical parameters exceed the preset security hard threshold, it unconditionally overrides the output action of the DRL agent. The real-time control module is used to deploy the optimal optimization strategy to the converter control system, generate control actions based on state space data, and work in conjunction with the safety protection strategy to perform real-time control of the converter and cooling system. The model calibration module is used to continuously compare the actual operating data with the predicted output of the digital twin model and to calibrate the key parameters of the digital twin model online.
[0016] The present invention provides an AI-based multi-objective optimization-based thermal safety control method and system for extreme-charge converters. By constructing a digital twin model and combining it with deep reinforcement learning (DRL) for multi-objective optimization decision-making, it achieves intelligent, precise, and adaptive control of the thermal safety of extreme-charge converters. While pursuing optimal system efficiency, the method provides hard protection through an independent safety protection strategy, ensuring absolute operational safety. Furthermore, the online self-calibration function of the model continuously maintains the high fidelity of the digital twin, thereby effectively improving the thermal management level, operating efficiency, reliability, and long-term adaptability to equipment aging and environmental changes of the converter system. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the extreme charge converter thermal safety control method based on AI multi-objective optimization provided in the embodiments of the present invention; Figure 2 This is a schematic diagram of the structure of the extreme charge converter thermal safety control system based on AI multi-objective optimization provided in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0020] Example 1: The following is combined with Figures 1-2 This invention describes an extreme charge converter thermal safety control method and system based on AI multi-objective optimization.
[0021] like Figure 1 As shown, the extreme charge converter thermal safety control method based on AI multi-objective optimization provided in this embodiment of the invention includes: S1: Real-time acquisition of the electrical and thermodynamic state variables of the extreme-charging converter, outputting state-space data. The extreme-charging converter is the core power conversion device for achieving efficient and fast charging, generating a large amount of heat during operation. For precise thermal management and safety control, it is essential to first comprehensively and in real-time acquire all key data reflecting its operating state. Electrical state variables describe the converter's energy conversion operation state, while thermodynamic state variables describe its heat generation and dissipation thermal state. The "state-space data" formed by processing this data constitutes a dataset that comprehensively characterizes the current state of the system, providing input for subsequent modeling, learning, and control.
[0022] The steps for outputting state-space data include: S11: Electrical state parameters are acquired in real time via the converter controller local area network bus and sensors at a sampling frequency of not less than 1Hz. These parameters include DC-side voltage and current, AC-side three-phase voltage and current, switching frequency, and power device drive signals. To ensure data accuracy and reliability, Hall effect sensors with an accuracy of not less than 0.5 should be used for DC and AC-side voltage and current measurements; drive signal acquisition must pass through an isolation barrier to ensure the safety of the control system.
[0023] S12: Thermodynamic state variables are acquired in real time using a temperature sensor and heat flow meter at a sampling frequency of no less than 0.5Hz. These variables include the power module junction temperature, radiator inlet and outlet water temperatures, coolant flow rate, and cooling fan speed. The power module is a power semiconductor device (such as IGBT or SiC MOSFET), and its junction temperature is estimated in real time using a negative temperature coefficient thermistor integrated within the module or by acquiring the on-state voltage drop of the device. The radiator inlet and outlet water temperatures are measured using a Pt100 platinum resistance temperature sensor with an accuracy of ±0.5°C; the coolant flow rate is measured using an electromagnetic or ultrasonic flow meter.
[0024] S13: Preprocess the electrical and thermodynamic state variables by performing timestamp alignment and filtering / denoising, and generate state-space data in time-series format. Timestamp alignment employs hardware triggering or a high-precision software synchronization algorithm to ensure strict time synchronization of data from different sources. Filtering / denoising uses digital filters based on signal characteristics (such as Butterworth low-pass filters) to eliminate high-frequency noise interference. The specific implementation of the high-precision software synchronization algorithm includes the following steps: Configure a network clock synchronization client based on a high-precision network time protocol or a precision time protocol for all data acquisition devices (such as converter controllers, temperature acquisition modules, and flow meter host computers) to synchronize the system clock from the same time server, and control the clock deviation of each acquisition node within milliseconds or even microseconds, providing a unified absolute time reference for subsequent software alignment.
[0025] At each data acquisition node, a local high-precision clock timestamp (called an "arrival timestamp") is added to each arriving data packet or sample value. For sensor data that does not have a timestamp by itself (such as temperature values read through an analog acquisition card), the acquisition card driver adds a timestamp to it the instant the analog-to-digital conversion is completed.
[0026] All timestamped data is sent to a central data processing unit. The data receiving thread stores data from different sources into their respective memory queues.
[0027] The synchronization processing thread uses a uniform, fixed time interval (e.g., a synchronization period of 10ms) as the target time grid. For each target time point t... b Iterate through all data queues, including: for high-speed data such as electrical quantities (sampling rate ≥ 1Hz), find the data whose timestamp is closest to t in its queue. b Two data points (t) c ,y c ) and (t c+1 ,y c+1 ), where t c ≤t b ≤t c+1 The linear interpolation algorithm is used to calculate t. b The estimated value at time t is expressed as: in, For the target time t b The estimated value, t c This is the forward sampling timestamp. This is at the target time t. b The previous closest to t b The time of an actual sampling point. c Let time t c The actual measured value, The sampling time interval is the time difference between two adjacent sampling points, representing the sampling period of the system. This is the difference between the measured values. It represents the change in the measured value within the sampling interval. c to t c+1 The trend (upward or downward) and rate of change during this period. This is the time scaling factor. It is a dimensionless coefficient between 0 and 1. It precisely represents the target time t. b The relative position between two consecutive sampling points. If t b Very close to t c The ratio is close to 0. For example, t b Very close to t c+1 The ratio is close to 1. If t c It's right in the middle, the ratio is 0.5.
[0028] For low-speed data such as thermodynamic quantities (sampling rate 0.5Hz), since the sampling period (2s) is much longer than the synchronization period (e.g., 10ms), a first-order hold strategy is usually adopted, which assumes that the sampled value remains unchanged until the next sampling point arrives. Therefore, the latest timestamp less than or equal to t is directly taken. b The sampled value is used as t b The value at any given moment.
[0029] Each target time point t b The values obtained from all data sources are combined into a complete state vector through interpolation or preservation, and a unified timestamp t is assigned to it. b This new, time-aligned dataset serves as the input for subsequent preprocessing and modeling.
[0030] S2: Construct a digital twin model of the extreme-charging converter based on state-space data. A digital twin model is a high-fidelity dynamic model of a physical entity in information space. It is not merely a simple simulation, but rather a virtual model that synchronously maps and interacts with the real extreme-charging converter by integrating physical laws, engineering knowledge, and real-time operational data. This model can simulate and predict the electrical and thermal behavior inside the converter under different operating conditions in real time, thus providing a safe, efficient, and infinitely repeatable "training ground" and "test bench" for subsequent AI algorithms, avoiding the safety hazards associated with directly exploring high-risk strategies on physical equipment.
[0031] The steps for constructing a digital twin model of an extreme-charge converter include: S21: Based on the converter topology, power device physical characteristics, and cooling system structure, a set of multiphysics coupled differential equations describing its electro-thermal dynamic characteristics is established as the model kernel of the digital twin model. The electrical part adopts an average model or a switching model, with the core being equations based on Kirchhoff's voltage and current laws, used to calculate system power losses.
[0032] The core of electrical modeling is writing equations based on the circuit topology. For the most common two-level three-phase voltage source converter, the modeling process is as follows: Write the equations for each phase loop on the AC side. Taking phase A as an example, its loop equation is: in, This is the output voltage at the midpoint of the A-phase bridge arm of the converter relative to the negative terminal of the DC bus. It is a discrete switching level (calculated by the switching model) or its time-averaged value (calculated by the averaging model). L and R are the inductance and resistance connecting the reactor. The induced voltage on the inductive element. This is the derivative of the A-phase current with respect to time, i.e., the rate of change of the current. The voltage drop across the resistive element is the voltage loss generated when current flows through the internal resistance of the reactor winding and the line resistance. d It is the instantaneous current of phase A, e d It is the electromotive force of phase A of the power grid, and the voltage at the connection point of phase A. It can be regarded as another voltage source applied in the circuit. During grid-connected operation, it is the output voltage of the converter. Overcoming this voltage and the voltage drop across the reactance is necessary to control the magnitude and direction of the current. This equation describes the dynamic relationship between the AC side voltage and current. Equations are then written for the DC bus nodes. The sum of the currents flowing into the nodes of the DC bus capacitor is zero, expressed as: in, It is the DC current flowing into the converter bridge arm. It is the current supplied by a DC source. It is a DC bus capacitor. This is the DC bus voltage. This equation describes the dynamic relationship between the DC side current and voltage.
[0033] Bridge arm output voltage It is determined by the state of the switching device. Define the switching function S. a (When the upper tube is on, S) a =1, S when the lower tube is conducting a If =0), then we have The conduction losses of the devices (IGBTs and diodes) are calculated using their on-state voltage drop, current, and switching function; the switching losses are calculated using the switching energy curve, DC voltage, and current. Summing up the losses of all devices yields the total power loss of the system, which is then used as a heat source input into the thermal model.
[0034] The thermal section employs a thermal equivalent circuit model based on Fourier's law and a thermal resistance-capacity network to simulate the heat transfer process from the junction to the shell, then to the radiator and coolant. The output (losses) of the electrical model serves as the input (heat source) of the thermal model, thus achieving electro-thermal coupling. Fourier's law states that heat flux density is proportional to the temperature gradient. In the equivalent circuit model, this is directly analogous to Ohm's law in circuits: Where ΔA is the temperature difference (analogous to voltage), P loss It is the heat flow rate or power loss (analogous to current), R th Thermal resistance (analogous to electrical resistance) characterizes the way a material impedes the flow of heat.
[0035] Heat capacity C th Characterizing a material's ability to store heat, analogous to a capacitor in a circuit. The rate of temperature change depends on the input heat flux and the heat capacity: Among them, C th For heat capacity, Let be the derivative of temperature with respect to time. A is the temperature rise in the RC circuit. This equation describes the dynamic change of temperature and is the core of the differential equation of the thermal model. A typical "junction-to-case-to-heat sink" thermal network model is an RC network composed of multiple cascaded thermal resistances and thermal capacities.
[0036] Junction-to-Case (JC): Determined by the material between the device chip and the substrate (such as silicon, solder, copper), denoted by R. th,jc and C th,jc describe.
[0037] Case to Heatsink (CS): Determined by thermal grease or thermal pads, using R th,cs and C th,cs describe.
[0038] Radiator to Coolant (SF): Determined by the radiator's geometry, material, and coolant flow rate, denoted by R. th,sf and C th,sf describe.
[0039] Finally, the junction temperature A j The calculation formula can be expressed as: Where ε is a dynamic term, i.e., determined by the heat capacity C. th It is described by the change in voltage (representing temperature) on the surface. A coolant R represents the coolant temperature, the temperature of the most fundamental cooling medium in the heat dissipation system, serving as the starting temperature reference point for the entire heat dissipation path. It is typically measured by a temperature sensor installed at the radiator inlet or outlet.th,jc Junction-to-case thermal resistance (JTR) represents the resistance to heat conduction from the chip's interior (junction) to the device's outer casing (substrate). Its magnitude is primarily determined by the materials and structure of the chip itself, the bonding layers, and the ceramic substrate. This is an important parameter inherent to the device and is usually found in the device datasheet. th,cs R represents the thermal resistance from the device casing to the heatsink. It describes the resistance to heat transfer from the device casing to the heatsink surface. Its magnitude is primarily determined by the thickness, thermal conductivity, and contact area of the thermal grease or pad. Applying thermal grease is intended to reduce this thermal resistance. th,sf This is the thermal resistance from the radiator to the coolant. It represents the resistance to heat transfer from the radiator surface to the coolant. Its magnitude is determined by the radiator's geometry (such as fin structure), material, and coolant flow rate. Generally, the higher the coolant flow rate, the lower the thermal resistance. This represents the steady-state temperature rise. This part calculates the temperature rise under constant power loss P. loss Below, the stable difference between the junction temperature and the coolant temperature. It represents the final temperature at which the cooling system reaches equilibrium. These thermal resistances R... th and heat capacity C th It is a key parameter for online calibration.
[0040] S22: The state-space data is divided into training and validation sets. A system identification algorithm is used to identify and fit the unknown parameters in the model kernel, outputting the parameterized model kernel. The system identification algorithm mainly targets parameters that are difficult to obtain accurately through theoretical calculations. For example, the least squares method is used to identify the thermal resistance and thermal capacity parameters in the thermal network, and an optimization method based on a genetic algorithm is used to fit the switching loss curve parameters of power devices to minimize the error between the model output and the measured data.
[0041] S23: Encapsulate the parameterized model kernel into an executable digital twin to complete the construction of the digital twin model.
[0042] S3: Based on the digital twin model, construct a deep reinforcement learning (DRL) training environment, and train the DRL agent according to the training environment to output the optimal optimization policy. Deep reinforcement learning (DRL) is an important branch of artificial intelligence that enables agents to learn optimal decision-making policies through continuous interaction with the environment. Here, we use the digital twin model constructed in the previous step as the "training environment" for the DRL agent. The agent tries different control actions (such as adjusting fan speed, water pump flow, etc.), observes the state changes and reward signals (i.e., scores evaluating the quality of the actions), and continuously learns through trial and error.
[0043] The steps to output the optimal optimization strategy include: S31: Using a digital twin model as the environment, define the state space, action space, and reward function of the DRL agent. The state space includes: current junction temperature, radiator water temperature, DC side current, AC side current, historical power loss sequence, etc. The action space is defined as continuous or discrete control commands to the cooling system actuators, such as fan speed percentage (0%~100%), water pump flow rate setting (low, medium, high), or pump duty cycle signal.
[0044] In step S31, the reward function is expressed as follows: In the formula, R t Let be the reward function, a scalar value used to evaluate the quality of the action taken by the agent at time t. A negative value indicates that it is a penalty term. To predict junction temperature, this is a key temperature value calculated by the digital twin model based on the current electrical state and cooling conditions; it is a core monitoring indicator for thermal management. j,max The junction temperature safety upper limit is an absolute safety threshold (e.g., 150°C) set according to device specifications and design margins. Once this temperature is exceeded, the device is at risk of permanent damage. Let A be the out-of-bounds penalty function. j0 ≤A j,max When A is within the safe temperature range, this value is 0, indicating no penalty is imposed; when A... j0 >A j,max At that time, this item is This refers to temperature values exceeding the safe range. This generates a significant negative reward (severe penalty), forcing the agent to learn to maintain the temperature below the safe threshold. loss P represents the total power loss of the system, generated by the power devices in the converter during switching and conduction, and signifies the system's operating efficiency. Lower losses indicate higher efficiency. loss,rBted This is the system's rated power loss, typically a baseline value. (B) t This is the action vector at time t, typically a control command for the cooling system, such as the percentage speed of the cooling fan, the flow rate of the water pump, or the duty cycle signal. t is the time node. This is the norm of the change in motion, typically calculated as the Euclidean distance (2-norm) or the absolute difference (1-norm) between the motion vectors at two consecutive moments. This term measures the degree of fluctuation in control commands. B range For the range of motion, for example, fan speed from 0% to 100%, its B rangeIt's 100. α is the weight of temperature change, β is the weight of system power loss, and γ is the weight of the action performed. α + β + γ = 1, typically α = 0.5, β = 0.3, and γ = 0.2. If the agent consistently cools significantly when the temperature is well below the safety threshold, leading to excessive energy consumption, α can be slightly reduced to 0.45 and β increased to, for example, 0.35.
[0045] S32: Proximal policy optimization is selected as the core algorithm for the DRL agent, and the policy network and value network of the DRL agent are initialized. Both the policy network and value network adopt a multilayer perceptron structure. The number of hidden layers and neurons is determined based on the complexity of the state and action space, typically 2-4 hidden layers with 64-256 neurons per layer. The network weights are initialized using the Xavier method. The core idea of Xavier initialization is to maintain the consistency of the signal variance during forward and backward propagation by initializing the weights to a normal or uniform distribution U[-sqrt(6 / (n_in+n_out)), sqrt(6 / (n_in+n_out))] with a mean of zero and a variance of 2 / (n_in+n_out). This method is mainly applicable to approximately linear activation functions such as Tanh and Sigmoid. Here, n_in is the number of neurons in the previous layer, and n_out is the number of neurons in the current layer.
[0046] S33: In the DRL training environment, the DRL agent and the digital twin model engage in extensive temporal interactions to collect empirical data. This empirical data is then used to iteratively update the policy network and value network, outputting the optimal policy. The interaction process is conducted in parallel under multiple randomly initialized conditions to fully explore the state-action space. Empirical data is stored in an experience replay buffer, from which small batches of data are randomly sampled during updates to break the correlation between data and improve training stability. Iterative updates continue until the policy performance converges, i.e., the reward function value stabilizes and reaches a high level.
[0047] S4: An independent safety protection strategy, independent of the DRL agent, is established. When critical system parameters exceed preset safety thresholds, the DRL agent's output actions are unconditionally overridden. Although the DRL agent undergoes extensive training, it is essentially a complex system based on probability and approximation functions. When encountering unprecedented extreme conditions, it may still output unsafe control commands. To ensure absolute safety and reliability, this method sets up an independent and highest-priority "safety protection strategy" (or safety firewall, supervisory controller). It continuously monitors the most critical safety parameters (such as maximum temperature, maximum current, etc.). Once any parameter exceeds the preset absolute safety threshold, this protection strategy immediately intervenes, unconditionally overriding the DRL agent's commands and executing predefined safety operations (such as forced cooling, power reduction, etc.), thus forming an unbreakable safety barrier and achieving a perfect combination of AI intelligent optimization and traditional hard safety logic.
[0048] Step S4, the steps for setting the security protection policy include: S41: Set preset safety hard thresholds for key system parameters, including power module junction temperature, radiator inlet and outlet water temperatures, DC-side voltage, and AC-side current. The preset safety hard thresholds include: Power module junction temperature: Find the maximum junction temperature of the device in the manual (set silicon IGBT to 150℃ or 175℃, SiC MOSFET may be higher, such as 175℃); the threshold must be set below this maximum value, set to 125℃.
[0049] Radiator inlet water temperature: If the inlet water temperature is too low, condensation may form on the radiator surface, causing a short circuit. Therefore, it should be determined according to the type of coolant and the piping design, and set to ≥5℃.
[0050] Radiator outlet water temperature: Excessive water temperature may cause local boiling of coolant or cavitation of pump, which will drastically reduce heat dissipation efficiency. At the same time, the temperature resistance of sealing rings and plastic joints should also be considered, and should be set to ≤65℃.
[0051] DC side voltage: Refer to the manual to determine the rated insulation voltage of the power devices and DC bus capacitors (e.g., the rated DC voltage for a 1200V module is set to 900V). The threshold should be lower than this rated value, and considering the turn-off overshoot caused by line inductance, it should be set to ≤800V.
[0052] AC side current: Refer to the datasheet to determine the maximum repeatable peak current of the power module. The threshold should be sufficient to prevent the current from reaching a level that could damage the device within a very short time (microseconds), and should be set to ≤2×I. rated , where I rated This is the rated output current of the converter.
[0053] S42: Construct a supervisory controller that operates independently of the DRL agent's strategy. This supervisory controller continuously monitors key system parameters. At the hardware level, this supervisory controller can be deployed within the programmable logic controller (PLC) or field-programmable gate array (FPGA) of the converter control system, ensuring its execution is independent of the main computing unit running the DRL algorithm, and possessing the highest operational priority and reliability.
[0054] S43: When any critical system parameter exceeds its corresponding preset safety hard threshold, the supervisory controller immediately triggers and outputs a predefined safety action. Predefined safety actions include frequency reduction, power reduction operation, and forced maximum cooling. The triggering logic is hard-decision, with no delay or filtering, ensuring immediate response. Forced maximum cooling includes setting the cooling fan speed to 100% and the water pump to maximum flow; the power reduction operation command is directly sent to the converter's upper-level controller, requiring it to reduce output power in preset steps until the critical parameter falls back below the safety threshold.
[0055] S5: Deploys the optimal optimization strategy to the converter control system, generates control actions based on state space data, and works in conjunction with the safety protection strategy to perform real-time control of the converter and cooling system. The converter control system reads the latest state space data in real time and instantly calculates the current optimal control action, outputting it to the actuators (such as fans, water pumps, converter switches, etc.). Simultaneously, the safety protection strategy runs in parallel, monitoring the system status in real time. The two work together to form an "AI optimization-led, safety protection as a safety net" operating mode, jointly achieving refined and intelligent real-time control of the converter power output and cooling system, ultimately achieving multiple goals such as improved efficiency, enhanced safety, and extended equipment lifespan. During deployment, the trained DRL intelligent agent strategy network is compiled and embedded into the real-time computing unit of the converter control system. Its control cycle is consistent with or an integer multiple of the state data acquisition cycle. Within each control cycle, the control system reads the latest preprocessed state data, inputs it into the strategy network, and the network performs forward calculations before outputting the recommended action command. Before the action command is finally sent to the actuator, it must pass the verification of the safety protection strategy module. If the security protection policy is not triggered, the DRL action instruction is allowed; if the security protection policy is triggered, its output security action will override the DRL instruction. This coordination mechanism is implemented through hardware logic or high-priority interrupts to ensure absolute security.
[0056] S6: Continuously compare actual operating data with the predicted output of the digital twin model to perform online calibration of key parameters of the digital twin model. Over time, the performance parameters of the extreme-charging converter and its cooling system may slowly change (e.g., due to device aging, scaling), causing the initially constructed digital twin model to gradually exhibit prediction deviations. This method continuously compares the model's predicted values with the actual values measured by sensors, calculating the error between the two in real time. When the error exceeds a certain tolerance, an online calibration process for the model parameters is automatically triggered, using the latest real data to correct and update the parameters within the twin model (e.g., thermal resistance, losses), ensuring that the mathematical model maintains a high degree of consistency with the physical entity. This guarantees the long-term accuracy and reliability of the model-based DRL optimization strategy and the entire control system.
[0057] The steps for online calibration of key parameters of a digital twin model include: S61: Continuously compare the actual state data collected during actual operation with the corresponding state data predicted by the digital twin model under the same input, and calculate their root mean square error. The comparison and error calculation are performed continuously at the same frequency as the data collection. To ensure fair comparison, the model prediction must use the actual state at the previous moment as the initial condition and apply the same control input as the current moment.
[0058] S62: Set the deviation tolerance threshold. If the error continues to exceed the preset tolerance for a period of time that reaches the preset time window, the model is determined to be mismatched and the model parameter self-learning algorithm is triggered.
[0059] In step S62, the steps for determining model mismatch include: S621: For different state variables predicted by the digital twin model, corresponding deviation tolerance thresholds are set, including: ±2°C for the power module junction temperature, ±1°C for the radiator inlet and outlet water temperatures, ±5V for the DC side voltage, and ±2A for the AC side current. These thresholds comprehensively consider sensor measurement errors, inherent model errors, and the system's allowable prediction deviations.
[0060] S622: Set a uniform decision time window with a length of 60 seconds. This time window is used to avoid false triggering caused by momentary interference or noise, ensuring the robustness of the decision.
[0061] S623: The trigger condition is defined as the root mean square error between the actual value and the predicted value of any state variable continuously exceeding its corresponding deviation tolerance threshold within a continuous time window.
[0062] S624: When the triggering condition is met, a model mismatch alarm signal is generated and the model parameter self-learning algorithm is started. The alarm signal will be reported to the system monitoring interface to prompt maintenance personnel to pay attention to the model status.
[0063] S63: The self-learning algorithm, based on the latest acquired real-time running data, uses recursive least squares to update and calibrate a predefined set of calibrable parameters in the digital twin model online. The calibrable parameter set includes the on-state resistance of power devices, switching energy loss, heat sink thermal resistance, heat sink thermal capacity, and equivalent heat capacity of the coolant. Online updating and calibration employs recursive least squares with a forgetting factor, ranging from 0.95 to 0.99, updating parameters at a frequency of at least 0.5 Hz. Recursive least squares recursively updates parameter estimates using new data, while the forgetting factor assigns higher weight to new data, enabling the algorithm to track the slow time-varying characteristics of the parameters. The parameter update process runs in a background thread, and the updated parameters are atomically replaced in the running digital twin model, ensuring the continuous accuracy of the model's predictive capabilities. During calibration, it is necessary to ensure that the data segment used covers different operating points to guarantee the observability and accuracy of parameter identification. The formula for recursive least squares is expressed as: Where y(k) is the observed output, i.e. the actual measured value. It is the regression vector at time k, and T is the transpose of the vector, i.e., the regression vector. Convert from column vector to row vector. θ is the parameter vector to be estimated, such as thermal resistance and heat capacity. e(k) is the observation noise, representing the difference between the actual measured value y(k) and the ideal linear model. The difference between them. k is the time index, representing the current time step or sampling moment.
[0064] The least squares cost function with a forgetting factor is expressed as: Where J(θ) is the cost function, and its value represents the weighted sum of squares of all prediction errors from the beginning to the current time k. i is the summation index, used to traverse all data from the first sampling time to the current k-th sampling time. λ is the forgetting factor, which assigns higher weight to new data, enabling the algorithm to track the slow changes in parameters. The term implies that the further away the data is from the current time k (the smaller i is), the smaller its weight. The squared error term represents the difference between the actual system output y(i) and the model predicted output at time i. The sum of squares is the square of the differences between the observations. Minimizing the sum of squares means finding the model parameters that best fit all the observed data as a whole.
[0065] In summary, the extreme-charge converter thermal safety control method and system based on AI multi-objective optimization provided by this invention achieves intelligent, precise, and adaptive control of extreme-charge converter thermal safety by constructing a digital twin model and combining it with deep reinforcement learning (DRL) for multi-objective optimization decision-making. While pursuing optimal system efficiency, this method provides hard protection through an independent safety protection strategy, ensuring absolute operational safety. Furthermore, the online self-calibration function of the model continuously maintains the high fidelity of the digital twin, thereby effectively improving the thermal management level, operating efficiency, reliability, and long-term adaptability to equipment aging and environmental changes of the converter system.
[0066] like Figure 2 As shown, the present invention also provides an AI-based multi-objective optimization-based extreme charge converter thermal safety control system, which includes: a data acquisition module, an electro-thermal digital twin model module, an intelligent agent optimization decision-making module, a safety protection and execution module, a real-time control module, and a model calibration module.
[0067] The data acquisition module is used to collect the electrical and thermodynamic state variables of the polarization converter in real time and output state-space data.
[0068] The electro-thermal digital twin model module is used to construct a digital twin model of the polarization converter based on state-space data.
[0069] The agent optimization decision module is used to build a deep reinforcement learning DRL training environment based on the digital twin model, and to train the DRL agent according to the training environment and output the optimal optimization strategy.
[0070] The security protection and execution module is used to set a security protection policy independent of the DRL agent. When the system's critical parameters exceed the preset security hard threshold, the output actions of the DRL agent are unconditionally overridden.
[0071] The real-time control module is used to deploy the optimal optimization strategy to the converter control system, generate control actions based on state space data, and work in conjunction with the safety protection strategy to perform real-time control of the converter and cooling system.
[0072] The model calibration module is used to continuously compare the actual operating data with the predicted output of the digital twin model and to calibrate the key parameters of the digital twin model online.
[0073] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A thermal safety control method for extreme charging and discharging of a converter based on AI multi-objective optimization, characterized in that, Comprise: S1: real-time acquisition of electrical state quantity and thermodynamic state quantity of the polar charge converter, output state space data; S2: based on the state space data, a digital twin model of the polar charge converter is constructed; S3: based on the digital twin model, a DRL training environment of deep reinforcement learning is constructed, and a DRL agent is trained according to the training environment, and an optimal optimization strategy is output; S4: set a safety guardian strategy independent of the DRL agent, when the system key parameter exceeds the preset safety hard threshold, the output action of the DRL agent is covered unconditionally; S5: deploy the optimal optimization strategy to the converter control system, generate control action according to the state space data, and work with the safety guardian strategy to realize real-time control of the converter and the cooling system; S6: continuously compare the actual running data with the predicted output of the digital twin model, and calibrate the key parameters of the digital twin model online.
2. The AI-based multi-objective optimization-based thermal safety control method for extreme charging and discharging of a power converter according to claim 1, characterized by, In step S1, the step of outputting state space data comprises: S11: through the converter controller local area network bus and the sensor, the electrical state quantity is collected in real time at a sampling frequency not less than 1Hz; the electrical state quantity includes DC side voltage and current, AC side three-phase voltage and current, switching frequency and power device driving signal; S12: through the temperature sensor and the heat flow meter, the thermodynamic state quantity is collected in real time at a sampling frequency not less than 0.5Hz, the thermodynamic state quantity includes power module junction temperature, radiator inlet and outlet water temperature, cooling liquid flow and radiator fan speed; S13: the electrical state quantity and the thermodynamic state quantity are pretreated, and the state space data in time series format is generated. 3.The AI-based multi-objective optimization-based thermal safety control method of extreme charging and discharging of a variable current, according to claim 1, wherein, In step S2, the step of constructing the digital twin model comprises: S21: based on the converter topology structure, the physical characteristics of the power device and the structure of the cooling system, a multi-physical field coupled differential equation set describing the electrical-thermal dynamic characteristics is established as the model kernel of the digital twin model; S22: the state space data is divided into training set and verification set, the unknown parameters in the model kernel are identified and fitted by using system identification algorithm, and the parameterized model kernel is output; S23: the parameterized model kernel is packaged into an executable digital twin body, and the construction of the digital twin model is completed. 4.The AI-based multi-objective optimization-based thermal safety control method of extreme charging and discharging of a variable current, according to claim 1, wherein, In step S3, the step of outputting the optimal optimization strategy comprises: S31: taking the digital twin model as the environment, defining the state space, action space and reward function of the DRL agent; S32: selecting proximal policy optimization as the core algorithm of the DRL agent, and initializing the policy network and value network of the DRL agent; S33: in the DRL training environment, the DRL agent and the digital twin model are interacted in time sequence, experience data is collected, the policy network and value network are updated iteratively by using the experience data, and the optimal optimization strategy is output.
5. The AI-based multi-objective optimization-based thermal safety control method of extreme charging and discharging of the electric vehicle of claim 4, wherein, In step S31, the formula of the reward function is: where R t is the reward function, is the predicted junction temperature, A j,max is the junction temperature safety upper limit, is the out-of-bound penalty function, P loss is the total system power loss, P loss,rated is the system rated power loss, B t is the action vector at time t, t is the time node, is the norm of the action change, B range is the action change range, a is the weight of temperature change, b is the weight of system power loss, g is the weight of action execution, a+b+g=1. 6.The AI-based multi-objective optimization-based thermal safety control method of extreme charging and discharging of a battery according to claim 1, wherein, In step S4, the step of setting the safety guardian strategy comprises: S41: set a preset safety hard threshold of a system key parameter, the system key parameter including a power module junction temperature, a radiator inlet and outlet water temperature, a direct current side voltage, and an alternating current side current; S42: construct a supervisory controller independent of a policy of the DRL agent, the supervisory controller continuously monitoring the system key parameter; S43: when any of the system key parameters exceeds the preset safety hard threshold corresponding thereto, the supervisory controller immediately triggers and outputs a predefined safety action; the predefined safety action includes frequency reduction, power reduction operation, and forced maximum cooling.
7. The AI-based multi-objective optimization-based thermal safety control method of extreme charging and discharging of the electric vehicle of claim 6, wherein, In step S6, the step of online calibration of the key parameters of the digital twin model includes: S61: continuously compare real state quantity data collected in actual operation with corresponding state quantity data predicted by the digital twin model under the same input, and calculate the root mean square error thereof; S62: set a deviation tolerance threshold, if the error continuously exceeds the preset tolerance for a time length reaching a preset time window, determine that the model is mismatched, and trigger a model parameter self-learning algorithm; S63: the self-learning algorithm updates and calibrates a pre-defined calibratable parameter set in the digital twin model based on a latest collected real-time operation data by using a recursive least square method.
8. The AI-based multi-objective optimization-based thermal safety control method of extreme charging and discharging of the electric vehicle of claim 7, wherein, In step S62, the step of determining that the model is mismatched includes: S621: for different state quantities predicted by the digital twin model, set a corresponding deviation tolerance threshold therefor, including: the deviation tolerance threshold of the power module junction temperature is set to ±2°C, the deviation tolerance threshold of the radiator inlet and outlet water temperature is set to ±1°C, the deviation tolerance threshold of the direct current side voltage is set to ±5V, and the deviation tolerance threshold of the alternating current side current is set to ±2A; S622: set a unified determination time window, the length of the time window being 60 seconds; S623: define a trigger condition as the root mean square error between the actual value and the predicted value of any state quantity continuously exceeding the corresponding deviation tolerance threshold in the continuous time window; S624: when the trigger condition is met, generate a model mismatch alarm signal and start the model parameter self-learning algorithm. 9.The AI-based multi-objective optimization-based thermal safety control method of extreme charging and discharging of a variable current, according to claim 7, wherein, In step S63, the calibratable parameter set includes a power device on-state resistance, a switching energy loss, a radiator thermal resistance, a radiator thermal capacity, and a cooling liquid equivalent heat capacity; the online updating and calibration uses a recursive least square method with a forgetting factor, the value of the forgetting factor being in a range of 0.95-0.99, and the parameter updating being performed at an updating frequency of not less than 0.5 Hz.
10. An AI multi-objective optimization based thermal safety control system for extreme fast reactor, which adopts the AI multi-objective optimization based thermal safety control method for extreme fast reactor according to any one of claims 1 to 9, characterized in that, The system includes: a data acquisition module, configured to acquire electrical state quantities and thermodynamic state quantities of the polar charging converter in real time, and output state space data; an electrical-thermal digital twin model module, configured to construct a digital twin model of the polar charging converter based on the state space data; an agent optimization decision module, configured to construct a DRL training environment of deep reinforcement learning based on the digital twin model, and train a DRL agent according to the training environment, and output an optimal optimization policy; a safety guardian and execution module, configured to set a safety guardian strategy independent of the DRL agent, and to unconditionally override the output action of the DRL agent when a system key parameter exceeds a preset safety hard threshold; a real-time control module, configured to deploy the optimal optimization strategy to a converter control system, to generate a control action according to the state space data, and to work with the safety guardian strategy to perform real-time control on the converter and the cooling system; a model calibration module, configured to continuously compare actual operation data with the predicted output of the digital twin model, and to perform online calibration on key parameters of the digital twin model.
Citation Information
Patent Citations
Adaptive maintaining method and system for time-varying consistency of digital twin components of power transformation equipment
CN120012595A
Method for preventing main steam pressure fluctuation based on control of fire coal amount
CN120332748A
Intelligent heat supply regulation and control system based on digital twinning and deep reinforcement learning
CN120627184A
Unicycle multi-agent dynamic coverage control method based on security reinforcement learning
CN120780023A
Multi-objective optimization capacitor bank intelligent exchange decision-making system
CN120824737A
Cited By
Chip test equipment temperature active regulation and control system and method based on multi-mode temperature sensing
CN121349223A
Mobile charging vehicle intelligent heat dissipation method and system based on two-phase immersion liquid cooling
CN121536193A
Intelligent heat dissipation method and system based on two-phase immersion liquid cooling mobile charging vehicle
CN121536193B