Electric drive assembly thermal failure risk reinforcement learning prediction and active collaborative inhibition method

By constructing a dynamic graph convolutional neural network model and a distributed control algorithm, the problem of predicting and suppressing the risk of over-temperature failure of the electric drive assembly was solved, realizing the thermal safety and energy efficiency optimization of the electric drive assembly and improving the safety and energy efficiency of the system.

CN122334050APending Publication Date: 2026-07-03TONGJI UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202610804242.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies cannot effectively predict the risk of overheating failure of electric drive assemblies, and the cooling system is slow to adjust, making it difficult to achieve rapid cooling and energy efficiency optimization, and cannot systematically consider the risk of global thermal failure.

Method used

A dynamic graph convolutional neural network model integrating long short-term memory network and threshold recursive unit is constructed. Temperature rise prediction is performed by combining sensor information, a global risk assessment framework is established, and a distributed model prediction control algorithm based on deep deterministic policy gradient algorithm and alternating direction multiplier method is used to achieve hierarchical decision-making and rolling time-domain optimization, so as to synergistically suppress the risk of thermal failure.

Benefits of technology

It achieves multi-objective collaborative optimization of thermal safety and energy efficiency of electric drive assembly, can predict over-temperature risks in advance, finely coordinate the actions of execution units, avoid control timing conflicts, and improve system safety and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122334050A_ABST
    Figure CN122334050A_ABST
Patent Text Reader

Abstract

This invention relates to a reinforcement learning-based prediction and active collaborative suppression method for thermal failure risk in electric drive assemblies. The method includes: constructing a dynamic graph convolutional neural network model for long-term and short-term temperature rise prediction, and combining sparse sensor information to estimate the temperature of key components at different future times; constructing a risk assessment framework of different levels based on global dynamic risk index weights using the entropy weight method; and establishing a hierarchical active collaborative suppression algorithm for reinforcement learning-based prediction of thermal failure risk in electric drive assemblies. Combining the temperature rise prediction and risk assessment results, the upper layer uses a deep deterministic strategy gradient algorithm to predict and generate thermal safety boundaries, while the lower layer applies a distributed model predictive control algorithm based on the alternating direction multiplier method to finely coordinate the actions of each execution unit. The upper and lower layer strategies achieve hierarchical decision-making and rolling time-domain optimization through time-domain decoupling. This invention has the advantages of strong thermal failure risk management capabilities, high environmental adaptability, and the ability to simultaneously achieve optimal energy efficiency and thermal safety under long-term operating conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of thermal failure prevention and control technology for automotive electric drive assemblies, specifically to a reinforcement learning-based active collaborative suppression method for predicting thermal failure risks in electric drive assemblies. Background Technology

[0002] As the integration of electric drive systems gradually increases, the contradiction between internal temperature rise and heat dissipation becomes more prominent. Especially under extreme operating conditions, the electric drive system experiences rapid temperature rise and sudden over-temperature events, leading to an increased risk of thermal failure. To address this potential over-temperature failure risk, the thermal management system of the electric drive system is often equipped with a large cooling capacity, resulting in high energy consumption. Therefore, developing a proactive method to suppress the thermal failure risk of electric drive systems that balances optimal energy efficiency and thermal safety is of great significance for ensuring the efficient and safe operation of electric drive systems.

[0003] In the prior art: CN111900833B discloses a power motor thermal management control method based on integrated vehicle control. When an abnormal motor temperature signal is detected, the method controls the speed of the motor cooling fan and cooling water pump by referring to the engine temperature, speed, gear, and motor output torque to reasonably adjust the cooling intensity and achieve energy saving and consumption reduction. CN119535989B discloses an adaptive control method for a multi-heat source thermal management system based on reinforcement learning model predictive control. Based on the temperature error output fed back by the system, the method optimizes the control strategy of the cooling system for devices such as fans, water pumps, and valves in real time by fusing strategy optimization of intelligent agents and traditional model predictive control, so that the system maintains efficient temperature control in dynamic changes. CN116315290B discloses a thermal management system and its thermal management method, thermal management device, and storage medium. By acquiring the mass flow rate of the coolant in the coolant circulation loop and the temperature difference between the output and input sides of the coolant in the heat exchange zone, the method calculates the heat difference between the target heat and the rated heat, and adjusts the coolant temperature on the input side of the heat exchange zone and the mass flow rate of the coolant in the coolant circulation loop.

[0004] The aforementioned existing technologies control the cooling system to regulate the system temperature by monitoring whether the temperature signal of the motor or the object being cooled is abnormal and calculating the difference between the target and rated heat of the thermal management system. However, all of these technologies rely on physical temperature sensors for monitoring, making it impossible to predict overheating and overheating failure risks in advance. When an overheating risk occurs suddenly, effective control is difficult to achieve. Furthermore, all technologies rely on the cooling system for regulation and do not directly suppress the power loss of the object being cooled, making it difficult to achieve rapid cooling. The heat dissipation potential still needs to be explored.

[0005] Furthermore, CN120160333A discloses a predictive control method for refrigeration units based on Bayesian reinforcement learning. This method constructs a Bayesian regularized neural network model, filters aligned multimodal data, generates corrected instantaneous cooling capacity prediction data, and dynamically adjusts state information using a deep deterministic gradient algorithm based on the prediction results, applying threshold constraints to temperature and pressure. CN113050426B discloses a thermal management control method and system incorporating a genetic ant colony algorithm. This method utilizes a support vector machine prediction model based on an improved genetic ant colony algorithm to determine the fan duty cycle matching the current operating conditions based on real-time sensing parameters such as ambient temperature, motor speed, motor torque, and motor outlet water temperature. This controls the electric fan speed, ensuring that the motor's heat dissipation requirements are met while minimizing fan energy consumption.

[0006] While the aforementioned existing technologies can predict temperature rise and implement corresponding cooling measures, they all rely on the cooling system for regulation or require driver intervention, failing to proactively suppress power loss in the object being cooled. Furthermore, they do not consider the overall over-temperature risk level, thus failing to achieve precise and efficient heat suppression.

[0007] CN119846490A discloses a method and related apparatus for assessing battery thermal runaway risk. It constructs state change curves using temperature, current, and voltage parameters to predict and assess thermal runaway risk, obtaining the assessment results. CN117454716A discloses a transformer overheating risk assessment method based on temperature rise simulation and fuzzy hierarchical analysis. It calculates and classifies parameters that significantly influence overheating risk based on numerical simulation of the transformer's internal thermal characteristics and operating load status, selects parameters with significant impact on overheating risk, and establishes an assessment model. The assessment level is evaluated based on the calculation results. Then, the weight coefficients of the corresponding parameters are obtained through fuzzy set theory and hierarchical analysis, and the FAHP expert scoring system is used to comprehensively score the overheating risk.

[0008] The aforementioned prior art assesses the risk level of overheating failure by analyzing the state parameters of batteries and transformers, but it cannot predict the overheating state in advance, nor can it assess the overall risk level at different times in the future, nor can it take targeted risk mitigation measures according to different overheating risk levels.

[0009] CN112363061A discloses a thermal runaway risk assessment method based on big data. This method predicts the current temperature rise rate by establishing a functional model relating various state parameters of a power battery, the average temperature value, and the maximum temperature rise value. The prediction rate is then compared with the actual temperature rise rate, and the prediction error is compared with two set thresholds to determine the thermal runaway risk level. While it can predict over-temperature conditions based on relevant parameters, it only considers the temperature rise rate, resulting in a single evaluation indicator. Furthermore, it cannot accurately pinpoint the location of over-temperature in real time and does not implement targeted risk mitigation measures based on different over-temperature risk levels.

[0010] The aforementioned existing technologies all assess thermal failure risks independently or address heat dissipation in localized areas, without systematically considering global thermal failure risks. Furthermore, when addressing over-temperature failure risks, none of these studies have conducted thermal failure risk assessment, temperature prediction, or thermal failure suppression analysis specifically for the electric drive assembly described in this invention. Moreover, due to physical limitations of actuators such as drive motors, water pumps, and oil pumps, and significant differences in their dynamic response speeds, cooling enhancement control commands in the aforementioned thermal management strategies cannot take effect immediately and exhibit control timing conflicts, leading to the continuous accumulation of thermal failure risks. In addition, the significant differences in the thermal characteristics of various components result in uneven temperature distribution in different areas, increasing the difficulty of coordinating thermal failure risk suppression. Therefore, constructing a dynamic real-time thermal failure risk assessment framework, considering thermal failure risk characteristics and operating conditions, and designing an active thermal failure risk suppression strategy that can balance the lowest short-term thermal failure risk with optimal long-term energy efficiency is a major research direction for the efficient and reliable control of electric drive assemblies. Summary of the Invention

[0011] The purpose of this invention is to overcome the shortcomings of the prior art by providing a reinforcement learning method for predicting and actively co-suppressing the thermal failure risk of electric drive systems that simultaneously achieves optimal energy efficiency and thermal safety.

[0012] The objective of this invention can be achieved through the following technical solutions: A reinforcement learning-based method for predicting and actively co-suppressing the thermal failure risk of electric drive assemblies is proposed to achieve multi-objective synergistic optimization of thermal safety, operational safety, and energy efficiency of electric drive assemblies. The method includes the following steps: Step 1: Construct a dynamic graph convolutional neural network temperature long-term prediction model and a short-term prediction model that respectively integrate Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU), and combine sensor information to estimate the temperature rise of key components at different future times; Step 2: Using the entropy weight method, construct risk assessment frameworks for different levels based on the weight distribution of global dynamic risk indicators; Step 3: Establish a hierarchical active collaborative suppression algorithm for predicting thermal failure risks in the electric drive assembly using reinforcement learning. Combining temperature rise prediction and risk assessment results, the upper layer uses a deep deterministic strategy gradient algorithm to predict and generate thermal safety boundaries, while the lower layer applies a distributed model predictive control algorithm based on the alternating direction multiplier method to finely coordinate the actions of each execution unit. Step 4: The upper and lower layer strategies achieve hierarchical decision-making and rolling time-domain collaborative optimization through time-domain decoupling, realizing multi-objective collaborative optimization of electric drive assembly thermal safety, operational safety and energy efficiency.

[0013] Further, step 1 specifically includes: Step 1-1: Collect sensor data, including motor speed, torque, oil temperature, and inverter cooling water temperature, based on historical sparse sensor data from the past 5 time steps. m t0 The maximum likelihood estimation method was used to identify the temperature of other key components (including the permanent magnet of the motor, windings, iron core, inverter IGBT (Insulated Gate Bipolar Transistor), assembly housing, on-board charger (OBC), and DC-DC converter) within five time steps at the temperature nodes. The maximum likelihood estimation model is shown below: in, express t Based on partial temperature sensor data m t0 Other estimated temperature node data, These are the model parameters for the maximum likelihood estimation model. Simultaneously, the estimation results of the maximum likelihood model are compared with sparse sensor historical data from the past five time steps. m t0 Find the union of the sets. t The node information for each time point is shown below: in, express t All node information at any given time; Steps 1-2: Construct dynamic graph convolutional neural network temperature long-term prediction model and short-term prediction model that respectively integrate LSTM and GRU. Based on the input information, predict the temperature rise of key components within the next 120 (long-term) or 5 (short-term) time steps.

[0014] Furthermore, step 2 specifically includes: Step 2-1: Determine the maximum temperature value based on the short-term or long-term temperature rise prediction results. m Max Number of overheating nodes n over Maximum temperature rise rate v T_max ; Step 2-2: Based on the characteristics of each component of the electric drive assembly, define dynamic risk indicators for thermal failure, including over-temperature thresholds. m thr The proportion of overheated nodes n over / n all With temperature rise rate threshold v T_thr ; Steps 2-3: Compare the highest temperature value, the percentage of overheated nodes, and the maximum temperature rise rate with the overheating threshold, the overheated node percentage threshold, and the temperature rise rate threshold, respectively, to determine whether they exceed the thresholds. Based on the comparison results, obtain scores for the three indicators. Simultaneously, using the entropy weight method, a high-medium-low risk level assessment framework is constructed based on the global dynamic risk indicator weight distribution. This comprehensively assesses the global thermal failure risk level at each time step in the future time domain, and predicts in real time the potential impact range and safety risks of overheating in key components.

[0015] Furthermore, step 3 specifically includes: Step 3-1: Based on the electric drive powertrain power domain control architecture, establish an active thermal failure risk reinforcement learning prediction hierarchical collaborative suppression algorithm, which includes a global optimization layer (upper layer) and a distributed execution layer (lower layer). Step 3-2: The global optimization layer (upper layer) handles thermal failure suppression control, integrating reinforcement learning and robust model predictive control. A Deep Deterministic Policy Gradient (DDPG) algorithm is used to dynamically correct the cooling and energy efficiency optimization boundaries of each subsystem. This predicts and generates thermal safety boundaries (including the maximum allowable rate of temperature rise and the maximum allowable torque derating), thermal equilibrium weights, and desired temperatures, which are then provided to the evaluation network and the long-term temperature prediction model.

[0016] The evaluation network is based on the prediction results of the long-term temperature prediction model and the actual component temperature fed back from the lower layer. Combined with the long-term thermal failure risk assessment level determined in step 2, a reward function is constructed to support the optimization of the value function of the evaluation network, so as to achieve effective evaluation of the output action of the action network.

[0017] The final determined optimal thermal safety boundary, thermal equilibrium weight, and expected temperature are encoded as global consistency constraints and provided to the lower distributed execution layer. Step 3-3: The distributed execution layer (lower layer) applies a distributed model predictive control algorithm based on the Alternating Direction Method of Multipliers (ADMM). Based on the thermal safety boundary input from the upper layer and the temperature predicted by the short-time temperature prediction model, the short-term thermal failure risk assessment level is determined in conjunction with Step 2. With the optimization objectives of minimizing over-temperature risk and energy consumption, rolling optimization is achieved by continuously updating the optimal control law. This enables real-time and precise coordination of the local control actions of each execution unit in the thermal management submodule and the motor control submodule within the short-time domain, and feeds back the actual temperature rise data and actuator status to the upper-level risk assessment module.

[0018] Furthermore, step 4 specifically includes: 1) Sampling period definition: Define the upper-layer sampling period as... T H The sampling period of the lower layer is T L , usually satisfy ,in, M 1 represents the time scale scaling factor; 2) Upper-layer thermal safety boundary update: Upper layer every The global thermal safety boundary is updated periodically. , ; 3) Lower-level scrolling optimization: The lower level executes within the same cycle as the upper level. M One short-term rolling optimization was performed to adjust the control quantities of the thermal management submodule and the motor control submodule.

[0019] The upper-level global optimization layer uses a longer sampling period and is responsible for long-term thermal failure risk prediction and thermal safety boundary generation. The lower-level distributed execution layer uses a shorter sampling period and is responsible for short-term over-temperature suppression and actuator coordination. The upper and lower layers achieve time-domain decoupling through a multi-sampling-period mechanism of "low-frequency decision-making at the upper level and high-frequency execution at the lower level," enabling global thermal safety boundary planning and local actuator rapid response to operate collaboratively at different time scales. Reinforcement learning does not directly control the fast actuators but outputs safety boundaries, improving system safety. ADMM improves system real-time performance and scalability by decomposing the coupled optimization problem into local problems.

[0020] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention constructs a dynamic graph convolutional neural network temperature long-term and short-term prediction model that integrates GRU and LSTM respectively. It can effectively estimate the temperature rise of key components at different times in the future by combining a small amount of sensor information, and provide effective input for thermal failure risk assessment and suppression. (2) This invention can systematically consider the global thermal failure risk characteristics and operating conditions of the electric drive assembly. The upper layer dynamically corrects the cooling and energy efficiency optimization boundaries of each subsystem through global optimization. It adopts targeted suppression strategies for the short-term and long-term risk levels at different times and proposes an active suppression method for thermal failure risk that can take into account both the lowest short-term thermal failure risk and the best long-term energy efficiency. (3) This invention fully considers the physical limitations of each actuator and the difference in dynamic response speed of each actuator. The lower layer uses a distributed model predictive control algorithm to finely coordinate the actions of each execution unit, and feedforward compensates for the error caused by the time lag effect of the upper layer strategy. This avoids the lag of cooling enhancement control command and power derating command and control timing conflict, so as to better cope with the risk of sudden thermal failure. (4) The upper and lower layer strategies proposed in this invention realize hierarchical decision-making and rolling time-domain optimization through time-domain decoupling, which can achieve multi-objective collaborative optimization of electric drive assembly thermal safety, operation safety and energy efficiency. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating the reinforcement learning prediction and active collaborative suppression method for the thermal failure risk of the electric drive assembly in an embodiment of the present invention. Figure 2 This is a schematic diagram of a dynamic graph convolutional neural network temperature short-time prediction model fused with LSTM in an embodiment of the present invention; Figure 3 This is a schematic diagram of a long-term temperature prediction model using a dynamic graph convolutional neural network incorporating GRU in an embodiment of the present invention. Figure 4 This is a schematic diagram of the dynamic thermal failure risk level determination process at different time steps in an embodiment of the present invention; Figure 5 This is a schematic diagram of the hierarchical active collaborative inhibition architecture for thermal failure risk reinforcement learning prediction in an embodiment of the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.

[0023] like Figure 1 As shown, the reinforcement learning prediction and active collaborative suppression method for thermal failure risk of electric drive assembly consists of four parts: Temperature information identification of key components in the electric drive assembly; Prediction of short- and long-term temperature rise of key components in electric drive assembly; Dynamic thermal failure risk level assessment of electric drive assembly; Reinforcement learning prediction and hierarchical collaborative inhibition of thermal failure risk in electric drive assembly.

[0024] The method of the present invention includes the following steps: Step 1: Construct a dynamic graph convolutional neural network temperature prediction model that integrates LSTM and GRU respectively, and combine it with sensor information to estimate the temperature rise of key components at different times in the future; Step 2: Construct risk assessment frameworks of different levels based on the weight distribution of local key components and global dynamic risk indicators using the entropy weight method; Step 3: Establish a reinforcement learning prediction and active hierarchical collaborative suppression algorithm for the thermal failure risk of the electric drive assembly. Combining the temperature rise prediction and risk assessment results, the upper layer uses a deep deterministic strategy gradient algorithm to predict and generate the thermal safety boundary, while the lower layer applies a distributed model predictive control algorithm based on the alternating direction multiplier method to finely coordinate the actions of each execution unit; Step 4: The upper and lower layer strategies achieve hierarchical decision-making and rolling time-domain optimization through time-domain decoupling, realizing multi-objective collaborative optimization of electric drive system thermal safety, operational safety and energy efficiency.

[0025] Step 1 specifically involves: Step 1-1: Based on the sparse sensing history data from the past 5 time steps m t0 (Including motor speed, torque, oil temperature, and inverter cooling water temperature), the maximum likelihood estimation method is used to identify the temperature nodes of other key components (including motor permanent magnets, windings, iron core, inverter IGBTs, assembly housing, on-board charger (OBC), DC-DC converter, etc.) within 5 time steps. .

[0026] The maximum likelihood estimation model is shown below: ; in, express t Based on partial temperature sensor data m t0 Other estimated temperature node data, These are the model parameters for the maximum likelihood estimation model. Simultaneously, the estimation results of the maximum likelihood model are compared with sparse sensor historical data from the past five time steps. m t0 Find the union of the sets. t The node information for each time point is shown below: in, express t Information on all nodes at each time step (the maximum likelihood estimation model method is an existing technology).

[0027] Steps 1-2: Construct dynamic graph convolutional neural network temperature prediction models that integrate LSTM or GRU respectively, and predict the temperature rise of key components within the next 120 or 5 time steps based on the input information.

[0028] Steps 1-2 are as follows: Step 1-2-1: Complete the input information after identification using the maximum likelihood estimation method. A dynamic graph convolutional neural network is used to learn the spatial correlation of each temperature node. The constructed dynamic graph convolutional neural network model is shown below: in, E 1 represents a learnable, randomly initialized dictionary of node embeddings. E 1∈R N×de , E Each line of 1 represents the embedding of a node. N Indicates the number of nodes. de This indicates the dimension of the node embedding. de The value range is [8, 64]. E 1 E 1 T Used to infer the spatial dependencies between each pair of nodes during training. E 1. It will automatically update to learn the hidden dependencies between nodes at different temperatures, thus obtaining the adjacency relationships. When adaptively embedding the adjacency matrix... A a When acquiring new adjacency relationships, the weights between the new adjacent nodes will change. ReLU (Rectified Linear Unit) and Softmax are both activation functions. ReLU is used to eliminate weak connections, while Softmax is used to achieve adaptive matrix normalization. D For degree matrix, A It is an adjacency matrix. I It is the identity matrix. W adap As weights. Generated directly. D 1 / 2 AD 1 / 2 This avoids repetitive calculations during the learning process. For the first l Layer adaptive embedding adjacency matrix, For the first l layer matrix The first in i Line 1 i Items in the column, i∈[0, n all ],j∈[0, N L ], n all For the total number of nodes, N L The number of steps during prediction. It is the first l The input matrix of the layer, , For the first l Layer i The items in the row, For the first l Layer j The items in the row, For nodes i The set of indexes of all neighboring nodes, For the first l Layer temperature node i The self-connect weight value, For the first l Layer temperature node i Weight values ​​on the hot path between other adjacent nodes , σ(·) represents the activation function. The proposed dynamic graph convolutional neural network fully considers the weights on each thermal path, thus it can well describe the spatial correlation characteristics of the transient temperature field; Step 1-2-2: Input the multidimensional signal output from the spatial correlation learning of the dynamic graph convolutional neural network into the LSTM to obtain the temporal correlation characteristics of each node, so as to establish a long-term temperature prediction model fused with the LSTM and the dynamic graph convolutional neural network (e.g., Figure 2 As shown in the figure, the model predicts the temperature rise over the next 120 time steps. The specific model expression is as follows: This represents the dynamic graph convolution process; see step 1-2-1 for a detailed explanation. f t , i t and o t They are at time t Forget gate, input gate, and output gate at the time. and These represent the candidate cell state and the renewed cell state, respectively. This represents the information retained in the old cell state. This indicates information about newly added cell states. h t Indicates at time t The hidden layer output,h t-1 Indicates at time t -1 is the hidden layer output. W f , W i , W c and W o All represent model weights. b f , b i , b c and b o The bias of the prediction model is represented by tanh(·), which represents the activation function. Steps 1-2-3: Input the multidimensional signal output from the spatial correlation learning of the dynamic graph convolutional neural network into the GRU to obtain the temporal correlation characteristics of each node, so as to establish a short-term temperature prediction model fused with the GRU (e.g., Figure 3 As shown in the figure, the model predicts the temperature rise over the next five time steps. The specific model expression is as follows: This represents the dynamic graph convolution process; see step 1-2-1 for a detailed explanation. u t , r t and c t They are at time t Update gates, reset gates, and memory gates in real time. h t Indicates at time t The hidden layer output, h t-1 Indicates at time t -1 is the hidden layer output. W u 、W r 、W c1 All represent the weights of the prediction model. b u 、b r 、b c1 This represents the bias of the prediction model.

[0029] Step 2 is as follows: Step 2-1: Determine the maximum temperature value based on the short-term or long-term temperature rise prediction results. m MaxNumber of overheating nodes n over Maximum temperature rise rate v T_max ; Step 2-2: Based on the characteristics of each component of the electric drive assembly, define dynamic risk indicators for thermal failure, including over-temperature thresholds. m thr The proportion of overheated nodes n over / n all With temperature rise rate threshold v T_thr ; Steps 2-3: According to Figure 4 The flowchart illustrating the dynamic thermal failure risk level determination process at different time steps shows that the highest temperature value, the proportion of overheated nodes, and the maximum temperature rise rate are compared sequentially with the overheating threshold, the overheated node proportion threshold, and the temperature rise rate threshold to determine whether they exceed the thresholds, and scores are obtained for the three indicators. Simultaneously, an entropy weight method is used to construct a high-medium-low risk level assessment framework based on the global dynamic risk indicator weight distribution, comprehensively assessing the global thermal failure risk level at each time step in the future time domain, and predicting the potential impact range and safety risks of overheating of key components in real time.

[0030] Steps 2-3 are as follows: Step 2-3-1: Calculate the highest temperature values ​​among key components of the electric drive assembly, such as the inverter IGBT junction temperature, on-board charger (OBC) temperature, DC / DC converter temperature, various motor components temperature, and oil temperature. m Max The proportion of overheated nodes n over / n all With the maximum rate of temperature rise v T_max sequentially with over-temperature threshold m thr Thresholds for the percentage of overheated nodes (defining two thresholds: 80% and 30%), and thresholds for the rate of temperature rise. v T_thr ( v T_thr according to m thr Comparison of 10% of the value: like m Max >m thr , n over / n all >80%, and v T_max>v T_thr Then it is defined as a pattern. m a ; like m Max >m thr , n over / n all >80%, and v T_max ≤ v T_thr Then it is defined as a pattern. m b ; like m Max >m thr 30% < n over / n all ≤80%, and v T_max >v T_thr Then it is defined as a pattern. m c ; like m Max >m thr 30% < n over / n all ≤80%, and v T_max ≤ v T_thr Then it is defined as a pattern. m d ; like m Max >m thr , n over / n all ≤30%, and v T_max >v T_thr Then it is defined as a pattern. m e ; like m Max >m thr , nover / n all ≤30%, and v T_max ≤ v T_thr Then it is defined as a pattern. m f ; like m Max ≤ m thr ,and v T_max >v T_thr Then it is defined as a pattern. m g ; like m Max ≤ m thr ,and v T_max ≤ v T_thr Then it is defined as a pattern. m h .

[0031] Calculate the scores of the three indicators under each mode. s 1, s 2, s 3, respectively corresponding to: 1) Does the highest temperature value exceed the over-temperature threshold? 2) Does the proportion of overheated nodes exceed the threshold for the proportion of overheated nodes? 3) Whether the maximum temperature rise rate exceeds the temperature rise rate threshold.

[0032] Table 1 shows an example of score mapping in this embodiment.

[0033] Table 1. Score chart for thermal failure risk assessment indicators; Step 2-3-2: Calculate the objective weight coefficients among the evaluation indicators within each part based on the entropy weight method, and plan for the future. N L The short- and long-term temperature rise prediction data for each time step can be used to determine three performance evaluation indicators for each time step. u mn For the first m Step 1 n The value of each performance metric, m ∈[1, N L ], n ∈[1,3], for short-time prediction, NL =5; for long-term forecasting N L =120.

[0034] right u mn Standardize to obtain , ū mn ∈[0,1]; Solve for the entropy of each performance metric , , ∈[0,1]; Calculate the weight of each performance metric W n , ; Step 2-3-3: The sum of the products of the different indicator scores obtained in Step 2-3-1 and the corresponding weight values ​​obtained in Step 2-3-2 is taken as the final score. The risk level is determined based on the final score. R sys (Values: High risk 3, Medium risk 2, Low risk 1).

[0035] For example: If it is determined to be high risk, then the risk level is... R sys =3; If it is determined to be of medium risk, then the risk level is... R sys =2; If it is determined to be low risk, then the risk level is... R sys =1, to construct a high-medium-low risk level assessment framework. At the same time, for the short-term and long-term predicted temperature rise results at each moment, steps 2-3 are used to determine the short-term and long-term over-temperature failure risk.

[0036] Step 3 specifically involves: Step 3-1: Based on the electric drive powertrain power domain control architecture, establish an active thermal failure risk reinforcement learning prediction hierarchical collaborative suppression algorithm, including a global optimization layer (upper layer) and a distributed execution layer (lower layer), such as... Figure 5 As shown; Step 3-2: The global optimization layer (upper layer) handles thermal failure suppression control, integrating reinforcement learning and robust model predictive control. Specifically, the Deep Deterministic Policy Gradient (DDPG) algorithm is used to dynamically correct the cooling and energy efficiency optimization boundaries of each subsystem. The predicted thermal safety boundaries (including the maximum allowable rate of temperature rise and the maximum allowable torque derating), thermal equilibrium weights, and desired temperatures are provided to the evaluation network and the long-term temperature prediction model.

[0037] The evaluation network is based on the prediction results of the long-term temperature prediction model and the actual component temperature fed back from the lower layer. Combined with the long-term thermal failure risk assessment level determined in step 2, a reward function is constructed to support the optimization of the value function of the evaluation network, so as to achieve effective evaluation of the output action of the action network.

[0038] The final determined optimal thermal safety boundary, thermal equilibrium weight, and expected temperature are encoded as global consistency constraints and provided to the lower distributed execution layer. Step 3-3: The distributed execution layer (lower layer) applies a distributed model predictive control algorithm based on the alternating direction multiplier method (ADMM). Based on the thermal safety boundary input from the upper layer and the temperature predicted by the short-time temperature prediction model, the short-term thermal failure risk assessment level is determined in conjunction with Step 2. With the optimization objectives of minimizing over-temperature risk and energy consumption, rolling optimization is achieved by continuously updating the optimal control law. This enables real-time and precise coordination of the local control actions of each execution unit in the thermal management submodule and the motor control module within the short-time domain, and feeds back the actual temperature rise data and actuator status to the upper-level risk assessment module.

[0039] Step 3-2 specifically involves: Step 3-2-1: The global optimization layer (upper layer) includes: robust model predictive control and a single-agent reinforcement learning framework based on deep deterministic policy gradient algorithm.

[0040] The single-agent reinforcement learning framework includes an action network and an evaluation network.

[0041] The robust model predictive control is based on a long-term temperature prediction model using a dynamic graph convolutional neural network fused with LSTM. It predicts the temperature rise of key components of the electric drive assembly in real time and constructs a robust feasible region by shrinking the temperature constraint based on the prediction error. Treat it as an action network Inputs and constraints also serve as evaluation criteria for the network. Q (·)enter.

[0042] Constructed robust feasible domain for: in, T i For the first i Temperature of key components T i,max For the first i The maximum temperature threshold for each key component can be determined according to the lower limit of the safe operating temperature of the corresponding component. For example, the maximum temperature threshold for the IGBT junction can be 150℃. For the first i The temperature safety margin for each component ranges from [5, 20]. , is the confidence coefficient, with a value range of [1.5, 3]. For the first i The standard deviation of the temperature prediction error for each component ranges from [1, 8]. For the rate of temperature rise, For the first i The maximum permissible global temperature rise rate for each component is within the range of [1.5, 3]. For electromagnetic torque, This is the maximum permissible torque derating factor, with a value of 0.5. The maximum permissible torque is determined by the motor's external characteristics.

[0043] The action network in the single-agent reinforcement learning framework Used to output continuous motion This includes thermal safety boundaries and lower-level optimization parameters: in To reinforce the observation vectors of the learning agent, This represents the maximum permissible rate of temperature rise generated in the upper layer. For the desired temperature trajectory, This is the thermal equilibrium weight matrix.

[0044] in, Temperature status of the electric drive assembly. This is the temperature rise rate state. N L For long-term prediction step size, As for the system risk level, For temperature prediction trajectory, , These are the motor control and thermal management control quantities from the previous moment.

[0045] Here are the parameters of the action network, and their update gradient is: in J This represents the objective function of the action network, which is the expected reward that can be obtained under the current policy. Represents the gradient of the action network parameters; To evaluate network parameters, M This refers to the number of samples in a small batch. This indicates the value of evaluating the network output. Q For action variables The vector of partial derivatives; M This represents the number of samples in a small batch.

[0046] The target evaluation network and target action network are soft-updated as follows: , in, This is the soft update coefficient. To evaluate network parameters for the target, The parameters are the target action network parameters.

[0047] The evaluation network in the single-agent reinforcement learning framework Q (·) Used to evaluate action networks Whether the output thermal safety boundary is good or not is expressed as follows: .

[0048] The network loss function is evaluated as follows , y t For DDPG target value, ,in, r t The current reward after performing the action. As a discount factor, , For the target evaluation network and the target action network, when o t+1 When it is in a terminated state, o t = r t .

[0049] Step 3-2-2: Dynamically adjust the optimization objective based on the deep deterministic strategy gradient algorithm in step 3-2-1, predict and generate the thermal safety boundary (including the maximum allowable temperature rise rate and the maximum allowable torque derating), thermal equilibrium weight, and desired temperature, and pass them as input variables to the evaluation network and the long-term temperature prediction model; The specific method for determining the maximum permissible torque derating is as follows: A dynamic performance test bench for the electric drive assembly was built to conduct motor torque derating tests. The output speed and torque of the motor and gearbox were measured to study the impact of motor torque derating on the amplitude of output torque fluctuation. Based on the analysis results of the degree of impact of motor torque derating on the amplitude of output torque fluctuation, the torque derating limit under different operating conditions was determined.

[0050] The thermal safety boundary is as follows: The maximum allowable rate of temperature rise is constrained as follows: ; in, , The sampling period; The torque derating boundary is: ; Furthermore, as thermal risk increases, the lower the allowable rate of temperature rise, the stronger the torque derating.

[0051] Step 3-2-3: Based on the prediction results of the long-term temperature prediction model and the actual component temperature fed back from the lower layer, the long-term thermal failure risk assessment level is determined in combination with Step 2. A cost (reward) function is constructed to support the evaluation network optimization value function, so as to realize the effective evaluation of the output action of the action network, and enable the single agent to obtain the long-term return of the system through short-term coordinated prediction optimization iteration.

[0052] The reward function also includes thermal risk, temperature deviation, auxiliary energy consumption, torque smoothness, and constraint violation penalties, as follows: in, For the weights of the reward function, Auxiliary energy consumption for water pumps and oil pumps. , For water pump power consumption, For oil pump power consumption, This represents the change in torque. To control the change in quantity, Penalties for violations of temperature and rate of temperature rise constraints ,in and This is the penalty coefficient.

[0053] Step 3-2-4: Encode the optimal thermal safety boundary, thermal equilibrium weight, and expected temperature generated in Step 3-2-1 based on the deep deterministic policy gradient algorithm into global consistency constraints for the lower layer, and input them into the lower layer control module.

[0054] Step 3-3 is as follows: Step 3-3-1: The distributed execution layer (lower layer) constructs a distributed model predictive control algorithm based on the alternating direction multiplier method for the motor vector control submodule and the thermal management submodule to address the short-term sudden over-temperature risk. The motor vector control submodule is used for active suppression of thermal failure risk, and its control quantity is: in, , For motor d / q shaft current, The inverter switching frequency, This submodule is used to actively adjust the inverter dead time. , , and While meeting torque, current, and voltage constraints, it actively adjusts the current and inverter control parameters to reduce motor and power device losses and actively suppress the risk of thermal failure.

[0055] The optimized variables for the motor vector control submodule are: ; The objective function of the motor vector control submodule is: The constraints of the motor vector control submodule are: ; in, Optimize the variable set for the motor control submodule. For torque tracking weights, To meet the requirement of electromagnetic torque, As a weight for loss suppression, For inverter losses, For current penalty weights, Penalty weights are applied to changes in switching frequency and dead time. The maximum allowable torque given by the upper layer. I max This is the maximum permissible current under the current temperature conditions. and These are the minimum and maximum allowable inverter switching frequencies, respectively. and These are the minimum and maximum allowable inverter dead time.

[0056] The thermal management submodule is used for passive cooling and energy efficiency optimization, and its control variables are: in, For water pump flow rate, For oil pump speed, The water pump speed, This is for adjusting the opening of cooling valves. The thermal management submodule regulates heat dissipation capacity through cooling actuators, reducing auxiliary energy consumption such as pumps / fans while suppressing overheating of critical components, thus achieving passive cooling and energy efficiency optimization.

[0057] The optimization variables for the thermal management submodule are: ; The objective function of the thermal management submodule is: ; The thermal management submodule is constrained as follows: ; in, Optimize the variable set for the thermal management submodule. As a penalty weight for the rate of temperature rise, Assigning energy consumption weights to water pumps and oil pumps. Weights for the rate of change of thermal management control actions.

[0058] Step 3-3-2: Decompose the global consistency constraints of the global optimization layer input and pass them as reference inputs to the reference trajectories such as the desired thermal safety boundary, desired thermal equilibrium weights, and desired temperature. Based on a dynamic graph convolutional neural network temperature short-term prediction model fused with GRU, the future short-term temperature rise status of key components of the electric drive assembly is predicted in real time. Step 3-3-3: Based on the thermal safety boundary input from the upper layer (global optimization layer) and the temperature predicted by the short-term temperature prediction model, and combined with Step 2, determine the short-term thermal failure risk assessment level. The optimization objectives are to achieve the highest cooling rate under high failure risk and the lowest energy consumption under low-to-medium failure risk. The optimal control law is searched using the alternating direction multiplier method, and rolling optimization is achieved by continuously updating the optimal control law. Through feedback correction and iterative solution of the local temperature rise and energy consumption optimization problem in each region, real-time and precise coordination of local control actions (inverter switching frequency and dead time adjustment, motor current / water pump flow / oil pump speed adjustment) of each execution unit in the short-term thermal management submodule and motor control module is achieved. The optimal control law search and rolling optimization process implemented in step 3-3-3 within each short sampling period using the alternating direction multiplier method is as follows: 1) Definition of coupling variables The motor vector control submodule and the thermal management submodule are coupled through losses, temperature, and temperature rise rate to provide... ,in For inverter losses, For motor losses, the motor vector control submodule predicts the coupling variable as follows: The thermal management submodule predicts the coupling variables as follows: The consistency constraint is ,Right now: , , ;in , and These represent the losses, temperature, and temperature rise rate predicted by the respective motor vector control submodules. , and These represent the losses, temperature, and rate of temperature rise used or estimated by the thermal management submodule in the temperature prediction model.

[0059] 2) Solving the local optimization problem of the motor vector control submodule in, r This represents the number of ADMM iterations. As an ADMM penalty factor, Z As a consensus variable, To scale the dual variable, This is the objective function for the motor vector control submodule.

[0060] 3) Solving the local optimization problem of the thermal management submodule in, To scale the dual variable, This is the objective function for the thermal management submodule.

[0061] 4) Update consensus variables Standard consensus updated to If thermal safety is a priority, a weighted consensus update can be adopted. in, Weights for the motor control submodule. This sets the weight for the thermal management submodule. When thermal safety needs to be prioritized, it can be set as follows: .

[0062] 5) Dual variable update The dual variables of the motor vector control submodule are updated as follows: ; The dual variables of the thermal management submodule are updated as follows: ; Dual variables can be understood as the cumulative compensation amount or coordination pressure for inconsistencies in predictions between submodules.

[0063] 6) Convergence judgment Original residual Defined as: ; Dual residuals Defined as: ; The stopping condition is: , ; in, and These are the thresholds for the original residual and the dual residual, respectively.

[0064] 7) Execute the optimal control law Only the first control variable in the prediction time domain is executed in each rolling cycle: , .

[0065] Step 3-3-4: Feed back the actual temperature rise data (including inverter IGBT junction temperature, on-board charger (OBC) temperature, DC-DC converter (DC / DC) temperature, motor component temperatures, and oil temperature) and actuator status (water pump / oil pump / motor power, motor speed, current, voltage, etc.) to the upper-level risk assessment module, triggering the reward function to update the support evaluation network's value function, thereby achieving effective evaluation of the action network's output action.

[0066] Step 4 specifically includes: 1) Sampling period definition: Define the upper-layer sampling period as... The sampling period of the lower layer is , usually satisfy ,in, M 1 represents the time scale scaling factor, with a value range of [5, 20]. 2) Upper-layer thermal safety boundary update: Upper layer every Update the global thermal safety boundary once , ; 3) Lower-level scrolling optimization: The lower level executes within the same cycle as the upper level. M One short-term rolling optimization was performed to adjust the control quantities of the motor vector control submodule and the thermal management submodule. and ,in .

[0067] The upper-level global optimization layer uses a longer sampling period and is responsible for long-term thermal failure risk prediction and thermal safety boundary generation. The lower-level distributed execution layer uses a shorter sampling period and is responsible for short-term over-temperature suppression and actuator coordination. The upper and lower layers achieve time-domain decoupling through a multi-sampling-period mechanism of "low-frequency decision-making at the upper level and high-frequency execution at the lower level," enabling global thermal safety boundary planning and local actuator rapid response to operate collaboratively at different time scales. Reinforcement learning does not directly control the fast actuators but outputs safety boundaries, improving system safety. ADMM improves system real-time performance and scalability by decomposing the coupled optimization problem into local problems.

[0068] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A reinforcement learning-based prediction and active collaborative suppression method for thermal failure risk of electric drive assemblies, used to achieve multi-objective collaborative optimization of thermal safety, operational safety, and energy efficiency of electric drive assemblies, characterized in that... Includes the following steps: Step 1: Construct a dynamic graph convolutional neural network for long-term and short-term temperature prediction, respectively, by integrating a long short-term memory network and a threshold recursive unit. Combine this with sensor information to estimate the temperature rise of key components at different future times. Step 2: Using the entropy weight method, construct risk assessment frameworks for different levels based on the weight distribution of global dynamic risk indicators; Step 3: Establish a hierarchical active collaborative suppression algorithm for predicting the thermal failure risk of the electric drive assembly using reinforcement learning; combining the temperature rise prediction and risk assessment results, the upper layer uses a deep deterministic strategy gradient algorithm to predict and generate the thermal safety boundary, and the lower layer applies a distributed model prediction control algorithm based on the alternating direction multiplier method to finely coordinate the actions of each execution unit. Step 4: The upper and lower layer strategies achieve hierarchical decision-making and rolling time-domain collaborative optimization through time-domain decoupling, realizing multi-objective collaborative optimization of electric drive assembly thermal safety, operational safety and energy efficiency.

2. The method according to claim 1, characterized in that, Step 1 is as follows: Step 1-1: Collect sensor data, including motor speed, torque, oil temperature, and inverter cooling water temperature, based on historical sparse sensor data from the past 5 time steps. μ t0 The maximum likelihood estimation method was used to identify the temperature of other key components within five time steps. The key components include the motor permanent magnet, windings, iron core, inverter IGBT, assembly housing, on-board charger, and DC / DC converter. The maximum likelihood estimation model is shown below: in, express t Based on partial temperature sensor data μ t0 Other estimated temperature node data, These are the model parameters for the maximum likelihood estimation model; simultaneously, the estimation results of the maximum likelihood estimation model are compared with the sparse sensing historical data from the past 5 time steps. μ t0 Find the union, and we get t The node information for each time point is as follows: in, express t All node information at any given time; Steps 1-2: Construct a long-term temperature prediction model and a short-term temperature prediction model that integrate LSTM and GRU respectively. Based on the input information, predict the temperature rise of key components within the next 120 or 5 time steps.

3. The method according to claim 1, characterized in that, Step 2 is as follows: Step 2-1: Determine the maximum temperature value based on the short-term or long-term temperature rise prediction results. μ Max Number of overheating nodes n over Maximum temperature rise rate v T_max ; Step 2-2: Based on the characteristics of each component of the electric drive assembly, define dynamic risk indicators for thermal failure, including over-temperature thresholds. μ thr The proportion of overheated nodes n over / n all With temperature rise rate threshold v T_thr ,in n all The total number of nodes; Steps 2-3: Compare the highest temperature value, the proportion of overheated nodes, and the maximum temperature rise rate with the overheating threshold, the overheated node proportion threshold, and the temperature rise rate threshold, respectively, to determine whether they exceed the thresholds, and obtain scores for the three indicators based on the comparison results; at the same time, use the entropy weight method to construct a high-medium-low risk level assessment framework based on the global dynamic risk indicator weight distribution, comprehensively assess the global thermal failure risk level at each time step in the future time domain, and predict the potential impact range and safety risks of overheating of key components in real time.

4. The method according to claim 3, characterized in that, Steps 2-3 are as follows: Step 2-3-1: Calculate the highest temperature values ​​among the key components of the electric drive assembly: inverter IGBT junction temperature, on-board charger (OBC) temperature, DC / DC converter temperature, various motor component temperatures, and oil temperature. μ Max The proportion of overheated nodes n over / n all With the maximum rate of temperature rise v T_max sequentially with over-temperature threshold μ thr Overheating node percentage threshold, temperature rise rate threshold v T_thr Compare: like μ Max >μ thr , n over / n all >80%, and v T_max >v T_thr Then it is defined as a pattern. m a ; like μ Max >μ thr , n over / n all >80%, and v T_max ≤ v T_thr Then it is defined as a pattern. m b ; like μ Max >μ thr 30% < n over / n all ≤80%, and v T_max >v T_thr Then it is defined as a pattern. m c ; like μ Max >μ thr 30% < n over / n all ≤80%, and v T_max ≤ v T_thr Then it is defined as a pattern. m d ; like μ Max >μ thr , n over / n all ≤30%, and v T_max >v T_thr Then it is defined as a pattern. m e ; like μ Max >μ thr , n over / n all ≤30%, and v T_max ≤ v T_thr Then it is defined as a pattern. m f ; like μ Max ≤ μ thr ,and v T_max >v T_thr Then it is defined as a pattern. m g ; like μ Max ≤ μ thr ,and v T_max ≤ v T_thr Then it is defined as a pattern. m h ; Calculate the scores of the three indicators under each mode. s 1, s 2, s 3, respectively corresponding to: 1) Does the highest temperature value exceed the over-temperature threshold? 2) Does the proportion of overheated nodes exceed the threshold for the proportion of overheated nodes? 3) Does the maximum temperature rise rate exceed the temperature rise rate threshold? Step 2-3-2: Calculate the objective weight coefficients among the evaluation indicators within each part based on the entropy weight method, and plan for the future. N L The short- and long-term temperature rise prediction data for each time step are used, and three performance evaluation indicators are determined for each time step. u mn For the first m Step 1 n The value of each performance metric, m ∈[1, N L ], n ∈[1,3], for short-time prediction, N L =5; for long-term forecasting N L =120; right u mn Standardize to obtain , ū mn ∈[0,1]; Solve for the entropy of each performance metric , , ∈[0,1]; Calculate the weight of each performance metric W n , ; Step 2-3-3: The sum of the products of the different indicator scores obtained in Step 2-3-1 and the corresponding weight values ​​obtained in Step 2-3-2 is taken as the final score. The risk level is determined based on the final score. R sys Values: High risk 3, Medium risk 2, Low risk 1.

5. The method according to claim 1, characterized in that, Step 3 specifically involves: Step 3-1: Based on the power domain control architecture of the electric drive assembly, establish an active thermal failure risk reinforcement learning prediction hierarchical collaborative suppression algorithm, which includes an upper global optimization layer and a lower distributed execution layer; Step 3-2: The global optimization layer handles thermal failure suppression control, which integrates reinforcement learning and robust model predictive control; a deep deterministic policy gradient algorithm is used to dynamically correct the cooling and energy efficiency optimization boundaries of each subsystem, and the thermal safety boundary, thermal equilibrium weight and expected temperature are generated from the prediction and provided to the evaluation network and the long-term temperature prediction model; the thermal safety boundary includes the maximum allowable temperature rise rate and the maximum allowable torque derating. The evaluation network is based on the prediction results of the long-term temperature prediction model and the actual component temperature fed back from the lower layer. Combined with the long-term thermal failure risk assessment level determined in step 2, a reward function is constructed to support the optimization of the value function of the evaluation network, so as to achieve effective evaluation of the output action of the action network. The final determined optimal thermal safety boundary, thermal equilibrium weight, and expected temperature are encoded as global consistency constraints and provided to the lower distributed execution layer. Step 3-3: The distributed execution layer applies a distributed model predictive control algorithm based on the alternating direction multiplier method. Based on the thermal safety boundary input from the upper layer and the temperature predicted by the short-time temperature prediction model, the short-term thermal failure risk assessment level is determined in conjunction with Step 2. With the optimization objectives of minimizing over-temperature risk and energy consumption, rolling optimization is achieved by continuously updating the optimal control law. This enables real-time and precise coordination of the local control actions of each execution unit in the thermal management submodule and the motor control submodule within the short-time domain, and feeds back the actual temperature rise data and actuator status to the upper-level risk assessment module.

6. The method according to claim 5, characterized in that, Step 3-2 specifically involves: Step 3-2-1: The global optimization layer includes: robust model predictive control and a single-agent reinforcement learning framework based on deep deterministic policy gradient algorithm; the single-agent reinforcement learning framework includes an action network and an evaluation network; The robust model predictive control is based on a long-term temperature prediction model using a dynamic graph convolutional neural network fused with LSTM. It predicts the temperature rise of key components of the electric drive assembly in real time and constructs a robust feasible region by shrinking the temperature constraint based on the prediction error. Treat it as an action network Inputs and constraints also serve as evaluation criteria for the network. Q (·)enter; Constructed robust feasible domain for: in, T i For the first i Temperature of key components T i,max For the first i Maximum temperature threshold of each key component For the first i Temperature safety margin of each component , Here is the confidence coefficient. For the first i Standard deviation of temperature prediction error for each component For the rate of temperature rise, For the first i The maximum permissible global temperature rise rate for each component For electromagnetic torque, This is the maximum permissible torque derating factor. Maximum permissible torque; The action network in the single-agent reinforcement learning framework Used to output continuous motion This includes thermal safety boundaries and lower-level optimization parameters: in To reinforce the observation vectors of the learning agent, This represents the maximum permissible rate of temperature rise generated in the upper layer. For the desired temperature trajectory, This is the thermal equilibrium weight matrix; in, Temperature status of the electric drive assembly. This is the temperature rise rate state. N L For long-term prediction step size, As for the system risk level, For temperature prediction trajectory, , These are the motor control and thermal management control quantities from the previous moment; Here are the parameters of the action network, and their update gradient is: in J This represents the optimization objective function of the action network, i.e., the expected reward that can be obtained under the current policy; Represents the gradient of the action network parameters; To evaluate network parameters, M This refers to the number of samples in a small batch. This indicates the value of evaluating the network output. Q For action variables The vector of partial derivatives; The target evaluation network and target action network are soft-updated as follows: , in, This is the soft update coefficient. To evaluate network parameters for the target, For the target action network parameters; The evaluation network in the single-agent reinforcement learning framework Q (·) Used to evaluate action networks Whether the output thermal safety boundary is good or not is expressed as follows: ; The network loss function is evaluated as follows , y t For DDPG target value, ,in, r t The current reward after performing the action. As a discount factor, , For the target evaluation network and the target action network, when o t+1 When it is in a terminated state, o t = r t ; Step 3-2-2: Dynamically adjust the optimization objective based on the deep deterministic policy gradient algorithm in step 3-2-1, predict and generate the thermal safety boundary, thermal equilibrium weight and expected temperature, and pass them as input variables to the evaluation network and the long-term temperature prediction model; The thermal safety boundary is as follows: The maximum allowable rate of temperature rise is constrained as follows: ; in, , The sampling period; The torque derating boundary is: ; Furthermore, as thermal risk increases, the lower the allowable rate of temperature rise, the stronger the torque derating. Step 3-2-3: Based on the prediction results of the long-term temperature prediction model and the actual component temperature fed back from the lower layer, combined with Step 2, determine the long-term thermal failure risk assessment level, construct a cost function to support the evaluation network optimization value function, realize the effective evaluation of the action network output action, and enable the single agent to obtain the long-term system return through short-term coordinated prediction optimization iteration. Step 3-2-4: Encode the optimal thermal safety boundary, thermal equilibrium weight, and expected temperature generated in Step 3-2-1 based on the deep deterministic policy gradient algorithm into global consistency constraints for the lower layer, and input them into the lower layer control module.

7. The method according to claim 6, characterized in that, The reward function also includes thermal risk, temperature deviation, auxiliary energy consumption, torque smoothness, and constraint violation penalties, as follows: in, For the weights of the reward function, Auxiliary energy consumption for water pumps and oil pumps. This is the change in torque. To control the change in quantity, Penalties for violations of temperature and rate of temperature rise constraints ,in and This is the penalty coefficient.

8. The method according to claim 5, characterized in that, Step 3-3 is as follows: Step 3-3-1: To address the risk of short-term sudden over-temperature, the lower distributed execution layer constructs a distributed model predictive control algorithm based on the alternating direction multiplier method for the motor vector control submodule and the thermal management submodule. The motor vector control submodule is used for active suppression of thermal failure risk, and its control quantity is: in, , For motor d / q shaft current, The inverter switching frequency, The inverter dead time; the motor vector control submodule actively adjusts... , , and While meeting torque, current and voltage constraints, it actively adjusts the current and inverter control parameters to reduce motor and power device losses and actively suppress the risk of thermal failure. The optimized variables for the motor vector control submodule are: ;in N L For long-term prediction step size; The objective function of the motor vector control submodule is: The constraints of the motor vector control submodule are: ; in, Optimize the variable set for the motor control submodule. For torque tracking weights, To meet the requirement of electromagnetic torque, As a weight for loss suppression, For inverter losses, For current penalty weights, Penalty weights are applied to changes in switching frequency and dead time. The maximum allowable torque given by the upper layer. I max This is the maximum permissible current under the current temperature conditions. and These are the minimum and maximum allowable inverter switching frequencies, respectively. and The minimum and maximum allowable inverter dead time; The thermal management submodule is used for passive cooling and energy efficiency optimization, and its control variables are: in, For water pump flow rate, For oil pump speed, The water pump speed, To cool the valve opening; the thermal management submodule adjusts the heat dissipation capacity through the cooling actuator, suppressing overheating of key components while reducing pump / fan auxiliary energy consumption, thus achieving passive cooling and energy efficiency optimization; The optimization variables for the thermal management submodule are: ; The objective function of the thermal management submodule is: ; The thermal management submodule is constrained as follows: ; in, Optimize the variable set for the thermal management submodule. As a penalty weight for the rate of temperature rise, Assigning energy consumption weights to water pumps and oil pumps. Weights for the rate of change of thermal management control actions; Step 3-3-2: Decompose the global consistency constraint input of the global optimization layer and pass it as a reference input to the desired thermal safety boundary, desired thermal equilibrium weight and desired temperature reference trajectory; based on the dynamic graph convolutional neural network temperature short-term prediction model with fused GRU, predict the future short-term temperature rise state of key components of the electric drive assembly in real time. Step 3-3-3: Based on the thermal safety boundary input from the upper layer and the temperature predicted by the short-term temperature prediction model, and combined with Step 2, determine the short-term thermal failure risk assessment level. With the highest cooling rate under high failure risk and the lowest energy consumption under low and medium failure risk as the optimization objectives, the alternating direction multiplier method is used to search for the optimal control law, and rolling optimization is achieved by continuously updating the optimal control law. Through feedback correction and iterative solution of the local temperature rise and energy consumption optimization problem in each region, real-time fine coordination of the local control actions of each execution unit in the thermal management submodule and motor control module in the short time domain is achieved. Step 3-3-4: Feed back the actual temperature rise data and actuator status to the upper-level risk assessment module, trigger the reward function to update the support evaluation network to optimize the value function, and realize the evaluation of the action network output action.

9. The method according to claim 8, characterized in that, The optimal control law search and rolling optimization process implemented in step 3-3-3 within each short sampling period using the alternating direction multiplier method is as follows: 1) Definition of coupling variables The motor vector control submodule and the thermal management submodule are coupled through losses, temperature, and temperature rise rate to provide... ,in For inverter losses, For motor losses, the motor vector control submodule predicts the coupling variable as follows: The thermal management submodule predicts the coupling variables as follows: Consistency constraint is ,Right now: , , ;in , and These represent the losses, temperature, and temperature rise rate predicted by the respective motor vector control submodules. , and These represent the losses, temperature, and rate of temperature rise used or estimated by the thermal management submodule in the temperature prediction model, respectively. 2) Solving the local optimization problem of the motor vector control submodule in, r This represents the number of ADMM iterations. As an ADMM penalty factor, Z As a consensus variable, To scale the dual variable, The objective function for the motor vector control submodule; 3) Solving the local optimization problem of the thermal management submodule in, To scale the dual variable, The objective function for the thermal management submodule; 4) Update consensus variables Standard consensus updated to If thermal safety is a priority, a weighted consensus update can be adopted. in, Weights for the motor control submodule. Assign weights to the thermal management submodule; when thermal safety needs to be prioritized, it can be set as follows: ; 5) Dual variable update The dual variables of the motor vector control submodule are updated as follows: ; The dual variables of the thermal management submodule are updated as follows: ; Dual variables can be understood as the cumulative compensation amount or coordination pressure for inconsistencies in predictions between submodules; 6) Convergence judgment Original residual Defined as: ; Dual residuals Defined as: ; The stopping condition is: , ; in, and These are the thresholds for the original residual and the dual residual, respectively. 7) Execute the optimal control law Only the first control variable in the prediction time domain is executed in each rolling cycle: , .

10. The method according to claim 1, characterized in that, Step 4 specifically involves: 1) Sampling period definition: Define the upper-layer sampling period as... T H The sampling period of the lower layer is T L ,satisfy ,in, M 1 represents the time scale scaling factor; 2) Upper-layer thermal safety boundary update: Upper layer every The global thermal safety boundary is updated periodically. , ; 3) Lower-level scrolling optimization: The lower level executes within the same cycle as the upper level. M One short-term rolling optimization was performed to adjust the control quantities of the thermal management submodule and the motor control submodule. The upper global optimization layer uses a longer sampling period and is responsible for long-term thermal failure risk prediction and thermal safety boundary generation; the lower distributed execution layer uses a shorter sampling period and is responsible for short-term over-temperature suppression and actuator coordination; the upper and lower layers achieve time-domain decoupling through a multi-sampling period mechanism of "low-frequency decision-making in the upper layer and high-frequency execution in the lower layer", enabling global thermal safety boundary planning and local actuator fast response to operate collaboratively at different time scales; reinforcement learning does not directly control fast actuators, but outputs safety boundaries to improve system safety; ADMM decomposes the coupled optimization problem into local problems.

Citation Information

Patent Citations

  • A thermal management control method for power motors based on integrated vehicle control

    CN111900833B

  • Thermal runaway risk assessment method based on big data

    CN112363061A

  • A thermal management control method and system integrating genetic ant colony algorithm

    CN113050426B

  • Thermal management system and its thermal management methods, thermal management equipment, storage media

    CN116315290B

  • Transformer overheating risk assessment method based on temperature rise simulation and fuzzy analytic hierarchy process

    CN117454716A