Thermal runaway prevention and heat dissipation optimization method and system based on capacitor module
By collecting and analyzing multi-source operation data of the dual-layer capacitor module, combining advanced neural networks and reinforcement learning technology, identifying causal relationships and conducting uncertainty evaluations, and dynamically adjusting the heat dissipation strategy, the problem of thermal management in the existing technology is solved, and efficient thermal runaway prevention and heat dissipation optimization is achieved.
Patent Information
- Application Number
- CN202510130815.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the thermal management method of the double-layer capacitor module lacks the refined evaluation of the state of a single capacitor, the thermal dissipation control strategy is not intelligent enough, and the function of warning potential faults, resulting in an increase in the risk of thermal runaway, performance decay and even safety accidents.
A thermal runaway prevention and heat dissipation optimization method based on capacitor module is adopted. By collecting multi-source operation data, combining improved hierarchical visual attention network and multi-level self-attention structure for feature extraction, learning position-encoded embedding time information is introduced, data completion and feature fusion is used using a bidirectional probability diffusion model and a multi-modal feature fusion network for data completion and feature fusion, causal relationships are identified and uncertainty evaluation is performed, a spatiotemporal dynamic graph neural network is constructed for temperature and stress prediction, and a layered reinforcement learning framework is used to dynamically adjust the heat dissipation strategy to achieve the optimal heat dissipation solution.
Effectively prevent thermal runaway, improve system safety, dynamically adjust heat dissipation strategies, achieve efficient heat dissipation, extend the life of capacitor modules, and provide scientific decision-making basis through multi-source data analysis and advanced prediction models to reduce maintenance costs.
Smart Images

Figure CN120180071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of thermal management of capacitor modules, and particularly to a method and system for preventing thermal runaway and optimizing heat dissipation based on capacitor modules. Background Art
[0002] As a new type of energy storage element, the electric double-layer capacitor has the advantages of high power density, fast charge and discharge speed, long cycle life, etc., and has broad application prospects in the fields of energy storage, transportation, power systems, etc. With the increasing complexity of application scenarios and the continuous improvement of system reliability requirements, the thermal management problem of electric double-layer capacitor modules has become increasingly prominent. A large amount of heat will be generated during the operation of the electric double-layer capacitor. If the heat dissipation is not timely, it will cause the module temperature to be too high, which will further lead to thermal runaway, performance degradation and even safety accidents;
[0003] In the prior art, the thermal management method of the electric double-layer capacitor module mainly relies on temperature sensors for temperature monitoring and adopts a simple control strategy, lacking refined evaluation of the state of single capacitors, the heat dissipation control strategy is not intelligent enough, and there is no function of warning potential faults;
[0004] Therefore, there is an urgent need for a solution to solve the problems existing in the prior art. Summary of the Invention
[0005] The embodiments of the present invention provide a method and system for preventing thermal runaway and optimizing heat dissipation based on capacitor modules, which can at least solve some problems existing in the prior art.
[0006] In the first aspect of the embodiments of the present invention, a method for preventing thermal runaway and optimizing heat dissipation based on capacitor modules is provided, including:
[0007] Collect multi-source operation data of single capacitors in the capacitor module, segment different types of data into data segments of the same size and perform linear projection by combining an improved hierarchical visual attention network, extract features by combining a multi-level self-attention structure and an attention mechanism, introduce learnable position encoding to embed time information, obtain initial features and add them to a bidirectional probability diffusion model, introduce conditional information during the diffusion process to complete data and add the complete data information to a multi-modal feature fusion network, calculate the correlation between modalities and adaptively modify the fusion weights by combining a dynamic weight allocation mechanism, obtain fusion features and add them to a causal structure discovery network, construct an initial causal graph structure based on expert knowledge and identify the causal relationship between different data by combining a causal discovery algorithm, obtain a time series dependence relationship and perform uncertainty evaluation, and obtain a performance analysis report corresponding to the single capacitor;
[0008] Construct a spatio-temporal dynamic graph neural network based on the performance analysis report, model the device temperature field distribution, obtain a dynamic prediction graph, and identify the evolution characteristics corresponding to the capacitor temperature and stress. Based on the evolution characteristics, combine hierarchical dilated convolution and self-attention mechanism to determine the parameter evolution law, determine the temperature change trend and stress accumulation pattern, and add them to a multi-task learning framework. Combine a tabular data neural network and a dynamic graph attention network to process the performance analysis report and the dynamic prediction graph to obtain a predicted state. Perform fault diagnosis through a large language model with a mixture of experts architecture and dynamic routing mechanism to obtain potential fault types and corresponding occurrence probabilities. Construct an uncertainty quantification module that combines a probabilistic graph model and Bayesian deep learning to calculate the confidence interval corresponding to each fault type. Combine pre-acquired expert knowledge to construct a hierarchical warning mechanism, set multi-level warning thresholds, and output graded warning signals;
[0009] Construct a hierarchical reinforcement learning framework based on the performance analysis report and the graded warning signals. Among them, the top-level policy network guides action selection according to the causal relationship, the middle-level coordination network evaluates the action impact based on the predicted state, and the bottom-level execution network performs device control to obtain an initial heat dissipation strategy. Add the initial heat dissipation strategy and the dynamic prediction graph to a hierarchical gated recurrent unit, and combine a soft attention mechanism to perform temporal state evaluation to obtain system state evolution characteristics. Combine the distributed advantage actor-critic algorithm to perform collaborative control on the heat dissipation unit. Combine the double deep Q-network and prioritized experience replay to determine a candidate control action sequence. Based on the heat dissipation effect, optimize the control action through the dynamic time warping algorithm and the deep deterministic policy gradient algorithm to obtain an optimal heat dissipation plan.
[0010] In an alternative embodiment,
[0011] Collect multi-source operation data of individual capacitors in the capacitor module. Combine an improved hierarchical visual attention network to segment different types of data into data segments of the same size and perform linear projection. Combine a multi-level self-attention structure and an attention mechanism to extract features. Introduce learnable positional encoding to embed time information to obtain initial features and add them to a bidirectional probabilistic diffusion model. Introduce conditional information during the diffusion process to complete the data to obtain complete data information and add it to a multi-modal feature fusion network. Calculate the correlation between modalities and combine a dynamic weight allocation mechanism to adaptively modify the fusion weights to obtain fusion features and add them to a causal structure discovery network. Construct an initial causal graph structure based on expert knowledge and combine a causal discovery algorithm to identify the causal relationship between different data to obtain a temporal dependence relationship and perform uncertainty evaluation. The performance analysis report corresponding to the individual capacitor includes:
[0012] Collect the voltage data, current data, ambient temperature data, housing temperature data, and charge and discharge state data of the individual capacitors in the acquisition capacitor module, and combine them to obtain multi-source operation data. Input the multi-source operation data into an improved hierarchical visual attention network. The improved hierarchical visual attention network includes a feature segmentation layer, a feature projection layer, and a feature enhancement layer. Among them, the feature segmentation layer uses an overlapping sliding window to segment the multi-source operation data to obtain data segments. The feature projection layer uses an independent projection matrix to map the data segments to a unified feature space to obtain mapped features. The feature enhancement layer enhances the mapped features through a channel-level attention mechanism and a spatial-level attention mechanism to obtain enhanced features;
[0013] Input the enhanced features into a multi-level self-attention structure. The multi-level self-attention structure includes a local attention unit and a global attention unit. The local attention unit uses a multi-head attention mechanism to extract local dependence features. The global attention unit uses a sparse attention mechanism to extract cross-segment association features. Embed the absolute position information and relative position information into the local dependence features and the cross-segment association features through a learnable position encoder to obtain initial features;
[0014] Input the initial features into a bidirectional probability diffusion model. During the training phase, construct a noise schedule to control the noise addition rate and train a denoising predictor. During the inference phase, convert the working condition information into a conditional vector, and perform data completion through multiple steps of iteration in combination with the conditional vector to obtain complete data information and input it into a multi-modal feature fusion network. Standardize the different modal features to obtain standardized features, construct an attention matrix to calculate the inter-modal correlation of the standardized features to obtain correlation features, and perform dynamic weight allocation on the correlation features based on a soft attention mechanism and introduce a cross-modal contrast learning strategy to obtain fusion features;
[0015] Construct an initial causal graph based on expert knowledge, including node attribute definitions and edge connection constraints. Input the fusion features into a causal structure discovery network, extract the trend features, periodic features, and mutation features of the time series feature sequence, evaluate the linear correlation and non-linear correlation between different feature sequences, determine the direction and weight of the causal edges based on the dependence strength to obtain causal relationships, extract the feature description sequences of the causal relationships under different time windows, generate multiple groups of evaluation samples through the bootstrap method and calculate the confidence scores, adopt an integration strategy to fuse different evaluation results to obtain the final confidence score, and generate a performance analysis report based on the final confidence score.
[0016] In an alternative embodiment,
[0017] The feature segmentation layer uses an overlapping sliding window to segment the multi-source operation data into data segments. The feature projection layer uses an independent projection matrix to map the data segments into a unified feature space to obtain mapped features. The feature enhancement layer enhances the mapped features through a channel-level attention mechanism and a spatial-level attention mechanism to obtain enhanced features, including:
[0018] Obtain the multi-source operation data and construct an adaptive segmentation mechanism. Detect the change rate of the multi-source operation data through a data change rate detection module, and dynamically adjust the sliding window parameters according to the change rate. Among them, when the change rate is higher than the first threshold, reduce the window length of the sliding window parameters to 80% of the reference window length. When the change rate is lower than the second threshold, increase the window length of the sliding window parameters to 120% of the reference window length;
[0019] Perform parallel processing on the multi-source operation data using sliding windows of different lengths. Among them, the first sliding window is used to capture transient change features to obtain a first feature segment, the second sliding window is used to extract local trend features to obtain a second feature segment, and the third sliding window is used to obtain macro-evolution features to obtain a third feature segment. Input the first feature segment, the second feature segment, and the third feature segment into a feature pyramid network for integration to obtain data segments;
[0020] Divide the original feature space into multiple subspaces. Each subspace is configured with a main projection matrix and an auxiliary projection matrix. Input the data segments into the main projection matrix to obtain basic features, input the data segments into the auxiliary projection matrix to obtain residual features, and perform feature fusion on the basic features and the residual features to obtain mapped features;
[0021] Construct a feature quality evaluation index. The feature quality evaluation index includes feature discrimination, feature stability, and information retention. Evaluate the mapped features based on the feature quality evaluation index to obtain an evaluation result, and optimize the main projection matrix and the auxiliary projection matrix according to the evaluation result;
[0022] Construct a hierarchical attention mechanism, including a feature channel attention mechanism, a time dimension attention mechanism, and a cross-modal attention mechanism. Among them, the feature channel attention mechanism calculates the inter-channel dependence relationship of the mapped features to obtain a channel attention map, extracts the channel feature statistical information of the mapped features to obtain channel importance, and performs adaptive weight fusion on the channel attention map and the channel importance to obtain channel attention weights;
[0023] The time - dimension attention mechanism extracts features of different time scales through the multi - head attention mechanism, and uses a deformable convolutional network to process the time series to obtain the time - dimension feature enhancement result. The cross - modal attention mechanism constructs different types of features into a graph structure and performs feature interaction through the message - passing mechanism to obtain enhanced features.
[0024] In an alternative embodiment,
[0025] Based on the performance analysis report, construct a spatio - temporal dynamic graph neural network and model the device temperature field distribution to obtain a dynamic prediction graph and identify the evolution features corresponding to the capacitor temperature and stress. Based on the evolution features, combine hierarchical dilated convolution and self - attention mechanism to determine the parameter evolution law, determine the temperature change trend and stress accumulation pattern and add them to the multi - task learning framework. Combine the tabular data neural network and the dynamic graph attention network to process the performance analysis report and the dynamic prediction graph to obtain the predicted state. Perform fault diagnosis through a large - language model with a mixture - of - experts architecture and a dynamic routing mechanism to obtain potential fault types and corresponding occurrence probabilities. Construct an uncertainty quantification module that combines a probabilistic graph model and Bayesian deep learning to calculate the confidence interval corresponding to each fault type. Combine the pre - obtained expert knowledge to construct a hierarchical early - warning mechanism, set multi - level early - warning thresholds and output hierarchical early - warning signals, including:
[0026] Receive the performance analysis report, divide the physical structure of the capacitor into a ten - by - ten grid - point array, collect the spatial position coordinates and real - time temperature values of each grid point in the grid - point array, construct spatial association edges based on the physical distance between the grid points, establish the connection relationship between grid points according to a preset distance threshold, and construct time - series samples in a sliding - window manner;
[0027] Based on the performance analysis report, construct a spatio - temporal dynamic graph neural network to model the device temperature field distribution. The spatio - temporal dynamic graph neural network includes a first graph convolutional layer, a second graph convolutional layer, and a third graph convolutional layer. The first graph convolutional layer performs spatial information aggregation to obtain the node weighted sum. The second graph convolutional layer controls the information flow through update gates and reset gates. The third graph convolutional layer generates candidate states and updates the node hidden state to obtain a dynamic prediction graph;
[0028] Construct a hierarchical dilated convolution structure, including a first dilated convolution branch, a second dilated convolution branch, and a third dilated convolution branch. The first dilated convolution branch uses a standard convolution kernel, the second dilated convolution branch uses a first dilated convolution kernel, and the third dilated convolution branch uses a second dilated convolution kernel. Input the output feature maps of the dilated convolution branches into the self - attention mechanism module to identify the evolution features corresponding to the capacitor temperature and stress in the dynamic prediction graph, and determine the temperature change trend and stress accumulation pattern;
[0029] Add the temperature change trend and the stress accumulation pattern to a multi-task learning framework, which includes a shared feature extraction layer, a temperature prediction branch, and a stress assessment branch. The temperature prediction branch uses a tabular data neural network to process the performance analysis report and the dynamic prediction graph, and the stress assessment branch uses a dynamic graph attention network to process the performance analysis report and the dynamic prediction graph to obtain a prediction state;
[0030] Construct a large language model with a mixture of experts architecture, which includes multiple expert modules and a dynamic routing mechanism. The dynamic routing mechanism calculates the allocation weights of each expert module based on the prediction state, determines the expert module for performing fault diagnosis according to the allocation weights, and generates a diagnosis result including potential fault types and corresponding occurrence probabilities;
[0031] Construct an uncertainty quantification module that integrates a probabilistic graphical model and Bayesian deep learning. The nodes in the probabilistic graphical model represent fault types, and the edge weights between the nodes represent the correlation degree between the fault types. Sampling is performed on the diagnosis result through a Bayesian deep learning framework to calculate the confidence interval corresponding to each fault type;
[0032] Construct a hierarchical early warning mechanism based on pre-acquired expert knowledge, divide the temperature parameters and stress parameters into multiple levels of early warning thresholds, and generate hierarchical early warning signals based on the fault type, the occurrence probability, and the confidence interval.
[0033] In an alternative embodiment,
[0034] The temperature prediction branch uses a tabular data neural network to process the performance analysis report and the dynamic prediction graph, and the stress assessment branch uses a dynamic graph attention network to process the performance analysis report and the dynamic prediction graph. The obtained prediction state includes:
[0035] Receive the performance analysis report and the dynamic prediction graph, perform standardized preprocessing on the tabular data in the performance analysis report to obtain preprocessed data, and the tabular data includes temperature parameters, voltage parameters, and current parameters;
[0036] Construct a tabular data neural network, the number of input layer nodes of the tabular data neural network corresponds to the feature dimension of the preprocessed data. The tabular data neural network includes a first hidden layer, a second hidden layer, and a third hidden layer. The first hidden layer and the second hidden layer use a rectified linear activation function and a dropout regularization layer, and the third hidden layer uses a hyperbolic tangent activation function. The dimensions of the first hidden layer, the second hidden layer, and the third hidden layer decrease layer by layer to form a funnel-shaped structure
[0037] Flatten the temperature field distribution information in the dynamic prediction graph in the time dimension and perform feature aggregation to obtain an extended feature vector, and perform dimensionality reduction processing on the extended feature vector through a two-layer fully connected network to obtain an extended feature representation that matches the feature dimension of the preprocessed data;
[0038] Construct a dynamic graph attention network, set the monitoring points as graph nodes, set the physical connection relationships between the monitoring points as graph edges, and fuse the stress-related indicators in the performance analysis report with the temperature distribution information in the dynamic prediction graph to form a node feature vector;
[0039] Construct a two-layer graph attention network structure. The first-layer graph attention network calculates the correlation degree between different nodes to obtain a feature transfer weight. The output of the feature transfer weight is processed through an exponential linear activation function and batch normalization. The second-layer graph attention network uses a multi-head attention mechanism to extract the interaction features between nodes. Construct a feature fusion module, adjust the output features of the tabular data neural network and the output features of the dynamic graph attention network to the same dimension through a dimension transformation network, and use a gating mechanism to adaptively fuse the output features to obtain a prediction state.
[0040] In an optional implementation manner,
[0041] Based on the performance analysis report and the hierarchical warning signal, construct a hierarchical reinforcement learning framework. Among them, the top-level policy network guides action selection according to the causal relationship, the middle-level coordination network evaluates the action impact based on the prediction state, and the bottom-level execution network performs device control to obtain an initial heat dissipation strategy. Add the initial heat dissipation strategy and the dynamic prediction graph to a hierarchical gated recurrent unit, combine a soft attention mechanism to perform temporal state evaluation, obtain system state evolution features, and combine a distributed advantage actor-critic algorithm to perform collaborative control on the heat dissipation unit. Combine a double deep Q network and priority experience replay to determine a candidate control action sequence, and optimize the control action based on the heat dissipation effect through a dynamic time warping algorithm and a deep deterministic policy gradient algorithm to obtain an optimal heat dissipation solution, including:
[0042] Receive a performance analysis report and a hierarchical warning signal. The performance analysis report includes temperature data streams, voltage data streams, and current data streams. The hierarchical warning signal includes temperature warning data and stress warning data. Perform causal analysis on the performance analysis report and the hierarchical warning signal through the top-level policy network to establish a state transition graph. The state transition graph records the mapping relationship from the device state to the heat dissipation effect, and the nodes of the state transition graph store state features;
[0043] Input the pre-determined causal relationship and prediction state data into the middle-layer coordination network. The prediction state data includes temperature field prediction data and stress prediction data. The state feature extractor of the middle-layer coordination network maps the temperature field data and stress data to the feature space to obtain feature data. The feature data is input into the action evaluation module, and the action evaluation module maintains a state-action evaluation table and dynamically updates it to obtain an action evaluation data stream;
[0044] The control instruction generator of the bottom-layer execution network receives the action evaluation data stream, queries the control instruction mapping table to generate an instruction data packet containing execution time, execution order, and control parameters, and parses the instruction data packet into a heat dissipation unit control signal through an instruction parser to form an initial heat dissipation strategy;
[0045] Input the initial heat dissipation strategy and dynamic prediction map data into the hierarchical gated recurrent unit. Extract the state features of the current time step through the input gate processor, filter the historical state information through the forget gate processor, integrate the current features and historical information through the output gate processor, calculate the weight coefficients of data at different time steps through the soft attention processor to generate a time series weighted feature stream, and update the system state based on the time series weighted feature stream through the state update processor to obtain the system state evolution feature;
[0046] Input the system state evolution feature into the distributed advantage actor-critic algorithm. Generate candidate action data and state value evaluation results through parallel actor-critic units, and collect learning results through an experience sharing processor and perform experience fusion to generate a cooperative control strategy;
[0047] Input the cooperative control strategy into the double deep Q network. Generate a Q-value estimation data stream through alternating processing of the main network and the target network, and allocate replay probabilities according to sample importance through the priority experience replay processor, and output a candidate control action sequence;
[0048] Input the candidate control action sequence and heat dissipation effect feedback data into the deep deterministic policy gradient algorithm. Generate a deterministic action through the policy network, evaluate the long-term benefit of the action through the value network, calculate the similarity of the control sequence through the dynamic time warping processor, and output the optimal heat dissipation plan through the policy optimization processor by combining the value evaluation and similarity scores.
[0049] In an alternative embodiment,
[0050] Input the cooperative control strategy into the double deep Q network. Generate a Q-value estimation data stream through alternating processing of the main network and the target network, and allocate replay probabilities according to sample importance through the priority experience replay processor. Outputting the candidate control action sequence includes:
[0051] Standardize the data flow of the collaborative control strategy, normalize the temperature data to a preset interval according to the preset maximum temperature value, normalize the stress data to a preset interval according to the preset maximum stress value, and normalize the state parameters of the heat dissipation unit according to the range to generate a standardized state vector;
[0052] Input the standardized state vector into the input layer of the main network, extract features through multiple neurons in the first hidden layer to obtain the first-layer feature vector, extract features of the first-layer feature vector through multiple neurons in the second hidden layer to obtain the second-layer feature vector, and generate a control action combination through multiple neurons in the output layer;
[0053] Input the standardized state vector into the target network. The target network has the same network structure as the main network. The target network updates its parameters according to the preset number of training iterations, and the parameter update method of the target network is to copy the parameter values from the main network;
[0054] Construct an experience replay pool to store training samples. The training samples include the current state vector, the executed action number, the reward value, the next state vector, and the termination flag, and calculate the time difference value and the reward amplitude value of the training samples;
[0055] Normalize the time difference value to a first preset interval to obtain a normalized time difference value, normalize the reward amplitude value to a second preset interval to obtain a normalized reward amplitude value, weight and sum the normalized time difference value and the normalized reward amplitude value according to a preset ratio to obtain a sample priority score, sort the training samples in the experience replay pool in descending order according to the sample priority score, assign a first preset sampling probability to the training samples with higher rankings, assign a second preset sampling probability to the training samples with middle rankings, and assign a third preset sampling probability to the training samples with lower rankings;
[0056] Randomly sample a preset number of the training samples from the experience replay pool based on the preset sampling probability to form a training batch, input the current state in the training batch into the main network to obtain an action value prediction, and input the next state in the training batch into the target network to obtain a target value;
[0057] Input the current state vector of the system into the trained main network to obtain an action value estimation score, sort the action value estimation scores in descending order, and select multiple actions with the highest scores to form a candidate control action sequence. Among them, each action in the candidate control action sequence includes the control parameters of the heat dissipation unit.
[0058] In the second aspect of the embodiments of the present invention, a thermal runaway prevention and heat dissipation optimization system based on a capacitor module is provided, including:
[0059] The first unit is used to collect multi-source operation data of individual capacitors in a capacitor module, segment different types of data into data segments of the same size and perform linear projection by combining an improved hierarchical visual attention network, extract features by combining a multi-level self-attention structure and an attention mechanism, introduce learnable position encoding to embed time information, obtain initial features and add them to a bidirectional probabilistic diffusion model, introduce conditional information during the diffusion process to complete the data to obtain complete data information and add it to a multi-modal feature fusion network, calculate the correlation between modalities and adaptively modify the fusion weights by combining a dynamic weight allocation mechanism, obtain fusion features and add them to a causal structure discovery network, construct an initial causal graph structure based on expert knowledge and identify the causal relationships between different data by combining a causal discovery algorithm, obtain temporal dependence relationships and perform uncertainty assessment, and obtain a performance analysis report corresponding to the individual capacitor;
[0060] The second unit is used to construct a spatio-temporal dynamic graph neural network based on the performance analysis report and model the device temperature field distribution, obtain a dynamic prediction graph and identify the evolution characteristics corresponding to the capacitor temperature and stress, determine the parameter evolution law based on the evolution characteristics by combining hierarchical dilated convolution and a self-attention mechanism, determine the temperature change trend and stress accumulation pattern and add them to a multi-task learning framework, process the performance analysis report and the dynamic prediction graph by combining a tabular data neural network and a dynamic graph attention network, obtain a predicted state, perform fault diagnosis through a large language model with a mixture of experts architecture and a dynamic routing mechanism, obtain potential fault types and corresponding occurrence probabilities, construct an uncertainty quantification module that combines a fusion probability graph model and Bayesian deep learning to calculate the confidence interval corresponding to each fault type, construct a hierarchical early warning mechanism by combining pre-acquired expert knowledge, set multi-level early warning thresholds and output hierarchical early warning signals;
[0061] The third unit is used to construct a hierarchical reinforcement learning framework based on the performance analysis report and the hierarchical early warning signals. Among them, the top-level policy network guides action selection according to the causal relationship, the middle-level coordination network evaluates the action impact based on the predicted state, and the bottom-level execution network performs device control to obtain an initial heat dissipation strategy. Add the initial heat dissipation strategy and the dynamic prediction graph to a hierarchical gated recurrent unit, perform temporal state evaluation by combining a soft attention mechanism, obtain system state evolution characteristics and perform collaborative control on the heat dissipation unit by combining a distributed advantage actor-critic algorithm, determine a candidate control action sequence by combining a double deep Q-network and prioritized experience replay, and optimize the control action based on the heat dissipation effect through a dynamic time warping algorithm and a deep deterministic policy gradient algorithm to obtain an optimal heat dissipation solution.
[0062] In the third aspect of the embodiments of the present invention,
[0063] A kind of electronic device is provided, including:
[0064] A processor;
[0065] A memory for storing instructions executable by the processor;
[0066] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0067] In the fourth aspect of the embodiments of the present invention,
[0068] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0069] In the present invention, through multi-source data analysis, causal structure discovery, and advanced prediction models, it is possible to accurately identify the potential fault types and their occurrence probabilities of individual capacitors in a capacitor module, and provide hierarchical warning signals, thereby effectively preventing the occurrence of thermal runaway, improving system safety. By adopting a hierarchical reinforcement learning framework, combining real-time state evaluation and multi-objective optimization algorithms, it is possible to dynamically adjust the heat dissipation strategy, achieve coordinated control of heat dissipation units, and finally obtain an optimal heat dissipation solution, improve heat dissipation efficiency, and extend the life of the capacitor module. By using improved deep learning models and graph neural network technologies, it is possible to efficiently process multi-source heterogeneous data, extract key features, and perform accurate fault diagnosis and prediction, providing a scientific decision-making basis for system operation and maintenance, and reducing maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a schematic flowchart of the method for preventing thermal runaway and optimizing heat dissipation of the capacitor module according to the embodiments of the present invention;
[0071] Figure 2 It is a schematic structural diagram of the system for preventing thermal runaway and optimizing heat dissipation of the capacitor module according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0073] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0074] Figure 1 This is a schematic flowchart of the method for preventing thermal runaway and optimizing heat dissipation based on a capacitor module according to an embodiment of the present invention. As Figure 1 shown, the method includes:
[0075] S1. Collect multi-source operation data of single capacitors in the capacitor module, segment different types of data into data segments of the same size and perform linear projection by combining an improved hierarchical visual attention network, extract features by combining a multi-level self-attention structure and an attention mechanism, introduce learnable positional encoding to embed time information, obtain initial features and add them to a bidirectional probability diffusion model, introduce conditional information during the diffusion process for data completion to obtain complete data information and add it to a multi-modal feature fusion network, calculate the correlation between modalities and adaptively modify the fusion weights by combining a dynamic weight allocation mechanism, obtain fusion features and add them to a causal structure discovery network, construct an initial causal graph structure based on expert knowledge and identify the causal relationships between different data by combining a causal discovery algorithm, obtain a temporal dependence relationship and perform uncertainty assessment, and obtain a performance analysis report corresponding to the single capacitor;
[0076] The single capacitor is an electronic component, usually composed of two electrodes and a dielectric (such as air, ceramic or electrolyte), and is used to store electrical energy. The hierarchical visual attention network is a neural network architecture that enhances image understanding ability through multi-level attention mechanisms. The learnable positional encoding is a method of embedding positional information into a neural network, usually used to process sequence data. The bidirectional probability diffusion model is a generative model that simulates the data generation and reverse denoising processes through a diffusion process. The causal structure discovery network is a machine learning model used to discover potential causal relationships from data. By modeling the dependence relationships between variables, the network can reveal causal relationships, thereby helping to infer the mechanisms behind the data and predict future events.
[0077] In an alternative embodiment,
[0078] Collect multi-source operation data of individual capacitors in a capacitor module, segment different types of data into data segments of the same size and perform linear projection by combining with an improved hierarchical visual attention network, extract features by combining a multi-level self-attention structure and an attention mechanism, introduce learnable positional encoding to embed time information, obtain initial features and add them to a bidirectional probability diffusion model, introduce conditional information during the diffusion process to complete data and obtain complete data information and add it to a multi-modal feature fusion network, calculate the correlation between modalities and adaptively modify the fusion weights by combining with a dynamic weight allocation mechanism, obtain fusion features and add them to a causal structure discovery network, construct an initial causal graph structure based on expert knowledge and identify the causal relationship between different data by combining with a causal discovery algorithm, obtain a time series dependence relationship and perform uncertainty assessment, and obtain the performance analysis report corresponding to the individual capacitor, including:
[0079] Collect the voltage data, current data, ambient temperature data, housing temperature data, and charge and discharge state data of individual capacitors in the capacitor module and combine them to obtain multi-source operation data. Input the multi-source operation data into an improved hierarchical visual attention network. The improved hierarchical visual attention network includes a feature segmentation layer, a feature projection layer, and a feature enhancement layer. Among them, the feature segmentation layer uses an overlapping sliding window to segment the multi-source operation data to obtain data segments. The feature projection layer uses an independent projection matrix to map the data segments to a unified feature space to obtain mapped features. The feature enhancement layer enhances the mapped features through a channel-level attention mechanism and a spatial-level attention mechanism to obtain enhanced features;
[0080] Input the enhanced features into a multi-level self-attention structure. The multi-level self-attention structure includes a local attention unit and a global attention unit. The local attention unit uses a multi-head attention mechanism to extract local dependence features. The global attention unit uses a sparse attention mechanism to extract cross-segment association features. Embed the absolute position information and relative position information into the local dependence features and the cross-segment association features through a learnable position encoder to obtain initial features;
[0081] Input the initial features into a bidirectional probability diffusion model. During the training stage, construct a noise schedule to control the noise addition rate and train a denoising predictor. During the inference stage, convert the working condition information into a conditional vector, and complete the data through multiple steps of iteration in combination with the conditional vector to obtain complete data information and input it into a multi-modal feature fusion network. Standardize different modal features to obtain standardized features, construct an attention matrix to calculate the inter-modal correlation of the standardized features to obtain correlation features, perform dynamic weight allocation on the correlation features based on a soft attention mechanism and introduce a cross-modal contrast learning strategy to obtain fusion features;
[0082] Construct an initial causal graph based on expert knowledge, including node attribute definitions and edge connection constraints. Input the fusion features into the causal structure discovery network, extract the trend features, periodic features, and mutation features of the time series feature sequence, evaluate the linear and non-linear correlations between different feature sequences, determine the direction and weight of the causal edges based on the dependence strength to obtain the causal relationship, extract the feature description sequence of the causal relationship under different time windows, generate multiple sets of evaluation samples by the bootstrap method and calculate the confidence scores, adopt an integration strategy to fuse different evaluation results to obtain the final confidence score, and generate a performance analysis report based on the final confidence score.
[0083] The independent projection matrix refers to a matrix used in mathematics and computing to project data from one space to another while preserving the independence of the data during the projection process. The sparse attention mechanism is a technique for reducing computational complexity, which improves computational efficiency by restricting the attention scope of each attention unit (for example, only focusing on a small part of the relevant inputs). The mutation feature refers to a feature that exhibits significant changes or jumps in a dataset, representing mutation points or outliers in time series data or classification tasks, and can reflect the critical turning points or important events of the system. The bootstrap method is a statistical method that generates multiple subsamples by randomly sampling from the original dataset and conducts statistical analysis based on these subsamples.
[0084] Collect multi-source operation data of individual capacitors in the capacitor module, including voltage data, current data, ambient temperature data, housing temperature data, and charge and discharge state data. For example, the above five types of data can be collected once every second, and these five types of data are combined to form a data point containing five dimensions. After continuous collection for a period of time, the multi-source operation data of the individual capacitor is obtained.
[0085] Process the collected multi-source operation data using an improved hierarchical visual attention network. The improved hierarchical visual attention network includes three sub-layers: a feature segmentation layer, a feature projection layer, and a feature enhancement layer. In the feature segmentation layer, the multi-source operation data is segmented using an overlapping sliding window. For example, a sliding window with a length of 60 seconds and a step size of 30 seconds is used to segment the multi-source operation data into multiple data segments with a length of 60 seconds. In the feature projection layer, the data points in each data segment are mapped to a unified feature space through an independent projection matrix. Each projection matrix is a learnable parameter, and through training and optimization, different types of data are converted to the same feature dimension. In the feature enhancement layer, the mapped features are enhanced using a channel-level attention mechanism and a spatial-level attention mechanism. The channel-level attention mechanism focuses on the importance between different data channels, and the spatial-level attention mechanism focuses on the importance of different time steps in the data segment. Through these two attention mechanisms, important feature information can be highlighted and the influence of noise can be suppressed.
[0086] The enhanced features are input into a multi-level self-attention structure, which includes a local attention unit and a global attention unit. The local attention unit uses the multi-head attention mechanism to extract local dependence features, focusing on the relationships between different time steps within a data segment. The global attention unit uses the sparse attention mechanism to extract cross-segment association features, focusing on the correlations between different data segments. In addition, the absolute position information and relative position information are embedded into the local dependence features and cross-segment association features through a learnable position encoder, enabling the model to learn temporal information. Finally, the initial feature representation of the single capacitor is obtained.
[0087] The initial features are input into a bidirectional probabilistic diffusion model. During the model training phase, a noise schedule is constructed to control the noise addition rate, and a denoising predictor is trained. During the inference phase, the operating conditions information, such as the load size, ambient temperature, etc., is converted into a conditional vector. Combining the conditional vector, data completion is performed through multi-step iteration to obtain complete data information. For example, if the voltage data at certain time points is missing, it can be completed through the diffusion model.
[0088] The complete data information is input into a multi-modal feature fusion network, where the features of different modalities are normalized, for example, using Z-score normalization, and an attention matrix is constructed to calculate the inter-modal correlations of the normalized features. For example, the correlation between the voltage feature and the current feature is calculated. Based on the soft attention mechanism, dynamic weight allocation is performed on the correlated features, and a cross-modal contrast learning strategy is introduced to obtain the fused features. In this way, the information of different modalities can be effectively fused together.
[0089] Construct an initial causal graph based on expert knowledge. The initial causal graph includes node attribute definitions and edge connection constraints. For example, define nodes such as voltage, current, temperature, etc., and constrain that there is a causal relationship between voltage and current. Input the fused features into the causal structure discovery network to extract the trend features, periodic features, and mutation features of the time series feature sequence. For example, extract the rising trend, periodic fluctuations, and mutation points of the voltage sequence. Evaluate the linear and non-linear correlations between different feature sequences. For example, calculate the Pearson correlation coefficient and mutual information between the voltage sequence and the current sequence. Determine the direction and weight of the causal edge based on the dependence strength to obtain the causal relationship. For example, if the voltage change leads the current change, then determine that voltage is the cause of current. Extract the feature description sequence of the causal relationship under different time windows. For example, within the past hour, the increase in voltage causes the increase in current. Generate multiple groups of evaluation samples by the bootstrap method and calculate the confidence scores. For example, randomly divide the data into multiple groups, conduct causal analysis separately, and calculate the confidence of each group's results. Adopt an integration strategy to fuse different evaluation results to obtain the final confidence score. Generate a performance analysis report based on the final confidence score. For example, report information such as the health status and remaining life of a single capacitor.
[0090] In this embodiment, through multi-source data fusion and causal structure discovery, the operating state of a single capacitor can be more comprehensively understood, thereby improving the accuracy of performance analysis. Through data completion and uncertainty evaluation, the impacts of data missing and noise can be effectively handled, enhancing the robustness of the model. Through the improved hierarchical visual attention network and multi-level self-attention structure, feature information can be efficiently extracted, improving the analysis efficiency.
[0091] In an alternative embodiment,
[0092] The feature segmentation layer uses an overlapping sliding window to segment the multi-source operation data to obtain data segments. The feature projection layer uses an independent projection matrix to map the data segments to a unified feature space to obtain mapped features. The feature enhancement layer enhances the mapped features through a channel-level attention mechanism and a spatial-level attention mechanism to obtain enhanced features, including:
[0093] Obtain the multi-source operation data and construct an adaptive segmentation mechanism. Detect the change rate of the multi-source operation data through a data change rate detection module, and dynamically adjust the sliding window parameters according to the change rate. Among them, when the change rate is higher than the first threshold, reduce the window length of the sliding window parameters to 80% of the reference window length. When the change rate is lower than the second threshold, increase the window length of the sliding window parameters to 120% of the reference window length.
[0094] The multi-source operation data is processed in parallel using sliding windows of different lengths. Among them, the first sliding window is used to capture transient change features to obtain a first feature segment, the second sliding window is used to extract local trend features to obtain a second feature segment, and the third sliding window is used to acquire macroscopic evolution features to obtain a third feature segment. The first feature segment, the second feature segment, and the third feature segment are input into a feature pyramid network for integration to obtain a data segment.
[0095] The original feature space is divided into multiple subspaces. Each subspace is configured with a main projection matrix and an auxiliary projection matrix. The data segment is input into the main projection matrix to obtain basic features, and the data segment is input into the auxiliary projection matrix to obtain residual features. The basic features and the residual features are fused in terms of features to obtain mapped features.
[0096] A feature quality evaluation index is constructed. The feature quality evaluation index includes feature discriminability, feature stability, and information retention. Based on the feature quality evaluation index, the mapped features are evaluated to obtain an evaluation result, and the main projection matrix and the auxiliary projection matrix are optimized according to the evaluation result.
[0097] A hierarchical attention mechanism is constructed, including a feature channel attention mechanism, a time dimension attention mechanism, and a cross-modal attention mechanism. Among them, the feature channel attention mechanism calculates the inter-channel dependence relationship of the mapped features to obtain a channel attention map, extracts the channel feature statistical information of the mapped features to obtain channel importance, and adaptively weights and fuses the channel attention map and the channel importance to obtain channel attention weights.
[0098] The time dimension attention mechanism extracts features of different time scales through a multi-head attention mechanism, and uses a deformable convolutional network to process time series to obtain a time dimension feature enhancement result. The cross-modal attention mechanism constructs different types of features into a graph structure and performs feature interaction through a message passing mechanism to obtain enhanced features.
[0099] The deformable convolutional network is a variant proposed on the basis of a traditional convolutional neural network, in which the shape and size of the convolutional kernel can be dynamically adjusted according to the characteristics of the input data.
[0100] Obtain multi-source operation data and construct an adaptive segmentation mechanism. The core of the adaptive segmentation mechanism is a data change rate detection module. This module continuously monitors the change rate of multi-source operation data and dynamically adjusts the parameters of the subsequent sliding window according to the change rate. For example, the benchmark window length is set to 100 data points. When the detected data change rate is higher than the preset first threshold (e.g., the change rate exceeds 0.05), the sliding window length is reduced to 80% of the benchmark window length, i.e., 80 data points, in order to more finely capture the characteristics of rapid changes. Conversely, when the data change rate is lower than the preset second threshold (e.g., the change rate is lower than 0.01), the sliding window length is increased to 120% of the benchmark window length, i.e., 120 data points, in order to better extract the characteristics of long-term trends.
[0101] Perform parallel processing on multi-source operation data using sliding windows of different lengths. For example, use a first sliding window with a length of 50 data points to capture transient change characteristics and obtain a first feature segment; use a second sliding window with a length of 100 data points to extract local trend characteristics and obtain a second feature segment; use a third sliding window with a length of 200 data points to obtain macro-evolution characteristics and obtain a third feature segment. Assume that the multi-source operation data includes three dimensions: temperature, pressure, and flow rate. Each sliding window extracts corresponding features in each dimension. Then, input these three feature segments into a feature pyramid network for integration, and finally obtain a data segment that fuses multi-scale information. The feature pyramid network finally obtains a unified data segment containing global and local information by upsampling layer by layer and fusing features of different scales.
[0102] Divide the original feature space into multiple subspaces, and configure a main projection matrix and an auxiliary projection matrix for each subspace. Assume that the dimension of the original feature space is 60, and it is divided into 3 subspaces, each with a dimension of 20. Input the data segment obtained in the previous step into the main projection matrix to obtain basic features; at the same time, input the data segment into the auxiliary projection matrix to obtain residual features. For example, the main projection matrix maps the 20-dimensional subspace features to 10 dimensions, and the auxiliary projection matrix also maps the 20-dimensional subspace features to 5 dimensions. Then, splice and fuse the 10-dimensional basic features and the 5-dimensional residual features to obtain 15-dimensional mapped features.
[0103] In order to optimize the projection matrix, construct feature quality evaluation indicators, including feature discrimination, feature stability, and information retention. For example, use mutual information to evaluate feature discrimination, use feature variance to evaluate feature stability, and use reconstruction error to evaluate information retention. Evaluate the mapped features based on these indicators, and optimize the main projection matrix and the auxiliary projection matrix according to the evaluation results. For example, use the gradient descent method to update the parameters of the projection matrix to maximize feature discrimination and stability, while minimizing the loss of information retention.
[0104] A hierarchical attention mechanism is constructed to enhance the mapped features. The hierarchical attention mechanism includes a feature channel attention mechanism, a time dimension attention mechanism, and a cross-modal attention mechanism. The feature channel attention mechanism calculates the inter-channel dependencies of the mapped features to obtain a channel attention map, and extracts channel feature statistical information, such as mean and variance, to obtain channel importance. Then, the channel attention map and the channel importance are adaptively weighted and fused to obtain channel attention weights. The time dimension attention mechanism extracts features at different time scales through a multi-head attention mechanism and processes the time series using a deformable convolutional network to obtain the enhanced result of the time dimension features. The cross-modal attention mechanism constructs different types of features into a graph structure and performs feature interaction through a message passing mechanism to obtain the final enhanced features.
[0105] In this embodiment, through the adaptive segmentation mechanism and multi-scale feature extraction, complex change patterns in the data can be captured more comprehensively. Through feature space projection and enhancement, the discriminability, stability, and information content of the features are improved. Through feature quality evaluation and projection matrix optimization, feature redundancy and noise interference can be effectively reduced, and the generalization ability of the model can be improved, enabling it to better adapt to data under different working conditions. The hierarchical attention mechanism can effectively fuse feature information from different channels, different time scales, and different modalities, thereby enhancing the model's ability to perceive and understand the operating state of complex systems.
[0106] S2. Based on the performance analysis report, a spatio-temporal dynamic graph neural network is constructed to model the device temperature field distribution, obtaining a dynamic prediction graph and identifying the evolution characteristics corresponding to the capacitor temperature and stress. Based on the evolution characteristics, combined with hierarchical dilated convolution and self-attention mechanism, the parameter evolution law is determined. The temperature change trend and stress accumulation pattern are determined and added to the multi-task learning framework. The performance analysis report and the dynamic prediction graph are processed by combining a tabular data neural network and a dynamic graph attention network to obtain a predicted state. Fault diagnosis is performed through a large language model with a mixture of experts architecture and a dynamic routing mechanism to obtain potential fault types and corresponding occurrence probabilities. An uncertainty quantification module that combines a probability graph model and Bayesian deep learning is constructed to calculate the confidence interval corresponding to each fault type. A hierarchical early warning mechanism is constructed by combining the pre-acquired expert knowledge, multi-level early warning thresholds are set, and a graded early warning signal is output;
[0107] The spatio-temporal dynamic graph neural network is a neural network model for processing spatio-temporal data, specifically designed to capture the dynamic relationships of data changing over time and space. By dynamically adjusting the network structure, it can handle the temporal dependencies and spatial correlations in spatio-temporal data and is widely used in fields such as video analysis and traffic prediction. The hierarchical dilated convolution is a method to expand the traditional convolution operation. By adding dilation factors to the convolution kernel, it can increase the receptive field while keeping the computational complexity relatively low. The tabular data neural network is a neural network architecture specifically designed to process structured tabular data (such as spreadsheets or databases). By combining the characteristics of neural networks and tabular data, it can automatically extract features and make effective predictions, and is widely used in data analysis in the fields of finance, healthcare, and business. The dynamic graph attention network is a network model that combines the attention mechanism and the dynamic graph structure. It can dynamically adjust the focus according to different input data, focusing on important time steps or nodes, and performs well in tasks such as graph data analysis and social network analysis. The mixture of experts architecture is a structure for task processing through the combination of multiple "expert" models. Each expert has specific advantages in different subtasks, and the system can improve the overall task processing efficiency by selecting appropriate experts. The dynamic routing mechanism is a strategy for realizing information flow and data transfer in neural networks. By dynamically selecting the "routing" paths in the network, it can improve the efficiency and performance of the model, especially when dealing with complex multi-modal data. Bayesian deep learning is a framework that combines Bayesian inference and deep learning, and can model the uncertainty of the model during the training process.
[0108] In an alternative embodiment,
[0109] Based on the performance analysis report, construct a spatio-temporal dynamic graph neural network and model the device temperature field distribution to obtain a dynamic prediction graph and identify the evolution characteristics corresponding to the capacitor temperature and stress. Based on the evolution characteristics, combine the hierarchical dilated convolution and the self-attention mechanism to determine the parameter evolution law, determine the temperature change trend and stress accumulation pattern and add them to the multi-task learning framework. Combine the tabular data neural network and the dynamic graph attention network to process the performance analysis report and the dynamic prediction graph to obtain the predicted state. Use a large language model with a mixture of experts architecture and a dynamic routing mechanism for fault diagnosis to obtain the potential fault types and the corresponding occurrence probabilities. Construct an uncertainty quantification module that combines a probabilistic graph model and Bayesian deep learning to calculate the confidence interval corresponding to each fault type. Combine the pre-acquired expert knowledge to construct a hierarchical early warning mechanism, set multiple early warning thresholds and output graded early warning signals including:
[0110] Receive a performance analysis report, divide the physical structure of the capacitor into a ten-by-ten grid point array, collect the spatial position coordinates and real-time temperature values of each grid point in the grid point array, construct spatial association edges based on the physical distances between the grid points, establish connection relationships between grid points according to a preset distance threshold, and construct time series samples in a sliding window manner;
[0111] Construct a spatio-temporal dynamic graph neural network based on the performance analysis report to model the device temperature field distribution. The spatio-temporal dynamic graph neural network includes a first graph convolutional layer, a second graph convolutional layer, and a third graph convolutional layer. The first graph convolutional layer performs spatial information aggregation to obtain a node weighted sum. The second graph convolutional layer controls information flow through an update gate and a reset gate. The third graph convolutional layer generates candidate states and updates the node hidden states to obtain a dynamic prediction graph;
[0112] Construct a hierarchical dilated convolutional structure, including a first dilated convolution branch, a second dilated convolution branch, and a third dilated convolution branch. The first dilated convolution branch uses a standard convolution kernel. The second dilated convolution branch uses a first dilated convolution kernel. The third dilated convolution branch uses a second dilated convolution kernel. Input the output feature maps of the dilated convolution branches into a self-attention mechanism module to identify the evolution features corresponding to the capacitor temperature and stress in the dynamic prediction graph, and determine the temperature change trend and stress accumulation pattern;
[0113] Add the temperature change trend and the stress accumulation pattern to a multi-task learning framework. The multi-task learning framework includes a shared feature extraction layer, a temperature prediction branch, and a stress assessment branch. The temperature prediction branch uses a tabular data neural network to process the performance analysis report and the dynamic prediction graph. The stress assessment branch uses a dynamic graph attention network to process the performance analysis report and the dynamic prediction graph to obtain a predicted state;
[0114] Construct a large language model with a mixture of experts architecture. The large language model includes multiple expert modules and a dynamic routing mechanism. The dynamic routing mechanism calculates the allocation weights of each expert module based on the predicted state, determines the expert module for performing fault diagnosis according to the allocation weights, and generates a diagnostic result including potential fault types and corresponding occurrence probabilities;
[0115] Construct an uncertainty quantification module that fuses a probabilistic graphical model and Bayesian deep learning. The nodes in the probabilistic graphical model represent fault types, and the edge weights between nodes represent the correlation degrees between fault types. Sample the diagnostic results through a Bayesian deep learning framework and calculate the confidence intervals corresponding to each fault type;
[0116] Construct a hierarchical early warning mechanism based on pre-acquired expert knowledge, divide the temperature parameter and stress parameter into multiple levels of early warning thresholds, and generate hierarchical early warning signals based on the fault type, the occurrence probability, and the confidence interval.
[0117] The grid point array is a two-dimensional or three-dimensional lattice structure for data representation, usually used to process spatial data such as images, geographical information, etc. It can accurately represent the characteristic values of each position and is the basis of many computer vision and image processing algorithms. The dilated convolution is a technique that introduces gaps in the convolution operation, captures a larger range of context information by increasing the receptive field of the convolution kernel, and is particularly suitable for processing tasks with long-range dependencies such as speech signal processing, image segmentation, etc.
[0118] Receive the performance analysis report of the capacitor to be analyzed and discretize the physical structure of the capacitor. For the convenience of analysis and modeling, the physical structure of the capacitor is divided into a 10x10 grid point array. For each grid point in the array, collect its spatial position coordinates (x, y) and the real-time temperature value T. For example, the coordinates of a certain grid point can be (2, 5), and the temperature value is 25°C. Then, construct spatial association edges based on the physical distance between grid points. Set a distance threshold, for example, 1.5. If the distance between two grid points is less than or equal to this threshold, establish a connection relationship between these two grid points to form a spatial graph structure. To capture the temporal variation of temperature, use a sliding window method to construct temporal samples. For example, use a time window with a length of 10, and take the grid point data of 10 consecutive moments as a sample.
[0119] Construct a spatio-temporal dynamic graph neural network to model the temperature field distribution of the device, so as to obtain a dynamic prediction graph, which includes three graph convolutional layers. The first graph convolutional layer performs spatial information aggregation, calculates the weighted sum of the temperature values of the neighbor nodes of each node, and the weights are determined by the connection relationship between nodes and the temperature difference. The second graph convolutional layer introduces an update gate and a reset gate mechanism to control the flow of information, similar to the gating mechanism in the recurrent neural network, and selectively retains and updates information. The third graph convolutional layer generates candidate states and updates the hidden state of the nodes, and finally outputs the temperature prediction values of each grid point at future moments to form a dynamic prediction graph. For example, the temperature values of each grid point at the next 5 moments can be predicted.
[0120] Construct a hierarchical dilated convolution structure and a self-attention mechanism to identify the evolution characteristics of capacitor temperature and stress in the dynamic prediction map. The dilated convolution structure consists of three branches, each using a convolution kernel of a different size. The first branch uses a standard convolution kernel, such as 3x3, to extract local features. The second branch uses the first dilated convolution kernel, such as 3x3, with a dilation rate of 2 to expand the receptive field and capture spatial information over a larger range. The third branch uses the second dilated convolution kernel, such as 3x3, with a dilation rate of 4 to further expand the receptive field. The output feature maps of the three branches are input into the self-attention mechanism module. The self-attention mechanism identifies the change patterns of temperature and stress over time and space by calculating the correlations between different positions in the feature map, such as the diffusion trend of the temperature rising area or the evolution law of the stress concentration area. Finally, the temperature change trend and stress accumulation pattern are determined.
[0121] Add the extracted temperature change trend and stress accumulation pattern to a multi-task learning framework, which includes a shared feature extraction layer for extracting common feature representations from the performance analysis report and the dynamic prediction map. The framework also includes two branches: a temperature prediction branch and a stress assessment branch. The temperature prediction branch uses a tabular data neural network to process temperature-related information in the performance analysis report and the dynamic prediction map to predict future temperatures. The stress assessment branch uses a dynamic graph attention network to process stress-related information in the performance analysis report and the dynamic prediction map to evaluate the current stress state. For example, the temperature prediction branch can predict the average temperature at the next 10 time instants, and the stress assessment branch can evaluate the current stress level, such as low, medium, or high.
[0122] Construct a large language model with a mixture of experts architecture for fault diagnosis. The large language model contains multiple expert modules, each focusing on a specific type of fault diagnosis. The dynamic routing mechanism calculates the assignment weights for each expert module based on the prediction state. For example, if the prediction state indicates that the temperature is too high, the weight of the expert module responsible for overheating fault diagnosis will be higher. Based on the assignment weights, one or more expert modules are selected to perform fault diagnosis and generate a diagnosis result containing potential fault types and their corresponding occurrence probabilities. For example, the diagnosis result may indicate that the probability of an "overheating" fault is 80%, and the probability of a "liquid leakage" fault is 15%.
[0123] Construct an uncertainty quantification module that fuses a probabilistic graphical model and Bayesian deep learning. The nodes in the probabilistic graphical model represent different fault types, and the edge weights between the nodes represent the degree of correlation between the fault types. For example, there may be a strong correlation between "overheating" and "short circuit" faults. Samples are taken from the diagnosis results through the Bayesian deep learning framework to calculate the confidence interval corresponding to each fault type. For example, if the probability of an "overheating" fault is 80%, its 95% confidence interval may be [70%, 90%].
[0124] Build a hierarchical early warning mechanism based on pre-acquired expert knowledge. Divide the temperature parameter and stress parameter into multiple levels of early warning thresholds. For example, the temperature early warning thresholds can be set as 40°C, 50°C, and 60°C, corresponding to the low, medium, and high early warning levels respectively. Generate hierarchical early warning signals based on the fault type, occurrence probability, and confidence interval. For example, if the probability of the "overheat" fault is 80%, the confidence interval is [70%, 90%], and the current temperature exceeds 50°C, then issue a medium-level early warning signal.
[0125] In this embodiment, a variety of advanced technologies such as spatio-temporal dynamic graph neural network, dilated convolution, self-attention mechanism, multi-task learning, mixture of experts architecture, and Bayesian deep learning are combined, which can more accurately predict the temperature field distribution of the capacitor, identify evolution characteristics, diagnose potential faults, and quantify the uncertainty of the diagnosis results, improving the accuracy and reliability of fault diagnosis. Through the hierarchical early warning mechanism and the setting of multi-level early warning thresholds, different levels of early warning signals can be issued in a timely manner according to the fault type, occurrence probability, and confidence interval, providing sufficient time to take preventive measures in advance to avoid accidents. The application of the probabilistic graphical model and the Bayesian deep learning framework makes the fault diagnosis results more transparent and interpretable, and can provide the correlation information between fault types and the confidence interval of each fault type, which helps to better understand the mechanism and risk level of fault occurrence.
[0126] In an alternative embodiment,
[0127] The temperature prediction branch uses a tabular data neural network to process the performance analysis report and the dynamic prediction graph, and the stress assessment branch uses a dynamic graph attention network to process the performance analysis report and the dynamic prediction graph. The obtained prediction states include:
[0128] Receive the performance analysis report and the dynamic prediction graph, perform standardized preprocessing on the tabular data in the performance analysis report to obtain preprocessed data, and the tabular data includes temperature parameters, voltage parameters, and current parameters;
[0129] Build a tabular data neural network, the number of input layer nodes of the tabular data neural network corresponds to the feature dimension of the preprocessed data. The tabular data neural network includes a first hidden layer, a second hidden layer, and a third hidden layer. The first hidden layer and the second hidden layer use the rectified linear activation function and the dropout regularization layer, and the third hidden layer uses the hyperbolic tangent activation function. The dimensions of the first hidden layer, the second hidden layer, and the third hidden layer decrease layer by layer to form a funnel-shaped structure
[0130] Flatten the temperature field distribution information in the dynamic prediction graph in the time dimension and perform feature aggregation to obtain an extended feature vector. Perform dimensionality reduction on the extended feature vector through a two-layer fully connected network to obtain an extended feature representation that matches the feature dimension of the preprocessed data;
[0131] Construct a dynamic graph attention network, set the monitoring points as graph nodes, set the physical connection relationships between the monitoring points as graph edges, and fuse the stress-related indicators in the performance analysis report and the temperature distribution information in the dynamic prediction graph to form a node feature vector;
[0132] Construct a two-layer graph attention network structure. The first-layer graph attention network calculates the degree of association between different nodes to obtain a feature transfer weight. The output of the feature transfer weight is processed through an exponential linear activation function and batch normalization. The second-layer graph attention network uses a multi-head attention mechanism to extract the interaction features between nodes, construct a feature fusion module, adjust the output features of the tabular data neural network and the output features of the dynamic graph attention network to the same dimension through a dimension transformation network, and use a gating mechanism to adaptively fuse the output features to obtain a prediction state.
[0133] The time dimension flattening is a technique that converts time series data into a one-dimensional vector and is commonly used to process data with time dependencies, such as video processing and dynamic time series analysis. The rectified linear activation function is an activation function that outputs the value when the input is positive and outputs zero when the input is negative, which can effectively solve the gradient vanishing problem and accelerate the training of neural networks and is commonly used in deep learning models.
[0134] Obtain the performance analysis report and dynamic prediction graph of the device. The performance analysis report contains various parameters during the operation of the device and is stored in tabular form, such as temperature, voltage, current, etc. The dynamic prediction graph shows the temperature field distribution information of the key components of the device over time. Here, taking a power transformer as an example, the performance analysis report includes temperature parameters such as transformer winding and oil temperature, as well as parameters such as operating voltage and current. The dynamic prediction graph shows the change of temperature distribution at different positions inside the transformer over time.
[0135] Preprocess the tabular data in the performance analysis report. To eliminate the influence of different parameter dimensions, standardize the tabular data. For example, use the min-max standardization method to scale the value of each parameter to between 0 and 1. Suppose the original value of the transformer winding temperature is 80 °C, the lowest value is 60 °C, and the highest value is 100 °C, then the standardized value is (80 - 60) / (100 - 60) = 0.5. Similar standardization processing is also performed on other parameters such as voltage and current.
[0136] Construct a neural network for tabular data. The number of input layer nodes of the neural network for tabular data corresponds to the dimension of the data features after preprocessing. Suppose the preprocessed data contains three features: temperature, voltage, and current, then the number of input layer nodes is 3. The network structure includes three hidden layers. The first and second layers use the rectified linear activation function and the dropout regularization layer to enhance the generalization ability of the model and prevent overfitting. The third layer uses the hyperbolic tangent activation function. The dimensions of these three hidden layers decrease layer by layer, forming a funnel-like structure. For example, the dimension of the first layer is 128, the dimension of the second layer is 64, and the dimension of the third layer is 32.
[0137] Meanwhile, process the dynamic prediction graph. Unfold the temperature field distribution information in the dynamic prediction graph along the time dimension, and aggregate the temperature distribution features of each time step to form an extended feature vector. For example, take the temperature distribution every 10 minutes as a time step, and unfold the temperature distribution data of one day into a sequence containing 24 * 6 = 144 time steps. Then, perform an average pooling operation on the temperature distribution features of each time step to obtain a feature vector representing the temperature distribution of that time step. Finally, concatenate the feature vectors of all time steps to form an extended feature vector.
[0138] To match the input dimension of the neural network for tabular data, use a two-layer fully connected network to reduce the dimension of the extended feature vector. For example, reduce the 144-dimensional extended feature vector to 3 dimensions.
[0139] Construct a dynamic graph attention network. Set the monitoring points as graph nodes and the physical connection relationships between the monitoring points as graph edges. Fuse the stress-related indicators in the performance analysis report with the temperature distribution information in the dynamic prediction graph to form a node feature vector. For example, fuse the stress indicator of the transformer winding and the temperature information at the corresponding position as the feature vector of that winding node.
[0140] This network adopts a two-layer graph attention network structure. The first layer calculates the correlation degree between different nodes to obtain the feature transfer weights. The output of the feature transfer weights is processed through the exponential linear activation function and batch normalization. The second layer uses the multi-head attention mechanism to extract the interaction features between nodes.
[0141] Construct a feature fusion module. Adjust the output features of the neural network for tabular data and the output features of the dynamic graph attention network to the same dimension through a dimension transformation network. For example, adjust the output features of both networks to 16 dimensions. Then, adopt a gating mechanism to adaptively fuse the adjusted features to obtain the final prediction state.
[0142] In this embodiment, by integrating the tabular data in the performance analysis report and the temperature field distribution information in the dynamic prediction graph, it can more comprehensively reflect the operating state of the device, thereby improving the prediction accuracy. By adopting a variety of neural network structures and regularization techniques, such as dropout and batch normalization, it can effectively prevent the model from overfitting and improve the robustness of the model. By adopting the graph attention network structure, it can effectively extract the interaction features between nodes, reduce the computational complexity, and improve the prediction efficiency.
[0143] S3. Construct a hierarchical reinforcement learning framework based on the performance analysis report and the hierarchical warning signals. Among them, the top-level policy network guides action selection according to the causal relationship, the middle-level coordination network evaluates the action impact based on the predicted state, and the bottom-level execution network performs device control to obtain an initial heat dissipation strategy. Add the initial heat dissipation strategy and the dynamic prediction graph to the hierarchical gated recurrent unit, and combine the soft attention mechanism to perform temporal state evaluation to obtain the system state evolution characteristics. Then, combine the distributed advantage actor-critic algorithm (A3C) to perform collaborative control on the heat dissipation unit, combine the double deep Q-network and the prioritized experience replay to determine the candidate control action sequence, and optimize the control action based on the heat dissipation effect through the dynamic time warping algorithm and the deep deterministic policy gradient algorithm to obtain the optimal heat dissipation plan.
[0144] The hierarchical reinforcement learning framework is a framework that divides the decision-making process in reinforcement learning into multiple levels, and each level is responsible for handling different tasks or subtasks. The soft attention mechanism is a mechanism that focuses on important parts by calculating the weighted sum of input information. The temporal state evaluation is a way to evaluate and predict the state of the system at a certain moment, usually combined with the value function in the reinforcement learning algorithm to quantify the quality of a state. The distributed advantage actor-critic algorithm (A3C) is a reinforcement learning algorithm that combines the actor-critic structure, aiming to accelerate the learning process through multi-threaded parallel training. The actor-critic structure consists of an "actor" responsible for generating action policies. The double deep Q-network is an algorithm that improves the overestimation problem in deep Q-learning. The prioritized experience replay is a method to improve the efficiency of reinforcement learning by assigning different priorities to the samples in the experience replay pool. By preferentially replaying the experiences that have a greater impact on learning, it can accelerate the learning process and improve the performance of the model, especially suitable for rare events or low-frequency samples in reinforcement learning. The dynamic time warping algorithm is an algorithm used to measure the similarity between two time series, commonly used in pattern recognition, speech recognition, and data mining.
[0145] In an alternative embodiment,
[0146] A hierarchical reinforcement learning framework is constructed based on the performance analysis report and the hierarchical warning signal, wherein the top-level policy network guides action selection according to the causal relationship, the middle-level coordination network evaluates the action impact based on the predicted state, and the bottom-level execution network performs equipment control to obtain an initial heat dissipation strategy, and the initial heat dissipation strategy and the dynamic prediction graph are added to the hierarchical gated recurrent unit, and the timing state evaluation is performed in combination with the soft attention mechanism to obtain the system state evolution characteristics and the heat dissipation unit is collaboratively controlled in combination with the distributed advantage actor critic algorithm, and the candidate control action sequence is determined in combination with the dual deep Q network and priority experience replay, and the control action is optimized based on the heat dissipation effect through the dynamic time warping algorithm and the deep deterministic policy gradient algorithm, and the optimal heat dissipation solution includes:
[0147] Receive a performance analysis report and a graded warning signal, wherein the performance analysis report includes a temperature data stream, a voltage data stream, and a current data stream, and the graded warning signal includes temperature warning data and stress warning data, perform causal analysis on the performance analysis report and the graded warning signal through a top-level strategy network, and establish a state transition diagram, wherein the state transition diagram records a mapping relationship between a device state and a heat dissipation effect, and a node of the state transition diagram stores state characteristics;
[0148] Inputting predetermined causal relationships and predicted state data into a middle-level coordination network, wherein the predicted state data includes temperature field prediction data and stress prediction data, and the state feature extractor of the middle-level coordination network maps the temperature field data and stress data into a feature space to obtain feature data, and the feature data is input into an action evaluation module, and the action evaluation module maintains a state-action evaluation table and dynamically updates it to obtain an action evaluation data stream;
[0149] The control instruction generator of the bottom execution network receives the action evaluation data stream, queries the control instruction mapping table to generate an instruction data packet including the execution time, execution order and control parameters, and parses the instruction data packet into a heat dissipation unit control signal to form an initial heat dissipation strategy through an instruction parser;
[0150] The initial heat dissipation strategy and dynamic prediction graph data are input into a hierarchical gated recurrent unit, the state features of the current time step are extracted through an input gate processor, the historical state information is filtered through a forget gate processor, the current features and historical information are integrated through an output gate processor, the weight coefficients of data at different time steps are calculated through a soft attention processor to generate a time series weighted feature stream, and the system state is updated based on the time series weighted feature stream through a state update processor to obtain a system state evolution feature;
[0151] Input the system state evolution characteristics into the distributed advantage actor-critic algorithm, generate candidate action data and state value evaluation results through parallel actor-critic units, and collect learning results through an experience sharing processor and perform experience fusion to generate a collaborative control strategy;
[0152] Input the collaborative control strategy into the double deep Q-network, generate a Q-value estimation data stream through alternating processing of the main network and the target network, allocate replay probabilities according to sample importance through a prioritized experience replay processor, and output a candidate control action sequence;
[0153] Input the candidate control action sequence and heat dissipation effect feedback data into the deep deterministic policy gradient algorithm, generate deterministic actions through a policy network, evaluate the long-term benefits of actions through a value network, calculate the similarity of the control sequence through a dynamic time warping processor, and output an optimal heat dissipation solution through a policy optimization processor by combining value evaluation and similarity scores.
[0154] The experience sharing processor is a mechanism that improves learning efficiency by sharing the experience obtained by a reinforcement learning agent in the environment. The Q-value is an index used to evaluate the value of an action in a certain state in reinforcement learning, indicating the expected cumulative reward that can be obtained after executing a certain action in a specific state.
[0155] Collect the performance analysis report and graded warning signals of the system. The performance analysis report includes temperature data stream, voltage data stream, and current data stream. For example, at a certain moment, the CPU temperature is 75 degrees Celsius, the voltage is 1.2 volts, and the current is 0.5 amperes. The graded warning signals include temperature warning data and stress warning data. For example, when the CPU temperature exceeds 80 degrees Celsius, a first-level temperature warning is triggered; when the system load exceeds 90%, a second-level stress warning is triggered. The top-level policy network performs causal analysis on the collected information and establishes a state transition diagram. This diagram records the mapping relationship from the device state to the heat dissipation effect. For example, when the CPU temperature rises from 70 degrees Celsius to 80 degrees Celsius, the corresponding rotation speed of the cooling fan increases from 1000 revolutions per minute to 2000 revolutions per minute. Each node of the state transition diagram stores state characteristics, such as temperature, voltage, current, warning level, etc.
[0156] Input the pre-determined causal relationship and prediction status data into the middle-layer coordination network. The prediction status data includes temperature field prediction data and stress prediction data. For example, it is predicted that the CPU temperature will reach 85 degrees Celsius and the system load will reach 95% in the next 5 minutes. The state feature extractor of the middle-layer coordination network maps the temperature field data and stress data into the feature space. For example, it converts the temperature and load data into a high-dimensional vector representation. These feature data are input into the action evaluation module. The action evaluation module maintains a state-action evaluation table and updates it dynamically. For example, in the current state, the evaluation value for increasing the speed of the cooling fan is 0.8, and the evaluation value for decreasing the speed of the cooling fan is 0.2.
[0157] The control instruction generator of the bottom-layer execution network receives the action evaluation data stream, queries the control instruction mapping table, and generates an instruction data packet containing the execution time, execution order, and control parameters. For example, according to the evaluation result, it generates an instruction: set the speed of the cooling fan to 3000 revolutions per minute, and the execution time is immediate execution. The instruction parser parses the instruction data packet into a control signal for the cooling unit to form an initial cooling strategy.
[0158] Input the initial cooling strategy and dynamic prediction graph data into the hierarchical gated recurrent unit. For example, input the initial fan speed setting and the temperature prediction data for the next 5 minutes. The input gate processor extracts the state features of the current time step, the forget gate processor filters the historical state information, the output gate processor integrates the current features and historical information, and the soft attention processor calculates the weight coefficients of the data at different time steps. For example, it assigns a higher weight to the data at the most recent time step. Finally, it generates a time-series weighted feature stream. The state update processor updates the system state based on the time-series weighted feature stream to obtain the system state evolution features.
[0159] Input the system state evolution features into the distributed advantage actor-critic algorithm. Multiple parallel actor-critic units generate candidate action data and state value evaluation results. For example, one actor unit suggests increasing the fan speed to 4000 revolutions per minute, and the corresponding critic unit evaluates the value of this action as 0.9. The experience sharing processor collects the learning results and performs experience fusion to generate a collaborative control strategy.
[0160] Then input the collaborative control strategy into the double deep Q network. The main network and the target network alternately process to generate the Q-value estimation data stream. For example, the main network estimates the Q value of increasing the fan speed to 4000 revolutions per minute as 0.85, and the target network estimates the Q value of this action as 0.92. The prioritized experience replay processor assigns replay probabilities according to the sample importance. For example, for the action experience that causes a significant decrease in the system temperature, it assigns a higher replay probability. Finally, it outputs a candidate control action sequence.
[0161] Input the candidate control action sequence and the heat dissipation effect feedback data into the deep deterministic policy gradient algorithm. The policy network generates deterministic actions. For example, set the fan speed to 3500 revolutions per minute. The value network evaluates the long-term benefits of the actions. The dynamic time warping processor calculates the similarity of the control sequences. For example, calculate the similarity between the current control sequence and the historical optimal control sequence. The policy optimization processor combines the value evaluation and the similarity score to output the optimal heat dissipation solution. For example, finally determine the optimal heat dissipation solution of setting the fan speed to 3500 revolutions per minute for a duration of 10 minutes.
[0162] In this embodiment, through the combination of a hierarchical reinforcement learning framework and multiple algorithms, the heat dissipation strategy can be dynamically adjusted to achieve precise control of the heat dissipation resources, thereby improving the heat dissipation efficiency, reducing energy consumption, predicting potential overheating risks according to the real-time changes of the system state, and taking preventive measures in advance, thereby enhancing the stability and reliability of the system. By optimizing the heat dissipation solution, the device temperature and stress can be effectively controlled, avoiding the damage caused to the device by long-term high-temperature operation, thereby extending the service life of the device.
[0163] In an alternative embodiment,
[0164] Input the collaborative control strategy into the double deep Q network, generate the Q-value estimation data stream through the alternating processing of the main network and the target network, and allocate the replay probability according to the sample importance through the prioritized experience replay processor. The output candidate control action sequence includes:
[0165] Perform normalization processing on the collaborative control strategy data stream, normalize the temperature data to a preset interval according to the preset maximum temperature value, normalize the stress data to a preset interval according to the preset maximum stress value, and normalize the heat dissipation unit state parameters according to the range to generate a normalized state vector;
[0166] Input the normalized state vector into the input layer of the main network, extract features through multiple neurons in the first hidden layer to obtain the first-layer feature vector, extract features of the first-layer feature vector through multiple neurons in the second hidden layer to obtain the second-layer feature vector, and generate a control action combination through multiple neurons in the output layer;
[0167] Input the normalized state vector into the target network. The target network has the same network structure as the main network. The parameters of the target network are updated according to the preset number of training iterations. The parameter update method of the target network is to copy the parameter values from the main network;
[0168] Construct an experience replay pool to store training samples. The training samples include the current state vector, the executed action number, the reward value, the next state vector, and the termination flag, and calculate the time difference value and the reward amplitude value of the training samples;
[0169] Normalize the time difference value to a first preset interval to obtain a normalized time difference value, normalize the reward amplitude value to a second preset interval to obtain a normalized reward amplitude value, sum the normalized time difference value and the normalized reward amplitude value according to a preset ratio to obtain a sample priority score, sort the training samples in the experience replay pool in descending order according to the sample priority score, assign a first preset sampling probability to the training samples ranked at the top, assign a second preset sampling probability to the training samples ranked in the middle, and assign a third preset sampling probability to the training samples ranked at the bottom;
[0170] Randomly sample a preset number of the training samples from the experience replay pool based on the preset sampling probability to form a training batch, input the current state in the training batch into the main network to obtain an action value prediction, and input the next state in the training batch into the target network to obtain a target value;
[0171] Input the current state vector of the system into the trained main network to obtain an action value estimation score, sort the action value estimation scores in descending order, and select multiple actions with the highest scores to form a candidate control action sequence, where each action in the candidate control action sequence includes control parameters of the heat dissipation unit.
[0172] Preprocess the collaborative control strategy data stream. Temperature data, stress data, and heat dissipation unit state parameters all need to be standardized. For example, assume that the set maximum temperature value is 100 °C and the preset normalization interval is [0, 1]. If the currently measured temperature value is 60 °C, then the normalized temperature value is 60 / 100 = 0.6. Similarly, assume that the set maximum stress value is 100 MPa and the preset normalization interval is also [0, 1]. If the currently measured stress value is 40 MPa, then the normalized stress value is 40 / 100 = 0.4. For the heat dissipation unit state parameters, such as the fan speed, assume its range is 0 - 2000 revolutions per minute. If the current speed is 1000 revolutions per minute, then the normalized value is 1000 / 2000 = 0.5. Combine these normalized data into a state vector, such as [0.6, 0.4, 0.5].
[0173] The standardized state vector is input into the deep Q-network. This network consists of a main network and a target network, and the two networks have the same structure. Assume that both the main network and the target network contain two hidden layers, with each layer containing 64 neurons. The state vector is first input into the first hidden layer of the main network. After being processed by the neurons in the first hidden layer, the first-layer feature vector is obtained. Then, the first-layer feature vector is input into the second hidden layer, and after being processed, the second-layer feature vector is obtained. Finally, the second-layer feature vector is input into the output layer, and each neuron in the output layer corresponds to a control action combination, such as [fan 1 speed, fan 2 speed].
[0174] The structure of the target network is the same as that of the main network, but its parameter update method is different. The parameters of the target network are periodically copied from the main network. For example, every 100 training iterations, the parameters of the main network are copied to the target network.
[0175] To train the network, an experience replay pool needs to be constructed to store training samples. Each training sample contains the current state vector, the action number executed, the reward value, the next state vector, and the termination flag. For example, a training sample can be: [0.6, 0.4, 0.5], 3, 10, [0.5, 0.3, 0.6], False, indicating that the current state is [0.6, 0.4, 0.5], the action numbered 3 is executed, a reward of 10 is obtained, the next state is [0.5, 0.3, 0.6], and this state is not a termination state.
[0176] For each training sample, the temporal difference value and the reward amplitude value need to be calculated. Then, the temporal difference value and the reward amplitude value are respectively normalized to a preset interval. For example, the temporal difference value is normalized to [0, 1], and the reward amplitude value is normalized to [-1, 1]. Then, the normalized temporal difference value and the reward amplitude value are weighted and summed according to a preset ratio to obtain the sample priority score. For example, assume that the weight of the temporal difference value is 0.8 and the weight of the reward amplitude value is 0.2. The normalized temporal difference value of a sample is 0.5, and the normalized reward amplitude value is 0.8. Then the priority score of this sample is 0.5 * 0.8 + 0.8 * 0.2 = 0.56. The training samples in the experience replay pool are sorted in descending order of the priority score.
[0177] Sample probabilities are assigned according to the priority scores of the samples. For example, the top 10% of the samples are assigned a sampling probability of 0.5, the samples ranked 10%-90% are assigned a sampling probability of 0.3, and the bottom 10% of the samples are assigned a sampling probability of 0.2. Then, a preset number of training samples are randomly sampled from the experience replay pool based on the preset sampling probabilities to form a training batch. The current state in the training batch is input into the main network to obtain the action-value prediction, and the next state in the training batch is input into the target network to obtain the target value. The parameters of the main network are updated using these values.
[0178] The current state vector of the system is input into the trained main network to obtain the action-value estimation score. The action-value estimation scores are sorted in descending order, and multiple actions with the highest scores are selected to form a candidate control action sequence. For example, the three actions with the highest scores are selected. Each action includes the control parameters of the heat dissipation unit, such as [Fan 1 speed = 1500 revolutions per minute, Fan 2 speed = 1200 revolutions per minute].
[0179] In this embodiment, through the deep reinforcement learning algorithm, a better collaborative control strategy for the heat dissipation unit can be learned, so as to more effectively reduce the system temperature and stress. The intelligent control strategy can avoid unnecessary activation of the heat dissipation unit, thereby reducing the overall energy consumption of the system. Effective temperature and stress control can slow down system aging and extend the service life.
[0180] Figure 2 This is a schematic structural diagram of the thermal runaway prevention and heat dissipation optimization system based on the capacitor module according to the embodiment of the present invention, as Figure 2 shown, the system includes:
[0181] The first unit is used to collect multi-source operation data of single capacitors in the capacitor module, segment different types of data into data segments of the same size and perform linear projection by combining an improved hierarchical visual attention network, extract features by combining a multi-level self-attention structure and an attention mechanism, introduce learnable position encoding to embed time information, obtain initial features and add them to a bidirectional probability diffusion model, introduce conditional information during the diffusion process for data completion to obtain complete data information and add it to a multi-modal feature fusion network, calculate the correlation between modalities and adaptively modify the fusion weights by combining a dynamic weight allocation mechanism, obtain fusion features and add them to a causal structure discovery network, construct an initial causal graph structure based on expert knowledge and identify the causal relationship between different data by combining a causal discovery algorithm, obtain the temporal dependence relationship and perform uncertainty assessment, and obtain the performance analysis report corresponding to the single capacitor;
[0182] A second unit is configured to build a spatio-temporal dynamic graph neural network based on the performance analysis report, model the device temperature field distribution to obtain a dynamic prediction graph, identify the evolution characteristics corresponding to the capacitor temperature and stress, determine the parameter evolution law based on the evolution characteristics by combining hierarchical dilated convolution and self-attention mechanism, determine the temperature change trend and stress accumulation pattern and add them to a multi-task learning framework, process the performance analysis report and the dynamic prediction graph by combining a tabular data neural network and a dynamic graph attention network to obtain a predicted state, perform fault diagnosis through a large language model with a mixture of experts architecture and dynamic routing mechanism to obtain potential fault types and corresponding occurrence probabilities, build an uncertainty quantification module that integrates a probability graph model and Bayesian deep learning to calculate the confidence interval corresponding to each fault type, build a hierarchical early warning mechanism by combining pre-acquired expert knowledge, set multi-level early warning thresholds and output hierarchical early warning signals;
[0183] A third unit is configured to build a hierarchical reinforcement learning framework based on the performance analysis report and the hierarchical early warning signals. Among them, the top-level policy network guides action selection according to the causal relationship, the middle-level coordination network evaluates the action impact based on the predicted state, and the bottom-level execution network performs device control to obtain an initial heat dissipation strategy. Add the initial heat dissipation strategy and the dynamic prediction graph to a hierarchical gated recurrent unit, perform temporal state evaluation by combining a soft attention mechanism to obtain system state evolution characteristics, and perform collaborative control on the heat dissipation unit by combining a distributed advantage actor-critic algorithm. Combine a double deep Q-network and prioritized experience replay to determine a candidate control action sequence, and optimize the control action based on the heat dissipation effect through a dynamic time warping algorithm and a deep deterministic policy gradient algorithm to obtain an optimal heat dissipation solution.
[0184] In the third aspect of the embodiments of the present invention,
[0185] Provided is an electronic device, including:
[0186] A processor;
[0187] A memory for storing instructions executable by the processor;
[0188] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0189] In the fourth aspect of the embodiments of the present invention,
[0190] Provided is a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0191] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for preventing thermal runaway and optimizing heat dissipation based on a capacitor module, characterized in that: include: Collect multi-source operation data of single capacitors in the capacitor module, combine the improved hierarchical visual attention network to divide different types of data into data segments of the same size and perform linear projection, combine the multi-level self-attention structure and attention mechanism to extract features, introduce learnable position encoding to embed time information, obtain initial features and add them to the bidirectional probability diffusion model, introduce conditional information to complete the data during the diffusion process to obtain complete data information and add it to the multimodal feature fusion network, calculate the correlation between modalities and modify the fusion weights adaptively in combination with the dynamic weight allocation mechanism, obtain fusion features and add them to the causal structure discovery network, build an initial causal graph structure based on expert knowledge and combine the causal discovery algorithm to identify the causal relationship between different data, obtain the temporal dependency and perform uncertainty assessment, and obtain the performance analysis report corresponding to the single capacitor; Based on the performance analysis report, a spatiotemporal dynamic graph neural network is constructed and the temperature field distribution of the equipment is modeled to obtain a dynamic prediction graph and identify the evolution characteristics corresponding to the capacitor temperature and stress. Based on the evolution characteristics, the parameter evolution law is determined in combination with the hierarchical dilated convolution and self-attention mechanism, the temperature change trend and stress accumulation pattern are determined and added to the multi-task learning framework, the performance analysis report and the dynamic prediction graph are processed in combination with the tabular data neural network and the dynamic graph attention network to obtain the prediction state, and fault diagnosis is performed through a large language model with a hybrid expert architecture and a dynamic routing mechanism to obtain potential fault types and corresponding occurrence probabilities, an uncertainty quantification module that integrates the probability graph model and Bayesian deep learning is constructed to calculate the confidence interval corresponding to each fault type, and a hierarchical warning mechanism is constructed in combination with pre-acquired expert knowledge, and a multi-level warning threshold is set and a graded warning signal is output; A hierarchical reinforcement learning framework is constructed based on the performance analysis report and the graded warning signal, wherein the top-level policy network guides action selection according to the causal relationship, the middle-level coordination network evaluates the action impact based on the predicted state, and the bottom-level execution network performs equipment control to obtain an initial heat dissipation strategy. The initial heat dissipation strategy and the dynamic prediction graph are added to the hierarchical gated recurrent unit, and the timing state evaluation is performed in combination with the soft attention mechanism to obtain the system state evolution characteristics and the heat dissipation unit is collaboratively controlled in combination with the distributed advantage actor critic algorithm. The candidate control action sequence is determined by combining the dual deep Q network and priority experience replay. Based on the heat dissipation effect, the control action is optimized by the dynamic time warping algorithm and the deep deterministic policy gradient algorithm to obtain the optimal heat dissipation solution.
2. The method according to claim 1, characterized in that Collect multi-source operation data of single capacitors in the capacitor module, combine the improved hierarchical visual attention network to divide different types of data into data segments of the same size and perform linear projection, combine the multi-level self-attention structure and attention mechanism to extract features, introduce learnable position encoding to embed time information, obtain initial features and add them to the bidirectional probability diffusion model, introduce conditional information to complete the data in the diffusion process to obtain complete data information and add it to the multimodal feature fusion network, calculate the correlation between modalities and modify the fusion weights adaptively in combination with the dynamic weight allocation mechanism, obtain fusion features and add them to the causal structure discovery network, build an initial causal graph structure based on expert knowledge and combine the causal discovery algorithm to identify the causal relationship between different data, obtain the temporal dependency and perform uncertainty assessment, and obtain the performance analysis report corresponding to the single capacitor including: The voltage data, current data, ambient temperature data, shell temperature data and charge-discharge state data of the single capacitor in the capacitor module are collected and combined to obtain multi-source operation data, and the multi-source operation data is input into an improved hierarchical visual attention network, wherein the improved hierarchical visual attention network includes a feature segmentation layer, a feature projection layer and a feature enhancement layer, wherein the feature segmentation layer uses an overlapping sliding window to segment the multi-source operation data to obtain data segments, the feature projection layer uses an independent projection matrix to map the data segments to a unified feature space to obtain mapping features, and the feature enhancement layer enhances the mapping features through a channel-level attention mechanism and a space-level attention mechanism to obtain enhanced features; The enhanced features are input into a multi-level self-attention structure, wherein the multi-level self-attention structure includes a local attention unit and a global attention unit, wherein the local attention unit uses a multi-head attention mechanism to extract local dependency features, and the global attention unit uses a sparse attention mechanism to extract cross-segment correlation features, and the absolute position information and the relative position information are embedded into the local dependency features and the cross-segment correlation features through a learnable position encoder to obtain an initial feature; The initial features are input into a bidirectional probability diffusion model, a noise scheduling table is constructed in the training phase to control the noise addition rate and train a denoising predictor, the operating condition information is converted into a conditional vector in the inference phase, data completion is performed through multiple steps of iteration in combination with the conditional vector to obtain complete data information and input into a multimodal feature fusion network, different modal features are standardized to obtain standardized features, an attention matrix is constructed to calculate the inter-modal correlation of the standardized features to obtain correlation features, dynamic weight allocation is performed on the correlation features based on a soft attention mechanism and a cross-modal comparative learning strategy is introduced to obtain fusion features; An initial causal graph including node attribute definitions and edge connection constraints is constructed based on expert knowledge, the fused features are input into the causal structure discovery network, the trend features, periodic features and mutation features of the time series feature sequence are extracted, the linear correlation and nonlinear correlation between different feature sequences are evaluated, the direction and weight of the causal edge are determined based on the dependency strength to obtain the causal relationship, the feature description sequence of the causal relationship is extracted under different time windows, multiple groups of evaluation samples are generated by the self-service method and the confidence scores are calculated, the integration strategy is adopted to fuse different evaluation results to obtain the final confidence score, and a performance analysis report is generated based on the final confidence score.
3. The method according to claim 2, characterized in that The feature segmentation layer uses overlapping sliding windows to segment the multi-source running data to obtain data segments, the feature projection layer uses an independent projection matrix to map the data segments to a unified feature space to obtain mapping features, and the feature enhancement layer enhances the mapping features through a channel-level attention mechanism and a space-level attention mechanism to obtain enhanced features including: Acquire the multi-source operation data and construct an adaptive segmentation mechanism, detect the change rate of the multi-source operation data through a data change rate detection module, and dynamically adjust the sliding window parameter according to the change rate, wherein when the change rate is higher than a first threshold, the window length of the sliding window parameter is reduced to eighty percent of the reference window length, and when the change rate is lower than a second threshold, the window length of the sliding window parameter is increased to one hundred and twenty percent of the reference window length; The multi-source operation data are processed in parallel using sliding windows of different lengths, wherein the first sliding window is used to capture transient change features to obtain a first feature segment, the second sliding window is used to extract local trend features to obtain a second feature segment, and the third sliding window is used to obtain macro-evolution features to obtain a third feature segment; the first feature segment, the second feature segment, and the third feature segment are input into a feature pyramid network for integration to obtain a data segment; The original feature space is divided into a plurality of subspaces, each of the subspaces is configured with a main projection matrix and an auxiliary projection matrix, the data fragments are input into the main projection matrix to obtain basic features, the data fragments are input into the auxiliary projection matrix to obtain residual features, and the basic features and the residual features are feature fused to obtain mapping features; Constructing a feature quality evaluation index, wherein the feature quality evaluation index includes feature discrimination, feature stability, and information retention, evaluating the mapping feature based on the feature quality evaluation index to obtain an evaluation result, and optimizing the main projection matrix and the auxiliary projection matrix according to the evaluation result; Constructing a hierarchical attention mechanism, including a feature channel attention mechanism, a time dimension attention mechanism, and a cross-modal attention mechanism, wherein the feature channel attention mechanism calculates the inter-channel dependency of the mapping feature to obtain a channel attention map, extracts the channel feature statistics of the mapping feature to obtain channel importance, and adaptively weights the channel attention map and the channel importance to obtain a channel attention weight; The time dimension attention mechanism extracts features of different time scales through a multi-head attention mechanism, and uses a deformable convolutional network to process time series to obtain time dimension feature enhancement results. The cross-modal attention mechanism constructs different types of features into a graph structure, and obtains enhanced features through feature interaction through a message passing mechanism.
4. The method according to claim 1, characterized in that: Based on the performance analysis report, a spatiotemporal dynamic graph neural network is constructed and the temperature field distribution of the equipment is modeled to obtain a dynamic prediction graph and identify the evolution characteristics corresponding to the capacitor temperature and stress. Based on the evolution characteristics, the parameter evolution law is determined in combination with the hierarchical dilated convolution and self-attention mechanism, the temperature change trend and stress accumulation pattern are determined and added to the multi-task learning framework, the performance analysis report and the dynamic prediction graph are processed in combination with the tabular data neural network and the dynamic graph attention network to obtain the prediction state, and the fault diagnosis is performed through a large language model of a hybrid expert architecture and a dynamic routing mechanism to obtain the potential fault type and the corresponding occurrence probability, and an uncertainty quantification module that integrates the probability graph model and Bayesian deep learning is constructed to calculate the confidence interval corresponding to each fault type, and a hierarchical warning mechanism is constructed in combination with the pre-acquired expert knowledge, and a multi-level warning threshold is set and a graded warning signal is output, including: receiving a performance analysis report, dividing the physical structure of the capacitor into a ten-by-ten grid point array, collecting the spatial position coordinates and real-time temperature value of each grid point in the grid point array, constructing a spatial correlation edge based on the physical distance between the grid points, establishing a connection relationship between the grid points according to a preset distance threshold, and constructing a time series sample using a sliding window method; Based on the performance analysis report, a spatiotemporal dynamic graph neural network is constructed to model the temperature field distribution of the equipment, wherein the spatiotemporal dynamic graph neural network includes a first graph convolution layer, a second graph convolution layer, and a third graph convolution layer, wherein the first graph convolution layer performs spatial information aggregation to obtain a weighted sum of nodes, the second graph convolution layer controls information flow through an update gate and a reset gate, and the third graph convolution layer generates a candidate state and updates a node hidden state to obtain a dynamic prediction graph; Constructing a hierarchical dilated convolution structure, including a first dilated convolution branch, a second dilated convolution branch, and a third dilated convolution branch, wherein the first dilated convolution branch adopts a standard convolution kernel, the second dilated convolution branch adopts a first dilated convolution kernel, and the third dilated convolution branch adopts a second dilated convolution kernel, inputting the output feature map of the dilated convolution branch into a self-attention mechanism module, identifying the evolution characteristics corresponding to the capacitor temperature and stress in the dynamic prediction map, and determining the temperature change trend and stress accumulation mode; Adding the temperature change trend and the stress accumulation pattern to a multi-task learning framework, the multi-task learning framework includes a shared feature extraction layer, a temperature prediction branch and a stress assessment branch, the temperature prediction branch uses a tabular data neural network to process the performance analysis report and the dynamic prediction graph, and the stress assessment branch uses a dynamic graph attention network to process the performance analysis report and the dynamic prediction graph to obtain a predicted state; Constructing a large language model of a hybrid expert architecture, the large language model comprising a plurality of expert modules and a dynamic routing mechanism, the dynamic routing mechanism calculating an allocation weight of each expert module based on the predicted state, determining an expert module to perform fault diagnosis according to the allocation weight, and generating a diagnosis result including potential fault types and corresponding occurrence probabilities; Construct an uncertainty quantification module that integrates a probabilistic graphical model and Bayesian deep learning, where the nodes in the probabilistic graphical model represent the fault type, and the edge weights between the nodes represent the degree of correlation between the fault types. The diagnosis results are sampled through a Bayesian deep learning framework, and the confidence interval corresponding to each fault type is calculated; A hierarchical warning mechanism is constructed based on pre-acquired expert knowledge, temperature parameters and stress parameters are divided into multi-level warning thresholds, and a graded warning signal is generated based on the fault type, the occurrence probability and the confidence interval.
5. The method according to claim 4, characterized in that The temperature prediction branch uses a table data neural network to process the performance analysis report and the dynamic prediction graph, and the stress assessment branch uses a dynamic graph attention network to process the performance analysis report and the dynamic prediction graph, and the prediction status obtained includes: receiving a performance analysis report and a dynamic prediction graph, and performing standardized preprocessing on the tabular data in the performance analysis report to obtain preprocessed data, wherein the tabular data includes temperature parameters, voltage parameters, and current parameters; Constructing a tabular data neural network, wherein the number of nodes in the input layer of the tabular data neural network corresponds to the feature dimension of the preprocessed data, the tabular data neural network comprises a first hidden layer, a second hidden layer and a third hidden layer, the first hidden layer and the second hidden layer use a rectified linear activation function and a random dropout regularization layer, the third hidden layer uses a hyperbolic tangent activation function, and the dimensions of the first hidden layer, the second hidden layer and the third hidden layer decrease layer by layer to form a funnel-shaped structure Flatten the temperature field distribution information in the dynamic prediction graph in the time dimension and perform feature aggregation to obtain an extended feature vector, and perform dimension reduction processing on the extended feature vector through a double-layer fully connected network to obtain an extended feature representation that matches the feature dimension of the preprocessed data; Constructing a dynamic graph attention network, setting monitoring points as graph nodes, setting physical connection relationships between the monitoring points as graph edges, and fusing stress-related indicators in the performance analysis report with temperature distribution information in the dynamic prediction graph to form a node feature vector; A two-layer graph attention network structure is constructed. The first-layer graph attention network calculates the degree of association between different nodes to obtain the feature transfer weight. The output of the feature transfer weight is processed by an exponential linear activation function and batch normalization. The second-layer graph attention network uses a multi-head attention mechanism to extract the interaction features between nodes, and constructs a feature fusion module. The output features of the tabular data neural network and the output features of the dynamic graph attention network are adjusted to the same dimension through a dimensionality transformation network. The output features are adaptively fused using a gating mechanism to obtain a predicted state.
6. The method according to claim 1, characterized in that A hierarchical reinforcement learning framework is constructed based on the performance analysis report and the hierarchical warning signal, wherein the top-level policy network guides action selection according to the causal relationship, the middle-level coordination network evaluates the action impact based on the predicted state, and the bottom-level execution network performs equipment control to obtain an initial heat dissipation strategy, and the initial heat dissipation strategy and the dynamic prediction graph are added to the hierarchical gated recurrent unit, and the timing state evaluation is performed in combination with the soft attention mechanism to obtain the system state evolution characteristics and the heat dissipation unit is collaboratively controlled in combination with the distributed advantage actor critic algorithm, and the candidate control action sequence is determined in combination with the dual deep Q network and priority experience replay, and the control action is optimized based on the heat dissipation effect through the dynamic time warping algorithm and the deep deterministic policy gradient algorithm, and the optimal heat dissipation solution includes: Receive a performance analysis report and a graded warning signal, wherein the performance analysis report includes a temperature data stream, a voltage data stream, and a current data stream, and the graded warning signal includes temperature warning data and stress warning data, perform causal analysis on the performance analysis report and the graded warning signal through a top-level strategy network, and establish a state transition diagram, wherein the state transition diagram records a mapping relationship between a device state and a heat dissipation effect, and a node of the state transition diagram stores state characteristics; Inputting predetermined causal relationships and predicted state data into a middle-level coordination network, wherein the predicted state data includes temperature field prediction data and stress prediction data, and the state feature extractor of the middle-level coordination network maps the temperature field data and stress data into a feature space to obtain feature data, and the feature data is input into an action evaluation module, and the action evaluation module maintains a state-action evaluation table and dynamically updates it to obtain an action evaluation data stream; The control instruction generator of the bottom execution network receives the action evaluation data stream, queries the control instruction mapping table to generate an instruction data packet including the execution time, execution order and control parameters, and parses the instruction data packet into a heat dissipation unit control signal to form an initial heat dissipation strategy through an instruction parser; The initial heat dissipation strategy and dynamic prediction graph data are input into a hierarchical gated recurrent unit, the state features of the current time step are extracted through an input gate processor, the historical state information is filtered through a forget gate processor, the current features and historical information are integrated through an output gate processor, the weight coefficients of data at different time steps are calculated through a soft attention processor to generate a time series weighted feature stream, and the system state is updated based on the time series weighted feature stream through a state update processor to obtain a system state evolution feature; Inputting the system state evolution characteristics into a distributed dominant actor-critic algorithm, generating candidate action data and state value evaluation results through parallel actor-critic units, collecting learning results through an experience sharing processor and performing experience fusion to generate a collaborative control strategy; The collaborative control strategy is input into a dual deep Q network, a Q-value estimation data stream is generated by alternating processing of the main network and the target network, a replay probability is assigned according to the importance of the sample through a priority experience replay processor, and a candidate control action sequence is output; The candidate control action sequence and heat dissipation effect feedback data are input into the deep deterministic policy gradient algorithm, deterministic actions are generated through the policy network, the long-term benefits of the actions are evaluated through the value network, the control sequence similarity is calculated through the dynamic time warping processor, and the optimal heat dissipation solution is output through the policy optimization processor combining the value evaluation and similarity score.
7. The method according to claim 6, characterized in that The collaborative control strategy is input into the dual deep Q network, and the Q value estimation data stream is generated by alternating processing of the main network and the target network. The replay probability is assigned according to the importance of the sample through the priority experience replay processor, and the candidate control action sequence is output, including: Standardize the data stream of the collaborative control strategy, normalize the temperature data to a preset interval according to a preset maximum temperature value, normalize the stress data to a preset interval according to a preset maximum stress value, normalize the heat dissipation unit state parameters according to the range, and generate a standardized state vector; The standardized state vector is input into the input layer of the main network, and features are extracted through multiple neurons in the first hidden layer to obtain a first layer feature vector, and the first layer feature vector is extracted through multiple neurons in the second hidden layer to obtain a second layer feature vector, and the second layer feature vector generates a control action combination through multiple neurons in the output layer; Inputting the standardized state vector into a target network, the target network having the same network structure as the main network, updating parameters of the target network according to a preset number of training iterations, and updating parameters of the target network by copying parameter values from the main network; Construct an experience replay pool to store training samples, wherein the training samples include a current state vector, an execution action number, a reward value, a next state vector, and a termination flag, and calculate a time difference value and a reward amplitude value of the training samples; Normalizing the time difference value to a first preset interval to obtain a normalized time difference value, normalizing the reward amplitude value to a second preset interval to obtain a normalized reward amplitude value, performing weighted summation of the normalized time difference value and the normalized reward amplitude value according to a preset ratio to obtain a sample priority score, arranging the training samples in the experience replay pool in descending order according to the sample priority scores, allocating a first preset sampling probability to the top-ranked training samples, allocating a second preset sampling probability to the middle-ranked training samples, and allocating a third preset sampling probability to the bottom-ranked training samples; Based on a preset sampling probability, a preset number of training samples are randomly sampled from the experience replay pool to form a training batch, the current state in the training batch is input into the main network to obtain an action value prediction, and the next state in the training batch is input into the target network to obtain a target value; The current state vector of the system is input into the trained main network to obtain an action value estimation score, the action value estimation scores are arranged in descending order, and multiple actions with the highest scores are selected to form a candidate control action sequence, wherein each action in the candidate control action sequence contains control parameters of the cooling unit.
8. A thermal runaway prevention and heat dissipation optimization system based on a capacitor module, used to implement any of the methods of claims 1-7, characterized in that: include: The first unit is used to collect multi-source operation data of single capacitors in the capacitor module, combine the improved hierarchical visual attention network to divide different types of data into data segments of the same size and perform linear projection, combine the multi-level self-attention structure and attention mechanism to extract features, introduce learnable position encoding to embed time information, obtain initial features and add them to the bidirectional probability diffusion model, introduce conditional information to complete the data in the diffusion process to obtain complete data information and add it to the multimodal feature fusion network, calculate the correlation between modalities and modify the fusion weights adaptively in combination with the dynamic weight allocation mechanism, obtain fusion features and add them to the causal structure discovery network, construct an initial causal graph structure based on expert knowledge and combine the causal discovery algorithm to identify the causal relationship between different data, obtain the temporal dependency and perform uncertainty assessment, and obtain the performance analysis report corresponding to the single capacitor; The second unit is used to construct a spatiotemporal dynamic graph neural network based on the performance analysis report and model the temperature field distribution of the equipment, obtain a dynamic prediction graph and identify the evolution characteristics corresponding to the capacitor temperature and stress, determine the parameter evolution law based on the evolution characteristics combined with the hierarchical dilated convolution and self-attention mechanism, determine the temperature change trend and stress accumulation mode and add them to the multi-task learning framework, process the performance analysis report and the dynamic prediction graph in combination with the tabular data neural network and the dynamic graph attention network to obtain the prediction state, perform fault diagnosis through a large language model of a hybrid expert architecture and a dynamic routing mechanism, obtain potential fault types and corresponding occurrence probabilities, construct an uncertainty quantification module that integrates the probability graph model and Bayesian deep learning to calculate the confidence interval corresponding to each fault type, and construct a hierarchical warning mechanism in combination with pre-acquired expert knowledge, set multi-level warning thresholds and output graded warning signals; The third unit is used to construct a hierarchical reinforcement learning framework based on the performance analysis report and the graded warning signal, wherein the top-level policy network guides action selection according to the causal relationship, the middle-level coordination network evaluates the action impact based on the predicted state, and the bottom-level execution network performs equipment control to obtain an initial heat dissipation strategy, and the initial heat dissipation strategy and the dynamic prediction graph are added to the hierarchical gated recurrent unit, and the timing state evaluation is performed in combination with the soft attention mechanism to obtain the system state evolution characteristics and the distributed advantage actor critic algorithm is combined to coordinate the heat dissipation unit, and the candidate control action sequence is determined by combining the dual deep Q network and priority experience replay, and the control action is optimized based on the heat dissipation effect through the dynamic time warping algorithm and the deep deterministic policy gradient algorithm to obtain the optimal heat dissipation solution.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Ship auxiliary equipment integrated monitoring method and system suitable for high-speed ship
CN120353181A
Cement quality detection method and system based on multi-modal characteristic data
CN120761421A
Test method and test system for thermal management system of electric vehicle
CN120764371A
Test method and test system of electric vehicle thermal management system
CN120764371B
Multivariable system control method based on causal reasoning and reinforcement learning
CN121097770A