Energy storage container distributed array active explosion venting control method

CN122221691APending Publication Date: 2026-06-16CHINA UNIV OF MINING & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF MINING & TECH
Filing Date
2026-05-14
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing explosion-proof and pressure relief technologies for energy storage containers suffer from problems such as lag, passivity, lack of flexibility and quantitative control. They cannot effectively predict and respond to explosions caused by thermal runaway. Furthermore, existing intelligent monitoring solutions have poor calculation timeliness and lack physical feedback mechanisms.

Method used

By employing multi-dimensional security situational awareness data acquisition, combined with a deep evaluation model based on physical information constraints and a deep reinforcement learning agent, a distributed array explosion relief control system is constructed. Through a dual-stream spatiotemporal attention network and a VSP explosion relief model, dynamic graded release of explosion energy and minimization of environmental hazards are achieved.

Benefits of technology

It achieves high-value perception and accurate decision-making in the early stages of an explosion, integrates physical constraints and reinforcement learning, has millisecond-level response capabilities, ensures container safety and reduces environmental hazards, and the system has autonomous learning and adaptive optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122221691A_ABST
    Figure CN122221691A_ABST
Patent Text Reader

Abstract

The application provides a kind of energy storage container distributed array active explosion venting control method, belongs to the technical field of energy storage explosion venting control based on reinforcement learning;Collecting energy storage container internal pressure high order derivative, voiceprint features and external environment vulnerability and other multi-source heterogeneous data, construct space-time feature tensor and input deep evaluation model based on physical information constraint;The model uses double-flow space-time attention network to capture the nonlinear evolution characteristics of thermal runaway, and ensures the physical interpretability of the prediction by constraining the residual error of the fluid mechanics equation, and outputs the thermal runaway energy index;Further, the index is used to dynamically correct the VSP model to calculate the theoretical venting flux, and the reinforcement learning agent is used to solve the minimum environmental hazard cost of the distributed valve array opening combination under the premise of meeting the flux constraint. The application realizes active perception, hierarchical release and guided control of explosion energy, and solves the technical problems of lagging response and uncontrollable secondary disasters in traditional passive explosion venting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of energy storage explosion control technology based on reinforcement learning, and particularly relates to an active explosion control method for distributed arrays of energy storage containers. Background Technology

[0002] As the global energy structure shifts towards a low-carbon model, electrochemical energy storage power stations centered on lithium-ion batteries have become a key infrastructure for building new power systems. However, the safety of energy storage containers, as high-energy-density enclosed carriers, is becoming increasingly prominent. In particular, when a battery experiences thermal runaway, it releases a large amount of high-temperature flammable gases (such as hydrogen and carbon monoxide) in a very short time. If these gases cannot be effectively contained in time, they will rapidly accumulate within the confined space of the container and reach their explosive limits. Once ignited by high-temperature particles or an electric arc, this will trigger a violent deflagration or even a detonation. This uncontrolled release of explosive energy not only causes the container itself to disintegrate, but the resulting shockwaves and flames can also affect surrounding equipment and personnel, causing severe secondary disasters and economic losses. Therefore, effectively controlling the release path and intensity of explosive energy under extreme conditions of thermal runaway-induced explosions has become a critical safety challenge that the energy storage industry urgently needs to address.

[0003] Currently, there are two main limitations to explosion-proof and depressurization technologies for energy storage containers: The first type is traditional passive mechanical pressure relief technology. This technology typically involves installing explosion-proof plates (rupture discs) or mechanical pressure relief valves with preset pressure thresholds on the top or side walls of the container. Its working principle is that when the pressure inside the container rises to a specific value, physical components rupture or open to release the pressure. However, this technology has significant limitations and is passive; it only responds to a single static pressure quantity and cannot detect the rate of pressure rise during thermal runaway. The explosion venting system, which monitors the combustion evolution trend, often only activates at the moment of the explosion, making early intervention difficult. Furthermore, this technology lacks flexibility, employing a threshold-triggered, one-size-fits-all approach that cannot adjust the venting flow rate according to the severity of the accident. The second category is "early intelligent monitoring" technology based on simulation. In recent years, some technical solutions have attempted to introduce environmental perception. For example, application number CN202311322712.0 constructs a risk database by establishing a three-dimensional geometric model and CFD fluid dynamics simulation, attempting to select the direction of explosion venting through a lookup table. However, such technologies have serious drawbacks in actual engineering: 1) Poor computational efficiency: Fluid simulation computation is huge, making it difficult to complete real-time inference within the millisecond-level sudden change window of thermal runaway. Most of them can only rely on offline databases and cannot cope with dynamic and complex working conditions; 2) Single control dimension: Such solutions are mostly limited to "direction selection" and lack quantitative control of "explosion flux" (i.e., opening area), failing to solve the problem of "how much to vent"; 3) Lack of physical closed loop: Relying only on feedforward control, there is a lack of physical feedback fallback mechanism after execution. Once the model prediction deviates, there is a lack of mandatory error correction methods. Summary of the Invention

[0004] To address the above problems, this invention proposes an active explosion venting control method for distributed arrays of energy storage containers, comprising the following processes: S1 collects multi-dimensional safety situation awareness data of the energy storage container under its operating status, including internal thermal runaway state data and external environmental vulnerability data, and preprocesses it; S2, the preprocessed data is input into the deep evaluation model based on physical information constraints. The model adopts a dual-stream spatiotemporal attention network architecture: it uses a bidirectional long short-term memory network to extract the nonlinear temporal evolution features of the internal thermal runaway state data, uses a multilayer perceptron or graph network to extract the spatial topological features of the external environment vulnerability data, and uses a multi-head self-attention mechanism to dynamically weight and fuse the multi-source features, and finally outputs a normalized thermal runaway energy index. S3. The thermal runaway energy index is substituted into the VSP explosion relief model as a dynamic correction factor to calculate the theoretical explosion relief flux required to maintain the safety of the container structure. An environmental hazard cost function containing personnel distance factor and wind direction factor is constructed. Under the premise of satisfying the theoretical explosion relief flux constraint, a deep reinforcement learning agent is used to solve the distributed valve array opening combination matrix that minimizes the environmental hazard cost. S4, the controller drives the valve to operate according to the matrix and continuously monitors the pressure rise rate feedback in the chamber. When the pressure rise rate does not show a downward inflection point, it triggers the forced skip mode.

[0005] Preferably, the internal thermal runaway state data includes the internal static pressure, pressurization rate, pressure acceleration, characteristic gas concentration, and acoustic signature. A high-frequency sensor array deployed within the container battery compartment synchronously collects internal static pressure and characteristic gas concentration data at a sampling frequency of at least 100Hz. The pressurization rate is obtained by performing first-order differential calculations on the internal static pressure data, and the pressure acceleration is obtained by performing second-order differential calculations. The pressure acceleration is used to characterize the critical trend of the thermal runaway combustion mode transitioning from deflagration to detonation. Acoustic signals within the compartment are collected using a microphone array, and Mel-frequency cepstral coefficients are extracted to obtain acoustic signature characteristics representing safety valve rupture or cell ejection events. Finally, the internal static pressure, pressurization rate, pressure acceleration, characteristic gas concentration, and acoustic signature are aligned according to timestamps.

[0006] Preferably, the external environment vulnerability data includes environmental wind speed, wind direction angle, and spatial distance coordinates of sensitive targets relative to each explosion relief surface; Real-time acquisition of ambient wind speed and wind direction angle is achieved using micro-meteorological instruments deployed on the exterior of the container; sensitive targets, including personnel, key equipment, and adjacent containers, are identified using a perimeter visual monitoring system, and the Euclidean distance and azimuth angle of each sensitive target relative to each explosion-proof surface of the container are calculated; a local environmental field vector centered on the container is constructed, which includes the wind field vector and the spatial coordinates of all sensitive targets; the maximum value of the ambient wind speed is normalized, and the distance data is transformed by reciprocal transformation to enhance the weight of nearby targets.

[0007] Preferably, the dual-stream spatiotemporal attention network architecture specifically includes a parallel temporal flow sub-network and a spatial flow sub-network: the temporal flow sub-network adopts a stacked bidirectional long short-term memory network structure, and the input is a temporal sequence of internal thermal runaway state data, used to capture the long-term dependence of pressure and gas parameters in the time dimension and the instantaneous mutation characteristics; the spatial flow sub-network adopts a graph attention network structure, modeling the container explosion relief surface and external sensitive targets as graph nodes, and the edge weights between nodes are determined by the spatial distance and wind direction factor, used to extract the spatial topological risk characteristics of the external environment; finally, through a multi-head self-attention mechanism layer, the attention weights are dynamically calculated according to the signal-to-noise ratio of each sensor channel, and the temporal flow features and spatial flow features are weighted and fused to generate a high-dimensional spatiotemporal hidden layer feature vector.

[0008] Preferably, the deep evaluation model introduces a physical residual constraint PINN based on the fluid dynamics conservation equation during the training phase. Specifically, a physical constraint term is introduced into the loss function of the deep evaluation model. The physical constraint term is constructed based on the ideal gas law and the mass conservation equation. The specific calculation method is as follows: the pressure value inside the chamber predicted by the model at the next moment is substituted into the physical equation, and the physical residual between it and the pressure, gas generation rate and explosion venting rate at the current moment is calculated. The total loss function is composed of the data fitting error and the physical residual weighted. By minimizing the total loss function, the model is trained, which forces the weight update direction of the neural network to conform to the fluid dynamics physical manifold.

[0009] Preferably, the VSP explosion relief model in S3 specifically involves: mapping the thermal runaway energy index output by the deep evaluation model to a dynamic correction factor; calculating the baseline explosion relief area based on the VSP model formula using the container's free volume, the current maximum pressurization rate, the design withstand pressure, and the flow coefficient; multiplying the baseline explosion relief area by the dynamic correction factor to obtain the theoretical explosion relief flux required to prevent the container structure from disintegrating at the current moment; the dynamic correction factor exhibits a nonlinear positive correlation with the thermal runaway energy index, and when the energy index reaches an extreme value, the correction factor tends to 1, thereby maximizing the theoretical explosion relief flux.

[0010] Preferably, the environmental hazard cost function in S3 is composed of a weighted sum of a distance cost term and a wind direction cost term; the distance cost term is inversely proportional to the square of the distance to the sensitive target, indicating that the closer the distance, the greater the hazard; the wind direction cost term is determined by the cosine of the angle between the explosion venting direction and the environmental wind direction, indicating the potential risk of the sensitive area being upwind or pointing downwind.

[0011] Preferably, in step S3, a deep reinforcement learning agent is trained using a near-end policy optimization algorithm, taking the current spatiotemporal hidden layer feature vector as the state input and outputting the opening combination of the distributed valve array as the action; in the reward function design, satisfying the theoretical required discharge flux is taken as a hard penalty constraint, and minimizing the environmental hazard cost is taken as a positive reward objective, thereby solving for the optimal opening combination matrix.

[0012] Preferably, in step S3, after executing the optimal opening combination matrix, a time window is set to continuously monitor the differential change in the pressurization rate inside the compartment; if the differential value of the pressurization rate is non-negative within the time window, it is determined that the current explosion relief strategy has failed or the risk has been underestimated, and the forced skip mode is immediately triggered; the logic of the forced skip mode is: ignoring the environmental hazard cost function, forcibly sending opening commands to all standby explosion relief valves that are in the closed state, realizing the maximum throughput release at the physical level, so as to prevent the disintegration of the container body structure as the highest priority protection target.

[0013] Preferably, the model also includes an online evolution step, specifically: packaging the entire process data of each data breach event, including the input multi-dimensional security situation awareness data, the decision actions output by the deep reinforcement learning agent, and the boost rate feedback results after execution, into an experience sample; if a forced skip mode is triggered, the sample is marked as a high-value negative sample and stored in the experience replay pool; the experience replay pool is used to perform online fine-tuning training on the deep evaluation model and the reinforcement learning agent to correct the model parameters.

[0014] Compared with the prior art, the present invention has the following innovative features and beneficial effects: (1) Enhancing the comprehensiveness and depth of situational awareness: Existing technologies rely solely on static pressure thresholds and cannot proactively predict explosions. This invention constructs a multi-dimensional perception system that internally senses detonation trends and externally assesses environmental risks: internally, it synchronously monitors static pressure, pressure rise rate, pressure acceleration, and acoustic signature characteristics to keenly capture the critical trend of thermal runaway transforming into detonation; externally, it quantifies the topological distance between the environmental wind field and sensitive targets in real time. This panoramic perception capability enables the system to acquire high-value safety situational inputs in the early stages of an explosion; (2) Integrating physical constraints and reinforcement learning to achieve highly interpretable assessment and hierarchical accurate decision-making: Addressing the issues of time-consuming traditional simulations and lack of quantitative control, this invention proposes a dual-driven architecture that integrates physical mechanisms and data. First, fluid dynamics residual constraints are introduced into the feature network to overcome the lack of physical interpretability in pure AI algorithms, outputting a hazard index that conforms to physical laws in milliseconds. Subsequently, this index is dynamically used to modify the VSP explosion relief model to calculate the theoretical flux, and it is used as a hard constraint. The optimal valve matrix is ​​then solved using a PPO reinforcement learning agent combined with an environmental cost function. This mechanism achieves seamless integration from hazard assessment to energy hierarchical release and minimizes environmental hazards. (3) Introducing closed-loop feedback and physical fallback to drive the continuous online evolution of the explosion relief system: Addressing the pain points of traditional intelligent solutions' open-loop control and failure due to inaccurate predictions, this invention designs a safety fallback based on physical feedback: After explosion relief is executed, the incremental pressure rate is monitored in real time. If the pressure is not suppressed, the highest priority forced bypass procedure is immediately triggered, and the backup array is fully activated as a fallback. Simultaneously, the system treats such extreme events as high-value negative samples, automatically correcting the PINN parameters and PPO decision boundary, enabling the system to possess the lifelong online evolution capability of autonomous learning and adaptive optimization. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the following description is only one embodiment of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating the overall technical route of the present invention.

[0017] Figure 2 This is a structural diagram of a deep evaluation model based on physical information constraints.

[0018] Figure 3 This is a structural diagram of a hierarchical explosion leakage decision and closed-loop control model based on reinforcement learning.

[0019] Figure 4 This is a schematic diagram showing the distribution of the explosion vent array in an energy storage container. Detailed Implementation

[0020] The present invention proposes an active explosion venting control method for a distributed array of energy storage containers, the overall technical route of which is shown in the flowchart below. Figure 1 As shown, the specific steps are as follows: S1. Construction of a multi-dimensional security situation awareness dataset for energy storage containers: Addressing the issues of existing technologies having a single perception dimension and lacking spatiotemporal evolution characteristics, this invention deploys a high-frequency sensor array inside the energy storage container to collect data on internal static pressure, pressurization rate, pressure acceleration, characteristic gas concentration, and acoustic signature features. Externally, a micro-meteorological instrument and a visual monitoring system are deployed to collect environmental wind speed, wind direction angle, and distance coordinates of sensitive targets. The collected multi-source heterogeneous data undergoes timestamp synchronization and nonlinear normalization processing, and a spatiotemporal feature tensor based on a sliding window is constructed to form a multi-dimensional security situation awareness dataset for model input.

[0021] S2 is a deep evaluation model design based on physical information constraints. The spatiotemporal feature tensors obtained in S1 are input into the deep evaluation model. First, through a dual-stream spatiotemporal attention network architecture, the nonlinear temporal evolution features of the internal thermal runaway state are extracted using the temporal subnetwork, and the topological risk features of the external environment are extracted using the spatial subnetwork. Second, the dynamic weighted fusion of internal and external features is achieved through a multi-head self-attention mechanism. Finally, physical residual constraints (PINN) based on fluid dynamics conservation equations are introduced during the model training phase to output a physically interpretable thermal runaway energy index.

[0022] S3, path optimization for graded release of explosion energy and minimization of hazards; using the thermal runaway energy index output by S2 as a dynamic correction factor, and substituting it into the VSP explosion venting model to calculate the theoretical required explosion venting flux under the current operating conditions; constructing an environmental hazard cost function that includes distance and wind direction factors; using a deep reinforcement learning agent (PPO), under the premise of satisfying the theoretical required explosion venting flux constraint, solving the distributed valve array opening combination matrix that minimizes the environmental hazard cost.

[0023] S4 executes control, closed-loop correction, and online model evolution; it transforms the optimal strategy generated by S3 into hardware-driven instructions and monitors the cabin pressurization rate feedback in real time after execution; if the feedback shows that the pressurization rate has not decreased, it immediately triggers a forced skip mode for physical fallback; at the same time, it packages the entire process data into experience samples and uses experience replay and gradient update mechanisms to fine-tune the deep evaluation model and reinforcement learning agent online, realizing the system's lifelong learning and continuous optimization.

[0024] The specific implementation process of the present invention will be described in detail below with reference to specific embodiments.

[0025] S1. Construction of a Multi-Dimensional Security Situation Awareness Dataset for Energy Storage Containers To address the challenges of rapid thermal runaway evolution and complex environments in energy storage containers, this invention constructs a multi-dimensional sensing system that detects both internally sensed detonation trends and externally sensed environmental vulnerabilities.

[0026] S1-1 Multi-source heterogeneous data acquisition and sensor array deployment: Simultaneously collect internal physical quantities reflecting the evolution characteristics of thermal runaway. High-frequency pressure sensors (sampling rate ≥1kHz) and multi-component gas sensors (CO, H2, VOCs) are deployed at key locations in the top, middle, and bottom of the container's battery compartment to collect the static pressure inside the compartment. The boost rate is obtained in real time through hardware differentiating circuits or edge computing. With pressure acceleration Pressure acceleration This is a key high-order physical quantity for determining whether thermal runaway has transitioned from deflagration to detonation. Simultaneously, a microphone array is deployed to collect acoustic signals to identify characteristic acoustic signatures of safety valve rupture or battery cell ejection. This invention introduces the higher-order derivative of pressure and acoustic signature features to construct an internal state parameter set vector. Its mathematical expression is: ; in, The average static pressure inside the cabin; and These are the pressure rise rate and pressure acceleration, respectively. This is the highest temperature of the battery module, used to help determine the source of thermal runaway; Characteristic gas concentrations are used to assist in early warning systems. The acoustic signature features are specifically extracted using Mel frequency cepstral coefficients to identify specific acoustic events such as safety valve rupture or battery cell ejection.

[0027] Ultrasonic micro-meteorological instruments are deployed on the outside of the container to obtain real-time wind speed. and wind direction angle Deploy binocular vision cameras or LiDAR to identify surrounding sensitive targets (personnel, equipment, buildings) and calculate the Euclidean distance of each target relative to each explosion-proof surface of the container. and azimuth .

[0028] .

[0029] S1-2 Data Cleaning, Signal and Noise Suppression, and Spatiotemporal Feature Tensor Construction Due to the presence of high-voltage electromagnetic interference inside the energy storage container, and the significant difference in sampling frequency between external environmental data and internal monitoring data (internal pressure sensor ≥1000Hz, external visual sensor approximately 30Hz), directly inputting the raw data into the model would lead to gradient explosion or feature misalignment. Therefore, this step performs rigorous signal conditioning and feature reconstruction, specifically including the following three sub-processes: (1) Signal smoothing and denoising based on Savitzky-Golay filtering: Regarding the static pressure data collected by S1-1 Because it is necessary to perform a second-order differential to obtain the pressure acceleration. Conventional mean filtering smooths out critical pressure spikes, leading to distortion of the differential signal. This embodiment employs a Savitzky-Golay digital filter for smoothing. This algorithm, based on least squares fitting of local polynomials, can filter out high-frequency random noise while preserving the transient characteristics of the early stages of thermal runaway to the greatest extent possible. Let the original pressure sequence be... Filtered output sequence The calculation formula is: ; Where 2m+1 is the length of the sliding window (in this embodiment, it is taken as...). =5, which means the window length is 11). As the normalization factor, These are the convolution coefficients. After processing with SG filtering, they are then... Obtained by performing difference operations and Its signal-to-noise ratio is improved by more than 20dB, ensuring the input accuracy of high-order physical quantities.

[0030] (2) Nonlinear normalization mapping of heterogeneous data: Considering that during thermal runaway, pressure values ​​increase exponentially (soaring from 101 kPa to 500 kPa in just a few hundred milliseconds), while temperature and wind speed change linearly, traditional linear normalization would compress early, subtle pressure rises to near zero, making it impossible for the model to recognize early signs. This invention designs a logarithmic nonlinear normalization algorithm for pressure parameters, as shown in the following formula: ; in, The ultimate pressure resistance designed for containers Standard atmospheric pressure To prevent the use of a small, logarithmically meaningless constant (taken as 1e-5), for external wind speed... and distance The standard maximum-minimum normalization method is adopted, which can uniformly map all heterogeneous data to the dimensionless interval [0, 1], while amplifying the weight of early fault features.

[0031] .

[0032] (3) Construction of the spatiotemporal feature tensor of the sliding window: To capture the temporal evolution of thermal runaway, this invention does not use single-point data, but instead constructs a time-sliding window. The length of the historical observation window is set to... (For example, 50 time steps), the normalized internal state vector External environment vector Perform feature concatenation to construct a multidimensional spatiotemporal feature tensor at time t. : ; The final generated tensor has a dimension of ,in This represents the total dimension of the sensor features. This tensor... It will serve as the direct input for the subsequent PINN in-depth evaluation model, incorporating both historical evolutionary trends and current external environmental constraints.

[0033] S2. Design of a Deep Evaluation Model Based on Physical Information Constraints To address the shortcomings of purely data-driven deep learning models, such as poor generalization ability and lack of physical interpretability in predictions when extreme explosion samples are lacking, this step constructs a dual-stream spatiotemporal attention network and introduces the fluid dynamics conservation equation as PINN to achieve control over the thermal runaway energy exponent. EI Accurate predictions. The model structure is as follows: Figure 2 As shown, it specifically includes the following three sub-processes: S2-1 Dual-Stream Spatiotemporal Attention Network Architecture Construction: The multidimensional spatiotemporal feature tensor output by S1-2 is used... As input, a parallel network architecture is constructed that includes internal temporal evolution flow and external spatial risk flow.

[0034] (1) Calculate the internal temporal evolution flow: This step aims to extract the nonlinear dynamic evolution characteristics of pressure and acoustic signature during thermal runaway. The tensor... Internal state components The input is fed into a bidirectional long short-term memory network (Bi-LSTM). Bi-LSTM consists of two LSTM layers, one forward and one backward, capable of simultaneously capturing the historical cumulative effects and future trends of thermal runaway. Let... t The input features at time t are (Right now The hidden layer state is Its update formula is: ; in, This represents a vector concatenation operation. It is used to enhance the representation of pressure acceleration. The focus is on the mutation point (i.e., the detonation critical point), which is addressed by incorporating a temporal attention mechanism after the Bi-LSTM. Attention weights. The calculation formula is: ; ; The final generated internal evolution feature vector The result of weighted summation: .

[0035] (2) Calculate the external space risk flow: This step aims to extract the vulnerability characteristics of the external environment. The tensor... External environment component (Including distance) With wind direction The data is mapped to a graph structure. A spatial relationship graph is constructed with containers as central nodes and sensitive targets as neighboring nodes. A graph attention network (GAT) is used to process this spatial feature. Attention coefficients between nodes are also considered. Euclidean distance With wind direction factor Joint decision: ; By aggregating neighbor node information, an external space risk feature vector is generated. This vector represents the potential collateral damage that a leak could cause in the current environment.

[0036] (3) Fusion of multimodal features: The internal features are fused through a multi-head self-attention layer. External features Dynamic fusion is performed to output high-dimensional hidden layer features. It is mapped to a thermal runaway energy index ranging from [0,1] through a fully connected layer (MLP). .

[0037] S2-2 Constructing a Physical Residual Constraint Model Based on Fluid Dynamics Conservation: In order to ensure that the model still conforms to physical laws under unseen extreme conditions, this invention innovatively introduces physical residual constraints into the loss function of model training.

[0038] Based on the ideal gas law and the law of conservation of mass, a physical equation for pressure change within the energy storage chamber is constructed. This assumes a free internal volume of... Constant, pressure change rate It should be determined by the gas production rate. With leakage rate Decision. Define the physical residual function. as follows: ; in, Automatic differentiation of pressure with respect to time for model prediction; The molar mass of the gas mixture; Based on energy index EI Estimated gas production rate function; This represents the equivalent natural leakage area coefficient for the current container. Based on this calculation... This allows the neural network to predict trajectories that closely approximate the actual fluid dynamic manifold.

[0039] S2-3 Hybrid Loss Function Optimization and Model Training: Constructing the Total Loss Function Data-driven prediction error Residuals driven by physics Weighted composition: ; in, This represents the total number of samples used in training the deep learning model. The physical constraint weight coefficients are set to 0.1 initially in this embodiment and dynamically increased with each training round. The AdamW optimizer is used to iteratively update the network parameters until the total loss function converges.

[0040] The trained deep evaluation model can receive the tensor output by S1 in real time. It outputs a physically explainable thermal runaway energy index within 10ms. This provides a quantitative input benchmark for subsequent tiered explosion relief decisions.

[0041] S3. Staged release of explosion energy and path optimization for minimizing hazards Physically Interpretable Thermal Runaway Energy Index Based on S2 Output This step addresses the joint optimization problem of discharge flux control and discharge path planning. This invention abandons the traditional coarse-grained approach of fully opening the discharge once a threshold is reached, and proposes a hierarchical decision-making strategy based on reinforcement learning. Figure 3 As shown, it specifically includes the following three sub-processes: S3-1 Theoretical Explosion Flux Calculation Based on AI Dynamically Corrected VSP Model To achieve the staged release of explosion energy, this embodiment does not rely entirely on a black-box model. Instead, it uses the classic chemical safety VSP (Vent Sizing Package) explosion venting model as its physical basis and dynamically corrects it using an AI-predicted energy index. First, the theoretical required explosion venting flux is constructed. Computational model: ; in, The free clearance volume of the energy storage container; This refers to the design value for the ultimate pressure resistance of the container structure. The flow coefficient of the explosion vent (usually taken as 0.6~0.8). The maximum boost rate predicted at the current moment is collected by S1. Combination This was deduced.

[0042] This innovatively introduces an AI dynamic correction factor. This factor is a non-linear mapping function that can be used to adjust the degree of explosive venting: ; in, The radical coefficient, The center offset represents a preset risk threshold, used to define the critical decision point for the system to switch from micro-normal fire suppression to active explosion venting protection. At that time, the correction factor As the value approaches 1, the system, according to the calculated theoretical maximum explosion discharge flux, drives a greater number of valves to open, in order to achieve rapid release of explosion energy and ensure the safety of the container structure; when At that time, the correction factor Approaching zero, the system limits the opening area of ​​the valve array, uses the slight positive pressure inside the compartment to suppress the spread of fire, and prevents violent secondary combustion caused by excessive oxygen leakage.

[0043] S3-2 Constructing an environmental hazard cost function that includes distance and wind direction factors. To quantify the potential damage to the external environment caused by the explosion relief operation, this step is based on the set of external environmental parameters collected by S1. Construct an environmental hazard cost function. Assume the distributed explosion relief array has a total of... The valve, the first The position vector of each valve is Define the cost of environmental harm. as follows: ; in, The total number of environmentally sensitive targets identified in the external environment of the energy storage container; Let V be the valve opening state vector. ; For the first The explosion vent and the first The Euclidean distance between several sensitive targets, such as the explosion vent array Figure 4 As shown, in the denominator To prevent the minimum value of division by zero; The angle between the wind direction and the target is the angle between the wind direction vector and the vector pointing from the vent to the target. This is an indicator function; it takes the value 1 when the target is in the downwind sector, and 0 otherwise. and These are distance weight and wind direction weight, respectively.

[0044] Through calculation When the distance to the target is closer ( (The target is large) or the target is located downwind ( The higher the value of the valve in that direction, the greater the cost.

[0045] S3-3 Solving the optimal strategy for valve arrays based on the PPO reinforcement learning algorithm After clarifying the flux constraints and environmental costs, this step models the cooperative control problem of the distributed valve array as a Markov decision process (MDP) and designs a deep reinforcement learning network to solve it.

[0046] (1) The state space of the agent Defined as a combined vector containing the high-dimensional spatiotemporal hidden layer features of the S2 output and the flux to be released at the current moment. Action space Defined as the opening combination vector of the distributed valve array. To guide the agent to minimize environmental impact while satisfying safety constraints, a reward function in the form of "soft reward + hard constraint" is designed: ; The first term aims to minimize environmental harm; the second term is a flux hard constraint penalty, which applies when the actual total activated area is less than the theoretical requirement. At that time, it generates huge negative rewards (due to the coefficient). (Amplification), forcing intelligent agents to prioritize structural safety.

[0047] (2) Network architecture and training objective function: This embodiment adopts an Actor-Critic dual network architecture. The Actor network is the policy network. Input status Output the probability distribution of each valve opening to generate decision actions. Critic networks are value networks. Input status It outputs the value assessment of the current state, which is used to guide the updating of the Actor network.

[0048] To ensure the stability of policy updates and avoid model collapse due to excessively large step sizes, the objective function of the PPO (Proximal Policy Optimization) algorithm is truncated. Conduct training: ; in, The ratio of the old to the new strategies; This is the cutoff range hyperparameter (usually set to 0.2); This is the advantage function estimate, used to measure the advantage of the current action compared to the average level. Through offline pre-training and online inference, the agent can output the optimal activation matrix within milliseconds based on real-time conditions. .

[0049] S4, Execution Control, Closed-Loop Correction and Online Model Evolution This step translates the digital decisions generated by S3 into millisecond-level actions in the physical world, and constructs a fallback mechanism based on physical feedback and a model evolution system based on data feedback to ensure the absolute safety and continuous optimization of the system under extreme conditions. The specific implementation process is as follows: S4-1 arrayed valve actuation and millisecond-level multi-system safety linkage: The controller receives the optimal enable combination matrix output by S3 via the high-speed CAN bus. Immediately afterwards, the multi-system parallel linkage logic is activated. At the valve actuation level, the binary status codes in the matrix are mapped to the solenoid valve actuation levels. Considering the potential back pressure resistance of high-pressure airflow inside the container during the initial explosion phase, the system employs a pulse-width modulation "high-pressure start-low-pressure hold" actuation strategy (e.g., 24V start, 12V hold) to ensure that all target valves complete their full-stroke opening to overcome back pressure within 10ms. At the electrical safety level, upon issuing the explosion venting command... At the synchronization point, the high-voltage interlock circuit of the battery management system is triggered, forcibly disconnecting the main circuit contactor to cut off the electrical connection of the battery cluster, thus blocking the continuous energy injection of the short-circuit current at the source. At the fire alarm linkage level, according to... Valve position index The system queries the fire protection map and activates the corresponding external water sprinkler or nitrogen curtain system. This operation aims to create a "media shielding layer" outside the explosion vent, reducing the probability of the ejected high-temperature, high-pressure gas igniting surrounding combustibles.

[0050] S4-2 is based on a forced skip-level physical fallback mechanism using boost rate differential feedback: Considering the potential probabilistic discrepancies between AI model predictions and actual fire development, this invention designs a physical closed-loop fallback logic independent of AI decision-making. The system sets the evaluation time window after the action is executed as follows: And monitor the change in the rate of pressurization inside the cabin in real time: ; Execute the closed-loop decision function based on this feedback. :like (in If the preset safety threshold is reached, the explosion venting is deemed effective, the energy inside the chamber is released, and the system remains in its current open state; if This indicates that the current opening area is insufficient to suppress the pressure rise (i.e., (The value was underestimated), and the system determined that the AI ​​decision-making had failed. At this point, a forced skip mode was immediately triggered. The controller will ignore the environmental hazard cost function in S3 and perform a set union operation. It forces the sending of opening commands to all standby valves that are in the closed state, establishing the prevention of the container body from disintegration as the highest and only priority for physical backup.

[0051] S4-3 Online Evolutionary Model System Based on Experience Replay and High-Value Samples: To enable the system to learn throughout its life, this step establishes an iterative mechanism for model parameters based on real-world data processing.

[0052] (1) Construction and hierarchical storage of experience samples: The system automatically packages the entire process data of this event into an experience sample. ,in This is a boolean flag indicating whether to trigger a forced overriding. If... The sample was marked as a "high-value negative sample" and stored in the priority experience pool. Otherwise, store in the regular experience pool. .

[0053] (2) Online model fine-tuning: During system standby, samples are drawn from the priority experience pool to perform backpropagation updates on the networks in S2 and S3. For the S2 deep evaluation model, the large physical residuals caused by negative samples are utilized. Update PINN parameters to correct the model's assessment of gas production rate. The physical cognitive bias. For the S3 reinforcement learning strategy, the large negative reward generated by negative samples (this value comes from the flux hard constraint penalty) is utilized. The PPO network parameters are updated using the policy gradient descent algorithm. : ; in The learning rate is used as the system operates. Through this mechanism, the system can continuously adjust its decision boundaries over time, reducing the probability of choosing a conservative strategy under similar conditions and achieving adaptive evolution of its intelligent decision-making capabilities.

[0054] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0055] While the above description illustrates specific embodiments of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for active explosion venting control of a distributed array of energy storage containers, characterized in that, The process includes the following: S1 collects multi-dimensional safety situation awareness data of the energy storage container under its operating status, including internal thermal runaway state data and external environmental vulnerability data, and preprocesses it; S2, the preprocessed data is input into the deep evaluation model based on physical information constraints. The model adopts a dual-stream spatiotemporal attention network architecture: it uses a bidirectional long short-term memory network to extract the nonlinear temporal evolution features of the internal thermal runaway state data, uses a multilayer perceptron or graph network to extract the spatial topological features of the external environment vulnerability data, and uses a multi-head self-attention mechanism to dynamically weight and fuse the multi-source features, and finally outputs a normalized thermal runaway energy index. S3. The thermal runaway energy index is substituted into the VSP explosion relief model as a dynamic correction factor to calculate the theoretical explosion relief flux required to maintain the safety of the container structure. An environmental hazard cost function containing personnel distance factor and wind direction factor is constructed. Under the premise of satisfying the theoretical explosion relief flux constraint, a deep reinforcement learning agent is used to solve the distributed valve array opening combination matrix that minimizes the environmental hazard cost. S4, the controller drives the valve to operate according to the matrix and continuously monitors the pressure rise rate feedback in the chamber. When the pressure rise rate does not show a downward inflection point, it triggers the forced skip mode.

2. The active explosion venting control method for a distributed array of energy storage containers as described in claim 1, characterized in that: The internal thermal runaway state data includes the internal static pressure, pressurization rate, pressure acceleration, characteristic gas concentration, and acoustic signature. A high-frequency sensor array deployed within the container battery compartment synchronously collects internal static pressure and characteristic gas concentration data at a sampling frequency of at least 100Hz. The pressurization rate is obtained by performing first-order differential calculations on the internal static pressure data, and the pressure acceleration is obtained by performing second-order differential calculations. The pressure acceleration is used to characterize the critical trend of the thermal runaway combustion mode transitioning from deflagration to detonation. Acoustic signals within the compartment are collected via a microphone array, and Mel-frequency cepstral coefficients are extracted to obtain acoustic signature characteristics characterizing safety valve rupture or battery cell ejection events. Finally, the static pressure, pressurization rate, pressure acceleration, characteristic gas concentration, and acoustic signature were aligned according to the timestamp.

3. The active explosion venting control method for a distributed array of energy storage containers as described in claim 1, characterized in that: The external environment vulnerability data includes environmental wind speed, wind direction angle, and spatial distance coordinates of sensitive targets relative to each explosion relief surface; Real-time acquisition of ambient wind speed and wind direction angle is achieved using micro-meteorological instruments deployed on the exterior of the container; sensitive targets, including personnel, key equipment, and adjacent containers, are identified using a perimeter visual monitoring system, and the Euclidean distance and azimuth angle of each sensitive target relative to each explosion-proof surface of the container are calculated; a local environmental field vector centered on the container is constructed, which includes the wind field vector and the spatial coordinates of all sensitive targets; the maximum value of the ambient wind speed is normalized, and the distance data is transformed by reciprocal transformation to enhance the weight of nearby targets.

4. The active explosion venting control method for a distributed array of energy storage containers as described in claim 1, characterized in that: The dual-stream spatiotemporal attention network architecture specifically includes a parallel temporal flow subnetwork and a spatial flow subnetwork: the temporal flow subnetwork adopts a stacked bidirectional long short-term memory network structure, and the input is a temporal sequence of internal thermal runaway state data, which is used to capture the long-term dependence of pressure and gas parameters in the time dimension and the instantaneous change characteristics; the spatial flow subnetwork adopts a graph attention network structure, which models the container explosion relief surface and external sensitive targets as graph nodes, and the edge weights between nodes are determined by the spatial distance and wind direction factor, which is used to extract the spatial topological risk characteristics of the external environment; Finally, through a multi-head self-attention mechanism layer, attention weights are dynamically calculated based on the signal-to-noise ratio of each sensor channel, and temporal flow features and spatial flow features are weighted and fused to generate a high-dimensional spatiotemporal hidden layer feature vector.

5. The active explosion venting control method for a distributed array of energy storage containers as described in claim 4, characterized in that: The deep evaluation model introduces a physical residual constraint PINN based on the fluid dynamics conservation equation during the training phase. Specifically, a physical constraint term is introduced into the loss function of the deep evaluation model. This physical constraint term is constructed based on the ideal gas law and the mass conservation equation. The specific calculation method is as follows: the pressure value inside the chamber predicted by the model at the next moment is substituted into the physical equation, and the physical residual between it and the current pressure, gas generation rate, and explosion venting rate is calculated. The total loss function is composed of the data fitting error and the physical residual weighted. By minimizing the total loss function, the model is trained, forcing the weight update direction of the neural network to conform to the fluid dynamics physical manifold.

6. The active explosion venting control method for a distributed array of energy storage containers as described in claim 1, characterized in that: The VSP explosion relief model in S3 specifically involves mapping the thermal runaway energy index output by the deep evaluation model to a dynamic correction factor; and calculating the baseline explosion relief area based on the VSP model formula using the container's free volume, the maximum current pressurization rate, the design withstand pressure, and the flow coefficient. Multiplying the baseline explosion relief area by the dynamic correction factor yields the theoretical explosion relief flux required to prevent the container structure from disintegrating at the current moment. The dynamic correction factor exhibits a nonlinear positive correlation with the thermal runaway energy index. When the energy index reaches its extreme value, the correction factor tends to 1, thereby maximizing the theoretical explosion relief flux.

7. The active explosion venting control method for a distributed array of energy storage containers as described in claim 1, characterized in that: The environmental hazard cost function in S3 is composed of a weighted sum of a distance cost term and a wind direction cost term. The distance cost term is inversely proportional to the square of the distance to the sensitive target, indicating that the closer the distance, the greater the hazard. The wind direction cost term is determined by the cosine of the angle between the explosion venting direction and the environmental wind direction, indicating the potential risk of the sensitive area being upwind or pointing downwind.

8. The active explosion venting control method for a distributed array of energy storage containers as described in claim 1, characterized in that: In S3, a deep reinforcement learning agent is trained using a near-end policy optimization algorithm. The current spatiotemporal hidden layer feature vector is used as the state input, and the opening combination of the distributed valve array is output as the action. In the reward function design, the theoretical required discharge flux is used as a hard penalty constraint, and the environmental hazard cost is used as a positive reward objective, thereby solving for the optimal opening combination matrix.

9. The active explosion venting control method for a distributed array of energy storage containers as described in claim 1, characterized in that: In S3, after executing the optimal opening combination matrix, a time window is set to continuously monitor the differential change of the pressurization rate inside the compartment. If the differential value of the pressurization rate is non-negative within the time window, it is determined that the current explosion relief strategy has failed or the risk has been underestimated, and the forced skip mode is immediately triggered. The logic of the forced skip mode is: ignoring the environmental hazard cost function, forcibly sending opening commands to all standby explosion relief valves that are in the closed state to achieve the maximum throughput release at the physical level, so as to prevent the disintegration of the container body structure as the highest priority protection target.

10. The active explosion venting control method for a distributed array of energy storage containers as described in claim 1, characterized in that: It also includes an online model evolution step, specifically: packaging the entire process data of each leak event, including the input multi-dimensional security situation awareness data, the decision actions output by the deep reinforcement learning agent, and the boost rate feedback results after execution, into an experience sample; if a forced skip mode is triggered, the sample is marked as a high-value negative sample and stored in the experience replay pool. This experience replay pool is used to fine-tune the deep evaluation model and reinforcement learning agent online, correcting the model parameters.

Citation Information

Patent Citations

  • Intelligent explosion venting monitoring and control system for energy storage device

    CN117078489A