Real-time monitoring and fault response control device of industrial and commercial liquid cooling energy storage system

Through multi-source data acquisition and feature engineering, combined with long and short-term memory networks and Weibull reliability models, the early warning threshold and maintenance time window are dynamically adjusted, which solves the problem of difficult detection and inaccurate maintenance timing of weak abnormal signals in liquid-cooled energy storage systems, and achieves accurate prediction and efficient response of faults.

CN120335425AActive Publication Date: 2025-07-18ZHEJIANG CHUANGQI NEW ENERGY TECH CO LTD

Patent Information

Application Number
CN202510414065.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

Liquid-cooled energy storage systems are difficult to accurately detect weak abnormal signals and predict faults under complex operating conditions, resulting in inaccurate maintenance timing and difficult traditional methods to adapt to dynamic operating environments and diverse fault modes.

Method used

Through multi-source data acquisition and feature engineering, long-term and short-term memory networks and Weibull reliability models are used, combined with multi-physics coupled models and graph neural networks, the early warning thresholds and maintenance time windows are dynamically adjusted to achieve accurate identification of weak abnormal signals and dynamic prediction of faults.

Benefits of technology

It improves the accuracy of fault detection and maintenance efficiency of liquid-cooled energy storage systems, reduces maintenance costs, and improves the reliability and response speed of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335425A_ABST
    Figure CN120335425A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time monitoring and fault response control device for an industrial and commercial liquid cooling energy storage system, and particularly relates to the technical field of liquid cooling energy storage fault management. Temperature, pressure, vibration and flow data of key components in a liquid cooling system are collected and preprocessed; a time sequence analysis and signal processing technology is utilized to extract feature vectors reflecting key component states from the multi-dimensional operation data set, the saliency of weak abnormal signals is enhanced, and a key feature set containing weak abnormal features is generated; constructing an anomaly recognition model by adopting a long-short-term memory network, inputting a key feature set, outputting an anomaly probability corresponding to the key component, and setting an early warning threshold value of the anomaly probability to trigger early warning; and predicting a fault maintenance time window of the key component based on a Weibull reliability model, and dynamically adjusting the priority of maintaining the key component and executing preventive maintenance based on a fault maintenance time window prediction result and an abnormal probability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of liquid-cooled energy storage fault management. More specifically, the present invention relates to a real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system. Background Technique

[0002] As an important part of modern energy management, industrial and commercial liquid-cooled energy storage systems are widely used in fields such as power peak shaving, renewable energy grid connection, and industrial load balancing. The system regulates the temperature of battery units and key components through liquid-cooling technology to ensure efficient operation and long-term stability. Key components such as battery units, cooling pipes, and pump valves operate under complex working conditions and need to withstand multiple stresses such as high temperature, high pressure, and vibration. The change of their states directly affects the safety and reliability of the system. To ensure the operation of the system, it is necessary to monitor multi-dimensional operation data such as temperature, pressure, vibration, and flow in real time, predict potential faults based on the monitoring results, and implement preventive maintenance. However, due to the dynamics of the operating environment and the diversity of fault modes, traditional monitoring and response mechanisms have significant limitations in dealing with weak abnormal signals and complex fault evolutions.

[0003] In the prior art, the liquid-cooled energy storage system has the problem of inaccurate maintenance timing due to complex and changeable working conditions. Specifically, traditional methods usually rely on static thresholds or single data sources to judge abnormalities, and it is difficult to accurately capture weak signals such as early deterioration of battery units or small fluctuations in coolant flow, resulting in a lag in anomaly detection; at the same time, the fault prediction model lacks the ability to dynamically adapt to real-time operation data, resulting in a large deviation in the prediction of the remaining service life, and it is difficult to match the maintenance time window with the actual fault process. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system. By integrating multi-source data acquisition and feature engineering technologies, it monitors the states of key components in real time, uses long short-term memory networks to accurately identify weak abnormal signals and output abnormal probabilities, combines the Weibull reliability model to dynamically predict the remaining service life and generate a fault maintenance time window, and at the same time optimizes the maintenance priority based on the abnormal probability and the prediction result of the time window to solve the problems raised in the above background technique.

[0005] To achieve the above object, the present invention provides the following technical solution: A real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system, comprising:

[0006] A multi-source data acquisition and preprocessing module: By deploying sensors, it collects the temperature, pressure, vibration, and flow data of key components in the liquid-cooled system in real time, and preprocesses the collected data to generate a denoised and standardized multi-dimensional operation data set;

[0007] Feature Engineering and Weak Signal Extraction Module: Using time series analysis and signal processing techniques, extract feature vectors reflecting the states of key components from multi-dimensional operation datasets, enhance the significance of weak abnormal signals, and generate a key feature set containing weak abnormal features;

[0008] Abnormal Pattern Recognition Module: Based on the key feature set, construct an abnormal recognition model using a long short-term memory network. Build a training set with historical operation data to obtain a trained abnormal recognition model; Output the abnormal probability corresponding to the key component, and set a warning threshold for the abnormal probability to trigger a warning;

[0009] Fault Warning and Optimization Module: Predict the remaining service life of key components based on the Weibull reliability model, generate a fault maintenance time window in combination with the confidence interval. Based on the prediction results of the fault maintenance time window and the abnormal probability, dynamically adjust the priority of maintaining key components and perform preventive maintenance; Use the operation data of key components after maintenance to update the abnormal recognition model and the Weibull reliability model to form a closed-loop optimization control.

[0010] Preferably, the method for obtaining the fault maintenance time window is as follows:

[0011] Data Acquisition: Obtain the average abnormal probability and historical fault data of key components;

[0012] Model Initialization: Initialize the parameters of the Weibull reliability model. Since it is suitable for describing the failure law of mechanical components, the parameters include the shape parameter β, which characterizes the failure mode; and the scale parameter η, which characterizes the life; Fit the Weibull model according to historical data and adjust the model parameters in combination with the current abnormal probability; Optimize the parameters through maximum likelihood estimation or the least squares method to ensure that the Weibull reliability model reflects the current state;

[0013] Life Prediction: Based on the adjusted Weibull reliability model and abnormal probability, calculate the remaining service life of the component through the following formula and generate a fault maintenance time window;

[0014]

[0015] Where t is time, F(t) is the abnormal probability, t is the remaining service life, and t is inversely deduced from the abnormal probability to obtain the remaining service life;

[0016] Set the confidence interval of the remaining service life, calculate the upper and lower limits of the remaining service life through the inverse function of the Weibull reliability model, and generate a fault maintenance time window; Verify that the trend of the remaining service life is consistent with the abnormal probability.

[0017] Preferably, the training process of the anomaly recognition model includes the following steps:

[0018] Step S1, Dataset Preparation and Partitioning: Prepare the historical dataset of key components, extract the key feature set, mark the fault time points, and partition it into a training set, a validation set, and a test set;

[0019] Step S2, Time Series Data Preprocessing and Input Construction: Perform time series processing on the historical dataset to generate an input format suitable for the long short-term memory network; Based on the key feature set, use the sliding window method to construct time series inputs, generating sequence data containing historical trends; Normalize the sequence data and map the feature values to the interval [0, 1]; Output the standardized time series input sequence to capture the cumulative effect of weak signals;

[0020] Step S3, Model Initialization and Training: Set the long short-term memory network structure, including an input layer, a hidden layer, and an output layer, and output the anomaly probability; Use the mean squared error as the loss function; Train the short-term memory network through the Adam optimizer and the loss function to obtain a trained anomaly recognition model;

[0021] Step S4, Adjust the hyperparameters of the anomaly recognition model through the validation set to reduce the loss function value; Use the cross-validation method to evaluate the generalization ability of the short-term memory network to ensure the accuracy of anomaly recognition; Determine the warning threshold of the anomaly probability according to the validation results.

[0022] Preferably, the device further includes: a warning threshold dynamic management module, which solves the problem of the fixed threshold failure caused by the working condition change of the liquid cooling system through warning threshold dynamic management, including:

[0023] Based on the multi-dimensional operation dataset, collect the multi-dimensional operation dataset of key components in real time to generate a time series data set containing multiple parameters;

[0024] After normalizing each parameter, calculate the fluctuation characteristics of each parameter, decompose the signal using the time-frequency analysis method based on wavelet transform, extract the high-frequency fluctuation component and the low-frequency trend feature, generate the multi-dimensional fluctuation feature vector v, and construct the dynamic threshold Y DT Formula:

[0025] Y DT = α·EWMA(v)+β·MAD(v)×(1+γ·CV(v)) where EWMA(v) represents the smoothed trend index, which is the weighted average of the fluctuation trends of each parameter; MAD(v) represents the dispersion; CV(v) adjusts the fluctuation intensity, which is the ratio of the standard deviation of each parameter to the mean of each parameter, and α, β, and γ are adjustment coefficients.

[0026] Preferably, the feature engineering and weak signal extraction module includes:

[0027] Decompose the signal of the multi-dimensional operation data set by wavelet transform to separate the high-frequency noise and the low-frequency trend component;

[0028] Extract the intrinsic mode functions of the weak anomaly signal based on empirical mode decomposition, calculate the instantaneous frequency and instantaneous amplitude of each intrinsic mode function through Hilbert transform, and generate a multi-dimensional feature vector reflecting the anomaly trend;

[0029] Reduce the dimension of the multi-dimensional feature vector through principal component analysis to form a key feature set.

[0030] Preferably, the training process of the anomaly recognition model further includes:

[0031] Introduce an attention mechanism into the long short-term memory network, and assign weights to each time step based on the characteristics of the weak anomaly signal to highlight the influence of the weak anomaly signal on the anomaly probability;

[0032] Set a bidirectional LSTM structure to capture historical and future trends simultaneously and dynamically adjust the number of hidden layer neurons;

[0033] Adopt Dropout regularization to prevent overfitting and combine with the early stopping strategy.

[0034] It should be explained that the attention mechanism (Attention Mechanism) is a calculation method commonly used in neural networks, aiming to enhance the model's ability to focus on important parts of the input sequence; in the present invention, the attention mechanism is used in the long short-term memory network (LSTM). By calculating the weights of the data at each time step, it highlights the contribution of the weak anomaly signal (such as small changes in temperature or vibration) to the prediction of the anomaly probability, and generates a context vector by weighted summation of the input features.

[0035] The specific steps for assigning weights to each time step are as follows:

[0036] The input is the time series data of the key feature set (such as feature vectors of temperature, pressure, etc.), denoted as X = [x1, x2,..., x T , where T is the total number of time steps;

[0037] The hidden layer of the LSTM outputs the hidden state ht (in vector form) at each time step;

[0038] Calculate the attention weight α t : Use a single-layer feedforward neural network (such as a neural network with a tanh activation function), and the input is ht and the context vector c at the previous moment t-1 The formula is

[0039] e t = W a ·ht +b a ,

[0040] where W a and b a are trainable parameters, and α t is the normalized weight; e t represents the attention score at any time step k in the time series, where k is the index of the time step, ranging from 1 to T;

[0041] Generate a context vector through weighted summation as the input for anomaly probability prediction, highlighting the time steps of weak anomaly signals (such as amplitude mutations) and increasing the corresponding weights; for example, if the temperature suddenly changes from 25°C to 28°C at a certain time step, the calculated α t may increase from 0.05 to 0.15, and the specific weight value is determined by the training data to ensure that the anomaly signal is amplified.

[0042] Explanation: The bidirectional LSTM is an improved LSTM structure. The core is to input the input sequence into two independent LSTM networks both forward and backward, generating forward hidden states and backward hidden states respectively, and finally concatenating them into the comprehensive hidden state for each time step, which is used for anomaly probability prediction; in the present invention, the complete dependency relationship of the sequence is captured through forward and backward processing, such as the cause and effect of pump pressure anomalies.

[0043] Explanation on how to capture historical and future trends: The bidirectional LSTM captures the historical trend from the past to the present (such as the cumulative effect of gradually increasing temperature) through the forward LSTM, and captures the trend from the future to the present (such as the reverse impact of abnormal fluctuations before a failure) through the backward LSTM. The historical trend from the past to the present and the trend from the future to the present constitute the forward and backward dependency relationship in the sequence, so as to more accurately predict the anomaly probability; for example, if the pump vibration suddenly increases at t = 10, the backward LSTM can trace the impact of the pressure anomaly at t = 11, and the forward LSTM captures the trend before t = 9, comprehensively improving the prediction accuracy.

[0044] Explanation on how to dynamically adjust the number of hidden layer neurons. Dynamically adjusting the number of hidden layer neurons means adjusting the network capacity of the LSTM according to the performance of the validation set to balance the model complexity and the risk of overfitting. The specific implementation method can be:

[0045] Initially set two hidden layers, with 64 neurons in each layer;

[0046] During the training process, use the validation set to evaluate the loss function (such as mean squared error). If the loss decreases slowly (less than 0.01 for 5 consecutive iterations), increase the number of neurons in each layer (such as increasing to 128); if there are signs of overfitting (such as the validation set loss increasing), reduce the number of neurons (such as reducing to 32); the adjustment is achieved through grid search within the range of [32, 64, 128], and retrain and validate after each adjustment; for example, if the validation set loss stagnates after 50 iterations, try increasing the number of neurons to 128 and observe the change in loss. The adjustment process can be automated through a Python script;

[0047] Explanation on how to prevent overfitting with Dropout regularization: Reduce the model's over - reliance on specific features by randomly discarding neurons during training (the dropout rate is set to 0.2, that is, 20% of the neurons are set to zero). The specific implementation method is as follows:

[0048] Add a Dropout layer after each hidden layer of the LSTM;

[0049] During training, randomly discard neurons with a dropout rate of 0.2 and their connections, and only retain 80% of the neurons to participate in the calculation;

[0050] During testing, all neurons participate in the calculation, but the output is multiplied by 0.8 to maintain the expected consistency; this method reduces parameter redundancy and avoids overfitting.

[0051] Explanation on how to combine the early stopping strategy. The early stopping strategy is a general regularization technique used to prevent over - training. The specific implementation method is as follows:

[0052] During the training process, calculate the validation set loss (such as mean squared error) after each iteration;

[0053] Set the patience value to 10, that is, if the validation set loss does not decrease for 10 consecutive times (the decrease threshold is set to 0.001), then stop training;

[0054] Record the current best model parameters (such as the weights when the loss is the lowest), and use this as the final anomaly recognition model; for example, if the loss at the 80th iteration is 0.05 and does not decrease in the subsequent 10 times, then save the model parameters at the 80th time to ensure that the model does not overfit.

[0055] Preferably, according to the anomaly probability output by the anomaly pattern recognition module, adjust the sampling frequency of the sensor.

[0056] Preferably, the early warning threshold dynamic management module further includes:

[0057] Predict the short - term fluctuation trend of the multi - dimensional operation data set based on the Markov chain, and calculate the state transition probability of each parameter in the future time;

[0058] Adjust the dynamic threshold formula in combination with the prediction results, introduce the trend factor δ. When the fluctuation trend is rising, δ < 1 to lower the warning threshold, and when it is falling, δ > 1 to raise the warning threshold.

[0059] Preferably, the device further includes: a fault maintenance time window optimization module for obtaining an optimized fault maintenance time window, including:

[0060] Obtain the abnormal probability of key components in the liquid cooling system through real-time monitoring, and generate a comprehensive input data set reflecting the component status in combination with historical operation data and the current multi-dimensional operation data set;

[0061] Utilize the multi-physics field coupling model to simulate the thermal, mechanical, and fluid mechanical behaviors of key components under the current abnormal probability based on the comprehensive input data set and the abnormal probability;

[0062] Generate a stress-strain distribution map through finite element analysis, predict the fatigue crack propagation time in combination with the abnormal probability, and generate a virtual life curve of the component status changing with time. The virtual life curve is used to characterize the dynamic deterioration process and potential fault characteristics of key components;

[0063] Combine the virtual life curve with the Weibull reliability model, fuse real-time monitoring data and historical operation data through the Bayesian update method, correct the parameters of the Weibull reliability model, and output an optimized Weibull reliability model;

[0064] Generate an optimized fault maintenance time window based on the optimized Weibull reliability model.

[0065] Explanation: The comprehensive input data set refers to the data set obtained by integrating real-time monitoring data, historical operation data, and abnormal probability for input into the multi-physics field coupling model. The specific content includes:

[0066] Real-time monitoring data: The temperature, pressure, vibration, and flow values (such as temperature 25°C, pressure 2 bar) of key components (such as pumps, valves) collected by sensors;

[0067] Historical operation data: The status records of component past operations (such as average life 10,000 hours, failure frequency distribution);

[0068] Abnormal probability: A value between 0 and 1 (such as 0.85) output by the abnormal pattern recognition module, reflecting the current fault risk;

[0069] Data integration method: Align the above data according to the timestamp to form a multi-dimensional vector (such as [t, T, P, V, F, P_anomaly]), and map it to the interval [0, 1] through normalization processing (such as Min-Max normalization) as the model input.

[0070] Explanation on how to couple a multi-physics coupling model: A multi-physics coupling model refers to simultaneously analyzing the interactions of thermal, mechanical, and fluid mechanics behaviors through numerical simulation methods (such as COMSOL Multiphysics software). The specific coupling steps are as follows:

[0071] Model definition: Establish a three-dimensional geometric model of key components (such as the impeller of a pump) and set material properties (such as density, thermal conductivity);

[0072] Physical field setting: Build a thermal field model based on the heat conduction equation and input the temperature boundary conditions; build a mechanical field based on the elastic mechanics equation and input the stress caused by vibration or pressure; build a fluid mechanics field model based on the Navier-Stokes equation and input the flow rate and pressure conditions;

[0073] Coupling mechanism: Achieve coupling through shared variables. For example, temperature changes affect material stress (thermal expansion coefficient), and fluid pressure affects stress distribution;

[0074] Solution: Use the finite element method for discretization and iteratively solve the steady-state or transient solutions of each physical field, and output the temperature field, stress field, and flow field.

[0075] Explanation on how to generate a virtual life curve through finite element analysis, including the following steps:

[0076] Stress-strain distribution: Based on the output of the multi-physics coupling model (such as stress 100 MPa), generate a stress-strain distribution diagram of the component through FEA (such as ANSYS software);

[0077] Prediction of fatigue crack growth: Adjust the initial defect assumption (such as crack length 0.1 mm) by combining the anomaly probability (such as 0.85), and calculate the crack growth rate using the Paris law;

[0078] Life calculation: Calculate the number of cycles N by integrating the crack growth equation from the initial crack length to the critical length (such as 1 mm);

[0079] Curve generation: Use time as the horizontal axis (convert the number of cycles to hours according to the operating conditions), and the remaining life as the vertical axis to plot a virtual life curve (such as decreasing from 1000 hours to 0 hours) to characterize the dynamic deterioration process.

[0080] Preferably, the device further includes a fault maintenance trigger condition management module for managing the trigger conditions of fault maintenance. The trigger conditions are the basis for determining the start time window of fault maintenance and include:

[0081] Obtain the dynamic simulation data of battery cells and key components generated based on the multi-physical field coupling model;

[0082] Analyze the fault propagation characteristics in the dynamic simulation data and construct an adaptive fault evolution network. The adaptive fault evolution network identifies the fault propagation paths and propagation intensities between components through a graph neural network and generates an evolution topology structure reflecting the multi-point fault synergy effect;

[0083] According to the evolution topology structure, fuse real-time sensor data and historical fault modes, and use a contrastive learning mechanism to judge the matching degree between the current operating state and potential fault modes, and generate a time series probability distribution of fault evolution;

[0084] If the time series probability distribution indicates that the potential fault mode exceeds the preset threshold, based on the collaborative optimization strategy of Monte Carlo simulation and reinforcement learning, dynamically adjust the trigger conditions of the fault maintenance time window.

[0085] Explanation: It is not clear which physical quantities are included in the dynamic simulation data. The dynamic simulation data refers to the time series data output by the multi-physical field coupling model, specifically including:

[0086] Battery cells: temperature, pressure, charge distribution;

[0087] Key components (such as pumps): stress, strain, flow rate, vibration acceleration;

[0088] Data form: multi-dimensional vectors recorded at time steps (such as every minute), reflecting the dynamic behavior of components.

[0089] Explanation: It is not clear what the fault propagation characteristics mean. The fault propagation characteristics refer to the law of fault transmission between components. For example, a pump fault causes abnormal valve pressure, and the specific manifestations are:

[0090] Propagation path: the causal relationship of the fault from the source component (such as a pump) to the affected component (such as a valve);

[0091] Propagation intensity: the degree of fault influence (such as a 10% increase in pressure);

[0092] Extraction method: analyze the time correlation of physical quantities in the dynamic simulation data (such as a stress mutation after a temperature increase).

[0093] Explanation: It is not clear what the fault evolution network means. The fault evolution network is a graph structure-based model used to characterize the dynamic process of fault propagation between components. Specifically:

[0094] Nodes: Key components in the liquid cooling system (such as pumps, valves);

[0095] Edges: Fault propagation relationships between components (such as pump failures affecting valves);

[0096] Construction method: Use graph neural networks to process dynamic simulation data and learn propagation paths and intensities.

[0097] Explanation: Regarding the meaning of the evolving topological structure and how it is generated is not clear. The evolving topological structure is a graphical representation of the fault evolution network over time, reflecting the synergistic effects of multi-point failures. The specific generation steps are as follows:

[0098] Input: Dynamic simulation data and real-time sensor data;

[0099] GNN processing: Update node features (such as pump stress) and edge weights (such as propagation intensity) through graph convolution. The formula is:

[0100]

[0101] Among them, is the adjacency matrix, H is the node feature, W is the weight, and σ is the activation function;

[0102] Output: A topological graph that evolves over time (such as the edge weight between the pump and the valve increasing from 0.3 to 0.8);

[0103] Explanation: Regarding how to use the contrastive learning mechanism to match the current operating state with potential fault modes is not clear; the contrastive learning mechanism matches the current operating state with potential fault modes by judging similarity, specifically including:

[0104] Input: The current operating state (real-time sensor data, such as a temperature of 30°C) and historical fault modes (such as a fault at a temperature of 40°C);

[0105] Feature extraction: Generate embedding vectors (such as 128-dimensional) of the current state and fault modes through GNN;

[0106] Loss function: Use the InfoNCE loss to calculate similarity: Output a matching degree score (from 0 to 1) and generate a fault probability distribution.

[0107] Explanation: Regarding how to optimize synergistically based on Monte Carlo simulation and reinforcement learning is not clear. Monte Carlo simulation (MonteCarlo) and reinforcement learning (RL) are combined to optimize the triggering conditions of the fault maintenance time window, achieving a balance between cost and risk by simulating the fault distribution and dynamically adjusting the threshold. The specific implementation method is as follows:

[0108] First, the Monte Carlo simulation is based on the failure probability distribution. It randomly samples multiple times (e.g., 1000 times) to simulate the failure occurrence time and generates a time window distribution of high risk (e.g., 10 - 14 hours). The specific implementation steps are as follows:

[0109] Input the current failure probability distribution, such as a discrete distribution where the probability changes with time (e.g., the probability is 0.7 at t = 10 hours and 0.9 at t = 12 hours);

[0110] Conduct multiple random samplings (e.g., 1000 times). Each sampling generates a failure occurrence time point from the probability distribution;

[0111] Statistically analyze the sampling results to generate a time window distribution of failure occurrences. For example, the results show that the failure may occur within 10 - 14 hours, with a mean of 12 hours and a standard deviation of 1 hour;

[0112] Output the time window distribution as a reference for reinforcement learning. For example, the time window [10, 14] hours represents a high - risk interval for potential failures;

[0113] Then, reinforcement learning uses the Q - learning algorithm to optimize the triggering condition and finds the optimal threshold through iterative learning. The specific steps are as follows:

[0114] Define the state as the current anomaly probability and time window, the action as adjusting the height of the triggering threshold, and the reward as maintaining the balance between the reduction of maintenance cost and the reduction of failure risk;

[0115] Initialize the Q - table to zero. Through multiple iterations, each time select an action according to the current state, calculate the reward, and update the Q - table value according to the Q - learning formula; in each update, repeat the iteration using the latest state and reward information until the Q - table converges, and output the triggering threshold with the maximum reward in each state as the optimal result;

[0116] Finally, select the action with the maximum reward in each state through the Q - table and output the optimized triggering threshold; for example, if the Q - table after training shows that the Q - value of action 0.75 is the highest in the state (probability 0.85, window [10, 14]), then set the triggering threshold to 0.75; this triggering threshold can be dynamically adjusted according to the system state. For example, when the probability rises to 0.9, it may be further optimized to 0.7 to ensure timely response.

[0117] Explanation: Regarding how to dynamically adjust the triggering condition of the failure maintenance time window is not clear. Based on updating the Q - table to dynamically adjust the triggering condition, the specific implementation method is as follows

[0118] Adjustment when the risk increases:

[0119] If Monte Carlo simulation shows an increased risk of failure (e.g., the mean of the probability distribution rises from 0.7 to 0.9) and the time window shrinks from [11, 15] hours to [9, 12] hours, then lower the trigger threshold to respond earlier; for example, adjust the original threshold of 0.8 to 0.75, and make a decision to maximize the Q-value based on the state (0.9, [9, 12]) in the Q-table.

[0120] Adjustment when the cost is too high:

[0121] If reinforcement learning optimization shows that the current threshold is too low, resulting in an increase in maintenance costs, for example, the threshold of 0.75 triggers too frequently and the cost rises from 80 yuan to 120 yuan and the reward value decreases, then raise the threshold to 0.85; specifically, evaluate according to the Q-table. For example, the Q-value of the action 0.85 in the state (0.85, [10, 14]) is higher than 0.75, reflecting a better balance of cost - risk.

[0122] Run the Monte Carlo simulation again every hour based on the latest sensor data and the probability of anomalies, update the time window distribution, and then iteratively adjust the Q-table through Q-learning to ensure that the trigger threshold adapts to the system state.

[0123] Technical effects and advantages of the present invention:

[0124] (1) The real-time monitoring and fault response control device for the industrial and commercial liquid-cooled energy storage system provided by the present invention extracts weak anomaly features through multi-source data acquisition and feature engineering, outputs the probability of anomalies using long short-term memory networks and attention mechanisms, dynamically adjusts the warning threshold in combination with Markov chains, optimizes the maintenance time window through the Weibull model and realizes closed-loop control, effectively solving the problems of difficult detection of weak anomaly signals, failure of fixed thresholds, and inaccurate maintenance timing.

[0125] (2) The real-time monitoring and fault response control device for the industrial and commercial liquid-cooled energy storage system provided by the present invention analyzes the thermal runaway risk using a multi-physical field coupling model and generates a virtual life curve, optimizes the maintenance time window through the Weibull model and Bayesian update; uses graph neural networks, contrastive learning, and Monte Carlo simulation and reinforcement learning to dynamically adjust the trigger conditions, outputs maintenance decisions for unknown faults, realizes accurate fault prediction and dynamic response, and improves the reliability and maintenance efficiency of the system. Description of the Drawings

[0126] Figure 1 It is a structural block diagram of the monitoring and control device for the liquid-cooled energy storage system of the present invention.

[0127] Figure 2 It is a flow chart of the fault maintenance time window optimization method of the present invention. Detailed Embodiments

[0128] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully communicated to those skilled in the art.

[0129] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0130] The following description of at least one exemplary embodiment is actually merely illustrative and in no way limits the present application or its application or use.

[0131] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the specification.

[0132] Example 1, referring to Figure 1 the structural block diagram of the liquid-cooled energy storage system monitoring and control device, the present invention provides a Figure 1 real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system as shown in

[0133] Multi-source data acquisition and preprocessing module: By deploying high-precision sensors, it collects the temperature, pressure, vibration, and flow data of key components in the liquid-cooled system in real time, and preprocesses the collected data to generate a denoised and standardized multi-dimensional operation dataset;

[0134] Feature engineering and weak signal extraction module: Using time series analysis and signal processing techniques, it extracts feature vectors reflecting the states of key components from the multi-dimensional operation dataset, enhances the significance of weak abnormal signals, and generates a key feature set containing weak abnormal features;

[0135] Explanation: In a liquid-cooled energy storage system, key components (such as pumps and valves) may only exhibit minor changes in temperature, pressure, or vibration before a failure. These changes are usually masked by noise or below the traditional monitoring threshold and are called "weak abnormal signals". Enhancing their significance is to highlight the features of these signals through feature engineering means; for example: differential analysis, higher-order statistics, frequency domain transformation;

[0136] Abnormal pattern recognition module: Based on the key feature set, it constructs an abnormal recognition model using a long short-term memory network, constructs a training set through historical operation data, and obtains a trained abnormal recognition model; outputs the abnormal probability (a value between 0 and 1) corresponding to the key component, and sets an early warning threshold for the abnormal probability to trigger an early warning;

[0137] Fault warning and optimization module: Based on the Weibull reliability model, predict the remaining service life of key components, generate a fault maintenance time window in combination with the confidence interval, and dynamically adjust the priority of maintaining key components and perform preventive maintenance based on the prediction results of the fault maintenance time window and the abnormal probability; Use the operating data of key components after maintenance to update the anomaly recognition model and the Weibull reliability model to form a closed-loop optimization control.

[0138] Explanation: By comprehensively considering the abnormal probability and the start time of the fault maintenance time window, determine the priority of each component; The specific method is to combine the high or low abnormal probability with the urgency of the maintenance time window, so that components with high risk and tight time obtain higher priority. This comprehensive method can flexibly adjust the relative importance of the two according to actual needs, but there is no need to specify a specific calculation method; Sort all key components according to the calculated priority, and give priority to preventive maintenance for components with higher priority; As the system operation data is updated (such as the abnormal probability increases or the time window shortens), the priority will be dynamically adjusted to ensure that the maintenance decision always adapts to the current state.

[0139] In the present invention, it is further necessary to explain that the method for obtaining the fault maintenance time window is as follows:

[0140] Data acquisition: Obtain the average abnormal probability (such as 0.85) and historical fault data (such as component average life, fault distribution) of key components;

[0141] Model initialization: Initialize the parameters of the Weibull reliability model. Since it is suitable for describing the failure law of mechanical components, the parameters include the shape parameter β, which characterizes the failure mode; and the scale parameter η, which characterizes the life; Fit the Weibul model according to historical data and adjust the model parameters in combination with the current abnormal probability; Optimize the parameters through maximum likelihood estimation or least squares method to ensure that the Weibull reliability model reflects the current state;

[0142] For example, if the abnormal probability is high, shorten the value of η, indicating that the life decay accelerates; If the abnormal probability is low (such as <0.3), keep or extend η, indicating that the component state is relatively stable;

[0143] Life prediction: Based on the adjusted Weibull reliability model and abnormal probability, calculate the remaining service life of the component through the following formula and generate a fault maintenance time window;

[0144]

[0145] Among them, t is time, F(t) is the abnormal probability, t is the remaining service life, and t is inversely deduced according to the abnormal probability (such as 0.85) to obtain the remaining service life;

[0146] Set the confidence interval of the remaining service life, calculate the upper and lower limits of the remaining service life through the inverse function of the Weibull reliability model, and generate a fault maintenance time window; for example, if the central value of the remaining service life is 12 hours and the confidence interval is ±2 hours, then the generated fault maintenance time window is 10 - 14 hours;

[0147] Verify that the remaining service life is consistent with the abnormal probability trend (such as the higher the probability, the shorter the remaining service life), and if abnormal, give an alarm in time.

[0148] In the present invention, it needs to be further explained that the training process of the abnormal recognition model includes the following steps:

[0149] Step S1, dataset preparation and division: Prepare the historical dataset of key components, extract the key feature set, mark the fault time points, and divide it into a training set (70%), a validation set (20%), and a test set (10%);

[0150] Step S2, time series data preprocessing and input construction: Perform time series processing on the historical dataset to generate an input format suitable for the long short-term memory network; based on the key feature set, use the sliding window method to construct time series inputs. For example, with a 30-minute window and a 5-minute step size, generate sequence data containing historical trends; perform normalization processing on the sequence data (such as Min-Max normalization) to map the feature values to the [0, 1] interval; output the standardized time series input sequence to capture the cumulative effect of weak signals;

[0151] Step S3, model initialization and training: Set the long short-term memory network structure, including an input layer, two hidden layers (64 neurons in each layer), and an output layer, and output the abnormal probability (a value between 0 and 1); use the mean square error as the loss function, which is used to measure the difference between the abnormal probability predicted by the model and the true label; train the short-term memory network through the Adam optimizer and the loss function, set the number of iterations to 100 times, and use the abnormal event labels in the historical operation data for supervised learning to obtain the trained abnormal recognition model;

[0152] Step S4, adjust the hyperparameters of the abnormal recognition model through the validation set to reduce the loss function value; use the cross-validation method to evaluate the generalization ability of the short-term memory network to ensure the accuracy of abnormal recognition; determine the warning threshold of the abnormal probability according to the validation results, for example, set it to 0.8, and trigger a warning when the abnormal probability exceeds this value.

[0153] In the present invention, it needs to be further explained that the device further includes: a warning threshold dynamic management module, which solves the problem of the fixed threshold failure caused by the change of working conditions in the liquid cooling system through warning threshold dynamic management, including:

[0154] Based on the multi-dimensional operation data set, the multi-dimensional operation data set of key components (including temperature, pressure, vibration and flow) is collected in real time to generate a time series data set containing multiple parameters;

[0155] After normalizing each parameter, the fluctuation characteristics of each parameter (such as temperature, pressure, vibration, flow) are calculated. The time-frequency analysis method based on wavelet transform is used to decompose the signal, and the high-frequency fluctuation component and the low-frequency trend feature are extracted to generate the multi-dimensional fluctuation feature vector v. For example, v = [v w , w p , v z , v f ; When the fluctuation feature extraction is completed, a dynamic threshold Y DT Formula:

[0156] Y DT = α·EWMA(v) + β·MAD(v)×(1 + γ·CV(v))

[0157] Wherein, EWMA(v) represents the smoothed trend index, which is the weighted average of the fluctuation trends of each parameter; MAD(v) measures the dispersion; CV(v) represents the adjusted fluctuation intensity, which is the ratio of the standard deviation of each parameter to the mean of each parameter. α, β, and γ are adjustment coefficients;

[0158] For the sake of understanding, an example is given to illustrate that the process of obtaining MAD(v) includes:

[0159] After normalizing each parameter, v = [0.1, 0.2, 0.15, 0.3] is obtained:

[0160] Median median(v) = 0.175;

[0161] Deviation |v i - 0.175| = [0.075, 0.025, 0.025, 0.125], where i is the sequence number;

[0162] MAD(v) = median(0.075, 0.025, 0.025, 0.125) = 0.05, indicating a stable distribution;

[0163] When the real-time anomaly probability exceeds the dynamic threshold Y DT , a warning signal is triggered.

[0164] In the present invention, it needs to be further explained that the warning threshold dynamic management module further includes:

[0165] Based on the Markov chain, predict the short-term fluctuation trend of the multi-dimensional operation data set, and calculate the state transition probability of each parameter within the future time (such as 1 hour);

[0166] Adjust the dynamic threshold formula in combination with the prediction results, introduce a trend factor δ (range 0.8 - 1.2), when the fluctuation trend rises, δ < 1 to lower the warning threshold, and when it falls, δ > 1 to raise the warning threshold.

[0167] In the present invention, it needs to be further explained that the present invention does not specifically limit the threshold update method: update the sliding window data every M hours (M is a predetermined time interval), and recalculate the dynamic warning threshold to adapt to system changes;

[0168] Or, set an adaptive update frequency (for example, update the threshold every 10 minutes when the abnormal probability change rate > 0.1).

[0169] In the present invention, it needs to be further explained that the feature engineering and weak signal extraction module includes:

[0170] Decompose the multi-dimensional operation data set by wavelet transform to separate high-frequency noise and low-frequency trend components;

[0171] Extract the intrinsic mode functions of weak abnormal signals based on empirical mode decomposition (the intrinsic mode functions are used to adaptively decompose signals, separate features of different scales, and thus extract weak abnormalities hidden in the signals). Calculate the instantaneous frequency and instantaneous amplitude of each intrinsic mode function through Hilbert transform to generate a multi-dimensional feature vector reflecting the abnormal trend; Instantaneous frequency: represents the local frequency of the signal at a certain moment and can reveal the change of the signal frequency over time. For example, if an abnormality occurs in the system, the instantaneous frequency of a certain IMF may suddenly increase or decrease;

[0172] Explanation: In the embodiments of the present invention, the wavelet transform uses Daubechies wavelet to decompose the signal to level 3 to separate high-frequency noise and low-frequency trend; empirical mode decomposition (EMD) extracts 3 intrinsic mode functions (IMFs), calculates the instantaneous frequency and instantaneous amplitude through Hilbert transform, and principal component analysis retains the features with the first 90% variance;

[0173] Instantaneous amplitude: reflects the intensity or energy of the signal at a certain moment. If an abnormality causes a sudden change in the signal amplitude (such as sudden amplification or weakening), the instantaneous amplitude can capture this change.

[0174] Reduce the dimension of the multi-dimensional feature vector through principal component analysis to form a key feature set. This enhances the accuracy and robustness of anomaly detection.

[0175] In the present invention, it needs to be further explained that the training process of the anomaly recognition model further includes:

[0176] Introduce an attention mechanism into the long short-term memory network, assign weights to each time step based on the characteristics of weak abnormal signals, and highlight the impact of weak abnormal signals on the abnormal probability;

[0177] Explanation: The attention mechanism calculates a weight value (usually between 0 and 1) for each time step in the sequence, indicating the relative importance of that time step to the final task (here, abnormal probability prediction);

[0178] Set up a bidirectional LSTM structure to capture both historical and future trends simultaneously, and dynamically adjust the number of hidden layer neurons;

[0179] Adopt Dropout regularization (dropout rate is 0.2) to prevent overfitting, combined with an early stopping strategy (stop training when the validation set loss has not decreased for 10 consecutive times).

[0180] In a possible embodiment, according to the abnormal probability output by the abnormal pattern recognition module, adjust the sampling frequency of the sensor; for example, when the abnormal probability exceeds 0.7, increase the sampling frequency to 10 times per second, otherwise keep it at 1 time per minute, and generate original multi-dimensional operation data with an adaptive frequency.

[0181] Summary: In the embodiments of the present invention, the multi-source data acquisition and preprocessing module is deployed to obtain the multi-dimensional operation data of the key components in the liquid-cooled energy storage system in real time. The wavelet transform and empirical mode decomposition techniques are used by the feature engineering and weak signal extraction module to extract weak abnormal features and generate a key feature set; the long short-term memory network and attention mechanism of the abnormal pattern recognition module are used to accurately output the abnormal probability, and the early warning threshold dynamic management module is incorporated to predict the fluctuation trend based on the Markov chain and dynamically adjust the threshold; finally, the Weibull reliability model of the fault warning and optimization module is used to predict the remaining life and optimize the maintenance time window, realizing closed-loop optimization control; effectively solving the problems of difficult detection of weak abnormal signals, failure of fixed thresholds, and inaccurate maintenance timing caused by the complex and changeable working conditions of the liquid-cooled energy storage system;

[0182] In addition, through multi-level signal processing, intelligent prediction, and dynamic optimization, the robustness of system monitoring and the timeliness of fault response are improved in this embodiment, significantly reducing the fault risk and maintenance cost, and providing an efficient solution for the stable operation of industrial and commercial liquid-cooled energy storage systems.

[0183] Embodiment 2. The difference between the embodiment of the present invention and Embodiment 1 is that the device further includes a fault maintenance time window optimization module for obtaining an optimized fault maintenance time window. Refer to Figure 2 The flowchart of the fault maintenance time window optimization method, including:

[0184] Obtain the abnormal probability of key components in the liquid cooling system through real-time monitoring, and generate a comprehensive input data set reflecting the component status by combining historical operation data and the current multi-dimensional operation data set;

[0185] Utilize the multi-physics coupling model, and based on the comprehensive input data set and the abnormal probability, simulate the thermal, mechanical, and fluid mechanical behaviors of the key components under the current abnormal probability;

[0186] In a possible embodiment, introduce electrochemical behavior simulation to analyze the temperature-pressure-charge distribution coupling effect of battery cells, and evaluate the influence of the non-uniform distribution of coolant flow rate on the thermal runaway risk by combining the liquid flow dynamic effect;

[0187] Generate a stress-strain distribution diagram through finite element analysis, predict the fatigue crack propagation time in combination with the abnormal probability, and adopt dynamic grid adaptive adjustment to simulate the accuracy, generating a virtual life curve of the component state changing with time. The virtual life curve is used to characterize the dynamic deterioration process and potential fault characteristics of the key component;

[0188] Combine the virtual life curve with the Weibull reliability model, fuse the real-time monitoring data and historical operation data through the Bayesian update method, correct the Weibull reliability model parameters (shape parameter β and scale parameter η), and control the fitting error of the virtual life curve within 5%, and output the optimized Weibull reliability model;

[0189] Generate an optimized fault maintenance time window based on the optimized Weibull reliability model, providing an optimized time basis for maintenance decision-making.

[0190] Explanation: The comprehensive input data set serves as the input of the multi-physics model, defining the starting point and boundary of the simulation; for example, if the abnormal probability is high, the model may assume a higher initial temperature or pressure; use the multi-physics model to simulate the dynamic physical behavior of the component. The simulation is not static but dynamic, reflecting the component's possible deterioration process from the current state (abnormal probability 0.85) to the future; extract deterioration indicators from the simulation results, generate time series data, plot the virtual life curve, and characterize the deterioration and fault characteristics; provide an optimized time basis for maintenance decision-making by generating an optimized fault maintenance time window.

[0191] In the embodiment of the present invention, it needs to be further explained that the device further includes a fault maintenance trigger condition management module for managing the trigger conditions of fault maintenance. The trigger conditions are the basis for determining to start the fault maintenance time window, usually a threshold or rule. For example, when the abnormal probability exceeds 0.8, trigger the calculation or execution of the maintenance window, including:

[0192] Explanation: The fault maintenance time window is not always active. Instead, a condition is required to determine when to calculate or apply this window. Dynamically adjusting the triggering condition directly affects the generation timing of the window;

[0193] Obtain the dynamic simulation data of battery cells and key components generated based on the multi-physics coupling model;

[0194] Analyze the fault propagation characteristics in the dynamic simulation data, construct an adaptive fault evolution network. The adaptive fault evolution network identifies the fault propagation paths and propagation intensities between components through a graph neural network, and generates an evolution topology structure reflecting the multi-point fault synergistic effect;

[0195] According to the evolution topology structure, fuse real-time sensor data and historical fault patterns, and use a contrast learning mechanism to judge the matching degree between the current operating state and potential fault patterns, and generate a time series probability distribution of fault evolution;

[0196] If the time series probability distribution indicates that the potential fault pattern exceeds the preset threshold, then based on the collaborative optimization strategy of Monte Carlo simulation and reinforcement learning, dynamically adjust the triggering condition of the fault maintenance time window, and output an optimized maintenance decision for unknown fault patterns.

[0197] In a possible embodiment, the device further includes a maintenance strategy fusion decision module, which dynamically generates a cross-component collaborative maintenance plan through an adaptive hybrid learning algorithm based on the triggering condition of the fault maintenance time window, the virtual life curve generated by the multi-physics coupling model, and the fault propagation topology between components; the maintenance strategy fusion decision module performs the following operations:

[0198] (1) Input the remaining life prediction value of the virtual life curve, the real-time anomaly probability, and the edge weights of the fault propagation topology into a graph convolutional neural network to generate the maintenance urgency scores of each component;

[0199] (2) Based on the current operating load of the liquid cooling system, the availability of external maintenance resources, and the preset shutdown tolerance threshold, construct a multi-objective optimization constraint condition including an economic penalty function and a risk diffusion suppression function;

[0200] (3) When detecting a sudden change in the component degradation rate or an increase in the cross-region fault correlation, fuse the latest failure mode features through an online incremental learning mechanism, reconstruct the decision boundary of the maintenance strategy, and update the weight parameters.

[0201] Explanation: Generating the maintenance urgency scores of each component through a graph convolutional neural network, the specific operations are as follows:

[0202] Step 1. Input data definition:

[0203] Node characteristics: Each component is a node in the graph structure, and its properties include:

[0204] Remaining life: The remaining effective working time of the component predicted by the virtual life curve (e.g. the remaining life of the pump body is 1200 hours);

[0205] Real-time anomaly probability: the current fault risk value output by the anomaly recognition model (e.g. 0.85 indicates high risk);

[0206] Edge weight: a connection relationship that reflects the strength of fault propagation between components (e.g., the probability of propagation from the pump to the valve is 0.6), provided by the fault propagation topology of claim 9;

[0207] Step 2: Graph convolution operation:

[0208] Neighborhood information aggregation: For each node, the weighted average of the features of all its neighboring nodes is calculated, and the weight is obtained by normalizing the edge weight. For example, when the pump node aggregates the features of the valve node, the contribution of the valve feature value is determined by the ratio of the pump→valve edge weight (0.6) to the sum of the weights of all adjacent edges.

[0209] Feature transformation and activation: The aggregated neighborhood features are superimposed on the node’s own features, linearly transformed through a trainable parameter matrix, and a nonlinear activation function (such as ReLU) is applied to generate updated node features;

[0210] Step 3: Output layer processing: After multi-layer graph convolution operations, the node features are finally mapped to the range of 0 to 1 through the Softmax function to generate the maintenance urgency score of each component.

[0211] Explanation: The economic penalty function includes maintenance labor costs, spare parts loss costs and downtime loss costs, and the weights are dynamically set according to the operating scenarios (for example, in industrial and commercial scenarios, the labor cost weight accounts for 60%, spare parts loss accounts for 30%, and downtime loss accounts for 10%); the risk diffusion suppression function includes the correlation evaluation of component abnormality probability and fault propagation intensity, which is calculated as follows: the real-time abnormality probability of each component is multiplied by the edge weight propagated to the key node, and the accumulated value is multiplied by the preset risk suppression coefficient to generate a system-level risk score; the economic penalty function prioritizes the controllability of maintenance costs, and the risk diffusion suppression function ensures that high-risk components are handled first. The two are dynamically balanced through preset weight coefficients, combined with real-time resource constraints and fault propagation path optimization decision boundaries to achieve global reliability optimization and economic benefit maximization of the liquid-cooled energy storage system maintenance plan.

[0212] When a sudden change in component degradation rate or an increase in cross-regional fault correlation is detected, the latest failure mode features are integrated through an online incremental learning mechanism to reconstruct the decision boundary of the maintenance strategy and update the weight parameters.

[0213] Explanation: The adaptive hybrid learning algorithm includes deep reinforcement learning and transfer learning, specifically including:

[0214] Deep reinforcement learning branch: Based on the virtual life curve data, calculate the long-term value return of maintenance actions (such as maintenance cost reduction rate, risk suppression gain) through the deep reinforcement learning algorithm;

[0215] Transfer learning branch: Utilize similar fault patterns in historical maintenance records (such as the pump-valve linkage failure case library), and accelerate the policy convergence in new scenarios through pre-trained model parameter transfer;

[0216] Dynamic generation logic: Input data: Component remaining life (virtual life curve), anomaly probability (LSTM output), fault propagation topology edge weight;

[0217] Output scheme: Generate a maintenance sequence (such as "replace the pump body first → delay the repair of the cooling pipeline"), and quantify the time window (such as 12 ± 2 hours) and execution granularity (such as local maintenance or system shutdown).

[0218] Summary: In the embodiments of the present invention, by adding a fault maintenance time window optimization module and a fault maintenance trigger condition management module on the basis of Embodiment 1, obtain the anomaly probability through real-time monitoring and combine historical and current multi-dimensional operation data to generate a comprehensive input data set; simulate the thermal, mechanical, and fluid mechanical behaviors of key components through a multi-physics field coupling model, introduce electrochemical and liquid flow dynamic effects to analyze the risk of battery thermal runaway, and generate a virtual life curve using finite element analysis; optimize the maintenance time window based on the Weibull reliability model and Bayesian update. At the same time, construct an adaptive fault evolution network through a graph neural network, fuse contrast learning, Monte Carlo simulation, and reinforcement learning, dynamically adjust the trigger conditions, and output maintenance decisions for unknown fault patterns. Achieve accurate fault prediction and dynamic response for the complex working conditions of the liquid-cooled energy storage system, solve the problems that traditional methods are difficult to handle unknown fault patterns and the maintenance timing is not optimal, and improve the reliability and maintenance efficiency of the system.

[0219] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system, characterized in that, Including: Multi-source data acquisition and preprocessing module: Real-time collect the temperature, pressure, vibration and flow data of key components in the liquid cooling system, and denoise and standardize the collected data to generate a multi-dimensional operation data set; Feature engineering and weak signal extraction module: Adopt wavelet transform and empirical mode decomposition technology to separate high-frequency noise and low-frequency trend components from the multi-dimensional operation data set, extract the instantaneous frequency and instantaneous amplitude features reflecting the state of key components, and generate a key feature set through principal component analysis for dimensionality reduction; Abnormal pattern recognition module: Build an abnormal recognition model based on the long short-term memory network, input the key feature set and output the abnormal probability of each key component, and set an early warning threshold for the abnormal probability to trigger an early warning; The long short-term memory network introduces an attention mechanism to assign weights to time steps and combines a bidirectional LSTM structure to capture historical and future trends; Fault early warning and optimization module: Predict the remaining service life of key components based on the Weibull reliability model, generate a fault maintenance time window in combination with the confidence interval, and dynamically adjust the maintenance priority according to the abnormal probability.

2. The real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system according to claim 1, wherein, The acquisition method of the fault maintenance time window is as follows: Obtain the average abnormal probability and historical fault data of key components; Initialize the parameters of the Weibull reliability model, the parameters include the shape parameter β; and the scale parameter η; Fit the Weibul model according to historical data, and adjust the model parameters in combination with the current abnormal probability; Based on the adjusted Weibull reliability model and abnormal probability, calculate the remaining service life t of the component through the following formula and generate a fault maintenance time window; Among them, t is time, F(t) is the abnormal probability, t is the remaining service life, and t is inversely deduced according to the abnormal probability to obtain the remaining service life; Set the confidence interval of the remaining service life, calculate the upper and lower limits of the remaining service life through the inverse function of the Weibull reliability model, and generate a fault maintenance time window; Verify that the remaining service life is consistent with the abnormal probability trend.

3. The real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system according to claim 1, characterized in that, The training process of the abnormal recognition model includes the following steps: Step S1, Prepare the historical data set of key components, extract the key feature set, mark the fault time points, and divide them into a training set, a validation set and a test set; Step S2, Perform time series processing on the historical data set to generate an input format suitable for the long short-term memory network; Based on the key feature set, use the sliding window method to construct a time series input to generate sequence data containing historical trends; Normalize the sequence data and map the feature values to the [0, 1] interval; Output the standardized time series input sequence to capture the cumulative effect of weak signals; Step S3, Set the long short-term memory network structure, including an input layer, a hidden layer and an output layer, and output the abnormal probability; Use the mean square error as the loss function; Train the short-term memory network through the Adam optimizer and the loss function to obtain a trained abnormal recognition model.

4. The real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system according to claim 1, characterized in that, It also includes: Early warning threshold dynamic management module, including: Based on the multi-dimensional operation data set, collect the multi-dimensional operation data set of key components in real time to generate a time series data set containing multiple parameters; After normalizing each parameter, calculate the fluctuation characteristics of each parameter. Use the time-frequency analysis method based on wavelet transform to decompose the signal, extract the high-frequency fluctuation component and the low-frequency trend feature, generate the multi-dimensional fluctuation feature vector v, and construct the dynamic threshold Y DT Formula: Y DT = α·EWMA(v) + β·MAD(v)×(1 + γ·CV(v)) Among them, EWMA(v) represents the smoothed trend index, which is the weighted average of the fluctuation trends of each parameter; MAD(v) represents the dispersion; CV(v) adjusts the fluctuation intensity, which is the ratio of the standard deviation of each parameter to the mean of each parameter, and α, β, and γ are adjustment coefficients.

5. The real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system according to claim 3, characterized in that, The training process of the anomaly recognition model further includes: Introducing an attention mechanism into the long short-term memory network, and based on the characteristics of weak anomaly signals, assigning weights to each time step to highlight the influence of weak anomaly signals on the anomaly probability; Setting a bidirectional LSTM structure to capture both historical and future trends simultaneously and dynamically adjust the number of neurons in the hidden layer; Using Dropout regularization to prevent overfitting and combining with an early stopping strategy.

6. The real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system according to claim 1, characterized in that, Adjusting the sampling frequency of the sensor according to the anomaly probability output by the anomaly pattern recognition module.

7. The real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system according to claim 4, wherein The early warning threshold dynamic management module further includes: Predicting the short-term fluctuation trend of the multi-dimensional operation dataset based on the Markov chain and calculating the state transition probability of each parameter in the future time; Combining the prediction results to adjust the dynamic threshold formula, introducing a trend factor δ, when the fluctuation trend rises, δ < 1 to lower the early warning threshold, and when it drops, δ > 1 to raise the early warning threshold.

8. A real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system according to any one of claims 1-7, characterized in that, It also includes: A fault maintenance time window optimization module for obtaining an optimized fault maintenance time window, including: Obtaining the anomaly probability of key components in the liquid cooling system through real-time monitoring, and combining historical operation data and the current multi-dimensional operation dataset to generate a comprehensive input dataset reflecting the component state; Using a multi-physics field coupling model, based on the comprehensive input dataset and the anomaly probability, to simulate the thermal, mechanical, and fluid mechanical behaviors of key components under the current anomaly probability; Generating a stress-strain distribution diagram through finite element analysis, combining the anomaly probability to predict the fatigue crack propagation time, and generating a virtual life curve of the component state changing with time; Combining the virtual life curve with the Weibull reliability model, fusing real-time monitoring data and historical operation data through the Bayesian update method, correcting the parameters of the Weibull reliability model, and outputting an optimized Weibull reliability model; Generating an optimized fault maintenance time window based on the optimized Weibull reliability model.

9. The real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system according to claim 8, characterized in that, The device further includes a fault maintenance trigger condition management module for managing the trigger conditions for fault maintenance. The trigger conditions are the basis for determining whether to start the fault maintenance time window, including: Obtaining the dynamic simulation data of the battery unit and key components generated based on the multi-physics field coupling model; Analyzing the fault propagation characteristics in the dynamic simulation data and constructing an adaptive fault evolution network. The adaptive fault evolution network identifies the fault propagation paths and propagation intensities between components through a graph neural network and generates an evolution topology structure reflecting the multi-point fault synergy effect; According to the evolution topology structure, fusing real-time sensor data and historical fault patterns, and using a contrast learning mechanism to judge the matching degree between the current operating state and potential fault patterns, and generating a time series probability distribution of fault evolution; If the time series probability distribution indicates that the potential fault pattern exceeds the preset threshold, then based on the collaborative optimization strategy of Monte Carlo simulation and reinforcement learning, dynamically adjust the trigger conditions for the fault maintenance time window.

10. The real-time monitoring and fault response control device for an industrial and commercial liquid-cooled energy storage system according to claim 9, characterized in that, It further includes a maintenance strategy fusion decision-making module, which dynamically generates a cross-component collaborative maintenance plan through an adaptive hybrid learning algorithm based on the trigger conditions of the fault maintenance time window, the virtual life curve, and the fault propagation topology structure among components, including: Inputting the remaining life prediction value, real-time anomaly probability of the virtual life curve, and the edge weights of the fault propagation topology into a graph convolutional neural network to evaluate the maintenance urgency of each component; Based on the current operating load of the liquid cooling system, the availability of external maintenance resources, and a preset shutdown tolerance threshold, constructing multi-objective optimization constraint conditions including an economic penalty function and a risk diffusion suppression function; When a sudden change in the component degradation rate or an enhanced cross-region fault correlation is detected, fuse the latest failure mode features through an online incremental learning mechanism, reconstruct the decision boundary of the maintenance strategy, and update the weight parameters.

Citation Information

Patent Citations

  • Online monitoring device and system of 3D printing equipment

    CN110370649A

  • Complex equipment key component fault diagnosis system based on LSTM coding network

    CN115307944A

  • Multi-objective optimization maintenance decision-making method for turbine blade of gas turbine

    CN119249850A

  • Parameter control method of ultrasonic motor

    CN119341433A

  • Intelligent power grid fault prediction and operation and maintenance method and system based on cloud computing, and medium

    CN119475131A

Cited By

  • Energy management and safety protection cooperation method for liquid cooling industrial and commercial energy storage system

    CN120354715A

  • Energy management and safety protection collaborative method for liquid-cooled industrial and commercial energy storage system

    CN120354715B

  • Centrifugal pump performance intelligent diagnosis method and system based on pressure characteristics

    CN120745513A

  • Expressway frame hydraulic monitoring control system based on internet of things

    CN120949591B

  • Intelligent prediction and data tracing management method for lubricating system

    CN121073247A