Algorithm fusion and fault early warning system in a battery safety management platform
The battery safety management platform, which uses multi-source data acquisition and time-series correlation modeling, solves the problems of delayed response and insufficient risk identification in existing technologies. It enables efficient fault early warning and refined control of battery systems, thereby improving the safety and reliability of battery systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DONGGUAN LITHIUM VALLEY ENERGY CO LTD
- Filing Date
- 2025-07-12
- Publication Date
- 2026-04-17
AI Technical Summary
Existing battery safety management systems lack the ability to dynamically fuse and process multi-source information in complex environments, resulting in delayed response times and insufficient accuracy in risk identification. In particular, they struggle to identify potential risks in a timely manner when there are no obvious faults, and the risk assessment results fail to form an efficient closed loop with control execution.
A multi-source data acquisition module is used to acquire multi-modal operating data such as voltage, current, temperature, humidity, acoustic signals, electrochemical impedance spectroscopy data, and local gas concentration. A reasoning chain is established based on time-series correlation modeling and causal feature recognition through a fusion computing engine to generate risk response signals. A fault early warning response module generates hierarchical control commands, and an execution interface module drives the physical controller to perform corresponding operations.
It significantly improves the ability to perceive non-obvious abnormal states inside the battery, can predict potential faults in advance and implement differentiated control, forming a complete closed loop from risk identification to execution, and improving the safety and operational reliability of the battery system.
Smart Images

Figure CN120949051B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of battery safety technology, and in particular to an algorithm fusion and fault early warning system in a battery safety management platform. Background Technology
[0002] In existing technologies, battery safety management systems typically rely on rule-based or fixed-threshold battery management systems (BMS) for monitoring and control. They primarily monitor single or limited parameters such as voltage, current, and temperature, and, combined with limited preset conditions, trigger alarms and protection actions for typical faults such as overvoltage, overtemperature, and overcurrent. Some systems incorporate data analysis models, but these are mostly limited to static data evaluation and lack the ability to dynamically fuse and process multi-source information in complex environments.
[0003] However, existing technologies generally suffer from problems such as delayed response time, insufficient accuracy in risk identification, and weak proactive intervention capabilities in actual battery operation scenarios. In particular, when facing inconspicuous faults such as thermal runaway, micro-leakage, or early internal resistance anomalies, traditional threshold mechanisms and single-layer judgment methods are unable to identify potential risks in a timely manner. At the same time, the risk judgment results in existing systems often fail to form an efficient closed loop with control execution, resulting in fault warnings being triggered, but the response behavior lacking specificity or timeliness.
[0004] In view of the above problems, there is an urgent need to propose a battery safety management platform that integrates multi-source data and has the ability to close the inference chain, so as to improve the depth of risk identification and the efficiency of control response. Summary of the Invention
[0005] This application provides an algorithm fusion and fault early warning system in a battery safety management platform to improve the depth of risk identification and the efficiency of control response.
[0006] This application provides an algorithm fusion and fault early warning system in a battery safety management platform, including:
[0007] The multi-source data acquisition module is used to acquire multi-modal operating data, including voltage, current, temperature, humidity, acoustic signals, electrochemical impedance spectroscopy data, and local gas concentration, during the operation of the battery system.
[0008] The fusion computing engine is used to receive the multimodal operation data and generate risk response signals to characterize battery safety risks based on a hierarchical algorithm fusion structure that establishes an inference chain between time-series correlation modeling, causal feature identification and risk trend inference.
[0009] The fault warning and response module is used to receive the risk response signal and generate hierarchical control instructions corresponding to different risk levels according to the preset response strategy. The hierarchical control instructions include one or more of the following: activating virtual circuit breaker logic, scheduling regional cooling resources, refreshing BMS operation mode, performing SOC valuation correction, or pushing remote risk alarm information.
[0010] The execution interface module is used to receive the hierarchical control commands and map the hierarchical control commands to the control interface connected to the physical controller or management unit of the battery system, thereby driving the actual execution of the corresponding control operations to realize online risk intervention and response scheduling of the battery system.
[0011] The beneficial effects of this application mainly include: (1) By simultaneously collecting multi-source information such as voltage, current, temperature, humidity, sound waves, electrochemical impedance spectroscopy and gas concentration, the perception and identification accuracy of non-obvious abnormal states inside the battery are significantly improved. (2) By constructing a reasoning chain based on time-series correlation modeling, causal feature identification and risk trend inference, a hierarchical risk analysis path is formed, which can effectively predict the evolution trend of potential thermal runaway, gas leakage or performance degradation, and intervene in possible complex faults in advance. (3) It has the ability to generate graded response and control commands, and can dynamically match diverse control strategies such as virtual fuse, cooling scheduling, BMS refresh, SOC correction or remote alarm according to different risk levels, so as to realize differentiated and refined safety prevention and control management. (4) By realizing linkage control with the physical controller through the execution interface module, a complete closed loop from risk identification, decision generation to intervention execution is formed, ensuring that the battery system can automatically trigger response measures in abnormal states, and significantly improving the safety and operational reliability of the battery system. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of an algorithm fusion and fault early warning system in a battery safety management platform provided in the first embodiment of this application. Detailed Implementation
[0013] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0014] The first embodiment of this application provides an algorithm fusion and fault early warning system in a battery safety management platform. Please refer to... Figure 1 This figure is a schematic diagram of the first embodiment of this application. The following is in conjunction with... Figure 1 The first embodiment of this application provides an algorithm fusion and fault early warning system in a battery safety management platform.
[0015] The algorithm fusion and fault early warning system in the battery safety management platform includes a multi-source data acquisition module 101, a fusion computing engine 102, a fault early warning response module 103, and an execution interface module 104.
[0016] The multi-source data acquisition module 101 is used to acquire multi-modal operating data, including voltage, current, temperature, humidity, acoustic signals, electrochemical impedance spectroscopy data, and local gas concentration, during the operation of the battery system.
[0017] In the battery safety management platform described in this invention, the multi-source data acquisition module 101 is used to acquire comprehensive and real-time data on the operating status of the battery system. This module forms the basis of the data-driven decision-making mechanism in the system. Its design must support the parallel access, synchronous control and high-frequency sampling of multiple types of sensors to ensure the timeliness and accuracy of subsequent risk identification.
[0018] Specifically, the multi-source data acquisition module 101 includes several sensing units for acquiring key battery operating parameters. These parameters include, but are not limited to, voltage, current, temperature, humidity, acoustic signals, electrochemical impedance spectroscopy data, and local gas concentration. Voltage and current sensors typically employ Hall effect devices or shunt solutions, enabling millisecond-level sampling. Temperature and humidity signals can be acquired using thermocouples, NTC thermistors, and capacitive humidity sensors, respectively. Acoustic signals are used to identify internal structural anomalies within the battery cell and can be achieved using ultrasonic transducers or MEMS microphones deployed at key locations on the battery pack casing. The acquisition of electrochemical impedance spectroscopy data is based on AC signal excitation and current response measurement, requiring the integration of a high-precision AC impedance measurement circuit to extract information such as battery polarization impedance and diffusion impedance. Local gas concentration data relies on electrochemical gas sensors, infrared gas sensors, or semiconductor gas-sensitive devices to detect potentially leaked components such as CO, CO2, and H2 from the battery cell.
[0019] Electrochemical impedance spectroscopy (EIS) data is a type of complex impedance information used to describe the changes in the electrochemical characteristics of a battery at different AC frequencies. It is typically obtained by measuring the battery's voltage-current response to a series of AC excitation signals. This data reflects the dynamic behavior of processes such as charge transfer, electrolyte diffusion, and electrode interface reactions within the battery, and is widely used to analyze battery health, aging, and potential failure modes. Its typical representations are Nyquist plots (real part versus imaginary part) and Bode plots (amplitude versus phase versus frequency), providing key parameters such as polarization impedance, diffusion impedance, and electrolyte resistance.
[0020] To obtain electrochemical impedance spectroscopy (EIS) data, a sinusoidal AC excitation signal with controlled amplitude needs to be applied to the battery, and the signal is scanned point by point within a preset frequency range (e.g., 10 mHz to 10 kHz). Specifically, either constant voltage excitation (Potentiostatic EIS) or constant current excitation (Galvanostatic EIS) can be used, capturing the corresponding response signals through voltage or current measurement loops. The measurement system should be equipped with a sinusoidal signal source with a phase-locked reference, a high-precision analog-to-digital converter (ADC), and a synchronous sampling device to extract the amplitude and phase response at each frequency point, and combine this with Fourier transform or phase-sensitive demodulation algorithms to obtain the complex impedance. Finally, the complex impedances at multiple frequency points are combined to form an electrochemical impedance spectroscopy curve for subsequent modeling analysis and fault identification.
[0021] Regarding data synchronization, the multi-source data acquisition module 101 is internally configured with a unified sampling clock and data buffering mechanism, enabling various asynchronous sensor data to be aligned under a unified time base and pushed sequentially to the fusion computing engine 102. Furthermore, to improve system reliability, this module preferably employs a redundant channel design and is configured with a preliminary data integrity verification mechanism to eliminate erroneous data input caused by sensor anomalies. All acquired data will be encapsulated into standardized multimodal operation data packets, each containing fields such as timestamp, sensor type identifier, sampled value, and data validity flag.
[0022] The multi-source data acquisition module 101 should also have environmental adaptability and electromagnetic compatibility design to ensure long-term stable operation in vehicle-mounted, energy storage station, or high-power industrial applications. Simultaneously, this module communicates with the fusion computing engine 102 via a high-speed data bus or communication interface (such as CAN, SPI, EtherCAT, RS485, or Ethernet) to ensure real-time data transmission capabilities under high-frequency sampling, meeting the overall dynamic risk assessment requirements of the platform.
[0023] Through the above structural and functional design, the multi-source data acquisition module 101 can provide the system with comprehensive, timely, and controllable operational data input, thereby supporting the subsequent risk perception and control decision-making process based on the inference chain.
[0024] Furthermore, the multi-source data acquisition module is also used for:
[0025] The synchronous time series corresponding to the acoustic signal channel and temperature signal channel collected during the operation of the battery system are obtained, and the two are timestamped to form a unified acoustic-temperature coupling input pair.
[0026] A short-time Fourier transform is performed on the acoustic signal channel to extract the spectral energy density distribution characteristics, and local spectral bands with abrupt increases in amplitude within the 1kHz to 5kHz frequency band are identified as candidate acoustic anomaly intervals.
[0027] Within the acoustic anomaly candidate interval, the temperature signal of the corresponding time period is subjected to first derivative analysis to screen out the segments in which the heating rate is significantly higher than the conventional change threshold within the time delay window, and construct the acoustic-temperature linkage anomaly segment.
[0028] Each anomalous sound-temperature linkage segment is represented as a three-dimensional feature tensor, with dimensions including spectral energy increase, temperature rise slope, and sound-temperature time difference. All segments are normalized and then input into the fusion computing engine.
[0029] The fusion computing engine determines whether the segment constitutes a precursor to thermal runaway based on the set linkage judgment rules or training model. If it is determined to constitute a precursor, the corresponding channel identifier, risk level label and thermoacoustic composite risk factor field are added to the risk response signal.
[0030] First, the multi-source data acquisition module needs to simultaneously acquire raw data from the acoustic signal channel and the temperature signal channel during battery system operation. The acoustic signal can be acquired through a high-sensitivity MEMS microphone or piezoelectric sensor, with a sampling frequency typically above 10kHz, used to capture weak features of high-frequency mechanical vibrations, gas release, or electrochemical noise. The temperature signal originates from thermocouples or RTD sensors attached to the cell casing or heat dissipation channels, with a sampling frequency set between 1Hz and 10Hz. Since the sampling frequencies of the two channels are inconsistent, the system needs to perform timestamp alignment processing on the two channels to achieve joint analysis. This processing can be achieved by linearly interpolating and resampling the temperature signal to form synchronized sample points under the acoustic signal time reference, ensuring the consistency of timing logic in subsequent analysis.
[0031] After time alignment, the system performs a Short-Time Fourier Transform (STFT) on the acoustic signal, converting the time-domain signal into a time-frequency domain matrix. The STFT window size can be set to 256~1024 points, with a sliding step size of 50% to balance time resolution and frequency resolution. Energy density is calculated on the resulting spectral matrix, extracting the energy distribution curve for each time slice within the target frequency band (1kHz to 5kHz). In practical applications, this frequency band covers the typical acoustic frequencies associated with most cell bulging, internal structural cracks, or minor gas ejection. If a sudden increase in spectral energy amplitude is detected in a time slice or several consecutive time slices within this frequency band (e.g., an instantaneous increase exceeding 300% of the normal average), the system marks it as a candidate region for acoustic anomalies.
[0032] For each candidate acoustic anomaly interval, the system will trace back its time axis, extract the temperature signal segment synchronized with it, and perform first-order derivative analysis on it, i.e., calculate the rate of temperature change over time. The criterion is whether the slope of the temperature change is significantly greater than a conventional threshold (e.g., more than twice the normal rate of rise) within the candidate interval or after a delay of τ seconds (where τ is an adjustable parameter between 1 and 10 seconds). If this condition is met, it can be identified as a "temperature rise accompanied by sonic boom" type of linked anomaly. This type is often a warning signal caused by the rapid release of heat accumulation inside the battery cell or the activation of side reactions, and is difficult to identify independently by a single physical quantity.
[0033] The system constructs a three-dimensional feature tensor for each segment identified as abnormal by the above judgment. The tensor has a fixed dimension of three dimensions, corresponding to: (1) the spectral energy amplification value, which represents the frequency domain intensity multiple of the current acoustic segment compared to the historical baseline; (2) the temperature rise slope, which is the maximum first derivative value of the temperature change within this time period; and (3) the sound-temperature time difference, which is the time delay between the occurrence of the sound wave mutation and the rapid rise in temperature, in seconds. The above three indicators are all continuous values, forming a linkage representation vector of cross-modal physical indicators. After forming the tensor, the system will normalize the indicators of all segments to eliminate the influence of the absolute value differences of different time periods and different sensors on the stability of the model.
[0034] The normalized feature tensor is fed as structured input into the fusion computing engine for identification and judgment. The fusion computing engine can choose to learn and identify using a threshold judgment model based on linkage rules (e.g., all three indicators exceeding a set threshold constitute a precursor) or a multimodal classifier built based on training samples (such as decision trees, support vector machines, or lightweight neural networks). The judgment output is a Boolean variable indicating whether the segment constitutes a precursor to thermal runaway.
[0035] If the system ultimately determines that a precursor to thermal runaway has occurred, the risk response signal generated by the fusion computing engine will be supplemented with the corresponding channel identifier (such as microphone number and temperature channel number), risk level label (such as "severe_warning"), and thermo-acoustic composite risk factor fields (such as linkage amplification index, signal correlation confidence level, etc.). This risk response signal will be transmitted as structured data to the fault warning response module, providing a more physically interpretable multimodal judgment basis for the system to finally generate hierarchical control commands.
[0036] Through the above processing flow, acoustic signals and temperature signals are fused in an engineering-feasible manner, and a complete logical chain is established from synchronous acquisition, frequency domain mutation identification, time-series linkage analysis, feature tensor construction to risk labeling output, which has high recognition rate, low false judgment rate and excellent scalability.
[0037] The fusion computing engine 102 is used to receive the multimodal operation data and generate a risk response signal to characterize battery safety risks based on a hierarchical algorithm fusion structure that establishes an inference chain between time-series correlation modeling, causal feature identification and risk trend inference.
[0038] The fusion computing engine 102 is used to perform structured processing and risk assessment on the multimodal operating data collected by the multi-source data acquisition module 101, and finally generate a risk response signal to drive the fault early warning and response module 103. To meet the needs of effective fusion, time series analysis and risk prediction of multidimensional data under complex working conditions, the fusion computing engine 102 adopts a three-layer algorithm fusion structure based on time series awareness, namely, a local feature extraction layer, a modal fusion and time series modeling layer, and a risk assessment and classification output layer.
[0039] In practical deployment, the multi-source data acquisition module 101 outputs data from multiple sensor channels, such as voltage, current, temperature, humidity, acoustic waves, electrochemical impedance spectroscopy, and gas concentration, with each channel representing a time series. The input to the fusion computing engine 102 is defined as a tensor of the form B, C, T, where B represents the batch size (e.g., 32), C represents the number of channels (e.g., 7 modal data types), and T represents the time step for sampling per channel (e.g., 120 steps, representing sampling once every 0.5 seconds within 60 seconds). This tensor can be constructed through a pre-data buffer module or dynamically generated from streaming acquisition using a sliding window approach.
[0040] In the local feature extraction layer, the data sequence for each channel is first transformed using a one-dimensional convolutional neural network (1D-CNN). The kernel size is typically set to 3 or 5, with a stride of 1 and identical padding, facilitating the extraction of dynamic patterns between adjacent time points. For example, for voltage signals, convolution can extract local trends such as abrupt changes and perturbation responses; for gas concentration signals, it can extract slowly accumulating upward trends. After each channel is convolved individually, the output remains the same. ,in The feature time step is then used. Max pooling is then employed to reduce the dimensionality, alleviating subsequent computational burden.
[0041] After entering the modal fusion and temporal modeling layer, modal weight parameters are first introduced for each channel feature, and importance scores are assigned according to the following strategy: if a channel has high historical stability (e.g., voltage), it is given a higher fusion weight; if a channel has a long response delay or high noise (e.g., sound waves), its weight is reduced. This weight can be set a priori during the initial training phase or learned automatically through the attention mechanism during subsequent training. The weighted channel features are concatenated into a two-dimensional feature vector sequence, as shown in the image. ,in To unify the time step (e.g., 60 steps), D is the feature dimension of each step after fusion (e.g., 128 dimensions).
[0042] The concatenated feature sequence is then input into a Gated Recurrent Unit (GRU) model to extract medium- to long-term temporal dependencies. The GRU typically has one to two layers, with each layer containing 128 to 256 hidden state dimensions. The core objective of this stage is to identify periodic fluctuations, slow degradation trends, or cross-modal resonance behaviors, such as the correlation between rising resistance and hysteretic temperature increases. The GRU output is either the last hidden state or the average pooling result across all time steps, forming a set of fused feature vectors, in the form of B, D′B, D'.
[0043] In the risk assessment and classification output layer, the fused feature vector is fed into a two-layer fully connected network. The first layer is typically a linear transformation of the ReLU activation function (e.g., input dimension 128, output dimension 64), and the second layer outputs a Sigmoid or Softmax function to generate the risk level classification result. The classification layer output can be set to multiple classes (e.g., normal, mild warning, moderate warning, severe warning), with each class corresponding to a risk label, and can also include the probability confidence level for each class.
[0044] This classification layer, based on trained or configured decision logic, divides the current state of the battery into several preset risk levels. Each risk level is called a "risk label," for example:
[0045] "normal": indicates that no abnormalities or risks were detected;
[0046] "mild_warning": This indicates that a minor anomaly has been detected, which is not enough to trigger a protection action, but observation is recommended.
[0047] "moderate_warning": indicates a moderate level of exception, requiring the system to respond;
[0048] "severe_warning": This indicates that a serious risk has been detected and protective measures must be taken immediately.
[0049] Each label's judgment result is accompanied by a probability confidence value, ranging from 0 to 1, which characterizes the reliability of the model or rule system for that classification result. For example, a confidence of 0.87 means that the system has 87% confidence in the current judgment of "moderate warning".
[0050] In addition, the system will output the main potential failure modes on which the current judgment is based. This field is named predicted_failure_type and its content comes from the system's internal failure feature extraction module. Possible types include, but are not limited to, "thermal_risk" (thermal runaway risk), "leakage_risk" (gas leakage risk), and "impedance_rise" (abnormal increase in internal resistance).
[0051] Finally, to enhance the timeliness of the response system, the system will also provide a time_horizon field to indicate the expected time range for risk development. For example, "within 5 minutes" means that the system predicts that the risk may reach a critical state within 5 minutes, triggering an actual failure.
[0052] The above fields will be combined and encapsulated into a single risk response signal, output in a standardized data structure for the fault warning and response module 103 to read and parse. An example is as follows:
[0053] {
[0054] "status":"moderate_warning",
[0055] "confidence": 0.87,
[0056] "predicted_failure_type":"thermal_risk",
[0057] "time_horizon":"within 5 minutes"
[0058] }
[0059] Internally, the risk response signal can be transmitted and stored in JSON format, as a structure, key-value pair object, or protocol frame. This structure is designed to enable the early warning module to formulate corresponding intervention measures based on risk level, confidence level, failure mode, and time window, achieving a logical closed loop from data analysis to safety control.
[0060] The risk response signal will be sent as a structured message to the fault warning response module 103, triggering the downstream hierarchical control logic.
[0061] The converged computing engine 102 can be deployed on edge gateways, embedded processors or BMS control units based on lightweight frameworks such as TensorFlow Lite, and its overall inference time can be controlled within 200 milliseconds to meet real-time requirements.
[0062] To further optimize the reliability and deployment adaptability of this system, a parameter update interface can be set for the fusion computing engine 102, allowing model iteration and online learning to be achieved by remotely distributing model weights.
[0063] In summary, the Fusion Computing Engine 102, through its modular, standardized, and time-aware structural design, constructs a complete path for data processing, feature fusion, and risk assessment, which can efficiently support battery safety analysis tasks driven by multi-source data.
[0064] Furthermore, the hierarchical algorithm fusion structure includes an intrinsic anomaly detection layer, a conditional causal attribution layer, and a time series generation and inversion layer;
[0065] The intrinsic anomaly detection layer is used to process the multimodal operating data through both the original feature encoding path and the perturbation enhancement path, whereby the perturbation enhancement path introduces a controlled perturbation before input. Information entropy difference calculation is performed on the encoding results of the original feature encoding path and the perturbation enhancement path to quantify the sensitivity of each feature channel to the perturbation. Based on a set anomaly response threshold, the information entropy difference calculation results are filtered to identify sensitive channels that show significant perturbation responses. Temporal change feature extraction is performed on the original operating sequence of the sensitive channels within a preset time window to generate feature deconstruction results representing the distribution of potential unstable factors.
[0066] The conditional causal attribution layer receives the feature deconstruction results and, based on the temporal variation features contained in the feature deconstruction results, constructs a causal path graph model to characterize the response mechanism between variables by combining the temporal correlation and lag response relationship between multiple channels. An adjustable time window is set to extract context from the causal path graph model and obtain the response background information of local paths. An intervention simulation strategy is used to simulate intervention experiments on specific paths in the causal path graph model, calculating the strength of the path's role in the unstable factors leading to the feature deconstruction results. Based on the intervention simulation results, feature combinations with strong causal associations to abnormal patterns in the causal path graph model are identified and output as a set of key triggering factors.
[0067] The time series generation and inversion layer receives the key triggering factor set and the feature deconstruction result, uses them as joint input, and generates multiple future state change sequences using a conditional generative adversarial network to simulate the operational evolution trend of the battery system under different disturbance scenarios. Statistical envelope analysis is performed on the future state change sequences to identify high-risk evolution periods that may occur during the prediction process, such as thermal runaway, abnormal gas release, or drastic changes in internal resistance. Based on the results of the statistical envelope analysis, combined with the abnormal response patterns reflected in the feature deconstruction result, the number of high-risk segments in the future state change sequences, and the importance score of the path corresponding to the key triggering factor set, a risk response signal reflecting the probability of failure, risk level, and urgency of response within multiple time stages is generated.
[0068] The hierarchical algorithm fusion structure described in this invention includes an intrinsic anomaly detection layer, a conditional causal attribution layer, and a time series generation and inversion layer. Each layer has a clearly defined input-output relationship, forming a bottom-up inference chain. This structure takes multimodal operational data as input, extracts dynamic behavioral patterns, causal features, and evolutionary trend information at different levels, and ultimately generates output results for risk response.
[0069] First, in the intrinsic anomaly detection layer, the multimodal operational data received by the system refers to synchronous time-series data from multiple physical sensors, including voltage, current, temperature, humidity, acoustic signals, electrochemical impedance spectroscopy data, and local gas concentration. Each modality's data is input into two parallel processing paths: the original feature encoding path and the perturbation enhancement path. In the original feature encoding path, the system uses a set of structures based on one-dimensional convolutional (1D-CNN) or transform encoders (such as multilayer perceptrons or sparse autoencoders) to perform feature mapping on the original time series of each modality, extracting its intrinsic variation patterns. In the perturbation enhancement path, before inputting the same multimodal time series, the system introduces amplitude-controlled perturbations. These perturbations are low-amplitude, low-frequency additive perturbations or signal offsets (such as a 5% amplitude sinusoidal perturbation or random noise), designed to simulate the sensitivity of each channel's response under boundary perturbation conditions.
[0070] The information entropy difference is calculated for the feature vectors output by the original feature encoding path and the perturbation enhancement path. Specifically, the information entropy is calculated for each path's output along the channel dimension, and the difference between the two entropy values yields the response sensitivity score for each channel. Shannon entropy can be used to calculate the uncertainty change of each channel before and after the perturbation based on the probability distribution of the output features. The system then filters out sensitive channels that show significant perturbation responses based on a set abnormal response threshold (e.g., a sensitivity score greater than 0.35). These channels exhibit significant changes in their encoded features when subjected to minor perturbations, indicating a significant ability to characterize system instability.
[0071] For the selected sensitive channels, the system will backtrack their original operating sequence and extract their temporal variation characteristics within a set time window (e.g., the most recent 60 seconds). This process can use time-domain statistics such as derivatives, volatility, trend direction, and energy density within a sliding window to form a set of temporal feature vectors. This result, called the feature deconstruction result, is used to characterize the spatial distribution and dynamic behavior of potential non-obvious instability factors in the system.
[0072] The feature deconstruction results are input into the conditional causal attribution layer. The core objective of this layer is to establish a temporal response mechanism graph between feature channels. First, based on the temporal variation characteristics recorded in the feature deconstruction results, the system analyzes the correlation and response lag between channels. This process typically uses multivariate cross-correlation, Granger causality tests, or dynamic time warping (DTW) methods to identify response paths between several channel pairs that have causal inference significance. The system then constructs a preliminary causal path graph model using these channels as nodes and lag relationships as edges.
[0073] Subsequently, by setting an adjustable time window (e.g., 5 seconds, 10 seconds, or 30 seconds), the background context information of each channel pair within that window is extracted, further enhancing the model's ability to perceive local changes in the path. For the initially constructed causal path graph model, the system will implement intervention simulation experiments on some paths based on an intervention simulation strategy. That is, in the simulation environment, the input of a certain channel is interrupted or amplified, and its impact on the output of downstream channels is observed, thereby evaluating the strength of the path's role in causing abnormal changes in the feature deconstruction results.
[0074] Based on the simulation results of the aforementioned intervention experiment, the system will identify the channel combinations with the strongest influence on unstable patterns in the causal path graph model and record these channel combinations as a set of key triggering factors. This set of key triggering factors is a crucial input for subsequent evolutionary trend prediction and control decisions.
[0075] Next, the time series generation and inversion layer receives the key triggering factor set and feature deconstruction results, and feeds them as joint input into a set of conditional generative adversarial networks. The generative network uses the triggering factors and their contextual features as conditional variables to output multiple potential future state change sequences, used to simulate the operating trend of the battery system under different perturbation scenarios (such as abnormal current, local overheating, and gas release). Each state change sequence is a set of time-modal tensors, representing the response trajectory of each channel of the system over a future period.
[0076] For each generated sequence of future state changes, the system will perform statistical envelope analysis, which calculates the confidence interval, variance boundary, and anomaly probability distribution of the prediction results at each time point to identify high-risk segments that may experience failures such as thermal runaway, abnormal gas release, or drastic changes in internal resistance during the prediction period. These segments will be recorded for their start and end times, channel locations, and anomaly types.
[0077] Finally, based on the results of statistical envelope analysis, combined with the abnormal response patterns reflected in the feature deconstruction results, the number of high-risk segments in the future state change sequence, and the importance scores of the paths corresponding to the key triggering factor set in the causal graph, a weighted fusion calculation is performed to generate a comprehensive and highly interpretable risk response signal. This risk response signal includes the current predicted risk level (e.g., moderate warning), fault type (e.g., thermal risk), expected occurrence time window (e.g., within 5 minutes), and risk confidence score (e.g., 0.87), and will be transmitted to the subsequent fault warning response module for further control decision generation.
[0078] The following is a reference implementation code for the layered algorithm fusion structure:
[0079] import torch
[0080] import torch.nn as nn
[0081] import numpy as np
[0082] # -------- Intrinsic Anomaly Detection Layer --------
[0083] class IntrinsicAbnormalityDetectionLayer(nn.Module):
[0084] def __init__(self, input_channels):
[0085] super().__init__()
[0086] self.encoder = nn.Conv1d(in_channels=input_channels, out_channels=64, kernel_size=3, padding=1)
[0087] self.perturbation_strength = 0.05 # Set the perturbation strength
[0088] def forward(self, x):
[0089] # x: Input multimodal data, in shape (batch_size, channels, time_steps)
[0090] original_encoded = self.encoder(x)
[0091] # Add controlled perturbation to the input data (using a sinusoidal perturbation as an example)
[0092] noise = self.perturbation_strength torch.sin(2 np.pi torch.rand_like(x))
[0093] perturbed_input = x + noise
[0094] perturbed_encoded = self.encoder(perturbed_input)
[0095] # Use standard deviation to approximate information entropy and calculate channel response sensitivity.
[0096] entropy_diff = torch.std(perturbed_encoded, dim=2) - torch.std(original_encoded, dim=2)
[0097] # Filter sensitive channels based on set thresholds
[0098] threshold = 0.35
[0099] sensitivity_mask = (entropy_diff>threshold)
[0100] # Feature deconstruction: using statistical indicators such as mean derivative, variance, and energy
[0101] time_diff = x[:, :, 1:] - x[:, :, :-1]
[0102] var = torch.var(x, dim=2, keepdim=True)
[0103] energy = torch.mean(x 2, dim=2, keepdim=True)
[0104] mean_derivative = torch.mean(time_diff, dim=2, keepdim=True)
[0105] decomposition_result = torch.cat([mean_derivative, var,energy], dim=2)
[0106] return decomposition_result, sensitivity_mask
[0107] # -------- Conditional Causal Attribution Level --------
[0108] class ConditionalCausalAttributionLayer(nn.Module):
[0109] def __init__(self, window_size=10):
[0110] super().__init__()
[0111] self.window_size = window_size # Adjustable time window
[0112] def compute_delayed_correlation(self, ci, cj, delay=1):
[0113] #Simulate hysteresis response: Manually move ci or cj backward.
[0114] if ci.shape[1] <= delay:
[0115] return torch.tensor(0.0)
[0116] return torch.mean(ci[:, :-delay] cj[:, delay:], dim=1)
[0117] def forward(self, decomposition_result, sensitivity_mask):
[0118] batch_size, channels, _ = decomposition_result.shape
[0119] causal_graph = torch.zeros((batch_size, channels, channels))
[0120] # Construct a causal path graph (using delayed correlation)
[0121] for i in range(channels):
[0122] for j in range(channels):
[0123] if i == j:
[0124] continue
[0125] ci = decomposition_result[:, i, :]
[0126] cj = decomposition_result[:, j, :]
[0127] delay_corr = self.compute_delayed_correlation(ci, cj,delay=1)
[0128] causal_graph[:, i, j] = delay_corr
[0129] # Context Extraction: Retrieve the background response within a local window along a causal path
[0130] # Here, the window mean is used as an example to approximate the background response intensity.
[0131] context_info = torch.mean(causal_graph, dim=2, keepdim=True)
[0132] # Intervention Simulation Strategy: Simulating the impact of disconnecting the path (simulated here using inverted relation values).
[0133] intervention_graph = 1.0 - causal_graph # Simplify the simulation mechanism for "path interruption"
[0134] path_impact = torch.abs(causal_graph - intervention_graph)
[0135] # Extract the strongest causal path combination (trigger factor) in the critical path
[0136] strongest_paths = torch.argmax(path_impact, dim=2)
[0137] trigger_factors = []
[0138] for b in range(batch_size):
[0139] factors = []
[0140] for i in range(channels):
[0141] if sensitivity_mask[b, i]:
[0142] j = strongest_paths[b, i]
[0143] factors.append((i.item(), j.item()))
[0144] trigger_factors.append(factors)
[0145] return causal_graph, trigger_factors, context_info
[0146] # -------- Time Series Generation and Inversion Layer --------
[0147] class TimeSeriesGenerationReversalLayer(nn.Module):
[0148] def __init__(self, input_dim, hidden_dim, forecast_steps):
[0149] super().__init__()
[0150] self.condition_embedding = nn.Linear(input_dim, hidden_dim)
[0151] self.generator = nn.GRU(input_size=hidden_dim, hidden_size=hidden_dim, batch_first=True)
[0152] self.decoder = nn.Linear(hidden_dim, input_dim)
[0153] self.forecast_steps = forecast_steps
[0154] def forward(self, decomposition_result, trigger_factors, context_info):
[0155] batch_size, channels, _ = decomposition_result.shape
[0156] cond_vec = decomposition_result.mean(dim=2)
[0157] cond_emb = self.condition_embedding(cond_vec)
[0158] # Initialize the prediction sequence
[0159] initial_input = torch.zeros(batch_size, self.forecast_steps,cond_emb.shape[1])
[0160] out, _ = self.generator(initial_input, cond_emb.unsqueeze(0))
[0161] predictions = self.decoder(out)
[0162] # Statistical envelope analysis: Calculate the variance and mean of the prediction interval for each channel.
[0163] variance = torch.var(predictions, dim=1)
[0164] mean_pred = torch.mean(predictions, dim=1)
[0165] # Risk assessment logic (using rules as an example): High risk is indicated by large variance or an abnormal mean.
[0166] risk_mask = (variance>0.2) | (mean_pred>1.5)
[0167] # Risk Response Signal Generation
[0168] risk_response = {
[0169] "status":"moderate_warning",
[0170] "confidence": 0.87,
[0171] "predicted_failure_type":"thermal_risk",
[0172] "time_horizon":"within 5 minutes",
[0173] "channels_at_risk": risk_mask.tolist()
[0174] }
[0175] return predictions, risk_response
[0176] Furthermore, the perturbation enhancement path in the hierarchical algorithm fusion structure of the fusion computing engine includes the following implementation steps:
[0177] Each channel of the multimodal operating data is input into the perturbation enhancement path, and an amplitude-controlled sinusoidal perturbation signal is introduced before the input to obtain the perturbed time series.
[0178] The disturbed time series is input into a feature representation model with the same feature encoding path as the original feature, and the corresponding perturbation feature encoding results are extracted.
[0179] The information entropy difference between the perturbation feature encoding result and the output result of the original feature encoding path is calculated for each channel, and a perturbation response sensitivity score is constructed based on the difference result.
[0180] First, the perturbation enhancement path processes multimodal operational data from the multi-source data acquisition module. This data includes time-series information from various sensor channels, such as voltage, current, temperature, humidity, acoustic signals, electrochemical impedance spectroscopy data, and local gas concentration. After receiving this multimodal operational data, the system needs to deconstruct the data according to the channel dimension, that is, input each channel as an independent sequence into the perturbation enhancement path to ensure that the perturbation process is controllable and traceable. Here, "each channel" refers to a continuous historical value sequence of a certain physical quantity acquired at the same timestamp. The data format is usually a two-dimensional tensor, such as (C, T), where C is the number of channels and T is the time step of each channel.
[0181] Before each channel's input perturbation enhancement path, the system introduces a controlled-amplitude sinusoidal perturbation signal into the original time series according to a unified perturbation injection strategy. The perturbation signal design must meet two requirements: first, the perturbation amplitude cannot exceed 10% of the channel mean to avoid altering the data trend; second, the perturbation frequency should be lower than 1 / 10 of the sampling frequency to ensure that the perturbation manifests as a low-frequency boundary disturbance rather than random noise. Taking a voltage channel as an example, if a channel has 120 sampling points within 60 seconds, the system can generate a sinusoidal wave with a period of 40 points and an amplitude of 5% of the channel standard deviation, which is then superimposed onto the original sequence to obtain the "perturbed time series." Mathematically, this can be expressed as:
[0182] ;
[0183] Where A is the controlled amplitude, f is the set frequency, and t is the time index.
[0184] After perturbation injection, the system inputs the perturbed time series of each channel into a feature representation model with the same original feature encoding path for unified processing. This model is typically a one-dimensional convolutional neural network (1D-CNN) or a sparse autoencoder, which extracts local variation patterns and overall trend features of the signal at different time scales. The convolution kernel parameters, activation functions, and pooling layer configurations in this network structure are consistent with the original path to ensure feature comparability. The output is the perturbed feature encoding result for each channel, in the format of a three-dimensional tensor (B, C, F), where B is the batch size, C is the number of channels, and F is the dimension of the encoded features per channel.
[0185] Simultaneously, the system retains the output results of the corresponding channels in the original feature encoding path and maps the perturbed feature encoding results one-to-one with the original feature encoding results, performing information entropy difference calculation along the channel dimension. Information entropy, as an indicator of feature uncertainty, is typically calculated using the Shannon entropy formula: ;
[0186] Where p(x) is the normalized probability density function of the feature in different dimensions. In practice, the encoded vector can be normalized first, and then the overall feature entropy of the channel can be calculated by taking the distribution entropy of each dimension. After calculating the information entropy of the original and perturbed encoded results respectively, the difference between the two under the same channel is taken to obtain the "information entropy difference".
[0187] Finally, the system uses the information entropy difference of each channel as a perturbation response sensitivity score. The higher the value, the more significant the increase in uncertainty of the channel's encoded features after being perturbed, i.e., the stronger the response to perturbation. The sensitivity score can be stored as a floating-point vector in the feature dictionary of the fusion computing engine and used for subsequent sensitive channel screening logic, abnormal region construction, or causal attribution path selection.
[0188] Through the three interconnected steps described above, this invention realizes a full-chain information processing structure in the perturbation enhancement path, from perturbation simulation and unified coding to sensitivity modeling. It has strong generalization and interpretability, which can significantly enhance the system's early identification capability of boundary anomalies and provide high-resolution risk starting point information for subsequent modules.
[0189] Furthermore, the fusion computing engine is also used for:
[0190] Based on the disturbance response sensitivity score, all feature channels are sorted according to their response intensity, and a set of sensitive channels that are significantly sensitive to disturbance responses is selected by combining the set abnormal response threshold.
[0191] Extract the original time series corresponding to the sensitive channel set from the multimodal operation data, and perform joint feature extraction operation based on the first derivative rate of change, short-time energy density and variance fluctuation intensity within a preset time window;
[0192] The results of the joint feature extraction operation are fused to generate a feature deconstruction result that represents the distribution of potential unstable factors under perturbation conditions. The feature deconstruction result is used as the input to the subsequent conditional causal attribution layer and time series generation inversion layer.
[0193] After obtaining the disturbance response sensitivity score, the fusion computing engine first performs a unified sorting operation on the scores of all feature channels. The disturbance response sensitivity score is calculated based on the information entropy difference between the disturbance enhancement path and the original feature encoding path. Each score corresponds to a feature channel, reflecting the severity of its encoded feature change under small disturbances. The sorting operation arranges all channels from highest to lowest score, aiming to prioritize the identification of channels most sensitive to disturbance fluctuations; these channels often reflect early faults. After sorting, the fusion computing engine introduces a set of preset abnormal response thresholds. These thresholds can be absolute values (e.g., 0.25) or dynamic quantile values based on statistical distribution (e.g., Top 20%). Using these thresholds, channels with scores higher than the threshold are selected from the sorted channels, forming a "set of sensitive channels with significant disturbance response." This set typically consists of 2-5 channels, and its number can be dynamically adjusted according to different operating conditions.
[0194] Next, the fusion computing engine will extract the time series corresponding to the aforementioned sensitive channel set from the original multimodal runtime data. Data for each channel is collected at a fixed time step, typically 2-10 samples per second, with the extraction range set to historical records within a preset time window, such as 60 or 120 seconds. Within this time window, the system will perform joint feature extraction on each sensitive channel. Joint feature extraction refers to simultaneously extracting three core features characterizing abnormal dynamics from different statistical dimensions: the rate of change of the first derivative, short-time energy density, and variance fluctuation intensity.
[0195] Among them, the rate of change of the first derivative refers to the mean or maximum value of the difference sequence between two adjacent sampling points in the sequence, used to capture abrupt change trends, and is suitable for detecting phenomena such as sharp voltage drops and rapid current increases; short-time energy density refers to the mean of the squared values of the sequence within a local sliding window, used to measure the power fluctuation amplitude within a short period of time, and is suitable for identifying high-frequency oscillations or acoustic interference; variance fluctuation intensity refers to the dispersion index of samples within this window, reflecting the stability level of the signal, and is suitable for detecting gas concentration fluctuations or abnormal temperature fluctuations. These three features will be extracted in the form of quantitative indicators to constitute a three-dimensional feature vector for each channel within this time period. To maintain dimensional uniformity and input standardization, all features will be normalized.
[0196] After concatenating the 3D feature vectors of all sensitive channels, the joint feature set at that moment is obtained. The fusion computing engine then performs feature fusion operations on these joint features. The fusion method can be simple concatenation followed by a fully connected layer, or a weighted fusion strategy (such as an attention mechanism) can be used to generate a single fused feature vector. This fused feature not only preserves the important local information of each channel, but also constructs the cooperative relationship between channels, which can comprehensively reflect the distribution pattern of potential unstable factors in the system under perturbation conditions across multiple channels.
[0197] The resulting fused feature vector constitutes the "feature deconstruction result" described in this invention. This feature deconstruction result is a highly semantic representation generated after perturbation-enhanced response discrimination, anomaly-sensitive channel identification, statistical feature linkage modeling, and multi-channel feature fusion processes. It possesses high discriminativeness and interpretability. This result serves as direct input to subsequent modules in the fusion computing engine, providing the conditional causal attribution layer with information to analyze the causal mechanisms between variables, and the time series generation and inversion layer with information to predict future risk trends, ensuring causal traceability throughout the entire inference path.
[0198] Furthermore, the intrinsic anomaly detection layer includes the following steps in identifying sensitive channels that respond significantly to disturbances and generating feature deconstruction results to represent the distribution of potential instability factors:
[0199] Based on the outputs of the original feature encoding path and the perturbation enhancement path, the information entropy of the encoded features for each channel is calculated, thereby obtaining the difference in information entropy of the perturbation response. ;
[0200] Extract the mean energy density E of each channel before and after the perturbation. i and its increment ΔE i ;
[0201] Calculate the Kullback–Leibler divergence between the original feature coding distribution and the perturbation coding distribution, denoted as KL. i This is used to measure the degree of morphological shift in the characteristic distribution response;
[0202] Each channel is quantitatively weighted according to the following perturbation response comprehensive scoring function: ;
[0203] in, Indicates channel The overall score of the disturbance response; This represents the difference in information entropy. Indicates channel The mean energy density of the original encoded data; ΔE i Indicates channel The increment of disturbance energy; Let represent the Kullback–Leibler divergence between the distributions of encoded features before and after the perturbation; α, β, and γ are non-negative weighting coefficients that satisfy . ;
[0204] All channels are sorted based on the scoring function results, and a response threshold is set. Filter out those that meet the requirements The channels constitute a set of sensitive channels;
[0205] The original running sequence of the sensitive channel set is extracted from the multimodal running data, local change features are extracted within a preset time window, and a multidimensional time series feature vector is constructed.
[0206] The multidimensional temporal feature vector and channel scoring results are input together into the subsequent conditional causal attribution layer to guide the initial edge weight setting and abnormal causal path identification of the causal path graph model.
[0207] In the process of performing intrinsic anomaly detection, the feature encoding results of each channel must first be obtained from the original feature encoding path and the perturbation enhancement path, respectively. The original feature encoding path refers to the representation vector of each channel obtained after processing the multimodal running data with a uniform neural network structure (such as one-dimensional convolution, autoencoder, or attention encoder) without perturbation input. The perturbation enhancement path applies a perturbation with controlled amplitude (such as low-frequency additive sine or random micro-fluctuations) before the input, and then encodes it using the same network structure as the original path.
[0208] For each channel Calculate the Shannon information entropy of the original feature code and the perturbed feature code respectively. The probability distribution can be approximated as follows: Perform softmax normalization on each encoded vector to obtain the probability distribution. Then through the formula The information entropy difference is obtained. The degree of uncertainty in the coding distribution before and after the perturbation is defined as:
[0209] ;
[0210] This represents the information entropy of the feature distribution obtained after encoding channel i under the original input data conditions. The acquisition method includes: feeding the original input (without perturbation) into a feature encoder (e.g., a 1D CNN or AutoEncoder) to obtain a real-valued vector; then using the softmax function to normalize this real-valued vector into a probability distribution; and finally calculating the Shannon entropy of this distribution to obtain the information entropy. .
[0211] This represents the information entropy of the encoded feature distribution of channel i under perturbed input conditions. The acquisition method is the same as the original case: after the perturbed input passes through the encoder, the softmax function is used to obtain the probability distribution, and then the Shannon entropy is calculated.
[0212] The information entropy difference represents the difference in information entropy before and after the disturbance. This difference is used to quantify the sensitivity of a channel to the disturbance signal. The larger the difference, the stronger the response of the channel's coding features to the disturbance.
[0213] Next, the system calculates the mean energy density for each channel, which is the average of the squared values of the feature encoding vectors, defined as:
[0214] ;
[0215] in, It is a passage The mean energy density of the original encoded features; For channel In feature encoding dimension The value on; It is the total dimension of the encoded feature vector, i.e. the dimension length, which may be 64, 128 or 256, depending on the output dimension of the encoder network structure used;
[0216] The disturbance energy increment is defined as the average energy after the disturbance input. Compared with the original energy mean The difference is denoted as This indicator reflects whether a disturbance causes an increase or sharp fluctuation in the channel's characteristic strength.
[0217] To further reflect changes in the morphology of the feature distribution, the system introduces the Kullback-Leibler divergence as a statistical distance index. For the same channel... The original distribution and perturbation distribution Defined as:
[0218] ;
[0219] in and This is the normalized probability density value, usually obtained by standardizing the feature vector using softmax. This divergence measure indicates whether the perturbation has caused a reconstruction or shift in the overall coding distribution; a larger value indicates a significant change in the way the channels are represented after the perturbation.
[0220] This represents the first (original) encoding before the perturbation. The normalized probability value corresponding to the dimensional feature;
[0221] This represents the normalized probability value corresponding to the j-th dimension feature after perturbation (perturbation coding).
[0222] To unify the above three factors into a single score, the system introduces a disturbance response comprehensive scoring function:
[0223] ;
[0224] This function integrates perturbation sensitivity (information entropy), relative change in energy increment (amplitude response), and distribution shift (KL divergence). Coefficients Used to adjust the weights of different indicators to satisfy the normalization constraint. Recommended settings are as follows: . A larger value helps to improve sensitivity to disturbances. A larger value indicates a greater emphasis on identifying distribution changes.
[0225] The system calculates for each channel. Then, the results are normalized and sorted, and those that meet the threshold conditions are selected. The channels are used as the set of sensitive channels. Threshold It can be a static value (such as 0.5) or a dynamic setting (such as the average plus the standard deviation).
[0226] The system then extracts the original running sequences of these sensitive channels from the raw multimodal running data within a recent time window (e.g., the past 60 seconds), calculates their local variation features (e.g., first derivative, local variance, short-time energy), and constructs them into three-dimensional or higher-dimensional temporal feature tensors. Each tensor contains a channel ID, a temporal dimension, and a local statistical dimension, along with a comprehensive score for the channel. Together they form a structured input.
[0227] The structured input will then be fed into the conditional causal attribution layer to initialize the node priorities and initial edge weights in the causal path graph, ensuring that sensitive channels are given priority during the graph construction process, which is beneficial for subsequent path backtracking and identification of key anomaly mechanisms.
[0228] In summary, this embodiment integrates multiple key indicators through a defined perturbation response scoring function, and constructs a channel sensitivity ranking mechanism and feature deconstruction expression based on this. It not only has a high recognition rate, but also cross-modal scalability.
[0229] The fault warning response module 103 is used to receive the risk response signal and generate hierarchical control instructions corresponding to different risk levels according to the preset response strategy. The hierarchical control instructions include one or more of the following: activating virtual circuit breaker logic, scheduling regional cooling resources, refreshing BMS operation mode, performing SOC valuation correction, or pushing remote risk alarm information.
[0230] The fault warning and response module 103 is the core decision-making component in the battery safety management platform. Its main function is to dynamically generate control response measures that match the risk level based on the risk classification judgment logic preset in the system after receiving the risk response signal output by the fusion computing engine 102. The output is in the form of structured hierarchical control instructions. These instructions will be received by the execution interface module 104 and transmitted to the physical controller or management unit of the battery system, thereby realizing online risk intervention.
[0231] In its specific design, the fault warning and response module 103 first parses the risk response signal from the fusion computing engine 102. This signal includes at least the current risk level label, possible fault types (such as thermal runaway, gas leakage, abnormal internal resistance, etc.), predicted risk window (such as expected to trigger within 5 minutes), and system confidence index. Based on this information, the module matches the corresponding control strategy in a preset response strategy table. The strategy table can be in the form of JSON structure, database table, or hash mapping table, and supports real-time updates and configuration.
[0232] To enhance the flexibility and accuracy of judgment, the fault early warning and response module integrates a simple and implementable hierarchical algorithm fusion structure, consisting of three layers: the first layer is a rule-based judgment layer, which uses condition matching to initially classify risk levels; the second layer is a parameter evaluation layer, which performs threshold quantization based on numerical values in the risk signal (such as the rate of increase in internal resistance and the magnitude of gas concentration increase); and the third layer is a strategy decision layer, which calculates the optimal control decision by comprehensively considering the weights of multiple fields such as risk type, level, and prediction window. This structure requires no complex model training, can run directly with configured rules, is suitable for deployment in embedded control chips or middleware, and has engineering feasibility.
[0233] In the hierarchical control instructions output by the fault early warning and response module 103, each control item has clear execution semantics and system interface requirements:
[0234] Activating the virtual fuse logic involves sending a virtual command to the battery management system (BMS) to disconnect the relay, simulating a power outage. This quickly cuts off the current path without requiring a physical fuse to activate, addressing short-term, high-risk electrical anomalies. This logic requires the BMS to have soft-break protection capabilities and a secure write-back mechanism to ensure the command is not accidentally triggered.
[0235] Scheduling regional cooling resources refers to utilizing the air-cooled, liquid-cooled, or phase-change material cooling components configured in the battery system. A cooling execution command is sent to the regional temperature control unit, which rapidly reduces the temperature of hot cells by increasing fan speed, adjusting coolant flow rate, or activating backup heat dissipation paths. This control command should include cooling intensity and duration parameters, and the target cooling area can be specified based on sensor number or cell serial number.
[0236] Refreshing the BMS operating mode refers to issuing an update command to the battery management system to update operating parameters, such as switching charging and discharging strategies, limiting maximum current output, or temporarily disabling the equalization function, in order to reduce system stress or adjust the operating status when a risk is anticipated. This command must include the target mode number and the update effective time to ensure synchronization with the system master controller.
[0237] Performing SOC valuation correction refers to re-estimating the current State of Charge (SOC) by invoking a backup valuation model or enabling a deep update mechanism when a deviation in the existing state estimate is detected during risk trend detection. This operation relies on the estimation module in the BMS or an external computing node and can be completed before the system goes into sleep mode or within a safety window, ensuring that subsequent control strategies are based on accurate state information.
[0238] Pushing remote risk alarm information refers to encapsulating the content of risk response signals into alarm data packets and sending them to a host computer or remote monitoring center via CAN bus, 4G / 5G module, LoRa, or Ethernet interface. The data packet should include the device number, risk type, predicted time, response level, and system status summary to facilitate backend operation and maintenance scheduling, alarm confirmation, or emergency shutdown.
[0239] Based on the aforementioned hierarchical algorithm fusion structure, the fault early warning and response module can combine preset response strategies to map risk classification results of different levels to corresponding hierarchical control instructions, and generate specific execution content in the structured decision-making process. The generation process of each control instruction is based on specific risk assessment results and parameter evaluation data, combined with the control response rules set in the strategy decision layer, to determine the action type, execution object, and parameter values. The generation processes of several typical control instructions are described below.
[0240] When the rule-based judgment layer identifies a risk level of "Severe Warning," and the parameter evaluation layer determines that a voltage change in a certain area within the battery pack, accompanied by abnormal acoustic signals and a sharp increase in gas concentration, reaches the threshold of pre-thermal runaway characteristics, the strategy decision layer will automatically match a high-priority fuse response strategy. This strategy requires immediate disconnection of the bus connection to prevent escalating risks; therefore, the generated control command is "Activate Virtual Fuse Logic." Its execution fields include: target location for disconnection (e.g., bus A), confirmation identifier, and operation priority.
[0241] In another scenario, when the system identifies a "moderate warning" and the internal resistance growth rate exceeds a set threshold but has not yet triggered a hardware alarm, the temperature signal shows a gradual upward trend. The strategy decision layer will select a milder intervention strategy based on the time window prediction results, namely "scheduling regional cooling resources." The control command will include the cooling method (such as increasing the coolant flow rate), the control target (such as the 4th–6th series cells), the cooling duration (such as 600 seconds), and the cooling intensity level.
[0242] If the assessment results show the current risk level as "mild warning," but simultaneously determine that the SOC estimate is inconsistent with the actual load behavior (e.g., the current SOC shows 90% but the output voltage is significantly lower), it indicates that the SOC estimate has drifted. In this case, the strategy layer will match the "execute SOC estimate correction" strategy. The generated control commands include enabling the backup estimation algorithm flag, correcting the window start mark, current integration parameters, and a forced refresh flag.
[0243] For example, when the risk level is at the critical point of "moderate warning" or "severe warning" and the current fluctuates drastically, the parameter evaluation layer identifies extreme load changes or unstable system operating environment. At this time, the strategy layer can match the strategy of "refreshing BMS operating mode". The control instruction will be specified to limit the maximum charging current to 80A, disable the equalization function and switch to the safety inspection mode to ensure that the system load remains stable before the risk subsides.
[0244] Furthermore, for all judgments reaching the warning level, regardless of whether physical intervention is triggered, the strategy layer can match a "push remote risk alarm information" command to report risk details to the remote monitoring center or cloud platform. The command content includes the device number, risk type, warning level, summary of the current system status, list of response actions, and signal generation timestamp.
[0245] Through the aforementioned hierarchical decision-making logic, the fault early warning and response module can transform complex multi-source risk information into targeted, parameter-configurable, and logically closed-loop control commands, achieving a highly automated response path from data judgment to control execution.
[0246] The fault warning and response module 103 employs an automatic closed-loop mechanism for its entire response process. All control commands are recorded in the system event log in real time after being output, and execution confirmation information or failure status is recorded through the feedback path after the execution interface module 104 completes its operation. This module supports multi-level command redundancy configuration and response rollback strategies to ensure that the system can achieve effective safety intervention with the shortest path and lowest cost when a fault risk occurs.
[0247] In summary, the fault warning and response module 103, through structured risk analysis, hierarchical response strategy generation, and diversified control command output mechanisms, realizes the core transformation logic from "risk identification" to "safety control" in the battery system. Those skilled in the art can, based on the above description and in conjunction with existing embedded controllers, BMS communication protocols, and industrial control logic, fully implement the hardware and software integration and deployment of this module.
[0248] Furthermore, the fault warning response module is specifically used for:
[0249] Extract the risk level label, fault type identifier, prediction time window, confidence score, and channel risk distribution information contained in the risk response signal to construct a structured risk metadata dataset;
[0250] Based on the risk level label, a set of candidate control strategies matching the current level is retrieved from the preset response strategy rule table. Each strategy in the response strategy rule table defines its applicable risk level range, target control object, execution timeliness, and operational intensity level.
[0251] Based on the prediction time window and the confidence score, combined with historical control strategy execution records, dynamic priority scores are assigned to each strategy in the candidate control strategy set, and a priority scheduling list for strategy screening is constructed.
[0252] Control strategies that meet the control time requirements and have an execution success rate score of not less than a set threshold are selected sequentially from the priority scheduling list to generate a preliminary set of control instructions that includes control type, target channel or device, execution parameters and response window.
[0253] The initial set of control instructions is subjected to conflict detection and compatibility analysis to eliminate control measures that cause resource conflicts or logical contradictions. The remaining instructions are then encapsulated with field structures to generate final hierarchical control instructions that include control target identifier, action instruction code, response level field, timeout processing parameters, and redundancy confirmation flag.
[0254] First, upon receiving the risk response signal from the fusion computing engine, the fault warning and response module performs structured parsing. This risk response signal consists of multiple fields, typically including risk level labels (such as normal, mild_warning, moderate_warning, severe_warning), fault type identifiers (such as thermal_risk, gas_release, impedance_rise, etc.), prediction time window (such as "within 5 minutes"), confidence score (such as 0.91), and channel risk distribution information (i.e., a Boolean array or probability value array indicating which channels are classified as high-risk during the prediction period). The system extracts these fields and constructs a set of structured metadata, assigning a specific data type and unit to each field, serving as the foundational data for subsequent strategy selection and parameter scheduling.
[0255] Based on risk level labels, the system will invoke a pre-defined response strategy rule table within the fault warning and response module. This rule table is organized in key-value pairs, with the index being the risk level label (e.g., "severe_warning") and the corresponding value being a set of control strategies. Each control strategy records its applicable risk level range (e.g., moderate and above), target control object (e.g., area cooling pump, main relay, BMS valuation submodule, etc.), execution timeliness (e.g., needing to be completed within 10 seconds), and operational intensity level (e.g., power outage, power limiting, model correction, etc.). This rule table supports configuration updates, ensuring the system can scalably respond to response needs in different scenarios.
[0256] Next, the system dynamically prioritizes the selected candidate strategies based on the predicted time window and confidence score in the current risk response signal. To avoid the strategy rigidity problem caused by static configuration, the system introduces historical control strategy execution records, including indicators such as average execution time, success rate, and system response deviation for each strategy in similar risk scenarios. The system compares the predicted time window with the strategy execution time, eliminating schemes with estimated completion times exceeding the response window; simultaneously, it weights the confidence score with the strategy execution risk level to calculate the priority score for each strategy. Finally, the system sorts all candidate strategies in descending order of score, forming a strategy scheduling list to guide subsequent strategy refinement and deployment order control.
[0257] From this priority scheduling list, the system will sequentially select strategies that meet the following two conditions: first, the response time is within the prediction window; and second, its historical execution success rate is not lower than a set threshold (e.g., 90%). The selected strategies will serve as a preliminary set of control instructions. Each control instruction will include a control type (e.g., "power off" or "cooling"), a target channel number or device ID (e.g., pack7.cell3 or cooling_unit_A), execution parameters (e.g., cooling intensity 50%, current limit 100A), and response window parameters (e.g., effective within 10 seconds and maintained for 60 seconds). This set constitutes the complete control candidate set.
[0258] However, before actual execution, the system needs to perform conflict detection and compatibility analysis on the aforementioned preliminary set of control commands. Conflict detection is used to identify situations where the same device or channel is scheduled by multiple control commands within the same time period, such as the same cooling pump being required to operate in two different fan speed modes simultaneously, or the same battery cell being simultaneously scheduled for power-off and voltage limiting. Compatibility analysis is used to identify contradictions between control logics, such as the simultaneous issuance of equalization function activation and SOC freeze. In this step, the system will construct a conflict graph structure to identify and eliminate commands with mutual exclusion relationships.
[0259] Finally, the retained control instructions will be repackaged into final executable hierarchical control instructions. Each control instruction is encapsulated according to a fixed structure of fields, including a control target identifier (e.g., hardware_id = “relay_07”), an action instruction code (e.g., code = “CUT_OFF”), a response level field (e.g., level = 3), a timeout handling parameter (e.g., timeout = 5000ms), and a redundancy acknowledgment flag (e.g., ack_required = true). This structure facilitates protocol interfacing with the physical controller and also enables subsequent execution interface modules to perform batch transmissions and acknowledgment matching.
[0260] The execution interface module 104 is used to receive the hierarchical control command and map the hierarchical control command to the control interface connected to the physical controller or management unit of the battery system, thereby driving the actual execution of the corresponding control operation to realize online risk intervention and response scheduling of the battery system.
[0261] The execution interface module 104 receives the hierarchical control commands output by the fault warning response module 103 and translates them into standard control commands that can be recognized and responded to by each physical controller of the battery system. Each type of control command corresponds to a clear control objective, data structure, and issuance process. Typical implementation examples for each control command are described below.
[0262] When the fault warning response module determines the current status to be a critical warning and outputs the "Activate Virtual Fuse Logic" command, the execution interface module interprets it as sending a disconnect command to the high-voltage relay controller in the battery system. For example, in a CAN communication-based system, this command will construct a data frame with ID 0x18FF30A1, where the first byte is set to 0x01 to represent the disconnect command, and subsequent bytes are filled with the controller address and redundancy check bits. After sending the command, the execution interface module will wait for the target controller to return a status code (e.g., 0xAA to confirm disconnection). If there is no response within a set time, a second attempt will be triggered or a report will be sent to a remote platform. This operation is suitable for critical protection scenarios such as controlling bus power outages and shutting down charging and discharging paths.
[0263] When the system identifies a localized overheating trend in the battery cell temperature distribution and outputs a "schedule regional cooling resources" command, the execution interface module will locate the corresponding liquid cooling module in that region based on the target location field in the risk response signal (e.g., the 7th battery cell). Subsequently, it sends a duty cycle increase command via the PWM control channel, increasing the coolant circulation pump's operating rate from 20% to 60%. If the Modbus RTU protocol is used, this command manifests as writing data 0x003C (i.e., 60% duty cycle) to register address 0x003C and enabling the "auto-hold" bit to maintain this cooling state for 10 minutes. If the system has multiple heat dissipation components, this module can also control fan speed or compressor cooling power in parallel to achieve multi-source combined cooling.
[0264] If the fusion computing engine identifies an anomaly in the estimated battery state of charge (SOC) that does not match the actual voltage and current characteristics, the warning module will output a "Execute SOC estimation correction" command. After parsing this command, the execution interface module sends control information to the SOC module of the BMS main controller, activating its backup estimation logic, such as enabling coulomb integration to re-estimate the charge. At this time, the control fields will include an "Estimation mode switch" flag, an "Estimation start point" (e.g., initial current integral value), and a "Correction window duration" (e.g., 600 seconds). Some platforms support writing control registers directly to the SOC estimation IC via the SPI interface (e.g., writing register 0x04 to 0xA5 to initiate the correction cycle). After the operation is complete, the module will monitor changes in the SOC value and determine whether it has returned to a reasonable range.
[0265] When a moderate risk occurs and the battery operating mode needs adjustment, the warning module outputs a "Refresh BMS Operating Mode" command. The execution interface module converts this command into a parameter update packet and sends it to the BMS main control unit, such as setting the maximum discharge current to reduce from 200A to 120A, disabling the equalization function, and switching to static monitoring mode. Specifically, three bytes can be written to the configuration register area via CAN frame ID 0x18EF2209: the first byte sets the current limit level (e.g., 0x78 for 120A), the second byte sets the equalization status (e.g., 0x00 for disabled), and the third byte sets the operating mode (e.g., 0x03 for monitoring mode). This operation is generally accompanied by configuration confirmation feedback; for example, receiving a 0xCC response frame indicates that the command has taken effect.
[0266] Finally, for the "push remote risk alarm information" command that needs to be synchronized to a remote system or dispatch backend, the execution interface module encodes and compresses the risk response signal (such as using protobuf or JSON format) and sends it to the specified URL interface or MQTT channel through the cellular communication module (4G / 5G) or Ethernet port.
[0267] The data will be sent to the remote platform interface via HTTP POST or published to the bms / alert channel via MQTT for storage, analysis, or further control and scheduling by the cloud platform.
[0268] The execution interface module is designed to not only support the execution of the aforementioned instructions, but also features a multi-threaded queue mechanism to ensure that when multiple control instructions of different types are received simultaneously, they can be processed sequentially according to priority, avoiding control conflicts. Furthermore, to ensure the traceability of the control process, each instruction and its execution result (success / failure, response delay, etc.) are written to an operation log and periodically uploaded to the security audit module.
[0269] Through the clear examples of the various control scenarios described above, the execution interface module 104 can effectively realize rapid closed-loop control from risk identification results to physical intervention actions, meeting the requirements of actual battery systems for response speed, safety assurance, and information feedback.
[0270] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
Claims
1. An algorithm fusion and fault early warning system in a battery safety management platform, characterized in that, include: The multi-source data acquisition module is used to acquire multi-modal operating data, including voltage, current, temperature, humidity, acoustic signals, electrochemical impedance spectroscopy data, and local gas concentration, during the operation of the battery system. The fusion computing engine is used to receive the multimodal operation data and generate risk response signals to characterize battery safety risks based on a hierarchical algorithm fusion structure that establishes an inference chain between time-series correlation modeling, causal feature identification and risk trend inference. The fault warning and response module is used to receive the risk response signal and generate hierarchical control instructions corresponding to different risk levels according to the preset response strategy. The hierarchical control instructions include one or more of the following: activating virtual circuit breaker logic, scheduling regional cooling resources, refreshing BMS operation mode, performing SOC valuation correction, or pushing remote risk alarm information. The execution interface module is used to receive the hierarchical control commands and map the hierarchical control commands to the control interface connected to the physical controller or management unit of the battery system, thereby driving the actual execution of the corresponding control operations to realize online risk intervention and response scheduling of the battery system.
2. The algorithm fusion and fault early warning system in the battery safety management platform according to claim 1, characterized in that, The hierarchical algorithm fusion structure includes an intrinsic anomaly detection layer, a conditional causal attribution layer, and a time series generation and inversion layer. The intrinsic anomaly detection layer is used to process the multimodal operating data through both the original feature encoding path and the perturbation enhancement path, whereby the perturbation enhancement path introduces a controlled perturbation before input. Information entropy difference calculation is performed on the encoding results of the original feature encoding path and the perturbation enhancement path to quantify the sensitivity of each feature channel to the perturbation. Based on a set anomaly response threshold, the information entropy difference calculation results are filtered to identify sensitive channels that show significant perturbation responses. Temporal change feature extraction is performed on the original operating sequence of the sensitive channels within a preset time window to generate feature deconstruction results representing the distribution of potential unstable factors. The conditional causal attribution layer receives the feature deconstruction results and, based on the temporal variation features contained in the feature deconstruction results, constructs a causal path graph model to characterize the response mechanism between variables by combining the temporal correlation and lag response relationship between multiple channels. An adjustable time window is set to extract context from the causal path graph model and obtain the response background information of local paths. An intervention simulation strategy is used to simulate intervention experiments on specific paths in the causal path graph model, calculating the strength of the path's role in the unstable factors leading to the feature deconstruction results. Based on the intervention simulation results, feature combinations with strong causal associations to abnormal patterns in the causal path graph model are identified and output as a set of key triggering factors. The time series generation and inversion layer receives the key triggering factor set and the feature deconstruction results, uses them as joint inputs, and generates multiple future state change sequences using a conditional generative adversarial network to simulate the operational evolution trend of the battery system under different disturbance scenarios. Statistical envelope analysis is performed on the future state change sequences to identify high-risk evolution periods that may occur during the prediction process, such as thermal runaway, abnormal gas release, or drastic changes in internal resistance. Based on the results of the statistical envelope analysis, combined with the abnormal response patterns reflected in the feature deconstruction results, the number of high-risk segments in the future state change sequences, and the importance score of the path corresponding to the key triggering factor set, a risk response signal reflecting the probability of failure, risk level, and urgency of response within multiple time stages is generated.
3. The algorithm fusion and fault early warning system in the battery safety management platform according to claim 2, characterized in that, The perturbation enhancement path in the hierarchical algorithm fusion structure of the fusion computing engine includes the following implementation steps: Each channel of the multimodal operating data is input into the perturbation enhancement path, and an amplitude-controlled sinusoidal perturbation signal is introduced before the input to obtain the perturbed time series. The disturbed time series is input into a feature representation model with the same feature encoding path as the original feature, and the corresponding perturbation feature encoding results are extracted. The information entropy difference between the perturbation feature encoding result and the output result of the original feature encoding path is calculated for each channel, and a perturbation response sensitivity score is constructed based on the difference result.
4. The algorithm fusion and fault early warning system in the battery safety management platform according to claim 3, characterized in that, The fusion computing engine is also used for: Based on the disturbance response sensitivity score, all feature channels are sorted according to their response intensity, and a set of sensitive channels that are significantly sensitive to disturbance responses is selected by combining the set abnormal response threshold. Extract the original time series corresponding to the sensitive channel set from the multimodal operation data, and perform joint feature extraction operation based on the first derivative rate of change, short-time energy density and variance fluctuation intensity within a preset time window; The results of the joint feature extraction operation are fused to generate a feature deconstruction result that represents the distribution of potential unstable factors under perturbation conditions. The feature deconstruction result is used as the input to the subsequent conditional causal attribution layer and time series generation inversion layer.
5. The algorithm fusion and fault early warning system in the battery safety management platform according to claim 3, characterized in that, The time series generation and inversion layer in the fusion computing engine is specifically used for: The feature deconstruction results are aligned with the key triggering factor set to construct a joint input tensor, which includes the temporal feature encoding of each channel and the corresponding causal path number, lag time label and trigger strength coefficient. The joint input tensor is input as a conditional variable into a conditional generative adversarial network. The conditional generative adversarial network generates multiple future state change sequences based on a trained perturbation response adversarial discrimination mechanism. The future state change sequences are used to simulate the operational evolution trend of the battery system under different perturbation scenarios. The future state change sequence is subjected to time-domain normalization to ensure the comparability of output sequences across channels in terms of amplitude, slope and fluctuation range, and to provide standardized input for subsequent risk boundary detection.
6. The algorithm fusion and fault early warning system in the battery safety management platform according to claim 5, characterized in that, The time series generation and inversion layer in the fusion computing engine is also used for: The future state change sequence is segmented by a sliding window, and a dynamic statistical envelope range for the evolution of the predicted value over time is constructed for each channel. The dynamic statistical envelope range includes a central trend curve, upper and lower boundary curves, and an anomaly threshold extrapolation interval. The dynamic statistical envelope range is matched with the historical evolution interval for overlap, and the risk density index of the current predicted evolution segment is calculated based on the boundary offset velocity, fluctuation area density and multi-channel cross-over overlap. Based on the risk density index, key time periods containing high-risk envelope features in future state change sequences are identified. Combining the feature deconstruction results with the path importance score of the key triggering factor set, a risk response signal with multiple time periods, multiple causal sources, and multiple response levels is generated.
7. The algorithm fusion and fault early warning system in the battery safety management platform according to claim 1, characterized in that, The fault early warning response module is specifically used for: Extract the risk level label, fault type identifier, prediction time window, confidence score, and channel risk distribution information contained in the risk response signal to construct a structured risk metadata dataset; Based on the risk level label, a set of candidate control strategies matching the current level is retrieved from the preset response strategy rule table. Each strategy in the response strategy rule table defines its applicable risk level range, target control object, execution timeliness, and operational intensity level. Based on the prediction time window and the confidence score, combined with historical control strategy execution records, dynamic priority scores are assigned to each strategy in the candidate control strategy set, and a priority scheduling list for strategy screening is constructed. Control strategies that meet the control time requirements and have an execution success rate score of not less than a set threshold are selected sequentially from the priority scheduling list, and a preliminary set of control instructions containing control type, target channel or device, execution parameters and response window is generated. The initial set of control instructions is subjected to conflict detection and compatibility analysis to eliminate control measures that cause resource conflicts or logical contradictions. The remaining instructions are then encapsulated with field structures to generate final hierarchical control instructions that include control target identifier, action instruction code, response level field, timeout processing parameters, and redundancy confirmation flag.
8. The algorithm fusion and fault early warning system in the battery safety management platform according to claim 1, characterized in that, The multi-source data acquisition module is also used for: The synchronous time series corresponding to the acoustic signal channel and temperature signal channel collected during the operation of the battery system are obtained, and the two are timestamped to form a unified acoustic-temperature coupling input pair. A short-time Fourier transform is performed on the acoustic signal channel to extract the spectral energy density distribution characteristics, and local spectral bands with abrupt increases in amplitude within the 1kHz to 5kHz frequency band are identified as candidate acoustic anomaly intervals. Within the acoustic anomaly candidate interval, the temperature signal of the corresponding time period is subjected to first derivative analysis to screen out the segments in which the heating rate is significantly higher than the conventional change threshold within the time delay window, and construct the acoustic-temperature linkage anomaly segment. Each anomalous sound-temperature linkage segment is represented as a three-dimensional feature tensor, with dimensions including spectral energy increase, temperature rise slope, and sound-temperature time difference. All segments are normalized and then input into the fusion computing engine. The fusion computing engine determines whether the segment constitutes a precursor to thermal runaway based on the set linkage judgment rules or training model. If it is determined to constitute a precursor, the corresponding channel identifier, risk level label and thermoacoustic composite risk factor field are added to the risk response signal.
Citation Information
Patent Citations
Big data-based energy storage power station battery health state diagnosis method and system
CN119147977A
Power equipment anomaly detection and early warning system and method
CN120127656A