A power battery fault diagnosis method and system, a terminal device, and a medium
Patent Information
- Application Number
- CN202611282913.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明要解决的技术问题在于,在动力电池故障诊断领域,真实车辆故障样本极为稀缺,现有方法大多仅能输出正常或异常的二值判断,无法进一步定位故障的具体类型及其对应的主导物理特征来源;同时,不同故障类型在早期共享异常偏离方向、后期才显现类型差异,传统单层表示难以有效捕捉并区分这种层级化演化结构
[0016]有益效果:本发明公开一种动力电池故障诊断方法、系统、终端设备及介质,涉及动力电池故障诊断技术领域。方法首先获取电池管理系统的实时监测数据,将所述监测数据划分为物理特征块,并基于各物理特征块构建监测窗口,其中,所述物理特征块的类型包括中心运行特征块、电压边缘特征块和温度边缘特征块。其后,将所述监测窗口输入预训练的确定性编码器,经前向映射得到全局潜在变量和局部潜在变量,并将二者拼接为联合潜在表示;其中,所述编码器为参数固定的确定性映射网络,对同一监测窗口始终输出相同的潜在变量。随后,计算所述全局潜在变量到预设正常基准点之间的距离作为异常分数,当所述异常分数超过预设阈值时,将所述监测窗口判定为异常窗口。针对所述异常窗口,提取经所述编码器映射得到的联合潜在表示,分别通过预设的故障原型集和预训练的分类器进行故障类型识别,并对两者的识别结果进行可靠度加权融合,得到故障类型定位结果。
Smart Images

Figure CN122815216A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power battery fault diagnosis technology, and in particular to a power battery fault diagnosis method, system, terminal equipment and medium. Background Technology
[0002] Lithium-ion power batteries have been widely used in new energy vehicles and energy storage systems, but safety faults such as internal short circuits, insulation abnormalities, and thermal runaway may occur during long-term operation.
[0003] In existing fault diagnosis methods, model-based methods rely on precise parameter identification and are computationally intensive; knowledge rule methods rely on manual thresholds and are difficult to cover complex and variable operating conditions; supervised data-driven methods require a large number of fault labels, while real vehicle fault samples are extremely scarce and severely imbalanced. Anomaly detection methods, represented by deep support vector data descriptions, can establish state boundaries using normal samples, but can only output binary anomaly alarms and cannot answer what type of fault it is or what physical feature it originates from. Furthermore, different fault types often deviate from the normal state along a common global anomaly direction in the early stages, only gradually differentiating into type-specific branches in later stages; traditional single-layer latent representations struggle to capture the evolutionary structure of these shared origins and divergent flows.
[0004] Therefore, there is an urgent need for a hierarchical diagnostic method that can maintain the boundaries of normal states and further locate anomalies to specific fault types and physical sources under the condition of scarce fault samples, so as to fill the gaps in existing technologies. Summary of the Invention
[0005] The technical problem this invention aims to solve is that, in the field of power battery fault diagnosis, real vehicle fault samples are extremely scarce. Most existing methods can only output a binary judgment of normal or abnormal, failing to further pinpoint the specific type of fault and its corresponding dominant physical characteristic source. Furthermore, different fault types share anomaly deviation direction in the early stages, only revealing type differences later; traditional single-layer representations struggle to effectively capture and distinguish this hierarchical evolutionary structure. Therefore, an effective solution is urgently needed to address these technical problems.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a method for diagnosing faults in a power battery, the method comprising: Real-time monitoring data from the battery management system is acquired, the monitoring data is divided into physical feature blocks, and a monitoring window is constructed based on each physical feature block; wherein, the types of physical feature blocks include center running feature blocks, voltage edge feature blocks, and temperature edge feature blocks; The monitoring window is input into a pre-trained deterministic encoder, and global and local latent variables are obtained through forward mapping. The two are then concatenated into a joint latent representation. The encoder is a deterministic mapping network with fixed parameters, which always outputs the same latent variables for the same monitoring window. The distance between the global latent variable and the preset normal baseline point is calculated as the anomaly score. When the anomaly score exceeds the preset threshold, the monitoring window is determined to be an anomaly window. For the abnormal window, the joint latent representation obtained by the encoder is extracted, and the fault type is identified by a preset fault prototype set and a pre-trained classifier, respectively. The identification results of the two are then weighted and fused based on reliability to obtain the fault type localization result.
[0007] In one implementation, dividing the monitoring data into physical feature blocks and constructing a monitoring window based on each physical feature block includes: The battery system operating parameters in the monitoring data are divided into the central operating feature block, the single cell voltage consistency index is divided into the voltage edge feature block, and the temperature consistency index is divided into the temperature edge feature block. The battery system operating parameters include at least current, voltage, and state of charge parameters; the cell voltage consistency index includes at least cell voltage range and cell voltage standard deviation; and the temperature consistency index includes at least temperature range and temperature rise rate. Denoising and normalization are performed on each physical feature block, and then low-dimensional feature vectors of each physical feature block are obtained by variance weighting and principal component dimensionality reduction. The low-dimensional feature vectors of each physical feature block at the same sampling time are concatenated to form the time feature; The time features of a preset number of consecutive sampling times are stacked in chronological order to form the monitoring window.
[0008] In one implementation, the deterministic encoder includes a shared backbone network, global branches, and local branches; The shared backbone network comprises multiple deep fully connected blocks connected in sequence. Each deep fully connected block consists of a fully connected layer, layer normalization, activation function, and proxy pulse gating. The global branch and the local branch each contain at least one fully connected layer and an activation function. The global branch outputs global latent variables of a preset dimension, and the local branch outputs local latent variables of a preset dimension. The output of the shared backbone network is fed into the global branch and the local branch respectively, so as to output the global latent variable and the local latent variable respectively.
[0009] In one implementation, calculating the distance between the global latent variable and a preset normal baseline point is used as an anomaly score. When the anomaly score exceeds a preset threshold, the monitoring window is determined to be an anomaly window, including: Calculate the Euclidean distance from the global latent variable to the normal reference point at each sampling time within the monitoring window, and use the maximum value of the distance at each time as the anomaly score; The preset threshold is determined based on the precision-recall curve of an independent validation set, or based on the high quantile of the abnormal score distribution of pure normal samples.
[0010] In one implementation, the fault type identification is performed using a preset fault prototype set and a pre-trained classifier, and the identification results from both are weighted and fused based on reliability to obtain the fault type localization result, including: The joint latent representation of the abnormal window obtained by the encoder is standardized using the one-dimensional mean and standard deviation of the preset support set; Calculate the distance from the standardized joint latent representation to each category of the fault prototype set, and determine the classification probability of each category based on the distance to obtain the first identification result; The standardized joint latent representation is input into the pre-trained classifier to obtain the probability distribution of each category as the second recognition result; The reliability of each is determined based on the distance interval of the first identification result and the probability entropy of the second identification result, respectively. The first identification result and the second identification result are then weighted and fused with a reliability normalization weight to obtain the fault type location result. Wherein, the support set is a set of joint latent representations obtained by mapping known fault type samples through the encoder; the fault prototype set is a set of cluster centers obtained by clustering the support set within each category; the fault type localization result includes at least one of fault type identifier, dominant physical feature block type, and diagnostic confidence.
[0011] In one implementation, the training steps of the deterministic encoder include: Obtain a training set consisting only of samples from the normal monitoring window; Construct a joint training architecture that includes the deterministic encoder and decoder; wherein the encoder maps the input normal monitoring window to global latent variables and local latent variables and concatenates them into a joint latent representation, and the decoder uses the joint latent representation as input to reconstruct the normal monitoring window; A joint loss function is constructed based on input reconstruction loss, recoding consistency loss, global single-class compactness loss, local diffusion smoothing loss, joint potential volume compression loss, and time-delay direction consistency loss, and the encoder and decoder are jointly trained. After training is complete, the encoder parameters are frozen, and the normal baseline is updated with the mean of the global latent variables of all normal training samples.
[0012] In one implementation, the steps of constructing the fault prototype set and the classifier include: After the deterministic encoder is trained and its parameters are frozen, samples of known fault types are collected, and the joint latent representation of each sample is extracted by the deterministic encoder and divided into support sets according to fault categories. The joint latent representation of the support set samples is standardized using the dimension-wise mean and standard deviation of the support set. Within each fault category, clustering is performed on the standardized joint latent representation, and the cluster centers within each category are used as fault prototypes for that category to form the fault prototype set. Based on the standardized support set used to construct the fault prototype set, the classifier is trained with the standardized joint latent representation as input and the known fault type as the label.
[0013] Secondly, embodiments of the present invention also provide a power battery fault diagnosis system, the system comprising: The monitoring window construction module is used to acquire real-time monitoring data of the battery management system, divide the monitoring data into physical feature blocks, and construct a monitoring window based on each physical feature block; wherein, the types of physical feature blocks include center running feature blocks, voltage edge feature blocks, and temperature edge feature blocks; The latent representation acquisition module is used to input the monitoring window into a pre-trained deterministic encoder, obtain global latent variables and local latent variables through forward mapping, and concatenate the two into a joint latent representation; wherein, the encoder is a deterministic mapping network with fixed parameters, which always outputs the same latent variables for the same monitoring window; An abnormal window judgment module is used to calculate the distance between the global latent variable and the preset normal benchmark point as an abnormal score. When the abnormal score exceeds the preset threshold, the monitoring window is judged as an abnormal window. The fault type localization module is used to extract the joint latent representation obtained by the encoder mapping for the abnormal window, identify the fault type through a preset fault prototype set and a pre-trained classifier, and perform reliability weighted fusion on the identification results of the two to obtain the fault type localization result.
[0014] Thirdly, embodiments of the present invention also provide a terminal device, the terminal device including a memory, a processor, and a power battery fault diagnosis program stored in the memory and executable on the processor, wherein when the processor executes the power battery fault diagnosis program, it implements the steps of the power battery fault diagnosis method described in any of the above schemes.
[0015] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a power battery fault diagnosis program, wherein when the power battery fault diagnosis program is executed by a processor, it implements the steps of the power battery fault diagnosis method described in any of the above schemes.
[0016] Beneficial Effects: This invention discloses a method, system, terminal device, and medium for diagnosing power battery faults, relating to the field of power battery fault diagnosis technology. The method first acquires real-time monitoring data from the battery management system, divides the monitoring data into physical feature blocks, and constructs a monitoring window based on each physical feature block. The types of physical feature blocks include central operating feature blocks, voltage edge feature blocks, and temperature edge feature blocks. Then, the monitoring window is input into a pre-trained deterministic encoder, which performs forward mapping to obtain global latent variables and local latent variables, and concatenates them into a joint latent representation. The encoder is a deterministic mapping network with fixed parameters, consistently outputting the same latent variables for the same monitoring window. Subsequently, the distance between the global latent variable and a preset normal reference point is calculated as an anomaly score. When the anomaly score exceeds a preset threshold, the monitoring window is determined to be an anomalous window. For the anomalous window, the joint latent representation obtained through the encoder mapping is extracted, and fault type identification is performed using a preset fault prototype set and a pre-trained classifier. The identification results from both are then weighted and fused based on reliability to obtain the fault type localization result.
[0017] This invention maps the original monitoring signal into a structured global and local joint latent representation through hierarchical physical feature block partitioning and deterministic latent representation. It establishes a compact normal state boundary using only normal samples pre-trained with the encoder, without relying on a large number of fault labels. Building upon anomaly detection, a dual-branch reliability-weighted fusion mechanism of the fault prototype set and classifier further locates the anomaly window to specific fault types, outputting the dominant physical feature block and diagnostic confidence score. This solves the technical bottleneck of existing methods that cannot locate faults in scenarios with scarce fault samples. Simultaneously, the fixed parameter characteristics of the deterministic mapping network ensure stable and reproducible model inference results, eliminating the need for random sampling and making it suitable for online deployment at both vehicle and edge computing. Attached Figure Description
[0018] Figure 1A flowchart illustrating a specific implementation of the power battery fault diagnosis method provided in this embodiment of the invention.
[0019] Figure 2 This is a schematic diagram of the overall process of the power battery fault diagnosis method provided in the embodiment of the present invention.
[0020] Figure 3 This is a schematic diagram of the hierarchical representation structure of the observation layer, physical block layer, potential structure layer, and event fault layer.
[0021] Figure 4 This is a flowchart illustrating the process of locating specific fault types based on multiple prototypes and MLP decoders after anomaly screening in the method of this embodiment of the invention.
[0022] Figure 5 This is a schematic diagram illustrating the overall diagnostic performance of the method in an embodiment of the present invention.
[0023] Figure 6 This is a schematic diagram of the power battery fault diagnosis system provided in an embodiment of the present invention.
[0024] Figure 7 This is a block diagram illustrating the internal structure of the terminal device provided in an embodiment of the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0026] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content, operations, or steps, nor does it require execution in the described order. For example, some operations or steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0027] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0028] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. For example, "first control information" and "second control information" are only used to distinguish different control information and do not limit their order.
[0029] Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or the order of execution, and that the words "first" and "second" do not necessarily imply that they are different.
[0030] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0031] Lithium-ion power batteries have advantages such as high energy density, long cycle life, and low maintenance costs, and have been widely used in new energy vehicles and energy storage systems. However, during long-term operation, batteries may experience internal short circuits, external short circuits, insulation abnormalities, abnormal aging, thermal abnormalities, and complex safety faults caused by the coupling of multiple factors. Due to the scarcity of real vehicle fault samples, random switching of operating conditions, and noise and missing monitoring signals, online fault diagnosis of power batteries still faces significant challenges.
[0032] Existing methods for diagnosing power battery faults mainly include model-based methods, data-driven methods, knowledge rule-based methods, and hybrid methods. Model-based methods rely on equivalent circuit models or electrochemical models and usually require accurate parameter identification; knowledge rule-based methods rely on manual thresholds and empirical rules, which are difficult to cover complex operating conditions; supervised data-driven methods require a large number of fault labels, while real vehicle-side fault labels are usually scarce and of uneven categories.
[0033] One type of anomaly detection and deep support vector data description method can establish normal state boundaries using normal operation samples and output anomaly scores based on the distance of the sample to the normal center (also known as the normal reference point). This type of method is suitable for scenarios where fault samples are scarce, but if only binary anomaly alarms are output, it is still difficult to answer which type of safety event, which type of physical feature block, and whether different fault types have common precursors or specific branches.
[0034] Existing global and local latent representation methods can improve anomaly detection capabilities to some extent by leveraging global compactness and local dynamics. However, in engineering diagnostics, simply using global boundaries and local trajectories as discriminative features at the same level is still insufficient. Power battery safety faults have a natural hierarchical structure: the bottom layer consists of observed signals such as voltage, current, temperature, SOC (State of Charge, i.e., the percentage of remaining battery capacity relative to rated capacity), and individual cell channels; the middle layer consists of physical blocks such as central operating characteristics, voltage edge characteristics, and temperature edge characteristics; the next layer consists of global deviations, local evolutions, and joint latent structures; and the top layer contains fault events, fault stages, and specific fault types.
[0035] Therefore, a new online fault diagnosis method for power batteries is needed. While retaining the ability to screen for anomalies at the normal state boundary, a hierarchical potential diagnostic representation should be constructed. After the abnormal sample enters the second stage, a lightweight fault type decoder that does not change the backbone representation should be used to locate the anomaly to the specific fault type and related physical feature block, thereby improving the interpretability and engineering usability of the diagnostic results.
[0036] This embodiment provides a method for diagnosing power battery faults, such as... Figure 1 As shown, the specific steps include the following: Step S100: Obtain real-time monitoring data from the battery management system, divide the monitoring data into physical feature blocks, and construct a monitoring window based on each physical feature block; wherein, the types of physical feature blocks include center running feature blocks, voltage edge feature blocks, and temperature edge feature blocks.
[0037] In this embodiment, the battery management system monitoring data is the full-dimensional sensor values of battery operation continuously collected by the battery management controller according to a fixed sampling period. Specifically, it can be a complete time-series dataset output by the new energy vehicle's on-board BMS (Battery Management System) every ten seconds, including total current, total voltage, remaining charge, voltage of each individual cell, temperature of each temperature probe, insulation resistance value, vehicle mileage and speed, etc. Each record in the dataset corresponds to all sensor readings at a single sampling time.
[0038] Physical feature blocks are feature sets categorized according to the battery physical failure mechanism and monitoring level corresponding to the sensing parameters. Based on the hierarchical differences of the monitoring objects, they are divided into three independent sets: central operating feature blocks that reflect the overall macroscopic operating conditions of the battery pack, voltage edge feature blocks that characterize the consistency deviation between cells, and temperature edge feature blocks that characterize the uniformity of the thermal field distribution. The three sets are independent of each other and there is no cross-reuse of parameters.
[0039] The central operating feature block contains all parameters that reflect the overall operating status of the battery system. Changes in these parameters directly correspond to macroscopic fluctuations in battery charging and discharging power and remaining capacity. Specifically, these parameters can include total circuit current, total battery voltage, state of charge value, vehicle speed, cumulative mileage, and insulation resistance monitoring value. These parameters do not reflect local anomalies in a single cell or a single-point temperature probe; they only reflect the overall operating level of the battery pack.
[0040] The voltage edge feature block contains derived indicators used to quantify the degree of voltage imbalance between cells. Abnormal indicator values indicate that some cells have polarization drift, cell aging, or potential local short circuits. Specifically, these indicators can be the range of voltage of all individual cells at the same time, the standard deviation of individual cell voltage distribution, or the extreme value of the rate of change of individual cell voltage within a continuous sampling interval. Based solely on these indicators, early signs of cell-level voltage deviation faults can be detected.
[0041] Temperature edge feature blocks contain derived indicators used to quantify the uniformity of the internal temperature field of the battery pack. Abnormal indicator values indicate local overheating, blockage of cooling channels, and early signs of thermal runaway of the battery cell. Specifically, these indicators can be the range of temperatures collected by all temperature probes, the standard deviation of temperature distribution, and the maximum single-point temperature rise rate within a continuous period. These indicators can locate the temperature measurement point range corresponding to thermal anomalies.
[0042] The monitoring window is a temporal feature matrix formed by stacking multiple sets of standardized low-dimensional features of continuous moments in chronological order. It is used to provide input samples with temporal correlation information to the subsequent network. The window contains a set number of complete feature vectors of continuous sampling moments. The window length can be flexibly adjusted according to the early warning amount. Specifically, 60 continuous sampling frames can be set as a monitoring window, corresponding to a continuous monitoring time series of 10 minutes.
[0043] By splitting the original monitoring data according to the physical failure dimension to construct a hierarchical feature structure, the physical semantic hierarchical isolation of the original sensing signals is achieved, avoiding global operating condition fluctuations from masking local anomalies in the battery cell and thermal field. This provides an input basis with clear physical interpretation for the extraction of hierarchical potential characterizations. At the same time, the stacking of time-series windows preserves the time-series characteristics of fault evolution, making up for the deficiency that single-moment sampling cannot capture gradual early faults.
[0044] In one implementation, dividing the monitoring data into physical feature blocks and constructing a monitoring window based on each physical feature block specifically includes the following steps: Step S110: The battery system operating parameters in the monitoring data are divided into the central operating feature block, the single cell voltage consistency index is divided into the voltage edge feature block, and the temperature consistency index is divided into the temperature edge feature block. The battery system operating parameters include at least current, voltage, and state of charge parameters; the cell voltage consistency index includes at least cell voltage range and cell voltage standard deviation; and the temperature consistency index includes at least temperature range and temperature rise rate. Step S120: Perform denoising and normalization processing on each physical feature block, and then obtain the low-dimensional feature vector of each physical feature block through variance weighting and principal component dimensionality reduction. Step S130: Concatenate the low-dimensional feature vectors of each physical feature block at the same sampling time to form a time feature; Step S140: Stack the time features of a preset number of consecutive sampling times in chronological order to form the monitoring window.
[0045] In this embodiment, the original monitoring parameters are first categorized into blocks. Specifically, all collected sensor values are divided into three independent sets according to their physical meaning. Parameters categorized into the central operating feature block include total circuit current, total battery voltage, state of charge, vehicle speed, cumulative mileage, and insulation resistance. Indicators categorized into the voltage edge feature block include single-cell voltage range, single-cell voltage standard deviation, and extreme values of single-cell voltage change rate. Indicators categorized into the temperature edge feature block include temperature range, temperature standard deviation, and maximum single-point temperature rise rate. The alarm level and comprehensive alarm flag fields generated by the BMS built-in alarm rules do not participate in the feature block process and are only used for fault labeling and diagnostic result verification. The time field is only used for time-series sorting and window splicing and is not used as a numerical feature in subsequent dimensionality reduction calculations.
[0046] Subsequently, a unified preprocessing transformation was performed on the three types of physical feature blocks, in the following order: robust denoising, feature normalization, variance-weighted mapping, and principal component dimensionality reduction. Robust denoising uses the quartile method to remove extreme outlier sampled values exceeding three times the interquartile range, avoiding instantaneous noise caused by vehicle bumps and electromagnetic interference. Normalization uses min-max standardization to map all features to the 0-1 interval, eliminating differences in the dimensions of different parameters. Variance-weighted mapping assigns feature weights according to the variance of a single feature in the dataset, amplifying the contribution of indicators with significant fluctuations and high fault sensitivity. Principal component dimensionality reduction retains principal components with a cumulative variance of up to 95%, compresses redundant feature dimensions, and generates low-dimensional feature vectors for each type of physical feature block.
[0047] Subsequently, the composite feature splicing at a single moment is completed. Specifically, the central low-dimensional vector, voltage edge low-dimensional vector, and temperature edge low-dimensional vector, which have undergone dimensionality reduction processing at the same sampling moment, are taken and spliced end to end in a fixed order to form a moment-specific feature vector. The three segments inside the vector correspond to the compressed features of the three types of physical blocks, respectively, preserving the hierarchical physical semantics.
[0048] Finally, the stacked timing window is executed. The time feature vectors generated from a fixed number of consecutive sampling times are selected. The samples are stacked vertically in order of sampling time from earliest to latest, forming a two-dimensional time-series matrix of monitoring windows. The window length can be adjusted according to the application scenario, such as... Figure 2 As shown, the hierarchical feature segmentation and temporal windowing construct the corresponding hierarchical representation system, transforming the observation layer into the physical block layer and providing structured input for potential representation extraction. Specifically, continuous features can be used to construct the physical block layer. The features at each moment are stacked in chronological order to obtain the first... One monitoring window .
[0049] Layered preprocessing is used to individually adapt the numerical distribution characteristics of three types of parameters: voltage, temperature, and system operation. Variance weighting and principal component dimensionality reduction filter out irrelevant redundant information, and time-series window stacking preserves the gradual time-series characteristics of faults. Compared with the processing method of directly splicing the original sensor data, this method can reduce the interference of irrelevant operating condition noise on the diagnostic model and improve the identification of early weak fault characteristics.
[0050] This paper demonstrates the construction and hierarchical feature organization of real vehicle data through a specific embodiment. Specifically, it uses real operational data from 20 new energy vehicles, containing approximately 8.6 million valid data points. Each data frame is sampled at a 10-second interval, and the number of outlier samples is approximately 3204, accounting for approximately 0.037% of all samples. Eight data tables covering approximately 95% of the outlier samples are selected as the main research objects, and a separate summary dataset is constructed.
[0051] The following data are selected from the original battery management system records: charging status, vehicle speed, cumulative mileage, total voltage, total current, state of charge, insulation resistance, voltage extremes, temperature extremes, extreme location, individual cell voltage channels, and temperature probe channels. Alarm levels and comprehensive alarm flags are only used for label generation and result verification, and are excluded from all input feature processing, normalization, variance weighting, dimensionality reduction, training, and anomaly scoring processes. The time field is used for sorting and window construction and is not directly used as a static numerical feature in dimensionality reduction.
[0052] like Figure 2 and Figure 3 As shown, the original monitoring signal is first converted into three physical feature blocks: central operation, voltage edge, and temperature edge. Based on these, a hierarchical diagnostic representation is constructed. This structure enables the model to both quickly screen for anomalies using global normal boundaries and output specific fault types and dominant physical blocks at the fault event layer.
[0053] Step S200: Input the monitoring window into a pre-trained deterministic encoder, obtain global latent variables and local latent variables through forward mapping, and concatenate the two into a joint latent representation; wherein, the encoder is a deterministic mapping network with fixed parameters, and always outputs the same latent variables for the same monitoring window.
[0054] In this embodiment, the standardized time-series matrix formed by the preprocessing of the monitoring window is completely fed into the pre-trained deterministic encoder. The encoder is a deep neural network structure that completes a fixed mapping relationship. All weight parameters of the network are frozen after training. There are no random sampling, random noise injection, or reparameterization perturbation modules in the inference stage. After the same set of input monitoring windows is fed into the network, no matter how many forward operations are repeated, the final output potential vector values are completely consistent. It has the engineering characteristics of reproducible results and stable inference, and is suitable for low-computing-power online deployment scenarios such as vehicle controllers and edge gateways.
[0055] The global latent variable is a low-dimensional vector output by the encoder's global branch. The vector dimension is used to characterize the macroscopic deviation of the overall battery operating state from the standard normal range. The global latent variable only extracts the overall operating condition deviation features shared by all time-series samples within the monitoring window, weakening the interference caused by local fluctuations at a single moment. Specifically, a four-dimensional vector can be set to carry the global state representation. Each dimension of the vector corresponds to four indicators: charge and discharge power deviation, capacity decay trend, overall insulation level, and long-term aging accumulation degree.
[0056] Local latent variables are low-dimensional vectors output by local branches of the encoder. The vector dimension is used to capture the details of temporal evolution within the window, abrupt changes in local features at a single moment, and temporal coupling changes between different physical feature blocks. It focuses on retaining the instantaneous abnormal fluctuation information brought about by voltage and temperature edge features. Specifically, an eight-dimensional vector can be set to carry the representation of local temporal evolution. Each dimension of the vector corresponds to subdivided temporal features such as single-unit voltage temporal drift, local temperature rise dynamics, and gradient of feature changes at multiple moments.
[0057] Joint latent representation is a high-dimensional composite vector formed by directly concatenating the global latent variables and local latent variables output at the same sampling time according to the feature dimension order. It integrates macroscopic overall offset information and microscopic temporal evolution information and serves as a standardized input feature for anomaly detection and fault classification. Specifically, a 4-dimensional global vector and an 8-dimensional local vector can be concatenated to obtain a 12-dimensional fixed-length joint latent vector. The first 4 dimensions of the vector carry global offset information, and the last 8 dimensions carry local temporal evolution information.
[0058] The encoder consists of three interconnected parts: a shared backbone network, a global output branch, and a local output branch. The shared backbone network completes the basic feature extraction of the original temporal window. After extraction, the backbone output is simultaneously split and sent to two independent branches to generate two types of latent vectors. The complete structural configuration of the deep fully connected blocks, gating units, activation functions, and layer normalization modules inside the network is described separately and will not be repeated in this section.
[0059] By using a global and local dual-branch deterministic mapping network, the two types of features, namely the overall macroscopic offset of the battery and the temporal local anomaly, are separated simultaneously. The deterministic network eliminates random jitter in inference, ensuring the stability of the on-board online diagnostic output. The hierarchical potential representation distinguishes between global fault precursors and local specific evolutions, thus solving the technical defect that single-layer features cannot distinguish the shared early abnormal trends of multiple faults.
[0060] In one implementation, the deterministic encoder includes a shared backbone network, global branches, and local branches; The shared backbone network comprises multiple deep fully connected blocks connected in sequence. Each deep fully connected block consists of a fully connected layer, layer normalization, activation function, and proxy pulse gating. The global branch and the local branch each contain at least one fully connected layer and an activation function. The global branch outputs global latent variables of a preset dimension, and the local branch outputs local latent variables of a preset dimension. The output of the shared backbone network is fed into the global branch and the local branch respectively, so as to output the global latent variable and the local latent variable respectively.
[0061] In this embodiment, the deterministic encoder is divided into three parts: a shared backbone network, a global output branch, and a local output branch. The three parts are connected in series to complete the feature mapping. The output features of the shared backbone network are simultaneously split and sent to two independent branches to generate global latent variables and local latent variables, respectively.
[0062] Specifically, the deterministic encoder refers to a fixed mapping that does not contain random latent variable sampling, reparameterization sampling, or random perturbations during inference. With fixed preprocessing and network parameters, the same input yields unique global and local latent variables. In one embodiment, a time-progressive feedforward spiking neural-inspired encoder is used instead of LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), or Transformer.
[0063] The encoder backbone consists of four deep pulse-inspired blocks, each composed of a fully connected layer, layer normalization, sigmoid activation, surrogate pulse gating with a threshold of 0.5, and Dropout (random deactivation). The hidden layer widths are 768, 512, 512, and 256, respectively, and the Dropout rate is 0.05. The surrogate pulses are only used to modulate continuous activations and train via surrogate gradients, without random sampling.
[0064] The global output branches share the output features of the backbone network, and finally output a four-dimensional global latent variable. Each dimension of the vector represents the overall operational offset of the battery, and only macroscopic state features shared by the window time series are extracted. The local output branches maintain the same hierarchical structure as the global branches, and finally output an eight-dimensional local latent variable. The vector fully captures the temporal dynamics within the window, the sudden changes in local features at a single moment, and the details of the coupling changes of multiple physical blocks.
[0065] The main output is fed into the global header and local header respectively. Each header contains, in sequence, a 256-dimensional fully connected layer, layer normalization, a sigmoid function, a latent output layer, and another sigmoid function. Specifically, it is represented as follows:
[0066]
[0067]
[0068]
[0069] In the formula, Assign a monitoring window number, Number the time within the window. For window length, for 3D physical block splicing characteristics and These are the global output header and the local output header after sharing the main trunk, respectively. As a global latent variable, For local latent variables, the semicolon indicates concatenation along the feature dimension. The joint latent variable is formed by concatenating the global latent variable and the local latent variable. In one embodiment, , Joint latent dimension .
[0070] The backbone network output features are undifferentiatedly split into global and local branches. These two types of branches independently perform mapping operations without interference. The global branch weakens instantaneous local fluctuations, while the local branch amplifies subtle temporal changes in voltage and temperature edge features. The two types of vectors are concatenated to form a complete joint latent representation, such as... Figure 3 As shown, the global and local vectors output by the encoder correspond to the potential structural layers of the hierarchical representation system, carrying two types of information: global normal boundary and local time trajectory, respectively.
[0071] In addition, decoder Using joint latent variables as input, The fully connected structure is used; except for the last layer, each layer sequentially employs layer normalization, LeakyReLU (slope of negative half axis 0.2), and Dropout of 0.05 to reconstruct the original structure. 3D time-series features. The encoder structure is responsible for forward mapping, but does not yet involve loss weighting and parameter updates.
[0072] Multi-layer pulse-inspired fully connected blocks enhance temporal feature extraction capabilities, dual-branch separation of macroscopic and microscopic feature mapping, deterministic non-random network structure ensures stable and reproducible online inference results for vehicles, and lightweight fully connected architecture adapts to the deployment requirements of low-computing-power vehicle controllers.
[0073] In one implementation, the training steps of the deterministic encoder specifically include: Step S201: Obtain a training set consisting only of samples from the normal monitoring window; Step S202: Construct a joint training architecture including the deterministic encoder and decoder; wherein, the encoder maps the input normal monitoring window to global latent variables and local latent variables and concatenates them into a joint latent representation, and the decoder uses the joint latent representation as input to reconstruct the normal monitoring window; Step S203: Construct a joint loss function based on input reconstruction loss, recoding consistency loss, global single-class compactness loss, local diffusion smoothing loss, joint potential volume compression loss and time delay direction consistency loss, and jointly train the encoder and decoder; Step S204: After training is completed, freeze the parameters of the encoder and update the normal reference point with the mean of the global latent variables of all normal training samples.
[0074] In this embodiment, the encoder and decoder are jointly trained using a normal monitoring window. The parameters of the encoder and decoder are updated using six types of constraints: input reconstruction, recoding consistency, global single-class compactness, local diffusion smoothing, joint latent volume compression, and time delay direction consistency. After training, the preprocessing mapping, encoder parameters, normal center, support set normalization parameters, and alarm threshold are fixed. During the online phase, the decoder is no longer invoked or backpropagation is performed; instead, a stable hierarchical latent representation is directly generated by freezing the encoder.
[0075] During the training phase, six independent constrained loss functions are constructed. All losses are superimposed and weighted to form a total loss, which is used for backpropagation to update network parameters. The six types of losses are described in detail below.
[0076] The input reconstruction loss, which constrains the decoder to reconstruct samples that are consistent with the original input, is expressed as:
[0077] In the formula, This represents the number of monitoring windows in a small batch. To reconstruct features for the decoder, Indicates decoder, It is a norm 2. Represents the number of sampling times per window. Features of the original time period Mean square error between compressed input and reconstructed input.
[0078] The recoding consistency loss, which constrains the re-encoded latent vector of the reconstructed sample to be unbiased from the original vector, is expressed as:
[0079] In the formula, To reconstruct the joint latent vector obtained by re-encoding the samples, This represents an encoder consisting of a global header and a local header. This indicates that the gradient operation is stopped, preventing the gradient from being backpropagated to the encoder backbone. The constrained recoding results are consistent with the original joint latent variables.
[0080] The global single-class compact loss constrains the global latent vector of normal samples to shrink towards the normal baseline center, and is expressed as:
[0081] In the formula, These are normal centers estimated solely from the global latent variables of the normal training samples. By employing the unsquared Euclidean norm, the global latent representation of normal samples is compacted towards the normal center, and the dominance of gradient by extreme samples is reduced, thereby enhancing the latent space clustering of normal samples.
[0082] The local diffusion smoothing loss, which constrains the local latent vector nearest neighbor density to be continuous without abrupt changes, is expressed as:
[0083]
[0084] In the formula, The number of joint potential vertices participating in graph construction within a mini-batch. For the first Potential points Nearest neighbor set For Gaussian similarity, For similarity bandwidth, For the first Local density estimation at points, is the diffusion coefficient. Local spikes, discrete fractures, and discontinuous collapses are suppressed by penalizing the density difference between adjacent potential points. In one embodiment, [the following is taken] , , .
[0085] The joint latent volume compression loss, which constrains normal samples to occupy a smaller volume in the joint latent space, is expressed as:
[0086]
[0087] In the formula, To extend to the center of the joint potential space, for Zero-dimensional vector For joint potential points To the center radial distance, Radial weight, Radial bandwidth, For the maximum radius of the small batch, To unify the latent dimension, For gamma function, for dimensional radius Hypersphere volume, This is the internal scaling factor for the volume term. Simultaneously penalizing latent points far from the center and those with excessively large occupancy radii, the compression is of the weighted radial occupancy volume of normal samples in the joint latent space, rather than the input physical space volume. In one embodiment, taking... , , .
[0088] The time-delay direction consistency loss, constraining the local potential temporal motion direction to match the time-delay resultant force, is expressed as:
[0089]
[0090]
[0091]
[0092] In the formula, For local potential trajectories from arrive The actual motion vector, For discrete time delay, For the current local latent variables and The difference of latent variables before time 1. To prevent the softening constant from having a denominator of zero, The time-delay direction vector, For Gaussian action gate, For the operating bandwidth, As the weight decays exponentially with time lag, For amplitude, It is the attenuation constant. The resultant force is in the direction of time delay. and They are respectively and The direction of unitization. The inconsistency between the actual local motion direction and the time-delayed resultant force direction is measured using a 1 minus cosine similarity. In one embodiment, we take... , , , , .
[0093] Based on the above six loss functions, the overall joint loss function adaptively balances the weights of the six types of constraints, and is expressed as follows:
[0094] In the formula, The main training losses , and It is a learnable logarithmic scaling parameter used for three purposes: adaptive balanced reconstruction, global single-class and recoding consistency. , and These represent the constraint weights for diffusion, volume, and time delay direction, respectively. In one embodiment, we take... , , .
[0095] After all loss iterations converge, the encoder is permanently frozen, all parameters are preprocessed and mapped, and the normal baseline center is updated using the mean of the global latent vectors of all normal training samples. The decoder only participates in constraint construction during the training phase, and the online diagnostic inference process no longer calls the decoder module.
[0096] A specific implementation demonstrates the training of the frozen latent representation backbone. Robust denoising, normalization, variance weighting, and principal component analysis are performed on the three physical feature blocks respectively. The principal components, retaining the main variance information, are concatenated into a unified low-dimensional feature vector. A monitoring window is formed using 60 consecutive data frames, with a training window stride of 5, a batch size of 128, 50 backbone training epochs, and an initial learning rate of... When the sampling interval is 10 seconds, it corresponds to a 10-minute local monitoring period. Under other deployment conditions, the length of the monitoring window can be adjusted according to the alarm lead time and computing resources, but training, validation, and deployment must use the same window definition.
[0097] The deterministic encoder described in the normal monitoring window has a global branch dimension of 4, used for normal centering and anomaly scoring, and a local branch dimension of 8, used for local temporal direction constraints. These two are concatenated to form a 12-dimensional joint latent representation. The encoder and decoder are trained based on the listed total loss. After each training round, the global latent mean of all normal training samples is used. Update normal center After training, the preprocessing parameters, encoder, support set normalization parameters, and normal centers are fixed; the decoder is only used to reconstruct constraint branches during training and is not called during online inference.
[0098] Step S300: Calculate the distance between the global latent variable and the preset normal reference point as the anomaly score. When the anomaly score exceeds the preset threshold, the monitoring window is determined to be an anomaly window.
[0099] In this embodiment, the global latent variable normal reference point is the center vector obtained by statistically analyzing a large number of global latent vectors containing only samples of normal operating conditions. The value of each dimension of the vector is equal to the arithmetic mean of the corresponding dimension of all normal training samples. The reference point represents the standard latent space coordinates of the battery when it is running smoothly without any faults. All vectors that deviate from this coordinate have varying degrees of abnormal risk. The reference point value is frozen after the encoder completes training and is no longer updated during the diagnostic phase.
[0100] Anomaly score is a scalar value that quantifies the degree of overall abnormal deviation of a single monitoring window. The value corresponds to the severity of the battery's deviation from the standard normal operating state. The higher the value, the more significant the potential fault. The anomaly score is calculated based on the spatial distance between the global potential vector and the normal benchmark point. It only relies on the global representation to complete the first-level rapid screening without calling local time-series features, thus reducing the computing power overhead of online inference.
[0101] The preset threshold is a fixed numerical boundary line that has been calibrated in advance. It is used to distinguish between normal monitoring windows and abnormal monitoring windows with potential faults. The threshold is divided into multiple warning intervals, and different intervals correspond to different fault warning levels. The threshold calibration is completed based on the verification dataset. After calibration, it is stored in the diagnostic terminal and does not change dynamically with the operating conditions during the operation phase.
[0102] An anomaly window is a time-series monitoring window where the anomaly score calculation result exceeds a preset threshold. After the judgment is completed, a secondary fault location process is executed to carry out detailed fault type identification. Monitoring windows that do not exceed the threshold are judged as normal data, and the data is stored and archived without entering the subsequent fault identification process. This enables rapid filtering of normal samples and reduces the continuous computing power occupation of the detailed fault identification module.
[0103] First-level anomaly screening is performed by using global potential vector space distance, lightweight and fast anomaly discrimination is performed based on low-dimensional global vectors, massive normal samples are filtered in advance, and a small number of suspected anomaly samples are sent to complex fault location branches, which greatly reduces the continuous inference computation of vehicle edge devices. At the same time, a benchmark center is generated based on normal sample statistics, and anomaly boundary construction can be completed based on normal data, which is suitable for the engineering status of scarce fault samples in real scenarios.
[0104] In one implementation, the step of calculating the distance between the global latent variable and a preset normal baseline point is used as an anomaly score. When the anomaly score exceeds a preset threshold, the monitoring window is determined to be an anomaly window. This specifically includes the following steps: Step S310: Calculate the Euclidean distance from the global latent variable to the normal reference point at each sampling time within the monitoring window, and use the maximum value of the distance at each time as the anomaly score; The preset threshold is determined based on the precision-recall curve of an independent validation set, or based on the high quantile of the abnormal score distribution of pure normal samples.
[0105] In this embodiment, during online operation, the current monitoring window is input into the freeze encoder, and the anomaly score is calculated solely based on the distance between the global latent variable and the normal center. The maximum, mean, or final value of the distance within the window can be used as the window's anomaly score, and multiple alarm thresholds can be set based on the validation set precision-recall curve or the high quantile of the normal distribution. When the anomaly score exceeds the threshold, the anomaly level is output, and the corresponding window is sent to the second-stage fault type localization module.
[0106] Specifically, the formula for calculating the distance of the global latent variables at a single time moment is expressed as:
[0107] In the formula, Representing the The first group monitoring window The single-point anomaly score corresponding to each sampling time. Representing the Group window The four-dimensional global latent variable output at each time step. This represents the four-dimensional normal baseline center vector obtained from the statistics of all normal samples. This represents the vector 2-norm operation, used to calculate the Euclidean linear distance between two sets of vectors in the latent space.
[0108] The comprehensive anomaly score for a single monitoring window is the maximum value of the distance between individual points at all sampling times within the window. The calculation formula is as follows:
[0109] In the formula, Representing the The group monitoring window is ultimately used to determine the window-level anomaly score. This indicates the total number of consecutive sampling times contained in a single monitoring window. The calculation is used to extract the time distance value that deviates from the normal baseline within the window, and the most severe offset point represents the degree of abnormality of the entire group of windows.
[0110] The exception determination rule is when The value is greater than the preset threshold When this occurs, the current window is marked as an abnormal window, and the process is transferred to the secondary fault location procedure, as shown below:
[0111] threshold There are two types of calibration implementation paths. The first type relies on the precision-recall curve of an independent validation dataset and selects the distance value of the optimal F1 index corresponding to the curve as a fixed threshold. The second type uses only a validation set composed of pure normal samples, calculates the distribution of abnormal scores in the window of all normal samples, and selects the distance value corresponding to the preset high quantile as the threshold. Either calibration method can be flexibly selected according to the strictness of the project's fault warning. Figure 2 It demonstrates the first-level anomaly screening step of the overall diagnostic process, which completes the filtering of normal samples.
[0112] The maximum distance of the window is used as the anomaly score to accurately capture instantaneous sudden faults within the window. Two threshold calibration methods are adapted to the false alarm and missed alarm control requirements of different projects. Moreover, the calculation is completed only based on a four-dimensional global vector, with extremely low computational load, and millisecond-level online anomaly screening can be achieved.
[0113] Step S400: For the abnormal window, extract the joint latent representation obtained by the encoder mapping, identify the fault type through a preset fault prototype set and a pre-trained classifier, and perform reliability weighted fusion on the identification results of the two to obtain the fault type localization result.
[0114] In this embodiment, the joint latent representation corresponding to the abnormal window is a standardized composite vector output by the encoder after freezing. The vector fully retains the dual features of global macroscopic offset and local temporal evolution, and carries the hierarchical abnormal information corresponding to the three types of physical feature blocks. It is a unified and standardized input carrier for subdivided fault identification. All fault prototypes and classifiers are trained based on the joint latent vector of the same dimension, ensuring the uniformity of input features of the two-level identification branches.
[0115] The fault prototype set is a set of cluster centers generated by clustering all known fault type samples. The set stores multiple sets of prototype vectors corresponding to each type of fault in groups according to the fault category. A single fault can correspond to one to three sets of independent prototype vectors, which can adapt to the multi-cluster non-spherical distribution characteristics formed by the same fault at different evolution stages and under different working conditions. The prototype vectors can be generated based on a small number of labeled fault samples, without the need for a large-scale fault labeling dataset.
[0116] The pre-trained classifier is a lightweight multilayer perceptron network that completes supervised training based on standardized joint latent vectors. The network input dimension is consistent with the joint latent vector dimension. The output layer outputs the normalized probability distributions corresponding to various known fault types. The classifier only operates on the frozen latent representation space and will not update the front-end encoder backbone network in reverse, thus protecting the normal state boundary formed by the first-level anomaly screening from being destroyed.
[0117] Fault type identification is a fault category prediction process based on two independent branches: distance matching of fault prototype sets and probability output of classifiers. The two branches use completely independent identification logic. The prototype branch completes matching based on the geometric distance of the latent space, while the classifier branch completes probability prediction based on the nonlinear mapping of the network. The output results of the two branches have complementary characteristics, which makes up for the discrimination defects of a single identification branch.
[0118] Reliability-weighted fusion generates differentiated weights based on the confidence levels of the outputs of the two identification branches. It then performs a weighted summation operation on the fault category probabilities output by the two branches to generate a unified fusion fault probability distribution. This eliminates the risk of misjudgment caused by low-confidence predictions from a single branch, and finally outputs standardized fault location information.
[0119] The fault type localization result is a standardized diagnostic information set output after weighted fusion calculation. The information set contains at least three items: the type identifier corresponding to the fault, the category of the dominant physical feature block that caused the fault, and the confidence value corresponding to the current diagnostic result. It can be directly read by the vehicle controller and cloud monitoring platform for fault alarm, safety protection strategy triggering, and maintenance tracing analysis.
[0120] By combining a prototype set with a multilayer perceptron using a weighted fusion of two branches, this method is adapted to scenarios where real vehicle fault samples are scarce. A small number of fault samples can be used to construct a prototype set for coarse classification. A lightweight classifier supplements the nonlinear discrimination capability, and the two-branch fusion reduces the probability of misidentification by a single branch. At the same time, it outputs the dominant physical feature block corresponding to the fault, providing an interpretable physical basis for fault tracing. This solves the technical shortcomings of existing anomaly detection methods, which can only output binary alarms and cannot locate specific fault types and anomaly sources.
[0121] In one implementation, the fault type identification is performed using a preset fault prototype set and a pre-trained classifier, and the identification results from both are weighted and fused based on reliability to obtain the fault type localization result. This specifically includes the following steps: Step S410: Standardize the joint latent representation of the abnormal window obtained by the encoder mapping using the mean and standard deviation of the preset support set. Step S420: Calculate the distance from the standardized joint latent representation to each category of the fault prototype set, and determine the classification probability of each category based on the distance to obtain the first identification result; Step S430: Input the standardized joint latent representation into the pre-trained classifier to obtain the probability distribution of each category as the second recognition result; Step S440: Determine the reliability of each based on the distance interval of the first identification result and the probability entropy of the second identification result respectively, and perform weighted fusion of the first identification result and the second identification result with reliability normalization weight to obtain the fault type location result; Wherein, the support set is a set of joint latent representations obtained by mapping known fault type samples through the encoder; the fault prototype set is a set of cluster centers obtained by clustering the support set within each category; the fault type localization result includes at least one of fault type identifier, dominant physical feature block type, and diagnostic confidence.
[0122] In this embodiment, after the anomaly window completes the first-level anomaly screening, it performs all the online real-time calculations for the second-level fault location. Specifically, it is based on the parallel computation of a fault prototype set that has been pre-trained and parameterized in the offline stage and a multilayer perceptron dual-branch model. There is no network backpropagation or parameter update operation. The entire set of computational logic is as follows: Figure 4 As shown.
[0123] In the secondary fault localization process, a standardization transformation operation is first performed. The purpose of standardization is to eliminate dimensional differences between different feature dimensions and avoid the interference of scale bias on spatial distance calculation and classification probability output. The operation targets the joint latent representation of the anomaly window obtained after the frozen deterministic encoder mapping. The joint latent representation of the anomaly window obtained by the encoder mapping is standardized using the dimensional mean and standard deviation of a preset support set. Here, the support set refers to the sample set formed by collecting all samples with explicit fault labels after the encoder is frozen, extracting the joint latent representation, and grouping them according to fault category. The standardization operation expression is:
[0124] In the formula, and These are the mean and standard deviation of all supporting samples, respectively. It is the numerical stability constant. This represents the joint latent vector after standardization. This represents the original joint latent vector directly output by the encoder.
[0125] After standardization, the fault prototype branch operation is initiated. This branch calculates the distance from the standardized joint latent representation to each category of prototype in the fault prototype set, and determines the classification probability based on these distances, yielding the first identification result. The fault prototypes are the cluster centers obtained from the offline stage clustering of standardized fault samples for each category. The total number of prototypes generated by clustering within a class of faults is denoted as . , represented as:
[0126] In the formula, For the first The class supports a certain number of samples. The minimum number of supported samples allowed for each prototype. The maximum number of prototypes allowed for a single fault type. For the first The actual number of class prototypes.
[0127] No. The first type of fault corresponding to the The standardized prototype vectors are denoted as follows: , represented as:
[0128] In the formula, For the first Class 1 K-means subclusters, The mean of the standardized potential vectors for this subcluster is the fault prototype.
[0129] Iterate through all prototype vectors and select the fault type corresponding to the prototype with the smallest Euclidean distance to the square of the normalized vector as the coarse identification result of the prototype branch. The corresponding solution expression is:
[0130] In the formula, This corresponds to the fault type of the most recent prototype.
[0131] The probability of attribution for a single type of fault is obtained by converting the distances of all prototypes, and is expressed as follows:
[0132] In the formula, This indicates that the current sample output by the prototype branch belongs to the first... The probability of a fault class, i.e., the first identification result. Indicates the first Total number of pre-stored prototypes for each type of fault. Indicates the first The first type of fault corresponding to the A standardized prototype vector, This represents the distance-temperature hyperparameter, used to adjust the influence of distance values on probability results. When it approaches 0, the classification result is equivalent to the nearest prototype classification. This represents the total number of all known fault types included in the offline modeling phase. This indicates the operation of natural exponents.
[0133] The reliability of the prototype branch is calculated based on the distance interval of the first identification result. The reliability is used to characterize the credibility of this prototype matching prediction and is expressed as:
[0134] In the formula, This represents the reliability of the prototype branch, with a fixed value range of 0 to 1. The closer the value is to 1, the higher the prediction reliability. and These are the squared distances from the query sample to the nearest and second-nearest class prototypes, respectively. This indicates a truncation operation, used to constrain the calculated original reliability value to... This avoids extreme distance differences that could cause numerical values to exceed limits.
[0135] Subsequently, online inference operations for the classification branch of the multilayer perceptron are executed in parallel. The network weights and activation rules of the multilayer perceptron are fixed offline; only forward mapping is performed online, without any training or update operations. The standardized joint latent representation is input into the pre-trained classifier to obtain the probability distribution of each class as the second recognition result. The original logistic values output by the classifier are normalized using Softmax to obtain the single-class fault probability, expressed as:
[0136]
[0137] In the formula, This indicates that the multilayer perceptron branch determines that the current sample belongs to the first... The probability of a fault class, i.e., the second identification result. To support a centralized number of known fault types, For parameters Two hidden layer MLP (Multi-Layer Perceptron). To output a vector of logical values, For the first Logical values, For the first Softmax-like probability.
[0138] The reliability of the classifier branch is determined based on the probability entropy of the second recognition result. The probability entropy is used to quantify the uncertainty of the classification output and is expressed as:
[0139]
[0140] In the formula, The entropy represents the information probability distribution corresponding to the output probability distribution of the multilayer perceptron. A higher entropy value indicates a more even probability of predicting various types of faults and a stronger uncertainty in the current prediction. For MLP branch reliability, This represents the total number of known fault types.
[0141] After completing the probability outputs and reliability calculations for both branches, a reliability-weighted fusion process is executed. Specifically, the prototype branch fusion weights are first calculated based on the normalized reliability of the two branches. The weight calculation formula is as follows:
[0142] In the formula, For prototype branch fusion weights, the remaining Weights are assigned to the branch probability results of the multilayer perceptron. It is a numerical stability constant to prevent the denominator from returning to zero when the reliability of both branches is zero at the same time.
[0143] The various fault probabilities output from the two branches are weighted and summed based on normalized weights to obtain a unified fault probability vector after fusion. The fusion operation is expressed as:
[0144] In the formula, This represents the normalized integrated probability vector corresponding to various faults after fusion.
[0145] By traversing all dimensions of the fusion probability vector, the fault code corresponding to the maximum probability is selected as the final fault type identifier. This maximum probability value is the comprehensive confidence level corresponding to this diagnosis. Simultaneously, the standardized drift contribution ratios of the three types of physical feature blocks are compared, and the category with the highest drift contribution is selected as the dominant physical feature block corresponding to this fault. The fault type identifier, diagnosis confidence level, and dominant physical feature block are integrated and output to form a complete fault type localization result.
[0146] Specifically, the fusion confidence level is expressed as:
[0147] In the formula, For the fusion of the first The probability of a type of failure. For the final fault category, To fuse confidence levels. When and Not lower than the confidence threshold Output high-confidence categories when the two branches are inconsistent; If the query category does not appear in the support set, the function outputs "Pending Review" and simultaneously provides the results of both branches, the distance interval, the probability entropy, and the dominant physical block. Therefore, the fusion function possesses directly executable probabilistic and decision logic, rather than being an undefined function.
[0148] The dual-branch weighted fusion operation can simultaneously leverage the advantages of prototype branch adaptation to multi-cluster non-spherical distribution of faults and multilayer perceptron to capture nonlinear fault boundaries. Based on predictive uncertainty, it adaptively allocates fusion weights, reducing the risk of misjudgment caused by low-confidence output of a single branch. At the same time, it outputs the physical anomaly source corresponding to the fault, providing interpretable diagnostic basis for battery fault tracing and vehicle safety management.
[0149] A specific example is used to illustrate the location of secondary fault types. For instance... Figure 4As shown, for monitoring windows where anomaly scores exceed the threshold, a frozen 12-dimensional joint latent representation is extracted and standardized using the support set's dimension-wise mean and standard deviation. One to three cluster mean prototypes are formed using K-means within the fault category, and a two-layer 256-dimensional MLP is used to form the Softmax class probability; the two branches are fused according to the reliability weighting rule in step 8. The multi-prototype branch handles multi-cluster and non-spherical structures within the same fault category, while the MLP branch learns the nonlinear class boundaries on the frozen latent representation. Neither branch updates the primary anomaly detection backbone in reverse.
[0150] Subsequently, the fault types corresponding to the anomaly labels are encoded according to the alarm categories defined in the dataset, such as fault types 16, 49152, and 8192. The multi-prototype branch outputs the nearest prototype category, the distance probability of all categories, and the nearest-nearest distance interval, while the MLP branch outputs the Softmax probability and normalized probability entropy of each fault type. If the results of the two branches are consistent and the fusion confidence reaches the threshold, a high-confidence fault type is output; if the two branches are inconsistent or the confidence is insufficient, the original results of the two branches are retained, and the results to be reviewed are output in combination with the alarm level, physical block drift contribution, and event stage.
[0151] In one implementation, the steps for constructing the fault prototype set and the classifier specifically include: Step S401: After the deterministic encoder has been trained and the parameters are frozen, samples of known fault types are collected, the joint latent representation of each sample is extracted by the deterministic encoder, and the samples are divided into support sets according to the fault category. Step S402: Standardize the joint latent representation of the support set samples using the dimension-wise mean and standard deviation of the support set; Step S403: Within each fault category, perform clustering on the standardized joint latent representation, and use the cluster centers within each category as the fault prototypes of that category to form the fault prototype set; Step S404: Based on the standardized support set used to construct the fault prototype set, the classifier is trained with the standardized joint latent representation as input and the known fault type as label.
[0152] In this embodiment, for samples that pass the first-level anomaly screening, the frozen joint latent representation is extracted, and the support set and query set are hierarchically divided according to the fault category. In one embodiment, the support set ratio is 0.5. The support set and query set are standardized using only the support set mean and standard deviation. A preferred approach to prototype partitioning is to perform K-means on the standardized joint latent representation within each fault category, rather than simultaneously using a combination of latent clustering, fault stage, and physical block contribution rules. Fault stage or physical block contribution is used only as an alternative grouping rule or explanatory label when reliable metadata is available, and is not combined with the default K-means. K-means is run 10 times with different initial centers, retaining the result with the minimum within-class squared error.
[0153] Within each fault category, K-means clustering is performed on the standardized joint latent representation to generate fault prototype sets for each category. In one embodiment, take... , .when Establish a mean prototype when; Automatic creation One prototype. Therefore, one to three sub-prototypes are supported for the same fault type. It is also possible to compare one, two, or three prototypes on independent validation sets and average them using macros. choose .
[0154] After the prototype set is constructed, a lightweight multilayer perceptron is trained on the joint latent representation of the same normalized support set as the prototype branch. In one embodiment, the input dimension is 12, containing two hidden layers, each with a width of 256; each hidden layer is, in sequence, fully connected, layer normalized, ReLU, and Dropout, with a Dropout rate of 0.10; the output layer dimension is equal to the number of known fault types in the support set; the Softmax layer outputs the normalized probability of each fault type. This MLP does not use the original monitoring signal as direct input, nor does it update the frozen backbone encoder.
[0155] MLP uses the AdamW optimizer with a learning rate of [missing information]. Weight decay The batch size is 64, and the maximum training time is 300 epochs. The early stopping condition is that the training loss does not improve for 40 consecutive epochs. Cross-entropy weights are assigned inversely proportional to the class frequency, and gradient norm is 1.0 for pruning.
[0156] The loss function uses cross-entropy loss, which assigns weights based on the inverse ratio of class frequencies. The loss calculation formula is as follows:
[0157]
[0158] In the formula, For MLP prediction categories, To support the sample size, For the first One supporting sample category, The class weights are obtained by inversely normalizing the class frequencies. The classification loss for MLP.
[0159] Throughout the training process, only the internal weight parameters of the multilayer perceptron are updated in reverse. The deterministic encoder and fault prototype set vectors are locked throughout the process and do not participate in gradient updates. After training converges, the standardized parameters of the support set, all fault prototype vectors, and the weights of the multilayer perceptron will be uniformly and permanently stored.
[0160] Offline modeling can build a dual-branch model with only a small number of labeled fault samples. The multi-prototype structure adapts to the spatial distribution characteristics of multiple evolution stages of faults. The lightweight multilayer perceptron learns the nonlinear discrimination boundary of the potential space and does not destroy the normal potential boundary constructed in the first-level anomaly screening stage. It is suitable for application scenarios where labeled fault samples of new energy vehicles and energy storage systems are scarce.
[0161] The technical effects of this embodiment are demonstrated through the results of the specific embodiments described above. For example... Figure 5 As shown, on the aforementioned real vehicle data, the overall true positive rate (TPR), positive predictive value (PPV), F1 score, and area under the curve (AUC) of the Level 1 anomaly detection model reached 0.9647, 0.9394, 0.9512, and 0.9596, respectively. Compared to the average results of the compared methods, the positive predictive value, F1 score, and AUC improved by 23.72%, 15.37%, and 37.97%, respectively. These values represent the performance of Level 1 anomaly detection and should not be confused with Level 2 multi-prototype or MLP fusion classification metrics.
[0162] In summary, compared with the prior art, this embodiment has the following outstanding technical effects: This invention extends the online fault diagnosis of power batteries from single-layer anomaly detection to a hierarchical diagnostic process involving observation layer, physical block layer, potential structure layer and event fault layer, and can simultaneously output anomaly level, fault type and physical explanation.
[0163] This invention retains the advantage of using normal samples to establish normal state boundaries, but uses this boundary as a first-level anomaly screening module, allowing faulty samples to enter the second stage for more granular localization.
[0164] The MLP branch of this invention only acts on the frozen joint latent representation and does not retrain the backbone anomaly detection network. Therefore, it can learn the nonlinear fault type boundary while avoiding the destruction of the stable latent structure formed by normal samples.
[0165] The multi-prototype branch of this invention can adapt to multi-stage, multi-cluster, and non-spherical trajectories of the same fault type, and the MLP branch can supplement the nonlinear discrimination capability. The fusion of the two can improve the robustness of specific fault type localization.
[0166] This invention can explain shared anomaly directions and fault-specific branches among different fault types. For example, different faults may deviate from the normal manifold in the early stages along a common global anomaly direction, while forming type-related branches in local potential spaces or physical block layers.
[0167] This invention can output the dominant physical feature block corresponding to the fault, such as center operation, voltage edge or temperature edge, thereby providing a basis for battery maintenance, fault review and subsequent safety strategies.
[0168] The online phase of this invention only requires sliding window updates, frozen encoder forward calculations, global distance scoring, and lightweight decoder inference, resulting in low computational complexity and making it easy to deploy on vehicle-side, edge-side, or cloud-side devices.
[0169] In real vehicle data implementations, the model of the present invention achieves high true accuracy, positive prediction value, F1 score and area under the curve in overall diagnostic performance, and can further display the branch structure and physical block location information of specific fault types.
[0170] Furthermore, this invention can be deployed on vehicle battery management system controllers, on-board computing units, edge gateways, or cloud monitoring platforms. When deployed on the vehicle side, preprocessing mapping, frozen encoders, normal center, alarm thresholds, multi-prototype sets, and MLP parameters can be stored as fixed model parameters, performing only sliding window updates, encoder forward computation, global distance scoring, and lightweight fault type decoding. For changes in vehicle model, chemistry system, or battery pack structure, the preprocessing parameters and normal center can be re-estimated using the corresponding object's normal operation data, and the second-stage prototype and MLP decoder can be updated using a small number of known fault samples.
[0171] This invention is not limited to lithium-ion power batteries for new energy vehicles, but also applicable to energy storage battery clusters, aviation power batteries, and other electrochemical energy storage systems with multi-channel voltage, temperature, current, or state monitoring signals. Any implementation that employs hierarchical global and local latent characterization, first-level anomaly screening, and multi-prototype and MLP-based fault type localization based on frozen latent representation falls under the equivalent implementation of this invention's technical concept.
[0172] like Figure 6As shown in the figure, an embodiment of the present invention provides a power battery fault diagnosis system, which includes: a monitoring window construction module 10, a potential representation acquisition module 20, an abnormal window judgment module 30, and a fault type location module 40.
[0173] Specifically, the monitoring window construction module 10 is used to acquire real-time monitoring data of the battery management system, divide the monitoring data into physical feature blocks, and construct a monitoring window based on each physical feature block; wherein, the types of physical feature blocks include center running feature blocks, voltage edge feature blocks, and temperature edge feature blocks; the latent representation acquisition module 20 is used to input the monitoring window into a pre-trained deterministic encoder, obtain global latent variables and local latent variables through forward mapping, and concatenate the two into a joint latent representation; wherein, the encoder is a deterministic mapping network with fixed parameters, and always outputs the same latent variables for the same monitoring window; the abnormal window judgment module 30 is used to calculate the distance between the global latent variable and the preset normal reference point as an abnormal score, and when the abnormal score exceeds a preset threshold, the monitoring window is judged as an abnormal window; the fault type localization module 40 is used to extract the joint latent representation obtained by the encoder mapping for the abnormal window, perform fault type identification through a preset fault prototype set and a pre-trained classifier respectively, and perform reliability weighted fusion of the identification results of the two to obtain the fault type localization result.
[0174] Based on the above embodiments, the present invention also provides a terminal device, the principle block diagram of which can be as follows: Figure 7 As shown, the terminal device includes a processor, memory, network interface, display screen, and temperature sensor connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for diagnosing power battery faults. The display screen can be an LCD screen or an e-ink screen. The temperature sensor is pre-installed inside the terminal device to detect the operating temperature of internal components.
[0175] Those skilled in the art will understand that Figure 7 The schematic diagram shown is only a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0176] In one embodiment, a terminal device is provided, including a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, the one or more programs including instructions for performing operations as described in the embodiments of the methods above.
[0177] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0178] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0179] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for diagnosing faults in a power battery, characterized in that, The method includes: Real-time monitoring data from the battery management system is acquired, the monitoring data is divided into physical feature blocks, and a monitoring window is constructed based on each physical feature block; wherein, the types of physical feature blocks include center running feature blocks, voltage edge feature blocks, and temperature edge feature blocks; The monitoring window is input into a pre-trained deterministic encoder, and global and local latent variables are obtained through forward mapping. The two are then concatenated into a joint latent representation. The encoder is a deterministic mapping network with fixed parameters, which always outputs the same latent variables for the same monitoring window. The distance between the global latent variable and the preset normal baseline point is calculated as the anomaly score. When the anomaly score exceeds the preset threshold, the monitoring window is determined to be an anomaly window. For the abnormal window, the joint latent representation obtained by the encoder mapping is extracted, and the fault type is identified by using a preset fault prototype set and a pre-trained classifier, respectively. The identification results of the two are then weighted and fused based on reliability to obtain the fault type localization result. The step of dividing the monitoring data into physical feature blocks and constructing a monitoring window based on each physical feature block includes: The battery system operating parameters in the monitoring data are divided into the central operating feature block, the single cell voltage consistency index is divided into the voltage edge feature block, and the temperature consistency index is divided into the temperature edge feature block. The battery system operating parameters include at least current, voltage, and state of charge parameters; the cell voltage consistency index includes at least cell voltage range and cell voltage standard deviation; and the temperature consistency index includes at least temperature range and temperature rise rate. Denoising and normalization are performed on each physical feature block, and then low-dimensional feature vectors of each physical feature block are obtained by variance weighting and principal component dimensionality reduction. The low-dimensional feature vectors of each physical feature block at the same sampling time are concatenated to form the time feature; The time features of a predetermined number of consecutive sampling times are stacked in chronological order to form the monitoring window; The process involves identifying fault types using a preset fault prototype set and a pre-trained classifier, and then performing a reliability-weighted fusion of the identification results to obtain the fault type localization result, including: The joint latent representation of the abnormal window obtained by the encoder is standardized using the one-dimensional mean and standard deviation of the preset support set; Calculate the distance from the standardized joint latent representation to each category of the fault prototype set, and determine the classification probability of each category based on the distance to obtain the first identification result; The standardized joint latent representation is input into the pre-trained classifier to obtain the probability distribution of each category as the second recognition result; The reliability of each is determined based on the distance interval of the first identification result and the probability entropy of the second identification result, respectively. The first identification result and the second identification result are then weighted and fused with a reliability normalization weight to obtain the fault type location result. Wherein, the support set is a set of joint latent representations obtained by mapping known fault type samples through the encoder; the fault prototype set is a set of cluster centers obtained by clustering the support set within each category; the fault type localization result includes at least one of fault type identifier, dominant physical feature block type, and diagnostic confidence.
2. The power battery fault diagnosis method according to claim 1, characterized in that, The deterministic encoder includes a shared backbone network, global branches, and local branches; The shared backbone network comprises multiple deep fully connected blocks connected in sequence. Each deep fully connected block consists of a fully connected layer, layer normalization, activation function, and proxy pulse gating. The global branch and the local branch each contain at least one fully connected layer and an activation function. The global branch outputs global latent variables of a preset dimension, and the local branch outputs local latent variables of a preset dimension. The output of the shared backbone network is fed into the global branch and the local branch respectively, so as to output the global latent variable and the local latent variable respectively.
3. The power battery fault diagnosis method according to claim 1, characterized in that, The calculation of the distance between the global latent variable and the preset normal baseline point is used as an anomaly score. When the anomaly score exceeds a preset threshold, the monitoring window is determined to be an anomaly window, including: Calculate the Euclidean distance from the global latent variable to the normal reference point at each sampling time within the monitoring window, and use the maximum value of the distance at each time as the anomaly score; The preset threshold is determined based on the precision-recall curve of an independent validation set, or based on the high quantile of the abnormal score distribution of pure normal samples.
4. The power battery fault diagnosis method according to claim 1, characterized in that, The training steps of the deterministic encoder include: Obtain a training set consisting only of samples from the normal monitoring window; Construct a joint training architecture that includes the deterministic encoder and decoder; wherein the encoder maps the input normal monitoring window to global latent variables and local latent variables and concatenates them into a joint latent representation, and the decoder uses the joint latent representation as input to reconstruct the normal monitoring window; A joint loss function is constructed based on input reconstruction loss, recoding consistency loss, global single-class compactness loss, local diffusion smoothing loss, joint potential volume compression loss, and time-delay direction consistency loss, and the encoder and decoder are jointly trained. After training is complete, the encoder parameters are frozen, and the normal baseline is updated with the mean of the global latent variables of all normal training samples.
5. The power battery fault diagnosis method according to claim 1, characterized in that, The steps for constructing the fault prototype set and the classifier include: After the deterministic encoder is trained and its parameters are frozen, samples of known fault types are collected, and the joint latent representation of each sample is extracted by the deterministic encoder and divided into support sets according to fault categories. The joint latent representation of the support set samples is standardized using the dimension-wise mean and standard deviation of the support set. Within each fault category, clustering is performed on the standardized joint latent representation, and the cluster centers within each category are used as fault prototypes for that category to form the fault prototype set. Based on the standardized support set used to construct the fault prototype set, the classifier is trained with the standardized joint latent representation as input and the known fault type as the label.
6. A power battery fault diagnosis system, characterized in that, The system, used to implement the power battery fault diagnosis method as described in any one of claims 1-5, comprises: The monitoring window construction module is used to acquire real-time monitoring data of the battery management system, divide the monitoring data into physical feature blocks, and construct a monitoring window based on each physical feature block; wherein, the types of physical feature blocks include center running feature blocks, voltage edge feature blocks, and temperature edge feature blocks; The latent representation acquisition module is used to input the monitoring window into a pre-trained deterministic encoder, obtain global latent variables and local latent variables through forward mapping, and concatenate the two into a joint latent representation; wherein, the encoder is a deterministic mapping network with fixed parameters, which always outputs the same latent variables for the same monitoring window; An abnormal window judgment module is used to calculate the distance between the global latent variable and the preset normal benchmark point as an abnormal score. When the abnormal score exceeds the preset threshold, the monitoring window is judged as an abnormal window. The fault type localization module is used to extract the joint latent representation obtained by the encoder mapping for the abnormal window, identify the fault type through a preset fault prototype set and a pre-trained classifier, and perform reliability weighted fusion on the identification results of the two to obtain the fault type localization result.
7. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a power battery fault diagnosis program stored in the memory and executable on the processor. When the processor executes the power battery fault diagnosis program, it implements the steps of the power battery fault diagnosis method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a power battery fault diagnosis program, which, when executed by a processor, implements the steps of the power battery fault diagnosis method as described in any one of claims 1-5.