Lithium-ion battery state of health estimation method based on physical information and data driving
Patent Information
- Application Number
- CN202610870437.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-04
AI Technical Summary
[0004]已经虽然已经公开了通过长短期的运行特征判断电池健康的判断,但是其没有公开如何有效提取多尺度局部时序模式,增强对短期充放电动态的感知能力;如何联合建模时间步重要性与特征通道重要性,实现两者的自适应深度融合;如何在数据驱动框架中嵌入电化学物理约束,保证预测结果的单调性、有界性和平滑性;如何稳定深层双向循环网络的训练过程,提升跨电池泛化能力
本发明在数据层面,将每次充放电循环的运行参数与实测健康度按时序严格对齐并形成序列,再通过滑动窗口划分和归一化,将全生命周期退化轨迹转化为大量有监督的局部时序样本,消除了多参数量纲差异对梯度传导的干扰。在特征提取层面,多尺度时间卷积模型通过三个不同尺寸卷积核的并行分支,在同一时间步上同步捕获从单次充放电的微观瞬态、跨若干次循环的中频退化节律到覆盖长程历史的宏观趋势,输出多分辨率融合的第一特征向量。在时序编码层面,深度双向长短期记忆网络通过前向和后向循环网络并行读取序列,并堆叠3层逐级抽象,同时引入层归一化稳定深层梯度传播,使得每个时间步的编码同时融合了历史累积成因和未来趋势预兆的双向长程依赖。在注意力融合层面,时间注意力分支通过Softmax权重对序列进行选择性浓缩,特征注意力分支通过压缩激励网络自动辨识并强化与退化强相关的物理参量通道,自适应门控网络根据输入内容动态调节时序聚焦与特征聚焦的融合比例,输出信噪比高且与当前退化阶段高度适配的第三特征向量。在回归预测层面,两层全连接网络将融合后的高层语义映射为单一健康度标量,并通过物理上限约束和相邻循环衰减幅度约束作为惩罚项加入损失函数,确保预测结果服从电池老化不可逆且渐进的基本规律。在不确定性量化层面,蒙特卡洛Dropout在推理阶段保持丢弃层激活并进行100次独立前向传播,通过预测分布的标准差量化置信度,并划分三级置信区间,使得单次评估同步输出健康度数值和可信度等级。整个方案以端到端联合训练的方式将数据驱动的特征学习与物理先验约束统一,实现了从窗口输入到带可靠性标签的健康度输出的一次性推理,兼顾了评估精度、物理一致性和工程可操作性。
Smart Images

Figure CN122690401A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy storage technology for new energy power systems, specifically to a method for assessing the health status of lithium-ion batteries based on physical information and data-driven approaches. Background Technology
[0002] Lithium-ion batteries have become the mainstream energy storage technology due to their high energy density and long cycle life. Battery state of health (SOH) is typically defined as the ratio of current maximum usable capacity to rated capacity, and is a key indicator for assessing battery aging. When the SOH drops below 70%-80%, the battery reaches the end of life (EOL) and must be replaced promptly to avoid safety risks.
[0003] A prior art method for estimating the health of a lithium-ion battery, disclosed in publication number CN120779255A, involves extracting feature vectors reflecting battery aging from charge-discharge cycle data to form a feature matrix, and preprocessing the feature matrix. Based on the preprocessed feature matrix, important features are selected to obtain an important feature matrix. Based on the important feature matrix, a Mogrifier-LSTM model is used to construct a battery health estimation model, and the model is trained to obtain a trained battery health estimation model. The trained model is then used to estimate the battery health.
[0004] While the method of judging battery health by long-term and short-term operational characteristics has been publicly disclosed, it has not disclosed how to effectively extract multi-scale local time-series patterns to enhance the perception of short-term charge and discharge dynamics; how to jointly model the importance of time steps and the importance of feature channels to achieve adaptive deep fusion of the two; how to embed electrochemical and physical constraints in the data-driven framework to ensure the monotonicity, boundedness, and smoothness of the prediction results; and how to stabilize the training process of deep bidirectional recurrent networks to improve cross-battery generalization ability.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a method for assessing the health status of lithium-ion batteries based on physical information and data-driven approaches, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A physical information and data-driven method for assessing the health status of lithium-ion batteries includes the following steps: Step 1: Cycle charge and discharge the sample battery and record the running data, and record the health status at the same time. Correspond the running data and the health status to form a charge and discharge sequence, and set a sliding window to form an input feature set in the charge and discharge sequence; Step 2: Input the input feature set into the multi-scale temporal convolution model to obtain the first feature vector of several duration features. Encode the first feature vector using deep bidirectional LSTM to obtain the second feature vector. Step 3: Perform dual attention fusion on the second feature vector and simultaneously perform fusion through adaptive gating to obtain the third feature vector. Input the third feature vector into a two-layer fully connected network and output it with health score as the label. Label the trained process as a health assessment model. Step 4: During the training of the health assessment model, obtain the weight coefficients for each training process, train synchronously in Monte Carlo Dropout, obtain the confidence score of each health assessment model output, and set the confidence interval. Step 5: Organize the actual charging and discharging operation data of the battery into an input feature set and input it into the health evaluation model. At the same time, use the weight coefficients summarized by the health evaluation model and the input features to obtain the confidence intervals corresponding to the health level and confidence level, and output them.
[0008] Furthermore, the sample batteries were subjected to cyclic charging and discharging, and the operating data of each battery was recorded during each charging and discharging cycle. In addition, the health status of each battery after each charging and discharging cycle was measured. The health status is the SOH value. The operating data experienced by each health status was recorded and sorted according to the order of acquisition time during the charging and discharging cycle, forming a charging and discharging sequence with the operating data as the sample and the corresponding health status as the label. Set the time window length and step size, slide the data on the charge-discharge sequence to obtain the input feature set, and normalize the data in the input feature set respectively.
[0009] Furthermore, the input feature set is input into a multi-scale temporal convolution model, which contains three convolutional branches. Each convolutional branch has a different kernel size, corresponding to short-term, medium-term, and long-term feature extraction, respectively. All convolutional branches have the same number of output channels. The outputs of the three convolutional branches are combined to form the first feature vector.
[0010] Furthermore, the first feature vector is encoded using a deep bidirectional LSTM, which has several layers. Each layer contains a forward LSTM and a backward LSTM. The hidden state distribution is stabilized by applying layer normalization, and the hidden states in the two directions are concatenated as the output of the layer. The input of the first layer is the first feature vector, and the input of each subsequent layer is the output of the previous layer, thus obtaining the second feature vector.
[0011] Furthermore, the second feature vector is subjected to dual attention fusion, which includes a temporal attention branch and a feature attention branch. The temporal attention branch obtains the temporal context vector by weighted summation of the hidden state data in the second feature vector through Softmax. The feature attention branch obtains the feature context vector by weighted summation of the hidden state data in each dimension through Squeeze-and-Excitation, with the number of hidden state data dimensions as the number of channels. The temporal context vector and the feature context vector are concatenated and input into a gating network. The gating network dynamically generates fusion weights to form adaptive gating fusion, which outputs a third feature vector.
[0012] Furthermore, the third feature vector is input into a two-layer fully connected network to output a health scalar. At the same time, constraints are set to form a penalty term. The constraints include a health score less than or equal to 1 and a health score greater than a preset health score change threshold in two adjacent charge-discharge cycles in the time series. The Adam optimizer is used to integrate the trained multi-scale temporal convolutional model, deep bidirectional LSTM encoding, dual attention fusion, and fully connected network into a health assessment model.
[0013] Furthermore, during the training of the health assessment model, the weight coefficients generated by the scale-time convolutional model, deep bidirectional LSTM encoding, dual attention fusion and fully connected network during operation are obtained and summarized. The summarized weight coefficients and input feature set are used to obtain the confidence of each output result through Monte Carlo Dropout. The confidence of each result and the health score of the health assessment model are output simultaneously. Define a confidence interval, which includes high reliability, medium reliability, and low reliability.
[0014] Furthermore, during the charging and discharging process of the battery, the operating data is recorded and organized into an input feature set, which is then input into the health evaluation model to obtain the health level and its corresponding confidence level. Based on the confidence level, the corresponding confidence interval is matched and combined with the health level for output.
[0015] Compared with the prior art, the beneficial effects of the present invention are: At the data level, this invention strictly aligns the operating parameters of each charge-discharge cycle with the measured health status in a time sequence to form a sequence. Then, through sliding window partitioning and normalization, the entire lifecycle degradation trajectory is transformed into a large number of supervised local time-series samples, eliminating the interference of multi-parameter dimensional differences on gradient propagation. At the feature extraction level, a multi-scale temporal convolutional model, through parallel branches of three convolutional kernels of different sizes, synchronously captures everything from the microscopic transients of a single charge-discharge cycle, the mid-frequency degradation rhythm across several cycles, to the macroscopic trends covering long-term history at the same time step, outputting a multi-resolution fused first feature vector. At the temporal encoding level, a deep bidirectional long short-term memory network reads the sequence in parallel through forward and backward recurrent networks, stacking three layers for progressive abstraction. Layer normalization is introduced to stabilize deep gradient propagation, ensuring that the encoding at each time step simultaneously integrates the bidirectional long-term dependence of historical cumulative causes and future trend predictions. At the attention fusion level, the temporal attention branch selectively condenses the sequence using Softmax weights, while the feature attention branch automatically identifies and strengthens physical parameter channels strongly correlated with degradation through a compression-excitation network. An adaptive gating network dynamically adjusts the fusion ratio of temporal and feature focusing based on the input content, outputting a third feature vector with high signal-to-noise ratio and a high degree of adaptation to the current degradation stage. At the regression prediction level, a two-layer fully connected network maps the fused high-level semantics to a single health scalar, and incorporates physical upper bound constraints and adjacent cycle decay magnitude constraints as penalty terms into the loss function to ensure that the prediction results conform to the fundamental law of irreversible and gradual battery aging. At the uncertainty quantification level, Monte Carlo Dropout maintains the activation of the dropout layer during the inference phase and performs 100 independent forward propagations. Confidence is quantified by the standard deviation of the prediction distribution, and three confidence intervals are defined, enabling simultaneous output of health value and confidence level in a single evaluation. The entire scheme unifies data-driven feature learning and physical prior constraints through end-to-end joint training, achieving one-time inference from window input to health output with reliability labels, balancing evaluation accuracy, physical consistency, and engineering operability. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the overall method flow of the present invention; Figure 2 This is a schematic diagram of the multi-scale temporal convolution model of the present invention; Figure 3 This is a schematic diagram of the deep bidirectional LSTM encoding of the present invention; Figure 4 This is a schematic diagram of dual attention fusion according to the present invention; Figure 5 This is a schematic diagram of the two-layer fully connected network of the present invention; Figure 6 This is a data comparison chart of a preferred embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0018] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly. Example
[0019] Please see Figures 1-6 The present invention provides a technical solution: A physical information and data-driven method for assessing the health status of lithium-ion batteries includes the following steps: Step 1: Cycle charge and discharge the sample battery and record the running data, and record the health status at the same time. Correspond the running data and the health status to form a charge and discharge sequence, and set a sliding window to form an input feature set in the charge and discharge sequence; Step 1 includes the following: The sample batteries were subjected to cyclic charging and discharging. The operating data of each battery was recorded during each charging and discharging cycle. The health status of each battery after each charging and discharging cycle was measured. The health status is the SOH value. The operating data experienced by each health status was recorded and sorted according to the order of acquisition time during the charging and discharging cycle, forming a charging and discharging sequence with the operating data as the sample and the corresponding health status as the label. Set the time window length and step size, slide the data on the charge-discharge sequence to obtain the input feature set, and normalize the data in the input feature set respectively.
[0020] In a preferred embodiment, the operating data consists of multi-dimensional time-series parameters that characterize the battery's current operating state and performance changes, collected in real-time by the battery testing system during the sample battery's cyclic charge-discharge process. Specifically, this includes the constant current charging current, constant voltage charging cutoff voltage, and charging duration during the charging phase; the constant current discharging current, discharging cutoff voltage, and discharging duration during the discharging phase; the temperature values at the battery surface or tabs throughout the entire charge-discharge cycle; the actual charged and discharged capacity calculated using the ampere-hour integration method or the coulomb counting method for this cycle; and the coulombic efficiency obtained from the ratio of discharged capacity to charged capacity. Furthermore, it includes characteristic points of the differential curve showing voltage changes with capacity during each charge-discharge cycle, such as the position and slope of the median point of the discharge voltage plateau, and the open-circuit voltage obtained by static measurement after each cycle.
[0021] In terms of recording method, the battery testing system creates a separate data recording unit for each charge-discharge cycle of each sample battery. This unit stores the numerical sequences of the aforementioned parameters sequentially according to the sampling timestamps. The sampling frequency is set based on the charge-discharge rate; for example, synchronous sampling is performed at 1-second intervals under 1C charge-discharge conditions to ensure strict time alignment of each parameter. After each charge-discharge cycle, the collected raw voltage and current data are integrated in ampere-hours to calculate the actual charge and discharge capacity for that cycle. Simultaneously, the voltage curve is differentiated to extract feature points. These derived parameters, along with the raw collected data, are stored in the data recording unit for that cycle.
[0022] All data recording units generated by all batteries across all cycles are arranged in series according to the chronological order of the charge-discharge cycles. Each data recording unit serves as an independent sample frame, containing a complete record of all operating parameters for that cycle. Frames are arranged sequentially in the form of time steps. Simultaneously, each sample frame is labeled with its corresponding health value, obtained by dividing the actual discharged capacity measured at the end of that cycle by the battery's nominal capacity. This forms a complete charge-discharge sequence with the sample frame sequence as input and the health sequence as the label.
[0023] The charge-discharge sequence is treated as a multi-dimensional time series. A sliding window length and step size are set. The window length is measured by the number of consecutive charge-discharge cycles it contains, for example, 32 cycles, and the step size is set to 1 cycle. The window slides progressively across the sequence from its starting position. At each step, 32 consecutive sample frames within the window are extracted as an input sample. This input sample forms a two-dimensional matrix, where the 32 rows correspond to the 32 consecutive charge-discharge cycles within the window, and each column corresponds to the sequence value of a certain operating parameter in those 32 cycles. All the two-dimensional matrix samples generated by the sliding extraction together constitute the input feature set. The label for each sample is the health score of the last cycle in the window. Subsequently, each column of parameters for each sample in the input feature set is normalized using minimum-maximum normalization or Z-score normalization, mapping parameters with different physical dimensions to a dimensionless numerical space, resulting in a standardized input feature set that can be directly used by subsequent multi-scale temporal convolutional models.
[0024] By strictly aligning and concatenating the complete operational data of each cycle with the measured SOH value at the end of the cycle according to the charge-discharge time sequence, the continuous and irreversible physical process of battery aging is transformed into an ordered causal inference structure. Each charge-discharge cycle's multidimensional signal is defined as the operating condition that induces capacity degradation, followed by SOH, which is the result label of the combined effect of this operating condition and historical accumulated damage. This arrangement solidifies the time-dependent unidirectional relationship of "operating parameters → health decay" at the data level, allowing the model to learn from the sequence how capacity degradation evolves with the accumulation of charge-discharge cycles, rather than simply treating SOH as a static output corresponding to the instantaneous operating condition. Simultaneously, the data organized in sequence form completely preserves the monotonic trend, local fluctuations, and potential acceleration inflection points of the degradation trajectory, retaining a continuous and unequal-length full-lifecycle degradation timeline for subsequent sliding window division. This allows subsequent models to capture long-range dependencies and phased patterns across cycles based on this timeline.
[0025] Based on the established complete charge-discharge degradation sequence, a sliding window with a set window length and step size is used to decompose a long sequence covering the entire life cycle into a large number of local subsequences with time shifts. Each subsequence represents a fixed-length segment of operational history starting from a certain initial cycle, and its corresponding label is still the health status measured at the end of that historical window. This partitioning method transforms SOH prediction from a regression problem of a single long sequence into a distribution learning problem of a large number of local historical segments. This allows the model to no longer be limited to fitting the entire degradation curve of a specific battery, but to learn to infer the current health status from any starting point of any aging stage, based solely on the operational characteristics of the most recent few charge-discharge cycles, greatly expanding the generalization boundary. The step size mechanism of the sliding window further results in a large amount of overlap between adjacent windows. The same segment of charge-discharge data appears in multiple samples starting from different historical positions. This data reuse not only expands the effective training sample size, but also forces the model to learn the differences in the contribution of the same operational characteristics to SOH prediction under different historical contexts, thereby enhancing the sensitivity to changes in the temporal context of the degradation trajectory. Data normalization within the window maps parameters with different physical dimensions and numerical scales, such as voltage, current, temperature, and capacity increments, to the same dimensionless space. This eliminates the offset in weight update direction caused by excessive differences in the original values of the parameters during gradient backpropagation. This allows each convolutional branch of subsequent multi-scale temporal convolutions to compare the temporal evolution patterns of different parameters on the same numerical benchmark. It also prevents high-amplitude parameters from obscuring low-amplitude but equally crucial degradation indicator signals, providing a dimensionally consistent numerical basis for the balanced extraction of temporal features from window data.
[0026] Step 2: Input the input feature set into the multi-scale temporal convolution model to obtain the first feature vector of several duration features. Encode the first feature vector using deep bidirectional LSTM to obtain the second feature vector. Step 2 includes the following: Step 201: Input the input feature set into the multi-scale temporal convolution model. The multi-scale convolution model contains three convolution branches, each with a different kernel size, corresponding to short-term, medium-term, and long-term feature extraction, respectively. All convolution branches have the same number of output channels. The outputs of the three convolution branches are combined to form the first feature vector.
[0027] like Figure 2As shown, the multi-scale temporal convolutional model is a parallel feature extraction structure consisting of three independent convolutional branches. Its starting point is to capture degradation phenomena occurring at completely different time scales from the charge-discharge history contained within a fixed-length sliding window, and then fuse this multi-level dynamic information into a unified first feature vector. The input is a two-dimensional data matrix after normalization in step one. The rows of the matrix correspond to the number of charge-discharge cycles contained within the sliding window, and the columns correspond to multiple operating parameters such as voltage, current, temperature, and capacity increment recorded in each cycle. After receiving this input matrix, the model simultaneously feeds it into the three parallel convolutional branches.
[0028] Within each convolutional branch, the core computational unit is a one-dimensional convolutional layer that slides unidirectionally along the time dimension, performing a local weighted summation across time on the parameters of each column of the input matrix. The only structural difference between the three branches lies in the kernel size of this one-dimensional convolutional layer: the kernel size of the first branch is set to 3, and its receptive field covers only 3 consecutive charge-discharge cycles. Its design purpose is to focus on microscopic anomalies such as instantaneous voltage fluctuations, current spikes, or short-term temperature rises that occur during a single charge-discharge process; the kernel size of the second branch is set to 5, and its receptive field expands to 5 cycles, enabling it to capture medium-granularity decay rhythms such as the average slope of capacity reduction or the gradual decrease in coulomb efficiency over several consecutive charge-discharge cycles; the kernel size of the third branch is set to 7, and its receptive field further spans 7 cycles, allowing direct observation of macroscopic trends over a complete maintenance cycle, such as the curvature change as it transitions from a stable decay phase to an accelerated decay phase. All convolutional layers have a fixed time step of 1 and employ a zero-padding strategy to ensure that the output sequence length after convolution is strictly consistent with the number of iterations in the input window. Each convolutional layer is followed by a batch normalization layer and a modified linear unit activation function in sequence. The batch normalization layer standardizes the feature responses across time steps within the same channel, eliminating internal covariate bias and making deep gradient propagation smoother. The modified linear unit introduces nonlinearity into the features, enabling the model to characterize complex fluctuations in the decay trajectory.
[0029] The output channels of the three convolutional branches are set to an identical 64. This means that regardless of whether the branch focuses on short-term, medium-term, or long-term features, each branch generates a 64-dimensional local semantic vector at each time step. Concatenating these 64-dimensional vectors generated by the three branches at their respective time steps along the channel dimension yields a first feature vector with the time step count remaining the same as the original window iteration count, but with the number of channels expanded to 192. At each time point, the first 64 channels of this first feature vector primarily respond to transient disturbances from the recent charge-discharge cycles within the window; the middle 64 channels primarily encode the mid-frequency rhythms spanning several charge-discharge cycles; and the last 64 channels primarily represent the long-term trend covering most of the historical interval. Degradation information from multiple time scales naturally coexists within a unified tensor structure, decomposing the overall picture of battery aging into a collection of different frequency components. This provides both microscopic details and macroscopic background for the subsequent bidirectional long short-term memory network.
[0030] Training is not performed independently, but rather as part of the front-end feature extraction stage of the entire health assessment process. It is connected to the bidirectional long short-term memory network in step 202, the dual attention fusion module in step 301, and the fully connected regression network in step 302, forming an end-to-end computational graph for joint optimization. At the start of training, the weights of all one-dimensional convolutional layers are randomly assigned using the Ho's initialization method, and the bias is initialized to 0. The training data consists of a large number of sliding window samples generated in step one. The input of each sample is a normalized running data matrix within the window, with the shape being the number of iterations multiplied by the number of parameters. The label is the measured health scalar corresponding to the end of that window. The entire joint model uses an adaptive moment estimation optimizer for parameter updates. The initial learning rate of the optimizer is set to 0.001, the first moment decay rate is 0.9, and the second moment decay rate is 0.99. Batch gradient descent is used during training, and the batch size is set to 32 based on available memory. The loss function consists of two parts: one is the mean squared error between the predicted health and the measured health, driving the model to approximate the true degradation trajectory; the other is the physical constraint penalty term set in step 302, which incurs a penalty when the predicted health exceeds one or the health decay in adjacent cycles exceeds a preset threshold, thus injecting the irreversible and gradual characteristics of battery aging as prior knowledge into the training process. After each forward propagation, the loss gradient is backpropagated from the fully connected layer to the attention module, the bidirectional long short-term memory network, and finally applied to the convolutional kernel weights of the three convolutional branches, causing convolutional kernels of different sizes to adjust their own perceptual modes in a direction that reduces the overall prediction error. Training continues until the overall loss on the validation set converges and no longer decreases. At this point, the multi-scale temporal convolution parameters are frozen, and the three convolutional branches have automatically learned the optimal convolutional kernel weight distribution. This makes the different channel groups in the first feature vector highly sensitive to the multi-scale dynamic features of the battery under different aging stages and operating conditions, and forms an efficient connection with the subsequent temporal coding network. The entire feature extraction process does not require manual division of frequency bands or design of filters. It is entirely driven by data to complete the adaptive decomposition and multi-resolution fusion of degradation information.
[0031] The normalized input feature set enters three parallel convolutional branches with different kernel sizes, which is equivalent to opening three time observation windows simultaneously in the same charge and discharge history: small-sized convolutional kernels only span a few adjacent sampling points, focusing on capturing local transient patterns with extremely fine granularity, such as sudden changes in voltage curves, current spikes, or instantaneous temperature disturbances during a single charge and discharge process. These transients often imply abrupt changes in internal resistance or precursors to micro-short circuits; medium-sized convolutional kernels cover several to a dozen charge and discharge cycles, and can abstract the stage-by-stage degradation performance within several consecutive windows, such as the slow but stable periodic decay slope of capacity and the gradual decrease in coulombic efficiency, into a medium-granular degradation rhythm; large-sized convolutional kernels have a receptive field spanning dozens or even hundreds of charge and discharge cycles, directly sensing large-scale trends across long periods, such as the global turning point of SOH from slow linear decay to accelerated decay inflection point. The three convolutional branches maintain the same number of channels in their outputs, ensuring that the dimension of the representation vectors generated at each time step is strictly consistent. These semantics, originating from different temporal perspectives, are equally weighted and merged in the same dimensional space. The first feature vector constructed in this way essentially performs spatiotemporal alignment and information fusion of microscopic transient perturbations, mesoscopic degradation rhythms, and macroscopic aging trends at each historical moment. This allows degradation features that were originally mixed at different time levels to coexist in a unified high-dimensional representation, solving the problem that a single fixed receptive field convolution can only be biased towards a certain scale, resulting in the loss of information at other scales. This multi-resolution fusion representation, as the initial input of the deep bidirectional LSTM in step 202, has already transformed the original running parameters from unstructured temporal values into a hierarchical feature sequence with both local details and global biases. When the bidirectional LSTM further models long-range dependencies on this basis, it no longer needs to re-explore cross-scale associations from the most original signal layer, but directly performs higher-order temporal inference on the already organized multi-scale semantics. This significantly reduces the uncertainty of the recurrent neural network in selecting the receptive field and improves the learning efficiency of multi-scale coupling effects in complex degradation trajectories.
[0032] Step 202: Encode the first feature vector using a deep bidirectional LSTM. The deep bidirectional LSTM has several layers. Each layer contains a forward LSTM and a backward LSTM. The hidden state distribution is stabilized by applying layer normalization and the hidden states in the two directions are concatenated as the output of the layer. The input of the first layer is the first feature vector, and the input of each subsequent layer is the output of the previous layer. Obtain the second feature vector.
[0033] like Figure 3As shown, deep bidirectional LSTM encoding is a deep recurrent encoding structure composed of multiple layers of bidirectional long short-term memory networks. Its design goal is to establish a complete contextual dependency relationship along the time direction of the charging and discharging cycle based on the first feature vector output in step 201, so that the encoding at each time step not only contains the historical accumulated information of all previous cycles, but also integrates the future trend information of all subsequent cycles. The input is the first feature vector formed by concatenating the three convolutional branches in step 201. This vector maintains the same length as the number of cycles in the sliding window in the time dimension, and has 192 channels in the feature dimension, which respectively carry the degenerate semantics of three time scales: short-term transients, medium-term rhythms, and long-term trends.
[0034] The first layer consists of a forward Long Short-Term Memory (LSTM) network and a backward LSM network operating in parallel. The forward LSM network progressively reads the first feature vector along the positive time direction of the charge-discharge cycle. At each time step, based on the current input and the hidden state of the previous step, it selectively retains, updates, and outputs information through a gating mechanism involving forget gates, input gates, and output gates, ultimately generating a positive hidden state sequence. The vector at each time step in this sequence is a compressed summary of all historical features preceding that step. Simultaneously, the backward LSM network progressively reads the first feature vector along the negative time direction of the charge-discharge cycle, tracing back from the end of the degenerate sequence to the beginning, generating a negative hidden state sequence. The vector at each time step in this sequence summarizes the state evolution pattern that will occur after that step. By concatenating the hidden states generated in the forward and reverse directions at each time step, the length of the output sequence is exactly the same as the input sequence. However, the feature dimension of each time step becomes the sum of the number of forward hidden units and the number of backward hidden units. In this way, the representation of each charge-discharge cycle moment is simultaneously endowed with the cumulative causes from the past and the trend predictions from the future. This allows the model to perceive the complete information chain of "how aging has progressed to this stage" and "where aging will go from this stage" when judging the health of a certain charge-discharge cycle.
[0035] The output sequence of the first layer becomes the input of the second layer. The second layer is also composed of a pair of forward long short-term memory networks and a backward long short-term memory network, with the same structure as the first layer. However, its task is no longer to process the original multi-scale convolutional branch outputs, but to perform higher-level abstraction based on the bidirectional temporal relationships established in the first layer. The forward network of the second layer, based on the bidirectional semantics already present at each time step of the first layer, continues to advance along time, gradually extracting higher-order dependency patterns across longer time intervals; the backward network operates in reverse, further establishing a deeper causal relationship between the state omens of the distant future and the current state. After the second layer encoding, the representation of each time step has the ability to characterize the macroscopic structural changes hidden across dozens or even hundreds of cycle spans during the degradation process.
[0036] In this model, the bidirectional long short-term memory network is stacked with three layers. The number of hidden units in each layer's forward long short-term memory unit is set to 128, and the number of hidden units in each layer's backward long short-term memory unit is also set to 128. Therefore, the feature dimension output after each layer is concatenated is 256. The three-layer stacking means that the original first feature vector undergoes three semantic refinements, from a concrete operational mode to an abstract degradation stage. The shallow layers are mainly responsible for perceiving short-term fluctuation patterns directly related to the input features; the middle layers begin to construct a phased degradation rhythm spanning several charge-discharge cycles; and the deep layers ultimately condense a long-range capacity decay trajectory spanning the entire sliding window history. To prevent gradient vanishing or exploding due to increased layer depth and sequence length, layer normalization is applied immediately after concatenating the forward and backward hidden states of each layer. This normalizes the activation values of the 256 hidden units at the same time step, adjusting their mean to zero and variance to one. This stabilizes the data distribution of each layer's output, ensuring that the gradient remains at a reasonable magnitude during backpropagation through each layer, thus guaranteeing the training stability of deep recurrent networks. After layer normalization, dropout regularization is applied with a dropout rate of 0.3%, randomly setting 30% of the hidden unit outputs to zero during training. This forces the network to make accurate inferences even with partially missing information, thereby suppressing over-reliance on specific time steps or feature dimensions and improving the generalization robustness of the encoding.
[0037] After progressive encoding through all three layers of bidirectional long short-term memory networks, the final output second feature vector maintains the same time dimension as the number of loops in the input window and has 256 dimensions in the feature dimension. Each time step in this vector sequence no longer merely represents the operating parameters of that charge-discharge cycle itself, but rather condenses all temporal relationships abstracted step-by-step across all stacked layers in two complete directions: from its historical starting point to the current moment and from the current moment to the end of the sequence. This forms a degenerate state sequence rich in multi-level contextual semantics. This sequence preserves the sequential structure between time steps, and the 256 feature channels of each time step have undergone sufficient interaction due to bidirectional loops and deep stacking. This provides a representational basis with both precise localization conditions and rich discriminative power for attention focusing in the time and feature dimensions in subsequent step 301.
[0038] This deep bidirectional long short-term memory (LSTM) encoding model is not trained as a standalone module, but is embedded in the complete end-to-end computation graph from the multi-scale temporal convolution in step 201 to the fully connected regression network in step 302 for joint optimization. During training, the convolution weights in step 201, all gating matrices and biases of the three-layer bidirectional LSM network of this model, the attention parameters in step 301, and the weights of the fully connected layers in step 302 simultaneously receive gradient updates from the loss function. The optimizer uses adaptive moment estimation, with an initial learning rate of 0.001, a first-order moment decay factor of 0.9, a second-order moment decay factor of 0.999, and a batch size of 32 or 64 depending on available computational resources. The loss function includes the mean squared error between the predicted health and the measured health, as well as the physical constraint penalty term set in step 302. When the validation set loss converges and no longer decreases, all parameters of the deep bidirectional LSM encoding are frozen. At this point, the model can automatically generate a complete and robust second feature vector for the charging and discharging history within any window during the inference phase.
[0039] The multi-scale temporal information carried by the first feature vector is fed layer by layer into a deep bidirectional recurrent encoding structure composed of forward and backward LSTMs stacked together. The forward LSTM reads the sequence step by step along the time progression direction of the charging cycle, compressing the running history before each cycle into a positive hidden state. The backward LSTM traces back from the degradation endpoint to the starting point in the reverse time direction, summarizing the state evolution that will occur after each cycle into a reverse hidden state. The two hidden states are concatenated at each time step, so that the representation of any charging and discharging moment no longer depends only on the previously accumulated historical conditions, but also obtains contextual supplementation from the future degradation direction. This bidirectional encoding enables the model to integrate the complete temporal vision of "how aging has developed to this point" and "where aging will go" when judging the current health status. Thus, for early weak signs such as the inflection point of accelerated degradation and capacity drop, the future state clues provided by the backward information can significantly enhance the recognition of the representation at that moment. The deeply stacked multi-layered structure abstracts this bidirectional temporal semantics step by step: the shallow layers mainly respond to the direct temporal correlation of short-term charging and discharging behaviors, while the deep layers gradually construct causal relationships between macroscopic degradation stages spanning dozens or even hundreds of cycles. The hidden states output by each layer have undergone a refinement from specific signals to abstract patterns compared to the inputs. Within each layer, normalization is applied to standardize the output distribution of each hidden unit at the same time step, constraining the amplitude of hidden states generated by different samples and at different time steps to a similar scale. This ensures that even under conditions of extremely long sequences and varied degradation stages, the gradients in the deep network can still propagate smoothly along both the forward and reverse paths without decaying or exploding, ultimately maintaining consistent coding quality across all time steps. The second feature vector obtained after passing through all LSTM layers has, while maintaining the original time step order, comprehensively expressed the local transients, mid-frequency prosody, and long-range trends extracted by multi-scale convolution into a high-order degenerative semantic rich in bidirectional long-range dependencies. The vector at each time step simultaneously contains contextual information of past accumulation and future trends, and the correlation between each feature dimension has been fully integrated. This provides a complete representation suitable for both time dimension focusing and feature dimension filtering for the next step of dual attention mechanism, enabling time attention to accurately locate the key cyclic intervals in the historical sequence that have the strongest explanatory power for capacity decay, while feature attention can distinguish the dominant parameters that truly carry degenerative information among the various dimensions.
[0040] Step 3: Perform dual attention fusion on the second feature vector and simultaneously perform fusion through adaptive gating to obtain the third feature vector. Input the third feature vector into a two-layer fully connected network and output it with health score as the label. Label the trained process as a health assessment model. Step 3 includes the following: Step 301: Perform dual attention fusion on the second feature vector. The dual attention fusion includes a temporal attention branch and a feature attention branch. The temporal attention branch obtains the temporal context vector by weighted summation of the hidden state data in the second feature vector through Softmax. The feature attention branch obtains the feature context vector by weighted summation of the hidden state data in each dimension through Squeeze-and-Excitation, with the number of dimensions of the hidden state data as the number of channels. The temporal context vector and the feature context vector are concatenated and input into a gating network. The gating network dynamically generates fusion weights to form adaptive gating fusion, which outputs a third feature vector.
[0041] like Figure 4 As shown, dual-attention fusion is a fusion structure composed of three parts: a temporal attention branch, a feature attention branch, and an adaptive gating network. Its goal is to automatically identify the most discriminative time position and physical parameter dimension for health prediction throughout the entire charging and discharging history, based on the deep bidirectional temporal semantics carried by the second feature vector. It then dynamically adjusts the fusion ratio of the two in a content-aware manner, outputting a highly condensed third feature vector. The input is the second feature vector output after encoding by a three-layer bidirectional long short-term memory network in step 202. This vector maintains the same length as the number of iterations in the sliding window in the temporal dimension and is 256-dimensional in the feature dimension. The 256-dimensional hidden state at each time step simultaneously contains the bidirectional long-range contextual information at that moment and high-order degenerate semantics after multiple layers of abstraction.
[0042] The model's temporal attention branch receives the complete second feature vector as input. This branch first inputs the second feature vector into a fully connected layer with an input dimension of 256 and an output dimension of 1. Its function is to compress the 256-dimensional hidden state at each time step into a scalar score, representing the original importance of that time step to the final prediction. The weight matrix of the fully connected layer is 256 rows and 1 column, with scalar biases. The weights are initialized using Ho's algorithm, and the biases are initialized to 0. The fully connected layer operates independently at each time step, outputting a score sequence of length equal to the number of window iterations. This score sequence is then fed into a Softmax function, which exponentially normalizes all scores along the time dimension, ensuring that the sum of the weights at all time steps equals 1. In the normalized weight sequence, time steps that contribute more to the health prediction are assigned weights closer to 1, while time steps that contribute less are assigned weights closer to 0. The original second feature vector is weighted and summed along the time dimension using this set of Softmax weights. Specifically, the 256-dimensional hidden states at each time step are multiplied by the corresponding weights and then summed, resulting in a fixed-dimensional time context vector of 256. This time context vector no longer carries the sequential information of the time steps. Instead, it selectively condenses information from all moments in the entire window history according to their relevance to health prediction, automatically highlighting key charge-discharge cycles that contain inflection points of accelerated capacity decline or abnormal operating modes, while suppressing irrelevant information introduced during long-term stable operation phases.
[0043] The model's feature attention branch also receives the complete second feature vector as input. This branch employs a compressed activation network structure, which operates in three steps. The first step is compression: global average pooling is performed on the second feature vector along the time dimension. This involves averaging the 256-dimensional hidden state values of each feature channel across all time steps, outputting a 256-dimensional global description vector. Each element in this vector represents the overall response intensity of the corresponding feature channel throughout the entire sequence, compressing all information from the time dimension and retaining only the global statistics of the channel dimension. The second step is activation: the 256-dimensional global description vector is input into a two-layer bottleneck fully connected network. The first fully connected layer of the bottleneck network has an input dimension of 256 and an output dimension of 16, compressing the original 256 dimensions to 16 dimensions, achieving dimensionality reduction and aggregation of information; the activation function is a modified linear unit. The second fully connected layer has an input dimension of 16 and an output dimension of 256, restoring the 16 dimensions to 256 dimensions, achieving dimensionality upscaling and restoration of information; the activation function is a sigmoid function, restricting the output values between 0 and 1, forming the activation weights for each of the 256 channels. Both fully connected layers use Ho's initialization for their weights, with biases initialized to 0. The third step is recalibration, where the 256-dimensional activation weights output by the Sigmoid function are multiplied channel-by-channel by the original second feature vector. Specifically, the first feature channel at each time step in the second feature vector is multiplied by the first activation weight, the second feature channel by the second activation weight, and so on. After recalibration, feature channels strongly correlated with health are assigned weights close to 1, thus being preserved or even enhanced, while redundant channels unrelated to degradation or containing noise are assigned weights close to 0, thus being suppressed. The recalibrated tensor is then subjected to global average pooling along the time dimension to compress the time-dimensional information, outputting a feature context vector with a fixed dimension of 256. This feature context vector does not focus on temporal sequence. Instead, it automatically identifies the feature channels corresponding to the physical parameters that truly play a dominant role in the battery aging process through a data-driven approach. For example, some channels may mainly respond to the cumulative effect of capacity decay, while others may mainly respond to the trend of increasing internal resistance. During training, the model adaptively adjusts the weights of the bottleneck network through gradient backpropagation, so that these key channels can obtain consistently high response weights.
[0044] After obtaining a 256-dimensional temporal context vector and a 256-dimensional feature context vector, the model concatenates them along the feature dimension to obtain a 512-dimensional pre-fusion vector. This 512-dimensional vector simultaneously contains a selective summary from the temporal dimension and a sensitive summary from the feature dimension. This 512-dimensional vector is fed into an adaptive gating network. The gating network is a two-layer fully connected structure. The first fully connected layer has an input dimension of 512 and an output dimension of 64, uses a modified linear unit as the activation function, and initializes the weights with Ho's algorithm and the biases to 0. The second fully connected layer has an input dimension of 64 and an output dimension of 512, uses a sigmoid activation function to constrain the output values between 0 and 1, and also initializes the weights with Ho's algorithm and the biases to 0. The gating network outputs a 512-dimensional fusion weight vector, which is multiplied element-wise with the concatenated 512-dimensional pre-fusion vector to form a weighted fusion vector, which also has a dimension of 512. At this point, each element in the weighted fusion vector is not simply equal to the original value in the temporal context vector or feature context vector, but rather the result of dynamic adjustment by the gating network based on the specific content of the current input: when the battery is in a stable linear decay stage, the information in the feature dimension contributes more to the prediction, and the gating network will automatically generate higher weight values for the 256 positions corresponding to the feature context vector; when the battery is in a transitional stage approaching the inflection point of accelerated decay, the before-and-after comparison information of the adjacent cycles in the temporal dimension is more critical for determining the inflection point position, and the gating network will correspondingly increase the weight values for the 256 positions corresponding to the temporal context vector. This adaptive mechanism makes the fusion strategy no longer a fixed linear combination, but a nonlinear fusion that is completely driven by the input content and dynamically changes with different aging stages and different operating conditions, avoiding the potential bias problem that fixed weights may cause in specific scenarios.
[0045] The weighted fusion vector is then processed by a fully connected layer for dimensionality reduction and integration. This fully connected layer has an input dimension of 512 and an output dimension of 256. The weights are initialized using the Ho's method, and the biases are initialized to 0. The final output is a third feature vector with a dimension of 256. This third feature vector is consistent with the second feature vector in terms of dimension, both being 256, and also maintains the same length in the time dimension as the number of window iterations. However, the information it carries has undergone a fundamental change: the original second feature vector is a bidirectional temporal encoding of multi-scale degradation features, with information distributed relatively evenly across various time steps and feature channels; while the third feature vector, through temporal focusing of time attention, channel filtering of feature attention, and adaptive balancing of the gating network, has highly concentrated the information on the few time positions and key feature dimensions that are most valuable for health prediction. Noise and redundancy are significantly compressed, and the signal-to-noise ratio and density of information are significantly improved.
[0046] This dual-attention fusion model, serving as an intermediate layer in the overall health assessment model, is not trained independently but is embedded in the end-to-end computation graph from the multi-scale temporal convolution in step 201 to the fully connected regression network in step 302 for joint optimization. During training, the weights and biases of the fully connected layers in the temporal attention branch, the two fully connected layers in the bottleneck network of the feature attention branch, the two fully connected layers in the gating network, and the weights and biases of the dimensionality reduction fully connected layers are all synchronously updated with gradients from the loss function, along with the kernel parameters of the front-end multi-scale temporal convolutional network, the gating matrix and biases of the mid-end deep bidirectional long short-term memory network, and the weights and biases of the back-end fully connected regression network. The optimizer employs adaptive moment estimation, with an initial learning rate of 0.001, a first-order moment decay factor of 0.9, a second-order moment decay factor of 0.999, and a batch size of 32 or 64 depending on available computational resources. The loss function includes the mean squared error between the predicted and measured health levels, as well as the physical constraint penalty term set in step 302. The weight allocation for temporal attention, the channel selection for feature attention, and the adaptive fusion ratio of the gating network are all learned automatically by data without any explicit rules during training. After training converges and the validation set loss no longer decreases, all parameters of the dual-attention fusion model are frozen. During the inference phase, the model can automatically perform temporal focusing, feature selection, and adaptive fusion operations on any input second feature vector, stably outputting a high-quality third feature vector. This provides the fully connected regression network in step 302 with a high signal-to-noise ratio, high information density, and a high degree of adaptation to the current degradation state as regression input.
[0047] The hidden state at each time step in the second feature vector carries both multi-scale degradation semantics and bidirectional long-range contextual information for that moment. Based on this, the temporal attention branch calculates the Softmax weight distribution for the hidden states of all time steps and performs a weighted summation, compressing the original variable-length, information-dispersed complete hidden state sequence into a fixed-dimensional temporal context vector. Each element in this vector is a comprehensive expression of the information of each time step in the original sequence after soft filtering according to importance. This automatically amplifies the encoding contribution of those charge-discharge cycle moments that have a decisive impact on SOH prediction, while suppressing redundant interference from long-term stable operation phases or noise intervals unrelated to degradation. Thus, in the temporal dimension, the originally implicit and difficult-to-quantify judgment problem of "which moments in the entire history are the most critical" is transformed into a weight allocation problem that can be directly calculated and optimized. The feature attention branch treats each feature dimension of the hidden state as an independent channel. It first performs global average compression on each channel through the Squeeze-and-Excitation mechanism to obtain the overall response intensity of that dimension across the entire sequence. Then, it models the dependencies between channels and generates activation weights for each dimension through a learnable fully connected layer. Finally, it recalibrates each feature dimension of the original hidden state according to the weights. The output feature context vector retains information from all time steps with the same dimensional structure, but strengthens the feature channels corresponding to key physical parameters that are highly correlated with the nature of capacity decay. At the same time, it weakens redundant dimensions that contribute little to the degradation modeling or even introduce interference. Thus, the prior knowledge of "which physical quantities play a dominant role in the degradation process" is transformed into a feature importance distribution that is automatically identified by data. The temporal context vector and feature context vector are concatenated and fed into the gating network. The gating network does not fuse the two in a fixed ratio, but dynamically generates fusion weights based on the specific content of the current input as the contribution of the temporal dimension and the feature dimension, respectively. This adaptive gating fusion mechanism can flexibly adjust the proportion of temporal focus and feature focus in the final representation according to the different aging stages and operating conditions of the battery: when the battery is in a stable linear decay stage, the importance of the feature dimension may be more prominent because there is a stable mapping between specific operating parameters and SOH; while when the battery is close to the inflection point of accelerated decay, the comparison of states near the cycle in the temporal dimension may be more discriminative than the feature value at a single moment, and the gating network can automatically increase the fusion weight of the temporal context vector.The third feature vector, generated by adaptive gating fusion after concatenation, dynamically balances the two complementary attention perspectives of "the most critical time period" and "the most sensitive physical parameter" in a content-aware manner at each prediction time. This forms a compact fusion representation that includes both accurate temporal localization and complete feature semantics. It provides a high-information-density, low-noise input that is highly adapted to the current prediction requirements for the subsequent fully connected regression network. This avoids the prediction bias caused by the fixed-ratio fusion neglecting a certain dimension in some scenarios. As a result, the entire health assessment model can adaptively adjust the attention strategy to obtain robust SOH estimates when facing different battery models, different charging and discharging strategies, and different aging modes.
[0048] Step 302: Input the third feature vector into a two-layer fully connected network and output a health scalar. At the same time, set constraints to form a penalty term. The constraints include a health score less than or equal to 1 and a health score greater than a preset health score change threshold in two adjacent charge-discharge cycles in the time series. The Adam optimizer is used to integrate the trained multi-scale temporal convolutional model, deep bidirectional LSTM encoding, dual attention fusion, and fully connected network into a health assessment model.
[0049] like Figure 5 As shown, after dual attention fusion, the third feature vector maintains a temporal length consistent with the number of sliding window iterations, with each time step representing a 256-dimensional high-level semantic representation. To map it to a single health scalar, global average pooling is first performed on the third feature vector along the temporal dimension, averaging the 256-dimensional features across all time steps and compressing them into a fixed 256-dimensional vector. This pooling operation gathers all key temporal and sensitive feature information focused by attention throughout the window history, forming a fixed-length descriptor whose length does not change with the number of iterations within the window, serving as the input to the fully connected regression network.
[0050] This fully connected regression network consists of two stacked fully connected layers. The first fully connected layer has an input dimension of 256 and an output dimension of 64. The weight matrix is initialized using the Ho's algorithm, and the bias is initialized to 0. It is followed by a modified linear unit activation function, which introduces nonlinearity into the feature transformation, enabling the network to fit the complex nonlinear mapping relationship between health status and multi-dimensional degradation features. The second fully connected layer has an input dimension of 64 and an output dimension of 1. The weights are also initialized using the Ho's algorithm, and the bias is initialized to 0. This layer does not have an activation function and directly outputs a continuous real value, which is the predicted battery health scalar. Its theoretical value range is between 0 and 1, representing the percentage of the current remaining battery capacity relative to the initial capacity.
[0051] During training, the loss function consists of two summed parts. The first part is the mean squared error between the predicted and measured health values, directly measuring regression accuracy and driving model parameters to update in the direction of reducing prediction error. The second part is a physical constraint penalty term, which consists of two constraints: when the model output health value exceeds 1, the square of the excess is added to the loss, forcing the predicted value to strictly not exceed the physical upper limit of 100%; simultaneously, for the same sliding window sample, if the predicted health values at two adjacent charge-discharge cycles on the time axis are compared, if the health value of the later cycle is higher than that of the previous cycle, or the decrease in health value between the two cycles exceeds the preset maximum allowable decay threshold, the square of the violation is added to the loss as a penalty term. This maximum allowable decay threshold is preset according to the battery model and aging characteristics, for example, set to 0.02, meaning that the decrease in health value between two adjacent charge-discharge cycles is not allowed to exceed 2%, thus transforming the irreversible and gradual evolution of battery aging into a calculable soft constraint, suppressing the model from producing non-physical prediction results such as capacity rebound or cliff-like jumps.
[0052] The parameters of the entire network, including the weights and biases of the two fully connected layers, are jointly trained end-to-end with the parameters of the front-end multi-scale temporal convolutional network, deep bidirectional long short-term memory network, and dual attention fusion module under an adaptive moment estimation optimizer. The initial learning rate of the optimizer is set to 0.001, the first moment decay rate is set to 0.9, the second moment decay rate is set to 0.999, and the training batch size is set to 32 or 64 depending on computational resources. During training, the parameters of all modules synchronously receive gradients from the composite loss function, without the need for staged pre-training or individual tuning. Training terminates when the total loss on the validation set converges and no longer decreases, and the proportion of predicted health sequences that meet the physical constraints reaches a preset requirement.
[0053] At this point, all computational steps in the entire data flow—from sliding window input to health scalar output—are solidified into an end-to-end health assessment model. This model receives normalized data, which undergoes multi-scale temporal convolution to extract multi-resolution features. This data is then encoded into bidirectional temporal semantics via a deep bidirectional long short-term memory network, and further focused on key time periods and sensitive parameters through dual attention fusion. Finally, global average pooling and two fully connected layers are used to regress the data into a health value. After receiving normalized windowed data, this model can complete feature extraction, temporal encoding, attention fusion, and regression prediction in a single step without any intermediate manual intervention. It directly outputs a battery health assessment value that conforms to physical priors, thus providing continuous, stable, and physically consistent health status estimates with high integration and low inference latency in engineering scenarios such as vehicle battery management or energy storage system monitoring.
[0054] The two-layer fully connected network compresses and maps the third feature vector, which integrates temporal focus and feature sensitivity information, into a scalar SOH value step by step, completing the regression transformation from a high-dimensional degenerate semantic space to a single-dimensional health index. This structure converges the complex temporal modeling results into a predictive output that can directly represent the remaining battery capacity in a compact parameterized manner, ensuring that the final output of the entire evaluation process is a set of clear and operable numerical results, rather than an intermediate representation that still requires manual interpretation. During the optimization process, by setting a physical upper limit constraint that the absolute value of the health status does not exceed 1 and a smoothness constraint that the rate of health status decay between adjacent charge and discharge cycles does not exceed a preset threshold, and by adding the deviation of violating these constraints as a penalty term to the loss function, it is equivalent to injecting the inherent irreversibility and asymptotic prior knowledge of the battery aging process into the data-driven gradient optimization direction. This forces the model to follow the basic physical law that capacity can only decrease monotonically and will not experience a sharp drop in the short term while adjusting parameters to fit the training data. This effectively suppresses output anomalies that violate physical common sense, such as predicted values above 1.0 or unreasonable capacity rebounds and instantaneous cliff drops between adjacent cycles caused by local data noise or model overfitting. This ensures that the predicted degradation trajectory has physical authenticity and engineering credibility in terms of numerical value. The Adam optimizer utilizes its adaptive estimation of first and second moments to correct the bias and adjust the step size of gradients for different parameters during training. This allows parameter updates to smoothly traverse the flat regions and sharp valleys of the complex loss surface when simultaneously optimizing the regression error term and the constraint penalty term in a multi-objective loss model. This accelerates convergence and reduces the risk of getting trapped in local optima, thereby obtaining a model parameter configuration that simultaneously achieves data fitting accuracy and physical constraint satisfaction within a limited number of training iterations. After training, the multi-scale temporal convolutional network, deep bidirectional LSTM encoder, dual attention fusion module, and two-layer fully connected regression network are solidified into an end-to-end integrated health assessment model. This model no longer needs to perform feature extraction, sequence encoding, attention fusion, and regression prediction in stages. Instead, it can directly receive the normalized sliding window input and output the corresponding health prediction value at once. This eliminates the interface alignment error and data transmission delay that may be introduced when deploying multiple modules separately. In engineering scenarios with limited resources and response time, such as in-vehicle BMS or cloud monitoring systems, it significantly improves the inference efficiency and deployment convenience of online assessment. At the same time, the unified end-to-end structure also provides a complete and unbroken computational graph for the synchronous extraction of weight coefficients of each layer and the Monte Carlo Dropout uncertainty estimation in step 4.
[0055] Step 4: During the training of the health assessment model, obtain the weight coefficients for each training process, train synchronously in Monte Carlo Dropout, obtain the confidence score of each health assessment model output, and set the confidence interval. Step 4 includes the following: During the training of the health assessment model, the weight coefficients generated by the scale-time convolutional model, deep bidirectional LSTM encoding, dual attention fusion and fully connected network during operation are obtained and summarized. The summarized weight coefficients and input feature set are used to obtain the confidence of each output result through Monte Carlo Dropout. The confidence of each result and the health score of the health assessment model are output simultaneously. Define a confidence interval, which includes high reliability, medium reliability, and low reliability.
[0056] During the training phase of the health assessment model, multiple layers have a built-in Dropout regularization mechanism. Specifically, in the deep bidirectional long short-term memory network in step 202, each layer has a Dropout layer with a dropout rate of 0.3 after layer normalization. Similarly, in the bottleneck fully connected network of the feature attention branch and the fully connected layers of the adaptive gating network in step 301, a Dropout layer with a dropout rate of 0.2 is set after the activation function. In the two fully connected regression network in step 302, a Dropout layer with a dropout rate of 0.2 is also set after the activation function of the modified linear unit in the first fully connected layer. During the forward propagation of each training iteration, these Dropout layers randomly set the output of a portion of neurons to zero according to the set dropout probability. The neurons with zeroed output do not participate in forward computation or receive gradient updates in that iteration, thus forcing the network to learn feature representations that are robust to the loss of any subset of neurons.
[0057] Once the health assessment model has converged, its parameters are saved. In step 4, the weight coefficients of each module recorded during each iteration of training are summarized. These weight coefficients fully describe the parameter state of the model at training convergence. The core operation of Monte Carlo Dropout is that during the inference phase, the activated states of all Dropout layers are maintained for the already trained model. That is, the Dropout layers are not switched to identity mappings during inference, but neurons are still randomly dropped with the same dropout probability as during training. For the same input sample, multiple independent forward propagation processes are performed. The subset of neurons randomly dropped in each forward propagation changes due to the different random seeds, resulting in slight differences in the actual effective computation graph for each propagation. Therefore, the health prediction value output by each propagation is also slightly different. This number of multiple samplings is set to 100, that is, during inference, for each input window, the model performs 100 forward propagations with random Dropout, resulting in 100 health prediction values.
[0058] These 100 predicted values constitute a health prediction distribution for the current input. A more concentrated distribution indicates a high degree of consistency in the predictions given by different neural pathways within the model, suggesting strong certainty and high reliability in the model's predictions. Conversely, a more dispersed distribution indicates that the input features lie in the fuzzy boundary region between the decision boundaries learned by different pathways, suggesting significant internal discrepancies and inherent uncertainty in the predictions. The statistical standard deviation of the 100 predicted values is used as a confidence measure for this prediction; a smaller standard deviation indicates higher confidence, and a larger standard deviation indicates lower confidence. Furthermore, the mean of the 100 predicted values is taken as the final health output value. Using the mean further smooths out any extreme biases that might be introduced by a single random dropout, making the output health point estimate more robust.
[0059] Based on this, two confidence thresholds are set to divide the prediction results into three confidence intervals. When the prediction confidence (standard deviation) is less than or equal to 0.03, the prediction is classified into the high-reliability interval, indicating that the model's prediction distribution is highly concentrated, and this health assessment value can be directly adopted as a decision-making basis. When the prediction confidence is greater than 0.03 but less than or equal to 0.07, it is classified into the medium-reliability interval, indicating that the prediction has a certain degree of dispersion, but is still within an acceptable range. It is recommended to use it after cross-validation with adjacent battery data or historical records. When the prediction confidence is greater than 0.07, it is classified into the low-reliability interval, indicating that there is a large discrepancy between multiple predictions of the model. This health estimate cannot be directly used as a decision-making basis and requires further manual review or offline precise detection.
[0060] During the training of the health assessment model, step 4 synchronously records the weight coefficients of each module in each iteration and performs the Monte Carlo Dropout sampling process described above on the validation or test set to calibrate the confidence interval threshold settings. After training is complete, the complete deployment package of the health assessment model includes not only the network parameters and forward computation graph fixed in steps 201 to 302, but also the calibrated confidence interval thresholds and the Monte Carlo Dropout sampling number settings. During actual inference, after receiving new battery operation data and organizing it into an input feature set in step 5, the health assessment model automatically performs 100 forward propagations with Dropout, synchronously outputting the health prediction value, confidence score, and the corresponding high, medium, or low reliability confidence interval levels. The entire process requires no additional manual parameter adjustment or post-processing modules, achieving end-to-end automation from raw data to health assessment with confidence labels.
[0061] In each iteration of the health assessment model training, the weights and bias coefficients generated from each layer of the multi-scale temporal convolutional network, deep bidirectional LSTM, dual attention fusion module and fully connected regression network are systematically collected and summarized. This operation unifies the trainable parameters that were originally scattered in different modules and at different depths into a complete parameter distribution view, so that all parameter combinations and their value patterns called by the model in one inference can be recorded and tracked, providing a complete parameter state snapshot for subsequent evaluation of prediction uncertainty. The aggregated weight coefficients are fed together with the current input feature set into the Monte Carlo Dropout mechanism. During the inference phase, the Dropout layer remains active and undergoes multiple independent forward propagations. Each propagation causes minor changes in the network's effective topology and parameter combinations due to the random discarding of different subsets of neurons. This generates a set of multiple SOH predictions around the same input. The dispersion of these predictions directly reflects the model's certainty regarding the prediction results under the current input and parameter states: when the SOH outputs from multiple forward propagations are highly consistent, it indicates that the model has reached a high degree of consensus on the prediction conclusions of each neural pathway for that input, and the prediction results are highly reliable; conversely, when the multiple outputs differ significantly, it indicates that some patterns in the input features are near the decision boundaries learned by different pathways, indicating significant internal disagreement within the model and inherent uncertainty in the prediction results. By statistically summarizing these predicted values to calculate the confidence level of each prediction, the deterministic model, which could only output a single SOH point estimate, is transformed into a probabilistic model that can simultaneously output the predicted value and the confidence level of that prediction. This overcomes the fundamental deficiency of point estimation in safety-critical applications, which cannot inform users "how reliable this prediction is." Based on this, three confidence intervals—high reliability, medium reliability, and low reliability—are set, mapping the confidence level of continuous values to directly interpretable discrete reliability levels. This allows battery management systems or maintenance personnel to immediately determine whether the current health assessment result should be adopted, cautiously considered, or further testing should be triggered without analyzing complex probability distributions. Thus, the quantification of uncertainty is transformed from theoretical indicators into actionable, tiered decision-making criteria. This ensures that differentiated handling strategies can be formulated based on reliability levels in different application scenarios such as battery reuse screening, in-service safety monitoring, and retirement timing determination, balancing the refinement of the assessment with the practical simplicity of engineering implementation.
[0062] Step 5: Organize the actual charging and discharging operation data of the battery into an input feature set and input it into the health evaluation model. At the same time, use the weight coefficients summarized by the health evaluation model and the input features to obtain the confidence intervals corresponding to the health level and confidence level, and output them.
[0063] As a preferred embodiment, Table 1 below provides a comparison of the health scores output by the health assessment models and the health score data output by the BMS system, as well as the confidence intervals corresponding to the health scores output by each health assessment model: Table 1 Data Comparison
[0064] Based on the data comparison in Table 1 above and the corresponding... Figure 6 The dotted-line graph shows that the health assessment model is more stable and more in line with the actual situation in terms of health status. In addition, when there are obvious errors in the health status output by the BMS, such as when the number of charging cycles is 450 and the health status output by the BMS system increases, the health assessment model can still identify the correct health status and maintain a moderately reliable confidence level.
[0065] During actual deployment, operational data such as voltage, current, temperature, and capacity changes generated after each charge-discharge cycle of the battery are processed in real time using a sliding window partitioning strategy and normalization parameters identical to those used in the training process. This forms an input feature set with the same distribution and structure as during model training. This strict consistency in data preprocessing ensures that the distribution drift from offline training to online inference is minimized, ensuring that the input received by the model under real operating conditions is highly consistent with the feature space learned during the training phase, avoiding prediction bias caused by differences in data preprocessing. After receiving this input feature set, the health assessment model utilizes the weight coefficients of each module solidified in step 4 and the synchronously recorded Monte Carlo Dropout multiple forward propagation mechanism to produce a set of SOH prediction values and their statistical distribution in parallel during a single inference. This simultaneously yields three levels of output information: health point estimation, predicted confidence value, and corresponding confidence interval level. All capabilities built during the offline training phase, such as multi-scale feature extraction, bidirectional temporal coding, dual attention fusion, physical constraint regression, and uncertainty quantification, are released to the online application scenario at once, without requiring any additional external analysis modules or post-processing steps. The method of directly outputting confidence interval levels transforms complex probability distributions and confidence values into three levels of reliability—high, medium, or low—that the battery management system can directly interpret. This allows on-site maintenance decisions to immediately know the reliability of the prediction while obtaining the SOH prediction value. When the model outputs a high reliability level, the system can directly use this health value for remaining life estimation or tiered utilization decisions. When a medium reliability level appears, the system can adopt the result after cross-validation with historical records or adjacent battery data. When a low reliability level is triggered, the system can automatically mark it as requiring manual review or trigger more refined offline detection. This forms a closed-loop automated evaluation mechanism from data acquisition to prediction to graded response. This method integrates health assessment and confidence judgment in the same forward inference process, which significantly shortens the delay from acquiring charge and discharge data to forming an executable assessment conclusion. This enables lithium-ion batteries to have engineering-feasible deployment efficiency in scenarios that require real-time or near-real-time health monitoring, such as energy storage power stations and electric vehicles. At the same time, the setting of three-level confidence intervals also provides a direct calculation basis for the automated hierarchical response of operation and maintenance strategies, reducing the risk of misjudgment and decision-making errors that may be caused by the lack of reliability judgment in single numerical prediction.
[0066] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0067] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0068] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0069] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for assessing the health status of lithium-ion batteries based on physical information and data-driven approaches, characterized in that, The specific steps include: Step 1: Cycle charge and discharge the sample battery and record the running data, and record the health status at the same time. Correspond the running data and the health status to form a charge and discharge sequence, and set a sliding window to form an input feature set in the charge and discharge sequence; Step 2: Input the input feature set into the multi-scale temporal convolution model to obtain the first feature vector of several duration features. Encode the first feature vector using deep bidirectional LSTM to obtain the second feature vector. Step 3: Perform dual attention fusion on the second feature vector and simultaneously perform fusion through adaptive gating to obtain the third feature vector. Input the third feature vector into a two-layer fully connected network and output it with health score as the label. Label the trained process as a health assessment model. Step 4: During the training of the health assessment model, obtain the weight coefficients for each training process, train synchronously in Monte Carlo Dropout, obtain the confidence score of each health assessment model output, and set the confidence interval. Step 5: Organize the actual charging and discharging operation data of the battery into an input feature set and input it into the health evaluation model. At the same time, use the weight coefficients summarized by the health evaluation model and the input features to obtain the confidence intervals corresponding to the health level and confidence level, and output them.
2. The lithium-ion battery health status assessment method based on physical information and data-driven approach according to claim 1, characterized in that: The sample batteries were subjected to cyclic charging and discharging. The operating data of each battery was recorded during each charging and discharging cycle. The health status of each battery after each charging and discharging cycle was measured. The health status is the SOH value. The operating data experienced by each health status was recorded and sorted according to the order of acquisition time during the charging and discharging cycle, forming a charging and discharging sequence with the operating data as the sample and the corresponding health status as the label. Set the time window length and step size, slide the data on the charge-discharge sequence to obtain the input feature set, and normalize the data in the input feature set respectively.
3. The lithium-ion battery health status assessment method based on physical information and data-driven approach according to claim 2, characterized in that: The input feature set is input into a multi-scale temporal convolution model, which contains three convolutional branches. Each convolutional branch has a different kernel size, corresponding to short-term, medium-term, and long-term feature extraction, respectively. All convolutional branches have the same number of output channels. The outputs of the three convolutional branches are combined to form the first feature vector.
4. The lithium-ion battery health status assessment method based on physical information and data-driven approach according to claim 3, characterized in that: The first feature vector is encoded using a deep bidirectional LSTM, which has several layers. Each layer contains a forward LSTM and a backward LSTM. The hidden state distribution is stabilized by applying layer normalization, and the hidden states in the two directions are concatenated as the output of the layer. The input of the first layer is the first feature vector, and the input of each subsequent layer is the output of the previous layer, thus obtaining the second feature vector.
5. The lithium-ion battery health status assessment method based on physical information and data-driven methods according to claim 4, characterized in that: The second feature vector is subjected to dual attention fusion, which includes a temporal attention branch and a feature attention branch. The temporal attention branch obtains the temporal context vector by weighted summation of the hidden state data in the second feature vector through Softmax. The feature attention branch obtains the feature context vector by weighted summation of the hidden state data in each dimension through Squeeze-and-Excitation, with the number of hidden state data dimensions as the number of channels. The temporal context vector and the feature context vector are concatenated and input into a gating network. The gating network dynamically generates fusion weights to form adaptive gating fusion, which outputs a third feature vector.
6. The lithium-ion battery health status assessment method based on physical information and data-driven approach according to claim 5, characterized in that: The third feature vector is input into a two-layer fully connected network, and the output is a health scalar. At the same time, constraints are set to form a penalty term. The constraints include a health score of less than or equal to 1 and a health score of two adjacent charge-discharge cycles in the time series that is greater than a preset health score change threshold. The Adam optimizer is used to integrate the trained multi-scale temporal convolutional model, deep bidirectional LSTM encoding, dual attention fusion, and fully connected network into a health assessment model.
7. The lithium-ion battery health status assessment method based on physical information and data-driven approach according to claim 6, characterized in that: During the training of the health assessment model, the weight coefficients generated by the scale-time convolutional model, deep bidirectional LSTM encoding, dual attention fusion and fully connected network during operation are obtained and summarized. The summarized weight coefficients and input feature set are used to obtain the confidence of each output result through Monte Carlo Dropout. The confidence of each result and the health score of the health assessment model are output simultaneously. Define a confidence interval, which includes high reliability, medium reliability, and low reliability.
8. The lithium-ion battery health status assessment method based on physical information and data-driven approach according to claim 7, characterized in that: During the charging and discharging process of the battery, the operating data is recorded and organized into an input feature set, which is then input into the health evaluation model to obtain the health level and its corresponding confidence level. Based on the confidence level, the corresponding confidence interval is matched and combined with the health level for output.
Citation Information
Patent Citations
Lithium ion battery health degree estimation method
CN120779255A