Battery energy storage system management using kolmogorov-arnold networks
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-05
- Publication Date
- 2026-08-13
Smart Images

Figure EP2026053070_13082026_PF_FP_ABST
Abstract
Description
[0001] 180028
[0002] Battery energy storage system management using Kolmogorov-Arnold Networks TECHNICAL FIELD
[0003] The invention is in the field of battery energy storage system (BESS) management.
[0004] BACKGROUND
[0005] In energy industry, the algorithm design of a Battery Energy Storage System (BESS) is essential to achieve energy transition. Artificial intelligence (Al) and energy integration highlight a remarkable development in beneficial applications such as diagnostics, charging management, predictive maintenance, state estimation, and Remaining Useful Lifetime (RUL). Different approaches are required to implement optimal technologies of energy systems, focusing on modeling and digitalization, with the electrochemical model, Equivalent Circuit Model (ECM), and Mathematical models being the most important to simulate, monitor, and predict the health and charge indicators of a BESS.
[0006] Not only Data Science techniques to deliver advances on several strategic topics such as digital twins, enterprise testing, and business consulting have been proposed, but also Deep Learning algorithms that explain the physics of a system and the behavior of Key Performance Indicators (KPIs) through Al-powered technology, all to support the experience provided by energy analysts. In building and operating a BESS, the State of Charge (SOC) and State of Health (SOH) are useful for improving the performance of renewable technologies, making algorithm design important for monitoring the efficiency of electrical machines and autonomous systems. Experimental measurements and battery testing track the charging and discharging, influenced by chemical structure, manufacturing characteristics, and actual usage, which are requested by user needs, industrial applications, and research goals. Accurate model evaluation is important for addressing the challenges of energy transition, ensuring effective performance monitoring, and supporting industrial procedures.
[0007] During the execution of Al methods in a BESS domain, there is an opportunity to optimize the architecture of a Neural Network (NN), thereby reducing the computational complexity in Deep Learning algorithms. However, when collecting different types of datasets and designing NNs, a challenge lies in explaining the physical and virtual entities of a BESS. This opacity regarding the internal mechanisms of the battery presents a significant obstacle when seeking to understand KPIs and user needs.
[0008] Providing better methods of battery management remains a challenge.
[0009] SUMMARY
[0010] According to a first aspect, the present disclosure provides a battery management system comprising a Kolmogorov-Arnold Network (KAN) model for real-time state estimation and predictive analytics, wherein the battery management system is configured to use the KAN model to dynamically update control parameters to optimize battery performance of a battery energy storage system.
[0011] Thus it will be seen that, by providing a battery management system utilizing a KAN model configured to update control parameters to optimize battery performance of a battery energystorage system, the battery management system is able to control the performance of a battery energy storage system by predicting real-time state of the battery energy storage system.
[0012] In some embodiments of the battery management system, the KAN model utilizes learnable activation functions.
[0013] In some embodiments of the battery management system, the KAN model is implemented using a combination of PyTorch and Keras and / or TensorFlow APIs for adaptability across different battery chemistries and industrial scenarios. This may advantageously provide for the monitoring of realtime interaction, retrieving or updating live data from external systems (e.g. the battery energy storage system).
[0014] According to a further aspect, the present disclosure provides a method for battery energy storage system (BESS) management using a KAN model, wherein the method comprises:
[0015] collecting experimental battery datasets under various operating conditions;
[0016] training the KAN model using univariate function decomposition for estimating one or more battery performance indicators;
[0017] using the KAN model to predict State of Charge (SOC) and / or State of Health (SOH) and / or Open Circuit Voltage (OCV) for the BESS; and
[0018] using the predicted State of Charge and / or State of Health to optimize charging management of the BESS by adjusting one or more control parameters based on real-time data.
[0019] In some embodiments, the method of BESS management comprises using the KAN model to predict Remaining Useful Lifetime (RUL) of battery cells of the BESS using historical charge and discharge cycles.
[0020] In some embodiments of the method of BESS management, the KAN model is integrated with one or more Equivalent Circuit Models for SOC and / or SOH estimation.
[0021] In some embodiments, the method of BESS management comprises using the KAN model to predict both State of Charge and State of Health for the BESS.
[0022] In some embodiments, the method of BESS management comprises using the KAN model to predict State of Charge and State of Health and Open Circuit Voltage for the BESS.
[0023] According to a further aspect, the present disclosure provides a method of optimizing charging profiles in a BESS, comprising using a KAN-based predictive model to adjust one or more charging parameters in response to real-time battery performance data.
[0024] In some embodiments, the method of optimizing charging profiles comprises using the KAN-based predictive model to adjust a C-rate and / or cut-off voltage and / or SOC limit in response to real-time battery performance data.
[0025] According to a further aspect, the present disclosure provides a method of machine learning-driven predictive maintenance for a battery system, the method comprising using a KAN model to detect performance degradation and / or failure conditions in the battery system.The present disclosure also provides a processing system configured to perform any embodiment of: the method of BESS management, the method of optimizing charging profiles, or the method of machine learning-driven predictive maintenance for a battery system.
[0026] The present disclosure also provides a software product comprising instructions that, when executed by a processing system, cause the processing system to perform any embodiment of: the method of BESS management, the method of optimizing charging profiles, or the method of machine learning-driven predictive maintenance fora battery system.
[0027] The present disclosure also provides a computer-readable medium storing instructions that, when executed by a processing system, cause the processing system to perform any embodiment of: the method of BESS management, the method of optimizing charging profiles, or the method of machine learning-driven predictive maintenance fora battery system.
[0028] The present disclosure relates to an artificial intelligence-based approach for optimizing Battery Energy Storage System (BESS) management, utilizing Kolmogorov-Arnold Networks (KANs). At least some embodiments of the disclosure provide an improved methodology for state estimation, and / or Remaining Useful Lifetime (RUL) prediction, and / or charging management of battery cells, which may significantly enhance the accuracy, and / or efficiency, and / or scalability of energy storage system monitoring and control. To address the limitations of existing neural network-based approaches, embodiments of the disclosure use KANs, which may leverage univariate function decomposition to improve model interpretability and / or reduce computational complexity. Unlike traditional deep learning methods such as Transformer Networks, Recurrent Neural Networks (RNNs), Convolutional Neural Networks (CNNs), and Multilayer Perceptrons (MLPs), KANs dynamically adjust activation functions during training, which may lead to superior adaptability and reduced training requirements. As will be established below, KANs offer better results when applied to electrochemical systems than Transformer Networks and the other NNs discussed herein. In particular, KANs provide a resilient framework for decomposing non-linear energy landscapes. This may advantageously enable a deeper understanding of critical battery processes, including charge transport, reaction kinetics, and degradation mechanisms.
[0029] The disclosure includes the following important components:
[0030] State Estimation: Some embodiments integrate KANs with Equivalent Circuit Models (ECMs) to estimate the State of Charge (SOC) and State of Health (SOH) of battery cells. The trained model may utilize real-world datasets obtained from experimental battery testing, ensuring high predictive accuracy across various operating conditions. In state estimation, prediction complexity can increase significantly due to noisy inputs and non-linear relationships; in this context, KANs may offer robustness by decomposing complex functions, isolating the influence of each input variable, and adapting to the various non-linearities in the data through univariate functions.
[0031] Remaining Useful Lifetime (RUL) Prediction: Some embodiments employ KAN-based models to analyze historical charge-discharge cycles and predict the remaining service life of batteries. This may enable efficient battery lifecycle management and early detection of potential failures. In the context of RUL prediction, KANs are well-suited for additive models, where the output is the sum of individual contributions from each input, such as battery capacity across cycles.Charging Management Optimization: Some embodiments optimize charging profiles by adjusting control parameters such as C-rate, cut-off voltage, and SOC limits based on real-time battery performance data. The use of KANs may ensure optimal charging strategies while preventing overcharging and degradation. KAN-based models may be used in some embodiments to estimate open circuit voltage (OCV).
[0032] Machine Learning-Driven Predictive Maintenance: In some embodiments, the presently-disclosed Al-driven framework enables real-time monitoring and early detection of performance degradation in BESS. By implementing predictive maintenance, operators can proactively address potential failures, thereby extend battery life and reduce operational costs.
[0033] Integration with Industrial Energy Storage Systems: Some embodiments facilitate Al-driven battery analytics integration into large-scale energy storage solutions. The KAN-based deep learning models may enhance decision-making and energy optimization for grid-scale and industrial BESS applications.
[0034] Efficient Computational Implementation: Some embodiments leverage TensorFlow and / or PyTorch and / or Keras APIs for implementing the KAN models, ensuring adaptability across different battery chemistries and industrial use cases. By minimizing computational complexity, some embodiments may significantly reduce the processing power required for real-time analytics and decision-making. Computer-Readable Medium Implementation: Some embodiments include a computer-readable medium storing instructions that, when executed by a processor, enable the system to perform any one or more of the methods disclosed herein. This may allow seamless integration with battery management systems and industrial software platforms.
[0035] The present disclosure introduces a transformative approach to battery management by employing KANs for improved state estimation, and / or predictive maintenance, and / or charging optimization. By overcoming the constraints of traditional deep learning methods, some embodiments of the disclosure may enable more accurate, and / or scalable, and / or energy-efficient solutions for nextgeneration BESS applications.
[0036] A further aspect of the present disclosure provides a method of analyzing health or performance of a battery energy storage system, wherein the method comprises:
[0037] providing empirical data, determined empirically from a battery energy storage system, as input to a trained Kolmogorov-Arnold network, implemented by a processing system, wherein the trained Kolmogorov-Arnold network is configured to model nonlinear dynamics of battery energy storage systems; and
[0038] operating the trained Kolmogorov-Arnold network to determine, as output data, one or more properties of the battery energy storage system indicative of a health or performance of the battery energy storage system.
[0039] In some embodiments, the method further comprises using the output data to control a charging or discharging of the battery energy storage system.
[0040] In some embodiments, the method further comprises using the output data to determine when to perform a maintenance action on the battery energy storage system.In some embodiments, the method further comprises generating the empirical data by measuring one or more physical parameters of the battery energy storage system. In some of these embodiments, the method may further comprise:
[0041] generating the empirical data continually or at intervals overtime; and
[0042] while generating the empirical data, operating the trained Kolmogorov-Arnold network to determine the one or more properties continually or at intervals over time, and using the output data to control a charging or discharging of the battery energy storage system continually or at intervals over time. In some of these embodiments, the method may further comprise controlling the charging or discharging of the battery energy storage system comprises adjusting a C-rate or cut-off voltage. In some embodiments, the empirical data comprises at least one electrical or thermal or temporal parameter.
[0043] In some embodiments, the empirical data comprises at least one parameter selected from the group consisting of: resistance, current, voltage, charge capacity, discharge capacity, temperature, and time.
[0044] In some embodiments, the empirical data comprises data collected over a plurality of charge and discharge cycles.
[0045] In some embodiments the one or more properties comprises: a state of charge (SOC), or a state of health (SOH), ora remaining useful lifetime (RUL), ora charge capacity, ora discharge capacity, or an open circuit voltage (OCV).
[0046] In some embodiments, the trained Kolmogorov-Arnold network comprises a set of univariate functions that has been trained to represent nonlinear dynamics of battery energy storage systems. In some embodiments, the trained Kolmogorov-Arnold network has been trained using battery cycling data determined empirically for one or more battery energy storage systems (which may optionally be of a same type or a similar type to the BESS whose health or performance is being analyzed).
[0047] In some embodiments, the method further comprises generating the trained Kolmogorov-Arnold network by training a Kolmogorov-Arnold network to model nonlinear dynamics of battery energy storage systems using battery cycling data determined empirically for one or more battery energy storage systems.
[0048] In some embodiments, the processing system comprises one or more processors and a storage medium, and wherein the storage medium stores software instructions that, when executed by the one or more processors, cause the processing system to implement and operate the trained Kolmogorov-Arnold network.
[0049] Another aspect of the disclosure provides a method of generating a trained Kolmogorov-Arnold network for use in analyzing health or performance of a battery energy storage system, wherein the method comprises:
[0050] accessing battery cycling data determined empirically for one or more battery energy storage systems; andusing the battery cycling data, by a processing system, to train a Kolmogorov-Arnold network to model nonlinear dynamics of battery energy storage systems for determining one or more properties of a battery energy storage system indicative of a health or performance of the battery energy storage system.
[0051] Another aspect of the disclosure provides an apparatus for analyzing health or performance of a battery energy storage system is provided, wherein the apparatus comprises:
[0052] a processing system; and
[0053] a memory system,
[0054] wherein the memory system is configured to store empirical data, determined empirically from a battery energy storage system, and
[0055] wherein the processing system is configured to:
[0056] read the empirical data from the memory system;
[0057] provide the empirical data as input to a trained Kolmogorov-Arnold network, implemented by the processing system, wherein the trained Kolmogorov-Arnold network is configured to model nonlinear dynamics of battery energy storage systems;
[0058] operate the trained Kolmogorov-Arnold network to determine, as output data, one or more properties of the battery energy storage system indicative of a health or performance of the battery energy storage system.
[0059] In some embodiments of the apparatus, the apparatus is a battery management system for the battery energy storage system, and is configured to use the output data to control a charging or discharging of the battery energy storage system.
[0060] In some embodiments, the apparatus is configured to use the output data to determine when to perform a maintenance action on the battery energy storage system, and to signal an operator to perform the maintenance action.
[0061] In some embodiments, the apparatus further comprises the battery energy storage system.
[0062] Another aspect of the disclosure provides a non-transitory computer-readable medium storing software instructions that, when executed by a processing system comprising one or more processors, cause the processing system to:
[0063] implement a trained Kolmogorov-Arnold network configured to model nonlinear dynamics of battery energy storage systems;
[0064] provide empirical data, determined empirically from a battery energy storage system, as input to the trained Kolmogorov-Arnold network; and
[0065] operate the trained Kolmogorov-Arnold network to determine, as output data, one or more properties of the battery energy storage system indicative of a health or performance of the battery energy storage system.
[0066] According to another aspect, the present disclosure provides a method for battery energy storage system (BESS) management using Kolmogorov-Arnold Networks (KANs), comprising:
[0067] o Collecting experimental battery datasets under various operating conditions;o Training a KAN model using univariate function decomposition to estimate battery performance indicators;
[0068] o Predicting State of Charge (SOC) and State of Health (SOH) with enhanced accuracy;
[0069] o Optimizing charging management by adjusting control parameters based on realtime data.
[0070] In some embodiments of the method, the KAN-based algorithm is used for predicting Remaining Useful Lifetime (RUL) of battery cells using historical charge and discharge cycles.
[0071] In some embodiments of the method, the KAN-based algorithm integrates with Equivalent Circuit Models (ECMs) to improve the precision of SOC and SOH estimation.
[0072] According to another aspect, the present disclosure provides a battery monitoring system incorporating a Kolmogorov-Arnold Network (KAN) for real-time state estimation and predictive analytics, wherein the system dynamically updates control parameters to optimize battery performance.
[0073] In some embodiments of the battery monitoring system, the KAN model reduces computational complexity by utilizing learnable activation functions to replace traditional neural network weight parameters.
[0074] In some embodiments of the battery monitoring system, the KAN-based model is implemented using a combination of PyTorch and Keras APIs to enhance adaptability across different battery chemistries and industrial scenarios.
[0075] According to another aspect, the present disclosure provides a method for optimizing charging profiles in a BESS, wherein a KAN-based predictive model adjusts C-rate, cut-off voltage, and other charging parameters in response to real-time battery performance data.
[0076] According to another aspect, the present disclosure provides a method for implementing machine learning-driven predictive maintenance in battery systems, wherein the KAN-based approach enables early detection of performance degradation and failure conditions.
[0077] According to another aspect, the present disclosure provides a software product integrating Al-driven battery analytics with industrial energy storage solutions according to any of the embodiments or aspects described above, wherein KAN-based deep learning models enhance decision-making for large-scale BESS applications.
[0078] According to a sixth aspect, the present disclosure provides a computer-readable medium storing instructions that, when executed by a processor, cause the system to perform the methods according to the embodiments and aspects described above.
[0079] Features of any aspect or embodiment described herein may, wherever appropriate, be applied to any other aspect or embodiment described herein. Where reference is made to different embodiments or sets of embodiments, it should be understood that these are not necessarily distinct but may overlap.BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Certain embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0081] Fig. 1. Schematic diagram of a system embodying the invention.
[0082] Fig. 2. Schematic diagram of software components of the system.
[0083] Fig. 3. Graphical predictions of the optimization algorithm to find the best-fit parameters.
[0084] Fig. 4. Bayesian optimization graph, representing the Gaussian process
[0085] Fig. 5. Bayesian optimization graph, representing the Expected Improvement plot
[0086] Fig. 6. Convergence plot showing surrogate models using Bayesian optimization. The X axis refers to the search number, the Y axis indicates the validation loss
[0087] Fig. 7. Partial dependence plot of a Fine-tuning process using Bayesian optimization in a CNN Fig. 8. Learning curve that shows an optimal fit in the Validation step
[0088] Fig. 9. Graphical predictions of the CALCE dataset. The X-axis indicates the number of discharged cycles, while the Y-axis represents the battery capacity. The horizontal line denotes the EOL criteria Fig. 10. Graphical predictions of the NASA dataset. The X-axis indicates the number of discharged cycles, while the Y-axis represents the battery capacity. The horizontal line denotes the EOL criteria Fig. 11. Graphical predictions of the ECM dataset. The X-axis indicates the amount of time, while the Y-axis represents the evolution of the SOC
[0089] Fig. 12. Graphical predictions of the CHRG dataset. The X-axis represents the SOC percentage, while the Y-axis indicates the OCV.
[0090] Fig. 13. Main contributions of the KANs in the Model evaluation for the different case studies.
[0091] DETAILED DESCRIPTION OF EMBODIMENTS
[0092] Exemplary embodiments of this disclosure address a critical gap in algorithm design for BESS applications by introducing and developing KANs, progressing from basic to advanced network architectures in the energy sector. A key feature of some embodiments lies in moving away from conventional ECM designs and traditional NNs, instead leveraging Transfer Learning on battery operating data. Some validation results are presented below, comparing the performance of the resulting KANs with the most optimal categories of NNs. The study employs a comprehensive computing approach and Deep Learning methodology using PyTorch and Keras as high-level APIs (other APIs such as TensorFlow can also be used). It rigorously tests the BESS under various conditions to assess charging management, Remaining Useful Lifetime (RUL), and state estimation across diverse datasets. By departing from previous research, this disclosure not only fills a vital gap in battery development and Al technology but also offers a novel perspective on energy systems through the innovative application of KANs.Figure 1 shows a schematic diagram of an exemplary apparatus 100 embodying the disclosure. The apparatus 100 comprises a battery energy storage system (BESS) 101 , and a processing system 102. The processing system 102 may, in some embodiments, form part or all of a battery management system (BMS) for the BESS, and may be configured to control a charging or discharging of the BESS. The BESS comprises one or more batteries for the storage of electrochemical energy and the release of electrical energy. The BESS may comprise batteries such as Li-ion batteries, or any other battery technology known to the skilled person. The processing system 102 comprises one or more processors 103 and a memory 104 (ora memory system distinct from the processing system 102, but in communication therewith). As will be appreciated by the skilled person, each processor may comprise a single processing core, or it may comprise a plurality of processing cores.
[0093] The memory 104 stores instructions which, when executed by the one or more processors 102, cause the one or more processors to carry out a method to analyze the health or performance of the battery energy storage system 101.
[0094] The processing system 102 comprises an input 105 for receiving empirical data relating to the BESS 101. As shown in Fig. 1, the BESS 101 is connected to the input 105. The empirical data is determined empirically from the BESS 101. The empirical data may be generated by measuring one or more physical parameters of the BESS 101 , e.g. by one or more sensors arranged to monitor the BESS 101. The empirical data in this embodiment relates to one or more parameters related to the BESS 101. For example, the empirical data may comprise at least one electrical or thermal or temporal parameter. The empirical data can comprise at least one parameter selected from the group consisting of: a resistance, current, voltage, charge capacity, discharge capacity, temperature, or time. The empirical data in this embodiment is collected over a plurality of charge and discharge cycle of the BESS 101 , though it will be appreciated that the performance data could also be collected over a single charge and discharge cycle.
[0095] The processing system 102 also comprises an output 106 for outputting output data from the processing system 102. As depicted in Fig. 1, the output 106 is connected to the BESS 101, The output 106 could be connected, in addition or alternatively, to other processing systems, an external server, etc.
[0096] The processing system 102 stores and implements a trained Kolmogorov-Arnold network (KAN). The trained KAN is configured to model nonlinear dynamics of battery energy storage systems (e.g. the BESS 101). The KAN comprises a set of univariate functions that have been trained to represent nonlinear dynamics of battery energy storage systems. The empirical data received by the processing system 102 is provided to the KAN, and the KAN determines one or more properties of the BESS 101 as output data. The properties are indicative of health or performance of the BESS 101. The properties of the BESS 101 that the KAN determines may include, but is not limited to, a state of charge, a charge capacity, a discharge capacity, an open circuit voltage, a state of health, remaining useful lifetime (RUL).
[0097] Once one or more properties of the BESS 101 is determined by KAN, the processing system 102 uses the output data to control a charging or discharging of the BESS. The control can be communicated to the BESS via the output 106.The empirical data can be generated continually or at intervals over time, and while the empirical data is being generated, the KAN can be operated to determine the one or more properties continually or at intervals over time. Then, the output data can be used to control charging or discharging of the BESS 101 continually or at intervals overtime. The controlling the BESS 101 may involve adjusting a C-rate or cut-off voltage.
[0098] The output data determined by the KAN can be used to help maintain the BESS 101. For example, based on the output data, the processing system 102 can determine when to perform a maintenance action on the BESS 101.
[0099] Figure 2 shows a schematic diagram 200 of the software components of the processing system 102. In particular, a location in memory stores empirical data (shown as the box labelled “Empirical Data”) received from a battery energy storage system. The empirical data is read and provided as an input to the KAN (shown in Fig. 2 as the box labelled “KAN”) which comprises a set of univariate functions. The KAN is implemented by the processing system 102. As described above, the KAN is operated to determine output data that comprises one or more properties of the battery energy storage system indicative of the health or performance of the battery energy storage system. The determined output data is then stored in a location in memory, which may be output for controlling an aspect of the operation of the battery energy storage system (i.e. a charging or discharging of the battery energy storage system).
[0100] Some embodiments additionally provide apparatus fortraining the KAN using empirically determined training data obtained from one or more BESS’s. The training system could be implemented on the same processing system 102 or on a separate processing system.
[0101] The KAN can be trained using a variety of training data sets (e.g. taken from public datasets or from experimental tests on battery cells), but in this embodiment, the KAN has been trained using battery cycling data determined empirically for one or more battery energy storage systems (i.e. not necessarily including the BESS 101). The one or more battery energy storage systems can be of the same type, or of different types. When the one or more battery energy storage systems are of similar or the same type, performance of the trained KAN model is improved, particularly when subsequently applied to a BESS of a similar or the same type.
[0102] The following sections provide further details of some example implementations of the disclosure and also provide a comparison between a KAN-based approach, according to the disclosure, and other machine-learning approaches.
[0103] Experimental data demonstrates the technical benefits of KAN-based systems according to the disclosure.
[0104] Case studies and applications
[0105] This section explains three different case studies of the relevant applications of a BESS, highlighting the level of complexity of algorithm design based on various user needs, battery properties, experimental tests, and operating conditions.
[0106] The case studies correspond to various datasets of a BESS subjected to experimental measurements through specific test criteria. Firstly, two Lithium-ion (Li-ion) cells are tested using a programmable DC electronic load to experimentally simulate the SOC by implementing a second-order ECM. Secondly, two of the most recognized datasets, namely CALCE and NASA, are processed and analyzed to evaluate the RUL of prismatic and Li-ion battery cells. Finally, using the programmable DC electronic load from the initial case study, a series of battery tests are performed with several Li-ion cells of their corresponding modules to evaluate optimal charging management. This section provides a high-level overview of example algorithm designs so that the optimal NN architecture can provide the most accurate performance and interpretability in multitasking areas for a BESS.
[0107] Case study 1: State estimation
[0108] Experimental
[0109] A Mitsubishi i-MiEV battery pack manufactured by GS Yuasa, consisting of 88 LEV50N type Li-ion cells, was collected from ISEAUTO, an innovative Estonian autonomous electric vehicle (AEV) project located on the Tallinn University of Technology (TalTech) campus. Due to topics of strategic importance, energy cooperation, and research management, the battery pack was dismantled into several battery cell modules. Fora comprehensive understanding of the LEV50N cells, Table 1 summarizes the cell parameters.
[0110] Table 1. LEV50N Battery cell parameters
[0111]
[0112] A total of ten battery tests for two Li-ion cells were conducted respectively to evaluate their charge indicators, all using a programmable DC electronic load through the UltraLoad Software for remote operation and monitoring. The programmable DC electronic load has an adjustable current rising speed from 0.001 A / ps to 5 A / ps, a list function that supports editing as many as 512 steps, dynamic mode up to 30 kHz, a readback resolution of 0.1 mV and 0.1 mA. Specifications of the battery tests consist of a resolution of 0.8%, a slew rate of 0.001 A / ps, a frequency of 1 Hz, and a step duration of 1 second.
[0113] An ECM is a useful engineering approach that can simulate the operation of a BESS. The first-order ECM considers a parallel capacitor, voltage source, ohm resistance, and polarization resistance, while the second-order ECM includes not only the same elements but also implements a polarization resistance, which is divided into diffusion resistance and electrochemical resistance. A second-order ECM is coded using Python as the programming language. Furthermore, parameter estimation is achieved using the SciPy optimized library, focusing on root finding by Local (multivariate) optimization.To visualize the results of the root findings in the battery tests, Fig. 3 shows a fitted curve and compares it with the experimental data.
[0114] To validate the results of the optimization and root finding algorithm in the parameter estimation for all battery tests, Table 2 provides the performance metrics consisting of Mean Absolute Error (MAE), Mean Square Error (MSE), and Root Mean Square Error (RMSE).
[0115] Table 2. Validation results of the optimization and root finding algorithm
[0116]
[0117] Finally, the best-fit parameters for each corresponding experimental measurement are collected and defined as input variables to run the second-order ECM, all to simulate the operation of the Li-ion cell and generate a new dataset.
[0118] The datasets corresponding to the state estimation application will be called “ECM dataset”, which contains SOC as a predicted variable, voltage, time, and current as input features, and which will be processed in the algorithm design for the following.
[0119] Case study 2: Remaining Useful Lifetime (RUL)
[0120] Two of the most recognized datasets focused on RUL and End-of-Life (EOL) criteria are collected, which are NASA and CALCE. The EOL criteria is defined as a concept in a BESS that accounts for failure in performance or functionality, which is typically associated with 70-80% of the full rated capacity.
[0121] The CALCE dataset, available from the Center for Advanced Life Cycle Engineering (CALCE), a research center at the University of Maryland, contains four prismatic cells cycled at a constant current of 1°C. The cells were tested on a charge profile that was a standard constant current / constant voltage protocol with a constant current rate of 0.5 C until the voltage reached 4.2 V and then held at 4.2 V until the charge current dropped below 0.05 A. The discharge cut-off voltage for these batteries was 2.7 V, and the total cell capacity was 1.1 Ah. In this dataset, the State of Health (SOH) represents the dependent variable, and the features used for the predictions are time, resistance, current, voltage, charge capacity, and discharge capacity.
[0122] The NASA dataset, available from the NASA Ames Research Center, contains the record of four commercial 18650 Li-ion batteries, with each battery repeating three operations: charge, discharge, and impedance measurements. In this specific case study, only the discharging process is considered to evaluate the RUL. Initially, charging was performed in constant current (CC) mode at 1.5 A until the battery voltage reached 4.2 V and then continued in constant voltage (CV) mode until the charging current dropped to 20 mA. Subsequently, discharge was performed at a constant current level of 2 A until the battery voltage dropped to 2.7 V and experiments were stopped if the cells reached 70% of the end-of-life criteria. The predicted or dependent variable is the SOH,however, compared to the CALCE dataset, temperature indicates an additional independent variable.
[0123] Case study 3: Charging management
[0124] This dataset corresponds to LEV50N cells, however, in this case study, sixteen cells from their respective modules are considered, conducting a total of 89 battery tests.
[0125] In this specific application, the charging management of Li-ion cells is evaluated using a programmable DC electronic load under several operating conditions, which are: 1) C rate, 2) operating voltage range, and 3) cut-off voltage range. The operating conditions of the LEV50N cells were exhaustively measured in various scenarios through the UltraLoad Software for remote operation and monitoring.
[0126] Regarding data acquisition, the independent variables are current, time, voltage, energy, resistance, and discharge capacity. Additional features based on the input variables such as SOC and power are obtained using the Coulomb counting method; similarly, open circuit voltage (OCV) is calculated using the ECM approach and is defined as the predicted or target variable in the network architecture.
[0127] For the purposes of this disclosure, the dataset referring to the charging management application will be called “CHRG dataset”, which, like the previous ones, will be processed to achieve the optimal algorithm design in the next sections.
[0128] Artificial Intelligence methods and Neural Network
[0129] This section presents a brief review of different Al methods, as used in the validation tests below, as well as in exemplary KAN-based embodiments, focused on NNs, their corresponding mechanism, architecture, and functionalities, highlighting the implemented algorithms: 1) Kolmogorov-Arnold Network (KAN), as implemented in embodiments of the disclosure,, 2) Multilayer Perceptron (MLP), and 3) Recurrent Neural Network (RNN) and Convolutional Neural Network (CNN), as used for comparison purposes.
[0130] The workflow of each NN has a sequential structure: Definition --> Network architecture --> Mechanism --> Functionalities --> Implementation.
[0131] This sequential structure will be used for discussing the proposed NNs, emphasizing a systematic review of their definition, network architecture, mechanism, functionalities, and implementation. Each stage is shortly explained as follows: The definition explains the fundamental idea of each NN and its role in solving computational problems. The structural aspects of the corresponding NNs, such as layers, nodes, hyperparameters, and connections, are highlighted by different types of architectures. In the mechanism, the key content is based on how each NN works, breaking down the processes that allow them to learn and make predictions. Functionality is described by the strengths and capabilities of NNs through their network mechanisms and architectures in Deep Learning. The selected NNs for implementation are discussed through the practical aspects of the algorithm design and coding.The following subsections describe some implemented NNs before introducing the algorithm design, used in certain embodiments, to provide understanding of the novel and innovative KAN employed in the embodiments of the disclosure.
[0132] Kolmogorov-Arnold Network (KAN)
[0133] KANs are computational models inspired by the Kolmogorov-Arnold representation theorem, whose implications in the field of Deep Learning are offering new alternatives for building NNs. Initially, KANs are considered a promising alternative to MLP. However, due to their high level of accuracy and interpretability, the opportunity to compete with other categories of NNs, such as CNNs, RNNs, and ANNs in several fields of Al is encouraged.
[0134] Compared to MLPs, which have fixed activation functions on nodes or neural units, a KAN has learnable activation functions on numerical values, also known as weights, which are associated with the connections between neurons or nodes across different layers of the network. Furthermore, in the KAN mechanism a univariate function parameterized as a spline has the task of replacing each weight parameter, consequently allowing the switch between coarse-grained and a finegrained grid. From a mathematical perspective, function composition also provides a powerful tool to explain the mechanism of a KAN by decomposing the multivariate function into univariate functions, this decomposition being a different approach than a traditional NN.
[0135] Basis functions are defined as coefficients of the building blocks to create more complex functions and play an important role in the network architecture, so as an added value, a KAN can learn these basis functions at the edges, resulting in highly flexible and interpretive activation functions at each connection. The remarkable functionalities of KAN are also manifested by replacing all weight parameters by the coefficients within the edge activation function, thereby eliminating traditional linear weights from the network.
[0136] Regarding the main hyperparameters of a KAN, the width, grid, and spline order “k” integrates the network architecture. The width refers to the number of basis functions used to build the activation functions within each layer, the grid defines the level of detail at which the interval over the activation function operates and captures in the network, finally, the "k" determines the degree of smoothness to parameterize the activation functions.
[0137] A summary of the KAN working framework, as implemented by embodiments of the disclosure, is explained in the following points:
[0138] • Input layer and processing. At initialization, the input layer of a KAN extracts the multivariate input to prepare it for processing through the corresponding hidden layers. After that, each input variable is individually transformed by a set of univariate functions.
[0139] • Hidden layers and univariate functions: The hidden layers of the KAN architecture extract the univariate functions from the initial step. To complete this step, a sum of univariate functions is obtained, which represent a combination of transformed input variables.
[0140] • Output layer: The output layer extracts the sum of univariate functions from the hidden layer. Function composition is applied to the combination of transformed input variables; the results are calculated and then summed up to obtain the KAN predictions.The KAN architecture relies heavily on applying univariate functions to each input variable individually and then to the summed outputs, thereby reducing the complexity of the multivariate function by decomposing it into sums of simpler univariate functions. The parallel computation of multiple univariate transformations for each input variable is a major advantage compared to other categories of NNs, making it very efficient for certain types of problems in the Deep Learning field.
[0141] Multilayer Perceptron (MLP)
[0142] A MLP is a type of artificial neural network (ANN) that consists of multiple layers of neurons. MLP architecture is characterized by the implementation of non-linear activation functions, which makes the network learn complex patterns in the data. MLPs are applied in machine learning, as by learning non-linear relationships in the data, MLPs could be used to model tasks such as regression, classification, and pattern recognition.
[0143] The MLP architecture comprises three layers: the input layer, the hidden layer, and the output layer. First, the input layer is the initial layer of the network, which processes the independent variables in the form of numbers. Second, is the hidden layer, which processes the information received from the input layer. Third, the output layer produces the results of the calculations applied to the network data.
[0144] The mechanism of MLP is summarized in the next steps.
[0145] • The activation rate of the hidden nodes is found using the inputs and the links from the input to the hidden layer. Each neuron in the hidden layer is connected to the neurons in the next layer.
[0146] • The corresponding weights of the neurons are updated with the help of the learning phase. The learning phase is repeated continuously until the error value exceeds the threshold level.
[0147] • Finally, the data is passed in a forward path from the input layer to the output layer, being the equivalent of a feed-forward that uses backpropagation to train all the nodes.
[0148] Recurrent Neural Network (RNN) and Convolutional Neural Network (CNN)
[0149] RNNs have an architecture based on recurrent connections and can model sequential data for sequence recognition and prediction. The mechanism of an RNN is given by high-dimensional hidden states with nonlinear dynamics, emphasizing that the state of the hidden layer at a given time is conditioned to its previous state.
[0150] Gated Recurrent Unit (GRU) and Long Short-Term Memory (LSTM) are types of RNNs designed to handle sequential data.
[0151] The LSTM is composed of three gates (forget, input, and output), whose beneficial mechanism allows RNNs to learn over many more time steps and control the flow of information to hidden neurons over long periods of time, thereby using more consistent errors. The corresponding gates of the LSTM retain the features extracted from previous time steps, regulating not only the amount of data entering at each corresponding time step, but also the number of weights being optimized. GRU has control units that modulate the flow of information within the unit, but without having separate memory cells. Unlike LSTM, GRU exposes the entire state at each step and computes a linear sum between the existing state and the newly computed state. Hidden layers containingmemory cells cover the main functions of GRU networks. Cell state changes and maintenance depend on two gates in the cell: a reset gate and an update gate.
[0152] CNN is a feed forward NN that applies convolutional operations to the input instead of general matrix multiplication, being able to extract features from the data with convolutional structures in at least one of its layers. Compared to other types of NNs, a CNN has different layers that make up its intrinsic architecture, whose main functionalities are to control overfitting in the model, introduce nonlinearity, perform normalization, and combine features to make accurate final predictions.
[0153] The RNNs implemented in the algorithm design are GRU and LSTM, while a CNN-1 D for CNN, whose performance will be compared and analyzed in the following sections.
[0154] Algorithm design
[0155] In this section, the algorithm design used in some embodiments will be described, illustrated, and executed to evaluate the Model Performance Analysis subsequently. In addition to KAN-based models, embodying the disclosure, various other NN-based approaches are described for comparison purposes.
[0156] In particular, a Deep Learning methodology proposed in R.A. Gilbert Zequera et al., “Deep Learning methodology for charging management applications in battery cells based on Neural Networks”, IEEE Transactions on Intelligent Vehicles, doi: 10.1109 / TIV.2024.3417216, is implemented as an added value to strengthen robustness. However, in the description below, not only charging management applications but also RUL and state estimation are considered. In addition, both Keras and PyTorch are defined as APIs to demonstrate a high level of adaptability and effectiveness. T ensorFlow is another API that may be used in some embodiments. T ensorFlow supports static and dynamic computational graphs. In terms of model building, TensorFlow provides many toolkits and libraries for the construction and deployment of neural networks. In addition, its computational usage offers efficient computation to different processors, such as CPU, GPU, and TPU.
[0157] Regarding the energy context, the use of TensorFlow promotes the adaptability and scalability of the implemented algorithms considering industrial and research needs, all due to the extensive ecosystem for model training, serving, deployment, and integration with other platforms. Examples of the ecosystems and rich set of tools include TensorFlow Hub, TensorFlow Extended (TFX), and TensorFlow Lite.
[0158] As for specific and advanced computing techniques, in the training process, an initial network architecture is designed to consider input features of the corresponding datasets and NN hyperparameters. Fine-tuning is then performed using Bayesian optimization to achieve minimal validation loss, and then the optimal hyperparameters of the networks are collected to execute Cross-validation and calculate performance metrics, avoiding later problems such as overfitting and underfitting. A summary of the algorithm design is explained in the following points:
[0159] Data acquisition is completed by considering the corresponding datasets for each BESS application and separating the training, validation, and testing sets.
[0160] • Exploratory data analysis (EDA), Feature Engineering, and Feature Selection are executed to complete the Data processing. In this step, data distribution, correlation matrix, and VarianceInflation Factor (VIF) are useful tools to achieve a solid understanding of the problem depending on the type of application and identifying KPIs.
[0161] • The initial network architecture is designed using the predictors and target variables from data processing, and the hyperparameters of the NN are selected. The training and validation sets are processed.
[0162] • Training is started and the loss is calculated on the validation set to obtain the initial performance of the NN. Fine-tuning is performed to achieve the minimum validation loss.
[0163] • Once the minimal validation loss is achieved, Fine-tuning is completed and optimal hyperparameters are collected. Cross-validation is performed to evaluate performance metrics in the validation step.
[0164] • PyTorch state dictionary (state_dict) and Hierarchical Data Format (HDF5) files are generated to store the resulting NNs. The testing set is processed, and final performance metrics are calculated in the Model evaluation step.
[0165] The interpretability of the algorithm design is given by both entities: the physical one, which refers to the operation of the BESS, and the virtual one, which is associated with the functionalities, architecture, and hyperparameters of the different NNs. This framework reflects and allows the user to understand the underlying factors that drive the predictions of the target variable in each application of the BESS.
[0166] Neural Network architecture and Training
[0167] After the successful Data processing completion, the training and validation sets were stored and consistently transformed into tensors to ensure alignment with time steps, features, and samples for algorithm design using Keras and PyTorch APIs. As noted previously, the TensorFlow API can also be used at this stage.
[0168] Activation functions, metrics, and kernel regularizers are the main arguments that integrate the programming interface in the network architecture, so different matrices are created with a set of hyperparameters to monitor the Training and validation steps. For RNNs, CNNs, and MLPs (described here for comparison purposes), batch size, learning rate, epochs, weight decay, and gamma regularizer are the selected hyperparameters, while for KANs (embodying the disclosure), width, k-spline order, and grid are additional hyperparameters to optimize.
[0169] The selection of the above hyperparameters is related to the impact on the performance and computational efficiency of the Training process, which affects the quantity of allocated error with which the NN weights are updated, and the amount of information the resulting architecture can capture and its suitability. As for the activation function, ReLu is selected for its sparsity and for being beneficial in reducing the probability of gradient vanishing.
[0170] Regularization techniques are used to implement and avoid overfitting or underfitting during the validation process, all to achieve convergence in the learning curves, so dropout, weight decay, and gamma regularizer are included in the hyperparameter set of the network architecture. Weight decay, also known as L2 regularization, is a technique applied to the weights of a neural network to minimize a loss function by implementing a penalty on the norm of the respective weights. Similarly,the gamma regularizer is defined as a multiplicative factor by which the learning rate decays at each epoch.
[0171] Additionally, the Adam optimizer is implemented to minimize losses and weights, stabilize training, and help NNs converge to optimal solutions, which, in parallel with an early stopping criterion and a learning rate schedule, integrates the final network architecture. Early stopping reduces the risk of overfitting and saves time and computational resources by simplifying the model and preserving the best weights, while a learning schedule is tasked with avoiding exceeding the minimum learning rate and fine-tuning the model parameters.
[0172] The notable differences in the Training step are based on the network architecture of KANs, which differ substantially from RNNs, CNNs, and MLPs due to the k-spline order, grid, and width, so these hyperparameters will be optimized first before obtaining the resulting NNs and completing the Fine-tuning process for mutual hyperparameters.
[0173] Fine-tuning and Bayesian optimization
[0174] In the field of Al and model interpretability, the goal of achieving Fine-tuning plays an important role in both Validation and Model Performance Analysis. One of the techniques to automate algorithm design is Neural Architecture Search (NAS), which helps to find the optimal architecture by searching over a large hyperparameter space.
[0175] In terms of programming and computer science, understanding the key differences between trainable parameters and hyperparameters is a fundamental task. Trainable parameters refer to the parameters learned by the algorithm during Training, such as the weights of the neural network, while on the other hand, hyperparameters are set before starting the learning process and are not updated in the learning step, thus dictating the overall structure and behavior of the model.
[0176] Fine-tuning using various optimization methods is a monumental challenge, but essential to ensure that models are accurate, efficient, and adaptable. Different types of hyperparameter optimization techniques include Evolutionary Algorithms, Reinforcement Learning, Grid Search, Bayesian optimization, and Random Search. In this research, due to its high level of convergence, computational efficiency, and fast convergence, Bayesian optimization was selected in the algorithm design.
[0177] Bayesian optimization assumes that a specific probability distribution underlies the performance and determines the next set of hyperparameters to be evaluated considering previously observed combinations. Two important concepts are considered in the implementation of Bayesian optimization: exploration and exploitation. Exploration refers to selecting a point in a region that currently has the best results, while exploitation means choosing the point with the highest uncertainty at each iteration.
[0178] In the context of NNs and BESSs, the main goal of Bayesian optimization is to minimize the loss during the Training and Validation steps, achieving convergence on the Learning curves and modeling the objective function. To successfully achieve this complex goal, a model called Gaussian Process is built, which refers to the infinite-dimensional natural analogue of the multidimensional Gaussian, and this process is used to model the unknown objective function and provides a posterior distribution over the function values given the observed data.In the algorithm design, an acquisition function called Expected Improvement is implemented to provide a balance between exploitation and exploration, which determines the next set of points to be evaluated in the search space and quantifies the potential convenience of sampling a particular point, all while considering both the predicted values of the surrogate model and the uncertainty associated with those predictions.
[0179] As mentioned earlier, the Fine-tuning process via Bayesian optimization is implemented using Keras and PyTorch APIs. However, to offer an innovative solution, several programming functions were developed from scratch using Python libraries designed for sequential model-based optimization. First, a fitness function is created to evaluate the performance of the NNs based on the selected hyperparameters, to minimize the loss. Next, a Gaussian process is constructed during the training phase, starting with an initial set of hyperparameters to learn the performance metrics of each NN. The Bayesian optimization search begins with these initial values and subsequently focuses on promising regions identified in earlier steps. In the third step, the model is sampled to maximize Expected Improvement, and the validation loss is calculated for each search iteration. Finally, the output function stores the set of hyperparameters that correspond to the minimum validation loss, and performance metrics are computed.
[0180] As for the KAN hyperparameters, the width, the grid and the k-spline order are the main elements to be optimized. In the case of the width, it is composed of the input size, the hidden size, and the output size, being the number of features and a single target variable the input and output size respectively, so only hidden size, grid, and k-spline order will be included in the Bayesian optimization for this type of NN.
[0181] Before obtaining the optimal and mutual hyperparameters of different NNs, Bayesian optimization is implemented to select the core hyperparameters of the KAN. In this specific procedure, a loop iterates through the diverse types of datasets based on case studies until convergence is achieved based on the minimum validation loss. The results show strong performance with values of 16, 5, and 3 for the hidden size, grid, and k-spline order respectively.
[0182] The Bayesian optimization graph is a powerful tool that helps visualize where the optimization algorithm is likely to sample next, providing a solid understanding of the trade-off between exploration and exploitation. The graph illustrates exploration by sampling in areas with high uncertainty, while exploitation is shown by sampling in areas expected to yield the best results. The above explanations are visualized in Fig. 4 and Fig. 5, showing a Bayesian optimization graph that is composed of the Gaussian process plot, and the Expected Improvement plot, considering the batch size of a KAN as an example hyperparameter.
[0183] In the Gaussian process plot, the Python fitness function provides observations along the data distribution of the selected hyperparameter with its associated validation loss, which is given by the Gaussian model. The X-axis plots the selected hyperparameter, while the Y-axis indicates the validation loss; the red points illustrate the observations provided by the Gaussian process.
[0184] Considering the Expected Improvement plot, the main objective is to analyze the different high-uncertainty sampling regions and the best results that exploration and exploitation yield. The X-axis indicates the selected hyperparameter, and the Y-axis represents the Expected Improvement,quantifying the best current known value that can be expected if the function is evaluated at the corresponding hyperparameter.
[0185] Analyzing the Gaussian process plot, exploitation is represented by low areas suggesting promising regions based on existing data and giving minimal validation loss, on the contrary, areas with high uncertainty are indicated by exploration indicating maximum validation loss values that could be targeted by the acquisition function for further sampling. As can be seen in the Expected Improvement plot, areas with great values indicate a high probability of improving the current hyperparameter, conversely, lower areas represent regions where the model is confident that little or no improvement will be achieved because these areas have already been well explored.
[0186] As for the optimal hyperparameters of each NN in Bayesian optimization, a search is performed until convergence on the minimum validation loss is achieved. Fig. 6 shows a plot illustrating the surrogate models in algorithm design.
[0187] At first, it can be observed that the NNs show a different level of validation loss that decreases with the amount of search, however, there is a similarity in the final that illustrates the effectiveness of Bayesian optimization. Due to their architecture and functionalities, GRU and LSTM show a similar trend, however, KAN also provides even lower validation loss than MLP and CNN, giving an idea about the promising results that could be delivered in the next steps.
[0188] A graph of all combinations of hyperparameter values is shown in Fig. 7, exemplifying the Bayesian optimization of the CNN. The vertical axis illustrates the influence of a single dimension on a fitness function, called a “Partial Dependence plot" and the horizontal axis the hyperparameters. It visualizes the effects of changing one or more variables in the algorithm design and shows how the approximate fitness value changes with different values in that dimension. The yellow regions show areas where the loss on the validation set is lower, as opposed to the darker regions. The star in the graph represents the location where the optimal value of the hyperparameter is found.
[0189] Fine-tuning through Bayesian optimization was performed on all types of NNs explained in Section 3, so CNNs, MLPs, RNNs, and KANs highlight different sets of optimized hyperparameters and performance metrics on the validation sets, which will be explored in the next subsection before providing the resulting NNs for each dataset in the Model Performance Analysis.
[0190] Validation
[0191] After Training and Fine-tuning are complete, the next step is to run cross-validation to measure the performance of each NN on the validation set. The goal of cross-validation is to provide an approximate performance of the model for data that will appear in the future. In addition, it is also important to consider balancing underfitting and overfitting.
[0192] Underfitting refers to poor model performance on both the training and test sets; on the other hand, overfitting indicates that the model was over-tuned during training, so it performs well on the training set but poorly on unseen being evaluated.
[0193] Due to the behavior of the training and validation losses at each epoch of the network architecture, overfitting, and underfitting can be identified from the learning curve. An underfitting plot shows high losses for both training and validation data at all epochs without significant improvement, hence it will be indicative of the lack of ability to learn the training set. On the other hand, in the case ofoverfitting, the training loss continues to decrease, while the validation loss starts to increase after reaching a minimum, thus the model fits the training data too closely, capturing noise along with the actual patterns.
[0194] In the programming framework, cross-validation is built in the objective function, in this case, the fitness function that was programmed in Python to run the Gaussian process. The objective function trains and evaluates the NN for each set of hyperparameters using cross-validation and returns the corresponding performance metrics. By using cross-validation in the objective function of Bayesian optimization, the selected hyperparameters are more likely to generalize well to unseen data, providing a more robust estimate of model performance and reducing the risk of overfitting across the search space.
[0195] A learning curve is a valuable tool for diagnosing the behavior of every NN in the Training and Validation steps. Key elements of an optimal fit consist of training losses that decrease to a plateau over epochs, while validation losses decrease, eventually plateau, and ideally remain close to the training loss. The numerical difference in the gap of a learning curve is the point at which the model's validation loss is slightly larger than the training loss when both curves plateau. The specific size of the gap depends on the problem, but in general, a gap of 1% is considered small and stable, indicating good generalization. Monitoring the gap of a learning curve during training helps make decisions in the algorithm design about model complexity, regularization, and when to stop training to avoid overfitting.
[0196] In Fig. 8, a good fit of the learning curve can be observed, showing a balance between bias and variance, indicating that the model generalizes well to new and unseen data. The curve illustrates the corresponding losses of a KAN, representing not only optimal tuning in Training and Validation but also convergence in a minimum number of epochs, thus reducing training time and computational complexity compared to the other NN categories.
[0197] Finally, the target variable is predicted for each dataset and performance metrics are obtained by calculating the MSE, MAE, and RMSE for the corresponding NNs. In addition, the Residual Sum of Squares (RSS) and Symmetric Mean Absolute Percentage Error (SMAPE) are the supported benchmarks employed to reflect the prediction's accuracy and robustness to different scales of values. As mentioned in the section “Artificial intelligence methods and neural networks”, GRU, LSTM, MLP, KAN, CNN-1 D are designed, implemented, and validated. Table 3, Table 4, Table 5, and Table 6 show the performance metrics in the Validation step.
[0198] Table 3. Validation results of the CALCE dataset
[0199]
[0200] Table 4. Validation results of the NASA dataset
[0201]
[0202] Table 5. Validation results of the ECM dataset
[0203]
[0204] Table 6. Validation results of the CHRG dataset
[0205]
[0206] From the performance metrics, the nature of the datasets and their main applications are evaluated in the Model evaluation step. If the dataset contains outliers or noise that influence the model performance metric, it is advisable to select MAE due to the robust prediction tasks. On the contrary, if the dataset includes scenarios where large errors are particularly critical, the effect of MSE will cause these large errors to have a larger impact on the metric, which is useful when the goal of the model is to minimize such significant deviations. Regarding the RMSE, it is convenient to use it to balance the penalty for larger errors and whose interpretation of the error metric is in the same units as the target variable, so that the model predictions are accurate and interpretable. As for the numerical values obtained in the validation results, the MAE is usually higher than the MSE, suggesting that the datasets have diminutive and uniformly distributed errors. This implies that the implemented NNs are generally accurate, without large deviations from the actual values,with operational variability being the core cause of error capture during battery testing. The implications of a larger MSE and RMSE than an MAE depend on the sensitivity of the model to outliers and greater errors, suggesting that predictions are heavily penalized by the MSE due to the quadrature effect, which could be beneficial information if large errors are particularly problematic in the BESS application.
[0207] The importance of the RSS relies on measuring the overall squared difference between the predicted and actual values, evaluating the error rate that the model accumulates over all data points, providing a holistic view of model accuracy, and handling continuous variables in the BESS operation. In the case of SMAPE, this benchmark normalizes errors relative to the actual and predicted values, making it robust to varying scales, therefore ensuring that the algorithm ensures that the algorithm does not disproportionately penalize small deviations for large values or overlook significant deviations for small values in the BESS performance predictions.
[0208] Considering the implementation of Deep Learning methodology, Regularization, Cross-validation, and Fine-tuning through Bayesian optimization are the testimony of the improvement and effectiveness of the different categories of NN. Among the most relevant benefits are the reduction of computational resources in terms of training time and knowledge transfer, global optimization due to the balance between exploration and exploitation, better generalization to ensure that hyperparameters are configured to maximize NN performance, and adaptability to a wide range of network architectures.
[0209] The validation results show that KANs provide a higher level of performance compared to CNNs and RNNs. A more detailed analysis and explanation will follow.
[0210] Model Performance Analysis
[0211] Optimal hyperparameters after Fine-tuning are provided to compare and analyze each NN approach on different datasets. In the Model evaluation stage, the final performance metric is calculated and the NNs are evaluated on different test sets. At the end of this section, a discussion and analysis of the results are carried out, mainly focusing on KANs in the framework of Al methods for relevant applications of BESS.
[0212] Optimal Network Hyperparameters
[0213] The Fine-tuning process was executed through Bayesian optimization, using the Expected Improvement as the acquisition function and finding a mutual balance between exploration and exploitation for each hyperparameter, after that, cross-validation is performed to obtain the optimal hyperparameters of the network.
[0214] Table 7, Table 8, Table 9, and Table 10 show the optimal set of hyperparameters of each NN for the corresponding datasets, after the successful completion of the Validation step.
[0215] Table 7. Optimal hyperparameters of the CALCE dataset
[0216]
[0217]
[0218] Table 8. Optimal hyperparameters of the NASA dataset
[0219]
[0220] Table 9. Optimal hyperparameters of the ECM dataset
[0221]
[0222] TABLE 10. Optimal hyperparameters of the CHRG dataset
[0223]
[0224] In the context of Bayesian optimization convergence, the optimal hyperparameters are similar for the NN categories, however, the main differences depend on the applications of a BESS.Regarding the RUL case study on the CALCE and NASA datasets, the learning rate and gamma of a KAN reach optimal values in a similar range as other types of NN, leading to a smooth and constant decrease in loss, however, significant differences are noted in the case of batch size and the most notable is found for the number of epochs and units, the latter being twice as small as the other NNs. The above result is explained based on the training time and knowledge transfer that the KAN manifests in the algorithm design due to its mathematical properties and functionalities, consequently impacting the dynamics of the network architecture, including training speed, memory usage, and faster convergence behavior.
[0225] In the case of state estimation and the ECM dataset, the complexity of the target variable reflects the convergence of the hyperparameters in a longer training time, providing high values of the epochs and units for the RNNs, however, in the case of MLP and CNN, there is mutual similarity which is given by the weight decay and gamma in the Regularization method. Analyzing the optimal hyperparameters of KAN, the epochs, learning rate, and number of units differ significantly from the other NN categories, so the algorithm design converges in a reasonable number of iterations without overshooting or oscillating and balancing training speed and stability.
[0226] Charging management for the CHRG dataset shows the most notable results in the case studies, mainly due to the large number of input features and correlation in the Data processing step. While the learning rate and weight decay are in a similar range, the rest of the hyperparameters differ considerably regarding convergence and parallel computation. The nature of the KAN architecture allows the network to approximate any multivariate continuous function, making it a powerful tool in function approximation; this behavior is explained by the transformation of the input variables through univariate functions, and the combination of these transformations into the final output using further univariate functions and summation.
[0227] From the BESS perspective, selecting a C-rate and a cut-off voltage range are important tasks that will determine how long it takes for the BESS to charge or discharge in the experimental test, consequently they have a direct impact on the Training and Fine-Tuning in the case studies, so the batch size number and the learning rate are the two core hyperparameters that will intrinsically impact the different applications, all due to the number of samples used in a forward and backward propagation through the network architecture. Similarly, the nature of the BESS dataset depends on the input features and their relationship with the target variable, so hyperparameters such as weight decomposition, gamma regularizer, epochs, and number of units play a critical role in the algorithm design to determine the complexity of predictions on each application, achieving balanced learning, effective convergence, and generalization in the Training and Validation steps.
[0228] After completing the Fine-tuning and cross-validation steps, the optimal hyperparameter networks with their respective architectures were processed and saved using PyTorch and Keras APIs to create different hierarchical data formats (HDF5) and PyTorch state dictionary (state_dict) files. In the following subsection, the final performance metrics on the testing sets for each case study will be analyzed, compared, and discussed to highlight the outstanding performance of KANs when applied to BESS’s.Model evaluation
[0229] The corresponding testing sets are processed for each case study. In the RUL application, three different CS2 prismatic cells composed of LiCo02 cathode integrate the CALCE dataset, while three commercial 18650 Li-ion batteries for the NASA dataset. Considering the state estimation, ten different ECM datasets are collected from an LEV50N cell integrating a Mitsubishi i-MiEV battery pack. In the case of the charging management application, a total of 89 battery tests were performed on sixteen LEV50N cells integrating four different battery modules.
[0230] To complete the Model evaluation step, all the different battery tests are stored in a Python list, processed, and transformed into a final set of matrices using Keras and PyTorch APIs to demonstrate the enormous level of adaptability and effectiveness in a programming environment. After that, the resulting network architectures were evaluated on the testing sets with the corresponding MSE, MAE, RMSE, SMAPE, and RSS to obtain the final performance metrics. Table 11 , Table 12, Table 13, and Table 14 show a summary of the results in the Model evaluation step for each NN type and case study.
[0231] T able 11. Model evaluation results of the CALCE dataset
[0232]
[0233] Table 12. Model evaluation results of the NASA dataset
[0234]
[0235] Table 13. Model evaluation results of the ECM dataset
[0236]
[0237]
[0238] Table 14. Model evaluation results of the CHRG dataset
[0239]
[0240] When analyzing the final performance metrics for each case study, NASA and CALCE datasets provide the highest level of accuracy in the Model evaluation, which is mainly due to the NNs’ property of identifying sequential and non-sequential behavior leading to a decreasing trend as a function of SOH in the RUL. For charging management, the large number of features and different operating conditions increase the level of complexity in the prediction results for modeling the temporal dynamics of a BESS, thus the NNs’ learning process takes considerable time to converge, and a higher error rate is obtained. The lowest rate of accurate predictions is calculated on the ECM datasets, all due to the high level of nonlinear behavior and complex objective functions, so that not only the convergence and training time increases but also the computation on charging and discharging, thus generating both forward and backward sequences in algorithm design.
[0241] Even under different operating conditions, scenarios, and BESS applications where battery testing was conducted, there is a tremendous accuracy of KAN outperforming the other NNs, with MSE less than 0.3% and MAE and RMSE less than 4.5% for all datasets. Considering the performance metrics at the validation step, KANs still provide the minimum error rate in the same numerical range beyond the initial expectations in the algorithm design process. In contrast, MLP and CNN-1 D have slightly higher accuracy in the model evaluation but are still lower than KAN. At the same time, LSTM and GRU maintain remarkable performance without showing significant final improvement. Regarding additional metrics, the different network architectures perform accurately across all the case studies with less than 6.03% and 9.10% for SMAPE and RSS beneath high levels of nonlinear behavior and health monitoring in a BESS, quantifying the overall error, and penalizing large deviations that ensure the algorithm learns effectively from the data.
[0242] The interpretation of KAN results is facilitated by its architecture, which simplifies high-dimensional problems by focusing on univariate functions. This approach is beneficial for modeling points where variable interactions are limited or can be modeled independently due to high correlations, such as the independent variables in a BESS. This reduces complexity and enhances performance in charging management applications.For RUL prediction, the objective function is represented by a curve that gradually decreases as battery usage time increases. In this context, KANs are well-suited for additive models, where the output is the sum of individual contributions from each input, such as battery capacity across cycles. In state estimation, prediction complexity increases significantly due to noisy inputs and non-linear relationships. However, KANs offer robustness by decomposing complex functions, isolating the influence of each input variable, and adapting to the various non-linearities in the data through univariate functions.
[0243] The OCV predictions with their corresponding SOC are visualized in Fig 9. Fig. 10 and Fig. 11 graphically represent the RUL for the CALCE and NASA datasets respectively, indicating the number of cycles and the battery capacity. In the state estimation and ECM datasets, Fig. 12 illustrates the evolution of the SOC over a certain time. For visualization purposes, only the most accurate NNs are shown in the corresponding graphs.
[0244] Considering the RUL in Fig. 10 and Fig. 11, KAN’s alignment with actual capacity curves shows its effectiveness in RUL prediction, leveraging its mathematical advantages across different datasets, and highlighting its capability of predicting the highly nonlinear and uncertain nature of battery degradation, thus offering a competitive option compared to other NNs. In Fig. 12, KAN predictions show a strong alignment with the actual SOC values across the entire period, especially during the intervals of both rapid transitions and stable phases of SOC, which is crucial for applications requiring real-time SOC estimation. Regarding Fig. 9, the KAN design allows it to adapt to static relationships like the OCV-SOC curve without relying on temporal or sequential dependencies, it maintains its accuracy across the entire SOC range, and superior performance ensures better reliability in SOC-OCV mapping, leading to more accurate and robust predictions in real-world applications.
[0245] To summarize this section, Fig. 13 provides a summary comparison of the different case studies, addressing the contributions of KANs in improving predictive accuracy and efficiency in BESS applications, and highlighting their expected impact on the final model performance.
[0246] Finally, the following subsection will present a brief discussion of some challenges addressed by certain embodiments of the disclosure.
[0247] Discussion
[0248] In the initial steps, managing the entire data lifecycle is an important task before initializing the algorithm design, so that understanding users' needs in the BESS application is achieved through EDA, Feature Engineering, Feature Selection, and Transformation. It was demonstrated that the predictors provide a level of interpretability depending on the case study, which helps to maximize and measure the predictive signals in further steps.
[0249] From the BESS domain, understanding the nature of the dataset allows the user to monitor the operational profile and physical interpretation of the BESS application, ensuring fairness and consistency in the selection of network hyperparameters. Regarding the architecture of the different KANs, a reciprocal relationship between hyperparameters was provided in a Fine-tuning process through Bayesian optimization for each case study, in which the learning rate and batch size monitor the behavior of predictions due to different step sizes and samples that update the weightsof the network, and the number of epochs and units dictates the network dynamics based on training speed and convergence.
[0250] Gaussian process and Expected Improvement are added values of this research through Bayesian optimization, not only explaining the balance between exploration and exploitation for each hyperparameter in the network but also stochastically determining the optimal architectures in the Fine-tuning process, whose performance and combination values are visualized in Convergence Partial dependence plots.
[0251] Regarding the performance metrics in Model evaluation, all NNs provided an accuracy greater than 94%, but the KANs showed the best results in all case studies with more than 96%. The outstanding performance of KANs is attributed to their intrinsic property of accurate approximation of complex nonlinear functions and faster training convergence. Although the nature of the datasets differs due to the type of application, KANs demonstrated their beneficial mechanism in BESS tasks where: 1 ) Based on Health and Charge indicators, a high level of correlation between the independent variables is found and their relationship with the target variable is highly nonlinear, 2) The entire sequence contains both past and future time steps, which requires context from both directions to understand the operation of a BESS under different profiles, 3) An amount of high-dimensional data is collected containing noisy inputs, or limited data is processed during specific charging scenarios. It will be appreciated that the disclosure also provides benefits in other BESS tasks, and that points 1-3 listed above are merely examples.
[0252] Compared to existing Deep Learning models, Machine Learning methods, and physics-based approaches, this research introduces the novel KANs in the energy framework, starting with algorithm design from a beginner level that initially familiarizes the reader with different case studies, until reaching an advanced network architecture that can make accurate predictions in the Model evaluation step. Furthermore, considering the BESSs applications, due to the high level of Al methods, NN categories, APIs, and programming tools, this research complements some proposed methodologies that not only focus on the aging state of a BESS through NNs, but also in the diagnostics of RUL, and state estimation using Kalman filtering, NNs, and Transformer models. In terms of accuracy, the proposed KANs improve existing methods for RUL, state estimation, and charging management, not only by comparing their resulting architectures with the most accurate NN categories such as CNN, RNN, and MLP, but also by explaining the design of these KAN architectures based on mathematical foundations, intrinsic properties, and computational functionalities. Regarding adaptability, the KANs provide the BESS framework to provide engineering behavior through Bayesian optimization in PyTorch and Keras APIs, whose future expectations can deploy the current methodology in a software environment through MLOps.
[0253] KANs have demonstrated superior performance compared with other NNs (including transformer Networks), achieving errors below 1.60% in a model performance analysis, outperforming other NNs that are considered to be accurate.
[0254] KANs are evaluated for RUL, state estimation, and charging management, whose programming configuration, mathematical properties, and network architecture can be further explored in the energy industry due to its exceptional performance metrics.KAN's continuous multivariate function property represents a superposition of continuous univariate functions, which allows it to model highly nonlinear systems with fewer layers than traditional network architectures. KAN's uniqueness achieves better accuracy with a more compact structure, reducing computational overhead, facilitating more informed decision-making, and providing better adaptability to dynamic patterns overtime. The KAN fills the technical gaps in this research not only by delivering competitive or superior performance with less than a 5.60% error rate in all performance metrics for each case study, but also by addressing key challenges such as efficiency, interpretability, and adaptability to energy storage systems.
[0255] In the context of the future opportunities of a KAN within an energy framework, the disclosure introduces the combination of Al methods to monitor the actual functioning of a BESS based on the user's needs, which vary significantly depending on the battery properties, available datasets, and experimental tests. Table 15 provides a summary of the challenges associated with predictive maintenance of a BESS addressed by this research.
[0256] Table 15. Challenges addressed by the current research
[0257]
[0258] In this description, the promising KANs were proposed, designed, compared to other NN categories, and employed to validate the experimental testing of different battery cells, marking a crucial step in the state estimation, charging management, and RUL applications of a BESS. This strategy was realized through the implementation of several computer science techniques that include Regularization, cross-validation, and Fine-tuning through Bayesian optimization to transform initial networks into several Al models, capable of emulating the intricacies of the BESSs. It was scientifically demonstrated that the design and execution of a KAN shows an optimal performance for battery development and Al-powered technology, answering the research question. Regarding the implications of the broader energy storage sector, this contribution leads to the initiative to promote ties of collaboration between different private administrations to establish the beginning ofremarkable agreements in the domains of energy storage systems, sustainability, model serving, Al, and interpretability, model resource management techniques, and High-Performance Modeling. From the quantitative perspective, the proposed KAN presents a unique and powerful alternative to traditional NNs for the algorithm design of a BESS, thus offering a balanced combination of simplicity, efficiency, and interpretability that engineers, researchers, and stakeholders can use to understand failure mechanisms, optimize performance, and make informed decisions about maintenance or replacement. Regarding the quantitative point of view, its compactness, robustness, and ability to generalize complex nonlinear systems make KAN particularly well-suited for predicting SOC, charging management, and RUL with less than 4.42% in terms of MSE, MAE, and RMSE, while a maximum of 5.56% and 2.12% forSMAPE and RSS.
[0259] For example, when a KAN-based model is applied to consider RUL, predictions closely track the actual capacity degradation trends, and both long-term degradation and sharp transitions in CALCE and NASA datasets are captured effectively. Considering state estimation, the SOC predictions over time align closely with the actual values, including transitions and gradual changes. In charging management, curve prediction is nearly identical to the actual data, demonstrating excellent accuracy that captures the nonlinear relationship between SOC and OCV, a matter of strategic importance for battery modelling.
[0260] In conclusion, this disclosure highlights KAN’s exceptional predictive prowess across a wide range of battery-related datasets. By adeptly capturing both long term trends and intricate nonlinear dynamics, KAN is a transformative tool in battery modelling and Time Series prediction.
[0261] Among the most important contributions of this disclosure are: 1) Proposing, validating, and evaluating KANs in the energy domain focused on a BESS, 2) Development of promising KAN architectures based on user needs, network hyperparameters, and battery chemistry, 3) Bayesian optimization that generates stochastic models for Health and Charge indicators in a BESS, 4) Algorithm design to improve RUL, state estimation, and charging management applications using KANs and a Deep Learning methodology in an energy framework.
[0262] It will be appreciated by those skilled in the art that one or more specific example embodiments have been described; however, other embodiments, including variations and modifications of the described examples, are possible within the scope of the disclosure.
Claims
- 32 - CLAIMS1. A battery management system comprising a Kolmogorov-Arnold Network (KAN) model for real-time state estimation and predictive analytics, wherein the battery management system is configured to use the KAN model to dynamically update control parameters to optimize battery performance of a battery energy storage system.
2. A battery management system as claimed in claim 1 , wherein the KAN model utilizes learnable activation functions.
3. A battery management system as claimed in claim 1 or 2, wherein the KAN model is implemented using a combination of PyTorch and Keras and optionally TensorFlow APIs for adaptability across different battery chemistries and industrial scenarios.
4. A method for battery energy storage system (BESS) management using a KAN model, wherein the method comprises:collecting experimental battery datasets under various operating conditions;training the KAN model using univariate function decomposition for estimating one or more battery performance indicators;using the KAN model to predict State of Charge (SOC) and / or State of Health (SOH) and / or Open Circuit Voltage (OCV) for the BESS; andusing the predicted State of Charge and / or State of Health to optimize charging management of the BESS by adjusting one or more control parameters based on real-time data.
5. A method as claimed in claim 4, comprising using the KAN model to predict Remaining Useful Lifetime (RUL) of battery cells of the BESS using historical charge and discharge cycles.
6. A method as claimed in claim 4 or 5, wherein the KAN model is integrated with one or more Equivalent Circuit Models for SOC and / or SOH estimation.
7. A method as claimed in any one of claims 4 to 6, comprising using the KAN model to predict both State of Charge and State of Health for the BESS.
8. A method as claimed in any one of claims 4 to 7, comprising using the KAN model to predict State of Charge and State of Health and Open Circuit Voltage for the BESS.
9. A method of optimizing charging profiles in a BESS, comprising using a KAN-based predictive model to adjust one or more charging parameters in response to real-time battery performance data.
10. A method as claimed in claim 8, comprising using the KAN-based predictive model to adjust a C-rate and / or cut-off voltage and / or SOC limit in response to real-time battery performance data.
11. A method of machine learning-driven predictive maintenance for a battery system, the method comprising using a KAN model to detect performance degradation and / or failure conditions in the battery system.
12. A processing system configured to perform a method as claimed in any one of claims 4 to- 33 - 13. A software product comprising instructions that, when executed by a processing system, cause the processing system to perform a method as claimed in any one of claims 4 to 11.
14. A computer-readable medium storing instructions that, when executed by a processing system, cause the processing system to perform a method as claimed in any one of claims 4 to 11.