Battery capacity data cleaning method and device, computer equipment and storage medium

Through the multi-model dynamic selection and iterative optimization strategies, outliers in the battery capacity data are automatically identified and removed, which solves the problem of noise and abnormal data collected in the data, and improves the accuracy of battery health status evaluation and life prediction.

CN120372182APending Publication Date: 2025-07-25SHENZHEN EACOMP TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510515170.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the process of battery health status evaluation and life expectancy prediction, the collected capacity data is mixed with noise and abnormal data, affecting the reliability of subsequent evaluation.

Method used

Using multi-model dynamic selection and iterative optimization strategies, we try multiple fit models and automatically select the best model based on the goodness of fit indicators, iteratively detect and remove outliers until the optimal fit and data cleaning effect is achieved.

Benefits of technology

It improves the accuracy of battery capacity data cleaning and the robustness of the model, improves the description accuracy of the capacity retention curve, and provides a more reliable data basis for battery life prediction and performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372182A_ABST
    Figure CN120372182A_ABST
Patent Text Reader

Abstract

The invention relates to a battery capacity data cleaning method and device, computer equipment and a storage medium. The method comprises the following steps: preprocessing collected original capacity data of a battery under a plurality of charge-discharge cycle times to obtain to-be-cleaned capacity data; selecting a fitting model with the highest goodness of fit from the plurality of fitting models according to the to-be-cleaned volume data, and performing abnormal data cleaning on the to-be-cleaned volume data according to a fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate volume data; and selecting the fitting model with the highest goodness of fit from the plurality of fitting models again according to the candidate capacity data, performing abnormal data cleaning on the candidate capacity data according to a fitting result corresponding to the fitting model with the highest goodness of fit to obtain new candidate capacity data, and iteratively executing the process until a preset iteration termination condition is met. And obtaining capacity data of the cleaned battery. By adopting the method, the accuracy of battery capacity data cleaning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of new energy technologies, and particularly to a method and device for cleaning battery capacity data, a computer device, and a storage medium. Background Art

[0002] As the performance of batteries has become the focus of the industry and scientific research, the assessment of the battery health state and the prediction of the battery life during its use are crucial.

[0003] Generally, during the assessment of the battery health state and the prediction of the battery life, it is necessary to collect the capacity data of the battery during the cyclic charge and discharge process. However, due to the instability of the collection device, the disturbance of environmental conditions, or the inconsistency between battery cells, etc., the collected capacity data often contains noise and abnormal data. If the collected capacity data is not cleaned, these abnormal data will affect the reliability of the subsequent prediction of the remaining battery life, maintenance decision-making, and safety assessment.

[0004] Therefore, how to clean the abnormal data in the collected battery capacity data has become an urgent technical problem to be solved. Summary of the Invention

[0005] Based on this, the present application provides a method and device for cleaning battery capacity data, a computer device, and a storage medium, which can improve the accuracy of cleaning battery capacity data.

[0006] In a first aspect, the present application provides a method for cleaning battery capacity data, the method comprising:

[0007] Preprocessing the original capacity data of the battery collected at multiple charge and discharge cycle numbers to obtain the capacity data to be cleaned;

[0008] Selecting the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned, and cleaning the abnormal data of the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data;

[0009] Re-selecting the fitting model with the highest goodness of fit from multiple fitting models according to the candidate capacity data, and cleaning the abnormal data of the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain new candidate capacity data, and iteratively executing this process until a preset iteration termination condition is satisfied to obtain the capacity data with the battery cleaning completed.

[0010] In some embodiments, preprocessing the original capacity data of the battery collected at multiple charge and discharge cycle numbers to obtain the capacity data to be cleaned includes:

[0011] Remove or interpolate the sampled data with empty capacity data in the original capacity data for multiple charge and discharge cycles to obtain the capacity data to be processed;

[0012] Determine the target capacity data for each charge and discharge cycle according to at least one capacity data for each charge and discharge cycle in the capacity data to be processed;

[0013] Perform format conversion on the target capacity data for each charge and discharge cycle to obtain the capacity data to be cleaned.

[0014] In some embodiments, selecting the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned includes:

[0015] Use multiple fitting models to fit the capacity data to be cleaned respectively to obtain multiple fitting results;

[0016] Determine the goodness of fit of each fitting result according to each fitting result and the capacity data to be cleaned;

[0017] Determine the fitting model corresponding to the maximum goodness of fit as the fitting model with the highest goodness of fit.

[0018] In some embodiments, determining the goodness of fit of each fitting result according to each fitting result and the capacity data to be cleaned includes:

[0019] Obtain the constant value range in each fitting model among multiple fitting models, and determine the sum of squared residuals of each fitting result according to each fitting result and the capacity data to be cleaned;

[0020] In the case where the sum of squared residuals of any fitting result is less than or equal to a preset threshold and the constant in any fitting result is within the corresponding constant value range, determine the goodness of fit of any fitting result according to any fitting result and the capacity data to be cleaned;

[0021] In the case where the sum of squared residuals of any fitting result is greater than the preset threshold, and / or the constant in any fitting result is outside the corresponding constant value range, determine the goodness of fit of any fitting result as the target value.

[0022] In some embodiments, performing abnormal data cleaning on the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data includes:

[0023] Determine the predicted capacity data for each charge and discharge cycle according to the fitting result corresponding to the fitting model with the highest goodness of fit;

[0024] Determine each deviation amount between each capacity data in the capacity data to be cleaned and each predicted capacity data;

[0025] Remove the abnormal capacity data and the corresponding charge and discharge cycle numbers from the capacity data to be cleaned according to each deviation amount, and obtain candidate capacity data.

[0026] In some embodiments, removing the abnormal capacity data and the corresponding charge and discharge cycle numbers from the capacity data to be cleaned according to each deviation amount to obtain candidate capacity data includes:

[0027] Obtain the highest residual threshold and the lowest residual threshold, and determine the target residual threshold according to the highest residual threshold, the lowest residual threshold, and each deviation amount, and determine at least two target deviation amounts greater than the target residual threshold from each deviation amount;

[0028] When the number of target deviation amounts is greater than the preset number, determine the capacity data corresponding to the largest preset number of target deviation amounts among the at least two target deviation amounts as the abnormal capacity data. When the number of target deviation amounts is less than or equal to the preset number, determine the capacity data corresponding to the at least two target deviation amounts as the abnormal capacity data;

[0029] Remove the abnormal capacity data and the corresponding charge and discharge cycle numbers from the capacity data to be cleaned, and obtain candidate capacity data.

[0030] In some embodiments, determining the target residual threshold according to the highest residual threshold, the lowest residual threshold, and each deviation amount includes:

[0031] Obtain the residual threshold adjustment parameter, and determine the standard deviation of the deviation amount according to each deviation amount;

[0032] Determine the dynamic threshold according to the residual threshold adjustment parameter and the standard deviation of the deviation amount;

[0033] Determine the larger value of the dynamic threshold and the lowest residual threshold as the target threshold;

[0034] Determine the smaller value of the highest residual threshold and the target threshold as the target residual threshold.

[0035] In a second aspect, the present application provides a battery capacity data cleaning device, and the device includes:

[0036] A preprocessing module, configured to preprocess the original capacity data of the battery collected at multiple charge and discharge cycle numbers to obtain capacity data to be cleaned;

[0037] A data cleaning module, configured to select the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned, and perform abnormal data cleaning on the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data;

[0038] The data cleaning module is also used to re-select the fitting model with the highest goodness of fit from multiple fitting models according to the candidate capacity data, and perform abnormal data cleaning on the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit, so as to obtain new candidate capacity data. This process is iteratively executed until the preset iteration termination condition is met, and the capacity data with the battery cleaning completed is obtained.

[0039] In a third aspect, the present application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of the first aspects are implemented.

[0040] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method according to any one of the first aspects are implemented.

[0041] In the technical solution provided by the embodiments of the present application, by re-selecting the fitting model with the highest goodness of fit from multiple fitting models according to the candidate capacity data, and performing abnormal data cleaning on the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit, new candidate capacity data is obtained. This process is iteratively executed until the preset iteration termination condition is met, and the capacity data with the battery cleaning completed is obtained. Thus, abnormal data can be automatically identified and removed in each iteration, continuously improving the data cleanliness and the robustness of the model. This iterative process not only helps to improve the description accuracy of the model for the capacity retention rate curve, but also provides a more reliable data basis for subsequent battery life prediction and performance evaluation. In addition, this technical solution that combines multi-model selection and iterative anomaly detection will help to promote the development of battery health management technology towards higher accuracy and higher intelligence. Therefore, the embodiments of the present application can improve the accuracy of battery capacity data cleaning. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0043] Figure 1 It is a schematic flowchart of a method for cleaning battery capacity data provided in the first embodiment;

[0044] Figure 2 It is a schematic flowchart of a method for cleaning battery capacity data provided in the second embodiment;

[0045] Figure 3Flow schematic diagram of a battery capacity data cleaning method provided for the third embodiment;

[0046] Figure 4 Flow schematic diagram of a battery capacity data cleaning method provided for the fourth embodiment;

[0047] Figure 5 Flow schematic diagram of a battery capacity data cleaning method provided for the fifth embodiment;

[0048] Figure 6 Structural schematic diagram of a battery capacity data cleaning device provided for some embodiments;

[0049] Figure 7 Structural schematic diagram of a computer device provided for some embodiments. Detailed implementation manners

[0050] The embodiments of the technical solution of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solution of the present application more clearly, and therefore are only examples and cannot be used to limit the protection scope of the present application.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion.

[0052] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of the present application, "a plurality of" means more than two unless otherwise specifically defined. In the description of the embodiments of the present application, "each" means each or every one of a plurality of unless otherwise specifically defined.

[0053] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0054] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates that the associated objects before and after are in an "or" relationship.

[0055] During the process of outlier cleaning, it usually relies on the fitting of a single model (such as a linear or simple exponential model) and simple statistical thresholds to determine outliers. However, the expression ability of a single model for complex decay laws is limited and cannot fully adapt to the diverse decay characteristics of different batteries under multiple usage conditions. In addition, it is difficult to completely eliminate all deviation points by cleaning data through a single outlier detection, resulting in unstable model fitting and low prediction accuracy. Some studies have explored machine learning models or methods such as clustering and isolation forests for outlier detection, but these methods still cannot accurately identify outlier data.

[0056] The detection methods for outlier data in the related art mainly have the following technical problems: Using a single model to fit the capacity decay data (the health state of the battery) is difficult to adapt to the complex decay laws of different batteries under multiple usage conditions, resulting in insufficient model accuracy and robustness; The outlier detection process is mostly one-time or lacks an iterative optimization mechanism and fails to optimize the remaining data again after removing some outliers, thus leading to insufficient data cleaning and it is difficult to improve the final model fitting accuracy and data quality; There is a lack of a mechanism for multi-model comparison and selection and it is unable to automatically select the optimal model according to the data characteristics, affecting the adaptability to complex data sets and reducing the accuracy of the final analysis and prediction.

[0057] The embodiments of the present application can continuously improve the data cleanliness and the robustness of the model by trying multiple fitting models (such as double exponential, linear exponential hybrid models, etc.) and automatically identifying and removing outliers using residual analysis in each iteration. This iterative process not only helps to improve the description accuracy of the model for the capacity retention rate curve, but also provides a more reliable data basis for subsequent battery life prediction and performance evaluation. This technical solution that integrates multi-model selection and iterative outlier detection will contribute to promoting the development of battery health management technology towards higher accuracy and higher intelligence.

[0058] The innovation point of the embodiments of the present application lies in introducing a strategy that combines multi-model dynamic selection and iterative optimization: By simultaneously trying multiple fitting models and automatically selecting the best model according to the data characteristics and goodness-of-fit indicators, and then detecting and removing outliers in an iterative manner based on this best model. Each round of iteration will re-fit and re-evaluate the model according to the updated data until the optimal fitting and data cleaning effects are achieved, thereby effectively improving the accuracy and robustness of the capacity retention rate data analysis.

[0059] The battery in the embodiments of the present application can be a lithium-ion battery or other types of batteries, and the embodiments of the present application do not limit this.

[0060] In the embodiments of the present application, the capacity data can be the current capacity value of the battery when it is fully charged (for example, the state of charge of the battery is 100%), or the capacity data can be the ratio of the current capacity value of the battery when it is fully charged (for example, the state of charge of the battery is 100%) to the rated capacity of the battery (also known as the capacity retention rate).

[0061] The computer device in the embodiments of the present application can be a combination of one or at least two of the following: battery management system, server, mobile phone (MobilePhone), tablet computer (Pad), computer with transceiver function, handheld computer, desktop computer, personal digital assistant, portable media player, smart speaker, navigation device, smart watch, smart glasses, wearable devices such as smart necklaces, pedometers, digital TV, virtual reality (VirtualReality, VR) devices, augmented reality (AugmentedReality, AR) devices, devices in industrial control (Industrial Control), devices in self-driving (Self Driving), devices in remote medical surgery (Remote Medical Surgery), devices in smart grid (Smart Grid), devices in transportation safety (Transportation Safety), devices in smart city (Smart City), devices in smart home (Smart Home), vehicles, in-vehicle devices, in-vehicle modules, and so on.

[0062] Figure 1 A flowchart of a method for cleaning battery capacity data provided for the first embodiment is shown as Figure 1 shown. This method is applied to a computer device, and the method includes:

[0063] S101. Preprocess the original capacity data of the battery collected at multiple charge and discharge cycles to obtain the capacity data to be cleaned.

[0064] In some embodiments, the battery mentioned in the present application can be a battery of one specification, or can be a battery of one specification produced by one production line, or can be a battery that has been developed. Exemplarily, data can be collected for one battery, or data can be collected for at least two batteries of the same specification.

[0065] A charge-discharge cycle refers to a process where a charging cycle includes the battery starting from a fully charged state, gradually discharging until the power is exhausted, and then being charged again to full. Exemplarily, if the battery discharges from a state of charge of 100% to 0% and then is charged back to a state of charge of 100%, this is a complete charge-discharge cycle. The number of charge-discharge cycles refers to the number of times the battery undergoes a complete charge-discharge process from being completely depleted of power to being completely charged during normal use. Among them, the battery being charged to full can mean that the state of charge of the battery is 100%.

[0066] Among them, each original capacity data includes a battery capacity data at each number of charge-discharge cycles. The numbers of charge-discharge cycles can be continuous or discontinuous.

[0067] In some embodiments, the number of battery capacity data at each number of charge-discharge cycles can be one or at least two. For example, the battery capacity data of each battery among multiple batteries at each number of charge-discharge cycles can be obtained. Exemplarily, the multiple batteries can be in the same environment. Also exemplarily, the multiple batteries can be in different environments. For example, at least one of the following is different in the environments where different batteries are located: temperature, humidity, charging rate, discharge rate. For example, if there are two batteries, the charging rate and discharge rate of these two batteries are both different.

[0068] Multiple numbers of charge-discharge cycles can correspond to a preset range of charge-discharge cycle numbers. Exemplarily, the range of charge-discharge cycle numbers can be from 0 to a preset number, and the preset number can be determined according to the attributes of the battery. For example, when the battery design is completed, it is necessary to determine the maximum number of charge-discharge cycles in the attributes of the battery. In this case, multiple capacity sampling data of the battery to be predicted can be collected in real time. When the coefficients of the piecewise polynomial functions of the last continuous N (N is an integer greater than or equal to 2) sampling intervals are the same or the change amount of the coefficients is within a preset range according to the piecewise polynomial functions of each sampling interval, it is determined that the change of the battery capacity data tends to be stable, and the number of the last charge-discharge cycle is determined as the right endpoint of the range of charge-discharge cycle numbers. Also for example, the preset number can be the maximum number of charge-discharge cycles of the battery. Also for example, the preset number can be the product of the maximum number of charge-discharge cycles of the battery and a preset multiple, and the preset multiple can be a real number greater than 0 and less than 1.

[0069] Exemplarily, the capacity data to be cleaned can include preset capacity data at multiple numbers of charge-discharge cycles. Exemplarily, the preset capacity data at each number of charge-discharge cycles is obtained by preprocessing the original capacity data at each number of charge-discharge cycles.

[0070] S102. Select the fitting model with the highest goodness of fit from multiple fitting models according to the data of the capacity to be cleaned, and perform abnormal data cleaning on the data of the capacity to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data.

[0071] The goodness of fit is an index used to measure the degree of closeness (difference degree) between the model value and the actual value when estimating the overall distribution by a regression equation or other models. It is also called the coefficient of determination or goodness of fit. The goodness of fit reflects the fitting degree of the regression line to the observed values, that is, the degree to which the change of the dependent variable can be explained by the independent variable. If the goodness of fit is larger, it indicates that the independent variable has a higher degree of explanation for the dependent variable, meaning that the fitting effect of the model on the data is better; on the contrary, the smaller the goodness of fit, the worse the fitting effect of the model. The goodness of fit of each fitting model can be determined according to the fitting result corresponding to each fitting model and the data of the capacity to be cleaned.

[0072] In some embodiments, the multiple fitting models may include at least two of the following: linear model, quadratic function model, cubic function model, quartic function model, exponential model, linear exponential model, quadratic function exponential model, double exponential model, Weibull distribution model, exponential distribution model, normal distribution model, lognormal distribution model, etc. For example, the multiple fitting models are a linear model, an exponential model, a linear exponential model, and a double exponential model. Exemplarily, the expression of the linear model is: , and the expression of the exponential model is , and the expression of the linear exponential model is , and the expression of the double exponential model is . Among them, , , , are the coefficients of the model.

[0073] The fitting result can be the function corresponding to the fitting model. For example, the expression of the linear model is: , then the function corresponding to this model is . Among them, the constant value of the constant term in the function can be obtained by fitting the data of the capacity to be cleaned. After each fitting model fits the data of the capacity to be cleaned, a corresponding fitting result can be obtained.

[0074] In some embodiments, cleaning the abnormal data from the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data may include: detecting at least one abnormal data from the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit, and removing the abnormal data from the capacity data to be cleaned to obtain candidate capacity data. Exemplarily, among the at least one abnormal data, the number of abnormal data is less than or equal to a preset number to avoid removing too many abnormal data at one time. Wherein, one abnormal data may be preset capacity data at a charge-discharge cycle number.

[0075] S103. Select the fitting model with the highest goodness of fit from multiple fitting models again according to the candidate capacity data, and clean the abnormal data from the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain new candidate capacity data, and iteratively execute this process until the preset iteration termination condition is met, and the capacity data of the battery after cleaning is obtained.

[0076] In some embodiments, meeting the preset iteration termination condition may include meeting any one of the following conditions: the number of iterations reaches the target number; the change amount between the goodness of fit of the fitting models with the highest goodness of fit in each adjacent two iterations in at least two iterations is less than or equal to a set threshold; among the deviation amounts between each capacity data and each predicted capacity data in the capacity data to be cleaned, there is no deviation amount exceeding the target residual threshold. Wherein, each predicted capacity data is determined according to the fitting result corresponding to the fitting model with the highest goodness of fit.

[0077] In some embodiments, the method for selecting the fitting model with the highest goodness of fit in different rounds is the same, and the method for cleaning abnormal data in different rounds is the same.

[0078] The following describes the implementation manner of the battery capacity data cleaning method in the embodiments of the present application: Using the capacity data to be cleaned as the initial data in the first round, selecting the fitting model with the highest goodness of fit from multiple fitting models according to the initial data in the first round, and cleaning the abnormal data from the initial data in the first round according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain the initial data in the second round, determining whether the preset iteration termination condition is met, if not, selecting the fitting model with the highest goodness of fit from multiple fitting models according to the initial data in the second round, and cleaning the abnormal data from the initial data in the second round according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain the initial data in the third round, determining whether the preset iteration termination condition is met, and so on, until the preset iteration termination condition is met, and the capacity data of the battery after cleaning is obtained.

[0079] In some embodiments, after selecting the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned, and performing abnormal data cleaning on the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data, the goodness of fit of the multiple fitting models can be sorted from high to low, and the first set number of fitting models can be obtained. In subsequent iterative processes, only the first set number of fitting models are used for fitting, thereby improving the efficiency of obtaining the capacity data after the battery cleaning is completed.

[0080] In the technical solution provided by the embodiments of the present application, by re-selecting the fitting model with the highest goodness of fit from multiple fitting models according to the candidate capacity data, and performing abnormal data cleaning on the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain new candidate capacity data, and iteratively executing this process until the preset iteration termination condition is met to obtain the capacity data after the battery cleaning is completed. Thus, abnormal data can be automatically identified and removed in each iteration, continuously improving the data cleanliness and the robustness of the model. This iterative process not only helps to improve the description accuracy of the model for the capacity retention rate curve, but also provides a more reliable data basis for subsequent battery life prediction and performance evaluation. In addition, this technical solution that combines multi-model selection and iterative anomaly detection will help to promote the development of battery health management technology towards higher accuracy and higher intelligence. Therefore, the embodiments of the present application can improve the accuracy of battery capacity data cleaning.

[0081] In some embodiments, in the case of obtaining the capacity data after the battery cleaning is completed, the curve of the fitting result corresponding to the fitting model with the highest goodness of fit obtained last time can be determined as the overall spline fitting curve.

[0082] In some other embodiments, the capacity data after the battery cleaning is completed includes multiple capacity sampling data; each of the capacity sampling data includes the battery capacity data at each charge-discharge cycle number, and the following steps can also be performed: determining multiple sampling nodes according to the charge cycle numbers in each of the capacity sampling data, and dividing the multiple capacity sampling data into multiple sampling intervals according to the multiple sampling nodes; performing polynomial function fitting on the capacity sampling data in each of the sampling intervals according to the preset function continuity condition and boundary condition to obtain the piecewise polynomial function of each of the sampling intervals; constructing the overall spline fitting curve corresponding to the multiple capacity sampling data according to the piecewise polynomial functions of each of the sampling intervals; the overall spline fitting curve represents the battery capacity data attenuation trajectory of the battery to be predicted.

[0083] In some embodiments, according to the overall spline fitting curve, the capacity data at the target charge-discharge cycle number can be determined. Wherein, the target charge-discharge cycle number is within the range of charge-discharge cycle numbers corresponding to the overall spline fitting curve. The target charge-discharge cycle number can be one or at least two charge-discharge cycle numbers.

[0084] In some embodiments, according to the capacity data at the target charge-discharge cycle number, the detection result of the battery to be predicted can be determined. For example, if the capacity data at the target charge-discharge cycle number is within the preset range, it is determined that the detection result of the battery to be predicted is passed, indicating that the performance of the battery to be predicted meets the requirements. On the contrary, if the capacity data at the target charge-discharge cycle number is not within the preset range, it is determined that the detection result of the battery to be predicted is not passed, indicating that the performance of the battery to be predicted does not meet the requirements.

[0085] In some embodiments, determining a plurality of sampling nodes according to the charge-discharge cycle numbers in each of the capacity sampling data, and dividing the plurality of capacity sampling data into a plurality of sampling intervals according to the plurality of sampling nodes includes: obtaining a plurality of capacity sampling data to be analyzed with the current sampling node as the starting sampling node, and determining the coefficient of variation of the plurality of capacity sampling data to be analyzed; the first current sampling node is the first capacity sampling data; according to the coefficient of variation, determining the next sampling node from the plurality of capacity sampling data to be analyzed; dividing the capacity sampling data between the current sampling node and the next sampling node into one sampling interval; using the next sampling node as the new current sampling node, and repeating the above steps to divide the plurality of capacity sampling data into the plurality of sampling intervals.

[0086] In some embodiments, determining the next sampling node from the plurality of capacity sampling data to be analyzed according to the coefficient of variation includes: in the case where the coefficient of variation is less than or equal to the preset coefficient, determining the last capacity sampling data among the plurality of capacity sampling data to be analyzed as the next sampling node; in the case where the coefficient of variation is greater than the preset coefficient, determining the next sampling node from the plurality of capacity sampling data to be analyzed according to the ratio between the coefficient of variation and the preset coefficient.

[0087] In some embodiments, the method further includes: determining that the first derivative value of the first sampling node in the first piecewise polynomial is the first target value and the second derivative value is the second target value, and the first derivative value of the last sampling node in the last piecewise polynomial is the third target value and the second derivative value is the fourth target value as the boundary conditions; determining that the first derivative values and the second derivative values of the common sampling nodes of two adjacent sampling intervals are the same in the two adjacent sampling intervals as the function continuity conditions.

[0088] In some embodiments, the method further includes: determining a first derivative value of the common sampling node according to the common sampling node and the previous capacity sampling data of the common sampling node; determining a second derivative value of the common sampling node according to the common sampling node, the previous capacity sampling data of the common sampling node, and the next capacity sampling data of the common sampling node.

[0089] In some embodiments, constructing the overall spline fitting curve corresponding to the plurality of capacity sampling data according to the piecewise polynomial functions of the respective sampling intervals includes: obtaining an extended interval of the plurality of sampling intervals; determining a piecewise polynomial function of the extended interval according to the piecewise polynomial functions of the last at least one sampling interval; and combining the curves corresponding to the piecewise polynomial functions of the respective sampling intervals and the curve corresponding to the piecewise polynomial function of the extended interval to obtain the overall spline fitting curve corresponding to the plurality of capacity sampling data.

[0090] Figure 2 A schematic flow chart of a battery capacity data cleaning method provided for the second embodiment is as Figure 2 shown, and this method is applied to a computer device. Figure 2 The difference between this embodiment and Figure 1 the embodiment is that S101 may include S1011 to S1013:

[0091] S1011, removing or interpolating the sampling data with empty capacity data in the original capacity data under multiple charge and discharge cycle numbers to obtain the to-be-processed capacity data.

[0092] Exemplarily, for the case of data missing, interpolation can be used for filling. For example, if the capacity data at the nth charge and discharge cycle is missing, the capacity data at the nth time can be estimated by using the linear interpolation method (such as taking the average value) according to the capacity data at the (n - 1)th and (n + 1)th times.

[0093] S1012, determining the target capacity data at each charge and discharge cycle number according to at least one capacity data at each charge and discharge cycle number in the to-be-processed capacity data.

[0094] Exemplarily, at least one capacity data for each charge-discharge cycle number in the capacity data to be processed can be integrally processed to obtain the target capacity data for each charge-discharge cycle number. Exemplarily, determining the target capacity data for each charge-discharge cycle number according to at least one capacity data for each charge-discharge cycle number in the capacity data to be processed may include: determining the average value, weighted sum value or weighted average value of at least one capacity data for each charge-discharge cycle number in the capacity data to be processed as the target capacity data for each charge-discharge cycle number. For example, the weight value corresponding to the battery in the standard environment may be greater than or equal to the weight value corresponding to the battery in other environments for weighted summation or weighted averaging. The standard environment may be the optimal operating temperature of the battery, or the standard environment temperature may be the average value of the actual operating temperature of the battery.

[0095] S1013. Perform format conversion on the target capacity data for each charge-discharge cycle number to obtain the capacity data to be cleaned.

[0096] Exemplarily, performing format conversion on the target capacity data for each charge-discharge cycle number to obtain a plurality of capacity data to be cleaned may include: converting each charge-discharge cycle number into data in integer format, and converting the target capacity data into data in floating-point format to obtain a plurality of capacity data of the battery to be cleaned.

[0097] In the technical solution provided by the embodiment of the present application, by sequentially performing removal or interpolation processing, integral processing, and format conversion processing on the original capacity data of the battery collected at multiple charge-discharge cycle numbers, the data quality of the data to be cleaned can be obtained.

[0098] Figure 3 It is a schematic flowchart of a method for cleaning battery capacity data provided for the third embodiment. As Figure 3 shown, this method is applied to a computer device. Figure 3 The difference between this embodiment and Figure 1 the embodiment is that S102 may include S1021 to S1024:

[0099] S1021. Use a plurality of fitting models to respectively fit the capacity data to be cleaned to obtain a plurality of fitting results.

[0100] For example, use a linear model to fit the capacity data to be cleaned to obtain the fitting result of the linear model; use an exponential model to fit the capacity data to be cleaned to obtain the fitting result of the exponential model; use a linear-exponential model to fit the capacity data to be cleaned to obtain the fitting result of the linear-exponential model; use a double-exponential model to fit the capacity data to be cleaned to obtain the fitting result of the double-exponential model.

[0101] In some embodiments, in order to improve the speed of determining the fitting result, the fitting of the data of the capacity to be cleaned by multiple fitting models can be performed in parallel. For example, multiple computing modules in a computer device can respectively use multiple fitting models to fit the data of the capacity to be cleaned, and obtain multiple fitting results.

[0102] S1022. Determine the goodness of fit of each fitting result according to each fitting result and the data of the capacity to be cleaned.

[0103] In some embodiments, the determination of the goodness of fit of each fitting result according to each fitting result and the data of the capacity to be cleaned can be achieved in the following manner: obtain the constant value range in each fitting model among multiple fitting models, and determine the sum of squared residuals of each fitting result according to each fitting result and the data of the capacity to be cleaned; in the case that the sum of squared residuals of any fitting result is less than or equal to a preset threshold and the constant in any fitting result is within the corresponding constant value range, determine the goodness of fit of any fitting result according to any fitting result and the data of the capacity to be cleaned; in the case that the sum of squared residuals of any fitting result is greater than the preset threshold and / or the constant in any fitting result is outside the corresponding constant value range, determine the goodness of fit of any fitting result as a target value. Exemplarily, the target value can be 0 or negative infinity.

[0104] By defining the upper and lower bounds of the constant value range for each model to limit the search space of the constant and prevent unreasonable constants from appearing during the fitting process. For example: linear model The upper and lower bounds of the constant value range are: . Also for example, exponential model The upper and lower bounds of the constant value range are: .

[0105] S1023. Determine the fitting model corresponding to the maximum goodness of fit as the fitting model with the highest goodness of fit.

[0106] Exemplarily, the goodness of fit is usually expressed as R², which reflects the reliability of the regression model in explaining the change of the dependent variable.

[0107] For example, the goodness of fit of the fitting result of the linear model is X1, the goodness of fit of the fitting result of the exponential model is X2, the goodness of fit of the fitting result of the linear-exponential model is X3, and the goodness of fit of the fitting result of the double-exponential model is X4. In the case that X3 is the maximum among X1 to X4, determine the linear-exponential model as the fitting model with the highest goodness of fit.

[0108] S1024. Perform abnormal data cleaning on the data of the capacity to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit, and obtain candidate capacity data.

[0109] In the technical solution provided by the embodiments of the present application, by determining the fitting model corresponding to the maximum goodness of fit as the fitting model with the highest goodness of fit, it is ensured that the model used in each iteration is the fitting model corresponding to the maximum goodness of fit. The fitting model corresponding to the maximum goodness of fit has the best response to the fitted data, thereby improving the accuracy of abnormal data cleaning.

[0110] Figure 4 It is a schematic flowchart of a method for cleaning battery capacity data provided in the fourth embodiment, as Figure 4 shown. This method is applied to a computer device. Figure 4 The difference between this embodiment and Figure 1 the embodiment is that S102 may include S1025 to S1028:

[0111] S1025. Select the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned.

[0112] Among them, for the selection of the fitting model with the highest goodness of fit, please refer to the description in the Figure 3 embodiment, which will not be elaborated here.

[0113] S1026. Determine the predicted capacity data at each charge-discharge cycle number according to the fitting result corresponding to the fitting model with the highest goodness of fit.

[0114] Among them, the predicted capacity data determined at each charge-discharge cycle number may be the predicted capacity data at each charge-discharge cycle number in the capacity data to be cleaned determined.

[0115] Exemplarily, each charge-discharge cycle number in the capacity data to be cleaned can be substituted into the fitting result corresponding to the fitting model with the highest goodness of fit to obtain the predicted capacity data at each charge-discharge cycle number.

[0116] S1027. Determine each deviation amount between each capacity data in the capacity data to be cleaned and each predicted capacity data.

[0117] Each deviation amount may be the absolute value of the difference between each capacity data and each predicted capacity data.

[0118] S1028. Remove the abnormal capacity data and the corresponding charge-discharge cycle numbers from the capacity data to be cleaned to obtain candidate capacity data.

[0119] In some embodiments, each deviation amount can be sorted from large to small, and the preset capacity data at the charge-discharge cycle numbers corresponding to the top preset number of deviation amounts can be determined as the abnormal capacity data, and the abnormal capacity data and the corresponding charge-discharge cycle numbers are removed from the capacity data to be cleaned to obtain candidate capacity data.

[0120] Exemplarily, the preset quantity can be determined according to the amount of collected data, where the amount of collected data can be the number of charge and discharge cycles collected. For example, for the original capacity data at 1000 charge and discharge cycles collected, the amount of collected data is 1000.

[0121] Exemplarily, the value range of the preset quantity can be a value from 1 to 10. For example, the preset quantity can be 1, 2, 3, 5, or 10, etc.

[0122] In the technical solution provided by the embodiment of the present application, first, according to the fitting result corresponding to the fitting model with the highest goodness of fit, the predicted capacity data at each charge and discharge cycle is determined. Then, the deviation amounts between each capacity data in the capacity data to be cleaned and each predicted capacity data are determined. Then, according to each deviation amount, the abnormal capacity data and the corresponding charge and discharge cycles are removed from the capacity data to be cleaned to obtain candidate capacity data. Thus, the removal of the abnormal capacity data and the corresponding charge and discharge cycles is determined according to each deviation amount between each capacity data and each predicted capacity data, improving the accuracy of detecting abnormal capacity data.

[0123] In some embodiments, removing the abnormal capacity data and the corresponding charge and discharge cycles from the capacity data to be cleaned to obtain candidate capacity data includes: obtaining a highest residual threshold and a lowest residual threshold, and determining a target residual threshold according to the highest residual threshold, the lowest residual threshold, and each deviation amount, and determining at least two target deviation amounts greater than the target residual threshold from each deviation amount; determining the abnormal capacity data from the capacity data to be cleaned according to the at least two target deviation amounts, and removing the abnormal capacity data and the corresponding charge and discharge cycles from the capacity data to be cleaned to obtain candidate capacity data.

[0124] In some embodiments, removing the abnormal capacity data and the corresponding charge and discharge cycles from the capacity data to be cleaned to obtain candidate capacity data includes: obtaining a highest residual threshold and a lowest residual threshold, and determining a target residual threshold according to the highest residual threshold, the lowest residual threshold, and each deviation amount, and determining at least two target deviation amounts greater than the target residual threshold from each deviation amount; when the number of target deviation amounts is greater than the preset quantity, determining the capacity data corresponding to the largest preset number of target deviation amounts among the at least two target deviation amounts as the abnormal capacity data, and when the number of target deviation amounts is less than or equal to the preset quantity, determining the capacity data corresponding to the at least two target deviation amounts as the abnormal capacity data; removing the abnormal capacity data and the corresponding charge and discharge cycles from the capacity data to be cleaned to obtain candidate capacity data.

[0125] In the technical solution provided by the embodiment of the present application, first determine the target residual threshold, then determine at least two target deviation amounts greater than the target residual threshold, and then determine the capacity data corresponding to the preset number of target deviation amounts with the largest values among the at least two target deviation amounts as the abnormal capacity data. Thus, the abnormal capacity data is the capacity data greater than the target residual threshold, thereby avoiding the misdetection of abnormal capacity data and improving the accuracy of the detection of abnormal capacity data.

[0126] In some embodiments, according to the highest residual threshold, the lowest residual threshold, and each deviation amount, determining the target residual threshold includes: obtaining a residual threshold adjustment parameter, and determining the standard deviation of the deviation amounts according to each deviation amount; determining a dynamic threshold according to the residual threshold adjustment parameter and the standard deviation of the deviation amounts; determining the larger value of the dynamic threshold and the lowest residual threshold as the target threshold; determining the smaller value of the highest residual threshold and the target threshold as the target residual threshold.

[0127] Exemplarily, the residual threshold adjustment parameter can be a preset fixed value. Exemplarily, the product of the residual threshold adjustment parameter and the standard deviation of the deviation amounts can be determined as the dynamic threshold.

[0128] In the embodiment of the present application, the target threshold is flexibly determined according to each deviation amount, and each deviation amount is determined according to each capacity data and each predicted capacity data in the capacity data to be cleaned. Thus, the determination of the target threshold can be flexibly determined according to the capacity data to be cleaned itself, thereby improving the accuracy of detecting abnormal capacity data.

[0129] In some embodiments, a specific process of an abnormal detection method for battery capacity evaluation is provided. This method effectively identifies and removes abnormal points in the data through multi-model fitting, residual analysis, and iterative optimization, thereby improving the accuracy and reliability of battery capacity data evaluation. The entire process includes five main steps: initialization and parameter setting, model fitting and selection, abnormal detection and removal, iterative optimization, and result analysis and visualization.

[0130] The initialization and parameter setting include the following:

[0131] A. Data input and preprocessing: The system accepts the original data set containing the number of battery cycles (Cycle) and the corresponding performance evaluation value, the state of health of the battery (including in the battery capacity data). The data should be stored in a structured format, and the integrity and accuracy of the data should be ensured. Before performing abnormal detection, it is necessary to perform preliminary cleaning on the input data (i.e., the above-mentioned preprocessing), including at least one of the following: removing missing values or performing missing value imputation, processing duplicate data points, ensuring the consistency of data types (such as the number of cycles is an integer, and the state of health of the battery is a floating point number). Among them, the state of health of the battery includes the battery capacity retention rate.

[0132] B. Model Configuration: The system pre - defines multiple mathematical models to adapt to different types of battery health state assessment data distributions, including but not limited to: linear models, exponential models, linear - exponential models, and double - exponential models. At the same time, upper and lower bounds of parameters (corresponding to the above - mentioned constants) are defined for each model to limit the parameter search space and prevent unreasonable parameter values from appearing during the fitting process.

[0133] C. Parameter Setting: Set a threshold factor (i.e., the residual threshold adjustment parameter mentioned above) to dynamically adjust the parameter of the residual threshold, and its default value is 2.0. Set the R² (goodness of fit) improvement threshold (corresponding to the set threshold mentioned above), a parameter that controls the iteration stop condition. When the R² improvement of the current model is less than this threshold, the iteration stops. For example, when the change in the goodness of fit between the best - fitting models in every two adjacent iterations in at least two iterations is less than or equal to the set threshold, the iteration stops. Exemplarily, the default value of R² is 0.001. Set the maximum number of iterations to limit the maximum number of iterations for anomaly detection and removal to prevent infinite loops. The default value is 10 times. Set the number of outlier points removed in each iteration: control the maximum number of outlier points to be removed in each iteration to avoid removing too many potentially normal data points at once, and the default value is 1 or 2.

[0134] In model fitting and selection, the non - linear least - squares method is used to fit each of the pre - defined models. This algorithm optimizes the model parameters by minimizing the sum of the squared residuals between the actual battery health state and the model prediction values. The process is as follows: First, traverse all pre - defined models. For each model, perform fitting using the least - squares method. If anomalies such as non - convergence or parameters exceeding the boundaries occur during the fitting process, capture the anomalies and record the fitting failure of the model, set R² to negative infinity, and set the prediction values to null values. For each successfully fitted model, calculate its goodness - of - fit index R², and the formula is as follows: . Where, is the actual battery health state (i.e., each capacity data in the capacity data to be cleaned mentioned above), is the model prediction value (i.e., each predicted capacity data mentioned above), is the mean of the battery health state, is the number of data points (i.e., the number of capacity data or the number of charge - discharge cycle times in the capacity data to be cleaned). Among them, non - convergence during the fitting process can include that the sum of the squared residuals of the fitting result is greater than the preset threshold.

[0135] After that, from all the successfully fitted models, select the model with the highest R² value as the best model for this round of iteration. At the same time, add the prediction values of the best model to the current cleaned dataset for subsequent residual calculation and anomaly detection.

[0136] In anomaly detection and removal, the residual of each data point is calculated (i.e., the difference between each capacity data in the data to be cleaned and each predicted capacity data). Then, the absolute value of the residual is calculated , which is used for sorting and threshold comparison. Then, a dynamic threshold can be obtained through the absolute value of the above residual (i.e., the above-mentioned target threshold), and the calculation formula of

[0137] is as follows: is the standard deviation of the absolute value of the residual, represents the threshold factor (i.e., the above-mentioned residual threshold adjustment parameter).

[0138] For this dynamic prediction, according to the decay characteristics of the battery health state, its lower limit can be set to: 0.002, ensuring that there is still a minimum threshold for anomaly detection even when the residual standard deviation is very small. The threshold is dynamically adjusted according to the standard deviation of the residual to ensure adaptability under different data distributions.

[0139] Then, anomaly point identification is performed. By comparing the absolute value of each residual with the dynamic threshold , all points whose absolute value of the residual exceeds the dynamic threshold can be identified ( ). Among them, . represents the index of the points in the data to be cleaned where the absolute value of the residual exceeds the dynamic threshold .

[0140] If the number of anomaly points removed in each iteration is set, then the top several points with the largest absolute value of the residual are selected from the detected anomaly points. After that. The identified anomaly points are removed from the current cleaned dataset to form a new cleaned dataset, preparing for the model fitting of the next iteration. In some embodiments, after removing the anomaly points, the index of the dataset is reset to ensure data consistency for subsequent operations.

[0141] In iterative optimization, during the anomaly detection process, the system gradually improves the fitting effect of the model and accurately identifies anomaly points through the method of iterative optimization. The iterative optimization process starts from the initialization stage. First, the original dataset (i.e., the original capacity data under the above-mentioned multiple charge and discharge cycle numbers) is copied as the dataset to be cleaned (i.e., the data to be cleaned mentioned above), or the dataset to be cleaned is determined according to the original dataset, and detailed information of each iteration is prepared and all anomaly points detected during the whole process are stored. To effectively evaluate the model performance, the system sets an initial R² value for comparing the improvement of the model in subsequent iterations.

[0142] In each iteration, the current cleaned data is first fitted with multiple models, and the model with the highest goodness of fit is selected as the best model for this round. Subsequently, the system records the R² value of this model and compares it with the R² value of the previous iteration to calculate the R² improvement for this round. If this improvement does not reach the preset threshold or the maximum number of predetermined iterations has been reached, the system will stop the iteration process. Otherwise, the system will continue with anomaly detection, identify and remove the detected anomaly points in this round, and record the relevant information of this iteration, including the name of the best model used, the current R² value, the number of removed anomaly points, and the prediction results of the model. Then, the system updates the recorded R² value, increments the iteration counter, and enters the next iteration. Among them, the iteration stop condition can be: stop the iteration when the R² improvement is less than the pre-set threshold (usually 0.001) or reach the maximum number of iterations, then stop the iteration.

[0143] During the entire iteration process, all detected and removed anomaly points will be aggregated and stored for subsequent analysis and reference. When the iteration process ends, the system will update and record the finally selected best model and its corresponding R² value, and output all detected anomaly points.

[0144] Figure 5 A flowchart of a method for cleaning battery capacity data provided for the fifth embodiment is shown as Figure 5 shown, and this method is applied to a computer device.

[0145] S501. Preprocess the original capacity data of the battery at multiple charge and discharge cycle times collected to obtain the capacity data to be cleaned.

[0146] S502. Set parameters and enter the iterative loop process.

[0147] S503. Model fitting and selection, calculate the deviation amounts between each capacity data and each predicted capacity data in the capacity data to be cleaned, and determine the target residual threshold.

[0148] Among them, model fitting and selection can include: fitting all predefined multiple models, calculating the R² value of each model, and selecting the best model with the highest R². The deviation amount can also be referred to as the absolute value of the residual.

[0149] S504. According to each deviation amount and the target residual threshold, detect and select anomaly points, and remove the anomaly points to obtain the capacity data after the i-th round of cleaning.

[0150] Among them, the anomaly points include abnormal capacity data and the corresponding charge and discharge cycle times.

[0151] The detection and selection of outliers may include: identifying data points corresponding to the absolute values of at least two target residuals whose absolute values of the residuals exceed the target residual threshold, and selecting the maximum preset number of data points from the identified data points.

[0152] S505. Check whether the iteration termination condition is satisfied.

[0153] Satisfying the iteration termination condition may include that the R² improvement is less than the threshold or the maximum number of iterations is reached.

[0154] If satisfied, execute S506; if not satisfied, go to S503 for the next round of outlier removal.

[0155] S506. The iteration ends, and the capacity data after the battery cleaning is completed is obtained.

[0156] In the battery capacity anomaly detection method provided by the embodiments of the present application, through iterative multi-model fitting and residual analysis, the accuracy and reliability of battery health state assessment are significantly improved. Specifically, when processing the original data, the system fits by predefining multiple mathematical models, and automatically selects the model with the highest goodness of fit as the best model, so as to effectively capture the true trend of the data. By dynamically setting the residual threshold, the method can flexibly adapt to different data distributions, ensuring the accurate identification and removal of outliers.

[0157] By introducing the iterative optimization mechanism, the anomaly detection process has the ability to gradually purify the data set, avoiding the error accumulation caused by removing too many normal data points at one time, and further improving the fitting quality of the model. Each iteration records the detailed model performance and outlier removal situation, providing comprehensive process transparency for users to monitor and adjust. In addition, the finally generated cleaned data set and all detected outliers provide a high-quality data basis for subsequent analysis and decision-making.

[0158] Through this method, the battery management system can more accurately evaluate the health state of the battery, timely detect potential anomalies, extend the service life of the battery, and improve safety. At the same time, the systematic anomaly detection and removal process has good robustness and adaptability, and is applicable to various different types and scales of data sets, providing strong technical support for multiple application fields such as battery performance research and manufacturing quality control.

[0159] In summary, the method of the present invention overcomes the deficiencies of traditional battery health state assessment methods in outlier processing through multi-model fitting, dynamic residual analysis and iterative optimization, significantly improves the accuracy and reliability of the assessment, and has broad application prospects and significant technical advantages.

[0160] Based on the same inventive concept, an embodiment of the present application further provides a battery capacity data cleaning device for implementing the battery capacity data cleaning method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the battery capacity data cleaning device provided below can refer to the limitations on the battery capacity data cleaning method in the above text, and will not be elaborated here.

[0161] In an exemplary embodiment, Figure 6 is a schematic structural diagram of a battery capacity data cleaning device provided for some embodiments, as Figure 6 shown. The battery capacity data cleaning device 600 includes:

[0162] A preprocessing module 601, configured to preprocess the original capacity data of the battery collected at multiple charge-discharge cycle times to obtain capacity data to be cleaned;

[0163] A data cleaning module 602, configured to select the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned, and perform abnormal data cleaning on the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data;

[0164] The data cleaning module 602 is further configured to re-select the fitting model with the highest goodness of fit from multiple fitting models according to the candidate capacity data, and perform abnormal data cleaning on the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain new candidate capacity data, and iteratively execute this process until a preset iteration termination condition is met to obtain the capacity data with the battery cleaning completed.

[0165] In some embodiments, the preprocessing module 601 includes a unit for obtaining capacity data to be processed, a unit for obtaining target capacity data, and a format conversion unit. Among them, the unit for obtaining capacity data to be processed is configured to remove or interpolate the sampling data with empty capacity data in the original capacity data at multiple charge-discharge cycle times to obtain the capacity data to be processed; the unit for obtaining target capacity data is configured to determine the target capacity data at each charge-discharge cycle time according to at least one capacity data at each charge-discharge cycle time in the capacity data to be processed; the format conversion unit is configured to perform format conversion on the target capacity data at each charge-discharge cycle time to obtain the capacity data to be cleaned.

[0166] In some embodiments, the data cleaning module 602 includes a fitting unit, a goodness-of-fit determination unit, and a fitting model determination unit. The fitting unit is configured to fit the to-be-cleaned capacity data using multiple fitting models respectively to obtain multiple fitting results. The goodness-of-fit determination unit is configured to determine the goodness of fit of each fitting result according to each fitting result and the to-be-cleaned capacity data. The fitting model determination unit is configured to determine the fitting model corresponding to the maximum goodness of fit as the fitting model with the highest goodness of fit.

[0167] In some embodiments, the goodness-of-fit determination unit is further configured to obtain the constant value range in each fitting model among the multiple fitting models, and determine the sum of squared residuals of each fitting result according to each fitting result and the to-be-cleaned capacity data. In the case where the sum of squared residuals of any fitting result is less than or equal to a preset threshold and the constant in any fitting result is within the corresponding constant value range, determine the goodness of fit of any fitting result according to any fitting result and the to-be-cleaned capacity data. In the case where the sum of squared residuals of any fitting result is greater than the preset threshold and / or the constant in any fitting result is outside the corresponding constant value range, determine the goodness of fit of any fitting result as a target value.

[0168] In some embodiments, the data cleaning module 602 includes a predicted capacity data determination unit, a deviation amount determination unit, and a candidate capacity data determination unit. The predicted capacity data determination unit is configured to determine the predicted capacity data at each charge-discharge cycle number according to the fitting result corresponding to the fitting model with the highest goodness of fit. The deviation amount determination unit is configured to determine the deviation amounts between each capacity data in the to-be-cleaned capacity data and each predicted capacity data. The candidate capacity data determination unit is configured to remove the abnormal capacity data and the corresponding charge-discharge cycle numbers from the to-be-cleaned capacity data according to the deviation amounts to obtain candidate capacity data.

[0169] In some embodiments, the candidate capacity data determination unit is further configured to obtain a highest residual threshold and a lowest residual threshold, determine a target residual threshold according to the highest residual threshold, the lowest residual threshold, and the deviation amounts, and determine at least two target deviation amounts greater than the target residual threshold from the deviation amounts. In the case where the number of target deviation amounts is greater than a preset number, determine the capacity data corresponding to the maximum preset number of target deviation amounts among the at least two target deviation amounts as the abnormal capacity data. In the case where the number of target deviation amounts is less than or equal to the preset number, determine the capacity data corresponding to the at least two target deviation amounts as the abnormal capacity data. Remove the abnormal capacity data and the corresponding charge-discharge cycle numbers from the to-be-cleaned capacity data to obtain candidate capacity data.

[0170] In some embodiments, the candidate capacity data determination unit is further configured to obtain a residual threshold adjustment parameter, and determine a standard deviation of the deviation amounts according to each deviation amount; determine a dynamic threshold according to the residual threshold adjustment parameter and the standard deviation of the deviation amounts; determine the larger value between the dynamic threshold and the lowest residual threshold as the target threshold; and determine the smaller value between the highest residual threshold and the target threshold as the target residual threshold.

[0171] The description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0172] Each module in the above battery capacity data cleaning device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of the processor, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0173] In an exemplary embodiment, Figure 7 FIG. 4 is a schematic structural diagram of a computer device provided for some embodiments. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through Wireless Fidelity (WIFI), a mobile cellular network, Near Field Communication (NFC), or other technologies. The computer program, when executed by the processor, implements a method for cleaning battery capacity data. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0174] Those skilled in the art can understand,Figure 7 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0175] For example, a computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method in any of the above embodiments are implemented.

[0176] For example, in an exemplary embodiment, when the processor is used to execute the computer program, it realizes: preprocessing the original capacity data of the battery collected under multiple charge and discharge cycle numbers to obtain the capacity data to be cleaned; selecting the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned, and performing abnormal data cleaning on the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data; reselecting the fitting model with the highest goodness of fit from multiple fitting models according to the candidate capacity data, and performing abnormal data cleaning on the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain new candidate capacity data, and iteratively executing this process until a preset iteration termination condition is met, and obtaining the capacity data after the battery cleaning is completed.

[0177] In one embodiment, a computer-readable storage medium is provided. When the computer program is executed by a processor, the steps of the method provided in any of the above embodiments are implemented.

[0178] For example, in an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are realized: preprocessing the original capacity data of the battery collected under multiple charge and discharge cycle numbers to obtain the capacity data to be cleaned; selecting the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned, and performing abnormal data cleaning on the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data; reselecting the fitting model with the highest goodness of fit from multiple fitting models according to the candidate capacity data, and performing abnormal data cleaning on the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain new candidate capacity data, and iteratively executing this process until a preset iteration termination condition is met, and obtaining the capacity data after the battery cleaning is completed.

[0179] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.

[0180] The processor, each functional module or each functional unit in any embodiment of the present application may include any one or more of the following integrations: general-purpose processor, application specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field programmable gate array (FPGA), central processing unit (CPU), graphics processing unit (GPU), embedded neural network processor (neural-network processing units, NPU), controller, microcontroller, microprocessor, programmable logic device, discrete gate or transistor logic device, discrete hardware component, quantum computing-based data processing logic unit, artificial intelligence (AI) processor, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0181] The memory or computer-readable storage medium in any embodiment of the present application may include at least one of non-volatile memory and volatile memory. The non-volatile memory includes the integration of one or more of the following: Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Ferromagnetic Random Access Memory (FRAM), Flash Memory, magnetic surface memory, optical disc, Compact Disc Read-Only Memory (CD-ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, Resistive Random Access Memory (ReRAM), Magnetoresistive Random Access Memory (MRAM), Ferroelectric Random Access Memory (FRAM), Phase Change Memory (PCM), graphene memory, volatile memory, etc. The volatile memory includes the integration of one or more of the following: Random Access Memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM), etc.

[0182] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.

[0183] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for cleaning battery capacity data, characterized in that, The method includes: Preprocessing the original capacity data of the battery collected at multiple charge-discharge cycle numbers to obtain the capacity data to be cleaned; Selecting the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned, and cleaning the abnormal data of the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data; Re-selecting the fitting model with the highest goodness of fit from the multiple fitting models according to the candidate capacity data, and cleaning the abnormal data of the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain new candidate capacity data, and iteratively executing this process until a preset iteration termination condition is satisfied to obtain the capacity data of the battery after cleaning is completed.

2. The method according to claim 1, wherein The preprocessing the original capacity data of the battery collected at multiple charge-discharge cycle numbers to obtain the capacity data to be cleaned includes: Removing or interpolating the sampling data with empty capacity data in the original capacity data at multiple charge-discharge cycle numbers to obtain the capacity data to be processed; Determining the target capacity data at each charge-discharge cycle number according to at least one capacity data at each charge-discharge cycle number in the capacity data to be processed; Performing format conversion on the target capacity data at each charge-discharge cycle number to obtain the capacity data to be cleaned.

3. The method according to claim 1, characterized in that, The selecting the fitting model with the highest goodness of fit from multiple fitting models according to the capacity data to be cleaned includes: Using the multiple fitting models to respectively fit the capacity data to be cleaned to obtain multiple fitting results; Determining the goodness of fit of each fitting result according to each fitting result and the capacity data to be cleaned; Determining the fitting model corresponding to the largest goodness of fit as the fitting model with the highest goodness of fit.

4. The method according to claim 3, characterized in that, The determining the goodness of fit of each fitting result according to each fitting result and the capacity data to be cleaned includes: Obtaining the constant value range in each fitting model among the multiple fitting models, and determining the sum of squared residuals of each fitting result according to each fitting result and the capacity data to be cleaned; When the sum of squared residuals of any fitting result is less than or equal to a preset threshold and the constant in the any fitting result is within the corresponding constant value range, determining the goodness of fit of the any fitting result according to the any fitting result and the capacity data to be cleaned; When the sum of squared residuals of any fitting result is greater than the preset threshold and / or the constant in the any fitting result is outside the corresponding constant value range, determining the goodness of fit of the any fitting result as the target value.

5. The method according to any one of claims 1 to 4, characterized in that, The cleaning the abnormal data of the capacity data to be cleaned according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data includes: Determining the predicted capacity data at each charge-discharge cycle number according to the fitting result corresponding to the fitting model with the highest goodness of fit; Determining each deviation amount between each capacity data in the capacity data to be cleaned and each predicted capacity data; Based on each of the deviation amounts, abnormal capacity data and corresponding charge and discharge cycle numbers are removed from the to-be-cleaned capacity data to obtain the candidate capacity data.

6. The method according to claim 5, wherein The removing abnormal capacity data and corresponding charge and discharge cycle numbers from the to-be-cleaned capacity data based on each of the deviation amounts to obtain the candidate capacity data includes: Obtaining a maximum residual threshold and a minimum residual threshold, determining a target residual threshold according to the maximum residual threshold, the minimum residual threshold, and each of the deviation amounts, and determining at least two target deviation amounts greater than the target residual threshold from each of the deviation amounts; When the number of the target deviation amounts is greater than a preset number, determining the capacity data corresponding to the maximum preset number of target deviation amounts among the at least two target deviation amounts as the abnormal capacity data, and when the number of the target deviation amounts is less than or equal to the preset number, determining the capacity data corresponding to the at least two target deviation amounts as the abnormal capacity data; Removing the abnormal capacity data and corresponding charge and discharge cycle numbers from the to-be-cleaned capacity data to obtain the candidate capacity data.

7. The method according to claim 6, characterized in that, The determining a target residual threshold according to the maximum residual threshold, the minimum residual threshold, and each of the deviation amounts includes: Obtaining a residual threshold adjustment parameter and determining a standard deviation of the deviation amounts according to each of the deviation amounts; Determining a dynamic threshold according to the residual threshold adjustment parameter and the standard deviation of the deviation amounts; Determining the larger value between the dynamic threshold and the minimum residual threshold as the target threshold; Determining the smaller value between the maximum residual threshold and the target threshold as the target residual threshold.

8. A battery capacity data cleaning device, characterized in that, The device includes: A preprocessing module for preprocessing the original capacity data of the battery at multiple charge and discharge cycle numbers collected to obtain to-be-cleaned capacity data; A data cleaning module for selecting a fitting model with the highest goodness of fit from multiple fitting models according to the to-be-cleaned capacity data, and performing abnormal data cleaning on the to-be-cleaned capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain candidate capacity data; The data cleaning module is further configured to re-select a fitting model with the highest goodness of fit from the multiple fitting models according to the candidate capacity data, and perform abnormal data cleaning on the candidate capacity data according to the fitting result corresponding to the fitting model with the highest goodness of fit to obtain new candidate capacity data, and iteratively execute this process until a preset iteration termination condition is satisfied to obtain the capacity data of the battery after cleaning is completed.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.