Chemical component data statistics method and device, electronic equipment and storage medium
By unifying the processing of multi-source lithium battery formation and capacity data, a data tracking and mapping relationship for the entire life cycle of the cell is established, solving the problem of fragmented multi-source data, realizing efficient multi-dimensional analysis and anomaly localization, and improving the process optimization and quality control capabilities of lithium battery production.
Patent Information
- Application Number
- CN202610875776.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-08-25
AI Technical Summary
The existing data statistics and analysis of lithium battery formation and capacity testing processes suffer from problems such as fragmented data from multiple sources, inefficient correlation, and lack of end-to-end traceability. This makes it impossible to achieve complete data tracking of a single cell throughout its entire production lifecycle, making it difficult to accurately match process parameters with battery performance results, and hindering accurate attribution analysis and process optimization.
By acquiring multi-source heterogeneous capacity and performance data, standardizing and integrating it, constructing a unified dataset, and writing it into a time-series database and a document database, we achieve unified format and dimension-aligned storage. Combined with sliding window incremental calculation and caching mechanisms, we perform segmented and accurate statistics, establish a precise mapping relationship between process parameters and battery performance, and support multi-dimensional analysis and anomaly tracing.
It enables complete data tracking throughout the entire lifecycle of battery cells, supports efficient multi-dimensional analysis and anomaly localization, improves the process optimization and quality control capabilities of lithium battery production, and adapts to the needs of intelligent and refined production.
Smart Images

Figure CN122633761A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, electronic device and storage medium for statistical analysis of fractionated and capacity-bound data. Background Technology
[0002] In the industrial production process of lithium batteries, formation and capacity testing are core processes that determine the battery's electrochemical performance, yield, and batch consistency. During formation and capacity testing, batteries continuously generate massive amounts of high-frequency production data, primarily including core process and performance parameters such as battery voltage, operating current, casing and tab temperature, charge / discharge timeline curves, actual capacity, and charge / discharge efficiency. With the increasing automation and scale of lithium battery production lines, the production data generated in the formation and capacity testing processes exhibits typical characteristics of high-frequency acquisition, multi-source heterogeneity, and strong time-series characteristics. Data acquisition accuracy can reach the second or even millisecond level, and data sources cover multiple independent channels, including production equipment terminals, MES production management systems, and on-site process configuration parameters.
[0003] Currently, data statistics and analysis for lithium battery formation and capacity testing processes in the industry generally adopt traditional manual and offline processing methods, mainly relying on relational databases, Excel tools, and simple scripts to complete data statistics and summarization. While this approach is simple to operate and highly versatile, in large-scale lithium battery production scenarios, it suffers from core technical defects such as fragmented multi-source data, inefficient correlation, and inability to achieve full-chain traceability. This severely restricts the accuracy and effectiveness of lithium battery production process optimization, quality control, and anomaly tracing. Specifically, existing data storage and management models employ fragmented storage logic, storing various types of production data independently according to single equipment and single production batches. Data from multiple data sources, including real-time equipment operating condition data, production ledger data retained by the MES system, and process parameter data configured on-site, lacks a unified correlation index and data linkage mechanism, resulting in a severe data silo problem prevalent in the industry. This core defect directly leads to the inability to achieve complete data tracking throughout the entire production lifecycle of a single cell, making it difficult to accurately match and construct a one-to-one correspondence between process parameters and the final battery performance results, and failing to fully understand the correlation patterns between different process configurations, production conditions, and battery yield and performance parameters. Because multi-source data cannot be effectively correlated and end-to-end traceability is lacking, the industry struggles to conduct accurate attribution analysis for differences in battery performance and fluctuations in production quality. This makes it difficult to provide complete and effective data support for process iteration and optimization of the formation and capacity testing process, refined control of production quality, and precise location of production anomalies, thus failing to meet the needs of intelligent and refined lithium battery production development. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this application provides a method, apparatus, electronic device, and storage medium for statistical analysis of formation and capacity testing data. This provides complete, accurate, and reliable data support for process iteration and optimization of lithium battery formation and capacity testing, refined control of production quality, and accurate attribution and rectification of production anomalies, effectively adapting to the needs of intelligent and refined large-scale production of lithium batteries.
[0005] A first aspect of this application provides a method for statistical analysis of composition volumetric data, the method comprising: A second aspect of this application provides a device for calculating composition and capacity data, the device comprising: A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the batching capacity data statistics method.
[0006] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method for calculating and statistically analyzing batch data.
[0007] In summary, the method, apparatus, electronic device, and storage medium for composition and capacity testing provided in this application acquire multi-source heterogeneous composition and capacity testing data from various channels, such as composition and capacity testing equipment terminals, MES production management systems, and on-site process configuration parameters. It standardizes and integrates raw data with inconsistent formats, dimensions, and standards, eliminating data barriers and format differences between different data sources. This results in a unified dataset with consistent format, aligned dimensions, and the ability to be linked and accessed, breaking the state of independent storage and fragmentation of multi-source data. Based on this, the unified dataset is written into a time-series database and a document database respectively, achieving complete retention of raw time-series data and orderly storage of statistical analysis data. This hierarchical and categorized storage model achieves standardized and integrated management of massive multi-source data, completely abandoning the traditional fragmented and distributed data storage model. Furthermore, this application relies on sliding window incremental calculation combined with a caching mechanism to perform segmented and precise statistics on the unified dataset according to different process stages of composition and capacity testing, efficiently obtaining production and performance statistical results corresponding to each process stage, ensuring that multi-source fused data can achieve refined and efficient statistical analysis. Furthermore, a multi-dimensional analysis model is constructed based on precise segmented statistical results to achieve in-depth cross-analysis of process parameters and battery performance data, effectively establishing a precise mapping relationship between different process configurations, production conditions, and performance indicators such as battery capacity and charge / discharge efficiency. In addition, intelligent judgment of production anomalies is completed by combining real-time statistical results with historical data trends. Relying on the unique cell identifier centrally bound to a unified dataset, abnormal data is traced back in all dimensions and throughout the entire process, accurately locating the corresponding process parameters, production stages, and production data, achieving complete data tracking throughout the entire battery cell production lifecycle. Attached Figure Description
[0008] Figure 1 This is a flowchart illustrating a method for statistical analysis of composition and capacity data in an embodiment of this application; Figure 2 This is a functional block diagram of a batching and capacity data statistics device shown in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application. Detailed Implementation
[0009] The present application will be further described below with reference to the accompanying drawings and embodiments.
[0010] The following will clearly and completely describe the concept, specific structure, and resulting technical effects of this application in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, features, and effects of this application. Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the scope of protection of this application. Furthermore, all connections / linkages involved in the patent do not simply refer to direct contact between components, but rather to the ability to form a better connection structure by adding or reducing connecting accessories according to specific implementation conditions. The various technical features in this application can be combined interactively without contradicting each other.
[0011] Terminology Explanation: Formation and Grading: Key processes in lithium-ion battery manufacturing, comprising two steps: formation and grading. Formation involves the initial charge and discharge cycle of the battery to activate its internal electrochemical system, while grading involves testing the battery's discharge capacity and classifying it according to capacity.
[0012] Reference Figure 1 The diagram shown is a flowchart illustrating a method for statistical analysis of formation and capacity data according to an embodiment of this application. The method includes the following steps.
[0013] S11, acquire multi-source chemical composition and capacity data collected by the chemical composition and capacity device, and perform standardization processing on the multi-source chemical composition and capacity data to obtain a unified dataset.
[0014] In some embodiments, during the lithium battery formation and capacity testing process, the time-series data of the battery cells during processing is collected in real time by the formation and capacity testing production equipment. This includes voltage data V(t), current data I(t), temperature data T(t), and the corresponding sampling timestamp t. Simultaneously, business attribute data corresponding to the battery cells is synchronously obtained from the MES system and the process configuration system. This includes the unique identifier of the battery cell (CellID), batch number (BatchID), product model (Model), and the process parameters used in this processing.
[0015] The collected multi-source heterogeneous data is integrated into the data processing unit of the electronic device. Before entering the storage and computing stages, it is standardized in a unified format to construct a unified data structure, for example...<CellID,Timestamp,V,I,T,BatchID,Model,ProcessParam> This allows for the association, storage, and computation of equipment-acquired data, MES data, and process parameter data within the same structure. The Timestamp is a millisecond-level Unix timestamp corresponding to the data sampling time, used to uniquely identify the acquisition time of each data point and serving as the time reference for timing alignment, stage division, and window calculation. The ProcessParam is a set of process configuration parameters used in the cell formation and capacity testing process, including at least the charge / discharge rate, charging cut-off voltage, discharging cut-off voltage, resting time, and temperature setpoint, used to characterize the cell's processing conditions.
[0016] In the standardization process, the first step is to align the timestamps of the multi-source data. All data sources are uniformly marked with millisecond-level Unix timestamps. When there is a clock deviation between different acquisition devices, the MES system clock is used as the reference clock to perform linear interpolation alignment on the voltage, current, and temperature data sequences, ensuring that the multi-dimensional data of the same cell at the same time maintains temporal consistency. Simultaneously, missing value handling is performed on the data sequences. If the voltage or current data of a single sampling point is missing, the average of the valid data from the adjacent times before and after that point is used to fill the gap, maintaining data continuity. If three or more consecutive sampling points from the same cell channel have missing data, the data for that period is deemed invalid, discarded, and a data anomaly alarm is triggered to the upper-level system. Based on this, a globally unique identifier, CellID, is generated according to preset coding rules. The coding structure is 8-digit batch number + 4-digit device number + 3-digit channel number + 8-digit production date (YYYYMMDD). This combination ensures that each cell has a unique identifiable identifier throughout the entire production line and process, providing a foundation for subsequent full lifecycle tracking. Finally, the process parameters were standardized and normalized. Parameters such as charge / discharge rate, cutoff voltage, and settling time in ProcessParam were converted into floating-point values and their units were normalized. For example, current was standardized to A, voltage to V, and temperature to °C, eliminating dimensional differences caused by different equipment and process configurations, and providing a standardized data foundation for subsequent unified calculations and multidimensional analysis.
[0017] After standardization, a unified dataset with consistent format, time alignment, unique index, and unified units is obtained, which can be directly used as input data for subsequent hierarchical storage, real-time computing, multidimensional analysis, and anomaly tracing. By constructing a unified data index model with a unique cell identifier and timestamp as its core, the defect of inefficient correlation between multi-source data is solved, achieving unified alignment of equipment data, MES data, and process parameters, laying the foundation for full lifecycle tracking of battery cells.
[0018] S12, the unified dataset is written into the time-series database for raw data storage.
[0019] To adapt to the massive, high-frequency, and highly time-series data storage characteristics of the batching and capacity-building process, and to ensure data writing efficiency, query performance, and subsequent full lifecycle tracking requirements, this step performs structured configuration of the time-series database and completes the persistent storage of the original data based on the configuration results.
[0020] In an optional implementation, the persistent storage of the unified dataset includes: Configure storage rules for the time-series database. The storage rules are based on the unique identifier of the battery cell as the partitioning basis and the timestamp corresponding to the data sampling time as the sorting basis. The unified dataset includes the unique identifier of the battery cell, the timestamp corresponding to the data sampling time, and the batch number. Configure a partitioning strategy for the storage rules. The partitioning strategy performs hash partitioning based on the unique identifier of the battery cell, dividing the data into a preset number of logical partitions. Each logical partition is time-sharded according to a preset time period. According to the storage rules and the partitioning strategy, an index structure is established in the time-series database. The index structure includes a first-level index that combines the unique identifier of the battery cell and the timestamp corresponding to the data sampling time, and a second-level index that combines the batch number and the timestamp corresponding to the data sampling time. The unified dataset is written into the time-series database for persistent storage according to the storage rules, the partitioning strategy, and the index structure.
[0021] After completing the collection and standardization of multi-source data from the formation and capacity testing process, a unified dataset with standardized format and fields is obtained. The unified dataset includes at least the unique identifier of the battery cell, the timestamp corresponding to the data sampling time, and the batch number.
[0022] In some embodiments, the electronic device can pre-configure the storage rules, partitioning strategy, index structure and writing method of the time series database, and write the unified dataset into the time series database according to the established rules after obtaining the unified dataset.
[0023] For storage rules, standardized raw data (e.g., standardized battery cell raw data) is partitioned based on the unique cell identifier (CellID) and sorted by the timestamp corresponding to the data sampling time. This ensures that all time-series data corresponding to the same battery cell are arranged continuously in physical storage, thereby improving the retrieval efficiency of single-cell data and providing basic support for tracking battery cell data throughout its entire lifecycle. A corresponding partitioning strategy is configured for the storage rules. This strategy performs a hash operation based on the unique cell identifier, and according to the operation result, all data to be stored is evenly divided into a preset number of logical partitions, avoiding query bottlenecks caused by centralized data distribution. Within each logical partition, the data is further sharded according to a preset time period, allowing data from the same time period to be stored centrally, effectively shortening the traversal range for time interval queries and improving the speed of statistical analysis under large data volumes. For example, 256 logical partitions are created by dividing the CellID hash value modulo 256, and each partition is sharded by Timestamp (one shard per hour). Furthermore, based on the established storage rules and partitioning strategy, a corresponding index structure is established in the time-series database. This index structure includes a primary index and a secondary index. The primary index is formed by combining the unique identifier of the battery cell and the timestamp corresponding to the data sampling time, i.e., (CellID, Timestamp), used to quickly locate all time-series data of a single battery cell within a specified time range. The secondary index is formed by combining the batch number and the timestamp corresponding to the data sampling time, i.e., (BatchID, Timestamp), used to enable batch data queries by production batch and by time range, meeting the statistical analysis needs at the production line and batch levels. Simultaneously, the database supports a high-concurrency write capability of no less than 10,000 records per second and supports data retrieval by any time interval. Finally, according to the configured storage rules, partitioning strategy, and index structure, the unified dataset is written to the time-series database record by record, completing the persistent storage of the original data.
[0024] By coordinating the settings of storage rules, partitioning strategies, and index structures, time-series databases can support high-concurrency data writing and efficient time interval queries. This enables them to adapt to the massive, high-frequency, and highly time-series data storage scenarios of batching and filling processes, providing a stable and reliable data foundation for subsequent real-time computing, multidimensional analysis, and anomaly tracing.
[0025] To further improve the writing efficiency and system stability of high-frequency time-series data in the formation and capacity-building process, electronic devices can also build and enable batch writing mode to achieve efficient persistent storage of a unified dataset.
[0026] In an optional implementation, the method further includes: A batch write mode is constructed, which is configured to trigger writing based on a preset data count threshold or a preset time threshold, and to upload data using streaming transmission; The unified dataset is written to the time-series database according to the batch writing method.
[0027] To optimize write performance, a combination of batch writing and streaming is employed. In this embodiment, a batch writing method can be pre-built before writing the unified dataset to the time-series database. This batch writing method flexibly triggers write operations based on a preset data count threshold or a preset time threshold. When the number of standardized data records accumulated in the system cache reaches the preset data count threshold, or the time interval since the last write operation reaches the preset time threshold, the current batch write is automatically triggered, avoiding the system overhead caused by frequent writing of single data records. Simultaneously, the batch writing method uses a streaming mechanism for data upload, continuously transmitting data to be added to the database through long connections, reducing the number of network handshakes and connection rebuilds, improving data transmission throughput, and ensuring no data loss or delay in high-concurrency acquisition scenarios. For example, a batch write operation is performed every 1000 standardized raw data records accumulated, or every 100-millisecond trigger cycle; simultaneously, a gRPC streaming upload mechanism is used for data transmission, reducing network interaction overhead and further improving the stability and throughput of data writing in high-concurrency scenarios.
[0028] After configuring the batch write mode, the system, in conjunction with the pre-defined storage rules, partitioning strategies, and index structure, synchronously writes the unified dataset to the time-series database according to the batch write method. During the write process, data is stored in partitions based on the unique identifier of the battery cell, arranged in order by timestamp, and quickly located and retrieved through primary and secondary indexes. By introducing batch write and streaming transmission mechanisms, the database I / O pressure and network load can be significantly reduced, enabling the system to stably support the massive, high-frequency, and high-concurrency data write requirements of the batch capacity production line, ensuring efficient and reliable storage of raw data.
[0029] By employing time-series databases, hash partitioning, time sharding, and high-concurrency write mechanisms, this approach addresses the performance bottleneck of existing database technologies under tens of thousands of concurrent connections, thus meeting the need for stable high-frequency data writing.
[0030] S13, The unified dataset is written into the document database for statistical data storage.
[0031] In some embodiments, while writing the unified dataset to a time-series database for persistent storage, the cell-level capacity assessment statistics obtained in real-time calculation can also be aggregated by cell dimension and written to a document database for persistent storage. The statistical results use the unique identifier of the cell as the core index, aggregating and storing the batch identifier, product model, process parameters, statistical indicators of each process stage, capacity assessment level, and the time range corresponding to the data, forming a structured cell statistical archive.
[0032] To facilitate understanding of the inventive concept of this application, the following is an example of a structure for storing statistical results aggregated by cell dimension: { "CellID":"B20240315_A01_001_20240315", "BatchID":"B20240315", Model:"NMC-27100", "ProcessParams":{ "ChargeRate":0.5, "DischargeRate": 1.0, "CutoffVoltage":4.2, "RestTime":600 }, "StageStats":{ "Charge":{"CapacityAh":2.85,"AvgVoltage":3.68,"MaxTemp":38.2}, "Rest":{"TempDropRate":0.05}, "Discharge":{"CapacityAh":2.76,"EffPercent":96.8} }, "Grade":"A", "TimestampRange":[1704067200000,1704070800000] } It should be noted that the time-series database and document database mentioned above are only examples. Other databases can be replaced according to the actual situation. For example, the time-series database can be replaced with any database system that supports time-series storage, and the document database can be replaced with a database that supports unstructured data storage.
[0033] In an optional implementation, the method further includes: Establish a one-to-one association mapping relationship between the unique cell identifiers of the statistical data in the document database and the unique cell identifiers of the original data in the time series database; When the storage time of the raw data exceeds the preset number of days, the raw data is automatically downsampled to aggregate the high-frequency sampled data into low-frequency sampled data. The maximum triangle three-bucket method is used to preserve the trend of downsampled time series data, and the Gorilla lossless compression algorithm is used to compress and store the processed time series data.
[0034] In some embodiments, to achieve rapid backtracking and locating from statistical results to raw data, a one-to-one association mapping relationship is established between the cell unique identifier field of statistical data in the document database and the cell unique identifier field of raw data in the time series database. This allows for direct location and retrieval of all corresponding raw time series data based on the statistical data of any cell, facilitating anomaly analysis, process traceability, and quality verification.
[0035] While storing statistical data, the system can also compress and downsample the raw data stored in the time-series database. When the storage time of the raw data exceeds the preset number of days (e.g., 30 days), the downsampling task is automatically started: the raw high-frequency sampled data is aggregated according to a preset period, for example, data at 1-second intervals are aggregated into 1-minute intervals, high-frequency sampling points are aggregated into low-frequency sampling points, and the minimum, maximum, and average values of voltage, current, and temperature are retained, thus reducing the amount of data while retaining key electrical characteristics and trends.
[0036] During downsampling, the Largest Triangle Three Buckets (LTTB) method is used to process the original time-series data, preserving key trend points of data fluctuations and ensuring that the downsampled data curves are consistent with the original curves. Furthermore, the Gorilla lossless compression algorithm (FB-series time series compression) is used to compress and store the downsampled time-series data, significantly reducing storage space usage without losing data information; for example, each data point occupies an average of only 1.37 bytes.
[0037] By employing the aforementioned statistical data storage, association mapping, downsampling, and compression mechanisms, efficient management and rapid traceability of statistical results are achieved, while effectively controlling the storage costs of long-term, massive amounts of raw data, thus improving the overall storage efficiency and query performance of the system. Through hierarchical storage and dual-database association mapping, the shortcomings of existing technologies—such as the inability to achieve full lifecycle tracking of battery cells and the inability to construct a mapping relationship between process parameters and performance results—are addressed.
[0038] S14, perform incremental calculations on the unified dataset based on a sliding window, and perform segmented statistics according to the formation and capacity-building process stages to obtain statistical results.
[0039] Throughout the real-time calculation and segmented statistics process, electronic devices can employ a caching mechanism to accelerate the statistical process. This caching mechanism can utilize memory caching or distributed caching to store time-series data and intermediate statistical values within the sliding window, thereby improving data read / write and statistical update efficiency and ensuring real-time calculation performance under high concurrency scenarios.
[0040] In an optional implementation, the incremental calculation of the unified dataset based on a sliding window, and the segmented statistics performed according to the formation and capacity-building process stages, yield statistical results including: The unified dataset is processed in real time according to a configurable sliding time window, and the window calculation is triggered at a fixed period. The statistical indicators are updated using an incremental calculation method based on newly added and removed data within the window. The statistical indicators include the average value and the cell capacity. Based on the current value, voltage change rate and duration in the unified dataset, the formation and capacity testing stages are identified and divided into charging stage, resting stage and discharging stage. Calculate the corresponding statistical indicators for the charging stage, the resting stage, and the discharging stage respectively to obtain the segmented statistical indicators for the corresponding stages; The statistical results are obtained by summarizing each of the segmented statistical indicators.
[0041] In some embodiments, after hierarchical storage of raw and statistical data in a standardized dataset, the electronic device can use a configurable sliding time window (e.g., 30 seconds / 60 seconds) to process the incoming standardized time-series data in real time, triggering calculations at fixed intervals (e.g., every second). This ensures that data is processed immediately upon entering the system, guaranteeing the real-time nature and continuity of the statistical process. During the window sliding process, statistical indicators are dynamically updated using incremental calculation based on newly added data (i.e., new data) and old data (i.e., removed data) within the window. These statistical indicators include at least the average data value and the battery cell charging / discharging capacity. This eliminates the need for repeated scanning of all data within the window, effectively reducing computational resource consumption and improving statistical efficiency. Specifically, for the average increment calculation, let the number of data points in the statistical window at the previous time step be n_old, the mean be μ_old, and the newly added data point be x_new. Then, the updated mean is expressed as μ_new = (μ_old * n_old + x_new) / (n_old + 1). When the window slides out of the old value x_out, n_new = n_old, and μ_new = (μ_old * n_old - x_out + x_new) / n_old. For the capacity increment calculation (current integral), let the sampling interval be Δt (unit: seconds), and the current sequence be I[0..k]. Then, the cumulative capacity (unit: Ah) is expressed as Capacity = (Σ_{i=1}^{k}(I[i] + I[i-1]) / 2 * Δt) / 3600. During the increment update, the trapezoidal area is calculated and accumulated only for the newly arrived interval (I_{prev}, I_{new}, Δt).
[0042] Simultaneously, based on the current and voltage values, their trends, and durations already included in the unified dataset, the formation and capacity testing stage of the battery cell is identified, automatically dividing the entire process into a charging stage, a resting stage, and a discharging stage. The charging stage is identified as current > 0.05C and a continuously rising voltage (slope > 0 for 5 consecutive points). The resting stage is identified as an absolute current value < 0.02C and a voltage change rate < 0.001V / second, lasting for more than 30 seconds. The discharging stage is identified as a current < -0.05C and a continuously decreasing voltage. After completing the process stage division, stage statistical indicators matching each stage are calculated for the charging, resting, and discharging stages, obtaining segmented statistical indicators for the corresponding stages. Specifically, during the charging phase, statistical indicators such as charging capacity (integral), coulombic efficiency (discharge capacity / charging capacity), and median voltage plateau are calculated. During the discharging phase, statistical indicators such as discharge capacity, median voltage, and discharge energy (voltage × current integral) are calculated. During the resting phase, statistical indicators such as temperature change rate (linear fitting) and voltage rebound rate (rebound voltage / discharge cutoff voltage) are calculated. Finally, the segmented statistical indicators corresponding to each process stage are summarized and organized to obtain statistical results that fully reflect the entire process of cell formation and capacity testing, providing accurate and reliable data support for subsequent multidimensional analysis and anomaly detection.
[0043] In the process of incremental calculation and segmented statistics on a unified dataset based on a sliding window, the system adopts a caching mechanism to accelerate the statistical process in order to improve the data read and write efficiency and statistical update speed in high-concurrency scenarios.
[0044] In other embodiments, the electronic device may also use a full calculation method to directly traverse and statistically analyze all time-series data within the current sliding window to complete the calculation of average value, capacity, and indicators at each stage, thereby meeting the requirements for statistical accuracy and implementation method selection in different scenarios.
[0045] In an optional implementation, the use of a caching mechanism to accelerate the statistical process includes: An ordered set is used to cache the sliding window timing data corresponding to the unique identifier of each cell, and the data sampling timestamp is used as the sorting basis. A hash structure is used to cache the statistical intermediate values of the corresponding sliding window. The statistical intermediate values include the number of data points, the cumulative voltage value, the cumulative current value, the cumulative temperature value, and the cumulative voltage square value. Upon receiving each new time-series data, the ordered set is updated based on the new time-series data, old time-series data exceeding the window range is removed, and the statistical intermediate values in the hash structure are updated synchronously. When the number of cached incremental data reaches a preset number or a preset time period, the statistical results in the hash structure are written into the document database, and the cached data corresponding to the current sliding window is cleared.
[0046] In some embodiments, the system uses the unique identifier CellID as an index and employs an ordered set (e.g., SortedSet) to cache and store the sliding window time-series data corresponding to each cell. The data sampling timestamps corresponding to the time-series data are used as the sorting criteria, ensuring that the data within the window is arranged in chronological order for easy maintenance and retrieval. For example, a SortedSet is used to store the sliding window data for each CellID: key="window:{CellID}", score=timestamp, member="V,I,T". Simultaneously, the system uses a hash structure to cache the statistical intermediate values corresponding to the current sliding window. These statistical intermediate values include the number of data points within the window, the accumulated voltage value, the accumulated current value, the accumulated temperature value, and the accumulated squared voltage value. For example, a Hash is used to store the statistical intermediate values of the current window: key="stats:{CellID}", with fields including count, sumV, sumI, sumT, and sumV2 (used to calculate the standard deviation).
[0047] During data processing, each time the system receives a new time-series data entry, it updates the corresponding ordered set (SortedSet) and removes older time-series data that exceeds the current sliding window's range to maintain the validity and accuracy of the window data. Simultaneously with updating the ordered set, the system synchronously updates the corresponding statistical intermediate values in the hash structure based on the new time-series data, completing real-time correction of the accumulated values. When the number of incremental data entries in the cache reaches a preset threshold (e.g., 500 entries), or the cache duration reaches a preset time period (e.g., 5 seconds), the system batch-writes the statistically completed results from the hash structure to the document database for persistent storage. After writing, it clears the cached data corresponding to the current sliding window, thereby releasing cache resources and providing available space for subsequent data processing.
[0048] The aforementioned caching mechanism significantly reduces the pressure of frequent database read / write operations, improving the response speed and processing stability of real-time statistics for segmented data. By employing a sliding window + incremental calculation + caching mechanism, traditional offline statistics and full scan methods are abandoned, significantly reducing computational complexity and achieving second-level real-time statistics to meet the high-concurrency real-time computing requirements of tens of thousands of channels.
[0049] S15, Construct a multidimensional analysis model based on the statistical results.
[0050] In some embodiments, to support flexible analysis of fragmented data, a multidimensional data cube (DataCube) is constructed, and dimensions and measures are organized using a star schema to obtain a multidimensional analysis model. The structure of the multidimensional analysis model is as follows: Fact Table: Each record corresponds to a summary of statistics for one battery cell at one process stage (charging / resting / discharging). Fields include: CellID (foreign key); BatchID (foreign key); Model (foreign key); Stage (enumeration: Charge / Rest / Discharge); TempRange (foreign key, such as 0-25, 25-45); RateRange (foreign key, such as 0.2C, 0.5C, 1.0C); CapacityAh (metric, floating point); AvgVoltage (metric); MaxTemp (metric); EnergyWh (a metric, the integral of voltage and current); Efficiency (a metric, discharge phase only); Grade (dimension, enumerating A / B / C).
[0051] Dimension Tables: Batch dimensions: BatchID, StartTime, EndTime, LineID, Operator; Model dimensions: Model, NominalCapacity, Chemistry (e.g., NCM, LFP); TempRange dimensions: RangeID, LowBound, HighBound; RateRange dimensions: RateID, RateValue, Type (Charge / Discharge).
[0052] Hierarchy: Time dimension: second → minute → hour → batch → production line → factory; Product Dimensions: Cell → Model → Material System; Process dimension: Stage → Process step → Equipment type.
[0053] Through a multidimensional analysis model, raw time-series data in a unified dataset can be pre-aggregated into a multidimensional data cube, supporting online analytical processing (OLAP) operations such as roll-up, drill-down, slicing, and dicing. Compared to existing technologies that only support unidimensional statistics, this application constructs a multidimensional analysis model of batch × temperature × rate × process parameters, providing data support for multidimensional cross-analysis and process optimization. S16, Perform cross-analysis of process parameters and performance data according to the multidimensional analysis model.
[0054] After constructing the multidimensional data model for capacity testing, multidimensional cross-analysis is performed based on the multidimensional data model to realize the correlation between process parameters and cell performance, parameter optimization, and abnormal batch identification.
[0055] In this embodiment, a temperature-capacity distribution analysis is first performed to evaluate the impact of different temperature ranges on battery discharge capacity. Statistical records for the discharge stage are selected from the fact table of the multidimensional data model and grouped according to temperature range. The mean, standard deviation, P10 quantile, P90 quantile, and capacity histogram distribution (e.g., 2.6~2.8Ah, 2.8~3.0Ah, etc.) of the discharge capacity of each group are calculated. The yield rate of each group with a capacity greater than or equal to a preset threshold is also calculated, forming a temperature-capacity cross-tabulation table that intuitively reflects the influence of temperature on cell discharge capacity and yield. The temperature-capacity cross-tabulation table is shown in Table 1 below.
[0056] Table 1:
[0057] Simultaneously, rate-yield variation analysis can be performed to evaluate the impact of different charge / discharge rates on the finished product yield (proportion of Class A capacity rating). The fact table is linked to the charge / discharge rate dimension table, and cells are grouped by charge or discharge rate (RateRange). The percentage of Class A capacity rating cells within each group is calculated, along with the percentages of Class B and Class C cells. Based on this, a rate-yield variation curve is generated, providing data for optimizing charge / discharge process rate parameters. The rate-yield relationship is illustrated in Table 2 below.
[0058] Table 2:
[0059] Simultaneously, it can also perform process parameter combination optimization analysis to screen the optimal process parameter combination that maximizes capacity. For example, it can perform multi-dimensional grouping and aggregation based on three dimensions: charging cut-off voltage, discharge rate, and temperature range, calculate the average capacity and standard deviation of the cells under each group combination, sort them from high to low average capacity, and output the Top-k optimal parameter combinations to achieve multi-objective collaborative optimization of process parameters.
[0060] Simultaneously, time series cross-analysis can be performed to compare capacity trends across different batches and identify production fluctuations. For example, statistical data from a fixed temperature range and a fixed discharge rate (e.g., 25~45℃, 1.0C) can be selected, grouped by batch number (BatchID), and the average capacity of each batch can be calculated. The batches can then be sorted by production date and a time series line chart can be generated. Furthermore, the 3σ principle is used for anomaly detection, automatically marking batches that exceed control limits, facilitating rapid identification of equipment aging, material fluctuations, or process drift issues.
[0061] By employing multidimensional data models to conduct temperature-capacity, rate-yield, multi-parameter combination, and batch-time-series cross-analysis, a mapping relationship between process parameters and performance results can be established. This allows for the rapid discovery of correlations between process parameters and cell performance, enabling the selection of optimal process parameters and automatic identification of abnormal batches. This effectively improves the efficiency of formation and capacity testing process optimization and quality control capabilities, while reducing production fluctuations and defect rates. S17. Based on the statistical results and historical data trends, anomaly determination is made, and the abnormal data is traced back based on the unique identifier of the battery cell in the unified dataset.
[0062] After completing real-time statistics, segmented calculations, and multidimensional analysis of the batching and capacity testing data, the system performs anomaly detection based on the real-time statistical results and historical statistical data within a sliding window. Upon identifying anomalies or abnormal trends, it performs end-to-end data traceability using the unique identifier of the battery cell. The anomaly detection process combines static threshold rules with dynamic trend prediction, and the traceability process relies on a hierarchical storage structure to achieve rapid location of statistical results from raw data, ensuring that production anomalies can be detected and located promptly.
[0063] In an optional implementation, the anomaly determination based on the statistical results and historical data trends includes: The statistical results are subjected to real-time anomaly detection based on preset anomaly rules. When the preset anomaly rules are met, it is determined that an anomaly exists. Based on statistical data within a historical sliding window, the trend of capacity change is determined. When the continuous decline in capacity exceeds a preset threshold, it is identified as an abnormal trend in advance.
[0064] In some embodiments, the system pre-configures anomaly detection rules, which may include temperature exceeding a threshold, capacity deviation exceeding a set range, and sudden voltage or current changes. The system compares the temperature, capacity deviation, voltage, and current data from the statistical results with the corresponding thresholds set in the anomaly detection rules. When the temperature exceeds the set upper limit, the capacity deviation exceeds the set range, or the voltage or current changes abruptly, the system immediately determines that the current cell has an anomaly. Simultaneously, the system reads the capacity statistics data within the historical sliding window and analyzes the capacity change trend. If the capacity shows a continuous decline and the magnitude exceeds a preset trend threshold, it is determined to be an abnormal trend before the final anomaly occurs, thus achieving early warning of anomalies.
[0065] In an optional implementation, the step of reverse tracing of abnormal data based on the unique identifier of the battery cell in the unified dataset includes: When an anomaly or abnormal trend is detected, the original data corresponding to the time-series database is queried using the unique identifier of the battery cell as an index. The time interval data corresponding to the original time series data is obtained, and the equipment information and process parameters corresponding to the abnormal battery cell are associated to realize the full-link reverse tracing from the statistical results to the original data.
[0066] In some embodiments, when the system detects an anomaly or an abnormal trend, it uses the unique identifier CellID of the battery cell in the unified dataset as a retrieval index to automatically query all the original time-series data corresponding to that battery cell in the time-series database and locate the time interval data where the anomaly occurred. Simultaneously, the system associates the production equipment information and formation / capacity testing process parameters corresponding to that battery cell, integrating statistical results, original time-series data, and equipment and process information to achieve complete reverse tracing from statistical results to original data, facilitating rapid identification of the cause of the anomaly.
[0067] Through the above optional implementation methods, by combining real-time anomaly detection, capacity trend prediction, and reverse tracing using unique cell identifiers, real-time anomaly detection and trend prediction during the formation and capacity testing process can be achieved. This avoids the lag of alarms that only occur after an anomaly. At the same time, relying on the unique cell identifiers enables rapid and accurate data traceability, effectively improving the quality control efficiency of the production process, reducing the defect rate, and providing a reliable basis for process optimization and equipment maintenance.
[0068] Compared to existing technologies, this application achieves stable writing of high-frequency data through a time-series database and hierarchical storage structure, demonstrating excellent high-concurrency processing capabilities; it achieves second-level data statistics through sliding windows and incremental calculations, significantly improving real-time statistical efficiency; it establishes a unique cell identification association system through a unified data model, enabling full lifecycle tracking of cells and correlation analysis between process parameters and performance results; it supports multi-dimensional combined statistics and cross-analysis based on a multi-dimensional data model, significantly improving the depth of process analysis and optimization capabilities; and it achieves rapid location and root cause tracing of abnormal batteries through real-time anomaly detection and trend prediction mechanisms combined with unique cell identification, demonstrating highly efficient anomaly location capabilities. In a lithium battery production scenario, the system connects to 10,000 channels of equipment, collecting data once per second and processing 10,000 data entries per second. After data collection and unified modeling, the data is written into a time-series database. The real-time calculation module calculates capacity and average parameters based on a 60-second sliding window, and the statistical results are stored in a document database. The multi-dimensional analysis module performs analysis by batch and temperature range. When the capacity deviation of a batch exceeds 5%, anomaly detection is automatically triggered. The data traceability module quickly locates the problematic cells and corresponding process parameters. The entire system achieves second-level statistical response and anomaly location time of less than 1 second, effectively improving the control efficiency of the formation and capacity testing process and product quality.
[0069] Reference Figure 2 The diagram shown is a functional block diagram of the formulation and capacity data statistics device according to an embodiment of this application.
[0070] In some embodiments, the capacity testing data statistics device 20 may include multiple functional modules composed of computer program segments. The computer programs for each program segment of the capacity testing data statistics device 20 may be stored in the memory of an electronic device and executed by at least one processor to perform (see details). Figure 1 This describes the function of breaking down data into statistical components. Based on its function, it can be divided into multiple functional modules. These modules may include: a data processing module 201, a hierarchical storage module 202, an incremental statistics module 203, a model building module 204, a multidimensional cross-analysis module 205, and an anomaly detection module 206. The term "module" in this application refers to a series of computer program segments that can be executed by at least one processor and perform a fixed function, stored in memory. In this embodiment, the functions of each module will be detailed in subsequent embodiments.
[0071] The data processing module 201 is used to acquire multi-source chemical composition and capacity data collected by the chemical composition and capacity device, and to standardize the multi-source chemical composition and capacity data to obtain a unified dataset.
[0072] The hierarchical storage module 202 is used to write the unified dataset into a time-series database for raw data storage and into a document database for statistical data storage.
[0073] The incremental statistics module 203 is used to perform incremental calculations on the unified dataset based on a sliding window, perform segmented statistics according to the batching and capacity-building process stage, and obtain statistical results; wherein a caching mechanism is used to accelerate the statistical process.
[0074] The model building module 204 is used to build a multidimensional analysis model based on the statistical results.
[0075] The multidimensional cross-analysis module 205 is used to perform cross-analysis of process parameters and performance data according to the multidimensional analysis model.
[0076] The anomaly determination module 206 is used to determine anomalies based on the statistical results and historical data trends, and to reverse trace the abnormal data based on the unique identifier of the battery cell in the unified dataset.
[0077] It should be understood that the various variations and specific embodiments of the formation and capacity data statistics method provided in the above embodiments are also applicable to the formation and capacity data statistics device of this embodiment. Through the foregoing detailed description of the formation and capacity data statistics method, those skilled in the art can clearly understand the implementation method of the formation and capacity data statistics device of this embodiment. For the sake of brevity, it will not be described in detail here.
[0078] See Figure 3 The diagram shown is a schematic representation of the structure of an electronic device according to an embodiment of this application. In a preferred embodiment of this application, the electronic device 3 includes a memory 31, at least one processor 32, and at least one communication bus 33.
[0079] Those skilled in the art should understand that Figure 3 The structure of the electronic device shown does not constitute a limitation of the embodiments of this application. It can be a bus structure or a star structure. The electronic device 3 may also include more or fewer other hardware or software than shown, or different component arrangements.
[0080] In some embodiments, the electronic device 3 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), digital processors, and embedded devices. The electronic device 3 may also include user equipment, which includes, but is not limited to, any electronic product capable of human-computer interaction with a user via a keyboard, mouse, remote control, touchpad, or voice control device, such as a personal computer, tablet computer, smartphone, or digital camera.
[0081] In the embodiments provided in this application, it should be understood that the disclosed methods, apparatuses, computer-readable storage media, and electronic devices can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple components or modules may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices, components, or modules may be electrical, mechanical, or other forms.
[0082] The components described as separate parts may or may not be physically separate. The components shown as components may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the components can be selected to achieve the purpose of this embodiment according to actual needs.
[0083] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each component can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0084] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drive, portable hard drive, read-only memory (ROM). Various media that can store program code, such as only memory, random access memory (RAM), magnetic disks, or optical disks.
[0085] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0086] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0087] The above is a detailed description of the preferred embodiments of this application. However, the invention of this application is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A method for statistical analysis of fractionation capacity data, characterized in that, The method includes: Acquire multi-source chemical composition and capacity data collected by the chemical composition and capacity device, and perform standardization processing on the multi-source chemical composition and capacity data to obtain a unified dataset; The unified dataset is written into a time-series database for raw data storage and into a document database for statistical data storage. Incremental calculations are performed on the unified dataset based on a sliding window, and segmented statistics are executed according to the stage of the batching and capacity-building process to obtain statistical results; a caching mechanism is used to accelerate the statistical process. A multidimensional analysis model is constructed based on the statistical results; Perform cross-analysis of process parameters and performance data based on the multidimensional analysis model; Anomaly determination is made based on the statistical results and historical data trends, and reverse tracing of abnormal data is achieved based on the unique cell identifier in the unified dataset.
2. The method for statistical analysis of composition and capacity data according to claim 1, characterized in that, Writing the unified dataset into a time-series database for raw data storage includes: Configure storage rules for the time-series database. The storage rules are based on the unique identifier of the battery cell as the partitioning basis and the timestamp corresponding to the data sampling time as the sorting basis. The unified dataset includes the unique identifier of the battery cell, the timestamp corresponding to the data sampling time, and the batch number. Configure a partitioning strategy for the storage rules. The partitioning strategy performs hash partitioning based on the unique identifier of the battery cell, dividing the data into a preset number of logical partitions. Each logical partition is time-sharded according to a preset time period. According to the storage rules and the partitioning strategy, an index structure is established in the time-series database. The index structure includes a first-level index that combines the unique identifier of the battery cell and the timestamp corresponding to the data sampling time, and a second-level index that combines the batch number and the timestamp corresponding to the data sampling time. The unified dataset is written into the time-series database for persistent storage according to the storage rules, the partitioning strategy, and the index structure.
3. The method for statistical analysis of composition and capacity data according to claim 2, characterized in that, The method further includes: A batch write mode is constructed, wherein the batch write mode is configured to trigger writing based on a preset data count threshold or a preset time threshold, and data is uploaded using streaming transmission; The unified dataset is written to the time-series database according to the batch writing method.
4. The method for statistical analysis of composition and capacity data according to claim 1, characterized in that, The incremental calculation of the unified dataset based on the sliding window, and the segmented statistics performed according to the formation and capacity-building process stages, yield the following statistical results: The unified dataset is processed in real time according to a configurable sliding time window, and the window calculation is triggered at a fixed period. The statistical indicators are updated using an incremental calculation method based on newly added and removed data within the window. The statistical indicators include the average value and the cell capacity. Based on the current value, voltage change rate and duration in the unified dataset, the formation and capacity testing stages are identified and divided into charging stage, resting stage and discharging stage. Calculate the corresponding statistical indicators for the charging stage, the resting stage, and the discharging stage respectively to obtain the segmented statistical indicators for the corresponding stages; The statistical results are obtained by summarizing each of the segmented statistical indicators.
5. The method for statistical analysis of composition and capacity data according to claim 1, characterized in that, The use of a caching mechanism to accelerate the statistical process includes: An ordered set is used to cache the sliding window timing data corresponding to the unique identifier of each cell, and the data sampling timestamp is used as the sorting basis. A hash structure is used to cache the statistical intermediate values of the corresponding sliding window. The statistical intermediate values include the number of data points, the cumulative voltage value, the cumulative current value, the cumulative temperature value, and the cumulative voltage square value. Upon receiving each new time-series data, the ordered set is updated based on the new time-series data, old time-series data exceeding the window range is removed, and the statistical intermediate values in the hash structure are updated synchronously. When the number of cached incremental data reaches a preset number or a preset time period, the statistical results in the hash structure are written into the document database, and the cached data corresponding to the current sliding window is cleared.
6. The method for statistical analysis of composition and capacity data according to claim 1, characterized in that, The anomaly determination based on the statistical results and historical data trends includes: The statistical results are subjected to real-time anomaly detection based on preset anomaly rules. When the preset anomaly rules are met, it is determined that an anomaly exists. Based on statistical data within a historical sliding window, the trend of capacity change is determined. When the continuous decline in capacity exceeds a preset threshold, it is identified as an abnormal trend in advance.
7. The method for statistical analysis of composition and capacity data according to claim 1, characterized in that, The method of reverse tracing of abnormal data based on the unique identifier of the battery cell in the unified dataset includes: When an anomaly or abnormal trend is detected, the original data corresponding to the time-series database is queried using the unique identifier of the battery cell as an index. The time interval data corresponding to the original time series data is obtained, and the equipment information and process parameters corresponding to the abnormal battery cell are associated to realize the full-link reverse tracing from the statistical results to the original data.
8. A device for calculating the composition and capacity data, characterized in that, The device includes: The data processing module is used to acquire multi-source chemical composition and capacity data collected by the chemical composition and capacity device, and to standardize the multi-source chemical composition and capacity data to obtain a unified dataset. The hierarchical storage module is used to write the unified dataset into a time-series database for raw data storage and into a document database for statistical data storage. The incremental statistics module is used to perform incremental calculations on the unified dataset based on a sliding window, and to perform segmented statistics according to the batching and capacity-deploying process stage to obtain statistical results; a caching mechanism is used to accelerate the statistical process. The model building module is used to build a multidimensional analysis model based on the statistical results. The multidimensional cross-analysis module is used to perform cross-analysis of process parameters and performance data based on the multidimensional analysis model. The anomaly detection module is used to detect anomalies based on the statistical results and historical data trends, and to reverse trace the abnormal data based on the unique identifier of the battery cell in the unified dataset.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the batching capacity data statistical method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for calculating the composition and capacity data according to any one of claims 1 to 7.