Power battery generation type data set quality evaluation method

By constructing a multi-dimensional evaluation framework and a standardized feature space, the problem of uncontrollable data quality in generated power batteries was solved, and the credibility and engineering value of the generated data in evaluation and application were improved.

CN121880918APending Publication Date: 2026-04-17CHINA AUTOMOTIVE ENG RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA AUTOMOTIVE ENG RES INST
Filing Date
2025-11-18
Publication Date
2026-04-17

Smart Images

  • Figure CN121880918A_ABST
    Figure CN121880918A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power battery management, and discloses a power battery generative dataset quality evaluation method, which comprises the following steps: mapping a power battery generative dataset to be evaluated to a standardized feature space to obtain a target generative dataset; performing first data processing on the target generation type data set to obtain a multi-dimensional evaluation result, including: calculating a statistical distribution closeness degree of the target generation type data set and a real reference data set to obtain a authenticity evaluation result; calculating the distribution breadth of the target generation type data set in the feature space to obtain a diversity evaluation result; calculating the multi-variable synchronization relation consistency of the target generation type data set under the same or similar working condition labels to obtain a consistency evaluation result; calculating the coverage integrity of the evaluation target generation type data set on the key working condition area to obtain a representative evaluation result; and performing second data processing on the multi-dimensional evaluation result to obtain a comprehensive evaluation result. According to the method, the power battery generation type data quality can be measured in multiple dimensions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power battery management technology, and more specifically to a method for quality assessment of generative datasets for power batteries. Background Technology

[0002] With the rapid development of the new energy vehicle industry, the performance, safety, and lifespan of power batteries, as core components, have received widespread attention. To improve the modeling, diagnosis, and prediction capabilities of power batteries, relying on high-quality, large-scale datasets has become an industry consensus. In recent years, with the development of artificial intelligence, simulation technology, and generative models, generative datasets (including data generated based on physical models, simulation platforms, machine learning algorithms, or diffusion models) have been increasingly adopted in the research and application of power batteries. Generative data has advantages such as low cost, strong controllability, and ease of constructing extreme operating conditions, and is becoming an important supplement to actual collected data.

[0003] However, existing technologies still have significant problems and shortcomings in using generative data, mainly in the following aspects: 1. Lack of a systematic quality assessment mechanism for generated data: Current power battery data quality assessments mainly focus on data collected from actual vehicle operation, paying attention to indicators such as data completeness, noise rate, sampling frequency, and consistency. For generated data, existing assessment methods cannot effectively reflect its authenticity, representativeness, diversity, and degree of matching with real-world scenarios, making it difficult to use generated data as a reliable basis for algorithm training or model validation.

[0004] 2. Incomplete or limited evaluation dimensions fail to accurately reflect problems in the generated data: The generated data may have the following issues: data distribution deviates from actual working conditions, feature redundancy or missing features, insufficient coverage of working conditions, and distortion of boundary data. If only traditional data statistical indicators are used for evaluation, these deep-seated structural or semantic defects often cannot be identified, leading to distorted results in the quality assessment of the generated data.

[0005] 3. Existing generative techniques are highly black-box and their quality is uncontrollable: Especially when using deep learning or generative models (such as GANs and diffusion models) for data augmentation or generation, the output data lacks interpretability. Although the generated samples are similar to real data in form, there may be situations where they are "seemingly reasonable but actually distorted." Due to the lack of an effective evaluation mechanism, such problems are difficult to detect in a timely manner.

[0006] 4. Lack of relevance between assessments and actual application tasks: Different tasks have different requirements for data quality. For example, models used for battery life prediction are more concerned with the consistency of long-term data trends, while models used for safety early warning are more concerned with the completeness of coverage of abnormal operating conditions. Currently, there is no universal mechanism to adaptively adjust the assessment focus, resulting in a disconnect between the generated data assessment results and specific application scenarios.

[0007] 5. Lack of a standard system leads to a lack of unified reference for data quality assessment: Currently, there are no national standards or industry specifications for quality assessment of power battery generated data. Different institutions and enterprises use custom rules for assessment, resulting in poor comparability of results and difficulty in supporting the sharing and reuse of generated data across platforms and models.

[0008] The main reasons for the above problems include: the semantic complexity and structural diversity of generative data itself make it difficult for traditional indicators to measure its essential quality; the evaluation methods do not combine the generation process and the characteristics of the target task, resulting in poor generalization ability; existing evaluation tools are mostly based on static rules and lack flexibility and intelligence; the industry lacks attention to the quality control of generated data and there is insufficient accumulation of relevant theories and practices.

[0009] Therefore, there is an urgent need to propose a quality assessment method for power battery generated data. Summary of the Invention

[0010] The present invention aims to provide a method for quality assessment of generative datasets for power batteries, which can comprehensively measure data quality across multiple dimensions and enhance its reliability in model training, algorithm verification, system testing and other stages.

[0011] The basic solution provided by this invention is: a method for quality assessment of generative datasets for power batteries, including: S1, map the generative dataset of the power battery to be evaluated to a standardized feature space to obtain the target generative dataset; the features include at least one of time series features, statistical features, working condition label embedding and deep features; S2 performs the first data processing on the target generative dataset to obtain multi-dimensional evaluation results, including: The statistical distribution of the target generative dataset is calculated to determine how closely it approximates that of the real reference dataset, thus obtaining the authenticity assessment results. Calculate the distribution breadth of the target generative dataset in the feature space to obtain the diversity assessment results; Calculate the consistency of multivariate synchronization relationships in the target generative dataset under the same or similar working condition labels, and obtain the consistency evaluation results; The completeness of the target generative dataset in covering key operating conditions is calculated to obtain representative evaluation results. S3 involves performing a second data processing step on the multi-dimensional evaluation results to obtain a comprehensive evaluation result.

[0012] The working principle and advantages of this invention are as follows: Addressing the technical bottleneck of uncontrollable quality in the current field of new energy power batteries despite the large-scale application of generative data, the advantages of this invention are: 1. A systematic evaluation framework for generated data from power batteries: Existing technologies mainly focus on quality control of measured data or use general generative model evaluation indicators (such as FID and IS) from the image / text domain. These technologies cannot effectively measure the quality of generated data such as power battery data, which is subject to strong physical constraints, multivariate coupling, and time sensitivity. This invention is the first to establish an evaluation system covering multiple dimensions such as authenticity, diversity, consistency, and representativeness for non-real-world data types such as "simulation generation, model synthesis, and augmented generation". It not only focuses on statistical similarity but also introduces exclusive dimensions such as compliance with physical laws and coverage of key operating conditions, solving the technical problem of determining the 'usability' of generated data and providing a basis for determining the reliability and usability of generated data.

[0013] 2. Scientifically designed multidimensional quality indicators comprehensively reflect potential data defects: In order to adapt to the characteristics of strong physical constraints and multivariate coupling of power battery data, this invention first maps the generated data to a standardized feature space containing time series, statistics and operating condition semantics to achieve a unified evaluation benchmark across models and operating conditions; this invention starts from the statistical characteristics, structural coverage, semantic stability and operating condition integrity of the generated data to construct four independent but complementary quality indicator systems, which can accurately reveal common problems such as redundancy, skewness, anomalies and sparsity in the data.

[0014] 3. Closely integrated with practical application tasks, enhancing the practicality of evaluation results: This invention introduces a multivariate synchronous relationship consistency assessment. By analyzing the residuals of the functional relationships between key parameters such as SOC, voltage, and temperature, it verifies whether the generated data conforms to electrochemical kinetics, effectively identifying invalid generated samples that are "mathematically reasonable but physically distorted," thus improving the interpretability of the evaluation. This is significantly superior to traditional methods that rely solely on purely statistical indicators such as KL divergence and Wasserstein distance. Simultaneously, the representative indicators designed in this invention are linked to actual battery operating conditions (such as low-temperature start-up, high-temperature fast charging, and long-cycle aging), deeply integrating the evaluation results with practical tasks such as BMS state estimation and fault warning. This ultimately forms a quantifiable and feedback-enabled comprehensive evaluation result, which can not only be used for quality judgment but also guide model design, sample generation construction, and deployment strategy adjustments, ensuring that the evaluation work serves real business needs.

[0015] 4. Introducing structure mapping and latent space computation to support complex data types: This invention supports the processing of complex data formats such as high-dimensional time series, nonlinear features, and multi-channel data by mapping multidimensional battery data to feature space or latent space (e.g., using dimensionality reduction encoders, feature compressors, etc.). Simultaneously, 5. Supports comprehensive scoring and visualization reports, with engineering implementation capabilities: This invention designs a configurable weighted comprehensive scoring mechanism and can output reports including fractal scoring, anomaly alerts, sample heatmaps, etc., and supports integration with existing enterprise data platforms and model evaluation systems.

[0016] 6. Possesses closed-loop capability of generation-evaluation-feedback: This invention supports running as an external quality gating module of the generation model, and can also be optionally linked with the generation model to drive the optimization of the generation strategy (such as resampling, penalty constraints, and condition completion) through quality feedback, thereby realizing a closed-loop adaptive optimization mechanism.

[0017] In summary, this invention not only solves the problems of lack of specificity, single evaluation dimensions, and poor engineering feasibility in the prior art, but also has many advantages such as comprehensive structure, advanced method, strong versatility, and high application value, and has broad prospects for promotion and industrialization potential. Attached Figure Description

[0018] Figure 1 A flowchart illustrating the power battery generative dataset quality assessment method provided in this embodiment of the invention. Figure 1 ; Figure 2 A flowchart illustrating the power battery generative dataset quality assessment method provided in this embodiment of the invention. Figure 2 . Detailed Implementation

[0019] The following detailed explanation illustrates the specific implementation methods: The basic implementation examples are as follows: Figure 1 and Figure 2 As shown: A method for quality assessment of generative datasets for power batteries, including: Before S1, there is also S0, which preprocesses the generative dataset of the power batteries to be evaluated: S01, Data Standardization and Normalization Z-score standardization is applied to numerical features to eliminate the influence of dimensions. Specifically, each feature column is calculated. mean and standard deviation Convert using the following formula:

[0020] in, This represents the original value of the i-th sample on the j-th feature; This represents the mean of all sample values ​​in the j-th feature column; It represents the standard deviation of all sample values ​​in the j-th feature column.

[0021] For bounded features such as voltage and current, Min-Max normalization is used to compress them to the [0,1] interval:

[0022] in, Represents the normalized eigenvalues; This represents the minimum value in the j-th feature column;

[0023] This represents the maximum value in the j-th feature column.

[0024] S02, Time Series Completion and Calibration Missing time points are filled in using linear interpolation:

[0025] in, Indicates the time point of missing data The state quantity at the location.

[0026] For data with time offsets, Dynamic Time Warping (DTW) is used to align with the real reference data, eliminating the phase difference between the generated and measured data. Dynamic time alignment is performed as follows:

[0027] in, This represents the i-th point in the generated data sequence; This represents the j-th point in the real reference data sequence; This is the optimal alignment path.

[0028] S03, Noise Filtering and Boundary Correction Smoothing is performed using a Savitzky-Golay filter (polynomial order p=3, window length L=11):

[0029] in, This represents the smoothed value at time t. This represents the value of the original time series at time t. This represents the filtering coefficients, which, after processing, effectively preserve the signal's trend characteristics while suppressing high-frequency noise.

[0030] Boundary value correction: Truncating values ​​that exceed the physical range.

[0031] S04, Align with the actual reference data (if it exists). S1 involves mapping the generative dataset of the power battery to be evaluated to a standardized feature space to obtain the target generative dataset. Features include at least one of time-series features, statistical features, operating condition label embeddings, and deep features. This step provides a unified representation basis for subsequent index calculations.

[0032] In S1, features are extracted from each data sample and uniformly mapped to a standardized feature space to construct a unified feature representation system to support subsequent multidimensional evaluation. 1) Extract physically meaningful time series features from the raw time series data, including fluctuation amplitude, slope, and (average) rate of change: Fluctuation range The range of parameter variation can be reflected by the following formula:

[0033] in, This represents the maximum value of the time series data. This represents the minimum value of the time series data.

[0034] slope The trend of change can be represented by least squares fitting:

[0035] in, Represents the time vector t and time series data covariance, It represents the variance of the time vector t.

[0036] Average rate of change The intensity of fluctuations can be calculated using the following formula:

[0037] Where T represents the total length of the time series data (number of time points). Let represent the value of the i-th sample at time t.

[0038] Constructing time series features as follows:

[0039] 2) Statistical characteristics, data calculation The statistical characteristics of each dimension of the data, including the median, mean, variance, and standard deviation. The structure is as follows:

[0040] 3) Embedding of working condition labels, such as semantic vector construction, can be achieved using one-hot encoding to generate working condition semantic vectors. :

[0041] 4) Deep feature generation: A variational autoencoder (VAE) is used to learn the essential structure of the data. The encoder outputs a latent vector. Compressing high-dimensional data into a low-dimensional semantic space:

[0042] in, This represents an encoder for variational autoencoders, with the following parameters: , This represents the input high-dimensional data sample. This represents the low-dimensional latent vector of the output.

[0043] 5) Feature fusion, combining time series features Statistical characteristics Working condition semantic vector and latent vectors To form a comprehensive feature matrix As a basis for evaluation:

[0044] S2 performs the first data processing on the target generative dataset to obtain multi-dimensional evaluation results, including: S21, calculate the similarity of the statistical distribution between the target generative dataset and the preset real reference dataset to obtain the authenticity assessment result.

[0045] Authenticity primarily assesses whether the generative dataset faithfully reflects the true distribution. Therefore, the design aims to verify, from a statistical distribution perspective, whether the generated data has "learned" the essential characteristics of the real data.

[0046] Indicators characterizing the similarity of statistical distributions include: distribution fit measures, such as KL divergence and Wasserstein distance, which measure the overall distribution similarity (global); differences in the mean and variance of multidimensional features, which check whether the first and second moments are aligned (local); higher-order difference indicators such as Mahalanobis distance and Bhattacharyya coefficient, which consider the feature covariance structure and are sensitive to correlation (higher-order); and nearest neighbor coverage, which assesses whether the generated samples "fall" into the real data-dense area (practical orientation).

[0047] By combining and evaluating multiple statistical measures in a hierarchical manner, the problem of "misjudging fidelity" in battery time-series data using a single indicator (such as FID) is solved.

[0048] The statistical similarity between generated data and real data is quantified, and the authenticity assessment results include an authenticity score.

[0049] 1) Distribution distance calculation: Wasserstein distance is used:

[0050] in, These represent the distributions of the generated dataset and the real dataset, respectively. It represents the infimum, which minimizes the overall cost of the solution.

[0051] The advantage of the Jensen-Shannon divergence in measuring overall distributional variance lies in its ability to reflect the geometric structure of the distribution. Calculate the Jensen-Shannon divergence:

[0052] in, Let these represent the probability distributions of the generated dataset and the real dataset, respectively. Evaluate probability density similarity.

[0053] 2) Local similarity verification: For each generated sample Calculate its average 5-nearest neighbor distance in the real dataset:

[0054] Statistical satisfaction The proportion of samples is used as the nearest neighbor coverage rate (threshold). (Take the median of the actual sample interval).

[0055] 3) Authenticity score After normalizing the above indicators, a weighted fusion is performed:

[0056] in, JSD represents the Wasserstein distance; JSD represents the Jensen-Shannon divergence. This represents the Mahalanobis distance considering covariance. Indicates the coverage rate of neighboring areas; This represents the weights, where the sum of all weights is 1. In this embodiment, , , The values ​​are 0.3, 0.3, 0.2, and 0.2, respectively.

[0057] S22, calculate the distribution breadth of the target generative dataset in the feature space to obtain the diversity assessment results.

[0058] Diversity primarily assesses the sample differences within a generative dataset, measuring the "diversity" and "deduplication" of the data to prevent the generative model from "pattern collapse," i.e. generating only a few types of samples; therefore, the design evaluates diversity from the perspective of spatial distribution.

[0059] Indicators characterizing the breadth of distribution include: mean Euclidean distance and its standard deviation, which measure the overall dispersion among samples; latent spatial coverage, such as the size of the t-SNE mapping region, which visualizes the latent spatial coverage; and local density, which is used to identify duplicate sample clusters and quantify the deduplication capability.

[0060] Introducing latent space analysis (such as t-SNE) to assess diversity breaks through the limitations of traditional methods that only consider sample size and is more suitable for high-dimensional nonlinear battery data.

[0061] The dataset is evaluated for its breadth of coverage and redundancy. The diversity assessment results include a diversity score.

[0062] 1) Spatial distribution analysis: Calculate the characteristic space variance entropy through principal component analysis (PCA). :

[0063] in, For the first Principal component variance contribution rate; a higher entropy value indicates a more uniform feature distribution.

[0064] Statistical mean Euclidean distance between samples and its standard deviation :

[0065] The ratio of the two It reflects the degree of dispersion of the distribution.

[0066] 2) Redundant sample detection: The density clustering algorithm DBSCAN (parameters) is used. Identify dense regions and define the proportion of duplicate samples. ,in This represents the number of samples that exceed the limit for clustering.

[0067] 3) Diversity score :

[0068] in, This represents the normalized local neighborhood entropy. This represents the normalized diversity density. The standard deviation representing the diversity density Indicates the diversity density value. The standard deviation representing the diversity density This indicates the proportion of redundant samples; high scores require data that is both widely distributed and avoids local clustering. This represents the weights, where the sum of all weights is 1. In this embodiment, , , The values ​​are 0.3, 0.3, 0.2, and 0.2, respectively.

[0069] S23, calculate the consistency of the multivariate synchronization relationship of the target generative dataset under the same or similar working condition labels, and obtain the consistency evaluation result.

[0070] Consistency primarily assesses whether the generated data adheres to the physical laws of batteries, that is, the consistency and stability of the generated data under the same or similar operating conditions, reflecting its ability to maintain physical laws.

[0071] Indicators characterizing the consistency of multivariate synchronization relationships include: cluster compactness among similar samples, i.e., checking whether the output is stable under similar operating conditions; similarity of generated outputs under small perturbations, such as the consistency of curve shape, i.e., checking whether small perturbations lead to drastic changes (robustness); trend analysis of operating condition continuity (avoiding jumps and oscillations), i.e. checking whether the time continuity is reasonable (no jumps or oscillations); and residual fitting of the functional relationship between synchronization features (such as SOC-voltage-temperature), i.e. checking the correctness of the multivariate synchronization relationship, such as SOC↑→voltage↑, temperature↑→internal resistance↓.

[0072] For power battery application scenarios, we focus on the constraints of their operating physical laws, take physical consistency as an independent evaluation dimension, and propose new indicators such as the fitting residual of the synchronous characteristic function relationship, which reflects the core difference in power battery evaluation.

[0073] Verify the extent to which the generated data conforms to physical laws; the consistency assessment results include a consistency score.

[0074] 1) In-condition stability: For a subset of samples under the same operating conditions Calculate the profile coefficient :

[0075] in, This represents a set of samples belonging to the same working condition. This represents the average distance between sample i and other samples in the same cluster (intra-cluster dissimilarity). The silhouette coefficient value represents the minimum average distance between sample i and all samples in other clusters (inter-cluster dissimilarity). A value close to 1 indicates that similar samples are tightly clustered in the feature space.

[0076] 2) Perturbation robustness test: Gaussian noise is added to the samples. Generate a perturbed version and calculate the dynamic time-warped distance ratio between the original and perturbed data. :

[0077] in, Represents the original data. This represents perturbation data, reflecting the model's stability to small perturbations.

[0078] 3) Verification of electrochemical laws: Calculate the residuals of the Nernst equation regarding the voltage-temperature coupling relationship.

[0079] in, This represents the voltage test value of the i-th sample. This represents the standard parameters in the electrochemical equation; the mean residual reflects the degree of conformity to physical laws. This item is used to check whether the voltage-temperature relationship conforms to physical laws.

[0080] 4) Consistency score Equal-weighted combination of three indicators:

[0081] in, This represents the contour coefficient score. This represents the distance or difference based on dynamic time warping. The value represents the fitting residual of the average function relationship; P represents the calculated coefficient, and in this embodiment, P is taken as 3.

[0082] S24, calculate the coverage integrity of the target generative dataset for the key operating condition area to obtain representative evaluation results.

[0083] The main evaluation criteria are whether the generated data can cover real-world usage scenarios and the engineering usability boundaries of the generated data.

[0084] Indicators characterizing coverage integrity include: sample density differences in different regions of the real operating space; sample proportion under key operating conditions (such as high-temperature fast charging and low-temperature degradation); coverage of abnormal operating conditions simulation, which can assess the ability of the generation model to generate abnormal scenarios; and integrity of time evolution trend.

[0085] By linking representativeness with the actual task requirements of BMS, the evaluation results can directly serve practical applications such as fault warning model training, control strategy verification under extreme conditions, and battery life prediction, thus making the relationship between evaluation and application close.

[0086] Assess coverage integrity in key operating areas; representative assessment results include representative scores.

[0087] 1) Key Scenario Detection: Define industrial scenarios of interest such as high-temperature fast charging and low-temperature discharge, and calculate coverage:

[0088] in, , These represent the number of samples for this scene in the generated set and the real set, respectively.

[0089] Take the lowest coverage rate As a weak point indicator.

[0090] 2) Lifetime evolution analysis: The battery life curve is divided into three stages: initial activation, mid-term stability, and final degradation. The sample proportion of each stage is statistically analyzed to see if it exceeds a threshold. Calculate the percentage of complete segments in SegCover.

[0091] 3) Representative Scoring Focus on key scenario coverage:

[0092] in, The key performance indicator (KPI) represents the representativeness score, while SegCover represents the segment coverage score. and Represents the weighting coefficients, which sum to 1 and > In this embodiment, and They are 0.6 and 0.4 respectively.

[0093] In S3, the second data processing step involves weighted summation of the scores from the multi-dimensional evaluation results to obtain the quality score of the generative dataset. :

[0094] in, The evaluation results are respectively based on authenticity, diversity, consistency, and representativeness. The weight parameters satisfy .

[0095] Specifically, the total score is calculated as follows:

[0096] This invention establishes a four-dimensional quality evaluation system. In the quality assessment system for power battery generation data, the allocation of weights for each dimension is based on profound domain expertise and rigorous engineering safety considerations. In this embodiment, the weight for authenticity... The highest weight of 0.4 is assigned because the power battery system relies absolutely on the accuracy of its electrochemical mechanisms—parameters such as voltage, temperature, and SOC must strictly follow physical laws such as the Nernst equation. Any tiny deviation can trigger a chain of safety failures, directly affecting the reliability of critical functions such as thermal runaway warning and overcharge protection. Consistency weight. Ranking second with a weight of 0.3 reflects the importance of the stability of multi-parameter coupling relationships. State estimation and lifetime prediction in the BMS algorithm are extremely sensitive to the coordinated changes in parameters such as voltage, current, and temperature. Data inconsistency will directly cause the SOC estimation error to exceed the 5% threshold required for automotive-grade standards. (Diversity weight) Setting it to 0.2 reflects a balance strategy prioritizing quality over quantity. While ensuring the accuracy of safety-critical data, it moderately pursues full-scenario coverage, avoiding diluting the density of critical samples by excessively generating data from ordinary operating conditions. Representativeness weight. Setting it to 0.1 serves as a safety redundancy guarantee, specifically ensuring the complete coverage of mandatory scenarios such as nail penetration testing and thermal runaway. Although these scenarios represent a small percentage of the total data, their safety value is extremely high. This weighting system has been validated through multiple rounds of experiments, demonstrating optimal balance characteristics in different systems such as ternary lithium batteries and lithium iron phosphate batteries. This improves the accuracy of the generated data in BMS algorithm training by 23%, while reducing the false negative rate in safety boundary scenarios to below 1%, fully reflecting a professional design philosophy based on the unique needs of power batteries.

[0097] Quality grading is shown in Table 1: Table 1 Quality Grading Table

[0098] In practical use: perform steps S1-S3 to assess the quality of the collected historical dataset.

[0099] Alternatively, the execution system and generative model corresponding to this method can be integrated and deployed. After the generative model generates the power battery generative dataset, it can be directly input into the execution system corresponding to this method to obtain a quality score. Then, the quality score is fed back into the generative model to adjust its generation strategy (such as diversity enhancement, anomaly penalty term optimization, etc.) to achieve a closed-loop iteration of generation-evaluation-optimization.

[0100] Specifically, iterative optimization is initiated when integrating the generative model: 1) Improved loss function: Add a quality penalty term to the GAN generator loss:

[0101] 2) Reinforcement learning strategy: Transforming quality scores into reward signals:

[0102] Guide the generative model to explore high-quality data regions.

[0103] 3) Sampling weight update: for low coverage conditions ,according to Increase the sampling probability.

[0104] This embodiment provides a method for quality assessment of generative datasets for power batteries, constructing a four-dimensional evaluation system encompassing authenticity, diversity, consistency, and representativeness to address the challenge of determining the "usability" of generated data. To adapt to the physical constraints and task requirements of battery data, a standardized feature space mapping is used to unify the representation of multi-source data; a novel multi-variable synchronization relationship consistency assessment is introduced to verify whether the generated data conforms to electrochemical laws; and representativeness is measured by the completeness of coverage of key operating conditions to ensure the data is suitable for tasks such as SOH and RUL. The evaluation results are quantifiable and feedback-enabled, significantly improving the credibility and engineering value of the generated data in battery simulation and AI training.

[0105] Example 2 Unlike Example 1, different generative models (such as GAN, VAE, and diffusion models, which affect the fidelity and diversity of generated data), different types of power batteries (such as ternary, lithium iron phosphate, and solid-state batteries, which determine the strength of physical constraints and risk sensitivity), and different battery application tasks (such as SOH estimation, RUL prediction, fault warning, and BMS control strategy training, which determine the focus) have different sensitivities to dimensions such as "authenticity" and "consistency." Therefore, the weights of the comprehensive score are dynamically adjusted accordingly to improve the accuracy of the evaluation.

[0106] The adaptive weight adjustment process in the comprehensive evaluation results includes: S1: Obtain three types of information, including the generation model type (GAN / VAE / diffusion), the target application scenario (SOH / RUL / fault warning), and the power battery type (NMC / LFP / sodium ion / solid). S2, based on the three types of information obtained, match the initial weights in the pre-built three-dimensional prior weight library; the weight prior knowledge base sets the initial weights according to different models and tasks, and can be adjusted according to actual use; S3, build a sensitivity feedback model to adaptively correct the initial weights. By introducing a sensitivity feedback model, if a certain dimension performs poorly, its influence on the total score is increased, causing the total score to drop significantly, highlighting its severity, and serving as a "weakness warning".

[0107] The sensitivity feedback model is constructed as follows:

[0108] in, These are the adjusted weights, used for the final weighted summation; These are the initial weights; Score the current dimension (i.e., authenticity score, consistency score, etc.). This is the threshold (if the score is below this value, such as below 0.7, it is considered a "serious defect"). The adjustment factor is 0.3 (recommended).

[0109] If this value is 0, the weight is not adjusted; the larger the difference, the stronger the penalty.

[0110] The power battery generative dataset quality assessment method provided in this embodiment proposes an adaptive assessment weight mechanism based on the generative model type, power battery type, and battery application task, breaking through the limitations of traditional static weighting. By constructing a prior weight template library and combining it with low-scoring item amplification feedback, the comprehensive score becomes more scenario-specific and has engineering guidance significance, effectively improving the intelligence and practicality of the assessment results.

[0111] The above descriptions are merely embodiments of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent.

Claims

1. A method for quality assessment of generative datasets for power batteries, characterized in that, include: S1, map the generative dataset of the power battery to be evaluated to the normalized feature space to obtain the target generative dataset; Features include at least one of time series features, statistical features, working condition label embeddings, and deep features; S2 performs the first data processing on the target generative dataset to obtain multi-dimensional evaluation results, including: The statistical distribution of the target generative dataset is calculated to determine how closely it approximates that of the real reference dataset, thus obtaining the authenticity assessment results. Calculate the distribution breadth of the target generative dataset in the feature space to obtain the diversity assessment results; Calculate the consistency of multivariate synchronization relationships in the target generative dataset under the same or similar working condition labels, and obtain the consistency evaluation results; The completeness of the target generative dataset in covering key operating conditions is calculated to obtain representative evaluation results. S3 involves performing a second data processing step on the multi-dimensional evaluation results to obtain a comprehensive evaluation result.

2. The method for quality assessment of power battery generative datasets according to claim 1, characterized in that, Indicators characterizing the similarity of statistical distributions include distribution fit measures, differences between the mean and variance of multidimensional features, higher-order difference indicators, and nearest neighbor coverage.

3. The method for quality assessment of power battery generative datasets according to claim 1, characterized in that, The authenticity assessment results include an authenticity score. : in, JSD represents the Wasserstein distance; JSD represents the Jensen-Shannon divergence. This represents the Mahalanobis distance considering covariance. Indicates the coverage rate of neighboring areas; This represents the weights, and the sum of all weights is 1.

4. The method for quality assessment of power battery generative datasets according to claim 1, characterized in that, Indicators characterizing the breadth of distribution include the mean Euclidean distance and its standard deviation, potential spatial coverage, and local density.

5. The method for quality assessment of power battery generative datasets according to claim 1, characterized in that, Diversity assessment results include diversity scores : in, Represents the local neighborhood entropy. This represents the normalized diversity density. The standard deviation representing the diversity density, Indicates the local density correlation coefficient; This represents the weights, and the sum of all weights is 1.

6. The method for quality assessment of power battery generative datasets according to claim 1, characterized in that, Indicators characterizing the consistency of multivariate synchronization relationships include cluster compactness among similar samples, similarity of generated outputs under small perturbations, trend of continuous operating conditions, and residual fit of functional relationships between synchronization features.

7. The method for quality assessment of power battery generative datasets according to claim 1, characterized in that, The conformity assessment results include a conformity score. : in, This represents the contour coefficient score. This represents the distance or difference based on dynamic time warping. represents the fitting residual of the average function relationship; P represents the calculated coefficient.

8. The method for quality assessment of power battery generative datasets according to claim 1, characterized in that, Indicators characterizing coverage integrity include the difference in sample density in different regions of the real working condition space, the proportion of samples under key working conditions, the coverage of abnormal working conditions simulation, and the integrity of the time evolution trend.

9. The method for quality assessment of power battery generative datasets according to claim 1, characterized in that, Representativeness assessment results include representativeness scores. : in, The key performance indicator (KPI) represents the representativeness score, while SegCover represents the segment coverage score. and Represents the weighting coefficients, which sum to 1 and > .

10. The method for quality assessment of power battery generative datasets according to claim 1, characterized in that, A target generated data quality assessment model is constructed to execute the power battery generated dataset quality assessment method according to any one of claims 1-9; the target generated data quality assessment model is integrated with the generation model to generate a power battery generated dataset through the generation model, which is directly input into the target generated data quality assessment model to obtain a comprehensive assessment result, and the comprehensive assessment result is fed back to the generation model to adjust its generation strategy.