Method, device, equipment and medium for extracting aging characteristic factors of power battery
By filtering and generating the initial sample set from the vehicle detailed data set of electric vehicles and filtering out the aging feature factors using the feature screening model, the problem of the feature factors losing their physical meaning in the existing technology is solved, and better model interpretability and feature reusability are achieved.
Patent Information
- Application Number
- CN202210809710.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-07-11
AI Technical Summary
When the prior art extracts battery aging characteristic factors from battery performance data, the physical meaning of the characteristics is lost, resulting in poor interpretability of subsequent model training, making it difficult to analyze the specific factors that affect battery aging.
By obtaining the vehicle detailed data set of electric vehicles, N groups of vehicle detailed data grouped by vehicle models are selected, the initial sample set is generated, and the aging characteristic factors affecting battery health are screened from the initial samples through the feature screening model. This feature screening model has the ability to record the use of features during training, and retains the physical meaning of sample features.
The physical meaning of retaining features in the training of battery health prediction model is realized, the explanatory nature of the model is improved, the specific reasons for the impact on battery aging can be better analyzed, and the extracted aging feature factor has high reusability and versatility.
Smart Images

Figure CN115184810B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of battery detection, and in particular, to a method, device, equipment, and medium for extracting aging characteristic factors of power batteries. Background Art
[0002] For electric vehicles, predicting the battery health using vehicle operation data and machine learning algorithms has become a mainstream trend because this method greatly reduces the cost and time of offline detection. However, the training of a battery health prediction model relies on a large number of samples containing effective features. How to quickly and effectively screen out effective features from data in numerous dimensions is crucial for the training of the model.
[0003] In the field of battery data feature engineering, the autoencoder algorithm is used to extract characteristic factors of battery aging from battery performance data. An autoencoder is a self-supervised model for pre-training. The features are first mapped to data with a lower dimension than the original dimension and then restored to be equal to the original dimension through the network. It is trained with the aim of minimizing the deviation between the original input features and the restored features. After training is completed, the intermediate low-dimensional data is taken as the extracted features for subsequent training. However, the characteristic factors extracted after autoencoder training lose their original physical meanings, and their interpretability is very poor during subsequent model training, making it difficult to analyze the specific factors affecting battery aging. Summary of the Invention
[0004] In view of the above technical problems in the prior art, the embodiments of this specification provide a method, device, equipment, and medium for extracting aging characteristic factors of power batteries.
[0005] In a first aspect, the embodiments of this specification provide a method for extracting aging characteristic factors of power batteries, including: obtaining a vehicle detail data set of an electric vehicle; screening out N groups of vehicle detail data grouped by vehicle models from the vehicle detail data set, where N is an integer greater than 1;
[0006] generating an initial sample set according to the N groups of vehicle detail data, where the initial sample set includes N groups of initial samples grouped by vehicle models; screening out aging characteristic factors affecting battery health from the N groups of initial samples through a feature screening model, and the feature screening model has the ability to record the feature usage situation during the training process.
[0007] Optionally, the method of obtaining a vehicle detail data set of an electric vehicle comprises: receiving vehicle operation data from a plurality of electric vehicles, wherein the vehicle operation data is collected and uploaded by electric vehicles of M types through their signal acquisition devices, each type of vehicle comprises a plurality of electric vehicles, and M is an integer greater than or equal to N; performing data cleaning on the vehicle operation data to obtain the vehicle detail data set.
[0008] Optionally, filtering out N groups of vehicle detail data grouped according to vehicle model from the vehicle detail data set includes: performing data processing on the vehicle detail data set to obtain M groups of vehicle detail data corresponding one-to-one to M types of vehicle models, wherein each group of vehicle detail data includes vehicle detail data of multiple electric vehicles of the same model; and filtering out the N groups of vehicle detail data from the M groups of vehicle detail data according to a preset vehicle number threshold.
[0009] Optionally, the data processing of the vehicle detail data set to obtain M groups of vehicle detail data corresponding one-to-one to M types of vehicle models includes: grouping the vehicle detail data set according to vehicle model to obtain M data groups corresponding one-to-one to the M types of vehicle models, each data group including vehicle detail data of each electric vehicle of the same model; and filtering the M data groups within the groups according to preset data filtering rules to filter out M groups of vehicle detail data corresponding one-to-one to the M types of vehicle models.
[0010] Optionally, the M data groups are respectively screened within the group according to preset data screening rules, including: taking each data group in the M data groups as a target data group; for the vehicle detail data of each electric vehicle in the target data group, if the mileage span of the electric vehicle is less than a preset mileage threshold, or the time span of the electric vehicle for collecting data is less than a preset duration threshold, then the vehicle detail data corresponding to the electric vehicle is eliminated; based on the vehicle detail data of the remaining electric vehicles in the M data groups, the M groups of vehicle detail data are obtained.
[0011] Optionally, generating an initial sample set based on the N groups of vehicle detail data includes: taking the vehicle detail data of each electric vehicle in the N groups of vehicle detail data as target vehicle detail data; extracting target data from the target vehicle detail data; binning the target data according to preset mileage intervals to obtain K boxes of data belonging to the same electric vehicle, where K is an integer greater than 1; and extracting samples from the K boxes of data to generate K initial samples corresponding to the same electric vehicle.
[0012] Optionally, the sample extraction of the K-box data to generate K initial samples corresponding to the same electric vehicle includes: sequentially calculating the cumulative feature factor and the real-time feature factor of each box of data in the K-box data; generating the K initial samples according to the cumulative feature factor and the real-time feature factor of the K-box data.
[0013] Optionally, the screening of the aging feature factors affecting the battery health degree from the N groups of initial samples by the feature screening model includes: respectively performing intra-group feature screening on the N groups of initial samples to obtain N groups of sample features corresponding to N vehicle models one by one; training the feature screening model by using the N groups of sample features; calculating the feature importance of each sample feature in the N groups of sample features by the trained feature screening model; and performing secondary feature screening according to the feature importance of each sample feature in the N groups of sample features to obtain the aging feature factors.
[0014] Optionally, the respectively performing intra-group feature screening on the N groups of initial samples to obtain N groups of sample features corresponding to N vehicle models one by one includes: respectively taking each group of initial samples in the N groups of initial samples as the target group of initial samples; normalizing the feature values of each sample feature in the target group of initial samples to obtain the normalized feature values of each sample feature in the target group of initial samples; for each sample feature in the target group of initial samples, calculating the feature variance of the sample feature according to the normalized feature value of the sample feature; and performing intra-group feature screening according to the feature variance of each sample feature in the target group of initial samples, and screening out each sample feature with a feature variance greater than a preset variance threshold as a group of sample features.
[0015] Optionally, the performing secondary feature screening according to the feature importance of each sample feature in the N groups of sample features to obtain the aging feature factors includes: respectively taking each group of sample features in the N groups of sample features as the target group of sample features, and performing intra-group feature ranking according to the feature importance of each sample feature in the target group of sample features; performing intra-group feature screening according to the intra-group feature ranking results of the N groups of sample features to screen out N groups of important sample features corresponding to the N vehicle models one by one; and performing inter-group feature comparison on the N groups of important sample features, and screening out the common sample features of the N vehicle models and each non-common sample feature with a feature importance greater than a preset threshold as the aging feature factors.
[0016] Optionally, the in-group feature ranking according to the feature importance of each sample feature in the target group sample features includes: ranking from largest to smallest according to the feature importance of each sample feature in the target group sample features; the in-group feature screening according to the in-group feature ranking results of the N groups of sample features to screen out N groups of important sample features corresponding to the N vehicle models one by one, including: for each group of important sample features in the N groups of important sample features, cumulatively add the feature importance of this group of sample features from largest to smallest until the cumulative result reaches the first preset weight threshold, then retain each sample feature participating in the cumulative calculation in this group of sample features, and remove the sample features with feature importance less than the second weight threshold from the sample features participating in the cumulative calculation, so as to obtain the important sample features in this group of sample features.
[0017] Optionally, it further includes: training a prediction model for predicting the health of a power battery using the aging feature factor, and / or analyzing the influence direction of the aging feature factor on the health of the power battery.
[0018] In a second aspect, an aging feature factor extraction device for a power battery provided by an embodiment of this specification includes: a data acquisition unit for acquiring a vehicle detail data set of an electric vehicle; a data screening unit for screening out N groups of vehicle detail data grouped by vehicle model from the vehicle detail data set, where N is an integer greater than 1; a sample extraction unit for generating an initial sample set according to the N groups of vehicle detail data, the initial sample set including N groups of initial samples grouped by vehicle model; a feature screening unit for screening out the aging feature factors affecting the battery health from the N groups of initial samples through a feature screening model, and the feature screening model has the ability to record the feature usage situation during the training process.
[0019] In a third aspect, an electronic device provided by an embodiment of this specification includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method according to any one of the embodiments in the first aspect.
[0020] In a fourth aspect, a computer-readable storage medium provided by an embodiment of this specification stores a computer program, and when the program is executed by a processor, it implements the method according to any one of the embodiments in the first aspect.
[0021] One or more technical solutions provided by the embodiments of this specification have at least the following technical effects or advantages:
[0022] In the embodiments of this specification, N groups of vehicle detail data grouped by vehicle models are screened out from the vehicle detail dataset of electric vehicles, and N groups of initial samples grouped by vehicle models are generated according to the N groups of vehicle detail data; aging characteristic factors affecting battery health are screened out from the N groups of initial samples through a feature screening model. Since the feature screening model used has the ability to record the feature usage situation during the training process, the screened sample features retain their original physical meanings, and have better interpretability when applied to the subsequent training of the battery health prediction model, which is also conducive to analyzing the specific reasons affecting battery aging.
[0023] Moreover, the synchronous training of the feature screening model based on multiple vehicle models ensures the generality of the features, so that the extracted aging characteristic factors have good reusability, and can avoid or reduce repeated training for extracting aging characteristic factors. Brief Description of the Drawings
[0024] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments, and are not considered to be a limitation of this specification. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0025] Figure 1 is a schematic diagram of the system architecture of the embodiments of this specification;
[0026] Figure 2 is a flowchart of the method for extracting the aging characteristics of the power battery in the embodiments of this specification;
[0027] Figure 3 is a schematic diagram of the process for generating an initial sample set according to the vehicle detail data in the embodiments of this specification;
[0028] Figure 4 is a schematic diagram of extracting multiple initial samples according to the vehicle detail data of an electric vehicle in the embodiments of this specification;
[0029] Figure 5 is a schematic diagram of the structure of the device for extracting the aging characteristics of the power battery in the embodiments of this specification;
[0030] Figure 6 is a schematic diagram of the structure of the electronic device in the embodiments of this specification. Detailed Embodiments
[0031] To better understand the above technical solutions, the technical solutions of the embodiments of this specification will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments of this specification and the embodiments are detailed descriptions of the technical solutions of the embodiments of this specification, rather than limitations on the technical solutions of this specification. Without conflict, the technical features in the embodiments of this specification and the embodiments can be combined with each other.
[0032] In a first aspect, an embodiment of this specification provides a method for extracting the aging characteristics of a power battery, which is applied to a system architecture as shown in Figure 1 and includes: a cloud server and multiple electric vehicles of multiple vehicle models that establish a data connection with the cloud server. Each electric vehicle collects and uploads its own vehicle operation data to the cloud server. The cloud server receives the vehicle operation data of each electric vehicle, obtains a vehicle detail data set of the electric vehicle according to the vehicle operation data uploaded by each electric vehicle, screens out N groups of vehicle detail data grouped by vehicle model from the vehicle detail data set, generates an initial sample set according to the N groups of vehicle detail data, and the initial sample set includes N groups of initial samples grouped by vehicle model; and filters out the aging characteristic factors that affect the battery health from the N groups of initial samples through a feature screening model.
[0033] Referring to Figure 2 as shown, the method for extracting the aging characteristics of the power battery provided by the embodiment of the present invention includes the following steps:
[0034] S101. Obtain a vehicle detail data set of the electric vehicle.
[0035] The cloud server receives the vehicle operation data from multiple electric vehicles. The vehicle operation data can be collected and uploaded by multiple electric vehicles of M vehicle models through their respective signal acquisition devices. Each vehicle model includes multiple electric vehicles. The cloud server performs data cleaning on the vehicle operation data uploaded by each electric vehicle received, and obtains a vehicle detail data set.
[0036] In a specific implementation process, the signal acquisition device of each electric vehicle collects and uploads the vehicle operation data to the cloud server in real time. The cloud server performs data cleaning on the vehicle operation data uploaded by each electric vehicle. Among them, data cleaning includes filtering abnormal data and repairing missing data to obtain the vehicle detail data of the electric vehicle and store it in the database. The cloud server continuously collects the vehicle detail data of each electric vehicle of various vehicle models through the above implementation method, so as to obtain a vehicle detail data set.
[0037] It should be understood that the signal acquisition device of the electric vehicle can be an in-vehicle T-BOX (Telematics Box, the intelligent network connection system of the vehicle). The signal acquisition device collects the CAN (Controller Area Network) bus signal data of the electric vehicle and analyzes it to obtain the vehicle operation data. The vehicle operation data collected by each electric vehicle includes dynamic signal fields and static signal fields. The dynamic fields include: the total current, total voltage, SOC (State of Charge), insulation resistance, the highest temperature of the battery pack, the lowest temperature of the battery pack, the highest voltage of the battery cell, the lowest voltage of the battery cell, the battery pack probe temperature array, and the battery pack single cell voltage array of the power battery on the electric vehicle, and also include: the driving speed, driving mileage, gear position, vehicle status, and charging status of the electric vehicle. The static signal fields include the rated capacity of the battery and the battery type.
[0038] The vehicle detail data can be high-frequency automotive data signals uploaded in accordance with GB / T 32960 "Technical Specifications for Electric Vehicle Remote Service and Management System", with a general frequency of once every 30s.
[0039] S102. Screen out N groups of vehicle detail data grouped by vehicle model from the vehicle detail data set, where N is an integer greater than 1.
[0040] It can be understood that the vehicle detail data set is processed to obtain M groups of vehicle detail data corresponding one-to-one to M vehicle models. Among them, each group of vehicle detail data includes the vehicle detail data of multiple electric vehicles from the same vehicle model; N groups of vehicle detail data are screened out from the M groups of vehicle detail data according to a preset vehicle number threshold.
[0041] In some embodiments, processing the vehicle detail data set includes: a data grouping step and a data screening step, and the order of the data grouping step and the data screening step can be swapped.
[0042] It should be understood that if the data grouping step is performed first and then the data screening step, it can be: first group the vehicle detail data set according to the vehicle model dimension to obtain M data groups corresponding one-to-one to M vehicle models, and each data group includes the vehicle detail data of each electric vehicle from the same vehicle model; perform in-group data screening on the M data groups respectively according to a preset data screening rule, so that each data group corresponds to a group of vehicle detail data obtained by screening, in order to obtain M groups of vehicle detail data corresponding one-to-one to M vehicle models.
[0043] Regarding performing in-group data screening on each data group according to a preset data screening rule, it can be performed according to the vehicle dimension, that is: according to a preset data screening rule, sequentially determine whether the vehicle detail data of each electric vehicle in the data group is to be excluded or retained.
[0044] Specifically, the preset data screening rules can be: if the driving mileage span of an electric vehicle is less than the preset mileage threshold, the vehicle detail data of this electric vehicle shall be excluded; if the time span of the collected data of an electric vehicle is less than the preset duration threshold, the vehicle detail data of this electric vehicle shall be excluded; if the driving mileage span of an electric vehicle is greater than or equal to the preset mileage threshold and the time span of the collected data is greater than or equal to the preset duration threshold, the vehicle detail data of this electric vehicle shall be retained.
[0045] Based on the above data screening rules, each of the M data groups is used as the target data group respectively; for the vehicle detail data of each electric vehicle within the target data group, if the driving mileage span of this electric vehicle is less than the preset mileage threshold, or the time span of the collected data of this electric vehicle is less than the preset duration threshold, the vehicle detail data corresponding to this electric vehicle shall be excluded; based on the vehicle detail data of the remaining electric vehicles in the M data groups, M groups of vehicle detail data are obtained.
[0046] For example, the preset mileage threshold is 20,000 kilometers, and the preset duration threshold is 365 days. The vehicle detail data of each electric vehicle with a driving mileage span ΔM≥20,000 kilometers and a data collection time span Δt≥365 days are screened out from the M data groups respectively, and M groups of vehicle detail data can be obtained.
[0047] Different from the above implementation, the vehicle detail data set can be screened first, and the vehicle detail data of each electric vehicle that does not meet the above data screening rules in the vehicle detail data set shall be excluded, and then the vehicle detail data of the remaining electric vehicles are grouped according to the vehicle models, and M groups of vehicle detail data can also be obtained.
[0048] It can be understood that screening out N groups of vehicle detail data from the M groups of vehicle detail data according to the preset vehicle number threshold can be: for each group of vehicle detail data in the M groups of vehicle detail data, if the number of electric vehicles from which this group of vehicle detail data comes is X and X is less than the preset vehicle number threshold, then this group of vehicle detail data shall be excluded, otherwise, this group of vehicle detail data shall be retained, so as to screen out the vehicle detail data of each vehicle model with a larger number of vehicles, and the vehicle detail data of the vehicle models with a smaller number of vehicles do not participate in the subsequent calculations.
[0049] For example, the preset vehicle number threshold is 100. After screening, it is necessary to ensure that the number of vehicles within each vehicle model is not less than 100, and the vehicle detail data corresponding to the vehicle models with less than 100 vehicles do not participate in the subsequent calculations.
[0050] S103. Generate an initial sample set based on N groups of vehicle detail data. The initial sample set includes N groups of initial samples grouped by vehicle models.
[0051] Reference Figure 3 As shown, to increase the sample size and at the same time ensure that the mileage span of each initial sample is the same, the steps of generating the initial sample set based on N groups of vehicle detail data include the following multiple refinement steps:
[0052] S1031: Respectively take the vehicle detail data of each electric vehicle in the N groups of vehicle detail data as the target vehicle detail data. Each group of vehicle detail data includes the vehicle detail data from multiple electric vehicles.
[0053] S1032: Extract target data from the target vehicle detail data.
[0054] It can be understood that the target data may include driving data and charging data. Since the vehicle speed and current direction of an electric vehicle are different under charging conditions and driving conditions, the vehicle detail data is sliced according to time points based on the vehicle speed and current direction, and the driving data and charging data are extracted from the target vehicle detail data.
[0055] S1033: Bin the target data at a preset mileage interval to obtain K bins of data belonging to the same electric vehicle, where K is an integer greater than 1.
[0056] It should be understood that the preset mileage interval needs to be a fraction of the preset mileage threshold. For example, if the preset mileage threshold is 10,000 kilometers, the preset mileage interval can be set to 2,500 kilometers, so that the vehicle detail data of each electric vehicle can be divided into multiple bins of data.
[0057] It should be understood that the number of bins that can be divided from the target data extracted from the same electric vehicle through binning is related to the driving mileage span of the electric vehicle, that is, the larger the driving mileage span of the electric vehicle, the larger the value of K and the more bins are divided.
[0058] S1034: Extract samples from the K bins of data to generate K initial samples corresponding to the same electric vehicle.
[0059] It should be understood that each initial sample includes multiple cumulative characteristic factors and multiple real-time characteristic factors. For the K bins of data of the same electric vehicle, calculate the cumulative characteristic factors and real-time characteristic factors of each bin of data in the K bins of data in turn; according to the cumulative characteristic factors and real-time characteristic factors of each bin of data in the K bins of data, generate K initial samples correspondingly, that is: one bin of data generates one initial sample.
[0060] Reference Figure 4As shown, the cumulative characteristic factors of the i-th initial sample are calculated from all the charging data and driving data in the data from the 1st bin to the i-th bin. The real-time characteristic factors of the i-th initial sample are only calculated from the charging data and driving data in the i-th bin, where i takes values from 1 to K in sequence.
[0061] It should be understood that it is also possible to first perform binning processing on the vehicle detail data of each electric vehicle according to a preset mileage interval, then extract the driving data and charging data from each bin of detail data, and generate a corresponding initial sample based on the driving data and charging data extracted from the detail data of this bin. In this way, multiple initial samples can also be generated from the vehicle detail data of the same electric vehicle.
[0062] Each initial sample includes multiple cumulative characteristic factors, specifically including: cumulative driving mileage, vehicle calendar days, cumulative parking charging duration, cumulative parking charging power, interval distribution of cumulative charging power, cumulative energy recovery power, cumulative discharging duration, cumulative discharging power, interval distribution of cumulative discharging power, interval distribution of cumulative current rate, interval distribution of cumulative SOC usage, cumulative average battery pack temperature distribution, cumulative battery pack temperature difference distribution, cumulative battery pack pressure difference distribution, interval distribution of cumulative insulation resistance value, cumulative battery shelving times, cumulative battery shelving duration, etc. It should be noted that the above various cumulative characteristic factors of each initial sample are all sample characteristics.
[0063] Each initial sample includes multiple real-time characteristic factors, specifically including: average battery temperature distribution, battery temperature difference distribution, battery pressure difference distribution, recent charging rate distribution, actual battery capacity (the battery capacity refers to the maximum capacity that the battery can charge during the charging process, with the unit of ampere-hour (Ah)), etc. It should be noted that the actual battery capacity in the real-time characteristic factors is the sample target value, and the other characteristic factors are all sample characteristics. Among them, the sample target value of the i-th initial sample of each electric vehicle is calculated using the charging data when the SOC is in the range of 40% - 60% in the i-th bin data of this electric vehicle, where i takes values from 1 to K in sequence, and the calculation method can adopt the current integration method.
[0064] Through the above steps S1031 - S1034, it is realized that multiple initial samples can be generated using the vehicle detail data of each electric vehicle. Thus, a large number of initial samples can be generated using the M groups of vehicle detail data, that is, an initial sample set is formed, and the initial sample set includes N groups of initial samples grouped by vehicle models.
[0065] S104. Screen out the aging characteristic factors that affect battery health from the N groups of initial samples through a feature screening model, and the feature screening model has the ability to record the usage of sample characteristics during the training process.
[0066] In the embodiments of this specification, in terms of having the ability to record the usage of sample features during the training process, the feature screening model can adopt the Xgboost (eXtreme Gradient Boosting) regression model. The Xgboost regression model constructs a stronger classifier with higher accuracy through multiple simple weak classifiers. Of course, other similar tree-structured models can also be used to replace the Xgboost regression model.
[0067] In order to screen out the aging feature factors that affect the battery health from N groups of initial samples, intra-group feature screening can be first performed at the vehicle model dimension, and then secondary feature screening can be combined with each vehicle model to ensure that the screened features have high generality. In the specific implementation process, the implementation process of performing the above two feature screenings includes the following steps:
[0068] S1041. Perform intra-group feature screening on each of the N groups of initial samples respectively to obtain N groups of sample features corresponding one-to-one to N vehicle models.
[0069] Respectively take each group of initial samples in the N groups of initial samples as the target group of initial samples; perform normalization processing on the feature values of each sample feature in the target group of initial samples to obtain the normalized feature values of each sample feature in the target group of initial samples; for each sample feature in the target group of initial samples, calculate the feature variance of this sample feature according to the normalized feature value of this sample feature; perform intra-group feature screening according to the feature variances of each sample feature in the target group of initial samples, and screen out each sample feature whose feature variance is greater than the preset variance threshold as a group of sample features. Through the above process, N groups of sample features corresponding one-to-one to N vehicle models can be screened out according to the N groups of initial samples.
[0070] It should be understood that each initial sample in the N groups of initial samples includes multiple sample features, that is: various cumulative feature factors and various real-time feature factors except the actual battery capacity. It is necessary to perform normalization processing on each sample feature of each initial sample respectively. Among them, the normalization processing of each sample feature follows the following formula:
[0071]
[0072] f is the feature value before normalization of this sample feature, f* is the feature value after normalization processing of this sample feature, and Fmin and Fmax are respectively the minimum value and the maximum value of this sample feature in this group of initial samples.
[0073] Taking N = 5 as an example, there are a total of 5 groups of initial samples. Taking the first group of initial samples as an example, assuming that the first group of initial samples has a total of 500 initial samples, for the sample feature "cumulative driving mileage", the maximum value of this sample feature in these 500 initial samples is 20,000 kilometers, and the minimum value is 12,000 kilometers. The feature value of this sample feature in the first initial sample is 15,000 kilometers. Normalize the feature value of this sample feature in the first initial sample: (15000 - 12000) / (20000 - 12000), and the normalized feature value is 0.375. Through the above example method, the normalized feature values of each sample feature in each initial sample can be calculated in turn.
[0074] Specifically, for each sample feature in the target group of initial samples, according to the normalized feature values of this sample feature in each initial sample in the target group of initial samples, calculate the feature variance of this sample feature within the target group of initial samples. The variance calculation formula for a single sample feature is as follows:
[0075]
[0076] f j * is the normalized feature value of the j-th initial sample of the sample feature in the target group of initial samples, μ is the average value of the sample feature in the target group of initial samples, and W is the number of samples in the target group of initial samples.
[0077] For N groups of initial samples, respectively screen out each sample feature with a feature variance greater than the preset variance threshold from each group of initial samples. For example, eliminate each sample feature with a feature variance less than 0.01 in each group of initial samples, so as to screen out N groups of sample features corresponding to N vehicle models one by one.
[0078] S1042. Use N groups of sample features to train the feature screening model.
[0079] For N groups of sample features corresponding to N vehicle models one by one, respectively divide each group of sample features into a training set and a test set according to a preset ratio. For example, divide each group of sample features into a training set and a test set according to an 8:2 ratio; input the training set samples of each group of sample features in the N groups of sample features into the feature screening model for training respectively. After the training is completed, input the test set samples of each group of sample features in the N groups of sample features into the trained feature screening model respectively to test the final effect of the trained feature screening model.
[0080] Taking the Xgboost regression model as an example, the sample features of each group are input into the Xgboost regression model for fitting and parameter tuning to complete the training of the Xgboost regression model, and the same technical parameters are used during the fitting and parameter tuning process. For example, the same technical parameters can be used during the fitting and parameter tuning process as follows:
[0081] 1. learning_rate (learning rate): 0.03;
[0082] 2. n_estimators (maximum number of subtrees): 4000;
[0083] 3. max_depth (maximum depth of the subtree): 13;
[0084] 4. min_child_weight (sum of sample weights of the minimum leaf nodes): 2;
[0085] 5. reg_alpha (weight of the L1 regularization term): 0.1
[0086] 6. reg_lambda (weight of the L2 regularization term): 0.2
[0087] S1043: Calculate the feature importance of each sample feature in the N groups of sample features through the trained feature screening model.
[0088] The formula for calculating the feature importance of each sample feature is:
[0089]
[0090] f weight is the feature importance of the sample feature, f times is the number of times the sample feature is used when all subtrees of the Xgboost regression model are split, and F split_time is the total number of splits of all subtrees of the Xgboost model.
[0091] S1044: Perform secondary feature screening based on the feature importance of each sample feature in the N groups of sample features to obtain the aging feature factors.
[0092] Respectively take each group of sample features in the N groups of sample features as the target group of sample features, and perform within-group feature sorting according to the feature importance of each sample feature in the target group of sample features; perform within-group feature screening according to the within-group feature sorting results of the N groups of sample features to screen out N groups of important sample features corresponding to N vehicle models one by one.
[0093] Specifically, the sample features in the target group are sorted from large to small according to their feature importance. For each group of important sample features in the N groups of important sample features, the feature importance of the group of sample features is accumulated from large to small, until the accumulation result reaches the first preset weight threshold, and the sample features in the group of sample features that participate in the accumulation calculation are retained, and the sample features whose feature importance is less than the second weight threshold are removed from the sample features that participate in the accumulation calculation, so as to obtain the important sample features in the group of sample features.
[0094] For ease of understanding, the following takes the first preset weight threshold as 0.95 and the second preset weight threshold as 0.05 as an example. A group of important sample features includes {feature 1 (feature importance 0.3), feature 2 (feature importance 0.29), feature 3 (feature importance 0.25), feature 4 (feature importance 0.1), feature 5 (feature importance 0.04), feature 6 (feature importance 0.03), feature 7 (feature importance 0.02)}, and an example of the secondary screening process is given: Since the cumulative weight value obtained by adding features 1 to 5 in sequence is: 0.3+0.29+0.25+0.1+0.04=0.98, which satisfies 0.98>0.95, features 6 and 7 need to be eliminated, and features 1 to 5 that participate in the cumulative calculation are retained. Then, since the feature importance of feature 5 is 0.03<0.05, feature 5 needs to be eliminated from features 1 to 5, thereby obtaining a group of important sample features as features 1 to 4.
[0095] After obtaining N groups of important sample features in the above manner, inter-group feature comparison is performed on the N groups of important sample features to screen out the common sample features of the N types of vehicle models and each non-common sample feature whose feature importance is greater than a preset threshold. The common sample features and each non-common sample feature whose feature importance is greater than a preset threshold are used as aging feature factors.
[0096] It should be noted that a shared sample feature refers to a sample feature that exists in each of the N groups of important sample features. If the sample feature does not exist in at least one group of important sample features, it is a non-shared sample feature. For example, if the preset threshold is 0.1, if the feature importance of any non-shared sample feature in the N groups of important sample features is less than 0.1, the non-shared sample feature is removed.
[0097] Through the above technical solution, the horizontal feature comparison of multiple vehicle models is realized. On the basis of ensuring the effectiveness of a single feature, the extracted feature dimensions have high generality among different vehicle models, and it can be ensured that the selected features have high generality. When the data sample changes, there is no need to retrain, and the reusability is good. Through training and fitting with the Xgboost regression model, using the proportion of the number of times each feature is used when the leaf nodes of a large number of weak learners in the Xgboost model are split as the screening criterion, it ensures the interpretability of battery parameters while the battery aging factor has a high contribution to the battery capacity, which provides great convenience for studying the impact of the battery aging factor on the battery capacity by using machine learning methods.
[0098] Further, after the aging feature factors are extracted, the aging feature factors can be used to train a prediction model for predicting the health of power batteries, and / or analyze the influence direction of the aging feature factors on the health of power batteries, so as to analyze the specific reasons for the impact on battery aging, and guide drivers to develop good vehicle usage habits.
[0099] Based on the same inventive concept, an embodiment of the present invention provides an apparatus for extracting aging feature factors of a power battery. Refer to Figure 5 As shown, it includes: a data acquisition unit 501, configured to acquire a vehicle detail data set of an electric vehicle; a data screening unit 502, configured to screen out N groups of vehicle detail data grouped by vehicle model from the vehicle detail data set, where N is an integer greater than 1; a sample extraction unit 503, configured to generate an initial sample set according to the N groups of vehicle detail data, and the initial sample set includes N groups of initial samples grouped by vehicle model; a feature screening unit 504, configured to screen out aging feature factors affecting the battery health from the N groups of initial samples through a feature screening model, and the feature screening model has the ability to record the feature usage situation during the training process.
[0100] In some embodiments, the data acquisition unit 501 is specifically configured to: receive vehicle operation data from multiple electric vehicles, where the vehicle operation data is collected and uploaded by electric vehicles of M vehicle models through their signal acquisition devices, each vehicle model includes multiple electric vehicles, and M is an integer greater than or equal to N; perform data cleaning on the vehicle operation data to obtain the vehicle detail data set.
[0101] In some embodiments, the data screening unit 502 is specifically configured to: perform data processing on the vehicle detail data set to obtain M groups of vehicle detail data corresponding to the M vehicle models one by one, where each group of vehicle detail data includes vehicle detail data of multiple electric vehicles of the same vehicle model; screen out the N groups of vehicle detail data from the M groups of vehicle detail data according to a preset vehicle number threshold.
[0102] In some embodiments, the data screening unit 502 is specifically used to: group the vehicle detail data set according to vehicle model to obtain M data groups corresponding one-to-one to the M types of vehicle models, each data group including vehicle detail data of each electric vehicle of the same model; and perform intra-group data screening on the M data groups according to preset data screening rules to screen out M groups of vehicle detail data corresponding one-to-one to the M types of vehicle models.
[0103] In some embodiments, the data screening unit 502 is specifically used to: respectively take each data group in the M data groups as a target data group; for the vehicle detailed data of each electric vehicle in the target data group, if the mileage span of the electric vehicle is less than a preset mileage threshold, or the time span of the electric vehicle collecting data is less than a preset duration threshold, then the vehicle detailed data corresponding to the electric vehicle is eliminated; based on the vehicle detailed data of the remaining electric vehicles in the M data groups, the M groups of vehicle detailed data are obtained.
[0104] In some embodiments, the sample extraction unit 503 is specifically used to: respectively use the vehicle detailed data of each electric vehicle in the N groups of vehicle detailed data as the target vehicle detailed data; extract the target data from the target vehicle detailed data; bin the target data according to a preset mileage interval to obtain K boxes of data belonging to the same electric vehicle, where K is an integer greater than 1; perform sample extraction on the K boxes of data to generate K initial samples corresponding to the same electric vehicle.
[0105] In some implementations, the sample extraction unit 503 is specifically used to: sequentially calculate the cumulative characteristic factor and the real-time characteristic factor of each box of data in the K boxes of data; and generate the K initial samples according to the cumulative characteristic factor and the real-time characteristic factor of the K boxes of data.
[0106] In some embodiments, the feature screening unit 504 is specifically used to: perform intra-group feature screening on the N groups of initial samples respectively to obtain N groups of sample features corresponding to N types of vehicle models; use the N groups of sample features to train the feature screening model; calculate the feature importance of each sample feature in the N groups of sample features through the trained feature screening model; perform secondary feature screening according to the feature importance of each sample feature in the N groups of sample features to obtain the aging feature factor.
[0107] In some embodiments, the feature screening unit 504 is specifically used to: respectively use each group of initial samples in the N groups of initial samples as the target group initial samples; normalize the feature values of each sample feature in the target group initial samples to obtain the normalized feature values of each sample feature in the target group initial samples; for each sample feature in the target group initial samples, calculate the feature variance of the sample feature according to the normalized feature value of the sample feature; perform intra-group feature screening according to the feature variance of each sample feature in the target group initial samples, and screen out each sample feature whose feature variance is greater than a preset variance threshold as a group of sample features.
[0108] In some embodiments, the feature screening unit 504 is specifically used to: respectively use each group of sample features in the N groups of sample features as a target group sample feature, and sort the features within the group according to the feature importance of each sample feature in the target group sample feature; perform intra-group feature screening based on the intra-group feature sorting results of the N groups of sample features to screen out N groups of important sample features corresponding one-to-one to the N types of vehicle models; perform inter-group feature comparison on the N groups of important sample features to screen out the common sample features of the N types of vehicle models, as well as each non-common sample feature whose feature importance is greater than a preset threshold, as the aging feature factor.
[0109] In some embodiments, the feature screening unit 504 is specifically used to: sort the sample features in the target group from large to small according to the feature importance of each sample feature; for each group of important sample features in the N groups of important sample features, successively accumulate the feature importance of the group of sample features from large to small until the accumulation result reaches a first preset weight threshold, retain the sample features in the group of sample features that participate in the sequential accumulation calculation, and eliminate the sample features whose feature importance is less than the second weight threshold from the sample features that participate in the sequential accumulation calculation, so as to obtain the important sample features in the group of sample features.
[0110] In some embodiments, a feature application unit is further included, which is specifically used to: use the aging characteristic factors to train a prediction model for predicting the health of the power battery, and / or analyze the impact direction of the aging characteristic factors on the health of the power battery.
[0111] Regarding the above-mentioned device, the specific functions of each module therein have been described in detail in the embodiment of the method for extracting the aging characteristic factor of the power battery provided in the embodiment of this specification, and will not be elaborated here.
[0112] Based on the same inventive concept, this specification also provides an electronic device, such as Figure 6As shown, it includes a memory 604, a processor 602, and a computer program stored on the memory 604 and executable on the processor 602. When the processor 602 executes the program, it implements the method for extracting the aging characteristic factors of the power battery described above.
[0113] Among them, in Figure 6 the bus architecture (represented by bus 600), bus 600 can include any number of interconnected buses and bridges. Bus 600 links together various circuits including one or more processors represented by processor 602 and memory represented by memory 604. Bus 600 can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and thus will not be further described herein. Bus interface 606 provides an interface between bus 600 and receiver 601 and transmitter 603. Receiver 601 and transmitter 603 can be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 602 is responsible for managing bus 600 and general processing, while memory 604 can be used to store data used by processor 602 when performing operations.
[0114] Based on the same inventive concept, this specification also provides a computer-readable storage medium with a computer program stored thereon. When the program is executed by a processor, it implements any one of the embodiments of the method for extracting the aging characteristic factors of the power battery described above.
[0115] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0116] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 or steps for realizing the functions specified in one block or a plurality of blocks.
[0118] Although the preferred embodiments of the present specification have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present specification.
[0119] Obviously, those skilled in the art can make various changes and modifications to the present specification without departing from the spirit and scope of the present specification. Thus, if these modifications and variations of the present specification fall within the scope of the claims of the present specification and their equivalent technologies, the present specification is also intended to include these modifications and variations.
Claims
1. A method for extracting aging characteristic factors of a power battery, characterized in that, include: Get the vehicle detail dataset of electric vehicles; Filtering out N groups of vehicle detail data grouped by vehicle type from the vehicle detail data set, where N is an integer greater than 1; Generate an initial sample set according to the N groups of vehicle detailed data, wherein the initial sample set includes N groups of initial samples grouped according to vehicle types; The aging characteristic factors that affect the health of the battery are screened out from the N groups of initial samples by means of a feature screening model, including: performing intra-group feature screening on the N groups of initial samples respectively to obtain N groups of sample features corresponding to N types of vehicle models; training the feature screening model using the N groups of sample features; calculating the feature importance of each sample feature in the N groups of sample features by means of the trained feature screening model; performing secondary feature screening according to the feature importance of each sample feature in the N groups of sample features to obtain the aging characteristic factors, wherein the feature screening model has the ability to record feature usage during the training process; The feature screening model is a tree structure model, and the feature importance of the sample feature is determined according to the ratio of the number of times the sample feature is used when all subtrees of the tree structure model are split to the total number of times all subtrees of the tree structure model are split.
2. The method according to claim 1, wherein The step of obtaining a vehicle detailed data set of an electric vehicle includes: Receiving vehicle operation data from a plurality of electric vehicles, wherein the vehicle operation data is collected and uploaded by electric vehicles of M types through their signal acquisition devices, each type of vehicle includes a plurality of electric vehicles, and M is an integer greater than or equal to N; The vehicle operation data is cleaned to obtain the vehicle detail data set.
3. The method according to claim 1, characterized in that, The step of filtering out N groups of vehicle detail data grouped by vehicle type from the vehicle detail data set includes: Processing the vehicle detail data set to obtain M groups of vehicle detail data corresponding to the M types of vehicle models, wherein each group of vehicle detail data includes vehicle detail data of multiple electric vehicles of the same model; According to a preset vehicle number threshold, the N groups of vehicle detail data are screened out from the M groups of vehicle detail data.
4. The method according to claim 3, characterized in that, The vehicle detail data set is processed to obtain M groups of vehicle detail data corresponding to M types of vehicle models, including: Grouping the vehicle detail data sets according to vehicle models to obtain M data groups corresponding to the M vehicle models, each data group including vehicle detail data of each electric vehicle of the same vehicle model; The M data groups are respectively screened within the group according to the preset data screening rules to screen out M groups of vehicle detailed data corresponding to the M types of vehicle models.
5. The method according to claim 4, wherein The performing intra-group data screening on the M data groups respectively according to the preset data screening rules includes: Taking each of the M data groups as a target data group; For the vehicle detailed data of each electric vehicle in the target data group, if the mileage span of the electric vehicle is less than a preset mileage threshold, or the time span of data collection of the electric vehicle is less than a preset time threshold, the vehicle detailed data corresponding to the electric vehicle is eliminated; Based on the vehicle detail data of the remaining electric vehicles in the M data groups, the M groups of vehicle detail data are obtained.
6. The method according to any one of claims 1-5, characterized in that, The generating an initial sample set according to the N groups of vehicle detail data includes: Respectively taking the vehicle detail data of each electric vehicle in the N groups of vehicle detail data as the target vehicle detail data; Extracting target data from the target vehicle detail data; Performing binning processing on the target data at a preset mileage interval to obtain K bins of data belonging to the same electric vehicle, where K is an integer greater than 1; Performing sample extraction on the K bins of data to generate K initial samples corresponding to the same electric vehicle.
7. The method according to claim 6, characterized in that The performing sample extraction on the K bins of data to generate K initial samples corresponding to the same electric vehicle includes: Sequentially calculating the cumulative characteristic factor and the real-time characteristic factor of each bin of data in the K bins of data; Generating the K initial samples according to the cumulative characteristic factor and the real-time characteristic factor of the K bins of data.
8. The method according to claim 1, characterized in that, The respectively performing intra-group feature screening on the N groups of initial samples to obtain N groups of sample features corresponding one-to-one to N vehicle models includes: Respectively taking each group of initial samples in the N groups of initial samples as the target group of initial samples; Performing normalization processing on the feature values of each sample feature in the target group of initial samples to obtain the normalized feature values of each sample feature in the target group of initial samples; For each sample feature in the target group of initial samples, calculating the feature variance of the sample feature according to the normalized feature value of the sample feature; Performing intra-group feature screening according to the feature variances of each sample feature in the target group of initial samples, and screening out each sample feature with a feature variance greater than a preset variance threshold as a group of sample features.
9. The method according to claim 1, characterized in that The obtaining the aging characteristic factor by performing secondary feature screening according to the feature importance of each sample feature in the N groups of sample features includes: Respectively taking each group of sample features in the N groups of sample features as the target group of sample features, and performing intra-group feature ranking according to the feature importance of each sample feature in the target group of sample features; Performing intra-group feature screening according to the intra-group feature ranking results of the N groups of sample features to screen out N groups of important sample features corresponding one-to-one to the N vehicle models; Performing inter-group feature comparison on the N groups of important sample features, and screening out the common sample features of the N vehicle models and each non-common sample feature with a feature importance greater than a preset threshold as the aging characteristic factor.
10. The method according to claim 9, wherein The performing intra-group feature ranking according to the feature importance of each sample feature in the target group of sample features includes: Performing sorting from large to small according to the feature importance of each sample feature in the target group of sample features; The performing intra-group feature screening according to the intra-group feature ranking results of the N groups of sample features to screen out N groups of important sample features corresponding one-to-one to the N vehicle models includes: For each group of important sample features among the N groups of important sample features, the feature importance degrees of the sample features in this group are accumulated in descending order until the accumulated result reaches the first preset weight threshold. Then, retain each sample feature participating in the successive accumulation calculation in this group of sample features, and remove the sample features with feature importance degrees less than the second weight threshold from the sample features participating in the successive accumulation calculation, so as to obtain the important sample features in this group of sample features.
11. The method according to claim 1, characterized in that, Further included are: training a prediction model for predicting the health state of a power battery using the aging feature factor, and / or analyzing the influence direction of the aging feature factor on the health state of the power battery.
12. An aging characteristic factor extraction device for a power battery, characterized in that, Including: a data acquisition unit for acquiring a vehicle detail data set of an electric vehicle; a data screening unit for screening out N groups of vehicle detail data grouped by vehicle models from the vehicle detail data set, where N is an integer greater than 1; a sample extraction unit for generating an initial sample set according to the N groups of vehicle detail data, and the initial sample set includes N groups of initial samples grouped by vehicle models; a feature screening unit for screening out the aging feature factors affecting the battery health state from the N groups of initial samples through a feature screening model, including: respectively performing in-group feature screening on the N groups of initial samples to obtain N groups of sample features corresponding one by one to N vehicle models; training the feature screening model using the N groups of sample features; calculating the feature importance degree of each sample feature in the N groups of sample features through the trained feature screening model; performing secondary feature screening according to the feature importance degrees of each sample feature in the N groups of sample features to obtain the aging feature factor, the feature screening model has the ability to record the feature usage situation during the training process, the feature screening model is a tree-structured model, and the feature importance degree of the sample feature is determined according to the ratio of the number of times the sample feature is used when all sub-trees of the tree-structured model are split to the total number of times all sub-trees of the tree-structured model are split.
13. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by the processor, it implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Gradient lifting tree modeling and prediction method for health condition of lithium battery
CN108896914A
Cited By
M-choline receptor agonist compound, preparation method therefor and use thereof
WO2023143575A1