A SOH prediction method for new energy batteries based on user driving habits
By standardizing and clustering driving information, user information and vehicle configuration information in the battery SOH prediction method, and analyzing user driving habit data with the LSTM model, the problem of inaccurate battery SOH prediction in the prior art is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202410994632.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-07-24
AI Technical Summary
In the prior art, the prediction of battery SOH is inaccurate, and it is not possible to effectively comprehensively analyze the battery usage scenarios and user driving habit data.
By obtaining driving information, user information and vehicle configuration information, standardization, data cleaning and normalization are carried out, and unsupervised clustering is carried out to obtain user habit classification, vehicle classification and user information classification. Then, an LSTM model is constructed and the sample data is analyzed using these classifications to improve the accuracy of battery SOH prediction.
By conducting a comprehensive analysis of user driving habit data, the accuracy of battery SOH prediction is improved, and the problem of inaccurate prediction in the prior art is overcome.
Smart Images

Figure CN118965040B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of new energy batteries, and in particular to a SOH prediction method for new energy batteries based on user driving habits. Background Art
[0002] In the widespread application of electric vehicles (EVs) and hybrid electric vehicles (HEVs), the state of health (SOH) of the battery system is crucial to the performance of the vehicle and the user's driving experience. SOH not only directly affects the vehicle's range and charging time, but is also closely related to the safety of the battery. Therefore, accurate prediction of the battery's SOH is of great significance for the maintenance and management of electric vehicles.
[0003] Chinese Patent Publication No.: CN116243197A discloses a battery SOH prediction method and device, including: S101 determining the temperature range where the surface temperature of the battery cell is located; S102 determining the charging current of the battery in the constant current stage of charging within the temperature range; S103 calculating the first impedance of the battery according to the real-time voltage, open circuit voltage and the charging current; S104 performing temperature normalization on the first impedance to obtain the first temperature normalized impedance; S105 repeating the above steps S101-S104 to obtain multiple first temperature normalized impedances under the set SOC; and S106 determining the battery SOH through the multiple first temperature normalized impedances. This invention realizes the analysis of the battery SOH based on the internal operation data of the battery, but does not realize the comprehensive analysis of the battery usage scenario and the user's driving habit data, and there is a problem of inaccurate battery SOH prediction. Summary of the invention
[0004] To this end, the present invention provides a SOH prediction method for a new energy battery based on a user's driving habits, so as to overcome the problem of inaccurate battery SOH prediction in the prior art.
[0005] To achieve the above object, the present invention provides a method for predicting the SOH of a new energy battery based on a user's driving habits, comprising:
[0006] Step S1, obtaining driving information, user information and vehicle configuration information;
[0007] Step S2, standardizing the driving information, user information and vehicle configuration information, and performing data cleaning on the standardized driving information;
[0008] Step S3, normalizing the cleaned driving information, standardized user information and vehicle configuration information;
[0009] Step S4, performing unsupervised clustering on the normalized driving information, user information and vehicle configuration information to obtain user habit classification, vehicle classification and user information classification;
[0010] Step S5, obtaining battery information, and analyzing sample data according to the battery information and user habit classification, vehicle classification, and user information classification;
[0011] Step S6, performing data cleaning and normalization processing on the sample data, and dividing the sample data into training data and test data;
[0012] Step S7, constructing an LSTM model according to the training data;
[0013] Step S8, testing the LSTM model according to the test data.
[0014] Furthermore, the data standardization and data cleaning method in step S2 includes:
[0015] Step S201: standardize driving information, user information and vehicle configuration information.
[0016] Step S202, extracting unique identification data from the standardized driving information, user information and vehicle configuration information;
[0017] Step S203, performing missing value processing on the standardized driving information, user information and vehicle configuration information according to the unique identification data;
[0018] Step S204, filtering abnormal data according to the standardized driving information;
[0019] Step S205: performing outlier processing on the standardized driving information according to the abnormal data.
[0020] Furthermore, the step S201 uses oneHot encoding to encode the optional data in the driving information, user information and vehicle configuration information, and uses numerical values to represent each data item;
[0021] The step S201 outputs the driving information in the order of time, vehicle speed, maximum vehicle speed, average vehicle speed, average engine speed, maximum engine speed, average brake pedal position, average accelerator pedal position, total mileage, average single mileage, number of accelerations, average number of accelerations per kilometer, number of brakes, average number of brakes per kilometer, idling time, average idling time each time, number of speeding, average number of speeding per kilometer, energy recovery, average energy recovery per kilometer, average battery voltage, average battery current, average SOC, number of fast charges, number of slow charges, charging time, charging amount, charging efficiency, average battery temperature, ambient temperature, driving mode, road type and SOH value. CSV file, output user information in the order of user ID, user city, user age, user gender, user frame number, user occupation, user driving experience, charging frequency, charging preference, average charging time, and charging location as a CSV file, and output vehicle configuration information in the order of vehicle frame number, vehicle brand, vehicle model, body type, production year, drive type, motor type, motor power, maximum torque, transmission system, battery type, battery capacity, battery voltage, number of battery modules, battery cooling system power, acceleration time per 100 kilometers, vehicle maximum speed, charging interface type, vehicle charging efficiency, braking system type, and autonomous driving type as a CSV file;
[0022] The step S202 extracts the vehicle frame number from the standardized driving information as unique identification data, extracts the user ID and the user vehicle frame number from the standardized user information as unique identification data, and extracts the vehicle frame number from the standardized vehicle configuration information as unique identification data;
[0023] The step S203 performs missing value processing on the standardized driving information, user information and vehicle configuration information according to the unique identification data, wherein:
[0024] When the unique identification data is missing, the standardized driving information, user information and vehicle configuration information corresponding to the currently analyzed unique identification data are deleted;
[0025] When the unique identification data is not a missing value, the standardized driving information, user information and vehicle configuration information corresponding to the currently analyzed unique identification data are retained;
[0026] The step S204 uses the DBScan algorithm to filter abnormal data according to the standardized driving information;
[0027] The step S205 performs abnormal value processing on the standardized driving information according to the abnormal data, and deletes the abnormal data in the standardized driving information.
[0028] Furthermore, the normalization processing method in step S3 includes:
[0029] Step S301, extracting quantified data from the cleaned driving information, standardized user information and vehicle configuration information;
[0030] Step S302, normalizing the quantized data;
[0031] Step S303, retaining the original values of the driving information, user information and vehicle configuration information that are not quantized data.
[0032] Further, the step S301 extracts the average vehicle speed, the maximum vehicle speed, the average engine speed, the maximum engine speed, the average brake pedal position, the average accelerator pedal position, the total mileage, the average single mileage, the number of accelerations, the average number of accelerations per kilometer, the number of brakes, the average number of brakes per kilometer, the idling time, the average idling time each time, the number of overspeeding, the average number of overspeeding per kilometer, the energy recovery amount, the average energy recovery amount per kilometer, the average battery voltage, the average battery current, the average SOC, the number of fast charges, the number of slow charges, the charging time, the charging amount, the charging efficiency, the average battery temperature, the ambient temperature and the SOH value from the driving information after data cleaning as quantitative data, extracts the charging frequency and the average charging time from the standardized user information as quantitative data, and extracts the motor power, the maximum torque, the battery capacity, the battery voltage, the number of battery modules, the acceleration time per 100 kilometers, the maximum speed of the vehicle and the charging efficiency of the vehicle from the standardized vehicle configuration information as quantitative data;
[0033] The step S302 performs normalization processing on the quantized data using a first normalization formula, and the first normalization formula is as follows:
[0034] X v =(x v -x vmin ) / (x vmax -x vmin );
[0035] Among them, X v Represents the normalized quantized data, v represents the quantized data number, v∈N + , x v Represents quantitative data, x vmin Represents the minimum value of the quantized data, x vmax Indicates the maximum value of the quantized data.
[0036] Furthermore, the step S4 uses the DBScan algorithm to perform unsupervised clustering on the normalized driving information, user information and vehicle configuration information;
[0037] In step S4, the number of other data points in the ε neighborhood of each data point in the normalized driving information, user information and vehicle configuration information is counted as the number of clustering neighborhood points, N ε (G(y,a))={G(y,b)∈D|dist(G(y,a),G(y,b))≤ε}, where G(y,a) represents the data in the yth column and the ath row of the normalized driving information, user information and vehicle configuration information, G(y,b) represents the data in the yth column and the bth row of the normalized driving information, user information and vehicle configuration information, y represents the data column number, y∈N + , a represents the first data row number, a∈N + , b represents the second data flight number, b∈N + , a≠b, ε represents the neighborhood radius, dist(G(y,a),G(y,b)) represents the distance between the data in the yth column and the ath row and the data in the yth column and the bth row, N ε (G(y,a)) represents the number of clustering neighborhood points, and D represents the normalized driving information, user information, and vehicle configuration information;
[0038] In step S4, when N ε When (G(y,a))≥MinPTs(y), the data points corresponding to the number of neighborhood points of the current analysis cluster are taken as core data points, the core data points are added to the cluster V, and the other data points in the neighborhood of the core data point ε are taken as data points to be analyzed, and the data points to be analyzed are set to G'(y,a). When N ε When (G'(y,a))≥MinPTs(y), the data point to be analyzed is set as visited, the data point to be analyzed is taken as the core data point, and added to the cluster V. ε When (G'(y,a))<MinPTs(y), the data point to be analyzed is set as visited, and the analysis of other data points in the neighborhood of the core data point ε is repeated until all the data points to be analyzed are set as visited, where V represents the core data point cluster set, MinPTs(y) represents the minimum sample threshold of the data point, and N ε (G'(y,a)) represents the number of cluster neighborhood points of the data point to be analyzed;
[0039] In the step S4, the clustered driving information is used as the user habit classification, the clustered user information is used as the user information classification, and the clustered vehicle configuration information is used as the vehicle classification.
[0040] Furthermore, the step S5 integrates the battery information with the user habit classification, vehicle classification and user information classification into a CSV file, and uses it as sample data. When integrating the CSV file, the integration is performed in the order of battery brand ID, vehicle series ID, battery model ID, age, user habit classification, vehicle classification, user information classification and measured SOH.
[0041] Furthermore, the method for processing the sample data in step S6 includes:
[0042] Step S601, performing data cleaning on sample data;
[0043] Step S602, normalizing the sample data after data cleaning;
[0044] Step S603, dividing the normalized sample data into training data and test data.
[0045] Furthermore, in step S601, data cleaning is performed on the default values in the sample data, wherein:
[0046] When the battery brand ID, vehicle ID, battery model ID and measured SOH in the sample data are default values, delete the current analysis sample data;
[0047] When the user habit classification, vehicle classification and user information classification are default values, the default value in the current sample data is set to the data with the highest weight in each classification;
[0048] When the useful life is the default value, no data cleaning is performed on the useful life;
[0049] The step S602 normalizes the measured SOH in the sample data after data cleaning by using a second normalization formula, and the second normalization formula is as follows:
[0050] S1=(SS min ) / (S max -S min );
[0051] Where S1 represents the measured SOH after normalization, S represents the measured SOH, and S min Indicates the minimum value of the measured SOH, S max It represents the maximum value among the measured SOH;
[0052] In step S603, the normalized sample data is processed in the order of data in the CSV file, with the first α data as training data and the last 1-α data as test data, where α represents the division threshold.
[0053] Furthermore, in step S7, the training data is input into the LSTM model for training, and the parameters in the LSTM model are set as follows:
[0054] Batch_size = 128, dropout = 0.8, lr = 0.001, use sigmoid function as activation function;
[0055] The step S8 obtains the predicted SOH of the three dimensions of user habit classification, vehicle classification and user information classification in the training data respectively, and sets the predicted SOH to s1, s2, s3, where s1 represents the predicted SOH of user habit classification, s2 represents the predicted SOH of vehicle classification, and s3 represents the predicted SOH of user information classification, and adjusts the neighborhood radius and the minimum sample threshold of data points of each classification according to the predicted SOH and the actual SOH, where:
[0056] When |s1 / S1-1|≤β, the predicted SOH of user habit classification is determined to be accurate, and the neighborhood radius of driving information and the minimum sample threshold of data points are not adjusted;
[0057] When |s1 / S1-1|>β, the predicted SOH of user habit classification is determined to be inaccurate, and the neighborhood radius of driving information and the minimum sample threshold of data points are adjusted;
[0058] When |s2 / S1-1|≤β, the predicted SOH of vehicle classification is determined to be accurate, and the neighborhood radius of vehicle configuration information and the minimum sample threshold of data points are not adjusted;
[0059] When |s2 / S1-1|>β, the predicted SOH of vehicle classification is determined to be inaccurate, and the neighborhood radius of vehicle configuration information and the minimum sample threshold of data points are adjusted;
[0060] When |s3 / S1-1|≤β, the predicted SOH of the user information classification is determined to be accurate, and the neighborhood radius of the user information and the minimum sample threshold of the data point are not adjusted;
[0061] When |s3 / S1-1|>β, the predicted SOH of user information classification is determined to be inaccurate, and the neighborhood radius of user information and the minimum sample threshold of data points are adjusted;
[0062] Where β represents the comparison threshold.
[0063] Compared with the prior art, the beneficial effects of the present invention are that, by acquiring driving information, user information and vehicle configuration information, the accuracy of information acquisition is improved, thereby improving the accuracy of battery SOH prediction; by standardizing and cleaning the driving information, user information and vehicle configuration information, the driving information, user information and vehicle configuration information are unified into the same data to ensure the uniformity of data format and remove abnormal data in each data, thereby improving the accuracy of battery SOH prediction; by normalizing the driving information after data cleaning and the user information and vehicle configuration information after standardization, the data of the driving information, user information and vehicle configuration information are made smoother, and the influence of data fluctuation on data analysis is reduced, thereby improving the accuracy of battery SOH prediction; by unsupervised clustering of the normalized driving information, user information and vehicle configuration information, the user habits are analyzed. Habit classification, vehicle classification and user information classification are performed to improve the accuracy of battery SOH prediction. By acquiring battery information, sample data composed of user driving data and battery data is constructed to improve the accuracy of battery SOH prediction. The sample data is cleaned and normalized to increase the accuracy of sample data to improve the accuracy of battery SOH prediction. The sample data is divided into training data and test data to improve the accuracy of battery SOH prediction. The training data is analyzed to construct an LSTM model to predict the battery SOH according to the user driving habit data, thereby improving the accuracy of battery SOH prediction. The test data is analyzed to test the LSTM model to determine the accuracy of the battery SOH prediction value, thereby improving the accuracy of battery SOH prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 Flow chart of the SOH prediction method of new energy batteries based on user driving habits in this embodiment.
[0065] Figure 2 Flowchart of the data standardization and data cleaning method of this embodiment.
[0066] Figure 3 Flow chart of the normalization processing method of this embodiment.
[0067] Figure 4 Flow chart of the method for processing sample data in this embodiment. DETAILED DESCRIPTION
[0068] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0069] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0070] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0071] See also Figure 1 As shown, it is a SOH prediction method of a new energy battery based on a user's driving habits in this embodiment, comprising:
[0072] Step S1, obtaining driving information, user information and vehicle configuration information, wherein the driving information includes time, vehicle speed, maximum vehicle speed, average vehicle speed, average engine speed, maximum engine speed, average brake pedal position, average accelerator pedal position, total mileage, average single mileage, number of accelerations, average number of accelerations per kilometer, number of brakes, average number of brakes per kilometer, idling time, average idling time per time, number of overspeeding, average number of overspeeding per kilometer, energy recovery amount, average energy recovery amount per kilometer, average battery voltage, average battery current, average SOC, number of fast charges, number of slow charges, charging time, charging amount, charging efficiency, average battery temperature, ambient temperature, driving mode, road type and SOH value, wherein the time is the time corresponding to each driving information during driving, the brake pedal position and the accelerator pedal position are the driver's pedaling depth, the energy recovery amount is the electric energy recovered when the vehicle decelerates, and the driving information is obtained by importing the backend data of the automobile driving data management platform, and the user information includes the user ID, the user's location City, user age, user gender, user vehicle frame number, user occupation, user driving experience, charging frequency, charging preference, average charging time, charging location, the charging preference includes fast charging and slow charging, the charging location includes private charging piles and public charging piles, the user information is obtained by user interactive input, the vehicle configuration information includes vehicle frame number, vehicle brand, vehicle model, body type, production year, drive type, motor type, motor power, maximum torque, transmission system, battery type, battery capacity, battery voltage, number of battery modules, battery cooling system power, acceleration time from 0 to 100 km / h, vehicle maximum speed, charging interface type, vehicle charging efficiency, braking system type and autonomous driving type, the body type includes sedan, SUV and MPV, the drive type includes front drive, rear drive and all-wheel drive, the motor type includes permanent magnet, asynchronous and induction, the braking system type includes with braking system and without braking system, the autonomous driving type includes with autonomous driving and without autonomous driving, and the vehicle configuration information is obtained by user interactive input;
[0073] Step S2, standardizing the driving information, user information and vehicle configuration information, and performing data cleaning on the standardized driving information;
[0074] Step S3, normalizing the cleaned driving information, standardized user information and vehicle configuration information;
[0075] Step S4, performing unsupervised clustering on the normalized driving information, user information and vehicle configuration information to obtain user habit classification, vehicle classification and user information classification;
[0076] Step S5, obtaining battery information, and analyzing sample data according to battery information and user habit classification, vehicle classification and user information classification, wherein the battery information includes battery brand ID, vehicle series ID, battery model ID, service life and measured SOH, wherein the battery brand ID is the ID of the brand to which the battery belongs, the vehicle series ID is the ID of the vehicle series in which the current battery is installed, and the battery model ID is the ID of the battery model under the current battery brand, and the battery information is obtained by user interactive input;
[0077] Step S6, performing data cleaning and normalization processing on the sample data, and dividing the sample data into training data and test data;
[0078] Step S7, constructing an LSTM model according to the training data;
[0079] Step S8, testing the LSTM model according to the test data.
[0080] See also Figure 2 As shown, the data standardization and data cleaning method in step S2 of this embodiment includes:
[0081] Step S201: standardize driving information, user information and vehicle configuration information.
[0082] Step S202, extracting unique identification data from the standardized driving information, user information and vehicle configuration information;
[0083] Step S203, performing missing value processing on the standardized driving information, user information and vehicle configuration information according to the unique identification data;
[0084] Step S204, filtering abnormal data according to the standardized driving information;
[0085] Step S205: performing outlier processing on the standardized driving information according to the abnormal data.
[0086] Specifically, in step S201 of this embodiment, the optional data in the driving information, user information and vehicle configuration information are encoded using a oneHot encoding method, and various data are represented by numerical values.
[0087] Specifically, in step S201 of this embodiment, when oneHot encoding is performed on the option data in the driving information, user information and vehicle configuration information, if the option data is charging preference, 1 is used to indicate that the charging preference is fast charging, and 2 is used to indicate that the charging preference is slow charging. When the option data is body type, 1 is used to indicate that the body type is a sedan, 2 is used to indicate that the body type is an SUV, and 3 is used to indicate that the body type is an MPV.
[0088] Specifically, the option data in this embodiment includes charging preference, charging location, vehicle body type, drive type, motor type, braking system type and autonomous driving type.
[0089] Specifically, step S201 in this embodiment processes the driving information in the order of time, vehicle speed, maximum vehicle speed, average vehicle speed, average engine speed, maximum engine speed, average brake pedal position, average accelerator pedal position, total mileage, average single mileage, number of accelerations, average number of accelerations per kilometer, number of brakes, average number of brakes per kilometer, idling time, average idling time per time, number of speeding, average number of speeding per kilometer, energy recovery, average energy recovery per kilometer, average battery voltage, average battery current, average SOC, number of fast charges, number of slow charges, charging time, charging amount, charging efficiency, average battery temperature, ambient temperature, driving mode, road type and SOH value. The structure is output as a CSV file, and the user information is output as a CSV file in the order of user ID, user city, user age, user gender, user frame number, user occupation, user driving experience, charging frequency, charging preference, average charging time, and charging location. The vehicle configuration information is output as a CSV file in the order of vehicle frame number, vehicle brand, vehicle model, body type, production year, drive type, motor type, motor power, maximum torque, transmission system, battery type, battery capacity, battery voltage, number of battery modules, battery cooling system power, acceleration time from 0 to 100 km / h, vehicle maximum speed, charging interface type, vehicle charging efficiency, braking system type, and autonomous driving type.
[0090] Specifically, step S202 described in this embodiment extracts the frame number in the standardized driving information as the unique identification data, extracts the user ID and user frame number in the standardized user information as the unique identification data, and extracts the vehicle frame number in the standardized vehicle configuration information as the unique identification data.
[0091] Specifically, step S203 in this embodiment performs missing value processing on the standardized driving information, user information and vehicle configuration information according to the unique identification data, wherein:
[0092] When the unique identification data is missing, the standardized driving information, user information and vehicle configuration information corresponding to the currently analyzed unique identification data are deleted;
[0093] When the unique identification data is not a missing value, the standardized driving information, user information and vehicle configuration information corresponding to the currently analyzed unique identification data are retained.
[0094] Specifically, step S204 in this embodiment uses the DBScan algorithm to filter abnormal data according to the standardized driving information.
[0095] Specifically, step S204 in this embodiment obtains driving parameters by matching the standardized driving information with time one by one, and sets the driving parameters as F j (i), i represents time, j represents the data number of driving information, j∈N + , and calculate the distance parameters through the distance analysis formula according to the driving parameters. The distance analysis formula is as follows:
[0096]
[0097] Among them, W j (i,k) represents the distance parameter, k represents the non-current analysis time, k≠i, T(i,k) represents the time interval between i and k, F j (k) represents the driving parameters at a time other than the current analysis time.
[0098] Specifically, the data number of the driving information described in this embodiment is a number used to distinguish the driving data currently being analyzed as a certain data, such as F1(i) represents the driving parameter corresponding to the average speed, F2(i) represents the driving parameter corresponding to the maximum speed, etc.
[0099] Specifically, in step S204 of this embodiment, the distance parameters corresponding to the driving parameters are counted to meet the W j The distance parameter of (i,k)≤w(j) is used as the driving neighborhood point, the number of driving neighborhood points is counted as the number of driving neighborhood points, and the driving parameters are classified according to the number of driving neighborhood points, where:
[0100] When N j When (i)≥MP(j), the driving parameter is determined to be the core object;
[0101] When N j When (i) < MP(j), the driving parameter is determined to be a non-core object;
[0102] Among them, w(j) represents the neighborhood threshold of driving information, N j (i) represents the number of driving neighborhood points, and MP(j) represents the minimum number of samples. It can be understood that in this embodiment, the driving information field threshold and the minimum number of samples are not specifically limited, and those skilled in the art can freely set them as long as they meet the classification of driving parameters.
[0103] Specifically, step S204 in this embodiment creates a new cluster C u , add the core object to C u , extract the jThe driving parameters corresponding to the non-current analysis time of the distance parameter (n, k) ≤ w(j) are taken as the direct density reachable neighborhood points. When the direct density reachable neighborhood points are core objects or N j When (m)≥MP(j), add the direct density reachability parameter to C u , and continue to search for the direct density reachable neighboring points of the current direct density reachable neighboring point, until there are no new direct density reachable neighboring points that can be added to C u Among them, C u Represents the cluster set, u represents the cluster set number, u∈N + , W j (n, k) represents the distance parameter of the core object, n represents the time of the core object, N j (m) represents the number of driving neighborhood points that can be reached by direct density, and m represents the time it takes to reach the neighborhood point by direct density.
[0104] Specifically, in step S204 of this embodiment, after traversing all core objects, the driving parameters that do not belong to any cluster are regarded as abnormal data.
[0105] Specifically, step S205 in this embodiment performs abnormal value processing on the standardized driving information according to the abnormal data, and deletes the abnormal data in the standardized driving information.
[0106] Specifically, in this embodiment, when the abnormal value processing is performed on the standardized driving information, the abnormal data can be output, and the abnormal data in the standardized driving information can be deleted after manual confirmation.
[0107] See also Figure 3 As shown, it is the normalization processing method of step S3 of this embodiment, including:
[0108] Step S301, extracting quantified data from the cleaned driving information, standardized user information and vehicle configuration information;
[0109] Step S302, normalizing the quantized data;
[0110] Step S303, retaining the original values of the driving information, user information and vehicle configuration information that are not quantized data.
[0111] Specifically, step S301 described in this embodiment extracts the average vehicle speed, maximum vehicle speed, average engine speed, maximum engine speed, average brake pedal position, average accelerator pedal position, total mileage, average single mileage, number of accelerations, average number of accelerations per kilometer, number of brakes, average number of brakes per kilometer, idling time, average idling time each time, number of speeding, average number of speeding per kilometer, energy recovery amount, average energy recovery amount per kilometer, average battery voltage, average battery current, average SOC, number of fast charges, number of slow charges, charging time, charging amount, charging efficiency, average battery temperature, ambient temperature and SOH value from the driving information after data cleaning as quantitative data, extracts the charging frequency and average charging time from the standardized user information as quantitative data, and extracts the motor power, maximum torque, battery capacity, battery voltage, number of battery modules, acceleration time per 100 kilometers, vehicle maximum speed and vehicle charging efficiency from the standardized vehicle configuration information as quantitative data.
[0112] Specifically, in step S302 of this embodiment, the quantized data is normalized by a first normalization formula, and the first normalization formula is as follows:
[0113] X v =(x v -x vmin ) / (x vmax -x vmin );
[0114] Among them, X v Represents the normalized quantized data, v represents the quantized data number, v∈N + , x v Represents quantitative data, x vmin Represents the minimum value of the quantized data, x vmax Indicates the maximum value of the quantized data.
[0115] Specifically, the quantized data number in this embodiment is a number used to distinguish the quantized data being analyzed as a certain data. For example, X1 can be used to indicate that the quantized data currently being analyzed is the average vehicle speed, and X2 can be used to indicate that the quantized data currently being analyzed is the maximum vehicle speed.
[0116] Specifically, step S4 in this embodiment uses the DBScan algorithm to perform unsupervised clustering on the normalized driving information, user information and vehicle configuration information.
[0117] Specifically, in step S4 of this embodiment, the number of other data points in the ε neighborhood of each data point in the driving information, user information and vehicle configuration information after the normalization processing is statistically processed is used as the number of cluster neighborhood points, N ε(G(y,a))={G(y,b)∈D|dist(G(y,a),G(y,b))≤}, where G(y,a) represents the data in the yth column and the ath row of the normalized driving information, user information and vehicle configuration information, G(y,b) represents the data in the yth column and the bth row of the normalized driving information, user information and vehicle configuration information, y represents the data column number, y∈N + , a represents the first data row number, a∈N + , b represents the second data flight number, b∈N + , a≠b, ε represents the neighborhood radius, and dist(G(y,a),G(y,b)) represents the distance between the data in the yth column and the ath row and the data in the yth column and the bth row. N ε (G(y,a)) represents the number of cluster neighborhood points, and D represents the normalized driving information, user information, and vehicle configuration information. It can be understood that the value of the domain radius is not specifically limited in this embodiment, and those skilled in the art can freely set it as long as it satisfies the analysis of the number of cluster neighborhood points.
[0118] Specifically, the data points described in this embodiment are the various data in the driving information, user information and vehicle configuration information after normalization. The driving information, user information and vehicle configuration information are CSV file data. The CSV file is a file used to store electronic spreadsheet information. Each cell in the table represents a data point. Each cell can be represented by a row and column position in the table. The data column number indicates the column number of the cell in the table, and the data row number indicates the row number of the cell in the table.
[0119] Specifically, in step S4 of this embodiment, when N ε When (G(y,a))≥MinPTs(y), the data points corresponding to the number of neighborhood points of the current analysis cluster are taken as core data points, the core data points are added to the cluster V, and the other data points in the neighborhood of the core data point ε are taken as data points to be analyzed, and the data points to be analyzed are set to G'(y,a). When N ε When (G'(y,a))≥MinPTs(y), the data point to be analyzed is set as visited, the data point to be analyzed is taken as the core data point, and added to the cluster V. ε When (G'(y,a))<MinPTs(y), the data point to be analyzed is set as visited, and the analysis of other data points in the neighborhood of the core data point ε is repeated until all the data points to be analyzed are set as visited, where V represents the core data point cluster set, MinPTs(y) represents the minimum sample threshold of the data point, and N ε(G'(y,a)) represents the number of clustered neighborhood points of the data point to be analyzed. It is understandable that the value of the minimum sample threshold of the data point is not specifically limited in this embodiment, and can be freely set by those skilled in the art as long as it satisfies the analysis of the core data points.
[0120] Specifically, in step S4 of this embodiment, the clustered driving information is used as the user habit classification, the clustered user information is used as the user information classification, and the clustered vehicle configuration information is used as the vehicle classification.
[0121] Specifically, in step S4 described in this embodiment, the matplotlib library and the scikit-learn library in Python can be used to visualize the clustered driving information, user information, and vehicle configuration information.
[0122] Specifically, step S5 described in this embodiment integrates the battery information with the user habit classification, vehicle classification and user information classification into a CSV file, and uses it as sample data. When integrating the CSV file, it is integrated in the order of battery brand ID, vehicle series ID, battery model ID, age, user habit classification, vehicle classification, user information classification and measured SOH.
[0123] See also Figure 4 As shown, it is the method for processing sample data in step S6 of this embodiment, including:
[0124] Step S601, performing data cleaning on sample data;
[0125] Step S602, normalizing the sample data after data cleaning;
[0126] Step S603, dividing the normalized sample data into training data and test data.
[0127] Specifically, in step S601 of this embodiment, data cleaning is performed on the default values in the sample data, wherein:
[0128] When the battery brand ID, vehicle ID, battery model ID and measured SOH in the sample data are default values, delete the current analysis sample data;
[0129] When the user habit classification, vehicle classification and user information classification are default values, the default value in the current sample data is set to the data with the highest weight in each classification;
[0130] When the useful life is the default value, no data cleaning is performed on the useful life.
[0131] Specifically, in step S602 of this embodiment, the measured SOH in the sample data after data cleaning is normalized by a second normalization formula, and the second normalization formula is as follows:
[0132] S1=(SS min ) / (S max -S min );
[0133] Where S1 represents the measured SOH after normalization, S represents the measured SOH, and S min Indicates the minimum value of the measured SOH, S max Indicates the maximum value among the measured SOH.
[0134] Specifically, in step S603 of this embodiment, the normalized sample data is processed in the order of data in the CSV file, and the first α data is used as training data, and the last 1-α data is used as test data, where α represents the division threshold, 30%≤α≤75%. It can be understood that the value of the division threshold is not specifically limited in this embodiment, and those skilled in the art can freely set it as long as the division of the sample data is satisfied. The division threshold can be set according to the number of data in the sample data. When the sample data is large, the division threshold can be appropriately increased for model training. In this embodiment, the optimal value of the division threshold is: α=70%.
[0135] Specifically, in step S7 of this embodiment, the training data is input into the LSTM model for training, and the parameters in the LSTM model are set as follows:
[0136] Batch_size=128, dropout=0.8, lr=0.001, and sigmoid function is used as activation function. The LSTM model can be established through the PyTorch library. Batch_size, dropout, and lr are the data required to establish the LSTM model. Inputting the training data into the LSTM model can automatically obtain other unspecified data. For example, when the training data is input, the feature dimension of the LSTM can be obtained as the number of columns of the training data CSV minus 1. The Batch_size represents the number of samples input into the model at one time, the dropout represents the probability of dropout, and the lr represents the learning rate, which is used for model training to determine the update amplitude data.
[0137] Specifically, step S8 in this embodiment respectively obtains the predicted SOH of the three dimensions of user habit classification, vehicle classification and user information classification in the training data, and sets the predicted SOH to s1, s2, s3, where s1 represents the predicted SOH of user habit classification, s2 represents the predicted SOH of vehicle classification, and s3 represents the predicted SOH of user information classification, and adjusts the neighborhood radius and the minimum sample threshold of data points of each classification according to the predicted SOH and the actual SOH, where:
[0138] When |s1 / S1-1|≤β, the predicted SOH of user habit classification is determined to be accurate, and the neighborhood radius of driving information and the minimum sample threshold of data points are not adjusted;
[0139] When |s1 / S1-1|>β, the predicted SOH of user habit classification is determined to be inaccurate, and the neighborhood radius of driving information and the minimum sample threshold of data points are adjusted;
[0140] When |s2 / S1-1|≤β, the predicted SOH of vehicle classification is determined to be accurate, and the neighborhood radius of vehicle configuration information and the minimum sample threshold of data points are not adjusted;
[0141] When |s2 / S1-1|>β, the predicted SOH of vehicle classification is determined to be inaccurate, and the neighborhood radius of vehicle configuration information and the minimum sample threshold of data points are adjusted;
[0142] When |s3 / S1-1|≤β, the predicted SOH of the user information classification is determined to be accurate, and the neighborhood radius of the user information and the minimum sample threshold of the data point are not adjusted;
[0143] When |s3 / S1-1|>β, the predicted SOH of user information classification is determined to be inaccurate, and the neighborhood radius of user information and the minimum sample threshold of data points are adjusted;
[0144] Wherein, β represents the comparison threshold, 0.03≤β≤0.07. It can be understood that the value of the comparison threshold is not specifically limited in this embodiment, and those skilled in the art can freely set it as long as it satisfies the analysis of the accuracy of the predicted SOH. The optimal value of the comparison threshold is: β=0.05.
[0145] Specifically, in step S8 of this embodiment, when adjusting the neighborhood radius and the minimum sample threshold of data points, the neighborhood radius and the minimum sample threshold of data points can be adjusted according to the ratio of the predicted SOH to the actual SOH. The adjusted neighborhood radius is ε1, and ε1 is set to ε×H. The adjusted minimum sample threshold of data points is MinPTs'(y), and MinPTs'(y) is set to MinPTs(y)×H, where H represents the predicted SOH and the actual SOH. It can be understood that the adjustment process of the neighborhood radius and the minimum sample threshold of data points is not specifically limited in this embodiment, and those skilled in the art can freely set it, such as adjusting the neighborhood radius and the minimum sample threshold of data points according to the variance of each value in the data, or adjusting the neighborhood radius and the minimum sample threshold of data points by manual judgment, etc.
[0146] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A SOH prediction method for new energy batteries based on user driving habits, characterized in that: include: Step S1, obtaining driving information, user information and vehicle configuration information; Step S2, standardizing the driving information, user information and vehicle configuration information, and performing data cleaning on the standardized driving information; Step S3, normalizing the cleaned driving information, standardized user information and vehicle configuration information; Step S4, performing unsupervised clustering on the normalized driving information, user information and vehicle configuration information to obtain user habit classification, vehicle classification and user information classification; Step S5, obtaining battery information, and analyzing sample data according to the battery information and user habit classification, vehicle classification, and user information classification; Step S6, performing data cleaning and normalization processing on the sample data, and dividing the sample data into training data and test data; Step S7, constructing an LSTM model according to the training data; Step S8, testing the LSTM model according to the test data; The step S4 uses the DBScan algorithm to perform unsupervised clustering on the normalized driving information, user information and vehicle configuration information; In step S4, the number of other data points in the ε neighborhood of each data point in the normalized driving information, user information and vehicle configuration information is counted as the number of clustering neighborhood points, N ε (G(y, a)) = {G(y, b)∈D|dist(G(y, a), G(y, b))≤ε}, where G(y, a) represents the data in the yth column and the ath row of the normalized driving information, user information, and vehicle configuration information, G(y, b) represents the data in the yth column and the bth row of the normalized driving information, user information, and vehicle configuration information, y represents the data column number, and y∈N + , a represents the first data row number, a∈N + , b represents the second data flight number, b∈N + , a≠b, ε represents the neighborhood radius, dist(G(y,a),G(y,b)) represents the distance between the data in the yth column and the ath row and the data in the yth column and the bth row, N ε (G(y,a)) represents the number of clustering neighborhood points, and D represents the normalized driving information, user information, and vehicle configuration information; In step S4, when N ε When (G(y,a))≥MinPTs(y), the data points corresponding to the number of neighborhood points of the current analysis cluster are taken as core data points, the core data points are added to the cluster V, and the other data points in the neighborhood of the core data point ε are taken as data points to be analyzed, and the data points to be analyzed are set to G'(y,a). When N ε When (G'(y,a))≥MinPTs(y), the data point to be analyzed is set as visited, the data point to be analyzed is taken as the core data point, and added to the cluster V. ε When (G'(y,a))<MinPTs(y), the data point to be analyzed is set as visited, and the analysis of other data points in the neighborhood of the core data point ε is repeated until all the data points to be analyzed are set as visited, where V represents the core data point cluster set, MinPTs(y) represents the minimum sample threshold of the data point, and N ε (G'(y,a)) represents the number of cluster neighborhood points of the data point to be analyzed; In the step S4, the clustered driving information is used as the user habit classification, the clustered user information is used as the user information classification, and the clustered vehicle configuration information is used as the vehicle classification.
2. The SOH prediction method of new energy batteries based on user driving habits according to claim 1 is characterized in that: The data standardization and data cleaning method in step S2 includes: Step S201: standardize driving information, user information and vehicle configuration information. Step S202, extracting unique identification data from the standardized driving information, user information and vehicle configuration information; Step S203, performing missing value processing on the standardized driving information, user information and vehicle configuration information according to the unique identification data; Step S204, filtering abnormal data according to the standardized driving information; Step S205: performing outlier processing on the standardized driving information according to the abnormal data.
3. The SOH prediction method of new energy batteries based on user driving habits according to claim 2 is characterized in that: The step S201 uses oneHot encoding to encode the optional data in the driving information, user information and vehicle configuration information, and uses numerical values to represent each data item; The step S201 outputs the driving information in the order of time, vehicle speed, maximum vehicle speed, average vehicle speed, average engine speed, maximum engine speed, average brake pedal position, average accelerator pedal position, total mileage, average single mileage, number of accelerations, average number of accelerations per kilometer, number of brakes, average number of brakes per kilometer, idling time, average idling time each time, number of speeding, average number of speeding per kilometer, energy recovery, average energy recovery per kilometer, average battery voltage, average battery current, average SOC, number of fast charges, number of slow charges, charging time, charging amount, charging efficiency, average battery temperature, ambient temperature, driving mode, road type and SOH value. CSV file, output user information in the order of user ID, user city, user age, user gender, user frame number, user occupation, user driving experience, charging frequency, charging preference, average charging time, and charging location as a CSV file, and output vehicle configuration information in the order of vehicle frame number, vehicle brand, vehicle model, body type, production year, drive type, motor type, motor power, maximum torque, transmission system, battery type, battery capacity, battery voltage, number of battery modules, battery cooling system power, acceleration time per 100 kilometers, vehicle maximum speed, charging interface type, vehicle charging efficiency, braking system type, and autonomous driving type as a CSV file; The step S202 extracts the vehicle frame number from the standardized driving information as unique identification data, extracts the user ID and the user vehicle frame number from the standardized user information as unique identification data, and extracts the vehicle frame number from the standardized vehicle configuration information as unique identification data; The step S203 performs missing value processing on the standardized driving information, user information and vehicle configuration information according to the unique identification data, wherein: When the unique identification data is missing, the standardized driving information, user information and vehicle configuration information corresponding to the currently analyzed unique identification data are deleted; When the unique identification data is not a missing value, the standardized driving information, user information and vehicle configuration information corresponding to the currently analyzed unique identification data are retained; The step S204 uses the DBScan algorithm to filter abnormal data according to the standardized driving information; The step S205 performs abnormal value processing on the standardized driving information according to the abnormal data, and deletes the abnormal data in the standardized driving information.
4. The SOH prediction method of new energy batteries based on user driving habits according to claim 1 is characterized in that: The normalization processing method in step S3 includes: Step S301, extracting quantified data from the cleaned driving information, standardized user information and vehicle configuration information; Step S302, normalizing the quantized data; Step S303, retaining the original values of the driving information, user information and vehicle configuration information that are not quantized data.
5. The SOH prediction method of new energy batteries based on user driving habits according to claim 4 is characterized in that: The step S301 extracts the average vehicle speed, maximum vehicle speed, average engine speed, maximum engine speed, average brake pedal position, average accelerator pedal position, total mileage, average single mileage, number of accelerations, average number of accelerations per kilometer, number of brakes, average number of brakes per kilometer, idling time, average idling time each time, number of overspeeding, average number of overspeeding per kilometer, energy recovery, average energy recovery per kilometer, average battery voltage, average battery current, average SOC, number of fast charges, number of slow charges, charging time, charging amount, charging efficiency, average battery temperature, ambient temperature and SOH value from the driving information after data cleaning as quantitative data, extracts the charging frequency and average charging time from the standardized user information as quantitative data, and extracts the motor power, maximum torque, battery capacity, battery voltage, number of battery modules, acceleration time per 100 kilometers, vehicle maximum speed and vehicle charging efficiency from the standardized vehicle configuration information as quantitative data; The step S302 performs normalization processing on the quantized data using a first normalization formula, and the first normalization formula is as follows: X v =(x v -x vmin ) / (x vmax -x vmin ); Among them, X v Represents the normalized quantized data, v represents the quantized data number, v∈N + , x v Represents quantitative data, x vmin Represents the minimum value of the quantized data, x vmax Indicates the maximum value of the quantized data.
6. The SOH prediction method of new energy batteries based on user driving habits according to claim 1 is characterized in that: The step S5 integrates the battery information with the user habit classification, vehicle classification and user information classification into a CSV file, and uses it as sample data. When integrating the CSV file, the integration is performed in the order of battery brand ID, vehicle series ID, battery model ID, age, user habit classification, vehicle classification, user information classification and measured SOH.
7. The SOH prediction method of new energy batteries based on user driving habits according to claim 1, characterized in that: The method for processing the sample data in step S6 includes: Step S601, performing data cleaning on sample data; Step S602, normalizing the sample data after data cleaning; Step S603, dividing the normalized sample data into training data and test data.
8. The SOH prediction method of new energy batteries based on user driving habits according to claim 7 is characterized in that: In step S601, data cleaning is performed on the default values in the sample data, wherein: When the battery brand ID, vehicle ID, battery model ID and measured SOH in the sample data are default values, delete the current analysis sample data; When the user habit classification, vehicle classification and user information classification are default values, the default value in the current sample data is set to the data with the highest weight in each classification; When the useful life is the default value, no data cleaning is performed on the useful life; The step S602 normalizes the measured SOH in the sample data after data cleaning by using a second normalization formula, and the second normalization formula is as follows: S1=(SS min ) / (S max -S min ); Where S1 represents the measured SOH after normalization, S represents the measured SOH, and S min Indicates the minimum value of the measured SOH, S max It represents the maximum value among the measured SOH; In step S603, the normalized sample data is processed in the order of data in the CSV file, with the first α data as training data and the last 1-α data as test data, where α represents the division threshold.
9. The SOH prediction method of new energy batteries based on user driving habits according to claim 8, characterized in that: In step S7, the training data is input into the LSTM model for training. The parameters in the LSTM model are set as follows: Batch_size = 128, dropout = 0.8, lr = 0.001, use sigmoid function as activation function; The step S8 obtains the predicted SOH of the three dimensions of user habit classification, vehicle classification and user information classification in the training data respectively, and sets the predicted SOH to s1, s2, s3, where s1 represents the predicted SOH of user habit classification, s2 represents the predicted SOH of vehicle classification, and s3 represents the predicted SOH of user information classification, and adjusts the neighborhood radius and the minimum sample threshold of data points of each classification according to the predicted SOH and the actual SOH, where: When |s1 / S1-1|≤β, the predicted SOH of user habit classification is determined to be accurate, and the neighborhood radius of driving information and the minimum sample threshold of data points are not adjusted; When |s1 / S1-1|>β, the predicted SOH of user habit classification is determined to be inaccurate, and the neighborhood radius of driving information and the minimum sample threshold of data points are adjusted; When |s2 / S1-1|≤β, the predicted SOH of vehicle classification is determined to be accurate, and the neighborhood radius of vehicle configuration information and the minimum sample threshold of data points are not adjusted; When |s2 / S1-1|>β, the predicted SOH of vehicle classification is determined to be inaccurate, and the neighborhood radius of vehicle configuration information and the minimum sample threshold of data points are adjusted; When |s3 / S1-1|≤β, the predicted SOH of the user information classification is determined to be accurate, and the neighborhood radius of the user information and the minimum sample threshold of the data point are not adjusted; When |s3 / S1-1|>β, the predicted SOH of user information classification is determined to be inaccurate, and the neighborhood radius of user information and the minimum sample threshold of data points are adjusted; Among them, β represents the comparison threshold.
Citation Information
Patent Citations
Method and device for predicting SOH (state of health) of battery
CN116243197A
Power battery remaining life prediction method based on data driving
CN112765772A
All-weather and multi-region power battery pack SOH (state of health) prediction method and system based on vehicle cloud collaboration
CN117743802A