A training method, device, and electronic device for a model based on vehicle data
Through multiple trainings combining the first network and the second network, data with low marking noise are screened out, which solves the problem of insufficient model accuracy in the prior art, and achieves high accuracy model training under complex operating conditions.
Patent Information
- Application Number
- CN202210176800.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-02-24
AI Technical Summary
In the prior art, the model training method based on experimental data sets is insufficient in model accuracy due to low operating conditions coverage, making it difficult to effectively apply to complex real-time operating conditions.
Using a multiple training method, the first network and the second network are combined, and the training data with less label noise is screened out, and the real vehicle data with more label noise is combined for model training, including multiple trainings and data fusion for the first network and the second network.
It improves the application accuracy of the model in actual scenarios, effectively utilizes real-vehicle data with high marking noise, supplements experimental data with low marking noise, and improves the overall accuracy of the model.
Smart Images

Figure CN114707631B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicles, and in particular to a training method, device and electronic equipment for a model based on vehicle data. Background Art
[0002] With the development of technology, vehicles have gradually become a part of people's daily lives. The battery management system (BMS) is an important component of vehicles, and the data it returns has attracted people's attention.
[0003] Due to the complexity of operating conditions, real-world vehicle data (data from real vehicles) often contains significant noise, making it difficult to use as accurate calibration values. This results in a significant amount of real-world data being unusable in subsequent analysis. Currently, research on BMS feedback data is typically based on experimental datasets. However, these datasets often contain experimental data with limited operating condition coverage, resulting in low model accuracy.
[0004] Therefore, there is an urgent need for a model training method to improve the accuracy of the model. Summary of the Invention
[0005] In view of this, the present application provides a training method, device and electronic device for a model based on vehicle data to improve the accuracy of the model.
[0006] In a first aspect, the present application provides a method for training a model based on vehicle data. The model includes a first network and a second network. The method includes training the model N times, where N = 2, 3, ..., wherein the i-th training of the model, where i = 1, 2, ..., N, includes:
[0007] The first network is trained using the i-th training data; the second network is trained using the j-th training data to obtain the j-th target data; wherein the j-th target data includes data in the j-th training data for which the loss function of the second network is less than the j-th preset value, i=j=1,2,…,N; wherein:
[0008] When i=j=1, the i-th training data includes the first data set, and the j-th training data includes the second data set, wherein the label noise of the first data set is less than the label noise of the second data set;
[0009] When i>1 and j>1, the i-th training data includes the first data set and the (j-1)-th target data, and the j-th training data includes the (j-1)-th training data.
[0010] In one possible implementation, the second network includes a first subnetwork and a second subnetwork;
[0011] The second network is trained using the j-th training data to obtain the j-th target data; wherein the j-th target data includes data in the j-th training data for which the loss function of the second network is less than the j-th preset value, i=j=1,2,…,N, including:
[0012] The first sub-network is trained using the j-th training data of the first sub-network to obtain the j-th data of the first sub-network; the j-th data of the first sub-network includes data in the j-th training data of the first sub-network where the loss function of the first sub-network is less than the j-th threshold of the first sub-network;
[0013] The second sub-network is trained using the j-th training data of the second sub-network to obtain the j-th data of the second sub-network; the j-th data of the second sub-network includes data in the j-th training data of the second sub-network where the loss function of the second sub-network is less than the j-th threshold of the second sub-network;
[0014] Fuse the j-th data of the first sub-network and the j-th data of the second sub-network to obtain the j-th target data;
[0015] Wherein, when j=1, the j-th training data of the first sub-network includes the third data set, the j-th training data of the second sub-network includes the fourth data set, and the second data set includes the third data set and the fourth data set;
[0016] When j>1, the j-th training data of the first sub-network includes the (j-1)-th training data of the second sub-network, and the j-th training data of the second sub-network includes the (j-1)-th training data of the first sub-network.
[0017] In a possible implementation, the first sub-network and the second sub-network are the same.
[0018] In a possible implementation, the first network includes transform.
[0019] In a possible implementation, the first sub-network includes an RNN, and the second sub-network includes an RNN.
[0020] In a possible implementation, the (j-1)th preset value is smaller than the jth preset value, where j=2,…,N.
[0021] In a second aspect, the present application provides a training device for a model based on vehicle data, wherein the model includes a first network and a second network, and the device is used to train the model N times, where N=2, 3, ... The device includes a first training unit and a second training unit.
[0022] The device is used to train the model for the i-th time, where i=1, 2, ..., N, and specifically includes:
[0023] The first training unit is used to train the first network using the i-th training data; the second training unit is used to train the second network using the j-th training data to obtain the j-th target data; wherein the j-th target data includes data in the j-th training data where the loss function of the second network is less than the j-th preset value, i=j=1, 2, ..., N;
[0024] Wherein, when i=j=1, the i-th training data includes the first data set, and the j-th training data includes the second data set, wherein the label noise of the first data set is less than the label noise of the second data set;
[0025] When i>1 and j>1, the i-th training data includes the first data set and the (j-1)-th target data, and the j-th training data includes the (j-1)-th training data.
[0026] In one possible implementation, the second network includes a first subnetwork and a second subnetwork;
[0027] The second training unit includes a first sub-training unit, a second sub-training unit and a data fusion unit, wherein:
[0028] a first sub-training unit, configured to train the first sub-network using the j-th training data of the first sub-network to obtain the j-th data of the first sub-network; the j-th data of the first sub-network including data in the j-th training data of the first sub-network where the loss function of the first sub-network is less than the j-th threshold of the first sub-network;
[0029] The second sub-training unit is configured to train the second sub-network using the j-th training data of the second sub-network to obtain the j-th data of the second sub-network; the j-th data of the second sub-network includes data in the j-th training data of the second sub-network where the loss function of the second sub-network is less than the j-th threshold of the second sub-network;
[0030] Wherein, when j=1, the j-th training data of the first sub-network includes the third data set, the j-th training data of the second sub-network includes the fourth data set, and the second data set includes the third data set and the fourth data set;
[0031] When j>1, the j-th training data of the first sub-network includes the (j-1)-th training data of the second sub-network, and the j-th training data of the second sub-network includes the (j-1)-th training data of the first sub-network;
[0032] The data fusion unit is used to fuse the j-th data of the first sub-network and the j-th data of the second sub-network to obtain the j-th target data.
[0033] In a possible implementation, the first sub-network and the second sub-network are the same.
[0034] In a third aspect, the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores code, and the processor is used to call the code stored in the memory to execute any of the above methods.
[0035] In a fourth aspect, the present application provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute any of the above methods.
[0036] Using the technical solution of this application, a second network is used to filter out training data with less labeling noise. Since the first dataset itself is training data with less labeling noise, the second network is used to filter out this training data, and the first dataset is used to train the first network. Using the technical solution of this embodiment, training data with more labeling noise (such as real vehicle data) is more effectively utilized to supplement training data with less labeling noise (such as experimental data, which covers fewer vehicle conditions). This allows the trained model to be better applied in real-world scenarios, thereby improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a flowchart of a method for training a model based on vehicle data provided in an embodiment of the present application;
[0038] Figure 2 A schematic diagram of training data for the model provided in an embodiment of the present application;
[0039] Figure 3 A schematic diagram of the structure of a training device for a vehicle data-based model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] Due to the complex operating conditions, real-world vehicle data (data from real vehicles) often contains a significant amount of noise, making it difficult to use as accurate calibration values. This results in a significant amount of real-world vehicle data being unusable in subsequent analysis. Currently, research on BMS feedback data is typically based on experimental datasets. However, the experimental data contained in these datasets has low coverage of operating conditions, resulting in low model accuracy. Therefore, a model training method is urgently needed to improve model accuracy.
[0041] Based on this, in an embodiment of the present application provided by the applicant, a trained model includes a first network and a second network. The training method includes training the model N times, where N = 2, 3…. The i-th training of the model, where i = 1, 2, …, N, includes the following process: training the first network using the i-th training data; training the second network using the j-th training data to obtain the j-th target data; wherein the j-th target data includes data in the j-th training data where the loss function of the second network is less than the j-th preset value, i = j = 1, 2, …, N; wherein: when i = j = 1, the i-th training data includes the first data set, and the j-th training data includes the second data set, wherein the label noise of the first data set is less than the label noise of the second data set; when i>1 and j>1, the i-th training data includes the first data set and the (j-1)-th target data.
[0042] The second network is used to filter out training data with less labeling noise. Since the first dataset itself is training data with less labeling noise, the second network is used to filter out this training data, and the first dataset is used to train the first network. The technical solution of this embodiment effectively utilizes training data with more labeling noise (e.g., real-world vehicle data) to supplement training data with less labeling noise (e.g., experimental data, which covers fewer vehicle conditions). This allows the trained model to be better applied in real-world scenarios, thereby improving the model's accuracy.
[0043] In order to facilitate understanding of the technical solutions provided by the embodiments of the present application, common application scenarios of the embodiments of the present application are first introduced.
[0044] Due to the complex working conditions, real vehicle data (data from real vehicles) usually contains a lot of noise, making it difficult to use as an accurate calibration value, resulting in a large amount of real vehicle data being unable to be used in subsequent analysis.
[0045] Examples of data provided by the BMS include: Vin number (vehicle frame number, which is a national standard), time, voltage, current, temperature, speed, acceleration, SOC, SOH… An explanation for data noise: The success of machine learning and deep neural networks relies on high-quality labeled training data. Labeling errors (noisy labels) in training data can significantly reduce the model's accuracy on clean test data. Unfortunately, large datasets almost always contain incorrect or inaccurate labels.
[0046] Machine learning, deep learning, and similar AI data model algorithms require relatively accurate labels during the training and accuracy evaluation process. If the label accuracy is low, the model's prediction results will also be inaccurate. For example, for battery capacity SOH estimation, if the battery capacity label accuracy is low—that is, if the SOH label contains errors—the battery capacity estimation model will not be able to produce accurate and convincing results.
[0047] The following is a brief introduction to the battery capacity estimation task.
[0048] In recent years, with the rapid development of electric vehicles, more and more research has been conducted on the aging mechanism and degradation law of lithium-ion batteries. Many methods and research results on SOH estimation have been published, such as offline testing method, electrochemical impedance spectroscopy analysis method, neural network estimation method, fuzzy logic estimation method, particle filter estimation method, life model estimation method, etc.
[0049] Capacity estimation involves calculating and estimating SOH. Battery aging is a gradual and complex process. Despite this, we still hope to find some quantifiable indicators to describe the degree of battery degradation. The principles for selecting such indicators are twofold: first, the indicator must be representative and reflect the degree of battery aging; second, it must be operational. For example, while the remaining number of cycles of a battery can reflect the degree of battery aging, using this number to define the degree of aging is not practical.
[0050] To facilitate understanding of the technical solution provided in the embodiments of the present application, a training method, device, and electronic device for a vehicle data-based model provided in the embodiments of the present application are described below with reference to the accompanying drawings.
[0051] Although the accompanying drawings show exemplary embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Based on the embodiments of the present application, other embodiments obtained by those skilled in the art without making any creative contribution shall fall within the scope of protection of the present application.
[0052] In the claims and description and drawings of this application, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions.
[0053] This application provides a method for training a model based on vehicle data.
[0054] See also Figure 1 , Figure 1 This is a flowchart of a vehicle data-based model training method provided in an embodiment of the present application.
[0055] In this embodiment, the model used for training includes a first network and a second network.
[0056] The vehicle data-based model training method in this embodiment trains the model multiple times, that is, the model is trained N times, where N=2, 3...
[0057] like Figure 1 As shown, Figure 1 The process of the i-th training is shown, i=1,2,…,N.
[0058] The training method of the vehicle data-based model in the embodiment of the present application includes S101-S102.
[0059] First, the first network is trained using the i-th training data; the second network is trained using the j-th training data to obtain the j-th target data; wherein the j-th target data includes the data in the j-th training data where the loss function of the second network is less than the j-th preset value, i=j=1,2,…,N.
[0060] When i=j=1, that is, the first training of the model, including S101.
[0061] S101. Train a first network using a first data set; train a second network using the jth training data to obtain jth target data, where the jth training data is the second data set, j=1; and the label noise of the first data set is less than the label noise of the second data set.
[0062] The first data set and the second data set both include training data, which are vehicle data and labels corresponding to the vehicle data.
[0063] Training data may contain labeling errors, also known as label noise. Label noise can reduce the accuracy of the model on clean test data. Large datasets often contain incorrect or inaccurate labels. For example, real-world vehicle data corresponds to complex operating conditions. Using real-world vehicle data as training data may result in lower model accuracy.
[0064] The first network and the second network each have their own loss functions.
[0065] The j-th target data may be part or all of the data in the second data set; for the training data in the second data set, the training data for which the loss function of the second network is less than the first preset value is the first target data.
[0066] In this embodiment, the training data used in the first training and subsequent rounds of training are usually different.
[0067] When i>1, that is, for the i-th training of the model, i=2, 3,…, N, including S102.
[0068] S102. Train the first network using the first data set and the (j-1)th target data; train the second network using the (j-1)th training data to obtain the jth target data; the jth target data includes data in the jth training data where the loss function of the second network is less than the jth preset value, j = 2,…,N.
[0069] Obtain the j-th target data, that is, in the training data used to train the second network, filter out the training data with less labeling noise (according to the loss function of the second sub-network).
[0070] The second data set used for the first training of the second network has relatively large label noise. However, in the second data set, there is also data with relatively small label noise. The above screening process also obtains this part of data with relatively small label noise.
[0071] The second network is used to filter out training data with less labeling noise. Since the first dataset itself is training data with less labeling noise, the second network is used to filter out this training data, and the first dataset is used to train the first network. The technical solution of this embodiment effectively utilizes training data with more labeling noise (e.g., real-vehicle data) to supplement training data with less labeling noise (e.g., experimental data, which covers fewer vehicle conditions). This allows a large amount of real-vehicle data to be used in analysis and model training, allowing the trained model to be better applied in real-world scenarios and improving model accuracy.
[0072] The following describes the specific implementation method.
[0073] The embodiment of the present application also provides a method for training a model based on vehicle data.
[0074] Taking the battery capacity estimation task as an example, the input of the model in the embodiment of the present application is vehicle data, and the output of the model is the estimated result of the battery capacity.
[0075] The model includes a first network and a second network, and the second network includes a first sub-network and a second sub-network.
[0076] The first sub-network and the second sub-network are both RNNs.
[0077] It can be understood that the model is divided into two layers, the first layer is the first network, and the second layer consists of two symmetrical networks.
[0078] The following describes the training of the model in the embodiment of the present application. The training method of the model based on vehicle data in this embodiment includes S201-S210.
[0079] S201: Acquire a first data set and a second data set.
[0080] The training data is used to train the network, the first data set is used to train the first network, and the second data set is used to train the second network.
[0081] The first data set includes: first vehicle data and a label of the first vehicle data; the second data set includes: second vehicle data and a label of the second vehicle data.
[0082] Taking the battery capacity estimation task as an example, the label of the first vehicle data and the label of the second vehicle data may be battery capacity.
[0083] The labeling noise of the first dataset is smaller than that of the second dataset.
[0084] Training data may contain labeling errors, also known as label noise. Label noise can reduce the accuracy of the model on clean test data. Large datasets often contain incorrect or inaccurate labels. For example, real-world vehicle data corresponds to complex operating conditions. Using real-world vehicle data as training data may result in lower model accuracy.
[0085] In a possible implementation, the first data set may be experimental data, and the second data set may be real vehicle data.
[0086] The real vehicle data corresponds to more complex working conditions and has larger labeling noise; the experimental data has lower working condition coverage and smaller labeling noise.
[0087] Taking the SOH estimation task as an example, the labeling noise of real vehicle data is large, which means that the mean and variance of the noise are large in the SOH distribution of the entire real vehicle data.
[0088] The labeling noise of the first dataset is relatively small, which can also be understood as the accuracy of the first dataset is relatively high; the labeling noise of the second dataset is relatively large, which can also be understood as the accuracy of the second dataset is relatively low.
[0089] The second network includes a first subnetwork and a second subnetwork. Therefore, the second data set can be used to train the first subnetwork and the second subnetwork respectively; or, the second data set can include two parts, and the two parts can be used to train the first subnetwork and the second subnetwork respectively.
[0090] The first network and the second network are trained using the first data set and the second data set, respectively. The following describes the model training process using three training rounds for the first network, three training rounds for the first sub-network, and three training rounds for the second sub-network as an example.
[0091] It is understandable that the above training rounds may also have other values.
[0092] S202: Perform a first training on the first network using the first data set.
[0093] The labeling noise of the first dataset is relatively small.
[0094] Compared to the second network, the first network may be a more complex network. In one possible implementation, the first network is a transform.
[0095] Because the first dataset has less labeling noise, the second network can use a network with higher information extraction capabilities. This means using a more complex network to train data with less labeling noise with higher accuracy. For deep neural networks, using a more complex network is generally more likely to produce better results when the training data is of high quality.
[0096] See also Figure 2 , Figure 2 A schematic diagram of the training data of the model provided in the embodiment of the present application.
[0097] There are three training processes from top to bottom.
[0098] The label noise of the first dataset is small. Figure 2 The first data set is shown as low-noise data set I.
[0099] S203: Perform a first training on the first sub-network using the third data set to obtain first target training data.
[0100] The second dataset includes a third dataset and a fourth dataset. Since the labeling noise of the second dataset is greater than that of the first dataset, the labeling noise of the third dataset and the fourth dataset is greater than that of the first dataset.
[0101] The second network includes a first sub-network and a second sub-network.
[0102] The second data set is used to train the second network. In this embodiment, the third data set is used to train the first sub-network, and the fourth data set is used to train the second sub-network.
[0103] In a possible implementation, the second dataset may be used to train the first sub-network and the second sub-network respectively.
[0104] In the third data set, training data in which the loss function of the first sub-network is less than a first preset value is determined to obtain first target training data.
[0105] The first target training data includes vehicle data and labels of the vehicle data.
[0106] The first target training data, that is, the training data with less labeled noise in the second data set (based on the loss function of the first sub-network).
[0107] S204: Perform a first training on the second sub-network using the fourth data set to obtain second target training data.
[0108] In the fourth data set, training data in which the loss function of the second sub-network is less than a second preset value is determined to obtain second target training data.
[0109] The second target training data includes vehicle data and labels of the vehicle data.
[0110] The second target training data, that is, the training data with less labeled noise in the second dataset (based on the loss function of the second sub-network).
[0111] S202-S203, that is, in the process of training the first sub-network and the second sub-network, obtain training data with less labeling noise in the training data (the second data set, that is, the third data set and the fourth data set).
[0112] The labeling noise of the second dataset is relatively large, that is, the difference between the labels of the vehicle data and the accurate labels is relatively large. However, in the second dataset, there are also data where the difference between the labels of the vehicle data and the accurate labels is relatively small.
[0113] The first target training data and the second target training data mentioned above can be understood as training data with high confidence (which can be called confidence data), that is, training data with small labeling noise.
[0114] The above process of obtaining the first target data and the second target data can be understood as a process of screening and obtaining confidence data from the second data set.
[0115] The first data set mentioned above is training data with less labeling noise, which can be understood as confidence data itself.
[0116] This embodiment does not limit the order of S202 - S204 .
[0117] In a possible implementation, the first sub-network is an RNN, and the second sub-network is an RNN.
[0118] Since the labeling noise of the second data set is greater than that of the first data set, for example, the second data set is real vehicle data.
[0119] Compared with the first network, the first sub-network and the second sub-network may have simpler structures.
[0120] For deep neural networks, using more complex networks is generally more likely to yield better results when the training data quality is high. When using real-world vehicle data with high labeling noise, a simpler network is used to ensure that a rough, but at least accurate, amount of information is extracted. This is because neural network algorithms prioritize fitting easy-to-fit samples when optimizing model parameters. If a complex structure is used in the second network, the network will increasingly learn from noise with varying amounts of information, making the extracted information essentially unusable and reducing the model's estimated accuracy. Simple network models consider samples with low individual loss to be high-quality samples in a noisy dataset.
[0121] like Figure 2 As shown in , since the label noise of the second data set is large, the second data set includes the third data set and the fourth data set. Figure 2 The third data set is shown as a high-noise data set I, which is used to train the first sub-network; the fourth data set is shown as a high-noise data set II, which is used to train the second sub-network.
[0122] S205 : Perform a second training on the first network using the first data set, the first target training data, and the second target training data.
[0123] The first data set, the first target training data, and the second target training data are confidence data.
[0124] S205 is to train the first network using the confidence data.
[0125] The first data set is experimental data, with less labeling noise and corresponding to fewer working conditions; the first target training data and the second target training data are both real vehicle data, with less labeling noise and corresponding to more working conditions.
[0126] In the process of training the first network, the training data includes experimental data with less labeling noise and real vehicle data with higher confidence (less labeling noise).
[0127] like Figure 2 As shown, the first target training data and the second target training data are added to the data used to train the first network (i.e., from Figure 2 The dotted box A is added to form the low-noise dataset II). That is, for the second training of the first network, the training data is the low-noise dataset II.
[0128] S206 : Perform a second training on the first sub-network using the fourth data set to obtain third target training data.
[0129] S206 is to train the first sub-network using the confidence data.
[0130] In the fourth data set, training data in which the loss function of the first sub-network is less than a third preset value is determined to obtain third target training data.
[0131] The obtained third target training data can be understood as confidence data.
[0132] The description of the third target training data is similar to the description of the first target training data in S203 and will not be repeated here.
[0133] like Figure 2 As shown, in Figure 2 The fourth data set is shown as the high noise data set II, which is used for the second training of the first sub-network.
[0134] S207: Perform a second training on the second sub-network using the third data set.
[0135] S207 is to train the second sub-network using the confidence data.
[0136] In the third data set, training data in which the loss function of the second sub-network is less than a fourth preset value is determined to obtain fourth target training data.
[0137] The obtained fourth target training data can be understood as confidence data.
[0138] The description of the fourth target training data is similar to the description of the second target training data in S204 and will not be repeated here.
[0139] like Figure 2 As shown, in Figure 2 The third data set is shown as a high-noise data set I, which is used for the second training of the second sub-network.
[0140] The second network includes a first sub-network and a second sub-network, and the second network is an interactive symmetrical structure.
[0141] This embodiment does not limit the order of S205 - S207 .
[0142] S208: Perform a third training on the first network using the first data set, the third target training data, and the fourth target training data.
[0143] The first data set, the third target training data, and the fourth target training data are confidence data.
[0144] S208 is to train the first network using the confidence data.
[0145] For the description of S208, please refer to the description of S205, which will not be repeated here.
[0146] like Figure 2As shown, the third target training data and the fourth target training data are added to the data used to train the first network (i.e., from Figure 2 The dotted box B is added to form the low-noise dataset III). That is, for the third training of the first network, the training data is the low-noise dataset III.
[0147] S209: Perform a third training on the first sub-network using the third data set.
[0148] The fourth target training data is confidence data.
[0149] like Figure 2 As shown, in Figure 2 The third data set is shown as a high-noise data set I, which is used for the third training of the first sub-network.
[0150] S210: Perform a third training on the second sub-network using the fourth data set.
[0151] The third target training data is confidence data.
[0152] like Figure 2 As shown, in Figure 2 The fourth data set is shown as the high noise data set II, which is used for the third training of the second sub-network.
[0153] For the description of S209-S210, please refer to the description of S206-S207, which will not be repeated here.
[0154] This embodiment uses the example of training the first network, the first sub-network, and the second sub-network three times. It is understood that the number of training rounds for the network can be other numbers. For example, when the first network, the first sub-network, and the second sub-network are trained four times, confidence data can be obtained through S209-S210, and the first network can be trained for the fourth time using the first data set and the confidence data obtained through S209-S210.
[0155] During the training of the first sub-network and the second sub-network, the training data is continuously screened to obtain training data with less annotation noise in the real vehicle data.
[0156] During the training of the first sub-network and the second sub-network, an interactive training process is formed.
[0157] After S201 - S210 , the first network, the first sub-network, and the second sub-network that have been trained three times are obtained, that is, the model trained using the first data set and the second data set is obtained.
[0158] The following is an implementation of the loss function of each network in the model provided in this embodiment.
[0159] The loss functions of the first network, the first sub-network, and the second sub-network are all basic MSE (Mean Squared Error). For the entire model, the joint loss function is:
[0160]
[0161] Here, m is the number of high-quality samples filtered out by the first network, n is the number of low-quality samples filtered out by the first network, and N is the number of high-quality samples transferred from the second network to the first network. In other words, N = m + n. High-quality samples are those for which the network's loss function is less than a preset value, while low-quality samples are those for which the network's loss function is greater than or equal to a preset value.
[0162] MSEa is the loss function of the first network, MSEi is the loss function of each subnetwork in the second network, and k is the number of subnetworks in the second network.
[0163] The above three networks control the weights in the overall network through the loss function.
[0164] The high and low quality data evaluated by the first network are weighted, and each sub-network in the second network is weighted equally.
[0165] During the continuous training process, the second network screens the second data set to obtain high-quality training data from the second data set.
[0166] The proportion of training data that is filtered can be reduced over iterations, that is, as the number of model training rounds increases, the confidence level in the loss function is mapped to the filtering ratio. The filtering ratio is inversely proportional to the number of training rounds, and the filtering ratio decreases faster as the number of training rounds approaches the set maximum number of rounds. In other words, as the number of training rounds increases, the amount of data that can be filtered from the training data (high-noise samples) to obtain high-quality samples decreases, and the filtering conditions become more stringent.
[0167] The screening ratio can be expressed as: R(t) = -log(Tt), where T is the total number of rounds and t is the current round number.
[0168] After each iteration (training), the loss function is sorted from low to high, and the training data with the highest R(t) ratio is taken to obtain high-quality data.
[0169] R(t) multiplied by the number of high-noise samples in the current round is N in the joint loss function mentioned above.
[0170] The vehicle data-based model training method provided in this embodiment can, on the one hand, screen out high-quality training data from real vehicle data; on the other hand, in the target task (such as the SOH prediction task), each network obtains a model with higher accuracy through weight matching.
[0171] The experimental data has high confidence, but covers fewer operating conditions; the real vehicle data covers a wider range of operating conditions, but has greater annotation noise. The technical solution of this embodiment adopts a semi-interactive model training method. For the training of the first network, a first data set with higher confidence is used, as well as training data with higher confidence screened by the second network. This method effectively utilizes the data with higher confidence in the real vehicle data, supplements the single data set (experimental data) training model, and solves the problem that the model is difficult to apply in actual scenarios, making the model have higher estimation accuracy.
[0172] In some possible implementations, when the model includes more network results, a technical solution similar to that of this embodiment can also be used to train the model. The principle is similar to that of this embodiment and will not be repeated here.
[0173] An embodiment of the present application also provides a training device for a model based on vehicle data.
[0174] The model includes a first network and a second network. The vehicle data-based model training device of an embodiment of the present application is used to train the model N times, where N=2, 3...
[0175] See also Figure 3 , Figure 3 A schematic diagram of the structure of a training device for a vehicle data-based model provided in an embodiment of the present application.
[0176] like Figure 3 As shown, the apparatus 300 includes a first training unit 301 and a second training unit 302 .
[0177] The apparatus 300 is used to train the model for the i-th time, where i=1, 2, ..., N, and includes:
[0178] The first training unit 301 is used to train the first network using the i-th training data; the second training unit 302 is used to train the second network using the j-th training data to obtain the j-th target data; wherein the j-th target data includes data in the j-th training data where the loss function of the second network is less than the j-th preset value, i=j=1, 2, ..., N;
[0179] Wherein, when i=j=1, the i-th training data includes the first data set, and the j-th training data includes the second data set, wherein the label noise of the first data set is less than the label noise of the second data set;
[0180] When i>1 and j>1, the i-th training data includes the first data set and the (j-1)-th target data, and the j-th training data includes the (j-1)-th training data.
[0181] The units included in the above-mentioned vehicle data-based model training device can achieve the same technical effects as the vehicle data-based model training method in the above embodiment. To avoid repetition, they will not be described here.
[0182] In some possible implementations, the second network includes a first sub-network and a second sub-network.
[0183] The second training unit includes a first sub-training unit, a second sub-training unit and a data fusion unit, wherein:
[0184] a first sub-training unit, configured to train the first sub-network using the j-th training data of the first sub-network to obtain the j-th data of the first sub-network; the j-th data of the first sub-network including data in the j-th training data of the first sub-network where the loss function of the first sub-network is less than the j-th threshold of the first sub-network;
[0185] The second sub-training unit is configured to train the second sub-network using the j-th training data of the second sub-network to obtain the j-th data of the second sub-network; the j-th data of the second sub-network includes data in the j-th training data of the second sub-network where the loss function of the second sub-network is less than the j-th threshold of the second sub-network;
[0186] Wherein, when j=1, the j-th training data of the first sub-network includes the third data set, the j-th training data of the second sub-network includes the fourth data set, and the second data set includes the third data set and the fourth data set;
[0187] When j>1, the j-th training data of the first sub-network includes the (j-1)-th training data of the second sub-network, and the j-th training data of the second sub-network includes the (j-1)-th training data of the first sub-network;
[0188] The data fusion unit is used to fuse the j-th data of the first sub-network and the j-th data of the second sub-network to obtain the j-th target data.
[0189] In some possible implementations, the first sub-network and the second sub-network are the same.
[0190] The units included in the above-mentioned vehicle data-based model training device can achieve the same technical effects as the vehicle data-based model training method in the above embodiment. To avoid repetition, they will not be described here.
[0191] An embodiment of the present application further provides an electronic device comprising a processor and a memory, wherein the memory stores code, and the processor is configured to call the code stored in the memory to execute any of the above vehicle data-based model training methods.
[0192] The units included in the above electronic device can achieve the same technical effects as the training method of the model based on vehicle data in the above embodiment. To avoid repetition, they will not be described here.
[0193] In an embodiment of the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the above-mentioned vehicle data-based model training method, and can achieve the same technical effect. To avoid repetition, it is not described here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0194] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for training a model based on vehicle data, characterized in that: The model is applied to a battery capacity estimation task, the model includes a first network and a second network, and the method includes training the model N times, where N=2, 3, ..., wherein the i-th training of the model, where i=1, 2, ..., N, includes: The first network is trained using the i-th training data; the second network is trained using the j-th training data to obtain j-th target data; wherein the j-th target data includes data in the j-th training data for which the loss function of the second network is less than a j-th preset value, where i=j=1, 2, …, N; Wherein: when i=j=1, the i-th training data includes a first data set, and the j-th training data includes a second data set, wherein the label noise of the first data set is less than the label noise of the second data set; the first data set includes first vehicle data and a label of the first vehicle data, and the second data set includes second vehicle data and a label of the second vehicle data, and the label of the first vehicle data and the label of the second vehicle data are battery capacity; When i>1 and j>1, the i-th training data includes the first data set and the j-1-th target data, and the j-th training data includes the j-1-th training data; The input of the model is vehicle data, and the output of the model is the estimated result of battery capacity; The second network is an interactive symmetrical structure, including a first subnetwork and a second subnetwork, the first subnetwork includes an RNN, and the second subnetwork includes an RNN; the jth target data is obtained by fusing the jth data of the first subnetwork and the jth data of the second subnetwork; wherein, when j=1, the jth training data of the first subnetwork includes the third data set, the jth training data of the second subnetwork includes the fourth data set, and the second data set includes the third data set and the fourth data set; when j>1, the jth training data of the first subnetwork includes the j-1th training data of the second subnetwork, and the jth training data of the second subnetwork includes the j-1th training data of the first subnetwork.
2. The method according to claim 1, characterized in that The second network is trained using the jth training data to obtain jth target data; wherein the jth target data includes data in the jth training data where the loss function of the second network is less than a jth preset value, i=j=1, 2, …, N, including: Training the first sub-network using the j-th training data of the first sub-network to obtain the j-th data of the first sub-network; the j-th data of the first sub-network includes data in the j-th training data of the first sub-network where the loss function of the first sub-network is less than the j-th threshold of the first sub-network; The second sub-network is trained using the j-th training data of the second sub-network to obtain the j-th data of the second sub-network; the j-th data of the second sub-network includes data in the j-th training data of the second sub-network in which the loss function of the second sub-network is less than the j-th threshold of the second sub-network.
3. The method according to claim 2, characterized in that The first sub-network and the second sub-network are the same.
4. The method according to claim 1, wherein The first network includes transform.
5. The method according to claim 1, wherein The j-1th preset value is less than the jth preset value, j = 2,…, N.
6. A training device for a model based on vehicle data, characterized in that: The model is applied to a battery capacity estimation task, the model includes a first network and a second network, and the device is used to train the model N times, where N=2, 3...; The apparatus comprises a first training unit and a second training unit; The device is used to train the model for an i-th time, where i=1, 2, …, N, and includes: The first training unit is used to train the first network using the i-th training data; the second training unit is used to train the second network using the j-th training data to obtain j-th target data; wherein the j-th target data includes data in the j-th training data for which the loss function of the second network is less than a j-th preset value, i=j=1, 2, …, N; Wherein, when i=j=1, the i-th training data includes a first data set, and the j-th training data includes a second data set, wherein the label noise of the first data set is less than the label noise of the second data set; the first data set includes first vehicle data and a label of the first vehicle data, and the second data set includes second vehicle data and a label of the second vehicle data, and the label of the first vehicle data and the label of the second vehicle data are battery capacity; When i>1 and j>1, the i-th training data includes the first data set and the j-1-th target data, and the j-th training data includes the j-1-th training data; The input of the model is vehicle data, and the output of the model is the estimated result of battery capacity; The second network is an interactive symmetrical structure, including a first subnetwork and a second subnetwork, the first subnetwork including an RNN, and the second subnetwork including an RNN; wherein, when j=1, the j-th training data of the first subnetwork includes the third data set, the j-th training data of the second subnetwork includes the fourth data set, and the second data set includes the third data set and the fourth data set; When j>1, the j-th training data of the first sub-network includes the j-1-th training data of the second sub-network, and the j-th training data of the second sub-network includes the j-1-th training data of the first sub-network; A data fusion unit is configured to fuse the j-th data of the first sub-network and the j-th data of the second sub-network to obtain the j-th target data.
7. The device according to claim 6, characterized in that The second training unit includes a first sub-training unit, a second sub-training unit and a data fusion unit, wherein: The first sub-training unit is configured to train the first sub-network using the j-th training data of the first sub-network to obtain the j-th data of the first sub-network; the j-th data of the first sub-network includes data in the j-th training data of the first sub-network where the loss function of the first sub-network is less than the j-th threshold of the first sub-network; The second sub-training unit is used to train the second sub-network using the j-th training data of the second sub-network to obtain the j-th data of the second sub-network; the j-th data of the second sub-network includes data in the j-th training data of the second sub-network where the loss function of the second sub-network is less than the j-th threshold of the second sub-network.
8. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores codes, and the processor is configured to call the codes stored in the memory to execute the method according to any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Image recognition model training method and device and image recognition method and device
CN112307860A
Battery SOH evaluation model construction method and battery SOH value evaluation method
CN113111580A