Fuel consumption prediction model training method and device, and electronic equipment
By segmenting and extracting features from the sample vehicle operation data set, and using anomaly index thresholds to determine fuel consumption type, the problem of high training difficulty and low accuracy of existing fuel consumption prediction models is solved, achieving efficient training and accurate prediction of fuel consumption prediction models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GREAT WALL MOTOR CO LTD
- Filing Date
- 2023-08-25
- Publication Date
- 2026-05-19
AI Technical Summary
Existing fuel consumption prediction models are difficult to train, have low prediction accuracy, and cannot accurately determine the fuel consumption level of sample data.
By segmenting the sample vehicle operation data set of multiple sample trips, the first positive sample and the first negative sample vehicle operation data set are obtained. Based on these data sets, feature extraction and training are performed. The abnormal index threshold is used to determine the fuel consumption type, thereby reducing the training difficulty and improving the prediction accuracy.
It has achieved accurate training of the fuel consumption prediction model, improved the accuracy and efficiency of fuel consumption prediction, and enabled better identification of vehicle fuel consumption types.
Smart Images

Figure CN117194978B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automotive fuel consumption prediction technology, and more specifically, to a training method, apparatus, and electronic device for a fuel consumption prediction model in the field of automotive fuel consumption prediction technology. Background Technology
[0002] With the development of society and the economy, automobiles are being used more and more widely, and their ownership is increasing year by year. Due to the shortage of oil and its rising price, fuel consumption of automobiles is receiving more and more attention.
[0003] Currently, there is a discrepancy between the fuel consumption displayed on a vehicle's dashboard and the actual fuel consumption. Furthermore, the process of manually calculating the actual fuel consumption is quite complex, making it difficult for drivers to obtain accurate information about their vehicle's fuel consumption. Therefore, to achieve accurate monitoring of vehicle fuel consumption, fuel consumption prediction models are typically used to predict fuel consumption.
[0004] In the training process of existing fuel consumption prediction models, the inability to accurately determine the level of fuel consumption in sample data leads to high training difficulty and low prediction accuracy. Summary of the Invention
[0005] This application provides a fuel consumption prediction method, a training method and apparatus for a fuel consumption prediction model, which can solve the problems of high training difficulty and low prediction accuracy of existing fuel consumption prediction models.
[0006] Firstly, a fuel consumption prediction method is provided, which includes:
[0007] Multiple sample vehicle operation data sets from multiple sample trips are segmented to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The different sample vehicle operation data sets for each sample trip are collected at different time points. Each sample vehicle operation data set includes multiple types of sample vehicle operation data. The actual fuel consumption of the first positive sample vehicle operation data set meets a first preset condition, and the actual fuel consumption of the first negative sample vehicle operation data set meets a second preset condition.
[0008] Feature extraction is performed on the multiple sample vehicle operation data sets for each sample trip to obtain sample trip features for each sample trip;
[0009] Based on the multiple sets of first positive sample vehicle operation data, the multiple sets of first negative sample vehicle operation data, and the sample trip characteristics of each sample trip, the fuel consumption prediction model is trained. The fuel consumption prediction model is used to predict the fuel consumption type, and the fuel consumption type is used to represent the abnormal index and abnormal index threshold of the actual fuel consumption of the target vehicle, as well as the relationship between the actual fuel consumption and the average fuel consumption of the target vehicle.
[0010] The first preset condition includes any one of the following: the abnormal index of actual fuel consumption is greater than the abnormal index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption; the abnormal index of actual fuel consumption is less than or equal to the abnormal index threshold.
[0011] The second preset condition includes: the abnormality index of the actual fuel consumption is greater than the abnormality index threshold, and the actual fuel consumption is greater than the average fuel consumption;
[0012] The anomaly index is used to indicate the degree of outlier of the actual fuel consumption relative to the actual fuel consumption of other sample vehicle operation data sets in the plurality of sample vehicle operation data sets, and the average fuel consumption is the average of multiple actual fuel consumptions of the sample trip to which the actual fuel consumption belongs.
[0013] In the above technical solution, firstly, the sample vehicle operation data sets of multiple sample trips are segmented to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The actual fuel consumption of the first positive sample vehicle operation data sets meets a first preset condition, which includes any one of the following: the anomaly index of the actual fuel consumption is greater than an anomaly index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption; or the anomaly index of the actual fuel consumption is less than or equal to an anomaly index threshold. The actual fuel consumption of the first negative sample vehicle operation data sets meets a second preset condition, which includes: the anomaly index of the actual fuel consumption is greater than an anomaly index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption. The process involves several steps: First, the fuel consumption is greater than the average fuel consumption. Second, feature extraction is performed on multiple sample vehicle operation data sets for each sample trip to obtain sample trip features for each sample trip. Then, based on multiple first positive sample vehicle operation data sets and the sample trip features of each first positive sample vehicle operation data set, multiple second positive sample vehicle operation data sets are obtained. Simultaneously, based on multiple first negative sample vehicle operation data sets and the sample trip features of each first negative sample vehicle operation data set, multiple second negative sample vehicle operation data sets are obtained. Finally, based on the multiple second positive sample vehicle operation data sets and the multiple second negative sample vehicle operation data sets, the fuel consumption prediction model is... Training is performed; this application divides multiple sample vehicle operation data sets into multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets based on the relationship between the anomaly index and the anomaly index threshold of the actual fuel consumption of each sample vehicle operation data set, and the relationship between the actual fuel consumption of each sample vehicle operation data set and the average fuel consumption of the sample trip to which the sample vehicle operation data set belongs. The first positive sample vehicle operation data sets correspond to low fuel consumption, and the first negative sample vehicle operation data sets correspond to high fuel consumption, thus achieving accurate judgment of the high or low fuel consumption of each sample vehicle operation data set. The second positive sample vehicle operation data set... The dataset is obtained based on the first positive sample vehicle operation dataset, and the second negative sample vehicle operation dataset is obtained based on the first negative sample vehicle operation dataset. Therefore, the fuel consumption type corresponding to the second positive sample vehicle operation dataset is low fuel consumption, and the fuel consumption type corresponding to the second negative sample vehicle operation dataset is high fuel consumption. When training the fuel consumption prediction model using multiple second positive sample vehicle operation datasets and multiple second negative sample vehicle operation datasets, since the fuel consumption types of the second positive sample vehicle operation datasets and the second negative sample vehicle operation datasets are known, the training difficulty of the fuel consumption prediction model is effectively reduced, and the prediction accuracy of the fuel consumption prediction model is improved.
[0014] In conjunction with the first aspect, in some possible implementations, before performing sample segmentation on the multiple sample vehicle operation data sets of multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets, the method further includes:
[0015] Acquire multiple sets of vehicle operation data from multiple sample vehicles, wherein each of the sample vehicles is of the same type as the target vehicle, and the target vehicle is the vehicle for which fuel consumption prediction is to be performed.
[0016] The empty vehicle operation data set in the plurality of vehicle operation data sets to be processed is filtered to obtain a plurality of sample vehicle operation data sets, wherein the empty vehicle operation data set is the vehicle operation data set to be processed where the actual fuel consumption and / or mileage of the sample vehicle is empty.
[0017] The multiple sample vehicle operation data sets are divided to obtain multiple sample trips, wherein each sample trip includes multiple sample vehicle operation data sets.
[0018] In the above technical solution, firstly, multiple sets of vehicle operation data for multiple sample vehicles are acquired; secondly, for each sample vehicle, vehicle operation data sets with empty actual fuel consumption and / or mileage are removed, resulting in multiple sets of vehicle operation data after removal. Each of these removed sets is then used as a sample vehicle operation data set, thus obtaining multiple sets of sample vehicle operation data, effectively ensuring the authenticity of the data in the sample vehicle operation data sets; finally, the multiple sets of sample vehicle operation data for each sample vehicle are divided into multiple sample trips, facilitating the subsequent acquisition of different sample trip features based on different sample trips of the same sample vehicle, effectively improving the accuracy of the sample trip features.
[0019] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of segmenting the multiple sample vehicle operation data sets of multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets includes:
[0020] Determine the anomaly index of actual fuel consumption in each of the sample vehicle operation data sets;
[0021] Determine the average fuel consumption for each of the sample trips;
[0022] Based on the anomaly index of actual fuel consumption in each of the multiple sample vehicle operation data sets, and the average fuel consumption of each sample trip, multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets are determined.
[0023] In the above technical solution, the anomaly index of actual fuel consumption in each sample vehicle operation data set of multiple sample trips is first determined. Then, the average fuel consumption of each sample trip is determined. Finally, based on the anomaly index of actual fuel consumption in each sample vehicle operation data set and the average fuel consumption of the sample trip to which the sample vehicle operation data set belongs, the sample type of the sample vehicle operation data set is determined. This results in multiple first positive sample vehicle operation data sets with low fuel consumption and multiple first negative sample vehicle operation data sets with high fuel consumption. This approach can accurately determine the fuel consumption level of the sample vehicle operation data sets, reduce the training difficulty of the fuel consumption prediction model, and facilitate its widespread use.
[0024] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, determining multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets based on the anomaly index of actual fuel consumption in each of the multiple sample vehicle operation data sets includes:
[0025] For each of the sample vehicle operation data sets, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold, and the actual fuel consumption in the sample vehicle operation data set is greater than the average fuel consumption, then the sample vehicle operation data set is determined to be the first negative sample vehicle operation data set.
[0026] For each of the sample vehicle operation data sets, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold, and the actual fuel consumption in the sample vehicle operation data set is less than or equal to the average fuel consumption, the sample vehicle operation data set is determined to be the first positive sample vehicle operation data set.
[0027] For each of the sample vehicle operation data sets, if the anomaly index of the actual fuel consumption in the sample vehicle operation data set is less than or equal to the anomaly index threshold, the sample vehicle operation data set is determined as the first positive sample vehicle operation data set.
[0028] In the above technical solution, for each sample vehicle operation data set in multiple sample vehicle operation data sets of multiple sample trips, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold and the actual fuel consumption in the sample vehicle operation data set is greater than the average fuel consumption, the sample vehicle operation data set is determined as the first negative sample vehicle operation data set; if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold and the actual fuel consumption in the sample vehicle operation data set is less than or equal to the average fuel consumption, or if the abnormality index of the sample vehicle operation data set is less than or equal to the abnormality index threshold, the sample vehicle operation data set is determined as the first positive sample vehicle operation data set. Based on the relationship between the abnormality index and the abnormality index threshold of the actual fuel consumption in each sample vehicle operation data set, and the relationship between the actual fuel consumption in each sample vehicle operation data set and the average fuel consumption of the sample trip in which the sample vehicle operation data set is located, the sample type of each sample vehicle operation data set can be accurately and quickly determined, thereby obtaining the fuel consumption type of each sample vehicle operation data set, effectively reducing the training difficulty of the fuel consumption prediction model.
[0029] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of extracting features from the plurality of sample vehicle operation data sets for each sample trip to obtain sample trip features for each sample trip includes:
[0030] Data fusion and data filtering are performed on the same type of sample vehicle operation data in the multiple sample vehicle operation data sets for each sample trip to obtain the sample trip characteristics.
[0031] In the above technical solution, by performing data fusion and data filtering on the same type of sample vehicle operation data from multiple sample vehicle operation data sets for each sample trip, the sample trip characteristics of the sample trip are obtained. This facilitates the subsequent generation of multiple second positive sample vehicle operation data sets based on multiple first positive sample vehicle operation data sets and the sample trip characteristics of the sample trips corresponding to each first positive sample vehicle operation data set, and the generation of multiple second negative sample vehicle operation data sets based on multiple first negative sample vehicle operation data sets and the sample trip characteristics of the sample trips corresponding to each first negative sample vehicle operation data set.
[0032] In combination with the first aspect and the above implementation methods, in some possible implementation methods, training the fuel consumption prediction model based on the plurality of first positive sample vehicle operation data sets, the plurality of first negative sample vehicle operation data sets, and the sample trip characteristics of each of the sample trips includes:
[0033] The multiple sets of first positive sample vehicle operation data and the multiple sets of first negative sample vehicle operation data are respectively fused with the sample trip features of each sample trip to obtain multiple sets of second positive sample vehicle operation data and multiple sets of second negative sample vehicle operation data.
[0034] The fuel consumption prediction model is trained based on the multiple sets of second positive sample vehicle operation data and the multiple sets of second negative sample vehicle operation data.
[0035] In the above technical solution, for each of the multiple first positive sample vehicle operation data sets, the sample trip features of the sample trip to which the first positive sample vehicle operation data set belongs are combined with the first positive sample vehicle operation data set to obtain a second positive sample vehicle operation data set. For each of the multiple first negative sample vehicle operation data sets, the sample trip features of the sample trip to which the first negative sample vehicle operation data set belongs are combined with the first negative sample vehicle operation data set to obtain a second negative sample vehicle operation data set. The fuel consumption prediction model is trained using multiple second positive sample vehicle operation data sets and multiple second negative sample vehicle operation data sets. The training effect is good, and the prediction accuracy of the trained fuel consumption prediction model is high.
[0036] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, after training the fuel consumption prediction model based on the plurality of first positive sample vehicle operation data sets, the plurality of first negative sample vehicle operation data sets, and the sample trip features of each sample trip, the method further includes:
[0037] Obtain multiple sets of first vehicle operation data for the target vehicle during the target journey;
[0038] Feature extraction is performed on the multiple sets of first vehicle operation data to obtain the travel features of the target trip;
[0039] The travel characteristics are fused with each of the first vehicle operation data sets to obtain multiple second vehicle operation data sets;
[0040] Based on the fuel consumption prediction model after training and the multiple sets of second vehicle operation data, the fuel consumption type of the target vehicle during the target trip is determined.
[0041] In the above technical solution, after obtaining the trained fuel consumption prediction model, firstly, multiple sets of first vehicle operation data for the target vehicle during the target trip are acquired, and features are extracted from the multiple sets of first vehicle operation data to obtain the trip features of the target trip; secondly, the trip features of the target trip are combined with each of the multiple sets of first vehicle operation data to obtain multiple sets of second vehicle operation data; finally, the multiple sets of second vehicle operation data are input into the trained fuel consumption prediction model, and the fuel consumption prediction model outputs the fuel consumption type of the target trip, thereby achieving prediction of the fuel consumption type of the target vehicle during the target trip with high accuracy.
[0042] Secondly, a training device for a fuel consumption prediction model is provided, the device comprising:
[0043] The sample segmentation module is used to segment multiple sample vehicle operation data sets from multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The different sample vehicle operation data sets for each sample trip are collected at different time points. Each sample vehicle operation data set includes multiple types of sample vehicle operation data. The actual fuel consumption of the first positive sample vehicle operation data set meets a first preset condition, and the actual fuel consumption of the first negative sample vehicle operation data set meets a second preset condition.
[0044] The first feature extraction module is used to extract features from the multiple sample vehicle operation data sets of each sample trip to obtain sample trip features for each sample trip.
[0045] The training module is used to train the fuel consumption prediction model based on the multiple first positive sample vehicle operation data sets, the multiple first negative sample vehicle operation data sets, and the sample trip features of each sample trip. The fuel consumption prediction model is used to predict the fuel consumption type, and the fuel consumption type is used to represent the abnormality index and abnormality index threshold of the actual fuel consumption of the target vehicle, as well as the relationship between the actual fuel consumption and the average fuel consumption of the target vehicle.
[0046] The first preset condition includes any one of the following: the abnormal index of actual fuel consumption is greater than the abnormal index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption; the abnormal index of actual fuel consumption is less than or equal to the abnormal index threshold.
[0047] The second preset condition includes: the abnormality index of the actual fuel consumption is greater than the abnormality index threshold, and the actual fuel consumption is greater than the average fuel consumption;
[0048] The anomaly index is used to indicate the degree of outlier of the actual fuel consumption relative to the actual fuel consumption of other sample vehicle operation data sets in the plurality of sample vehicle operation data sets, and the average fuel consumption is the average of multiple actual fuel consumptions of the sample trip to which the actual fuel consumption belongs.
[0049] In conjunction with the second aspect, in some possible implementations, the device further includes:
[0050] The first acquisition module is used to acquire multiple sets of vehicle operation data to be processed from multiple sample vehicles, wherein each of the sample vehicles is of the same type as the target vehicle, and the target vehicle is the vehicle whose fuel consumption is to be predicted.
[0051] The filtering module is used to filter the empty vehicle operation data set in the plurality of vehicle operation data sets to be processed, so as to obtain a plurality of sample vehicle operation data sets, wherein the empty vehicle operation data set is the vehicle operation data set to be processed where the actual fuel consumption and / or mileage of the sample vehicle is empty.
[0052] The partitioning module is used to partition the multiple sample vehicle operation data sets of the multiple sample vehicles to obtain multiple sample trips, wherein each sample trip includes multiple sample vehicle operation data sets.
[0053] In combination with the second aspect and the above implementation methods, in some possible implementations, the sample segmentation module is specifically used for:
[0054] Determine the anomaly index of actual fuel consumption in each of the sample vehicle operation data sets;
[0055] Determine the average fuel consumption for each of the sample trips;
[0056] Based on the anomaly index of actual fuel consumption in each of the multiple sample vehicle operation data sets, and the average fuel consumption of each sample trip, multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets are determined.
[0057] In combination with the second aspect and the above implementation methods, in some possible implementations, the sample segmentation module is specifically used for:
[0058] For each of the sample vehicle operation data sets, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold, and the actual fuel consumption in the sample vehicle operation data set is greater than the average fuel consumption, then the sample vehicle operation data set is determined to be the first negative sample vehicle operation data set.
[0059] For each of the sample vehicle operation data sets, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold, and the actual fuel consumption in the sample vehicle operation data set is less than or equal to the average fuel consumption, the sample vehicle operation data set is determined to be the first positive sample vehicle operation data set.
[0060] For each of the sample vehicle operation data sets, if the anomaly index of the actual fuel consumption in the sample vehicle operation data set is less than or equal to the anomaly index threshold, the sample vehicle operation data set is determined as the first positive sample vehicle operation data set.
[0061] In combination with the second aspect and the above implementation methods, in some possible implementation methods, the first feature extraction module is specifically used for:
[0062] Data fusion and data filtering are performed on the same type of sample vehicle operation data in the multiple sample vehicle operation data sets for each sample trip to obtain the sample trip characteristics.
[0063] In combination with the second aspect and the above implementation methods, in some possible implementations, the training module is specifically used for:
[0064] The multiple sets of first positive sample vehicle operation data and the multiple sets of first negative sample vehicle operation data are respectively fused with the sample trip features of each sample trip to obtain multiple sets of second positive sample vehicle operation data and multiple sets of second negative sample vehicle operation data.
[0065] The fuel consumption prediction model is trained based on the multiple sets of second positive sample vehicle operation data and the multiple sets of second negative sample vehicle operation data.
[0066] In conjunction with the second aspect and the above-described implementations, in some possible implementations, the apparatus further includes:
[0067] The second acquisition module is used to acquire multiple sets of first vehicle operation data of the target vehicle during the target journey;
[0068] The second feature extraction module is used to extract features from the plurality of first vehicle operation data sets to obtain the travel features of the target trip;
[0069] The fusion module is used to fuse the trip features with each of the first vehicle operation data sets to obtain multiple second vehicle operation data sets;
[0070] The fuel consumption prediction module is used to determine the fuel consumption type of the target vehicle during the target trip based on the fuel consumption prediction model after training and the multiple sets of second vehicle operation data.
[0071] Thirdly, an electronic device is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the electronic device to perform the methods of the first aspect or any possible implementation thereof.
[0072] Fourthly, a computer program product is provided, comprising: computer program code, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0073] Fifthly, a computer-readable storage medium is provided that stores computer program code, which, when executed on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof. Attached Figure Description
[0074] Figure 1 This is a schematic diagram of the implementation environment for a training method for a fuel consumption prediction model provided in an embodiment of this application;
[0075] Figure 2 This is a schematic flowchart illustrating a training method for a fuel consumption prediction model provided in an embodiment of this application;
[0076] Figure 3 This is a schematic diagram of the structure of a training device for a fuel consumption prediction model provided in an embodiment of this application;
[0077] Figure 4 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application. Detailed Implementation
[0078] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0079] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0080] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0081] Vehicle Type: Vehicles with similar characteristics can be classified into different types, either by vehicle function or by vehicle model. There are multiple standards for classifying vehicle types, and this embodiment does not impose any specific limitations. For example, vehicle types can be distinguished by vehicle model; all vehicles of the same type have the same model.
[0082] Vehicle operating data: Specific parameters that may be involved in the vehicle's operation, including driving parameters, actual fuel consumption parameters, and vehicle-related parameters; among which, driving parameters include vehicle speed, longitudinal acceleration, and lateral acceleration; actual fuel consumption parameters include the vehicle's actual fuel consumption; vehicle-related parameters include engine speed, tire pressure, coolant temperature, transmission fluid temperature, and remaining fuel level.
[0083] Unsupervised methods: There is no training set, only a set of data, and the goal is to find patterns within that set of data.
[0084] Isolation Forest algorithm: This is an unsupervised anomaly detection method that detects outliers by isolating them, and it is constructed from multiple isolated trees. When there are few outliers in a sample and they are far from normal samples, outliers in the data space are quickly isolated, and in this case, outliers are closer to the root node, while normal values are more sparsely distributed from the root node. Compared with other methods, Isolation Forest has many advantages in anomaly detection. First, it only needs to obtain a small number of samples from a large dataset to obtain a fast and highly generalizable detection algorithm. Second, it does not require the training dataset to contain outliers, which is more advantageous for detection tasks with few outliers. Third, the tree depth is the basis for determining the distance threshold of anomalies, and this threshold is not necessarily related to the size of the dataset dimension.
[0085] Isolation forests introduce an outlier function s(x,n) to evaluate whether a specific sample x (which can be denoted as an observation point) is an outlier, as shown in the following formula (1):
[0086]
[0087] Where c(n) = H(n-1) - (2(n-1) / n); H(n-1) = ln(n-1) + Euler; E(h(x)) is the average path length from a specific observation point x to the root node of an isolated tree in multiple isolated trees. c(n) is the average path length of all isolated trees in an isolated forest model constructed from a dataset containing n samples, specifically defined as the parameter used to standardize the record E(h(x)). H(n-1) is the harmonic number, and Euler represents Euler's constant.
[0088] In an isolated tree, the path length from an observation point to the root node can be defined as the number of other nodes between the root node and the observation point plus 1. If an observation point is directly connected to the root node, there are no other nodes between the observation point and the root node, the number of other nodes is 0, and the path length is 0 plus 1, which equals 1. If an observation point is connected to the root node through three other nodes, then the path length of the observation point is the number of other nodes between it and the root node, 3 plus 1, which equals 4.
[0089] The process of using the Isolation Forest algorithm to perform anomaly detection on a test dataset containing n samples (all variables in the samples are continuous values) is as follows:
[0090] The following process is performed to generate an isolation tree for the dataset to be tested:
[0091] First, determine a root node. Then, randomly select a variable of a certain dimension of the sample and determine the classification value of the variable. The classification value can be the average of all samples in the dataset on the variable, or it can be the median of all samples on the variable. The specific determination method is not limited. After determining the classification value, the dataset can be split into two subsets according to whether the value of the selected variable is greater than the classification value. The two subsets are then assigned to the two child nodes of the root node.
[0092] Then, for each child node's subset, if the subset contains more than one sample, the above random classification process is repeated for this subset. A variable is determined, and the corresponding classification value is determined based on the value of each sample in this subset on the variable. Then, based on whether the value of the sample on the variable is greater than the classification value, this subset is further split into two smaller subsets and assigned to the two child nodes of the current node. This process is repeated until each leaf node's subset contains only one sample, thus obtaining an isolated tree composed of several nodes.
[0093] By repeating the process of generating isolated trees K times on the dataset to be tested, we can obtain K isolated trees. The model composed of these K isolated trees is the isolated forest model. The process of generating these K isolated trees is equivalent to the process of training an isolated forest model using the dataset to be tested.
[0094] As can be seen from the above process of generating an isolated tree, for any sample in the dataset to be tested, a unique leaf node corresponding to that sample can be found in each isolated tree. Then, the number of other nodes between the leaf node corresponding to the sample and the root node can be determined, thereby calculating the path length of the sample in the isolated tree.
[0095] Furthermore, for a specific sample in the dataset to be tested, after calculating the lengths of the K paths of the sample in the K isolated trees, the anomaly index s(x,n) of the sample can be calculated according to formula (1), and then the sample can be determined as an anomalous sample based on the magnitude of the anomaly index.
[0096] The implementation environment of the technical solutions provided in the embodiments of this application will be described below. Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for a fuel consumption prediction model provided in this application embodiment. See also... Figure 1 The implementation environment may include an in-vehicle terminal 110 and a server 120.
[0097] The vehicle-mounted terminal 110 is connected to the server 120 via a wireless network. The vehicle-mounted terminal 110 includes various types of sensors, which acquire vehicle operation data while the driver is operating the vehicle. Accordingly, the vehicle-mounted terminal 110 also includes a processor and a memory. The memory stores the information collected by the sensors, and the processor processes the information stored in the memory. The vehicle-mounted terminal 110 has an application installed and running that supports acquiring multiple vehicle operation data points for a target journey.
[0098] Server 120 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0099] After introducing the implementation environment provided by the embodiments of this application, the application scenarios of the embodiments of this application will be introduced below.
[0100] The technical solutions provided in this application can be applied to various fuel consumption prediction scenarios. Using the technical solutions provided in this application to train a fuel consumption prediction model is easy to train, and the trained fuel consumption prediction model can more accurately predict the fuel consumption type of the target trip based on multiple vehicle operation data sets of the target vehicle's target trip.
[0101] After introducing the implementation environment and application scenarios of the embodiments of this application, the technical solutions provided by the embodiments of this application will be described below. See also... Figure 2 , Figure 2 This is a schematic flowchart of a training method for a fuel consumption prediction model provided in an embodiment of this application. Taking the server as the executing entity as an example, the training method 200 for the fuel consumption prediction model includes steps 202 to 206.
[0102] Step 202: The server performs sample segmentation on multiple sample vehicle operation data sets for multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The different sample vehicle operation data sets for each sample trip are collected at different time points. Each sample vehicle operation data set includes multiple types of sample vehicle operation data. The actual fuel consumption of the first positive sample vehicle operation data set meets the first preset condition, and the actual fuel consumption of the first negative sample vehicle operation data set meets the second preset condition.
[0103] The first preset condition includes any one of the following: the abnormal index of actual fuel consumption is greater than the abnormal index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption; the abnormal index of actual fuel consumption is less than or equal to the abnormal index threshold.
[0104] The second preset condition includes: the abnormal index of actual fuel consumption is greater than the abnormal index threshold, and the actual fuel consumption is greater than the average fuel consumption.
[0105] Anomaly index is used to indicate the degree to which actual fuel consumption is out of the ordinary range relative to the actual fuel consumption of other sample vehicle operation data sets in multiple sample vehicle operation data sets. Average fuel consumption is the average of multiple actual fuel consumptions for the sample trip to which the actual fuel consumption belongs.
[0106] It should be noted that the anomaly index of different sample vehicle operation data sets can be the same or different. For any sample vehicle operation data set among multiple sample vehicle operation data sets, the anomaly index of that sample vehicle operation data set is the degree of outlier of the actual fuel consumption of that sample vehicle operation data set relative to the actual fuel consumption of other sample vehicle operation data sets among multiple sample vehicle operation data sets; other sample vehicle operation data sets refer to all sample vehicle operation data sets among multiple sample vehicle operation data sets except for that sample vehicle operation data set.
[0107] Each sample vehicle operation data set includes at least nine types of sample vehicle operation data: vehicle speed, longitudinal acceleration, lateral acceleration, engine speed, actual fuel consumption, tire pressure, coolant temperature, transmission fluid temperature, and remaining fuel level. Furthermore, the sample vehicle operation data set may also include other types of sample vehicle operation data; this embodiment does not specifically limit this.
[0108] Each sample trip includes multiple sample vehicle operation data sets, and the multiple sample vehicle operation data sets for the same sample trip are collected at different time points.
[0109] It should be noted that before step 202, method 200 also includes steps 2011 to 2013.
[0110] Step 2011: The server obtains multiple sets of vehicle operation data from multiple sample vehicles. Each sample vehicle is of the same type as the target vehicle, which is the vehicle whose fuel consumption is to be predicted.
[0111] It should be noted that the server contains a vehicle operation database. For any vehicle, its onboard terminal can collect a set of data to be processed from the vehicle's historical journey at a first preset time interval, and upload the collected data set to the server for storage in the vehicle operation database. Therefore, the vehicle operation database stores data sets to be processed for various types of vehicles, and there are multiple data sets for each type of vehicle. Each vehicle's multiple data sets to be processed are identified by its Vehicle Identification Number (VIN). Since each vehicle's VIN is unique, using the VIN as the identifier for each vehicle's multiple data sets to be processed allows the server to retrieve multiple data sets to be processed from the same sample vehicle based on the identifier of each data set.
[0112] In this embodiment, the vehicle type of each sample vehicle is the same as that of the target vehicle. The server retrieves the data set to be processed of the sample vehicles with the same vehicle type as the target vehicle from the vehicle operation database according to the vehicle type of the target vehicle. The number of data sets to be processed for each sample vehicle is multiple. The number of data sets to be processed for different sample vehicles can be the same or different. This embodiment does not make a specific limitation on this.
[0113] The first preset time interval can be set by the developer, and this embodiment does not impose specific limitations on it. For example, the first preset time interval can be 1 minute.
[0114] Step 2012: The server filters the empty vehicle operation data set from the multiple unprocessed vehicle operation data sets to obtain multiple sample vehicle operation data sets. The empty vehicle operation data set is the unprocessed vehicle operation data set where the actual fuel consumption and / or mileage of the sample vehicle is empty.
[0115] Specifically, for each sample vehicle, multiple sets of vehicle operation data to be processed are processed. The sets of vehicle operation data with empty actual fuel consumption and / or mileage are removed, and each set of vehicle operation data after removal is treated as a sample vehicle operation data set.
[0116] It's important to note that vehicle operation datasets lacking mileage and actual fuel consumption data cannot provide effective information and do not contribute to the modeling process. They may even mislead the fuel consumption prediction model, potentially misinterpreting these missing values as specific numerical values or features, thus affecting the final prediction accuracy. Therefore, removing empty vehicle operation datasets can improve the accuracy of the fuel consumption prediction model, avoid misleading it, and ensure that the model can better perform fuel consumption type prediction tasks.
[0117] In this embodiment of the application, for each sample vehicle, the multiple sample vehicle operation data sets of the sample vehicle are obtained by filtering the multiple unprocessed vehicle operation data sets of the sample vehicle. The number of sample vehicle operation data sets of the sample vehicle is less than or equal to the unprocessed vehicle operation data sets of the sample vehicle. The unprocessed vehicle operation data sets may be vehicle operation data sets with empty actual fuel consumption and / or mileage, while the actual fuel consumption and mileage of the sample vehicle operation data sets are not empty.
[0118] Step 2013: The server divides the multiple sample vehicle operation data sets of multiple sample vehicles into multiple sample trips, where each sample trip includes multiple sample vehicle operation data sets.
[0119] Specifically, for each sample vehicle, multiple sample vehicle operation data sets are arranged according to the chronological order of data collection time to obtain the sample sequence of the sample vehicle. If the data collection time interval between two adjacent sample vehicle operation data sets in the sample sequence is greater than a second preset time interval, then the two adjacent sample vehicle operation data sets are divided into different sample trips.
[0120] For each sample vehicle, multiple sample trips are assigned a trip number, and different sample trips of the same sample vehicle have different trip numbers.
[0121] In this embodiment of the application, the trip number of the sample trip can be composed of the vehicle identification number (VIN) and the numerical number of the sample vehicle corresponding to the sample trip.
[0122] For example, if the vehicle identification number (VIN) of the sample vehicle is 5YJSA1CN2DFP12345, and this sample trip is the 20th trip of the sample vehicle, then the trip number of this sample trip can be 5YJSA1CN2DFP12345-000020.
[0123] The second preset time interval can be set by the developer, and this embodiment does not impose specific limitations on it. For example, the second preset time interval can be 15 minutes.
[0124] For example, if the data collection times for multiple sample vehicle operation data sets of the same sample vehicle are 12:59, 1:00, 1:01, 1:02, 1:03, 1:04, 1:05, 1:07, 1:25, 1:26, 1:27, 1:28, 1:29, and 1:31 on May 1, 2023, and the preset time interval is 15 minutes, since the time interval between 1:07 and 1:25 is 18 minutes, then the multiple sample vehicle operation data sets corresponding to the data collection times of 12:59, 1:00, 1:01, 1:02, 1:03, 1:04, 1:05, and 1:07 are divided into the first sample trip of the sample vehicle, and the multiple sample vehicle operation data sets corresponding to the data collection times of 1:25, 1:26, 1:27, 1:28, 1:29, and 1:31 are divided into the second sample trip of the sample vehicle.
[0125] It should be noted that the number of sample vehicle operation data sets in each sample trip is no less than 3. If the number of sample vehicle operation data sets in a sample trip is less than 3, the sample trip will be removed. This is to avoid the sample trip characteristics from having a large gap with the actual characteristics of the sample vehicles corresponding to the sample trip due to insufficient sample vehicle operation data sets. This can effectively ensure the accuracy of the training samples for training the fuel consumption prediction model.
[0126] In this embodiment of the application, by dividing the operation data set of multiple sample vehicles of the same sample vehicle into multiple sample trips, it is convenient to obtain different sample trip features based on different sample trips of the same sample vehicle, which can effectively improve the accuracy of sample trip features.
[0127] In this embodiment, the server first acquires multiple sets of vehicle operation data for multiple sample vehicles. Second, it removes sets of vehicle operation data for each sample vehicle whose actual fuel consumption and / or mileage are empty, resulting in multiple sets of vehicle operation data after removal. Each of these removed sets is then used as a sample vehicle operation data set, thus obtaining multiple sets of sample vehicle operation data, effectively ensuring the authenticity of the data in the sample vehicle operation data sets. Finally, the multiple sets of sample vehicle operation data for each sample vehicle are divided into multiple sample trips, facilitating the subsequent acquisition of different sample trip features based on different sample trips of the same sample vehicle, effectively improving the accuracy of the sample trip features.
[0128] In step 202, the server performs sample segmentation on the multiple sample vehicle operation data sets of multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets, specifically including steps 2021 to 2023.
[0129] Step 2021: The server determines the anomaly index of actual fuel consumption in each sample vehicle operation data set.
[0130] Specifically, when the server determines the anomaly index of any sample vehicle operation data set, it includes steps 20211 to 20215.
[0131] Step 20211: Set the maximum depth h of the isolated forest, where h = ceil(log2n), and n is the total number of sample vehicle operation data sets.
[0132] In the isolated forest model, the maximum depth h of the isolated forest is also called the height of the tree.
[0133] It should be noted that, for a given dataset, the maximum depth h of each tree in an isolated forest model depends on the number of samples m in the dataset. The formula h = ceil(log2m) uses the log2 function to calculate the logarithm to base 2, and then applies the floor function ceil to ensure that the maximum depth h is an integer. This allows for better control over the depth of the trees in the isolated forest, making it suitable for the size of the dataset.
[0134] The determination of the maximum depth h of the isolated forest is based on the following two factors: First, when the number of samples in the dataset is large, the depth of the tree needs to be increased appropriately in order to better capture the features of outliers; Second, the maximum depth h of the isolated forest should not be too large, otherwise it will lead to overfitting and reduce the generalization ability of the isolated forest model.
[0135] Step 20212: Perform P random samplings on the actual fuel consumption of the n sample vehicle operation data sets. Each time, construct a random binary tree by sampling the actual fuel consumption of the Q sample vehicle operation data sets. Determine the isolated forest model based on the constructed P random binary trees, where P > 0, Q > 0, and P and Q are both positive integers.
[0136] The construction method of P random binary trees is the same. When constructing any random binary tree, the actual fuel consumption in the sample vehicle operation data set is randomly selected as the dividing value. If the actual fuel consumption in the sample vehicle operation data set used to construct the random binary tree is less than the dividing value, it is divided into the left child node; otherwise, it is divided into the right child node. The process stops when the preset conditions are met, and a random binary tree is generated. The preset conditions may be that the sample vehicle operation data set used to construct the random binary tree cannot be further divided or the random binary tree reaches its maximum depth.
[0137] Step 20213: Determine the distance from the actual fuel consumption in each sample vehicle operation data set to the root node of each isolated tree in the isolated forest model.
[0138] In this model, the sample vehicle operation data set is equivalent to a sample. As mentioned earlier, in each isolated tree of the isolated forest model, a sample can uniquely correspond to a leaf node in that isolated tree. Therefore, for a specific sample vehicle operation data set containing the actual fuel consumption x, for each isolated tree, we can first determine the leaf node corresponding to x in that isolated tree, then count the number of other nodes between the leaf node corresponding to x and the root node of that isolated tree, and add 1 to the count to obtain the distance (i.e., path length) from x to the root node in that isolated tree.
[0139] Step 20214: Determine the average path length of the actual fuel consumption in each sample vehicle operation data set based on the distance from the actual fuel consumption in each sample vehicle operation data set to the root node of each isolated tree.
[0140] Assuming the isolated forest model includes K isolated trees, step 20213 can be used to obtain the path length of the actual fuel consumption x in each isolated tree in the sample vehicle operation data set, that is, to obtain the K path lengths of x. The average of these K path lengths can be calculated to obtain the mean path length of the actual fuel consumption x in the sample vehicle operation data set.
[0141] Step 20215: Determine the anomaly index of the sample vehicle operation data set based on the average path length of actual fuel consumption in the sample vehicle operation data set.
[0142] The mean path length of the actual fuel consumption x in the sample vehicle operation data set is used as the formula (1). E(h(x)) in the sample vehicle operation data set is obtained by calculating c(n) in formula (1) based on the total number n of the sample vehicle operation data set. Thus, the anomaly index s(x,n) of the sample vehicle operation data set x can be calculated by formula (1).
[0143] Step 2022: Determine the average fuel consumption for each sample trip.
[0144] The average fuel consumption for different sample trips can be the same or different.
[0145] It should be noted that for each sample trip, the sample trip includes multiple sample vehicle operation data sets, and each sample vehicle operation data set includes an actual fuel consumption data. The average fuel consumption of the sample trip is the average of the actual fuel consumption of all sample vehicle operation data sets for that sample trip.
[0146] Step 2023: The server determines multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets based on the anomaly index of the actual fuel consumption in each sample vehicle operation data set and the average fuel consumption of each sample trip.
[0147] Among them, the server uses the isolated forest algorithm to quickly detect anomalies in the sample vehicle operation data set, which has the advantages of high efficiency and speed.
[0148] In this embodiment, the server first determines the anomaly index of the actual fuel consumption in each sample vehicle operation data set of multiple sample trips, then determines the average fuel consumption of each sample trip, and finally determines the sample type of the sample vehicle operation data set based on the anomaly index of the actual fuel consumption in each sample vehicle operation data set and the average fuel consumption of the sample trip to which the sample vehicle operation data set belongs. This results in multiple first positive sample vehicle operation data sets with low fuel consumption and multiple first negative sample vehicle operation data sets with high fuel consumption. This can accurately determine the fuel consumption level of the sample vehicle operation data set, reduce the training difficulty of the fuel consumption prediction model, and facilitate its widespread use.
[0149] It should be noted that in step 2023, the server determines multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets based on the abnormal index of actual fuel consumption in each sample vehicle operation data set and the average fuel consumption of each sample trip. Specifically, this includes steps 20231 to 20233.
[0150] Step 20231: For each sample vehicle operation data set, if the abnormal index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormal index threshold, and the actual fuel consumption in the sample vehicle operation data set is greater than the average fuel consumption, the server determines the sample vehicle operation data set as the first negative sample vehicle operation data set.
[0151] Specifically, for each sample vehicle operation data set, if the anomaly index of the actual fuel consumption in the sample vehicle operation data set is greater than the anomaly index threshold, it indicates that the actual fuel consumption in the sample vehicle operation data set is preliminarily determined to be abnormal data by the isolated forest model; if the actual fuel consumption in the sample vehicle operation data set is greater than the average fuel consumption of the sample trip to which the sample vehicle operation data set belongs, it indicates that the actual fuel consumption in the sample vehicle operation data set exceeds the average fuel consumption of its trip. Therefore, the actual fuel consumption in the sample vehicle operation data set is determined to be abnormal data, and the sample vehicle operation data set is designated as the first negative sample vehicle operation data set.
[0152] It should be noted that a variable can be added to each sample vehicle operation data set to indicate whether the actual fuel consumption in that sample vehicle operation data set is abnormal. The variable corresponding to that sample vehicle operation data set is "1", which means that the actual fuel consumption in that sample vehicle operation data set belongs to abnormal fuel consumption data, and that sample vehicle operation data set is the first negative sample vehicle operation data set.
[0153] Step 20232: For each sample vehicle operation data set, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold, and the actual fuel consumption in the sample vehicle operation data set is less than or equal to the average fuel consumption, the server determines the sample vehicle operation data set as the first positive sample vehicle operation data set.
[0154] Specifically, for each sample vehicle operation data set, if the anomaly index of the actual fuel consumption in the sample vehicle operation data set is greater than the anomaly index threshold, it indicates that the actual fuel consumption in the sample vehicle operation data set is preliminarily determined to be abnormal data by the isolated forest model; if the actual fuel consumption in the sample vehicle operation data set is less than or equal to the average fuel consumption of the sample trip to which the sample vehicle operation data set belongs, it indicates that the actual fuel consumption in the sample vehicle operation data set does not exceed the average fuel consumption of its trip. Therefore, the actual fuel consumption in the sample vehicle operation data set is determined to be normal data, and the sample vehicle operation data set is the first positive sample vehicle operation data set.
[0155] It should be noted that a variable can be added to each sample vehicle operation data set to indicate whether the actual fuel consumption in that sample vehicle operation data set is abnormal. The variable corresponding to that sample vehicle operation data set is "0", which means that the actual fuel consumption in that sample vehicle operation data set is normal fuel consumption data, and that sample vehicle operation data set is the first positive sample vehicle operation data set.
[0156] Step 20233: For each sample vehicle operation data set, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is less than or equal to the abnormality index threshold, the server determines the sample vehicle operation data set as the first positive sample vehicle operation data set.
[0157] Specifically, for each sample vehicle operation data set, if the anomaly index of the actual fuel consumption in the sample vehicle operation data set is less than or equal to the anomaly index threshold, it means that the actual fuel consumption in the sample vehicle operation data set can be determined to be normal data by the isolated forest model, and the sample vehicle operation data set is the first positive sample vehicle operation data set.
[0158] In this embodiment, for each sample vehicle operation data set in multiple sample vehicle operation data sets of multiple sample trips, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold and the actual fuel consumption in the sample vehicle operation data set is greater than the average fuel consumption, the sample vehicle operation data set is determined as the first negative sample vehicle operation data set; if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold and the actual fuel consumption in the sample vehicle operation data set is less than or equal to the average fuel consumption, or if the abnormality index of the sample vehicle operation data set is less than or equal to the abnormality index threshold, the sample vehicle operation data set is determined as the first positive sample vehicle operation data set. Based on the relationship between the abnormality index and the abnormality index threshold of the actual fuel consumption in each sample vehicle operation data set, and the relationship between the actual fuel consumption in each sample vehicle operation data set and the average fuel consumption of the sample trip in which the sample vehicle operation data set is located, the sample type of each sample vehicle operation data set can be accurately and quickly determined, thereby obtaining the fuel consumption type of each sample vehicle operation data set, effectively reducing the training difficulty of the fuel consumption prediction model.
[0159] Step 204: The server extracts features from the set of multiple sample vehicle operation data for each sample trip to obtain the sample trip features for each sample trip.
[0160] In this embodiment of the application, the server performs data fusion and data filtering on the same type of sample vehicle operation data from multiple sample vehicle operation data sets for each sample trip to obtain the sample trip characteristics.
[0161] The sample journey features of each sample journey include a first feature value, a second feature value, a third feature value, a fourth feature value, a fifth feature value, a sixth feature value, and a seventh feature value.
[0162] In the embodiments of this application, the methods for determining the sample journey characteristics of multiple sample journeys are all the same. The method for determining the sample journey characteristics of any one sample journey specifically includes steps 2041 to 2046.
[0163] Step 2041: Sum the vehicle speed, longitudinal acceleration, lateral acceleration, engine speed, average fuel consumption, tire pressure, coolant temperature and transmission oil temperature from the multiple sample vehicle operation data sets of the sample trip to obtain the first feature value of the sample trip.
[0164] The first characteristic value includes the sum of vehicle speed, longitudinal acceleration, lateral acceleration, engine speed, average fuel consumption, tire pressure, water temperature, and transmission oil temperature for the sample trip.
[0165] Step 2042: Calculate the average of vehicle speed, longitudinal acceleration, lateral acceleration, engine speed, average fuel consumption, tire pressure, coolant temperature, and transmission oil temperature from multiple sample vehicle operation data sets for the sample trip to obtain the second feature value of the sample trip.
[0166] The second characteristic value includes the average vehicle speed, average longitudinal acceleration, average lateral acceleration, average engine speed, average fuel consumption, average tire pressure, average coolant temperature, and average transmission oil temperature for the sample trip.
[0167] Step 2043: Compare and calculate the vehicle speed, longitudinal acceleration, lateral acceleration, engine speed, average fuel consumption, tire pressure, water temperature and transmission oil temperature in the multiple sample vehicle operation data sets of the sample trip to obtain the third feature value and the fourth feature value of the sample trip.
[0168] The third characteristic value includes the maximum vehicle speed, maximum longitudinal acceleration, maximum lateral acceleration, maximum engine speed, maximum average fuel consumption, maximum tire pressure, maximum water temperature, and maximum transmission oil temperature for the sample trip.
[0169] The fourth characteristic value includes the minimum vehicle speed, minimum longitudinal acceleration, minimum lateral acceleration, minimum engine speed, minimum average fuel consumption, minimum tire pressure, minimum water temperature, and minimum transmission oil temperature for the sample trip.
[0170] Step 2044: Perform standard deviation calculations on the vehicle speed, longitudinal acceleration, lateral acceleration, engine speed, average fuel consumption, tire pressure, water temperature, and transmission oil temperature from the multiple sample vehicle operation data sets of the sample trip to obtain the fifth feature value of the sample trip.
[0171] The fifth feature value includes the standard deviation of vehicle speed, longitudinal acceleration, lateral acceleration, engine speed, average fuel consumption, tire pressure, coolant temperature, and transmission oil temperature for the sample trip.
[0172] Step 2045: Determine the sixth feature value based on the water temperature of the multiple sample vehicle operation data sets of the sample trip.
[0173] Among them, if the water temperature of each sample vehicle operation data set is higher than the preset water temperature, the sample vehicle operation data is recorded as the water temperature abnormal sample vehicle operation data set; the sixth characteristic value is the number of water temperature abnormal sample vehicle operation data sets among multiple sample vehicle operation data sets.
[0174] It should be noted that the preset water temperature can be set by the developer, and this embodiment does not impose any specific limitations on it.
[0175] Step 2046: Determine the seventh feature value based on the remaining fuel amount of multiple sample vehicle operation data sets for this sample trip.
[0176] Among them, if the remaining fuel in each sample vehicle operation data set is higher than the preset fuel level, the sample vehicle operation data set is recorded as the fuel level abnormal sample vehicle operation data set; the seventh characteristic value is the number of fuel level abnormal sample vehicle operation data sets among multiple sample vehicle operation data sets.
[0177] It should be noted that the preset oil volume can be set by the developer, and this embodiment does not impose any specific limitations on it.
[0178] In this embodiment of the application, for each sample travel feature, a travel feature number is set. The travel feature number of the sample travel feature can be composed of the travel number of the sample travel corresponding to the sample travel feature and the feature character.
[0179] For example, if the trip number of a sample trip is 5YJSA1CN2DFP12345-000020, then the trip feature number of the sample trip feature of this sample trip can be 5YJSA1CN2DFP12345-000020-0.
[0180] In this embodiment of the application, by performing data fusion and data filtering on the same type of sample vehicle operation data from multiple sample vehicle operation data sets for each sample trip, the sample trip characteristics of the sample trip are obtained. This facilitates the subsequent generation of multiple second positive sample vehicle operation data sets based on multiple first positive sample vehicle operation data sets and the sample trip characteristics of the sample trips corresponding to each first positive sample vehicle operation data set, and the generation of multiple second negative sample vehicle operation data sets based on multiple first negative sample vehicle operation data sets and the sample trip characteristics of the sample trips corresponding to each first negative sample vehicle operation data set.
[0181] Step 206: The server trains the fuel consumption prediction model based on multiple sets of first positive sample vehicle operation data, multiple sets of first negative sample vehicle operation data, and the sample trip features of each sample trip. The fuel consumption prediction model is used to predict the fuel consumption type, and the fuel consumption type is used to represent the abnormal index and abnormal index threshold of the actual fuel consumption of the target vehicle, as well as the relationship between the actual fuel consumption and the average fuel consumption of the target vehicle.
[0182] Fuel consumption types can include high fuel consumption and low fuel consumption.
[0183] It should be noted that the fuel consumption prediction model can specifically be a gradient boosting machine (LightGBM) classification model.
[0184] In this embodiment, firstly, the sample vehicle operation data sets of multiple sample trips are segmented to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The actual fuel consumption of the first positive sample vehicle operation data sets meets a first preset condition, which includes any one of the following: the anomaly index of the actual fuel consumption is greater than an anomaly index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption; or the anomaly index of the actual fuel consumption is less than or equal to an anomaly index threshold. The actual fuel consumption of the first negative sample vehicle operation data sets meets a second preset condition, which includes: the anomaly index of the actual fuel consumption is greater than an anomaly index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption. The fuel consumption is higher than the average fuel consumption. Secondly, feature extraction is performed on multiple sample vehicle operation data sets for each sample trip to obtain sample trip features for each sample trip. Then, based on multiple first positive sample vehicle operation data sets and the sample trip features of each first positive sample vehicle operation data set, multiple second positive sample vehicle operation data sets are obtained. Simultaneously, based on multiple first negative sample vehicle operation data sets and the sample trip features of each first negative sample vehicle operation data set, multiple second negative sample vehicle operation data sets are obtained. Finally, based on the multiple second positive sample vehicle operation data sets and the multiple second negative sample vehicle operation data sets, the fuel consumption prediction model is... Training is performed; this application divides multiple sample vehicle operation data sets into multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets based on the relationship between the anomaly index and the anomaly index threshold of the actual fuel consumption of each sample vehicle operation data set, and the relationship between the actual fuel consumption of each sample vehicle operation data set and the average fuel consumption of the sample trip to which the sample vehicle operation data set belongs. The first positive sample vehicle operation data sets correspond to low fuel consumption, and the first negative sample vehicle operation data sets correspond to high fuel consumption, thus achieving accurate judgment of the high or low fuel consumption of each sample vehicle operation data set. The second positive sample vehicle operation data set... The dataset is obtained based on the first positive sample vehicle operation dataset, and the second negative sample vehicle operation dataset is obtained based on the first negative sample vehicle operation dataset. Therefore, the fuel consumption type corresponding to the second positive sample vehicle operation dataset is low fuel consumption, and the fuel consumption type corresponding to the second negative sample vehicle operation dataset is high fuel consumption. When training the fuel consumption prediction model using multiple second positive sample vehicle operation datasets and multiple second negative sample vehicle operation datasets, since the fuel consumption types of the second positive sample vehicle operation datasets and the second negative sample vehicle operation datasets are known, the training difficulty of the fuel consumption prediction model is effectively reduced, and the prediction accuracy of the fuel consumption prediction model is improved.
[0185] In step 206, the server trains the fuel consumption prediction model based on multiple sets of first positive sample vehicle operation data, multiple sets of first negative sample vehicle operation data, and the sample trip characteristics of each sample trip. Specifically, this includes steps 2061 and 2062.
[0186] Step 2061: The server fuses multiple sets of first positive sample vehicle operation data and multiple sets of first negative sample vehicle operation data with the sample trip features of each sample trip to obtain multiple sets of second positive sample vehicle operation data and multiple sets of second negative sample vehicle operation data.
[0187] Specifically, for each first positive sample vehicle operation data set, the server combines the sample trip features of the sample trip to which the first positive sample vehicle operation data set belongs with the first positive sample vehicle operation data set to obtain a second positive sample vehicle operation data set; for each first negative sample vehicle operation data set, the server combines the sample trip features of the sample trip to which the first negative sample vehicle operation data set belongs with the first negative sample vehicle operation data set to obtain a second negative sample vehicle operation data set.
[0188] It should be noted that, based on the sample trip feature number of each sample trip, each sample trip is concatenated into its corresponding first positive sample vehicle operation data set and first negative sample vehicle operation data set to obtain the second positive sample vehicle operation data set and the second negative sample vehicle operation data set.
[0189] In this embodiment, the second positive sample vehicle operation data set and the second negative sample vehicle operation data set have a wider data dimension than the first positive sample vehicle operation data set and the first negative sample vehicle operation data set. The server trains the fuel consumption prediction model based on the second positive sample vehicle operation data set and the second negative sample vehicle operation data set, which can effectively improve the prediction accuracy of the fuel consumption prediction model.
[0190] Step 2062: The server trains the fuel consumption prediction model based on multiple sets of second positive sample vehicle operation data and multiple sets of second negative sample vehicle operation data.
[0191] The server trains the fuel consumption prediction model using multiple sets of second positive sample vehicle operation data and multiple sets of second negative sample vehicle operation data until the fuel consumption prediction model converges, at which point the trained fuel consumption prediction model can be obtained.
[0192] In this embodiment, each set of second positive sample vehicle operation data and the second set of second negative sample vehicle operation data are input into the fuel consumption prediction model during the training process. The fuel consumption prediction model is adjusted based on the ideal results until it converges. The condition for convergence can be a pre-set number of training rounds or a stopping condition determined during training. The stopping condition can be that the loss function of the fuel consumption prediction model converges to the expected value, or that the loss function reaches a stable value and then shows discrepancies.
[0193] For example, in an embodiment of this application, the training stopping condition for the fuel consumption prediction model is that the loss function of the fuel consumption prediction model converges to 0.9.
[0194] It should be noted that the training process can include transfer learning, multi-task learning, and adversarial training, performing data augmentation on multiple sets of second positive sample vehicle operation data and multiple sets of second negative sample vehicle operation data. Transfer learning uses a model trained on a similar task as the initial model and retrains it on the original task. By sharing the knowledge learned by the model, transfer learning can accelerate the model's learning efficiency and improve its generalization ability. Multi-task learning also uses a model trained on a similar task as the initial model and retrains it on the original task. By sharing the knowledge learned by the model, transfer learning can accelerate the model's learning efficiency and improve its generalization ability. Data augmentation includes a series of techniques for generating new training samples. These techniques achieve this by randomly jittering and perturbing the original data while keeping the class labels unchanged. The goal of applying data augmentation is to increase the model's generalization ability. Adversarial training is an important method for enhancing model robustness. During adversarial training, small perturbations are added to multiple sets of second positive sample vehicle operation data and multiple sets of second negative sample vehicle operation data to cause the fuel consumption prediction model to make mistakes. This allows the fuel consumption prediction model to adapt to the perturbations during training, thereby enhancing the robustness of the fuel consumption prediction model.
[0195] Furthermore, after the fuel consumption prediction model is trained, it can be tested using a test set. The test set can be obtained by using a predetermined proportion of sample trips from the multiple sample trips obtained in step 2013 as the test set, and the remaining sample trips (excluding those in the test set) as training samples to execute steps 202-206 to train the fuel consumption prediction model. The predetermined proportion can be set by the developer, and this embodiment does not impose a specific limitation. For example, the predetermined proportion can be 10%.
[0196] In this embodiment, for each of the multiple first positive sample vehicle operation data sets, the sample trip features of the sample trip to which the first positive sample vehicle operation data set belongs are combined with the first positive sample vehicle operation data set to obtain a second positive sample vehicle operation data set. For each of the multiple first negative sample vehicle operation data sets, the sample trip features of the sample trip to which the first negative sample vehicle operation data set belongs are combined with the first negative sample vehicle operation data set to obtain a second negative sample vehicle operation data set. The fuel consumption prediction model is trained using the multiple second positive sample vehicle operation data sets and the multiple second negative sample vehicle operation data sets. The training effect is good, and the prediction accuracy of the trained fuel consumption prediction model is high.
[0197] In one possible implementation, the method 200 further includes steps 2081 to 2084.
[0198] Step 2081: The server obtains a set of multiple first vehicle operation data for the target vehicle during the target journey.
[0199] The target trip can be either the trip that the target vehicle is currently undertaking or the trip that the target vehicle has already completed; the multiple sets of first vehicle operation data for the target trip are collected by the server at different collection times.
[0200] Step 2082: The server extracts features from multiple sets of first vehicle operation data to obtain the travel features of the target trip.
[0201] The method by which the server obtains the trip characteristics of the target trip in step 2082 is the same as the method by which the server obtains the sample trip characteristics of each sample trip in step 204, and will not be described again in this embodiment.
[0202] Step 2083: The server merges the trip features with each first vehicle operation data set to obtain multiple second vehicle operation data sets.
[0203] Specifically, for each of the multiple first vehicle operation data sets, the server combines the trip characteristics of the target trip obtained in step 2082 with the first vehicle operation data set to obtain the second vehicle operation data set.
[0204] Step 2084: Based on the fuel consumption prediction model after training and multiple sets of second vehicle operation data, the server determines the fuel consumption type of the target vehicle during the target trip.
[0205] The server inputs multiple sets of second vehicle operation data into the trained fuel consumption prediction model. The trained fuel consumption prediction model predicts the results of the multiple sets of second vehicle operation data and outputs the fuel consumption type of each set of second vehicle operation data. Then, the fuel consumption type of the target trip is determined based on the number of second vehicle operation data sets with high fuel consumption and the number of second vehicle operation data sets with low fuel consumption.
[0206] It should be noted that if the number of second vehicle operation data sets with high fuel consumption is greater than the number of second vehicle operation data sets with low fuel consumption, the fuel consumption type of the target trip is determined to be high fuel consumption; if the number of second vehicle operation data sets with high fuel consumption is less than or equal to the number of second vehicle operation data sets with low fuel consumption, the fuel consumption type of the target trip is determined to be low fuel consumption.
[0207] In this embodiment, after obtaining the trained fuel consumption prediction model, firstly, multiple sets of first vehicle operation data for the target vehicle during the target trip are acquired, and features are extracted from the multiple sets of first vehicle operation data to obtain the trip features of the target trip; secondly, the trip features of the target trip are combined with each of the multiple sets of first vehicle operation data to obtain multiple sets of second vehicle operation data; finally, the multiple sets of second vehicle operation data are input into the trained fuel consumption prediction model, and the fuel consumption prediction model outputs the fuel consumption type of the target trip, thereby achieving prediction of the fuel consumption type of the target vehicle during the target trip with high accuracy.
[0208] Figure 3 This is a schematic diagram of the structure of a training device for a fuel consumption prediction model provided in an embodiment of this application. For example, as shown... Figure 3 As shown, the device 300 includes a sample segmentation module 301, a first feature extraction module 302, and a training module 303.
[0209] The sample segmentation module 301 is used to segment multiple sample vehicle operation data sets of multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The different sample vehicle operation data sets of each sample trip are collected at different time points. Each sample vehicle operation data set includes multiple types of sample vehicle operation data. The actual fuel consumption of the first positive sample vehicle operation data set meets a first preset condition, and the actual fuel consumption of the first negative sample vehicle operation data set meets a second preset condition.
[0210] The first feature extraction module 302 is used to extract features from the set of multiple sample vehicle operation data for each sample trip to obtain the sample trip features for each sample trip.
[0211] Training module 303 is used to train the fuel consumption prediction model based on multiple sets of first positive sample vehicle operation data, multiple sets of first negative sample vehicle operation data, and sample trip features of each sample trip. The fuel consumption prediction model is used to predict the fuel consumption type, which represents the abnormal index and abnormal index threshold of the actual fuel consumption of the target vehicle, as well as the relationship between the actual fuel consumption and the average fuel consumption of the target vehicle.
[0212] The first preset condition includes any one of the following: the abnormal index of actual fuel consumption is greater than the abnormal index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption; the abnormal index of actual fuel consumption is less than or equal to the abnormal index threshold.
[0213] The second preset condition includes: the abnormal index of actual fuel consumption is greater than the abnormal index threshold, and the actual fuel consumption is greater than the average fuel consumption.
[0214] Among them, the anomaly index is used to indicate the degree of outlier of the actual fuel consumption relative to the actual fuel consumption of other sample vehicle operation data sets in multiple sample vehicle operation data sets, and the average fuel consumption is the average of multiple actual fuel consumptions of the sample trip to which the actual fuel consumption belongs.
[0215] In one possible implementation, the device 300 further includes a first acquisition module, a filtering module, and a segmentation module.
[0216] The first acquisition module is used to acquire multiple sets of vehicle operation data to be processed from multiple sample vehicles. Each sample vehicle is of the same type as the target vehicle, which is the vehicle whose fuel consumption is to be predicted.
[0217] The filtering module is used to filter the empty vehicle operation data set in multiple vehicle operation data sets to obtain multiple sample vehicle operation data sets. The empty vehicle operation data set is the vehicle operation data set to be processed where the actual fuel consumption and / or mileage of the sample vehicle is empty.
[0218] The partitioning module is used to partition the multiple sample vehicle operation data sets of multiple sample vehicles to obtain multiple sample trips, where each sample trip includes multiple sample vehicle operation data sets.
[0219] In one possible implementation, the sample segmentation module 301 is specifically used to: determine the anomaly index of actual fuel consumption in each sample vehicle operation data set; determine the average fuel consumption of each sample trip; and determine multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets based on the anomaly index of actual fuel consumption in each sample vehicle operation data set and the average fuel consumption of each sample trip.
[0220] In one possible implementation, the sample segmentation module 301 is specifically configured to: for each sample vehicle operation data set, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold, and the actual fuel consumption in the sample vehicle operation data set is greater than the average fuel consumption, determine the sample vehicle operation data set as a first negative sample vehicle operation data set; for each sample vehicle operation data set, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold, and the actual fuel consumption in the sample vehicle operation data set is less than or equal to the average fuel consumption, determine the sample vehicle operation data set as a first positive sample vehicle operation data set; for each sample vehicle operation data set, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is less than or equal to the abnormality index threshold, determine the sample vehicle operation data set as a first positive sample vehicle operation data set.
[0221] In one possible implementation, the first feature extraction module 302 is specifically used to: perform data fusion and data filtering on the same type of sample vehicle operation data in multiple sample vehicle operation data sets for each sample trip, to obtain the sample trip features of the sample trip.
[0222] In one possible implementation, the training module 303 is specifically used to: fuse multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets with the sample trip features of each sample trip to obtain multiple second positive sample vehicle operation data sets and multiple second negative sample vehicle operation data sets; and train the fuel consumption prediction model based on the multiple second positive sample vehicle operation data sets and multiple second negative sample vehicle operation data sets.
[0223] In one possible implementation, the device 300 further includes a second acquisition module, a second feature extraction module, a fusion module, and a fuel consumption prediction module.
[0224] The second acquisition module is used to acquire multiple sets of first vehicle operation data of the target vehicle during the target journey.
[0225] The second feature extraction module is used to extract features from multiple sets of first vehicle operation data to obtain the travel features of the target trip.
[0226] The fusion module is used to fuse the travel features with each first vehicle operation data set to obtain multiple second vehicle operation data sets.
[0227] The fuel consumption prediction module is used to determine the fuel consumption type of the target vehicle during the target trip based on the fuel consumption prediction model after training and multiple sets of second vehicle operation data.
[0228] It should be noted that the training device for the fuel consumption prediction model provided in the above embodiments is only illustrated by the division of the above functional modules during fuel consumption prediction model training. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the fuel consumption prediction model provided in the above embodiments and the training method embodiments for the fuel consumption prediction model belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0229] The technical solution provided in this application firstly involves segmenting multiple sample vehicle operation data sets from multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The actual fuel consumption of the first positive sample vehicle operation data sets meets a first preset condition, which includes any one of the following: the anomaly index of the actual fuel consumption is greater than an anomaly index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption; or the anomaly index of the actual fuel consumption is less than or equal to an anomaly index threshold. The actual fuel consumption of the first negative sample vehicle operation data sets meets a second preset condition, which includes: the anomaly index of the actual fuel consumption is greater than an anomaly index threshold. Furthermore, the actual fuel consumption is greater than the average fuel consumption. Secondly, feature extraction is performed on multiple sample vehicle operation data sets for each sample trip to obtain sample trip features for each sample trip. Then, based on multiple first positive sample vehicle operation data sets and the sample trip features of the sample trips to which each first positive sample vehicle operation data set belongs, multiple second positive sample vehicle operation data sets are obtained. Simultaneously, based on multiple first negative sample vehicle operation data sets and the sample trip features of the sample trips to which each first negative sample vehicle operation data set belongs, multiple second negative sample vehicle operation data sets are obtained. Finally, based on the multiple second positive sample vehicle operation data sets and the multiple second negative sample vehicle operation data sets, fuel consumption is... The prediction model is trained. Based on the relationship between the anomaly index and the anomaly index threshold of the actual fuel consumption of each sample vehicle operation data set, and the relationship between the actual fuel consumption of each sample vehicle operation data set and the average fuel consumption of the sample trip to which that sample vehicle operation data set belongs, this application divides multiple sample vehicle operation data sets into multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The first positive sample vehicle operation data sets correspond to low fuel consumption, and the first negative sample vehicle operation data sets correspond to high fuel consumption, thus achieving accurate judgment of the fuel consumption level of each sample vehicle operation data set. The second positive sample vehicle... The operating data set is obtained based on the first positive sample vehicle operating data set, and the second negative sample vehicle operating data set is obtained based on the first negative sample vehicle operating data set. Therefore, the fuel consumption type corresponding to the second positive sample vehicle operating data set is also low fuel consumption, and the fuel consumption type corresponding to the second negative sample vehicle operating data set is also high fuel consumption. When training the fuel consumption prediction model using multiple second positive sample vehicle operating data sets and multiple second negative sample vehicle operating data sets, since the fuel consumption types of the second positive sample vehicle operating data sets and the second negative sample vehicle operating data sets are known, the training difficulty of the fuel consumption prediction model is effectively reduced, and the prediction accuracy of the fuel consumption prediction model is improved.
[0230] Figure 4 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application.
[0231] For example, such as Figure 4 As shown, the vehicle 400 includes a memory 401 and a processor 402. The memory 401 stores executable program code 4011, and the processor 402 is used to call and execute the executable program code 4011 to perform a training method for a fuel consumption prediction model.
[0232] Furthermore, embodiments of this application also protect an apparatus that may include a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code to perform a training method for a fuel consumption prediction model provided in embodiments of this application.
[0233] This embodiment can divide the device into functional modules based on the above method example. For example, each module can correspond to a separate function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0234] When each functional module is divided according to its corresponding function, the device may further include a sample segmentation module, a first feature extraction module, and a training module. It should be noted that all relevant content regarding the steps involved in the above method embodiments can be referenced to the functional descriptions of the corresponding functional modules, and will not be repeated here.
[0235] It should be understood that the apparatus provided in this embodiment is used to execute the training method of the above-described fuel consumption prediction model, and therefore can achieve the same effect as the above-described implementation method.
[0236] When using an integrated unit, the device may include a processing module and a storage module. When the device is applied to a vehicle, the processing module can be used to control and manage the vehicle's movements. The storage module can be used to support the vehicle in executing program code, etc.
[0237] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits as disclosed in this application. The processor may also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and microprocessors, etc., and the storage module may be a memory.
[0238] In addition, the device provided in the embodiments of this application may specifically be a chip, component or module. The chip may include a connected processor and a memory. The memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute a training method for a fuel consumption prediction model provided in the above embodiments.
[0239] This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the above-described related method steps to implement the training method for a fuel consumption prediction model provided in the above embodiment.
[0240] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement a training method for a fuel consumption prediction model provided in the above embodiment.
[0241] In this embodiment, the device, computer-readable storage medium, computer program product, or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0242] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0243] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0244] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A training method for a fuel consumption prediction model, characterized in that, The method includes: Multiple sample vehicle operation data sets from multiple sample trips are segmented to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The different sample vehicle operation data sets for each sample trip are collected at different time points. Each sample vehicle operation data set includes multiple types of sample vehicle operation data. The actual fuel consumption of the first positive sample vehicle operation data set meets a first preset condition, and the actual fuel consumption of the first negative sample vehicle operation data set meets a second preset condition. Feature extraction is performed on the multiple sample vehicle operation data sets for each sample trip to obtain sample trip features for each sample trip; Based on the multiple sets of first positive sample vehicle operation data, the multiple sets of first negative sample vehicle operation data, and the sample trip characteristics of each sample trip, the fuel consumption prediction model is trained. The fuel consumption prediction model is used to predict the fuel consumption type, and the fuel consumption type is used to represent the abnormal index and abnormal index threshold of the actual fuel consumption of the target vehicle, as well as the relationship between the actual fuel consumption and the average fuel consumption of the target vehicle. The first preset condition includes any one of the following: the abnormal index of actual fuel consumption is greater than the abnormal index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption; the abnormal index of actual fuel consumption is less than or equal to the abnormal index threshold. The second preset condition includes: the abnormality index of the actual fuel consumption is greater than the abnormality index threshold, and the actual fuel consumption is greater than the average fuel consumption; The anomaly index is used to indicate the degree of outlier of the actual fuel consumption relative to the actual fuel consumption of other sample vehicle operation data sets in the multiple sample vehicle operation data sets, and the average fuel consumption is the average of multiple actual fuel consumptions of the sample trip to which the actual fuel consumption belongs. The step of training the fuel consumption prediction model based on the plurality of first positive sample vehicle operation data sets, the plurality of first negative sample vehicle operation data sets, and the sample trip characteristics of each sample trip includes: The multiple sets of first positive sample vehicle operation data and the multiple sets of first negative sample vehicle operation data are respectively fused with the sample trip features of each sample trip to obtain multiple sets of second positive sample vehicle operation data and multiple sets of second negative sample vehicle operation data. The fuel consumption prediction model is trained based on the multiple sets of second positive sample vehicle operation data and the multiple sets of second negative sample vehicle operation data.
2. The method according to claim 1, characterized in that, Before performing sample segmentation on multiple sample vehicle operation data sets from multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets, the method further includes: Acquire multiple sets of vehicle operation data from multiple sample vehicles, wherein each of the sample vehicles is of the same type as the target vehicle, and the target vehicle is the vehicle for which fuel consumption prediction is to be performed. The empty vehicle operation data set in the plurality of vehicle operation data sets to be processed is filtered to obtain a plurality of sample vehicle operation data sets, wherein the empty vehicle operation data set is the vehicle operation data set to be processed where the actual fuel consumption and / or mileage of the sample vehicle is empty. The multiple sample vehicle operation data sets are divided to obtain multiple sample trips, wherein each sample trip includes multiple sample vehicle operation data sets.
3. The method according to claim 1, characterized in that, The step of segmenting multiple sample vehicle operation data sets from multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets includes: Determine the anomaly index of actual fuel consumption in each of the sample vehicle operation data sets; Determine the average fuel consumption for each of the sample trips; Based on the anomaly index of actual fuel consumption in each of the multiple sample vehicle operation data sets, and the average fuel consumption of each sample trip, multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets are determined.
4. The method according to claim 3, characterized in that, The step of determining multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets based on the anomaly index of actual fuel consumption in each of the multiple sample vehicle operation data sets includes: For each of the sample vehicle operation data sets, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold, and the actual fuel consumption in the sample vehicle operation data set is greater than the average fuel consumption, then the sample vehicle operation data set is determined to be the first negative sample vehicle operation data set. For each of the sample vehicle operation data sets, if the abnormality index of the actual fuel consumption in the sample vehicle operation data set is greater than the abnormality index threshold, and the actual fuel consumption in the sample vehicle operation data set is less than or equal to the average fuel consumption, the sample vehicle operation data set is determined to be the first positive sample vehicle operation data set. For each of the sample vehicle operation data sets, if the anomaly index of the actual fuel consumption in the sample vehicle operation data set is less than or equal to the anomaly index threshold, the sample vehicle operation data set is determined as the first positive sample vehicle operation data set.
5. The method according to claim 1, characterized in that, The step of extracting features from the plurality of sample vehicle operation data sets for each sample trip to obtain sample trip features for each sample trip includes: Data fusion and data filtering are performed on the same type of sample vehicle operation data in the multiple sample vehicle operation data sets for each sample trip to obtain the sample trip characteristics.
6. The method according to claim 1, characterized in that, After training the fuel consumption prediction model based on the plurality of first positive sample vehicle operation data sets, the plurality of first negative sample vehicle operation data sets, and the sample trip features of each sample trip, the method further includes: Obtain multiple sets of first vehicle operation data for the target vehicle during the target journey; Feature extraction is performed on the multiple sets of first vehicle operation data to obtain the travel features of the target trip; The travel characteristics are fused with each of the first vehicle operation data sets to obtain multiple second vehicle operation data sets; Based on the fuel consumption prediction model after training and the multiple sets of second vehicle operation data, the fuel consumption type of the target vehicle during the target trip is determined.
7. A training device for a fuel consumption prediction model, characterized in that, The device includes: The sample segmentation module is used to segment multiple sample vehicle operation data sets from multiple sample trips to obtain multiple first positive sample vehicle operation data sets and multiple first negative sample vehicle operation data sets. The different sample vehicle operation data sets for each sample trip are collected at different time points. Each sample vehicle operation data set includes multiple types of sample vehicle operation data. The actual fuel consumption of the first positive sample vehicle operation data set meets a first preset condition, and the actual fuel consumption of the first negative sample vehicle operation data set meets a second preset condition. The first feature extraction module is used to extract features from the multiple sample vehicle operation data sets of each sample trip to obtain sample trip features for each sample trip. The training module is used to train the fuel consumption prediction model based on the multiple first positive sample vehicle operation data sets, the multiple first negative sample vehicle operation data sets, and the sample trip features of each sample trip. The fuel consumption prediction model is used to predict the fuel consumption type, and the fuel consumption type is used to represent the abnormality index and abnormality index threshold of the actual fuel consumption of the target vehicle, as well as the relationship between the actual fuel consumption and the average fuel consumption of the target vehicle. The first preset condition includes any one of the following: the abnormal index of actual fuel consumption is greater than the abnormal index threshold, and the actual fuel consumption is less than or equal to the average fuel consumption; the abnormal index of actual fuel consumption is less than or equal to the abnormal index threshold. The second preset condition includes: the abnormality index of the actual fuel consumption is greater than the abnormality index threshold, and the actual fuel consumption is greater than the average fuel consumption; The anomaly index is used to indicate the degree of outlier of the actual fuel consumption relative to the actual fuel consumption of other sample vehicle operation data sets in the multiple sample vehicle operation data sets, and the average fuel consumption is the average of multiple actual fuel consumptions of the sample trip to which the actual fuel consumption belongs. The training module is specifically used for: The multiple sets of first positive sample vehicle operation data and the multiple sets of first negative sample vehicle operation data are respectively fused with the sample trip features of each sample trip to obtain multiple sets of second positive sample vehicle operation data and multiple sets of second negative sample vehicle operation data. The fuel consumption prediction model is trained based on the multiple sets of second positive sample vehicle operation data and the multiple sets of second negative sample vehicle operation data.
8. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the electronic device to perform the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 6.