A data analysis method for diesel engine fuel consumption test
Through the improved LOF algorithm and sliding window technology, combined with the fluctuation of diesel engine fuel consumption data, the K value is dynamically adjusted, and the abnormal detection and misjudgment problem caused by the difference in fuel consumption data density of diesel engines under different working conditions is solved, achieving more accurate and efficient abnormal detection.
Patent Information
- Application Number
- CN202510415479.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The fuel consumption performance and data density of diesel engines under different operating conditions have differences, resulting in limitations in traditional LOF algorithms, which leads to misjudgment in abnormal detection.
By collecting multi-dimensional diesel engine fuel consumption data, integrating it into a data set, and dividing sliding windows of multiple lengths, calculating the fluctuation of the data in the sliding window, using the improved LOF algorithm to calculate the weighted LOF score of the sample, and dynamically adjusting the K value to adapt to the sparseness of the data set.
It realizes more accurate abnormal detection of diesel engine fuel consumption data, reduces misjudgment, and improves the accuracy and adaptability of detection.
Smart Images

Figure CN119939478B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing. More specifically, the present invention relates to a data analysis method for diesel engine fuel consumption tests. Background Art
[0002] A vehicle diesel engine is a power device based on an internal combustion engine, mainly used in the automotive field, and generates power through the combustion of diesel. Among them, after a vehicle diesel engine undergoes four thermodynamic processes of intake, compression, expansion (work), and exhaust, it can return to the initial state to continuously generate mechanical work. The performance indicators of a vehicle diesel engine mainly include power performance indicators (such as effective torque, effective power, speed, etc.) and fuel consumption rate, and these indicators directly affect the performance and efficiency of the diesel engine. By analyzing the fuel consumption test data, deficiencies in the fuel consumption of the vehicle diesel engine can be discovered, thereby guiding the design optimization of the engine.
[0003] In the abnormal detection of fuel consumption data of vehicle diesel engines, the LOF algorithm can identify abnormal fuel consumption data points that are significantly different from most data points. These abnormal points may be caused by reasons such as engine failures, abnormal driving behaviors, or environmental factors.
[0004] However, there are significant differences in the fuel consumption performance of diesel engines under different working conditions. For example, under heavy load or full load conditions, the diesel engine needs to output maximum power and torque, and the fuel consumption is usually high; while under light load or idle conditions, the fuel consumption is relatively low; in addition, transient conditions with turbocharging intervention, low-temperature starting conditions, and high-altitude or high-altitude conditions will also affect the fuel consumption of the diesel engine. The data density of fuel consumption data under these conditions (i.e., the distribution and aggregation degree of data points) is also different. Normal fuel consumption data under a certain working condition may be misjudged as abnormal points by the LOF algorithm due to the significant difference in its density from the data under other working conditions. Conversely, some data points that seem abnormal under other working conditions are regarded as normal under specific working conditions, resulting in misjudgment phenomena in the abnormal detection of the traditional LOF algorithm. Summary of the Invention
[0005] To solve the above technical problems that the fuel consumption performance and data density of diesel engines vary under different working conditions, resulting in limitations of the traditional LOF algorithm and misjudgment in abnormal detection, the present invention provides the following technical solutions.
[0006] A data analysis method for diesel engine fuel consumption tests, comprising:
[0007] Collect the fuel consumption data of the diesel engine in multiple dimensions at set times, integrate the multi-dimensional fuel consumption data corresponding to the same time into a sample, and sort and integrate the samples corresponding to each collected time into a data set in chronological order;
[0008] The dataset is partitioned into sliding windows of multiple lengths. For a sliding window, the degree of fluctuation of the data within the sliding window is calculated. For a sample, the LOF score of the sample within each sliding window is calculated using an improved LOF algorithm, and in combination with the degree of fluctuation of each sliding window, the weighted LOF score of the sample is obtained through weighted summation; the value of K in the improved LOF algorithm is positively correlated with the sparsity of the dataset;
[0009] Samples with a weighted LOF score greater than or equal to a preset threshold are marked as abnormal samples, and the abnormal samples are processed.
[0010] The present invention first integrates and analyzes multi-dimensional fuel consumption data, partitions the formed dataset into sliding windows of multiple lengths, thereby capturing the local characteristics of the data in different time periods, further calculating the degree of fluctuation of the data within the sliding window, and at the same time dynamically adjusting the value of K of the LOF algorithm according to the change of data density. Then, the LOF score of the sample within each sliding window is calculated using the improved LOF algorithm, and the weighted LOF score of the sample is obtained through weighted summation in combination with the degree of fluctuation of each sliding window, which can more accurately identify abnormal samples. The traditional LOF algorithm may have problems such as inaccurate local density estimation or insufficient sensitivity to anomalies at different scales. The improved algorithm can better adapt to the characteristics of different datasets and improve the accuracy of anomaly detection by considering the positive correlation between the value of K and the sparsity of the dataset and the method of weighted summation.
[0011] Preferably, the multi-dimensional fuel consumption data includes one or more of fuel consumption, rotational speed, torque, and coolant temperature.
[0012] Preferably, the sparsity of the dataset satisfies the relational expression:
[0013] ; where is the sparsity of the dataset, is the number of dimensional fuel consumption data included in the dataset, is the number of samples included in the dataset, is the th dimensional fuel consumption data in the th sample and the mean Euclidean distance between the dimensional fuel consumption data in all other samples, represents normalization processing.
[0014] By calculating the average of the average Euclidean distances of each dimension in the dataset, the sparsity of the dataset can be evaluated. If the sparsity of the dataset is large, it indicates that the samples in the dataset are sparsely distributed in each dimension; if the sparsity of the dataset is small, it indicates that the samples in the dataset are densely distributed in each dimension.
[0015] Preferably, the value of K in the improved LOF algorithm satisfies the relational expression:
[0016] ; where represents the value of K in the improved LOF algorithm, is the preset initial value of K, is the sparsity of the dataset, is the logarithmic function with base e.
[0017] For different datasets, especially datasets with uneven density distributions, a fixed value of K may not be suitable for all situations. By introducing the sparsity of the dataset, the improved value of K can be dynamically adjusted according to the distribution characteristics of the data. When the dataset is sparser, the value of the sparsity of the dataset will be larger, thus increasing the value of K to better capture the outliers in the sparse region.
[0018] Preferably, the partitioning of the dataset into sliding windows of multiple lengths includes:
[0019] The dataset is partitioned into sliding windows of different lengths, with the minimum length being half of the number of samples in the dataset and the maximum length being the number of samples in the dataset.
[0020] By setting the minimum length to half of the number of samples in the dataset, it can be ensured that each window contains enough data points to reflect some basic characteristics of the data and avoid feature loss caused by too short windows; by setting the maximum length to the number of samples in the dataset, it can avoid low computational efficiency and resource waste caused by too long windows. At the same time, this also ensures that each data point is included in at least one window, thus making full use of the dataset.
[0021] Preferably, the process of obtaining the degree of fluctuation includes:
[0022] In the current sliding window, for each dimension of fuel consumption data, calculate the mean of the fuel consumption data of this dimension in the current sliding window. For each sample in the current sliding window, calculate the absolute value of the difference between the fuel consumption data of this dimension and its mean. Sum up the absolute values of the differences of all samples in the current sliding window to obtain the total fluctuation of this dimension in the current sliding window;
[0023] Within the current sliding window, calculate the slope of the fuel consumption data for this dimension, and calculate the average value of the slopes of the fuel consumption data for this dimension across all samples;
[0024] Sum the products of the total fluctuations of the fuel consumption data for each dimension and the average slope, and normalize to obtain the degree of data fluctuation within the sliding window.
[0025] By calculating the absolute value of the difference between the fuel consumption data of each sample and the mean within the current sliding window, and summing to obtain the total fluctuation, this step measures the degree of dispersion of the fuel consumption data for this dimension within the current window. The larger the total fluctuation, the greater the degree of deviation of the data points from the mean, that is, the greater the degree of data fluctuation; calculating the slope of the fuel consumption data and averaging can reflect the change trend of the data within a certain time range. The larger the average slope, the more obvious the change trend of the data.
[0026] Sum the products of the total fluctuations of the fuel consumption data for each dimension and the average slope, and normalize to obtain the overall degree of data fluctuation within the sliding window, comprehensively considering the degree of dispersion and change trend of the data, so as to more comprehensively and accurately evaluate the data fluctuation situation.
[0027] Preferably, the weighted LOF score satisfies the relational expression:
[0028] ; In the formula, is the weighted LOF score of the sample, is the number of sliding windows containing this sample, represents the degree of data fluctuation within the th sliding window, represents the LOF score of this sample in the
[0029] By performing weighted averaging on the LOF scores within multiple sliding windows, a comprehensive, smooth and reliable anomaly score is provided, so as to better capture and reflect the anomaly situations of the sample at different time periods, and at the same time reduce the noise impact caused by data fluctuations.
[0030] Preferably, the processing of abnormal samples includes:
[0031] According to the characteristics of the abnormal sample, conduct fault troubleshooting, and based on the troubleshooting results, adjust or maintain the diesel engine.
[0032] Preferably, the sparsity degree of the data set satisfies the relational expression:
[0033] ; In the formula, is the sparsity degree of the data set, is the number of fuel consumption data dimensions included in the data set, is the information entropy of the fuel consumption data for the th dimensionality.
[0034] Preferably, after collecting the fuel consumption data of the diesel engine under multiple dimensions at the set time, all the fuel consumption data is subjected to standardization or normalization processing.
[0035] The beneficial effects of the present invention are as follows:
[0036] Through multi-dimensional data integration, dynamic anomaly detection, adaptive K-value selection, multiple sliding window partitions, and the combination of the degree of fluctuation and weighted LOF scores, the present invention realizes comprehensive, accurate, and efficient analysis of diesel engine fuel consumption data, providing strong support for the performance evaluation, fault diagnosis, and maintenance of diesel engines. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a flowchart of the method for steps S1 - S3 in a method for analyzing data of a diesel engine fuel consumption test according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0039] The application scenario of the present invention is: using an improved LOF algorithm to perform anomaly detection on the fuel consumption data of a diesel engine.
[0040] Referring to Figure 1 , a method for analyzing data of a diesel engine fuel consumption test includes steps S1 - S3, specifically as follows:
[0041] S1: Collect the fuel consumption data of the diesel engine under multiple dimensions at the set time, integrate the multi-dimensional fuel consumption data corresponding to the same time into a sample, and sort and integrate the samples corresponding to each collected time into a dataset in chronological order.
[0042] In one embodiment, a number of key data of the diesel engine are collected simultaneously according to the set time, including but not limited to: fuel consumption, rotational speed, torque, coolant temperature, exhaust temperature, intake pressure, and exhaust flow rate. Specifically, a fuel consumption meter is used to measure the fuel consumption of the diesel engine, a rotational speed sensor is used to monitor the rotational speed of the diesel engine, a torque sensor is used to measure the torque of the diesel engine, a temperature sensor is used to record the coolant temperature, a thermocouple is used to monitor the exhaust temperature, an intake pressure sensor is used to measure the intake pressure, and a vortex flowmeter is used to measure the exhaust flow rate.
[0043] Integrate the above 7 characteristic data collected at the same moment into one sample. Each sample contains a complete description of the working state of the diesel engine at that moment. Arrange the samples corresponding to each moment in sequence according to the collection time order. Since the 7 characteristic data collected have different units and formats, in order to conduct unified analysis and processing, we need to standardize these data. Standardization processing usually includes data cleaning (removing outliers, filling missing values, etc.), data transformation (such as logarithmic transformation, Z-score standardization, etc.), and data normalization (scaling the data to a specific range, such as between 0 and 1). The data after standardization processing are integrated into a structured dataset for diesel engine fuel consumption test.
[0044] To sum up, this dataset not only contains the fuel consumption data of the diesel engine at each moment, but also contains information in multiple dimensions such as rotational speed, torque, temperature, pressure, and flow rate related to it, and reflects the dynamic changes of the diesel engine during the entire test working process.
[0045] S2: Divide the dataset into sliding windows of various lengths. For a sliding window, calculate the degree of data fluctuation within the sliding window. For a sample, use the improved LOF algorithm to calculate its LOF scores within various sliding windows, and combine the degree of fluctuation of each sliding window to obtain the weighted LOF score of the sample through weighted summation; the K value in the improved LOF algorithm is positively correlated with the sparsity of the dataset.
[0046] In the diesel engine fuel consumption test, the fuel consumption data densities collected in different test stages (such as cold start, steady state, transient state) are different. For example, the data density in the steady state stage is large, the data is relatively stable, and the change is small; the data density in the transient state stage is small, the data changes greatly, and the fluctuation is obvious; the data in the cold start stage may be between the steady state and the transient state, or it may have its own unique fluctuation.
[0047] The traditional LOF algorithm uses a fixed K value (i.e., the number of nearest neighbors) for outlier detection. However, this method with a fixed K value has the following problems in scenarios with different data densities:
[0048] In areas with large data density (such as the steady state stage), a larger K value may cause the LOF scores of outlier data or noise to be too low, so that they are misidentified as normal data; in areas with small data density (such as the transient state stage), a smaller K value may cause the LOF scores of normal data to be too high, so that they are misidentified as outlier data.
[0049] Therefore, calculate the sparsity of the data to adaptively adjust the K value. The sparser the dataset (i.e., the smaller the density of the dataset), the larger the K value should be used to improve the accuracy of outlier detection.
[0050] Specifically, first, for the fuel consumption data of each dimension and each sample in the dataset, analyze the degree of difference of the fuel consumption data of each dimension among different samples, and then quantify the average difference of the fuel consumption data of the dimension among different samples, and average the average sparsity of all features to obtain the sparsity of the entire dataset.
[0051] In one embodiment, the sparsity of the dataset satisfies the relationship:
[0052]
[0053] In the formula, is the sparsity of the dataset, is the number of fuel consumption data of dimensions included in the dataset, is the number of samples included in the dataset, is the th th fuel consumption data of the th dimension in the th sample and the mean Euclidean distance between the th fuel consumption data of the
[0054] where reflects the degree of difference of the th fuel consumption data of the dimension among different samples, reflects the average difference of the th fuel consumption data of the dimension among different samples, is to average the average sparsity of all features to obtain the sparsity of the entire dataset.
[0055] The greater the sparsity, the smaller the density of the dataset. This means that the feature values in the dataset vary greatly among different samples and the data points are more dispersed; conversely, the smaller the sparsity, the greater the density of the dataset. This means that the feature values in the dataset vary less among different samples and the data points are more concentrated.
[0056] In another embodiment, the sparsity of the dataset also satisfies the relationship:
[0057]
[0058]
[0059] In the formula, is the sparsity of the dataset, is the number of fuel consumption data of dimensions included in the dataset, is the th information entropy of the fuel consumption data of the dimension, is the number of samples included in the dataset, For the th dimension fuel consumption data in the th sample and the th dimension fuel consumption data in all other samples, the average Euclidean distance is is the natural logarithm function.
[0060] Furthermore, the K value in the LOF algorithm is adaptively adjusted according to the sparsity of the dataset, and then the improved LOF algorithm is obtained. The K value in the improved LOF algorithm satisfies the relational expression:
[0061]
[0062] In the formula, represents the K value in the improved LOF algorithm, is the preset initial K value, is the sparsity of the dataset, is the natural logarithm function.
[0063] Among them, by multiplying , the initial K value can be appropriately adjusted according to the sparsity of the dataset. If the dataset is very sparse (i.e., its sparsity is large), then will also be large, so that the final value is also larger; on the contrary, when the dataset is not so sparse, is close to 1, making the final value close to the initial K value.
[0064] In another embodiment, the K value in the improved LOF algorithm also satisfies the relational expression:
[0065]
[0066] In the formula, represents the K value in the improved LOF algorithm, is the preset initial K value, is the sparsity of the dataset, and are both constants obtained by fitting the dataset.
[0067] After obtaining the improved LOF algorithm, the collected dataset is divided into sliding windows of different lengths, the fluctuation degree of the data in each sliding window is calculated, and weighted according to the fluctuation degree and the LOF score of the samples in the sliding window to obtain the anomaly score of the samples.
[0068] Specifically, first set the length of the sliding window, the minimum length is half of the number of samples in the dataset, and the maximum length is the number of samples in the dataset.
[0069] Then, within the current sliding window, for each dimension's fuel consumption data, calculate the mean of the fuel consumption data for that dimension within the current sliding window. For each sample in the current sliding window, calculate the absolute value of the difference between the fuel consumption data for that dimension and its mean. Sum up the absolute values of the differences for all samples within the current sliding window to obtain the total fluctuation of that dimension within the current sliding window. Within the current sliding window, calculate the slope of the fuel consumption data for that dimension (i.e., the change in the fuel consumption data for that dimension between adjacent samples), and further calculate the average value of the slopes of the fuel consumption data for that dimension over all samples.
[0070] Furthermore, measure the degree of data fluctuation within the sliding window based on the slope of the fuel consumption data for each dimension and the absolute difference between the fuel consumption data for that dimension within the sliding window and the mean of the fuel consumption data for that dimension in all samples. That is, the relationship is satisfied as follows:
[0071]
[0072] In the formula, is the degree of data fluctuation within the sliding window, is the number of dimensions of fuel consumption data included in the dataset, is the value of the -th sample in the sliding window for the -th dimension of fuel consumption data, is the mean of the -th dimension of fuel consumption data within the sliding window, is the average value of the slopes of the -th dimension of fuel consumption data within the sliding window, is the number of samples included in the sliding window, represents normalization processing.
[0073] The above degree of fluctuation reflects the magnitude of data variation within the sliding window. When the degree of fluctuation is large, it means that there are large variations in the data within the sliding window. This may cause abnormal samples to be far from the dense region, resulting in their local density being much lower than that of their neighbors, thus leading to a high LOF score, and normal samples may be misjudged as abnormal samples. When the degree of fluctuation is small, it means that the data changes little within the sliding window, and abnormal samples may still be in a relatively dense region, making their local density similar to that of their neighbors, thus resulting in a low LOF score, and abnormal samples may be misjudged as normal samples.
[0074] According to the above formula for calculating the degree of data fluctuation within the sliding window, the degree of data fluctuation within each sliding window can be calculated in the same way.
[0075] In one embodiment, for the position of each sample in each sliding window, its LOF (Local Outlier Factor) score is calculated, and then the LOF score of each sample is adjusted according to the degree of fluctuation of the data within each sliding window to obtain the weighted LOF score of the sample, that is, the relational expression is satisfied as:
[0076]
[0077] In the formula, is the weighted LOF score of the sample, is the number of sliding windows containing the sample, represents the th degree of fluctuation of the data within the sliding window, represents the LOF score of the sample in the th sliding window.
[0078] When the degree of fluctuation is larger, is smaller, so the weight of the LOF score of this sliding window is also smaller. This means that the area with large fluctuations (which may be the normal area) has less influence on the final anomaly score, avoiding misjudging normal samples as anomalies. When the degree of fluctuation is smaller, is larger, so the weight of the LOF score of this sliding window is also larger. This means that the area with small fluctuations (which may be the abnormal area) has a greater influence on the final anomaly score, avoiding missing the judgment of abnormal samples.
[0079] According to the above calculation formula of the weighted LOF score of the sample, the weighted LOF scores of all samples can be calculated.
[0080] S3: Mark the samples with weighted LOF scores greater than or equal to the preset threshold as abnormal samples and process the abnormal samples.
[0081] In one embodiment, the threshold is set to 1.5, and the samples with weighted LOF scores greater than or equal to the preset threshold are marked as abnormal samples. Further, the detailed information of the abnormal samples is recorded, including the timestamp, abnormal features, etc. According to the features of the abnormal samples, fault troubleshooting is carried out. For example, check whether the sensors of the diesel engine are faulty and whether the fuel system is abnormal, etc.
[0082] According to the troubleshooting results, corresponding adjustments or maintenance are carried out on the diesel engine. For example, replace faulty components, adjust fuel injection parameters, etc.
[0083] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of this invention patent shall be subject to the appended claims.
Claims
1. A data analysis method for diesel engine fuel consumption test, characterized in that: include: Collect the fuel consumption data of the diesel engine in multiple dimensions at a set time, integrate the multi-dimensional fuel consumption data corresponding to the same time into one sample, and sort the samples corresponding to each collected time into one data set in chronological order; The data set is divided into sliding windows of various lengths. For a sliding window, the fluctuation degree of the data in the sliding window is calculated. For a sample, the LOF score in each sliding window is calculated using an improved LOF algorithm, and the weighted LOF score of the sample is obtained by weighted summation in combination with the fluctuation degree of each sliding window. The K value in the improved LOF algorithm satisfies the relationship: ; In the formula, represents the K value in the improved LOF algorithm, is the preset initial K value, is the sparsity of the dataset, It is a logarithmic function with the natural constant e as the base; The sparsity of the data set satisfies the relationship: ; In the formula, is the sparsity of the dataset, is the number of dimensional fuel consumption data contained in the dataset, is the number of samples contained in the data set, For the The first The fuel consumption data of the dimension is compared with the first The mean Euclidean distance between the fuel consumption data of each dimension, Indicates normalization processing; The samples whose weighted LOF scores are greater than or equal to the preset threshold are marked as abnormal samples and processed.
2. A diesel engine fuel consumption test data analysis method according to claim 1, characterized in that: The multi-dimensional fuel consumption data includes one or more of fuel consumption, rotation speed, torque, and coolant temperature.
3. A diesel engine fuel consumption test data analysis method according to claim 2, characterized in that: The step of dividing the data set into sliding windows of various lengths comprises: The data set is divided into sliding windows of different lengths, with the minimum length being half the number of samples in the data set and the maximum length being the number of samples in the data set.
4. A diesel engine fuel consumption test data analysis method according to claim 3, characterized in that: The process of obtaining the fluctuation degree includes: In the current sliding window, for each dimension of fuel consumption data, calculate the mean value of the dimension of fuel consumption data in the current sliding window. For each sample in the current sliding window, calculate the absolute value of the difference between the dimension of fuel consumption data and its mean value. Sum the absolute values of the differences of all samples in the current sliding window to obtain the sum of fluctuations of the dimension in the current sliding window. In the current sliding window, the slope of the fuel consumption data of this dimension is calculated, and the average slope of the fuel consumption data of this dimension over all samples is calculated; The product of the sum of the fluctuations of the fuel consumption data of each dimension and the mean of the slope is summed and normalized to obtain the degree of fluctuation of the data in the sliding window.
5. A diesel engine fuel consumption test data analysis method according to claim 4, characterized in that: The weighted LOF score satisfies the relationship: ; In the formula, is the weighted LOF score of the sample, is the number of sliding windows that contain the sample, Indicates The degree of fluctuation of the data in a sliding window Indicates that the sample is in LOF score of a sliding window.
6. A diesel engine fuel consumption test data analysis method according to claim 5, characterized in that: Processing of abnormal samples includes: According to the characteristics of abnormal samples, fault troubleshooting is carried out, and according to the troubleshooting results, the diesel engine is adjusted or maintained.
7. A diesel engine fuel consumption test data analysis method according to claim 2, characterized in that: The sparsity of the data set satisfies the relationship: ; In the formula, is the sparsity of the dataset, is the number of dimensional fuel consumption data contained in the dataset, For the The information entropy of fuel consumption data in three dimensions.
8. A diesel engine fuel consumption test data analysis method according to claim 1, characterized in that: After collecting the fuel consumption data of the diesel engine in multiple dimensions at a set time, all the fuel consumption data are standardized or normalized.
Citation Information
Patent Citations
Data stream anomaly detection method based on local vector dot product density
CN108667684A
Network security detection method and system
CN117914629A