A Vehicle-mounted Data Processing Method and System

Through the isolated forest algorithm, the accuracy of the abnormal monitoring of multi-dimensional vehicle data is solved by using feature combination and density feature evaluation, and more efficient abnormal data screening and processing is achieved.

CN120105311BActive Publication Date: 2025-07-18MAIWEI TECH (GUANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510570898.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-18
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

When building an abnormality detection model in the prior art, the vehicle-related data is large and multi-dimensional, resulting in low accuracy of the abnormality monitoring results of on-board data.

Method used

The isolated forest algorithm is used to process on-vehicle data, and by obtaining the minimum correlation between each feature, the isolated forest is constructed, the density characteristics and weights are analyzed, and the abnormal scores of the data points are calculated.

Benefits of technology

It improves the accuracy of vehicle-mounted data abnormality monitoring, reduces errors caused by a single feature segmentation point, and ensures accurate positioning and timely processing of abnormal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105311B_ABST
    Figure CN120105311B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and in particular, to a vehicle-mounted data processing method and system. The method includes the steps of: taking any feature of the data points in the vehicle-mounted data set as a target feature, and obtaining a correlation sequence of the time series of the target feature; calculating a combined value of the target feature of the data points, and selecting a segmentation point in the time series of the combined value of the target feature to construct an isolation forest of the target feature; obtaining a density feature of the isolation forest of the target feature through the difference between each data point in the isolation forest of the target feature and the median; obtaining the weight of the isolation forest of the target feature through the density feature of the isolation forest of the target feature; obtaining the average path length of the data point in each isolation forest of the target feature, and obtaining the anomaly score of the data point, so as to realize vehicle-mounted data processing, effectively improving the accuracy of vehicle-mounted data anomaly monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a vehicle data processing method and system. Background Art

[0002] Vehicle intelligence refers to improving the performance of automobiles to a new level through cutting-edge information technology, data communication, sensor technology, control technology, computer technology, etc. Among them, vehicle networking technology enables automobiles to communicate and exchange data with other vehicles, infrastructure, network services, etc. through the combination of vehicles and the Internet, so as to realize the abnormal monitoring of vehicle data.

[0003] The prior art focuses on constructing a neural network monitoring model for abnormal monitoring of vehicle data. For example, the patent application document with the publication number CN117292462A discloses an intelligent processing method and system for vehicle abnormal information. This application monitors the running information of the vehicle in real time; constructs an abnormal detection model based on a machine learning algorithm to analyze the running information to obtain the abnormal information of the vehicle; classifies and processes the abnormal information to determine the abnormal type; determines the corresponding abnormal handling scheme based on the abnormal type and executes the abnormal handling scheme.

[0004] The above prior art can realize the abnormal monitoring of vehicle data by constructing an abnormal detection model. However, the amount of vehicle-related data is relatively large and usually multi-dimensional data. During the process of constructing the abnormal detection model, the data processing amount is large and it is easily affected by the curse of dimensionality, resulting in a low accuracy of the obtained abnormal monitoring results.

[0005] Based on this, how to accurately obtain the abnormal monitoring results of vehicle data is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0006] In order to solve the technical problem of how to accurately obtain the abnormal monitoring results of vehicle data, the present invention provides a vehicle data processing method and system.

[0007] In a first aspect, the present invention provides a vehicle data processing method, adopting the following technical solution:

[0008] Take any feature of the data points in the vehicle-mounted dataset as the target feature, and take the other feature time series with the smallest absolute value of the Pearson correlation coefficient between the time series of the target feature as the correlation series of the target feature; calculate the combined value of the target feature of the data point according to the values of the data point in the time series of the target feature and the correlation series, and the importance of the time series of the target feature of the data point to the correlation series and the time series pair of the correlation series; form the combined value time series corresponding to the same type of target features, and select a split point in the combined value time series of the target feature to construct the target feature isolation forest; obtain the density feature of the target feature isolation forest through the difference between each data point in the target feature isolation forest and the median; obtain the weight of the target feature isolation forest through the density feature of the target feature isolation forest, and the weight of the target feature isolation forest is negatively correlated with the density feature; obtain the product of the average path length and the weight of the data point in each target feature isolation forest, and normalize the mean of all products to obtain the anomaly score of the data point; realize vehicle-mounted data processing based on the comparison result between the anomaly score of the data point and the preset threshold.

[0009] The present invention processes multi-dimensional vehicle-mounted data through the isolation forest algorithm, and can accurately screen out abnormal data from the vehicle-mounted dataset. In this process, the present invention takes into account that the correlations between different features of the vehicle-mounted data are different, but the isolation forest algorithm only uses a single feature when selecting a split point each time, resulting in unbalanced feature use and affecting the accuracy of the construction of the isolation tree. Based on this, the present invention combines each feature with the feature with the smallest correlation, and completely uses the combined value of this feature as the split point to construct the corresponding feature isolation forest, and analyzes the density feature of the feature isolation forest to evaluate the isolation effect to realize the weighting of the average path length of the data points, reducing the influence of the single feature splitting the isolation tree on the abnormal monitoring result, accurately obtaining the anomaly score of the data point, and effectively improving the accuracy of vehicle-mounted data abnormal monitoring.

[0010] According to a vehicle-mounted data processing method provided by the present invention, before taking any feature of the data points in the vehicle-mounted dataset as the target feature, it further includes: preprocessing the multiple vehicle-mounted feature values obtained at each acquisition moment as a data point to obtain the data points in the vehicle-mounted dataset.

[0011] The present invention takes into account that there may be situations such as data missing in the originally collected vehicle-mounted data, so the overall quality of the data is improved through preprocessing to facilitate subsequent data processing.

[0012] According to a vehicle-mounted data processing method provided by the present invention, the method for obtaining the importance of the relevant sequence of the target feature to the time series sequence includes: taking the ratio of the Pearson correlation coefficient between the relevant sequence of the target feature and the time series sequence to the sum of the Pearson correlation coefficients between the time series sequence of the target feature and all feature time series sequences as the importance of the relevant sequence of the target feature to the time series sequence.

[0013] According to a vehicle-mounted data processing method provided by the present invention, calculating the combined value of the target feature of the data point includes:

[0014] ; is the combined value of the target feature of the i-th data point, , are respectively the importance of the relevant sequence of the target feature to the time series sequence and the importance of the time series sequence to the relevant sequence, , are respectively the values of the i-th data point in the time series sequence and the relevant sequence of the target feature.

[0015] According to a vehicle-mounted data processing method provided by the present invention, selecting a splitting point in the combined value time series sequence of the target feature to construct an isolation forest of the target feature includes: randomly selecting a preset number of data points in the vehicle-mounted data set as a sub-data set; randomly selecting the combined value of the target feature of a data point in the sub-data set as the splitting point, and dividing the sub-data set into a left branch and a right branch; continuing to randomly select splitting points in the branches for division until a preset termination condition is reached to obtain an isolation tree corresponding to the sub-data set; and forming an isolation forest of the target feature by combining multiple isolation trees.

[0016] The present invention obtains an isolation forest of the target feature by using the combined value of a single target feature as the splitting point, so that the obtained isolation forest of the target feature completely depends on the data performance of the combined value itself, and the corresponding weights can be accurately calculated on the corresponding target feature when analyzing the isolation results of the isolation forest of the target feature subsequently.

[0017] According to a vehicle-mounted data processing method provided by the present invention, the density feature of the isolation forest of the target feature satisfies the relational expression: ; is the density feature of the isolation forest of the target feature, is the number of data points in the isolation forest of the target feature, is the median of the combined values of the target feature in the isolation forest of the target feature, is the combined value of the target feature of the j-th data point in the isolation forest of the target feature, is the exponential function with base e, is the absolute value symbol.

[0018] The present invention provides an accurate method for calculating the density feature of the target feature isolation forest. By obtaining the difference between each data point and the median, the proximity between each data point and the median can be accurately obtained. The smaller the difference, the higher the proximity, and the greater the corresponding density feature.

[0019] According to a vehicle-mounted data processing method provided by the present invention, the method for obtaining the weight of the target feature isolation forest includes: taking the sum of the importance of the relevant sequence of the target feature to the time series and the importance of the time series to the relevant sequence as the feature importance of the target feature isolation forest corresponding to the target feature; ;

[0020] is the weight of the target feature isolation forest, is the natural constant, is the density feature of the target feature isolation forest, 、 are respectively the importance of the relevant sequence of the target feature to the time series and the importance of the time series to the relevant sequence, is the maximum value of the feature importance of all target feature isolation forests, is the linear normalization function.

[0021] According to a vehicle-mounted data processing method provided by the present invention, the vehicle-mounted data processing based on the comparison result between the anomaly score of the data point and the preset threshold includes: if the anomaly score of the data point is greater than the preset threshold, then the data point is abnormal vehicle-mounted data; otherwise, the data point is normal vehicle-mounted data.

[0022] According to a vehicle-mounted data processing method provided by the present invention, after the vehicle-mounted data processing is implemented, it further includes: in response to the data point being abnormal vehicle-mounted data, prompting the abnormal vehicle-mounted data.

[0023] The present invention takes into account that abnormal vehicle-mounted data may affect the safety of the vehicle and the users. Therefore, a prompt is sent out in time after detecting abnormal vehicle-mounted data for timely processing.

[0024] In a second aspect, the present invention provides a vehicle-mounted data processing system, adopting the following technical solution:

[0025] A vehicle-mounted data processing system includes: a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned vehicle-mounted data processing method is implemented.

[0026] By adopting the above technical solution, a computer program is generated from the above-mentioned vehicle-mounted data processing method and stored in a memory to be loaded and executed by a processor, so as to manufacture a terminal device according to the memory and the processor, which is convenient to use.

[0027] The present invention has the following technical effects:

[0028] Based on the above technical solution, when processing vehicle-mounted data, the present invention processes multi-dimensional vehicle-mounted data through the isolation forest algorithm, and can screen out abnormal data from the vehicle-mounted data set. In this process, the present invention takes into account that the correlations between different features of vehicle-mounted data are different, but the isolation forest algorithm only uses a single feature when selecting a splitting point each time, resulting in unbalanced feature use and affecting the accuracy of the construction of isolation trees. Based on this, the present invention combines each feature with the feature having the smallest correlation, and completely uses the combined value of the features as the splitting point to construct the corresponding feature isolation forest, and analyzes the density feature of the feature isolation forest to evaluate the isolation effect to realize the weighting of the average path length of data points, so as to accurately obtain the abnormal score of data points, effectively improving the accuracy of vehicle-mounted data anomaly monitoring. Description of the Drawings

[0029] By reading the following detailed description with reference to the drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become easily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts.

[0030] Figure 1 It is a schematic flowchart of a vehicle-mounted data processing method provided by an embodiment of the present invention. Detailed Embodiments

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0032] It should be understood that when the claims, specifications and drawings of the present invention use terms such as "first" and "second", they are only used to distinguish different objects, rather than to describe a specific order. The terms "including" and "comprising" used in the specifications and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0033] Vehicle intelligence refers to enhancing various performance of automobiles to a new level through cutting-edge information technology, data communication, sensor technology, control technology, computer technology, etc. Among them, the vehicle networking technology enables the vehicle to communicate and exchange data with other vehicles, infrastructure, network services, etc. by combining the vehicle with the Internet, so as to realize the abnormal monitoring of in-vehicle data.

[0034] In the prior art, the method of constructing a neural network monitoring model for in-vehicle data abnormal monitoring has a large amount of data processing and is easily affected by the curse of dimensionality, resulting in low accuracy of the obtained abnormal monitoring results.

[0035] The Isolation Forest algorithm is an unsupervised learning abnormal monitoring algorithm. It isolates abnormal points by randomly splitting the data space and can effectively process data sets with a large number of features without being affected by the curse of dimensionality.

[0036] Based on this, the embodiments of the present invention disclose a method for processing in-vehicle data. This method processes the in-vehicle data set through the Isolation Forest algorithm, can accurately identify the abnormal data points in the in-vehicle data set, and effectively improves the accuracy of in-vehicle data processing. Specifically, refer to Figure 1 as shown Figure 1 which is a schematic flowchart of a method for processing in-vehicle data provided by the embodiments of the present invention. This method specifically includes the following steps.

[0037] S1: Obtain the features corresponding to each data point in the in-vehicle data set.

[0038] Among them, the features of the in-vehicle data can be the engine speed of the vehicle, the displacement, speed and angular velocity of the vehicle, etc., and can be specifically set according to actual needs. The embodiments of the present invention do not limit this too much here.

[0039] Specifically, the engine speed of the vehicle is obtained by installing a speed sensor on the engine; the acceleration, speed and displacement of the vehicle are obtained by using a satellite positioning system; the angular velocity of the vehicle is obtained by an inertial measurement unit.

[0040] Among them, the acquisition frequency can be set to 30Hz, and the acquisition duration can be 6 hours; the acquisition duration and acquisition frequency can be specifically set according to actual needs.

[0041] Exemplarily, in the embodiments of the present invention, obtaining the features corresponding to each data point in the in-vehicle data set includes: preprocessing multiple in-vehicle feature values obtained at each acquisition moment as a data point to obtain the data points in the in-vehicle data set, and each data point corresponds to multiple features.

[0042] Among them, the preprocessing can be wavelet transform data denoising, missing data interpolation, data format conversion, etc.; the specific preprocessing method can be set according to actual needs, and the embodiments of the present invention do not limit too much here.

[0043] It can be understood that the values of each vehicle feature at the current moment are collected at each collection moment, and the values of each feature are collected in time sequence. Therefore, the same type of features can be arranged in the order of the collection time sequence to obtain the time sequence of the feature, and the feature values in the time sequence of each feature correspond one by one.

[0044] Based on the above steps, a vehicle-mounted data set and the time sequences of each feature in the vehicle-mounted data set can be obtained. However, in the process of dividing the data set by the conventional isolation forest algorithm, random data in one of the dimensional features is randomly selected for division, so as to obtain an isolation tree of a single feature. In the process of randomly selecting the division point, some features will be overused while some features will be ignored, resulting in low accuracy of the isolated abnormal data.

[0045] Based on this, the embodiments of the present invention construct a multi-feature isolation tree to obtain an isolation forest by obtaining the correlation between features, so as to reduce the error caused by dividing with a single feature and improve the accuracy of anomaly detection, that is, continue to execute the following steps.

[0046] S2: Take any feature of the data points in the vehicle-mounted data set as the target feature, and take the time sequence of other features with the smallest absolute value of the Pearson correlation coefficient between the time sequences of the target feature as the correlation sequence of the target feature.

[0047] Among them, the specific steps of obtaining the Pearson correlation coefficient between the time sequences of each feature can be realized by the prior art, and the embodiments of the present invention do not elaborate here.

[0048] It should be noted that in the process of randomly selecting features as the division point by the isolation forest algorithm, if there is a high correlation between some features in the data set, then such features may show similar patterns in the data distribution. For example, among the features obtained in the above steps, acceleration is a measure of the rate of change of velocity, so the correlation between acceleration and velocity is strong; displacement is the integral of velocity with respect to time, so the correlation between displacement and velocity is strong; angular velocity is a physical quantity that describes the speed of rotation of an object, and the correlation with velocity, acceleration, and displacement is weak. When the algorithm randomly selects features, these features with high correlation are more likely to be selected simultaneously because of similar data distributions, resulting in a higher possibility that features with high correlation are selected as the division point.

[0049] Based on this, in the embodiments of the present invention, by obtaining the feature time series with the least correlation as the correlation sequence of the current target feature time series, the isolated result error caused by the repeated use of the feature time series with a larger correlation is avoided. And based on the data performance of the time series of the current target feature and the correlation sequence, the combined value of the target feature of the data point is obtained.

[0050] It can be understood that the Pearson correlation coefficient between the time series of features can measure the degree of correlation between the time series. The larger the absolute value of the Pearson correlation coefficient, the greater the degree of correlation between the time series; the smaller the absolute value of the Pearson correlation coefficient, the smaller the degree of correlation between the time series.

[0051] Based on this, in the embodiments of the present invention, by obtaining the importance between the correlation sequence of the target feature and the target feature time series, the combined value of the target feature of the data point can be obtained.

[0052] Specifically, according to the values of the data point in the time series of the target feature, the correlation sequence, and the importance of the correlation sequence of the target feature of the data point to the time series and the time series to the correlation sequence, the combined value of the target feature of the data point is calculated.

[0053] Exemplarily, in the embodiments of the present invention, the method for obtaining the importance of the correlation sequence of the target feature to the time series includes: taking the ratio of the Pearson correlation coefficient between the correlation sequence of the target feature and the time series and the sum of the Pearson correlation coefficients between the time series of the target feature and all feature time series as the importance of the correlation sequence of the target feature to the time series.

[0054] Among them, the larger the Pearson correlation coefficient between the correlation sequence of the target feature and the time series, the greater the degree of correlation between the correlation sequence of the target feature and the time series, and the higher the importance of the correlation sequence to the time series.

[0055] After obtaining the importance of any time series to another time series based on the above method, the combined value of the target feature of the data point can be calculated based on the following steps.

[0056] Exemplarily, in the embodiments of the present invention, to determine the combined value of the target feature of the data point, the following relational expression can be specifically referred to:

[0057] ;

[0058] is the combined value of the target feature of the i-th data point, is the importance of the correlation sequence of the target feature to the time series, is the importance of the time series of the target feature to the correlation sequence, is the value of the i-th data point in the time series of the target feature, is the value of the i-th data point in the relevant sequence of the target feature.

[0059] In the above formula, represents the ratio of the importance of the relevant sequence of the target feature to the time series sequence to the importance of the relevant sequence of the target feature to the time series sequence and the importance of the time series sequence of the target feature to the relevant sequence and the value. The larger this value is, the higher the proportion of the importance of the relevant sequence to the time series sequence, and the higher the contribution of the value of the i-th data point in the relevant sequence to the combined value.

[0060] Similarly, represents the ratio of the importance of the time series sequence of the target feature to the relevant sequence to the importance of the relevant sequence of the target feature to the time series sequence and the importance of the time series sequence of the target feature to the relevant sequence and the value. The larger this value is, the higher the proportion of the importance of the time series sequence to the relevant sequence, and the higher the contribution of the value of the i-th data point in the time series sequence to the combined value.

[0061] After obtaining the combined values of the target features of each data point based on the above steps, the combined values of the target features of each data point can be used as splitting points to obtain the isolation forest corresponding to the target feature. By analyzing the data distribution in the isolation forest of each target feature, evaluating the isolation effect of the isolation forest of each target feature and calculating the weights, the anomaly scores of each data point can be accurately obtained, that is, continue to execute the following steps.

[0062] S3: Select splitting points in the time series sequence of the combined values of the target feature to construct the isolation forest of the target feature; obtain the density feature of the isolation forest of the target feature through the difference between each data point in the isolation forest of the target feature and the median.

[0063] Among them, the combined values corresponding to the same type of target features are combined into a time series sequence of combined values.

[0064] It should be noted that the conventional isolation forest algorithm randomly selects feature values as splitting points for division in randomly selected feature dimensions, resulting in a large difference in the depth distribution between different isolation trees and a low accuracy of the anomaly scores of the obtained data points.

[0065] Based on this, the embodiments of the present invention construct a time series sequence of combined values by obtaining the combined values of the target features of each data point, and randomly select data points as splitting points for division in the time series sequence of the combined values of the target feature to obtain an isolation forest of the target feature composed of multiple decision trees.

[0066] Exemplarily, in the embodiment of the present invention, constructing an isolation forest for target features by selecting split points in the time series of combined values of target features includes: randomly selecting a preset number of data points from the vehicle dataset as a sub-dataset; randomly selecting the combined value of the target feature of a data point in the sub-dataset as a split point, and dividing the sub-dataset into a left branch and a right branch; continuing to randomly select split points in the branches for division until a preset termination condition is reached to obtain an isolation tree corresponding to the sub-dataset; and forming an isolation forest for target features by combining multiple isolation trees.

[0067] Among them, the preset number can be set to 300; specifically, the number can be set according to actual needs.

[0068] Exemplarily, the preset termination condition can be that all data points are isolated and / or the preset tree depth is reached.

[0069] Among them, the tree depth can be specifically set according to actual needs, and the embodiment of the present invention does not limit it too much here.

[0070] It should be further noted that during the process of splitting to obtain the isolation forest of target features, outliers are usually isolated earlier, that is, outliers are closer to the root node of the isolation tree, while normal values are farther from the root node of the isolation tree. Therefore, the greater the density of the isolation forest of target features, the longer the path required to isolate abnormal data, corresponding to a worse isolation effect; on the contrary, the smaller the density of the isolation forest of target features, the shorter the path to isolate abnormal data, corresponding to a better isolation effect.

[0071] Based on this, the embodiment of the present invention evaluates the isolation effect of the isolation forest of target features by calculating the density feature of the isolation forest of target features.

[0072] Exemplarily, in the embodiment of the present invention, to determine the density feature of the isolation forest of target features, the following relational expression can be specifically referred to:

[0073] ;

[0074] is the density feature of the isolation forest of target features, is the number of data points in the isolation forest of this target feature, is the median of the combined values of target features in the isolation forest of this target feature, is the combined value of the target feature of the j-th data point in the isolation forest of this target feature, is the exponential function with base e, is the absolute value symbol.

[0075] In the above formula, It represents the mean difference between the target feature combination values of all data points in the target feature isolation forest and the median of the target feature combination values. The smaller this value is, the higher the degree of proximity between the target feature combination values of all data points in the target feature isolation forest and the median, the closer the target feature combination values of the data points in the corresponding target feature isolation forest, and the greater the density feature of the target feature isolation forest.

[0076] After obtaining the density features of each target feature isolation forest based on the above steps, continue to execute the following steps.

[0077] S4: Obtain the weight of the target feature isolation forest through the density feature of the target feature isolation forest.

[0078] Among them, the weight of the target feature isolation forest is negatively correlated with the density feature.

[0079] It should be noted that after obtaining the density features of each target feature isolation forest based on the above steps, the isolation effect of each target feature isolation forest can be evaluated. The smaller the density feature of the target feature isolation forest, the better the isolation effect, the higher the contribution degree of the corresponding target feature isolation forest to calculating the anomaly score of the data point, and the greater the weight. Based on this, the weights of each target feature isolation forest in calculating the anomaly score can be accurately obtained.

[0080] It can be understood that each target feature has a uniquely corresponding target feature isolation forest.

[0081] Exemplarily, in the embodiments of the present invention, when determining the weight of the target feature isolation forest, the sum value of the importance of the relevant sequence of the target feature to the time series and the importance of the time series to the relevant sequence can be obtained as the feature importance of the target feature isolation forest corresponding to the target feature; calculate the weight of the target feature isolation forest, and specifically refer to the following relational expression:

[0082] ;

[0083] is the weight of the target feature isolation forest, is the natural constant, is the density feature of the target feature isolation forest, is the importance of the relevant sequence of the target feature to the time series, is the importance of the time series of the target feature to the relevant sequence, is the maximum value of the feature importance of all target feature isolation forests, is the linear normalization function.

[0084] In the above formula, The feature importance of the target feature isolation forest. The larger this value is, the higher the importance of the target feature isolation forest, and the greater the weight when calculating the anomaly score of the data point.

[0085] The larger the density feature of the target feature isolation forest is, the worse its isolation effect. In order to reduce the impact of the target feature isolation forest with a poor isolation effect on the calculation of the anomaly score of the data point, it is necessary to reduce the weight of this data point; conversely, the smaller the density feature of the target feature isolation forest is, the better its isolation effect. When calculating the anomaly score of the data point, a larger weight needs to be set for the target feature isolation forest with a better isolation effect.

[0086] After obtaining the weights of each target feature isolation forest based on the above steps, the path length of the data point can be weighted based on the weights of each target feature isolation forest, so as to accurately obtain the anomaly score of the data point, that is, continue to execute the following steps.

[0087] S5: Obtain the product of the average path length and the weight of the data point in each target feature isolation forest, and normalize the mean of all products to obtain the anomaly score of this data point; realize vehicle-mounted data processing based on the comparison result between the anomaly score of the data point and the preset threshold.

[0088] Exemplarily, in the embodiment of the present invention, when calculating the anomaly score of a data point, the average path length of the data point in each target feature isolation forest can be obtained, and the corresponding average path length is weighted based on the weights of each target feature isolation forest, so as to obtain the anomaly score of this data point.

[0089] Exemplarily, to calculate the anomaly score of a data point, specifically, the following relational expression can be referred to:

[0090] ;

[0091] is the anomaly score of the data point in the vehicle-mounted dataset, is the number of features corresponding to this data point, is the weight of the th target feature isolation forest, is the average path length of this data point in the th target feature isolation forest,

[0092] is the linear normalization function.

[0093] Exemplarily, in the embodiments of the present invention, vehicle data processing is implemented based on the comparison result between the anomaly score of a data point and a preset threshold, including: if the anomaly score of the data point is greater than the preset threshold, then the data point is abnormal vehicle data; otherwise, the data point is normal vehicle data.

[0094] Among them, the preset threshold can be set to 0.6; the preset threshold can be specifically set according to actual needs, and the embodiments of the present invention do not impose too many restrictions here.

[0095] It can be understood that when the anomaly score of a data point is greater than the preset threshold, it indicates that the feature at the acquisition moment corresponding to the current data point is abnormal. In this case, it is necessary to promptly remind of the abnormal data to reduce potential safety hazards.

[0096] Exemplarily, in the embodiments of the present invention, after implementing vehicle data processing, it further includes: in response to the data point being abnormal vehicle data, prompting the abnormal vehicle data.

[0097] Among them, the prompting method can be a signal lamp or icon flashing prompt, or a voice prompt, which can be specifically set according to actual needs, and the embodiments of the present invention do not impose too many restrictions here.

[0098] It can be seen that in the embodiments of the present invention, when processing vehicle data, any feature of the data points in the vehicle data set can be used as the target feature, and the other feature time series with the smallest absolute value of the Pearson correlation coefficient between the time series of the target feature is used as the correlation series of the target feature; according to the values of the data point in the time series of the target feature and the correlation series, and the importance of the time series of the correlation series of the target feature of the data point to the time series of the target feature and the time series pair of the correlation series, calculate the combined value of the target feature of the data point; form a combined value time series with the combined values corresponding to the same type of target features, and select a splitting point in the combined value time series of the target feature to construct an isolation forest of the target feature; obtain the density feature of the isolation forest of the target feature through the difference between each data point in the isolation forest of the target feature and the median; obtain the weight of the isolation forest of the target feature through the density feature of the isolation forest of the target feature, and the weight of the isolation forest of the target feature is negatively correlated with the density feature; obtain the product of the average path length and the weight of the data point in each isolation forest of the target feature, and normalize the mean of all products to obtain the anomaly score of the data point; implement vehicle data processing based on the comparison result between the anomaly score of the data point and the preset threshold.

[0099] In this way, the embodiments of the present invention can screen out abnormal data from the vehicle-mounted data set by processing multi-dimensional vehicle-mounted data through the Isolation Forest algorithm. In this process, the embodiments of the present invention take into account that the correlations between different features of the vehicle-mounted data are different. However, the Isolation Forest algorithm only uses a single feature each time when selecting a splitting point, resulting in unbalanced feature usage and affecting the accuracy of the construction of isolation trees. Based on this, the embodiments of the present invention combine each feature with the feature having the smallest correlation, and completely use the combined value of the features as the splitting point to construct the corresponding feature isolation forest, and analyze the density characteristics of each feature isolation forest to evaluate the isolation effect, so as to realize the weighting of the average path length of data points, accurately obtain the abnormal score of data points, and effectively improve the accuracy of abnormal monitoring of vehicle-mounted data.

[0100] The embodiments of the present invention also disclose a vehicle-mounted data processing system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a vehicle-mounted data processing method provided by the present invention is implemented.

[0101] The above system also includes other components well-known to those skilled in the art such as a communication bus and a communication interface. Their settings and functions are known in the art, so they will not be described in detail here.

[0102] In the present invention, the aforementioned memory can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, the computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), etc., or any other medium that can be used to store the required information and can be accessed by an application program, module, or both. Any such computer storage medium can be part of the device or accessible or connectable to the device.

[0103] Although this specification has shown and described multiple embodiments of the present invention, it is obvious to those skilled in the art that such embodiments are provided only by way of example. Those skilled in the art will think of many changes, alterations, and alternative ways without departing from the spirit and idea of the present invention. It should be understood that various alternative solutions to the embodiments of the present invention described herein can be adopted in the process of practicing the present invention.

[0104] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A vehicle-mounted data processing method, characterized in that, Including: Taking any feature of the data points in the vehicle-mounted dataset as the target feature, and taking the other feature time series with the smallest absolute value of the Pearson correlation coefficient between the time series of the target feature as the correlation series of the target feature; Calculating the combined value of the target feature of the data point according to the values of the data point in the time series of the target feature and the correlation series, and the importance of the correlation series of the target feature to the time series and the time series pair of the correlation series, including: ; is the combined value of the target feature for the i-th data point, , are respectively the importance of the relevant sequence pair of the target feature to the time series sequence and the time series sequence pair to the relevant sequence, , are respectively the values of the i-th data point in the time series sequence and the relevant sequence of the target feature; Forming the combined value time series with the combined values corresponding to the same type of target features, selecting split points in the combined value time series of the target feature to construct the target feature isolation forest; obtaining the density feature of the target feature isolation forest through the difference between each data point in the target feature isolation forest and the median; obtaining the weight of the target feature isolation forest through the density feature of the target feature isolation forest, and the weight of the target feature isolation forest is negatively correlated with the density feature; Obtaining the product of the average path length and the weight of each data point in the target feature isolation forest, normalizing the mean of all products to obtain the anomaly score of the data point; realizing vehicle-mounted data processing based on the comparison result between the anomaly score of the data point and the preset threshold; The method for obtaining the importance of the correlation series of the target feature to the time series, including: Taking the ratio of the Pearson correlation coefficient between the correlation series of the target feature and the time series and the sum of the Pearson correlation coefficients between the time series of the target feature and all feature time series as the importance of the correlation series of the target feature to the time series.

2. The vehicle-mounted data processing method according to claim 1, wherein Before taking any feature of the data points in the vehicle-mounted dataset as the target feature, it further includes: Preprocessing the multiple vehicle-mounted feature values obtained at each acquisition moment as a data point to obtain the data points in the vehicle-mounted dataset.

3. The vehicle-mounted data processing method according to claim 1, characterized in that, Selecting split points in the combined value time series of the target feature to construct the target feature isolation forest, including: Randomly selecting a preset number of data points in the vehicle-mounted dataset as the sub-dataset; randomly selecting the combined value of the target feature of a data point in the sub-dataset as the split point, and dividing the sub-dataset into a left branch and a right branch; continuing to randomly select split points in the branches for division until the preset termination condition is reached to obtain the isolation tree corresponding to the sub-dataset; forming the target feature isolation forest with multiple isolation trees.

4. A vehicle-mounted data processing method according to claim 1, characterized in that, The density feature of the target feature isolation forest satisfies the relational expression: ; is the density feature of the target feature isolation forest, is the number of data points in the target feature isolation forest, is the median of the target feature combination values in the target feature isolation forest, is the target feature combination value of the j-th data point in the target feature isolation forest, is the exponential function with base e, is the absolute value symbol.

5. The vehicle-mounted data processing method according to claim 2, characterized in that, The method for obtaining the weight of the target feature isolation forest, including: Taking the sum of the importance of the correlation series of the target feature to the time series and the importance of the time series to the correlation series as the feature importance of the target feature isolation forest corresponding to the target feature; ; is the weight of the target feature isolation forest, is the natural constant, is the density feature of the target feature isolation forest, and are the importance of the relevant sequence to the time series and the importance of the time series to the relevant sequence of the target feature respectively, is the maximum value of the feature importance of all target feature isolation forests, is the linear normalization function.

6. The vehicle-mounted data processing method according to claim 5, wherein Realizing vehicle-mounted data processing based on the comparison result between the anomaly score of the data point and the preset threshold, including: If the anomaly score of the data point is greater than the preset threshold, then the data point is abnormal vehicle-mounted data; otherwise, the data point is normal vehicle-mounted data.

7. The vehicle-mounted data processing method according to claim 6, wherein After realizing vehicle-mounted data processing, it further includes: In response to the data point being abnormal vehicle-mounted data, prompting the abnormal vehicle-mounted data.

8. A vehicle-mounted data processing system, characterized in that, Including: A processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a vehicle-mounted data processing method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Intelligent processing method and system for vehicle abnormal information

    CN117292462A

  • Abnormal sample detection method based on improved isolated forest and related equipment

    CN113420073A