Abnormal Detection Method, Device and Equipment for CAN Bus Data
Detecting abnormal data in CAN bus data through local outlier factor model and preset thresholds, the problem of inability to effectively detect CAN bus data abnormalities in the prior art is solved, and the security and detection efficiency of the CAN bus are improved.
Patent Information
- Application Number
- CN202111454384.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-01
AI Technical Summary
The prior art cannot effectively detect abnormal data in CAN bus data, resulting in threats to automobile information security.
The local outlier factor model and preset abnormal data detection threshold are used to identify abnormal data in the CAN bus data by combining the CAN bus data sequence and calculating the local abnormal factor.
Quickly identify abnormal data in CAN bus data, improving the security and abnormal detection efficiency of CAN bus, ensuring data integrity and reliability.
Smart Images

Figure CN114253779B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicles, and in particular, to a method, device and equipment for detecting anomalies in CAN bus data. Background Art
[0002] CAN (Controller Area Network) is an internationally standardized serial communication protocol, and data can be transmitted between various actuators inside a vehicle using the CAN bus. Data sent by different types or different actuators can be identified using different IDs (Identity Documents). With the emergence of automotive information security issues, malicious attacks penetrate directly into the in-vehicle CAN bus network through external interfaces, seriously endangering the personal and property safety of vehicle occupants. However, the current CAN bus lacks a basic information security mechanism. To improve the security of the CAN bus, it is a very important link to implement anomaly detection for CAN bus data. Therefore, how to quickly detect whether there is abnormal data in CAN bus data is an urgent problem to be solved. Summary of the Invention
[0003] The present invention provides a method, device and equipment for detecting anomalies in CAN bus data, which can efficiently detect whether there is abnormal data in CAN bus data.
[0004] The specific technical solutions are as follows:
[0005] In a first aspect, an embodiment of the present invention provides a method for obtaining N CAN bus data to be detected, where the identity identification numbers (IDs) of different CAN bus data among the N CAN bus data are different, and N is a positive integer;
[0006] The N CAN bus data are combined into a CAN bus data sequence to be detected;
[0007] The CAN bus data sequence to be detected is input into a local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected, where the local outlier factor model is a model trained based on multiple CAN bus data sequences for calculating the local outlier factor, and the number of CAN bus data and the type of IDs in each CAN bus data sequence among the multiple CAN bus data sequences are both N;
[0008] If the local outlier factor is greater than a preset abnormal data detection threshold, it is determined that there is abnormal data among the N CAN bus data.
[0009] Optionally, before inputting the CAN bus data sequence to be detected into the local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected, the method further includes:
[0010] Collect T CAN bus data, where the T CAN bus data includes CAN bus data of N different IDs;
[0011] Construct an initialized CAN bus data sequence with all data contents being 0, where the number of data in the initialized CAN bus data sequence is N*X, X is the maximum number of bytes among the bytes of each CAN bus data in the T CAN bus data, and every X adjacent data in the initialized CAN bus data sequence represents CAN bus data of one ID;
[0012] Starting from the first CAN bus data in the T CAN bus data, sequentially replace the CAN bus data of the corresponding ID in the initialized CAN bus data sequence to generate T CAN bus data sequences, where the i-th CAN bus data sequence is obtained by replacing the CAN bus data of the corresponding ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data in the T CAN bus data, and i≥2;
[0013] If there are duplicate CAN bus data sequences in the T CAN bus data sequences, perform deduplication processing on the T CAN bus data sequences, and obtain CAN bus data sequences of a first preset ratio from the deduplicated T CAN bus data sequences as training samples;
[0014] Input the training samples into the local outlier factor system for model training to obtain a local outlier factor model.
[0015] Optionally, after inputting the training samples into the local outlier factor system for model training to obtain a local outlier factor model, the method further includes:
[0016] Select CAN bus data sequences of a second preset ratio from the training samples;
[0017] Input the CAN bus data sequences of the second preset ratio into the local outlier factor model to obtain the local outlier factor of each CAN bus data sequence in the CAN bus data sequences of the second preset ratio;
[0018] Determine the maximum local outlier factor among the local outlier factors of each CAN bus data sequence in the CAN bus data sequences of the second preset ratio as the preset abnormal data detection threshold.
[0019] Optionally, after inputting the training samples into the local outlier factor system for model training to obtain the local outlier factor model, the method further includes:
[0020] According to the deduplicated T CAN bus data sequences and the preset abnormal data construction rule, construct a test sample including multiple test CAN bus data sequences, where the multiple test CAN bus data sequences include CAN bus data sequences without abnormal data and CAN bus data sequences with abnormal data;
[0021] Input the test sample into the local outlier factor model to obtain the local outlier factor of each test CAN bus data sequence;
[0022] Obtain the abnormal data test result by comparing the local outlier factors of multiple test CAN bus data sequences with the preset abnormal data detection threshold, where the abnormal data test result includes the test result of whether there is abnormal data in each test CAN bus data sequence;
[0023] Calculate the abnormal detection accuracy rate by comparing the abnormal data test result with the CAN bus data sequences with preset abnormal data in the test sample;
[0024] If the abnormal detection accuracy rate is less than the preset accuracy rate threshold, increase the sample size of the training samples, and continue to train the local outlier factor model with the training samples with increased sample size until the abnormal detection accuracy rate is greater than or equal to the preset accuracy rate threshold, then obtain the final required local outlier factor model.
[0025] Optionally, according to the deduplicated T CAN bus data sequences and the preset abnormal data construction rule, constructing a test sample including multiple test CAN bus data sequences includes:
[0026] Obtain all CAN bus data sequences except the training samples from the deduplicated T CAN bus data sequences as the test CAN bus data sequences without abnormal data in the test sample;
[0027] Obtain the j-th byte of each CAN bus data from the deduplicated T CAN bus data sequences, where j is a positive integer;
[0028] Statistically determine whether the values corresponding to all the j-th bytes include all the values within the preset value range;
[0029] If all the values within the preset value range are included, select a preset number of values from the preset value range as the abnormal data;
[0030] If all the values within the preset numerical range are not included, then a preset number of values selected from the values not included in the preset numerical range are determined as abnormal data;
[0031] Select at least one CAN bus data sequence from the T CAN bus data sequences after deduplication as the target CAN bus data sequence;
[0032] Replace the j-th byte of the preset number of CAN bus data in the target CAN bus data sequence with the preset number of abnormal data to generate a test CAN bus data sequence including abnormal data;
[0033] After generating multiple test CAN bus data sequences including abnormal data for multiple different bytes, the multiple test CAN bus data sequences including abnormal data and the test CAN bus data sequence without abnormal data form a test sample.
[0034] Optionally, replacing the i-th CAN bus data in the T CAN bus data with the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence includes:
[0035] If the number of bytes of the i-th CAN bus data in the T CAN bus data is less than X, then pad zeros at the end of the i-th CAN bus data so that the number of bytes after padding is equal to X, and then replace the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence with the padded i-th CAN bus data;
[0036] If the number of bytes of the i-th CAN bus data is equal to X, then the i-th CAN bus data replaces the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence.
[0037] Optionally, obtaining N CAN bus data to be detected includes:
[0038] Collect N CAN bus data to be detected;
[0039] If there is non-decimal CAN bus data among the N CAN bus data to be detected collected, then convert the non-decimal CAN bus data into decimal CAN bus data.
[0040] Optionally, after determining that there is abnormal data among the N CAN bus data, the method further includes:
[0041] For each of the N CAN bus data, use the anomaly data detection model corresponding to the CAN bus data to determine whether the CAN bus data is anomaly data. Among them, different CAN bus data with different IDs correspond to different anomaly data detection models, and the anomaly data detection model is a model trained based on multiple CAN bus data with the same ID and used to identify whether CAN bus data is abnormal.
[0042] In a second aspect, an embodiment of the present invention provides an anomaly detection device for CAN bus data, and the device includes:
[0043] A first acquisition unit, configured to acquire N CAN bus data to be detected, where different CAN bus data among the N CAN bus data have different identity identification numbers ID, and N is a positive integer;
[0044] A combination unit, configured to combine the N CAN bus data into a CAN bus data sequence to be detected;
[0045] A second acquisition unit, configured to input the CAN bus data sequence to be detected into a local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected. Among them, the local outlier factor model is a model trained according to multiple CAN bus data sequences and used to calculate the local outlier factor. The number of CAN bus data and the type of ID in each CAN bus data sequence among the multiple CAN bus data sequences are both N;
[0046] A determination unit, configured to determine that there is anomaly data in the N CAN bus data if the local outlier factor is greater than a preset anomaly data detection threshold.
[0047] Optionally, the device further includes:
[0048] An acquisition unit, configured to acquire T CAN bus data before inputting the CAN bus data sequence to be detected into the local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected. Among the T CAN bus data, there are N types of CAN bus data with different IDs, and T is a positive integer;
[0049] A first construction unit, configured to construct an initialized CAN bus data sequence with all data contents being 0. Among them, the number of data in the initialized CAN bus data sequence is N*X, X is the maximum number of bytes among the number of bytes of each CAN bus data in the T CAN bus data, and every X adjacent data in the initialized CAN bus data sequence represents a CAN bus data of one ID;
[0050] A replacement unit, configured to sequentially replace the CAN bus data corresponding to the ID in the initialized CAN bus data sequence starting from the first CAN bus data among the T CAN bus data, so as to generate T CAN bus data sequences. Wherein, the i-th CAN bus data sequence is obtained by replacing the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data among the T CAN bus data, and i ≥ 2;
[0051] A deduplication unit, configured to perform deduplication processing on the T CAN bus data sequences if there are duplicate CAN bus data sequences among the T CAN bus data sequences, and obtain a first preset proportion of CAN bus data sequences from the deduplicated T CAN bus data sequences as training samples;
[0052] A training unit, configured to input the training samples into a local outlier factor system for model training to obtain a local outlier factor model.
[0053] Optionally, the apparatus further includes:
[0054] A selection unit, configured to select a second preset proportion of CAN bus data sequences from the training samples after inputting the training samples into a local outlier factor system for model training to obtain a local outlier factor model;
[0055] The second acquisition unit is further configured to input the second preset proportion of CAN bus data sequences into the local outlier factor model to obtain the local outlier factors of each CAN bus data sequence in the second preset proportion of CAN bus data sequences;
[0056] The determination unit is further configured to determine the largest local outlier factor among the local outlier factors of each CAN bus data sequence in the second preset proportion of CAN bus data sequences as the preset abnormal data detection threshold.
[0057] Optionally, the apparatus further includes:
[0058] A second construction unit is further configured to, after inputting the training samples into a local outlier factor system for model training to obtain a local outlier factor model, construct a test sample including a plurality of test CAN bus data sequences according to the deduplicated T CAN bus data sequences and a preset abnormal data construction rule, where the plurality of test CAN bus data sequences include CAN bus data sequences without abnormal data and CAN bus data sequences with abnormal data;
[0059] The second obtaining unit is further configured to input the test sample into the local outlier factor model to obtain the local outlier factor of each test CAN bus data sequence;
[0060] The apparatus further includes:
[0061] A comparison unit, configured to obtain an abnormal data test result by comparing the local outlier factors of multiple test CAN bus data sequences with the preset abnormal data detection threshold, where the abnormal data test result includes the test result of whether there is abnormal data in each test CAN bus data sequence;
[0062] A calculation unit, configured to calculate the abnormal detection accuracy rate by comparing the abnormal data test result with the CAN bus data sequences with preset abnormal data in the test sample;
[0063] The training unit is configured to, if the abnormal detection accuracy rate is less than the preset accuracy rate threshold, increase the sample size of the training sample, and continue to train the local outlier factor model with the training sample with the increased sample size until the abnormal detection accuracy rate is greater than or equal to the preset accuracy rate threshold, so as to obtain the finally required local outlier factor model.
[0064] Optionally, the second construction unit includes:
[0065] An obtaining module, configured to obtain all CAN bus data sequences except the training sample from the T CAN bus data sequences after duplicate removal as the test CAN bus data sequences without abnormal data in the test sample; obtain the j-th byte of each CAN bus data from the T CAN bus data sequences after duplicate removal, where j is a positive integer;
[0066] A statistics module, configured to count whether the values corresponding to all the j-th bytes include all the values within the preset value range;
[0067] A selection module, configured to, if all the values within the preset value range are included, select a preset number of values from the preset value range as the abnormal data, and if not all the values within the preset value range are included, determine the preset number of values selected from the values not included in the preset value range as the abnormal data;
[0068] A selection module, configured to select at least one CAN bus data sequence from the T CAN bus data sequences after duplicate removal as the target CAN bus data sequence;
[0069] A first replacement module, configured to replace the j-th byte of a preset number of CAN bus data in the target CAN bus data sequence with a preset number of abnormal data, and generate a test CAN bus data sequence including the abnormal data;
[0070] A construction module, configured to, after generating a plurality of test CAN bus data sequences including abnormal data for a plurality of different bytes, form a test sample from the plurality of test CAN bus data sequences including abnormal data and the test CAN bus data sequence without abnormal data.
[0071] Optionally, the replacement unit includes:
[0072] A supplement module, configured to, if the number of bytes of the i-th CAN bus data among the T CAN bus data is less than X, pad zeros at the end of the i-th CAN bus data so that the number of bytes after padding is equal to X, and then replace the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence with the padded i-th CAN bus data;
[0073] A second replacement module, configured to, if the number of bytes of the i-th CAN bus data is equal to X, replace the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data.
[0074] Optionally, the first acquisition unit includes:
[0075] An acquisition module, configured to acquire N CAN bus data to be detected;
[0076] A conversion module, configured to, if there is non-decimal CAN bus data among the N CAN bus data to be detected acquired, convert the non-decimal CAN bus data into decimal CAN bus data.
[0077] Optionally, the device further includes:
[0078] A judgment unit, configured to, after determining that there is abnormal data in the N CAN bus data, for each CAN bus data in the N CAN bus data, use the abnormal data detection model corresponding to the CAN bus data to judge whether the CAN bus data is abnormal data, where different IDs of CAN bus data correspond to different abnormal data detection models, and the abnormal data detection model is a model trained based on a plurality of CAN bus data with the same ID and used to identify whether CAN bus data is abnormal.
[0079] In a third aspect, an embodiment of the present invention provides a storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor implements the method described in the first aspect.
[0080] In a fourth aspect, an embodiment of the present invention provides an electronic device, including:
[0081] One or more processors;
[0082] A storage device for storing one or more programs,
[0083] wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in the first aspect.
[0084] As can be seen from the above, the method, apparatus, and device for detecting anomalies in CAN bus data provided by the embodiments of the present invention can, after obtaining N CAN bus data to be detected, first combine the N CAN bus data into a CAN bus data sequence to be detected, and then input the CAN bus data sequence to be detected into a local outlier factor model trained based on multiple CAN bus data sequences to obtain the local outlier factor of the CAN bus data sequence to be detected. Finally, when the local outlier factor is greater than a preset anomaly data detection threshold, it can be determined that there is anomaly data in the N CAN bus data; conversely, it can be determined that there is no anomaly data in the N CAN bus data. Thus, compared with the prior art that cannot detect anomalies in CAN bus data, the present invention can quickly identify whether there is anomaly data in multiple CAN bus data with different IDs by combining the local outlier factor model and the preset anomaly data detection threshold, so as to quickly process the anomaly data and further improve the security of the CAN bus.
[0085] The embodiments of the present invention can at least achieve the following technical effects:
[0086] 1. By first using the local outlier factor model and the preset anomaly data detection threshold to quickly identify whether there is anomaly data in multiple CAN bus data (a coarse-grained anomaly detection method), when there is anomaly data, then using different anomaly data detection models corresponding to CAN bus data with different IDs to further detect anomalies in multiple CAN bus data to determine which specific CAN bus data is anomaly data (a fine-grained anomaly detection method), and when there is no anomaly data, there is no need to continue anomaly detection. Therefore, this method of first using the coarse-grained method for anomaly detection and then using the fine-grained method for anomaly detection can improve the anomaly detection efficiency from an overall perspective compared with directly using the fine-grained method for anomaly detection.
[0087] 2. Construct multiple CAN bus data sequences from multiple CAN bus data, and use a part of the sequences among the multiple CAN bus data sequences as training samples for local outlier factor model training, and use another part of the sequences to construct test samples, and use the test samples to test the trained local outlier factor model. Since both the training samples and the test samples are determined based on real CAN bus data, the accuracy of calculating the local outlier factor by the local outlier factor model trained with the training samples and test samples is higher than that trained with artificially constructed training samples and test samples.
[0088] 3. By padding zeros at the end of the CAN bus data with a smaller number of bytes, the number of bytes of all CAN bus data used for training the local outlier factor model can be made the same, which can improve the accuracy of the local outlier factor model.
[0089] 4. By converting non - decimal CAN bus data among the N CAN bus data to be detected into decimal CAN bus data, the efficiency of calculating the local outlier factor by the local outlier factor model can be improved.
[0090] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above - mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0092] Figure 1 It is a flowchart of an abnormal detection method for CAN bus data provided by an embodiment of the present invention;
[0093] Figure 2 It is a flowchart of a training method for a local outlier factor model provided by an embodiment of the present invention;
[0094] Figure 3 It is a block diagram of a composition of an abnormal detection device for CAN bus data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0095] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0096] It should be noted that the terms "including" and "having" and any variations thereof in the embodiments of the present invention and the accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.
[0097] The present invention provides a method, device, and equipment for detecting anomalies in CAN bus data to solve the problem that the prior art cannot detect anomalies in CAN bus data. The method provided by the embodiments of the present invention can be applied to any electronic device with computing capabilities, and the electronic device can be a terminal or a server. In one implementation, the functional software for implementing this method can exist in the form of a separate client software or in the form of a plugin for the current relevant client software.
[0098] The embodiments of the present invention will be described in detail below.
[0099] Figure 1 FIG. is a schematic flowchart of a method for detecting anomalies in CAN bus data provided by an embodiment of the present invention. The method may include the following steps:
[0100] S100: Obtain N CAN bus data to be detected.
[0101] Among them, the IDs (Identity Document) of different CAN bus data among the N CAN bus data are different, that is, the N CAN bus data are N CAN bus data with different IDs, and N is a positive integer.
[0102] Since the efficiency of using decimal calculation is higher when calculating the local outlier factor in the local outlier factor model, after collecting the N CAN bus data to be detected, if there are non-decimal CAN bus data among the collected N CAN bus data to be detected, then convert the non-decimal CAN bus data into decimal CAN bus data.
[0103] S110: Combine the N CAN bus data into a CAN bus data sequence to be detected.
[0104] Among them, when the N CAN bus data are combined into a CAN bus data sequence to be detected, the sorting in the sequence is combined according to the sorting of each ID in the CAN bus data sequence in the local outlier factor model training sample. For example, it can be sorted in ascending order of ID to form a CAN bus data sequence to be detected.
[0105] S120: Input the CAN bus data sequence to be detected into the local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected.
[0106] Wherein, the local outlier factor model is a model for calculating the local outlier factor obtained by training according to multiple CAN bus data sequences. For each CAN bus data sequence in the multiple CAN bus data sequences, the number of data and the type of ID of the CAN bus data are both N.
[0107] The specific process of the local outlier factor model calculating the local outlier factor includes:
[0108] (1) Calculate the K-nearest distance of the CAN bus data sequence to be detected. Calculate the distances between a point and multiple surrounding points, and sort these distances from small to large. The k-th distance is the K-nearest distance of this point, denoted as K-distance. Therefore, the CAN bus data sequence to be detected can be regarded as a point P to be detected, calculate the distances between this point P to be detected and multiple points around it in the local outlier factor model, and determine the K-nearest distance from them. Among them, which specific multiple surrounding points are those is the best range found by the local outlier factor model through multiple trainings. For example, the range with point P as the center and R as the radius is the range where multiple surrounding points are located. The value of K can be determined according to specific circumstances. For example, it can be taken as 100.
[0109] Suppose there are N points around point P, and the set form of these points is as follows: [P1, P2,..., P N .
[0110] Sort the distances from these N points to point P from small to large, and the distance sequence between these N points and point P is: k ∈ [1, N], wherein, To Increase in turn. k represents the k-th position after sorting from small to large. Then Is the K-nearest distance of point P, denoted as K-distance(P).
[0111] (2) Calculate the reachable distance of point P. The reachable distance is related to the K-nearest distance. When a parameter k is given, the reachable distance from point P to other points (assumed to be point O) is the maximum value of the K-nearest distance of point O and the direct distance between point P and point O. When the reachable distance between point P and point O is denoted as reach_dist k (P, O), and the direct distance between point P and point O is denoted as d(P, O), the formula expression form of the reachable distance between point P and point O is:
[0112] reach_dist k (P, O) = max{K-distance(O), d(P, O)}, k ∈ [1, N].
[0113] (3) Calculate the local reachability density of point P. For point P, those points whose distance to point P is less than or equal to the K-nearest neighbor distance K-distance(P) of point P are the K-nearest neighbors of point P, denoted as N k (P), and the number of K-nearest neighbors is denoted as |N k (P)|, |N k (P)| ≥ k, because there may be different points with the same K-nearest neighbor distance to P. The local reachability density of point P is the reciprocal of the average reachability distance of its K-nearest neighbor data points, denoted as lrd k (P), and the specific expression is
[0114]
[0115] (4) Calculate the local outlier factor of point P. The local outlier factor algorithm measures the degree of abnormality of data points, not by looking at its absolute local density, but by looking at its relative density with respect to the surrounding neighboring data points. The advantage of doing this is that it allows for uneven data distributions and different densities. The local outlier factor is defined using local relative density. The local relative density (i.e., the local outlier factor) of point P is the ratio of the average local reachability density of the K-nearest neighbors of point P to the local reachability density of point P, that is, first find the local reachability density of the K-nearest neighbors of point P according to the method of finding the local reachability density of point P and take the average, and then find the ratio of this average to the local reachability density of point P. The local outlier factor of point P is denoted as LOF k (P), and its expression is:
[0116]
[0117] S130: If the local outlier factor is greater than the preset outlier data detection threshold, it is determined that there is outlier data in the N CAN bus data.
[0118] The preset abnormal data detection threshold is a critical value used to determine whether the CAN bus data corresponding to the local outlier factor is abnormal data. When the local outlier factor of the CAN bus data sequence to be detected is greater than the preset abnormal data detection threshold, it can be determined that there is abnormal data among the N CAN bus data that make up the CAN bus data sequence to be detected; when the local outlier factor of the CAN bus data sequence to be detected is less than or equal to the preset abnormal data detection threshold, it can be determined that there is no abnormal data among the N CAN bus data that make up the CAN bus data sequence to be detected. For the method of determining the preset abnormal data detection threshold, refer to the method embodiments below and will not be elaborated here.
[0119] When it is determined that there is abnormal data among the N CAN bus data, in order to further determine which specific CAN bus data is abnormal data, for each CAN bus data among the N CAN bus data, the abnormal data detection model corresponding to the CAN bus data can be used to determine whether the CAN bus data is abnormal data.
[0120] Among them, CAN bus data with different IDs correspond to different abnormal data detection models. The abnormal data detection model is a model trained based on multiple CAN bus data with the same ID and used to identify whether CAN bus data is abnormal. The method for training the abnormal data detection model includes: (1) a training method based on supervised learning, which inputs existing normal CAN bus data and artificially constructed abnormal CAN bus data into a supervised learning model for training; (2) a training method based on unsupervised learning, which only inputs normal CAN bus data into the model for training. Both of these methods complete training with data of a single ID in the CAN bus during the training process, thereby obtaining a training model.
[0121] The abnormal detection method for CAN bus data provided by the embodiments of the present invention can, after obtaining the N CAN bus data to be detected, first combine the N CAN bus data into a CAN bus data sequence to be detected, then input the CAN bus data sequence to be detected into a local outlier factor model trained based on multiple CAN bus data sequences to obtain the local outlier factor of the CAN bus data sequence to be detected. Finally, when the local outlier factor is greater than the preset abnormal data detection threshold, it can be determined that there is abnormal data among the N CAN bus data, and vice versa, it can be determined that there is no abnormal data among the N CAN bus data. It can be seen that compared with the prior art that cannot perform abnormal detection on CAN bus data, the present invention can quickly identify whether there is abnormal data among multiple CAN bus data with different IDs by combining the local outlier factor model and the preset abnormal data detection threshold, so as to quickly process the abnormal data, thereby improving the safety of the CAN bus.
[0122] Optionally, the local outlier factor model training method in the above method embodiments is as follows Figure 2 shown, and the method mainly includes:
[0123] S200: Collect T CAN bus data.
[0124] Among them, the T CAN bus data includes CAN bus data of N different IDs. In the embodiments of the present invention, one data can be collected every time a data is generated in the CAN bus. When N different IDs of CAN bus data are collected and reach T, the collection can be stopped. Among them, since the update cycles of CAN bus data with different IDs are different, the number of IDs in the T CAN bus data may be more for some and less for others.
[0125] When sorting the T CAN bus data in the collection order to form a data set D, the expression can be: D = [D1, D2,..., D i , i ∈ [1, T], D i represents the i-th CAN bus data.
[0126] Each CAN bus data is composed of multiple bytes. Assuming that the maximum number of bytes is 8, the expression of each CAN bus data can be D i = [d 1* , d 2* ,..., d j* , i ∈ [1, T], j* ∈ [1, 8], d j* represents the value corresponding to the j*-th byte of the i-th CAN bus data.
[0127] The original data of the CAN bus data is generally hexadecimal data. For the convenience of subsequent processing, the hexadecimal data can be converted into decimal data. The expression of each CAN bus data after the base conversion can be D i = [d1, d2,..., d j , i ∈ [1, T], j ∈ [1, 8], d j represents the value corresponding to the j-th byte of the i-th CAN bus data.
[0128] S210: Construct an initialized CAN bus data sequence with all data contents being 0.
[0129] Among them, the number of data in the initialized CAN bus data sequence is N * X, where X is the maximum number of bytes in each CAN bus data among the T CAN bus data, and every X adjacent data in the initialized CAN bus data sequence represents CAN bus data of one ID.
[0130] In specific implementation, it is possible to first count how many types of IDs there are among the T CAN bus data collected in S200, and the statistical result is N types of IDs. Then, a sequence containing N empty positions can be created first. The sorting of the elements in this sequence can be determined according to requirements. For example, it can be sorted in ascending order of ID. Then, each empty position is expanded into X empty positions, and finally each empty position is initialized to 0, obtaining an initialized CAN bus data sequence containing N * X zeros.
[0131] When X = 8, the expression of the initialized CAN bus data sequence can be N ∈ [1, T], j ∈ [1, 8], d Nj represents the j-th byte of the N-th CAN bus data. d Nj The initial value is 0.
[0132] S220: Starting from the first CAN bus data among the T CAN bus data, sequentially replace the CAN bus data corresponding to the ID in the initialized CAN bus data sequence to generate T CAN bus data sequences.
[0133] Among them, the i-th CAN bus data sequence is obtained by replacing the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data among the T CAN bus data, where i ≥ 2.
[0134] Exemplarily, assume T = 3, and the IDs of these three CAN bus data are ID1, ID2, and ID1 respectively, and the corresponding CAN bus data are represented by respectively. The initialized CAN bus data sequence is where the ID sorting is ID1 and ID2. Replace sequentially with to obtain the sequence and
[0135] The following are the expressions of each sequence:
[0136]
[0137] In practical applications, the number of bytes of different CAN bus data may be different. To improve the accuracy of the local outlier factor model, the number of bytes of all CAN bus data can be extended to the maximum number of bytes. Specifically, if the number of bytes of the i-th CAN bus data among the T CAN bus data is less than X, then zeros are padded to the tail of the i-th CAN bus data so that the number of bytes after padding is equal to X, and then the padded i-th CAN bus data replaces the CAN bus data with the corresponding ID in the (i - 1)-th CAN bus data sequence; if the number of bytes of the i-th CAN bus data is equal to X, then the i-th CAN bus data replaces the CAN bus data with the corresponding ID in the (i - 1)-th CAN bus data sequence.
[0138] S230: If there are duplicate CAN bus data sequences in the T CAN bus data sequences, then perform a deduplication process on the T CAN bus data sequences, and obtain a first preset proportion of the CAN bus data sequences from the deduplicated T CAN bus data sequences as training samples.
[0139] Since in practical applications, when the actual value of the CAN bus data is 0 and the corresponding CAN bus data in the i-th CAN bus data sequence is also 0, the obtained i-th CAN bus data sequence and the (i + 1)-th CAN bus data sequence are the same sequence, there may be duplicate CAN bus data sequences in the T CAN bus data sequences. Therefore, the T CAN bus data sequences can be deduplicated first, and then a first preset proportion of the CAN bus data sequences are obtained from the deduplicated T CAN bus data sequences as training samples. The first preset proportion can be determined according to actual experience, for example, it can be 80%.
[0140] S240: Input the training samples into the local outlier factor system for model training to obtain a local outlier factor model.
[0141] Input the training samples into the local outlier factor system for model training. By successively calculating the K-nearest neighbor distance, reachability distance, and local reachability density of each CAN bus data sequence in the training samples, the local outlier factor of each CAN bus data sequence is finally obtained. And in each model training process, the local range of each CAN bus data sequence is continuously updated to obtain the calculation accuracy of the optimized local outlier factor, and finally a local outlier factor model that meets the preset accuracy requirements is obtained.
[0142] After obtaining the local outlier factor model based on the training samples, the obtained local outlier factor model can be tested to determine whether the local outlier factor model needs to be further optimized. Specifically, according to the T CAN bus data sequences after deduplication and the preset abnormal data construction rules, a test sample including multiple test CAN bus data sequences can be constructed, where the multiple test CAN bus data sequences include CAN bus data sequences without abnormal data and CAN bus data sequences with abnormal data; input the test sample into the local outlier factor model to obtain the local outlier factor of each test CAN bus data sequence; obtain the abnormal data test result by comparing the local outlier factors of the multiple test CAN bus data sequences with the preset abnormal data detection threshold, where the abnormal data test result includes the test result of whether there is abnormal data in each test CAN bus data sequence; calculate the abnormal detection accuracy rate by comparing the abnormal data test result with the CAN bus data sequences with preset abnormal data in the test sample; if the abnormal detection accuracy rate is less than the preset accuracy rate threshold, increase the sample size of the training sample, and continue to train the local outlier factor model with the training sample with the increased sample size until the abnormal detection accuracy rate is greater than or equal to the preset accuracy rate threshold, and obtain the finally required local outlier factor model.
[0143] Among them, the construction process of the test sample includes: obtaining all CAN bus data sequences except the training samples from the T CAN bus data sequences after deduplication as the test CAN bus data sequences without abnormal data in the test sample; obtaining the j-th byte of each CAN bus data from the T CAN bus data sequences after deduplication, where j is a positive integer; counting whether the values corresponding to all the j-th bytes include all the values within the preset value range; if all the values within the preset value range are included, select a preset number of values from the preset value range as the abnormal data; if not all the values within the preset value range are included, determine the preset number of values not included in the preset value range as the abnormal data; select at least one CAN bus data sequence from the T CAN bus data sequences after deduplication as the target CAN bus data sequence; replace the j-th byte of the preset number of CAN bus data in the target CAN bus data sequence with the preset number of abnormal data to generate a test CAN bus data sequence including abnormal data; after generating multiple test CAN bus data sequences including abnormal data for multiple different bytes, form a test sample by combining the multiple test CAN bus data sequences including abnormal data and the test CAN bus data sequences without abnormal data. Among them, the preset number is an empirical value, for example, the preset number is 1.
[0144] Exemplarily, if the T de-duplicated CAN bus data sequences include 100,000 CAN bus data sequences, then the first 80,000 CAN bus data sequences can be selected as training samples, and the last 20,000 CAN bus data sequences that are test CAN bus data sequences without abnormal data. Select the first byte of each CAN bus data from the 100,000 CAN bus data sequences, and count whether the values corresponding to all the first bytes include all the values within [0, 255]. If all the values within [0, 255] are included, then select a preset number of values from [0, 255] as abnormal data (such as selecting 5 as abnormal data). If only all the values within [0, 200] are included and the values within [201, 255] are not included, then a preset number of values can be selected from [201, 255] as abnormal data (such as selecting 234 as abnormal data). After obtaining the abnormal data, a CAN bus data sequence can be selected from the 100,000 CAN bus data sequences as the target CAN bus data sequence, and then a CAN bus data is selected from the target CAN bus data sequence. Replace the first byte of the selected CAN bus data with the abnormal data 5 or 234. The generated CAN bus data sequence after replacement is the test CAN bus data sequence containing abnormal data in the test sample. By repeatedly executing the above process for different bytes multiple times, multiple CAN bus data sequences containing abnormal data can be obtained. Finally, the 20,000 test CAN bus data sequences without abnormal data and the multiple test CAN bus data sequences with abnormal data are shuffled to form the test sample.
[0145] In addition, obtaining the abnormal data test result by comparing the local abnormal factors of multiple test CAN bus data sequences with the preset abnormal data detection threshold includes: respectively comparing the local abnormal factors of multiple test CAN bus data sequences with the preset abnormal data detection threshold. If the local abnormal factor of the test CAN bus data sequence to be compared is greater than the preset abnormal data detection threshold, it is determined that there is abnormal data in the test CAN bus data sequence to be compared. If the local abnormal factor of the test CAN bus data sequence to be compared is less than or equal to the preset abnormal data detection threshold, it is determined that there is no abnormal data in the test CAN bus data sequence to be compared.
[0146] Corresponding to the above method embodiments, an embodiment of the present invention provides an abnormal detection device for CAN bus data, as Figure 3 shown, the device includes:
[0147] The first acquisition unit 30 is configured to acquire N CAN bus data to be detected, where the identity identification numbers ID of different CAN bus data among the N CAN bus data are different, and N is a positive integer;
[0148] The combination unit 32 is configured to combine the N CAN bus data into a CAN bus data sequence to be detected;
[0149] The second acquisition unit 34 is configured to input the CAN bus data sequence to be detected into the local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected, where the local outlier factor model is a model trained according to multiple CAN bus data sequences for calculating the local outlier factor, and the number of data and the type of ID of each CAN bus data sequence in the multiple CAN bus data sequences are both N;
[0150] The determination unit 36 is configured to determine that there is abnormal data in the N CAN bus data if the local outlier factor is greater than a preset abnormal data detection threshold.
[0151] Optionally, the apparatus further includes:
[0152] The acquisition unit is configured to acquire T CAN bus data before inputting the CAN bus data sequence to be detected into the local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected, where the T CAN bus data include N types of CAN bus data with different IDs, and T is a positive integer;
[0153] The first construction unit is configured to construct an initialized CAN bus data sequence with all data contents being 0, where the number of data in the initialized CAN bus data sequence is N*X, X is the maximum number of bytes of each CAN bus data among the T CAN bus data, and every X adjacent data in the initialized CAN bus data sequence represent a CAN bus data of one ID;
[0154] The replacement unit is configured to start from the first CAN bus data of the T CAN bus data and sequentially replace the CAN bus data with corresponding IDs in the initialized CAN bus data sequence to generate T CAN bus data sequences, where the i-th CAN bus data sequence is obtained by replacing the CAN bus data with the corresponding ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data among the T CAN bus data, and i≥2;
[0155] A deduplication unit, configured to, if there are duplicate CAN bus data sequences among the T CAN bus data sequences, perform deduplication processing on the T CAN bus data sequences, and obtain a first preset proportion of CAN bus data sequences from the deduplicated T CAN bus data sequences as training samples;
[0156] A training unit, configured to input the training samples into a local outlier factor system for model training to obtain a local outlier factor model.
[0157] Optionally, the device further includes:
[0158] A selection unit, configured to, after inputting the training samples into a local outlier factor system for model training to obtain a local outlier factor model, select a second preset proportion of CAN bus data sequences from the training samples;
[0159] The second obtaining unit 34 is further configured to input the second preset proportion of CAN bus data sequences into the local outlier factor model to obtain the local outlier factor of each CAN bus data sequence in the second preset proportion of CAN bus data sequences;
[0160] The determination unit 36 is further configured to determine the largest local outlier factor among the local outlier factors of each CAN bus data sequence in the second preset proportion of CAN bus data sequences as the preset abnormal data detection threshold.
[0161] Optionally, the device further includes:
[0162] A second construction unit, further configured to, after inputting the training samples into a local outlier factor system for model training to obtain a local outlier factor model, construct a test sample including a plurality of test CAN bus data sequences according to the deduplicated T CAN bus data sequences and a preset abnormal data construction rule, where the plurality of test CAN bus data sequences include CAN bus data sequences without abnormal data and CAN bus data sequences with abnormal data;
[0163] The second obtaining unit 34 is further configured to input the test sample into the local outlier factor model to obtain the local outlier factor of each test CAN bus data sequence;
[0164] The device further includes:
[0165] A comparison unit, configured to obtain an abnormal data test result by comparing the local outlier factors of a plurality of test CAN bus data sequences with the preset abnormal data detection threshold, where the abnormal data test result includes the test result of whether there is abnormal data in each test CAN bus data sequence;
[0166] A calculation unit, configured to calculate the correct rate of anomaly detection by comparing the anomaly data test result with a CAN bus data sequence with preset anomaly data in the test samples.
[0167] The training unit is configured to, if the correct rate of anomaly detection is less than a preset correct rate threshold, increase the sample size of the training samples, and continue to train the local outlier factor model with the training samples with the increased sample size until the correct rate of anomaly detection is greater than or equal to the preset correct rate threshold, so as to obtain the finally required local outlier factor model.
[0168] Optionally, the second construction unit includes:
[0169] An acquisition module, configured to acquire, from the T CAN bus data sequences after duplicate removal, all CAN bus data sequences except the training samples as test CAN bus data sequences without anomaly data in the test samples; and acquire the j-th byte of each CAN bus data from the T CAN bus data sequences after duplicate removal, where j is a positive integer.
[0170] A statistics module, configured to count whether the values corresponding to all the j-th bytes include all the values within a preset value range.
[0171] A selection module, configured to, if all the values within the preset value range are included, select a preset number of values from within the preset value range as anomaly data, and if not all the values within the preset value range are included, determine a preset number of values selected from the values not included in the preset value range as anomaly data.
[0172] A selection module, configured to select at least one CAN bus data sequence from the T CAN bus data sequences after duplicate removal as a target CAN bus data sequence.
[0173] A first replacement module, configured to replace the j-th byte of a preset number of CAN bus data in the target CAN bus data sequence with a preset number of anomaly data to generate a test CAN bus data sequence including anomaly data.
[0174] A construction module, configured to, after generating multiple test CAN bus data sequences including anomaly data for multiple different bytes, form a test sample with the multiple test CAN bus data sequences including anomaly data and the test CAN bus data sequences without anomaly data.
[0175] Optionally, the replacement unit includes:
[0176] A supplementary module, configured to, if the number of bytes of the i-th CAN bus data among the T CAN bus data is less than X, pad zeros at the tail of the i-th CAN bus data until the number of bytes after padding equals X, and then replace the CAN bus data with the corresponding ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data after padding zeros;
[0177] A second replacement module, configured to, if the number of bytes of the i-th CAN bus data equals X, replace the CAN bus data with the corresponding ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data.
[0178] Optionally, the first acquisition unit 30 includes:
[0179] An acquisition module, configured to acquire N CAN bus data to be detected;
[0180] A conversion module, configured to, if there is non-decimal CAN bus data among the N CAN bus data to be detected acquired, convert the non-decimal CAN bus data into decimal CAN bus data.
[0181] Optionally, the apparatus further includes:
[0182] A judgment unit, configured to, after determining that there is abnormal data among the N CAN bus data, for each CAN bus data among the N CAN bus data, use the abnormal data detection model corresponding to the CAN bus data to judge whether the CAN bus data is abnormal data, where different CAN bus data with different IDs correspond to different abnormal data detection models, and the abnormal data detection model is a model trained based on multiple CAN bus data with the same ID and used to identify whether CAN bus data is abnormal.
[0183] Based on the above method embodiments, another embodiment of the present invention provides a storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor implements the method as described above.
[0184] Based on the above method embodiments, another embodiment of the present invention provides an electronic device, including:
[0185] One or more processors;
[0186] A storage device, configured to store one or more programs,
[0187] wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.
[0188] The above system and device embodiments correspond to the method embodiments and have the same technical effects as the method embodiments. For specific descriptions, please refer to the method embodiments. The device embodiments are obtained based on the method embodiments. For specific descriptions, please refer to the method embodiment section and will not be elaborated here. Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0189] Those of ordinary skill in the art can understand that the modules in the devices in the embodiments can be distributed in the devices in the embodiments as described in the embodiments, or can be correspondingly changed and located in one or more devices different from the present embodiment. The modules in the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An abnormal detection method for CAN bus data, characterized in that The method includes: Obtain N CAN bus data to be detected, where the identity identification numbers ID of different CAN bus data among the N CAN bus data are different, and N is a positive integer; Combine the N CAN bus data into a CAN bus data sequence to be detected; Input the CAN bus data sequence to be detected into a local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected, where the local outlier factor model is a model trained based on multiple CAN bus data sequences and used to calculate the local outlier factor. For each CAN bus data sequence among the multiple CAN bus data sequences, the number of CAN bus data and the type of ID are both N; If the local outlier factor is greater than a preset abnormal data detection threshold, determine that there is abnormal data among the N CAN bus data; Before inputting the CAN bus data sequence to be detected into the local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected, the method further includes: Collect T CAN bus data, where the T CAN bus data include N types of CAN bus data with different IDs; Construct an initialized CAN bus data sequence with all data contents being 0, where the number of data in the initialized CAN bus data sequence is N*X, X is the maximum number of bytes among the bytes of each CAN bus data in the T CAN bus data, and every X adjacent data in the initialized CAN bus data sequence represent a CAN bus data of one ID; Starting from the first CAN bus data in the T CAN bus data, sequentially replace the CAN bus data of the corresponding ID in the initialized CAN bus data sequence to generate T CAN bus data sequences, where the i-th CAN bus data sequence is obtained by replacing the CAN bus data of the corresponding ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data in the T CAN bus data, and i≥2; If there are duplicate CAN bus data sequences among the T CAN bus data sequences, perform deduplication processing on the T CAN bus data sequences, and obtain a first preset proportion of CAN bus data sequences from the deduplicated T CAN bus data sequences as training samples; Input the training samples into a local outlier factor system for model training to obtain a local outlier factor model; After determining that there is abnormal data among the N CAN bus data, the method further includes: For each CAN bus data among the N CAN bus data, use the abnormal data detection model corresponding to the CAN bus data to determine whether the CAN bus data is abnormal data. Among them, different CAN bus data with different IDs correspond to different abnormal data detection models, and the abnormal data detection model is a model trained based on multiple CAN bus data with the same ID and used to identify whether CAN bus data is abnormal.
2. The method according to claim 1, characterized in that, After inputting the training samples into the Local Outlier Factor (LOF) system for model training to obtain the LOF model, the method further includes: Selecting CAN bus data sequences with a second preset ratio from the training samples; Inputting the CAN bus data sequences with the second preset ratio into the LOF model to obtain the local outlier factors of each CAN bus data sequence in the CAN bus data sequences with the second preset ratio; Determining the maximum local outlier factor among the local outlier factors of each CAN bus data sequence in the CAN bus data sequences with the second preset ratio as the preset abnormal data detection threshold; 3. The method according to claim 1, wherein After inputting the training samples into the LOF system for model training to obtain the LOF model, the method further includes: Constructing a test sample including a plurality of test CAN bus data sequences according to the deduplicated T CAN bus data sequences and the preset abnormal data construction rule, where the plurality of test CAN bus data sequences include CAN bus data sequences without abnormal data and CAN bus data sequences with abnormal data; Inputting the test sample into the LOF model to obtain the local outlier factors of each test CAN bus data sequence; Obtaining the abnormal data test result by comparing the local outlier factors of the plurality of test CAN bus data sequences with the preset abnormal data detection threshold, where the abnormal data test result includes the test result of whether there is abnormal data in each test CAN bus data sequence; Calculating the abnormal detection accuracy rate by comparing the abnormal data test result with the CAN bus data sequences with abnormal data preset in the test sample; If the abnormal detection accuracy rate is less than the preset accuracy rate threshold, increase the sample size of the training samples, and continue to train the LOF model with the training samples with the increased sample size until the abnormal detection accuracy rate is greater than or equal to the preset accuracy rate threshold, and then obtain the finally required LOF model.
4. The method according to claim 3, wherein Constructing a test sample including a plurality of test CAN bus data sequences according to the deduplicated T CAN bus data sequences and the preset abnormal data construction rule, including: Obtaining all CAN bus data sequences except the training samples from the deduplicated T CAN bus data sequences as the test CAN bus data sequences without abnormal data in the test sample; Obtaining the j-th byte of each CAN bus data from the deduplicated T CAN bus data sequences, where j is a positive integer; Counting whether the values corresponding to all the j-th bytes include all the values within the preset value range; If all the values within the preset value range are included, selecting a preset number of values from the preset value range as the abnormal data; If not all the values within the preset value range are included, determining a preset number of values selected from the values not included in the preset value range as the abnormal data; Select at least one CAN bus data sequence from the T CAN bus data sequences after duplicate removal as the target CAN bus data sequence; Replace the j-th byte of the preset number of CAN bus data in the target CAN bus data sequence with the preset number of abnormal data to generate a test CAN bus data sequence including abnormal data; After generating multiple test CAN bus data sequences including abnormal data for multiple different bytes, form a test sample with the multiple test CAN bus data sequences including abnormal data and the test CAN bus data sequence without abnormal data.
5. The method according to claim 1, wherein Replacing the i-th CAN bus data in the T CAN bus data with the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence includes: If the number of bytes of the i-th CAN bus data in the T CAN bus data is less than X, pad zeros at the end of the i-th CAN bus data so that the number of bytes after padding is equal to X, and then replace the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence with the padded i-th CAN bus data; If the number of bytes of the i-th CAN bus data is equal to X, replace the CAN bus data corresponding to the ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data.
6. The method according to claim 1, wherein Obtaining the N CAN bus data to be detected includes: Collect the N CAN bus data to be detected; If there is non-decimal CAN bus data among the collected N CAN bus data to be detected, convert the non-decimal CAN bus data into decimal CAN bus data.
7. An abnormal detection device for CAN bus data, characterized in that, The device includes: A first acquisition unit for acquiring N CAN bus data to be detected, where the identity identification numbers ID of different CAN bus data among the N CAN bus data are different, and N is a positive integer; A combination unit for combining the N CAN bus data into a CAN bus data sequence to be detected; A second acquisition unit for inputting the CAN bus data sequence to be detected into a local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected, where the local outlier factor model is a model trained based on multiple CAN bus data sequences for calculating the local outlier factor, and the number of CAN bus data and the type of ID of each CAN bus data sequence in the multiple CAN bus data sequences are both N; A determination unit for determining that there is abnormal data in the N CAN bus data if the local outlier factor is greater than a preset abnormal data detection threshold; The device further includes: An acquisition unit for acquiring T CAN bus data before inputting the CAN bus data sequence to be detected into the local outlier factor model to obtain the local outlier factor of the CAN bus data sequence to be detected, where the T CAN bus data include N types of CAN bus data with different IDs, and T is a positive integer; A first construction unit for constructing an initialized CAN bus data sequence with all data contents being 0, where the number of data in the initialized CAN bus data sequence is N*X, X is the maximum number of bytes among the number of bytes of each CAN bus data in the T CAN bus data, and every X adjacent data in the initialized CAN bus data sequence represent CAN bus data of one ID; A replacement unit for starting from the first CAN bus data of the T CAN bus data and sequentially replacing the CAN bus data of the corresponding ID in the initialized CAN bus data sequence to generate T CAN bus data sequences, where the i-th CAN bus data sequence is obtained by replacing the CAN bus data of the corresponding ID in the (i - 1)-th CAN bus data sequence with the i-th CAN bus data in the T CAN bus data, and i≥2; A deduplication unit for, if there are duplicate CAN bus data sequences among the T CAN bus data sequences, performing deduplication processing on the T CAN bus data sequences and obtaining a first preset proportion of CAN bus data sequences from the deduplicated T CAN bus data sequences as training samples; A training unit for inputting the training samples into a local outlier factor system for model training to obtain a local outlier factor model; A judgment unit for, after determining that there is abnormal data among the N CAN bus data, for each CAN bus data among the N CAN bus data, using the abnormal data detection model corresponding to the CAN bus data to judge whether the CAN bus data is abnormal data, where different ID CAN bus data correspond to different abnormal data detection models, and the abnormal data detection model is a model trained based on multiple CAN bus data of the same ID for identifying whether CAN bus data is abnormal; 8. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, where, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 - 6.
Citation Information
Patent Citations
Anomaly detection method and device for vehicle-mounted CAN bus
CN112491920A
Virtual machine cluster anomaly detection method based on outlier factor, equipment and medium
CN113191432A