Method for Evaluating the Richness of In-vehicle Network Intrusion Detection Datasets Based on Improved SOM
Through the improved SOM, multi-layer splitting and clustering of vehicle network data sets is solved, the problem of insufficient evaluation dimensions in the existing technology is achieved, and more accurate data set richness evaluation is achieved, which improves the robustness of the model.
Patent Information
- Application Number
- CN202310356707.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-04
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-04-04
AI Technical Summary
The existing evaluation methods of in-vehicle network intrusion detection data sets cover limited evaluation dimensions, which leads to inaccurate evaluation results and makes it difficult to accurately compare the richness of different data sets.
The improved self-organized mapping network (SOM) is used to perform multi-layer splitting and clustering of the on-board network data sets. By setting clustering parameters and distance correlation indicators, the richness of the data set is evaluated, including Euclidean distance, Hamming distance, Pearson coefficient and Spearma coefficient, etc., to form a richness evaluation matrix.
It improves the accuracy of data set richness evaluation, can more accurately compare the richness of different data sets, and enhances the robustness of the model in practical applications.
Smart Images

Figure CN116471065B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle networking information security testing, and particularly relates to a method for evaluating the richness of an in-vehicle network intrusion detection data set based on an improved SOM. Background Art
[0002] Currently, with the development of automotive intelligence and networking, the interfaces between vehicles and the outside world are gradually increasing, and at the same time, the dependence on communication and perception data is also gradually increasing. Once a vehicle encounters an information security attack by hackers, it may affect the functions of the vehicle and endanger the lives and property safety of drivers and passengers. In-vehicle network intrusion detection technology is to detect malicious intrusions in the in-vehicle network. However, the training of intrusion detection models usually requires a large amount of data as support. Existing data sets are usually simulated by computers or extracted from real vehicles, and some malicious attacks in the data sets are also inserted artificially. Moreover, the vehicle models, working states, scenarios, attack types, and types of inserted attacks involved in these data sets are different. Therefore, it is often difficult to directly compare intrusion detection models trained and tested based on different data sets through surface accuracy and other indicators. Therefore, it is of great significance to evaluate the richness of the data set. Generally speaking, the larger the richness of the data set, the more complete the scenarios it contains. After the model trained based on this data set reaches a certain accuracy, its robustness in actual applications is higher. However, in the existing technology, the detection of the richness of the data set is usually evaluated based on a single piece of information, resulting in limited evaluation dimensions covered and inaccurate evaluation results. Summary of the Invention
[0003] In view of the above analysis, the present invention provides a method for evaluating the richness of an in-vehicle network intrusion detection data set based on an improved SOM, which solves the problems of limited evaluation dimensions covered and inaccurate evaluation results in the existing evaluation methods for richness evaluation.
[0004] A method for evaluating the richness of an in-vehicle network intrusion detection data set based on an improved SOM according to the present invention is characterized by including the following specific steps:
[0005] Step 1: Obtain an in-vehicle network data set; split the data packets in the data set layer by layer according to the characteristics of the in-vehicle network protocol; separate the data packets according to the data label type;
[0006] Step 2: Obtain the multi-layer headers and payloads of the data packets; extract the data features of all headers according to the data packet type;
[0007] Step 3, set SOM clustering parameters; perform SOM clustering on the dataset, data features and payloads of the packets and headers of normal types and all attack types respectively based on the SOM clustering parameters; obtain the clustering results of the in-vehicle network dataset; obtain the richness evaluation matrix of the in-vehicle network dataset based on the clustering results.
[0008] Optionally, the data label types include normal labels and attack labels; based on the data label types, the packet types of the in-vehicle network dataset are divided into normal types and the i-th attack type, where i = 0, 1, 2, …, m, and m is the total number of attack categories.
[0009] Optionally, the specific steps of Step 3 are as follows:
[0010] Set SOM clustering parameters;
[0011] Set the initial number of neurons n*n, where n*n represents the initial number of SOM categories, and n = 1, 2, 3, …;
[0012] Initialize the weight vector of each neuron;
[0013] Obtain the input vector D(t) of the t-th data of the dataset, packets, headers and payloads respectively;
[0014] Obtain the weight vector W v (s) of the v-th neuron in the s-th iteration, where v = 0, 1, 2, 3, … n*n and s = 0, 1, 2, 3, … S, and S is the total number of iterations;
[0015] Initialize the iterative weight vector W v (0);
[0016] Based on the SOM clustering parameters, obtain the distance and correlation between the input vector D(t) of the t-th data of the dataset, packets, headers and payloads respectively and the weight vector W v (s) of each SOM neuron;
[0017] Obtain the winning neuron u; update the iterative weight vector based on the winning neuron u;
[0018] After S iterations, cluster the dataset on n*n neurons;
[0019] If the distance between two neurons is less than the distance threshold ts, classify these two neurons as the same type of neurons; retain the neurons whose passing data is greater than and / or equal to the clustering threshold; obtain the minimum number of clusters under the distance threshold according to the retained neurons;
[0020] Obtain the richness evaluation matrix of the in-vehicle network dataset based on the minimum number of clusters.
[0021] Optionally, in step 3, when clustering the data packets, headers, and payloads of the entire data set, n > i + 1; when clustering the data packets, headers, and payloads of all attack types, n > i.
[0022] Optionally, when performing SOM clustering, the methods for obtaining distances include the Euclidean distance method and the Hamming distance method; the methods for obtaining correlations include the Pearson coefficient method, the cosine correlation method, and the Spearma coefficient method.
[0023] Optionally, the SOM clustering parameters include a distance metric and a correlation metric.
[0024] Optionally, the distance metrics include the Euclidean distance and the Hamming distance; the correlation metrics include the Pearson coefficient, the cosine correlation coefficient, and the Spearma coefficient.
[0025] Optionally, it further includes step 4, obtaining an evaluation matrix of the richness of multiple in-vehicle network data sets; comparing the evaluation matrices of the richness of multiple in-vehicle network data sets to obtain a comparison result.
[0026] The present invention has at least the following beneficial effects: Based on the improved Self-Organizing Maps (SOM), the present invention clusters various information, and compares the richness of the same parts of different data sets through the number of categories, thereby improving the evaluation accuracy.
[0027] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can be made obvious from the description, or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained from the content specifically pointed out in the description and the drawings. Description of the Drawings
[0028] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs represent the same components.
[0029] Figure 1 It is a flowchart of the evaluation method of the present invention.
[0030] Figure 2 It is a flowchart of the multi-data set richness comparison of the evaluation method of the present invention. Detailed Embodiments
[0031] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings, where the drawings form a part of the present invention and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0032] A specific embodiment of the present invention, as Figure 1 shown, discloses a method for evaluating the richness of an in-vehicle network intrusion detection dataset based on an improved SOM, including the following specific steps:
[0033] Step 1: Obtain the in-vehicle network dataset; split the data packets in the dataset layer by layer according to the characteristics of the in-vehicle network protocol; separate the data packets according to the data label type;
[0034] Optionally, the in-vehicle network dataset includes a real vehicle dataset and / or a virtual dataset; the in-vehicle network dataset is used for in-vehicle network intrusion detection;
[0035] Optionally, the data label types include normal labels and attack labels; based on the data label types, the data packet types of the in-vehicle network dataset are divided into normal types and the i-th type of attack type, where i = 0, 1, 2, …, m, and m is the total number of attack categories.
[0036] Step 2: Obtain the multi-layer headers and payloads of the data packets; extract the data features of all headers according to the data packet types;
[0037] Furthermore, the data features in the headers include the timestamp of the CAN bus, the arbitration domain ID of the CAN bus, the message ID of the SOME / IP dataset, the message length of the SOME / IP dataset, and the message type of the SOME / IP dataset.
[0038] Step 3: Set the data types and clustering methods of the dataset, data packets, headers, and payloads; perform preprocessing such as base conversion and missing value supplementation on the data; set the SOM clustering parameters; perform SOM clustering on the dataset, data packets of normal types and all attack types, the data features and payloads of the headers respectively based on the SOM clustering parameters; obtain the clustering results of the in-vehicle network dataset; obtain the richness evaluation matrix of the in-vehicle network dataset based on the clustering results;
[0039] The specific steps are as follows:
[0040] Step 31: Obtain the data types of the dataset, data packets, headers, and payloads, as well as the clustering method;
[0041] The data types are binary, decimal, and hexadecimal; perform base conversion on the input data bit by bit according to the set data types; for data with missing values, supplement with -1, thereby reducing the change in the final clustering results caused by data processing.
[0042] Step 32: Set the SOM clustering parameters;
[0043] Optionally, the SOM clustering parameters include a distance metric and a correlation metric.
[0044] Optionally, the distance metrics include Euclidean distance and Hamming distance; the correlation metrics include Pearson coefficient, cosine correlation coefficient, and Spearman coefficient.
[0045] Step 33: Set the initial number of neurons to n*n, where n*n represents the initial number of SOM categories, and n = 1, 2, 3...;
[0046] Optionally, when clustering the dataset, all types of data packets, data types of headers, and payloads, n > i + 1; when clustering all types of attack-type data packets, data types of headers, and payloads, n > i.
[0047] Step 34: Initialize the weight vector of each neuron;
[0048] Optionally, the dimension in the weight vector of the neuron is the same as the dimension of the data to be clustered. For example, when clustering the payload part of the CAN bus intrusion detection dataset, the maximum input data length is 8 bytes. If clustering its binary data using Hamming distance, 8 * 8 = 64 dimensions are required. If clustering its hexadecimal data using Euclidean distance, 8 * 2 = 16 dimensions are required.
[0049] It can be understood that the data to be clustered includes the dataset, data packets, headers, and payloads.
[0050] Step 35: Obtain the input vector D(t) of the t-th data of the dataset, data packets, headers, and payloads to be clustered respectively;
[0051] Step 36: Obtain the weight vector W v (s) of the v-th neuron in the s-th iteration, where v = 0, 1, 2, 3... n*n, s = 0, 1, 2, 3... S, and S is the total number of iterations; initialize the iterative weight vector W v (0);
[0052] Step 37: Based on the SOM clustering parameters, obtain the distances and correlations between the input vector D(t) of the t-th data of the dataset, data packets, headers, and payloads to be clustered respectively and the weight vector W v (s) of each SOM neuron;
[0053] Optionally, when performing SOM clustering, the methods for obtaining distances include Euclidean distance method and Hamming distance method, and the methods for obtaining correlations are Pearson coefficient method, cosine correlation method, and Spearman coefficient method.
[0054] Preferably, t = s + 1; when the data type is binary data type, the Hamming distance is used to extract the data features of the header; when the data type is decimal data or hexadecimal data type, the Euclidean distance is used to extract the data features of the header; binary data, decimal data, and hexadecimal data can be converted to each other;
[0055] Among them, the expression for obtaining the Euclidean distance by the Euclidean distance method is:
[0056]
[0057] In the formula, d O represents the Euclidean distance, x O k represents the k-th group of decimal and hexadecimal data in the input vector, and y O k represents the k-th group of decimal and hexadecimal data in the SOM neuron weight vector;
[0058] The expression for obtaining the Hamming distance by the Hamming distance method is:
[0059]
[0060] In the formula, d H represents the Hamming distance, x H k represents the k-th bit of binary data in the input vector, and y H k represents the k-th bit of binary data in the SOM neuron weight vector, represents the exclusive OR operation;
[0061] Exemplarily, the binary data in the input vector is 11101010, and the binary data in the SOM neuron weight vector is 11011010. It can be calculated that the Hamming distance between the two groups of binary data is 2;
[0062] The expression for obtaining the cosine correlation by the cosine correlation method is:
[0063]
[0064] In the formula, r cos represents the cosine correlation, x C k represents the k-th bit of binary data in the input vector, and y C k represents the k-th bit of binary data in the SOM neuron weight vector.
[0065] Step 38, obtain the winning neuron u; update the iterative weight vector based on the winning neuron u, and the expression: W v (s + 1) = Wv (s) + h(u, s)g(D(t) - W v (s)), where h is the neighborhood and learning rate function;
[0066] Step 39: After S iterations, cluster the data on n * n neurons;
[0067] Step 310: If the distance between two neurons is less than the distance threshold ts, classify the two neurons as the same type of neurons; retain the neurons with data passing through greater than and / or equal to the clustering threshold; obtain the minimum number of clusters under the distance threshold based on the retained neurons;
[0068] Optionally, delete the neurons through which no data passes or the data passing through is less than the clustering threshold.
[0069] It can be understood that if the distance between the weight vectors W1(s) and W2(s) of the first neuron and the second neuron is less than the distance threshold ts, the first neuron and the second neuron are classified as the same type of neurons.
[0070] Step 310: Obtain the evaluation matrix of the richness of the in - vehicle network dataset based on the minimum number of clusters.
[0071] Optionally, the clustering result is the number of clusters.
[0072] Furthermore, evaluate each part of the dataset, including the header, payload, normal data, attack data, and specific header features (such as: timestamp, source ID, destination ID), etc., and finally obtain the evaluation matrix.
[0073]
[0074] In the formula, M in the matrix j represents the j - th type of message represented by the input vector. When j = 1, it represents all the messages represented by the input vector. When j = 2, it represents the CAN header represented by the input vector. When j = 3, it represents the CAN data represented by the input vector. When j = 4, it represents the timestamp represented by the input vector. When j = 5, it represents the arbitration field represented by the input vector. When j = 6, it represents the control field represented by the input vector. When j = 7, it represents the CRC represented by the input vector. When j = 8, it represents the ACK represented by the input vector. N represents the j - th type of normal message represented in the input vector, (j = 1, 2, L, 8), A represents the j - th (j = 1, 2, L, 8) type of abnormal message represented in the input vector, (j = 1, 2, L, 8), a1 j represents the first type of attack in the abnormal messages for the j - th (j = 1, 2, L, 8) type of message, a2 jIndicates the second type of attack in the abnormal messages for the j-th (j = 1, 2, L, 8) type of messages, aT j Indicates the n-th type of attack in the abnormal messages for the j-th (j = 1, 2, L, 8) type of messages, d O Indicates the Euclidean distance, d H Indicates the Hamming distance, r Pearson Indicates the Pearson correlation coefficient, r Spearma Indicates the Spearma correlation coefficient, and δ is the clustering result corresponding to each row and column of the matrix.
[0075] Step 4: Obtain different evaluation matrices for the richness of in-vehicle network datasets based on Steps 1 to 3; compare different evaluation matrices for the richness of in-vehicle network datasets to obtain a comparison result.
[0076] Optionally, the elements included in the evaluation matrix for the richness of in-vehicle network datasets are the categories of the clustering results under the same index. The more categories there are, the stronger and higher the richness of the dataset, thereby realizing the richness evaluation between different datasets.
[0077] Embodiment 1
[0078] Use an improved SOM based on distance metrics and correlation metrics to cluster the in-vehicle network intrusion detection dataset, and evaluate its richness through the clustering results. This method includes four steps: the first three steps are respectively dataset input and segmentation, data feature extraction, and SOM clustering. Then, perform the first three steps on different datasets, compare the clustering results under the same clustering threshold, and thus conduct a richness evaluation. If the number of categories in the clustering results is large, it is considered that the richness is strong.
[0079] Step 1: Obtain the Car-hacking-dataset-v2 dataset of the CAN bus, and divide the dataset into 5 parts according to the labels in the dataset as normal, Flooding, Spoofing, Replay, and Fuzzing.
[0080] Step 2: Split the data packets in the dataset according to the characteristics of the CAN protocol, divide them into multiple layers of headers and payloads, and extract the data features in the headers one by one according to the CAN protocol type: timestamp, arbitration domain ID, and DLC (download content).
[0081] Step 3: Use SOM clustering based on the Hamming distance metric for the headers, payloads, and the whole data packets of the normal type and various attack types of data respectively. Set the initial number of neurons to 3 * 3, initialize the weights of each neuron. After the dataset is input, when the t-th data of the dataset is input, an input vector D(t) is formed, and calculate the weight vector of the v-th neuron at the s-th iteration as W v(s), calculate the distance between the input vector and the SOM neuron weight vector using the Hamming distance, then calculate the winning neuron u, traverse each neuron in the SOM, and update the iterative weight vector to: W v (s + 1) = W v (s) + h(u, s)g(D(t) - W v (s)), where h is the neighborhood and learning rate function. Finally, obtain the hit counts of 9 neurons and the distances between neurons, set the Hamming distance threshold l0, merge the data and classes with Hamming distance less than the threshold, and finally obtain the clustering results of different parts of the dataset. Summarize the number of clusters into a matrix to obtain the richness evaluation matrix of Car-hacking-dataset-v2 under the above clustering parameters.
[0082] Step 4, perform the first three steps on another in-vehicle network Car-hacking-dataset-v1 dataset. During the clustering process, select the same Hamming distance statistical index and the same clustering parameters respectively to obtain the richness evaluation matrix of Car-hacking-dataset-v1. Compare the number of clusters in their common parts, such as the whole, normal messages, and Fuzzing, respectively, to realize the richness comparison of the two datasets.
[0083] As described above, only the specific embodiments of the present invention are preferred, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. An evaluation method for the richness of an in-vehicle network intrusion detection dataset based on an improved SOM, characterized in that It includes the following specific steps: Step 1: Obtain the in-vehicle network dataset; split the data packets in the dataset layer by layer according to the characteristics of the in-vehicle network protocol; separate the data packets according to the data label type; Step 2: Obtain the multi-layer headers and payloads of the data packets; extract the data characteristics of all headers according to the data packet type; Step 3: Set the SOM clustering parameters; Based on the SOM clustering parameters, perform SOM clustering on the dataset, the data packets of the normal type and all attack types, the data characteristics and payloads of the headers respectively; obtain the clustering results of the in-vehicle network dataset; obtain the richness evaluation matrix of the in-vehicle network dataset based on the clustering results. The specific steps are as follows: Set the SOM clustering parameters; Set the initial number of neurons n*n, where n*n represents the initial number of SOM categories, and n = 1, 2, 3…; Initialize the weight vectors of each neuron; Obtain the input vector D(t) of the t-th data of the dataset, data packets, headers and payloads respectively; Obtain the weight vector \(W^{(s)}\) of the \(v\)-th neuron, where \(v = 0, 1, 2, 3,\cdots,n\times n\) and \(s = 0, 1, 2, 3,\cdots,S\), and \(S\) is the total number of iterations. v (s), \(v = 0, 1, 2, 3,\cdots,n\times n\), \(s = 0, 1, 2, 3,\cdots,S\), where \(S\) is the total number of iterations; Initialize the iterative weight vector W v (0); Based on the SOM clustering parameters, obtain the distance and correlation between the input vector D(t) of the t-th data of the dataset, data packet, header, and payload respectively and each SOM neuron weight vector W v (s); Obtain the winning neuron u; update and iterate the weight vector based on the winning neuron u; After S iterations, cluster the dataset on n*n neurons; If the distance between two neurons is less than the distance threshold ts, classify these two neurons as the same type of neurons; retain the neurons whose passed data is greater than and / or equal to the clustering threshold; Obtain the minimum number of clusters under the distance threshold according to the retained neurons; Obtain the richness evaluation matrix of the in-vehicle network dataset based on the minimum number of clusters.
2. The evaluation method according to claim 1, wherein The data label types include normal labels and attack labels; based on the data label types, the data packet types of the in-vehicle network dataset are divided into normal types and the i-th type of attack types, where i = 0, 1, 2, …, m, and m is the total number of attack categories.
3. The evaluation method according to claim 1, characterized in that In Step 3, when clustering the data packets, headers and payloads of the entire dataset, n > i + 1; when clustering the data packets, headers and payloads of all attack types, n > i.
4. The evaluation method according to claim 2, wherein When performing SOM clustering, the methods for obtaining distances include the Euclidean distance method and the Hamming distance method; the methods for obtaining correlations include the Pearson coefficient method, the cosine correlation method and the Spearma coefficient method.
5. The evaluation method according to claim 1, wherein The SOM clustering parameters include distance metrics and correlation metrics.
6. The evaluation method according to claim 5, wherein The distance metrics include the Euclidean distance and the Hamming distance; the correlation metrics include the Pearson coefficient, the cosine correlation coefficient and the Spearma coefficient.
7. The evaluation method according to any one of claims 1 to 6, characterized in that It also includes Step 4: Obtain the richness evaluation matrices of multiple in-vehicle network datasets; compare the richness evaluation matrices of multiple in-vehicle network datasets to obtain the comparison results.
Citation Information
Patent Citations
Intrusion detection method based on incremental GHSOM (Growing Hierarchical Self-organizing Maps) neural network
CN102789593A
Vehicle-mounted network intrusion detection method based on message sequence prediction
CN110149345A
Power line communication noise identification method and device based on self-organizing mapping neural network
CN113489514A