Graph neural network multi-classification attack detection method based on automobile CAN benign traffic
By adopting a multi-classified attack detection scheme based on CAN benign traffic in automotive networks, the problem of dependence on real attack data in the existing technology is solved, more efficient and accurate attack detection is achieved, operation and maintenance costs are reduced, and higher-level security guarantees are provided.
Patent Information
- Application Number
- CN202510076286.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing multi-classified attack detection system based on deep learning is difficult to obtain enough labeled real attack data in automotive networks, resulting in high cost of model training and maintenance and affecting detection accuracy.
The multi-classification attack detection scheme of graph neural network based on benign CAN traffic is adopted. The information collection module extracts normal CAN ID sequences and statistical information from benign traffic through the information collection module. The attack sequence generation module generates different types of attack sequences. The data preprocessing module converts the sequence into undirected and unrighteous graph data. The graph attention neural network model extracts features from the graph data and classifies them.
Reduces dependence on real attack data, reduces the cost of model training, improves the accuracy of the detection system, and has more refined attack analysis capabilities to provide higher-level security guarantees.
Smart Images

Figure CN120017329A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automobile network security, and in particular to a multi-classification attack detection method for an automobile control area network based on deep learning. Background Art
[0002] With the continuous development of automobile intelligence and networking technology, modern cars are gradually equipped with a large number of electronic control units (ECUs) based on the Controller Area Network (CAN). These ECUs achieve efficient data exchange through the CAN bus to ensure the various functions and safety of the car. However, with the popularization of network technology, the network security issues of automobiles are becoming increasingly prominent, especially in vehicle intelligent control systems. The risk of network attacks cannot be ignored. Network attackers may tamper with, interfere with or destroy the vehicle control system by invading the CAN bus, thereby threatening the safety of the vehicle and the life of the driver. Therefore, how to effectively detect and defend against various attacks in the automotive network has become a technical problem that needs to be solved urgently.
[0003] At present, intrusion detection technology based on deep learning has become a mainstream method in the field of network security, especially for intrusion detection of CAN buses. Traditional multi-classification attack detection systems based on deep learning usually rely on a large amount of attack data for training. However, attack data of automotive networks is often difficult to obtain, and there are many types of attacks with high concealment and variability, which makes the collection of real attack data extremely complicated. Due to the inability to obtain enough real attack data with labels, existing deep learning models have to frequently retrain the trained models, which not only increases the cost of training and maintenance of attack detection models, but also affects the accuracy of the detection system. Summary of the invention
[0004] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a graph neural network multi-classification attack detection solution based on CAN benign traffic. The technical solution of the present invention is as follows:
[0005] A graph neural network multi-classification attack detection scheme based on CAN benign traffic, characterized by comprising: an information collection module, an attack sequence generation module, a data preprocessing module, and a graph neural network model with a graph attention mechanism; wherein the information collection module is used to extract normal CAN ID sequences and some statistical information from the benign traffic of the automobile control area network;
[0006] The attack sequence generation module is used to generate a spoof attack, impersonation attack, and pause attack CAN ID sequence for each high-frequency CAN ID;
[0007] The data preprocessing module is used to convert various types of CAN ID sequences into undirected unweighted graph data, and to represent the complex timing information of the vehicle control area network by utilizing the structural characteristics of the graph;
[0008] The graph attention neural network model is used to extract important features from different types of graph data and classify the graph data according to the extracted features;
[0009] Furthermore, the information collection module is used to extract normal CANID sequences from benign traffic in the vehicle control area network and statistically analyze high-frequency CAN IDs and their corresponding frequency information, specifically including:
[0010] ① Input the benign traffic log data of the vehicle control area network into the information collection module. The information collection module first counts the average number p of CAN messages contained in the traffic log data every t milliseconds, and then extracts the arbitration field of each CAN message, i.e., CAN ID, in sequence, so as to obtain the normal CAN ID sequence S Normal ={ID1,ID2,…,ID N}.
[0011] ② For normal CAN ID sequence S Normal ={ID1,ID2,…,ID N}Remove duplicate CAN ID values to obtain the normal CAN ID node list L contained in this type of vehicle Normal ={ID1,ID2,…,ID n}.
[0012] ③ Divide the normal CAN ID sequence into N / p units, and then count the normal CAN ID node list L Normal The average value of each ID in each unit is used to obtain the frequency information of each CAN ID, generate frequency information, and then filter out the high-frequency CAN ID according to the set high-frequency threshold T, and finally use the high-frequency CAN ID value ID h and its corresponding frequency information f h Stored into statistical information and input into the attack sequence generation module.
[0013] Furthermore, the attack sequence generation module is used to generate CAN ID sequences of various attack types according to the normal CAN ID sequence and statistical information provided by the information collection module, specifically including:
[0014] ① ID for high frequency CAN ID h The generation process of the spoofing attack sequence is as follows: first, the normal CAN ID sequence S Normal Divide into groups, and then insert the value ID into a random position in each group h Finally, concatenate the CAN ID sequences of all groups in sequence to generate the ID h Deception attack sequence.
[0015] ② ID for high frequency CAN ID h The generation process of the pause attack sequence is as follows: replace the S Normal All values are ID h Delete the CAN ID value and generate the ID h Pause attack sequence.
[0016] ③ ID for high frequency CAN ID h The generation process of the impersonation attack sequence is as follows: follow the method of step ① to h Insert multiple IDs into the pause attack sequence h , thereby generating h Impersonation attack sequence.
[0017] ④ Repeat steps ①②③ until spoofing, impersonation, and pause attack sequences are generated for all high-frequency CAN ID values.
[0018] Furthermore, the data preprocessing module is used to convert different types of CAN ID sequences into undirected unweighted graphs, specifically including:
[0019] ① The data preprocessing module first divides each CAN ID sequence into a group sequence containing p CAN ID numbers, and labels each group sequence with a corresponding label according to the type of CAN ID sequence.
[0020] ② In the control area network traffic log data, the CAN ID usually consists of three hexadecimal digits. In order to facilitate the processing of the graph neural network model, it needs to be converted into a three-dimensional feature vector. The specific conversion method is: first convert the three hexadecimal digits into three-dimensional decimal digits, and then normalize each digit by dividing each decimal digit by 16, and then splice the processed data to obtain a three-dimensional feature vector. For example, the CAN ID with a value of 69E becomes [0.375, 0.5625, 0.875] after vectorization processing.
[0021] ③ In order to convert each group sequence into undirected unweighted graph data, first convert all CAN ID values in the group sequence into three-dimensional feature vectors according to step ②, then use the converted CAN ID as the node in the graph data, and then establish edges between the CAN ID nodes of the graph data according to the adjacency relationship between the group sequences. If there are repeated edges, no processing will be performed. Finally, the graph data is labeled with corresponding labels according to the label information provided in step ①.
[0022] Furthermore, the graph attention neural network model is used to learn and extract features of different types of generated graph data, and classify them according to the extracted features, specifically including:
[0023] ① A 4-layer or 5-layer graph attention neural network model is used to extract the relationship between each node and its neighbor nodes in undirected unweighted graph data. The self-attention mechanism assigns different weights to each node neighbor, so that the feature aggregation of each node not only depends on the neighbor nodes, but also can better express the relationship between nodes by dynamically adjusting the weights.
[0024] ② A convergence layer uses a global average pooling operation to average the features of all nodes in the graph to obtain a global vector representation of the graph.
[0025] ③A Dropout layer is used to randomly "discard" some neurons during the training process to prevent the model from overfitting on the training data. By allowing the network to use only some neurons at a time, Dropout forces the network to learn more robust features, enabling it to have better generalization capabilities on new, unseen data.
[0026] ④A linear layer is used to map the global features of the graph to the output space of different categories and generate predictions for the categories of the graph data.
[0027] The advantages and beneficial effects of the present invention are as follows:
[0028] Compared with the prior art, the invention has the following three advantages:
[0029] (1) The present invention relies only on benign traffic log data in the vehicle control area network (CAN) to achieve multi-classification attack detection on the vehicle CAN bus system. This innovative design effectively reduces the reliance on real attack data, thereby significantly reducing the cost of collecting real attack data during model training. Since traditional intrusion detection systems usually require a large amount of labeled attack data for training, and the acquisition of attack data is very difficult and costly, the present invention breaks this bottleneck by using benign traffic to generate different types of attack data, reducing the demand for real attack data and reducing the operation and maintenance costs of the intrusion detection system.
[0030] (2) The present invention introduces a self-attention mechanism, which enables the model to automatically and dynamically focus on the most representative and critical parts when processing input data. This mechanism enables the model to more effectively extract valuable information from complex input features, thereby achieving more accurate feature learning and attack classification.
[0031] (3) The multi-classification attack detection scheme proposed in the present invention has a more refined attack analysis capability. Security personnel can not only determine which CAN ID the attacker has launched an attack against based on the detection results, but can also further infer whether the ECU that originally sent the CAN signal has been compromised. This feature enables the present invention to provide a higher level of security protection, helping security personnel to identify and respond to complex attack behaviors in a timely manner. Through real-time detection and multi-dimensional attack analysis, security personnel can take further effective measures to enhance the overall security of the automotive CAN bus system. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 The present invention generally relates to a flow chart;
[0033] Figure 2 This is an example of network traffic log data for the car control area;
[0034] Figure 3 It is a structural diagram of the graph neural network model of the present invention; DETAILED DESCRIPTION
[0035] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and in detail describe the technical solutions in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention. Figure 1-3 , the specific implementation methods of the present invention are as follows:
[0036] 1. Collect the network traffic data of the control area under various driving scenarios of the car and classify it according to Figure 2 The data is collected in an exemplary manner, and the collected data content specifically includes the timestamp of each CAN data received, the arbitration field of the CAN data frame, namely, the CAN ID, and the data domain part of the CAN data frame.
[0037] 2. Input the collected vehicle control area network traffic data into the information collection module. The information collection module extracts the CAN ID sequence and statistical data information from the normal control area network benign traffic log data. The specific implementation steps are as follows:
[0038] ① The information collection module first counts the average number of CAN messages contained in the traffic log data every 80 milliseconds, and then extracts the arbitration field of each CAN message, namely the CAN ID, in sequence, thereby obtaining the normal CAN ID sequence S Normal ={ID1,ID2,…,ID N}.
[0039] ② For normal CAN ID sequence S Normal ={ID1,ID2,…,ID N}Remove duplicate CAN ID values to obtain the normal CAN ID node list L contained in this type of vehicle Normal ={ID1,ID2,…,ID n}.
[0040] ③ Divide the normal CAN ID sequence into N / p units, and then count the normal CAN ID node list L Normal The average value of each ID in each unit is used to obtain the frequency information of each CAN ID, generate frequency information, and then filter out the high-frequency CAN ID according to the set high-frequency threshold 3, and finally use the high-frequency CAN ID value ID h and its corresponding frequency information f h Stored in statistics.
[0041] 3. The attack sequence generation module generates various types of attack CAN ID sequences based on the normal CAN ID sequence and statistical information provided by the information collection module. The specific implementation steps are as follows:
[0042] ① ID for high frequency CAN ID h The generation process of the spoofing attack sequence is as follows: first, the normal CAN ID sequence S Normal Divide into groups, and then insert the value ID into a random position in each group h Finally, concatenate the CAN ID sequences of all groups in order to generate the ID h Deception attack sequence.
[0043] ② ID for high frequency CAN ID h The generation process of the pause attack sequence is as follows: replace the S of the normal CAN ID sequence Normal All values are ID h Delete the CAN ID value and generate the ID h Pause attack sequence.
[0044] ③ ID for high frequency CAN ID hThe generation process of the impersonation attack sequence is as follows: follow the method of step ① to h Insert multiple IDs into the pause attack sequence h , thereby generating h Impersonation attack sequence.
[0045] ④ Repeat steps ①②③ until spoofing, impersonation, and pause attack sequences are generated for all high-frequency CAN ID values.
[0046] 4. Divide the CAN ID sequence transmitted by the attack sequence generation module into group sequences containing the same number of CAN IDs, and convert each group sequence into undirected unweighted graph data and label it accordingly. The specific steps are as follows: Data preprocessing module
[0047] ① The data preprocessing module first divides each CAN ID sequence into group sequences containing pp CAN IDs, and labels each group sequence according to the type of CAN ID sequence.
[0048] ② To facilitate graph neural network processing, the CAN ID (usually a three-digit hexadecimal number) in the control area network traffic log data needs to be converted into a three-dimensional feature vector. The specific conversion method is: first convert the hexadecimal number into a three-dimensional decimal number; then normalize each decimal number by dividing it by 16; finally, splice the processed numbers into a three-dimensional feature vector. For example, the CAN ID value of "69E" becomes [0.375, 0.5625, 0.875] after conversion.
[0049] ③ After converting the CAN IDs in each CAN ID sequence into three-dimensional feature vectors, use these feature vectors as nodes in the graph data and establish edges between the nodes based on the adjacency relationship of the group sequence. If duplicate edges appear, no processing is done. According to the label information provided in step (1), the graph data is labeled accordingly.
[0050] 5. Use Python's PyTorch and PyTorch Geometric libraries to build a graph attention neural network (GAT) model. The network model architecture is as follows Figure 3 As shown in the figure, the specific structure includes 4 to 5 layers of graph attention convolution layers (GATConv), which are used to extract important features from graph data. The configuration of each layer is as follows:
[0051] The first graph attention convolution layer: the number of input neural units is 3, the number of hidden neural units is 128, and the number of attention heads is 4;
[0052] For the remaining graph attention convolutional layers, the number of input neural units is 128×4, the number of hidden neural units is 128, and the number of attention heads is 4;
[0053] A pooling layer uses the global_mean_pool function to compress the graph data into a one-dimensional feature vector;
[0054] A Dropout layer to improve the generalization ability of the model;
[0055] A linear layer is used to perform classification based on the extracted graph data features. The number of input neural units is 128×4, and the number of output neural units is the number of attack categories to be detected.
[0056] 6. During the training process of the graph attention neural network model, the graph data set generated by the data preprocessing module is divided into 80% as a training set and 20% as a test set to train the graph attention neural network model. The number of training epochs is 200.
[0057] 7. After the model training is completed, the data preprocessing module and the graph attention neural network model are deployed to the CAN bus system of the same vehicle model. When the data preprocessing module receives a quantity of p CAN data, it converts it into graph data and inputs it into the trained graph attention neural network model. By analyzing the output results of the model, it is determined whether there is an attack in the control area network traffic data and the type of attack is identified.
[0058] The above embodiments should be understood to be only used to illustrate the present invention and not to limit the protection scope of the present invention. After reading the contents of the present invention, technicians can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A graph neural network multi-classification attack detection scheme based on automobile CAN benign traffic, characterized in that: include: An information collection module, an attack sequence generation module, a data preprocessing module, and a graph neural network model with a graph attention mechanism; wherein the information collection module is used to extract normal CAN ID sequences and related statistical information from benign traffic in a vehicle control area network; wherein the attack sequence generation module is used to generate CAN ID sequences such as spoofing attacks, impersonation attacks, and pause attacks for each high-frequency CAN ID; the data preprocessing module is used to convert various types of CAN ID sequences into undirected and unweighted graph data, and use the structural features of the graph to represent complex timing information in the vehicle control area network; the graph attention neural network model is used to extract important features from the graph data and classify the graph data based on the extracted features.
2. According to claim 1, a graph neural network multi-classification attack detection scheme based on automobile CAN benign traffic is characterized in that: Used to extract normal CAN ID sequences and related statistical information from benign traffic in the automotive control area network, including: Input the benign traffic log data of the vehicle control area network and count the average number p of CAN messages in each t millisecond time period; Extract the arbitration field CAN ID of each CAN message, obtain the normal CAN ID sequence, and remove duplicate CAN IDs to obtain a normal CAN ID node list; The frequency information of each CAN ID is counted, and the high-frequency CAN ID and its frequency information are filtered out according to the set high-frequency threshold T for use by the attack sequence generation module.
3. According to claim 1, a graph neural network multi-classification attack detection scheme based on automobile CAN benign traffic is characterized in that: The attack sequence generation module is used to generate CAN ID sequences of various attack types according to the normal CAN ID sequence and statistical information provided by the information collection module, specifically including: By randomly inserting high-frequency CAN IDs into the normal CAN ID sequence, a spoofing attack sequence targeting high-frequency CAN IDs is generated; By removing the high-frequency CAN ID from the normal CAN ID sequence, a pause attack sequence targeting the high-frequency CAN ID is generated; By inserting a high-frequency CAN ID into a pause attack sequence, an impersonation attack sequence targeting a high-frequency CAN ID is generated.
4. According to claim 1, a graph neural network multi-classification attack detection solution based on benign automobile traffic is characterized in that: The data preprocessing module converts a fixed-size CAN ID sequence into an undirected unweighted graph, specifically including: Divide different types of CAN ID sequences into group sequences containing p CAN IDs, and label each group sequence accordingly; Convert each CAN ID into a three-dimensional feature vector, and convert the group sequence into undirected unweighted graph data based on the adjacency relationship; Label each graph data according to the category to which the CAN ID sequence belongs.
5. According to claim 1, a graph neural network multi-classification attack detection scheme based on automobile CAN is characterized in that: The graph attention neural network model classifies each graph data according to the learned graph data features, specifically including: 4 or 5 layers of graph attention convolutional layers, which assign different weights to each node and its neighboring nodes through the self-attention mechanism, thereby extracting the relationship between nodes; The global average pooling layer is used to aggregate the features of all nodes in the graph and generate a global feature representation of the graph; The Dropout layer is used to prevent overfitting and improve the generalization ability of the model; The linear layer maps the global features of the graph to the category output space and generates predictions for the categories of the graph data. Through this graph neural network model, important features are extracted from the graph data, and the graph data is classified into multiple categories based on these features.
Citation Information
Patent Citations
Vehicle-mounted network variant attack intrusion detection method and system based on domain adversarial neural network
CN114157469A
Vehicle-mounted network intrusion detection method supporting data privacy protection
CN116471062A
CAN-FD anomaly detection method based on time sequence content attention and long and short term memory network
CN117176421A
Automatic driving vehicle intrusion detection method based on time and space
CN118677669A
Automobile CAN bus message hybrid attack simulation method based on large model
CN118677686A