In-vehicle CAN bus intrusion detection system and method based on artificial intelligence

Through the in-vehicle CAN bus intrusion detection system based on artificial intelligence, the convolutional length and short-term memory network and convolutional neural network are used to solve the problem of CAN bus safety hazards in intelligent connected vehicles, and high accuracy and real-time intrusion detection are achieved.

CN120012077APending Publication Date: 2025-05-16SUN YAT SEN UNIV

Patent Information

Application Number
CN202510110324.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The security risks of CAN buses in intelligent connected vehicles have increased significantly, and the lack of effective security protection mechanisms has led to malicious attackers being able to eavesdrop on or inject malicious information through remote or physical access, interfering with vehicle control.

Method used

The in-vehicle CAN bus intrusion detection system is adopted based on artificial intelligence. Through data preprocessing, feature extraction and time series analysis technology, the intrusion behavior is identified using the convolutional length and short-term memory network (ConvLSTM) and the convolutional neural network (CNN) and the inflection neural network (CNN) to output the intrusion probability value for detection.

Benefits of technology

It improves the accuracy and real-time nature of CAN bus intrusion detection, enhances the ability to identify key features, can effectively process time-dependent CAN data, and provides comprehensive protection measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012077A_ABST
    Figure CN120012077A_ABST
Patent Text Reader

Abstract

The invention discloses an in-vehicle CAN bus intrusion detection system and method based on artificial intelligence. The method comprises the following steps: S1, converting original data collected by an in-vehicle CAN bus into a CAN image sequence; s2, extracting key features of the CAN image sequence; and S3, performing time sequence analysis on the features output by the feature extraction module by using a convolutional long-short-term memory network, and outputting an intrusion probability value as a judgment result of intrusion detection. Compared with the prior art, the deep learning technology is utilized, and the accuracy and real-time performance of intrusion detection are effectively improved through data preprocessing, feature extraction and time sequence analysis technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent automobile safety technology, and more specifically, to an in-vehicle CAN bus intrusion detection system and method based on artificial intelligence. Background Art

[0002] With the rapid development of intelligent connected vehicle technology, the networking and intelligence level of vehicles has been significantly improved. Intelligent connected vehicles are becoming more and more popular and gradually becoming an important development direction of the future automotive industry. According to the forecast of International Data Corporation (IDC), the global sales of connected vehicles are expected to reach 78.3 million in 2025, which fully demonstrates the growth trend of its market demand. However, at the same time, automotive network security has also become an issue that needs to be solved urgently.

[0003] As the core underlying communication protocol of the vehicle network, the Controller Area Network (CAN) is widely used in the information exchange of various electronic control units (ECUs) in the vehicle. The CAN bus has the advantages of simple structure, good real-time performance, and high reliability. Multiple ECU nodes in the vehicle are connected in parallel through a pair of twisted pairs and transmit data in a broadcasting manner, which makes the CAN bus an important choice for traditional vehicle networks. However, when the CAN bus was first designed, it was mainly aimed at the internal communication scenarios of traditional vehicles. Its communication protocol only introduced a cyclic redundancy check (CRC) to detect the integrity of the data frame, but did not provide a security protection mechanism against malicious intrusion. This design is not a significant problem in the context of traditional vehicle use, because traditional vehicles lack networking functions, the in-vehicle network structure is relatively simple, and it is difficult for external attackers to access the vehicle's CAN bus. Therefore, the security risks of the CAN bus are relatively low.

[0004] However, the automotive industry has ushered in the rapid development of networking and intelligence. Intelligent connected vehicles have multiple networking functions such as Bluetooth, Wi-Fi, and cellular communications. They can be connected to the vehicle through mobile phone apps. At the same time, the number of electronic control units (ECUs) inside the vehicle has also increased rapidly, making the communication scenarios of the CAN bus more complex. Although the popularity of these networking and intelligent functions has brought more convenient and intelligent experiences to drivers and passengers, it has also greatly increased the exposure of the CAN bus and provided malicious attackers with multiple potential intrusion paths. Attackers can eavesdrop on the communication data of the vehicle network through remote access or physical connection, and even inject malicious information, thereby achieving the purpose of controlling the vehicle or interfering with its normal operation. This makes the security risks of the CAN bus increasingly significant, especially in the context of the widespread application of intelligent connected vehicles, and its security issues have become an important challenge restricting the development of the industry.

[0005] Based on the above background, intrusion detection technology for CAN bus has received widespread attention in recent years. Compared with adding security protection measures at the hardware level, software-based intrusion detection methods are more feasible and flexible because they do not require changes to the hardware architecture of existing vehicles, can be quickly deployed and cover a wide range of application scenarios. At present, CAN bus intrusion detection technology is mainly divided into traditional methods and artificial intelligence-based methods. Traditional methods include frequency detection, statistical analysis, etc., while artificial intelligence-based methods combine deep learning technologies such as convolutional neural networks (CNN) and long short-term memory networks (LSTM), showing stronger capabilities in extracting features and identifying attack behaviors. With the continuous deepening of information security research on intelligent connected vehicles, CAN bus intrusion detection technology based on artificial intelligence will become a key means to ensure vehicle network security.

[0006] The prior art discloses a CAN bus attack detection method based on ResNet and AGRU, including: preprocessing CAN data set data; mapping the characteristic values ​​of timestamp (Timestamp), identifier (CAN ID) and data field (CANData) in the CAN data frame to the values ​​of image pixels to generate image data; inputting the generated image data into a convolutional neural network (ResNet) to extract spatial features; inputting the extracted spatial features into a gated recurrent unit (GRU) for time series modeling, adding an attention mechanism, and weighted summing the time series data through attention weights; the final result is mapped to the range of 0-1 through an activation function, and judging whether the data frame has intrusion behavior through a threshold.

[0007] To this end, in combination with the above requirements and the defects of the prior art, the present application proposes an in-vehicle CAN bus intrusion detection system and method based on artificial intelligence. Summary of the invention

[0008] The present invention provides an in-vehicle CAN bus intrusion detection system and method based on artificial intelligence, which utilizes deep learning technology to effectively improve the accuracy and real-time performance of intrusion detection through data preprocessing, feature extraction and time series analysis technology.

[0009] The primary purpose of the present invention is to solve the above technical problems. The technical solution of the present invention is as follows:

[0010] The first aspect of the present invention provides an in-vehicle CAN bus intrusion detection method based on artificial intelligence, and the method comprises the following steps:

[0011] S1. Convert the raw data collected by the CAN bus in the vehicle into a CAN image sequence.

[0012] S2. Extract key features of the CAN image sequence.

[0013] S3. Use a convolutional long short-term memory network to perform time series analysis on the features output by the feature extraction module, and output an intrusion probability value as a judgment result of intrusion detection.

[0014] Furthermore, the step S1 includes the following specific steps: first, receiving a data frame on the CAN bus, cleaning the data frame, extracting the ID and the corresponding data field in hexadecimal format to obtain a CAN frame list, and converting it line by line into a binary format to obtain a coded vector, and superimposing the obtained coded vector to reconstruct a two-dimensional CAN image; wherein, in the process of converting into a coded vector, if the length of the data field is less than 64 bits, 0 is added to the end of the field to ensure that the final length of the vector is 76 bits; the CAN image is a two-dimensional image in the shape of [n, 76], where n represents a hyperparameter that controls the image length.

[0015] Furthermore, in the process of creating the CAN image sequence, a sliding window is moved on the original frame list with a step size of 10 frames each time to construct a CAN image sequence of length l, where l is a hyperparameter that controls the amount of contextual information provided by the network. The larger the l is, the longer the image sequence is and the more information it contains. After creating the CAN image sequence, it is also necessary to set a classification label for it. When at least one frame in the frame list corresponding to the CAN image sequence is an attack frame, the label is set to "attack"; if all frames are normal, it is set to "normal".

[0016] Furthermore, in step S1, the data preprocessing process is specifically as follows:

[0017] Extract the ID and data fields:

[0018] ID i =extract_id(f i )

[0019] D i =extract_data_field(f i )

[0020] where f i Represents several raw CAN data frames, and then converts the data into a binary vector:

[0021] B i = to_binary(ID i ,12)+to_binary(D i ,64)

[0022] Where ID i Convert to 12-bit binary number, Di Convert to a binary number with 8 bits per byte, a total of 64 bits, and fill any bits less than 64 bits with 0; convert the generated binary vector to an image:

[0023] C img =[B1,B2,…,B n ]

[0024] The serialized binary vector B i Further converted into CAN image, accumulating n B i To form the final CAN image, where n is a hyperparameter that determines the final size of the image; finally, construct the CAN image sequence:

[0025] Seq(k)=[C img,k ,C img,k+S ,…,C img,k+(L-1)S ]

[0026] Where Seq(k) represents an image sequence of length L starting from the kth image, and each C img,j is an independent CAN image, and k is the starting index of the sequence.

[0027] Furthermore, in step S2, the key features are extracted by using a convolutional neural network to extract key features for intrusion detection from the CAN image sequence; the convolutional neural network includes two convolutional layers, an activation layer and a pooling layer, and a convolutional block attention module is introduced into the convolutional neural network, and the convolutional block attention module enhances the feature map through channel attention focusing and spatial attention focusing; wherein, the channel attention is implemented by global average pooling and global maximum pooling, and the spatial attention aggregates the channel information of the feature map by calculating the average and maximum values ​​of the feature map along the direction of the channel axis, and the vectors are connected and input into a convolutional layer with a kernel size of 7×7 to obtain attention in the spatial dimension.

[0028] Furthermore, the feature extraction process in step S2 is:

[0029] Feature_Seq(k)=CBAM(CNN(Seq(k)))

[0030] The convolutional neural network includes multiple layers of convolution and activation layers, and the first layer of convolution is:

[0031] F1 = LeakyReLU(Conv(F in ))

[0032] Where F1 is the output feature map of the first convolution layer, F in is a single CAN image input; the second layer of convolution is:

[0033] F2 = LeakyReLU(Conv(F1))

[0034] Where F2 is the output feature map of the second convolution layer; the average pooling layer connected thereafter is:

[0035] F out ={AvgPool}(F2)

[0036] Among them, F out is the output feature map after average pooling, the output F out Input to the convolution block attention module for processing to obtain the enhanced feature map; the calculation expression of the convolution block attention module is:

[0037] F′=W SAM ⊙(W CAM ⊙ F)

[0038] W CAM =σ(MLP(AvgPool(F))+MLP(MaxPool(F)))

[0039] W SAM =σ(Conv 7×7 [Avg c (W CAM ⊙F); Max c (W CAM ⊙F)])

[0040] Among them, F and F′ represent the input feature map and enhanced feature map respectively, W SAM With W CAM They represent the weights calculated by the CAM and SAM submodules respectively, σ represents the sigmoid activation function, the subscript c represents the operation along the channel axis, and ⊙ represents the Hadamard product.

[0041] Furthermore, in step S3, the convolutional long short-term memory network is composed of a plurality of convolutional long short-term memory layers, each of which includes an input gate, a forget gate, and an output gate, and the calculation expression of the convolutional long short-term memory layer is as follows:

[0042]

[0043] in W it , W ft , W ot They represent the input vector, unit state, unit hidden state, input gate, forget gate, and output gate at time t, respectively. i 、b f 、b c 、b oIndicates the corresponding unit bias, the symbol * is the convolution operator, ⊙ is the Hadamard product, σ is the sigmoid function as the activation function, and tanh is the hyperbolic tangent function as the activation function; in the last part of the convolutional long short-term memory network, two fully connected layers are set as the classification decision network, and a random drop mechanism is equipped, and the Sigmoid function outputs the probability value as the judgment result of the entire model.

[0044] Furthermore, the specific calculation process of the time series analysis is as follows: for each time step t and each convolutional long short-term memory layer m, the input feature map F m,t Receive data from the output of the previous layer or the state of the previous time step:

[0045]

[0046] X in,t is the network input at time step t, F m-1,t,out is the output of the previous layer at the same time step; each layer uses the ConvLSTM unit to update its cell state C m,t And calculate the output F m,t,out :

[0047] C m,t , F m,t,out =ConvLSTM(F m-1,t,out , S l,t-1 )

[0048] Among them, F m-1,t,out is the output of the previous layer at the same time step, S m,t-1 is the state of the same layer at the previous time step; at the last layer of each time step, the output F M,t,out Can be used as the final network output or passed to the next time step:

[0049] F out,t =F M,t,out

[0050] Among them, M is the last layer, F out,t is the final output at this time step.

[0051] Furthermore, after the time series analysis, a classification decision network is required to output the final classification result to determine whether there is an intrusion behavior in the current sequence. The specific process is: first, the input feature map is flattened to convert the multi-dimensional features into a one-dimensional vector; then, the feature transformation and dimensionality reduction are performed through two fully connected layers, and the classification result is output. The predicted classification result is compared with the actual label assigned by the data preprocessing module in the early stage, and the binary cross entropy function is used to calculate the loss value of the current model. Based on the calculated loss value, the model parameters in step S2 and step S3 are automatically adjusted through the back propagation algorithm. The specific calculation process is as follows:

[0052] y1=Dropout(LeakyReLU(W1x+b1), p=0.2)

[0053] Classification=Softmax(W2y1+b2)

[0054] Where y1 represents the output of the first linear layer, the flattened input is linearly transformed, and the LeakyReLU activation function is directly applied, and then the Dropout operation is performed, W1 is the weight matrix, b1 is the bias vector, x is the flattened input vector, and p is the Dropout probability; the second linear layer receives the output y1 of the first linear layer, performs a linear transformation, and applies the Softmax activation function to output the final classification probability, where W2 and b2 are the weight matrix and bias vector of the second linear layer.

[0055] The second aspect of the present invention provides an in-vehicle CAN bus intrusion detection system based on artificial intelligence, which is used for the in-vehicle CAN bus intrusion detection method based on artificial intelligence, and includes: a data preprocessing module, a feature extraction module and a time series analysis module.

[0056] The data preprocessing module converts the raw data collected by the in-vehicle CAN bus into a CAN image sequence and inputs it into the feature extraction module. The feature extraction module uses a convolutional neural network to extract the key features of the CAN image sequence. The time series analysis module uses a convolutional long short-term memory network to perform time series analysis on the features output by the feature extraction module, and outputs an intrusion probability value as the judgment result of intrusion detection.

[0057] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0058] The present invention provides an in-vehicle CAN bus intrusion detection system and method based on artificial intelligence. The adopted convolutional long short-term memory network and time series analysis method can enhance the model's recognition ability for key features and process CAN data with time dependence. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 The present invention is a flow chart of an in-vehicle CAN bus intrusion detection method based on artificial intelligence.

[0060] Figure 2 A schematic diagram of CAN image creation in an embodiment of the present invention.

[0061] Figure 3 The figure is a schematic diagram of the generation process of the CAN image sequence in one embodiment of the present invention.

[0062] Figure 4 Schematic diagram of the structure of a convolutional block attention module in one embodiment of the present invention.

[0063] Figure 5 The present invention is a schematic diagram of an in-vehicle CAN bus intrusion detection system based on artificial intelligence.

[0064] Figure 6 It is a schematic diagram of the interaction between various modules of the system in an embodiment of the present invention.

[0065] Figure 7 This is a confusion matrix image diagram in one embodiment of the present invention.

[0066] Figure 8 It is a ROC curve and AUC image diagram of the model in one embodiment of the present invention. DETAILED DESCRIPTION

[0067] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0068] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0069] Example 1

[0070] like Figure 1 As shown, the present invention provides an in-vehicle CAN bus intrusion detection method based on artificial intelligence, and the method comprises the following steps:

[0071] S1. Convert the raw data collected by the CAN bus in the vehicle into a CAN image sequence.

[0072] S2. Extract key features of the CAN image sequence.

[0073] S3. Use a convolutional long short-term memory network to perform time series analysis on the features output by the feature extraction module, and output an intrusion probability value as a judgment result of intrusion detection.

[0074] The step S1 includes the following specific steps: first, receiving a data frame on the CAN bus, cleaning the data frame, extracting the ID and the corresponding data field in hexadecimal format to obtain a CAN frame list, and converting it line by line into a binary format to obtain a coded vector, and superimposing the obtained coded vector to reconstruct a two-dimensional CAN image; wherein, in the process of converting to a coded vector, if the length of the data field is less than 64 bits, 0 is added to the end of the field to ensure that the final length of the vector is 76 bits; the CAN image is a two-dimensional image in the shape of [n, 76], where n represents a hyperparameter that controls the image length.

[0075] In a specific embodiment, the process of creating a CAN image is as follows: Figure 2 As shown in the figure, through this conversion, the CAN bus data is reconstructed into a two-dimensional image. The converted data not only retains most of the useful information of the original CAN bus data, but also provides rich visual information for subsequent feature extraction through graphical representation, which is easier for subsequent feature extraction and analysis.

[0076] like Figure 3 As shown in the figure, in the process of creating the CAN image sequence, a sliding window is moved on the original frame list, and the step length of each movement is 10 frames to construct a CAN image sequence of length, where is a hyperparameter that controls the amount of context information provided by the network. The larger it is set, the longer the image sequence is and the more information it contains. After creating the CAN image sequence, it is also necessary to set a classification label for it. When at least one frame in the frame list corresponding to the CAN image sequence is an attack frame, the label is set to "attack". If all frames are normal, it is set to "normal".

[0077] It should be noted that in real-world CAN injection attacks, the distribution of attack frames is usually much sparser than that of normal frames, which makes it difficult for machine learning algorithms to accurately identify the type of attack. Therefore, two types of classification labels are used to distinguish between those that require alarm when intrusion is discovered and those that do not require alarm in normal situations to achieve better performance.

[0078] In step S1, the data preprocessing process is specifically as follows:

[0079] Extract the ID and data fields:

[0080] ID i =extract_id(f i )

[0081] Di =extract_data_field(f i )

[0082] where f i Represents several raw CAN data frames, and then converts the data into a binary vector:

[0083] B i = to_binary(ID i ,12)+to_binary(D i , 64)

[0084] Where ID i Convert to 12-bit binary number, D i Convert to a binary number with 8 bits per byte, a total of 64 bits, and fill any bits less than 64 bits with 0; convert the generated binary vector to an image:

[0085] C img =[B1, B2, ..., B n ]

[0086] The serialized binary vector B i Further converted into CAN image, accumulating n B i To form the final CAN image, where n is a hyperparameter that determines the final size of the image; finally, construct the CAN image sequence:

[0087] Seq(k)=[C img,k ,C img,k+S ,…,C img,k+(L-1)S ]

[0088] Where Seq(k) represents an image sequence of length L starting from the kth image, and each C img,j is an independent CAN image, and k is the starting index of the sequence.

[0089] In step S2, the key features are extracted by using a convolutional neural network to extract the key features for intrusion detection from the CAN image sequence; the convolutional neural network includes two convolutional layers, an activation layer and a pooling layer, and Figure 4 As shown in the figure, a convolutional block attention module is introduced in the convolutional neural network, and the convolutional block attention module enhances the feature map through channel attention focusing and spatial attention focusing; wherein, the channel attention is implemented by global average pooling and global maximum pooling, and the spatial attention aggregates the channel information of the feature map by calculating the average and maximum values ​​of the feature map along the direction of the channel axis, and the vectors are connected and input into the convolutional layer with a kernel size of 7×7 to obtain attention in the spatial dimension.

[0090] It should be noted that a multi-layer CNN architecture is adopted, which includes 2 convolutional layers, an activation layer and a pooling layer. The convolutional layer is responsible for extracting spatial features, the activation layer (such as ReLU) introduces nonlinearity, and the pooling layer is used to reduce the feature dimension and enhance the translation invariance of the feature. In order to adapt to the characteristics of CAN images, the present invention designs specific convolution kernels. The size, number and depth of these convolution kernels are carefully selected and adjusted to maximize the effect of feature extraction. Experiments show that the convolution kernel designed by the present invention can effectively capture abnormal patterns in CAN bus data.

[0091] In the feature extraction process, the depth and width of the feature map extracted by the CNN model have an important impact on the subsequent downstream module effects. The present invention balances the complexity and performance of the model by adjusting the depth and width of the model output, ensuring that the model maintains efficient operation while extracting rich features.

[0092] The feature extraction process in step S2 is:

[0093] Feature_Seq(k)=CBAM(CNN(Seq(k)))

[0094] The convolutional neural network includes multiple layers of convolution and activation layers, and the first layer of convolution is:

[0095] F1 = LeakyReLU(Conv(F in ))

[0096] Where F1 is the output feature map of the first convolution layer, F in is a single CAN image input; the second layer of convolution is:

[0097] F2 = LeakyReLU(Conv(F1))

[0098] Where F2 is the output feature map of the second convolution layer; the average pooling layer connected thereafter is:

[0099] F out ={AvgPool}(F2)

[0100] Among them, F out is the output feature map after average pooling, the output F out Input to the convolution block attention module for processing to obtain the enhanced feature map; the calculation expression of the convolution block attention module is:

[0101] F′=W SAM ⊙(W CAM ⊙F)

[0102] W CAM=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))

[0103] W SAM =σ(Conv 7×7 [Avg c (W CAM ⊙F); Max c (W CAM ⊙F)])

[0104] Among them, F and F′ represent the input feature map and enhanced feature map respectively, W SAM With W CAM They represent the weights calculated by the CAM and SAM submodules respectively, σ represents the sigmoid activation function, the subscript c represents the operation along the channel axis, and ⊙ represents the Hadamard product.

[0105] It should be noted that, in the present invention, the balance between computational efficiency and detection performance is taken into consideration, and the CBAM module is integrated after the feature extraction module, and its output is used to enhance the subsequent time series analysis module. Experimental verification shows that the introduction of the CBAM module significantly improves the model's ability to identify abnormal behavior of the CAN bus, especially when facing complex and variable network attacks.

[0106] In step S3, the convolutional long short-term memory network is composed of multiple convolutional long short-term memory layers, each of which includes an input gate, a forget gate, and an output gate. The calculation expression of the convolutional long short-term memory layer is as follows:

[0107]

[0108] in W it , W ft , W ot They represent the input vector, unit state, unit hidden state, input gate, forget gate, and output gate at time t, respectively. i 、b f 、b c 、b o Indicates the corresponding unit bias, the symbol * is the convolution operator, ⊙ is the Hadamard product, σ is the sigmoid function as the activation function, and tanh is the hyperbolic tangent function as the activation function; in the last part of the convolutional long short-term memory network, two fully connected layers are set as the classification decision network, and a random drop mechanism is equipped, and the Sigmoid function outputs the probability value as the judgment result of the entire model.

[0109] It should be noted that the input gate is given by the formula W it Description, which controls the current input (input data at time step t) and the hidden state at the previous time step How to affect the current cell state. Using a sigmoid activation function, this gate determines how much new information should be added to the cell state. Represents the modulation of the input gate by the previous state of the unit, which can help the model remember or forget previous information.

[0110] The forget gate is given by the formula W ft Description, which determines how much of the previous cell state is retained The sigmoid function is also used to achieve this, so that the network can "forget" irrelevant information and only retain information that is useful for future predictions.

[0111] The cell status is updated by the formula Description, unit status Updated according to the output of the forget gate and the input gate. Through the forget gate, part of the old state is retained, while the input gate controls the generation of new states, which are processed by the tanh activation function. and This allows the cell state to forget irrelevant information while incorporating new relevant information.

[0112] The forget gate is given by the formula W ot Description, which determines how much of the current cell state Should be output to the hidden state This is also adjusted by the sigmoid function to ensure that the output information is selective and depends on the current input. Previous hidden state and the current cell status

[0113] The hidden state is updated by the formula Description, hidden state is the output gate W ot and cell status The combined result of the network is processed by the tanh function to maintain the nonlinear characteristics of the network, and then combined with the result of the output gate W ot This step ensures that the network is able to pass useful information to the next time step or the next layer of the network.

[0114] The specific calculation process of the time series analysis is as follows: for each time step t and each convolutional long short-term memory layer m, the input feature map F m,t Receive data from the output of the previous layer or the state of the previous time step:

[0115]

[0116] X in,t is the network input at time step t, F m-1,t,out is the output of the previous layer at the same time step; each layer uses the ConvLSTM unit to update its cell state C m,t And calculate the output F m,t,out :

[0117] C m,t , F m,t,out =ConvLSTM(F m-1,t,out , S l,t-1 )

[0118] Among them, F m-1,t,out is the output of the previous layer at the same time step, S m,t-1 is the state of the same layer at the previous time step; at the last layer of each time step, the output F M,t,out Can be used as the final network output or passed to the next time step:

[0119] F out,t =F M,t,out

[0120] Among them, M is the last layer, F out,t is the final output at this time step.

[0121] After time series analysis, a classification decision network is required to output the final classification result to determine whether there is intrusion in the current sequence. The specific process is as follows: first, the input feature map is flattened to convert the multi-dimensional features into one-dimensional vectors; then, the features are transformed and reduced in dimension through two fully connected layers, and the classification results are output. The predicted classification results are compared with the actual labels assigned by the data preprocessing module in the early stage, and the binary cross entropy function is used to calculate the loss value of the current model. Based on the calculated loss value, the model parameters in steps S2 and S3 are automatically adjusted through the back propagation algorithm. The specific calculation process is as follows:

[0122] y1=Dropout(LeakyReLU(W1x+b1), p=0.2)

[0123] Classification=Softmax(W2y1+b2)

[0124] Where y1 represents the output of the first linear layer, the flattened input is linearly transformed, and the LeakyReLU activation function is directly applied, and then the Dropout operation is performed, W1 is the weight matrix, b1 is the bias vector, x is the flattened input vector, and p is the Dropout probability; the second linear layer receives the output y1 of the first linear layer, performs a linear transformation, and applies the Softmax activation function to output the final classification probability, where W2 and b2 are the weight matrix and bias vector of the second linear layer.

[0125] According to the above technical features, the data preprocessing process of the present invention not only converts CAN bus data into images, but more importantly, optimizes the data structure of the image in an intelligent way, so that it can be more effectively adapted to the needs of subsequent deep learning models. This includes intelligent screening and recoding of CAN data frames to maximize the amount of information retained and availability in the image data, thereby improving the sensitivity and accuracy of the intrusion detection system. The design of the adaptive feature extraction network uses a customized convolutional neural network (CNN) structure, combined with a convolutional block attention mechanism (CBAM), which can automatically identify and emphasize key features. This structure shows excellent performance when processing complex vehicle network data. In this way, the model not only extracts the spatial features of the image, but also automatically adjusts the learning focus according to the importance of different channels and spatial positions, thereby effectively improving the accuracy and efficiency of intrusion behavior recognition. The time series analysis method used uses a convolutional long short-term memory network (ConvLSTM). This design uniquely combines the advantages of convolutional neural networks and long short-term memory networks, and can simultaneously process and analyze the spatial features of image data and its temporal dynamics. This combination effectively solves the challenge of difficult simultaneous processing of high-dimensional space and time dependency problems in traditional CAN bus data intrusion detection.

[0126] Overall, the whole-link optimization from data acquisition to processing and analysis is taken into consideration. Through the close integration and interaction of various modules, all-round protection of the vehicle CAN bus can be achieved, especially under the dual requirements of real-time and accuracy, showing excellent performance.

[0127] Example 2

[0128] like Figure 5 As shown, the present invention also provides an in-vehicle CAN bus intrusion detection system based on artificial intelligence, which is used for the in-vehicle CAN bus intrusion detection method based on artificial intelligence, and includes: a data preprocessing module, a feature extraction module and a time series analysis module.

[0129] The data preprocessing module converts the raw data collected by the in-vehicle CAN bus into a CAN image sequence and inputs it into the feature extraction module. The feature extraction module uses a convolutional neural network to extract the key features of the CAN image sequence. The time series analysis module uses a convolutional long short-term memory network to perform time series analysis on the features output by the feature extraction module, and outputs an intrusion probability value as the judgment result of intrusion detection.

[0130] In this embodiment, Figure 6 The figure shows the overall structure of the model described in the present invention, and the complete system architecture is implemented in the Pytorch deep learning framework. First, the data preprocessing module is built, which is responsible for the processing and conversion of the original data; the second is the feature extraction module, which is used to capture the key features in the data; and finally the time series analysis module, which realizes the in-depth analysis of the sequence data.

[0131] The implementation process of the system of the present invention is divided into two main stages: training and deployment, specifically:

[0132] In the training phase, the present invention uses a vehicle intrusion detection dataset collected from a real vehicle for model training. Specifically, the "CAR HACKING: TTACK&DEFENSE CHALLENGE 2020 DATASET" dataset is used, which contains a large amount of real vehicle communication data and intrusion behavior data.

[0133] The data preprocessing module adopts the following processing flow: First, read the content of the data set in units of lines, referring to Figure 2 and Figure 3 In the processing method shown, the system removes redundant information and irrelevant fields in the data. Subsequently, the preprocessed CAN data frame is converted into a standardized CAN image format through a specific algorithm. These images will be further processed and spliced ​​into image sequences of a fixed length of L. For each sequence, the system will assign a corresponding label to identify whether the sequence contains intrusion information. This label information is crucial for the subsequent supervised learning process.

[0134] The process of data preprocessing is shown in the following formula:

[0135] 1.1.1 Data field extraction and conversion

[0136] ID Extraction: ID i =extract_id(f i )

[0137] Data field extraction: D i =extract_data_field(f i ), where f iRepresents several original CAN data frames

[0138] 1.1.2 Data is converted into binary vectors:

[0139] Binary conversion: B i = to_binary(ID i ,12)+to_binary(D i , 64)

[0140] The ID to be extracted i and D i Convert to binary format. i Convert to 12-bit binary number, D i Convert to a binary number with 8 bits per byte, for a total of 64 bits. Any bits less than 64 bits are filled with 0s.

[0141] 1.1.3. Convert binary vector to image:

[0142] Image generation: C img =[B1, B2, ..., B n ]

[0143] The serialized binary vector B i Further converted into CAN image, accumulating n B i To form the final CAN image, where n is a hyperparameter that determines the final size of the image.

[0144] 1.1.4. Construct CAN image sequence:

[0145] Image sequence generation: Seq(k) = [C img,k , C img,k+S , ..., C img,k+(L-1)·S ]

[0146] Where Seq(k) represents an image sequence of length L starting from the kth image, and each C img,j is an independent CAN image, and k is the starting index of the sequence.

[0147] After preprocessing, the image sequence is transferred to the feature extraction module for processing. In this module, the system treats each CAN image as an independent time step and processes it in sequence through two key components: first, the CNN network is used to extract the spatial features of the image; followed by the CBAM module, which is used to enhance important features and suppress irrelevant features. The series processing of these two components will generate a feature map containing rich feature information. The feature extraction process can be expressed as: Feature_Seq(k)=CBAM(CNN(Seq(k)))

[0148] 1.2.1 CNN Structure

[0149] The model consists of a series of multiple convolutional and activation layers, which include layer-by-layer feature extraction and nonlinear introduction.

[0150] The first convolution layer: F1 = LeakyReLU (Conv (F in ), where F1 is the output feature map of the first convolution layer, F in is the input single CAN image.

[0151] The second convolution layer: F2 = LeakyReLU (Conv (F1)), where F2 is the output feature map of the second convolution layer.

[0152] Average pooling layer: F out ={AvgPool}(F2),F out is the output feature map after average pooling. This step aims to reduce the feature dimension and enhance the generalization ability of the model. out It will be processed by the following CBAM to obtain the enhanced feature map.

[0153] 1.2.2 Structure of CBAM

[0154] The CBAM calculation expression in the feature extraction module can be written as:

[0155] F′=W SAM ⊙(W CAM ⊙F)#(2.1)

[0156] W CAM =σ(MLP(AvgPool(F))+MLP(MaxPool(F)))#(2.2)

[0157] W SAM =σ(Conv 7×7 [Avg c (W CAM ⊙F); Max c (W CAM ⊙F)])#(2.3)

[0158] Where F and F′ represent the input feature map and enhanced feature map respectively, W SAM With W CAM They represent the weights calculated by the CAM and SAM submodules respectively, σ represents the sigmoid activation function, the subscript c represents the operation along the channel axis, and ⊙ represents the Hadamard product.

[0159] As shown in formula (2.1), the input feature map passes through CAM and SAM successively, that is, the weight W calculated by the two sub-modules is obtained by multiplying them successively. CAMWith W SAM , and finally we get the enhanced feature map, W. CAM As shown in formula (2.2), the input feature map is processed by global average pooling and global maximum pooling, respectively, and then calculated by the multi-layer perceptron and added together, and finally the channel attention weight is calculated by the sigmoid activation function. SAM As shown in formula (2.3), the input feature map is weighted by the channel attention weight, and its average and maximum values ​​are calculated. It passes through a convolutional layer with a kernel size of 7×7, and finally the spatial attention weight is calculated by the sigmoid activation function to obtain the required feature map F′.

[0160] In the time series analysis stage, the system uses a multi-layer stacked ConvLSTM structure to process the feature map sequence. The feature map is input into the structure in the order of time steps, and each layer of ConvLSTM will further extract features and perform time series analysis on the input data. When the system completes the processing of l CAN images, it means that the processing of a complete sequence has ended. At this time, the hidden state of the top ConvLSTM layer contains the complete spatial and temporal feature information of the sequence. The system extracts the hidden state and uses it as the input feature map of the classification decision network.

[0161] The specific calculation steps are as follows:

[0162] For each time step t and each ConvLSTM layer m, the input feature map F m,t Receive data from the output of the previous layer or the state of the previous time step: Among them, X in,t is the network input at time step t, F m-1,t,out is the output of the previous layer at the same time step.

[0163] Each layer uses ConvLSTM units to update its cell state C m,t And calculate the output F m,t,out :C m,t , F m,t,out =ConvLSTM(F m-1,t,out , S l,t-1 ). Among them, F m-1,t,out is the output of the previous layer at the same time step, Figure 1 It can be seen that the hidden state of the ConvLSTM unit is inherited from bottom to top, which reflects the increase in network depth; S m,t-1 is the state of the same layer at the previous time step, which reflects the temporal dependency.

[0164] At the last layer of each time step, the output F M,t,out Can be used as the final network output or passed to the next time step: Fout,t =F M,t,out , where M is the last layer and F out,t is the final output at this time step.

[0165] The classification decision network adopts the following structure: first, the input feature map is flattened to convert the multi-dimensional features into one-dimensional vectors; then, the features are transformed and reduced in dimension through two fully connected layers, and finally the classification results are output to determine whether there is intrusion in the current sequence. The system compares the predicted classification results with the actual labels assigned by the data preprocessing module in the early stage, and uses the binary cross entropy function to calculate the loss value of the current model. Based on the calculated loss value, the system automatically adjusts the model parameters in the feature extraction module and the time series analysis module through the back propagation algorithm to continuously improve the detection accuracy.

[0166] The specific calculation steps are as follows:

[0167] First linear layer: linearly transform the flattened input, directly apply the LeakyReLU activation function, and then perform the Dropout operation: y1 = Dropout(LeakyReLU(W1x+b1), p = 0.2). Where W1 is the weight matrix, b1 is the bias vector, x is the flattened input vector, and p is the probability of Dropout.

[0168] The second linear layer: accepts the output y1 of the first layer, performs a linear transformation, and applies the Softmax activation function to output the final classification probability: Classification = Softmax (W2y1 + b2). Where W2 and b2 are the weight matrix and bias vector of the second linear layer.

[0169] When the model is trained and reaches the expected performance indicators, the system will save the trained model parameters. In actual deployment, you only need to load the saved model on the adapted hardware device. Figure 2 As shown in Figure 2, the data processing flow in the deployment phase is basically the same as that in the training phase. The operation process is as follows: Figure 2 As shown, first prepare the CAN raw data frame that meets the requirements, and input it into the model after the same preprocessing steps. The model can then give a prediction result on whether there is intrusion behavior in the current input data.

[0170] Example 3

[0171] Based on the above-mentioned embodiment 1 and embodiment 2, combined Figure 7-Figure 8 , this embodiment describes in detail the specific process of the model deployment stage of the present invention.

[0172] Figure 7 and Figure 8The confusion matrix image, ROC curve and AUC image of the model of the present invention are shown under the conditions of image sequence length l=8, number of ConvLSTM unit layers 6, and dimensions of each hidden layer [64, 64, 32, 16, 8, 1]. It can be seen that Figure 7 It is a confusion matrix image of the model of the present invention under the conditions of image sequence length l=8, number of ConvLSTM unit layers 6 layers, and dimensions of each hidden layer [64, 64, 32, 16, 8, 1]. The confusion matrix is ​​a table used to evaluate the performance of a classification model, which shows the number of true positive TP (True Positive), true negative TN (True Negative), false positive FP (False Positive) and false negative FN (False Negative) for each class. In this figure, normal samples are regarded as positive classes, then there are 49911 TPs, 36320 TNs, 994 FPs, and 1607 FNs.

[0173] Figure 8 It is the ROC curve and AUC image of the model of the present invention under the conditions of image sequence length 1=8, ConvLSTM unit layer number 6 layers, and hidden layer dimensions of each layer [64,64,32,16,8,1]. ROC curve refers to the receiver operating characteristic curve, which is a graphical representation of the trade-off between the true positive rate and the false positive rate when the classification threshold changes. AUC refers to the area under the ROC curve, which is a classifier performance indicator. The higher the AUC, the better the performance of the classifier. In this figure, the average AUC value is 0.9943, the AUC value for normal samples is 0.9951, and the AUC value for normal samples is 0.9945.

[0174] In the embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above method embodiments; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media capable of storing program codes.

[0175] Alternatively, if the above-mentioned embodiments of the present invention are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media capable of storing program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0176] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. The icons describing the structural positional relationships in the accompanying drawings are only used for exemplary purposes and are not to be construed as limiting the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and is not possible to list all the embodiments here. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the claims of the present invention.

Claims

1. An artificial intelligence-based in-vehicle CAN bus intrusion detection system, characterized in that: The following steps are involved: S1, converting the raw data collected by the CAN bus in the vehicle into a CAN image sequence; S2, extracting key features of the CAN image sequence; S3. Use a convolutional long short-term memory network to perform time series analysis on the features output by the feature extraction module, and output an intrusion probability value as a judgment result of intrusion detection.

2. The in-vehicle CAN bus intrusion detection method based on artificial intelligence according to claim 1 is characterized in that: The step S1 includes the following specific steps: first, receiving a data frame on the CAN bus, cleaning the data frame, extracting the ID and the corresponding data field in hexadecimal format to obtain a CAN frame list, and converting it line by line into a binary format to obtain a coded vector, and superimposing the obtained coded vector to reconstruct a two-dimensional CAN image; wherein, in the process of converting to a coded vector, if the length of the data field is less than 64 bits, 0 is added to the end of the field to ensure that the final length of the vector is 76 bits; the CAN image is a two-dimensional image in the shape of [n, 76], where n represents a hyperparameter that controls the image length.

3. The in-vehicle CAN bus intrusion detection method based on artificial intelligence according to claim 2 is characterized in that: In the process of creating the CAN image sequence, a sliding window is moved on the original frame list with a step length of 10 frames each time to construct a CAN image sequence of length l, where l is a hyperparameter that controls the amount of contextual information provided by the network. The larger the value, the longer the image sequence and the more information it contains. After creating the CAN image sequence, it is also necessary to set a classification label for it. When at least one frame in the frame list corresponding to the CAN image sequence is an attack frame, the label is set to "attack". If all frames are normal, it is set to "normal".

4. The in-vehicle CAN bus intrusion detection method based on artificial intelligence according to claim 3 is characterized in that: In step S1, the data preprocessing process is specifically as follows: Extract the ID and data fields: ID i =extract_id(f i ) D i =extract_data_field(f i ) where f i Represents several raw CAN data frames, and then converts the data into a binary vector: B i =to_binary(ID i ,12)+to_binary(D i ,64) Where ID i Convert to 12-bit binary number, D i Convert to a binary number with 8 bits per byte, a total of 64 bits, and fill any bits less than 64 bits with 0; convert the generated binary vector to an image: C img =[B1,B2,…,B n ] The serialized binary vector B i Further converted into CAN image, accumulating n B i To form the final CAN image, where n is a hyperparameter that determines the final size of the image; finally, construct the CAN image sequence: Seq(k)=[C img,k ,C img,k+S ,…,C img,k+(L-1)S ] Where Seq(k) represents an image sequence of length L starting from the kth image, and each C img,j is an independent CAN image, and k is the starting index of the sequence.

5. The in-vehicle CAN bus intrusion detection method based on artificial intelligence according to claim 1 is characterized in that: In the step S2, the key features are extracted by using a convolutional neural network to extract the key features for intrusion detection from the CAN image sequence; the convolutional neural network includes two convolutional layers, an activation layer and a pooling layer, and a convolutional block attention module is introduced into the convolutional neural network, and the convolutional block attention module enhances the feature map through channel attention focusing and spatial attention focusing; wherein, the channel attention is implemented by global average pooling and global maximum pooling, and the spatial attention aggregates the channel information of the feature map by calculating the average and maximum values ​​of the feature map along the direction of the channel axis, and the vectors are connected and input into a convolutional layer with a kernel size of 7×7 to obtain attention in the spatial dimension.

6. The in-vehicle CAN bus intrusion detection method based on artificial intelligence according to claim 5 is characterized in that: The feature extraction process in step S2 is: Feature_Seq(k)=CBAM(CNN(Seq)k))) The convolutional neural network includes multiple layers of convolution and activation layers, and the first layer of convolution is: F1=LeakyReLU(Conv(F in )) Where F1 is the output feature map of the first convolution layer, F in is a single CAN image input; the second layer of convolution is: F2 = LeakyReLU(Conv(F1)) Where F2 is the output feature map of the second convolution layer; the average pooling layer connected thereafter is: <h2 style=";text-align:left;direction:ltr">F<h2 style=";text-align:left;direction:ltr"> out <h2 style=";text-align:left;direction:ltr"> (F2) where F out is the output feature map after average pooling, the output F out Input to the convolution block attention module for processing to obtain the enhanced feature map; the calculation expression of the convolution block attention module is: F’=W SAM ⊙(W CAM ⊙F) W CAM =σ(MLP(AvgPool(F))+MLP(MaxPool(F))) IN SAM =σ(Conv 7×7 [Avg c (IN CAM ⊙F);Max c (IN CAM ⊙F)]) Among them, F and F' represent the input feature map and enhanced feature map respectively, W SAM With W CAM They represent the weights calculated by the CAM and SAM submodules respectively, σ represents the sigmoid activation function, the subscript c represents the operation along the channel axis, and ⊙ represents the Hadamard product.

7. The in-vehicle CAN bus intrusion detection method based on artificial intelligence according to claim 1 is characterized in that: In step S3, the convolutional long short-term memory network is composed of multiple convolutional long short-term memory layers, each of which includes an input gate, a forget gate, and an output gate. The calculation expression of the convolutional long short-term memory layer is as follows: in W it , W ft , W ot They represent the input vector, unit state, unit hidden state, input gate, forget gate, and output gate at time t, respectively. i 、b f 、b c 、b o Indicates the corresponding unit bias, the symbol * is the convolution operator, ⊙ is the Hadamard product, σ is the sigmoid function as the activation function, and tanh is the hyperbolic tangent function as the activation function; in the last part of the convolutional long short-term memory network, two fully connected layers are set as the classification decision network, and a random drop mechanism is equipped, and the Sigmoid function outputs the probability value as the judgment result of the entire model.

8. The in-vehicle CAN bus intrusion detection method based on artificial intelligence according to claim 7 is characterized in that: The specific calculation process of the time series analysis is as follows: for each time step t and each convolutional long short-term memory layer m, the input feature map F m,t Receive data from the output of the previous layer or the state of the previous time step: X in,t is the network input at time step t, F m-1,t,out is the output of the previous layer at the same time step; each layer uses the ConvLSTM unit to update its cell state C m,t And calculate the output F m,t,out : C m,t ,F m,t,out =ConvLSTM(F m-1,t,out ,S l,t-1 ) Among them, F m-1,t,out is the output of the previous layer at the same time step, S m,t-1 is the state of the same layer at the previous time step; at the last layer of each time step, the output F M,t,out Can be used as the final network output or passed to the next time step: F out,t =F M,t,out Among them, M is the last layer, F out,t is the final output at this time step.

9. The in-vehicle CAN bus intrusion detection method based on artificial intelligence according to claim 8 is characterized in that: After time series analysis, a classification decision network is required to output the final classification result to determine whether there is intrusion in the current sequence. The specific process is as follows: first, the input feature map is flattened to convert the multi-dimensional features into one-dimensional vectors; then, the features are transformed and reduced in dimension through two fully connected layers, and the classification results are output. The predicted classification results are compared with the actual labels assigned by the data preprocessing module in the early stage, and the binary cross entropy function is used to calculate the loss value of the current model. Based on the calculated loss value, the model parameters in steps S2 and S3 are automatically adjusted through the back propagation algorithm. The specific calculation process is as follows: y1=Dropout(LeakyReLU(W1x+b1),p=0.2) Classification=Softmax(W2y1+b2) Where y1 represents the output of the first linear layer, the flattened input is linearly transformed, and the LeakyReLU activation function is directly applied, and then the Dropout operation is performed, W1 is the weight matrix, b1 is the bias vector, x is the flattened input vector, and p is the probability of Dropout; The second linear layer receives the output y1 of the first linear layer, performs a linear transformation, and applies the Softmax activation function to output the final classification probability, where W2 and b2 are the weight matrix and bias vector of the second linear layer.

10. An in-vehicle CAN bus intrusion detection system based on artificial intelligence, characterized in that: It includes: data preprocessing module, feature extraction module and time series analysis module; The data preprocessing module converts the raw data collected by the in-vehicle CAN bus into a CAN image sequence and inputs it into the feature extraction module. The feature extraction module uses a convolutional neural network to extract the key features of the CAN image sequence. The time series analysis module uses a convolutional long short-term memory network to perform time series analysis on the features output by the feature extraction module, and outputs an intrusion probability value as the judgment result of intrusion detection.

Citation Information

Patent Citations

  • CAN bus intrusion detection method and device based on neural network, and medium

    CN117278296A

  • Internet of vehicles CAN bus intrusion detection method based on spatial-temporal characteristics

    CN117478410A

Cited By

  • Intelligent error detection method and system for measurement while drilling data

    CN120929927A

  • An intelligent error detection method and system for measurement-while-drilling data

    CN120929927B