An in-vehicle intrusion detection method and system based on CAN bus data frames

The intrusion detection model constructed through One-hot encoding and GAN algorithm solves the problem that existing systems cannot identify complex attack methods in real time, and realizes high-precision intrusion detection of automotive CAN bus networks to protect vehicle safety.

CN115314311BActive Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210969625.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-12
Publication Date
2025-06-27
Estimated Expiration
2042-08-12

AI Technical Summary

Technical Problem

Most existing vehicle intrusion detection systems rely on developer documents and cannot effectively identify and defend against complex attack methods, especially the changes in data frame content in the CAN bus network cannot be monitored in real time.

Method used

One-hot encoding is used to preprocess the data packets and messages transmitted by the ECU of the automotive CAN bus network, and an intrusion detection model of the generator network and discriminator network is constructed based on the GAN algorithm. Through the training model, the abnormal data frame is recognized, and real-time intrusion detection of the CAN bus network is realized.

Benefits of technology

It realizes intrusion detection that does not rely on developer documents, can identify abnormal data frames with high accuracy, protect vehicle safety, and has high detection accuracy, and is suitable for different vehicle platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115314311B_ABST
    Figure CN115314311B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle intrusion detection method and system based on CAN bus data frames, which preprocess the data messages of the automotive CAN bus network and the messages transmitted by each ECU using One-hot encoding; construct an intrusion detection model including a generator network and a discriminator network based on the GAN algorithm, and train the intrusion detection model; input the preprocessed One-hot encoded data messages of the automotive CAN bus network and the messages transmitted by each ECU into the trained intrusion detection model to complete intrusion detection. The present invention directly obtains the data frames of the CAN bus from the vehicle OBD-II port, and the acquisition of data is convenient and fast; it does not occupy the CAN bus bandwidth and computing resources, and is directly applied to the vehicle to monitor the CAN network data transmission in real time and protect the vehicle safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent connected vehicles, and particularly relates to a vehicle-mounted intrusion detection method and system based on CAN bus data frames. Background Art

[0002] With the rapid development of digital processes such as artificial intelligence and 5G, more and more advanced technologies have been applied to automobiles. While these technologies improve the user driving experience, they also expose a large number of interfaces inside the vehicle to the outside world, making the CAN bus more vulnerable to attacks by malicious actors. As the most common bus in the automotive field, the CAN bus itself has almost no data encryption and authentication means.

[0003] For this reason, researchers have designed different types of intrusion detection systems to protect vehicle safety, but most of them focus on a specific form of attack defense and are not universal for existing advanced attack means. For example, IDS based on parameter monitoring cannot identify changes in the content of data frames in the bus network by attackers. IDS based on information entropy relies on a large number of message changes in the network and cannot identify spoofing attacks. IDS based on fingerprint recognition relies on developer documents provided by manufacturers to construct the quantity information of in-vehicle electronic control units. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a vehicle-mounted intrusion detection method and system based on CAN bus data frames, which can achieve intrusion detection without relying on developer documents in view of the above deficiencies in the prior art.

[0005] The present invention adopts the following technical solutions:

[0006] A vehicle-mounted intrusion detection method based on CAN bus data frames includes the following steps:

[0007] S1. Preprocess the data messages of the vehicle CAN bus network and the messages transmitted by each ECU using one-hot encoding;

[0008] S2. Construct an intrusion detection model including a generator network and a discriminator network based on the GAN algorithm, and train the intrusion detection model;

[0009] S3. Input the one-hot encoding preprocessed in step S1 for the data messages of the vehicle CAN bus network and the messages transmitted by each ECU into the intrusion detection model trained in step S2 to complete intrusion detection.

[0010] Specifically, in step S1, before encoding, data segments of different lengths are padded to 8 bytes, and a 64×64 one-hot encoding map is generated for every 16 CAN data frames.

[0011] Specifically, in step S2, the generator network includes two encoders and one decoder. The encoder network has 5 layers. The convolution kernel size of each layer is 4×4, the number of channels is 3, the stride size is 2, the padding is 1, and the ReLU() activation function is used between each layer.

[0012] The decoder uses transposed convolution. The encoder network has 5 layers. The convolution kernel size of each layer is 4×4, the number of channels is 3, the stride size is 2, the padding is 1, and the Leaky ReLU() activation function is used between each layer.

[0013] Specifically, in step S2, the training process of the intrusion detection model is as follows:

[0014] The generator network reads the input image x and forwards the input image x to the encoder Encoder1; the encoder Encoder1 calculates the input image x through the convolutional layer, then performs Leaky ReLU() activation function and normalization processing, compresses the input image x into the latent vector z, and the decoder learns the distribution of the input image x in the latent vector z to generate the reconstructed image Reconstructed image Extracts typical features through the encoder Encoder2 to form the feature vector The discriminator network discriminates between the input image x and the reconstructed image to determine the real image and the reconstructed image. The generator network re-learns the features to reconstruct the real image according to the feedback of the discriminator network. The discriminator network learns the difference between the real image and the reconstructed image, and the training is completed when the discriminator network recognizes the reconstructed image.

[0015] Furthermore, the reconstruction loss function L of the input image x and the generated image is: con where real

[0016]

[0017] is the real image, fake images is the reconstructed image, l1_loss(·) is the L1 distance loss function, n is the number of images, x images is the i-th input image, i is the i-th reconstructed image. is the i-th reconstructed image.

[0018] Furthermore, the encoder loss L between the latent vector z and the feature vector is: enc where

[0019]

[0020] Among them, l2_loss(·) is the L2 distance loss function, bottleneck1 is the latent vector, bottleneck2 is the feature vector, and n is the number of images. is the i-th feature vector.

[0021] Furthermore, the loss function L of the discriminator network adv is:

[0022]

[0023] Among them, bec_loss(input images ) is the cross-entropy loss function.

[0024] Specifically, in step S3, calculate the encoder loss L of all sample data enc , select the threshold that can maximize the classification effect as the discrimination criterion, and obtain the scoring value score in the interval [0,1] by comparing the distance feature difference between the feature vectors output by the two encodings and the latent vector z, and calculate the threshold that maximizes the bipartite state. For the sample x input to the system, if the score is determined as abnormal data, if the score is determined as normal data.

[0025] Furthermore, the scoring value score is:

[0026]

[0027] Second, the embodiment of the present invention provides a vehicle intrusion detection system based on CAN bus data frames, including:

[0028] A data module for preprocessing the data messages of the automotive CAN bus network and the messages transmitted by each ECU using One-hot encoding;

[0029] A training module for constructing an intrusion detection model including a generator network and a discriminator network based on the GAN algorithm and training the intrusion detection model;

[0030] A detection module for inputting the One-hot encoding preprocessed by the data module for the data messages of the automotive CAN bus network and the messages transmitted by each ECU into the intrusion detection model trained by the training module to complete intrusion detection.

[0031] Compared with the prior art, the present invention has at least the following beneficial effects:

[0032] A vehicle intrusion detection method based on CAN bus data frames. Since the input of the GAN algorithm is in the form of images, to facilitate the image conversion process, One-hot encoding is used to effectively preprocess the content status of data packets and messages. In actual situations, the instances of vehicle attacks are very few, so the normal vehicle CAN data is much more than the abnormal vehicle CAN data. And deep learning models often highly depend on the quality of input data. The difference in the quantity and distribution of the two types of data often makes the model unable to well learn the distribution characteristics of abnormal vehicle CAN data, resulting in low detection accuracy. Therefore, an intrusion detection model containing a generator network and a discriminator network is constructed based on the GAN algorithm. This model can only learn the characteristics of normal vehicle CAN data during training, and can accurately identify abnormal vehicle CAN data during testing, and then detect intrusions. When conducting model training, a large number of data packets of the automotive CAN bus network and messages transmitted by each ECU are required. Therefore, they are input into the intrusion detection model constructed based on the GAN algorithm containing a generator network and a discriminator network for training. After training is completed, testing is performed with abnormal CAN bus data to complete intrusion detection.

[0033] Furthermore, since the data volume transmitted by the CAN bus data frame is 0 to 8 bytes, in order to uniformly process each data frame, data segments of different lengths need to be padded to 8 bytes. The driving and being attacked of the vehicle both occur within a continuous period of time. Therefore, the intrusion detection model needs to learn the features of data within a period of time. Generating a 64×64 One-hot encoded image from every 16 CAN data frames can well preserve the changes in data within the continuous time.

[0034] Furthermore, the size of the input image is 64×64. The encoder performs downsampling on it, and the output will be a pixel point of 1×1 containing the features of the original image. After passing through the first layer of the network, the size becomes 32×32, after the second layer, it becomes 15×15, after the third layer, it becomes 7×7, after the fourth layer, it becomes 3×3, and after the fifth layer, it becomes 1×1. Using the ReLU() activation function for each layer can accelerate the training speed of the network, prevent the gradient disappearance of the network, increase the non-linearity of the network, and prevent overfitting of the results. The convolutional kernel, channel size, stride, and padding settings of the decoder are the same as those of the encoder. Therefore, the operations are the same as those of the encoder, except that transposed convolution is used for upsampling. The LeakyReLU() activation function almost has all the advantages of the ReLU() activation function, except that it gives all negative values a non-zero slope, which is convenient for the backpropagation of the network.

[0035] Furthermore, the intrusion detection model first uses the encoder Encoder1 in the generator network to learn the feature distribution of the input image, which is represented as a latent vector z. z is the latent feature space that best represents the image x and has the smallest dimension. Subsequently, the decoder Decoder re-expands the latent vector z to generate a reconstructed image. The encoder Encoder2 then learns the features in the reconstructed image. and represents them as feature vectors. The latent vector z and the feature vectors are used during the training process of the generator network and the discriminator network, so as to more accurately distinguish the authenticity of the image.

[0036] Furthermore, in order to optimize the discriminator network in the direction of being able to more comprehensively learn the context information of the input data, the L1 distance between the input image x and the image reconstructed by the generator network is used as the loss function. The L1 distance function can generate fewer blurred results.

[0037] Furthermore, in order to enable the generator network to better learn how to encode the features of the reconstructed image, this loss function is applied to minimize the distance between the latent vector z and the feature vectors.

[0038] Furthermore, in order to reduce the instability of the GAN model training, a feature matching loss is used, and the cross-entropy between the input image and the reconstructed image generated by the generator network is used to optimize the GAN model.

[0039] Furthermore, the intrusion detection model determines whether the vehicle is attacked by detecting whether there is abnormal data. Therefore, a criterion for determining whether the current data is abnormal is required. The operating principle of this criterion is to assign a score score to the sample and find the corresponding threshold. Compare the sample score score and the threshold. Based on the relationship between them, it can be determined whether the current sample is normal data or abnormal data.

[0040] Furthermore, the purpose of setting score is to facilitate the subsequent judgment of abnormal data. The difference between the two is reflected by calculating the L2 distance between the sample latent vector z and the feature vectors.

[0041] It can be understood that the beneficial effects of the second aspect above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.

[0042] In summary, the data of the present invention is convenient to obtain, can monitor the CAN network data transmission in real time, protect the vehicle safety, has high detection accuracy, and can directly detect intrusion.​​

[0043] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0044] Figure 1 It is a one-hot encoding diagram of a CAN bus data frame;

[0045] Figure 2 It is the internal network structure diagram of the present invention;

[0046] Figure 3 It is the ROC curve diagram of the present invention in the Luxgen model;

[0047] Figure 4 It is the ROC curve diagram of the present invention in the Buick model;

[0048] Figure 5 It is the detection accuracy diagram of the present invention under various attack frequencies;

[0049] Figure 6 It is the detection accuracy diagram of the present invention for different ECUs in Luxgen;

[0050] Figure 7 It is the detection accuracy diagram of the present invention for different ECUs in Buick. Detailed Embodiments

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0052] In the description of the present invention, it should be understood that the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0053] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0054] It should be further understood that the term "and / or" as used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates an "or" relationship between the associated objects before and after.

[0055] It should be understood that although terms such as first, second, third, etc. may be used to describe preset ranges in the embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range can also be referred to as the second preset range, and similarly, the second preset range can also be referred to as the first preset range.

[0056] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0057] Various structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures and their relative sizes and positional relationships are only exemplary, and may actually deviate due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0058] The present invention provides a vehicle-mounted intrusion detection method based on CAN bus data frames. Utilizing the characteristics of the data frame sequence and content in the automotive CAN bus, a vehicle-mounted intrusion detection system based on CAN bus data frames is designed, which is divided into three parts: acquisition and preprocessing of data frames, construction of an intrusion detection model, and intrusion detection.

[0059] A vehicle-mounted intrusion detection method based on CAN bus data frames according to the present invention includes the following steps:

[0060] S1. Acquisition and preprocessing of data frames

[0061] Data is transmitted in the form of data frames in the CAN bus network. Connect the CAN_H and CAN_L of the CANalyst-II device to the corresponding wires of the automotive OBD-II respectively. Utilize the characteristic of data frames being broadcast and transmitted on the CAN bus network to obtain the data packets of the automotive CAN bus network and the messages transmitted by each ECU.

[0062] Table 1 Data Frame Format in CAN Bus

[0063]

[0064] Table 1 shows the data frame format in the CAN bus. The length of the data field is 0 - 8 bytes.

[0065] Preprocess the data using One-hot encoding. Before encoding, pad the data segments of different lengths to 8 bytes. Since each data sample has 8 bytes and each byte can store 2 hexadecimal digits, one 64×64 One-hot encoding graph can be generated from every 16 CAN data frames. The CAN data frame encoding graph is as Figure 1 shown.

[0066] S2. Construct an intrusion detection model

[0067] To detect unknown attacks, the present invention uses normal data as training samples to train a model to judge the data frames in the CAN bus and distinguish between normal data frames and intrusion data frames. The present invention constructs an intrusion detection model based on the GAN algorithm. The internal network structure of this intrusion detection model includes two parts: a generator network and a discriminator network.

[0068] Please refer to Figure 2 , the generator network consists of two encoders (Encoder1 and Encoder2) and a decoder (Decoder).

[0069] The network of the encoder has 5 layers. The size of the convolution kernel of each layer is 4×4, the number of channels is 3, the stride size is 2, the padding is 1, and the ReLU() activation function is used between each layer.

[0070] The decoder (Decoder) uses transposed convolution and has the same number of layers as the encoder network. The size of the convolution kernel of each layer is 4×4, the number of channels is 3, the stride size is 2, the padding is 1, and the Leaky ReLU() activation function is used between each layer.

[0071] The discriminator network D is only used for classification and discrimination; it has a structure similar to Encoder1 and serves to meet the data input consistency.

[0072] Finally, when training the model, the Decoder first reads the image x from the input module and forwards the image to the Encoder1;

[0073] Next, after the calculation is completed through the convolutional layer, the input image x is compressed into the latent vector z using the Leaky ReLU() activation function and normalization process. At the same time, the latent vector z is the minimum dimension that can ensure the best performance of the input image x. When passing through the Decoder, z is amplified so that the input image x is reconstructed into the image The discriminator network D makes a judgment on the input image x and the generated image During continuous training, the ability to discriminate anomalies is maximized. During the training process, the generator continuously optimizes its "forgery" ability and reduces the feature gap between the generated image and the input image x.

[0074] The reconstruction loss function of the input image x and the generated image is:

[0075]

[0076] The encoder loss L between the latent vector z and the feature vector enc is:

[0077]

[0078] The loss function of the discriminator network D is:

[0079]

[0080] S3. Intrusion detection

[0081] In the intrusion detection stage, the encoding loss L enc ;

[0082] After the model training converges, calculate the L enc value of all sample data, and select the threshold that can maximize the classification effect as the discrimination criterion.

[0083] For the input of sample data, by comparing the distance feature differences between the bottleneck features of the two encoding outputs and z, a score value score in the range of [0,1] is obtained.

[0084] The scoring formula is defined as follows:

[0085]

[0086] Analyze the distribution of normal and abnormal samples of the output CAN protocol, and the threshold for maximizing the binary state can be calculated. For the sample x input to the system, if the score then it is determined as abnormal data. Conversely, if the score then it is normal data.

[0087] In another embodiment of the present invention, a vehicle intrusion detection system based on CAN bus data frames is provided. This system can be used to implement the above-mentioned vehicle intrusion detection method based on CAN bus data frames. Specifically, the vehicle intrusion detection system based on CAN bus data frames includes a module, a module, a module, a module, and a module.

[0088] Among them, the data module preprocesses the data messages of the automotive CAN bus network and the messages transmitted by each ECU using One-hot encoding;

[0089] The training module constructs an intrusion detection model including a generator network and a discriminator network based on the GAN algorithm, and trains the intrusion detection model;

[0090] The detection module inputs the One-hot encoded data messages of the automotive CAN bus network and the messages transmitted by each ECU preprocessed by the data module into the intrusion detection model trained by the training module to complete intrusion detection.

[0091] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0092] Experimental environment:

[0093] (1) Test vehicle and data frame acquisition:

[0094] The experiment was conducted on two cars, a Buick and a Luxgen. CANalyst-II was used to collect CAN bus data frames and analyze their internal instructions. Then, the collected data packets were modified to construct intrusion data packets, demonstrating the traffic on the CAN bus during a vehicle intrusion. Among them, the data frames were divided into normal data (data frames during normal vehicle operation) and intrusion data (attack data frames sent by attackers).

[0095] (2) Hardware and software environment of the experiment:

[0096] This invention is designed based on an improved GAN algorithm and developed using the Python language, the Pytorch framework, and Jupyter notebooks. The computer hardware used in the experiment was a CPU of Intel Xeon Gold 5218R, 128GB of memory, and a graphics card of NVIDIA GTX 3080.

[0097] (3) Training samples and parameter settings:

[0098] A sample set of n×16 was obtained by extracting features from the training data. The core parameter learn rate of the GAN network was set to 0.001, the Batch size was set to 128, and the number of sample training times was set to 200.

[0099] Experimental results:

[0100] (1) ROC curves in different vehicle models

[0101] Before the start of the detection experiment, we first judged the training effect of the model. Figure 3 and Figure 4 ROC curve graphs of the CAN protocol intrusion detection model on two vehicle platforms, the Luxgen U5 and the Buick Regal, were plotted respectively. Each graph contains four attack ROC curves for such vehicles. Among them, the abscissa is the false positive rate, indicating that the closer the value is to zero, the better the discrimination effect; the ordinate is the true positive rate, which directly reflects the accuracy of the detection rate, and the larger its value, the higher the accuracy.

[0102] It can be seen from the graph that the ROC curves of the Luxgen and Buick cars both reach the best performance point (the upper left corner of the curve) when the FPR is less than 0.1, and the TPR values at this point of each curve are all greater than 0.9, which indicates that the system model of the present invention has good detection performance for the four different attacks on the two cars and there is no overfitting training situation.

[0103] In Luxgen cars, the AUC value of the ROC curve for spoofing attacks is 0.9941, the AUC value for Bus-off is 0.9787, the AUC value for masquerade attacks is 0.9768, and the AUC value for SOME attacks is 0.9766. In Buick cars, the AUC values of the ROC curves for the four types of attacks are 0.9982, 0.9778, 0.9665, and 0.9662 respectively. This shows that the present invention has good training results on different cars and can effectively detect various attacks.

[0104] (2) Detection accuracy in different vehicles

[0105] Table 2 and Table 3 are respectively the experimental detection results of the intrusion detection system of the present invention for two cars, Luxgen U5 and Buick Regal.

[0106] Table 2 Detection accuracy of Luxgen

[0107]

[0108]

[0109] Table 3 Detection accuracy of Buick

[0110]

[0111] Among them, the detection accuracy Accuracy can reflect the probability that the overall test set samples are correctly recognized. The detection accuracy in Luxgen is above 96%, and the detection accuracy in Buick is above 93.13%. The Precision results show that it is above 97% in Luxgen and above 95.15% in Buick. This shows that the intrusion detection system of the present invention has excellent effects in Luxgen and Buick, can correctly classify normal data and abnormal data, and has a high classification accuracy.

[0112] In addition, among the four types of attacks, the detection accuracies of the two cars for spoofing attacks are both close to 1, and the overall detection performance of Luxgen is better than that of Buick. Compared with the detection accuracy of spoofing attacks approaching 100%, the masquerade attacks and the same-source attacks (SOME) are slightly lower, but still have a detection effect above 90%. The F1 value illustrates the evaluation results of Precision and Recall from the side. FNR and FPR indicate that the misjudgment rate of the model for incorrect data is relatively low, and the closer the value is to 0, the better the detection performance of the model.

[0113] (3) Detection time consumption in different vehicles

[0114] Table 4 Detection time consumption in different vehicles

[0115] Luxgen U5 Buick Regal Total number of samples 200000 200000 Sample size 64×64 64×64 Total time consumed (s) 24.005 27.064 Average time consumed (ms) 0.12 0.13 Overall variance <![CDATA[1.08×10 -4 > <![CDATA[2.34×10 -5 > 95% confidence interval [0.03,0.24] [0.05,0.19]

[0116] Table 4 shows the measurement of the detection sample time for Luxgen and Buick. The results show that the average time required for this model to classify and determine the data frames in the CAN protocols of Luxgen and Buick is 0.12 milliseconds and 0.13 milliseconds respectively, meeting the millisecond-level transmission requirements of data frames in the CAN bus network. The overall variances of the Luxgen and Buick test samples are both below 0.000108, indicating that the degree of dispersion of the time consumed during detection is relatively small. For the 95% confidence interval, the distribution of the time consumed by Luxgen is between 0.03 and 0.24, while Buick is relatively more concentrated, between 0.05 and 0.19. The results of both the method and the confidence interval show that the system detection has a certain stability, and the detection time consumption for the discrimination of multiple samples is close.

[0117] (4) Detection performance at various attack frequencies

[0118] To discuss whether different attack frequencies will affect the accuracy of the detection system, different-frequency attack samples were generated for multiple CAN IDs in the Luxgen samples. The frequency of the attack samples is 1 to 10 times the frequency of the attacked ID. The attacked CAN IDs are 316 (T = 10ms), 38C (T = 10ms), and 3AC (T = 10ms). In Buick, the selected attacked IDs are: 1C8 (T = 12.5ms), 1E9 (T = 25ms), and 34A (T = 50ms). The attack frequency is also 1 to 10 times the attacked ID.

[0119] As Figure 5 shown, the abscissa represents the production frequency of spoofing attack samples, and the ordinate represents the detection accuracy of the test. Looking at the whole graph, in the frequency range [1, 5], the detection accuracy of the model will slightly decrease. Except for the experimental result of CAN ID 1E9 in Buick at an attack frequency of 7, the detection accuracy of the remaining test experiments is above 95%. We determined that the reason for this situation is that the data frame corresponding to ID 1E9 accounts for a relatively low proportion in the samples, resulting in a deviation. This indicates that the model has better detection performance for spoofing attacks at different frequencies among different-frequency IDs in the CAN protocol, and the detection accuracy of this model will not be affected by the data transmission frequency of the ECU.

[0120] (5) Detection performance for different attack IDs

[0121] To verify whether the intrusion data frames of different IDs in the CAN bus will cause deviations to the detection system, in this experiment, 5 ECUs representing different functions were selected from each of the two vehicles, Luxgen U5 and Buick Regal. Based on four types of means such as spoofing attacks, an abnormal data set of 100,000 was constructed respectively to conduct performance tests on the intrusion detection of each ECU of the vehicle. Figure 6Shows the comparison of the detection accuracy of the model for 5 different CAN IDs in Luxgen cars. The abscissa is the 5 different CAN IDs in Luxgen, and the ordinate is the detection accuracy. It can be seen from the figure that when the attack ID is 1C8, the detection rates for the four attacks such as spoofing attacks are all higher than 95.47%; when the ID is 0F1, the detection rates are all higher than 94.92%; when the ID is 238, the detection rates are all higher than 94.06%; when the ID is 0D1, the detection rates are all higher than 93.86%; when the ID is 1E7, the detection rates are all higher than 95.77%; Therefore, the performance of our intrusion detection system will not degrade with the change of the attack ID. Figure 7 Is the result of the detection accuracy of 5 different CAN IDs in Buick cars. The chart shows that the model can detect spoofing attacks, Bus-off attacks, masquerade attacks, and SOME attacks against different CAN IDs in Buick cars.

[0122] Therefore, the present invention is not affected by the attack ID. Regardless of the value of the ID, the present invention can detect it.

[0123] (5) Tests of drivers with different behaviors

[0124] Drivers with different behaviors often have differences in habits when driving a car. To avoid differences in the recognition results due to the driver's behavior, the relationship between multiple drivers with different characteristics and the detection accuracy was also studied in the experiment. In this experiment, 4 drivers were selected. Among them, Driver 1 and Driver 2 are male, with ages of 49 and 25 years respectively; Driver 3 and Driver 4 are female, with ages of 22 and 26 years respectively. In the experiment, the average recognition accuracies of the 4 drivers are: 95.05%, 96.78%, 95.62%, 96.94%, all above 95%. Among the four types of attacks, the spoofing attack has the best detection effect, and its average detection rate among the 4 drivers is 98.71%. Both the masquerade attack and the homologous attack are attacks generated by modifying the content of the data frame. The results of the detection accuracies of the two are quite close in the same driver, and the average detection rate is above 95%.

[0125] Table 5 Detection accuracy for different drivers

[0126]

[0127]

[0128] Combined with the results in Table 5, the detection model of the present invention will not reduce the detection accuracy due to the change of the driver.

[0129] In summary, the present invention, a vehicle-mounted intrusion detection method and system based on CAN bus data frames, has the following characteristics:

[0130] 1. The data frame of the CAN bus is directly obtained from the automotive OBD-II port, facilitating quick and easy data acquisition.

[0131] 2. It does not occupy the CAN bus bandwidth and computing resources, can be directly applied to vehicles, and monitors the CAN network data transmission in real time to protect vehicle safety.

[0132] 3. Tests were conducted on two different real vehicles. The experimental results show that the present invention is applicable to different vehicles and has good robustness.

[0133] 4. It can accurately distinguish various advanced attack means such as spoofing attacks, Bus-off attacks, masquerade attacks, and same-source attacks (SOME), with a detection accuracy of over 95%.

[0134] 5. Compared with other vehicle intrusion detection methods, the detection accuracy of the present invention is not affected by the attack frequency, the types of attack CAN IDs, and different drivers, and is only related to the sequence and content of the data frames in the automotive CAN bus. Once an attacker injects or stops sending CAN ID messages into the CAN bus, the present invention can directly detect the intrusion.

[0135] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0136] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0137] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.

[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.

[0139] The above is only to illustrate the technical idea of the present invention and should not be construed as limiting the scope of the present invention. Any modifications made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the scope of the claims of the present invention.

Claims

1. An in-vehicle intrusion detection method based on CAN bus data frames, characterized in that, The following steps are involved: S1. Use one-hot coding to pre-process the data messages of the automobile CAN bus network and the messages transmitted by each ECU. Before coding, pad the data segments of different lengths to 8 bytes, and generate a 64×64 one-hot coding map for every 16 CAN data frames; S2. Based on the GAN algorithm, an intrusion detection model including a generator network and a discriminator network is constructed, and the intrusion detection model is trained. The training process of the intrusion detection model is as follows: The generator network reads the input image and forwards the input image to Encoder1; Encoder1 computes the input image x through convolutional layers, then performs Leaky ReLU() activation function and normalization processing, and compresses the input image into a latent vector The decoder learns the distribution of the input image in the latent vector to generate a reconstructed image The reconstructed image extracts typical features through Encoder2 to form a feature vector The discriminator network discriminates between the input image and the reconstructed image to determine the real image and the reconstructed image. The generator network re-learns the features to reconstruct the real image according to the feedback of the discriminator network. The discriminator network learns the difference between the real image and the reconstructed image. The training is completed when the discriminator network recognizes the reconstructed image; S3, input the data message of the automobile CAN bus network and the message transmitted by each ECU into the intrusion detection model trained in step S2 by the one-hot encoding preprocessed in step S1, and complete the intrusion detection.

2. The vehicle intrusion detection method based on CAN bus data frames according to claim 1, characterized in that In step S2, the generator network includes two encoders and one decoder. The encoder network has 5 layers. The convolution kernel size of each layer is 4×4, the number of channels is 3, the step size is 2, the padding is 1, and the ReLU() activation function is used between each layer. The decoder uses transposed convolution, and the encoder network has 5 layers. The convolution kernel size of each layer is 4×4, the number of channels is 3, the step size is 2, the padding is 1, and the Leaky ReLU() activation function is used between each layer.

3. The vehicle intrusion detection method based on CAN bus data frames according to claim 1, characterized in that, Input image and the reconstructed loss function of the generated image is as follows For: Among them, is the real image, is the reconstructed image, is the L1 distance loss function, is the number of images, is the th input image, is the th reconstructed image.

4. The vehicle intrusion detection method based on CAN bus data frames according to claim 1, characterized in that Latent vector and the feature vector The encoder loss between is as follows: Among them, is the L2 distance loss function, is the latent vector, is the feature vector, is the number of images, is the th feature vector.

5. The vehicle intrusion detection method based on CAN bus data frames according to claim 1, characterized in that Loss function of the discriminator network is as follows: Among them, is the cross-entropy loss function.

6. The vehicle intrusion detection method based on CAN bus data frames according to claim 1, wherein In step S3, calculate the encoder loss of all sample data , select the threshold that can maximize the classification effect as the discrimination criterion, and obtain the scoring value in the range of [0, 1] by comparing the feature vector of the two encoding outputs and the latent vector to calculate the threshold that maximizes the bipartite state . For the sample input to the system , if the score , it is determined as abnormal data, and if the score , it is determined as normal data .

7. The vehicle intrusion detection method based on CAN bus data frames according to claim 6, characterized in that Scoring value is: 。 8. An in-vehicle intrusion detection system based on CAN bus data frames, characterized in that, include: The data module is used to pre-process the data messages of the vehicle CAN bus network and the messages transmitted by each ECU using one-hot encoding. Before encoding, the data segments of different lengths are padded to 8 bytes, and a 64×64 one-hot encoding map is generated for every 16 CAN data frames; The training module is used to build an intrusion detection model including a generator network and a discriminator network based on the GAN algorithm, and train the intrusion detection model. The training process of the intrusion detection model is as follows: The generator network reads the input image , and forwards the input image to Encoder1; Encoder1 computes the input image x through convolutional layers, then performs Leaky ReLU() activation function and normalization processing, and compresses the input image into a latent vector . The decoder learns the distribution of the input image in the latent vector to generate a reconstructed image . The reconstructed image extracts typical features through Encoder2 to form a feature vector . The discriminator network distinguishes between the input image and the reconstructed image to determine the real image and the reconstructed image. The generator network re-learns features to reconstruct the real image according to the feedback of the discriminator network. The discriminator network learns the difference between the real image and the reconstructed image. When the discriminator network recognizes the reconstructed image, the training is completed; The detection module is used to input the data messages of the automobile CAN bus network and the messages transmitted by each ECU preprocessed by the data module into the intrusion detection model trained by the training module to complete the intrusion detection.