Hybrid multi-stage in-vehicle network intrusion detection method
By adopting a hybrid multi-stage intrusion detection method in the on-board network, using convolutional neural network and attention mechanism for known attack detection, and using automatic encoder and generative adversarial network for unknown attack detection, the problem of difficulty in detecting both known and unknown attacks in the on-board network is solved, and efficient and fast attack detection and reduced missed rate is achieved.
Patent Information
- Application Number
- CN202510174838.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
AI Technical Summary
Due to the lack of identity authentication and data encryption, the on-board networks in modern smart cars are vulnerable to network threats such as forged message injection and denial of service (DoS) attacks. Traditional attack prevention methods are difficult to take into account the detection needs of known attacks and unknown attacks.
采用一种混合多阶段车载网络入侵检测方法,利用基于卷积神经网络和注意力机制的轻量化监督学习模型在第一阶段对车载网络消息进行已知类型攻击检测,并通过基于自动编码器和生成对抗网络的无监督模型在第二阶段进行未知类型攻击检测。
It realizes rapid detection of known and unknown attacks, reduces the missed rate of potential attacks, adapts to the computing resource limitations of on-board equipment, and shortens detection and response time.
Smart Images

Figure CN120034374A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle network security technology, and in particular to a hybrid multi-stage vehicle network intrusion detection method. Background Art
[0002] With the continuous advancement of technology, the electronic control unit (ECU), as the "brain" of the vehicle system, plays an important role in modern smart cars, responsible for monitoring and controlling various functions from engine management to safety systems. With the increase in automobile functions, the number of ECUs has also grown significantly, and the key to achieving efficient communication between ECUs lies in the in-vehicle network (IVN). At present, the controller area network (CAN) is widely adopted due to its efficiency and simplicity, becoming the most common communication protocol in IVN.
[0003] Although the CAN protocol provides an efficient data exchange mechanism in IVN, its design does not take security into consideration and lacks basic security features such as authentication and data encryption. This makes the vehicle network vulnerable to network threats such as forged message injection and denial of service (DoS) attacks. As vehicles are increasingly connected to external networks, the risks posed by these threats have increased significantly, which may endanger the normal operation of the vehicle and even the lives of drivers and passengers. In response to these security issues, traditional anti-attack measures have been unable to simultaneously take into account the detection needs of known and unknown attacks. Summary of the invention
[0004] In order to alleviate the above problems, the present application provides a hybrid multi-stage vehicle network intrusion detection method, including:
[0005] A lightweight supervised learning model based on convolutional neural networks and attention mechanisms detects known types of attacks on vehicle network messages in the first stage;
[0006] Based on an unsupervised model constructed by an autoencoder and a generative adversarial network, unknown type attacks are detected on the vehicle network messages in the second stage.
[0007] Optionally, the lightweight supervised learning model based on the convolutional neural network and the attention mechanism comprises, before the step of performing known type attack detection on the vehicle network message in the first stage:
[0008] Based on the feature extraction network, a lightweight supervised learning model based on a convolutional neural network is designed, and the lightweight supervised learning model is trained based on pre-labeled message data to classify and detect input vehicle network messages.
[0009] Optionally, the hybrid multi-stage vehicle network intrusion detection method further includes, during the intrusion detection process:
[0010] By quickly classifying the vehicle network messages in a first stage to identify known types of attack messages;
[0011] If the first stage determines that the vehicle network message is a normal message, the second stage will be entered to determine the unknown type of attack message;
[0012] If the in-vehicle network message is determined to be a normal message in the second stage, the in-vehicle network message is determined to be a normal message.
[0013] Optionally, a lightweight supervised learning model based on a convolutional neural network and an attention mechanism includes the following steps in the first phase of detecting known types of attacks on vehicle network messages:
[0014] Converting the received in-vehicle network message into a message image, performing feature extraction on the message image based on the convolutional neural network, and obtaining a message feature graph;
[0015] The message image is classified and detected according to the message feature graph to identify attack messages of known types.
[0016] Optionally, in the process of converting the received vehicle network message into a message image, the feature value of each vehicle network message is used as a pixel in the RGB image, multiple vehicle network messages constitute a 9x9 image, and a 3-channel color image is generated through continuous messages.
[0017] Optionally, feature extraction is performed on the message image based on a convolutional neural network. In the process of obtaining the message feature map, a lightweight multi-branch convolutional neural network is used, each branch uses a convolution kernel of a different size, and a spatial attention mechanism is used to enhance the lightweight supervised learning model's attention to important features, extract features from the message image, and integrate the features extracted by each branch through a compression and excitation network to form a final feature map.
[0018] Optionally, the unsupervised model constructed based on the autoencoder and the generative adversarial network includes, before the step of performing unknown type attack detection on the in-vehicle network message in the second stage:
[0019] A generator and a discriminator are constructed based on a generative adversarial network, and an encoder and a decoder are constructed based on an autoencoder, wherein the generator is used to generate a first image from low-dimensional noise, the discriminator is used to determine the probability that the first image is a normal message image, the encoder is used to convert the first image into a low-dimensional feature vector, and the decoder is used to reconstruct the low-dimensional feature vector into a second image;
[0020] The generator and the discriminator are trained by normal message data, and whether the input message image is of an attack type is judged by reconstruction error and evaluation of the discriminator to help the training of the encoder and the decoder.
[0021] Optionally, the generator and the discriminator are trained by normal message data, and in the process of judging whether the input message image is of an attack type through reconstruction error and evaluation of the discriminator, a minimized loss function is obtained based on a first preset expression, wherein the loss function includes the deviation between the distribution of the first image and the distribution of the second image, and a gradient penalty term derived from the second image.
[0022] Optionally, the encoder includes a first encoder and a second encoder; and the process of using the discriminator to assist in the training of the encoder and the decoder includes:
[0023] Mapping the first image to a first feature vector in a latent space using the first encoder, and reconstructing the feature vector using the decoder to generate the second image;
[0024] The second image is mapped to a second feature vector using the second encoder, and the first image and the second image are sent to the discriminator for similarity determination.
[0025] Optionally, the unsupervised model constructed based on the autoencoder and the generative adversarial network, in the second stage, performs unknown type attack detection on the in-vehicle network message, including:
[0026] Using the discriminator to obtain an image deviation value between the first image and the second image, obtaining an evaluation deviation value of the discriminator and a vector deviation value of a latent space vector;
[0027] The image deviation value, the evaluation deviation value and the vector deviation value are combined to calculate and obtain an abnormality score, so as to determine whether the in-vehicle network message is an attack message based on a preset score threshold.
[0028] The present application discloses a hybrid multi-stage vehicle network intrusion detection method. A lightweight supervised learning model based on a convolutional neural network and an attention mechanism is used to perform known type attack detection on vehicle network messages in the first stage. An unsupervised model based on an autoencoder and a generative adversarial network is used to perform unknown type attack detection on the vehicle network messages in the second stage. Through a multi-stage architecture, known attacks can be quickly detected in the first stage, reducing the burden of processing unknown attacks in the second stage. It can cover known and unknown attacks and reduce the underreporting rate of potential attacks. The lightweight design concept also adapts to the computing resource limitations of vehicle-mounted equipment and shortens the detection response time. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0030] Figure 1 It is a flowchart of a hybrid multi-stage in-vehicle network intrusion detection method for this application.
[0031] Figure 2 It is a schematic diagram of CAN message image generation in an embodiment of this application;
[0032] Figure 3 It is a schematic diagram of the structures of FANet, LWNet and their components of a hybrid multi-stage in-vehicle network intrusion detection system for this application;
[0033] Figure 4 It is a schematic diagram of the GAE model design of a hybrid multi-stage in-vehicle network intrusion detection system for this application;
[0034] Figure 5 It is a schematic diagram of two stages of known attack detection and unknown attack detection in an embodiment of this application.
[0035] Figure 6 It is a schematic diagram of the CHD and SAD datasets in an embodiment of this application.
[0036] Figure 7 It is a schematic diagram of the number of CAN message images of each type in the CHD and SAD datasets in an embodiment of this application.
[0037] Figure 8 It is a schematic diagram of the number of each data type in the training set, validation set and test set in an embodiment of this application.
[0038] Fig. 9 It is a schematic diagram of the detection result of the first stage in an embodiment of this application.
[0039] Fig.10 It is a schematic diagram of the search process of the optimal threshold in an embodiment of this application.
[0040] Fig.11 It is a schematic diagram of the detection result of the second stage in an embodiment of this application.
[0041] Fig.12 It is a schematic diagram of the ablation experiment result in an embodiment of this application.
[0042] Fig.13Schematic diagram of comparative experimental performance of HMS-IDS according to an embodiment of the present application.
[0043] The realization of the purpose, functional features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. The above-mentioned drawings have shown clear embodiments of this application, which will be described in more detail later. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0044] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0045] It should be understood that, although the terms first, second, third, etc. may be used to describe various information in this article, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this article, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination". Furthermore, as used in this article, the singular forms "one", "one" and "the" are intended to also include plural forms, unless there is an opposite indication in the context. It should be further understood that the terms "comprising" and "including" indicate that there are the described features, steps, operations, elements, components, projects, kinds, and / or groups, but do not exclude the existence, occurrence or addition of one or more other features, steps, operations, elements, components, projects, kinds, and / or groups. The terms "or", "and / or", "including at least one of the following" etc. used in this application can be interpreted as inclusive, or mean any one or any combination. For example, “comprising at least one of the following: A, B, C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”, and for another example, “A, B or C” or “A, B and / or C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”. An exception to this definition will only occur when a combination of elements, functions, steps or operations are inherently mutually exclusive in some manner.
[0046] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0047] First embodiment
[0048] This application first provides a hybrid multi-stage vehicle network intrusion detection method. Figure 1 A flow chart of a hybrid multi-stage vehicle network intrusion detection method is provided for this application.
[0049] like Figure 1 As shown, a hybrid multi-stage vehicle network intrusion detection method includes:
[0050] S10: A lightweight supervised learning model built based on convolutional neural networks and attention mechanisms to detect known types of attacks on in-vehicle network messages in the first stage.
[0051] The number and functionality of electronic control units (ECUs) in in-vehicle networks (IVNs) play a crucial role in driving experience and vehicle functionality. ECUs are considered the brain of the vehicle system, responsible for monitoring and controlling a variety of functions from engine management to safety systems. Modern vehicles are equipped with a large number of ECUs, and their number is expected to continue to increase. As the number of ECUs increases, the controller area network (CAN) is widely adopted to achieve efficient communication of in-vehicle network messages between individual ECUs. As a standard protocol for IVNs, CAN provides a reliable framework for data exchange between ECUs, which greatly facilitates the integration and collaboration of automotive systems. CAN is a standard protocol for in-vehicle networks (IVNs), so most existing IVN intrusion detection datasets consist of continuous CAN messages. For example, lightweight supervised learning models are a solution proposed in the field of machine learning and deep learning to cope with resource constraints and efficiency requirements. Lightweight supervised learning models refer to supervised learning models that minimize model complexity and computational requirements while maintaining model performance. These models are usually lightweight by optimizing algorithms, reducing model parameters, and using more efficient computing resources.
[0052] For example, Convolutional Neural Networks (CNN) is a type of feedforward neural network with deep structure that contains convolution calculations. It is one of the important algorithms for deep learning. The design of CNN is inspired by the ability of animal visual systems to process information hierarchically. It extracts local features of input data through convolution operations, forms complex feature representations through multi-layer convolution and pooling operations, and finally performs tasks such as classification or regression through fully connected layers. Attention Mechanism is a technology that enables AI models to dynamically focus on the most relevant parts when processing input data. It enables the model to capture important information more effectively by assigning different weights, thereby improving the completion of tasks. The attention mechanism imitates the selective attention ability of humans when processing information, allowing the model to dynamically adjust its attention weights when processing input data, thereby highlighting important information and ignoring unimportant information.
[0053] S20: Based on the unsupervised model constructed by the autoencoder and the generative adversarial network, in the second stage, unknown type attack detection is performed on the in-vehicle network message.
[0054] For example, Generative Adversarial Networks (GAN) is a deep learning model that generates new data similar to training data through adversarial training of two neural networks. GAN consists of a generator and a discriminator. The goal of the generator is to generate samples that look real, while the goal of the discriminator is to distinguish between real samples and generated samples.
[0055] Through a multi-stage architecture, this embodiment can quickly detect known attacks in the first stage, reducing the burden of processing unknown attacks in the second stage; it can cover known attacks and unknown attacks, reducing the missed reporting rate of potential attacks; the lightweight design concept also adapts to the computing resource limitations of vehicle-mounted equipment and shortens the detection response time.
[0056] Please refer to Figure 1-Figure 5 Optionally, the lightweight supervised learning model based on the convolutional neural network and the attention mechanism comprises, before the step of performing known type attack detection on the vehicle network message in the first stage:
[0057] Based on the feature extraction network, a lightweight supervised learning model based on a convolutional neural network is designed, and the lightweight supervised learning model is trained based on pre-labeled message data to classify and detect input vehicle network messages.
[0058] For example, the hybrid multi-stage vehicle network intrusion detection system (HMS-IDS) aims to quickly and efficiently detect known attacks in the vehicle network in the first stage of detection. Figure 2 As shown, based on the feature extraction network FANet (feature extractor), a lightweight supervised learning model LWNet based on the convolutional neural network CNN can be designed to classify (detect) the input message image. The first stage model relies on a large amount of labeled data for training, so it is very effective against known attacks.
[0059] Exemplarily, LWNet uses the feature maps extracted by FANet to classify the input CAN image. Specifically, it fuses the feature maps in the channel dimension through point-by-point convolution, and then flattens the fused feature maps into a vector. Finally, the vector is mapped to the classification result through a fully connected layer (FC).
[0060] Optionally, the hybrid multi-stage vehicle network intrusion detection method further includes, during the intrusion detection process:
[0061] By quickly classifying the vehicle network messages in a first stage to identify known types of attack messages;
[0062] If the first stage determines that the vehicle network message is a normal message, the second stage will be entered to determine the unknown type of attack message;
[0063] If the in-vehicle network message is determined to be a normal message in the second stage, the in-vehicle network message is determined to be a normal message.
[0064] For example, the schematic diagram of the structure of HMS-IDS is shown as follows: Figure 5As shown. The constructed intrusion detection system is divided into at least two stages: the first stage is used to detect known and common attacks, while the second stage is used to detect unknown and low-frequency attacks. The two stages are independent of each other during the training process. After each training is completed, attack detection can be performed. During the detection process, the CAN message image is generated by the CAN image generator and first sent to the first stage for classification. If the first stage identifies the image as an attack, the corresponding vehicle network message is immediately classified as an attack. However, if the first stage determines that the image is normal, it will be forwarded to the second stage for further evaluation. In summary, HMS-IDS classifies a CAN message image as normal only when both the first and second stages classify it as normal. Conversely, once any stage identifies it as an attack, HMS-IDS classifies it as an attack. This multi-stage intrusion detection system has many advantages: First, it can detect both known and unknown attacks, thereby ensuring that unknown attacks are not missed, while an intrusion detection system with only the first stage may miss unknown attacks. In addition, since the first stage is more lightweight than the second stage and performs better in detecting known attacks, a large number of known attacks can be efficiently detected in the first stage, while only a small number of unknown attacks need to be processed in the second stage. Therefore, the intrusion detection system proposed in this application has a shorter attack response time and better overall detection performance than the intrusion detection system with only the second stage.
[0065] Optionally, a lightweight supervised learning model based on a convolutional neural network and an attention mechanism includes the following steps in the first phase of detecting known types of attacks on vehicle network messages:
[0066] Converting the received in-vehicle network message into a message image, performing feature extraction on the message image based on the convolutional neural network, and obtaining a message feature graph;
[0067] The message image is classified and detected according to the message feature graph to identify attack messages of known types.
[0068] Exemplarily, a CAN image generator is designed to construct multiple independent CAN messages into a CAN image, and the image is used as the input of an intrusion detection system (IDS). In the field of IVN intrusion detection, CAN ID and payload can be used as features because anomalies are usually associated with them. Through the designed CAN image generator, CAN messages can be converted into CAN images, which can be used in the training and testing phases of the intrusion detection system (IDS). In addition, in actual deployment scenarios, real CAN data can also be processed by this image generator.
[0069] In order to extract the features of CAN images, a lightweight feature extraction network (FANet) is designed. This network can be used in the first and second stages of HMS-IDS. The extracted features are spliced to generate the final feature map, thereby obtaining a richer and more diverse CAN message image feature representation. The first stage of HMS-IDS aims to quickly and accurately detect known attacks using supervised learning models. During the training process of supervised learning, since a large amount of annotated attack data is required to learn the feature representation of attack data, the model of the first stage is sensitive to known attacks. If the first stage identifies that the image is an attack, it is immediately classified as an attack. However, if the first stage determines that the image is normal, it will be forwarded to the second stage for further evaluation.
[0070] Optionally, in the process of converting the received vehicle network message into a message image, the feature value of each vehicle network message is used as a pixel in the RGB image, multiple vehicle network messages constitute a 9x9 image, and a 3-channel color image is generated through continuous messages.
[0071] For example, Figure 2 As shown, the CAN image generator also uses the 9 feature values in the CAN message to construct a CAN message image, including the CAN ID and the 8-byte payload. The feature values can be converted to decimal integers between 0 and 256 for RGB encoding. Specifically, the image generator uses the 9 feature values of a single CAN message to construct a row of the CAN image, and each feature value represents a pixel in the row. Therefore, with 9 CAN messages, a 9×9 image can be constructed, where each row corresponds to a message. In order to include more information, the CAN image generator uses 27 consecutive CAN messages to construct the image, and the resulting image size is 9×9×3. Finally, by applying RGB encoding to three different channels of the feature values at the same position, we can get a colored CAN image. If all the CAN messages that make up the image are normal messages, the CAN message image is labeled "normal"; if any of the CAN messages is an attack message, the CAN message image is labeled "attack". Through the designed CAN image generator, we can convert CAN messages into CAN message images, which can be used in the training and testing phases of the intrusion detection system (IDS).
[0072] Optionally, during the process of extracting features from the message image based on a convolutional neural network to obtain a message feature map, a lightweight multi-branch convolutional neural network is used. Each branch uses convolutional kernels of different sizes, and a spatial attention mechanism is adopted to enhance the lightweight supervised learning model's attention to important features, extract features from the message image, and integrate the features extracted by each branch through a squeeze-and-excitation network to form the final feature map.
[0073] Exemplarily, the main idea of FANet is to extract features through three branches, and then splice these features to generate the final feature map, thereby obtaining a richer and more diverse CAN image feature representation. By using different convolutional kernel sizes in different branches, FANet can capture feature information of various scales at the same time. Larger convolutional kernels have a larger receptive field and can capture more extensive features, while smaller convolutional kernels have a smaller receptive field and are more suitable for capturing local fine features. Specifically, larger convolutional kernels can focus on the connections between messages in the CAN image, while smaller convolutional kernels tend to focus on the features of individual messages.
[0074] To ensure that FANet is lightweight enough, the depthwise separable convolution (DSC) idea in MobileNet can be adopted. DSC consists of two parts: depthwise convolution and pointwise convolution. In depthwise convolution, the number of input channels, output channels, and convolutional kernels is the same. It uses an independent convolutional kernel for each channel of the input feature map, and then splices the outputs of all convolutional kernels to form the final result. Pointwise convolution is essentially a 1×1 convolution and has two functions in DSC: First, it enables DSC to flexibly change the number of output channels; second, it performs channel fusion on the feature map output by depthwise convolution.
[0075] Considering that the specification of the CAN message image is 9×9×3, depthwise separable convolutions with convolutional kernel sizes of 3, 5, and 7 can be used in the three branches respectively to extract features of different scales. In addition, to further enable FANet to focus on important spatial positions in the CAN image, a spatial attention module (SAM) is added after the depthwise convolution of each branch. Specifically, SAM captures the importance of the feature map in the spatial dimension and strengthens the features at important spatial positions. After obtaining the features extracted by the three branches, they can be spliced to form the total feature map. At the same time, a Squeeze-and-Excitation module (SEM) is used to enable the final feature map to effectively integrate the features extracted by each branch. Specifically, SEM adaptively learns channel weights to adjust the influence of each channel, so that FANet can make more effective use of the feature information of each branch. The specific structures of FANet, SAM, and SEM are as Figure 3 as shown
[0076] Optionally, before the step of performing unknown type attack detection on the vehicle network message in the second stage, the unsupervised model constructed based on the autoencoder and the generative adversarial network includes:
[0077] Construct a generator and a discriminator based on the generative adversarial network, and construct an encoder and a decoder based on the autoencoder, where the generator is used to generate a first image from low-dimensional noise, the discriminator is used to determine the probability that the first image is a normal message image, the encoder is used to convert the first image into a low-dimensional feature vector, and the decoder is used to reconstruct the low-dimensional feature vector into a second image;
[0078] Train the generator and the discriminator with normal message data, and determine whether the input message image is of an attack type by evaluating the reconstruction error and the discriminator to assist in the training of the encoder and the decoder.
[0079] Exemplarily, an unsupervised model (GAE) based on a generative adversarial network (GAN) and an autoencoder (AE) is used as the second stage to detect unknown attacks. The model consists of multiple components: a generator, a discriminator, an encoder, and a decoder. The role of the generator is to generate a first image x from low-dimensional noise z, the role of the discriminator is to determine the probability that the input x is a normal CAN message image, and the encoder is used to convert the first image x into a low-dimensional feature vector The role of the decoder is to convert the low-dimensional feature vector into a second image where the low-dimensional noise z can be randomly generated. The discriminator of GAN can be used to better train AE, and at the same time, the result of the discriminator is used as part of the anomaly score to more effectively detect unknown attacks. The overall structure of the unsupervised model and the specific structure of each component are as Figure 4 shown
[0080] Optionally, in the process of training the generator and the discriminator with normal message data and determining whether the input message image is of an attack type by evaluating the reconstruction error and the discriminator, a minimized loss function is obtained based on the following preset expression, and the loss function includes the deviation between the distribution of the first image and the distribution of the second image, and a gradient penalty term derived from the second image:
[0081]
[0082] where L is the loss function, x and represent the first image and the second image respectively, and is an interpolated message image, which is a weighted combination of x and pr is the distribution of the first image, which is a real image; p g is the distribution of the second image, which is a generated image. D(x) is the discrimination result of the discriminator for the original image (i.e., the first image x), and λ is the calculated weight coefficient.
[0083] Exemplarily, during the entire training process, normal data can be used for training. As Figure 4 shown, first, a GAN is trained to obtain a discriminator with good enough performance to determine whether the data comes from the normal data distribution. The GAN used in this application can be an improved version of WGAN-GP, which can be stably trained and converge quickly. When training the GAN, the goal is to minimize the loss function as shown in formula (1).
[0084] where x and represent the real message image and the generated message image respectively, and is the interpolated message image, which is a weighted combination of x and . Therefore, this loss function includes the deviation between the distribution p r of the real image and the distribution p g of the generated image, as well as the gradient penalty term derived from .
[0085] Optionally, the encoder includes a first encoder and a second encoder; the process of using the discriminator to assist in the training of the encoder and decoder includes:
[0086] using the first encoder to map the first image to a first feature vector in the latent space, and using the decoder to reconstruct the second image by reconstructing the feature vector;
[0087] using the second encoder to map the second image to a second feature vector, and sending the first image and the second image into the discriminator for similarity judgment.
[0088] Exemplarily, the discriminator obtained by training the GAN can determine whether the data comes from the normal data distribution. Then, this discriminator can be used to assist in the training of the encoder and decoder, as Figure 4 shown. When a CAN message image is input as the first image x into the GAE model, the process involves several steps. First, encoder 1 maps x to a low-dimensional feature vector in the latent space, i.e., the first feature vector z. Then the decoder uses this low-dimensional feature vector z to generate a reconstructed image, i.e., the second image Finally, encoder 2 maps the reconstructed image, i.e., the second image to another vector, i.e., the second feature vector In this process, the original image (i.e., the first image x) and the reconstructed image (i.e., the second image ) are fed into the discriminator for discrimination. If the encoder and decoder are effective, the original image and the reconstructed image should be very similar, and the vectors z and should be very similar. This means that the process should approximate an identity mapping. Therefore, the discriminator should produce very similar results when evaluating the original image and the reconstructed image.
[0089] Therefore, the loss function can be composed of three parts: the original image (i.e., the first image x) and the reconstructed image (i.e., the second image ) between the reconstruction error L rec (As shown in formula (2)), the discriminator classifies the original image (i.e., the first image x) and the reconstructed image (i.e., the second image )’s discrimination deviation L dis (as shown in formula (3)), and the first eigenvector z and the second eigenvector of the latent space vector The vector deviation L between z (As shown in formula (4)).
[0090]
[0091] Among them, D(x) is the discriminator's discriminant result on the original image (i.e., the first image x), The discriminator reconstructs the image (i.e., the second image )’s judgment result.
[0092] Therefore, the total loss function L is a weighted combination of these three parts, as shown in formula (5).
[0093] L=λ rec L rec +λ dis L dis +λ z L z (5)
[0094] Among them, λ rec is the reconstruction error L rec The weight coefficient, λ dis is the discrimination bias L dis The weight coefficient, λ z is the vector deviation L z The goal is to minimize this loss, which means that the smaller the loss, the greater the similarity between the original image and the reconstructed image. Ultimately, after the model training is completed, this process will approximate the identity mapping of a normal CAN image.
[0095] Optionally, the unsupervised model constructed based on the autoencoder and the generative adversarial network, in the second stage, performs unknown type attack detection on the in-vehicle network message, including:
[0096] Using the discriminator to obtain an image deviation value between the first image and the second image, obtaining an evaluation deviation value of the discriminator and a vector deviation value of a latent space vector;
[0097] The image deviation value, the evaluation deviation value, and the vector deviation value are combined based on the following expression to calculate the abnormality score Score, and whether the vehicle network message is an attack message is determined based on a preset score threshold:
[0098]
[0099] Among them, Score rec is the first image x and the second image The image deviation value between dis is the evaluation deviation value of the discriminator, Score z is the latent space vector deviation value.
[0100] For example, after all modules are trained, they can be used to generate an anomaly score for each CAN message image. Figure 4 As shown, similar to the loss function, the anomaly score is also divided into three parts: the original image x and the reconstructed image The image deviation value Score rec , the evaluation deviation value Score of the discriminator dis And the vector deviation value Score of the latent space vector z Therefore, the anomaly score is a combination of these three parts, as shown in formula (6).
[0101] Since the model is trained based on normal CAN message images, if the test CAN message image is normal, the scores from the three parts should be low (ideally close to 0, as the process approximates the identity mapping of normal CAN message images). On the contrary, for abnormal images containing attacks, the GAE model cannot reconstruct them well, resulting in a large deviation between the original image and the reconstructed image. Therefore, the anomaly score of each part should be high. Therefore, there will be a significant difference in the anomaly score between normal and attack images.
[0102] For example, a preset scoring threshold η can be defined to determine whether such a difference can be considered as an attack. Using the preset scoring threshold η, the input can be classified: when the abnormal score Score of the input CAN message image exceeds the preset scoring threshold η, it can be considered as an attack; otherwise, it is judged to be normal. This classification criterion is shown in formula (7).
[0103]
[0104] Second embodiment
[0105] On the basis of the first embodiment, in order to verify the high performance of the proposed HMS-IDS, this application uses two authoritative in-vehicle network (IVN) intrusion detection datasets for evaluation. The first dataset is the Car Hacking Dataset (CHD). CHD is constructed by recording CAN traffic from real vehicles through the OBD-II interface, during which message injection attacks occur. The attack types of CHD include fuzzy attack, denial of service attack, spoofing the RPM Gauge, and spoofing the Drive Gear. The dataset contains 300 injections of each attack, each attack lasts 3 to 5 seconds, and a total of 30 to 40 minutes of CAN traffic is recorded
[12] . The other dataset is the Survival Analysis Dataset (SAD), which focuses on the following three attack scenarios: flooding attack, fuzzy attack, and malfunction attack. These scenarios will immediately and severely affect the vehicle functions, or deepen the intensity and damage of the attack
[25] . In summary, these two datasets cover common attacks in IVN, including DoS, fuzzy attacks, forgery attacks, flooding attacks, and functional failure attacks.
[0106] Figure 6 The number of normal and attack messages in the two datasets CHD and SAD is listed. Figure 7 The number of CAN message images of each type in the CHD and SAD datasets is shown, showing the number of CAN images obtained after the builder processes the messages. These CAN images are the input to our intrusion detection system. We allocate 70% of the data for training, 10% for validation, and 20% for testing. Figure 8 This is an illustration of the number of each data type in the training set, validation set, and test set, and lists the number of each type of data used in training, validation, and testing after the data set is divided.
[0107] The experimental system proposed in this application is developed based on Python 3.7.16 and PyTorch 1.13.1. The development environment uses PyCharm 2022.2.2Community version. The computer hardware configuration includes an Intel Xeon E5-2678 v3@2.50GHz CPU, 64GB memory, and an NVIDIA RTX 2080SUPER GPU with 8GB VRAM.
[0108] To evaluate the performance of our intrusion detection system (IDS), we used several common evaluation metrics in the field, including accuracy, precision, recall, and F1-score.
[0109] The definitions of these indicators are as follows:
[0110] Accuracy: The percentage of correctly predicted samples, as a proportion of the total number of samples.
[0111] Precision: The proportion of predicted positive samples that are actually positive samples.
[0112] Recall rate: The proportion of samples that are actually positive samples that are correctly predicted to be positive.
[0113] F1-score: The harmonic mean of precision and recall, used to provide a balanced assessment of classifier performance.
[0114] The definition of each indicator is shown in formulas (8), (9), (10) and (11), respectively, where TP is the number of samples correctly predicted as positive samples, TN is the number of samples correctly predicted as negative samples, FP is the number of samples incorrectly predicted as positive samples, and FN is the number of samples incorrectly predicted as negative samples.
[0115]
[0116] The first-stage model LWNet is trained and evaluated using the divided CHD and SAD datasets. The high performance of LWNet is verified by accuracy, precision, recall, and F1 score. Since the training set contains all types of attacks, all attacks can be identified as known attacks. The experimental results are shown in Figure 2. Fig. 9 As shown in Figure 2, LWNet performs well in detecting all types of known attacks, with accuracy, precision, recall, and F1 score close to 1.
[0117] In order to more intuitively demonstrate the performance of LWNet, the detection results can be displayed using a confusion matrix. The confusion matrix shown in the experiment of this application shows that the number of misclassified samples is extremely small, which further illustrates the excellent performance of LWNet in detecting known attacks. Based on the above observations, the first stage of the HMS-IDS proposed in this application shows strong performance in detecting known attacks.
[0118] For the second-stage model GAE, a threshold that can make the model performance most balanced can be found. We call the threshold that makes the performance the best the optimal threshold η. In intrusion detection, the F1 score is a comprehensive indicator because it considers both the precision and recall of the model. When the F1 score reaches a high enough value, the performance of GAE is very good. Therefore, the F1 score can be used to determine the optimal threshold η, that is, the threshold that maximizes the F1 score. The search process to determine the optimal threshold is as follows Fig.10 shown.
[0119] The second-stage model GAE is trained and evaluated using the divided CHD and SAD datasets. The high performance of GAE is verified by accuracy, precision, recall, and F1 score. Since GAE is an unsupervised model, only normal data in the training set is used in its training process. Therefore, any attack in the CHD and SAD datasets can be considered as an unknown attack. The experimental results are shown in Figure 2. Fig.11 As shown in the figure. GAE performs very well in detecting unknown attacks. On the CHD dataset, the overall accuracy is 0.986 and the F1 score is 0.991; on the SAD dataset, the overall accuracy is 0.985 and the F1 score is 0.987. In order to intuitively demonstrate the excellent performance of the GAE model, the confusion matrix can be used to display the detection results. According to the experimental results, the second stage of the HMS-IDS proposed in this application performs very well in detecting unknown attacks.
[0120] In order to verify the effectiveness of each part of the GAE loss function, a set of ablation experiments can be performed. The experimental results are as follows Fig.12 As shown in Figure 1, by using only specific parts of the loss function or the full loss function, it is possible to observe that the best performance of the model occurs when the full loss function is used. This suggests that each part of the loss function contributes to the model performance.
[0121] After verifying the effectiveness of the first and second phases respectively, we also need to verify the effectiveness of the entire multi-phase IDS. The main principle of the multi-phase IDS we proposed is to detect known attacks in phase 1 and unknown attacks in phase 2. For phase 1 based on supervised learning, the attack types included in the training set can be regarded as known attacks, while the attack types not included in the training set are regarded as unknown attacks. For phase 2 based on unsupervised learning, since it only uses normal data for training, all attack types are regarded as unknown attacks.
[0122] To evaluate our proposed HMS-IDS, we conducted a series of experiments in which one of the attack types in the dataset was considered as an unknown attack and the other attack types were considered as known attacks. For example, in the CHD dataset, if DoS is considered an unknown attack, then Fuzzy, Gear, and RPM are considered known attacks. Therefore, we conducted a total of four sets of experiments on the CHD dataset and three sets of experiments on the SAD dataset, and took the average of these experimental results as the final result. In these experiments, the training data of the stage 1 model included normal data and known attack data, while the training data of the stage 2 model only included normal data. The test set contains normal data, known attack data, and unknown attack data.
[0123] To demonstrate the superior performance of the proposed HMS-IDS, it is compared with a large number of existing state-of-the-art IVN IDSs, including three supervised learning-based IDSs, five unsupervised learning-based IDSs, and three synthetic IDSs. Fig.13 The results show that when IDS needs to detect both known and unknown attacks, the hybrid IDS designs different models for known and unknown attacks, so its overall performance is better than that of IDS based only on supervised learning or unsupervised learning.
[0124] The HMS-IDS proposed in this application uses a supervised learning model to detect known attacks in the first stage, and a pure unsupervised learning model to detect unknown attacks in the second stage, thus achieving optimal performance in all indicators. On the CHD dataset, the F1 score is as high as 0.996 and the accuracy is 0.993; on the SAD dataset, the F1 score is 0.993 and the accuracy is 0.992. These results are far superior to IDS based only on supervised learning or unsupervised learning and other hybrid IDS.
[0125] It should be noted that in the present application, step codes such as S10, S20, etc. are used for the purpose of expressing the corresponding content more clearly and concisely, and do not constitute a substantial limitation on the sequence. When implementing the step, those skilled in the art may execute S20 first and then S10, etc., but these should all be within the scope of protection of the present application.
[0126] In the system embodiment provided in the present application, all technical features of any of the above-mentioned method embodiments may be included, and the expanded and explained contents of the specification are basically the same as those of the above-mentioned method embodiments, and will not be repeated here.
[0127] The embodiment of the present application further provides a computer program product, which includes a computer program code. When the computer program code runs on a computer, the computer executes the methods in the above various possible implementation modes.
[0128] An embodiment of the present application also provides a chip, including a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device equipped with the chip executes the methods in various possible implementation modes as described above.
[0129] In the present application, the same or similar terminology concepts, technical solutions and / or application scenario descriptions are generally described in detail only the first time they appear. When they appear again later, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of the present application, for the same or similar terminology concepts, technical solutions and / or application scenario descriptions that are not described in detail later, reference can be made to the previous related detailed descriptions.
[0130] In the present application, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0131] The various technical features of the technical solution of the present application can be arbitrarily combined. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present application.
[0132] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A hybrid multi-stage vehicle network intrusion detection method, characterized in that: include: A lightweight supervised learning model based on convolutional neural networks and attention mechanisms detects known types of attacks on vehicle network messages in the first stage; Based on an unsupervised model constructed by an autoencoder and a generative adversarial network, unknown type attacks are detected on the vehicle network messages in the second stage.
2. A hybrid multi-stage vehicle network intrusion detection method according to claim 1, characterized in that: The lightweight supervised learning model based on the convolutional neural network and the attention mechanism includes, before the step of detecting known types of attacks on the vehicle network messages in the first stage: Based on the feature extraction network, a lightweight supervised learning model based on a convolutional neural network is designed, and the lightweight supervised learning model is trained based on pre-labeled message data to classify and detect input vehicle network messages.
3. A hybrid multi-stage vehicle network intrusion detection method according to claim 2, characterized in that: The hybrid multi-stage vehicle network intrusion detection method further includes, during the intrusion detection process, comprising: By quickly classifying the vehicle network messages in a first stage to identify known types of attack messages; If the first stage determines that the vehicle network message is a normal message, the second stage will be entered to determine the unknown type of attack message; If the in-vehicle network message is determined to be a normal message in the second stage, the in-vehicle network message is determined to be a normal message.
4. A hybrid multi-stage vehicle network intrusion detection method according to claim 3, characterized in that: The lightweight supervised learning model based on convolutional neural networks and attention mechanisms includes the following steps in the first phase of detecting known types of attacks on vehicle network messages: Converting the received in-vehicle network message into a message image, performing feature extraction on the message image based on the convolutional neural network, and obtaining a message feature graph; The message image is classified and detected according to the message feature graph to identify attack messages of known types.
5. A hybrid multi-stage vehicle network intrusion detection method according to claim 4, characterized in that: In the process of converting the received vehicle network message into a message image, the feature value of each vehicle network message is used as a pixel in the RGB image, multiple vehicle network messages constitute a 9x9 image, and a 3-channel color image is generated through continuous messages.
6. A hybrid multi-stage vehicle network intrusion detection method according to claim 5, characterized in that: The message image is feature extracted based on a convolutional neural network. In the process of obtaining the message feature map, a lightweight multi-branch convolutional neural network is used, each branch uses a convolution kernel of a different size, and a spatial attention mechanism is used to enhance the lightweight supervised learning model's attention to important features, extract features from the message image, and integrate the features extracted by each branch through a compression and excitation network to form a final feature map.
7. A hybrid multi-stage vehicle network intrusion detection method according to claim 6, characterized in that: The unsupervised model based on the autoencoder and the generative adversarial network comprises, before the step of performing unknown type attack detection on the vehicle network message in the second stage: A generator and a discriminator are constructed based on a generative adversarial network, and an encoder and a decoder are constructed based on an autoencoder, wherein the generator is used to generate a first image from low-dimensional noise, the discriminator is used to determine the probability that the first image is a normal message image, the encoder is used to convert the first image into a low-dimensional feature vector, and the decoder is used to reconstruct the low-dimensional feature vector into a second image; The generator and the discriminator are trained by normal message data, and whether the input message image is of an attack type is judged by reconstruction error and evaluation of the discriminator to help the training of the encoder and the decoder.
8. A hybrid multi-stage vehicle network intrusion detection method according to claim 7, characterized in that: The generator and the discriminator are trained by normal message data. In the process of judging whether the input message image is of an attack type by reconstructing the error and evaluating the discriminator, a minimized loss function is obtained based on a first preset expression. The loss function includes the deviation between the distribution of the first image and the distribution of the second image, and the gradient penalty term derived from the second image.
9. A hybrid multi-stage vehicle network intrusion detection method according to claim 8, characterized in that: The encoder includes a first encoder and a second encoder; The process of using the discriminator to assist the training of the encoder and decoder includes: Mapping the first image to a first feature vector in a latent space using the first encoder, and reconstructing the feature vector using the decoder to generate the second image; The second image is mapped to a second feature vector using the second encoder, and the first image and the second image are sent to the discriminator for similarity determination.
10. A hybrid multi-stage vehicle network intrusion detection method according to claim 9, characterized in that: The unsupervised model based on the autoencoder and the generative adversarial network, in the second stage, performs unknown type attack detection on the vehicle network message, including: Using the discriminator to obtain an image deviation value between the first image and the second image, obtaining an evaluation deviation value of the discriminator and a vector deviation value of a latent space vector; An abnormality score is calculated based on the image deviation value, the evaluation deviation value, and the vector deviation value, and whether the in-vehicle network message is an attack message is determined based on a preset score threshold.
Citation Information
Patent Citations
Smart switching apparatus and control method thereof
KR102914422B1