Lightweight method for anomaly detection in internet of vehicles based on deep learning and knowledge distillation
By training a lightweight student network using reconstruction-based unsupervised learning and knowledge distillation methods, the problem of anomaly detection on resource-constrained devices in the Internet of Vehicles is solved, achieving efficient detection of unknown anomalies.
Patent Information
- Application Number
- CN202411475169.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-10-22
AI Technical Summary
In the Internet of Vehicles (IoV), anomaly detection is difficult to deploy on resource-constrained devices, especially due to the difficulty in labeling abnormal data and the insufficient ability to detect unknown anomalies. Furthermore, existing deep learning methods are difficult to run efficiently on resource-constrained devices.
A reconstruction-based unsupervised learning method is used to train the teacher network, and a lightweight student network is trained through knowledge distillation. The teacher network is constructed using Transformer and LSTM, and the knowledge distillation method based on feature transfer is combined to achieve efficient detection by the lightweight student network.
It enables high-precision detection of unknown anomalies on resource-constrained devices, reduces the number of neural network parameters, improves detection speed, and solves the deployment problem on resource-constrained devices.
Smart Images

Figure CN119475153B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vehicle networking anomaly detection, and in particular to a vehicle networking anomaly detection lightweight method based on deep learning and knowledge distillation. BACKGROUND
[0002] Vehicle networking is a new wireless network technology that is attracting more and more attention. Vehicle networking connects vehicles, roads and traffic infrastructure, providing vehicles with instant communication, information services and traffic management functions. By monitoring vehicle location, speed and traffic conditions in real time, vehicle networking can improve the safety, efficiency and convenience of road traffic, thereby providing strong support for the construction of intelligent transportation systems.
[0003] However, with the rapid development of vehicle networking, network security problems are also emerging. Since vehicle communication involves a large amount of private data and basic safety information such as location and travel route, vehicle networking is threatened by various malicious attacks. For example, hackers may cause traffic accidents by tampering with the information sent by vehicles, fake vehicle identity information to create false congestion, or consume network resources by sending high-frequency messages, thereby affecting the entire network. Therefore, anomaly detection systems are used to detect these anomalies, so that some countermeasures can be taken against the abnormal messages, such as not forwarding messages considered to be abnormal.
[0004] Deep learning has great potential in anomaly detection. Previous anomaly detection schemes for vehicle networking mostly focus on detecting one or a few types of anomalies, while deep neural networks can detect multiple types of anomalies through training, with stronger robustness. However, in reality, anomalies are usually rare and are masked by a large number of normal messages, making data labeling difficult and expensive. Therefore, we need to study vehicle networking anomaly detection in an unsupervised setting. The heterogeneity of anomalies makes the problem difficult, which means we not only have to identify known anomalies, but also try our best to detect never-before-seen anomalies.
[0005] Deep learning anomaly detection also faces a serious problem: most devices in vehicle networking are resource-constrained devices, and we must reduce the resources required by the anomaly detection method as much as possible while ensuring detection performance, in order to make deployment possible. Knowledge distillation, as an effective model lightweight method in deep learning, can help us lightweight the deep learning model and facilitate deployment on resource-constrained devices.
[0006] In summary, the present application has high practical application value. SUMMARY
[0007] The purpose of the present application is to overcome the shortcomings and deficiencies of the prior art, and to provide a lightweight vehicle networking anomaly detection method based on deep learning and knowledge distillation, which trains a high-precision teacher network through a reconstruction-based unsupervised learning method, and then trains a lightweight student network through knowledge distillation while maintaining high detection accuracy, and realizes fast detection on the time window extracted from the vehicle time series.
[0008] To achieve the above purpose, the technical scheme provided by the present application is: a lightweight vehicle networking anomaly detection method based on deep learning and knowledge distillation, comprising the following steps:
[0009] 1) Use the public vehicle networking data set VeReMi Extension to preprocess the time series messages of each vehicle in the data set using a sliding window strategy to obtain a time window, and divide the preprocessed data set into a training set and a test set, wherein the training set only contains normal time windows, and the test set contains normal time windows and abnormal time windows;
[0010] 2) Construct a teacher network based on Transformer and LSTM, wherein the Transformer is used as an encoder to capture the dependency between each position in the input time window to obtain a compressed representation of the time window information, and the LSTM is used as a decoder to effectively maintain long-distance dependence through a memory unit and a gating mechanism, and the compressed information of the Transformer encoder is recovered into the input data through decoding; the teacher network uses a fully connected neural network as a linear embedding layer to replace the embedding layer in the original Transformer, which maps discrete input message features to a continuous vector space and expands the feature dimension of the input, so that the hidden relationship between the features can be better captured subsequently; the teacher network only uses the encoder of the Transformer to extract important information in the time window, and does not need the decoder to generate sequence output, so the decoder in the original Transformer is removed; the teacher network introduces a feedforward neural network layer after the Transformer and a fully connected network layer after the LSTM, the purpose of introducing the feedforward neural network layer after the Transformer is to integrate the features learned by the encoder and perform nonlinear transformation through layer-by-layer transmission to map high-dimensional data to low-dimensional representation, and the purpose of introducing the fully connected network layer after the LSTM is to map the decoded output to the expected output dimension; a large number of normal time windows in the training set are used to train the teacher network using a reconstruction-based unsupervised learning method, so that the teacher network learns the distribution of the normal time window and can restore the normal time window, but cannot restore the abnormal time window that has not been seen during training, thereby achieving the purpose of distinguishing abnormal time windows;
[0011] 3) Construct a lightweight student network similar to the teacher network structure, but with reduced number of layers, units and parameters, and use the trained teacher network to guide the lightweight student network to implement knowledge distillation on the normal time windows in the training set, increase the ability of the lightweight student network to reconstruct the normal time windows, and obtain a trained lightweight student network;
[0012] 4) Test the trained lightweight student network using the test set, select the threshold with the highest detection accuracy on the test set as the preset threshold, and then deploy the preset threshold and the trained lightweight student network to the roadside unit in the Internet of Vehicles for anomaly detection task, reconstruct the time window through the lightweight student network, take the reconstruction error of the time window as its anomaly score, and finally compare the anomaly score with the preset threshold, if the anomaly score exceeds the preset threshold, it is determined as abnormal, otherwise it is determined as normal.
[0013] Further, the step 1) comprises the following steps:
[0014] 1.1) Obtain the public Internet of Vehicles dataset VeReMi Extension and perform data labeling work, mark the messages in the dataset that are not changed due to noise or attack as normal messages, otherwise mark them as abnormal messages;
[0015] 1.2) In the dataset, a basic safety message of Internet of Vehicles carries the following contents: sending time, pseudonym ID, position x, position y, speed x, speed y, acceleration x, acceleration y, direction x and direction y, and the following preprocessing is performed on the above contents: only select sending time, pseudonym ID, position x, position y, speed x and speed y as input features for each message, for all basic safety messages, use z-score standardization strategy, and use sliding window strategy, use sliding window with window size of 50 and step size of 10 to intercept time window, if the time window contains abnormal messages, it is marked as abnormal time window, otherwise it is marked as normal time window;
[0016] 1.3) Divide the preprocessed dataset into training set and test set, wherein 90% of the normal time windows are used to construct the training set, and the remaining normal time windows and abnormal time windows are constructed according to the proportion of the original dataset to construct the test set.
[0017] Further, in step 2), the teacher network is constructed and trained, comprising the following steps:
[0018] 2.1) Build a teacher network with large model capacity as a time window reconstructor, which has the following modules from low to high: 1 layer of 16 head Transformer encoder, feedforward neural network layer, 4 layers of LSTM decoder and fully connected network layer; wherein the input normal time window W inputThe six features include sending time, pseudonym ID, position x, position y, speed x, and speed y, and the time window size is 50, so W input has a dimension of 50*6. After passing through the linear embedding layer in the Transformer encoder, it is mapped to a high-dimensional representation of 50*128. After passing through the entire Transformer encoder, the dimension is still 50*128. The feedforward neural network layer includes two fully connected layers, and LeakyReLU activation function is used between the two fully connected layers. The first fully connected layer expands the dimension to 50*512, and the second fully connected layer compresses the dimension back to 50*128. The LSTM decoder contains 256 units per layer. After passing through the LSTM decoder, the dimension is expanded to 50*256. Finally, after passing through the fully connected network layer, the dimension is mapped to 50*2. The final output reconstructs the time window W output only contains two features of position x and position y.
[0019] 2.2) Since the dimensions of the input normal time window W input and the output reconstructed time window W output are different, only the positions x and y of the six features in the input normal time window W input are used to calculate the reconstruction loss with the output reconstructed time window W output . Among them, the normal time window of the positions x and y of the six features in the input normal time window W input is denoted as with a dimension of 50*2; and the output reconstructed time window is directly used as the reconstructed window of the input normal time window, denoted as with a dimension of 50*2; and the target function L T of the teacher network based on the unsupervised learning training of the reconstruction is constructed.
[0020]
[0021] In the formula, MAE represents the mean absolute error, The smaller the MAE is, the closer the reconstructed time window is to the input normal time window. When training the teacher network, the parameters of the teacher network are optimized by minimizing the above target function L T to enhance the ability of the teacher network to reconstruct the normal time window.
[0022] Further, the step 3) includes the following steps:
[0023] 3.1) A lightweight student network with few parameters is built as a lightweight time window reconstructor, and the modules from low to high are: 1 layer of 2-head Transformer encoder, feedforward neural network layer, 1 layer of LSTM decoder, and fully connected network layer; wherein the input normal time window Winput including six features of sending time, pseudonym ID, position x, position y, speed x and speed y, and the size of the time window is 50, so W input The dimension of W is 50*6, and after the linear embedding layer in the Transformer encoder, it is mapped to a high-dimensional representation of 50*64, and then passes through the entire Transformer encoder, and the dimension is still 50*64; The feedforward neural network layer includes two fully connected layers, and LeakyReLU activation function is used between the two fully connected layers, the first fully connected layer expands the dimension to 50*256, and the second fully connected layer compresses the dimension back to 50*64; The LSTM decoder contains 256 units, and after passing through the LSTM decoder, the dimension is expanded to 50*256; Finally, after passing through the fully connected network layer, the dimension is mapped to 50*2, and the final output reconstructs the time window W output only contains two features of position x and position y;
[0024] 3.2) Since the dimensions of the input normal time window W input and the output reconstructed time window W output are different, only the position x and the position y of the six features in the input normal time window are used to calculate the reconstruction loss with the output reconstructed time window, wherein the normal time window of the position x and the position y of the six features in the input normal time window W input is denoted as with a dimension of 50*2; and the output reconstructed time window is directly used as the reconstructed window of the input normal time window, denoted as with a dimension of 50*2; the target function of the reconstruction-based unsupervised learning training of the student network
[0025]
[0026] wherein MAE represents the mean absolute error, The smaller the MAE is, the closer the reconstructed time window is to the input normal time window; when training the lightweight student network, the parameters of the lightweight student network are optimized by minimizing the above target function to enhance the ability of the lightweight student network to reconstruct the normal time window;
[0027] The lightweight student network not only needs to learn the ability to reconstruct the normal time window, but also needs to accept the knowledge transferred from the teacher network. First, a knowledge distillation method based on feature transfer is defined, which is called FKD. Three kinds of features are defined: the output of the input data after passing through the Transformer encoder and the feedforward neural network layer is defined as the encoding feature EF, the output of the data after passing through the LSTM decoder is defined as the decoding feature DF, and the output of the data after passing through the fully connected network layer is defined as the output feature OF. The objective function of the lightweight student network knowledge distillation method FKD is constructed
[0028]
[0029] In the formula, MSE is the mean square error, EF T represents the encoding feature of the teacher network, EF S represents the encoding feature of the lightweight student network, DF T represents the decoding feature of the teacher network, DF S represents the decoding feature of the lightweight student network, OF T represents the output feature of the teacher network, OF S represents the output feature of the lightweight student network; since the mean square error of three kinds of features is used, the coefficient
[0030] Finally, the objective function of the lightweight student network based on unsupervised learning training of reconstruction is combined with the objective function of the knowledge distillation method FKD in proportion to construct the total objective function L of the lightweight student network training S :
[0031]
[0032] In the formula, alpha and beta are hyperparameters, used to adjust the weight between the objective function and the objective function , and alpha + beta = 1; by minimizing the above total objective function L S , the lightweight student network can learn to reconstruct the normal time window while accepting the feature knowledge from the output of the teacher network.
[0033] Further, in step 4), the roadside unit processes the time series sent by each vehicle in the coverage range to convert it into a time window that can be input into the lightweight student network, and reconstructs the time window through the lightweight student network, taking the reconstruction error of the time window as its anomaly score
[0034]
[0035] MAE represents the average absolute error, represents the time window to be measured; represents After processing, only the time window of the two features of position x and position y is retained; represents the reconstructed time window containing only the two features of position x and position y after the input data is reconstructed by the lightweight student network; The normal time window conforms to the distribution of the normal data used when training the lightweight student network, so it can be well reconstructed, and the abnormal score is low; And the abnormal time window does not conform to the distribution of the normal data used when training the lightweight student network, so it will deviate from the original data after reconstruction, resulting in a large abnormal score;
[0036] Finally, by comparing the abnormal score with the preset threshold value theta, if the abnormal score of the time window exceeds the preset threshold value theta, it is determined to be abnormal, otherwise it is determined to be normal, finally, the roadside unit will not forward the time window determined to be abnormal.
[0037] Compared with the prior art, the present application has the following advantages and beneficial effects:
[0038] 1、The present application adopts a reconstruction-based unsupervised learning method to train the teacher network and the lightweight student network, solving the problem of difficult labeling and high labeling cost of abnormal data in the supervised learning method. At the same time, after training using the reconstruction-based unsupervised learning method, unknown abnormalities can be detected, and the generalization ability for new and unseen abnormalities is achieved, solving the problem that traditional vehicle networking anomaly detection methods cannot detect unknown abnormalities.
[0039] 2、The present application proposes a feature-based knowledge distillation method FKD, which enhances the ability of the lightweight student network to reconstruct normal time windows by transferring the encoder features, decoder features and output features when the input data passes through the teacher network, solving the problem that traditional knowledge distillation methods are not suitable for reconstruction-based unsupervised learning.
[0040] 3、The present application combines the reconstruction-based unsupervised learning method with the feature-based knowledge distillation method FKD, and deploys the optimized lightweight student network to the roadside unit of the vehicle networking for anomaly detection task. The lightweight student network has the ability to detect unknown abnormalities, and at the same time, due to the reduction of the number of neural network layers or units and the reduction of the number of parameters, the prediction speed is improved, solving the technical problem that the deep neural network has large parameters and is difficult to be deployed to the resource-limited device in the vehicle networking, achieving the technical effect of ensuring detection accuracy while improving detection speed. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1Flowchart of the method of the present application.
[0042] Figure 2 Training schematic diagram of teacher network and lightweight student network.
[0043] Figure 3 Flowchart of time window anomaly detection.
[0044] Figure 4 Schematic diagram of lightweight student network calculating anomaly score. DETAILED DESCRIPTION
[0045] The present application will be further described in conjunction with the embodiments and the accompanying drawings, but the embodiments of the present application are not limited thereto.
[0046] The experimental platform of the present embodiment is Python 3.9, Pytorch 1.13.1, and the computer configuration is: the CPU model is Intel(R) Xeon(R) Gold 6230, the memory is 12 GB, and the graphics card model is NVIDIA GeForce RTX 2080Ti.
[0047] As shown in Figures 1 to 4 The present embodiment discloses a lightweight method for vehicle networking anomaly detection based on deep learning and knowledge distillation, which is roughly divided into four stages: the first stage is data set acquisition and data preprocessing, mainly downloading public data set, using sliding window method to generate vehicle time window, selecting the required features, and performing standardization operation; the second stage is the construction and training of teacher network, mainly constructing a teacher network with large model capacity, and performing unsupervised learning training based on reconstruction; the third stage is the construction and training of lightweight student network, mainly constructing lightweight student network, and migrating the knowledge of teacher network to lightweight student network by using knowledge distillation; the fourth stage is the deployment of model, mainly deploying the trained lightweight student network to the roadside unit for anomaly detection. The specific steps are as follows:
[0048] 1) Using the existing vehicle networking data set VeReMi Extension, the time series of different vehicles are preprocessed to generate time windows, and the data set is divided into training set and test set, which includes the following steps:
[0049] 1.1) Obtain the public vehicle networking data set VeReMi Extension, and perform data labeling work, the messages in the data set that are not changed due to noise or attack are marked as normal messages, otherwise they are marked as abnormal messages;
[0050] 1.2) In the data set, the content carried by a basic safety message of vehicle networking includes: sending time, pseudonym ID, position x, position y, speed x, speed y, acceleration x, acceleration y, direction x and direction y. The following preprocessing is performed on the above-mentioned content: only sending time, pseudonym ID, position x, position y, speed x and speed y are selected as input features for each message, z-score standardization strategy is used for all basic safety messages, and sliding window strategy is used. A sliding window with a window size of 50 and a step size of 10 is used to intercept a time window. If the time window contains an abnormal message, it is marked as an abnormal time window, otherwise it is a normal time window;
[0051] 1.3) The preprocessed data set is divided into a training set and a test set, wherein 90% of the normal time windows are used to construct the training set, and the remaining normal time windows and abnormal time windows are constructed in proportion to the original data set to construct the test set.
[0052] 2) A teacher network based on Transformer and LSTM is constructed, wherein the Transformer is used as an encoder to capture the dependency between each position in the input time window and obtain a compressed representation of the time window information, and the LSTM is used as a decoder to effectively maintain long-distance dependence through a memory unit and a gating mechanism. The compressed information of the Transformer is recovered into the input data through decoding; the teacher network uses a fully connected neural network as a linear embedding layer to replace the embedding layer in the original Transformer, maps discrete input message features to a continuous vector space, and expands the feature dimension of the input at the same time, so as to better capture the hidden relationship between the features in the subsequent process; the teacher network only uses the encoder of the Transformer to extract important information in the time window, and does not need the decoder to generate sequence output, so the decoder in the original Transformer is removed; the teacher network introduces a feedforward neural network layer after the Transformer and a fully connected network layer after the LSTM. The feedforward neural network layer after the Transformer integrates the features learned by the encoder and performs nonlinear transformation through layer-by-layer transmission to map high-dimensional data to low-dimensional representation. The fully connected network layer after the LSTM maps the decoded output to the dimension of the expected output; a large number of normal time windows in the training set are used to train the teacher network using an unsupervised learning method based on reconstruction, so that the teacher network learns the distribution of the normal time window and can restore the normal time window, but cannot restore the abnormal time window that has not been seen during training, thereby achieving the purpose of distinguishing the abnormal time window; which includes the following steps:
[0053] 2.1) The teacher network with large model capacity is built as a time window reconstructor, and its modules from low to high level are: 1 layer of 16 heads of Transformer encoder, feedforward neural network layer, 4 layers of LSTM decoder and fully connected network layer; wherein, the input normal time window W input includes six features of sending time, pseudonym ID, position x, position y, speed x and speed y, and the size of the time window is 50, so W input has a dimension of 50*6. After passing through the linear embedding layer in the Transformer encoder, it is mapped to a high-dimensional representation of 50*128, and then passes through the entire Transformer encoder, and the dimension is still 50*128. The feedforward neural network layer includes two fully connected layers, and LeakyReLU activation function is used between the two fully connected layers. The first fully connected layer expands the dimension to 50*512, and the second fully connected layer compresses the dimension back to 50*128. The LSTM decoder contains 256 units per layer, and after passing through the LSTM decoder, the dimension is expanded to 50*256. Finally, after passing through the fully connected network layer, the dimension is mapped to 50*2, and the final output normal time window W output only contains two features of position x and position y;
[0054] 2.2) Since the dimensions of the input normal time window W input and the output normal time window W output are different, only the positions x and y of the six features in the input normal time window W input are used to calculate the reconstruction loss with the output normal time window W output , wherein, in the input normal time window W input , only the normal time window of the positions x and y of the six features is taken as with a dimension of 50*2; and the output reconstruction time window is directly used as the reconstruction window of the input normal time window, recorded as with a dimension of 50*2; the target function L T of the reconstruction-based unsupervised learning training of the teacher network is constructed:
[0055]
[0056] wherein, MAE represents the mean absolute error, the smaller the value is, the closer the reconstructed normal time window is to the input normal time window; when training the teacher network, the parameters of the teacher network are optimized by minimizing the above target function L T , so as to enhance the ability of the teacher network to reconstruct the normal time window.
[0057] 3) Construct a lightweight student network similar to the teacher network structure, but with reduced number of layers, units and parameters, and use the trained teacher network to guide the lightweight student network to implement knowledge distillation on the normal time window in the training set, increase the ability of the lightweight student network to reconstruct the normal time window, and obtain an optimized lightweight student network; which includes the following steps:
[0058] 3.1) Build a lightweight student network with less parameters as a lightweight time window reconstructor, which from low to high level modules are: 1 layer 2 head Transformer encoder, feedforward neural network layer, 1 layer LSTM decoder and fully connected network layer; wherein, the input normal time window W input includes six features of sending time, pseudonym ID, position x, position y, speed x and speed y, and the time window size is 50, so W input The dimension is 50*6, after the linear embedding layer in the Transformer encoder, it is mapped to a high-dimensional representation of 50*64, and then passes through the entire Transformer encoder, the dimension is still 50*64; The feedforward neural network layer includes two fully connected layers, and LeakyReLU activation function is used between the two fully connected layers, the first fully connected layer expands the dimension to 50*256, and the second fully connected layer compresses the dimension back to 50*64; The LSTM decoder contains 256 units, and after passing through the LSTM decoder, the dimension is expanded to 50*256; Finally, after passing through the fully connected network layer, the dimension is mapped to 50*2, and the final output normal time window W output only contains two features of position x and position y;
[0059] 3.2) Because the dimensions of the input normal time window W input and the output normal time window W output are different, only the position x and position y in the six features of the input normal time window are used to calculate the reconstruction loss with the output normal time window, wherein, in the input normal time window W input , only the normal time window of the position x and position y in the six features is taken as The dimension is 50*2; while the output reconstructed time window is directly used as the reconstructed window of the input normal time window, denoted as The dimension is 50*2; the target function of the unsupervised learning training of the student network based on reconstruction is
[0060]
[0061] In the formula, MAE represents the mean absolute error, The smaller, the closer the reconstructed normal time window is to the input normal time window; in training the lightweight student network, the parameters of the lightweight student network are optimized by minimizing the above objective function to enhance the ability of the lightweight student network to reconstruct the normal time window;
[0062] The lightweight student network not only needs to learn the ability to reconstruct the normal time window, but also needs to accept the knowledge transferred from the teacher network. First, a feature transfer-based knowledge distillation method, called FKD, is defined. Three features are defined: the output of the input data after passing through the Transformer encoder and the feedforward neural network layer is defined as the encoding feature EF, the output of the data after passing through the LSTM decoder is defined as the decoding feature DF, and the output of the data after passing through the fully connected network layer is defined as the output feature OF. The objective function of the lightweight student network knowledge distillation method FKD is constructed
[0063]
[0064] In the formula, MSE is the mean square error, EF T represents the encoding feature of the teacher network, EF S represents the encoding feature of the lightweight student network, DF T represents the decoding feature of the teacher network, DF S represents the decoding feature of the lightweight student network, OF T represents the output feature of the teacher network, OF S represents the output feature of the lightweight student network; since the mean square error of three features is used, the coefficient
[0065] Finally, the objective function of the unsupervised learning training based on reconstruction in the lightweight student network is combined with the objective function of the knowledge distillation method FKD in proportion to construct the total objective function L of the lightweight student network training S :
[0066]
[0067] In the formula, α and β are hyperparameters for adjusting the weights between the objective function and the objective function , and α+β1; by minimizing the above total objective function L S , the lightweight student network can learn to reconstruct the normal time window while accepting the feature knowledge from the output of the teacher network.
[0068] The trained teacher network and the lightweight student network are tested on the test set multiple times, and the best threshold is selected, and the experimental results of the experiment are described in detail as follows:
[0069] According to the final detection result of the network, the accuracy and speed of the network are evaluated from the accuracy, precision, recall rate and F1 score indicators, and the results are shown in Table 1.
[0070] Table 1 Comparison of network performance
[0071]
[0072]
[0073] The above table results show that the comprehensive performance of the teacher network and the lightweight student network in the application is better than that of the previous vehicle networking anomaly detection model, and high detection performance is achieved.
[0074] The parameter quantity and prediction speed of the teacher network, the lightweight student network and the previous model are compared as shown in Table 2.
[0075] Table 2 Comparison of parameter quantity and prediction speed
[0076]
[0077] The above table results show that after knowledge distillation, the parameter quantity of the student network is reduced, the reasoning time is significantly reduced, and the reasoning speed is improved.
[0078] 4) The trained lightweight student network is tested using the test set, and the threshold with the highest detection accuracy on the test set is selected as the preset threshold θ, and the lightweight student network is deployed together with the roadside unit in the vehicle networking for real-time anomaly detection.
[0079] The roadside unit creates a time window with a window size of 50 and a step of 10 for each message sent by a vehicle, and only retains six fields of sending time, pseudonym ID, position x, position y, speed x and speed y. Finally, the time window is subjected to z-score standardization operation.
[0080] As shown in Figure 4 , the test time window is input into the lightweight student network, and the anomaly score of the test time window is calculated after the reconstruction of the lightweight student network. The reconstruction error of the time window is taken as the anomaly score of the time window
[0081]
[0082] In the formula, MAE represents the average absolute error, represents the test time window; represents After processing, only the time window of the two features of position x and position y is reserved; The reconstructed time window of the input data after being reconstructed by the lightweight student network, and the output only contains the two features of position x and position y; the normal time window conforms to the distribution of the normal data used when training the lightweight student network, and can be well reconstructed, so the abnormal score is low; and the abnormal time window does not conform to the distribution of the normal data used when training the lightweight student network, and deviates from the original data after reconstruction, resulting in a larger abnormal score;
[0083] The abnormal score is compared with the preset threshold θ, if the abnormal score of the time window exceeds the preset threshold θ, it is determined to be abnormal, otherwise it is determined to be normal, finally, the roadside unit performs non-forwarding processing on the time window determined to be abnormal.
[0084] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments, any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application should be equivalent replacement methods, and are all included in the protection scope of the present application.
Claims
1. A lightweight method for anomaly detection in Internet of Vehicles based on deep learning and knowledge distillation, characterized in that, The method comprises the following steps: 1) using the disclosed VeReMi Extension dataset, preprocessing the time series messages of each vehicle in the dataset using a sliding window strategy to obtain time windows, and dividing the preprocessed dataset into a training set and a test set, wherein the training set only contains normal time windows, and the test set contains normal time windows and abnormal time windows; 2) constructing a teacher network based on Transformer and LSTM, wherein the Transformer is used as an encoder to capture the dependency between each position in the input time window and obtain a compressed representation of the time window information, and the LSTM is used as a decoder to effectively maintain long-distance dependence through a memory unit and a gating mechanism, and the compressed information of the Transformer encoding is recovered into the input data through decoding; the teacher network uses a fully connected neural network as a linear embedding layer to replace the embedding layer in the original Transformer, maps discrete input message features to a continuous vector space, and expands the feature dimension of the input to facilitate better capture of hidden relationships between features; the teacher network only uses the encoder of the Transformer to extract important information in the time window, and does not need the decoder to generate sequence output, so the decoder in the original Transformer is removed; the teacher network introduces a feedforward neural network layer after the Transformer and a fully connected network layer after the LSTM, the feedforward neural network layer after the Transformer integrates the features learned by the encoder and performs nonlinear transformation through layer-by-layer transmission to map high-dimensional data to low-dimensional representation, and the fully connected network layer after the LSTM maps the decoded output to the dimension of the expected output; using a large number of normal time windows in the training set, an unsupervised learning method based on reconstruction is used to train the teacher network, so that the teacher network learns the distribution of the normal time window and can restore the normal time window, but cannot restore the abnormal time window that is not seen during training, thereby achieving the purpose of distinguishing abnormal time windows; 3) constructing a lightweight student network similar in structure to the teacher network, but with reduced number of layers, number of units and parameter quantity, using the trained teacher network to guide the lightweight student network to implement knowledge distillation on the normal time windows in the training set, increasing the ability of the lightweight student network to reconstruct the normal time window, and obtaining a trained lightweight student network; 4) using the trained lightweight student network to test the test set, selecting the threshold with the highest detection accuracy on the test set as the preset threshold, and then deploying the preset threshold and the trained lightweight student network to the roadside unit in the Internet of Vehicles for anomaly detection tasks, reconstructing the time window by the lightweight student network, taking the reconstruction error of the time window as the anomaly score, and finally comparing the anomaly score with the preset threshold, if the anomaly score exceeds the preset threshold, it is determined as abnormal, otherwise it is determined as normal.
2. The deep learning and knowledge distillation based lightweight anomaly detection method for Internet of Vehicles according to claim 1, characterized in that, The step 1) comprises the following steps: 1.1) Obtain the public vehicle networking dataset VeReMi Extension and perform data annotation work, mark the messages in the dataset that are not changed due to noise or attack as normal messages, otherwise mark them as abnormal messages; 1.2) In the dataset, a basic safety message of vehicle networking carries the following contents: sending time, pseudonym ID, position x, position y, speed x, speed y, acceleration x, acceleration y, direction x and direction y, the above contents are preprocessed as follows: only select sending time, pseudonym ID, position x, position y, speed x and speed y as input features for each message, for all basic safety messages, z-score standardization strategy is adopted, and a sliding window strategy is used, a sliding window with a window size of 50 and a step size of 10 is used to intercept a time window, if the time window contains an abnormal message, it is marked as an abnormal time window, otherwise it is a normal time window; 1.3) Divide the preprocessed dataset into a training set and a test set, wherein 90% of the normal time windows are used to construct the training set, and the remaining normal time windows and abnormal time windows are constructed according to the proportion of the original dataset.
3. The deep learning and knowledge distillation based lightweight anomaly detection method for Internet of Vehicles according to claim 1, characterized in that, In step 2), the teacher network is constructed and trained, including the following steps: 2.1) The teacher network with large capacity is built as the time window reconstructor, and its modules from low to high level are: 1 layer of 16 heads of Transformer encoder, feedforward neural network layer, 4 layers of LSTM decoder and fully connected network layer; wherein, the input normal time window W input includes six features of sending time, pseudonym ID, position x, position y, speed x and speed y, and the size of the time window is 50, so the dimension of W input is 50*6. After the linear embedding layer in the Transformer encoder, it is mapped to a high-dimensional representation of 50*128, and then the dimension is still 50*128 after the entire Transformer encoder. The feedforward neural network layer includes two fully connected layers, and LeakyReLU activation function is used between the two fully connected layers. The first fully connected layer expands the dimension to 50*512, and the second fully connected layer compresses the dimension back to 50*128. Each layer of the LSTM decoder contains 256 units, and the dimension is expanded to 50*256 after the LSTM decoder. Finally, the dimension is mapped to 50*2 through the fully connected network layer, and the final output reconstructed time window W output only contains two features of position x and position y. 2.2) Since the dimensions of the input normal time window W input and the output reconstruction time window W output are different, only the position x and the position y of the six features in the input normal time window W input are used to calculate the reconstruction loss with the output reconstruction time window W output , wherein the normal time window of the position x and the position y of the six features in the input normal time window W input is denoted as with the dimension of 50*2; and the output reconstruction time window is directly taken as the reconstruction window of the input normal time window, denoted as with the dimension of 50*2; and the target function L T of the teacher network based on the unsupervised learning training of the reconstruction is constructed. In the formula, MAE represents the mean absolute error, The smaller the MAE is, the closer the reconstructed time window is to the normal input time window; in training the teacher network, the parameters of the teacher network are optimized by minimizing the objective function L T to enhance the ability of the teacher network to reconstruct the normal time window.
4. The deep learning and knowledge distillation based lightweight anomaly detection method for Internet of Vehicles according to claim 1, characterized in that, The step 3) includes the following steps: 3.1) A lightweight student network with a small number of parameters is built as a lightweight time window reconstructor, whose modules from low to high are: 1 layer 2-head Transformer encoder, feedforward neural network layer, 1 layer LSTM decoder and fully connected network layer; among them, the input normal time window W input includes six features: sending time, pseudonym ID, position x, position y, speed x and speed y, and the time window size is 50, so W input The dimension is 50*6, which is mapped to a high-dimensional representation of 50*64 after the linear embedding layer in the Transformer encoder, and the dimension is still 50*64 after the entire Transformer encoder. The feedforward neural network layer includes two fully connected layers, and LeakyReLU activation function is used between the two fully connected layers. The first fully connected layer expands the dimension to 50*256, and the second fully connected layer compresses the dimension back to 50*64. The LSTM decoder contains 256 units, and the dimension is expanded to 50*256 after the LSTM decoder. Finally, the dimension is mapped to 50*2 through the fully connected network layer, and the final output reconstructed time window W output only contains two features: position x and position y. 3.2) Since the dimensions of the input normal time window W input and the output reconstruction time window W output are different, only the position x and the position y of the six features in the input normal time window are used to calculate the reconstruction loss with the output reconstruction time window, wherein, in the input normal time window W input , only the normal time window of the position x and the position y of the six features is taken as with the dimension of 50*2; and the output reconstruction time window is directly taken as the reconstruction window of the input normal time window, recorded as with the dimension of 50*2; the target function of the reconstruction-based unsupervised learning training of the student network In the formula, MAE represents the mean absolute error, The smaller the MAE is, the closer the reconstruction time window is to the input normal time window; in training the lightweight student network, the parameters of the lightweight student network are optimized by minimizing the objective function to enhance the ability of the lightweight student network to reconstruct the normal time window. The lightweight student network not only needs to learn the ability to reconfigure the normal time window, but also needs to accept the knowledge transferred from the teacher network. First, a knowledge distillation method based on feature transfer is defined, which is called FKD. Three features are defined: the output of the input data after passing through the Transformer encoder and the feedforward neural network layer is defined as the encoding feature EF, the output of the data after passing through the LSTM decoder is defined as the decoding feature DF, and the output of the data after passing through the fully connected network layer is defined as the output feature OF. The objective function of the lightweight student network knowledge distillation method FKD is constructed where MSE is the mean square error, EF T denotes the encoding features of the teacher network, EF S denotes the encoding features of the lightweight student network, DF T denotes the decoding features of the teacher network, DF S denotes the decoding features of the lightweight student network, OF T denotes the output features of the teacher network, OF S denotes the output features of the lightweight student network; since the mean square error of three features is used, the preceding is multiplied by a factor Finally, the target function of the unsupervised learning training based on reconstruction in the lightweight student network The target function of the knowledge distillation method FKD The total target function L of the lightweight student network training is constructed by proportional combination S : where a and β are hyperparameters that adjust the weights between the objective functions and the objective function , and a + β = 1; by minimizing the total objective function L S , the lightweight student network can learn to reconstruct normal time windows while accepting feature knowledge from the teacher network output.
5. The deep learning and knowledge distillation based lightweight anomaly detection method for Internet of Vehicles according to claim 1, characterized in that, In step 4), the roadside unit processes the time series sent by each vehicle in the coverage area into time windows that can be input into the lightweight student network, reconstructs the time windows through the lightweight student network, and takes the reconstruction error of the time windows as the anomaly score of the vehicle In the formula, MAE represents the mean absolute error, represents the time window to be measured; represents After processing, only the time window of the two features of position x and position y is retained; represents the reconstructed time window containing only the two features of position x and position y after the input data is reconstructed by the lightweight student network; the normal time window conforms to the distribution of the normal data used when training the lightweight student network, and can be well reconstructed, so the abnormal score is low; and the abnormal time window does not conform to the distribution of the normal data used when training the lightweight student network, and deviates from the original data after reconstruction, resulting in a large abnormal score; Finally, by comparing the abnormal score with the preset threshold θ, if the abnormal score of the time window exceeds the preset threshold θ, it is determined to be abnormal, otherwise it is determined to be normal, finally, the roadside unit performs non-forwarding processing on the time window determined to be abnormal.
Citation Information
Patent Citations
Multilayer neural network language model training method and device based on knowledge distillation
CN111611377A
Real-time communication system and method suitable for non-stationary network environment
CN117749775A