Lightweight vehicle-mounted network intrusion detection method based on BERT
By dynamically adjusting attention weight design, the VehicleBERT series model solves the problem of resource limitation in vehicle network intrusion detection, and realizes efficient and lightweight detection, which is suitable for vehicle environments.
Patent Information
- Application Number
- CN202510432400.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
The existing in-vehicle network intrusion detection technology based on deep learning is difficult to balance detection accuracy and resource occupation in resource-constrained environments. The BERT model parameters are too large and the computing resource requirements are high, which cannot meet the real-time requirements.
By adjusting dynamic attention weights, VehicleBERT series models are designed, including VehicleBERT and VehicleBERT-DA models, which are lightweight, reduce computing and storage requirements, while maintaining high detection accuracy.
In vehicle network intrusion detection, VehicleBERT and VehicleBERT-DA models significantly reduce computing overhead and resource usage, adapt to the real-time requirements of vehicle networks, while maintaining high detection accuracy.
Smart Images

Figure CN120358052A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of in-vehicle network intrusion detection, and relates to a lightweight model based on BERT (Bidirectional Encoder Representations from Transformers), named VehicleBERT model, which can perform intrusion detection on in-vehicle networks (IVN networks). Background Art
[0002] Deep learning technologies have been widely applied in fields such as natural language processing and image recognition. These technologies also form more abstract and non-linear high-level representations by combining low-level features, and then mine the features and the input-output relationships between data, and have achieved good results in the field of intrusion detection. In addition, in recent years, most in-vehicle CAN bus intrusion detection technologies have been developed based on CNN and RNN networks and have made remarkable progress. However, deep learning-based methods, such as models combining CNN and RNN, although excellent in feature extraction and sequence modeling, have problems such as high model complexity, slow inference speed, and inability to meet real-time requirements in resource-constrained in-vehicle devices. Especially, the BERT model performs excellently in feature extraction, but its parameter scale is too large to be directly applied to in-vehicle environments.
[0003] In paper [1], BERT was first proposed. Existing natural language processing technologies, such as models like ELMo and GPT, although having made breakthroughs in context understanding, still have problems of insufficient modeling of long-distance dependencies and inability to fully capture bidirectional context information. BERT has achieved remarkable improvements in context semantic understanding by introducing a deep bidirectional Transformer architecture and combining two pre-training tasks of Masked Language Model and NextSentence Prediction. Represented by BERT-Base (110 million parameters) and BERT-Large (340 million parameters), the BERT model performs excellently in various NLP tasks. However, its large parameter scale and high computational resource requirements make it difficult to be directly applied in scenarios with resource constraints and high real-time requirements such as in-vehicle network intrusion detection, exposing the deficiencies of existing technologies in terms of lightweight, efficiency, and adaptability.
[0004] In the paper [2], a vehicle network intrusion detection system CAN-BERT based on the BERT language model was proposed, aiming to detect abnormal behaviors and network attacks by deeply learning the CAN bus message sequence. Experimental results show that CAN-BERT outperforms the existing technologies in terms of real-time detection ability and high accuracy. However, the paper did not elaborate on how to achieve model lightweighting, nor did it consider the resource constraints in the vehicle network environment. Therefore, this method may face high computational overhead in practical applications, limiting its effectiveness and application prospects in resource-constrained scenarios.
[0005] Existing technologies generally face the problem of being difficult to balance detection accuracy and resource occupancy, and there is an urgent need for a lightweight and efficient vehicle network intrusion detection solution.
[0006] [1] Kenton, J.D.M.W.C. and Toutanova, L.K., 2019, June. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT (Vol. 1, p. 2).
[0007] [2] Alkhatib, N., Mushtaq, M., Ghauch, H. and Danger, J.L., 2022, December. Can-bert do it? controller area network intrusion detection system based on bert language model. In 2022 IEEE / ACS 19th International Conference on Computer Systems and Applications (AICCSA) (pp. 1-8). IEEE. Summary of the Invention
[0008] In view of the above problems, the purpose of the present invention is to provide a lightweight vehicle network intrusion detection method based on BERT, which is used to balance the loss brought by lightweighting by adjusting dynamic attention weights -> finally design the VehicleBERT series of models to be applicable to vehicle network intrusion detection and overcome the problem that existing deep learning-based vehicle network intrusion detection technologies generally face difficulties in balancing detection accuracy and resource occupancy.
[0009] A lightweight vehicle network intrusion detection method based on BERT provided by the present invention specifically includes the following steps: S1. In the data preprocessing stage, convert the labeled CAN samples into text samples, including: loading data, label mapping, feature column processing, feature splicing, dataset division, word segmentation, and encoding; S2. In the model training stage, input the labeled text samples into the VehicleBERT series models for training, including: setting training hyperparameters and passing the processed data to the VehicleBERT series models for training; S3. In the model test and evaluation stage, conduct inference tests on the trained VehicleBERT series models.
[0010] As an optimization of the present invention, the following steps are further included in step S1: S101. Load data. The data used is the Car-hacking public dataset, which is generated by transmitting CAN data packets to the CAN bus of a real vehicle and includes CAN data packets collected from the OBD-II port; S102. Label mapping. Create a label mapping dictionary to map the original labels to digital labels: 'R' → 0, 'RPM' → 1, 'gear' → 2, 'DoS' → 3, 'Fuzzy' → 4. Check the unique values in the Label column to ensure that all labels are within the mapping range, replace the values in the Label column with digital labels, and check whether there are null values in the Label column after label replacement; S103. Feature column processing. Check whether there are null values in the CAN ID and the data columns DATA[0]-DATA[7], and delete the rows with null values in the CAN ID and Label columns; S104. Feature splicing. Splice the CAN ID and the data columns DATA[0]-DATA[7] into text format, check whether there are null values in the spliced text column, and ensure that the number of spliced text and labels is consistent; S105. Dataset division. Divide the training set, validation set, and test set according to the ratio of 7:2:1 and save them; S106. Word segmentation and encoding. Before the VehicleBERT series models start training and testing, use the BERT tokenizer to convert the text into the format required for input to the VehicleBERT series models.
[0011] As an optimization of the present invention, the following steps are further included in step S106: (1)Tokenization breaks a text into smaller units. The tokenization method used by BERT is WordPiece, which breaks words into the smallest lexical units. The goal of this process is to handle out-of-vocabulary words and reduce the size of the vocabulary. The formula is as follows: ; Among them, represents the original text, represents the tokenizer, represents the result after tokenization; (2)Add special tokens. [CLS] and [SEP] are added at the beginning and end of the tokenization result respectively. The formula is as follows: ; Among them, represents the flag for the classification task. The output result of the VehicleBERT series of models is based on the representation at this position, represents the separator, which is used to distinguish sentence pairs or the end of the text, represents each token unit obtained by the above tokenizer, represents the final tokenization result after adding special tokens; (3)After tokenization, the next step is encoding, which converts the tokenized text into a numerical format that the VehicleBERT series of models can understand. BERT uses two main inputs: input_ids: the integer ID corresponding to each token. The formula is as follows: ; Among them, represents the integer representation of the token unit, which is a list where each integer corresponds to the position of a token unit in the vocabulary, represents the th token unit, which is a word or subword obtained from the tokenization result, represents the function that maps the token unit to the index in the vocabulary, represents the flag for the classification task. The output result of the VehicleBERT series of models is based on the representation at this position, represents the separator, which is used to distinguish sentence pairs or the end of the text; is used to indicate which positions of the VehicleBERT series of models need to be attended to and which positions to ignore. The purpose is to help the VehicleBERT series of models handle variable-length inputs and calculate the padding for the padded parts. The formula is as follows: ; Among them, represents the attention mask, which is a list with the same length as Consistent, used to indicate which positions are valid text and which are padding, Indicates the th mask value, with a value of 1 or 0, where, if is valid text , = 1, if is padding , then = 0.
[0012] As a preference of the present invention, the following steps are further included in step S2: S201. Set training hyperparameters. In the VehicleBERT series of models, AdamW is used as the optimizer. AdamW introduces weight decay and improves the defect of Adam in regularization by directly reducing the value of the parameters during the gradient update stage; S202. Pass the processed data to the VehicleBERT series of models for training. The VehicleBERT series of models is based on BERT. The VehicleBERT-DA model and the VehicleBERT model in the VehicleBERT series of models include an embedding layer, an encoder layer, a pooling layer, and an output layer; Embedding layer: Convert the input vocabulary index into a word vector representation; Encoder layer: Extract the context features of the sequence through self-attention and a feed-forward network; Pooling layer: Extract global features from the encoder output; Output layer: Map the features of the pooling layer to the category space for classification prediction.
[0013] As a preference of the present invention, the following steps are further included in step S202: 1.1. Word embedding calculation. The role of word embedding is to map each discrete token ID to a high-dimensional vector space; The formula is as follows: ; Where, is the word embedding matrix, is the embedding vector of the th token, is each discrete token ID; 1.2. Position embedding calculation. To introduce the position information of each token in the sequence, the embedding layer adds a position embedding. The position embedding is a fixed sine embedding or a learnable embedding. Here, a learnable embedding is selected. The formula is as follows: ; Where, is the position of the token in the sequence, is the learnable position embedding matrix, is the positional embedding vector of the th token; 1.3. Paragraph Embedding Calculation. If the VehicleBERT series model is used for sentence pair tasks, it is necessary to distinguish tokens belonging to different sentences, which is achieved through paragraph embedding. The formula is as follows: ; where is the paragraph embedding matrix, is the type of the sentence to which the th token belongs. Since the vehicle network intrusion detection task is single-sentence text classification, the VehicleBERT series model will default to setting it to all 0s; 1.4. Combination of Embeddings. Add the word embedding, positional embedding, and paragraph embedding element by element. The formula is as follows: ; where is the output of the embedding layer, containing word information , positional information , and sentence type information .
[0014] As a preference of the present invention, the following steps are further included in step S202: 2.1. Multi-Head Self-Attention Mechanism. Calculate the query (Q), key (K), and value (V). The formula is as follows: ; where is the input matrix, which generates the query matrix , key matrix , and value matrix through three linear transformations. , , are the learned weight matrices; Calculate the attention for each head. The formula is as follows: ; where is the dimension of the key, which is the same as the dimension of the query. is used to scale the dot product to prevent gradient explosion or disappearance. is the query matrix, is the transposed matrix of the key matrix, is the value matrix, is the attention result for each head; The dynamic attention heads are introduced here. In the standard multi-head attention mechanism, the weights of all heads are static. In the VehicleBERT-DA model, the dynamic attention heads will dynamically adjust the weights of each attention head according to the global features of each layer. The generation of the dynamic weights is as follows: ; where is the global feature, and are the weights and biases, respectively, used to adjust the generated dynamic weights. Through and the output range is restricted and normalized so that the weight of each head is between 0 and 1; Then, the dynamic weights are applied to the output of each head: ; where is the result of applying the dynamic weights to the output of each head, is the output of each head, is the dynamically generated weight; Concatenate the outputs of all heads. The weighted outputs of all heads are concatenated: ; where represents the result of concatenating the outputs of all heads, represents the concatenation function, represents the result of applying the dynamic weights to the output of each head. Finally, the final multi-head attention output is obtained through a linear transformation, is the output weight matrix, and its shape is where is the number of heads, is the dimension of each head, is the dimension of the VehicleBERT-DA model or the dimension of the hidden layer, representing the size of the feature vector of each layer in the VehicleBERT-DA model; 2.2 Feed-Forward Network. The feed-forward network consists of two linear layers. The output of the first linear layer is non-linearly transformed through an activation function, and the second linear layer maps the output back to the dimension of the VehicleBERT series model. The formula is as follows: ; where represents the input, , represents the weights and biases of the first linear layer, , Denote the weights and biases of the second linear layer, and finally perform lightweight design on it. By setting intermediate_size = 256, reduce the dimension of the intermediate layer of the feed-forward network. In the VehicleBERT series of models, compress intermediate_size to 256 to reduce the computational burden and memory consumption of the network; 2.3. Residual connection, which is used to solve the problem of gradient vanishing. The residual connection directly adds the input to the output for the flow of information. The formula is as follows: ; Among them, is the input, is the output after processing. For each layer, the residual connection will directly add the input to the output after transformation. In the VehicleBERT series of models, the residual connection is used to add the output of the multi-head attention to the input and add the output of the feed-forward network to the input; 2.4. Layer normalization, which is used to stabilize training. The output of each layer is within a set range to avoid gradient explosion or vanishing. The formula is as follows: ; Among them, represents the input, represents the mean of the input , represents the standard deviation of the input , represents the scaling factor, represents the bias term. The scaling factor γ is automatically initialized to 1, and the bias β is set to 0 and optimized during training. Layer normalization in the VehicleBERT series of models is mainly used to perform layer normalization on the results after the residual connection.
[0015] As a preference of the present invention, the following steps are further included in step S202: 3.1.1. Representation of the encoder layer output. The pooling layer receives the output of the encoder as input. This output is the hidden state from each token, which is a tensor. The output of the encoder layer represents the context information of each token in the input sequence and is the hidden state after being processed by the Transformer encoder; 3.1.2. Extract the output of the [CLS] token. Extract the hidden state corresponding to the first token as the representation of the entire input sequence. 3.1.3. Pooling output. Extract the representation of the [CLS] token from the output of the encoder as the final pooling output; 3.1.4. Lightweight design: The VehicleBERT series of models only focus on the [CLS] features after dynamic optimization. The VehicleBERT series of models complete the extraction of [CLS] features through a simple [:, 0, :] indexing operation, and the design of the dynamic adjustment module is mainly based on low-dimensional operations, supporting low-resource environments.
[0016] As a preference of the present invention, the following steps are further included in step S202: 4.1.1. Input the [CLS] features from the pooling layer; 4.1.2. Perform feature mapping through a linear layer to map the [CLS] features from the hidden_size dimension to the dimension of the target category; 4.1.3. Obtain the probability distribution of the final classification through linear layer calculation and the Softmax function; Lightweight design: The VehicleBERT series of models set the hidden layer dimension to 64, the number of parameters in the output layer is reduced, the size of the weight matrix of the linear layer is (64, 5), and the number of parameters is 320.
[0017] As a preference of the present invention, the following steps are further included in step S3: S301. Set the evaluation metrics to accuracy, precision, recall, and F1-score respectively, and their calculation formulas are as follows: ; ; ; ; Among them, represents the number of normal samples correctly predicted by the VehicleBERT series of models, represents the number of abnormal samples correctly predicted by the VehicleBERT series of models, represents the number of abnormal samples wrongly predicted as normal samples by the VehicleBERT series of models, represents the number of normal samples wrongly predicted as abnormal samples by the VehicleBERT series of models; S302. Test BERT-Base, VehicleBERT, and VehicleBERT-DA (where VehicleBERT-DA is the VehicleBERT with dynamic self-attention) using the same software and hardware environment; (1) Confusion matrix test, (2) Accuracy, precision, recall, and F1-score tests; (3)Testing of other lightweight metrics, where the lightweight metrics include: the number of parameters, the space size, the GPU occupancy rate, and the CPU occupancy rate; (4)Experimental tests of BERT - Base, VehicleBERT, and VehicleBERT - DA, and comparison results.
[0018] As a preference of the present invention, the VehicleBERT series of models refers to the VehicleBERT model and the VehicleBERT - DA model.
[0019] The beneficial effects of the present invention are as follows: The traditional intrusion detection model of the present invention is complex and inconvenient, and cannot be directly deployed in a resource - constrained vehicle - mounted network environment. At the same time, although the BERT model has good accuracy in vehicle - mounted network intrusion detection, its consumption of computing resources is large, especially in terms of GPU and CPU occupancy, and it has high requirements for hardware. This makes the application of BERT may be restricted in actual deployment, especially in resource - limited embedded systems or scenarios with low - latency requirements. The present invention conducts lightweight design based on the BERT model, and balances the loss caused by lightweight by adjusting the dynamic attention weights -> finally designs the VehicleBERT model to be applicable to vehicle - mounted network intrusion detection. For this reason, the lightweight models proposed by the present invention, such as VehicleBERT and VehicleBERT - DA, provide a better solution. While maintaining a high detection accuracy, they effectively reduce the computational overhead and resource occupancy, and can better meet the real - time requirements of vehicle - mounted networks. Description of the Drawings
[0020] Through the following description with reference to the drawings, and with a more comprehensive understanding of the present invention, other objects and results of the present invention will become clearer and easier to understand. In the drawings: Figure 1 is the overall flowchart in the present invention; Figure 2 is the diagram of the VehicleBERT series of models in the present invention; Figure 3 is the BERT - Base confusion matrix diagram in the present invention; Figure 4 is the VehicleBERT confusion matrix diagram in the present invention; Figure 5 is the VehicleBERT - DA confusion matrix diagram in the present invention. Detailed Embodiments
[0021] Refer to Figures 1-5, an embodiment of the present invention provides a lightweight in-vehicle network intrusion detection method based on BERT, including the following steps: S1. In the data preprocessing stage, convert the labeled CAN samples into text samples; S101. Load data. The data used is the Car-hacking public dataset, which is generated by transmitting CAN data packets to the CAN bus of a real vehicle and includes CAN data packets collected from the OBD-II port; S102. Label mapping. Create a label mapping dictionary to map the original labels to digital labels: 'R' → 0, 'RPM' → 1, 'gear' → 2, 'DoS' → 3, 'Fuzzy' → 4. Check the unique values in the Label column to ensure that all labels are within the mapping range. Replace the values in the Label column with digital labels. After replacing the labels, check if there are any null values in the Label column; S103. Feature column processing. Check if there are any null values in the CAN ID and data columns DATA[0]-DATA[7], and delete the rows with null values in the CAN ID and Label columns; S104. Feature concatenation. Concatenate the CAN ID and data columns DATA[0]-DATA[7] into text format. Check if there are any null values in the concatenated text column and ensure that the number of concatenated texts is the same as the number of labels; S105. Dataset division. Divide the training set, validation set, and test set in a ratio of 7:2:1 and save them; S106. Tokenization and encoding. Before starting the training and testing of the VehicleBERT series models, use the BERT tokenizer to convert the text into the format required for input to the VehicleBERT series models; (1) Tokenization is to break a piece of text into smaller units (words or sub-words). The tokenization method used by BERT is WordPiece. By breaking words into the smallest lexical units, the main goal of this process is to handle unknown vocabulary problems and reduce the size of the vocabulary. The formula is as follows: ; Among them, represents the original text, represents the tokenizer, represents the result after tokenization; (2) Add special Tokens. Add [CLS] and [SEP] at the beginning and end of the tokenization result respectively. The formula is as follows: ; Among them, The flag bit indicating the classification task. The output result of the VehicleBERT series of models is based on the representation at this position. Indicates the separator, used to distinguish sentence pairs or the end of the text. Indicates each token unit obtained by the above tokenizer. Indicates the final tokenization result after adding special tokens; (3) After tokenization is completed, the next step is encoding, which converts the tokenized text into a numerical format that the VehicleBERT series of models can understand. BERT uses two main inputs: input_ids: the integer ID corresponding to each token (indicating the position of the token in the vocabulary), and its formula is as follows: ; Where, Indicates the integer representation of the token unit, which is a list, and each integer in it corresponds to the position of a token unit in the vocabulary. Indicates the th token unit, which is a word or sub-word obtained from the tokenization result (Tokens). Indicates the index function that maps the token unit to the vocabulary. The flag bit indicating the classification task. The output result of the VehicleBERT series of models is based on the representation at this position. Indicates the separator, used to distinguish sentence pairs or the end of the text; Used to indicate which positions of the VehicleBERT series of models need to be attended to and which positions can be ignored. The purpose is to help the VehicleBERT series of models process variable-length inputs and avoid calculating the padding part. Its formula is as follows: ; Where, Indicates the attention mask, which is a list with the same length as , used to indicate which positions are valid text and which are padding. Indicates the th mask value, taking values of 1 or 0. Among them, if is valid text (non-padding), = 1, if is padding , then = 0.
[0022] S2. In the model training stage, the labeled text samples are input into the VehicleBERT series of models for training. S201. Set training hyperparameters. Hyperparameter Name Set Value Learning Rate 1e-4 Batch Size 32 Number of Training Epochs 3 Optimizer AdamW Loss Function CrossEntropyLoss As shown in the above table, it shows the setting of hyperparameters in the training process of the VehicleBERT series of models. In the VehicleBERT series of models, AdamW is used as the optimizer. AdamW introduces weight decay and improves the regularization defect of the original Adam by directly reducing the value of the parameters during the gradient update stage. S202. Pass the processed data to the VehicleBERT series of models for training. The VehicleBERT series of models is based on BERT. The VehicleBERT-DA model and the VehicleBERT model in the VehicleBERT series of models include an embedding layer, an encoder layer, a pooling layer, and an output layer. Embedding layer: Convert the input vocabulary index into a word vector representation. 1.1. Word embedding calculation. The role of word embedding is to map each discrete token ID to a high-dimensional vector space. The formula is as follows: ; Among them, is the word embedding matrix, is the th embedding vector of the token, is each discrete token ID. 1.2. Position embedding calculation. To introduce the position information of each token in the sequence, the embedding layer adds position embedding. The position embedding is a fixed sine embedding or a learnable embedding. Here, a learnable embedding is selected. The formula is as follows: ; Among them, is the position of the token in the sequence, is the learnable position embedding matrix, is the th position embedding vector of the token; 1.3. Paragraph embedding calculation. If the VehicleBERT series of models is used for sentence pair tasks (such as classification or question answering), it is necessary to distinguish tokens belonging to different sentences, which is achieved through paragraph embedding. The formula is as follows: ; Among them, is the paragraph embedding matrix, is the The type of the sentence to which a token belongs (0 indicates sentence A, and 1 indicates sentence B). Since the vehicle network intrusion detection task is single-sentence text classification, the VehicleBERT series of models will default to setting it all to 0; 1.4. Embedding combination: Add word embeddings, position embeddings, and segment embeddings element-wise. The formula is as follows: ; Among them, is the output of the embedding layer, containing word information , position information and sentence type information ; Encoder layer: Extract the context features of the sequence through self-attention and feed-forward networks.
[0023] 2.1. Multi-head self-attention mechanism: Calculate queries (Q), keys (K), and values (V). The formula is as follows: ; Among them, is the input matrix, which generates the query matrix , key matrix , value matrix , , , are the learned weight matrices; Calculate the attention for each head. The formula is as follows: ; Among them, is the dimension of the key, which is the same as the dimension of the query, is used to scale the dot product to prevent gradient explosion or vanishing, is the query matrix, is the transposed matrix of the key matrix, is the value matrix, is the attention result for each head; In particular, a dynamic attention head is introduced here. In the standard multi-head attention mechanism, the weights of all heads are static. In the VehicleBERT-DA model, the dynamic attention head dynamically adjusts the weights of each attention head according to the global features of each layer (the output from the [CLS] token). The generation of the dynamic weights is as follows: ; Among them, is the global feature (the output of the [CLS] token), and are the weights and biases, which are used to adjust the generated dynamic weights respectively. By and restrict the output range and normalize it to ensure that the weights of each head are between 0 and 1; Then, apply the dynamic weights to the output of each head: ; where is the result of applying the dynamic weights to the output of each head, is the output of each head, is the dynamically generated weight; Concatenate the outputs of all heads, and concatenate the weighted outputs of all heads: ; where represents the result of concatenating the outputs of all heads, represents the concatenation function, represents the result of applying the dynamic weights to the output of each head. Finally, obtain the final multi-head attention output through a linear transformation. is the output weight matrix, and its shape is where is the number of heads, is the dimension of each head, is the dimension of the VehicleBERT-DA model or the dimension of the hidden layer, representing the size of the feature vector of each layer in the VehicleBERT-DA model; Finally, perform a lightweight design on it. By reducing the configurations of num_attention_heads (the number of attention heads) and hidden_size (the dimension of the hidden layer), resource savings are achieved. BERT uses 12 heads and a hidden dimension of 768. In this VehicleBERT series of models, 2 heads and a hidden dimension of 64 are used to reduce the computational and memory overhead; 2.2. Feed-forward network. The feed-forward network consists of two linear layers. The output of the first linear layer is non-linearly transformed through an activation function (ReLU), and the second linear layer maps the output back to the dimension of the VehicleBERT series model. The formula is as follows: ; where, represents the input, , represents the weights and biases of the first linear layer, , Denote the weights and biases of the second linear layer, and finally perform lightweight design on it. By setting intermediate_size (the dimension of the intermediate layer) = 256, reduce the dimension of the intermediate layer of the feed-forward network. The intermediate layer dimension of the standard BERT's feed-forward network is 3072. In this VehicleBERT series of models, the intermediate_size is compressed to 256, thereby reducing the computational burden and memory consumption of the network; 2.3. Residual connection, which is used to avoid the problem of gradient vanishing. The residual connection adds the input directly to the output to ensure the flow of information. The formula is as follows: ; Among them, is the input, is the output after processing. For each layer, the residual connection will directly add the input to the output after transformation to ensure the flow of information and improve the stability of training. In this VehicleBERT series of models, the residual connection is mainly used for adding the output of the multi-head attention to the input and adding the output of the feed-forward network to the input; 2.4. Layer normalization, which is used to stabilize training, ensure that the output of each layer is within a set range, and avoid gradient explosion or vanishing. The formula is as follows: ; Among them, represents the input, represents the mean of the input , represents the standard deviation of the input , represents the scaling factor, represents the bias term. The scaling factor γ is automatically initialized to 1, and the bias β is set to 0 and optimized during training. Layer normalization in the VehicleBERT series of models is mainly used to perform layer normalization on the results after the residual connection; Finally, perform lightweight design on the entire encoder. The standard BERT model uses 12 layers of encoders, while in this VehicleBERT series of models, only 1 layer of encoder is used. By reducing the number of encoder layers, reduce the depth of the VehicleBERT series of models, reduce the training time and computational complexity, and also reduce the memory occupancy; 3.1. Pooling layer: Extract global features from the encoder output, 3.1.1. The representation output by the encoder layer. The pooling layer receives the output of the encoder as input. This output is the hidden state from each token, which is a tensor. The output of the encoder layer represents the context information of each token in the input sequence and is the hidden state after being processed by the Transformer encoder. 3.1.2. Extract the output of the [CLS] token. Extract the hidden state corresponding to the first token as the representation of the entire input sequence. The VehicleBERT series models of the present invention extract the [CLS] feature from the output after dynamic optimization. 3.1.3. Pooling output. Extract the representation of the [CLS] token from the output of the encoder as the final pooling output. 3.1.4. Lightweight design. Some models may further enhance the feature extraction ability by stacking multiple layers of pooling (such as max-pooling or average-pooling). The VehicleBERT series models of the present invention only focus on the [CLS] feature after dynamic optimization, avoid complex multi-layer processing, and maintain lightweight. Traditional pooling needs to cooperate with multi-head attention calculation, global feature extraction, and high-dimensional feature integration. The VehicleBERT series models complete the [CLS] feature extraction with a simple [:, 0, :] indexing operation, and the design of the dynamic adjustment module is mainly based on low-dimensional operations, supporting low-resource environments. 4.1. Output layer: Map the features of the pooling layer to the category space for classification prediction. 4.1.1. Input the [CLS] feature from the pooling layer. 4.1.2. Perform feature mapping through a linear layer, mapping the [CLS] feature from the hidden_size dimension to the dimension of the target category (i.e., the number of num_labels categories. Define num_labels = 5, so that the output dimension of the VehicleBERT series models changes from 2 categories of traditional BERT to 5 categories for use in in-vehicle network intrusion detection).
[0024] 4.1.3. Obtain the probability distribution of the final classification through linear layer calculation and the Softmax function. Lightweight design: The hidden_size (hidden layer dimension) of traditional BERT is relatively large (such as 768), while the hidden layer dimension of the VehicleBERT series models of the present invention is set to 64, reducing the number of parameters in the output layer. The size of the weight matrix of the linear layer is (64, 5), and the number of parameters is 320, which is significantly lower than that of traditional BERT. S3. In the model testing and evaluation phase, perform inference tests on the trained VehicleBERT series models; S301. Set the evaluation metrics to accuracy, precision, recall, and F1-score respectively, and their calculation formulas are as follows: ; ; ; ; Among them, represents the number of normal samples correctly predicted by the VehicleBERT series models, represents the number of abnormal samples correctly predicted by the VehicleBERT series models, represents the number of abnormal samples wrongly predicted as normal samples by the VehicleBERT series models, represents the number of normal samples wrongly predicted as abnormal samples by the VehicleBERT series models; S302. Test BERT-Base, VehicleBERT, and VehicleBERT-DA using the same software and hardware environment, where VehicleBERT-DA is VehicleBERT with dynamic self-attention; (1) Confusion matrix test, as Figure 3 (BERT-Base), Figure 4 (VehicleBERT), Figure 5 shown (VehicleBERT-DA); (2) Accuracy, precision, recall, and F1-score tests; Model Accuracy (%) Precision (%) Recall (%) F1-Score (%) BERT-Base 100.00 100.00 99.97 99.98 VehicleBERT 99.98 99.99 99.86 99.93 VehicleBERT-DA 100.00 100.00 100.00 100.00 (3) Other lightweight index tests, where the lightweight indexes include: number of parameters, space size, GPU occupancy rate, and CPU occupancy rate; Model Number of Parameters Space Size (MB) GPU Occupancy Rate (%) CPU Occupancy Rate (%) VehicleBERT 2040901 7.79 5-15% 5-10% VehicleBERT-DA 2041159 7.79 5-15% 5-10% BERT-Base 109486085 417.66 70-95% 40-60% (4) Experimental tests of BERT-Base, VehicleBERT, and VehicleBERT-DA, and comparison results.
[0025] In the in-vehicle network intrusion detection task; BERT-Base, although showing excellent classification performance (accuracy, precision, recall, and F1-score are all 1.0000), its huge number of parameters (109,486,085) and high resource consumption (417.66MB, GPU occupancy 70 - 95%, CPU occupancy 40 - 60%) make it difficult to apply in resource-constrained in-vehicle environments; VehicleBERT reduces the number of parameters significantly to 2,040,901 and the storage space to 7.79 MB, while still maintaining performance close to that of BERT-Base (accuracy 0.9998, precision 0.9999, recall 0.9986, F1 score 0.9993), and significantly reducing resource consumption (GPU occupancy 5 - 15%, CPU occupancy 5 - 10%), making it suitable for in-vehicle devices. The further optimized VehicleBERT-DA achieves perfect classification performance (accuracy, precision, recall, and F1 score are all 1.0000), while maintaining extremely low computational and storage overheads (number of parameters 2,041,159, storage space 7.79 MB, GPU occupancy 5 - 15%, CPU occupancy 5 - 10%), demonstrating the best lightweight and efficiency.
[0026] In summary, although BERT-Base performs excellently in terms of performance, its high resource requirements limit its practical applications. The VehicleBERT and VehicleBERT-DA in the VehicleBERT series of models provide solutions more suitable for in-vehicle environments with lower resource consumption. In particular, VehicleBERT-DA is the most excellent model in this embodiment and serves as the final VehicleBERT model.
[0027] The above is only a specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A lightweight in-vehicle network intrusion detection method based on BERT, characterized in that, It includes the following steps: S1. Data preprocessing stage: Convert the labeled CAN samples into text samples, including: loading data, label mapping, feature column processing, feature concatenation, dataset division, word segmentation, and encoding; S2. Model training stage: Input the labeled text samples into the VehicleBERT series models for training, including: setting training hyperparameters and passing the processed data to the VehicleBERT series models for training; S3. Model testing and evaluation stage: Conduct inference testing on the trained VehicleBERT series models.
2. The lightweight vehicle network intrusion detection method based on BERT according to claim 1, characterized in that, In step S1, it also includes the following steps: S101. Load data: Use the Car-hacking public dataset as the data. The Car-hacking public dataset is generated by transmitting CAN data packets to the CAN bus of a real vehicle and includes CAN data packets collected from the OBD-II port; S102. Label mapping: Create a label mapping dictionary to map the original labels to digital labels: 'R' → 0, 'RPM' → 1, 'gear' → 2, 'DoS' → 3, 'Fuzzy' → 4. Check the unique values in the Label column to ensure that all labels are within the mapping range. Replace the values in the Label column with digital labels. After replacing the labels, check if there are any null values in the Label column; S103. Feature column processing: Check if there are any null values in the CAN ID and the data columns DATA[0]-DATA[7], and delete the rows with null values in the CAN ID and Label columns; S104. Feature concatenation: Concatenate the CAN ID and the data columns DATA[0]-DATA[7] into text format, check if there are any null values in the concatenated text column, and ensure that the number of concatenated text and labels is the same; S105. Dataset division: Divide the training set, validation set, and test set in a ratio of 7:2:1 and save them; S106. Word segmentation and encoding: Before the VehicleBERT series models start training and testing, use the BERT tokenizer to convert the text into the format required for input to the VehicleBERT series models.
3. The lightweight vehicle network intrusion detection method based on BERT according to claim 2, characterized in that, In step S106, it also includes the following steps: (1) Word segmentation is to break a piece of text into smaller units. The tokenization method used by BERT is WordPiece. By breaking words into the smallest lexical units, the goal of this process is to handle the problem of unknown words and reduce the size of the vocabulary. The formula is as follows: ; Among them, represents the original text, represents the tokenizer, represents the result after tokenization; (2)Add special tokens, and add [CLS] and [SEP] at the beginning and end of the tokenization result respectively. The formula is as follows: ; Among them, A flag bit indicating a classification task. The output result of the VehicleBERT series of models is based on the representation at this position. Indicates a separator used to distinguish sentence pairs or the end of text. Indicates each token unit obtained by the above tokenizer. Indicates the final tokenization result after adding special tokens; (3)After tokenization is completed, the next step is encoding, which converts the tokenized text into a numerical format that can be understood by the VehicleBERT series of models. BERT uses two main inputs: input_ids: the integer ID corresponding to each token, and its formula is as follows: ; Among them, represents the integer representation of the tokenization unit, which is a list where each integer corresponds to the position of a tokenization unit in the vocabulary, represents the th tokenization unit, which is a word or subword obtained from the tokenization result, represents the index function that maps the tokenization unit to the vocabulary, represents the flag bit for the classification task, and the output result of the VehicleBERT series of models is based on the representation at this position, represents the delimiter used to distinguish sentence pairs or the end of the text; It is used to indicate which positions of the VehicleBERT series of models need to be attended to and which positions to be ignored, aiming to help the VehicleBERT series of models process variable-length inputs and calculate the padding for the padded parts. The formula is as follows: ; Among them, denotes an attention mask, which is a list with a length consistent with , and is used to indicate which positions are valid text and which are padding. denotes the -th mask value, taking a value of 1 or 0. Among them, if is valid text , = 1; if is padding , then = 0.
4. A lightweight in-vehicle network intrusion detection method based on BERT according to claim 1, characterized in that In step S2, it also includes the following steps: S201. Set training hyperparameters: In the VehicleBERT series models, use AdamW as the optimizer. AdamW introduces weight decay, which improves the defect of Adam in regularization by directly reducing the values of parameters during the gradient update stage; S202. Transfer the processed data to the VehicleBERT series of models for training. The VehicleBERT series of models is based on BERT. The VehicleBERT-DA model and the VehicleBERT model in the VehicleBERT series include an embedding layer, an encoder layer, a pooling layer, and an output layer. The embedding layer: converts the input vocabulary index into a word vector representation. The encoder layer: extracts the context features of the sequence through self-attention and a feed-forward network. The pooling layer: extracts global features from the encoder output. The output layer: maps the features of the pooling layer to the class space for classification prediction.
5. The lightweight vehicle network intrusion detection method based on BERT according to claim 4, characterized in that, In step S202, the following steps are also included: 1.
1. Word embedding calculation. The role of word embedding is to map each discrete token ID to a high-dimensional vector space. The formula is as follows: ; Among them, is the word embedding matrix, is the embedding vector of the th token, and is each discrete token ID; 1.
2. Position embedding calculation. To introduce the position information of each token in the sequence, the embedding layer adds a position embedding. The position embedding is a fixed sine embedding or a learnable embedding. Here, a learnable embedding is selected. The formula is as follows: ; Among them, is the position of the token in the sequence, is the learnable position embedding matrix, is the position embedding vector of the 1.
3. Paragraph embedding calculation. If the VehicleBERT series of models is used for sentence pair tasks, it is necessary to distinguish tokens belonging to different sentences, which is achieved through paragraph embedding. The formula is as follows: ; Among them, is the paragraph embedding matrix, is the type of the sentence to which the -th token belongs. Since the vehicle network intrusion detection task is single-sentence text classification, the VehicleBERT series of models will default to setting it to all 0s; 1.
4. Combination of embeddings. Add the word embedding, position embedding, and paragraph embedding element-wise. The formula is as follows: ; Among them, is the output of the embedding layer, containing word information , position information and sentence type information .
6. A lightweight vehicle network intrusion detection method based on BERT according to claim 4, characterized in that, In step S202, it also includes the following steps: 2.
1. Multi-head self-attention mechanism. Calculate the query (Q), key (K), and value (V). The formula is as follows: ; Among them, is the input matrix, and the query matrix is generated through three linear transformations , the key matrix , and the value matrix . , , are the learned weight matrices; Calculate the attention of each head. The formula is as follows: ; Among them, is the dimension of the key, which is the same as the dimension of the query, used to scale the dot product to prevent gradient explosion or vanishing, is the query matrix, the transpose matrix of the key matrix, is the value matrix, is the attention result of each head; Here, a dynamic attention head is introduced. In the standard multi-head attention mechanism, the weights of all heads are static. In the VehicleBERT-DA model, the dynamic attention head dynamically adjusts the weights of each attention head according to the global features of each layer. The generation of dynamic weights is as follows: ; Among them, is the global feature, and are the weights and biases, respectively used to adjust the generated dynamic weights, and through and limit the output range and normalize it so that the weights of each head are between 0 and 1; Then, apply the dynamic weights to the output of each head: ; wherein is the result of applying dynamic weights to the output of each head, is the output of each head, is the dynamically generated weight; Concatenate the outputs of all heads. Concatenate the weighted outputs of all heads: ; Among them represents the output result of concatenating all heads, represents the concatenation function, represents the result of applying dynamic weights to the output of each head. Finally, the final multi-head attention output is obtained through a linear transformation, is the output weight matrix, and the shape is , where is the number of heads, is the dimension of each head, is the dimension of the VehicleBERT-DA model or the dimension of the hidden layer, representing the size of the feature vector of each layer in the VehicleBERT-DA model; 2.
2. Feed-forward network. The feed-forward network consists of two linear layers. The output of the first linear layer undergoes a non-linear transformation through an activation function, and the second linear layer maps the output back to the dimension of the VehicleBERT series of models. The formula is as follows: ; Among them, represents the input, , represents the weights and biases of the first linear layer, , represents the weights and biases of the second linear layer, and finally, a lightweight design is carried out. By setting intermediate_size = 256, the dimension of the intermediate layer of the feed-forward network is reduced. In the VehicleBERT series of models, intermediate_size is compressed to 256, reducing the computational burden and memory consumption of the network; 2.
3. Residual connection. Used to solve the problem of gradient vanishing. The residual connection directly adds the input to the output for the flow of information. The formula is as follows: ; Among them, is the input, is the output after processing. For each layer, the residual connection directly adds the input to the output after transformation. In the VehicleBERT series of models, the residual connection is used to add the output of the multi-head attention to the input and add the output of the feed-forward network to the input; 2.
4. Layer normalization. Layer normalization is used to stabilize training. The output of each layer is within a set range to prevent gradient explosion or vanishing. The formula is as follows: ; Among them, represents the input, represents the input mean value of, represents the input standard deviation of, represents the scaling factor, represents the bias term. The scaling factor γ is automatically initialized to 1, and the bias β is set to 0 and optimized during the training process. Layer normalization is mainly used in the VehicleBERT series of models to perform layer normalization on the results after the residual connection.
7. A lightweight vehicle network intrusion detection method based on BERT according to claim 4, characterized in that, In step S202, it also includes the following steps: 3.1.
1. Representation of the encoder layer output. The pooling layer receives the output of the encoder as input. This output is the hidden state from each token, which is a tensor. The output of the encoder layer represents the context information of each token in the input sequence and is the hidden state after being processed by the Transformer encoder. 3.1.
2. Extract the output of the [CLS] token, and extract the hidden state corresponding to the first token as the representation of the entire input sequence. 3.1.
3. Pooling output: Extract the representation of the [CLS] token from the output of the encoder as the final pooling output. 3.1.
4. Lightweight design: The VehicleBERT series of models only focus on the dynamically optimized [CLS] features. The VehicleBERT series of models complete the [CLS] feature extraction with a simple [:, 0, :] indexing operation, and the design of the dynamic adjustment module is mainly based on low-dimensional operations, supporting low-resource environments.
8. A lightweight in-vehicle network intrusion detection method based on BERT according to claim 4, characterized in that, In step S202, it also includes the following steps: 4.1.
1. Input the [CLS] features from the pooling layer. 4.1.
2. Perform feature mapping through a linear layer, mapping the [CLS] features from the hidden_size dimension to the dimension of the target category. 4.1.
3. Obtain the probability distribution of the final classification through linear layer calculation and the Softmax function. Lightweight design: The VehicleBERT series of models set the hidden layer dimension to 64, reduce the number of parameters in the output layer, the size of the weight matrix of the linear layer is (64, 5), and the number of parameters is 320.
9. A lightweight in-vehicle network intrusion detection method based on BERT according to claim 1, characterized in that, In step S3, it also includes the following steps: S301. Set the evaluation metrics to accuracy, precision, recall, and F1-score respectively, and their calculation formulas are as follows: ; ; ; ; Among them, represents the number of normal samples correctly predicted by the VehicleBERT series models, represents the number of abnormal samples correctly predicted by the VehicleBERT series models, represents the number of abnormal samples wrongly predicted as normal samples by the VehicleBERT series models, represents the number of normal samples wrongly predicted as abnormal samples by the VehicleBERT series models; S302. Test BERT-Base, VehicleBERT, and VehicleBERT-DA using the same software and hardware environment, where VehicleBERT-DA is the VehicleBERT with dynamic self-attention. (1) Confusion matrix test, (2) Accuracy, precision, recall, and F1-score tests; (3) Other lightweight metric tests, where the lightweight metrics include: number of parameters, space size, GPU occupancy rate, and CPU occupancy rate; (4) Experimental tests of BERT-Base, VehicleBERT, and VehicleBERT-DA, and compare the results.
10. A lightweight in-vehicle network intrusion detection method based on BERT according to any one of claims 1-9, characterized in that, The VehicleBERT series of models refer to the VehicleBERT model and the VehicleBERT-DA model.
Citation Information
Cited By
Industrial internet network intrusion detection method based on pre-trained large language model
CN121644231A