Segmentation learning vehicle networking anomaly detection method and system based on large language model
By adopting a segmentation learning method based on large language models in the Internet of Vehicles, the context dependence of the BSM sequence is extracted using the LLM model of the two-way self-attention mechanism, high-precision anomaly detection is achieved, and the computing load and privacy leakage risks are reduced, and the problems of insufficient detection accuracy and high privacy leakage risks in the existing technology are solved.
Patent Information
- Application Number
- CN202510279852.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, the Internet of Vehicles abnormality detection method has insufficient detection accuracy, large computing resources occupies, and high privacy leakage risk, making it difficult to effectively detect potential security threats in dynamic information sharing among vehicles.
The segmented learning vehicle network abnormal detection method is adopted based on a large language model. The BSM sequence is generated through a sliding window, and the Token encoder is used to map it into a high-dimensional token representation, and the token is uploaded to the cloud. The context dependence relationship between the tokens is extracted using the LLM model of the two-way self-attention mechanism, and the message sequence list matrix is generated. The classifier on the vehicle calculates the probability of anomalies based on the representation matrix to complete real-time detection.
It has achieved the effects of high detection accuracy, low computing load and strong privacy protection, and solved the problems of insufficient detection accuracy, large computing resource utilization, and high privacy leakage risks in the existing technology.
Smart Images

Figure CN120123879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Internet of Vehicles (IoV) security, and particularly to a split learning-based IoV anomaly detection method (IoVSL) and system using a large language model. Background Art
[0002] With the popularization of intelligent connected vehicles, the IoV enables dynamic information sharing among vehicles through vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communication technologies. Vehicles broadcast basic safety messages (BSMs) to transmit key data such as location, speed, and timestamp to support functions such as cooperative driving and collision warning. However, the openness of the IoV exposes it to serious security threats: malicious vehicles may forge BSMs, leading to network anomalies (such as false traffic congestion signals and incorrect path planning), and even causing traffic accidents. Therefore, anomaly detection technology has become a core requirement for IoV security.
[0003] In the prior art, traditional machine learning methods (such as support vector machines and random forests) rely on manual feature engineering and are difficult to capture the context dependencies of BSM sequences, resulting in limited detection accuracy. Deep learning methods (such as CNN and LSTM) can automatically extract features, but are limited by the computing power of in-vehicle devices and require data to be transmitted to the cloud for processing, leading to a risk of privacy leakage. In addition, existing models generally ignore the impact of intra-class similarity and inter-class differences on detection performance and cannot balance computational efficiency and privacy protection. For example, methods based on large language models (LLMs) such as MistralBSM need to convert BSMs into text inputs, resulting in loss of context information and insufficient detection accuracy. Summary of the Invention
[0004] The present invention provides a split learning-based IoV anomaly detection method and system using a large language model, which have the advantages of high detection accuracy, low computational load, and strong privacy protection, and solve the problems of insufficient detection accuracy, large consumption of computing resources, and high risk of privacy leakage in the prior art.
[0005] A split learning-based IoV anomaly detection method using a large language model, the method comprising the following steps:
[0006] Step 1: Generate BSM sequences based on a sliding window and map the BSM sequences to high-dimensional Token representations;
[0007] Step 2: Upload the Tokens to the cloud and use an LLM model with a bidirectional self-attention mechanism to extract the context dependencies between the Tokens and generate a message sequence representation matrix;
[0008] Step 3: A classifier on the vehicle calculates the anomaly probability based on the representation matrix to complete real-time detection.
[0009] The present invention also provides a split learning vehicle networking anomaly detection system based on a large language model. This system is used to execute a split learning vehicle networking anomaly detection method based on a large language model as described in the claims; this anomaly detection system is implemented through an IoV anomaly detection model based on a large language model composed of a Token encoder, an LLM model with a bidirectional self-attention mechanism, and a classifier; the IoV anomaly detection model based on a large language model is used to detect anomalies in the BSM sequence; the Token encoder and the classifier are deployed on the vehicle, and the LLM model with a bidirectional self-attention mechanism is deployed in the cloud.
[0010] Advantages of the present invention:
[0011] 1. For the vehicle networking anomaly detection method described in the present invention, through the combination of a vehicle-cloud collaborative architecture and a bidirectional self-attention mechanism, the context dependence relationship of the BSM sequence is fully extracted, thereby achieving the effect of high detection accuracy and solving the problem of insufficient detection accuracy in the prior art.
[0012] 2. For the vehicle networking anomaly detection method described in the present invention, through split learning, computationally intensive tasks are offloaded to the cloud, and the module on the vehicle only requires 568 KB of storage space, thereby achieving the effects of low computational load and strong privacy protection, and solving the problems of large computational resource occupation and high privacy leakage risk in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of the overall architecture of a split learning vehicle networking anomaly detection method based on a large language model described in the present invention;
[0014] Figure 2 It is a schematic diagram of the principle of the bidirectional self-attention mechanism in the method of the present invention;
[0015] Figure 3 It is a performance comparison diagram of the method of the present invention on the VeReMi extended dataset; (a) is the comparison effect diagram of precision; (b) is the comparison effect diagram of recall; (c) is the comparison effect diagram of F1 score; (d) is the comparison effect diagram of accuracy. DETAILED DESCRIPTION OF THE INVENTION
[0016] DETAILED DESCRIPTION OF THE INVENTION I. In combination with Figure 1Describe this embodiment, a split learning Internet of Vehicles (IoV) anomaly detection system based on a large language model. This detection system is implemented based on the IoV anomaly detection model (IoVLLM) of the large language model. IoVLLM includes a Token encoder, an LLM model with a bidirectional self-attention mechanism, and a classifier. The lightweight Token encoder and classifier are deployed on the vehicle, and the LLM model with the bidirectional self-attention mechanism is deployed on the cloud; it is used to detect anomalies in the BSM sequence. In the training phase, the parameters of the LLM are fine-tuned using the labeled BSM sequence, and the parameters of the tokenizer and classifier are updated. In the online detection phase, the trained IoVSL is used to detect new BSM sequences.
[0017] Specific Embodiment 2. In combination with Figures 1 to 3 Describe this embodiment. This embodiment is an anomaly detection method (IoVSL) based on a large language model for split learning of the Internet of Vehicles anomaly detection system described in Specific Embodiment 1. This method splits the IoV anomaly detection model, reducing the computing and storage pressure on the vehicle and the risk of privacy leakage.
[0018] The anomaly detection method described in this embodiment is implemented by the following steps:
[0019] Step 1: Generate a BSM sequence based on a sliding window and map it to a high-dimensional Token representation through a Token encoder;
[0020] The BSM sequence M = {M 1 , M 2 ,..., M m} is generated in a sliding window manner. Each element in M represents each BSM in the sequence. m defines the window length, and the window slides with a specified step size r. It is mapped to a high-dimensional Token representation through a Token encoder;
[0021] In this embodiment, the Token encoder includes a linear layer and a batch normalization layer;
[0022] The linear layer is used to encode the BSM sequence into the semantic space of the LLM, achieving the modality alignment of sequence data and text data, as shown in the following formula:
[0023] Linear out = x BSM × W BSM + b BSM
[0024] where Linear out represents the representation of the BSM sequence in the semantic space of the LLM, x BSM ∈ R m×nis the initial feature matrix of the BSM sequence M, W BSM ∈R n×d , b BSM ∈R m×d are the learnable parameters of the linear layer.
[0025] To accelerate the training process, improve the convergence speed and generalization ability of the model, a batch normalization layer is used to normalize Linear out to obtain the Token representation E ∈ R m×d , as follows:
[0026]
[0027] where γ and β are the learnable parameters of the batch normalization layer, ∈ is a small constant, and μ and σ 2 represent the mean and standard deviation of Linear out respectively. In addition, the i-th row of the Token representation E ∈ R m×d also represents the Token representation of the i-th BSM in the BSM sequence, that is, E = (e 1 , e 2 , …, e m , ) ∈ R m×d .
[0028] In this embodiment, since the LLM cannot directly process data patterns other than text. Therefore, a Token encoder is designed to convert the constructed BSM sequence into a Token representation E = (e 1 , e 2 , …, e m , ) ∈ R m×d within the semantic space of the LLM, where e i ∈ R d represents the Token representation of the i-th BSM in the sequence.
[0029] Step 2: Upload the Token to the cloud, and use the LLM model with bidirectional self-attention mechanism to extract the context dependencies between Tokens and generate a message sequence representation matrix;
[0030] In this embodiment, in order to fully capture the context dependencies between BSMs, the Llama3.1-8B large model in the large language model LLM is introduced as the backbone network for anomaly detection. Llama3 is one of the state-of-the-art open-source large language models, providing enhanced understanding and reasoning capabilities. Compared with earlier models, Llama3.1-8B has a larger parameter size and a more diverse training dataset, thus optimizing its ability to understand deeper semantic nuances and model complex contexts. This makes the model more effective in processing long texts and large-scale data. Llama3.1-8B adopts a pure decoder architecture based on Transformer and uses a masking mechanism to implement a unidirectional self-attention mechanism. Therefore, the generation of tokens in Llama3.1-8B can only depend on the previously generated tokens. However, there is a context relationship between each BSM in the sequence and all other BSMs. Therefore, in order to enable the representation of each BSM to fully focus on the information of its preceding and succeeding BSMs, the present invention removes the causal masking mechanism of Llama3.1-8B to implement a bidirectional self-attention mechanism; that is: improving Llama3.1-8B using the bidirectional self-attention mechanism to construct an LLM model with a bidirectional self-attention mechanism, which can fully extract the context relationship between each BSM in the sequence and all other BSMs to obtain a message sequence representation matrix Obtain a message sequence representation matrix, such as Figure 2 shown, where the matrix represents the i-th BSM.
[0031] In this embodiment, the LLM model with a bidirectional self-attention mechanism can fully extract the context dependencies between BSMs, realize the similarity representation between BSM sequences with the same label, that is, intra-class similarity, and the difference representation between BSM sequences with different labels, that is, inter-class difference. Therefore, in order to enhance intra-class similarity and inter-class difference, an intra-class loss function L INTRA and an inter-class loss function L INTER are designed as follows:
[0032]
[0033] L INTER = ||P 0 - P 1 || 2
[0034] where ||·|| 2 is the L2 norm, and P 0 , P 1 are the prototypes corresponding to normal and abnormal BSM sequences. The prototypes are obtained by averaging the elements at the same positions in the representation matrices of BSM sequences with the same label. For example, is the normalized representation matrix of the normal BSM sequence. The prototype P of the normal BSM sequence 0 is calculated as follows:
[0035]
[0036] where (P 0 ) j,i represents the element in the j-th row and i-th column of the prototype P 0 , and represents the element in the j-th row and i-th column of the representation matrix . Similarly, P 1 can also be calculated. For example, is the normalized representation matrix of the abnormal BSM sequence. The prototype P of the abnormal BSM sequence 1 is calculated as follows:
[0037]
[0038] where (P 1 ) j,i represents the element in the j-th row and i-th column of the prototype P 1 , and represents the element in the j-th row and i-th column of the representation matrix .
[0039] Step 3: The classifier on the vehicle calculates the abnormal probability based on the characterization matrix to complete real-time detection.
[0040] In this embodiment, the goal of anomaly detection is to determine whether the BSM sequence is a normal class or an abnormal class. Therefore, a binary classification model including an average pooling layer, a batch normalization layer, and a softmax linear layer is designed.
[0041] First, use the average pooling layer to calculate the representation vector V ∈ R d of the BSM sequence as follows:
[0042]
[0043] where V i is the position of the i-th element in the representation vector V, and is the element in the j-th row and i-th column of the message sequence characterization matrix E * ∈ R m×d .
[0044] Then, use the batch normalization layer to normalize the representation vector of the BSM sequence as follows:
[0045] V norm = BN(V)
[0046] Finally, a linear layer with softmax is used to perform binary classification on the BSM sequence as follows:
[0047] classV = softmax(V norm ×W cf +b cf )
[0048] where classV = (P abnormal , P normal ) ∈ R 2 The elements in represent the anomaly probability and the normal probability respectively. W cf ∈ R d×2 and b cf ∈ R 2 are the learnable parameters of the linear layer. In the detection stage, P abnormal ≥ P normal indicates an IoV anomaly, otherwise it is normal.
[0049] In this embodiment, in order to more precisely quantify the difference between the probability distribution of the detection model and the true probability distribution, cross-entropy loss is introduced as follows:
[0050] L CE = -y × log(P abnormal ) + (1 - y) × log(P normal )
[0051] where y ∈ {0, 1} is the true label of the BSM sequence.
[0052] Therefore, combining the intra-class loss function, the inter-class loss function, and the cross-entropy loss function, the total loss function is defined as:
[0053] L TOTAL = L CE + L INTRA - L INTER
[0054] In the training stage, by minimizing L TOTAL , the parameters of the LLM model of the bidirectional self-attention mechanism in Llama 3.1 - 8B are fine-tuned, and the parameters of the Token encoder and the classifier are updated. In the online detection stage, the trained IoVLLM is used to detect new BSM sequences.
[0055] As Figure 3As shown, to verify the effectiveness of the method proposed in this embodiment, the IoVSL method of the present invention was compared with six baseline methods: VeReMi Extension, DCLE, DeepADV, DRL, DLE, and MistralBSM. The effectiveness of IoVSL was analyzed through metrics such as precision, recall, F1-score, and accuracy.
[0056] Precision measures the proportion of true anomalies among all detected anomalies. The comparison of the precision of IoVSL with other baseline methods is shown in Figure 3 (a) of. The precision of IoVSL is 100%, which is 0.19% higher than the best baseline method. Since VeReMi Extension relies on predefined rules and threshold selection, the threshold selection will affect the detection precision. Although DLE, DCLE, DeepADV, and DRL use CNN or LSTM to extract features from BSM or BSM sequences, they cannot fully capture the context dependencies between BSMs in the sequence. Although MistralBSM introduces LLM, converting the BSM sequence into a prompt will result in a significant loss of context dependencies between BSMs. Therefore, DLE, DCLE, DeepADV, DRL, and MistralBSM misclassify many normal BSM sequences that are hardly distinguishable from abnormal BSM sequences as abnormal. The LLM model using the bidirectional self-attention mechanism accurately extracts the context dependencies between BSMs and maximizes the inter-class loss to clarify the feature differences between abnormal and normal message sequences. In addition, these deep learning baseline methods need to transmit BSMs to the cloud or edge for anomaly detection, which brings the risk of privacy leakage.
[0057] Recall measures the percentage of detected anomalies compared to the total number of actual anomalies. Figure 3 (b) in shows the comparison of IoVSL with other benchmark methods in terms of recall. The recall of IoVSL is 99.92%, which is 0.22% higher than the best benchmark method. IoVSL can automatically extract the context relationships between BSMs and accurately represent the BSM sequence. VeReMiExtension ignores anomaly messages that bypass the predefined threshold. DLE, DCLE, DeepADV, DRL, and MistralBSM cannot fully extract the context relationships between BSMs. They also ignore the similarities of BSM sequences belonging to the same label and the differences between sequences belonging to different labels. Therefore, they produce inaccurate representations and incorrect classifications of anomaly message sequences. In summary, our method captures the differences between normal and abnormal BSM sequences and correctly classifies abnormal BSM sequences.
[0058] The F1-score is the harmonic mean of the model's precision and recall.Figure 3 Among them, (c) shows the comparison of IoVSL with other benchmark detection methods in terms of F1 score. The F1 score of IoVSL reaches 99.96%, which is 1.51% higher than the best benchmark method. VeReMi Extension cannot detect behaviors that bypass the fixed threshold. In contrast, IoVSL automatically extracts features without a threshold. This results in a significantly higher F1 score of IoVSL than VeReMi Extension. In addition, IoVSL is more effective than DLE, DCLE, DeepADV, DRL, and MistralBSM. It extracts the context dependencies between BSM messages more comprehensively. It also captures the intra-class similarity between normal or abnormal BSM sequences belonging to the same class, as well as the inter-class differences between sequences belonging to different classes. At the same time, the method of the present invention uses split learning to reduce the computational and storage overhead of vehicles while reducing the risk of privacy leakage.
[0059] Since the accuracy rate considers both abnormal and normal BSM sequence samples, it provides a balanced evaluation of the normal class and the abnormal class. Figure 3 Among them, (d) shows the comparison of IoVSL with other benchmark methods in terms of accuracy. The accuracy rate of anomaly detection of this method is 99.96%, which is 1.14% higher than the best baseline method. The reasons for the highest accuracy rate of IoVSL are: (1) It accurately represents the BSM sequence using the LLM model; (2) It uses the intra-class similarity loss and the inter-class difference loss respectively to ensure the similar representation of BSM sequences of the same class and the different representation of BSM sequences of different classes. At the same time, IoVSL uses split learning to achieve vehicle-cloud collaborative anomaly detection.
[0060] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0061] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it cannot be understood as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.
Claims
1. A segmentation learning method for vehicle network anomaly detection based on a large language model, characterized by: The method is implemented by the following steps: Step 1: Generate a BSM sequence based on a sliding window, and map the BSM sequence into a high-dimensional Token representation; Step 2: Upload the token to the cloud, use the LLM model of the bidirectional self-attention mechanism to extract the contextual dependencies between tokens, and generate a message sequence representation matrix; Step 3: The on-board classifier calculates the abnormal probability according to the message sequence representation matrix to complete real-time detection.
2. The method for detecting anomalies in Internet of Vehicles based on segmentation learning of a large language model according to claim 1, characterized in that: In step 1, the BSM sequence is mapped to a high-dimensional Token representation through a Token encoder; the BSM sequence M={M1,M2,...,M m } is generated in a sliding window manner, m is the defined window length, and the window slides with a specified step size; the BSM sequence is converted into a Token representation E = (e1, e2, ..., e m ,),e i Token representation of the i-th BSM in the sequence.
3. The method for detecting anomalies in Internet of Vehicles based on segmentation learning of a large language model according to claim 2 is characterized by: The Token encoder includes a linear layer and a batch normalization layer; The linear layer realizes the modality alignment of BSM sequence data and text data, which is expressed as follows: Linear out =x BSM ×W BSM +b BSM Where, Linear out is the representation of the BSM sequence in the semantic space of LLM, x BSM is the initial feature matrix of the BSM sequence M, W BSM , b BSM are the learnable parameters of the linear layer; Use batch normalization layer to Linear out Normalized, it can be expressed as follows: Where γ and β are learnable parameters of the batch normalization layer, ∈ is a constant, μ and σ 2 Linear out The mean and standard deviation of .
4. The method for detecting anomalies in Internet of Vehicles based on segmentation learning of a large language model according to claim 1, characterized in that: In step 2, the LLM model based on the bidirectional self-attention mechanism is used to extract the contextual relationship between each BSM in the BSM sequence and all other BSMs, and obtain the message sequence representation matrix E * ; According to the characterization matrix E * , design the intra-class loss L INTRA and the inter-class loss L INTER , which can be expressed as: L INTER =||P0-P1||2 Where ||·||2 is the L2 norm, P0 and P1 are the prototypes corresponding to the normal BSM sequence and the abnormal BSM sequence.
5. The method for detecting anomalies in Internet of Vehicles based on segmentation learning of a large language model according to claim 4 is characterized by: In step 3, the classifier is designed as a binary classification model including an average pooling layer, a batch normalization layer, and a softmax linear layer; First, the average pooling layer is used to calculate the representation vector V of the BSM sequence; then, the batch normalization layer is used to normalize the representation vector V of the BSM sequence; finally, the linear layer with softmax is used to perform binary classification on the BSM sequence.
6. The method for detecting anomalies in Internet of Vehicles based on segmentation learning of a large language model according to claim 5 is characterized by: The position of the i-th element in the representation vector V is expressed as follows: Where V i represents the position of the i-th element in the vector V, is the message sequence feature matrix E * The element in the j-th row and i-th column of ; The softmax linear layer is used to classify the BSM sequence into two categories, which can be expressed as follows: classV=softmax(V norm ×W cf +b cf ) In the formula, classV=(P abnormal ,P normal ) are abnormal probability and normal probability respectively, W cf and b cf is the learnable parameter of the softmax linear layer; In the detection phase, P abnormal ≥P normal Indicates that the IoV is abnormal, otherwise it is normal.
7. The method for detecting anomalies in Internet of Vehicles based on segmentation learning of a large language model according to claim 6 is characterized by: The cross entropy loss function is introduced into the binary classification model, which is expressed as follows: L CE =-y×log(P abnormal )+(1-y)×log(P normal ) Where y∈{0,1} is the true label of the BSM sequence; Combining intra-class loss, inter-class loss and cross entropy loss, the total loss function is: L TOTAL =L CE +L INTRA -L INTER During the training phase, by minimizing L TOTAL , fine-tune the parameters of the LLM model of the bidirectional self-attention mechanism, and update the parameters of the Token encoder and classifier.
8. A segmentation learning vehicle network anomaly detection system based on a large language model, characterized by: The system is used to execute a segmentation learning vehicle network anomaly detection method based on a large language model as described in any one of claims 1-7. The anomaly detection system is implemented by an IoV anomaly detection model based on a large language model composed of a Token encoder, an LLM model of a bidirectional self-attention mechanism, and a classifier; and the IoV anomaly detection model based on a large language model is used to detect anomalies in a BSM sequence.
9. The segmentation learning vehicle network anomaly detection system based on a large language model according to claim 8 is characterized by: The Token encoder and classifier are deployed on the vehicle, and the LLM model of the bidirectional self-attention mechanism is deployed in the cloud.