A constraint generation and multi-layer discrimination method for vehicle CAN abnormal traffic

CN122845207APending Publication Date: 2026-09-29HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610989376.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0007]针对现有技术中的缺陷,本发明提供一种面向车载CAN异常流量的约束生成与多层判别方法,以解决车载CAN网络中攻击样本稀缺、生成样本不符合CAN物理约束、少数类异常流量难以区分以及现有分类模型缺少类别结构约束的问题

Benefits of technology

(1)本发明通过构建SC-CPD模型和CAN约束算子,在训练和采样阶段同时约束字节取值范围、有效字节掩码、仲裁ID、DLC和车辆状态,使生成样本具备协议合法性和语义可信度,解决了传统生成式增强方法容易生成DLC之外字节非零、载荷越界或与车辆状态不一致的样本;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845207A_ABST
    Figure CN122845207A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of information security of Internet of Vehicles, and provides a constraint generation and multi-layer discrimination method for abnormal CAN flow of vehicles, comprising: acquiring vehicle CAN bus flow data as vehicle network flow data set, and preprocessing the data set to divide training set and test set; constructing SC-CPD model, and using the completed SC-CPD model to conditionally generate the attack category with less sample number in the training set to obtain an enhanced training set; constructing PCCT model, and using the enhanced training set to train the PCCT model; and discriminating based on the category probability output by the completed PCCT model, and selecting the attack category with the maximum probability as the final detection result. The present application can improve the precision of abnormal flow detection while improving the adaptability of the model to the category imbalance scene, and is convenient for deployment in vehicle gateway, domain controller or security monitoring terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle network information security technology, specifically to a constraint generation and multi-layer discrimination method for abnormal CAN traffic in vehicles. Background Technology

[0002] With the development of intelligent, connected, and software-defined vehicles, the number of in-vehicle electronic control units (ECUs) is constantly increasing. As the core communication protocol in the vehicle network, the CAN bus has long been responsible for transmitting critical control messages related to powertrain, chassis, body, and diagnostics. Initially, the CAN protocol focused primarily on real-time performance, reliability, and anti-interference capabilities, without incorporating mechanisms for message source authentication, encrypted transmission, and integrity protection. This allows attackers to potentially forge, replay, or flood CAN messages once they access the in-vehicle network via the OBD-II interface, T-Box, in-vehicle entertainment system, or compromised external ECUs, thereby affecting vehicle control safety.

[0003] To defend against such attacks, vehicle intrusion detection systems based on machine learning and deep learning have become a research hotspot. Traditional methods typically rely on statistical features, rule-based thresholds, or shallow classifiers, offering some detection capability against certain attack types. However, in real-world vehicle environments, CAN communication exhibits strong periodicity, strong protocol constraints, and high class imbalance, leading to significant limitations of existing methods in handling minority class attacks and cross-state generalization scenarios. First, there is a severe imbalance between the scarcity of attack samples and the data distribution. In real-world vehicle operation, the vast majority of CAN frames are normal periodic messages or event-triggered messages, with a low proportion of abnormal frames. Ordinary supervised learning models tend to favor the majority class of normal samples during training, compressing the decision boundaries for the minority class of abnormal samples and thus reducing their ability to identify low-frequency abnormal traffic.

[0004] Second, generative enhancement methods are prone to violating CAN protocol constraints. Existing GAN, VAE, or ordinary diffusion models often only focus on the similarity of numerical distribution when generating tabular CAN samples, while ignoring protocol structure constraints such as byte value range, valid DLC bytes, arbitration ID semantics, and vehicle state semantics. This can easily generate physically illegitimate or semantically unreliable pseudo-samples, which in turn interfere with the training of the detection model.

[0005] Third, existing methods are insufficient in modeling the internal structure and contextual relationships of CAN frames. CAN messages are not ordinary independent table features; there are close relationships between their arbitration ID, DLC, vehicle status, and payload bytes. If the detection model only treats each field as ordinary numerical features, it is difficult to fully capture the intra-frame byte dependencies, inter-field constraints, and semantic differences under different vehicle states, resulting in insufficient generalization ability of the model in complex and abnormal scenarios.

[0006] Fourth, deep models lack minority class structure constraints in their discrimination boundaries. While Transformers, CNNs, or LSTMs can learn complex nonlinear features, in imbalanced scenarios, minority class samples in the latent space may overlap with normal or other anomalous samples. Relying solely on the final classification head cannot guarantee that each class forms a stable and separable prototype structure in the latent space, affecting the model's ability to stably identify minority class anomalous traffic. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a constraint generation and multi-layer discrimination method for abnormal traffic in vehicle CAN networks, which solves the problems of scarce attack samples, generated samples not conforming to CAN physical constraints, difficulty in distinguishing a few types of abnormal traffic, and lack of category structure constraints in existing classification models.

[0008] This invention provides a constraint generation and multi-layer discrimination method for abnormal CAN bus traffic in vehicles, comprising: The vehicle CAN bus traffic data is acquired as the vehicle network traffic dataset, and the dataset is preprocessed and divided into training and testing sets according to a preset ratio. The fields of the vehicle CAN bus traffic data include timestamp, arbitration ID, DLC, vehicle status, payload bytes, effective byte mask, and attack category label. A state-constrained counterfactual load diffusion generation model is constructed, and the trained state-constrained counterfactual load diffusion generation model is used to conditionally generate attack categories with a small number of samples in the training set to obtain an enhanced training set. A prototype calibration CAN Transformer classification model is constructed, and the PCCT model is trained using the enhanced training set. The prototype calibration CAN Transformer classification model is used to map arbitration ID, vehicle status, DLC, effective byte mask and payload bytes to a unified feature space. Intra-frame byte dependency and state context information are extracted through the Transformer encoder, and the probability of each attack category is output through the main classification head, class prototype distance calibration head and training auxiliary classification head. Based on the trained prototype calibration CAN Transformer classification model, the output class probabilities are used for discrimination, and the attack class with the highest probability is selected as the final detection result.

[0009] Optionally, the state-constrained counterfactual load diffusion generation model includes: The conditional fusion layer is used to concatenate the noisy payload, effective byte mask, category embedding vector, arbitration ID embedding vector, vehicle state embedding vector, and time step embedding vector to obtain the conditional fusion vector. Among them, the noisy payload is obtained by adding diffused noise to the payload bytes, the category embedding vector, arbitration ID embedding vector, and vehicle state embedding vector are obtained by embedding and encoding the category label, arbitration ID, and vehicle state respectively, and the time step embedding vector is obtained by performing nonlinear mapping on the diffused time step. The conditional denoising backbone network is used to extract conditional denoising features from the conditional fusion vector; A noise prediction head is used to output a continuous noise estimate based on conditional denoising features; The byte discrete prediction header is used to output a discrete probability distribution of values ​​from 0 to 255 at each valid byte position based on the conditional denoising features. The CAN constraint operator is used to project the generated payload obtained from the noise prediction header and the byte discrete prediction header onto the feasible region that satisfies the CAN protocol constraints.

[0010] Optionally, the CAN constraint operator is specifically: , Among them, continuous load estimation , For noisy loads, For continuous noise estimation, The conditional denoising backbone network is based on the conditional denoising features extracted from the conditional fusion vector. For weights and biases, For the Sigmoid function, the truncation function, or an equivalent bounded function, This indicates element-wise multiplication, where m is the effective byte mask.

[0011] Optionally, the total loss function of SC-CPD is: , Among them, diffusion denoising loss , The noise is Gaussian, and m is the effective byte mask; Byte Discrete Cross-Entropy Loss , For the first Each actual byte value; The CAN constraint loss is: ; Semantic distribution loss ,in, and These are the byte mean and standard deviation of the generated sample within the g-th semantic group, respectively. and These are the byte mean and standard deviation of the real sample within this semantic group, respectively. This is the balance coefficient; Counterfactual target loss Among them, the generated samples are The prototype of the target attack category is The normal category prototype is The effective byte mask is m, and the effective byte weighted distance is... , This is the interval threshold. The normal prototype is far from the constraint weights; These are preset weighting coefficients.

[0012] Optionally, the prototype calibrates the CAN Transformer classification model, including: The token fusion layer is used to fuse payload byte embedding with position encoding, valid byte mask mapping, and context embedding to form a context-enhanced byte token sequence. The transformer encoder is used to model the byte token sequence to obtain an enhanced token representation, and then obtains the CAN frame-level feature vector through pooling operations; The main classification header outputs the main classification probability based on the CAN frame-level feature vector. Train auxiliary classification heads, including normal / attack binary classification heads and attack family classification heads. The normal / attack binary classification heads output binary classification probabilities based on CAN frame-level feature vectors, and the attack family classification heads output attack family probabilities based on CAN frame-level feature vectors. The prototype distance calibration head includes learnable prototype vectors corresponding to the number of attack categories. It is used to obtain prototype distance scores and output prototype probabilities based on CAN frame-level feature vectors and learnable prototype vectors.

[0013] Optionally, the final classification probability output by the prototype calibration CAN Transformer classification model is obtained by jointly calibrating the main classification probability and the prototype probability, specifically as follows: , in, To preset the fusion weights, This indicates a normalization operation. Primary classification probability, This is the weight matrix. For bias, For prototype probability, The prototype distance score is given, where h represents the CAN frame-level feature vector. Represents the prototype vector of the k-th class; This serves as the final detection result.

[0014] Optionally, the prototype calibrates the total training loss of the CAN Transformer classification model. ; Among them, the primary classification loss y represents the true category label. Represents the cross-entropy loss function; Prototype Distance Calibration Loss ; Normal / Attack Binary Classification Loss , , Labels indicating normal / attack status; Attack family classification loss , , Tag for attacking groups; This is the weight matrix. For bias; These are the preset weighting coefficients.

[0015] By adopting the above technical solution, this application has the following beneficial effects: (1) By constructing the SC-CPD model and CAN constraint operator, this invention simultaneously constrains the byte value range, effective byte mask, arbitration ID, DLC and vehicle status during the training and sampling stages, so that the generated samples have protocol legality and semantic credibility, solving the problem that traditional generative enhancement methods are prone to generating samples with non-zero bytes outside DLC, load out of bounds or inconsistent with vehicle status. (2) The SC-CPD model provided by the present invention generates samples based on attack category, arbitration ID and vehicle status. It uses semantic distribution loss to maintain statistical consistency within semantic units of the same state and arbitration ID. It uses counterfactual target loss to make the generated attack samples close to the target attack category prototype and far away from the normal category prototype. It effectively alleviates the problem of sample duplication caused by ordinary oversampling and category semantic ambiguity caused by ordinary generation models, and significantly improves the enhancement quality of minority class attack samples. (3) The SC-CPD provided by the present invention introduces counterfactual target constraints, which not only require the production samples to meet the CAN protocol format, but also further constrain the category attribution of the generated samples in the prototype space, so that the generated samples have more explicit attack category semantics, thereby providing higher quality training samples for subsequent classification models. (4) The PCCT model provided by this invention maps the payload byte, arbitration ID, vehicle status, DLC and effective byte mask to the feature space in a unified manner. It extracts the intra-frame byte dependency and field context relationship through the Transformer encoder, and forms a multi-layer discrimination structure through the normal / attack classification head, attack family classification head, main classification head and class prototype distance calibration head, thereby enhancing the separability of minority class attacks in the latent space. (5) The present invention realizes a closed loop from data preprocessing, constraint generation, enhancement training to prototype calibration and discrimination, which can improve the accuracy of abnormal traffic detection and enhance the model's adaptability to class imbalance scenarios, making it easy to deploy in vehicle gateways, domain controllers or security monitoring terminals. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0017] Figure 1 The flowchart illustrates a constraint generation and multi-layer discrimination method for abnormal traffic in vehicle CAN provided by an embodiment of the present invention; Figure 2 A schematic diagram of the state-constrained counterfactual load diffusion generation model provided in an embodiment of the present invention is shown; Figure 3 A schematic diagram of the prototype calibration CAN Transformer classification model provided in an embodiment of the present invention is shown. Detailed Implementation

[0018] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of the present invention and are therefore merely examples, and should not be construed as limiting the scope of protection of the present invention.

[0019] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application should have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0020] In one embodiment, such as Figure 1 As shown, a constraint generation and multi-layer discrimination method for abnormal CAN bus traffic in vehicles is provided, including: S1. Obtain the vehicle CAN bus traffic data as the vehicle network traffic dataset, and preprocess the dataset.

[0021] First, the data field acquisition is performed: the vehicle CAN bus traffic data is processed in the form of a single CAN message, including timestamp, arbitration ID, DLC, vehicle status, payload bytes, valid byte mask, and attack category label.

[0022] The following explains the correspondence and source of each field in the vehicle CAN bus traffic data. The timestamp is the acquisition time information appended by the CAN acquisition device when receiving or recording a CAN frame; the arbitration ID corresponds to the identifier field of the CAN frame; the DLC corresponds to the data length code field of the CAN frame, with Data_Byte_0 to Data_Byte_7 corresponding to the data field of the CAN frame; the vehicle status field is obtained from vehicle operation logs, data acquisition platform annotations, experimental records, or vehicle status signal parsing; the payload bytes include Data_Byte_0 to Data_Byte_7, and the valid byte mask includes Data_Valid_0 to Data_Valid_7; the valid byte mask Data_Valid_0 to Data_Valid_7 is generated during the preprocessing stage based on the DLC field; the category label is obtained from attack injection experimental records, dataset annotations, or manual annotations; the category label is not used as model input during the online detection stage.

[0023] The data was then cleaned and encoded. The system cleaned the raw data, removing records with missing category labels, abnormal DLC fields, missing payload fields, or physical fields that were out of bounds; the DLC field was limited to the range of 0 to 8, the payload bytes were limited to the range of 0 to 255, and the effective byte mask was limited to 0 or 1; the character-type arbitration ID was uniformly converted to uppercase strings, and arbitration ID encoders, vehicle status encoders, and category label encoders were constructed respectively; subsequently, while maintaining the proportion of each category, the dataset was divided into training and test sets.

[0024] It should be noted that the attack category can be determined based on a specific vehicle CAN attack dataset or actual attack annotation system, and is not limited to a fixed category. In one embodiment, the attack category label includes a normal category and four other categories: Flooding attack, Fuzzing attack, Replay attack, and Spoofing attack. Among them, Flooding attack indicates high-frequency injected messages causing bus congestion, Fuzzing attack indicates random or malformed payload injection, Replay attack indicates replay of historical legitimate CAN frames, and Spoofing attack indicates deceptive messages that forge specific arbitration IDs or status semantics.

[0025] S2. Construct a state-constrained counterfactual load diffusion generation model SC-CPD, and use the trained state-constrained counterfactual load diffusion generation model to conditionally generate attack categories with a small number of samples in the training set to obtain an enhanced training set.

[0026] like Figure 2 As shown, SC-CPD is used to generate conditions for attack categories with a small number of samples in the training set. The SC-CPD model includes a payload perturbation input, conditional embedding, diffusion time step embedding, effective byte mask, conditional fusion layer, conditional denoising backbone network, noise prediction head, byte discrete prediction head, and CAN constraint operator.

[0027] Let the true normalized load be The effective byte mask is The target attack category is Arbitration ID is The vehicle status is The diffusion time step is .

[0028] During the training phase, the system first selects minority class attack samples from the training set, normalizes their 8-byte true payloads into continuous payload vectors, and adds Gaussian noise to this payload vector at random diffusion time steps to obtain a noisy payload vector. During the generation phase, the system starts with the random noise payload and performs multi-step reverse denoising sampling based on the target attack category, arbitration ID, vehicle status, and effective byte mask. For the true payload... By adding diffused noise, we obtain a noisy load: .in, It is Gaussian noise. The noise intensity is related to the diffusion time step.

[0029] The conditional embedding layer is used to embed and encode the target attack category, arbitration ID, and vehicle status, specifically as follows: The target attack category indicates the anomaly type of the generated sample, the arbitration ID limits the CAN message semantic unit to which the generated payload belongs, and the vehicle state constrains the vehicle operating context corresponding to the generated payload. Through conditional embedding, SC-CPD can generate payloads with specified categories, arbitration IDs, and vehicle states, instead of generating them randomly without conditions.

[0030] The diffusion time step embedding layer is used to perform a nonlinear mapping on the current diffusion time step to obtain the time step embedding vector. Since different diffusion time steps correspond to different noise intensities, time step embedding enables the conditional denoising backbone network to perceive the current denoising stage and thus adaptively adjust the denoising amplitude.

[0031] The effective byte mask is determined based on the DLC field and is used to identify which positions in the 8-byte payload are effective payload bytes and which are invalid padding bytes. For example, when DLC is 6, the effective byte mask can be represented as the first 6 positions being 1 and the last 2 positions being 0. This effective byte mask is input to the conditional fusion layer, causing the model to focus on the positions of effective bytes during denoising; it is also input to the CAN constraint operator, used to force the positions of invalid bytes to zero during the output generation stage.

[0032] The conditional fusion layer is used to concatenate or fuse the noisy payload vector, target attack category embedding, arbitration ID embedding, vehicle state embedding, diffusion time step embedding, and effective byte mask to form a conditional fusion vector. The conditional fusion vector simultaneously includes the current state of the payload to be denoised, the target category semantics, the CAN message context, the diffusion stage information, and the DLC valid byte information, and serves as the input to the conditional denoising backbone network.

[0033] The conditional denoising backbone network employs a multilayer perceptron structure, including fully connected layers, nonlinear activation functions, and normalization layers. In one embodiment, the conditional denoising backbone network may adopt a structure of "fully connected layer—SiLU activation function—LayerNorm layer—fully connected layer—SiLU activation function" to extract conditional denoising features from the conditional fusion vector. This structure can conditionally model noisy loads based on target category, arbitration ID, vehicle state, and diffusion time step information. The conditional denoising backbone network extracts features from the conditional fusion vector: .in, This represents a conditional denoising backbone network consisting of fully connected layers, activation functions, and normalization layers. This is a conditional denoising feature. The noise prediction head outputs a continuous noise estimate based on the conditional denoising feature: .

[0034] The denoising features output by the conditional denoising backbone network are input into the noise prediction head and the byte discrete prediction head, respectively.

[0035] The noise prediction head outputs a continuous noise estimate, which guides the diffusion model to recover the clean load from the noisy load. The noise prediction head yields the continuous load estimate: .

[0036] The discrete byte prediction header outputs the discrete probability distribution of values ​​from 0 to 255 at each valid byte position, enabling the model to learn the discrete value patterns of CAN payload bytes. Through a dual-branch structure of continuous noise prediction and discrete byte prediction, SC-CPD can simultaneously model the continuous disturbance patterns and discrete byte distributions of the payload. The discrete byte prediction header outputs the discrete prediction distribution for each byte position based on conditional denoising features. , .in, Indicates the first 256-dimensional logits corresponding to each byte position Indicates the first The value of each byte is The probability, .

[0037] The CAN constraint operator projects the generated payload obtained from the noise prediction header and the byte discrete prediction header onto the feasible region that satisfies the CAN protocol constraints. The expression is: .in, For the Sigmoid function, the truncation function, or an equivalent bounded function, This indicates element-wise multiplication. The generated payload, generated through this projection, satisfies the byte value range constraint and the DLC effective byte mask constraint.

[0038] CAN constraint operators include byte range constraints, valid byte mask constraints, and state consistency constraints. The byte range constraint ensures that the generated payload, after reconstruction, falls within the range of 0 to 255. The valid byte mask constraint ensures that invalid bytes outside the DLC are at zero positions. The state consistency constraint ensures that the arbitration ID, vehicle status, and DLC fields remain unchanged during generation, so that the generated sample still belongs to the corresponding CAN semantic space.

[0039] Before being used to generate samples, SC-CPD needs to be trained on minority class attack samples in the training set. Its training objective is a weighted average of diffusion denoising loss, byte discrete cross-entropy loss, CAN constraint loss, semantic distribution loss, and counterfactual objective loss. .

[0040] Among them, diffusion denoising loss Used to constrain the model to recover diffused noise, byte discrete cross-entropy loss. Used to constrain the model to recover the true byte values, CAN constraint loss. Used to constrain the generated payload to meet the physical requirements of the CAN protocol, semantic distribution loss. Counterfactual target loss is used to maintain statistical consistency between samples generated with the same arbitration ID and vehicle status and real samples. This is used to make the generated samples closer to the target attack class prototype and further away from the normal class prototype, thereby enhancing the class discriminative power of minority class samples. These are preset weighting coefficients.

[0041] Diffusion denoising loss , The noise is Gaussian, and m is the effective byte mask; Byte Discrete Cross-Entropy Loss , For the first Each actual byte value; The CAN constraint loss is: ; Semantic distribution loss ,in, and These are the byte mean and standard deviation of the generated sample within the g-th semantic group, respectively. and These are the byte mean and standard deviation of the real sample within this semantic group, respectively. This is the balance coefficient; Counterfactual target loss Among them, the generated samples are The prototype of the target attack category is The normal category prototype is The effective byte mask is m, and the effective byte weighted distance is... , This is the interval threshold. The normal prototype is far from the constraint weights.

[0042] After training, the SC-CPD model parameters are fixed, and the trained SC-CPD is used to oversample attack categories with a small number of samples in the training set.

[0043] Specifically, the system first selects the target attack category, then selects the corresponding arbitration ID, vehicle status, and DLC from the real samples of that category as conditional constraints, and performs reverse denoising sampling starting from random noise loads. Each sampling step combines the CAN constraint operator to project the generated load, ultimately obtaining legitimate CAN minority class augmented samples. The generated samples are merged with the original training samples to form an augmented training set with a more balanced class distribution.

[0044] Specifically, by setting a preset absolute quantity threshold τ, attack categories with fewer than the threshold τ can be identified as minority attack categories.

[0045] S3. Construct the prototype calibration CAN Transformer classification model PCCT and train the PCCT model using the enhanced training set. The prototype calibration CAN Transformer classification model PCCT is used to map the arbitration ID, vehicle status, DLC, effective byte mask and payload bytes to a unified feature space. It extracts intra-frame byte dependencies and state context information through the Transformer encoder and outputs hierarchical discrimination results through the main classification head, class prototype distance calibration head and training auxiliary classification head.

[0046] like Figure 3As shown, the classification model PCCT first processes the payload byte input, valid byte mask, and context input to obtain payload byte embedding and position encoding, valid byte mask mapping, context embedding, and input token fusion layer.

[0047] The following sections explain the processing of payload byte input, valid byte mask, and context input.

[0048] The payload byte embedding and position encoding are obtained. PCCT uses 8 payload bytes from each CAN frame as byte token input. The value of each payload byte ranges from 0 to 255, and is first mapped to a fixed-dimensional vector representation through a byte embedding layer. Since different byte positions have different semantics in the CAN payload, the system further superimposes position encoding on each byte token, enabling the model to distinguish different payload byte positions. Let the 8-byte payload of a CAN frame be: The valid byte mask is: Arbitration ID is The vehicle status is DLC is The byte embedding layer maps each byte to a byte vector and overlays positional encoding: .in, For byte embedding functions, For the first The position encoding of each byte position.

[0049] Obtain the valid byte mask mapping. The valid byte mask is used to indicate the payload location corresponding to the DLC. PCCT injects the valid byte mask into the corresponding byte token after linear mapping, enabling the model to distinguish between the true payload location and the padding location outside the DLC. Through this design, the Transformer encoder can reduce the interference of invalid padding bytes on the classification results when modeling byte dependencies. Valid byte mask injected into the corresponding byte token after linear mapping: .in, and These are the mask mapping parameters.

[0050] Obtain the context embedding. The context input includes the arbitration ID, vehicle status, and DLC. PCCT embeds and encodes the arbitration ID, vehicle status, and DLC separately, and then fuses them to form a frame-level context vector. This context vector represents the message semantic environment to which the current CAN frame belongs, ensuring that subsequent byte tokens not only contain payload value information but also the corresponding ID, status, and length information. The context embedding layer encodes the arbitration ID, vehicle status, and DLC separately to form the frame-level context vector: .in, , , These are respectively: Arbitration ID embedding, Vehicle Status embedding, and DLC embedding.

[0051] The token fusion layer fuses payload byte embedding with position encoding, valid byte mask mapping, and context embedding to form a context-enhanced byte token sequence. In one implementation, a frame-level context vector is injected into each byte token, making each token simultaneously contain byte value, byte position, validity, and frame-level protocol context information. The token fusion layer injects the frame-level context vector into each byte token: This results in a context-enhanced token sequence: .

[0052] The Transformer encoder models the byte token sequence to obtain an enhanced token representation, which is then used to obtain a CAN frame-level feature vector through pooling. The Transformer encoder includes a self-attention layer, a feedforward network, and a normalization layer to capture the dependencies between different payload bytes, as well as the contextual relationships between payload bytes and arbitration ID, vehicle status, and DLC. Since a typical CAN payload contains 8 bytes, in one implementation, the Transformer encoder takes 8 context-enhanced byte tokens as input and outputs 8 enhanced token representations. After average pooling, attention pooling, or equivalent aggregation operations, the multiple enhanced tokens output by the Transformer encoder yield a comprehensive feature representation of the entire CAN frame, i.e., the CAN frame-level feature vector. This feature vector is used for subsequent multi-layer classification heads and class-prototype distance calibration heads for discrimination. The Transformer encoder models the token sequence as follows: .in, This represents the 8 enhanced tokens output by the Transformer. A pooling operation is then performed to obtain the CAN frame-level feature vector. .

[0053] The main classification head outputs the main classification probability based on the CAN frame-level feature vector. The main classification head receives the CAN frame-level feature vector and outputs fine-grained class probabilities to determine whether the current CAN frame belongs to the normal category or a certain abnormal category. The main classification head is the fundamental discriminant branch for PCCT's final classification. The main classification head outputs fine-grained class probabilities: Normal / Attack binary classification head outputs binary classification probabilities: Attack family classification head output attack family probability: , This is the weight matrix. For bias.

[0054] The training auxiliary classification head includes a normal / attack binary classification head and an attack family classification head. The normal / attack binary classification head outputs binary classification probabilities based on CAN frame-level feature vectors, while the attack family classification head outputs attack family probabilities based on CAN frame-level feature vectors. The normal / attack binary classification head is used during the training phase to constrain the CAN frame-level feature vectors to have a coarse-grained ability to distinguish between normal and abnormal categories; the attack family classification head is used during the training phase to constrain the formation of a clearer, fine-grained structure among abnormal categories. The training auxiliary classification head participates in loss calculation during the training phase and updates the parameters of the Transformer encoder through backpropagation; however, it does not need to directly participate in the final probability fusion during the inference phase.

[0055] The class prototype distance calibration head includes learnable prototype vectors corresponding to the number of attack categories. It is used to obtain prototype distance scores based on CAN frame-level feature vectors and learnable prototype vectors, and outputs the prototype probabilities from these distance scores. Through class prototype distance calibration, the model can bring similar samples closer to their corresponding class prototypes in the latent space, while keeping samples from different classes further apart, thereby enhancing the separability of minority class anomalies. Let the prototype vector of the k-th class be... The prototype distance score is: The prototype probability is: .

[0056] The PCCT training phase employs a multi-task joint loss function: .in, Primary classification loss, For prototype distance calibration loss, The loss is classified as normal / attack. Classify the losses for the attacking group. These are the preset weighting coefficients. Each loss can be expressed as: , , , Where y is the true category label, For normal / attack indication labels, The label is for the attack group; when the sample is of the normal class, it is not calculated. Or reset its rights to zero.

[0057] PCCT is trained on an enhanced training set, and its total loss is a weighted sum of the main classification loss, class prototype distance calibration loss, normal / attack binary classification loss, and attack family classification loss. Through multi-task joint training, PCCT can simultaneously learn fine-grained class discrimination, normal / abnormal coarse-grained boundaries, and latent space prototype structures, thereby improving the ability to detect anomalous traffic in a minority of classes.

[0058] For the trained PCCT model, a test set is used for prediction and evaluation. Each CAN sample first undergoes byte embedding, mask mapping, and context embedding to form a context-enhanced token sequence, which is then processed by a Transformer encoder to obtain the CAN frame-level feature vector. Subsequently, the main classification head outputs the main classification probability, and the class prototype distance calibration head outputs the prototype probability. The system jointly calibrates the main classification probability and the prototype probability to obtain the final classification probability, and selects the class with the highest probability as the final detection result. Finally, the model performance is evaluated using metrics such as accuracy, precision, recall, F1 score, and confusion matrix.

[0059] S4. Based on the trained prototype, calibrate the CAN Transformer classification model, determine the output class probabilities, and select the attack class with the highest probability as the final detection result.

[0060] During the inference phase, the final classification probability of PCCT is obtained by jointly calibrating the main classification probability and the prototype probability: .in, To preset the fusion weights, This indicates a normalization operation. Final choice: As a result of the test.

[0061] The following combination Figure 1 A specific application scenario is provided to verify the effectiveness of the embodiments of the present invention. The steps are as follows: 1. Dataset Selection: This embodiment uses the Car Hacking: Attack & Defense Challenge 2020 dataset as the experimental dataset. This Car-Hacking dataset originates from the 2020 In-Vehicle Network Attack & Defense Challenge, collected from the CAN bus of a real vehicle's Controller Area Network (CAN). It includes background CAN traffic under normal driving conditions, as well as typical attack traffic constructed by injecting abnormal messages into the CAN bus. Each CAN frame record in the dataset includes fields such as timestamp, arbitration ID, DLC, vehicle status, 8-byte payload, valid byte flag, and category label, reflecting the differences between normal and abnormal messages in in-vehicle CAN communication in terms of protocol structure, payload distribution, and state context. This embodiment divides the detection categories into normal and various abnormal attack categories to verify the effectiveness of the present invention in detecting abnormal traffic in in-vehicle CAN and in scenarios of class imbalance.

[0062] 2. Data Preprocessing: Traverse the raw CAN traffic data, removing records with missing category labels, abnormal DLC fields, missing payload fields, or physical fields exceeding preset ranges; limit the DLC field to the range of 0 to 8, limit the 8 payload byte fields to the range of 0 to 255, and limit the valid byte flag to 0 or 1; generate or correct an 8-dimensional valid byte mask based on the DLC field, marking invalid payload positions outside the DLC as 0; uniformly convert character-type arbitration IDs to uppercase strings, and construct arbitration ID encoders, vehicle status encoders, and category label encoders respectively, converting discrete fields into numerical codes that the model can input. While maintaining the proportion of normal and abnormal categories, the dataset is hierarchically divided to obtain training and test sets.

[0063] 3. Construction and Training of SC-CPD: Based on the idea of ​​state-constrained counterfactual payload diffusion generation, an SC-CPD model is constructed for minority class attack sample augmentation. This model takes the target attack category, arbitration ID, vehicle state, and diffusion time step as conditional inputs, and an 8-byte payload vector and a valid byte mask as generation constraint information. During the training phase, the real payload bytes are first normalized to a continuous value range, and noise is added under random diffusion time steps to obtain a noisy payload vector. Then, the noisy payload, category embedding, arbitration ID embedding, vehicle state embedding, time step embedding, and valid byte mask are input into the conditional denoising backbone network. SC-CPD learns the continuous denoising process through a noise prediction head, learns the byte value distribution from 0 to 255 through a byte discrete prediction head, and ensures that the generated payload meets the requirements of byte range, DLC valid mask, and state consistency through the CAN constraint operator. The training objective is composed of a weighted average of diffusion denoising loss, byte discrete cross-entropy loss, CAN constraint loss, semantic distribution loss, and counterfactual target loss.

[0064] 4. Enhanced Training Set Generation: After SC-CPD training is completed, its model parameters are fixed, and oversampling is performed on anomaly categories with fewer samples in the training set. Specifically, the system selects corresponding real samples as condition sources according to the target attack category, keeping their arbitration ID, vehicle state, and DLC fields unchanged. Multi-step reverse denoising sampling is performed starting from random noise payloads, and projection constraints are applied to the generated payloads using CAN constraint operators after each sampling step. This ultimately generates minority class enhanced samples that satisfy the physical constraints and state semantic consistency of the CAN protocol. The generated samples are then merged with the original training samples to form an enhanced training set with a more balanced class distribution.

[0065] 5. Construction and Training of the PCCT Model: The Prototype Calibration CAN Transformer Classification Model (PCCT) was constructed as the final detection model. This model uses 8 payload bytes as byte token inputs, obtaining byte-level representations through a byte embedding layer and a position encoding layer. The effective byte mask is linearly mapped and injected into the corresponding byte token to distinguish between payload and padding positions. The arbitration ID, vehicle status, and DLC are embedded and encoded separately, fused to form a frame-level context vector, and injected into each byte token. Subsequently, the context-enhanced byte token sequence is input into the Transformer encoder, where a self-attention mechanism is used to extract intra-frame byte dependencies and field context relationships, which are then pooled to obtain the CAN frame-level feature vector. During the training phase, PCCT undergoes multi-task joint optimization using a main classification head, a normal / attack binary classification head, an attack family classification head, and a class prototype distance calibration head. This allows the model to simultaneously learn fine-grained class discrimination, coarse-grained normal / abnormal boundaries, and the latent space prototype structure.

[0066] 6. Experimental Setup: The Car-Hacking dataset was divided into training and test sets in a 7:3 ratio. Accuracy, F1-score, and average inference time were used as the main evaluation metrics. Furthermore, to verify the effectiveness of each module of this invention, an ablation experiment was conducted in this embodiment. The ablation experiment included: First, removing the SC-CPD generation enhancement module and training the PCCT model using only the original training set; Second, removing the CAN constraint operator and using a generation model without byte range constraints, DLC effective mask constraints, and state consistency constraints for enhancement; Third, removing the counterfactual target loss and retaining only the diffusion denoising loss, byte discretization loss, and constraint loss; Fourth, removing the class prototype distance calibration head and using only the PCCT main classification head for prediction; Fifth, using the complete SC-CPD and PCCT combined model as the complete method of this invention.

[0067] 7. Experimental results: Tables 1 and 2 show the results based on the Car-Hacking dataset and the ablation experiment, respectively.

[0068] As shown in Table 1, the comparative experimental results demonstrate that the method proposed in this invention achieves optimal performance in vehicle network intrusion detection tasks. In terms of accuracy, precision, recall, and weighted average F1 score, this invention achieves 99.60%, 99.62%, 99.58%, and 99.60%, respectively, comprehensively outperforming the comparative methods.

[0069] Compared to baseline methods, this invention achieves an improvement of approximately 2.5 percentage points over KNN and also shows significant advantages over deep learning methods DBN-LSTM and LSTM-AE. This is mainly due to two core technological innovations in the scheme: State-Constrained Counterfactual Payload Diffusion (SC-CPD), which effectively alleviates the data imbalance problem by learning the conditional distribution of the minority class and combining it with physical constraints for sample augmentation; and Prototype Calibration CAN Transformer (PCCT), which employs a hierarchical classification head and prototype calibration mechanism to ensure that minority attack classes maintain good separability even during training dominated by normal frames.

[0070] Table 1. Performance comparison with existing methods on the Car-Hacking dataset.

[0071] As shown in Table 2, the experimental data reveals that the model's classification performance decreased to varying degrees after each module was removed. Specifically, removing the generation enhancement module resulted in a 0.30 percentage point decrease in accuracy, indicating that the enhanced samples generated by the diffusion model effectively improved the learning performance for a few attack categories. Removing the CAN constraint operator further reduced accuracy by 0.10 percentage points, demonstrating that protocol-level constraints such as byte range, DLC mask, and state consistency enhanced the authenticity of the generated data. Removing the counterfactual loss resulted in a 0.30 percentage point decrease in accuracy, reflecting the effectiveness of this loss function in guiding generated samples to approximate the target attack prototype. Removing the prototype calibration head still resulted in a 0.20 percentage point decrease in accuracy, indicating that prototype feature calibration and hierarchical classification strategies help enhance the separability of the feature space. Overall, there is a synergistic effect among the SC-CPD generation enhancement module, CAN constraint operator, counterfactual target loss, and prototype calibration head. The complete combination scheme achieved optimal performance in accuracy, precision, recall, and weighted F1 score, with an accuracy of 99.60%.

[0072] Table 2 shows the ablation experimental results on the Car-Hacking dataset.

[0073] The above embodiments are only used to provide a detailed description of the technical solutions of this application. However, the descriptions of the above embodiments are only for the purpose of helping to understand the methods of the embodiments of the present invention and should not be construed as limiting the embodiments of the present invention. Any variations or substitutions that can be easily conceived by those skilled in the art should be covered within the protection scope of the embodiments of the present invention.

Claims

1. A constraint generation and multi-layer discrimination method for abnormal CAN bus traffic in vehicles, characterized in that, include: The vehicle CAN bus traffic data is acquired as the vehicle network traffic dataset, and the dataset is preprocessed and divided into training and test sets according to a preset ratio. The fields of the vehicle CAN bus traffic data include timestamp, arbitration ID, DLC, vehicle status, payload bytes, valid byte mask, and attack category label; A state-constrained counterfactual load diffusion generation model is constructed, and the trained state-constrained counterfactual load diffusion generation model is used to conditionally generate attack categories with a small number of samples in the training set to obtain an enhanced training set. A prototype calibration CAN Transformer classification model is constructed and trained using the enhanced training set. The prototype calibration CAN Transformer classification model is used to map arbitration ID, vehicle status, DLC, effective byte mask and payload bytes to a unified feature space. Intra-frame byte dependency and state context information are extracted through the Transformer encoder, and the probability of each attack category is output through the main classification head, the prototype distance calibration head and the training auxiliary classification head. The attack category with the highest probability is selected as the final detection result based on the class probability output by the prototype calibration CAN Transformer classification model after training.

2. The method according to claim 1, characterized in that, State-constrained counterfactual load diffusion generation models include: The conditional fusion layer is used to concatenate the noisy payload, effective byte mask, category embedding vector, arbitration ID embedding vector, vehicle state embedding vector, and time step embedding vector to obtain the conditional fusion vector. Among them, the noisy payload is obtained by adding diffused noise to the payload bytes, the category embedding vector, arbitration ID embedding vector, and vehicle state embedding vector are obtained by embedding and encoding the category label, arbitration ID, and vehicle state respectively, and the time step embedding vector is obtained by performing nonlinear mapping on the diffused time step. The conditional denoising backbone network is used to extract conditional denoising features from the conditional fusion vector; A noise prediction head is used to output a continuous noise estimate based on conditional denoising features; The byte discrete prediction header is used to output a discrete probability distribution of values ​​from 0 to 255 at each valid byte position based on the conditional denoising features. The CAN constraint operator is used to project the generated payload obtained from the noise prediction header and the byte discrete prediction header onto the feasible region that satisfies the CAN protocol constraints.

3. The method according to claim 2, characterized in that, The CAN constraint operator is as follows: , Among them, continuous load estimation , For noisy loads, For continuous noise estimation, The conditional denoising backbone network is based on the conditional denoising features extracted from the conditional fusion vector. For weights and biases, For the Sigmoid function, the truncation function, or an equivalent bounded function, This indicates element-wise multiplication, where m is the effective byte mask.

4. The method according to claim 3, characterized in that, The total loss function of SC-CPD is: , Among them, diffusion denoising loss , The noise is Gaussian, and m is the effective byte mask; Byte Discrete Cross-Entropy Loss , For the first Each actual byte value; The CAN constraint loss is: ; Semantic distribution loss ,in, and These are the byte mean and standard deviation of the generated sample within the g-th semantic group, respectively. and These are the byte mean and standard deviation of the real sample within this semantic group, respectively. This is the balance coefficient; Counterfactual target loss Among them, the generated samples are The prototype of the target attack category is The normal category prototype is The effective byte mask is m, and the effective byte weighted distance is... , This is the interval threshold. The normal prototype is far from the constraint weights; These are preset weighting coefficients.

5. The method according to claim 3, characterized in that, Prototype calibration of the CAN Transformer classification model, including: The token fusion layer is used to fuse payload byte embedding with position encoding, valid byte mask mapping, and context embedding to form a context-enhanced byte token sequence. The transformer encoder is used to model the byte token sequence to obtain an enhanced token representation, and then obtains the CAN frame-level feature vector through pooling operations; The main classification header outputs the main classification probability based on the CAN frame-level feature vector. Train auxiliary classification heads, including normal / attack binary classification heads and attack family classification heads. The normal / attack binary classification heads output binary classification probabilities based on CAN frame-level feature vectors, and the attack family classification heads output attack family probabilities based on CAN frame-level feature vectors. The prototype distance calibration head includes learnable prototype vectors corresponding to the number of attack categories. It is used to obtain prototype distance scores and output prototype probabilities based on CAN frame-level feature vectors and learnable prototype vectors.

6. The method according to claim 5, characterized in that, The final classification probability output by the prototype calibration CAN Transformer classification model is obtained by jointly calibrating the main classification probability and the prototype probability, specifically as follows: , in, To preset the fusion weights, This indicates a normalization operation. Primary classification probability, This is the weight matrix. For bias, For prototype probability, The prototype distance score is given, where h represents the CAN frame-level feature vector. Represents the prototype vector of the k-th class; This serves as the final detection result.

7. The method according to claim 5, characterized in that, Total training loss of the prototype calibration CAN Transformer classification model ; Among them, the primary classification loss y represents the true category label. Represents the cross-entropy loss function; Prototype Distance Calibration Loss ; Normal / Attack Binary Classification Loss , , Labels indicating normal / attack status; Attack family classification loss , , Tag for attacking groups; This is the weight matrix. For bias; These are the preset weighting coefficients.