A lightweight internet of vehicles intrusion detection method and system based on isomorphic knowledge distillation
Patent Information
- Application Number
- CN202611317456.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]为解决现有技术中车载入侵检测大模型算力开销大、复杂度高而无法在车载边缘部署,以及小模型轻量化压缩技术导致隐蔽威胁感知精度(尤其是召回率指标)大幅坍塌的问题,本发明提供一种基于同构知识蒸馏的轻量化车联网入侵检测方法及系统,旨在提出一种兼顾线性复杂度基座构建与同构暗知识蒸馏量化的联合训练主动防御方案(ViLKD-IDS),在不损失多维攻击特征(宏观纹理与微观比特跳变)的前提下,保证入侵检测精度,同时显著降低模型参数规模、显存开销与推理时延,实现模型轻量化,有效提升其在车载资源受限环境下的实时部署能力
1.大幅度的模型参数压缩:本发明通过同构暗知识迁移,将模型参数总量由教师模型的 22.31 M锐减至 5.85 M,参数缩减比例高达 73.78%(实现 3.81 倍的规模压缩),大幅降低了车载微处理器的存储和 SRAM 缓存占用。
Smart Images

Figure CN122845295A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle network security technology, specifically relating to a lightweight vehicle network intrusion detection method and system based on isomorphic knowledge distillation. More particularly, it relates to a lightweight distributed vehicle intrusion detection method and system based on high-fidelity reconstruction of multi-dimensional feature spatiotemporal images, construction of a linear complexity visual extended long short-term memory network (Vision-xLSTM) baseline architecture, distillation of general visual prior isomorphic dark knowledge, and spatial quantization perception optimization. Background Technology
[0002] Against the backdrop of the rapid development of intelligent connected vehicle technology, the security of internal and external communications within the Internet of Vehicles (IoV) has become a core element in ensuring reliable vehicle operation and protecting the lives of occupants. However, traditional intrusion detection systems (IDS) face numerous technical bottlenecks when processing Controller Area Network (CAN) bus traffic: traditional methods rely excessively on labor-intensive manual feature engineering, making it difficult to capture complex nonlinear spatiotemporal characteristics in the traffic; existing algorithms lack sensitivity to long-distance temporal dependencies, easily leading to missed detections of covert collaborative attacks. Furthermore, while existing deep learning intrusion detection models have significantly improved performance in terms of detection accuracy, they still generally suffer from high computational costs, difficulty in deployment on in-vehicle edge devices, or a surge in missed detection rates for covert attacks after conventional lightweighting.
[0003] To address the aforementioned pain points, this invention utilizes the Vision-xLSTM model and the isomorphic knowledge distillation framework to research and design a lightweight intrusion detection scheme for the Internet of Vehicles. Summary of the Invention
[0004] To address the issues of high computational cost and complexity of large-scale vehicle intrusion detection models, which prevent their deployment at the vehicle edge, and the significant collapse in the accuracy of covert threat perception (especially recall rate) caused by lightweight compression techniques for small models, this invention provides a lightweight vehicle-to-everything (V2X) intrusion detection method and system based on isomorphic knowledge distillation. It aims to propose a joint training active defense scheme (ViLKD-IDS) that balances linear complexity base construction with isomorphic dark knowledge distillation quantization. This scheme ensures intrusion detection accuracy without sacrificing multidimensional attack features (macro-texture and micro-bit transitions), while significantly reducing model parameter size, memory overhead, and inference latency, thus achieving model lightweighting and effectively improving its real-time deployment capability in resource-constrained vehicle environments.
[0005] To achieve the above objectives, the present invention provides the following solution: A lightweight intrusion detection method for connected vehicles based on isomorphic knowledge distillation, the method comprising: Based on multi-dimensional feature fusion, robust quantile normalization and spatial topology reconstruction mapping, the discrete CAN bus message sequence acquired in real time is reconstructed into a three-channel pseudo-color spatiotemporal image. Based on the Vision-xLSTM architecture, a benchmark intrusion detection network, ViL-IDS, is constructed as a teacher model; Based on the Vision-xLSTM architecture and isomorphic knowledge distillation, a lightweight intrusion detection network ViLKD-IDS is constructed as a student model; The student model is trained using the teacher model to obtain the trained lightweight intrusion detection network ViLKD-IDS; The three-channel pseudo-color spatiotemporal image is input into the trained lightweight intrusion detection network ViLKD-IDS to obtain intrusion detection results.
[0006] Preferably, the method for reconstructing a three-channel pseudo-color spatiotemporal image from a real-time acquired discrete CAN bus message sequence based on multi-dimensional feature fusion, robust quantile normalization, and spatial topology reconstruction mapping includes: The discrete CAN bus message sequence is segmented using a sliding window to obtain overlapping CAN data blocks, wherein each CAN data block contains multiple consecutive frames of CAN message messages. Extract the identifier field, complete data payload field, high-precision global timestamp, and frequency indicator reflecting injection density of each CAN message in the CAN data block, and perform multi-dimensional feature fusion to construct the original feature matrix of the CAN data block; The original feature matrix is transformed into a three-channel pseudo-color spatiotemporal image with spatial topological consistency by employing robust quantile normalization and spatial topological reconstruction mapping.
[0007] Preferably, the method for training the student model using the teacher model to obtain the trained lightweight intrusion detection network ViLKD-IDS includes: The teacher model was pre-trained using a general natural image dataset; By utilizing a joint knowledge distillation mechanism based on relative entropy and cross entropy, the soft label distribution knowledge output by the pre-trained teacher model is transferred to the student model. The weight matrix of the student model is subjected to fixed-point truncated symmetric quantization mapping to obtain the trained lightweight intrusion detection network ViLKD-IDS.
[0008] Preferably, the method for inputting the three-channel pseudo-color spatiotemporal image into the trained lightweight intrusion detection network ViLKD-IDS to obtain intrusion detection results includes: The three-channel pseudo-color spatiotemporal image generated by reconstructing the discrete CAN bus message sequence is divided into image blocks to form an image block sequence. Linear projection is performed on each image block, and positional encoding is superimposed to obtain the initial feature identifier sequence of the image block sequence; The initial feature identifier sequence is input into the Vision-xLSTM encoder for feature modeling to obtain the final feature identifier sequence; By fusing the preset image block features of the final feature identifier sequence, a fused feature is obtained, wherein the preset image block features include the first image block feature and the last image block feature; The fused features are input into the classification layer, and after linear classification and mapping by the Softmax function, the classification prediction probability is obtained, thereby realizing intrusion detection.
[0009] The present invention also provides a lightweight vehicle network intrusion detection system based on isomorphic knowledge distillation. The system is used to implement the aforementioned method and includes: an image reconstruction module, a teacher module, a student module, a training module, and a detection module. The image reconstruction module is used to reconstruct a three-channel pseudo-color spatiotemporal image from a real-time acquired discrete CAN bus message sequence based on multi-dimensional feature fusion, robust quantile normalization, and spatial topology reconstruction mapping. The teacher module is used to construct a benchmark intrusion detection network ViL-IDS as a teacher model based on the Vision-xLSTM architecture. The student module is used to construct a lightweight intrusion detection network ViLKD-IDS as a student model based on the Vision-xLSTM architecture and isomorphic knowledge distillation. The training module is used to train the student model using the teacher model to obtain the trained lightweight intrusion detection network ViLKD-IDS. The detection module is used to input the three-channel pseudo-color spatiotemporal image into the trained lightweight intrusion detection network ViLKD-IDS to obtain intrusion detection results.
[0010] Preferably, the image reconstruction module includes: a CAN data segmentation unit, a first feature fusion unit, and a conversion and reconstruction unit; The CAN data segmentation unit is used to segment the discrete CAN bus message sequence using a sliding window to obtain overlapping CAN data blocks, wherein each CAN data block contains multiple consecutive CAN message frames. The first feature fusion unit is used to extract the identifier field, complete data payload field, high-precision global timestamp, and frequency indicator reflecting injection density of each frame of CAN message in the CAN data block, and perform multi-dimensional feature fusion to construct the original feature matrix of the CAN data block. The transformation and reconstruction unit is used to convert the original feature matrix into a three-channel pseudo-color spatiotemporal image with spatial topological consistency by employing robust quantile normalization and spatial topological reconstruction mapping.
[0011] Preferably, the training module includes: a pre-training unit, a knowledge distillation unit, and a truncation quantization unit; The pre-training unit is used to pre-train the teacher model using a general natural image dataset; The knowledge distillation unit is used to transfer the soft label distribution knowledge output by the pre-trained teacher model to the student model using a joint knowledge distillation mechanism based on relative entropy and cross entropy. The truncated quantization unit is used to perform fixed-point truncated symmetric quantization mapping on the weight matrix of the student model to obtain the trained lightweight intrusion detection network ViLKD-IDS.
[0012] Preferably, the detection module includes: an image segmentation unit, a precoding unit, a sequence coding unit, a second feature fusion unit, and a classification unit; The image segmentation unit is used to divide the three-channel pseudo-color spatiotemporal image generated by reconstructing the discrete CAN bus message sequence into image blocks to form an image block sequence. The precoding unit is used to perform linear projection on each image block and superimpose positional codes to obtain an initial feature identifier sequence of the image block sequence; The sequence encoding unit is used to input the initial feature identifier sequence into the Vision-xLSTM encoder for feature modeling to obtain the final feature identifier sequence; The second feature fusion unit is used to fuse preset image block features of the final feature identifier sequence to obtain fused features, wherein the preset image block features include the first image block feature and the last image block feature; The classification unit is used to input the fused features into the classification layer, and obtain the classification prediction probability after linear classification and softmax function mapping to realize intrusion detection.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Significant model parameter compression: This invention reduces the total number of model parameters from 22.31 M in the teacher model to 5.85 M through isomorphic dark knowledge transfer, with a parameter reduction ratio of up to 73.78% (achieving a 3.81-fold scale compression), which greatly reduces the storage and SRAM cache usage of the vehicle microprocessor.
[0014] 2. High-efficiency inference latency (hard real-time performance): Due to the linear computational complexity and quantization-aware optimization of Vision-xLSTM, the average inference latency per data packet is compressed to the level of 1.165 ms, and the inference energy efficiency is significantly improved by 25.2%, which fully meets the real-time defense requirements of high-frequency communication on the vehicle bus.
[0015] 3. Improved anomaly detection recall against the trend: When faced with the 2020 Attack & Defense Challenge dataset, which is characterized by highly imbalanced data and complex attack patterns, thanks to the high-dimensional visual prior provided by the teacher model, the student model of this invention significantly reduced its size while increasing the anomaly detection recall from 81.38% of the ViL-IDS benchmark model to 90.85%, effectively alleviating the defect that "traditional lightweighting inevitably leads to a surge in false negatives". Attached Figure Description
[0016] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the overall workflow of the ViLKD-IDS intrusion detection method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the data preprocessing process according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the three-stage process of homogeneous knowledge transfer for vehicle networks according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the Vision-xLSTM lightweight model training architecture based on isomorphic knowledge distillation, according to an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] Example 1 This invention provides a lightweight intrusion detection method for connected vehicles based on isomorphic knowledge distillation, comprising: Based on multi-dimensional feature fusion, robust quantile normalization and spatial topology reconstruction mapping, the discrete CAN bus message sequence acquired in real time is reconstructed into a three-channel pseudo-color spatiotemporal image. Based on the Vision-xLSTM architecture, a benchmark intrusion detection network, ViL-IDS, is constructed as a teacher model; Based on the Vision-xLSTM architecture and isomorphic knowledge distillation, a lightweight intrusion detection network ViLKD-IDS is constructed as a student model; The student model is trained using the teacher model to obtain the trained lightweight intrusion detection network ViLKD-IDS; The three-channel pseudo-color spatiotemporal image is input into the trained lightweight intrusion detection network ViLKD-IDS to obtain intrusion detection results.
[0021] The lightweight vehicle network intrusion detection framework designed in this invention mainly includes the following core entities: 1. On-board unit (OBU) / On-board electronic control unit (ECU): Deployed inside intelligent connected vehicles, with limited computing resources, it is responsible for collecting real-time CAN bus traffic, obtaining discrete CAN bus message sequences, and running an extremely lightweight student model (ViLKD-IDS) for online hard real-time active defense.
[0022] 2. Vehicle-to-Everything (V2X) Control Center / Central Computing Platform (TA): Possesses high-performance computing resources and is responsible for pre-training and refining complex teacher models (ViL-IDS), generating lightweight student model weights through isomorphic knowledge distillation, and finally distributing them to the vehicle edge.
[0023] Under the premise of a homogeneous network with fully aligned underlying operators, a joint knowledge distillation mechanism based on relative entropy and cross-entropy is used to transfer the dark knowledge contained in the teacher model to the lightweight student model, and then fine-tunes it for the vertical domain of automotive protocols. On this basis, a fixed-point truncation symmetric quantization mapping strategy is introduced to further compress model parameters. The resulting lightweight intrusion detection network, ViLKD-IDS, is then deployed to the automotive edge microprocessor to perform high-efficiency, high-sensitivity hard real-time online defense.
[0024] Combination Figure 1 The specific implementation process of this invention is described below: Step 1: Environment Initialization Configure key parameters for the intrusion detection system on the vehicle-to-everything (V2X) central computing platform. Initialize the teacher and student models.
[0025] Step 2: Three-channel pseudo-color spatiotemporal image reconstruction based on multi-dimensional feature fusion, robust quantile normalization, and spatial topology reconstruction mapping The vehicle-mounted communication node or cloud gateway intercepts a continuous discrete CAN bus message sequence. To overcome the limitations of single-field dependency, a sliding window is used to segment the discrete CAN bus message sequence, resulting in overlapping CAN data blocks. Multidimensional features of each frame of CAN message in the CAN data block are extracted, and a high-dimensional spatiotemporal feature matrix covering message identifier, complete data payload field, high-precision global timestamp, and frequency indicator reflecting injection density is constructed as the original feature matrix of the CAN data block.
[0026] This invention abandons the min-max scaling method, which is extremely sensitive to extreme outliers, and instead employs robust quantile normalization, which has strict mathematical sorting logic, to normalize the original feature matrix. Furthermore, it utilizes spatial topological reconstruction mapping to transform the normalized feature matrix into one with spatial topological consistency. Three-channel pseudo-color spatiotemporal images.
[0027] This three-channel pseudo-color spatiotemporal image faithfully transforms the macroscopic "high-frequency flow oscillation" into the "high-contrast periodic stripe texture" in the image, and maps the microscopic "bit-level load tampering" into the "sharp single-pixel color difference edge" in visual space.
[0028] Step 3: Construct a benchmark system for Visual Extended Long Short-Term Memory Networks (ViL-IDS) with linear computational complexity. This step is based on the Vision-xLSTM architecture, and a high-performance benchmark intrusion detection network ViL-IDS is built on a high-performance computing platform as a teacher model.
[0029] Step 4: Refining the Teacher Model's General Visual Priors Will have A general-purpose natural image dataset of physical size is input into the teacher model ViL-IDS for pre-training. The encoder module parameters are frozen to preserve the ability to extract general spatiotemporal features. Then, progressive unfreezing and fine-tuning are performed on the target vehicle network CAN bus dataset.
[0030] Step 5: Distillation of Isomorphic Dark Knowledge Transfer Construct a student model ViLKD-IDS with a small parameter scale, ensuring that its underlying execution operators are completely isomorphically aligned with those of the teacher model ViL-IDS with a large parameter scale.
[0031] By utilizing a joint knowledge distillation mechanism that combines soft prediction relative entropy distillation loss and true label-hard prediction cross-entropy loss, the soft label distribution knowledge (dark knowledge) output by the teacher model after pre-training in step four is transferred to the student model.
[0032] During the training of vehicle network data, the difference in probability distribution between the teacher model and the student model in the Softmax output layer (smoothed by temperature coefficient) is calculated. The distillation loss function is constructed using relative entropy (KL divergence) to guide the student model to inherit the hidden knowledge contained in the teacher model while the number of parameters is sharply reduced.
[0033] Step Six: Weighted Truncation Symmetric Quantization Perceptual Training The weight matrix of the student model obtained in step five is subjected to fixed-point truncated symmetric quantization mapping, and the quantization loss function is optimized to obtain the trained lightweight intrusion detection network ViLKD-IDS.
[0034] Step 7: Edge Deployment and Online Hard Real-Time Detection The weights of the finally trained lightweight intrusion detection network ViLKD-IDS are deployed to the vehicle edge microprocessor. The OBU reconstructs the discrete CAN bus message sequences collected in real time into... The three-channel pseudo-color spatiotemporal image is then uniformly divided into non-overlapping segments of size [missing information]. The local image patch is flattened to obtain the length. N A sequence of 64 image patches is used, and then this sequence is fed into the lightweight detection model ViLKD-IDS. O ( N With a linear computational complexity, it completes single-packet inference within milliseconds and outputs normal or abnormal classification results.
[0035] Example 2 This embodiment combines Figures 2-4The specific implementation process of the method described in the foregoing embodiments is explained in detail below: I. Multidimensional Feature Fusion and Quantile Normalization Spatiotemporal Reconstruction To accurately capture the hidden nonlinear spatiotemporal characteristics in CAN traffic, this embodiment establishes a sliding window in the on-board unit (OBU) at the vehicle edge. The original discrete CAN bus message sequence is denoted as... ,in s l ( l =0,1,2,..., L ) indicates the first l Frame CAN message message, L This is the message sequence length. Assume the total length of the sliding window is... w Each sliding step is The sequence of data blocks extracted by the sliding window is denoted as... Each data block D i satisfy: .
[0036] At any moment ,right n A continuous data block Perform data transformation, such as Figure 2 As shown. First, regarding the first... i Data blocks D i For each frame of the message, extract its 29-bit or 11-bit identifier field (ID), 8-byte complete data payload, high-precision global timestamp, and frequency indicator reflecting the injection density to form a 1× d The single-frame feature row vector, d This represents the characteristic dimensions of a single CAN message frame. Next, the data blocks... D i All w The feature vectors of the frame message are stacked in chronological order to form a data block. D i The original feature matrix (It is worth considering) w =27、 d =9 (Scenario).
[0037] To avoid the problem that the min-max scaling method is too sensitive to extreme anomaly attack values, this invention introduces robust quantile normalization for the original feature matrix. F i Perform normalization mapping to obtain the normalized feature matrix. Then, according to the preset channel mapping rules, The image is reconstructed into a three-channel pseudo-color image of a preset size (e.g., 9×9×3); further, a spatial topology reconstruction strategy is employed to map the three-channel pseudo-color image to a preset input size. (For example, 32×32×3, where, H For high, W For width, C Three-channel pseudo-color spatiotemporal image (number of channels) This improves image resolution while preserving the spatial topological relationships of the original features.
[0038] For three-channel pseudo-color spatiotemporal images Perform the following preprocessing: Divide it uniformly into spatial resolutions of [missing information]. (For example, 4×4, P Given an image patch (with a side length of [number]), obtain [the image patch]. Image patches, image patch sequence X i Recorded as: .
[0039] Obviously Then, each image patch... ( j =0,1,2,..., N ) Flatten The row vectors are projected by the linear projection matrix. Mapped to M 3D embedding space, concatenating globally learnable classification label vectors and superimposed position encoding matrix The initial feature identifier sequence of the image patch sequence is obtained. : .
[0040] Initial feature identifier sequence It is used as input to the Vision-xLSTM encoder for deep feature modeling.
[0041] In summary, using the above methods, the macroscopic "high-frequency flow oscillation" in the discrete CAN bus message sequence is transformed into the "contrast texture" of the image, and the microscopic "bit-level modification" is transformed into "single-pixel jump edge".
[0042] II. Linear Complexity Backbone Architecture Based on Vision-xLSTM To eliminate the quadratic computational complexity inherent in the global self-attention mechanism of the traditional Transformer architecture This embodiment employs the Vision-xLSTM architecture. Each mLSTM module in the Vision-xLSTM encoder uses a forward linear recursive state update mechanism. Under the control of the forget gate and the input gate, the memory matrix is updated stepwise recursively. The update equation is as follows: ; in, C t This indicates the mLSTM module processing the input feature identifier sequence of the th element. t The hidden state matrix after each feature identifier; C t-1 This indicates the mLSTM module processing the input feature identifier sequence of the th element. t- The hidden state matrix after one feature identifier; f t and i t These are the forgetting gate scalar and the input gate scalar, respectively, reconstructed using the exponential activation function; v t and k t These are the value and key vectors of the current input, respectively, and the feature identifier sequence from the input to the current mLSTM module. (The input to the first mLSTM module is) The first in ) t The first feature identifier (i.e., the first feature identifier) t (Line) generated; Representing vectors k t The transpose of the current mLSTM module outputs a new feature identifier sequence after encoding. (where h) class h1, h2, ..., h N Each of these represents a new feature identifier, which is then used as the input to the next mLSTM module. After encoding by all mLSTM modules, the final feature identifier sequence is obtained. The first image patch feature in the final feature identifier sequence is then used as the input to the next mLSTM module. and end image patch features The tokens are concatenated to obtain the fused features. This information is then passed to subsequent linear classification layers. The linear classification layers are processed by a learnable weight matrix. and learnable bias vector Its function is to obtain the unnormalized category score Logits output. ,in , z k Indicates the first kEach category score, k=1,2,... K , K This represents the total number of categories for the classification task.
[0043] The above update process does not require constructing a global self-attention matrix; instead, it uses a forward linear recursive scan to complete the state update. Therefore, the time and space complexity of its sequence modeling can be reduced to the same as the sequence length. N The linear relationship ensures the real-time response basis of the algorithm when processing high-frequency sequences of the vehicle bus.
[0044] III. Distillation of General Visual Priors and Isomorphic Relative Entropy Dark Knowledge like Figure 3 As shown, in the general visual prior pre-training stage, the teacher model (ViL-IDS) is first pre-trained on the physically aligned CIFAR-10 general natural image dataset to obtain rich visual representation capabilities. After pre-training, all parameters of the teacher model are frozen and kept unchanged throughout the knowledge distillation training process. Subsequently, in the knowledge distillation stage, the reconstructed three-channel pseudo-color spatiotemporal image is simultaneously input into both the frozen teacher model and the student model to be trained. The teacher model extracts high-dimensional visual prior features and soft labels to provide dark knowledge supervision for the student model; during training, only the student model parameters are updated, ultimately obtaining a student model that combines the advantages of general visual priors and lightweight deployment. Specifically, in the isomorphic knowledge distillation stage, dark knowledge transfer distillation is performed between the isomorphic student model (ViLKD-IDS), whose underlying operators and execution pipeline are fully aligned, and the pre-trained teacher model. Figure 4 As shown, assume that the unnormalized logistic column vectors (Logits) output by the ViL-IDS teacher model and the ViLKD-IDS student model are respectively and Where (t) represents the teacher model and (s) represents the student model. Assume the knowledge distillation temperature coefficient is... T After mapping using the temperature Softmax function, the soft predictions of Logits are obtained, which are in the form of a column vector of the distribution function. and : .
[0045] Distillation loss is defined using the Kullback-Leibler (KL) divergence between the output distribution functions of the teacher and student models in soft prediction: .
[0046] Knowledge of distillation temperature coefficient T It is often a constant greater than 1, when the temperature coefficient is set to... TWhen =1, the hard prediction result obtained by the student model is denoted as . Compare it with real labels y i By comparison, we define the cross-entropy loss: .
[0047] The weighted combination of the two yields the total loss function: , in As a weighting coefficient, it controls the proportion of the student model's absorption of the true distribution and the teacher model's soft prediction relative to the hidden knowledge, enabling the hidden knowledge contained in the teacher model to be transferred to the small-scale student model with high fidelity.
[0048] IV. Weighted Truncated Symmetric Quantization Perceptual Training To further compress hardware storage and reduce the algorithm's footprint on the vehicle edge hardware gateway's Flash and SRAM caches, this embodiment adjusts the weight matrix of the student model. The mathematical pseudo-quantization mapping function of the fixed-point truncation symmetric quantization mapping strategy is expressed as follows: , in, This is a scaling factor for dynamic statistics. Represents the weight matrix The total number of elements in the array; This is a boundary truncation function designed to forcibly constrain out-of-range abnormal weights within a specified range. Within the range; This represents the rounding operator to the nearest integer. The quantized student model weight matrix is denoted as... .
[0049] To measure and constrain the feature representation distortion caused by the quantization process, this embodiment compares the feature probability distributions generated by the full-precision student model and the quantized student model under the same input. and And calculate the second-order Wasserstein distance between them. Further optimize the quantization loss function based on the second-order Wasserstein distance: , in, This represents the set of learnable parameters of the network; This represents the area partition of a high-dimensional unit sphere. The formula utilizes the second-order Wasserstein distance. By dividing the area of a high-dimensional sphere Below are the network parameters Optimize and minimize the full-precision feature probability distribution. With quantized feature probability distribution The geometric differences between them are analyzed. The above quantization loss is optimized to obtain the trained lightweight intrusion detection network ViLKD-IDS, which is used for CAN message intrusion detection.
[0050] V. Experimental Results and Feasibility Verification To verify the effectiveness and engineering feasibility of the proposed solution, the ViLKD-IDS generated after training was deployed on a mobile NVIDIA RTX 4060 graphics card platform for simulation verification. The core purpose of choosing this platform is to simulate the actual operating environment of current mainstream edge-side in-vehicle intelligent driving computing units (such as the NVIDIA DRIVE OrinX platform with 254 TOPS computing power) with high fidelity through reasonable computing power constraints (both platforms are completely consistent in their underlying software ecosystem, and both natively support the standard CUDA core operator set and TensorRT inference acceleration engine).
[0051] A quantitative comparison of the parameter sizes of the teacher model (ViL-IDS) and the distilled student model (ViLKD-IDS) was performed. Model lightweighting not only directly affects the Flash storage usage of the vehicle ECU, but also determines the SRAM bandwidth pressure and dynamic power consumption during inference.
[0052] As shown in Table 1, the experiment statistically analyzed the total number of parameters, compression ratio, and parameter reduction ratio of the teacher and student models. Analysis of Table 1 shows that by introducing the isomorphic knowledge distillation paradigm, the lightweight scheme proposed in this invention achieves significant size reduction. Compared to the 22.31M parameter size of the basic ViL-IDS architecture, the parameter size of the distillation-optimized ViLKD-IDS model is reduced to 5.85M, a reduction of approximately 73.78% in the number of parameters, achieving a compression ratio of 3.81.
[0053] From an engineering implementation perspective, the significant reduction in the number of parameters means that the space occupied by the model weight file in persistent storage is reduced by about three-quarters, which greatly alleviates the resource requirements of typical automotive gateway controllers for non-volatile storage. Simultaneously, the lightweight model requires significantly fewer weight parameters to be loaded during inference, effectively reducing the instantaneous pressure on memory bus bandwidth. Experimental results show that this distillation strategy successfully retains the high-dimensional feature representation capabilities of the teacher model while significantly reducing model redundancy, laying the physical foundation for ViLKD-IDS to achieve hard real-time detection on computationally limited automotive embedded platforms.
[0054] Table 1. Comparison of total parameters, compression ratio, and storage space usage between the teacher and student models. To fully verify the effectiveness of the proposed Vision-xLSTM and isomorphic knowledge distillation joint scheme (ViLKD-IDS) in the actual vehicle networking environment, especially its detection performance and training and testing efficiency, it is also compared with the ViL-IDS scheme without distillation optimization: their experimental results are compared on three major benchmark datasets: Car-Hacking, M-CAN, and 2020 Attack & Defense Challenge.
[0055] The experimental results are shown in Table 2. Overall, the ViLKD-IDS optimized by knowledge distillation not only improves training efficiency and inference speed, but also effectively enhances recall in complex hybrid attack scenarios, alleviates the performance degradation caused by feature dimension compression during the lightweighting process, and achieves synergistic optimization of model lightweighting and detection performance.
[0056] Table 2 Experimental Results Specifically, on the Car-Hacking and M-CAN benchmark datasets, due to the high-frequency message injection characteristics of injection attacks, the three-channel pseudo-color spatiotemporal images generated by these attacks exhibit good feature separability. The ViLKD-IDS of this invention, while maintaining high detection accuracy, reduces inference latency by 11.4% and 25.2% respectively, effectively improving the model's real-time detection capability and applicability for vehicle deployment. On the complex, highly dynamic, extremely imbalanced data distribution, and covert collaborative tampering attack dataset 2020 Attack & Defense Challenge, traditional lightweight models are prone to performance degradation and missed anomaly samples due to model capacity limitations. The ViLKD-IDS scheme proposed in this invention employs an isomorphic knowledge distillation strategy, transferring the discriminative knowledge of the teacher model to the lightweight student model and utilizing distillation constraints to optimize its feature representation capability, thereby improving detection performance under complex attack scenarios while maintaining model lightweightness. Experimental results show that this invention not only reduces the average inference latency per packet to the 1.165 ms level and improves the global classification accuracy to 97.27%, but also increases the anomaly detection recall rate from 81.38% of the ViL-IDS benchmark model to 90.85%, effectively enhancing anomaly detection capabilities and vehicle-mounted real-time deployment performance in complex attack scenarios. Based on the above experimental analysis, the method proposed in this invention, through a combination of isomorphic knowledge distillation and quantized perception optimization, effectively reduces model inference latency and resource consumption while ensuring detection accuracy. It achieves a good balance between detection performance, real-time performance, and resource consumption, providing technical support for the lightweight deployment of vehicle-to-everything (V2X) intrusion detection systems.
[0057] Example 3 Based on the same inventive concept, the present invention also provides a lightweight vehicle network intrusion detection system based on isomorphic knowledge distillation, for implementing the method described in the foregoing embodiments. The system includes: an image reconstruction module, a teacher module, a student module, a training module, and a detection module. The image reconstruction module, based on multi-dimensional feature fusion, robust quantile normalization and spatial topology reconstruction mapping, reconstructs the real-time acquired discrete CAN bus message sequence into a three-channel pseudo-color spatiotemporal image. The teacher module is used to construct a benchmark intrusion detection network ViL-IDS as a teacher model based on the Vision-xLSTM architecture. The student module is used to construct a lightweight intrusion detection network ViLKD-IDS as a student model based on the Vision-xLSTM architecture and isomorphic knowledge distillation. The training module is used to train the student model using the teacher model to obtain the trained lightweight intrusion detection network ViLKD-IDS. The detection module is used to input the three-channel pseudo-color spatiotemporal image into the trained lightweight intrusion detection network ViLKD-IDS to obtain intrusion detection results.
[0058] Furthermore, in this embodiment, the image reconstruction module includes: a CAN data segmentation unit, a first feature fusion unit, and a conversion and reconstruction unit; The CAN data segmentation unit is used to segment the discrete CAN bus message sequence using a sliding window to obtain overlapping CAN data blocks, wherein each CAN data block contains multiple consecutive CAN message frames. The first feature fusion unit is used to extract the identifier field, complete data payload field, high-precision global timestamp, and frequency indicator reflecting injection density of each frame of CAN message in the CAN data block, and perform multi-dimensional feature fusion to construct the original feature matrix of the CAN data block. The transformation and reconstruction unit is used to convert the original feature matrix into a three-channel pseudo-color spatiotemporal image with spatial topological consistency by employing robust quantile normalization and spatial topological reconstruction mapping.
[0059] Furthermore, in this embodiment, the training module includes: a pre-training unit, a knowledge distillation unit, and a truncation quantization unit; The pre-training unit is used to pre-train the teacher model using a general natural image dataset; The knowledge distillation unit is used to transfer the soft label distribution knowledge output by the pre-trained teacher model to the student model using a joint knowledge distillation mechanism based on relative entropy and cross entropy. The truncated quantization unit is used to perform fixed-point truncated symmetric quantization mapping on the weight matrix of the student model to obtain the trained lightweight intrusion detection network ViLKD-IDS.
[0060] Furthermore, in this embodiment, the detection module includes: an image segmentation unit, a precoding unit, a sequence coding unit, a second feature fusion unit, and a classification unit; The image segmentation unit is used to divide the three-channel pseudo-color spatiotemporal image generated by reconstructing the discrete CAN bus message sequence into image blocks to form an image block sequence. The precoding unit is used to perform linear projection on each image block and superimpose positional codes to obtain an initial feature identifier sequence of the image block sequence; The sequence encoding unit is used to input the initial feature identifier sequence into the Vision-xLSTM encoder for feature modeling to obtain the final feature identifier sequence; The second feature fusion unit is used to fuse preset image block features of the final feature identifier sequence to obtain fused features, wherein the preset image block features include the first image block feature and the last image block feature; The classification unit is used to input the fused features into the classification layer, and obtain the classification prediction probability after linear classification and softmax function mapping to realize intrusion detection.
[0061] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A lightweight intrusion detection method for vehicle networks based on isomorphic knowledge distillation, characterized in that, The method includes: Based on multi-dimensional feature fusion, robust quantile normalization and spatial topology reconstruction mapping, the discrete CAN bus message sequence acquired in real time is reconstructed into a three-channel pseudo-color spatiotemporal image. Based on the Vision-xLSTM architecture, a benchmark intrusion detection network, ViL-IDS, is constructed as a teacher model; Based on the Vision-xLSTM architecture and isomorphic knowledge distillation, a lightweight intrusion detection network ViLKD-IDS is constructed as a student model; The student model is trained using the teacher model to obtain the trained lightweight intrusion detection network ViLKD-IDS; The three-channel pseudo-color spatiotemporal image is input into the trained lightweight intrusion detection network ViLKD-IDS to obtain intrusion detection results.
2. The lightweight vehicle network intrusion detection method based on isomorphic knowledge distillation according to claim 1, characterized in that, Methods for reconstructing real-time acquired discrete CAN bus message sequences into three-channel pseudo-color spatiotemporal images based on multi-dimensional feature fusion, robust quantile normalization, and spatial topology reconstruction mapping include: The discrete CAN bus message sequence is segmented using a sliding window to obtain overlapping CAN data blocks, wherein each CAN data block contains multiple consecutive frames of CAN message messages. Extract the identifier field, complete data payload field, high-precision global timestamp, and frequency indicator reflecting injection density of each CAN message in the CAN data block, and perform multi-dimensional feature fusion to construct the original feature matrix of the CAN data block; The original feature matrix is transformed into a three-channel pseudo-color spatiotemporal image with spatial topological consistency by employing robust quantile normalization and spatial topological reconstruction mapping.
3. The lightweight vehicle network intrusion detection method based on isomorphic knowledge distillation according to claim 1, characterized in that, The method for training the student model using the teacher model to obtain the trained lightweight intrusion detection network ViLKD-IDS includes: The teacher model was pre-trained using a general natural image dataset; By utilizing a joint knowledge distillation mechanism based on relative entropy and cross entropy, the soft label distribution knowledge output by the pre-trained teacher model is transferred to the student model. The weight matrix of the student model is subjected to fixed-point truncated symmetric quantization mapping to obtain the trained lightweight intrusion detection network ViLKD-IDS.
4. The lightweight vehicle network intrusion detection method based on isomorphic knowledge distillation according to claim 3, characterized in that, The method for obtaining intrusion detection results by inputting the three-channel pseudo-color spatiotemporal image into the trained lightweight intrusion detection network ViLKD-IDS includes: The three-channel pseudo-color spatiotemporal image generated by reconstructing the discrete CAN bus message sequence is divided into image blocks to form an image block sequence. Linear projection is performed on each image block, and positional encoding is superimposed to obtain the initial feature identifier sequence of the image block sequence; The initial feature identifier sequence is input into the Vision-xLSTM encoder for feature modeling to obtain the final feature identifier sequence; By fusing the preset image block features of the final feature identifier sequence, a fused feature is obtained, wherein the preset image block features include the first image block feature and the last image block feature; The fused features are input into the classification layer, and after linear classification and mapping by the Softmax function, the classification prediction probability is obtained, thereby realizing intrusion detection.
5. A lightweight vehicle network intrusion detection system based on isomorphic knowledge distillation, the system being used to implement the method described in any one of claims 1-4, characterized in that, The system includes: an image reconstruction module, a teacher module, a student module, a training module, and a detection module; The image reconstruction module is used to reconstruct a three-channel pseudo-color spatiotemporal image from a real-time acquired discrete CAN bus message sequence based on multi-dimensional feature fusion, robust quantile normalization, and spatial topology reconstruction mapping. The teacher module is used to construct a benchmark intrusion detection network ViL-IDS as a teacher model based on the Vision-xLSTM architecture; The student module is used to construct a lightweight intrusion detection network ViLKD-IDS as a student model based on the Vision-xLSTM architecture and isomorphic knowledge distillation. The training module is used to train the student model using the teacher model to obtain the trained lightweight intrusion detection network ViLKD-IDS. The detection module is used to input the three-channel pseudo-color spatiotemporal image into the trained lightweight intrusion detection network ViLKD-IDS to obtain intrusion detection results.
6. The lightweight vehicle network intrusion detection system based on isomorphic knowledge distillation according to claim 5, characterized in that, The image reconstruction module includes: a CAN data segmentation unit, a first feature fusion unit, and a conversion and reconstruction unit; The CAN data segmentation unit is used to segment the discrete CAN bus message sequence using a sliding window to obtain overlapping CAN data blocks, wherein each CAN data block contains multiple consecutive CAN message frames. The first feature fusion unit is used to extract the identifier field, complete data payload field, high-precision global timestamp, and frequency indicator reflecting injection density of each frame of CAN message in the CAN data block, and perform multi-dimensional feature fusion to construct the original feature matrix of the CAN data block. The transformation and reconstruction unit is used to convert the original feature matrix into a three-channel pseudo-color spatiotemporal image with spatial topological consistency by employing robust quantile normalization and spatial topological reconstruction mapping.
7. The lightweight vehicle network intrusion detection system based on isomorphic knowledge distillation according to claim 5, characterized in that, The training module includes: a pre-training unit, a knowledge distillation unit, and a truncation quantization unit; The pre-training unit is used to pre-train the teacher model using a general natural image dataset; The knowledge distillation unit is used to transfer the soft label distribution knowledge output by the pre-trained teacher model to the student model using a joint knowledge distillation mechanism based on relative entropy and cross entropy. The truncated quantization unit is used to perform fixed-point truncated symmetric quantization mapping on the weight matrix of the student model to obtain the trained lightweight intrusion detection network ViLKD-IDS.
8. The lightweight vehicle network intrusion detection system based on isomorphic knowledge distillation according to claim 7, characterized in that, The detection module includes: an image segmentation unit, a precoding unit, a sequence encoding unit, a second feature fusion unit, and a classification unit; The image segmentation unit is used to divide the three-channel pseudo-color spatiotemporal image generated by reconstructing the discrete CAN bus message sequence into image blocks to form an image block sequence. The precoding unit is used to perform linear projection on each image block and superimpose positional codes to obtain an initial feature identifier sequence of the image block sequence; The sequence encoding unit is used to input the initial feature identifier sequence into the Vision-xLSTM encoder for feature modeling to obtain the final feature identifier sequence; The second feature fusion unit is used to fuse preset image block features of the final feature identifier sequence to obtain fused features, wherein the preset image block features include the first image block feature and the last image block feature; The classification unit is used to input the fused features into the classification layer, and obtain the classification prediction probability after linear classification and softmax function mapping to realize intrusion detection.