Internet of vehicles multi-layer intrusion detection method and system, electronic device and storage medium
Patent Information
- Application Number
- CN202611013587.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明旨在至少解决相关技术中存在的单一检测手段难以兼顾已知攻击的高精度检测和未知攻击的有效识别的技术问题
[0018]第四方面,基于同一发明构思,本发明还提出一种计算机可读存储介质,所述计算机可读存储介质上存储有计算机程序,所述计算机程序被处理器执行时实现如本发明上述实施例所述的车联网多层入侵检测方法的步骤。
Smart Images

Figure CN122845220A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle network security technology, specifically relating to a multi-layer intrusion detection method, system, electronic device, and storage medium for vehicle networks. Background Technology
[0002] Vehicle-to-everything (V2X) technology enables electronic control units within a vehicle's internal network to interconnect via vehicular networks such as the CAN (Controller Area Network) bus. Simultaneously, vehicles connect more closely to external networks through V2X (Vehicle to Everything) technology. Attack scenarios include in-vehicle threats such as CAN bus injection and DoS (Denial of Service) attacks, as well as attacks targeting the entire V2X ecosystem, such as fake base station spoofing and man-in-the-middle attacks. Therefore, it is crucial to protect the integrity, confidentiality, and availability of communication and data exchange within the V2X ecosystem.
[0003] Signature-based intrusion detection systems use supervised machine learning models to detect known attack patterns and perform well in detecting known attacks; anomaly-based intrusion detection systems use unsupervised learning methods to distinguish abnormal behavior from normal data and can identify unknown attacks.
[0004] However, signature-based methods cannot identify unknown attacks, while anomaly-based methods can detect unknown attacks but have a high false alarm rate. A single detection method is difficult to achieve both high-precision detection of known attacks and effective identification of unknown attacks. Summary of the Invention
[0005] This invention aims to at least address the technical problem in related technologies where a single detection method is insufficient to simultaneously achieve high-precision detection of known attacks and effective identification of unknown attacks. Therefore, the purpose of this invention is to propose a multi-layered intrusion detection method, system, electronic device, and storage medium for vehicle-to-everything (V2X) networks.
[0006] In a first aspect, the present invention proposes a multi-layer intrusion detection method for vehicle-to-everything (V2X) networks, comprising the following steps: acquiring network traffic data in the V2X environment; performing cluster-based hierarchical sampling, oversampling for imbalanced categories, and numerical normalization on the network traffic data to generate a balanced and standardized training dataset; sequentially performing information gain-based initial feature selection, correlation analysis-based redundant feature removal, and kernel principal component analysis-based nonlinear dimensionality reduction on the training dataset to obtain an optimized feature set; inputting the optimized feature set into an ensemble learning detection layer composed of multiple tree-based supervised learning models, and fusing the outputs of the multiple supervised learning models through a stacked ensemble strategy to identify and output the type of known attacks; inputting network traffic data not identified as known attack types as suspicious instances into an anomaly detection layer based on clustering labels, wherein the anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and anomalous attack instances based on clustering probabilities, and outputs the detection result of unknown attacks.
[0007] According to an embodiment of the present invention, a multi-layer intrusion detection method for vehicle-to-everything (V2X) networks first acquires network traffic data in the V2X environment; secondly, the network traffic data undergoes hierarchical sampling based on clustering, oversampling for imbalanced categories, and numerical normalization to generate a balanced and standardized training dataset; then, the training dataset is sequentially subjected to initial feature selection based on information gain, redundant feature removal based on correlation analysis, and nonlinear dimensionality reduction based on kernel principal component analysis to obtain an optimized feature set; next, the optimized feature set is input into an ensemble learning detection layer composed of multiple tree-based supervised learning models, and the outputs of multiple supervised learning models are fused through a stacked ensemble strategy to identify and output the types of known attacks; finally, network traffic data not identified as known attack types is used as suspicious instances and input into an anomaly detection layer based on clustering labels, where the anomaly detection layer uses an unsupervised clustering method to further analyze the suspicious instances. This approach involves clustering and labeling data, identifying normal instances and anomalous attack instances based on clustering probability, and outputting detection results for unknown attacks. Through close collaboration between data preprocessing, three-level feature optimization, ensemble learning detection, and cluster anomaly detection, a complete technical chain is constructed, from data acquisition to known attack identification and unknown attack detection. This enables unified detection and classification output of known attacks, unknown attacks, and normal data packets within the same vehicle-to-everything (V2X) scenario, effectively addressing the challenges of large-scale, class-imbalanced, high-dimensional redundancy, and noisy features in V2X traffic data. Furthermore, the dual-layer detection design balances high-accuracy detection of known attacks with effective identification of unknown attacks. All steps can be completed offline for model training, requiring only low-complexity forward computation during online detection, meeting the real-time and resource constraints of the in-vehicle environment and providing an efficient and comprehensive technical solution for V2X security protection.
[0008] In addition, the multi-layer intrusion detection method for vehicle networks according to embodiments of the present invention may also have the following additional technical features: Furthermore, the network traffic data originates from the vehicle's internal controller area network bus data and / or the vehicle's communication data with external networks; the process of performing cluster-based stratified sampling, oversampling for imbalanced categories, and numerical normalization on the network traffic data includes: dividing the network traffic data into multiple clusters based on the K-means clustering method, and randomly sampling data samples from each cluster according to a preset ratio to form a representative data subset, wherein the number of clusters in the K-means clustering method is determined by Bayesian optimization with the silhouette coefficient as the objective function, and the preset ratio is adjusted according to the limitations of the target vehicle's computing resources; based on the representative data subset, a synthetic minority class oversampling method is used to generate new minority class samples based on linear interpolation of minority class samples and their neighboring samples to balance the number of samples of each category in the representative data subset, thereby obtaining a class-balanced training dataset; numerically encoding the classification features in the class-balanced training dataset, and standardizing all numerical features to ensure that each feature has a uniform dimension, thereby obtaining the training dataset.
[0009] Further, the step of sequentially performing initial feature selection based on information gain, redundant feature removal based on correlation analysis, and nonlinear dimensionality reduction based on kernel principal component analysis on the training dataset includes: calculating the information gain value of each feature in the training dataset relative to its classification label; sorting the features from high to low according to the information gain value; selecting a predetermined number of features at the top of the sort; removing features with importance below a preset threshold to obtain an initial feature set; based on the initial feature set, calculating the correlation value between each pair of features; when the correlation value is greater than a preset correlation threshold, calculating the correlation degree between each feature and the classification label, and removing the feature with a lower correlation degree to the classification label to remove redundant features, thereby obtaining a feature set after redundancy removal, wherein the preset correlation threshold is determined by Bayesian optimization; and using kernel principal component analysis on the feature set after redundancy removal, mapping the original features to a high-dimensional space through a kernel function before extracting principal components to reduce feature dimensionality and suppress noisy features, thereby obtaining the optimized feature set, wherein the number of extracted principal components and the type of kernel function are determined by Bayesian optimization with verification accuracy as the objective function.
[0010] Further, the step of inputting the optimized feature set into an ensemble learning detection layer composed of multiple tree-based supervised learning models, and fusing the outputs of the multiple supervised learning models through a stacked ensemble strategy to identify and output the type of known attack, includes: using the optimized feature set as input, training multiple tree-based supervised learning models as base learners to obtain each base learner and its prediction output for training samples, wherein the multiple tree-based supervised learning models include decision tree models, random forest models, extremely random tree models, and extreme gradient boosting models; for each base learner, using a Bayesian optimization method based on a tree-structured Parzen estimator, with the detection performance of the model on the validation set as the objective, determining the optimal hyperparameter combination for each base learner, and retraining each base learner based on the optimal hyperparameter combination; obtaining the predicted labels output by each of the retrained base learners for the same input data, using each predicted label as a new feature, and training a meta-learner, which is used to fuse the prediction results of each base learner and output the final known attack type determination.
[0011] Further, the step of inputting network traffic data not identified as known attack types as suspicious instances into a clustering-based anomaly detection layer, where the anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and anomalous attack instances based on clustering probabilities, and outputs the detection results of unknown attacks, includes: using network traffic data not identified as known attack types as suspicious instances, clustering the suspicious instances using a K-means-based clustering labeling model, assigning each data point to the nearest neighbor cluster, and assigning a normal or attack label to the cluster based on the known labels of the majority of data samples in each cluster, thereby determining the initial classification result and corresponding clustering probability of each data point, wherein the number of clusters and the distance metric in the clustering labeling model are determined by Bayesian optimization based on Gaussian processes with classification accuracy as the objective function; and excluding clusters with clustering probabilities lower than a preset probability threshold. The suspicious instance is determined as an uncertain instance, wherein the preset probability threshold is determined by Bayesian optimization based on Gaussian process with verification accuracy as the objective function; a first bias classifier and a second bias classifier are pre-trained, wherein the first bias classifier is trained by false negative samples generated by the clustering labeling model and randomly sampled normal data, and the first bias classifier is used to reduce the false negative rate; the second bias classifier is trained by false positive samples generated by the clustering labeling model and randomly sampled attack data, and the second bias classifier is used to reduce the false positive rate; for the uncertain instance, if its initial classification result is normal, it is input into the first bias classifier for reclassification; if its initial classification result is attack, it is input into the second bias classifier for reclassification; the reclassification result is used as the final output of the uncertain instance to obtain the detection result of unknown attack.
[0012] Furthermore, when identifying and outputting the type of known attack, the type of known attack is directly output as the final detection result of the corresponding network traffic data; after identifying normal instances and abnormal attack instances based on clustering probability, normal data packets or unknown attacks are output as the final detection result of the corresponding network traffic data according to the identification results.
[0013] Furthermore, the distance metric includes one of Euclidean distance, Manhattan distance, or Mahalanobis distance.
[0014] Secondly, based on the same inventive concept, this invention also proposes a multi-layer intrusion detection system for vehicle-to-everything (V2X) networks, comprising: a data acquisition module for acquiring network traffic data in a V2X environment; a data processing module for performing cluster-based hierarchical sampling, oversampling for imbalanced categories, and numerical normalization on the network traffic data to generate a balanced and standardized training dataset; a feature engineering module for sequentially performing information gain-based initial feature selection, correlation analysis-based redundant feature removal, and kernel principal component analysis-based nonlinear dimensionality reduction on the training dataset to obtain an optimized feature set; a known attack detection module for inputting the optimized feature set into an ensemble learning detection layer composed of multiple tree-based supervised learning models, fusing the outputs of the multiple supervised learning models through a stacked ensemble strategy to identify and output the type of known attacks; and an unknown attack detection module for inputting network traffic data not identified as known attack types as suspicious instances into an anomaly detection layer based on clustering labels, wherein the anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and anomalous attack instances based on clustering probabilities, and outputs the detection result of unknown attacks.
[0015] The vehicle-to-everything (V2X) multi-layer intrusion detection system provided in the second aspect has the same technical features as the V2X multi-layer intrusion detection method provided in the first aspect, and therefore also has similar technical effects as the first aspect, which will not be described in detail here.
[0016] Thirdly, based on the same inventive concept, the present invention also proposes an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store a computer program; and the processor is used to execute the computer program stored in the memory to implement the steps of the multi-layer intrusion detection method for vehicle networking as described in the above embodiments of the present invention.
[0017] The electronic device provided in the third aspect has the same technical features as the multi-layer intrusion detection method for vehicle networking provided in the first aspect, and therefore also has similar technical effects as the first aspect, which will not be described in detail here.
[0018] Fourthly, based on the same inventive concept, the present invention also proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the multi-layer intrusion detection method for vehicle networking as described in the above embodiments of the present invention.
[0019] The computer-readable storage medium provided in the fourth aspect has the same technical features as the multi-layer intrusion detection method for the Internet of Vehicles provided in the first aspect, and therefore also has similar technical effects as the first aspect, which will not be described in detail here.
[0020] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0021] The above and additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a schematic diagram of a multi-layered hybrid intrusion detection architecture for vehicle networking according to a specific embodiment of the present invention; Figure 2 This is a flowchart of a multi-layer intrusion detection method for vehicle networking according to an embodiment of the present invention; Figure 3 This is a structural block diagram of a multi-layer intrusion detection system for vehicle networking according to an embodiment of the present invention; Figure 4 This is a structural block diagram of an electronic device according to an embodiment of the present invention.
[0022] Figure label: 100 - Vehicle-to-everything (V2X) multi-layer intrusion detection system; 110 - Data acquisition module; 120 - Data processing module; 130 - Feature engineering module; 140 - Known attack detection module; 150 - Unknown attack detection module. Detailed Implementation
[0023] The embodiments of the present invention are described in detail below, and the embodiments described with reference to the accompanying drawings are exemplary.
[0024] To address the problem that single detection methods in related technologies cannot simultaneously achieve high-precision detection of known attacks and effective identification of unknown attacks, this invention provides a multi-layered intrusion detection method, system, electronic device, and storage medium for vehicle networks. The following references... Figures 1-4 This invention describes a multi-layer intrusion detection method, system, electronic device, and storage medium for vehicle networks according to embodiments of the present invention.
[0025] Figure 1 This is a schematic diagram of a multi-layered hybrid intrusion detection architecture for vehicle networking according to a specific embodiment of the present invention. Figure 1As shown in the specific embodiment, the multi-layered hybrid intrusion detection architecture for the Internet of Vehicles includes four main layers. The first layer utilizes four tree-based supervised learning models, namely decision tree, random forest, extremely random tree, and extreme gradient boosting, to develop a signature-based intrusion detection system for known attack detection. The second layer employs stacked ensemble and Bayesian optimization to optimize the first-layer model. The third layer utilizes an unsupervised clustering label CL-K-means (Clustered Label K-means) model to develop an anomaly-based intrusion detection system for unknown attack detection. The fourth layer employs Bayesian optimization to optimize the third-layer model. Specifically, the first layer performs data preprocessing and feature engineering after collecting in-vehicle and external network traffic datasets. Data preprocessing includes K-means-based clustering sampling and synthetic minority class oversampling. Feature engineering includes information gain-based feature selection, correlation-based feature selection, and kernel principal component analysis. The second layer uses a Bayesian optimization method based on tree Parzen estimators and a stacked ensemble model, combining the outputs of the four basic learners in the first layer for model optimization. The third layer passes suspicious instances to the clustering labeling model to separate attack samples from normal samples. The fourth layer uses a BO-GP (Bayesian Optimization-Gaussian Process) method to reduce the classification error of CL-K-means. Finally, three detection results are output: known attacks and their types, unknown attacks, and normal data packets.
[0026] It should be noted that the multi-layer intrusion detection method for vehicle networking in the following specific embodiments of the present invention can all be based on, for example, Figure 1 The architecture shown is implemented in detail. The specific structure and functions of this architecture will be described in detail in the corresponding parts of the following specific embodiments, and will not be repeated here.
[0027] Figure 2 This is a flowchart of a multi-layer intrusion detection method for vehicle networking according to an embodiment of the present invention. Figure 2 As shown, a multi-layer intrusion detection method for vehicle networking according to an embodiment of the present invention includes the following steps: Step S1: Obtain network traffic data in the vehicle networking environment.
[0028] In a specific embodiment, in-vehicle network traffic data can be collected through the CAN (Controller Area Network) bus interface deployed inside the vehicle, while communication data between the vehicle and external networks (such as roadside units, other vehicles, cloud servers, etc.) can be collected through the V2X (Vehicle to Everything) communication module. This provides comprehensive network traffic data covering both in-vehicle and external networks. Specifically, the network traffic data sources include in-vehicle CAN bus data and / or vehicle-to-external network communication data, enabling detection to simultaneously cover in-vehicle CAN bus attacks and external V2X attacks, achieving comprehensive vehicle network security protection. V2X includes: V2V (Vehicle to Vehicle), V2S (Vehicle to Sensors), V2P (Vehicle to Personal Devices), V2N (Vehicle to Network), V2C (Vehicle to Cloud), and V2R (Vehicle to Road Side Units).
[0029] Specifically, the embodiments of the present invention first acquire network traffic data in the vehicle network environment; thereby providing a raw data source for subsequent intrusion detection, ensuring that the detection can simultaneously cover in-vehicle threats such as CAN bus injection and DoS (Denial of Service) attacks, as well as V2X attacks such as fake base station spoofing and man-in-the-middle hijacking, thus constructing a comprehensive security monitoring data foundation from inside the vehicle to outside the vehicle.
[0030] Step S2: Perform cluster-based stratified sampling, oversampling for imbalanced categories, and numerical normalization on the network traffic data to generate a balanced and standardized training dataset.
[0031] In a specific embodiment, the K-Means (K-means Clustering Algorithm) method can be used to divide the original network traffic data into multiple clusters, and data samples are randomly extracted from each cluster according to a preset ratio to form a representative data subset. On this basis, the SMOTE (Synthetic Minority Oversampling Technique) method is used to generate new minority class samples based on linear interpolation of minority class samples and their nearest neighbor samples to balance the number of each class. Finally, the classification features are numerically encoded, and Z-Score Normalization (Z-Score Normalization) is performed on all numerical features to generate a balanced and standardized training dataset.
[0032] Specifically, the embodiments of the present invention further perform cluster-based stratified sampling, oversampling for imbalanced classes, and numerical normalization on network traffic data to generate a balanced and standardized training dataset. This effectively solves the problems of large scale, class imbalance, and differences in feature dimensions in vehicle network traffic data: cluster-based stratified sampling preserves the distribution representativeness of the original data, significantly reduces the amount of training data, the K value is automatically determined through Bayesian optimization to avoid uncertainty caused by manual setting, and the sampling ratio can be flexibly adapted to the computing power of different vehicle devices; the synthetic samples generated by SMOTE oversampling have diversity and data authenticity, effectively solve the class imbalance problem, avoid model bias towards the majority class and overfitting caused by simple replication, and improve the detection sensitivity of attack samples; Z-Score normalization eliminates the influence of dimensions between different features, so that each feature contributes weights on a uniform scale, improving the convergence speed and detection accuracy of model training.
[0033] Step S3: Perform initial feature selection based on information gain, redundant feature removal based on correlation analysis, and nonlinear dimensionality reduction based on kernel principal component analysis on the training dataset in sequence to obtain the optimized feature set.
[0034] In a specific embodiment, the information gain value of each feature relative to the classification label can be calculated first. The computational complexity of information gain can quickly obtain an importance score for each feature. After sorting by importance, high-importance features are selected, and features with importance below a preset threshold are removed to obtain an initial feature set. Then, the symmetric uncertainty between pairs of features is calculated. When the correlation value is high, it is determined to be redundant. The correlation degree between two features and the classification label is calculated separately, and the feature with a lower correlation degree with the classification label is removed to eliminate redundant features. Finally, KPCA (Kernel Principal Component Analysis) can be used to map the original features to a high-dimensional space through the kernel function and then perform principal component extraction to reduce the feature dimension and suppress noisy features. The number of extracted principal components and the type of kernel function are determined by Bayesian optimization based on Gaussian process with the verification accuracy as the objective function.
[0035] Specifically, in this embodiment of the invention, the training dataset is then subjected to a series of steps: initial feature selection based on information gain, redundant feature removal based on correlation analysis, and nonlinear dimensionality reduction based on kernel principal component analysis, to obtain an optimized feature set. Thus, through this three-level feature optimization chain of information gain initial selection, correlation filtering, and kernel principal component analysis, irrelevant, redundant, and noisy features are efficiently removed, significantly reducing the feature dimensionality. Information gain ensures that the features most relevant to the classification are retained, correlation filtering eliminates information duplication between features, and kernel principal component analysis captures the nonlinear relationships between features through nonlinear dimensionality reduction, further suppressing noise. The three levels of optimization work together seamlessly, with the correlation threshold, number of principal components, and kernel type all adaptively determined through Bayesian optimization, ensuring optimal feature optimization results. Therefore, the synergistic effect of these techniques extracts an optimized feature set rich in information, low in redundancy, and with suppressed noise from the original high-dimensional features, providing high-quality feature input for subsequent detection models, significantly improving model training efficiency and detection accuracy, and reducing the risk of overfitting.
[0036] Step S4: Input the optimized feature set into an ensemble learning detection layer consisting of multiple tree-based supervised learning models, and fuse the outputs of multiple supervised learning models through a stacked ensemble strategy to identify and output the type of known attack.
[0037] In a specific implementation, the optimized feature set can be input into four tree-based models: DT (Decision Tree), RF (Random Forest), ET (Extremely Randomized Trees), and XGBoost (eXtreme Gradient Boosting) for training. For each base learner, a Bayesian optimization method based on TPE (Tree-structured Parzen Estimator) is used. With the model's detection performance on the validation set as the objective, the optimal hyperparameter combination (such as maximum tree depth, minimum number of samples per leaf node, learning rate, etc.) is determined for each base learner, and each base learner is retrained based on the optimal hyperparameter combination. Then, the predicted labels output by each base learner for the same input data are used as new features to train a meta-learner (such as a logistic regression model). The meta-learner then fuses the prediction results of each base learner to output the final known attack type determination.
[0038] Specifically, in this embodiment of the invention, the optimized feature set is then input into an ensemble learning detection layer composed of multiple tree-based supervised learning models. The outputs of multiple supervised learning models are fused using a stacked ensemble strategy to identify and output the type of known attack. By selecting four tree-based models as base learners, the interpretability of decision trees, the variance reduction capability of random forests (Bagging, Bootstrap aggregating), the overfitting suppression effect of extremely random trees, and the high-precision prediction advantage of extreme gradient boosting are fully utilized. These four models capture attack patterns from different angles, exhibiting strong complementarity. TPE Bayesian optimization efficiently searches the high-dimensional hyperparameter space within a limited number of iterations, ensuring that each base learner reaches its optimal performance state. Stacked ensemble learns the combined weights and decision rules of the predicted labels from each base learner through a meta-learner, integrating the advantages of each model and compensating for the shortcomings of a single model. Thus, the synergistic effect of these technical features constructs a high-precision, highly robust known attack detection layer, achieving high-accuracy identification of multiple known attack types.
[0039] Step S5: Network traffic data that is not identified as a known attack type is taken as a suspicious instance and input into the cluster-based anomaly detection layer. The anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and anomalous attack instances based on clustering probability, and outputs the detection results of unknown attacks.
[0040] In a specific implementation, suspicious instances detected as "normal" in the previous layer can be input into the CL-K-means (K-means clustering labeling) model. The K-means-based clustering labeling model is used to cluster the suspicious instances, assigning each data point to its nearest neighbor cluster. Based on the known labels of the majority of data samples in each cluster, the cluster is labeled as either normal or aggressor, thus determining the initial classification result and corresponding cluster probability (i.e., the proportion of majority class samples in that cluster). Specifically, the number of clusters and the distance metric (Euclidean distance, Manhattan distance, or Mahalanobis distance, covariance distance) of the CL-K-means model can be determined through Bayesian optimization based on Gaussian processes, with classification accuracy as the objective function.
[0041] Specifically, in this embodiment of the invention, network traffic data not identified as known attack types is input as suspicious instances into an anomaly detection layer based on clustering labels. The anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and anomalous attack instances based on clustering probabilities, and outputs the detection results of unknown attacks. Thus, unsupervised clustering using CL-K-means achieves the detection of unknown attacks without relying on a prior pattern library. The number of clusters and distance metrics are adaptively determined through Bayesian optimization to ensure optimal clustering quality. The clustering probability threshold is optimized to achieve the best balance between excluding low-confidence predictions and maintaining detection coverage. Therefore, through the synergistic effect of the above technical features, effective identification of new and variant attacks is achieved.
[0042] Therefore, the multi-layer intrusion detection method for vehicle-to-everything (V2X) networks according to embodiments of the present invention first acquires network traffic data in the V2X environment; secondly, it performs cluster-based hierarchical sampling, oversampling for imbalanced categories, and numerical normalization on the network traffic data to generate a balanced and standardized training dataset; then, it sequentially performs information gain-based initial feature selection, correlation analysis-based redundant feature removal, and kernel principal component analysis-based nonlinear dimensionality reduction on the training dataset to obtain an optimized feature set; next, it inputs the optimized feature set into an ensemble learning detection layer composed of multiple tree-based supervised learning models, and fuses the outputs of multiple supervised learning models through a stacked ensemble strategy to identify and output the types of known attacks; finally, it inputs network traffic data not identified as known attack types as suspicious instances into an anomaly detection layer based on clustering labels, and the anomaly detection layer uses an unsupervised clustering method to classify suspicious instances. For example, clustering and labeling are performed, and normal instances and anomalous attack instances are identified based on clustering probability, with the detection results of unknown attacks output. Through the close collaboration of data preprocessing, three-level feature optimization, ensemble learning detection, and cluster anomaly detection, a complete technical chain is constructed from data acquisition to known attack identification and unknown attack detection. This achieves unified detection and classification output of known attacks, unknown attacks, and normal data packets in the same vehicle-to-everything (V2X) scenario, effectively solving the problems of large-scale, class-imbalanced, high-dimensional redundancy, and noisy features in V2X traffic data. Simultaneously, the dual-layer detection design ensures both high-accuracy detection of known attacks and effective identification of unknown attacks. Furthermore, all steps can be completed offline for model training, requiring only low-complexity forward computation during online detection, meeting the real-time and resource constraints of the in-vehicle environment, and providing an efficient and comprehensive technical solution for V2X security protection.
[0043] In one embodiment of the present invention, network traffic data originates from vehicle internal controller local area network bus data and / or vehicle communication data to external networks.
[0044] Specifically, the network traffic data obtained in this embodiment of the invention comes from the vehicle's internal controller local area network bus data and / or the vehicle's communication data with external networks, enabling the detection to simultaneously cover in-vehicle CAN bus attacks and external V2X attacks, achieving comprehensive vehicle network security protection.
[0045] In one embodiment of the present invention, network traffic data is subjected to cluster-based stratified sampling, oversampling for imbalanced categories, and numerical normalization processing, including: dividing network traffic data into multiple clusters based on the K-means clustering method, and randomly sampling data samples from each cluster according to a preset ratio to form a representative data subset, wherein the number of clusters in the K-means clustering method is determined by Bayesian optimization with the silhouette coefficient as the objective function, and the preset ratio is adjusted according to the limitation of the target vehicle's computing resources; based on the representative data subset, a synthetic minority class oversampling method is used to generate new minority class samples based on linear interpolation of minority class samples and their neighboring samples to balance the number of samples of each category in the representative data subset, thereby obtaining a class-balanced training dataset; numerically encoding the classification features in the class-balanced training dataset, and performing standardization on all numerical features so that each feature has a uniform dimension, thereby obtaining the training dataset.
[0046] In a specific embodiment, a sampling method based on K-means clustering is first used to generate a highly representative subset. Specifically, the original data points are first divided into multiple clusters, and then data is extracted proportionally from each cluster to form a highly representative and efficient data subset. After forming k clusters by K-means clustering of the original data samples, random sampling is applied to each cluster, selecting 10% of the data as the sample set. The data sampling percentage can be adjusted according to the data scale and resource constraints.
[0047] K-means clustering aims to minimize the sum of squared distances between all data points and their respective cluster centroids, which can be expressed as: min Σ|| x i - u j In the formula ||², min represents minimization, and Σ represents accumulation. x i For the first i Each input sample data point u j For the first j The centroid (cluster center) of each cluster, || x i - u j ||²for x i To the centroid of its cluster u j The square of the Euclidean distance; its linear time complexity is O(nkt), where n is the data size, k is the number of clusters, and t is the number of iterations; the main hyperparameter k of K-means clustering is adjusted by Bayesian optimization based on Gaussian process, and the silhouette coefficient is used as the objective function.
[0048] In network traffic data, the proportion of normal samples is usually much larger than that of attack samples, which may lead to model bias and reduced detection rate. Therefore, the SMOTE oversampling technique can be used to create high-quality instances for the minority class. Specifically, for minority class instances... X Assuming X i The new synthetic instance is a sample randomly selected from its k nearest neighbors. X n = X + rand (0,1)×( X i - X In the formula, rand (0,1) represents a random number between 0 and 1.
[0049] The categorical features are converted into numerical features using a label encoder, and then the network dataset is standardized using the Z-Score method, so that the feature mean is 0 and the standard deviation is 1. The standardized feature values are represented as follows: x n = ( x - μ ) / σ In the formula, x These are the original eigenvalues. μ and σ These are the mean and standard deviation, respectively.
[0050] Specifically, this invention employs K-means clustering sampling to ensure that the sampled dataset retains the distribution characteristics of the original data, avoiding the loss of key attack patterns caused by random sampling. The number of clusters is automatically determined through Bayesian optimization, avoiding uncertainties caused by manual settings. The sampling ratio can be flexibly adapted to the computing power of different in-vehicle devices, enhancing the adaptability and flexibility of the method. The synthetic samples generated by SMOTE oversampling have diversity and data authenticity, effectively solving the class imbalance problem and avoiding overfitting caused by simple replication. Z-Score standardization eliminates the influence of different feature dimensions, enabling each feature to contribute weights on a uniform scale, thus improving the stability of model training. This ensures the generation of a high-quality training dataset that is both representative and class-balanced, laying a solid data foundation for subsequent feature optimization and detection.
[0051] In one embodiment of the present invention, the training dataset is sequentially subjected to initial feature selection based on information gain, redundant feature removal based on correlation analysis, and nonlinear dimensionality reduction based on kernel principal component analysis. This includes: calculating the information gain value of each feature in the training dataset relative to its classification label; sorting features from high to low based on their information gain values; selecting a predetermined number of features at the top of the sorted list; removing features with importance below a preset threshold to obtain an initial feature set; calculating the correlation value between each pair of features based on the initial feature set; when the correlation value is greater than a preset correlation threshold, calculating the correlation degree between each feature and the classification label; and removing the feature with a lower correlation degree to the classification label to remove redundant features, resulting in a feature set after redundancy removal. The preset correlation threshold is determined through Bayesian optimization. For the feature set after redundancy removal, kernel principal component analysis is used to map the original features to a high-dimensional space using a kernel function before extracting principal components to reduce feature dimensionality and suppress noisy features, resulting in an optimized feature set. The number of extracted principal components and the type of kernel function are determined through Bayesian optimization with verification accuracy as the objective function.
[0052] In a specific embodiment, IG (Information Gain) is used to select important features and measure the amount of information that a feature can provide for the target variable. T and random variables X The characteristics represented, information gain is expressed as IG ( T | X )= H ( T )- H ( T | X In the formula, IG ( T | X ) is a random variable X For target variable T Information gain H ( T ) is the target variable T entropy, H ( T | X ) is in random variables X target variable T The entropy and information gain have a computational complexity of O(n), which can quickly obtain an importance score for each feature.
[0053] Fast Correlation-Based Filter (FCBF) removes redundant features by calculating the correlation between input features. This method effectively removes redundant features and preserves informative features in high-dimensional datasets, with a time complexity of O(nlogn), where nlogn is the linear logarithm. The SU (Symmetric Uncertainty) is calculated by normalizing the information gain value, expressed as: SU ( X , Y )=2[ IG ( X | Y ) / ( H ( X )+ H ( Y ))], where, SU ( X , Y ) as a feature X and characteristics Y The symmetric uncertainty between them IG ( X | Y ) for features Y Features under the conditions X The information gain it possesses H ( X )and H ( Y ) are features X and characteristics Y The entropy of the features is calculated; when the correlation value between two features is greater than a threshold α, the feature with higher importance is retained and the other is discarded. Specifically, the threshold α is optimized using Bayesian optimization based on a Gaussian process.
[0054] After implementing information gain and fast correlation filtering, kernel principal component analysis is used to improve anomaly-based intrusion detection. Kernel principal component analysis learns nonlinear functions or decision boundaries through kernel tricks to reduce the dimensionality of nonlinear data and reduce computational complexity, overfitting risk, and diffuse noise. The number of extracted features and kernel type are optimized through Bayesian optimization based on Gaussian processes, and accuracy is validated as the objective function.
[0055] Specifically, this invention rapidly eliminates irrelevant features through information gain, significantly reducing feature dimensionality; it effectively identifies and removes redundant features with overlapping information through correlation filtering, avoiding interference from multicollinearity in model training; and it uses kernel principal component analysis to capture nonlinear relationships between features, further suppressing noise and reducing dimensionality. The three levels of feature optimization work together seamlessly, with the correlation threshold, number of principal components, and kernel type all adaptively determined through Bayesian optimization, ensuring optimal feature optimization results. This ensures that an optimized feature set rich in information, low in redundancy, and with suppressed noise can be extracted from the original high-dimensional features, helping to improve the training efficiency and detection accuracy of subsequent detections and reducing the risk of overfitting.
[0056] In one embodiment of the present invention, an optimized feature set is input into an ensemble learning detection layer composed of multiple tree-based supervised learning models. The outputs of multiple supervised learning models are fused through a stacked ensemble strategy to identify and output the type of known attack. This includes: using the optimized feature set as input, training multiple tree-based supervised learning models as base learners to obtain each base learner and its prediction output for training samples. The multiple tree-based supervised learning models include decision tree models, random forest models, extremely random tree models, and extreme gradient boosting models. For each base learner, a Bayesian optimization method based on a tree-structured Parzen estimator is used to determine the optimal hyperparameter combination for each base learner, with the detection performance of the model on the validation set as the objective. The base learners are then retrained based on the optimal hyperparameter combination. The predicted labels output by each retrained base learner for the same input data are obtained. Each predicted label is used as a new feature to train a meta-learner. The meta-learner is used to fuse the prediction results of each base learner and output the final known attack type determination.
[0057] In a specific embodiment, after data preprocessing and feature engineering, the labeled dataset is trained using an ensemble learning model to develop a signature-based intrusion detection system. The system selects four tree-based machine learning methods as base learners: decision trees, random forests, extremely random trees, and extreme gradient boosting; these methods are used to identify and detect known attack patterns, and accuracy and robustness are improved by combining multiple models.
[0058] After obtaining four tree-based machine learning models, a stacking method is used to combine them; the stack uses the output labels estimated by the four base learners as input features to train a meta-learner to make the final prediction; the important hyperparameters of the four tree-based machine learning methods are optimized through TPE Bayesian optimization.
[0059] For tree-based methods, the training time complexities of decision trees, random forests, extremely random trees, and extreme gradient boosting are O(n²f), O(n²ft), O(nft), and O(nft), respectively, where f is the number of features and t is the number of decision trees in the ensemble model.
[0060] Specifically, this embodiment of the invention selects four tree-based models as base learners, fully leveraging the interpretability of decision trees, the variance reduction capability of random forests, the overfitting suppression effect of extremely random trees, and the high-precision prediction advantage of extreme gradient boosting. The four models capture attack patterns from different angles and are highly complementary. TPE Bayesian optimization efficiently searches the high-dimensional hyperparameter space within a limited number of iterations, enabling each base learner to reach its optimal performance state. Stacked ensemble learns the combined weights and decision rules of the predicted labels of each base learner through meta-learners, combining the advantages of each model and making up for the shortcomings of a single model. This ensures that a high-precision and highly robust known attack detection layer can be constructed.
[0061] In one embodiment of the present invention, network traffic data not identified as a known attack type is used as a suspicious instance and input into an anomaly detection layer based on clustering labels. The anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and anomalous attack instances based on clustering probabilities, and outputs the detection result of unknown attacks. This includes: using network traffic data not identified as a known attack type as a suspicious instance, clustering the suspicious instances using a K-means-based clustering labeling model, assigning each data point to its nearest neighbor cluster, and assigning a normal or attack label to the cluster based on the known labels of the majority of data samples in each cluster, thereby determining the initial classification result and corresponding clustering probability of each data point. The number of clusters and the distance metric in the clustering labeling model are determined using Bayesian optimization based on Gaussian processes with classification accuracy as the objective function; the clustering probabilities are then... Suspicious instances with probabilities below a preset threshold are classified as uncertain instances. The preset probability threshold is determined using Bayesian optimization based on a Gaussian process, with verification accuracy as the objective function. A first bias classifier and a second bias classifier are pre-trained. The first bias classifier is trained using false negative samples generated by the clustering labeling model and randomly sampled normal data, and is used to reduce the false negative rate. The second bias classifier is trained using false positive samples generated by the clustering labeling model and randomly sampled attack data, and is used to reduce the false positive rate. For uncertain instances, if their initial classification result is normal, they are input into the first bias classifier for reclassification; if their initial classification result is attack, they are input into the second bias classifier for reclassification. The reclassification result is used as the final output of the uncertain instance to obtain the detection result of the unknown attack.
[0062] In a specific embodiment, signature-based intrusion detection can detect a variety of known attacks, but attackers can still execute unknown attacks not included in the known attack patterns and may be misclassified as normal. Therefore, instances marked as "normal" by the signature-based intrusion detection system are considered suspicious instances and are further passed to the anomaly-based intrusion detection system.
[0063] The system uses an optimized dataset to train an anomaly-based intrusion detection system. The main process of CL-K-means is as follows: split the dataset into a sufficient number of clusters using K-means; label each cluster according to the data sample label; label the test sample as normal or attack based on the cluster label it is assigned to; calculate the percentage of majority class samples for each test cluster as the confidence score or clustering probability; optimize the number of clusters and distance metric to obtain the optimal model through Bayesian optimization based on Gaussian process.
[0064] To improve the detection rate and reduce the false positive rate of CL-K-means, the system uses two bias classifiers to reduce FN (False Negative) and FP (False Positive) respectively. First, false negatives and false positives obtained by CL-K-means are collected in the training set. Then, the best-performing single supervised learning model is selected from the signature-based intrusion detection. Next, the first bias classifier B1 is trained with the same number of randomly sampled normal data to reduce false negatives. The second bias classifier B2 is trained with the same number of randomly sampled attack data to reduce false positives.
[0065] After implementing CL-K-means, each data sample with a cluster probability less than a probability threshold is considered an uncertain instance. Specifically, the probability threshold is a continuous variable and is optimized to 0.901 using a Bayesian method based on a Gaussian process, with verification accuracy as the objective function. Uncertain instances are passed to B1 if they are marked as normal by CL-K-means, and to B2 if they are marked as attack, to obtain the final classification result.
[0066] Specifically, this embodiment of the invention achieves unknown attack detection without relying on a prior pattern library through unsupervised clustering using CL-K-means. The number of clusters and distance metrics are adaptively determined through Bayesian optimization to ensure optimal clustering quality. The clustering probability threshold is optimized to achieve the best balance between excluding low-confidence predictions and maintaining detection coverage. The dual-bias classifiers B1 and B2 have clearly defined roles: B1 specifically compensates for false negatives to improve attack recall, while B2 specifically compensates for false positives to reduce false alarm rate, and performs differentiated secondary judgment on low-confidence ambiguous samples. This ensures that effective identification of novel and variant attacks can be achieved.
[0067] In one embodiment of the present invention, when identifying and outputting the type of a known attack, the type of the known attack is directly output as the final detection result of the corresponding network traffic data; after identifying normal instances and abnormal attack instances based on clustering probability, normal data packets or unknown attacks are output as the final detection result of the corresponding network traffic data according to the identification results.
[0068] In a specific embodiment, when the first-layer integrated detection model determines that a data packet is a known attack (such as a DoS attack, FTP (File Transfer Protocol) brute-force attack, etc.), it directly outputs the attack type label as the final detection result, and the data packet no longer enters the anomaly detection layer; when the data packet is determined to be normal in the first layer, it enters the anomaly detection layer. If the CL-K-means and bias classifier determine that it is an attack, it outputs the "unknown attack" label; if it is determined to be normal, it outputs the "normal" label.
[0069] Specifically, the embodiments of the present invention ensure that known attacks directly output type information through a clear final judgment mechanism, avoiding unnecessary computational overhead caused by their continued entry into the anomaly detection layer, thereby improving the overall detection efficiency of the system; at the same time, it clearly distinguishes between normal communication and unknown threats, providing a clear basis for subsequent security responses; thus ensuring that the three types of data (known attacks, unknown attacks, and normal) each have a clear output path and final judgment result in the detection process, avoiding result conflicts or duplicate judgments between multiple layers of detection.
[0070] In one embodiment of the present invention, the distance metric includes one of Euclidean distance, Manhattan distance, or Mahalanobis distance.
[0071] In specific implementations, CL-K-means clustering can use Euclidean distance (suitable for data with a spherically distributed feature space), Manhattan distance (good robustness to outliers), or Mahalanobis distance (considering the covariance structure between features, suitable for cases with strong feature correlation). The specific distance metric used is automatically selected and determined by Bayesian optimization based on Gaussian processes with classification accuracy as the objective function. For K-means clustering, the distance metric can be used to divide data points according to Euclidean distance, Manhattan distance, or Mahalanobis distance. The number of clusters and the distance metric are both obtained as the main hyperparameters by Bayesian optimization based on Gaussian processes.
[0072] Specifically, this embodiment of the invention provides multiple distance metric options, enabling CL-K-means to adapt to network traffic data with different data distribution characteristics. Bayesian optimization automatically selects the optimal metric method to ensure that the clustering effect and classification accuracy are optimal, thereby improving the adaptability and accuracy of unknown attack detection.
[0073] The advantages and benefits of the multi-layer intrusion detection method for vehicle networking described above in this invention will be explained below with reference to a specific embodiment.
[0074] In a specific embodiment, to fully verify the detection accuracy, generalization ability, unknown attack identification effect, and real-time vehicle operation feasibility of the proposed multi-layer intrusion detection method for vehicle networking, a standardized experimental environment was built. Four types of authoritative publicly available industry datasets were used for multi-dimensional comparative testing. The training and test sets were split using a uniform 80% / 20% sample partitioning method. Known attack detection was evaluated using 10-fold cross-validation, while unknown attack detection was evaluated using hold-out validation. Accuracy, Precision, Recall, F1-Score, True Positive Rate (TPR), and False Positive Rate (FPR) were selected as unified evaluation metrics. The complete experimental configuration and test results are as follows.
[0075] Table 1 Test Environment Setup Table
[0076] Table 2 Detailed Configuration Table of CICIDS2017 Dataset
[0077] Table 3. Detailed Configuration Table of CSE-CIC-IDS-2018 Dataset
[0078] Table 4. Detailed Configuration Table of CIC-DDoS-2019 Dataset
[0079] Table 5 Detailed Configuration Table of CAN-intrusion-dataset
[0080] Table 6. Experimental Results on the CICIDS2017 Dataset
[0081] Table 7 Experimental Results on the CSE-CIC-IDS-2018 Dataset
[0082] Table 8. Experimental Results on the CIC-DDoS-2019 Dataset
[0083] Table 9. Performance Comparison of MIDS with Other Machine Learning Models on the CICIDS2017 Dataset
[0084] Table 10 Evaluation Results of Unknown Attack Detection on CAN-intrusion-dataset
[0085] Table 11 Evaluation results of unknown attack detection on the CICIDS2017 dataset
[0086] Table 1 above shows the test environment setup, which records the unified hardware and software configuration for this experiment to eliminate the interference of environmental differences on the test results and ensure fair and reproducible performance comparisons across multiple models and datasets. Specifically, the hardware is equipped with a 12-core Xeon server CPU and an RTX 3080 Ti accelerated graphics card. The software is based on an Ubuntu 22.04 64-bit system and Python 3.12 to build a machine learning runtime environment. Sufficient computing power is used for model training and large-scale traffic sample preprocessing, while also verifying that the algorithm can be adapted to the real-time inference requirements of lightweight automotive hardware.
[0087] This experiment uses four standard datasets in the field of intrusion detection, covering three typical scenarios: general Internet traffic, large-scale DDoS traffic, and CAN bus communication traffic in vehicle networking. Tables 2-5 record the total number of samples, the number of training / test set partitions, and the attack category labels for each dataset. Among them, the class label is the sample classification identifier, the attack type is the attack behavior category of the corresponding traffic, the original sample number is the original total number of samples in the dataset, the number of training set samples is the samples used for model feature learning and hyperparameter optimization under 80% partition ratio, and the number of test set samples is the remaining 20% of the samples, used to test the model's generalization ability. Table 2 above is a detailed configuration table of the CICIDS2017 dataset, which is a general benchmark dataset for network intrusion detection. It includes samples of more than ten mainstream network attacks, such as normal traffic BENIGN and botnets, DDoS denial-of-service attacks, various DoS attacks, Heartbleed vulnerability attacks, FTP / SSH brute-force attacks, port scanning, web injection, and XSS cross-site scripting.
[0088] Table 3 above is a detailed configuration table of the CSE-CIC-IDS-2018 dataset. As an upgraded dataset of CICIDS2017, it expands the DDoS attack samples such as HOIC and LOIC-UDP to verify the stability of model recognition under complex high-volume attacks.
[0089] Table 4 above is a detailed configuration table of the CIC-DDoS-2019 dataset, which covers various direct / reflective DDoS attacks such as SYN floods, TFTP, NTP reflection, LDAP, DNS, MSSQL, and NetBIOS, to verify the model's ability to finely distinguish between different types of DDoS traffic.
[0090] Table 5 above is a detailed configuration table of the CAN-intrusion-dataset dataset, adapted to the vehicle application scenario of this invention. It includes five types of typical vehicle traffic: CAN bus normal messages, bus DoS attacks, fuzzy random message attacks, RPM speed spoofing, and gear position spoofing, used to verify the effectiveness of vehicle internal bus intrusion detection.
[0091] Table 6 above shows the experimental results on the CICIDS2017 dataset, Table 7 above shows the experimental results on the CSE-CIC-IDS-2018 dataset, and Table 8 above shows the experimental results on the CIC-DDoS-2019 dataset. As shown in Tables 6-8, the multi-class detection metrics of the fully optimized MIDS model of this invention and four classic ensemble learning models, namely XGBoost, RF Random Forest, DT Decision Tree, and ET Extreme Random Tree, are recorded on the CICIDS2017, CSE-CIC-IDS-2018, and CIC-DDoS-2019 datasets. At the same time, two sets of MIDS control versions are set: MIDS (Without FS & HPO) is the basic stacked model without the introduction of FS (Feature Selection) and HPO (Hyperparameter Optimization); MIDS (Multi-Class Model) is the fully optimized multi-layer intrusion detection model of this invention.
[0092] As shown in Tables 6-8, the test results demonstrate that on the general network dataset CICIDS2017, the fully optimized MIDS model of this invention achieves an accuracy of 99.989%, precision of 99.978%, recall of 99.971%, and an F1 score of 0.99997, comprehensively outperforming all traditional tree models. On the high-volume complex DDoS dataset CSE-CIC-IDS-2018, the fully optimized MIDS model of this invention achieves an accuracy of 98.873% and an F1 score of 0.98874, leading the benchmark model in generalization performance. On the specialized DDoS dataset CIC-DDoS-2019, the fully optimized MIDS model of this invention achieves an accuracy of 95.65% and an F1 score of 0.9551, maintaining stable and high-precision identification against multiple types of differentiated DDoS attacks, proving that feature selection and hyperparameter optimization can significantly improve the classification accuracy of multi-layer stacked models.
[0093] Table 9 above is a performance comparison table of MIDS and other machine learning models on the CICIDS2017 dataset. As shown in Table 9, the complete optimized MIDS model of the present invention is compared with the existing mainstream intrusion detection algorithms such as KNN, SU-IDS, STDeepGraph, GAN-RF, and Multi-SVM on the CICIDS2017 dataset. All comparison models maintain the same testing environment, dataset, and evaluation metrics.
[0094] As shown in Table 9, the comparison results show that the accuracy of the existing best comparison model GAN-RF is only 99.83%, while the accuracy, precision, recall and F1 score of the complete optimized MIDS model of this invention are all surpassed, and the false positive and false negative levels are lower, with significant technical advantages in overall detection performance.
[0095] The real-world connected vehicle environment continuously generates new attacks that are not trained on. Therefore, an unknown attack detection and verification is added, and a retention verification method is adopted: the training set only contains normal traffic and known attack samples, while the verification set introduces new unknown attack samples. The TPR true positive rate (unknown attack detection rate), FPR false positive rate (normal traffic false alarm rate), and F1 score are used to evaluate the model's anomaly recognition capability.
[0096] Table 10 above shows the evaluation results of unknown attack detection on the CAN-intrusion-dataset. As shown in Table 10, for the four types of unknown attacks on the CAN bus: DoS, Fuzzy fuzziness, speed spoofing, and gear spoofing, the complete optimized MIDS model of this invention has a TPR of 94.861%, an average FPR of only 0.241%, and an average F1 score of 0.96451. Among them, the detection rate of DoS, speed spoofing, and gear spoofing attacks is close to 100%. Only the Fuzzy fuzziness attack has a slight decrease in recognition effect due to the random values of the message and the overlap of some features with normal messages, which is a common recognition difficulty in the industry. Table 11 above shows the evaluation results of unknown attack detection on the CICIDS2017 dataset. As shown in Table 11, it covers 14 types of unknown network attacks, including Bots, DDoS, various types of DoS, brute-force attacks, port scanning, Web injection, and XSS. The average TPR of the fully optimized MIDS model of this invention is 75.943%, the average FPR is 13.882%, and the average F1 score is 0.80013. Most attacks can be effectively identified. Only Web Attack-XSS cross-site scripting has poor detection performance because its traffic distribution is highly similar to normal Web traffic. Further optimization can be achieved by refining the Web layer features.
[0097] The detection process of this invention is divided into an offline training stage for the server and an online testing stage for the vehicle terminal. The online inference stacked model consists of four types of tree-based learners, a CL-K-means clustering module, and a bias classifier. The time complexity of each module is as follows: O(dft) for the tree-based signature detection module, O(fk) for the CL-K-means clustering module, and O(dft) for the bias classifier. The overall maximum runtime complexity is O(2dft+fk), where d is the maximum tree depth, f is the number of features, t is the number of trees, and k is the number of cluster centers. The overall complexity is linear, with low computational overhead, which can adapt to the computing power constraints of embedded hardware such as vehicle T-BOX and meet the response requirements for real-time detection and real-time alarm of vehicle network messages.
[0098] Multiple datasets, multiple model comparisons, and dual testing results for known and unknown attacks confirm that the multi-layer intrusion detection method of this invention has extremely high recognition accuracy for general network traffic, large-scale DDoS traffic, and vehicle CAN bus traffic. Compared with existing traditional machine learning and deep learning intrusion detection algorithms, it comprehensively leads in all core performance indicators. It can maintain a high detection rate and an extremely low false alarm rate when facing novel and unknown attacks not covered by the training set. The algorithm has low inference time complexity, balances detection accuracy and vehicle real-time performance, and can effectively solve the defects of existing vehicle network intrusion detection solutions, such as insufficient accuracy, weak generalization ability, and inability to adapt to the real-time operation of vehicle hardware.
[0099] In summary, the key points and protection points of the multi-layer intrusion detection method for vehicle networking provided by this invention include the following six aspects: (1) Four-layer multi-layer hybrid detection architecture. It includes a four-layer detection architecture consisting of a tree-based supervised learning model, stacked ensemble and Bayesian optimization, CL-K-means anomaly detection and Gaussian process-based Bayesian optimization, which realizes unified output of known attacks, unknown attacks and normal data packets.
[0100] (2) Data preprocessing link for vehicle network data. This includes a combined process of data sampling based on K-means clustering, synthetic minority class oversampling, label encoding, and Z-Score standardization to generate representative and balanced datasets.
[0101] (3) IG-FCBF-KPCA feature optimization framework. This includes a feature engineering framework that first removes unimportant features through information gain, then removes redundant features through fast correlation filtering, and finally reduces dimensionality and noisy features through kernel principal component analysis.
[0102] (4) Signature-based stacked ensemble detection mechanism. This includes a combination of base learners consisting of decision trees, random forests, extremely random trees and extreme gradient boosting, as well as methods to improve the performance of known attack detection through stacked learning and Bayesian optimization based on tree Parzen estimators.
[0103] (5) Unknown attack detection mechanism combining CL-K-means and bias classifier. This includes a process for detecting unknown attacks by clustering suspicious instances, judging uncertain instances by cluster probability, and feeding uncertain instances into two bias classifiers to reduce false negatives and false positives respectively.
[0104] (6) Low runtime complexity design. Includes a lightweight hybrid detection mechanism with an overall runtime complexity of O(2dft + fk) to meet the real-time response requirements of the vehicle environment.
[0105] In other words, the multi-layer intrusion detection method for vehicle networking provided by this invention has at least the following beneficial effects compared with the prior art: (1) Collaborative protection against known and unknown attacks. A four-layer detection framework is used to achieve comprehensive protection from in-vehicle network to vehicle-to-everything communication. The first and second layers use optimized machine learning algorithms to identify known threats such as controller area network bus attacks and protocol spoofing. The third and fourth layers use improved clustering methods to detect unknown attacks.
[0106] (2) Data quality improvement. K-means-based clustering sampling and synthetic minority class oversampling techniques are used to solve the problems of high-dimensional sparsity and class imbalance. Information gain, fast correlation filtering and kernel principal component analysis are combined to effectively extract the characteristics of vehicle network attacks.
[0107] (3) Improved model optimization and generalization capabilities. Stacked ensemble and Bayesian optimization methods are adopted to reduce system overhead while ensuring detection accuracy. The model is tested on multiple standard datasets to verify its effectiveness in different vehicle network attack scenarios.
[0108] (4) Real-time advantage. It adopts a lightweight design with a maximum overall runtime complexity of O(2dft + fk), which meets the real-time response requirements of the vehicle environment.
[0109] (5) Detection performance advantages. In the detection of known attacks, the present invention achieves an accuracy of 99.989% on the CICIDS2017 dataset; and achieves an accuracy of over 95% on both the CSE-CIC-IDS-2018 and CIC-DDoS-2019 datasets; for the detection of unknown attacks, the present invention achieves an average F1 score of 0.96451 on the CAN-intrusion-dataset and an average F1 score of 0.80013 on the CICIDS2017 dataset.
[0110] In addition to the above solutions, the following alternative solutions can also achieve the purpose of this invention without departing from the technical content of this invention: (1) Data sampling ratio substitution: After K clusters are formed by K-means clustering, the original scheme randomly selects 10% of the data in each cluster as the sample set. This data sampling percentage can be adjusted according to the data scale and resource constraints (e.g., 5%-30%).
[0111] (2) Distance metric substitution: K-means clustering can divide data points according to Euclidean distance, Manhattan distance or Mahalanobis distance; in CL-K-means, the number of clusters K and the distance metric are both obtained as the main hyperparameters by Bayesian optimization based on Gaussian process.
[0112] (3) Sampling and oversampling mechanism alternatives: The class imbalance problem can be solved by resampling methods, including random sampling and synthetic minority oversampling techniques. The original scheme chose synthetic minority oversampling techniques because it synthesizes high-quality instances based on the K nearest neighbor concept, which is different from the simple copy instance method that may lead to overfitting.
[0113] (4) Clustering method replacement: If time and budget permit, K-means can be replaced with other clustering methods with the same clustering labeling technique, depending on the specific data shape and distribution, to further improve system performance.
[0114] (5) Mini-batch K-means substitution: To further reduce model training time, mini-batch K-means can be used, in which a random sample subset is used as the mini-batch in each training iteration.
[0115] (6) Online learning alternative: Online learning technology, which can continuously update the learning model according to new attack patterns, may also improve the versatility and accuracy of intrusion detection systems.
[0116] (7) Other anomaly detection methods as controls or alternatives: Isolation forests and single-class support vector machines can be compared as other unsupervised anomaly detection methods; they can be used as controls or alternative candidates in equivalent implementations, but CL-K-means with bias classifiers has advantages in accuracy and efficiency.
[0117] A further embodiment of the present invention discloses a multi-layer intrusion detection system for vehicle networking. Figure 3 This is a structural block diagram of a multi-layer intrusion detection system for vehicle networking according to an embodiment of the present invention. Figure 3 As shown, in one embodiment of the present invention, the vehicle network multi-layer intrusion detection system 100 includes: a data acquisition module 110, a data processing module 120, a feature engineering module 130, a known attack detection module 140, and an unknown attack detection module 150.
[0118] Specifically, the data acquisition module 110 is used to acquire network traffic data in the vehicle networking environment.
[0119] The data processing module 120 is used to perform cluster-based hierarchical sampling, oversampling for imbalanced categories, and numerical normalization on network traffic data to generate a balanced and standardized training dataset.
[0120] The feature engineering module 130 is used to sequentially perform initial feature selection based on information gain, redundant feature removal based on correlation analysis, and nonlinear dimensionality reduction based on kernel principal component analysis on the training dataset to obtain an optimized feature set.
[0121] The known attack detection module 140 is used to input the optimized feature set into an ensemble learning detection layer consisting of multiple tree-based supervised learning models, and to fuse the outputs of multiple supervised learning models through a stacked ensemble strategy to identify and output the type of known attack.
[0122] The unknown attack detection module 150 is used to input network traffic data that is not identified as a known attack type as a suspicious instance into the clustering-based anomaly detection layer. The anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and anomalous attack instances based on clustering probability, and outputs the detection results of unknown attacks.
[0123] It should be noted that the specific implementation of the vehicle network multi-layer intrusion detection system 100 in this embodiment of the invention is similar to the specific implementation of the vehicle network multi-layer intrusion detection method described in the above embodiment of the invention, and therefore has similar technical effects. For details, please refer to the description of the vehicle network multi-layer intrusion detection method section. To reduce redundancy, it will not be repeated here.
[0124] Further embodiments of the present invention also disclose an electronic device, Figure 4 This is a structural block diagram of an electronic device according to an embodiment of the present invention. Figure 4 As shown, in one embodiment of the present invention, the electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; when the processor executes the computer programs stored in the memory, it implements the vehicle network multi-layer intrusion detection method as described in any of the above embodiments of the present invention.
[0125] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0126] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0127] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0128] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0129] The method provided in this invention can be applied to electronic devices. Specifically, the electronic device can be a desktop computer, a portable computer, a smart mobile terminal, a server, etc. No limitation is made herein; any electronic device that can implement this invention falls within the protection scope of this invention.
[0130] It should be noted that the specific implementation of the electronic device in the embodiments of the present invention is similar to the specific implementation described in the above embodiments of the vehicle network multi-layer intrusion detection method of the present invention, and therefore has similar technical effects. For details, please refer to the description of the vehicle network multi-layer intrusion detection method section. In order to reduce redundancy, it will not be repeated here.
[0131] A further embodiment of the present invention discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the multi-layer intrusion detection method for vehicle networking as described in any of the above embodiments of the present invention.
[0132] It should be noted that the specific implementation of the computer-readable storage medium in the embodiments of the present invention is similar to the specific implementation described in the above embodiments of the vehicle network multi-layer intrusion detection method of the present invention, and therefore has similar technical effects. For details, please refer to the description of the vehicle network multi-layer intrusion detection method section. In order to reduce redundancy, it will not be repeated here.
[0133] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example.
[0134] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A multi-layer intrusion detection method for vehicle-to-everything (V2X) networks, characterized in that, Includes the following steps: Acquire network traffic data in the connected vehicle environment; The network traffic data is subjected to cluster-based stratified sampling, oversampling for imbalanced categories, and numerical normalization to generate a balanced and standardized training dataset. The training dataset is subjected to a series of steps, including initial feature selection based on information gain, redundant feature removal based on correlation analysis, and nonlinear dimensionality reduction based on kernel principal component analysis, to obtain an optimized feature set. The optimized feature set is input into an ensemble learning detection layer consisting of multiple tree-based supervised learning models. The outputs of the multiple supervised learning models are fused through a stacked ensemble strategy to identify and output the types of known attacks. Network traffic data that is not identified as a known attack type is treated as suspicious instances and input into an anomaly detection layer based on clustering labels. The anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and anomalous attack instances based on clustering probabilities, and outputs the detection results of unknown attacks.
2. The multi-layer intrusion detection method for vehicle networking according to claim 1, characterized in that, The network traffic data originates from the vehicle's internal controller LAN bus data and / or the vehicle's communication data with external networks. The process of performing cluster-based stratified sampling, oversampling for imbalanced categories, and numerical normalization on the network traffic data includes: Based on the K-means clustering method, the network traffic data is divided into multiple clusters, and data samples are randomly extracted from each cluster according to a preset ratio to form a representative data subset. The number of clusters in the K-means clustering method is determined by Bayesian optimization with the silhouette coefficient as the objective function, and the preset ratio is adjusted according to the limitation of the target vehicle's computing resources. Based on the representative data subset, a synthetic minority oversampling method is used to generate new minority samples by linear interpolation based on minority samples and their neighboring samples, so as to balance the number of samples of each class in the representative data subset and obtain a class-balanced training dataset. The classification features in the class-balanced training dataset are numerically encoded, and all numerical features are standardized so that each feature has a uniform dimension, thus obtaining the training dataset.
3. The multi-layer intrusion detection method for vehicle networking according to claim 1, characterized in that, The process of sequentially performing initial feature selection based on information gain, redundant feature removal based on correlation analysis, and nonlinear dimensionality reduction based on kernel principal component analysis on the training dataset includes: Calculate the information gain value of each feature in the training dataset relative to its classification label, sort the features from high to low according to the information gain value, select a predetermined number of features with the highest ranking, remove features with importance below a preset threshold, and obtain an initial feature set. Based on the initial feature set, the correlation value between each pair of features is calculated. When the correlation value is greater than the preset correlation threshold, the correlation degree between each feature and the classification label is calculated, and the feature with a lower correlation degree with the classification label is removed to eliminate redundant features and obtain the feature set after redundancy removal. The preset correlation threshold is determined by Bayesian optimization. For the feature set after redundancy removal, kernel principal component analysis is used to map the original features to a high-dimensional space through the kernel function before principal component extraction is performed to reduce the feature dimension and suppress noisy features, thus obtaining the optimized feature set. The number of extracted principal components and the type of kernel function are determined by Bayesian optimization with the accuracy of verification as the objective function.
4. The multi-layer intrusion detection method for vehicle networking according to claim 1, characterized in that, The step of inputting the optimized feature set into an ensemble learning detection layer composed of multiple tree-based supervised learning models, and fusing the outputs of the multiple supervised learning models through a stacked ensemble strategy to identify and output the type of known attack includes: Using the optimized feature set as input, multiple tree-based supervised learning models are trained as base learners to obtain each base learner and its prediction output for training samples. The multiple tree-based supervised learning models include decision tree model, random forest model, extremely random tree model and extreme gradient boosting model. For each base learner, a Bayesian optimization method based on a tree-structured Parzen estimator is used to determine the optimal hyperparameter combination for each base learner with the detection performance of the model on the validation set as the objective, and each base learner is retrained based on the optimal hyperparameter combination. Obtain the predicted labels output by each of the retrained base learners for the same input data, use each predicted label as a new feature, train a meta-learner, and use the meta-learner to fuse the prediction results of each base learner and output the final known attack type determination.
5. The multi-layer intrusion detection method for vehicle networking according to claim 1, characterized in that, The process involves inputting network traffic data not identified as known attack types as suspicious instances into a clustering-based anomaly detection layer. This anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and anomalous attack instances based on clustering probabilities, and outputs the detection results for unknown attacks. This includes: Network traffic data that is not identified as a known attack type is used as suspicious instances. The suspicious instances are clustered using a K-means-based clustering labeling model. Each data point is assigned to the nearest neighbor cluster, and the cluster is labeled as normal or attack based on the known labels of the majority of data samples in each cluster. This determines the initial classification result and the corresponding clustering probability of each data point. The number of clusters and the distance metric in the clustering labeling model are determined by Bayesian optimization based on Gaussian process with classification accuracy as the objective function. Suspicious instances whose clustering probability is lower than a preset probability threshold are classified as uncertain instances, wherein the preset probability threshold is determined by Bayesian optimization based on Gaussian process with verification accuracy as the objective function; A first bias classifier and a second bias classifier are pre-trained, wherein the first bias classifier is trained by false negative samples generated by the clustering labeling model and randomly sampled normal data, and the first bias classifier is used to reduce the false negative rate; the second bias classifier is trained by false positive samples generated by the clustering labeling model and randomly sampled attack data, and the second bias classifier is used to reduce the false positive rate. For the uncertain instance, if its initial classification result is normal, it is input into the first bias classifier for reclassification; if its initial classification result is attack, it is input into the second bias classifier for reclassification. The reclassification result is used as the final output of the uncertain instance to obtain the detection result of the unknown attack.
6. The multi-layer intrusion detection method for vehicle networking according to claim 1, characterized in that, When identifying and outputting the type of a known attack, the type of the known attack is directly output as the final detection result of the corresponding network traffic data. After identifying normal instances and abnormal attack instances based on clustering probability, the system outputs normal data packets or unknown attacks as the final detection results of the corresponding network traffic data based on the identification results.
7. The multi-layer intrusion detection method for vehicle networking according to claim 5, characterized in that, The distance metric includes one of Euclidean distance, Manhattan distance, or Mahalanobis distance.
8. A multi-layer intrusion detection system for vehicle networking, characterized in that, include: The data acquisition module is used to acquire network traffic data in the vehicle networking environment; The data processing module is used to perform cluster-based hierarchical sampling, oversampling for imbalanced categories, and numerical normalization on the network traffic data to generate a balanced and standardized training dataset. The feature engineering module is used to sequentially perform initial feature selection based on information gain, redundant feature removal based on correlation analysis, and nonlinear dimensionality reduction based on kernel principal component analysis on the training dataset to obtain an optimized feature set. The known attack detection module is used to input the optimized feature set into an ensemble learning detection layer composed of multiple tree-based supervised learning models, and to fuse the outputs of the multiple supervised learning models through a stacked ensemble strategy to identify and output the type of known attack. The unknown attack detection module is used to input network traffic data that is not identified as a known attack type as a suspicious instance into an anomaly detection layer based on clustering labels. The anomaly detection layer clusters and labels the suspicious instances using an unsupervised clustering method, identifies normal instances and abnormal attack instances based on clustering probability, and outputs the detection result of unknown attacks.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the steps of the multi-layer intrusion detection method for vehicle networking as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the multi-layer intrusion detection method for vehicle networking as described in any one of claims 1-7.