Illegal access detection and positioning method and device for CAN bus
Through dual sampling point data acquisition and differential processing combined with deep learning model, the problem of time-consuming and resource-consuming CAN bus malicious access detection is solved, efficient and accurate malicious access detection and positioning is achieved, and the security of the CAN bus and the overall security performance of intelligent connected vehicles are improved.
Patent Information
- Application Number
- CN202510200234.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-04
AI Technical Summary
The existing CAN bus is insecurity, malicious access detection and positioning technology is time-consuming and resource consumption, making it difficult to meet the real-time requirements of vehicles.
The CAN bus voltage data is monitored by dual-sampling point data acquisition technology, and the access detection and positioning is carried out through differential processing and feature extraction, and the access detection and positioning is carried out by multi-layer perceptron model. The data of the horizontal and vertical dimensions in the CAN bus and its changes in timing are captured, and the access detection and access point positioning is carried out through deep learning models.
It realizes accurate and efficient malicious access detection and positioning with relatively low resource consumption, improves detection accuracy and real-timeness, reduces resource consumption, and is suitable for the real-time requirements of on-board systems.
Smart Images

Figure CN120263440A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network security mainly for intelligent connected vehicles, and particularly to a method and device for detecting illegal access and locating access points for a CAN bus based on voltage characteristics. Background Art
[0002] Intelligent connected vehicles have become the core trend, integrating cutting-edge technologies in multiple fields and promoting the transformation of the automotive industry towards intelligence, networking, electrification, and sharing. Among them, the CAN (Controller Area Network) bus is a key infrastructure for in-vehicle communication, connecting numerous electronic control units and ensuring the rapid and accurate interaction of information between various systems.
[0003] However, when the CAN bus was designed, emphasis was placed on communication efficiency, with insufficient consideration for security. Its communication protocol has no identity authentication mechanism, and any access node can send information without verification, allowing malicious attackers to easily access and eavesdrop on sensitive information such as vehicle location, driving habits, and system status, seriously infringing on user privacy. They can also manipulate key vehicle systems such as acceleration, braking, and steering by injecting forged messages, directly threatening driving safety. At the same time, the CAN bus is not equipped with an encryption mechanism, and communication data is transmitted in plaintext, making it easy to be monitored and tampered with. Moreover, the message length limit makes it difficult to embed sufficient security metadata for encryption and authentication operations. In addition, the CAN bus lacks a built-in intrusion detection mechanism and has low physical interface security. For example, the OBD-II port (On-Board Diagnostics II) is easily invaded.
[0004] Currently, research on CAN bus security mainly focuses on intrusion detection systems, including clock-based and voltage-based methods. For example, clock-based methods utilize the periodic characteristics of CAN message transmission to detect intrusion by analyzing message intervals or frequency changes. Another example is voltage-based methods, which rely on the voltage output characteristic differences between ECU transceivers to detect intrusion and identify the source. However, these methods have limitations. The main reason is that most existing CAN bus access detection and location technology solutions require analyzing multiple data frames of multiple ECU nodes, which is time-consuming and relies on high sampling frequencies and complex processing procedures. This increases hardware costs and computational resource consumption, prolongs data processing time, and is difficult to meet the strict real-time requirements of in-vehicle systems. Even a short detection delay during high-speed vehicle driving can lead to serious consequences.
[0005] Therefore, how to design a more accurate, efficient, and relatively low-resource-consuming malicious access detection and location scheme has become a research topic that needs to be addressed. Summary of the Invention
[0006] An embodiment of the present invention provides a method and device for detecting and locating illegal access to a CAN bus, which can detect malicious access accurately, efficiently and with relatively low resource consumption, and realize the location of the malicious access point.
[0007] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:
[0008] A method for detecting and locating illegal access to a CAN bus includes:
[0009] S1. The monitoring system samples the original data through two independent CAN bus data sampling points.
[0010] Among them, through two independent CAN bus data sampling points, the CANH and CANL wire harnesses are respectively monitored, and the voltage data of the data frame being transmitted on the bus is obtained as the original data.
[0011] S2. Preprocess the collected multi-dimensional voltage data, and the obtained preprocessed data includes: longitudinal differential data and transverse differential data.
[0012] Among them, the dominant level is extracted by using the voltage data of CANH and CANL respectively; differential processing is performed on the collected original data, including: obtaining the transverse differential data between different access points at the same moment, and the differential data on the two wire harnesses of the same access point.
[0013] S3. Use the preprocessed data for feature extraction to obtain a feature matrix reflecting the degree of data dispersion and distribution.
[0014] Among them, the overall trend features are extracted, including: average statistical features and standard deviation statistical features; the features of the degree of dispersion are extracted, including: the 25th percentile, median and 75th percentile in the differential vector of the preprocessed data.
[0015] S4. Train a multi-classification model with a multi-layer perceptron (MLP) as the core.
[0016] Among them, the MLP model includes three hidden layers, the ReLU activation function is used in the hidden layers, the softmax activation function is used in the output layer, the Adam optimizer is selected and the categorical cross-entropy is used as the loss function, and the accuracy is used as the performance evaluation index; in the process of training the multi-classification model in S4, iterative learning is performed for 100 epochs, and 30% of the training data is reserved as the validation set.
[0017] S5. Input the feature matrix obtained in S3 into the trained multi-classification model to obtain the location information of the illegal access point.
[0018] Among them, a detection module is established based on the electrical characteristics of the CAN bus, and the feature matrix is input into the detection module. In the detection module, malicious access in the bus topology is detected by analyzing a CAN frame; when the detection module detects malicious access, the access point is located.
[0019] An illegal access detection and positioning device for a CAN bus, comprising:
[0020] An acquisition module, configured to monitor the system through two independent CAN bus data sampling points;
[0021] A preprocessing module, configured to preprocess the data collected from the sampling points, and the obtained preprocessed data includes: original data, longitudinal difference data, and transverse difference data;
[0022] A feature analysis module, configured to extract features by using the preprocessed data to obtain a feature matrix reflecting the dispersion degree and distribution of the data;
[0023] A model training module, configured to train a multi-classification model with a multi-layer perceptron (MLP) as the core;
[0024] An analysis module, configured to input the feature matrix into the trained multi-classification model to obtain the position information of the illegal access point.
[0025] The illegal access detection and positioning method and device for a CAN bus provided by the embodiments of the present invention utilize the dual-sampling point data acquisition technology to capture the data in the horizontal and vertical dimensions in the CAN bus and their changes in time series, so as to analyze the level fluctuations caused by the changes in the bus topology, and combine the deep learning model to perform access detection and access point positioning. It realizes accurate and efficient detection of malicious access with relatively low resource consumption and locates the malicious access point. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0027] Figure 1 It is a schematic diagram of a dual-sampling point data acquisition system provided by an embodiment of the present invention.
[0028] Figure 2 It is a data acquisition and processing flow chart provided by an embodiment of the present invention.
[0029] Figure 3 It is a feature extraction flow chart provided by an embodiment of the present invention.
[0030] Figure 4 This is the access detection and positioning flowchart provided by the embodiments of the present invention.
[0031] Figure 5 This is the schematic diagram of a clean CAN bus provided by the embodiments of the present invention.
[0032] Figure 6 This is the schematic diagram of a CAN bus with malicious access provided by the embodiments of the present invention.
[0033] Figure 7 This is the overall process schematic diagram provided by the embodiments of the present invention. Among them, mean() in the figure represents mean calculation, and std() represents variance calculation. In this embodiment, they are used to represent the extracted feature parameter values. Detailed implementation manners
[0034] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any unit and all combinations of one or more related listed items. Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless defined as herein.
[0035] An embodiment of the present invention provides an illegal access detection and positioning method for a CAN bus. The main design objective is to provide an efficient, accurate, and low-resource-consuming illegal access detection and access point positioning method for the CAN bus, so as to solve the problems of untimely malicious access detection, inaccurate positioning, and high resource consumption in the prior art, and achieve effective security protection for the CAN bus. The main design idea is to fuse a dual-sampling-point data acquisition system with a deep learning model to explore the possibility of completing timely access detection and access point positioning using only the bus data of one data frame, which is expected to solve the above problems and improve the security of the CAN bus and the overall safety performance of intelligent connected vehicles.
[0036] As Figure 7 shown, the method process includes:
[0037] S1. The monitoring system samples the original data through two independent CAN bus data sampling points.
[0038] S2. Preprocess the collected multi-dimensional voltage data to obtain preprocessed data including longitudinal differential data and transverse differential data.
[0039] S3. Use the preprocessed data for feature extraction to obtain a feature matrix reflecting the dispersion degree and distribution of the data.
[0040] S4. Train a multi-classification model with a multi-layer perceptron (MLP) as the core.
[0041] S5. Input the feature matrix obtained in S3 into the trained multi-classification model to obtain the position information of the illegal access point.
[0042] In this embodiment, S1 includes: monitoring the CANH and CANL two wire harnesses respectively through two independent CAN bus data sampling points, and obtaining the voltage data of the data frame being transmitted on the bus as the original data, Data raw {CANH1, CANL1, CANH2, CANL2}, where Data raw represents the original data, and CANH1, CANL1, CANH2, and CANL2 respectively represent the voltage data on the CANH and CANL wire harnesses at the two sampling point positions in the dual-sampling-point system.
[0043] Among them, two independent CAN bus data sampling points are set to monitor the CANH and CANL two wire harnesses in real time and three-dimensionally, obtain the multi-dimensional voltage data of the data frame being transmitted on the bus, and comprehensively capture the voltage change trend of the bus.
[0044] Specifically, the designed dual-sampling-point data acquisition system adopts the dual-sampling-point data acquisition method. Through two independent CAN bus data sampling points set, it monitors the CANH and CANL wire harnesses in real time and three-dimensionally, obtains the multi-dimensional voltage data of the data frames being transmitted on the bus, and comprehensively captures the voltage change trend of the bus. This method can not only capture the longitudinal voltage changes at a single acquisition point in the time dimension, but also record the transverse voltage data at different positions at the same moment, thus providing a more comprehensive perspective for analyzing the topological state of the CAN bus.
[0045] Optionally, the sampling frequency of the sampling point is not greater than 20MS / s.
[0046] In this embodiment, S2 includes:
[0047] Using the respective voltage data of CANH and CANL, extract the dominant levels of CANH1, CANL1, CANH2, and CANL2, where
[0048]
[0049] Perform differential processing on the sequential sequence of the extracted dominant levels to obtain the differential data of the dominant levels, including: obtaining the transverse differential data Diff H 、Diff L ,Diff H represents the voltage difference between two access points on the CANH wire harness, Diff L represents the voltage difference between two access points on the CANL wire harness, and the differential data Diff1 and Diff2 on the two wire harnesses of the same access point.
[0050]
[0051] Among them, in data preprocessing, make full use of the unique advantages of the horizontal and vertical two-way perspectives of the dual-sampling-point data acquisition technology to perform differential processing on the original data. From the two dimensions of the horizontal time series and the vertical comparison of different sampling points, deeply analyze the internal correlation and potential characteristics of the data, and gradually transform the original data into a more valuable and more convenient format for subsequent operations.
[0052] Specifically, it includes: Extracting the dominant level: Among the dominant level and recessive level of the differential bus, the dominant level contains more level characteristic information of the data transmitted on the bus, including more longitudinal differential data characteristics that can be used for access detection and positioning. Therefore, the differential of CANH and CANL is used to determine the existence of the dominant level and extract it. Differential processing: Perform differential processing on the collected original data, and utilize the advantages of the dual-sampling-point data acquisition technology to comprehensively process the data from both horizontal and vertical dimensions. Specifically, it includes calculating the horizontal differential data between different access points at the same moment and the differential data on the two wire harnesses of the same access point, so as to reduce the dependence of the data on absolute values and obtain characteristic information that can better reflect the changes in the bus topology and network state adjustment.
[0053] In this embodiment, S3 includes: Extracting the overall trend characteristics, including:
[0054] Extracting the overall trend characteristics, including: the average statistical characteristic and the standard deviation statistical characteristic;
[0055]
[0056] Among them, represents the average, σ represents the standard deviation, n represents the number of statistical values, and x i represents the statistical value; Extracting the degree of dispersion characteristics, including: the 25th percentile Q1, the median Q2, and the 75th percentile Q3 in the differential vector of the preprocessed data,
[0057] Among them, in the feature extraction, a statistical feature extraction method is adopted to deeply mine the preprocessed data vector and differential vector, and refine those key features that can accurately reflect the essence of the data and hide clues of illegal access, providing strong support for the accurate judgment of the model.
[0058] Specifically, it includes: Extracting the overall trend characteristics: Adopt a statistical feature extraction method to deeply mine the data vector and differential vector. Calculate the average and standard deviation statistical characteristics, which can comprehensively capture the dynamic behavior of the CAN bus and help improve the detection accuracy of the system for malicious access behavior. For example, the average reflects the central tendency of the data, and the standard deviation reveals the overall fluctuation range of the data. Extracting the degree of dispersion characteristics: Adopt statistical characteristics for the degree of data dispersion, and calculate the 25th percentile, median, and 75th percentile in the data vector and differential vector. The 25th percentile and 75th percentile describe the degree of dispersion and distribution of the data, and the median can more accurately reflect the central position of the data when the data distribution is uneven or there are outliers.
[0059] In the model selection and training of this embodiment, a multi-layer perceptron (MLP) is used as the core multi-classification model, and it is trained in a systematic and refined manner to make it familiar with the data patterns in various legal and illegal access scenarios and have a powerful discrimination ability.
[0060] Specifically, model selection: Select a multi-layer perceptron (MLP) as the multi-classification model and train it systematically. The training set and test set can be obtained by conducting multiple experimental verifications and collecting data from the experimental scenarios constructed in the laboratory. First, reshape the feature matrices of the training set and test set to meet the model input specification requirements; then, adjust the data mean to 0 and the standard deviation to 1 through standardization processing to accelerate model convergence and improve training efficiency; next, convert the label data into one-hot encoding form; design an MLP model with three hidden layers, where the hidden layers use the ReLU activation function, the output layer uses the softmax activation function, select the Adam optimizer, and use categorical cross-entropy as the loss function and accuracy as the main performance evaluation metric. Model training: During the training process, perform iterative learning for 100 epochs, and reserve 30% of the training data as the validation set to monitor the model training progress and performance in real time; finally, evaluate the model performance on the test set, calculate the test loss and accuracy, and report the test accuracy to reflect the generalization ability of the model.
[0061] In this embodiment, S5 includes: establishing a detection module for detecting malicious access on the bus based on the electrical characteristics of the CAN bus, and inputting the feature matrix into the detection module. In the detection module, malicious access in the bus topology is detected by analyzing a CAN frame.
[0062] Among them, malicious access on the bus will cause changes in the circuit topology on the bus, resulting in a slight change in the bus voltage. This change is used as the electrical characteristic and can be used to detect whether there is malicious access on the bus and determine its specific access location. Specifically, features can be extracted from the voltage data of the CAN frame currently transmitted on the bus and sent into the trained MLP model. The parameters output by the model represent whether the current bus is in a normal state or a malicious access state.
[0063] When the detection module detects malicious access, locate the access point.
[0064] Among them, if the model prediction result is a malicious access state, map the prediction parameters to the access environments of different preset access points, and the specific access point location can be obtained. Input the feature matrix into the multi-classification model that has been fully trained in advance. The model quickly conducts analysis and prediction based on the learned knowledge, accurately determines whether the access is illegal, and precisely locks the specific location information of the access point, thereby ensuring the safety of the CAN bus.
[0065] Specifically, it includes: Access detection: In this stage, the feature matrix generated by step (3) is used as the input. Based on the electrical characteristics of the CAN bus, the module detects malicious access in the bus topology by analyzing a CAN frame, regardless of which ECU the frame comes from. A classifier is trained for each ECU. When the classifier receives the ID and the feature matrix, the feature matrix is input into the pre-trained classification model for analysis and prediction. If the prediction result shows no malicious access, the access detection process is terminated and the communication data is continuously monitored; if malicious access is detected, go to step (52) to locate the specific access position of the malicious node and generate a warning message. Access point location: In the access detection stage, the model generates a multi-class code as the prediction result based on the input feature matrix. Subsequently, according to the mapping relationship established in the training stage, during the access point location process, the prediction result is accurately mapped to the real access point to determine its exact physical location. In this way, the accurate location of the malicious access node can be achieved, providing key information for taking effective security measures in a timely manner and ensuring the safe and stable operation of the CAN bus.
[0066] This embodiment also provides an illegal access detection and location device for the CAN bus, including:
[0067] An acquisition module, used to monitor the system through two independent CAN bus data sampling points;
[0068] A preprocessing module, used to preprocess the data collected from the sampling points, and the obtained preprocessed data includes: original data, longitudinal difference data, and transverse difference data;
[0069] A feature analysis module, used to extract features using the preprocessed data to obtain a feature matrix reflecting the dispersion degree and distribution of the data;
[0070] A model training module, used to train a multi-classification model with a multi-layer perceptron (MLP) as the core;
[0071] An analysis module, used to input the feature matrix into the trained multi-classification model to obtain the location information of the illegal access point.
[0072] The following further describes this embodiment with an actual application scenario as an example:
[0073] First, a prototype system is constructed, and an experimental prototype system is constructed with Raspberry Pi 4B and MCP2515 as core components. Raspberry Pi 4B assumes the role of a microcomputer motherboard and uses the ARM architecture to provide powerful computing power; MCP2515, as a high-performance independent CAN protocol controller, complies with the CAN V2.0B technical specification to ensure compatibility and reliability with CAN bus communication. MCP2515 is connected to the microcontroller unit (MCU) via SPI to simplify the system architecture and enhance communication efficiency. The oscilloscope MSO5104 is used to accurately capture the voltage signal in the CAN bus. It has a real-time sampling rate of up to 4 GSa / s, and can capture continuously and rapidly changing voltage signals under high-speed data transmission of 500k, meeting the requirements of the dual sampling point data acquisition method, and can simultaneously capture four serial data streams (such as Figure 1 As shown in Figure 2), it provides a comprehensive data view for analyzing CAN bus communication data. The dual sampling point data acquisition system layout is as follows: Figure 1 shown.
[0074] In the key link of scenario design, a variety of network environments with different configurations were built to carry out experimental evaluation work in an all-round and multi-angle manner. Among them, it covers both normal operation of CAN networks without malicious access (such as Figure 5 In the scenario 0 shown in the figure, each ECU node exchanges information in an orderly manner according to the established rules, providing a benchmark reference for the normal communication status of the CAN bus. At the same time, a variety of malicious access scenarios are also designed (such as Figure 6 Scene 1 to Scene 6 shown in the figure simulate various intrusion methods that malicious attackers may take. In a normally operating CAN network, if there are N normally operating ECU nodes, there are N+1 potential malicious access points. These network access locations are meticulously abstractly numbered, which provides an important basis for the subsequent accurate identification and positioning of malicious access points. For malicious access scenarios, malicious ECUs are cleverly inserted into different positions of the CAN bus for simulation. Malicious nodes may use a variety of cunning methods in this process. They may be like "spies" hiding in the dark, disguised as normal ECU nodes, quietly lurking in the bus, waiting for an opportunity to steal sensitive information; or like "saboteurs", actively injecting forged data, trying to interfere with the normal communication order and destroy the stable operation of the CAN bus. By designing these different types of malicious access scenarios, the performance of the detection and positioning system in the face of various complex attacks can be comprehensively and deeply evaluated, so as to continuously optimize and improve the system's defense capabilities and ensure the security and reliability of the CAN bus in practical applications.
[0075] 1. Data collection. In the data collection stage, the collection of the training set mainly starts from the legal CAN network Scene0. For each ECU, 500 frames of level signals are collected, and these signals comprehensively cover the typical data states in the normal communication process. At the same time, for the attack scenarios Scene1 to Scene6 with malicious access nodes, for each ECU in each attack scenario, an additional 500 frames of level signals are collected. These signals contain the abnormal changes in the level caused by the access of malicious nodes, thus simulating effective attack samples for model training. When collecting the test set, 300 frames of level signals are collected for each ECU from the legal CAN network Scene0 to evaluate the performance of the model in the normal communication scenario, and its expected prediction result is a specific identifier. From the malicious access attack scenarios Scene1 to Scene6, an additional 500 frames of level signals are collected for each network with a malicious access scenario to evaluate the detection ability of the model when facing different malicious access attacks. The expected prediction result is the specific access location of the corresponding malicious access network, and the malicious access nodes used in the test stage are from different manufacturers and have different transceivers, so as to comprehensively test the detection performance of the model.
[0076] 2. Data preprocessing. Given that the actual voltage signals of the CAN bus are affected by various complex factors, resulting in a non-ideal state, it is necessary to perform differential processing operations on the collected raw data. In the specific operation process, according to the designed double-sampling-point data collection method, the level signals are accurately captured, and then the differential vector calculation work is carried out on the collected data. As Figure 2 shown, by making full use of the unique advantages of the double-sampling-point data collection, Diff_1 (the difference between CAN_H and CAN_L at sampling point 1), Diff_2 (the difference between CAN_H and CAN_L at sampling point 2), Diff_H (the difference between sampling point 1 and sampling point 2 on the CAN_H line), and Diff_L (the difference between sampling point 1 and sampling point 2 on the CAN_L line) are calculated. The data obtained through differential calculation effectively eliminates the dependence on absolute values and significantly enhances the perception ability of the changes in the bus topology structure and network state adjustment. Among them, the horizontal difference can clearly reflect the mutual relationship between nodes at different positions, and the vertical difference can accurately reflect the dynamic changes of voltage in the time dimension, thus providing a more valuable data basis for the subsequent feature extraction work and strongly promoting the smooth progress of the entire detection and positioning process.
[0077] 3. Feature extraction. In the feature extraction link, as Figure 3As shown, it focuses on mining valuable feature information from the preprocessed multi-dimensional data. The extracted features cover several statistical features such as the mean, standard deviation, 25th percentile, median, and 75th percentile. They describe the overall characteristics of the data from a macro level. The mean can accurately present the central tendency of the data, the standard deviation effectively reflects the degree of dispersion of the data, and the percentiles clearly show the distribution of the data. The differential vector features focus on capturing the subtle details of voltage changes. Through a comprehensive analysis of these diverse features, the operating state of the CAN bus can be understood more deeply and comprehensively, thus enabling the accurate identification of malicious access behaviors. The rich and diverse features provide sufficient and key information for the subsequent classification model, greatly improving the accuracy of model detection and laying a solid foundation for ensuring the safe and stable operation of the CAN bus.
[0078] 4. Model Selection and Training. As an integrated unit, the access detection and localization module has the ability to directly output detection results. Its output results are divided into two cases: one is to indicate no malicious access; the other is to accurately point out the specific location of the malicious access node. The input data of this module is the feature matrix generated based on the electrical characteristics of the CAN bus. These feature matrices will be fed into a multi-layer perceptron (MLP) model with three hidden layers for classification. Before model training, the dataset is reasonably divided in a ratio of 7:3. Among them, 70% of the data is assigned as the training set for model learning and parameter optimization, and the remaining 30% of the data is used as the test set to evaluate the model performance. During the training process, the various parameters of the model are continuously adjusted until the optimal state is reached. After a series of designed experiments and evaluations, the final experimental results show that within the entire test set range, this method performs extremely well in terms of access recognition and access point localization. For example, the classification accuracy is as high as 99.2%, the precision exceeds 99.3%, the recall rate reaches 99.0%, and the F1 score is 99.2%. These excellent data fully prove that the model has excellent performance and reliability in the detection and localization of illegal access to the CAN bus, and can provide a solid guarantee for the safe and stable operation of the CAN bus.
[0079] 5. Access Detection and Localization. During the actual operation process, the data of the CAN bus will first go through processes such as preprocessing and feature extraction, and then the processed data will be input into the pre-trained MLP model. The model will conduct in-depth analysis of the data in real time to accurately predict whether there is malicious access. As Figure 4As shown, once malicious access is detected, the system will immediately take a series of security measures such as alarm and disconnection based on the access point location information output by the model, and continuously monitor the bus status to ensure that the CAN bus is always in a safe and stable operating state. If no malicious access is detected, the process returns to the initial state and starts a new round of detection work again. This cycle repeats continuously to guard the safety of the CAN bus without interruption.
[0080] The illegal access detection and location method for the CAN bus provided by the embodiments of the present invention uses dual-sampling-point data acquisition technology to capture data in two dimensions (horizontal and vertical) in the CAN bus and its temporal changes, so as to analyze the level fluctuations caused by bus topology changes, and combines a deep learning model for access detection and access point location. It realizes accurate and efficient detection of malicious access with relatively low resource consumption and locates the malicious access point. Compared with the prior art, the present invention can improve the detection accuracy: through dual-sampling-point data acquisition and differential vector calculation, it can more sensitively perceive the bus topology changes and voltage fluctuations caused by the access of malicious nodes, and the extracted key statistical features help to accurately identify malicious access behaviors, thereby improving the detection accuracy; and it realizes precise positioning: a classifier is trained for each ECU, which can accurately locate the access position when malicious access is detected, providing strong support for taking safety measures in a timely manner; at the same time, it also realizes the purpose of reducing resource consumption: high-precision detection can be achieved by using a relatively low sampling frequency (20MS / s), effectively reducing the resource consumption compared with the prior art, and is more suitable for resource-constrained environments such as in-vehicle CAN networks. In addition, it also improves the real-time performance: the detection and location work can be completed with only one data frame. Compared with the previous solutions that rely on multiple data frames, the real-time performance of the system is greatly improved, and it can respond to potential security threats more timely.
[0081] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments. As described above, only the specific implementation manners of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An illegal access detection and location method for a CAN bus, characterized in that, Including: S1. The monitoring system samples data through two independent CAN bus sampling points; S2. Preprocess the data collected from the sampling points, and the preprocessed data includes: original data, longitudinal differential data, and lateral differential data; S3. Use the preprocessed data for feature extraction to obtain a feature matrix reflecting the degree of data dispersion and distribution; S4. Train a multi-classification model with a multi-layer perceptron (MLP) as the core; S5. Input the feature matrix obtained in S3 into the trained multi-classification model to obtain the location information of illegal access points.
2. The method according to claim 1, wherein S1 includes: Monitor the CANH and CANL wire harnesses respectively through two independent CAN bus data sampling points, and obtain the voltage data of the data frame being transmitted on the bus as the original data, Data raw {CANH1, CANL1, CANH2, CANL2}, where Data raw represents the original data, and CANH1, CANL1, CANH2, and CANL2 respectively represent the voltage data on the CANH and CANL wire harnesses at two sampling point positions in the dual-sampling point system.
3. The method according to claim 1, wherein S2 Including: Using the voltage data of CANH and CANL respectively, extract the dominant levels of CANH1, CANL1, CANH2, and CANL2, where Differentiate the sequential sequence of the extracted dominant levels to obtain the differential data of the dominant levels, including: obtaining the horizontal differential data Diff between different access points at the same moment H , Diff L , Diff H indicates the voltage difference between two access points on the CANH harness, and Diff L indicates the voltage difference between two access points on the CANL harness, and the differential data Diff1 and Diff2 on the two harnesses of the same access point 4. The method according to claim 1, wherein S3 includes: Extract overall trend features, including: average statistical features and standard deviation statistical features; Among them, represents the average, σ represents the standard deviation, n represents the number of statistical values, and x i represents the statistical value; Extract features of the degree of dispersion, including: the 25th percentile Q1, median Q2, and 75th percentile Q3 in the difference vector of the preprocessed data 5. The method according to claim 1, wherein The MLP model includes three hidden layers, the hidden layers use the ReLU activation function, the output layer uses the softmax activation function, selects the Adam optimizer, and uses categorical cross-entropy as the loss function and accuracy as the performance evaluation index; During the process of training the multi-classification model in S4, perform iterative learning through 100 epochs and reserve 30% of the training data as the validation set.
6. The method according to claim 1, wherein S5 Including: Based on the electrical characteristics of the CAN bus, establish a detection module for detecting malicious access on the bus, and input the feature matrix into the detection module. The detection module detects malicious access in the bus topology by analyzing a CAN frame; When the detection module detects malicious access, locate the access point.
7. The method according to claim 1, characterized in that, The sampling frequency of the sampling point is not greater than 20MS / s.
8. An illegal access detection and positioning device for a CAN bus, characterized in that, Including: Adoption module, used for the monitoring system to sample data through two independent CAN bus sampling points; Preprocessing module, used to preprocess the data collected from the sampling points, and the preprocessed data includes: original data, longitudinal differential data, and lateral differential data; Feature analysis module, used to perform feature extraction using the preprocessed data to obtain a feature matrix reflecting the degree of data dispersion and distribution; Model training module, used to train a multi-classification model with a multi-layer perceptron (MLP) as the core; Analysis module, used to input the feature matrix into the trained multi-classification model to obtain the location information of illegal access points.