CAN bus data detection method based on improved auto-encoder
Through the improved autoencoder, the detection of vehicle CAN bus data has been solved, and the existing system's shortcomings in detecting mixed ID segments and Data segment data has been achieved, efficient, real-time and accurate detection has been achieved, improving the safety of the car.
Patent Information
- Application Number
- CN202510261117.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-13
AI Technical Summary
The existing vehicle CAN bus intrusion detection system is difficult to accurately detect hybrid ID segments and Data segment data. It has high real-time requirements but cumbersome operation and low accuracy, which affects the safety of the car.
The CAN bus data detection method based on the improved autoencoder is adopted. By acquiring CAN bus data in real time, integrating ID hybrid data segments into a two-dimensional matrix, inputting the improved autoencoder for feature extraction and classification detection, real-time detection and attack type annotation are realized.
It improves detection rate and speed, ensures the accuracy and real-time detection, and can accurately label the attack types of abnormal data, enhancing the safety of the car.
Smart Images

Figure CN120151003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data detection, and in particular to a CAN bus data detection method based on an improved autoencoder. Background Art
[0002] Currently, in the vehicle CAN bus intrusion detection system, the problems existing at home and abroad are as follows: (1) Some detection systems cannot detect the data in the mixed ID segment and Data segment, and it is difficult to accurately detect and label all attack messages. (2) Intelligent connected vehicles have extremely high requirements for real-time performance. However, the current intrusion detection systems have some deficiencies, including cumbersome operations and interference with normal CAN bus messages, which make real-time detection difficult. (3) The accuracy rates of most detection systems are not high, making it difficult to ensure the safety of vehicles. Summary of the Invention
[0003] To solve the above problems, the present invention proposes a CAN bus data detection method based on an improved autoencoder. This detection system has a high detection rate, fast speed, high accuracy, and can perform real-time detection. If an anomaly is detected, it will also label which known attack type the specific abnormal data belongs to.
[0004] To achieve the above object, the technical solution adopted by the present invention is:
[0005] A CAN bus data detection method based on an improved autoencoder includes the following steps:
[0006] S1. Real-time obtain the vehicle CAN bus data. The normal data is obtained from the OBD interface of the vehicle during driving, and the abnormal data is simulated by CANoe.
[0007] S2. In the abnormal data, integrate the ID and Data segments into an n*m two-dimensional matrix with n frames in a group, where the m dimension is the sum of the k bits of the ID and the n bits of the Data.
[0008] S3. Input the processed CAN bus data containing the ID segment and DATA segment into the improved autoencoder for feature extraction and subsequent classification detection model, so as to obtain the expected classification detection result.
[0009] Preferably, the feature structure of the autoencoder is an encoder composed of a padding layer, three convolutional layers, and a bottleneck layer arranged in sequence according to the data transmission direction; where
[0010] The padding layer is used to ensure that the lengths of all input n*m two-dimensional matrices are the same;
[0011] The convolutional layer is used to increase the number of channels, enabling it to expand the receptive field and capture the feature relationships within a larger time window;
[0012] The bottleneck layer is used to compress the high-dimensional features of the data into low dimensions and remove redundant and unimportant information, only retaining the most representative key information
[0013] Preferably, the training method for the autoencoder is as follows:
[0014] S31. Use the mean squared error as the loss function to calculate the difference between the input data and the output of the decoder;
[0015] S32. Calculate the gradient of the loss function with respect to the model parameters and update the model parameters using an optimizer;
[0016] S33. Repeat the above S31 and S32 until the loss function converges.
[0017] Preferably, the classification and detection module includes a position encoding layer, a sparse attention mechanism layer, a Transform encoding layer, a pooling layer, and a fully connected layer arranged in sequence according to the data transmission direction.
[0018] Preferably, the position encoding layer uses sine-cosine encoding. The sine-cosine encoding formula corresponding to the encoding value of the i-th dimension at the pos-th time step is as follows:
[0019]
[0020] where pos is the position of the time step; i is the feature dimension index; and d is the dimension of the feature vector.
[0021] Preferably, the steps for the sparse attention mechanism layer to dynamically select global and local attention are as follows:
[0022] S41. Calculate the probability distribution of the attention weights: Use the following formula:
[0023]
[0024] where P ij represents the correlation between the i-th Query and the j-th Key;
[0025] S42. Select global and local attention: Calculate the correlation probability between each time step's Query and its local neighborhood, and select the time step with the largest probability value for calculation. The global time steps include: the start, end, and the time steps where abnormal signals are located in the sequence.
[0026] Preferably, in the abnormal data,
[0027] Inject a large number of messages with a value of 0 as a DoS attack;
[0028] Inject messages similar to normal messages as a Fuzzy attack;
[0029] Inject error messages for gear rotation speed as the RPM\Gear attack type.
[0030] The beneficial effects of using the present invention are:
[0031] This method can detect the data of the vehicle CAN bus in real time, determine whether there is an abnormality in the message, and can also display the attack type of the abnormal message. Brief Description of the Drawings
[0032] Figure 1 It is a flowchart of a CAN bus data detection method based on an improved autoencoder according to the present invention;
[0033] Figure 2 It is a schematic structural diagram of a CAN bus data detection method based on an improved autoencoder according to the present invention. Detailed Embodiment
[0034] To make the purpose, technical solution and advantages of the present technical solution clearer, the present technical solution will be further described in detail below in conjunction with the specific embodiments. It should be understood that these descriptions are exemplary and not intended to limit the scope of the present technical solution.
[0035] As Figure 1 and Figure 2 shown. This embodiment proposes a CAN bus data detection method based on an improved autoencoder, including:
[0036] Obtain vehicle CAN bus messages in real time. Normal data is obtained from the OBD interface of the vehicle during driving, and abnormal data is formed by CANoe experiment simulation;
[0037] Integrate and process the ID mixed Data segment into a two-dimensional matrix of 64 frames in a group of 64*93, where the 93 dimensions refer to the sum of 29 bits of the ID and 64 bits of the Data;
[0038] Input the processed CAN bus data including the ID segment and the DATA segment into the improved autoencoder for feature extraction and subsequent classification detection model, so as to obtain the expected classification detection result.
[0039] The autoencoder feature structure consists of an encoder composed of a padding layer, three convolutional layers and a bottleneck layer, and the decoder is the reverse engineering of the encoder. The specific process and operation steps of the input data passing through the feature extraction of the autoencoder are as follows:
[0040] The role of the padding layer is that 64*94 with an even dimension helps the CNN process.
[0041] The role of the convolutional layer is to increase the number of channels, enabling it to expand the receptive field and capture the feature relationships within a larger time window. The specific calculation formula is:
[0042]
[0043] where k h and k w are the height and width of the convolutional kernel, Input is the input feature map, Kernel is the convolutional kernel weight, and Bias is the bias term. For the ReLU function, the specific calculation formula is:
[0044] ReLU(x) = max(0, x),
[0045] This function is adopted because compared with activation functions such as Sigmoid or Tanh, ReLU does not require complex exponential operations, so it is faster during the training and inference processes. Moreover, the gradient in the positive interval is constantly 1 and does not decay as the number of network layers increases, thus effectively alleviating the vanishing gradient problem and making the training of deep networks more stable. Consider the above operations as one convolutional layer.
[0046] The bottleneck layer after convolutional operations of more than three layers is actually also a type of convolutional calculation, except that the convolutional kernel has changed:
[0047] Conv2d(384, 128, kernel_size = 1, stride = 1)
[0048] Its main role is: to compress high-dimensional features into a lower dimension, remove redundant and unimportant information, and only retain the most representative key information.
[0049] The decoder structure is the reverse engineering of the encoder, and its convolutional calculation formula is:
[0050]
[0051] The sigmod activation function is used:
[0052]
[0053] Since the goal of the autoencoder is to reconstruct the input data, the output of the decoder needs to be as close as possible to the original input. Using Sigmoid can ensure that the output value of the decoder matches the range of the input data, facilitating the calculation of the reconstruction error (such as the mean square error).
[0054] The goal of autoencoder training is to minimize the error between the input data and the reconstructed data, so that the data representation learned by the model can retain important features as much as possible while removing irrelevant information. During the training process, the mean squared error is used as the loss function:
[0055]
[0056] Calculate the difference between the input data and the output of the decoder. Where N is the number of samples, Input i is the original input data, Reconstructed i is the reconstructed data output by the decoder (obtained by the sigmoid function). The mean squared error (MSE) is used to calculate the difference between the input data and the reconstructed data. The smaller the value of MSE, the closer the reconstructed data is to the original input data. Backpropagation: Calculate the gradient of the loss function with respect to the model parameters and update the model parameters using an optimizer (Adam). Repeat the above steps until the loss function converges.
[0057] The classification detection module is a deep learning model for long time series, including a positional encoding layer, a sparse attention mechanism layer, a Transform encoding layer, a pooling layer, and a fully connected layer. Its core is the sparse self-attention mechanism, which aims to efficiently model the global dependencies and pattern changes in long sequence data. This mechanism can effectively capture long-distance dependencies by allocating weights between different positions in the input data, thereby realizing the recognition of complex patterns in time series. This has significant advantages for processing CAN bus data, especially in anomaly detection and fault warning, because CAN bus data usually contains a large amount of time series information and complex temporal patterns.
[0058] The specific steps for designing the detection classification module are as follows:
[0059] The positional encoding uses sine-cosine encoding. Its main purpose is to help the model distinguish the order of each time step or symbol in the sequence. In time series data, it helps the model capture temporal dependencies. The sine-cosine encoding formula is as follows: The encoding value of the i-th dimension at the pos-th time step is: where pos is the position of the time step; i is the feature dimension index (such as 0-255); d is the dimension of the feature vector (such as 256), specifically as follows:
[0060]
[0061] In the above formula, the inputs of the sine and cosine functions determine their frequencies. For example, sin(x) will repeat periodically as x increases. The denominator 10000 in the formula 2i / dThe range of frequency variation is controlled. When i = 0, the frequency is the lowest (corresponding to a larger period); when i increases, the frequency gradually rises (a smaller period). Taking 10000 as the base, the low-dimensional encoding of the feature vector can capture the global patterns of long time steps (low-frequency variations), and the high-dimensional encoding can capture the local details of short time steps (high-frequency variations). If the base is too small, when the position value grows rapidly, the variation of the high-dimensional frequency will be too drastic, resulting in the model being unable to capture the relationships between long-distance time steps. If the base is too large, the frequency difference between adjacent time steps in the sequence will become too small, and the perception of the position encoding for short-distance step lengths will be weakened.
[0062]
[0063] The input of the sine and cosine functions is, where the exponential scaling 2i / d dynamically adjusts the size of the denominator according to the feature dimension d: when i is small, the denominator is close to 1, and the input changes of sin and cos are large (capturing local high-frequency features). When i is large, the denominator becomes larger, and the input changes are small (capturing global low-frequency features). Taking 10000 as the base enables this scaling mechanism to change smoothly and cover both global and local information.
[0064] The calculation formula of the traditional attention mechanism is: where Q, K, and V are the Query, Key, and Value matrices. Here, Q, K, and V are the Query, Key, and Value matrices. If the dimension is [Q, K]. This requires calculating the dot product of Q at all time steps with K at all time steps (A = QK T , Q: a matrix of shape [L, d] with a d-dimensional vector for each time step, K T : the transpose matrix of shape [d, L]), and the result is an attention matrix A of shape [L, L] with a complexity of O(L 2 ).
[0065]
[0066] It is assumed that most values in the attention matrix are insignificant (close to 0), and only a small number of highly correlated values contribute to the final result.
[0067] In this embodiment, a probabilistic method is adopted. According to the norm of Query or other heuristic rules, a part of the "important Queries" are selected, and only the attention weights corresponding to these important Queries are calculated, while other Queries are ignored. In this way, the attention matrix A changes from the original fully connected [L, L] to a sparse [L, k], where k << L (usually k = logL). The specific steps for dynamically selecting global and local attention are as follows:
[0068] Step 1: Calculate the probability distribution of attention weights: For each Query, the attention weight distribution can be calculated through a softmax operation to obtain which Keys the Query should focus on. This distribution can be achieved through the following probability calculation:
[0069]
[0070] where P ij represents the correlation between the i-th Query and the j-th Key (i.e., the probability of attention allocation).
[0071] Step 2: Select global and local attention: Calculate the correlation probability between each time-step Query and its local neighborhood, and select the time-step with a larger probability value for calculation. In this embodiment, important global time-steps are dynamically selected, such as the start, end, and time-steps where abnormal signals are located in the sequence.
[0072] Through the above probability distribution, in this embodiment, it can be determined which Keys each time-step Query should interact with. At this time, the attention matrix constructed in this embodiment will become sparse, that is, only a part of the elements are non-zero, and the other elements are zero. Local attention: For each Query, in this embodiment, only the attention calculation is performed with its adjacent k time-steps, and the correlation of other time-steps is zero. Global attention: In this embodiment, several globally important time-steps (such as abnormal points) are selected to interact with all time-steps, and the correlation of other time-steps is still zero. Therefore, the final attention matrix is a sparse matrix, which contains key local and global dependencies. The computational complexity is reduced from the traditional O(L 2 ) to O(L·k).
[0073] Finally, the specific operations and functions of the Transformer encoding layer, pooling layer, and fully connected layer: The core role of the Transformer encoder is to transform the input sequence into a context-sensitive feature representation through the multi-head attention mechanism and the feed-forward network, capturing global and local dependencies, so as to provide efficient and accurate input for subsequent tasks (anomaly detection). The core of the pooling layer is to compress the size of the input data, extract the main features, while reducing the computational amount and noise interference. It makes the model more robust to the position changes of the input data, and at the same time avoids overfitting. In time-series data, it is especially suitable for highlighting global trends or abnormal signals.
[0074] The fully connected layer and the Softmax activation function are in the last layer of the classification task, and their functions can be divided into two parts: feature integration (fully connected layer) and probability distribution generation (Softmax activation function). The formula of Softmax:
[0075]
[0076] Fully connected layer: Integrates the features of the input data to generate scores for each category. Softmax activation function: Converts these scores into a probability distribution, thus intuitively indicating the likelihood that the input belongs to each category, thereby completing the detection task.
[0077] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0078] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0079] In addition, in each embodiment of this application, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0080] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes. The above content is only the preferred embodiment of the present invention. For those of ordinary skill in the art, many changes can be made in the specific implementation manners and application scopes according to the idea of the technical content of this application. As long as these changes do not depart from the concept of the present invention, they all fall within the protection scope of this patent.
Claims
1. A CAN bus data detection method based on an improved autoencoder, characterized in that: The steps include: S1. Real-time acquisition of vehicle CAN bus data. Normal data is acquired from the OBD interface of the vehicle, and abnormal data is simulated by CANoe. S2. Integrate the ID mixed Data segments in the abnormal data into a two-dimensional matrix of n*m with n frames as a group, where the m dimension is the sum of the k bits of the ID and the n bits of the Data; S3. The processed CAN bus data including the ID segment and the DATA segment is input into the improved autoencoder for feature extraction and subsequent classification detection model, so as to obtain the expected classification detection result.
2. The CAN bus data detection method based on the improved autoencoder according to claim 1 is characterized in that: The characteristic structure of the autoencoder is an encoder consisting of a padding layer, three convolutional layers and a bottleneck layer arranged in sequence according to the data transmission direction; The padding layer is used to ensure that the lengths of all input n*m two-dimensional matrices are consistent; The convolutional layer is used to increase the number of channels, so that it can expand the receptive field and capture feature relationships within a larger time window; The bottleneck layer is used to compress the high-dimensional features of the data into low dimensions and remove redundant and unimportant information, retaining only the most representative key information.
3. The CAN bus data detection method based on the improved autoencoder according to claim 1 is characterized in that: The training method of the autoencoder training is as follows: S31, using mean square error as a loss function, calculating the difference between the input data and the decoder output; S32, calculating the gradient of the loss function with respect to the model parameters, and updating the model parameters using the optimizer; S33. Repeat the above S31 and S32 until the loss function converges.
4. The CAN bus data detection method based on the improved autoencoder according to claim 1 is characterized in that: The classification detection module includes a position encoding layer, a sparse attention mechanism layer, a Transform encoding layer, a pooling layer and a fully connected layer which are arranged in sequence according to the data transmission direction.
5. The CAN bus data detection method based on the improved autoencoder according to claim 4 is characterized in that: The position coding layer adopts sine and cosine coding, and the sine and cosine coding formula corresponding to the coding value of the i-th dimension at the pos-th time step is as follows: Among them, pos is the position of the time step; i is the feature dimension index; d is the dimension of the feature vector.
6. The CAN bus data detection method based on the improved autoencoder according to claim 4 is characterized in that: The steps of dynamically selecting global and local attention in the sparse attention mechanism layer are as follows: S41. Calculate the probability distribution of attention weights: Use the following formula: Among them, P ij Indicates the correlation between the i-th Query and the j-th Key; S42. Select global and local attention: Calculate the probability of correlation between the query and the local neighborhood for each time step, and select the time step with the largest probability value for calculation, where the global time step includes: the beginning, end and abnormal signal time step of the sequence.
7. The CAN bus data detection method based on the improved autoencoder according to claim 1 is characterized in that: In the abnormal data, Inject a large number of zero messages as a DoS attack; Injecting messages similar to normal messages as fuzzy attacks; Injects error messages targeting gear speed as an RPM\Gear attack type.