Semi-supervised in-vehicle CAN bus anomaly detection method and device based on dynamic graph
By adopting a semi-supervised method based on dynamic graphs in the detection of in-vehicle CAN bus abnormality, undirected time dynamic graphs are constructed and graph embedded, the problem of unawareness and high false alarm rate in the prior art is solved, and the effective detection of CAN network attacks and the effect of reducing false alarm rate is achieved.
Patent Information
- Application Number
- CN202510118247.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The prior art has defects in the detection of in-vehicle CAN bus abnormality, which cannot detect unknown or new attacks. Signature-based methods require continuous update of signature lists, and there is a delay in extracting attack features in real time on mobile vehicles; the abnormal detection mechanism has a high false alarm rate.
The semi-supervised in-vehicle CAN bus anomaly detection method based on dynamic graphs is adopted. By constructing an undirected dynamic graph of time, the time graph attention network is used to convert the timestamp into a time embedding vector, the low-dimensional representation of nodes and edges is learned, the abnormal score is predicted, and the abnormal score is recorded through the time memory database, the statistical distribution of normal nodes is calculated, the pseudo-label is generated for supervised learning, and the abnormal detection model is trained.
Without the need to know the vehicle CAN IDs, CAN network attacks on real vehicles can be detected, reducing the problem that abnormal detection depends on CAN IDs, and reducing the false alarm rate by leveraging the statistical distribution of unlabeled samples.
Smart Images

Figure CN120017342A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anomaly detection, and in particular to a method and device for detecting anomalies of a semi-supervisory in-vehicle CAN bus based on a dynamic graph. Background Art
[0002] With the advancement of the times, the networking and informatization of automobiles have become the main development directions of the Internet of Vehicles (IoV), greatly enriching vehicle functionality and improving the user experience. Modern vehicles have become mobile information platforms composed of a large number of interconnected embedded subsystems. Within each subsystem, an ECU (Electronic Control Unit) controls corresponding mechanical components and interconnects with vehicle body networks such as the CAN (Controller Area Network) and the LIN (Local Interconnect Network). As vehicles become increasingly intelligent and networked, the number of internal electronic control units (ECUs) increases, and the electronic control systems become increasingly complex. The amount of information exchange between these onboard electronic devices and ECUs and the external world is also increasing. Most of these onboard devices and ECUs are connected to the vehicle's internal bus network, making the vehicle itself a complex network system. The in-vehicle bus network facilitates information exchange between ECUs within the vehicle. The ECUs in the vehicle's control system communicate via different bus networks, which are connected by gateways to form the in-vehicle bus network system. The CAN bus developed by BOSH has become the most widely used in-vehicle bus network technology due to its advantages such as strong real-time performance, long transmission distance, strong anti-electromagnetic interference ability and low cost.
[0003] When the CAN bus was first designed, information security was not considered. Therefore, it only has basic integrity checks and lacks information security protection technologies and methods. Attackers can penetrate the vehicle's CAN bus network, which connects to key control units, by accessing physical interfaces (such as OBD-II (On-Board Diagnostics-II) and media players) or wireless access interfaces (such as Wi-Fi and Bluetooth). Attackers can then send malicious attack messages through the CAN bus network, interfering with vehicle operation and seriously endangering the life, property, and information security of the driver, passengers, and other road users. Therefore, the information security of the vehicle's CAN bus network is a key issue in the information security of connected vehicles.
[0004] The reason for this is that the CAN bus design has the following vulnerabilities: First, the CAN bus lacks information identity authentication, that is, the CAN bus data frame has no address field. Any malicious node can disguise itself as a normal node and send messages on the bus, and the receiving node cannot determine whether the sender is a genuine node. Second, the data sent by the node is not encrypted, making it easy to sniff, spoof, modify, and replay it. Then, the CAN bus uses an ID-based arbitration mechanism. Malicious attackers can exploit this mechanism to construct higher-priority data frames and continuously send them to the bus, causing a DoS attack. Finally, the bandwidth on the CAN bus is limited. The high-speed CAN bus has a data transmission rate of only 500kbit / s, and the maximum data payload length is only 64 bits. This limits the CAN bus protocol's ability to provide strong access control functions, making it easier for attackers to attack ECUs.
[0005] Researchers both domestically and internationally have conducted in-depth research on these issues. Current in-vehicle network security solutions can be broadly categorized into three approaches: message authentication, message encryption, and intrusion detection. While message authentication and encryption have proven effective for Internet security and provide a certain degree of in-vehicle security, their application in in-vehicle networks is limited by the computational performance and real-time requirements of the ECU, making them ineffective in protecting bus security. In contrast, intrusion detection methods operate passively within the vehicle, requiring no changes to existing network and protocol specifications compared to some authentication and encryption methods. This makes them a key research area for automotive bus intrusion detection.
[0006] The two main research directions of intrusion detection are: signature-based methods and anomaly-based methods. The following will introduce and summarize their shortcomings respectively.
[0007] The signature-based anomaly detection method uses known attack behaviors to extract features from the data set and design an intrusion detection model. Jin et al. [Shiyi Jin, Jin-Gyun Chung, and Yinan Xu, Signature-Based Intrusion Detection System (IDS) for In-Vehicle CAN Bus Network. In 2021 IEEE International Symposium on Circuits and Systems (ISCAS), 22-28 May 2021, Daegu, Korea.] proposed a lightweight signature-based intrusion detection system and analyzed the impact of drop attacks (Drop Attack), replay attacks (Replay Attack) and tampering attacks (Tampering Attack) on the CAN bus. In order to detect the above attacks, the article selected five signatures: ID, time interval, correlation, context change amplitude and value range. For ID signatures, a whitelist mechanism is used; for context change amplitude and value range signatures, threshold statistical analysis is used; for time interval signatures, the time interval between message arrivals is calculated. By using CANoe software to simulate a real vehicle CAN bus network and using the replay block to simulate different attack modes, experimental results show that the proposed IDS can effectively detect drop and replay attacks with detection rates of 100% and 98.2% respectively. However, for tampering attacks, the detection rate is only 66.2%.
[0008] The anomaly-based detection mechanism is modeled according to the characteristics of normal traffic on the bus, and the characteristics that deviate from normal traffic are defined as abnormal traffic. In the current in-vehicle network environment where various types of attacks exist, the anomaly-based method is more suitable for attack detection than the signature-based method due to its stable baseline characteristics. Song et al. [Hyun Min Song, Jiyoung Woo, HKKim, “In-vehicle network intrusion detection using deepconvolutional neural network”, Vehicular Communications, vol. 21, pp. 100-198, 2020.] proposed an IDS (Intrusion Detection System) based on DCNN (Deep Convolutional Neural Network). DCNN learns network traffic patterns and detects malicious traffic without manually designing features. Experimental results show that compared with traditional machine learning algorithms, the IDS has lower false negative rates and error rates. Han et al. [MLHan, BIKwak, and HKKim, “Event-triggered interval-based anomaly detection and attack identification methods for an in-vehicle network,” IEEE Trans. Inf. Forensics Security, vol. 16, pp. 2941–2956, 2021.] detect and identify vehicle network anomalies by using the periodic event-triggered interval of CAN messages, by considering different attack scenarios and three types of machine learning models. The results show that when a tree-based machine learning model is used as a classifier, the proposed attack identification method can achieve an accuracy of more than 94%. However, this method is only applicable to periodic events. For non-periodic events or periodic events but the transmission rate on the bus depends on multiple variables, this method still has certain limitations. Summary of the Invention
[0009] To address the limitations of existing technologies, including the following: 1) Signature-based anomaly detection cannot detect unknown or new attacks, requiring constant updating of the list of signatures considered as attack items. Furthermore, there is a high latency in extracting attack features in real time on a moving vehicle. 2) Anomaly-based detection mechanisms typically have a high false alarm rate. The present invention provides a semi-supervised in-vehicle CAN bus anomaly detection method and device based on dynamic graphs. The technical solution is as follows:
[0010] On the one hand, a semi-supervised in-vehicle CAN bus anomaly detection method based on a dynamic graph is provided. The method is implemented by an in-vehicle CAN bus anomaly detection device, and the method includes:
[0011] S1. Construct an undirected time dynamic graph based on the message flow sequence and message content of the CAN bus.
[0012] S2. The timestamps in the undirected temporal dynamic graph are converted into temporal embedding vectors through the time encoder, and the low-dimensional representation of nodes and edges in the undirected temporal dynamic graph is learned through the graph embedding module to obtain the node embedding of each node in the undirected temporal dynamic graph.
[0013] S3. Predict the anomaly score of each node based on the node embedding and anomaly detection network of each node. Record the anomaly score of each node and the time corresponding to the anomaly score through the time memory library, and calculate the statistical distribution of normal nodes.
[0014] S4. Calculate the reference distribution based on the anomaly score of each node and the statistical distribution of normal nodes, calculate the deviation score based on the reference distribution, and calculate the deviation loss based on the deviation score.
[0015] S5. Generate a pseudo label for each node according to the deviation score, train the anomaly detection model according to the deviation loss and the pseudo label of the node, and obtain a trained anomaly detection model.
[0016] S6. Obtain the message stream sequence and message content of the CAN bus to be detected, input them into the trained anomaly detection model, and obtain the CAN bus anomaly detection result.
[0017] Optionally, the undirected time dynamic graph in S1 is as shown in the following formula (1):
[0018] G=(V,E) (1)
[0019] Where G represents an undirected time-dynamic graph, V = v i represents the set of nodes involved in all CAN message flows, i represents the number of nodes, E = {δ(t1), δ(t2), ..., δ(t m )) represents the message stream sequence, m represents the number of observed messages, and the event δ(t)=(vi , v j , t, x ii ) indicates that at time t, i To the target node v j A message interaction occurs with edge feature x ij .
[0020] Optionally, the node embedding of each node in S2 is as shown in the following formula (2):
[0021] z i (t) = h (K) i (t) (2)
[0022] in,
[0023]
[0024] Where z i (t) represents the node v i Node embedding, i represents the number of nodes, t represents the time, K represents the number of layers in the neural network, h (k) i (t) represents the node v i The intermediate representation at time t of the k-th GNN layer, COMBINE(·) represents the function for combining the representations from the neighbors and their previous layer representations, N i Represents node v i The set of neighbor nodes at time t, Indicates that during the k-th layer aggregation process at time t, node v i The neighbor node set N i , AGG(·) represents the aggregation function, h j (k-1) (t) represents the node v j The representation of time t at the k-1th layer, x ij Represents node v i With node v j The associated edge features between them, φ(·) represents the relative time encoder based on cosine transform, Δt represents the relative time span of two timestamps, v j ∈N(v i , t) represents the node v i The set of first-order neighboring nodes occurring before t.
[0025] Optionally, the anomaly score of each node is predicted based on the node embedding and anomaly detection network in S3, including:
[0026] A feedforward neural network is used as an anomaly detector to map the node embedding of each node to the anomaly score space and predict the anomaly score of each node, as shown in the following formula (5):
[0027]
[0028] Where s i (t) represents the one-dimensional anomaly score, t represents the time, i represents the number of nodes, represents a feedforward neural network, θ a represents the parameter set of the anomaly detector, z i (t) represents the node v i , ReLU(·) represents the activation function, and W1, W2, b1, and b2 represent the learnable parameters of the anomaly detector.
[0029] Optionally, the time memory in S3 is as shown in the following formula (6):
[0030] m=(s i (t), t), ify i (t)=0 or -1 or 1 (6)
[0031] Where m represents the information stored in the time memory bank, s i (t) represents the one-dimensional anomaly score, t represents the time, i represents the number of nodes, y i (t) represents the node v i Label information at time t.
[0032] Optionally, in S4, the reference distribution is calculated based on the anomaly score of each node and the statistical distribution of normal nodes, as shown in the following equations (7)-(8):
[0033]
[0034] Where μr(t) represents the average value of the reference score of the statistical distribution of normal nodes at time t, k′ represents the number of samples randomly drawn from the memory bank, Represents the weighted term of each statistical sample, t i represents the storage time of abnormality score, r i represents the i-th anomaly score, σ r (t) represents the standard deviation of the reference score of the statistical distribution of normal nodes at time t.
[0035] The deviation score is calculated based on the reference distribution as shown in the following formula (9):
[0036]
[0037] Where, dev(v i, t) represents the deviation score, s i (t) represents the one-dimensional anomaly score, μ r (t) represents the average value of the reference score of the statistical distribution of normal nodes at time t, σ r (t) represents the standard deviation of the reference score of the statistical distribution of normal nodes at time t.
[0038] The deviation loss is calculated based on the deviation score, as shown in the following formula (10):
[0039] L dev =(1-y i (t))·|dev(v i , t)|+y i (t)·max(0,m′-|dev(v i , t)|) (10)
[0040] Where, L dev Denotes the deviation loss, y i (t) represents the node v i The label information at time t, dev(v i , t) represents the deviation score, and m′ represents the threshold parameter.
[0041] Optionally, S5 generates a pseudo label for each node based on the deviation score, including:
[0042] The deviation score distance between nodes is calculated based on the deviation score, the nodes are grouped according to the deviation score distance, and a pseudo label is generated for each node based on the grouping result.
[0043] Among them, the supervised contrastive learning loss of the node is as shown in the following formula (11):
[0044]
[0045] Where, represents the supervised contrastive learning loss, N represents the batch size of batch training samples, j represents the number of nodes, Δd ij represents the deviation fraction distance, z i (t i ) represents node v i At time t i The embedding representation of z j (t j ) represents node v j At time t j The embedding representation is, τ represents the scalar temperature parameter, k represents the number of nodes, z k (t k ) represents node v k At time t kEmbedding representation of .
[0046] On the other hand, a semi-supervisory in-vehicle CAN bus anomaly detection device based on a dynamic graph is provided. The device is applied to a semi-supervisory in-vehicle CAN bus anomaly detection method based on a dynamic graph. The device includes:
[0047] The dynamic graph construction module is used to construct an undirected time dynamic graph according to the message flow sequence and message content of the CAN bus.
[0048] The temporal graph attention network module is used to convert the timestamps in the undirected temporal dynamic graph into temporal embedding vectors through the temporal encoder, and learn the low-dimensional representation of nodes and edges in the undirected temporal dynamic graph through the graph embedding module to obtain the node embedding of each node in the undirected temporal dynamic graph.
[0049] The anomaly detection and temporal memory module is used to predict the anomaly score of each node based on the node embedding and anomaly detection network of each node, record the anomaly score of each node and the time corresponding to the anomaly score through the temporal memory, and calculate the statistical distribution of normal nodes.
[0050] The deviation loss network module is used to calculate the reference distribution based on the anomaly score of each node and the statistical distribution of normal nodes, calculate the deviation score based on the reference distribution, and calculate the deviation loss based on the deviation score.
[0051] The supervised contrastive learning module is used to generate a pseudo label for each node according to the deviation score, and train the anomaly detection model based on the deviation loss and the pseudo label of the node to obtain a trained anomaly detection model.
[0052] The output module is used to obtain the message stream sequence and message content of the CAN bus to be detected, input them into the trained anomaly detection model, and obtain the CAN bus anomaly detection results.
[0053] Optionally, an undirected time dynamic graph is as shown in the following formula (1):
[0054] G=(V,E) (1)
[0055] Where G represents an undirected time-dynamic graph, V = v i represents the set of nodes involved in all CAN message flows, i represents the number of nodes, E = {δ(t1), δ(t2), ..., δ(t m )) represents the message stream sequence, m represents the number of observed messages, and the event δ(t)=(v i , v j , t, x ij ) indicates that at time t, i To the target node v j A message interaction occurs with edge feature xij .
[0056] Optionally, the node embedding of each node is as shown in the following formula (2):
[0057] z i (t) = h (K) i (t) (2)
[0058] in,
[0059]
[0060] Where z i (t) represents the node v i Node embedding, i represents the number of nodes, t represents the time, K represents the number of layers in the neural network, h (k) i (t) represents the node v i The intermediate representation at time t of the k-th GNN layer, COMBINE(·) represents the function for combining the representations from the neighbors and their previous layer representations, and Ni represents the node v i The set of neighbor nodes at time t, Indicates that during the k-th layer aggregation process at time t, node v i The neighbor node set N i , AGG(·) represents the aggregation function, h j (k-1) (t) represents the node v j The representation of time t at the k-1th layer, x ij Represents node v i With node v j The associated edge features between them, φ(·) represents the relative time encoder based on cosine transform, Δt represents the relative time span of two timestamps, v j ∈N(v i , t) represents the node v i The set of first-order neighboring nodes occurring before t.
[0061] Optionally, the anomaly detection and time memory module is further configured to:
[0062] A feedforward neural network is used as an anomaly detector to map the node embedding of each node to the anomaly score space and predict the anomaly score of each node, as shown in the following formula (5):
[0063]
[0064] Where s i (t) represents the one-dimensional anomaly score, t represents the time, i represents the number of nodes, represents a feedforward neural network, θ a represents the parameter set of the anomaly detector, z i (t) represents the node v i , ReLU(·) represents the activation function, and W1, W2, b1, and b2 represent the learnable parameters of the anomaly detector.
[0065] Optionally, the time memory library is as shown in the following formula (6):
[0066] m=(s i (t), t), ify i (t)=0 or -1 or 1 (6)
[0067] Where m represents the information stored in the time memory bank, s i (t) represents the one-dimensional anomaly score, t represents the time, i represents the number of nodes, y i (t) represents the node v i Label information at time t.
[0068] Optionally, the reference distribution is calculated based on the anomaly score of each node and the statistical distribution of normal nodes, as shown in the following equations (7)-(8):
[0069]
[0070] Where μr(t) represents the average value of the reference score of the statistical distribution of normal nodes at time t, k′ represents the number of samples randomly drawn from the memory bank, Represents the weighted term of each statistical sample, t i represents the storage time of abnormality score, r i represents the i-th anomaly score, σ r (t) represents the standard deviation of the reference score of the statistical distribution of normal nodes at time t.
[0071] The deviation score is calculated based on the reference distribution as shown in the following formula (9):
[0072]
[0073] Where, dev(v i , t) represents the deviation score, s i (t) represents the one-dimensional anomaly score, μ r (t) represents the average value of the reference score of the statistical distribution of normal nodes at time t, σ r (t) represents the standard deviation of the reference score of the statistical distribution of normal nodes at time t.
[0074] The deviation loss is calculated based on the deviation score, as shown in the following formula (10):
[0075] L dev =(1-y i (t))·|dev(v i , t)|+y i (t)·max(0,m′-|dev(v i , t)|) (10)
[0076] Where, L dev Denotes the deviation loss, y i (t) represents the node v i The label information at time t, dev(v i , t) represents the deviation score, and m′ represents the threshold parameter.
[0077] Optionally, the supervised contrastive learning module is further configured to:
[0078] The deviation score distance between nodes is calculated based on the deviation score, the nodes are grouped according to the deviation score distance, and a pseudo label is generated for each node based on the grouping result.
[0079] Among them, the supervised contrastive learning loss of the node is as shown in the following formula (11):
[0080]
[0081] Where, represents the supervised contrastive learning loss, N represents the batch size of batch training samples, j represents the number of nodes, Δd ij represents the deviation fraction distance, z i (t i ) represents node v i At time t i The embedding representation of z j (t j ) represents node v j At time t j The embedding representation is, τ represents the scalar temperature parameter, k represents the number of nodes, z k (t k ) represents node v k At time t k Embedding representation of .
[0082] On the other hand, an in-vehicle CAN bus anomaly detection device is provided, which includes: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned semi-supervised in-vehicle CAN bus anomaly detection methods based on dynamic graphs is implemented.
[0083] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned semi-supervisory in-vehicle CAN bus anomaly detection methods based on dynamic graphs.
[0084] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0085] This paper proposes a temporal dynamic graph based on CAN message flows, converting timestamps into temporal embedding vectors using a temporal graph attention network. As the graph embedding module continuously learns low-dimensional representations of nodes and edges in the graph, capturing the graph's structural features and node state changes, it can detect CAN network attacks on real vehicles without requiring knowledge of the vehicle's CAN IDs, effectively resolving the issue of in-vehicle CAN network anomaly detection relying on CAN IDs.
[0086] In real-world scenarios, anomalous samples are often rare and difficult to obtain, leading to extremely unbalanced datasets. This paper uses the statistical distribution of unlabeled samples as a reference distribution for loss calculation and generates corresponding pseudo-labels for supervised learning. This fully exploits the potential of unlabeled samples and effectively addresses the problem of highly unbalanced training sets caused by the difficulty of obtaining the target data in the real world. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0088] Figure 1 This is a flow chart of a method for detecting anomalies of a semi-supervised in-vehicle CAN bus based on a dynamic graph provided by an embodiment of the present invention;
[0089] Figure 2 This is a design diagram of a semi-supervisory in-vehicle CAN bus anomaly detection method based on a dynamic graph provided by an embodiment of the present invention;
[0090] Figure 3 This is a block diagram of a semi-supervisory in-vehicle CAN bus anomaly detection device based on a dynamic graph provided by an embodiment of the present invention;
[0091] Figure 4 The present invention provides a schematic structural diagram of an in-vehicle CAN bus anomaly detection device. DETAILED DESCRIPTION
[0092] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0093] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.
[0094] In the embodiments of the present invention, the terms "image" and "picture" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same. The terms "of," "corresponding," and "corresponding" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same.
[0095] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0096] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0097] The embodiment of the present invention provides a semi-supervisory in-vehicle CAN bus anomaly detection method based on dynamic graph, which can be implemented by an in-vehicle CAN bus anomaly detection device, which can be a terminal or a server. Figure 1 、 Figure 2 The flowchart of the semi-supervisory in-vehicle CAN bus anomaly detection method based on dynamic graph is shown. The processing flow of the method may include the following steps:
[0098] S1. Construct an undirected time dynamic graph based on the message flow sequence and message content of the CAN bus.
[0099] In one feasible implementation, dynamic graphs, compared to static graphs, can be updated over time or with data changes, reflecting real-time data or the results of interactive operations. Compared to message sequences and static graphs, dynamic graphs can further embed message content and changes in node status over time.
[0100] Furthermore, in the anomaly detection mechanism of the present invention, the dynamic graph is represented as:
[0101] G=(V,E) (1)
[0102] Where G represents an undirected time-dynamic graph, V = v i represents the set of nodes involved in all CAN message flows, i represents the number of nodes, E is the message flow sequence, and it is usually assumed that E = {δ(t1), δ(t2), ..., δ(t m )) is the message flow of the generated temporal network, m is the number of observed messages, and the event δ(t)=(v i , v j , t, x ij ) indicates that at time t, i To the target node v j A message interaction occurs, accompanied by edge feature x ij In the present invention, edge features refer to the data fields carried in the message. Note that there may be multiple links between the same pair of node identifiers at different timestamps.
[0103] S2. The timestamps in the undirected temporal dynamic graph are converted into temporal embedding vectors through the time encoder, and the low-dimensional representation of nodes and edges in the undirected temporal dynamic graph is learned through the graph embedding module to obtain the node embedding of each node in the undirected temporal dynamic graph.
[0104] In one implementation, a temporal graph encoder based on time-coded information is used to convert timestamps into embedding vectors that reflect temporal periodicity and trends. These embedding vectors are then combined with node and edge features through a multi-layer attention mechanism to aggregate information about neighboring nodes and, combined with the time coding, update the node's feature representation.
[0105] Specifically, in dynamic graphs, node embeddings should include not only static node features (such as inherent node attributes), but also node features and topological structures that evolve over time. The Temporal Graph Attention Network module is used to identify node embeddings as cosine functions with respect to time, and can infer and observe node embeddings as the graph changes.
[0106] The temporal graph attention network mainly consists of two key parts: the time encoder and the graph embedding module. The time encoder converts the timestamp into a time embedding vector through the cosine periodic function. These vectors can capture the periodicity and trend of time. The graph embedding module is responsible for learning low-dimensional representations of nodes and edges in the graph. These representations can capture the structural characteristics of the graph and the dynamic characteristics that change over time. This module is a number of GNN (Graph Neural Network) layers, which take the graph G constructed at time t as input and send a message to all v in G through the message passing mechanism. i Extract node representation z i (t). Formally, the forward propagation of the kth layer is described as follows:
[0107]
[0108] Where z i (t) represents the node v i Node embedding, i represents the number of nodes, t represents the time, K represents the number of layers in the neural network, h (k) i (t) represents the node v i The intermediate representation at the time (timestamp) t of the k-th GNN layer, COMBINE(·) represents the function for combining the representations from the neighbors and their previous layer representations, N i Represents node v i The set of neighbor nodes at time t, these neighbor nodes are in the graph G with node v i directly connected nodes, Indicates that during the k-th layer aggregation process at time t, node v i The neighbor node set N i , AGG(·) represents the aggregation function, which is used to propagate and aggregate messages from neighbors, and h j (k-1) (t) represents the node v j The representation (or embedding) of time t at the k-1th layer, x ij Represents node v i With node v j The associated edge features between them, φ(·) represents the relative time encoder based on cosine transform, φ(·) converts the time information into a set of vectors that can represent time features, which helps to capture the periodic pattern on the dynamic graph, Δt=tt ij Indicates the relative time span between two timestamps, t ij Represents the edge (v i , v j ) is added to the graph, v j ∈N(v i , t) represents the node v i The set of first-order adjacent nodes that occurred before t. Finally, we can get node v i Node embedding: z i (t) = h (K) i (t).
[0109] S3. Predict the anomaly score of each node based on the node embedding and anomaly detection network of each node. Record the anomaly score of each node and the time corresponding to the anomaly score through the time memory library, and calculate the statistical distribution of normal nodes.
[0110] In one possible implementation, the encoded data is fed into an anomaly detection network that predicts an anomaly score for each node. The IDS also uses a temporal memory to record each predicted anomaly score and the node's temporal information.
[0111] Optionally, the anomaly score of each node is predicted based on the node embedding and anomaly detection network in S3, including:
[0112] In order to distinguish abnormal samples from normal samples, this paper introduces an anomaly detector that embeds the learned nodes into z i (t) is mapped to anomaly score space, where normal samples are clustered and abnormal samples deviate from the set. Specifically, the present invention uses a simple feedforward neural network As an anomaly detector:
[0113]
[0114] Where s i (t) represents the one-dimensional anomaly score, t represents the time, i represents the number of nodes, represents a feedforward neural network, θ a represents the parameter set of the anomaly detector, z i (t) represents the node v i , ReLU(·) represents the activation function, and W1, W2, b1, b2 represent the learnable parameters of the anomaly detector. In the framework of the present invention, s i (t) is a one-dimensional anomaly score whose value ranges from -∞ to +∞. Ideally, the anomaly scores of normal samples will be clustered in a specific interval, and it is easy to find anomalies (outliers) based on the overall distribution of anomaly scores.
[0115] Furthermore, based on the assumption that most unlabeled data are normal samples, the present invention generates this distribution by using a temporal memory library to record the historical anomaly scores of samples. The messages that need to be stored in the memory library are described as follows:
[0116] m=s i (t),ify i (t)=0 or -1 or 1 (5)
[0117] Among them, y i (t) is the node v iThe label information at time point t, with values of -1, 0, and 1, respectively, indicates that the node is an unlabeled sample, a normal sample, or an anomaly sample at the current time point. Since the attributes of nodes in a dynamic graph are constantly evolving, the labels and anomaly scores also change accordingly. Furthermore, samples with longer time intervals should have less impact on the current statistical distribution. Therefore, the present invention also records the time t corresponding to each anomaly score in the time memory library.
[0118] m=(s i (t), t), ify i (t)=0 or -1 or 1 (6)
[0119] At the same time, the temporal memory is designed as a first-in, first-out queue of size M. Because outdated samples produce less gain in the change of model parameters during model training, the temporal memory controls the queue size by discarding sufficiently old samples.
[0120] S4. Calculate the reference distribution based on the anomaly score of each node and the statistical distribution of normal nodes, calculate the deviation score based on the reference distribution, and calculate the deviation loss based on the deviation score.
[0121] In a feasible implementation, the information generated by steps S2-S3 is used to generate the statistical distribution of normal samples as prior knowledge, and to guide subsequent network learning by calculating the deviation loss.
[0122] Specifically, the present invention uses deviation loss [Guansong Pang, Chunhua Shen, and Anton vandenHengel. Deep anomaly detection with deviation networks. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, August 4-8, 2019, pages 353-362, 2019.] as the main learning objective, which aims to enable the model to distinguish normal nodes from abnormal nodes. Deviation loss is calculated by comparing the anomaly score of each node with the statistical distribution of normal samples stored in the memory bank.
[0123] Furthermore, based on the assumption that the distribution of anomaly scores satisfies the Gaussian distribution, the present invention randomly extracts a set of samples of size Ms from the time memory to calculate the reference score. In order to improve the robustness of the model, the present invention adds random perturbations in the experiment. At the same time, the present invention introduces the effect of time decay, using the time mark t and the anomaly score storage time t i The relative time span between the two is used to calculate the weighted term of each statistical sample The average value μ of the reference score of the normal sample statistical distribution at time point t r (t) and standard deviation σ r (t) is calculated as follows:
[0124]
[0125] Where μ r (t) represents the average value of the reference score of the statistical distribution of normal nodes at time (time point) t, and k′ represents the number of samples randomly drawn from the memory bank. These samples are used to calculate the reference score at timestamp t (including the mean μ r (t) and standard deviation σ r (t)), Represents the weighted term of each statistical sample, t i represents the storage time of abnormality score, r i represents the i-th anomaly score, σ r (t) represents the standard deviation of the reference score of the statistical distribution of normal nodes at time point t.
[0126] Furthermore, the deviation score measures the predicted anomaly score s of a node at time t i (t) and the statistical distribution of normal sample scores in the memory bank (mean μ r (t) and standard deviation σ r The deviation score is calculated as follows:
[0127]
[0128] Where, dev(v i , t) represents the deviation score, s i (t) represents the one-dimensional anomaly score, μ r (t) represents the average value of the reference score of the statistical distribution of normal nodes at time point t, σ r (t) represents the standard deviation of the reference score of the statistical distribution of normal nodes at time point t.
[0129] Furthermore, for each node v i , the deviation score dev(v i , t) is used to calculate the deviation loss Ldev . L dev It is a contrastive loss, which aims to optimize the model parameters. It takes into account the label information y of the node i (t), if the node is a normal node (y i (t)=0), the loss function encourages the deviation score to be close to 0; if the node is an abnormal node (y i (t)=1), then the loss function encourages the deviation score to be away from 0. The calculation formula of deviation loss is as follows:
[0130] L dev =(1-y i (t))·|dev(v i , t)|+y i (t)·max(0,m′-|dev(v i , t)|) (10)
[0131] Where, L dev Denotes the deviation loss, y i (t) represents the node v i The label information at time point t, dev(v i , t) represents the deviation score, m′ represents the threshold parameter, which is equivalent to the Z-Score confidence interval.
[0132] Furthermore, by minimizing the bias loss, the node representation learned by the model will make the anomaly scores of normal samples close to the normal sample distribution in the temporal memory library, while the anomaly scores of abnormal samples will be far away from this distribution. In this way, the model can effectively detect abnormal nodes in dynamic graphs.
[0133] S5. Generate a pseudo label for each node according to the deviation score, train the anomaly detection model according to the deviation loss and the pseudo label of the node, and obtain a trained anomaly detection model.
[0134] In one feasible implementation, to further leverage the potential of large amounts of unlabeled data, the framework of this invention introduces a novel pseudo-label contrastive learning module. This module uses the predicted anomaly scores to generate pseudo-labels for nodes by calculating the fractional distance between nodes. Nodes with closer distances are grouped into the same pseudo-group and assigned the same label. Nodes within the same pseudo-group form positive pairs for contrastive learning.
[0135] Specifically, in order to fully utilize the potential of unlabeled samples, the present invention uses the existing deviation score dev(v i ,t) Generate a pseudo label y for each sample i(t) and incorporate them into the training of the graph encoder network. Here, the present invention designs a supervised contrastive learning task by calculating the deviation score distance Δd between sample pairs ij =|dev(v i , t i )-dev(v j , t j )| for grouping. If the deviation score distance between two samples is less than 1 standard deviation, they are grouped into the same group and are considered similar. Based on the assumption that similar samples should have similar feature representations, nodes with closer deviation scores have more similar node representations, while nodes with larger deviation score differences have larger representation differences. The goal of this task is consistent with the deviation loss of the present invention, which separates the differences between normal and abnormal samples in both the representation space and the anomaly score space. Therefore, for a single sample v i , whose supervised contrastive learning loss is:
[0136]
[0137] Where, represents the supervised contrastive learning loss, N represents the number of samples in the batch training sample, that is, the batch size, j represents the number of nodes, Δd ij represents the deviation fraction distance, z i (t i ) represents node v i At time t i The embedding representation of z j (t j ) represents node v j At time t j The embedding representation of , τ represents the scalar temperature parameter used to control the sensitivity of the loss function, k represents the number of nodes, z k (t k ) represents node v k At time t k The embedding representation of is an indicator function that is 1 when the deviation score distance between two samples is less than 1, and 0 otherwise.
[0138] This classification (grouping) is based on the similarities between samples rather than on predefined category labels. This approach is particularly suitable for semi-supervised learning scenarios, where the utilization of unlabeled data is crucial to improving model performance.
[0139] S6. Obtain the message stream sequence and message content of the CAN bus to be detected, input them into the trained anomaly detection model, and obtain the CAN bus anomaly detection result.
[0140] In this embodiment, a temporal dynamic graph based on CAN message flows is proposed, which converts timestamps into temporal embedding vectors using a temporal graph attention network. As the graph embedding module continuously learns low-dimensional representations of nodes and edges in the graph, capturing the graph's structural features and changes in node states, it can detect CAN network attacks on real vehicles without requiring knowledge of the vehicle's CAN IDs, effectively resolving the problem of in-vehicle CAN network anomaly detection relying on CAN IDs.
[0141] In real-world scenarios, anomalous samples are often rare and difficult to obtain, leading to extremely unbalanced datasets. This paper uses the statistical distribution of unlabeled samples as a reference distribution for loss calculation and generates corresponding pseudo-labels for supervised learning. This fully exploits the potential of unlabeled samples and effectively addresses the problem of highly unbalanced training sets caused by the difficulty of obtaining the target data in the real world.
[0142] Figure 3 This is a block diagram of a semi-supervisory in-vehicle CAN bus anomaly detection device based on a dynamic graph according to an exemplary embodiment. The device is used in a semi-supervisory in-vehicle CAN bus anomaly detection method based on a dynamic graph. Figure 3 The device includes a dynamic graph construction module 310, a time graph attention network module 320, an anomaly detection and time memory library module 330, a deviation loss network module 340, a supervised contrastive learning module 350 and an output module 360. Among them:
[0143] The dynamic graph construction module 310 is used to construct an undirected time dynamic graph according to the message flow sequence and message content of the CAN bus.
[0144] The time graph attention network module 320 is used to convert the timestamps in the undirected time dynamic graph into time embedding vectors through the time encoder, and learn the low-dimensional representation of nodes and edges in the undirected time dynamic graph through the graph embedding module to obtain the node embedding of each node in the undirected time dynamic graph.
[0145] The anomaly detection and time memory module 330 is used to predict the anomaly score of each node based on the node embedding of each node and the anomaly detection network, record the anomaly score of each node and the time corresponding to the anomaly score through the time memory, and calculate the statistical distribution of normal nodes.
[0146] The deviation loss network module 340 is configured to calculate a reference distribution based on the anomaly score of each node and the statistical distribution of normal nodes, calculate a deviation score based on the reference distribution, and calculate a deviation loss based on the deviation score.
[0147] The supervised contrastive learning module 350 is used to generate a pseudo label for each node according to the deviation score, and train the anomaly detection model according to the deviation loss and the pseudo label of the node to obtain a trained anomaly detection model.
[0148] The output module 360 is used to obtain the message stream sequence and message content of the CAN bus to be detected, input them into the trained anomaly detection model, and obtain the CAN bus anomaly detection result.
[0149] In this embodiment, a temporal dynamic graph based on CAN message flows is proposed, which converts timestamps into temporal embedding vectors using a temporal graph attention network. As the graph embedding module continuously learns low-dimensional representations of nodes and edges in the graph, capturing the graph's structural features and changes in node states, it can detect CAN network attacks on real vehicles without requiring knowledge of the vehicle's CAN IDs, effectively resolving the problem of in-vehicle CAN network anomaly detection relying on CAN IDs.
[0150] In real-world scenarios, anomalous samples are often rare and difficult to obtain, leading to extremely unbalanced datasets. This paper uses the statistical distribution of unlabeled samples as a reference distribution for loss calculation and generates corresponding pseudo-labels for supervised learning. This fully exploits the potential of unlabeled samples and effectively addresses the problem of highly unbalanced training sets caused by the difficulty of obtaining the target data in the real world.
[0151] Figure 4 FIG. 1 is a structural diagram of a vehicle CAN bus abnormality detection device provided by an embodiment of the present invention. Figure 4 As shown, the vehicle CAN bus abnormality detection device may include the above Figure 3 The semi-supervisory in-vehicle CAN bus anomaly detection device based on dynamic graph is shown. Optionally, the in-vehicle CAN bus anomaly detection device 410 may include a first processor 2001.
[0152] Optionally, the in-vehicle CAN bus abnormality detection device 410 may further include a memory 2002 and a transceiver 2003 .
[0153] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.
[0154] The following combination Figure 4 The components of the vehicle CAN bus abnormality detection device 410 are described in detail:
[0155] The first processor 2001 is the control center of the in-vehicle CAN bus anomaly detection device 410 and can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).
[0156] Optionally, the first processor 2001 can perform various functions of the in-vehicle CAN bus abnormality detection device 410 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.
[0157] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 CPU0 and CPU1 are shown in FIG.
[0158] In a specific implementation, as an embodiment, the vehicle CAN bus abnormality detection device 410 may also include multiple processors, such as Figure 4 1 and 2. The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0159] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be repeated here.
[0160] Alternatively, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and access the first processor 2001 through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0161] The transceiver 2003 is used to communicate with a network device or a terminal device.
[0162] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 4 (not shown separately in the figure). The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0163] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be connected to the vehicle through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0164] It should be noted that Figure 4 The structure of the in-vehicle CAN bus abnormality detection device 410 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0165] In addition, the technical effects of the in-vehicle CAN bus anomaly detection device 410 can refer to the technical effects of the semi-supervised in-vehicle CAN bus anomaly detection method based on dynamic graphs described in the above method embodiment, and will not be repeated here.
[0166] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0167] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0168] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0169] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0170] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0171] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0172] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0173] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0174] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.
[0175] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0176] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0177] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0178] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A semi-supervised in-vehicle CAN bus anomaly detection method based on dynamic graph, characterized in that: The method comprises: S1. Construct an undirected time dynamic graph according to the message flow sequence and message content of the CAN bus; S2. Convert the timestamp in the undirected time dynamic graph into a time embedding vector through a time encoder, and learn the low-dimensional representation of nodes and edges in the undirected time dynamic graph through a graph embedding module to obtain a node embedding of each node in the undirected time dynamic graph; S3. Predict the anomaly score of each node according to the node embedding and the anomaly detection network of each node, record the anomaly score of each node and the time corresponding to the anomaly score through the time memory library, and calculate the statistical distribution of normal nodes; S4. Calculate a reference distribution according to the abnormal score of each node and the statistical distribution of normal nodes, calculate a deviation score according to the reference distribution, and calculate a deviation loss according to the deviation score; S5. Generate a pseudo label for each node according to the deviation score, and train an anomaly detection model according to the deviation loss and the pseudo label of the node to obtain a trained anomaly detection model; S6. Obtain the message stream sequence and message content of the CAN bus to be detected, input them into the trained anomaly detection model, and obtain the CAN bus anomaly detection result.
2. The semi-supervised in-vehicle CAN bus anomaly detection method based on dynamic graph according to claim 1 is characterized in that: The undirected time dynamic graph in S1 is shown in the following formula (1): G=(V,E) (1) Where G represents an undirected time-dynamic graph, V = v i represents the set of nodes involved in all CAN message flows, i represents the number of nodes, E = {δ(t1),δ(t2),…,δ(t m )} represents the message stream sequence, m represents the number of observed messages, and event δ(t)=(v i ,v j ,t,x ij ) indicates that at time t, from the source node v i To the target node v j A message interaction occurs with edge feature x ij .
3. The semi-supervised in-vehicle CAN bus anomaly detection method based on dynamic graph according to claim 1 is characterized in that: The node embedding of each node in S2 is shown in the following formula (2): z i (t)=h (K) i (t) (2) in, In the formula, z i (t) represents the node v i Node embedding, i represents the number of nodes, t represents the time, K represents the number of layers in the neural network, and h represents the number of nodes in the neural network. (k) i (t) represents the node v i The intermediate representation at time t of the kth GNN layer, COMBINE(·) represents the function used to combine the representations from neighbors and their previous layer representations, N i Represents node v i The set of neighbor nodes at time t, Indicates that during the k-th layer aggregation process at time t, node v i The neighbor node set N i , AGG(·) represents the aggregation function, h j (k-1) (t) represents the node v j The representation of time t at the k-1th layer, x ij Represents node v i With node v j The associated edge features between them, φ(·) represents the relative time encoder based on cosine transform, Δt represents the relative time span of two timestamps, v j ∈N(v i ,t) represents the node v i The set of first-order neighboring nodes occurring before t.
4. The semi-supervised in-vehicle CAN bus anomaly detection method based on dynamic graph according to claim 1 is characterized in that: The step of predicting the anomaly score of each node according to the node embedding of each node and the anomaly detection network in S3 includes: A feedforward neural network is used as an anomaly detector to map the node embedding of each node to the anomaly score space and predict the anomaly score of each node, as shown in the following formula (5): In the formula, s i (t) represents the one-dimensional anomaly score, t represents the time, i represents the number of nodes, represents a feedforward neural network, θ a represents the parameter set of the anomaly detector, z i (t) represents the node v i , ReLU(·) represents the activation function, and W1, W2, b1, b2 represent the learnable parameters of the anomaly detector.
5. The semi-supervised in-vehicle CAN bus anomaly detection method based on dynamic graph according to claim 1 is characterized in that: The time memory library in S3 is shown in the following formula (6): m=(s i (t),t),if y i (t)=0 or -1 or 1 (6) Where m represents the information stored in the time memory bank, s i (t) represents the one-dimensional anomaly score, t represents the time, i represents the number of nodes, and y i (t) represents the node v i Label information at time t.
6. The semi-supervised in-vehicle CAN bus anomaly detection method based on dynamic graph according to claim 1 is characterized in that: The reference distribution is calculated in S4 according to the abnormal score of each node and the statistical distribution of normal nodes, as shown in the following equations (7)-(8): In the formula, μ r (t) represents the average value of the reference score of the statistical distribution of normal nodes at time t, k′ represents the number of samples randomly drawn from the memory bank, represents the weighted term of each statistical sample, t i represents the storage time of abnormality score, r i represents the i-th anomaly score, σ r (t) represents the standard deviation of the reference score of the statistical distribution of normal nodes at time t; The deviation score is calculated according to the reference distribution, as shown in the following formula (9): Where, dev(v i ,t) represents the deviation score, s i (t) represents the one-dimensional anomaly score, μ r (t) represents the average value of the reference score of the statistical distribution of normal nodes at time t, σ r (t) represents the standard deviation of the reference score of the statistical distribution of normal nodes at time t; The deviation loss is calculated according to the deviation score, as shown in the following formula (10): L dev =(1-y i (t))·|dev(v i ,t)|+y i (t)·max(0,m′-|dev(v i ,t)|) (10) Where, L dev represents the deviation loss, y i (t) represents the node v i The label information at time t, dev(v i ,t) represents the deviation score, and m′ represents the threshold parameter.
7. The semi-supervised in-vehicle CAN bus anomaly detection method based on dynamic graph according to claim 1 is characterized in that: Generating a pseudo label for each node according to the deviation score in S5 includes: Calculating the deviation score distance between nodes according to the deviation score, grouping the nodes according to the deviation score distance, and generating a pseudo label for each node according to the grouping result; Among them, the supervised contrastive learning loss of the node is as shown in the following formula (11): In the formula, represents the supervised contrastive learning loss, N represents the batch size of batch training samples, j represents the number of nodes, Δd ij represents the deviation fraction distance, z i (t i ) represents node v i At time t i The embedding representation of z j (t j ) represents node v j At time t j The embedding representation is, τ represents the scalar temperature parameter, k represents the number of nodes, z k (t k ) represents node v k At time t k The embedded representation of .
8. A semi-supervisory in-vehicle CAN bus anomaly detection device based on a dynamic graph, the semi-supervisory in-vehicle CAN bus anomaly detection device based on a dynamic graph is used to implement the semi-supervisory in-vehicle CAN bus anomaly detection method based on a dynamic graph as claimed in any one of claims 1 to 7, characterized in that: The device comprises: A dynamic graph construction module is used to construct an undirected time dynamic graph according to the message flow sequence and message content of the CAN bus; A time graph attention network module is used to convert the timestamp in the undirected time dynamic graph into a time embedding vector through a time encoder, and learn the low-dimensional representation of nodes and edges in the undirected time dynamic graph through a graph embedding module to obtain a node embedding of each node in the undirected time dynamic graph; An anomaly detection and time memory library module, used to predict the anomaly score of each node according to the node embedding of each node and the anomaly detection network, record the anomaly score of each node and the time corresponding to the anomaly score through the time memory library, and calculate the statistical distribution of normal nodes; A deviation loss network module, used to calculate a reference distribution according to the abnormal score of each node and the statistical distribution of normal nodes, calculate a deviation score according to the reference distribution, and calculate a deviation loss according to the deviation score; A supervised contrastive learning module is used to generate a pseudo label for each node according to the deviation score, and train an anomaly detection model according to the deviation loss and the pseudo label of the node to obtain a trained anomaly detection model; The output module is used to obtain the message stream sequence and message content of the CAN bus to be detected, and input them into the trained anomaly detection model to obtain the CAN bus anomaly detection result.
9. A vehicle CAN bus abnormality detection device, characterized in that: The in-vehicle CAN bus abnormality detection device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
CAN anomaly detection method and system based on semi-supervised learning, and storage medium
CN115865735A
Dynamic graph anomaly detection method and system based on GNN and LSTM
CN117315331A
Semi-supervised graph anomaly detection method based on multi-view comparative learning
CN118734212A
CAN bus network anomaly detection method and system
CN118827187A
Cited By
Network security protection method for identifying abnormal traffic by using intelligent algorithm
CN120880758A